{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30823,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# İmport DATA  and basic ","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport warnings\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.impute import SimpleImputer\nwarnings.filterwarnings('ignore')\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport pandas as pd\n\npd.set_option(\"display.max_columns\", 100)\nfrom sklearn.linear_model import LinearRegression, SGDRegressor, Ridge, Lasso, ElasticNet\nfrom sklearn.neighbors import KNeighborsRegressor\nfrom sklearn.ensemble import GradientBoostingRegressor, AdaBoostRegressor\nfrom sklearn.tree import ExtraTreeRegressor, DecisionTreeRegressor\nfrom xgboost import XGBRegressor\nfrom sklearn.svm import SVR\nfrom sklearn.neural_network import MLPRegressor\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:51:36.952363Z","iopub.execute_input":"2025-01-03T21:51:36.952701Z","iopub.status.idle":"2025-01-03T21:51:37.005744Z","shell.execute_reply.started":"2025-01-03T21:51:36.95267Z","shell.execute_reply":"2025-01-03T21:51:37.004964Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Regression with an Insurance\n### Goal: The objectives of this challenge is to predict insurance premiums based on various factors.","metadata":{}},{"cell_type":"code","source":"df=pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ntest_df=pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\n\ndf.sample(6)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:51:40.063118Z","iopub.execute_input":"2025-01-03T21:51:40.063477Z","iopub.status.idle":"2025-01-03T21:51:46.673421Z","shell.execute_reply.started":"2025-01-03T21:51:40.063447Z","shell.execute_reply":"2025-01-03T21:51:46.672335Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 📊 Dataset Columns and Descriptions  \n<hr>\n\n1. **`id`** 🆔  \n   - **English:** Unique identifier for each customer.  \n   - **Türkçe:** Her müşteri için benzersiz kimlik numarası.  \n\n2. **`Age`** 🎂  \n   - **English:** Customer's age.  \n   - **Türkçe:** Müşterinin yaşı.  \n\n3. **`Gender`** 🧑‍🤝‍🧑  \n   - **English:** Customer's gender (e.g., Male, Female).  \n   - **Türkçe:** Müşterinin cinsiyeti (ör. Erkek, Kadın).  \n\n4. **`Annual Income`** 💵  \n   - **English:** Yearly income of the customer.  \n   - **Türkçe:** Müşterinin yıllık geliri.  \n\n5. **`Marital Status`** 💍  \n   - **English:** Whether the customer is single, married, or divorced.  \n   - **Türkçe:** Müşterinin medeni durumu (bekar, evli, boşanmış).  \n\n6. **`Number of Dependents`** 👨‍👩‍👧‍👦  \n   - **English:** Number of dependents relying on the customer.  \n   - **Türkçe:** Müşteriye bağlı kişilerin sayısı.  \n\n7. **`Education Level`** 🎓  \n   - **English:** The highest level of education completed by the customer.  \n   - **Türkçe:** Müşterinin tamamladığı en yüksek eğitim seviyesi.  \n\n8. **`Occupation`** 🧑‍💼  \n   - **English:** The type of work the customer does.  \n   - **Türkçe:** Müşterinin mesleği.  \n\n9. **`Health Score`** 🩺  \n   - **English:** A score indicating the overall health status of the customer.  \n   - **Türkçe:** Müşterinin genel sağlık durumunu gösteren bir puan.  \n\n10. **`Location`** 📍  \n    - **English:** Customer's geographical location.  \n    - **Türkçe:** Müşterinin coğrafi konumu.  \n\n11. **`Policy Type`** 📜  \n    - **English:** Type of insurance policy the customer holds.  \n    - **Türkçe:** Müşterinin sahip olduğu sigorta poliçesi türü.  \n\n12. **`Previous Claims`** 🛠️  \n    - **English:** Number of claims made by the customer in the past.  \n    - **Türkçe:** Müşterinin geçmişte yaptığı hasar taleplerinin sayısı.  \n\n13. **`Vehicle Age`** 🚗  \n    - **English:** The age of the vehicle insured by the customer.  \n    - **Türkçe:** Sigortalı aracın yaşı.  \n\n14. **`Credit Score`** 💳  \n    - **English:** Customer's financial creditworthiness score.  \n    - **Türkçe:** Müşterinin finansal kredi skoru.  \n\n15. **`Insurance Duration`** ⏳  \n    - **English:** Duration for which the insurance policy is active.  \n    - **Türkçe:** Sigorta poliçesinin aktif olduğu süre.  \n\n16. **`Policy Start Date`** 📅  \n    - **English:** Date when the insurance policy began.  \n    - **Türkçe:** Sigorta poliçesinin başlangıç tarihi.  \n\n17. **`Customer Feedback`** 🗣️  \n    - **English:** Feedback or reviews provided by the customer.  \n    - **Türkçe:** Müşterinin verdiği geri bildirimler.  \n\n18. **`Smoking Status`** 🚬  \n    - **English:** Whether the customer smokes or not.  \n    - **Türkçe:** Müşterinin sigara içip içmediği.  \n\n19. **`Exercise Frequency`** 🏋️  \n    - **English:** How often the customer exercises.  \n    - **Türkçe:** Müşterinin egzersiz yapma sıklığı.  \n\n20. **`Property Type`** 🏠  \n    - **English:** Type of property owned by the customer (e.g., house, apartment).  \n    - **Türkçe:** Müşterinin sahip olduğu mülk türü (ör. ev, daire).  \n\n\n\n    ### Target Columns\n22. **`Premium Amount`** 💰  \n    - **English:** Amount the customer pays for the insurance policy.  \n    - **Türkçe:** Müşterinin sigorta poliçesi için ödediği prim miktarı.\n      \n<hr>","metadata":{}},{"cell_type":"code","source":"print(df.columns,'\\n',test_df.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:51:46.674434Z","iopub.execute_input":"2025-01-03T21:51:46.674673Z","iopub.status.idle":"2025-01-03T21:51:46.680124Z","shell.execute_reply.started":"2025-01-03T21:51:46.674652Z","shell.execute_reply":"2025-01-03T21:51:46.679153Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# EDA(Exploratory Data Analysis)","metadata":{}},{"cell_type":"code","source":"df.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:51:46.681943Z","iopub.execute_input":"2025-01-03T21:51:46.682227Z","iopub.status.idle":"2025-01-03T21:51:47.213406Z","shell.execute_reply.started":"2025-01-03T21:51:46.682204Z","shell.execute_reply":"2025-01-03T21:51:47.212434Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_df.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:51:47.214695Z","iopub.execute_input":"2025-01-03T21:51:47.215Z","iopub.status.idle":"2025-01-03T21:51:47.562603Z","shell.execute_reply.started":"2025-01-03T21:51:47.214967Z","shell.execute_reply":"2025-01-03T21:51:47.561751Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# YProfiling\n\n#### **YProfiling** is a tool used to accelerate exploratory data analysis (EDA). It quickly visualizes and summarizes the key statistics, distributions, missing values, correlations, and other important analyses of your dataset. Similar to Pandas Profiling, it provides a quick overview of the data and speeds up the data analysis process.\n","metadata":{}},{"cell_type":"code","source":"from ydata_profiling import ProfileReport\n\nprofile = ProfileReport(df, title=\"YData Profiling Report\")\n#profile.to_file(\"rapor.html\")  # Raporu Saving html\nprofile","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:52:55.19826Z","iopub.execute_input":"2025-01-03T21:52:55.199068Z","iopub.status.idle":"2025-01-03T21:54:23.78141Z","shell.execute_reply.started":"2025-01-03T21:52:55.199029Z","shell.execute_reply":"2025-01-03T21:54:23.779486Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## DATA Preprocessing ","metadata":{}},{"cell_type":"code","source":"categorical_columns = df.select_dtypes(include=['object', 'category']).columns.tolist()\nnumerical_columns = df.select_dtypes(include=['number']).columns.tolist()\n\nprint('Categorcla :',categorical_columns,'\\n','Numerical  :',numerical_columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:54:32.764035Z","iopub.execute_input":"2025-01-03T21:54:32.764368Z","iopub.status.idle":"2025-01-03T21:54:32.951334Z","shell.execute_reply.started":"2025-01-03T21:54:32.764343Z","shell.execute_reply":"2025-01-03T21:54:32.950306Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df[numerical_columns].info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:54:34.35839Z","iopub.execute_input":"2025-01-03T21:54:34.358712Z","iopub.status.idle":"2025-01-03T21:54:34.418007Z","shell.execute_reply.started":"2025-01-03T21:54:34.358687Z","shell.execute_reply":"2025-01-03T21:54:34.417003Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df[categorical_columns].info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:54:34.715732Z","iopub.execute_input":"2025-01-03T21:54:34.716087Z","iopub.status.idle":"2025-01-03T21:54:35.381274Z","shell.execute_reply.started":"2025-01-03T21:54:34.716054Z","shell.execute_reply":"2025-01-03T21:54:35.380267Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Uniq Values of Categorical\nfor col in categorical_columns:\n    print(col,'\\n',df[col].unique(),'\\n',\"-\" * 30)  ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:54:35.382507Z","iopub.execute_input":"2025-01-03T21:54:35.382762Z","iopub.status.idle":"2025-01-03T21:54:36.146408Z","shell.execute_reply.started":"2025-01-03T21:54:35.382739Z","shell.execute_reply":"2025-01-03T21:54:36.145485Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Uniq Values of \nfor col in numerical_columns:\n    print(col,'\\n',df[col].unique(),'\\n',\"-\" * 30)  ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:55:01.255536Z","iopub.execute_input":"2025-01-03T21:55:01.255903Z","iopub.status.idle":"2025-01-03T21:55:01.45622Z","shell.execute_reply.started":"2025-01-03T21:55:01.255859Z","shell.execute_reply":"2025-01-03T21:55:01.455165Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df['Premium Amount'].describe().T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:05:51.992667Z","iopub.execute_input":"2025-01-03T22:05:51.993153Z","iopub.status.idle":"2025-01-03T22:05:52.047317Z","shell.execute_reply.started":"2025-01-03T22:05:51.993106Z","shell.execute_reply":"2025-01-03T22:05:52.046279Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Clean the data with a function :)","metadata":{}},{"cell_type":"code","source":"def Categorical_Encode(df):\n    # Select categorical columns\n    categorical_columns = df.select_dtypes(include=['object']).columns\n    \n    # If no categorical columns are found, print a warning and return the DataFrame\n    if len(categorical_columns) == 0:\n        print(\"Warning: No categorical columns found in the DataFrame.\")\n        return df\n    \n    # Fill missing values with the most frequent value using SimpleImputer\n    imputer = SimpleImputer(strategy='most_frequent')\n    df[categorical_columns] = imputer.fit_transform(df[categorical_columns])\n    \n    # Apply Label Encoding to each categorical column\n    for col in categorical_columns:\n        le = LabelEncoder()\n        df[col] = le.fit_transform(df[col])\n    \n    return df\n\n\ndef fix_dates(df, date_column='Policy Start Date'):\n    # Check if the specified date column exists in the DataFrame\n    if date_column not in df.columns:\n        print(f\"Warning: '{date_column}' column not found in the DataFrame.\")\n        return df\n    \n    # Convert the date column to datetime format, coercing errors to NaT (Not a Time)\n    df[date_column] = pd.to_datetime(df[date_column], errors='coerce')\n    \n    # Check for missing or invalid dates in the date column\n    if df[date_column].isnull().any():\n        print(f\"Warning: Missing or invalid dates found in the '{date_column}' column.\")\n    \n    # Extract year, month, and day from the date column\n    df['Year'] = df[date_column].dt.year\n    df['Month'] = df[date_column].dt.month\n    df['Day'] = df[date_column].dt.day\n    \n    # Drop the original date column\n    df.drop(date_column, axis=1, inplace=True)\n    \n    return df\n\n\n\ndef fill_missing_with_iqr(df):\n    \n    numeric_columns = df.select_dtypes(include=['number']).columns\n    \n    if len(numeric_columns) == 0:\n        print(\"Uyarı: DataFrame'de numeric kolon bulunamadı.\")\n        return df\n    \n    for col in numeric_columns:\n        if df[col].isnull().any():\n            Q1 = df[col].quantile(0.25)\n            Q3 = df[col].quantile(0.75)\n            IQR = Q3 - Q1\n            \n            lower_bound = Q1 - 1.5 * IQR\n            upper_bound = Q3 + 1.5 * IQR\n            \n            random_values = np.random.uniform(lower_bound, upper_bound, size=df[col].isnull().sum())\n            df.loc[df[col].isnull(), col] = random_values\n    \n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:55:02.149256Z","iopub.execute_input":"2025-01-03T21:55:02.149576Z","iopub.status.idle":"2025-01-03T21:55:02.158142Z","shell.execute_reply.started":"2025-01-03T21:55:02.149551Z","shell.execute_reply":"2025-01-03T21:55:02.156972Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#date\nclean_train = fix_dates(df)\nclean_test = fix_dates(test_df)\n\n\n# categorical\nclean_train = Categorical_Encode(clean_train)\nclean_test  =Categorical_Encode(clean_test)\n\n# numeric\nclean_train = fill_missing_with_iqr(clean_train) \nclean_test  = fill_missing_with_iqr(clean_test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:55:02.882365Z","iopub.execute_input":"2025-01-03T21:55:02.882672Z","iopub.status.idle":"2025-01-03T21:55:12.722832Z","shell.execute_reply.started":"2025-01-03T21:55:02.882647Z","shell.execute_reply":"2025-01-03T21:55:12.722056Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:32:22.01681Z","iopub.execute_input":"2025-01-03T22:32:22.017145Z","iopub.status.idle":"2025-01-03T22:32:22.035043Z","shell.execute_reply.started":"2025-01-03T22:32:22.017118Z","shell.execute_reply":"2025-01-03T22:32:22.034229Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"clean_train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:55:21.564086Z","iopub.execute_input":"2025-01-03T21:55:21.564452Z","iopub.status.idle":"2025-01-03T21:55:21.583598Z","shell.execute_reply.started":"2025-01-03T21:55:21.564421Z","shell.execute_reply":"2025-01-03T21:55:21.582461Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"clean_train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:55:22.319646Z","iopub.execute_input":"2025-01-03T21:55:22.320043Z","iopub.status.idle":"2025-01-03T21:55:22.366628Z","shell.execute_reply.started":"2025-01-03T21:55:22.320004Z","shell.execute_reply":"2025-01-03T21:55:22.365739Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"clean_train.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:55:24.680067Z","iopub.execute_input":"2025-01-03T21:55:24.680378Z","iopub.status.idle":"2025-01-03T21:55:24.714851Z","shell.execute_reply.started":"2025-01-03T21:55:24.680354Z","shell.execute_reply":"2025-01-03T21:55:24.713798Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Distributions of current data ","metadata":{}},{"cell_type":"code","source":"important_columns = ['Annual Income', 'Health Score', 'Credit Score', 'Premium Amount']\n\nfor col in important_columns:\n    fig, axes = plt.subplots(1, 3, figsize=(18, 5))\n    \n    # Histogram\n    sns.histplot(clean_train[col], kde=True, color='blue', ax=axes[0])\n    axes[0].set_title(f'Histogram of {col}')\n    axes[0].set_xlabel(col)\n    axes[0].set_ylabel('Frequency')\n    \n    # Boxplot\n    sns.boxplot(x=clean_train[col], color='red', ax=axes[1])\n    axes[1].set_title(f'Boxplot of {col}')\n    axes[1].set_xlabel(col)\n    \n    # Violin Plot\n    sns.violinplot(x=clean_train[col], color='green', ax=axes[2])\n    axes[2].set_title(f'Violin Plot of {col}')\n    axes[2].set_xlabel(col)\n    \n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T21:55:29.091263Z","iopub.execute_input":"2025-01-03T21:55:29.091619Z","iopub.status.idle":"2025-01-03T21:56:00.657109Z","shell.execute_reply.started":"2025-01-03T21:55:29.091588Z","shell.execute_reply":"2025-01-03T21:56:00.656145Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(12, 8))\nsns.heatmap(clean_train.corr(), annot=True, cmap='coolwarm', fmt='.2f')\nplt.title('Corelasyon Heatmap')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T23:01:09.670228Z","iopub.execute_input":"2025-01-03T23:01:09.670557Z","iopub.status.idle":"2025-01-03T23:01:12.497854Z","shell.execute_reply.started":"2025-01-03T23:01:09.670532Z","shell.execute_reply":"2025-01-03T23:01:12.496865Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Modeling","metadata":{}},{"cell_type":"markdown","source":"### Find the best model for data","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LinearRegression, Ridge, Lasso, SGDRegressor\nfrom sklearn.ensemble import GradientBoostingRegressor, RandomForestRegressor\nfrom xgboost import XGBRegressor\nfrom lightgbm import LGBMRegressor\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import r2_score, mean_squared_error, mean_absolute_error\nimport pandas as pd\nimport numpy as np\nimport time\n\ndef algo_test_large(x, y, sample_size=None):\n   \n    # Use a sample of the data if sample_size is provided, otherwise use the full dataset\n    if sample_size:\n        sample_data = x.sample(n=sample_size, random_state=42)\n        sample_target = y.loc[sample_data.index]\n    else:\n        sample_data = x\n        sample_target = y\n\n    # Define models to test\n    models = {\n        'Linear': LinearRegression(n_jobs=-1),\n        'Ridge': Ridge(),\n        'Lasso': Lasso(),\n        'Gradient Boosting': GradientBoostingRegressor(),\n        'XGBoost': XGBRegressor(n_jobs=-1),\n        'LightGBM': LGBMRegressor(n_jobs=-1),\n        'SGD': SGDRegressor(),\n        'Random Forest': RandomForestRegressor(n_jobs=-1)\n    }\n\n    # Split the data into training and testing sets\n    x_train, x_test, y_train, y_test = train_test_split(sample_data, sample_target, test_size=0.1, random_state=42)\n    results = []\n\n    # Train and evaluate each model\n    for name, model in models.items():\n        start_time = time.time()\n        model.fit(x_train, y_train)\n        y_pred = model.predict(x_test)\n        r2 = r2_score(y_test, y_pred)\n        rmse = np.sqrt(mean_squared_error(y_test, y_pred))\n        mae = mean_absolute_error(y_test, y_pred)\n        training_time = time.time() - start_time\n        results.append((name, r2, rmse, mae, training_time))\n\n    # Create a DataFrame with the results\n    result_df = pd.DataFrame(results, columns=['Model', 'R_Squared', 'RMSE', 'MAE', 'Training Time (s)'])\n    return result_df.sort_values('R_Squared', ascending=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:12:10.727418Z","iopub.execute_input":"2025-01-03T22:12:10.727738Z","iopub.status.idle":"2025-01-03T22:12:10.735899Z","shell.execute_reply.started":"2025-01-03T22:12:10.727711Z","shell.execute_reply":"2025-01-03T22:12:10.734751Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Train and Testing","metadata":{}},{"cell_type":"code","source":"x = clean_train.drop(columns=['id', 'Premium Amount'])\ny = clean_train['Premium Amount']\n\ntest = clean_test.drop(columns=['id'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:12:14.478796Z","iopub.execute_input":"2025-01-03T22:12:14.479171Z","iopub.status.idle":"2025-01-03T22:12:14.580255Z","shell.execute_reply.started":"2025-01-03T22:12:14.479136Z","shell.execute_reply":"2025-01-03T22:12:14.579337Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(x.shape,test.shape)\nprint(x.columns.tolist(),'\\n',test.columns.tolist()) # same  , its good","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:12:17.240526Z","iopub.execute_input":"2025-01-03T22:12:17.240866Z","iopub.status.idle":"2025-01-03T22:12:17.247662Z","shell.execute_reply.started":"2025-01-03T22:12:17.240838Z","shell.execute_reply":"2025-01-03T22:12:17.246402Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#  If the columns are different, you can follow this step to make them the same:\ntest = test.reindex(columns=x.columns, fill_value=0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:12:18.202624Z","iopub.execute_input":"2025-01-03T22:12:18.202969Z","iopub.status.idle":"2025-01-03T22:12:18.29608Z","shell.execute_reply.started":"2025-01-03T22:12:18.202932Z","shell.execute_reply":"2025-01-03T22:12:18.295314Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"results = algo_test_large(x, y)\nprint(results)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:12:21.809518Z","iopub.execute_input":"2025-01-03T22:12:21.809932Z","iopub.status.idle":"2025-01-03T22:32:22.015453Z","shell.execute_reply.started":"2025-01-03T22:12:21.809894Z","shell.execute_reply":"2025-01-03T22:32:22.014329Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Best model ->\n\n Model     R_Squared          RMSE           MAE  \\\n LightGBM  3.570845e-02  8.488941e+02  6.484566e+02   \n XGBoost  3.272428e-02  8.502066e+02  6.487245e+02 ","metadata":{}},{"cell_type":"markdown","source":"## final stage","metadata":{}},{"cell_type":"code","source":"from lightgbm import LGBMRegressor\nmodel = LGBMRegressor(n_jobs=-1, random_state=42)\nmodel.fit(x, y)\n\nprediction = model.predict(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:43:51.605838Z","iopub.execute_input":"2025-01-03T22:43:51.60625Z","iopub.status.idle":"2025-01-03T22:43:57.696807Z","shell.execute_reply.started":"2025-01-03T22:43:51.606219Z","shell.execute_reply":"2025-01-03T22:43:57.694745Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"prediction","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:44:21.297478Z","iopub.execute_input":"2025-01-03T22:44:21.297793Z","iopub.status.idle":"2025-01-03T22:44:21.303846Z","shell.execute_reply.started":"2025-01-03T22:44:21.297766Z","shell.execute_reply":"2025-01-03T22:44:21.302842Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/playground-series-s4e12/sample_submission.csv')\n\nsubmission['Premium Amount'] = prediction\n\nsubmission.to_csv('result.csv', index=False)\n\nprint(\"Submission dosyası kaydedildi: 'submission.csv'\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T22:47:26.85206Z","iopub.execute_input":"2025-01-03T22:47:26.85241Z","iopub.status.idle":"2025-01-03T22:47:28.419328Z","shell.execute_reply.started":"2025-01-03T22:47:26.852378Z","shell.execute_reply":"2025-01-03T22:47:28.418226Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\nfeature_importance = model.feature_importances_\nfeature_names = x.columns\n\nplt.figure(figsize=(10, 6))\nsns.barplot(x=feature_importance, y=feature_names)\nplt.title('Feature Importance')\nplt.xlabel('Importance')\nplt.ylabel('Features')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T23:08:37.135128Z","iopub.execute_input":"2025-01-03T23:08:37.135444Z","iopub.status.idle":"2025-01-03T23:08:37.408968Z","shell.execute_reply.started":"2025-01-03T23:08:37.13542Z","shell.execute_reply":"2025-01-03T23:08:37.408145Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# hyper parametres. ","metadata":{}},{"cell_type":"code","source":"from lightgbm import LGBMRegressor\n\nmodel = LGBMRegressor(\n    n_estimators=100,\n    learning_rate=0.1,\n    max_depth=5,\n    num_leaves=31,\n    min_data_in_leaf=20,\n    lambda_l1=0.0,\n    lambda_l2=0.0,\n    random_state=42,\n    n_jobs=-1\n)\n\nmodel.fit(x, y)\n\nnew_pred = model.predict(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T23:15:19.127706Z","iopub.execute_input":"2025-01-03T23:15:19.128051Z","iopub.status.idle":"2025-01-03T23:15:25.763396Z","shell.execute_reply.started":"2025-01-03T23:15:19.128023Z","shell.execute_reply":"2025-01-03T23:15:25.762585Z"},"jupyter":{"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/playground-series-s4e12/sample_submission.csv')\n\nsubmission['Premium Amount'] = new_pred\n\nsubmission.to_csv('result_New.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-03T23:16:08.287859Z","iopub.execute_input":"2025-01-03T23:16:08.288246Z","iopub.status.idle":"2025-01-03T23:16:09.855712Z","shell.execute_reply.started":"2025-01-03T23:16:08.288217Z","shell.execute_reply":"2025-01-03T23:16:09.854966Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### from lightgbm import LGBMRegressor\n\nmodel = LGBMRegressor(\n    n_estimators=100,\n    learning_rate=0.1,\n    max_depth=5,\n    num_leaves=31,\n    min_data_in_leaf=20,\n    lambda_l1=0.0,\n    lambda_l2=0.0,\n    random_state=42,\n    n_jobs=-1\n)\n\nmodel.fit(x, y)\n\npredictions = model.predict(test)\nsubmission = pd.DataFrame({'id': test_data['id'], 'Premium Amount': predictions})\nsubmission.to_csv('submission.csv', index=False)Finally, we made it! 🎉 I hope this journey has been helpful and insightful for you. Whether you're just starting out or already deep into data science, I trust that the code, explanations, and tips provided will serve you well in your projects. Keep coding, keep learning, and remember: every line of code brings you one step closer to mastery. Thank you for trusting me to be part of your journey! 🚀","metadata":{}}]}