{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"},{"sourceId":214226817,"sourceType":"kernelVersion"}],"dockerImageVersionId":30823,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# First Place - Single Model - CV 1.016 LB 1.016 - Wow!\nI'm excited to share this notebook which achieves `Gold Medal Single Model` in Kaggle's December 2024 Insurance playground competition! Often in Kaggle competitons, it requires an ensemble of many models to achieve gold medal. Therefore it is an exciting accomplishment to build a single model which achieves gold medal!\n\nThe secret to this model's success is using [RAPIDS cuDF-Pandas][2] to engineer features. Using a for-loop, we searched thousands of combinations of original columns to find combinations that when combined with Target Encoding and Count Encoding improved CV and LB score. I shared the concept of combining columns with Target Encoding in September's competition [here][1]. (In that competition, we used all combinations of 2 and 3 because there were only 286 total).\n\nIn this competition, there were 145k combinations (of 2,3,4,5,6), so we can't use them all. We must search for the best ones. Searching was done using [RAPIDS cuDF-Pandas][2] and GPU. In total there are 145,000 combinations of 2,3,4,5,6 columns that can be extracted from the original 23 columns. For each of these 145k combinations, we must combine the dataframe columns and compute groupby aggregate target encoding and count encoding. Then we have to a train a model with the new features added. This requires lots of time and was made possible by using the blazing speed of GPU! We found 170 powerful combos and share the top 20 in this notebook.\n\n# How To Improve\nThis notebook takes under 2 hours to run on Kaggle's T4 GPU 16GB and achieves `CV = 1.019` and `LB = 1.019`. We can improve this model by adding more features. My final submission for December's 2024 playground competition adds 400 more engineered features. We can also improve CV and LB by decreasing learning rate by a factor of 10x and increasing number of iterations by a factor of 10x. And when performing TE, we can use 10 inner nested folds instead of 5. With these changes, my final submission achieves `CV = 1.016` and `LB = 1.016`, WOW!\n\n[1]: https://www.kaggle.com/code/cdeotte/rapids-cuml-lasso-lb-0-72500-cv-0-72800\n[2]: https://rapids.ai/cudf-pandas/","metadata":{}},{"cell_type":"markdown","source":"# Zero Code Change GPU Acceleration!\nThe following single magic command converts all subsequent CPU Pandas calls to GPU RAPIDS cuDF calls! This gives us a huge speedup! Woohoo! (Note in order to import RAPIDS cuDF 24.12, we attach `Utility Script` [here][1] to our notebook)\n\n[1]: https://www.kaggle.com/code/cdeotte/rapids-cudf-24-12-cuml-24-12","metadata":{}},{"cell_type":"code","source":"%load_ext cudf.pandas","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T02:43:16.80436Z","iopub.execute_input":"2024-12-30T02:43:16.804656Z","iopub.status.idle":"2024-12-30T02:43:16.809747Z","shell.execute_reply.started":"2024-12-30T02:43:16.804635Z","shell.execute_reply":"2024-12-30T02:43:16.809016Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Load Train and Test\nWe load train and test data. And we convert datatime column into year, month, day, day of week, and elapsed seconds since 1970.","metadata":{}},{"cell_type":"code","source":"import numpy as np, pandas as pd\npd.set_option('display.max_columns', 500)\n\ntrain = pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\ntrain[\"Policy Start Date\"] = pd.to_datetime( train[\"Policy Start Date\"] )\ntrain[\"year\"] = train[\"Policy Start Date\"].dt.year.astype(\"float32\")\ntrain[\"month\"] = train[\"Policy Start Date\"].dt.month.astype(\"float32\")\ntrain[\"day\"] = train[\"Policy Start Date\"].dt.day.astype(\"float32\")\ntrain[\"dow\"] = train[\"Policy Start Date\"].dt.dayofweek.astype(\"float32\")\ntrain[\"seconds\"] = (train[\"Policy Start Date\"].astype(\"int64\") // 10**9).astype(\"float32\")\ntrain[\"y\"] = np.log1p( train[\"Premium Amount\"] )\nprint( train.shape )\ntrain.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T02:43:16.810646Z","iopub.execute_input":"2024-12-30T02:43:16.810939Z","iopub.status.idle":"2024-12-30T02:43:17.091968Z","shell.execute_reply.started":"2024-12-30T02:43:16.810908Z","shell.execute_reply":"2024-12-30T02:43:17.091291Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test = pd.read_csv(\"/kaggle/input/playground-series-s4e12/test.csv\")\ntest[\"Policy Start Date\"] = pd.to_datetime( test[\"Policy Start Date\"] )\ntest[\"year\"] = test[\"Policy Start Date\"].dt.year.astype(\"float32\")\ntest[\"month\"] = test[\"Policy Start Date\"].dt.month.astype(\"float32\")\ntest[\"day\"] = test[\"Policy Start Date\"].dt.day.astype(\"float32\")\ntest[\"dow\"] = test[\"Policy Start Date\"].dt.dayofweek.astype(\"float32\")\ntest[\"seconds\"] = (test[\"Policy Start Date\"].astype(\"int64\") // 10**9).astype(\"float32\")\nprint( test.shape )\ntest.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T02:43:17.140941Z","iopub.execute_input":"2024-12-30T02:43:17.141219Z","iopub.status.idle":"2024-12-30T02:43:17.333315Z","shell.execute_reply.started":"2024-12-30T02:43:17.141195Z","shell.execute_reply":"2024-12-30T02:43:17.332484Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Label Encode Categorical Features\nWe identify and label encode all categorical features. We convert all data to either `int32` or `float32` for reduced memory utilization. We identify which numerical or categorical features have cardinality greater than 8 for special attention later.","metadata":{}},{"cell_type":"code","source":"RMV = [\"id\",\"Policy Start Date\",\"Premium Amount\",\"y\"]\nFEATURES = [c for c in train.columns if not c in RMV]\ncombined = pd.concat([train,test],axis=0,ignore_index=True)\n\nCATS = []\nHIGH_CARDINALITY = []\nprint(f\"THE {len(FEATURES)} BASIC FEATURES ARE:\")\n\nfor c in FEATURES:\n    ftype = \"numerical\"\n    if combined[c].dtype==\"object\":\n        CATS.append(c)\n        combined[c] = combined[c].fillna(\"NAN\")\n        combined[c],_ = combined[c].factorize()\n        combined[c] -= combined[c].min()\n        ftype = \"categorical\"\n    if combined[c].dtype==\"int64\":\n        combined[c] = combined[c].astype(\"int32\")\n    elif combined[c].dtype==\"float64\":\n        combined[c] = combined[c].astype(\"float32\")\n        \n    n = combined[c].nunique()\n    print(f\"{c} ({ftype}) with {n} unique values\")\n    if n>=9: HIGH_CARDINALITY.append(c)\n    \ntrain = combined.iloc[:len(train)].copy()\ntest = combined.iloc[len(train):].reset_index(drop=True).copy()\n\nprint(\"\\nTHE FOLLOWING HAVE 9 OR MORE UNIQUE VALUES:\", HIGH_CARDINALITY )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T02:43:17.334326Z","iopub.execute_input":"2024-12-30T02:43:17.334674Z","iopub.status.idle":"2024-12-30T02:43:17.843316Z","shell.execute_reply.started":"2024-12-30T02:43:17.334648Z","shell.execute_reply":"2024-12-30T02:43:17.842599Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Feature Engineering 120 Features!\nIn another notebook, we ran a `for-loop` to try different combinations of (2,3,4,5,or 6) columns. Each iteration,we added a random new combination of columns and evaluated whether it improved CV score or not. If it improved CV score we saved it to a list. This is called `forward feature selection`. By using GPU (RAPIDS cuDF-Pandas), we were able to try thousands of combinations quickly! We found 20 powerful column combinations defined below!\n\nFor each of these column combinations, we will create 6 features, namely TE (target encode) mean, TE median, TE min, TE max, TE nunique, and CE (count encode). This generates 120 features!","metadata":{}},{"cell_type":"code","source":"lists2 = [['Annual Income', 'Health Score'], ['Credit Score', 'Health Score'], ['Customer Feedback', 'Gender', 'Marital Status', 'Occupation', 'Smoking Status', 'year'], ['Exercise Frequency', 'Health Score'], ['Health Score', 'Marital Status'], ['Education Level', 'Gender', 'Health Score'], ['Health Score', 'Occupation'], ['Age', 'Health Score'], ['Health Score', 'dow'], ['Age', 'Exercise Frequency', 'Location'], ['Health Score', 'Smoking Status', 'month'], ['Health Score', 'Location', 'Policy Type'], ['Health Score', 'Insurance Duration'], ['Health Score', 'Number of Dependents'], ['Customer Feedback', 'Exercise Frequency', 'Previous Claims', 'Property Type', 'dow'], ['Customer Feedback', 'Health Score'], ['Health Score', 'Property Type'], ['Health Score', 'day', 'seconds'], ['Health Score', 'year'], ['Age', 'Gender', 'Insurance Duration', 'year']]\nprint(f\"We have {len(lists2)} powerful combination of columns!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T01:07:53.787905Z","iopub.execute_input":"2024-12-23T01:07:53.788153Z","iopub.status.idle":"2024-12-23T01:07:53.806571Z","shell.execute_reply.started":"2024-12-23T01:07:53.788116Z","shell.execute_reply":"2024-12-23T01:07:53.805793Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Target Encoding Function\nThis function is written using CPU Pandas (based on code [here][2]). However when we add `%load_ext cudf.pandas` magic at the start of our notebook, this notebook will be executed using GPU [RAPIDS cuDF][7] and made faster! Based on the number of rows of train data we are transforming, we can witness speed ups varying from 10x to 100x! (See notebook [here][1] to see how much speed up we can expect using GPU vs CPU for various dataframe sizes when executing various dataframe operations).\n\nFor more information about Target Encoding, see blog [here][2] and tutorial [here][3]. [KGMON][6] Theo published a notebook [here][4] and [KGMON][6] will present a live workshop at NVIDIA GTC in March 2025 [here][5]. We hope to see you there!\n\n[1]: https://www.kaggle.com/code/cdeotte/compare-cpu-dataframes-to-gpu-dataframes\n[2]: https://medium.com/rapids-ai/target-encoding-with-rapids-cuml-do-more-with-your-categorical-data-8c762c79e784\n[3]: https://github.com/rapidsai/deeplearning/blob/main/RecSys2020Tutorial/03_3_TargetEncoding.ipynb\n[4]: https://www.kaggle.com/code/theoviel/explaining-and-accelerating-target-encoding\n[5]: https://www.nvidia.com/gtc/sessions/gtc-data-science/\n[6]: https://www.nvidia.com/en-us/ai-data-science/kaggle-grandmasters/\n[7]: https://docs.rapids.ai/install/","metadata":{}},{"cell_type":"code","source":"def target_encode(train, valid, test, col, target=\"y\", kfold=5, smooth=20, agg=\"mean\"):\n\n    train['kfold'] = ((train.index) % kfold)\n    col_name = '_'.join(col)\n    train[f'TE_{agg.upper()}_' + col_name] = 0.\n    for i in range(kfold):\n        \n        df_tmp = train[train['kfold']!=i]\n        if agg==\"mean\": mn = train[target].mean()\n        elif agg==\"median\": mn = train[target].median()\n        elif agg==\"min\": mn = train[target].min()\n        elif agg==\"max\": mn = train[target].max()\n        elif agg==\"nunique\": mn = 0\n        df_tmp = df_tmp[col + [target]].groupby(col).agg([agg, 'count']).reset_index()\n        df_tmp.columns = col + [agg, 'count']\n        if agg==\"nunique\":\n            df_tmp['TE_tmp'] = df_tmp[agg] / df_tmp['count']\n        else:\n            df_tmp['TE_tmp'] = ((df_tmp[agg]*df_tmp['count'])+(mn*smooth)) / (df_tmp['count']+smooth)\n        df_tmp_m = train[col + ['kfold', f'TE_{agg.upper()}_' + col_name]].merge(df_tmp, how='left', left_on=col, right_on=col)\n        df_tmp_m.loc[df_tmp_m['kfold']==i, f'TE_{agg.upper()}_' + col_name] = df_tmp_m.loc[df_tmp_m['kfold']==i, 'TE_tmp']\n        train[f'TE_{agg.upper()}_' + col_name] = df_tmp_m[f'TE_{agg.upper()}_' + col_name].fillna(mn).values  \n    \n    df_tmp = train[col + [target]].groupby(col).agg([agg, 'count']).reset_index()\n    if agg==\"mean\": mn = train[target].mean()\n    elif agg==\"median\": mn = train[target].median()\n    elif agg==\"min\": mn = train[target].min()\n    elif agg==\"max\": mn = train[target].max()\n    elif agg==\"nunique\": mn = 0\n    df_tmp.columns = col + [agg, 'count']\n    if agg==\"nunique\":\n        df_tmp['TE_tmp'] = df_tmp[agg] / df_tmp['count']\n    else:\n        df_tmp['TE_tmp'] = ((df_tmp[agg]*df_tmp['count'])+(mn*smooth)) / (df_tmp['count']+smooth)\n    df_tmp_m = valid[col].merge(df_tmp, how='left', left_on=col, right_on=col)\n    valid[f'TE_{agg.upper()}_' + col_name] = df_tmp_m['TE_tmp'].fillna(mn).values\n    valid[f'TE_{agg.upper()}_' + col_name] = valid[f'TE_{agg.upper()}_' + col_name].astype(\"float32\")\n\n    df_tmp_m = test[col].merge(df_tmp, how='left', left_on=col, right_on=col)\n    test[f'TE_{agg.upper()}_' + col_name] = df_tmp_m['TE_tmp'].fillna(mn).values\n    test[f'TE_{agg.upper()}_' + col_name] = test[f'TE_{agg.upper()}_' + col_name].astype(\"float32\")\n\n    train = train.drop('kfold', axis=1)\n    train[f'TE_{agg.upper()}_' + col_name] = train[f'TE_{agg.upper()}_' + col_name].astype(\"float32\")\n\n    return(train, valid, test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T01:07:53.807211Z","iopub.execute_input":"2024-12-23T01:07:53.807462Z","iopub.status.idle":"2024-12-23T01:07:53.821387Z","shell.execute_reply.started":"2024-12-23T01:07:53.807442Z","shell.execute_reply":"2024-12-23T01:07:53.820622Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Train XGBoost\nWe train a 20-fold XGBoost model using the 23 basic features and 120 feature engineered features. We will also encode the basic 23 features with some TE and CE. If you want to see the train and validation scores from the training process, click the \"Show hidden output\" at the end of the second code cell below.","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBRegressor\nimport xgboost as xgb, time\nprint(f\"Using XGBoost version\",xgb.__version__)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T01:07:53.822157Z","iopub.execute_input":"2024-12-23T01:07:53.822387Z","iopub.status.idle":"2024-12-23T01:07:55.996108Z","shell.execute_reply.started":"2024-12-23T01:07:53.822352Z","shell.execute_reply":"2024-12-23T01:07:55.995336Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n\nFOLDS = 20\nfrom sklearn.model_selection import KFold\nkf = KFold(n_splits=FOLDS, shuffle=True, random_state=42)\n\noof = np.zeros(len(train))\npred = np.zeros(len(test))\n\nfor i, (train_index, test_index) in enumerate(kf.split(train)):\n\n    print(\"#\"*25)\n    print(f\"### Fold {i+1}\")\n    print(\"#\"*25)\n    \n    x_train = train.loc[train_index,FEATURES+[\"y\"] ].copy()\n    y_train = train.loc[train_index,\"y\"]\n    x_valid = train.loc[test_index,FEATURES].copy()\n    y_valid = train.loc[test_index,\"y\"]\n    x_test = test[FEATURES].copy()\n\n    start = time.time()\n    print(f\"FEATURE ENGINEER {len(FEATURES)} COLUMNS and {len(lists2)} GROUPS: \",end=\"\")\n    for j,f in enumerate(FEATURES+lists2):\n\n        if j<len(FEATURES): c = [f]\n        else: c = f \n        print(f\"({j+1}){c}\",\", \",end=\"\")\n\n        # LOW CARDINALITY FEATURES - TARGET ENCODE MEAN AND MEDIAN\n        x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=20, agg=\"mean\")\n        x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"median\")\n\n        # HIGH CARDINALITY FEATURES - TE MIN, MAX, NUNIQUE and CE\n        if (j>=len(FEATURES)) | (c[0] in HIGH_CARDINALITY):\n            x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"min\")\n            x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"max\")\n            x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"nunique\")\n    \n            # COUNT ENCODING (USING COMBINED TRAIN TEST)\n            tmp = combined.groupby(c).y.count()\n            nm = f\"CE_{'_'.join(c)}\"; tmp.name = nm\n            x_train = x_train.merge(tmp, on=c, how=\"left\")\n            x_valid = x_valid.merge(tmp, on=c, how=\"left\")\n            x_test = x_test.merge(tmp, on=c, how=\"left\")\n            x_train[nm] = x_train[nm].astype(\"int32\")\n            x_valid[nm] = x_valid[nm].astype(\"int32\")\n            x_test[nm] = x_test[nm].astype(\"int32\")\n            \n    end = time.time()\n    elapsed = end-start\n    print(f\"Feature engineering took {elapsed:.1f} seconds\")\n    x_train = x_train.drop(\"y\",axis=1)\n\n    model = XGBRegressor(\n        device=\"cuda\",\n        max_depth=8, \n        colsample_bytree=0.9, \n        subsample=0.9, \n        n_estimators=2_000, \n        learning_rate=0.01, \n        early_stopping_rounds=25,  \n        eval_metric=\"rmse\",\n    )\n    model.fit(\n        x_train, y_train,\n        eval_set=[(x_valid, y_valid)],   \n        verbose=100\n    )\n\n    # INFER OOF\n    oof[test_index] = model.predict(x_valid)\n    # INFER TEST\n    pred += model.predict(x_test)\n\n    m = np.sqrt(np.mean( (y_valid.to_numpy() - oof[test_index])**2.0 )) \n    print(f\" => Fold {i+1} RMSLE = {m:.5f}\")\n\n# COMPUTE AVERAGE TEST PREDS\npred /= FOLDS","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T01:21:06.761265Z","iopub.execute_input":"2024-12-23T01:21:06.761609Z","iopub.status.idle":"2024-12-23T01:33:52.07809Z","shell.execute_reply.started":"2024-12-23T01:21:06.761581Z","shell.execute_reply":"2024-12-23T01:33:52.077276Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Compute CV Score\nOur model has a local CV score of 1.019, wow!","metadata":{}},{"cell_type":"code","source":"m = np.sqrt(np.mean( (train.y.values - oof)**2.0 )) \nprint(f\"Overall CV RMSLE = {m:.5f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T01:34:33.323965Z","iopub.execute_input":"2024-12-23T01:34:33.324265Z","iopub.status.idle":"2024-12-23T01:34:33.332805Z","shell.execute_reply.started":"2024-12-23T01:34:33.32424Z","shell.execute_reply":"2024-12-23T01:34:33.332036Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# XGB Top 25 Feature Importance\nHere are the top 25 features of our XGBoost model based on \"weight\". (Alternatively, we can display top 25 by \"gain\" or \"cover\").","metadata":{}},{"cell_type":"code","source":"# Plot top 25 features by importance\nimport matplotlib.pyplot as plt\nfig, ax = plt.subplots(figsize=(10, 10))  # Adjust the figure size if needed\nxgb.plot_importance(\n    model,\n    ax=ax,\n    max_num_features=25,  # Display only the top 25 features\n    importance_type=\"weight\",  # Options: 'weight', 'gain', 'cover', 'total_gain', 'total_cover'\n)\nplt.title(\"XGB Top 25 Feature Importances\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T01:34:34.837689Z","iopub.execute_input":"2024-12-23T01:34:34.837984Z","iopub.status.idle":"2024-12-23T01:34:35.27976Z","shell.execute_reply.started":"2024-12-23T01:34:34.837963Z","shell.execute_reply":"2024-12-23T01:34:35.278925Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Create Submission CSV\nWe write our submission CSV file to disk and submit to competition.","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(\"/kaggle/input/playground-series-s4e12/sample_submission.csv\")\nsub[\"Premium Amount\"] = np.exp( pred )-1\nsub.to_csv(\"submission.csv\",index=False)\nprint( sub.shape )\nsub.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T01:34:37.917266Z","iopub.execute_input":"2024-12-23T01:34:37.917564Z","iopub.status.idle":"2024-12-23T01:34:39.507134Z","shell.execute_reply.started":"2024-12-23T01:34:37.91754Z","shell.execute_reply":"2024-12-23T01:34:39.506257Z"}},"outputs":[],"execution_count":null}]}