{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"},{"sourceId":216741403,"sourceType":"kernelVersion"}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduce","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    <strong>The purpose of this notebook is,</strong> I want to know where is the gap between me and the first place? I conducted research by trying different ways to compare CV scores.\n</p>\n\n<p style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    <strong>If you want to boost your score even further,</strong> try improving your CV and LB by reducing your learning rate by a factor of 10 and increasing the number of iterations by a factor of 10. When performing TE(target encoding), we can use 10 internally nested folds instead of 5. I didn't try it because it would have taken too long.\n</p>","metadata":{}},{"cell_type":"markdown","source":"# cuDF","metadata":{}},{"cell_type":"markdown","source":"### GPUT4×2 is needed to run it","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    这里我使用了 <strong>@cdeotte</strong> 发布的 <a href=\"https://www.kaggle.com/code/cdeotte/rapids-cudf-24-12-cuml-24-12\"><strong>[cuDF]</strong></a>\n</p>","metadata":{}},{"cell_type":"code","source":"%load_ext cudf.pandas","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:07.178779Z","iopub.execute_input":"2025-01-10T02:03:07.179147Z","iopub.status.idle":"2025-01-10T02:03:07.183725Z","shell.execute_reply.started":"2025-01-10T02:03:07.179115Z","shell.execute_reply":"2025-01-10T02:03:07.182786Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Date-time processing","metadata":{}},{"cell_type":"code","source":"import numpy as np, pandas as pd\npd.set_option('display.max_columns', 500)\n\ntrain = pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\ntrain[\"Policy Start Date\"] = pd.to_datetime( train[\"Policy Start Date\"] )\ntrain[\"year\"] = train[\"Policy Start Date\"].dt.year.astype(\"float32\")\ntrain[\"month\"] = train[\"Policy Start Date\"].dt.month.astype(\"float32\")\ntrain[\"day\"] = train[\"Policy Start Date\"].dt.day.astype(\"float32\")\ntrain[\"dow\"] = train[\"Policy Start Date\"].dt.dayofweek.astype(\"float32\")\ntrain[\"seconds\"] = (train[\"Policy Start Date\"].astype(\"int64\") // 10**9).astype(\"float32\")\n\ntrain[\"y\"] = np.log1p( train[\"Premium Amount\"] )\nprint( train.shape )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:07.198756Z","iopub.execute_input":"2025-01-10T02:03:07.198957Z","iopub.status.idle":"2025-01-10T02:03:07.464434Z","shell.execute_reply.started":"2025-01-10T02:03:07.198932Z","shell.execute_reply":"2025-01-10T02:03:07.463639Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test = pd.read_csv(\"/kaggle/input/playground-series-s4e12/test.csv\")\ntest[\"Policy Start Date\"] = pd.to_datetime( test[\"Policy Start Date\"] )\ntest[\"year\"] = test[\"Policy Start Date\"].dt.year.astype(\"float32\")\ntest[\"month\"] = test[\"Policy Start Date\"].dt.month.astype(\"float32\")\ntest[\"day\"] = test[\"Policy Start Date\"].dt.day.astype(\"float32\")\ntest[\"dow\"] = test[\"Policy Start Date\"].dt.dayofweek.astype(\"float32\")\ntest[\"seconds\"] = (test[\"Policy Start Date\"].astype(\"int64\") // 10**9).astype(\"float32\")\nprint( test.shape )\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:07.465744Z","iopub.execute_input":"2025-01-10T02:03:07.466066Z","iopub.status.idle":"2025-01-10T02:03:07.653093Z","shell.execute_reply.started":"2025-01-10T02:03:07.466016Z","shell.execute_reply":"2025-01-10T02:03:07.652247Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Numerical feature processing","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    <strong>I applied FOLDS = 20 when I ran all the steps.</strong>\n</p>\n<p style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    <strong>When this treatment is not used,</strong> Overall CV RMSLE = 1.01931\n</p>\n<p style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    <strong>When this method was used,</strong> the Overall CV RMSLE = 1.01852\n</p>\n<p style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    <strong>This is a treatment I explored during the game,</strong> which I post in <a href=\"https://www.kaggle.com/competitions/playground-series-s4e12/discussion/553025\"><strong>[this discussion]</strong></a>, and is helpful for scoring\n</p>\n<p style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    <strong>Wow, this is really great.</strong>\n</p>","metadata":{}},{"cell_type":"code","source":"\ndef log_transform_features(df, num_cols):\n    X_num = df[num_cols]\n    X_num_log = pd.DataFrame(np.log1p(X_num), columns=X_num.columns)\n    df[num_cols] = X_num_log\n    return df\n    \nlog_transform_cols = ['Annual Income', 'Previous Claims']\n\n\ntrain = log_transform_features(train, log_transform_cols)\ntest = log_transform_features(test, log_transform_cols)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:07.653988Z","iopub.execute_input":"2025-01-10T02:03:07.654323Z","iopub.status.idle":"2025-01-10T02:03:07.85721Z","shell.execute_reply.started":"2025-01-10T02:03:07.65429Z","shell.execute_reply":"2025-01-10T02:03:07.856533Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# classification","metadata":{}},{"cell_type":"code","source":"RMV = [\"id\",\"Policy Start Date\",\"Premium Amount\",\"y\"]\nFEATURES = [c for c in train.columns if not c in RMV]\ncombined = pd.concat([train,test],axis=0,ignore_index=True)\n\nCATS = []\nHIGH_CARDINALITY = []\nprint(f\"THE {len(FEATURES)} BASIC FEATURES ARE:\")\n\nfor c in FEATURES:\n    ftype = \"numerical\"\n    if combined[c].dtype==\"object\":\n        CATS.append(c)\n        combined[c] = combined[c].fillna(\"NAN\")\n        combined[c],_ = combined[c].factorize()\n        combined[c] -= combined[c].min()\n        ftype = \"categorical\"\n    if combined[c].dtype==\"int64\":\n        combined[c] = combined[c].astype(\"int32\")\n    elif combined[c].dtype==\"float64\":\n        combined[c] = combined[c].astype(\"float32\")\n        \n    n = combined[c].nunique()\n    print(f\"{c} ({ftype}) with {n} unique values\")\n    if n>=9: HIGH_CARDINALITY.append(c)\n    \ntrain = combined.iloc[:len(train)].copy()\ntest = combined.iloc[len(train):].reset_index(drop=True).copy()\n\nprint(\"\\nTHE FOLLOWING HAVE 9 OR MORE UNIQUE VALUES:\", HIGH_CARDINALITY )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:07.857801Z","iopub.execute_input":"2025-01-10T02:03:07.858043Z","iopub.status.idle":"2025-01-10T02:03:08.312983Z","shell.execute_reply.started":"2025-01-10T02:03:07.858024Z","shell.execute_reply":"2025-01-10T02:03:08.31214Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# New feature","metadata":{}},{"cell_type":"code","source":"lists2 = [['Annual Income', 'Health Score'], ['Credit Score', 'Health Score'], \n          ['Customer Feedback', 'Gender', 'Marital Status', 'Occupation', 'Smoking Status', 'year'], \n          ['Exercise Frequency', 'Health Score'], ['Health Score', 'Marital Status'], \n          ['Education Level', 'Gender', 'Health Score'], ['Health Score', 'Occupation'], \n          ['Age', 'Health Score'], ['Health Score', 'dow'], ['Age', 'Exercise Frequency', 'Location'], \n          ['Health Score', 'Smoking Status', 'month'], ['Health Score', 'Location', 'Policy Type'], \n          ['Health Score', 'Insurance Duration'], ['Health Score', 'Number of Dependents'], \n          ['Customer Feedback', 'Exercise Frequency', 'Previous Claims', 'Property Type', 'dow'], \n          ['Customer Feedback', 'Health Score'], ['Health Score', 'Property Type'], \n          ['Health Score', 'day', 'seconds'], ['Health Score', 'year'], ['Age', 'Gender', 'Insurance Duration', 'year']]\nprint(f\"We have {len(lists2)} powerful combination of columns!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:08.313817Z","iopub.execute_input":"2025-01-10T02:03:08.314133Z","iopub.status.idle":"2025-01-10T02:03:08.319733Z","shell.execute_reply.started":"2025-01-10T02:03:08.314108Z","shell.execute_reply":"2025-01-10T02:03:08.318988Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Target encode","metadata":{}},{"cell_type":"code","source":"def target_encode(train, valid, test, col, target=\"y\", kfold=5, smooth=20, agg=\"mean\"):\n\n    train['kfold'] = ((train.index) % kfold)\n    col_name = '_'.join(col)\n    train[f'TE_{agg.upper()}_' + col_name] = 0.\n    for i in range(kfold):\n        \n        df_tmp = train[train['kfold']!=i]\n        if agg==\"mean\": mn = train[target].mean()\n        elif agg==\"median\": mn = train[target].median()\n        elif agg==\"min\": mn = train[target].min()\n        elif agg==\"max\": mn = train[target].max()\n        elif agg==\"nunique\": mn = 0\n        df_tmp = df_tmp[col + [target]].groupby(col).agg([agg, 'count']).reset_index()\n        df_tmp.columns = col + [agg, 'count']\n        if agg==\"nunique\":\n            df_tmp['TE_tmp'] = df_tmp[agg] / df_tmp['count']\n        else:\n            df_tmp['TE_tmp'] = ((df_tmp[agg]*df_tmp['count'])+(mn*smooth)) / (df_tmp['count']+smooth)\n        df_tmp_m = train[col + ['kfold', f'TE_{agg.upper()}_' + col_name]].merge(df_tmp, how='left', left_on=col, right_on=col)\n        df_tmp_m.loc[df_tmp_m['kfold']==i, f'TE_{agg.upper()}_' + col_name] = df_tmp_m.loc[df_tmp_m['kfold']==i, 'TE_tmp']\n        train[f'TE_{agg.upper()}_' + col_name] = df_tmp_m[f'TE_{agg.upper()}_' + col_name].fillna(mn).values  \n    \n    df_tmp = train[col + [target]].groupby(col).agg([agg, 'count']).reset_index()\n    if agg==\"mean\": mn = train[target].mean()\n    elif agg==\"median\": mn = train[target].median()\n    elif agg==\"min\": mn = train[target].min()\n    elif agg==\"max\": mn = train[target].max()\n    elif agg==\"nunique\": mn = 0\n    df_tmp.columns = col + [agg, 'count']\n    if agg==\"nunique\":\n        df_tmp['TE_tmp'] = df_tmp[agg] / df_tmp['count']\n    else:\n        df_tmp['TE_tmp'] = ((df_tmp[agg]*df_tmp['count'])+(mn*smooth)) / (df_tmp['count']+smooth)\n    df_tmp_m = valid[col].merge(df_tmp, how='left', left_on=col, right_on=col)\n    valid[f'TE_{agg.upper()}_' + col_name] = df_tmp_m['TE_tmp'].fillna(mn).values\n    valid[f'TE_{agg.upper()}_' + col_name] = valid[f'TE_{agg.upper()}_' + col_name].astype(\"float32\")\n\n    df_tmp_m = test[col].merge(df_tmp, how='left', left_on=col, right_on=col)\n    test[f'TE_{agg.upper()}_' + col_name] = df_tmp_m['TE_tmp'].fillna(mn).values\n    test[f'TE_{agg.upper()}_' + col_name] = test[f'TE_{agg.upper()}_' + col_name].astype(\"float32\")\n\n    train = train.drop('kfold', axis=1)\n    train[f'TE_{agg.upper()}_' + col_name] = train[f'TE_{agg.upper()}_' + col_name].astype(\"float32\")\n\n    return(train, valid, test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:08.320574Z","iopub.execute_input":"2025-01-10T02:03:08.320857Z","iopub.status.idle":"2025-01-10T02:03:08.337159Z","shell.execute_reply.started":"2025-01-10T02:03:08.320828Z","shell.execute_reply":"2025-01-10T02:03:08.336333Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:08.339434Z","iopub.execute_input":"2025-01-10T02:03:08.339651Z","iopub.status.idle":"2025-01-10T02:03:08.392944Z","shell.execute_reply.started":"2025-01-10T02:03:08.339632Z","shell.execute_reply":"2025-01-10T02:03:08.392086Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:08.394172Z","iopub.execute_input":"2025-01-10T02:03:08.394457Z","iopub.status.idle":"2025-01-10T02:03:08.442891Z","shell.execute_reply.started":"2025-01-10T02:03:08.394429Z","shell.execute_reply":"2025-01-10T02:03:08.442264Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:08.443674Z","iopub.execute_input":"2025-01-10T02:03:08.443896Z","iopub.status.idle":"2025-01-10T02:03:08.462727Z","shell.execute_reply.started":"2025-01-10T02:03:08.443865Z","shell.execute_reply":"2025-01-10T02:03:08.462089Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# XGB","metadata":{}},{"cell_type":"markdown","source":"### If you want to try it, please drop the \"\"\" after it.","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBRegressor\nimport xgboost as xgb, time\nprint(f\"Using XGBoost version\",xgb.__version__)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:08.463414Z","iopub.execute_input":"2025-01-10T02:03:08.463613Z","iopub.status.idle":"2025-01-10T02:03:08.468251Z","shell.execute_reply.started":"2025-01-10T02:03:08.463595Z","shell.execute_reply":"2025-01-10T02:03:08.467436Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-size:1.2em; line - height:1.5; color:#333;\">\n    <strong>Here, I removed the newly created features from lists2 and processed 23 columns of features to observe the CV score gap between them.</strong>\n</p>\n<p style=\"font-size:1.2em; line - height:1.5; color:#333;\">\n    <strong>When I applied FOLDS = 2,</strong> the scores for all 20 features were the same with or without Folds = 1.02840. Does not reflect the effect of the 20 new features added.\n</p>\n<p style=\"font-size:1.2em; line - height:1.5; color:#333;\">\n    <strong>So when I used FOLDS = 20,</strong> I was able to reflect the effects of the 20 new features.\n</p>\n<p style=\"font-size:1.2em; line - height:1.5; color:#333;\">\n    <strong>When using 20 new features,</strong> Overall CV RMSLE = 1.01931.\n</p>\n<p style=\"font-size:1.2em; line - height:1.5; color:#333;\">\n    <strong>When the 20 new features are not used,</strong> Overall CV RMSLE = 1.02279.\n</p>","metadata":{}},{"cell_type":"code","source":"\"\"\"\n%%time\n\n#You can set FOLDS = 20 or FOLDS = 2\n\n#FOLDS = 20\nFOLDS = 2\nfrom sklearn.model_selection import KFold\nkf = KFold(n_splits=FOLDS, shuffle=True, random_state=42)\n\noof = np.zeros(len(train))\npred = np.zeros(len(test))\n\nfor i, (train_index, test_index) in enumerate(kf.split(train)):\n\n    print(\"#\"*25)\n    print(f\"### Fold {i+1}\")\n    print(\"#\"*25)\n    \n    x_train = train.loc[train_index,FEATURES+[\"y\"] ].copy()\n    y_train = train.loc[train_index,\"y\"]\n    x_valid = train.loc[test_index,FEATURES].copy()\n    y_valid = train.loc[test_index,\"y\"]\n    x_test = test[FEATURES].copy()\n\n    start = time.time()\n    print(f\"FEATURE ENGINEER {len(FEATURES)} COLUMNS: \",end=\"\")\n    for j,f in enumerate(FEATURES):\n\n        c = [f]\n        print(f\"({j+1}){c}\",\", \",end=\"\")\n\n        # LOW CARDINALITY FEATURES - TARGET ENCODE MEAN AND MEDIAN\n        x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=20, agg=\"mean\")\n        x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"median\")\n\n        # HIGH CARDINALITY FEATURES - TE MIN, MAX, NUNIQUE and CE\n        if c[0] in HIGH_CARDINALITY:\n            x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"min\")\n            x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"max\")\n            x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"nunique\")\n    \n            # COUNT ENCODING (USING COMBINED TRAIN TEST)\n            tmp = combined.groupby(c).y.count()\n            nm = f\"CE_{'_'.join(c)}\"; tmp.name = nm\n            x_train = x_train.merge(tmp, on=c, how=\"left\")\n            x_valid = x_valid.merge(tmp, on=c, how=\"left\")\n            x_test = x_test.merge(tmp, on=c, how=\"left\")\n            x_train[nm] = x_train[nm].astype(\"int32\")\n            x_valid[nm] = x_valid[nm].astype(\"int32\")\n            x_test[nm] = x_test[nm].astype(\"int32\")\n            \n    end = time.time()\n    elapsed = end-start\n    print(f\"Feature engineering took {elapsed:.1f} seconds\")\n    x_train = x_train.drop(\"y\",axis=1)\n\n    model = XGBRegressor(\n        device=\"cuda\",\n        max_depth=8, \n        colsample_bytree=0.9, \n        subsample=0.9, \n        n_estimators=2_000, \n        learning_rate=0.01, \n        early_stopping_rounds=25,  \n        eval_metric=\"rmse\",\n    )\n    model.fit(\n        x_train, y_train,\n        eval_set=[(x_valid, y_valid)],   \n        verbose=100\n    )\n\n    # INFER OOF\n    oof[test_index] = model.predict(x_valid)\n    # INFER TEST\n    pred += model.predict(x_test)\n\n    m = np.sqrt(np.mean( (y_valid.to_numpy() - oof[test_index])**2.0 )) \n    print(f\" => Fold {i+1} RMSLE = {m:.5f}\")\n\n# COMPUTE AVERAGE TEST PREDS\npred /= FOLDS\n\"\"\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:08.469143Z","iopub.execute_input":"2025-01-10T02:03:08.469464Z","iopub.status.idle":"2025-01-10T02:03:08.482494Z","shell.execute_reply.started":"2025-01-10T02:03:08.46943Z","shell.execute_reply":"2025-01-10T02:03:08.481691Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"在这里我将FOLDS = 20修改为了FOLDS = 2.\n\n当FOLDS=20时，Overall CV RMSLE = 1.01931\n\n当FOLDS=2时，Overall CV RMSLE = 1.02840\n对'Annual Income', 'Previous Claims'进行log转化，Overall CV RMSLE = 1.02457","metadata":{}},{"cell_type":"code","source":"\"\"\"\n%%time\n#FOLDS = 20\nFOLDS = 2\nfrom sklearn.model_selection import KFold\nkf = KFold(n_splits=FOLDS, shuffle=True, random_state=42)\n\noof = np.zeros(len(train))\npred = np.zeros(len(test))\n\nfor i, (train_index, test_index) in enumerate(kf.split(train)):\n\n    print(\"#\"*25)\n    print(f\"### Fold {i+1}\")\n    print(\"#\"*25)\n    \n    x_train = train.loc[train_index,FEATURES+[\"y\"] ].copy()\n    y_train = train.loc[train_index,\"y\"]\n    x_valid = train.loc[test_index,FEATURES].copy()\n    y_valid = train.loc[test_index,\"y\"]\n    x_test = test[FEATURES].copy()\n\n    start = time.time()\n    print(f\"FEATURE ENGINEER {len(FEATURES)} COLUMNS and {len(lists2)} GROUPS: \",end=\"\")\n    for j,f in enumerate(FEATURES+lists2):\n\n        if j<len(FEATURES): c = [f]\n        else: c = f \n        print(f\"({j+1}){c}\",\", \",end=\"\")\n\n        # LOW CARDINALITY FEATURES - TARGET ENCODE MEAN AND MEDIAN\n        x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=20, agg=\"mean\")\n        x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"median\")\n\n        # HIGH CARDINALITY FEATURES - TE MIN, MAX, NUNIQUE and CE\n        if (j>=len(FEATURES)) | (c[0] in HIGH_CARDINALITY):\n            x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"min\")\n            x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"max\")\n            x_train, x_valid, x_test = target_encode(x_train, x_valid, x_test, c, smooth=0, agg=\"nunique\")\n    \n            # COUNT ENCODING (USING COMBINED TRAIN TEST)\n            tmp = combined.groupby(c).y.count()\n            nm = f\"CE_{'_'.join(c)}\"; tmp.name = nm\n            x_train = x_train.merge(tmp, on=c, how=\"left\")\n            x_valid = x_valid.merge(tmp, on=c, how=\"left\")\n            x_test = x_test.merge(tmp, on=c, how=\"left\")\n            x_train[nm] = x_train[nm].astype(\"int32\")\n            x_valid[nm] = x_valid[nm].astype(\"int32\")\n            x_test[nm] = x_test[nm].astype(\"int32\")\n            \n    end = time.time()\n    elapsed = end-start\n    print(f\"Feature engineering took {elapsed:.1f} seconds\")\n    x_train = x_train.drop(\"y\",axis=1)\n\n    model = XGBRegressor(\n        device=\"cuda\",\n        max_depth=8, \n        colsample_bytree=0.9, \n        subsample=0.9, \n        n_estimators=2_000, \n        learning_rate=0.01, \n        early_stopping_rounds=25,  \n        eval_metric=\"rmse\",\n    )\n    model.fit(\n        x_train, y_train,\n        eval_set=[(x_valid, y_valid)],   \n        verbose=100\n    )\n\n    # INFER OOF\n    oof[test_index] = model.predict(x_valid)\n    # INFER TEST\n    pred += model.predict(x_test)\n\n    m = np.sqrt(np.mean( (y_valid.to_numpy() - oof[test_index])**2.0 )) \n    print(f\" => Fold {i+1} RMSLE = {m:.5f}\")\n\n# COMPUTE AVERAGE TEST PREDS\npred /= FOLDS\n\"\"\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T02:03:08.483398Z","iopub.execute_input":"2025-01-10T02:03:08.483689Z","iopub.status.idle":"2025-01-10T03:46:43.569163Z","shell.execute_reply.started":"2025-01-10T02:03:08.483661Z","shell.execute_reply":"2025-01-10T03:46:43.568212Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Verification score","metadata":{}},{"cell_type":"code","source":"\"\"\"\nm = np.sqrt(np.mean( (train.y.values - oof)**2.0 )) \nprint(f\"Overall CV RMSLE = {m:.5f}\")\n\"\"\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-10T03:46:43.569874Z","iopub.execute_input":"2025-01-10T03:46:43.570117Z","iopub.status.idle":"2025-01-10T03:46:43.588362Z","shell.execute_reply.started":"2025-01-10T03:46:43.570097Z","shell.execute_reply":"2025-01-10T03:46:43.587686Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Conclusion","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    <strong>I learned that the gap with the first place has these aspects:</strong>\n</p>\n<ol style=\"font-size:1.2em; line-height:1.5; color:#333;\">\n    <li>He added a new feature list</li>\n    <li>He uses TE(object coding) for features, which is a great way to do it, and of course, the effect is very good.</li>\n    <li>His verification method fits perfectly</li>\n</ol>","metadata":{}}]}