{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"},{"sourceId":9178166,"sourceType":"datasetVersion","datasetId":5547076},{"sourceId":210584397,"sourceType":"kernelVersion"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **FOREWORD**","metadata":{}},{"cell_type":"markdown","source":"This kernel is a second part to my [ML baseline public materials](https://www.kaggle.com/code/ravi20076/playgrounds4e11-imports-v1). My older kernel for Baseline models became a bit clunky and needed revision. <br>\nThis kernel is divided into 2 parts with scripts- <br>\n1. Utility script with relevant imports, package installations and model training script for all types of Playground assignments <br>\n2. Model kernel using the previous step as imported script and execution of the model <br>\n\nI wish to extend sincere thanks to the Kaggle community for the long standing support over the years! Thanks for all the feedback on my kernels and many thanks for your collective generosity!\n\n### **WHAT IS DIFFERENT HERE** <br>\n1. Usage of a utility script for imports and a general training class for all types of Playground datasets with(out) the original data <br>\n2. Common training class for regression, multiclass and binary problems <br>\n3. Ability to train single models one at a time <br>\n4. Ability to enter a pipeline object instead of a model object for training <br>\n5. Separate ensemble with facility for Optuna blending with normalised weights as output. User has full choice to implement his/ her ensemble method <br>\n6. Ability to skip early stopping if needed in the pipeline <br> \n7. Compatible with any scikit-learn model/ classical ML model <br>\n8. Ability to perform online full fit <br>\n9. Can be used for code competitions as well, returns fitted models as one of the outputs <br>\n10. Ability to load the dataset as per the choice of original dataset included/ excluded in the CV scheme <br>\n","metadata":{}},{"cell_type":"markdown","source":"### **COMPETITION AND DATASET DETAILS** <br>\n\nThis is a regression problem for the [Playground Series S4-E12](https://www.kaggle.com/competitions/playground-series-s4e12) competition. <br> **RMSLE score** is the evaluation metric and needs to be minimized <br>\n\nIn this baseline kernel, I start off with simple feature engneering and ML models to initiate the process. Let's delve deeper into the challenge as we move along and improve the process! <br>\n\nAll the best!","metadata":{}},{"cell_type":"markdown","source":"# **IMPORTS**","metadata":{}},{"cell_type":"code","source":"%%time \n\n!pip install -q -r /kaggle/input/playgrounds4e12-public-imports-v1/req_kaggle.txt\n\nexec(open('/kaggle/input/playgrounds4e12-public-imports-v1/myimports.py','r').read())\nexec(open('/kaggle/input/playgrounds4e12-public-imports-v1/training.py','r').read())\n\n%matplotlib inline\nprint()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T12:11:25.224371Z","iopub.execute_input":"2024-12-02T12:11:25.224881Z","iopub.status.idle":"2024-12-02T12:12:07.466898Z","shell.execute_reply.started":"2024-12-02T12:11:25.224829Z","shell.execute_reply":"2024-12-02T12:12:07.465656Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **CONFIGURATION**","metadata":{}},{"cell_type":"code","source":"%%time \n\nclass CFG:\n    \"\"\"\n    Configuration class for parameters and CV strategy for tuning and training\n    Some parameters may be unused here as this is a general configuration class\n    \"\"\";\n\n    # Data preparation:-\n    version_nb  = 1\n    model_id    = \"V1_5\"\n    model_label = \"ML\"\n\n    test_req           = False\n    test_sample_frac   = 0.01\n\n    gpu_switch         = \"OFF\"\n    state              = 42\n    target             = f\"PremiumAmount\"\n    grouper            = f\"\"\n    tgt_mapper         = {}\n\n    ip_path            = f\"/kaggle/input/playground-series-s4e12\"\n    op_path            = f\"/kaggle/working\"\n    orig_path          = f\"/kaggle/input/insurance-premium-prediction/Insurance Premium Prediction Dataset.csv\"\n\n    dtl_preproc_req    = True\n    ftre_plots_req     = True\n    ftre_imp_req       = True\n\n    nb_orig            = 0\n    orig_all_folds     = False\n\n    # Model Training:-\n    pstprcs_oof        = True\n    pstprcs_train      = True\n    pstprcs_test       = True\n    \n    ML                 = True\n    test_preds_req     = False\n\n    pseudo_lbl_req     = \"N\"\n    pseudolbl_up       = 0.975\n    pseudolbl_low      = 0.00\n\n    n_splits           = 3 if test_req == True else 10\n    n_repeats          = 1\n    nbrnd_erly_stp     = 0\n    mdlcv_mthd         = 'KF'\n\n    # Ensemble:-\n    ensemble_req       = True\n    optuna_req         = False\n    metric_obj         = 'minimize'\n    ntrials            = 10 if test_req == True else 300\n\n    # Global variables for plotting:-\n    grid_specs = {'visible'  : True,\n                  'which'    : 'both',\n                  'linestyle': '--',\n                  'color'    : 'lightgrey',\n                  'linewidth': 0.75\n                 }\n\n    title_specs = {'fontsize'   : 9,\n                   'fontweight' : 'bold',\n                   'color'      : '#992600',\n                  }\n\nPrintColor(f\"\\n---> Configuration done!\\n\")\n\ncv_selector = \\\n{\n \"RKF\"   : RKF(n_splits = CFG.n_splits, n_repeats= CFG.n_repeats, random_state= CFG.state),\n \"RSKF\"  : RSKF(n_splits = CFG.n_splits, n_repeats= CFG.n_repeats, random_state= CFG.state),\n \"SKF\"   : SKF(n_splits = CFG.n_splits, shuffle = True, random_state= CFG.state),\n \"KF\"    : KFold(n_splits = CFG.n_splits, shuffle = True, random_state= CFG.state),\n \"GKF\"   : GKF(n_splits = CFG.n_splits)\n}\n\ncollect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T12:12:07.468874Z","iopub.execute_input":"2024-12-02T12:12:07.469434Z","iopub.status.idle":"2024-12-02T12:12:07.680436Z","shell.execute_reply.started":"2024-12-02T12:12:07.4694Z","shell.execute_reply":"2024-12-02T12:12:07.679117Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"|Configuration parameter| Explanation| Data type| Sample values |  \n| ---------------------- | ------------------------------- | --------------------- | --------------- |\n| version_nb    | Version Number | int | 1 | \n| model_id      | Model ID    | string | V1_1 | \n| model_label   | Model Label | string | ML | \n| test_req      | Test Required| bool | True / False | \n| test_sample_frac| Test sampled fraction | int | 1000 |\n| gpu_switch      | Do we need GPU support | bool | True / False |\n| state           | Random state | int | 42 |\n| target          | Target column | str |  |\n| grouper         | CV grouper column | str |  |\n| ip_path, op_path | Data paths  | str | |\n| pstprcs_* | Do we need post-processing  | bool |True / False |\n| ML| Do we need machine learning models  | bool |True / False |\n| test_preds_req| Do we need test set predictions (training in inference kernel)  | bool |True / False |\n| pseudo_lbl_req| Pseudo label required?  | bool |True / False |\n| pseudo_lbl_* | Pseudo label cutoff | float | |\n| n_splits/ n_repeats | N-splits and repeats for CV scheme | int | 3/5/10|\n| nbrnd_erly_stp | Early stopping rounds | int | 40|\n| mdlcv_mthd | Model CV method | str | RSKF|\n| ensemble_req | Do we need ensemble | bool | True / False |\n| optuna_req   | Do we need optuna | bool | True / False |\n| metric_obj   | Metric direction | str | minimize/ maximize |\n| ntrials      | Trials | int | 300 |","metadata":{}},{"cell_type":"markdown","source":"# **PREPROCESSING**","metadata":{}},{"cell_type":"code","source":"%%time \n\nexec(open('/kaggle/input/playgrounds4e12-public-imports-v1/pp.py','r').read())\npp = Preprocessor()\npp.DoPreprocessing();","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T12:12:07.682428Z","iopub.execute_input":"2024-12-02T12:12:07.683036Z","iopub.status.idle":"2024-12-02T12:12:25.112819Z","shell.execute_reply.started":"2024-12-02T12:12:07.682968Z","shell.execute_reply":"2024-12-02T12:12:25.11165Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **DATA TRANSFORMS**","metadata":{}},{"cell_type":"code","source":"%%time \n\nclass FeatureMaker:\n    \"This class develops features as per the requirement and cleans up the dataset off outliers\"\n\n    def __init__(self):\n        self.target = CFG.target\n\n    def fit(self, X, y = None, **fit_params):\n        return self\n\n    def transform(\n        self, \n        X        : pd.DataFrame, y = None,  \n        cat_cols : list = [],\n        mode     : str  = \"Train\",\n        **params\n    ):\n        \"This method transforms the dataset based on additional columns and data cleaning\"\n\n        df = X.copy()\n\n        df[\"PolicyStartDate\"] = pd.to_datetime(df[\"PolicyStartDate\"])\n        \n        df[\"Month\"]       = df[\"PolicyStartDate\"].dt.month\n        df[\"Day\"]         = df[\"PolicyStartDate\"].dt.day\n        df[\"Week\"]        = df[\"PolicyStartDate\"].dt.isocalendar().week\n        df[\"Weekday\"]     = df[\"PolicyStartDate\"].dt.weekday\n        df['DaySin']      = np.sin(2 * np.pi * df['Day'] / 30)  \n        df['DayCos']      = np.cos(2 * np.pi * df['Day'] / 30)\n        df['WeekdaySin']  = np.sin(2 * np.pi * df['Weekday'] / 7)\n        df['WeekdayCos']  = np.cos(2 * np.pi * df['Weekday'] / 7)\n        \n        df['DaysSinceStart']  = \\\n        np.ceil(\n            (pd.to_datetime(\"12-31-2024\") - df[\"PolicyStartDate\"])/ pd.Timedelta(1, \"d\")\n        )\n        \n        df[\"Ratio_IncomeAge\"]   = np.clip(df[\"AnnualIncome\"] / df[\"Age\"], a_min = 1e-6, a_max = 1e9)\n        df[\"Score\"]             = df[\"CreditScore\"] + df[\"HealthScore\"]\n        \n        df = df.drop(\"PolicyStartDate\", axis=1, errors = \"ignore\")\n        \n        if mode == \"Train\" :\n            cat_cols = \\\n            (df.\n             drop([\"Source\", \"PolicyStartDate\"], axis=1, errors = \"ignore\").\n             select_dtypes([\"object\", pd.StringDtype, \"category\"]).\n             columns\n            )\n\n        for col in cat_cols:\n            try:\n                df[col] = \\\n                df[col].astype(pd.StringDtype).fillna(\"missing\").astype(\"category\")\n            except:\n                df[col] = df[col].fillna(\"missing\").astype(\"category\")\n        \n        return (df, list(cat_cols))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T12:12:25.115139Z","iopub.execute_input":"2024-12-02T12:12:25.115477Z","iopub.status.idle":"2024-12-02T12:12:25.127733Z","shell.execute_reply.started":"2024-12-02T12:12:25.115446Z","shell.execute_reply":"2024-12-02T12:12:25.126514Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time \n\nXtrain = pp.train.copy()\n\nif CFG.test_req:\n    Xtrain = Xtrain.groupby([\"Source\", CFG.target], as_index = False).sample(frac = CFG.test_sample_frac)\n    Xtrain.index = range(len(Xtrain))\n    PrintColor(f\"---> Syntax check mode - shape = {Xtrain.shape}\", color = Fore.RED)\n    \nytrain = Xtrain[CFG.target]\nXtrain = Xtrain.drop([CFG.target, CFG.grouper], axis= 1, errors = \"ignore\")  \n\nxform = FeatureMaker()\nxform.fit(Xtrain, ytrain);\n\nXtrain, cat_cols  = xform.transform(Xtrain, cat_cols = [], mode = \"Train\")\nXtest, _          = xform.transform(pp.test, cat_cols = cat_cols, mode = \"Test\")\n\nPrintColor(f\"\\n---> Shapes = {Xtrain.shape} {ytrain.shape} {Xtest.shape}\")\n\n# Initializing the cv scheme:-\ncv = cv_selector[CFG.mdlcv_mthd]\n\nif CFG.nb_orig > 0:\n    all_df = []\n    \n    for mysource in [\"Competition\", \"Original\"]:\n        df = pd.concat([Xtrain.loc[Xtrain.Source == mysource], ytrain], axis=1, join = \"inner\")\n        df.index = range(len(df))\n        for fold_nb, (_, dev_idx) in enumerate(cv.split(df, df[CFG.target])):\n            df.loc[dev_idx, \"fold_nb\"] = fold_nb\n            \n        all_df.append(df)      \n    ygrp = pd.concat(all_df, axis=0, ignore_index = True)[\"fold_nb\"].astype(np.uint8)\n                      \nelse:\n    df = Xtrain.loc[Xtrain.Source == \"Competition\"]\n    df.index = range(len(df))\n    \n    for fold_nb, (_, dev_idx) in enumerate(cv.split(df, ytrain.iloc[df.index])):\n        df.loc[dev_idx, \"fold_nb\"] = fold_nb \n    ygrp = df[\"fold_nb\"].astype(np.uint8)\n\nytrain = np.log1p(ytrain)\n\ncollect();\nprint();\n\n_ = utils.CleanMemory()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T12:12:25.129229Z","iopub.execute_input":"2024-12-02T12:12:25.129631Z","iopub.status.idle":"2024-12-02T12:12:29.681008Z","shell.execute_reply.started":"2024-12-02T12:12:25.129592Z","shell.execute_reply":"2024-12-02T12:12:29.67971Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **MODEL TRAINING**","metadata":{}},{"cell_type":"code","source":"%%time \n\ntry:\n    l = MyLogger()\n    l.init(logging_lbl = \"lightgbm_custom\")\n    lgb.register_logger(l)\nexcept:\n    pass\n\ndef mymetric(ytrue, ypred):\n    \"\"\"\n    Custom evaluation metric to track RMSLE during training.\n    \n    Args:\n        ytrue - true values\n        ypred - prediction values\n    \"\"\"\n    return ('RMSLE', utils.ScoreMetric(ytrue, np.clip(ypred, 20, 4999)), False)\n    \nMdl_Master = \\\n{   \n f'CB1R' : CBR(**{    \"loss_function\"         : \"RMSE\",\n                      'task_type'             : \"GPU\" if CFG.gpu_switch == \"ON\" else \"CPU\",\n                      'learning_rate'         : 0.035,\n                      'iterations'            : 200,\n                      'max_bin'               : 16384,\n                      'max_depth'             : 8,\n                      'colsample_bylevel'     : 0.70,\n                      'l2_leaf_reg'           : 0.25,\n                      'random_strength'       : 0.20,\n                      'verbose'               : 0,\n                      'random_state'          : CFG.state,\n                      'cat_features'          : cat_cols,\n                     }\n                  ),\n    \n f'LGBM1R' : LGBMR(**{\"objective\"           : \"regression_l2\",\n                      'device'              : \"gpu\" if CFG.gpu_switch == \"ON\" else \"cpu\",\n                      'metric'              : \"custom\",\n                      'learning_rate'       : 0.03,\n                      'n_estimators'        : 200,\n                      'max_bin'             : 8192,\n                      'max_depth'           : 8,\n                      'num_leaves'          : 128,\n                      'min_data_in_leaf'    : 128,\n                      'colsample_bytree'    : 0.70,\n                      'lambda_l1'           : 0.001,\n                      'lambda_l2'           : 0.01,\n                      'verbosity'           : -1,\n                      'random_state'        : CFG.state,\n                     }\n                  ),\n\n f'LGBM2R' : LGBMR(**{\"objective\"           : \"regression_l2\",\n                      'device'              : \"gpu\" if CFG.gpu_switch == \"ON\" else \"cpu\",\n                      'metric'              : \"custom\",\n                      'data_sample_strategy': 'goss',\n                      'learning_rate'       : 0.035,\n                      'n_estimators'        : 150,\n                      'max_bin'             : 4096,\n                      'max_depth'           : 8,\n                      'num_leaves'          : 64,\n                      'min_data_in_leaf'    : 64,\n                      'colsample_bytree'    : 0.75,\n                      'lambda_l1'           : 0.01,\n                      'lambda_l2'           : 0.01,\n                      'verbosity'           : -1,\n                      'random_state'        : CFG.state,\n                     }\n                  ),\n\n f'XGB1R' : XGBR(**{  \"objective\"             : \"reg:squarederror\",\n                      'device'                : \"cuda\" if CFG.gpu_switch == \"ON\" else \"cpu\",\n                      'disable_default_eval_metric' : True,\n                      'learning_rate'       : 0.035,\n                      'n_estimators'        : 200,\n                      'max_bin'             : 8192,\n                      'max_depth'           : 8,\n                      'colsample_bytree'    : 0.75,\n                      'colsample_bylevel'   : 0.70,\n                      'lambda_l1'           : 0.01,\n                      'lambda_l2'           : 0.01,\n                      'verbose'             : 0,\n                      'random_state'        : CFG.state,\n                      'enable_categorical'  : True,\n                     }\n                  ),\n}\n\n# Initializing model outputs\nOOF_Preds    = {}\nMdl_Preds    = {}\nFittedModels = {}\nFtreImp      = {}\nSelMdlCols   = {}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T12:12:29.682462Z","iopub.execute_input":"2024-12-02T12:12:29.682829Z","iopub.status.idle":"2024-12-02T12:12:29.70005Z","shell.execute_reply.started":"2024-12-02T12:12:29.682794Z","shell.execute_reply":"2024-12-02T12:12:29.69861Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n\n# Model training:-\ndrop_cols = [\"Source\", \"id\", \"Id\", \"Label\", CFG.target, \"fold_nb\"]\n\nfor method, mymodel in tqdm(Mdl_Master.items()):\n\n    PrintColor(f\"\\n{'=' * 20} {method.upper()} MODEL TRAINING {'=' * 20}\\n\")\n\n    md = \\\n    ModelTrainer(\n        problem_type   = \"regression\",\n        es             = CFG.nbrnd_erly_stp,\n        target         = CFG.target,\n        orig_req       = True if CFG.nb_orig > 0 else False,\n        orig_all_folds = CFG.orig_all_folds,\n        metric_lbl     = \"rmsle\",\n        drop_cols      = drop_cols,\n        pp_preds       = CFG.pstprcs_oof,\n        )\n\n    sel_mdl_cols = list(Xtest.columns) \n    PrintColor(\n        f\"Selected columns = {len(sel_mdl_cols) :,.0f}\", \n        color = Fore.RED\n    )\n    SelMdlCols[method] = (sel_mdl_cols, cat_cols)\n\n    Xtrain_ = Xtrain.copy()\n    Xtest_  = Xtest.copy()\n\n    if \"CB\" in method :\n        Xtrain_[cat_cols] = Xtrain_[cat_cols].astype(pd.StringDtype)\n        Xtest_[cat_cols]  = Xtest_[cat_cols].astype(pd.StringDtype)\n    else:\n        pass\n\n    fitted_models, oof_preds, test_preds, ftreimp, mdl_best_iter =  \\\n    md.MakeOfflineModel(\n        Xtrain_,\n        ytrain,\n        ygrp,\n        Xtest_,\n        clone(mymodel),\n        method,\n        test_preds_req   = True,\n        ftreimp_plot_req = CFG.ftre_imp_req,\n        ntop = 50,\n    )\n\n    OOF_Preds[method]    = oof_preds\n    Mdl_Preds[method]    = test_preds\n    FittedModels[method] = fitted_models\n    FtreImp[method]      = ftreimp\n\n    del fitted_models, oof_preds, test_preds, ftreimp, sel_mdl_cols, Xtrain_, Xtest_\n    print()\n    collect();\n\n_ = utils.CleanMemory();\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T12:12:29.701584Z","iopub.execute_input":"2024-12-02T12:12:29.701995Z","iopub.status.idle":"2024-12-02T12:15:31.749086Z","shell.execute_reply.started":"2024-12-02T12:12:29.701947Z","shell.execute_reply":"2024-12-02T12:15:31.747751Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **ENSEMBLE**","metadata":{}},{"cell_type":"code","source":"%%time \n\noof_preds = pd.DataFrame(OOF_Preds)\nmdl_preds = pd.DataFrame(Mdl_Preds)\n\noof_ens_preds = oof_preds.mean(axis=1).values\ntest_preds    = mdl_preds.mean(axis=1).values\n\nscore = \\\nutils.ScoreMetric(\n    md.PostProcessPreds(ytrain.values), \n    oof_ens_preds\n)\n\nPrintColor(f\"\\n---> Final Ensemble Score = {score :,.6f}\\n\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T12:15:31.751032Z","iopub.execute_input":"2024-12-02T12:15:31.751516Z","iopub.status.idle":"2024-12-02T12:15:31.902301Z","shell.execute_reply.started":"2024-12-02T12:15:31.751465Z","shell.execute_reply":"2024-12-02T12:15:31.901252Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **CLOSURE**","metadata":{}},{"cell_type":"code","source":"%%time \n\ntry:\n    oof_preds.assign(**{\"Ensemble\": oof_ens_preds}).\\\n    to_parquet(\n        os.path.join(CFG.op_path, f\"OOF_Preds_{CFG.model_label}{CFG.model_id}.parquet\")\n    )\n\n    mdl_preds.assign(**{\"Ensemble\": test_preds}).\\\n    to_parquet(\n        os.path.join(CFG.op_path, f\"Mdl_Preds_{CFG.model_label}{CFG.model_id}.parquet\")\n    )\n    \nexcept:\n    oof_preds.\\\n    to_parquet(\n        os.path.join(CFG.op_path, f\"OOF_Preds_{CFG.model_label}{CFG.model_id}.parquet\")\n    )  \n    \n    mdl_preds.\\\n    to_parquet(\n        os.path.join(CFG.op_path, f\"Mdl_Preds_{CFG.model_label}{CFG.model_id}.parquet\")\n    )\n\npp.sub_fl[f\"Premium Amount\"] = test_preds\npp.sub_fl.to_csv(\n    os.path.join(CFG.op_path, f\"submission.csv\"), index = None\n)\n\n\nprint()\n!ls\nprint()\n!head submission.csv\n\n_ = utils.CleanMemory()\nprint()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T12:15:31.903892Z","iopub.execute_input":"2024-12-02T12:15:31.904274Z","iopub.status.idle":"2024-12-02T12:15:36.578428Z","shell.execute_reply.started":"2024-12-02T12:15:31.904242Z","shell.execute_reply":"2024-12-02T12:15:36.576557Z"}},"outputs":[],"execution_count":null}]}