{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **The Starter Pack**\n\nThis notebook serves as a template to provide foundational functionalities for this kaggle competition submission. It includes the following key sections:\n\n## **Purpose**  \nTo provide a structured framework for:  \n- Loading data  \n- Basic Exploratory Data Analysis (EDA)\n- Basic Transformation \n- Building a basic Machine Learning (ML) model  \n- Evaluating the model performance  \n- Preparing results for submission  \n\nThis will help us to have the basic modules to start with then its all about make it complex each step.\n\n---","metadata":{}},{"cell_type":"markdown","source":"## Importing libraries","metadata":{}},{"cell_type":"code","source":"# Loading required libraries\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os # os functions\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-13T11:03:21.927282Z","iopub.execute_input":"2024-12-13T11:03:21.927669Z","iopub.status.idle":"2024-12-13T11:03:22.432358Z","shell.execute_reply.started":"2024-12-13T11:03:21.927624Z","shell.execute_reply":"2024-12-13T11:03:22.431111Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Basic EDA\nLet's have a look at the dataset files available","metadata":{}},{"cell_type":"code","source":"# checking files and their sizes\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        file_path = os.path.join(dirname, filename)\n        print(file_path,\" | \", round(os.path.getsize(file_path)/(1024 ** 2),2),'MB')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T11:03:22.435143Z","iopub.execute_input":"2024-12-13T11:03:22.435622Z","iopub.status.idle":"2024-12-13T11:03:22.445387Z","shell.execute_reply.started":"2024-12-13T11:03:22.435584Z","shell.execute_reply":"2024-12-13T11:03:22.444247Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Before loading data into a dataframe lets have a look at the data structure.","metadata":{}},{"cell_type":"code","source":"# checking data structure\nprint(\"\\n train.csv \")\n!head -3 /kaggle/input/playground-series-s4e12/train.csv\nprint(\"\\n test.csv\")\n!head -3 /kaggle/input/playground-series-s4e12/test.csv\nprint(\"\\n submission.csv\")\n!head -3 /kaggle/input/playground-series-s4e12/sample_submission.csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T11:03:22.447187Z","iopub.execute_input":"2024-12-13T11:03:22.447658Z","iopub.status.idle":"2024-12-13T11:03:26.05494Z","shell.execute_reply.started":"2024-12-13T11:03:22.447608Z","shell.execute_reply":"2024-12-13T11:03:26.053536Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Lets load the files in to pandas dataframe","metadata":{}},{"cell_type":"code","source":"# train dataframe\ntrain_df = pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\n\nwith pd.option_context('display.max_columns', None): # setting the max rows\n    display(train_df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T11:03:26.057647Z","iopub.execute_input":"2024-12-13T11:03:26.058031Z","iopub.status.idle":"2024-12-13T11:03:31.252871Z","shell.execute_reply.started":"2024-12-13T11:03:26.057979Z","shell.execute_reply":"2024-12-13T11:03:31.251541Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Basic Transformation\nIn this section, we will apply the transformations identified during the EDA step. For the starter pack, let’s remove the date column.","metadata":{}},{"cell_type":"code","source":"# removing the date column\ntrain_df.drop(columns=['Policy Start Date'], inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T11:03:31.254315Z","iopub.execute_input":"2024-12-13T11:03:31.25465Z","iopub.status.idle":"2024-12-13T11:03:31.476987Z","shell.execute_reply.started":"2024-12-13T11:03:31.254616Z","shell.execute_reply":"2024-12-13T11:03:31.475673Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Basic AutoML model\nH2O AutoML is an open-source machine learning platform that provides an automated approach to building machine learning models. It supports a variety of algorithms for classification, regression, and other tasks, with the goal of making machine learning accessible to both beginners and experts by automating many of the processes involved.","metadata":{}},{"cell_type":"code","source":"# install h2o if not installed\n!pip install h2o","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T11:03:31.47864Z","iopub.execute_input":"2024-12-13T11:03:31.479143Z","iopub.status.idle":"2024-12-13T11:03:41.627724Z","shell.execute_reply.started":"2024-12-13T11:03:31.479089Z","shell.execute_reply":"2024-12-13T11:03:41.626592Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# running autoML\n\nimport h2o\nimport pandas as pd\nfrom h2o.automl import H2OAutoML\nfrom sklearn.metrics import mean_squared_log_error\n\n# Initialize H2O cluster\nh2o.init()\n\n# Convert pandas DataFrame to H2OFrame\ntrain_hf = h2o.H2OFrame(train_df)\n\n# Define the target and feature columns\ny = 'Premium Amount'\nX = train_hf.columns\nX.remove('Premium Amount')\n\n# Split the data into training and testing sets\ntrn_hf, val_hf = train_hf.split_frame(ratios=[.8], seed=42)\n\n# Initialize H2O AutoML model\naml = H2OAutoML(max_models=20, seed=1, max_runtime_secs=300)\n\n# Train the model\naml.train(x=X, y=y, training_frame=trn_hf)\n\n# Get the leader model (best model from AutoML run)\nleader_model = aml.leader\n\n# Print the leader model details\n#print(leader_model)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T11:03:41.631478Z","iopub.execute_input":"2024-12-13T11:03:41.63186Z","iopub.status.idle":"2024-12-13T11:09:22.104404Z","shell.execute_reply.started":"2024-12-13T11:03:41.631823Z","shell.execute_reply":"2024-12-13T11:09:22.102974Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Model evaluation\nRoot Mean Squared Logarithmic Error (RMSLE) is not directly available in scikit-learn, so we use Mean Squared Logarithmic Error (MSLE) and then take the square root to compute RMSLE.","metadata":{}},{"cell_type":"code","source":"# Make predictions on the validation set\nval_hf['y_pred'] = leader_model.predict(val_hf)\n\n# Convert the predictions to pandas DataFrame for easier inspection\nval_df = val_hf.as_data_frame(use_multi_thread=True)\ny_pred = val_df.y_pred\ny_true = val_df['Premium Amount'].values\n\n# evaluate model performance (e.g., Root Mean Squared Log Error)\n\n# Compute the Mean Squared Logarithmic Error (MSLE)\nmsle = mean_squared_log_error(y_true, y_pred)\n\n# Compute the Root Mean Squared Logarithmic Error (RMSLE)\nrmsle = np.sqrt(msle)\nprint('\\n')\nprint(f'Root Mean Squared Log Error: {rmsle:.2f}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T11:09:22.106557Z","iopub.execute_input":"2024-12-13T11:09:22.107289Z","iopub.status.idle":"2024-12-13T11:09:27.508659Z","shell.execute_reply.started":"2024-12-13T11:09:22.107237Z","shell.execute_reply":"2024-12-13T11:09:27.507399Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## Submission file preperation\nLet's prepare the competition submission file.","metadata":{}},{"cell_type":"code","source":"# loading test file\ntest_df = pd.read_csv(\"/kaggle/input/playground-series-s4e12/test.csv\")\ntest_hf = h2o.H2OFrame(test_df)\n\ntest_hf['Premium Amount'] = leader_model.predict(test_hf)\nsubmission_df = test_hf[['id','Premium Amount']].as_data_frame(use_multi_thread=True)\nsubmission_df['Premium Amount'] = submission_df['Premium Amount'].round(decimals=3)\n\nsubmission_df.to_csv('submission.csv', index=False)\n\n# Shutdown H2O cluster after use\nh2o.cluster().shutdown()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T11:09:27.510457Z","iopub.execute_input":"2024-12-13T11:09:27.510978Z","iopub.status.idle":"2024-12-13T11:10:00.208621Z","shell.execute_reply.started":"2024-12-13T11:09:27.510926Z","shell.execute_reply":"2024-12-13T11:10:00.20737Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}}]}