{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30822,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:10:13.782905Z","iopub.execute_input":"2024-12-24T16:10:13.783331Z","iopub.status.idle":"2024-12-24T16:10:13.792967Z","shell.execute_reply.started":"2024-12-24T16:10:13.783302Z","shell.execute_reply":"2024-12-24T16:10:13.791665Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Importing Essential Libraires","metadata":{}},{"cell_type":"code","source":"import pandas as pd  \nimport numpy as np   \nimport matplotlib.pyplot as pyt  ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:10:13.794431Z","iopub.execute_input":"2024-12-24T16:10:13.794894Z","iopub.status.idle":"2024-12-24T16:10:13.817928Z","shell.execute_reply.started":"2024-12-24T16:10:13.794849Z","shell.execute_reply":"2024-12-24T16:10:13.816536Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Importing the Dataset (Train)","metadata":{}},{"cell_type":"code","source":"ds = pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\nds.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:10:13.820368Z","iopub.execute_input":"2024-12-24T16:10:13.820761Z","iopub.status.idle":"2024-12-24T16:10:18.806591Z","shell.execute_reply.started":"2024-12-24T16:10:13.8207Z","shell.execute_reply":"2024-12-24T16:10:18.805522Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#checking for missing values \nprint(ds.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:10:18.80831Z","iopub.execute_input":"2024-12-24T16:10:18.808678Z","iopub.status.idle":"2024-12-24T16:10:19.424513Z","shell.execute_reply.started":"2024-12-24T16:10:18.80865Z","shell.execute_reply":"2024-12-24T16:10:19.423415Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Filling Missing values with Median and Mode","metadata":{}},{"cell_type":"code","source":"# replacing the null values with Median \nds = ds.fillna(ds.median(numeric_only=True))\n\n# Replace empty strings with NaN in the entire dataset\nds = ds.replace(r'^\\s*$', pd.NA, regex=True)\n\n# Handle categorical columns (strings) with missing values\ncategorical_cols = ds.select_dtypes(include=['object', 'category'])\n\n# Fill missing values (including replaced empty strings) with the mode\nds[categorical_cols.columns] = categorical_cols.apply(lambda col: col.fillna(col.mode().iloc[0]))\n\n# Recheck \nds.isnull().sum()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:10:19.425579Z","iopub.execute_input":"2024-12-24T16:10:19.42588Z","iopub.status.idle":"2024-12-24T16:10:31.415908Z","shell.execute_reply.started":"2024-12-24T16:10:19.425854Z","shell.execute_reply":"2024-12-24T16:10:31.414643Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Encoding the Categorical Column","metadata":{}},{"cell_type":"code","source":"# Encoding with LabelEncoder For columns Like Gender and Smoking Status\nfrom sklearn.preprocessing import LabelEncoder\n\n# Initialize LabelEncoder\nlabel_encoder = LabelEncoder()\n\n# Encode the 'gender' column\nds['Gender'] = label_encoder.fit_transform(ds['Gender'])\n\n# Encode the 'Smoking status' column\nds['Smoking Status'] = label_encoder.fit_transform(ds['Smoking Status'])\nds.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:10:31.417056Z","iopub.execute_input":"2024-12-24T16:10:31.417484Z","iopub.status.idle":"2024-12-24T16:10:31.861535Z","shell.execute_reply.started":"2024-12-24T16:10:31.417447Z","shell.execute_reply":"2024-12-24T16:10:31.860628Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.compose import ColumnTransformer\nfrom sklearn.preprocessing import OneHotEncoder\n\n# Apply OneHotEncoder to specified columns\nct = ColumnTransformer(transformers=[('encoder', OneHotEncoder(), [4, 6, 7, 9, 10, 16, 18, 19])], remainder='passthrough')\nds_transformed = np.array(ct.fit_transform(ds))\n\n# Convert back to a pandas DataFrame\nds = pd.DataFrame(ds_transformed)\n\n# Display the first few rows\nprint(ds.head())\n# Drop the column at index 36\nds.drop(columns=ds.columns[36], inplace=True)\n\n# Display the first few rows\nprint(ds.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:10:31.862715Z","iopub.execute_input":"2024-12-24T16:10:31.863116Z","iopub.status.idle":"2024-12-24T16:10:39.357781Z","shell.execute_reply.started":"2024-12-24T16:10:31.863087Z","shell.execute_reply":"2024-12-24T16:10:39.356704Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Correlation Matrix and HeatMap","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\n\n\n# Calculate the correlation matrix\ncorrelation_matrix = ds.corr()\n\n# Show the correlation matrix with respect to \"LBM change (kg)\"\nlbm_corr = correlation_matrix[38].sort_values(ascending=False)\n# Plot the correlation matrix\npyt.figure(figsize=(40, 25))\nsns.heatmap(correlation_matrix, annot=True, cmap=\"coolwarm\", fmt=\".2f\", linewidths=0.7)\npyt.title(\"Correlation Matrix\")\npyt.show()\n\n# Display the correlation of features with \"LBM change (kg)\"\nlbm_corr\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:10:39.360911Z","iopub.execute_input":"2024-12-24T16:10:39.361188Z","iopub.status.idle":"2024-12-24T16:11:06.893595Z","shell.execute_reply.started":"2024-12-24T16:10:39.361164Z","shell.execute_reply":"2024-12-24T16:11:06.892429Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Observations from Data","metadata":{}},{"cell_type":"markdown","source":"### The next highest is 0.039394, which is extremely weak.\n### Most other variables have correlations near 0, both positive and negative.\n### Some variables even have slightly negative correlations (e.g., -0.024471).","metadata":{}},{"cell_type":"markdown","source":"# Conclusion ","metadata":{}},{"cell_type":"markdown","source":"### The correlations suggest that none of the independent variables (X) have a strong relationship with the dependent variable (y).\n### This implies that the dataset's X and y are largely independent or that the relationship might be non-linear, which is not captured by the correlation coefficient.","metadata":{}},{"cell_type":"code","source":"X = ds.iloc[:, :-1].values\n# X = X[:,:-1]\nX\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:11:06.895103Z","iopub.execute_input":"2024-12-24T16:11:06.895643Z","iopub.status.idle":"2024-12-24T16:11:06.902902Z","shell.execute_reply.started":"2024-12-24T16:11:06.895592Z","shell.execute_reply":"2024-12-24T16:11:06.901746Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y = ds.iloc[:,[-1]].values \ny","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:11:06.904175Z","iopub.execute_input":"2024-12-24T16:11:06.904546Z","iopub.status.idle":"2024-12-24T16:11:06.983928Z","shell.execute_reply.started":"2024-12-24T16:11:06.90451Z","shell.execute_reply":"2024-12-24T16:11:06.983039Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Fealture Scaling \nfrom sklearn.preprocessing import StandardScaler\nsc = StandardScaler()\nX = sc.fit_transform(X)\nX","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:11:06.984839Z","iopub.execute_input":"2024-12-24T16:11:06.985111Z","iopub.status.idle":"2024-12-24T16:11:10.58633Z","shell.execute_reply.started":"2024-12-24T16:11:06.985091Z","shell.execute_reply":"2024-12-24T16:11:10.585219Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y = y.reshape(-1, 1)\ny = y.ravel() \n\n# from sklearn.svm import SVR\n# reg = SVR(kernel=\"rbf\")\n# reg.fit(X,y)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:11:10.593209Z","iopub.execute_input":"2024-12-24T16:11:10.59351Z","iopub.status.idle":"2024-12-24T16:11:10.612285Z","shell.execute_reply.started":"2024-12-24T16:11:10.593472Z","shell.execute_reply":"2024-12-24T16:11:10.611002Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Analysing Scatter Plot ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\n\n# Assuming X and y are numpy arrays\n# Convert X to a DataFrame for easier handling\nX_df = pd.DataFrame(X)\n\n# Convert y to a Series (if it's a single column)\ny_series = pd.Series(y.flatten())  # Flatten y to ensure it's 1D\n\n# Visualize each column in X against y\nfor col in X_df.columns:\n    plt.figure(figsize=(6, 4))\n    sns.scatterplot(x=X_df[col], y=y_series)\n    plt.title(f\"Scatter Plot: Feature {col} vs Target\")\n    plt.xlabel(f\"Feature {col}\")\n    plt.ylabel(\"Target (y)\")\n    plt.grid(True)\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T16:11:10.613624Z","iopub.execute_input":"2024-12-24T16:11:10.614006Z","iopub.status.idle":"2024-12-24T16:13:15.35145Z","shell.execute_reply.started":"2024-12-24T16:11:10.613961Z","shell.execute_reply":"2024-12-24T16:13:15.350117Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Conclusion","metadata":{}},{"cell_type":"markdown","source":"## Analysis of Feature-Target Relationships\nAfter analyzing the dataset for the Kaggle competition, the following insights were derived:\n\n### 1. Heatmap Analysis\nA heatmap was generated to examine the correlation of independent variables (features, X) with the dependent variable (y). The results revealed that:\n\nCorrelation values were extremely low for most features.\nThis indicates a very weak linear relationship between the features and the target variable.\n\n### 2. Scatter Plot Observations\nTo further understand the relationships, scatter plots were drawn for each feature in X against y. The scatter plots revealed the following:\n\nTwo vertical and parallel lines for every feature.\nThis suggests that:\nThe features are binary or take on a small set of discrete values (e.g., 0 and 1).\nThere is no visible pattern or trend between the features and the target variable.\nThis implies that these features have very little or no predictive power for y.\n### 3. Conclusion\nBoth the heatmap and scatter plot analyses confirm that:\n\nThe features in X have minimal or no relationship with the target variable y.\nThis indicates that:\nThe current feature set may not be sufficient to model the target variable effectively.\nFeature engineering or the inclusion of additional, more relevant features might be necessary to improve model performance.\n## Potential Next Steps:\nFeature Selection: Use methods like mutual information or tree-based feature importance to identify any potentially useful features.\nFeature Engineering: Explore additional data sources or derive new features to enhance the dataset.\nModeling: Experiment with non-linear models such as:\nRandom Forest\nXGBoost\nNeural Networks These models could better capture complex relationships that linear correlation analysis fails to detect.\n","metadata":{}}]}