{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# S4E12 Lazy Regressor","metadata":{}},{"cell_type":"markdown","source":"https://www.kaggle.com/code/gauravduttakiit/rice-msc-lazy-predict","metadata":{}},{"cell_type":"markdown","source":"LazyPredict has multiple models to make predictions and lets you easily determine which model performs best. A separate model must be created to predict test data.","metadata":{}},{"cell_type":"markdown","source":"The main classification and regression models are listed below:\n\n### Classification models (LazyClassifier):\n\nLogisticRegression\nSGDClassifier\nPassiveAggressiveClassifier\nRandomForestClassifier\nGradientBoostingClassifier\nExtraTreesClassifier\nAdaBoostClassifier\nKNeighborsClassifier\nDecisionTreeClassifier\nExtraTreeClassifier\nLinearSVC\nSVC\nNuSVC\nLGBMClassifier\nXGBClassifier\nCatBoostClassifier\nBaggingClassifier\nGaussianNB\nBernoulliNB\nMultinomialNB\nQuadraticDiscriminantAnalysis\nLinearDiscriminantAnalysis\nMLPClassifier\nRidgeClas sifier\nPerceptron\n\n### Regression models (LazyRegressor):\n\nLinearRegression\nRidge\nLasso\nElasticNet\nLars\nLassoLars\nOrthogonalMatchingPursuit\nBayesianRidge\nARDRegression\nSGDRegressor\nPassiveAggressiveRegressor\nRANSACRegressor\nTheilSenRegressor\nHuberRegressor\nKernelRidge\nSVR\nNuSVR\nKNeighborsRegressor\nDecisionTreeRegressor\nRandomForestRegressor\nExtraTreesRegressor\nAdaBoostRegressor\nGradientBoostingRegressor\nMLPRegressor\nXGBRegressor\nLGBMRegressor\nCatBoostRegressor\nExtraTreeRegressor","metadata":{}},{"cell_type":"code","source":"!pip install openpyxl\n!pip install lazypredict","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-12-04T14:04:51.768138Z","iopub.execute_input":"2024-12-04T14:04:51.768697Z","iopub.status.idle":"2024-12-04T14:05:13.868985Z","shell.execute_reply.started":"2024-12-04T14:04:51.768585Z","shell.execute_reply":"2024-12-04T14:05:13.867378Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import random\nimport pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom lazypredict.Supervised import LazyRegressor\nfrom sklearn.metrics import classification_report","metadata":{"execution":{"iopub.status.busy":"2024-12-04T14:05:13.872774Z","iopub.execute_input":"2024-12-04T14:05:13.873328Z","iopub.status.idle":"2024-12-04T14:05:13.881293Z","shell.execute_reply.started":"2024-12-04T14:05:13.873273Z","shell.execute_reply":"2024-12-04T14:05:13.879845Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train0=pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\nprint(len(train0))\nN=list(range(len(train0)))\nrandom.shuffle(N)\ntrain=train0.iloc[0:100000]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T14:05:13.882858Z","iopub.execute_input":"2024-12-04T14:05:13.883236Z","iopub.status.idle":"2024-12-04T14:05:20.800665Z","shell.execute_reply.started":"2024-12-04T14:05:13.883203Z","shell.execute_reply":"2024-12-04T14:05:20.799233Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(train.columns.tolist())","metadata":{"execution":{"iopub.status.busy":"2024-12-04T14:05:20.80318Z","iopub.execute_input":"2024-12-04T14:05:20.8038Z","iopub.status.idle":"2024-12-04T14:05:20.810041Z","shell.execute_reply.started":"2024-12-04T14:05:20.80376Z","shell.execute_reply":"2024-12-04T14:05:20.808657Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train=train.dropna()","metadata":{"execution":{"iopub.status.busy":"2024-12-04T14:05:20.814166Z","iopub.execute_input":"2024-12-04T14:05:20.814716Z","iopub.status.idle":"2024-12-04T14:05:20.836924Z","shell.execute_reply.started":"2024-12-04T14:05:20.814663Z","shell.execute_reply":"2024-12-04T14:05:20.835681Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\ndef labelencoder(df):\n    for c in df.columns:\n        if df[c].dtype=='object': \n            df[c] = df[c].fillna('N')\n            lbl = LabelEncoder()\n            lbl.fit(list(df[c].values))\n            df[c] = lbl.transform(df[c].values)\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-12-04T14:05:20.838657Z","iopub.execute_input":"2024-12-04T14:05:20.839096Z","iopub.status.idle":"2024-12-04T14:05:20.846682Z","shell.execute_reply.started":"2024-12-04T14:05:20.839049Z","shell.execute_reply":"2024-12-04T14:05:20.845429Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train=labelencoder(train)\ndisplay(train.info())","metadata":{"execution":{"iopub.status.busy":"2024-12-04T14:05:20.848154Z","iopub.execute_input":"2024-12-04T14:05:20.848536Z","iopub.status.idle":"2024-12-04T14:05:20.979076Z","shell.execute_reply.started":"2024-12-04T14:05:20.848499Z","shell.execute_reply":"2024-12-04T14:05:20.97767Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y = train['Premium Amount']\nX = train.drop('Premium Amount',axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-12-04T14:05:20.980962Z","iopub.execute_input":"2024-12-04T14:05:20.981463Z","iopub.status.idle":"2024-12-04T14:05:20.997763Z","shell.execute_reply.started":"2024-12-04T14:05:20.981392Z","shell.execute_reply":"2024-12-04T14:05:20.996241Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42,shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2024-12-04T14:05:20.999135Z","iopub.execute_input":"2024-12-04T14:05:20.999506Z","iopub.status.idle":"2024-12-04T14:05:21.011606Z","shell.execute_reply.started":"2024-12-04T14:05:20.999451Z","shell.execute_reply":"2024-12-04T14:05:21.010159Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"clf = LazyRegressor(verbose=0,predictions=True)\nmodels,predictions = clf.fit(X_train, X_test, y_train, y_test)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-12-04T14:05:21.013306Z","iopub.execute_input":"2024-12-04T14:05:21.013785Z","iopub.status.idle":"2024-12-04T14:05:47.109667Z","shell.execute_reply.started":"2024-12-04T14:05:21.01375Z","shell.execute_reply":"2024-12-04T14:05:47.108485Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"* Adjusted R-Squared: The closer to 1, the better the model.\n* R-Squared (coefficient of determination): The closer to 1, the better the model.\n* RMSE (Root Mean Squared Error): The closer to 0, the better the model.","metadata":{}},{"cell_type":"code","source":"print(models)","metadata":{"_kg_hide-output":false,"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T14:05:47.111144Z","iopub.execute_input":"2024-12-04T14:05:47.111522Z","iopub.status.idle":"2024-12-04T14:05:47.121175Z","shell.execute_reply.started":"2024-12-04T14:05:47.111487Z","shell.execute_reply":"2024-12-04T14:05:47.119868Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}