{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nBriefly, in this notebook I try to solve the Spaceship Titanic competition JUST by statistical methods and probability and not Machine Learning.\n\n## Statistical Analysis\nYou may hear Bayesian Inference or probabily work with it. Here I used this method but in a slightly different way;\nDefining Bayesian Score instead of Bayesian Probability.\n\n## Bayes Theorem\nThis theorem describes the probability of an event, based on prior knowledge of conditions that might be related to the event.\n### Simple Form:\n${\\displaystyle P(A\\mid B)={\\frac {P(B\\mid A)P(A)}{P(B)}}}$\n### Extended Form:\nOften, for some partition ${A_j}$ of the sample space, the event space is given in terms of $P(A_j)$ and $P(B | A_j)$. It is then useful to compute P(B) using the law of total probability:\n\n${\\displaystyle P(B)={\\sum _{j}P(B|A_{j})P(A_{j})}}$\n\nOr equivalently:\n${\\displaystyle P(B)={\\sum _{j}P(B \t\\cap A_{j})}}$\n\n## Bayesian Score\nAs in our case the event we want to predict (Transported) is not completely consists of the dataset's features, the bayesian theorem assumptions doesn't satisfy. Therefore, I use the idea behind this great theorem and define a new random variable, named Bayesian Score.\n\nTo understand the restriction of using this theorem, consider the $Ω$ as sample space, then we have:\n\n$Columns = \\big \\{PassengerId, HomePlanet, CryoSleep, Cabin, Destination, Age,VIP, RoomService, FoodCourt, ShoppingMall, Spa, VRDeck,Name\\big \\}$\n\n\n${\\displaystyle Ω \\neq \\bigcup_{\\alpha \\in Columns} \\alpha }$\n\n\nTherefore:\n\n${\\displaystyle P(Transported=0 \\,or\\, 1) \\neq {\\sum _{\\alpha \\in Columns}P(Transported \\cap \\alpha)}}$\n\n### BS: Bayesian Score\n\n${\\displaystyle BS(B)={\\prod _{j}P(B|A_{j})P(A_{j})}}$\n\nOr equivalently:\n${\\displaystyle BS(B)={\\prod _{j}P(B \t\\cap A_{j})}}$\n\n### Finally\n\n${\\displaystyle BS\\big (Transported = 0 \\,or\\, 1 \\big )={\\prod _{\\alpha \\in Columns}P \\big ( (Transported = 0 \\,or\\, 1) \\cap \\alpha \\big)}}$\n\n## Advantage\n\nAs you can see in the following, one of the benefit of this type of statistical analysis is that, **You don't need to fill missing values** in the data which is almost one of the most time consuming and chanllenging task in data analysis.","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nfrom matplotlib import pyplot\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nwarnings.filterwarnings('ignore')\nimport re\nimport math\nimport collections","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:29:13.424045Z","iopub.execute_input":"2022-09-23T15:29:13.424534Z","iopub.status.idle":"2022-09-23T15:29:13.431302Z","shell.execute_reply.started":"2022-09-23T15:29:13.424492Z","shell.execute_reply":"2022-09-23T15:29:13.430114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(r'../input/spaceship-titanic/train.csv')\ntest = pd.read_csv(r'../input/spaceship-titanic/test.csv')\nsubmission = pd.read_csv(r'../input/spaceship-titanic/sample_submission.csv')\nAll = pd.concat([train, test], sort=False).reset_index(drop=True)\nAll","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:32:21.975263Z","iopub.execute_input":"2022-09-23T15:32:21.975792Z","iopub.status.idle":"2022-09-23T15:32:22.073254Z","shell.execute_reply.started":"2022-09-23T15:32:21.97575Z","shell.execute_reply":"2022-09-23T15:32:22.071927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preparation\n\nIn this section I prepare the data including extract meaningful part from some of columns, seperate them, encoding categorical features to numeric and so on.","metadata":{}},{"cell_type":"code","source":"def family(data):\n    \n    data['Family'] = train['Name']\n\n    for i in range(data.shape[0]):\n        if (data['Name'].isnull()[i] ==False):\n            data['Family'][i] = data['Name'][i].split(' ')[1]\n    return data\nfamily(All)","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:32:26.085262Z","iopub.execute_input":"2022-09-23T15:32:26.086316Z","iopub.status.idle":"2022-09-23T15:32:40.065046Z","shell.execute_reply.started":"2022-09-23T15:32:26.086269Z","shell.execute_reply":"2022-09-23T15:32:40.063822Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def passengerID(data):\n    \n    data['family_id'] = data['PassengerId']\n    data['second_id'] = data['PassengerId']\n    for i in range(data.shape[0]):\n        \n        data['family_id'][i] = data['PassengerId'][i][:4]\n        data['second_id'][i] = data['PassengerId'][i][5:]\n    return data\n\npassengerID(All)","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:32:44.151903Z","iopub.execute_input":"2022-09-23T15:32:44.152337Z","iopub.status.idle":"2022-09-23T15:32:54.587881Z","shell.execute_reply.started":"2022-09-23T15:32:44.1523Z","shell.execute_reply":"2022-09-23T15:32:54.586637Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def separate_cabin(data):\n    data['Cabin_deck'] = data['Cabin']\n    #data['Cabin_num'] = data['Cabin']\n    data['Cabin_side'] = data['Cabin']\n    for i in range(data.shape[0]):\n\n        if data['Cabin'].isnull()[i] ==False:\n\n            data['Cabin_deck'][i] = data['Cabin'][i][0]\n     #       data['Cabin_num'][i] = int(re.findall(r'\\d+', data['Cabin'][i])[0])\n            data['Cabin_side'][i] = data['Cabin'][i][-1]\n\n    data.drop(['Cabin'],axis=1,inplace=True)\n    return data\nseparate_cabin(All)","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:33:04.345014Z","iopub.execute_input":"2022-09-23T15:33:04.345442Z","iopub.status.idle":"2022-09-23T15:33:22.997591Z","shell.execute_reply.started":"2022-09-23T15:33:04.345408Z","shell.execute_reply":"2022-09-23T15:33:22.996783Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"All['HomePlanet'].replace(['Earth', 'Europa','Mars'], [0, 1,2], inplace=True)\nAll['CryoSleep'].replace([False, True], [0, 1], inplace=True)                       \nAll['Destination'].replace(['TRAPPIST-1e', '55 Cancri e','PSO J318.5-22'], [0, 1,2], inplace=True)\nAll['VIP'].replace([False, True], [0, 1], inplace=True)\nAll['Transported'].replace([False, True], [0, 1], inplace=True)\nAll['Cabin_deck'].replace(['F', 'G','E','B','C','D','A','T'], [0, 1,2,3,4,5,6,7], inplace=True)\nAll['Cabin_side'].replace(['S', 'P'], [0, 1], inplace=True)\nAll","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:33:28.455096Z","iopub.execute_input":"2022-09-23T15:33:28.455527Z","iopub.status.idle":"2022-09-23T15:33:28.553187Z","shell.execute_reply.started":"2022-09-23T15:33:28.455481Z","shell.execute_reply":"2022-09-23T15:33:28.552192Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Numeric Cols Peparation\n\nDue to large number of unique values of some columns like $\\big \\{VRDeck,Spa,FoodCourt,RoomService,ShoppingMall \\big \\}$ we need to partitioning them to smaller groups. I did this by exploring each of this columns and their relation to the target (Transported).\nIn other words, for better partitioning, we need to group each interval of values which have less diversity in order of their Transported values.\nAs an example if the label values of an interal has $n$ False and $0$ True and another interval has $\\dfrac {m}{2}$ False and $\\dfrac {m}{2}$ True, the former is the better choice, beacuse it separates the feature to more distinct part and therefore high variance probability. \n","metadata":{}},{"cell_type":"code","source":"sns.relplot(data=All.head(train.shape[0]), kind=\"line\",x=\"Age\", y=\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:30:41.509244Z","iopub.execute_input":"2022-09-23T15:30:41.509878Z","iopub.status.idle":"2022-09-23T15:30:43.605707Z","shell.execute_reply.started":"2022-09-23T15:30:41.509828Z","shell.execute_reply":"2022-09-23T15:30:43.604778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1 = All.head(train.shape[0])\nVR_Tran = []\nfor i in range(train_1.shape[0]):\n    \n    if train_1['VRDeck'][i] > 4000:\n        VR_Tran.append(train_1[\"Transported\"][i])\nprint(\"Transported=False:\",VR_Tran.count(0)/len(VR_Tran))        ","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:33:39.822098Z","iopub.execute_input":"2022-09-23T15:33:39.822513Z","iopub.status.idle":"2022-09-23T15:33:39.8913Z","shell.execute_reply.started":"2022-09-23T15:33:39.822479Z","shell.execute_reply":"2022-09-23T15:33:39.889676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(All.shape[0]):\n    \n    if All['VRDeck'][i] > 1000:\n        All['VRDeck'][i]=1\n    elif All['VRDeck'][i] > 100:\n        All['VRDeck'][i]=2\n    elif All['VRDeck'][i] > 50:\n        All['VRDeck'][i]=3\n    elif All['VRDeck'][i] > 10:\n        All['VRDeck'][i]=4\n    elif All['VRDeck'][i] > 0:\n        All['VRDeck'][i]=5\n    elif All['VRDeck'][i] ==0:\n        All['VRDeck'][i]=6","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:33:44.397572Z","iopub.execute_input":"2022-09-23T15:33:44.397992Z","iopub.status.idle":"2022-09-23T15:33:48.608592Z","shell.execute_reply.started":"2022-09-23T15:33:44.397958Z","shell.execute_reply":"2022-09-23T15:33:48.607459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_1 = All.head(train.shape[0])\nSpa_Tran = []\nfor i in range(train_1.shape[0]):\n    \n    if train_1['Spa'][i] > 3000:\n        Spa_Tran.append(train_1[\"Transported\"][i])\nprint(\"Transported=False:\",Spa_Tran.count(0)/len(Spa_Tran))        ","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:33:59.181986Z","iopub.execute_input":"2022-09-23T15:33:59.183203Z","iopub.status.idle":"2022-09-23T15:33:59.253446Z","shell.execute_reply.started":"2022-09-23T15:33:59.18316Z","shell.execute_reply":"2022-09-23T15:33:59.252447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(All.shape[0]):\n    \n    if All['Spa'][i] > 1500:\n        All['Spa'][i]=1\n    elif All['Spa'][i] > 1000:\n        All['Spa'][i]=2\n    elif All['Spa'][i] > 500:\n        All['Spa'][i]=3\n    elif All['Spa'][i] > 100:\n        All['Spa'][i]=4\n    elif All['Spa'][i] > 10:\n        All['Spa'][i]=5\n    elif All['Spa'][i] > 1:\n        All['Spa'][i]=6\n    elif All['Spa'][i] ==1:\n        All['Spa'][i]=7\n    elif All['Spa'][i] ==0:\n        All['Spa'][i]=8","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:34:06.989699Z","iopub.execute_input":"2022-09-23T15:34:06.990129Z","iopub.status.idle":"2022-09-23T15:34:11.318824Z","shell.execute_reply.started":"2022-09-23T15:34:06.990093Z","shell.execute_reply":"2022-09-23T15:34:11.317816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Food_Tran = []\nfor i in range(train_1.shape[0]):\n    \n    if train_1['FoodCourt'][i] > 17000:\n        \n        Food_Tran.append(train_1[\"Transported\"][i])\n        \nprint(\"Transported=False:\",Food_Tran.count(1)/len(Food_Tran))   ","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:35:59.682053Z","iopub.execute_input":"2022-09-23T15:35:59.682492Z","iopub.status.idle":"2022-09-23T15:35:59.749234Z","shell.execute_reply.started":"2022-09-23T15:35:59.682457Z","shell.execute_reply":"2022-09-23T15:35:59.747977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(All.shape[0]):\n    \n    if All['FoodCourt'][i] > 15000:\n        All['FoodCourt'][i]=1\n    elif All['FoodCourt'][i] > 8760:\n        All['FoodCourt'][i]=2\n    elif All['FoodCourt'][i] > 1500:\n        All['FoodCourt'][i]=3\n    elif All['FoodCourt'][i] > 500:\n        All['FoodCourt'][i]=4\n    elif All['FoodCourt'][i] > 100:\n        All['FoodCourt'][i]=5\n    elif All['FoodCourt'][i] > 10:\n        All['FoodCourt'][i]=6\n    elif All['FoodCourt'][i] > 0:\n        All['FoodCourt'][i]=7\n    elif All['FoodCourt'][i]==0:\n        All['FoodCourt'][i]=8","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:36:09.733637Z","iopub.execute_input":"2022-09-23T15:36:09.734033Z","iopub.status.idle":"2022-09-23T15:36:14.022227Z","shell.execute_reply.started":"2022-09-23T15:36:09.734001Z","shell.execute_reply":"2022-09-23T15:36:14.021028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Room_Tran = []\nfor i in range(train_1.shape[0]):\n    \n    if train_1['RoomService'][i] > 3682:\n        Room_Tran.append(train_1[\"Transported\"][i])\nprint(\"Transported=False:\",Room_Tran.count(0)/len(Room_Tran))     ","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:38:35.389768Z","iopub.execute_input":"2022-09-23T15:38:35.3902Z","iopub.status.idle":"2022-09-23T15:38:35.45827Z","shell.execute_reply.started":"2022-09-23T15:38:35.390165Z","shell.execute_reply":"2022-09-23T15:38:35.457141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(All.shape[0]):\n    \n    if All['RoomService'][i] > 1200:\n        All['RoomService'][i]=1\n    elif All['RoomService'][i] > 500:\n        All['RoomService'][i]=2\n    elif All['RoomService'][i] > 50:\n        All['RoomService'][i]=3\n    elif All['RoomService'][i] > 10:\n        All['RoomService'][i]=4\n    elif All['RoomService'][i] > 0:\n        All['RoomService'][i]=5\n    elif All['RoomService'][i] == 0:\n        All['RoomService'][i]=6\n","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:38:53.542703Z","iopub.execute_input":"2022-09-23T15:38:53.54317Z","iopub.status.idle":"2022-09-23T15:38:57.640616Z","shell.execute_reply.started":"2022-09-23T15:38:53.543129Z","shell.execute_reply":"2022-09-23T15:38:57.639016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Shop_Tran = []\nfor i in range(train_1.shape[0]):\n    \n    if train_1['ShoppingMall'][i] > 5000:\n        \n        Shop_Tran.append(train_1[\"Transported\"][i])\n        \nprint(\"Transported=True:\",Shop_Tran.count(1)/len(Shop_Tran))   ","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:43:39.95822Z","iopub.execute_input":"2022-09-23T15:43:39.958666Z","iopub.status.idle":"2022-09-23T15:43:40.025785Z","shell.execute_reply.started":"2022-09-23T15:43:39.95863Z","shell.execute_reply":"2022-09-23T15:43:40.024953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(All.shape[0]):\n    \n    if All['ShoppingMall'][i] > 2000:\n        All['ShoppingMall'][i]=1\n    elif All['ShoppingMall'][i] > 1000:\n        All['ShoppingMall'][i]=2\n    elif All['ShoppingMall'][i] > 500:\n        All['ShoppingMall'][i]=3\n    elif All['ShoppingMall'][i] > 0:\n        All['ShoppingMall'][i]=4\n    elif All['ShoppingMall'][i] == 0:\n        All['ShoppingMall'][i]=5\n","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:45:06.316549Z","iopub.execute_input":"2022-09-23T15:45:06.317007Z","iopub.status.idle":"2022-09-23T15:45:10.365788Z","shell.execute_reply.started":"2022-09-23T15:45:06.31697Z","shell.execute_reply":"2022-09-23T15:45:10.364554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(All.shape[0]):\n    \n    try:\n         All['Family'][i] = All['Family'].value_counts()[All['Family'][i]]\n    except Exception:\n        pass\nAll ","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:46:39.502611Z","iopub.execute_input":"2022-09-23T15:46:39.503113Z","iopub.status.idle":"2022-09-23T15:47:14.33935Z","shell.execute_reply.started":"2022-09-23T15:46:39.503077Z","shell.execute_reply":"2022-09-23T15:47:14.338252Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Label = All['Transported']\nAll.drop(['family_id','Transported','Name','PassengerId'],axis=1,inplace=True)\n\nAll['second_id'] = All['second_id'].tolist()\nfor i in range(All.shape[0]):\n    All['second_id'][i] = int(All['second_id'][i])\nAll['Transported'] = Label\nAll","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:48:21.582874Z","iopub.execute_input":"2022-09-23T15:48:21.58332Z","iopub.status.idle":"2022-09-23T15:48:26.373651Z","shell.execute_reply.started":"2022-09-23T15:48:21.583285Z","shell.execute_reply":"2022-09-23T15:48:26.372567Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def BS_Prob(col):\n    data = All.head(train.shape[0])\n    values = data[col].value_counts().index.tolist()\n    class_dist_0 = []\n    class_dist_1 = []\n    probability_0 = {}\n    probability_1 = {}\n    for i in values:\n        \n        try:\n            \n            a = All.head(train.shape[0]).groupby(col)['Transported'].value_counts()[i][0]\n            b = All.head(train.shape[0]).groupby(col)['Transported'].value_counts()[i][1]\n            class_dist_0.append(a)\n            class_dist_1.append(b)\n            \n        except Exception:\n            pass\n      \n        if (a*b)!=0:\n            p = a/(a+b)\n            probability_0[i]= p #probability of col=i and Transported=0\n            probability_1[i] = (1-p)\n        elif a == 0:\n            probability_0[i] = 0.001\n            probability_1[i] = (1-p)\n        elif b == 0:\n            probability_0[i] = 0.999\n            probability_1[i] = (1-p)\n\n    return probability_0,probability_1,class_dist_0,class_dist_1,values","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:48:59.074123Z","iopub.execute_input":"2022-09-23T15:48:59.07454Z","iopub.status.idle":"2022-09-23T15:48:59.085287Z","shell.execute_reply.started":"2022-09-23T15:48:59.074497Z","shell.execute_reply":"2022-09-23T15:48:59.084381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Return = BS_Prob('Age')\ndf = pd.DataFrame({'Transported_0': Return[2],'Transported_1': Return[3]}, index=Return[3])\nax = df.plot.bar(rot=0)\nReturn[0]","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:49:07.693992Z","iopub.execute_input":"2022-09-23T15:49:07.695213Z","iopub.status.idle":"2022-09-23T15:49:09.382803Z","shell.execute_reply.started":"2022-09-23T15:49:07.695166Z","shell.execute_reply":"2022-09-23T15:49:09.381662Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols = All.columns.tolist()[:-1]\ndef predict(record):\n\n    p0,p1 = 1,1\n    for i in cols:\n        if np.isnan(record[i]) == False:\n            \n            prob = BS_Prob(i)\n            \n            prob1 = round(prob[0][int(record[i])],2)\n            prob2 = round(prob[1][int(record[i])],2)\n        else:\n            prob1 = 1\n            prob2 = 1\n            \n        p0 = p0 * prob1\n        p1 = p1 * prob2\n    return p0,p1","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:49:29.829522Z","iopub.execute_input":"2022-09-23T15:49:29.829975Z","iopub.status.idle":"2022-09-23T15:49:29.837053Z","shell.execute_reply.started":"2022-09-23T15:49:29.829936Z","shell.execute_reply":"2022-09-23T15:49:29.836162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nresult_test = []\nTrue_prob = []\nFalse_prob = []\ntest = All.tail(test.shape[0])\n\nfor i in range(test.shape[0]):\n    Pre_test = predict(test.iloc[i])\n    \n    True_prob.append(Pre_test[1])\n    False_prob.append(Pre_test[0])\n    \n    print(i,Pre_test[0] , Pre_test[1])\n    \n    if Pre_test[0] > Pre_test[1]:\n        result_test.append(0)\n    else:\n         result_test.append(1)","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission['Transported'] = result_test\nsubmission.Transported = submission.Transported.replace({1:True, 0:False})                   ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Improvement\n\nHere again by exploring some of features, we can obtain high confindence in prediction just by single columns, but in small interval of values.\nActually, as you see the probabilities which are greater than %98, we can predict Transpoted value for records, that their values on some features are in the high confindence intervals.","metadata":{}},{"cell_type":"code","source":"train0 = pd.read_csv(r'../input/spaceship-titanic/train.csv')\ntest0 = pd.read_csv(r'../input/spaceship-titanic/test.csv')\nAll0 = pd.concat([train0, test0], sort=False).reset_index(drop=True)\ntest_1 = All0.tail(test.shape[0])\ntest_1=test_1.reset_index(drop=True)\ntest_1","metadata":{"execution":{"iopub.status.busy":"2022-09-23T15:51:10.957965Z","iopub.execute_input":"2022-09-23T15:51:10.95839Z","iopub.status.idle":"2022-09-23T15:51:11.051039Z","shell.execute_reply.started":"2022-09-23T15:51:10.958354Z","shell.execute_reply":"2022-09-23T15:51:11.049763Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(test_1.shape[0]):\n    \n    if test_1['Spa'][i] > 3000:\n        \n        print(i,submission['Transported'][i])\n        \n        submission['Transported'][i] = False","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(test_1.shape[0]):\n    \n    if test_1['VRDeck'][i] > 4000:\n        \n        print(i,submission['Transported'][i])\n        \n        submission['Transported'][i] = False","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"VR_Tran = []\nfor i in range(test_1.shape[0]):\n    \n    if test_1['RoomService'][i] > 3682: \n        \n        print(i,submission['Transported'][i])\n        \n        submission['Transported'][i] = False","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(test_1.shape[0]):\n    \n    if test_1['FoodCourt'][i] > 17000:\n        \n        print(i,test_1['FoodCourt'][i],submission['Transported'][i])\n        \n        submission['Transported'][i] = True","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(test_1.shape[0]):\n    \n    if test_1['ShoppingMall'][i] > 5000:\n        \n        print(i,submission['Transported'][i])\n        \n        submission['Transported'][i] = True","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","metadata":{},"execution_count":null,"outputs":[]}]}