{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":56537,"databundleVersionId":8015876,"sourceType":"competition"},{"sourceId":8437587,"sourceType":"datasetVersion","datasetId":5025835},{"sourceId":177791851,"sourceType":"kernelVersion"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](https://cdn-images-1.medium.com/max/1000/1*dmj7e_r42GkIKSB2KfF2Mw.png)","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>♨️Description</p></div>\n\n- The participants of the challenges, in order to get the best score, Ensembling the results of their notebooks or others and usually use the trial and error method to determine the best coefficients. But Ensembling the best scores is not always successful. That is, in most cases, only by Ensembling some results, the score will improve. Also, the only option for most participants to choose the right results for Ensembling is the trial and error method.\n\n- When the number of answer columns is more than one, the darkness increases. Because it is not known that after finding the right coefficient for Ensembling the first columns, the same coefficient is optimal for Ensembling the next columns. Correlation of projected columns, when you want to Ensembling two columns together, like a candle in the dark can help you reach your destination.\n\n- For example, when the correlation between column A and column B is negative, Ensembling should not be performed on these two columns. That is, if the public score for result A is better than the public score for result B, the first choice is result A alone and the second choice is result B alone, and no linear combination between these two columns can be good.\n\n- It is obvious that when the correlation of column A and column B is a positive number, it is still necessary to see how far the correlation value is from zero and how close it is to one. In addition, it should be seen how much the public score of result A is better than the public score for result B. For example, if the correlation is very close to one and the public score of result A is much better, result A alone is a good option and Ensembling cannot help in improving the score.\n\n<p style=\"border-bottom: 10px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nimport random\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\nfrom glob import glob\nfrom tqdm import tqdm\nfrom scipy import stats\nfrom pathlib import Path\nfrom itertools import groupby\n# ..................................\nimport matplotlib.pyplot as plt\nimport plotly.figure_factory as ff\nimport plotly.express as px\n%matplotlib inline","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-05-17T08:09:21.909894Z","iopub.execute_input":"2024-05-17T08:09:21.91086Z","iopub.status.idle":"2024-05-17T08:09:25.801641Z","shell.execute_reply.started":"2024-05-17T08:09:21.910821Z","shell.execute_reply":"2024-05-17T08:09:25.800563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install datatable","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-05-17T08:09:25.804019Z","iopub.execute_input":"2024-05-17T08:09:25.804652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import datatable as dt \nfrom sklearn.metrics import r2_score\n!ls ../input/*","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsample = dt.fread('../input/leap-atmospheric-physics-ai-climsim/sample_submission.csv').to_pandas()\nidx = sample.pop('sample_id')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = sample.astype(float)\nsample.iloc[:3]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>♨️Import Files</p></div>\n\n<p style=\"border-bottom: 10px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"%%time\n# My Notebook - LEAP♨️Ensembling & Correlation\nsub_54731 = dt.fread('../input/leap-ensembling-correlation/submission.csv').to_pandas().iloc[:, 1:].copy()\nsub_54731.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# through MLP - Thanks to: @farisalahmdi\nsub_52173 = dt.fread('../input/leap52173/submission.csv').to_pandas().iloc[:, 1:].copy()\nsub_52173.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>♨️Auxiliary Functions</p></div>\n\n<p style=\"border-bottom: 10px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"# :::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::: 1\ndef df_creator(dfx, dfy, n):\n    df_sub = pd.DataFrame(columns=['dfx','dfy'])\n    \n    df_sub['dfx'] = dfx.iloc[:, n].copy()\n    df_sub['dfy'] = dfy.iloc[:, n].copy()\n        \n    return df_sub  \n\n# :::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::: 2 \ndef df_corr(df_sub):   \n    corr = df_sub.corr(numeric_only=True).round(3)  \n    \n    corr_list = list(corr.iloc[0])[1:]\n    return corr_list\n\n# :::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::: 3\ndef heatmap(dfx, dfy):\n    N = random.randrange(dfx.shape[1])\n    \n    print('\\n\\nHeatmap (For a random column)')\n    print(':' *40)\n    print('Column Name :', list(dfx.columns)[N])\n    print('Column Number :', N)\n    print(':' *40, '\\n')\n    \n    df_sub = df_creator(dfx, dfy, N)\n    corr_matrix = df_sub.corr()\n    fig = plt.figure(figsize=(4,3));\n\n    cmap=sns.diverging_palette(240, 10, s=75, l=50, sep=1, n=6, center='light', as_cmap=False);\n    sns.heatmap(corr_matrix, center=0, annot=True, cmap=cmap, linewidths=2);\n    plt.suptitle(f'Heatmap (N={N})', y=0.95, fontsize=12, c='darkred');\n    plt.show()\n\n# :::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::: 4\ndef ensembling_histograms(submission, dfx, dfy):\n    N = random.randrange(dfx.shape[1])\n    \n    print('\\n\\nEnsembling Histograms (For a random column)')\n    print(':' *50)\n    print('Column Name :', list(dfx.columns)[N])\n    print('Column Number :', N)\n    print(':' *50)\n    \n    hist_data = [submission.iloc[:, N], dfx.iloc[:, N], dfy.iloc[:, N]]\n    group_labels = ['Generated', 'First Results (Main)', 'Second Results (Support)']\n    \n    fig = ff.create_distplot(hist_data, group_labels, bin_size=.2, show_hist=False, show_rug=False)\n    fig.show()    \n\n# :::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::: 5\ndef info_corr(corr_limit, counter_x, counter_y, dfx, dfy):\n    counter_z = dfx.shape[1] - (counter_x + counter_y)\n    \n    print('\\nCorrelation information (For all columns)')\n    print(':' *70)\n    print('A Percent =', round(counter_y / dfx.shape[1], 3)) \n    print('A: Correlation less than Zero')\n    print('The number of columns evaluated with only the second result =', counter_y) \n    print('-  ' *24)\n    print('B Percent =', round(counter_z / dfx.shape[1], 3)) \n    print('B: Correlation more than Zero and less than corr_limit')\n    print('The number of columns evaluated by Ensembling =', counter_z)\n    print('-  ' *24)\n    print('C Percent =', round(counter_x / dfx.shape[1], 3))\n    print('C: Correlation more than corr_limit')\n    print('The number of columns evaluated with only the first result =', counter_x)\n    print('-  ' *24)   \n    print('The correlation limit that was considered as the basis =', corr_limit)\n    print(':' *70 ,'\\n\\n')\n    \n    columns = ['A: Correlation less than Zero','B: Correlation more than Zero and less than corr_limit','C: Correlation more than corr_limit']\n    data = [[ round(counter_y / dfx.shape[1], 3), round(counter_z / dfx.shape[1], 3), round(counter_x / dfx.shape[1], 3)]]\n    de_data = pd.DataFrame(data=data , columns=columns)\n    \n    sns.set()\n    de_data.plot(kind='barh', stacked=True, figsize=(10,1), color=['pink','violet','purple'])\n    plt.gca().set_facecolor('lightyellow')\n    plt.legend(fontsize=10, loc=3, bbox_to_anchor=(0, 1))\n    plt.show() \n\n# :::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::: 6\ndef ensembling_scatter(submission, dfx, dfy):   \n    \n    X  = dfx  # main\n    Y1 = dfy  # support\n    Y2 = submission\n    \n    plt.figure(figsize=(10, 5), facecolor='lightyellow')\n    plt.title('Scatter Graph (For all columns)\\n', fontsize=12)   \n\n    plt.scatter(X, Y1, s=2.5, label='dfy - Support', c='darkcyan')    \n    plt.scatter(X, Y2, s=2.5, label='Generated', c='red')\n    plt.scatter(X, X , s=4.0, label='dfx - Main(X=Y)', c='orange')\n     \n    plt.gca().set_facecolor('lightgreen')\n    plt.legend(fontsize=10, loc=4)\n    \n    # plt.savefig('scatter101.png')\n    plt.show()     \n\n# :::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::: 7\ndef plot_corr(tcorr):  \n    display(tcorr.iloc[:, :4].style.background_gradient(cmap='Pastel1', axis=None, vmin=0, vmax=1.0))\n\n    plt.figure(figsize=(10, 5), facecolor='lightyellow')\n    plt.title('Percentage of States\\n', fontsize=12)   \n    plt.xlabel('Correlation Limit')\n    plt.xticks(range(11), round(tcorr.iloc[:,0], 1))\n\n    plt.plot(tcorr.iloc[:,1], color='orange', lw=2, label='A: Correlation less than Zero')\n    plt.plot(tcorr.iloc[:,2], color='darkcyan', lw=2, label='B: Correlation more than Zero and less than corr_limit')\n    plt.plot(tcorr.iloc[:,3], color='red', lw=2, label='C: Correlation more than corr_limit')\n     \n    plt.gca().set_facecolor('lightgreen')\n    plt.legend(fontsize=10, loc=0)\n    \n    # plt.savefig('plot101.png')\n    plt.show()  \n    ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>♨️Correlation & Tuning Functions</p></div>\n\n<p style=\"border-bottom: 10px solid navy\"></p>","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>(1) generate_corr_coeff (dfx, dfy, corr_limit, coeff)</p></div>\n\n- This function takes the regression results and prediction of two notebooks, each of these results can have thousands of columns. The first result (dfx) is called \"Main\" and the second result (dfy) is called \"Support\".\n\n- This function ignores the column corresponding to \"Main\" when the correlation of two similar columns in \"Main\" and \"Support\" becomes negative. We can also set \"corr_limit\" to be less than one (and greater than zero). When the correlation of two similar columns in \"Main\" and \"Support\" is greater than \"corr_limit\", this function ignores the column corresponding to \"Support\".\n\n- This function performs Ensembling with a coefficient that we define as \"coeff\", only for columns whose correlation is between zero and \"corr_limit\".\n\n- It is obvious that if we change the places of \"Main\" and \"Support\", the result of Ensembling will change and the score may be better with the new setting of \"corr_limit\" and \"coeff\".","metadata":{}},{"cell_type":"code","source":"# Ensemble for two results by determining \"corr_limit & coeff\"\ndef generate_corr_coeff(dfx, dfy, corr_limit, coeff):\n    submission = sample.copy()\n    \n    counter_x = 0\n    counter_y = 0\n    for n in range(dfx.shape[1]):   \n        df_sub = df_creator(dfx, dfy, n)\n        corr_list = df_corr(df_sub)\n        \n        submission.iloc[:, n] = (dfx.iloc[:, n] * coeff) + (dfy.iloc[:, n] * (1.- coeff))\n        \n        if (corr_list[0] > corr_limit):  \n            submission.iloc[:, n] = dfx.iloc[:, n]\n            counter_x += 1  \n            \n        if (corr_list[0] < 0):\n            submission.iloc[:, n] = dfy.iloc[:, n]\n            counter_y += 1\n            \n    print('\\n\\n', ':. ' *12, 'Ensembling (Different coefficients for different columns)', '.: ' *12, '\\n\\n')\n    \n    # heatmap(dfx, dfy)\n    # ensembling_histograms(submission, dfx, dfy)\n    \n    ensembling_scatter(submission, dfx, dfy)\n    info_corr(corr_limit, counter_x, counter_y, dfx, dfy)\n    \n    # display(submission)\n    return submission","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>(2) generate_corr (dfx, dfy, corr_limit)</p></div>\n\n- This function is very similar to the previous function, except that you do not need to specify the \"coeff\". This function calculates a coefficient for Ensembling for both columns whose correlation is between zero and \"corr_limit\". That is, it considers \"coeff\" equal to the correlation value of two columns.\n\n- Certainly, the results of this function are not as good as the first function, but trying this function can be a good image to determine \"coeff\" and \"corr_limit\".","metadata":{}},{"cell_type":"code","source":"# Ensemble for two results by determining \"corr_limit\"\ndef generate_corr(dfx, dfy, corr_limit):\n    submission = sample.copy()\n    \n    counter_x = 0\n    counter_y = 0\n    for n in range(dfx.shape[1]):   \n        df_sub = df_creator(dfx, dfy, n)\n        corr_list = df_corr(df_sub)\n        \n        submission.iloc[:, n] = (dfx.iloc[:, n] * corr_list[0]) + (dfy.iloc[:, n] * (1.- corr_list[0]))\n        \n        if (corr_list[0] > corr_limit):  \n            submission.iloc[:, n] = dfx.iloc[:, n]\n            counter_x += 1  \n            \n        if (corr_list[0] < 0):\n            submission.iloc[:, n] = dfy.iloc[:, n]\n            counter_y += 1\n            \n    print('\\n\\n', ':. ' *12, 'Ensembling (Different coefficients for different columns)', '.: ' *12, '\\n\\n')\n    \n    # heatmap(dfx, dfy)\n    # ensembling_histograms(submission, dfx, dfy)\n    \n    ensembling_scatter(submission, dfx, dfy)\n    info_corr(corr_limit, counter_x, counter_y, dfx, dfy)\n    \n    # display(submission)\n    return submission","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>(3) tuning_corr_limit (dfx, dfy)</p></div>\n\n- Before setting the parameters of the above functions, you can use the following function to see the correlation status of \"Main\" and \"Support\" in the table and graph.\n\n- If the number of prediction columns is thousands of numbers, this function must perform a lot of calculations and the execution of this function takes time.","metadata":{}},{"cell_type":"code","source":"def tuning_corr_limit(dfx, dfy): \n    '''\n    A : Less than Zero\n    B : More than Zero and less than corr_limit\n    C : more than corr_limit\n    '''\n    cname = ['corr_limit','Percentage of A','Percentage of B','Percentage of C','The number of A','The number of B','The number of C']\n    \n    clist = np.arange(0, 1.1, 0.1)\n    cdata = np.zeros((len(clist), 7), dtype=float)\n    tcorr = pd.DataFrame(data=cdata, columns=cname)\n    \n    for c in range(len(clist)):\n        counter_a = 0\n        counter_b = 0\n        counter_c = 0\n        \n        for n in range(dfx.shape[1]): \n            df_sub = df_creator(dfx, dfy, n)\n            corr_list = df_corr(df_sub)\n        \n            if (corr_list[0] < 0):  \n                counter_a += 1 \n            if (corr_list[0] > 0) and (corr_list[0] < clist[c]):\n                counter_b += 1 \n            if (corr_list[0] > clist[c]):\n                counter_c += 1\n        \n        tcorr.iloc[c,0] = clist[c]\n        tcorr.iloc[c,1] = round(counter_a / dfx.shape[1], 3)\n        tcorr.iloc[c,2] = round(counter_b / dfx.shape[1], 3)\n        tcorr.iloc[c,3] = round(counter_c / dfx.shape[1], 3)\n        tcorr.iloc[c,4] = counter_a\n        tcorr.iloc[c,5] = counter_b\n        tcorr.iloc[c,6] = counter_c\n        \n    plot_corr(tcorr)\n    # display(tcorr)\n    return tcorr","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>♨️Ensembling with Correlation Guidance - ECG</p></div>\n\n<p style=\"border-bottom: 10px solid navy\"></p>","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>T u n i n g</p></div>","metadata":{}},{"cell_type":"code","source":"%%time\ntuning_corr_limit(sub_54731, sub_52173)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>E n s e m b l i n g</p></div>","metadata":{}},{"cell_type":"code","source":"%%time\nsub = generate_corr_coeff(sub_54731, sub_52173, 0.85, 0.95)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>S u b m i s s i o n</p></div>","metadata":{}},{"cell_type":"code","source":"%%time\nsub.insert(0, 'sample_id', idx)\nsub.to_csv('submission.csv', index=False)\n!ls","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"border-bottom: 30px solid pink\"></p>\n<p style=\"border-bottom: 25px solid lightcyan\"></p>","metadata":{}}]}