{
  "id": 514698,
  "title": "Visualization of overfitting (?)",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/514698",
  "author_name": "chumajin",
  "post_date": "2024-06-25T05:53:45.689000",
  "votes": 31,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I found the analysis of my model's validation results interesting and decided to share it. It is likely overfitting.</p>\n<p>As discussed <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490\" target=\"_blank\">here</a>, ptend_q0003_15 has a significant outlier with a risk of 1270σ, making it one of the targets that poses a major risk.</p>\n<p>In my model, at epoch 7 during training, the metric r2  of ptend_q0003_15 was around 0.764, but as training progressed (epoch 13), it became -0.10186. (Despite other ptend_q0003 R2 values increasing like this.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fbe9e9bf15881356292b8935db21d96ec%2FClipboard04.jpg?generation=1719305382236571&amp;alt=media\"></p>\n<p>Furthermore, as I proceeded with the analysis, focusing exclusively on ptend_q0003_15, the relationship between epoch and R2 is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F631277899f17a837767e1a5688abed19%2FClipboard05.jpg?generation=1719305399783920&amp;alt=media\"></p>\n<p>It is evident that from around epoch 10, the R2 value clearly starts to decrease. </p>\n<p>To visualize this, I compared the ground truth and predicted values at epoch 7 and epoch 13, as shown in the following figure.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F8a6556ec83150798f54adbcb06074578%2FClipboard06.jpg?generation=1719305423607927&amp;alt=media\"></p>\n<p>The results at epoch 13, which appear to be overfitting, show that the predicted values increase as the ground truth values become larger.</p>\n<p>Therefore, I tried replacing only the results of ptend_q0003_15 at epoch 13 with the results at epoch 7, thinking that both the cross-validation (CV) and leaderboard (LB) scores would improve.</p>\n<p>The results are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>epoch 7 only</td>\n<td>0.725157</td>\n<td>0.72251</td>\n</tr>\n<tr>\n<td>epoch 13 only</td>\n<td>0.736532</td>\n<td>0.73701</td>\n</tr>\n<tr>\n<td>replaced result</td>\n<td>0.738868</td>\n<td>0.7366</td>\n</tr>\n</tbody>\n</table>\n<p>As a result of the replacement, the CV score naturally improved, but the Public LB score worsened. In other words, this means that epoch 13, which has an R2 score of -0.10186 for ptend_q0003_15 in validation, performed better than epoch 7, which has a score of 0.764. Considering the graph of the ground truth and predictions above, it suggests that there might be no outlier plots in the public LB data, and the plots where the predictions and ground truth partially match better at epoch 13 might be present in the public LB.</p>\n<p>However, it is uncertain what the results will be on the Private LB (a little shake will be occur because of this?) As <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> stated in previous <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490\" target=\"_blank\">post</a>.</p>\n<p>If anyone has experienced a similar situation and found a method to improve this overfit, please teach me !<br>\nenjoy!</p>",
  "messages": [
    {
      "id": 2888891,
      "postDate": "2024-06-25T05:53:45.690Z",
      "content": "<p>I found the analysis of my model's validation results interesting and decided to share it. It is likely overfitting.</p>\n<p>As discussed <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490\" target=\"_blank\">here</a>, ptend_q0003_15 has a significant outlier with a risk of 1270σ, making it one of the targets that poses a major risk.</p>\n<p>In my model, at epoch 7 during training, the metric r2  of ptend_q0003_15 was around 0.764, but as training progressed (epoch 13), it became -0.10186. (Despite other ptend_q0003 R2 values increasing like this.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fbe9e9bf15881356292b8935db21d96ec%2FClipboard04.jpg?generation=1719305382236571&amp;alt=media\"></p>\n<p>Furthermore, as I proceeded with the analysis, focusing exclusively on ptend_q0003_15, the relationship between epoch and R2 is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F631277899f17a837767e1a5688abed19%2FClipboard05.jpg?generation=1719305399783920&amp;alt=media\"></p>\n<p>It is evident that from around epoch 10, the R2 value clearly starts to decrease. </p>\n<p>To visualize this, I compared the ground truth and predicted values at epoch 7 and epoch 13, as shown in the following figure.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F8a6556ec83150798f54adbcb06074578%2FClipboard06.jpg?generation=1719305423607927&amp;alt=media\"></p>\n<p>The results at epoch 13, which appear to be overfitting, show that the predicted values increase as the ground truth values become larger.</p>\n<p>Therefore, I tried replacing only the results of ptend_q0003_15 at epoch 13 with the results at epoch 7, thinking that both the cross-validation (CV) and leaderboard (LB) scores would improve.</p>\n<p>The results are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>epoch 7 only</td>\n<td>0.725157</td>\n<td>0.72251</td>\n</tr>\n<tr>\n<td>epoch 13 only</td>\n<td>0.736532</td>\n<td>0.73701</td>\n</tr>\n<tr>\n<td>replaced result</td>\n<td>0.738868</td>\n<td>0.7366</td>\n</tr>\n</tbody>\n</table>\n<p>As a result of the replacement, the CV score naturally improved, but the Public LB score worsened. In other words, this means that epoch 13, which has an R2 score of -0.10186 for ptend_q0003_15 in validation, performed better than epoch 7, which has a score of 0.764. Considering the graph of the ground truth and predictions above, it suggests that there might be no outlier plots in the public LB data, and the plots where the predictions and ground truth partially match better at epoch 13 might be present in the public LB.</p>\n<p>However, it is uncertain what the results will be on the Private LB (a little shake will be occur because of this?) As <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> stated in previous <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490\" target=\"_blank\">post</a>.</p>\n<p>If anyone has experienced a similar situation and found a method to improve this overfit, please teach me !<br>\nenjoy!</p>",
      "rawMarkdown": "I found the analysis of my model's validation results interesting and decided to share it. It is likely overfitting.\n\nAs discussed [here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490), ptend_q0003_15 has a significant outlier with a risk of 1270σ, making it one of the targets that poses a major risk.\n\nIn my model, at epoch 7 during training, the metric r2  of ptend_q0003_15 was around 0.764, but as training progressed (epoch 13), it became -0.10186. (Despite other ptend_q0003 R2 values increasing like this.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fbe9e9bf15881356292b8935db21d96ec%2FClipboard04.jpg?generation=1719305382236571&alt=media)\n\nFurthermore, as I proceeded with the analysis, focusing exclusively on ptend_q0003_15, the relationship between epoch and R2 is as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F631277899f17a837767e1a5688abed19%2FClipboard05.jpg?generation=1719305399783920&alt=media)\n\nIt is evident that from around epoch 10, the R2 value clearly starts to decrease. \n\nTo visualize this, I compared the ground truth and predicted values at epoch 7 and epoch 13, as shown in the following figure.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F8a6556ec83150798f54adbcb06074578%2FClipboard06.jpg?generation=1719305423607927&alt=media)\n\nThe results at epoch 13, which appear to be overfitting, show that the predicted values increase as the ground truth values become larger.\n\nTherefore, I tried replacing only the results of ptend_q0003_15 at epoch 13 with the results at epoch 7, thinking that both the cross-validation (CV) and leaderboard (LB) scores would improve.\n\nThe results are as follows:\n\n| model           | CV       | LB      |\n|-----------------|----------|---------|\n| epoch 7 only    | 0.725157 | 0.72251 |\n| epoch 13 only   | 0.736532 | 0.73701 |\n| replaced result | 0.738868 | 0.7366  |\n\n\n\nAs a result of the replacement, the CV score naturally improved, but the Public LB score worsened. In other words, this means that epoch 13, which has an R2 score of -0.10186 for ptend_q0003_15 in validation, performed better than epoch 7, which has a score of 0.764. Considering the graph of the ground truth and predictions above, it suggests that there might be no outlier plots in the public LB data, and the plots where the predictions and ground truth partially match better at epoch 13 might be present in the public LB.\n\nHowever, it is uncertain what the results will be on the Private LB (a little shake will be occur because of this?) As @tatamikenn stated in previous [post](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490).\n\nIf anyone has experienced a similar situation and found a method to improve this overfit, please teach me !\nenjoy!\n",
      "votes": 29
    },
    {
      "id": 2888968,
      "postDate": "2024-06-25T07:07:15.080Z",
      "content": "<p>Amazing report… Thanks for sharing. I also come to the same conclusion that first few targets after non-zero weighted ones are the most problematic.</p>\n<p>I'm tracking mse and mae too and they seem to be more stable compared to r2 score. I guess your loss is stable too but few outliers in those columns are messing up the r2 score. I think there might be a shake happening due to this but I can't predict the scale of it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Ff4a731836bb4ea435148cd8b3a55b811%2Fptend_q0003_scores.png?generation=1719298852296657&amp;alt=media\" alt=\"q0003\"></p>",
      "rawMarkdown": "Amazing report... Thanks for sharing. I also come to the same conclusion that first few targets after non-zero weighted ones are the most problematic.\n\nI'm tracking mse and mae too and they seem to be more stable compared to r2 score. I guess your loss is stable too but few outliers in those columns are messing up the r2 score. I think there might be a shake happening due to this but I can't predict the scale of it.\n\n![q0003](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Ff4a731836bb4ea435148cd8b3a55b811%2Fptend_q0003_scores.png?generation=1719298852296657&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 2889374,
          "postDate": "2024-06-25T13:00:51.917Z",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> Thank you for your comment! Indeed, MSE and MAE seem to be much more stable compared to the R2 score.</p>\n<p>This time, it seems that my validation data for ptend_q0003_15 happened to contain outliers, but it is also possible that other columns in the private dataset may contain outliers and could have a negative impact.</p>",
          "rawMarkdown": "@gunesevitan Thank you for your comment! Indeed, MSE and MAE seem to be much more stable compared to the R2 score.\n\nThis time, it seems that my validation data for ptend_q0003_15 happened to contain outliers, but it is also possible that other columns in the private dataset may contain outliers and could have a negative impact.",
          "votes": 1,
          "replies": [
            {
              "id": 2889979,
              "postDate": "2024-06-25T19:39:11.090Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2890756,
      "postDate": "2024-06-26T09:32:42.903Z",
      "content": "<p>Thanks for sharing your results. I tried something similar to this a while ago but failed🤣</p>",
      "rawMarkdown": "Thanks for sharing your results. I tried something similar to this a while ago but failed🤣",
      "votes": 1,
      "replies": [
        {
          "id": 2891154,
          "postDate": "2024-06-26T14:10:23.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/ajobseeker\" target=\"_blank\">@ajobseeker</a> Thank you for your comment and your trial report! While the extent of improvement is uncertain, this type of post-processing might work if done well.</p>",
          "rawMarkdown": "@ajobseeker Thank you for your comment and your trial report! While the extent of improvement is uncertain, this type of post-processing might work if done well."
        }
      ]
    },
    {
      "id": 2892758,
      "postDate": "2024-06-27T12:09:13.183Z",
      "content": "<p>Super helpful report… thanks for sharing.</p>",
      "rawMarkdown": "Super helpful report… thanks for sharing.\n\n\n"
    },
    {
      "id": 2891890,
      "postDate": "2024-06-27T01:34:37.903Z",
      "content": "<p>What happened with a mean of epochs 7 and 13 on ptend_q0003_15 / CV+LB result?</p>",
      "rawMarkdown": "What happened with a mean of epochs 7 and 13 on ptend_q0003_15 / CV+LB result?"
    },
    {
      "id": 2891867,
      "postDate": "2024-06-27T00:28:25.293Z",
      "content": "<p>Super helpful report… thanks for sharing.</p>",
      "rawMarkdown": "Super helpful report... thanks for sharing."
    },
    {
      "id": 2896487,
      "postDate": "2024-06-29T20:32:59.497Z",
      "content": "<p>This is really useful. Thank you!</p>",
      "rawMarkdown": "This is really useful. Thank you!"
    }
  ],
  "comments": [
    {
      "id": 2888968,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-06-25T07:07:15.080000",
      "content": "<p>Amazing report… Thanks for sharing. I also come to the same conclusion that first few targets after non-zero weighted ones are the most problematic.</p>\n<p>I'm tracking mse and mae too and they seem to be more stable compared to r2 score. I guess your loss is stable too but few outliers in those columns are messing up the r2 score. I think there might be a shake happening due to this but I can't predict the scale of it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Ff4a731836bb4ea435148cd8b3a55b811%2Fptend_q0003_scores.png?generation=1719298852296657&amp;alt=media\" alt=\"q0003\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2889374,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-06-25T13:00:51.917000",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> Thank you for your comment! Indeed, MSE and MAE seem to be much more stable compared to the R2 score.</p>\n<p>This time, it seems that my validation data for ptend_q0003_15 happened to contain outliers, but it is also possible that other columns in the private dataset may contain outliers and could have a negative impact.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2889979,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-25T19:39:11.090000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2890756,
      "author_name": "HB",
      "author_url": "",
      "post_date": "2024-06-26T09:32:42.903000",
      "content": "<p>Thanks for sharing your results. I tried something similar to this a while ago but failed🤣</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2891154,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-06-26T14:10:23.100000",
          "content": "<p><a href=\"https://www.kaggle.com/ajobseeker\" target=\"_blank\">@ajobseeker</a> Thank you for your comment and your trial report! While the extent of improvement is uncertain, this type of post-processing might work if done well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2892758,
      "author_name": "MrSimple",
      "author_url": "",
      "post_date": "2024-06-27T12:09:13.183000",
      "content": "<p>Super helpful report… thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2891890,
      "author_name": "Rob Freeman",
      "author_url": "",
      "post_date": "2024-06-27T01:34:37.903000",
      "content": "<p>What happened with a mean of epochs 7 and 13 on ptend_q0003_15 / CV+LB result?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2891867,
      "author_name": "Joe Li",
      "author_url": "",
      "post_date": "2024-06-27T00:28:25.293000",
      "content": "<p>Super helpful report… thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2896487,
      "author_name": "DraganPinsent98",
      "author_url": "",
      "post_date": "2024-06-29T20:32:59.497000",
      "content": "<p>This is really useful. Thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2888891": "I found the analysis of my model's validation results interesting and decided to share it. It is likely overfitting.\n\nAs discussed [here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490), ptend_q0003_15 has a significant outlier with a risk of 1270σ, making it one of the targets that poses a major risk.\n\nIn my model, at epoch 7 during training, the metric r2  of ptend_q0003_15 was around 0.764, but as training progressed (epoch 13), it became -0.10186. (Despite other ptend_q0003 R2 values increasing like this.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fbe9e9bf15881356292b8935db21d96ec%2FClipboard04.jpg?generation=1719305382236571&alt=media)\n\nFurthermore, as I proceeded with the analysis, focusing exclusively on ptend_q0003_15, the relationship between epoch and R2 is as follows:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F631277899f17a837767e1a5688abed19%2FClipboard05.jpg?generation=1719305399783920&alt=media)\n\nIt is evident that from around epoch 10, the R2 value clearly starts to decrease. \n\nTo visualize this, I compared the ground truth and predicted values at epoch 7 and epoch 13, as shown in the following figure.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F8a6556ec83150798f54adbcb06074578%2FClipboard06.jpg?generation=1719305423607927&alt=media)\n\nThe results at epoch 13, which appear to be overfitting, show that the predicted values increase as the ground truth values become larger.\n\nTherefore, I tried replacing only the results of ptend_q0003_15 at epoch 13 with the results at epoch 7, thinking that both the cross-validation (CV) and leaderboard (LB) scores would improve.\n\nThe results are as follows:\n\n| model           | CV       | LB      |\n|-----------------|----------|---------|\n| epoch 7 only    | 0.725157 | 0.72251 |\n| epoch 13 only   | 0.736532 | 0.73701 |\n| replaced result | 0.738868 | 0.7366  |\n\n\n\nAs a result of the replacement, the CV score naturally improved, but the Public LB score worsened. In other words, this means that epoch 13, which has an R2 score of -0.10186 for ptend_q0003_15 in validation, performed better than epoch 7, which has a score of 0.764. Considering the graph of the ground truth and predictions above, it suggests that there might be no outlier plots in the public LB data, and the plots where the predictions and ground truth partially match better at epoch 13 might be present in the public LB.\n\nHowever, it is uncertain what the results will be on the Private LB (a little shake will be occur because of this?) As @tatamikenn stated in previous [post](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490).\n\nIf anyone has experienced a similar situation and found a method to improve this overfit, please teach me !\nenjoy!\n",
    "2888968": "Amazing report... Thanks for sharing. I also come to the same conclusion that first few targets after non-zero weighted ones are the most problematic.\n\nI'm tracking mse and mae too and they seem to be more stable compared to r2 score. I guess your loss is stable too but few outliers in those columns are messing up the r2 score. I think there might be a shake happening due to this but I can't predict the scale of it.\n\n![q0003](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Ff4a731836bb4ea435148cd8b3a55b811%2Fptend_q0003_scores.png?generation=1719298852296657&alt=media)",
    "2890756": "Thanks for sharing your results. I tried something similar to this a while ago but failed🤣",
    "2892758": "Super helpful report… thanks for sharing.\n\n\n",
    "2891890": "What happened with a mean of epochs 7 and 13 on ptend_q0003_15 / CV+LB result?",
    "2891867": "Super helpful report... thanks for sharing.",
    "2896487": "This is really useful. Thank you!"
  }
}