{
  "id": 500042,
  "title": "Some good practices for experiment tracking",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/500042",
  "author_name": "Gunes Evitan",
  "post_date": "2024-05-04T05:17:24.102000",
  "votes": 18,
  "comment_count": 0,
  "views": 0,
  "content": "<p>It's not easy to track your experiments when there are 368 targets so I'll share how I do it in this competition.</p>\n<p>I always make this visualization in my projects. It shows how much validation scores deviate between folds and how much mean validation score is different than oof score. In this case mean r2 score is different than oof r2 score so I needed to do some investigation.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fd33e37356148d30dc7f65ea73cdff432%2Fglobal_scores.png?generation=1714796594337547&amp;alt=media\"></p>\n<p>Same visualization can be done for multiple targets. This image is so small when it is embedded but it can be downloaded to see the details. It shows the same thing for all single targets but targets are bars and each metric is another bar plot. We can see how unstable each single target is on fold level.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fecc8993a199fe918dcbdbab264a1f849%2Fsingle_target_scores.png?generation=1714796719649587&amp;alt=media\"></p>\n<p>I think this is one of the most important visualizations since it will cover 360 of the targets. I saw this on ClimSim paper and it's a good way to display scores of target groups. X-axis is the 60 vertical levels and y-axis is the score (rmse, mae or r2 score). Blue line is the mean fold score, blue area is its confidence interval (std). Orange line is the oof score.</p>\n<p>It can be seen that while rmse and mae is stable on all levels, r2 score deviates between folds. This could be happening due to random split changing the target mean in that fold since oof score is stable.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F8d82ff44564bf92a3dd400c9f6c89518%2Fptend_q0001_scores.png?generation=1714796895704005&amp;alt=media\"></p>\n<p>We can also see that r2 score deviation between folds is happening because of ptend_q_0002 somewhere around 28.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F2bbfb44daa91ad4b25ad0657b925511f%2Fptend_q0002_scores.png?generation=1714796905823493&amp;alt=media\"></p>\n<p>Finally, we can visualize each target, prediction and test prediction on top of each other to compare them. ptend_q_0002 was the only one with negative r2 score so that was the culprit.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fb062808d64a3e3f1dcd29c72bf30502c%2Fcam_out_FLWDS.png?generation=1714799569903939&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F0e686c759add71dc3d6b02e0fe96d2bf%2Fptend_q0002_25.png?generation=1714799604025334&amp;alt=media\"></p>",
  "messages": [
    {
      "id": 2792150,
      "postDate": "2024-05-04T05:17:24.103Z",
      "content": "<p>It's not easy to track your experiments when there are 368 targets so I'll share how I do it in this competition.</p>\n<p>I always make this visualization in my projects. It shows how much validation scores deviate between folds and how much mean validation score is different than oof score. In this case mean r2 score is different than oof r2 score so I needed to do some investigation.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fd33e37356148d30dc7f65ea73cdff432%2Fglobal_scores.png?generation=1714796594337547&amp;alt=media\"></p>\n<p>Same visualization can be done for multiple targets. This image is so small when it is embedded but it can be downloaded to see the details. It shows the same thing for all single targets but targets are bars and each metric is another bar plot. We can see how unstable each single target is on fold level.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fecc8993a199fe918dcbdbab264a1f849%2Fsingle_target_scores.png?generation=1714796719649587&amp;alt=media\"></p>\n<p>I think this is one of the most important visualizations since it will cover 360 of the targets. I saw this on ClimSim paper and it's a good way to display scores of target groups. X-axis is the 60 vertical levels and y-axis is the score (rmse, mae or r2 score). Blue line is the mean fold score, blue area is its confidence interval (std). Orange line is the oof score.</p>\n<p>It can be seen that while rmse and mae is stable on all levels, r2 score deviates between folds. This could be happening due to random split changing the target mean in that fold since oof score is stable.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F8d82ff44564bf92a3dd400c9f6c89518%2Fptend_q0001_scores.png?generation=1714796895704005&amp;alt=media\"></p>\n<p>We can also see that r2 score deviation between folds is happening because of ptend_q_0002 somewhere around 28.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F2bbfb44daa91ad4b25ad0657b925511f%2Fptend_q0002_scores.png?generation=1714796905823493&amp;alt=media\"></p>\n<p>Finally, we can visualize each target, prediction and test prediction on top of each other to compare them. ptend_q_0002 was the only one with negative r2 score so that was the culprit.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fb062808d64a3e3f1dcd29c72bf30502c%2Fcam_out_FLWDS.png?generation=1714799569903939&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F0e686c759add71dc3d6b02e0fe96d2bf%2Fptend_q0002_25.png?generation=1714799604025334&amp;alt=media\"></p>",
      "rawMarkdown": "It's not easy to track your experiments when there are 368 targets so I'll share how I do it in this competition.\n\nI always make this visualization in my projects. It shows how much validation scores deviate between folds and how much mean validation score is different than oof score. In this case mean r2 score is different than oof r2 score so I needed to do some investigation.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fd33e37356148d30dc7f65ea73cdff432%2Fglobal_scores.png?generation=1714796594337547&alt=media)\n\nSame visualization can be done for multiple targets. This image is so small when it is embedded but it can be downloaded to see the details. It shows the same thing for all single targets but targets are bars and each metric is another bar plot. We can see how unstable each single target is on fold level.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fecc8993a199fe918dcbdbab264a1f849%2Fsingle_target_scores.png?generation=1714796719649587&alt=media)\n\nI think this is one of the most important visualizations since it will cover 360 of the targets. I saw this on ClimSim paper and it's a good way to display scores of target groups. X-axis is the 60 vertical levels and y-axis is the score (rmse, mae or r2 score). Blue line is the mean fold score, blue area is its confidence interval (std). Orange line is the oof score.\n\nIt can be seen that while rmse and mae is stable on all levels, r2 score deviates between folds. This could be happening due to random split changing the target mean in that fold since oof score is stable.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F8d82ff44564bf92a3dd400c9f6c89518%2Fptend_q0001_scores.png?generation=1714796895704005&alt=media)\n\nWe can also see that r2 score deviation between folds is happening because of ptend_q_0002 somewhere around 28.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F2bbfb44daa91ad4b25ad0657b925511f%2Fptend_q0002_scores.png?generation=1714796905823493&alt=media)\n\nFinally, we can visualize each target, prediction and test prediction on top of each other to compare them. ptend_q_0002 was the only one with negative r2 score so that was the culprit.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fb062808d64a3e3f1dcd29c72bf30502c%2Fcam_out_FLWDS.png?generation=1714799569903939&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F0e686c759add71dc3d6b02e0fe96d2bf%2Fptend_q0002_25.png?generation=1714799604025334&alt=media)\n",
      "votes": 18
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2792150": "It's not easy to track your experiments when there are 368 targets so I'll share how I do it in this competition.\n\nI always make this visualization in my projects. It shows how much validation scores deviate between folds and how much mean validation score is different than oof score. In this case mean r2 score is different than oof r2 score so I needed to do some investigation.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fd33e37356148d30dc7f65ea73cdff432%2Fglobal_scores.png?generation=1714796594337547&alt=media)\n\nSame visualization can be done for multiple targets. This image is so small when it is embedded but it can be downloaded to see the details. It shows the same thing for all single targets but targets are bars and each metric is another bar plot. We can see how unstable each single target is on fold level.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fecc8993a199fe918dcbdbab264a1f849%2Fsingle_target_scores.png?generation=1714796719649587&alt=media)\n\nI think this is one of the most important visualizations since it will cover 360 of the targets. I saw this on ClimSim paper and it's a good way to display scores of target groups. X-axis is the 60 vertical levels and y-axis is the score (rmse, mae or r2 score). Blue line is the mean fold score, blue area is its confidence interval (std). Orange line is the oof score.\n\nIt can be seen that while rmse and mae is stable on all levels, r2 score deviates between folds. This could be happening due to random split changing the target mean in that fold since oof score is stable.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F8d82ff44564bf92a3dd400c9f6c89518%2Fptend_q0001_scores.png?generation=1714796895704005&alt=media)\n\nWe can also see that r2 score deviation between folds is happening because of ptend_q_0002 somewhere around 28.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F2bbfb44daa91ad4b25ad0657b925511f%2Fptend_q0002_scores.png?generation=1714796905823493&alt=media)\n\nFinally, we can visualize each target, prediction and test prediction on top of each other to compare them. ptend_q_0002 was the only one with negative r2 score so that was the culprit.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Fb062808d64a3e3f1dcd29c72bf30502c%2Fcam_out_FLWDS.png?generation=1714799569903939&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F0e686c759add71dc3d6b02e0fe96d2bf%2Fptend_q0002_25.png?generation=1714799604025334&alt=media)\n"
  }
}