{
  "id": 384408,
  "title": "What is your biggest gap on the validation dataset and on LB? Mine 0.77 vs 0.32:-)",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/384408",
  "author_name": "Andrij",
  "post_date": "2023-02-07T19:46:46.074000",
  "votes": 0,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hello everybody! <br>\nI'm only a few days into the competition, but a quick review showed its \"difficulty\". <br>\nPersonally, I predict a shake-up, but I could be wrong, we'll see. I built an approach that allowed me to get pF1=0.77 on the validation data using one model, on lb I got 0.32, I'm a bit upset about that, but it just says that my approach is wrong, or it didn't work well on the lb part of the data. After looking through the public notebooks and comments, I realized that the gap between val_test and lb is huge. <br>\nSo, who has the biggest gap, and what do you think about it in general. In my opinion, our problem lies precisely in the imbalance of classes. I met with a similar problem for the NLP problem and the bag of words model, then the best solution was to consider only the features (ngrams) that indicate a positive class, such a trick will not work here.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F7ec945d2b2918d7d2ce7ce31ba377417%2F1.PNG?generation=1675799168302642&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2134251,
      "postDate": "2023-02-07T21:15:12.883Z",
      "content": "<p>The metric is indeed quite shaky, but your validation score \"looks to good to be true\". Especially, compared to the values reported by other competitors. Most of the time, when this happens, there is a leak in your validation. Did you not split by patient ID maybe? Is this a super small subset?</p>",
      "rawMarkdown": "The metric is indeed quite shaky, but your validation score \"looks to good to be true\". Especially, compared to the values reported by other competitors. Most of the time, when this happens, there is a leak in your validation. Did you not split by patient ID maybe? Is this a super small subset?",
      "votes": 7,
      "replies": [
        {
          "id": 2134297,
          "postDate": "2023-02-07T21:49:34.027Z",
          "content": "<p>There is such a chance, but what I tried should theoretically eliminate the problem with the leakage of part of the data. In my opinion, a deep neural network (it doesn't matter which one you use) cannot distinguish features that are specific to the positive and negative classes (there is not enough data for the positive class). Most of the features of the negative class found by the network are irrelevant and can be discarded. In general, we do not need features that characterize the negative class, we only need features that characterize the positive class. For me, a black box is a disadvantage of the network in a situation where it is necessary to correctly interpret the results of its work. I tried to work around this limitation.</p>",
          "rawMarkdown": "There is such a chance, but what I tried should theoretically eliminate the problem with the leakage of part of the data. In my opinion, a deep neural network (it doesn't matter which one you use) cannot distinguish features that are specific to the positive and negative classes (there is not enough data for the positive class). Most of the features of the negative class found by the network are irrelevant and can be discarded. In general, we do not need features that characterize the negative class, we only need features that characterize the positive class. For me, a black box is a disadvantage of the network in a situation where it is necessary to correctly interpret the results of its work. I tried to work around this limitation.",
          "replies": [
            {
              "id": 2134667,
              "postDate": "2023-02-08T07:10:41.213Z",
              "content": "<p>I checked everything again, you are absolutely right, it was a date leak. Good luck</p>",
              "rawMarkdown": "I checked everything again, you are absolutely right, it was a date leak. Good luck",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2137542,
      "postDate": "2023-02-10T05:51:46.597Z",
      "content": "<p>Thanks sharing.</p>\n<p>Do  you use some external data?<br>\nOr  pretrain use train data? Then you should take care about  making CV folds.<br>\nAnyway I feel there can be a leak.</p>",
      "rawMarkdown": "Thanks sharing.\n\nDo  you use some external data?\nOr  pretrain use train data? Then you should take care about  making CV folds.\nAnyway I feel there can be a leak.",
      "votes": 1,
      "replies": [
        {
          "id": 2137581,
          "postDate": "2023-02-10T06:51:23.247Z",
          "content": "<p>Hello! I had a data leak. pF1 on the \"correct\" validation sets 0.25) I didn't remove the topic so people wouldn't lose their medals for commenting. Yes, cross-validation helped me. At first, the result on lb was worse compared to the public model, but now everything is fine. I don't trust lb in this contest, it can be accurate, but its less likely than usual to be accurate. In the end, the winners will compete for the item, good luck!</p>",
          "rawMarkdown": "Hello! I had a data leak. pF1 on the \"correct\" validation sets 0.25) I didn't remove the topic so people wouldn't lose their medals for commenting. Yes, cross-validation helped me. At first, the result on lb was worse compared to the public model, but now everything is fine. I don't trust lb in this contest, it can be accurate, but its less likely than usual to be accurate. In the end, the winners will compete for the item, good luck!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2134684,
      "postDate": "2023-02-08T07:32:48.320Z",
      "content": "<p>I just started and made a single submission.</p>\n<pre><code>OOF ROC AUC \nOOF PF1: \nLB PF1: \n</code></pre>",
      "rawMarkdown": "I just started and made a single submission.\n\n```python\nOOF ROC AUC 0.8502\nOOF PF1: 0.3171\nLB PF1: 0.39\n```",
      "votes": 1
    },
    {
      "id": 2134534,
      "postDate": "2023-02-08T04:10:47.390Z",
      "content": "<p>How about your ROC_AUC ??</p>",
      "rawMarkdown": "How about your ROC_AUC ??",
      "votes": 1
    },
    {
      "id": 2134174,
      "postDate": "2023-02-07T20:03:33.787Z",
      "content": "<p>For me single model with val scores pF1: 0.41, F1: 0.57(t=0.5), ROCAUC: 0.84 gives me LB 0.51(t=0.5)</p>\n<p>you can find more about results here<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/380657\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/380657</a><br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333</a></p>",
      "rawMarkdown": "For me single model with val scores pF1: 0.41, F1: 0.57(t=0.5), ROCAUC: 0.84 gives me LB 0.51(t=0.5)\n\nyou can find more about results here\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/380657\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333",
      "votes": 1,
      "replies": [
        {
          "id": 2134176,
          "postDate": "2023-02-07T20:04:30.750Z",
          "content": "<p>Thank you! I'll take a look</p>",
          "rawMarkdown": "Thank you! I'll take a look"
        }
      ]
    },
    {
      "id": 2134840,
      "postDate": "2023-02-08T09:31:34.990Z",
      "content": "<p>0.77 seems too high to be accurate. In my case, the situation is different. I have a validation accuracy of 20% with F1 scores usually between 0.3 and 0.4, which gave me a best score of 0.54. Also, I do not use cross-validation.</p>",
      "rawMarkdown": "0.77 seems too high to be accurate. In my case, the situation is different. I have a validation accuracy of 20% with F1 scores usually between 0.3 and 0.4, which gave me a best score of 0.54. Also, I do not use cross-validation.",
      "votes": 2,
      "replies": [
        {
          "id": 2135013,
          "postDate": "2023-02-08T11:56:53.890Z",
          "content": "<p>Excellent result. I checked my pipeline again and realized that the problem was a data leak. Regarding the differences between lb and val_pF1, I use and as I understand many people also clearly overfitting model. Therefore, it is better to use kfold. Good luck!</p>",
          "rawMarkdown": "Excellent result. I checked my pipeline again and realized that the problem was a data leak. Regarding the differences between lb and val_pF1, I use and as I understand many people also clearly overfitting model. Therefore, it is better to use kfold. Good luck!",
          "votes": 2,
          "replies": [
            {
              "id": 2135114,
              "postDate": "2023-02-08T13:10:24.383Z",
              "content": "<p>Great to hear that! My model also overfits a lot, but I haven't found a solution that doesn't harm my score yet. I'll try to improve my data aug.</p>\n<p>K-Fold cross-validation would definitely be better. Have you tried it? Does it consume significantly more computational time? In my case, I want to conserve my TPU credits as much as possible to try different ideas. Currently, a single training takes me between 1 to 1h30.</p>",
              "rawMarkdown": "Great to hear that! My model also overfits a lot, but I haven't found a solution that doesn't harm my score yet. I'll try to improve my data aug.\n\nK-Fold cross-validation would definitely be better. Have you tried it? Does it consume significantly more computational time? In my case, I want to conserve my TPU credits as much as possible to try different ideas. Currently, a single training takes me between 1 to 1h30."
            },
            {
              "id": 2135141,
              "postDate": "2023-02-08T13:30:36.723Z",
              "content": "<p>For testing, I prepared a dataset with less expansion, then training is faster (models are less accurate, but more models can be tried). We have too small data set for positive class, so we can't trust one model, read other discussions, I see somewhere that kfold improved the result by lb. But taking into account all the above, I don't trust public lb in this competition (more precisely, I trust less than usual).</p>",
              "rawMarkdown": "For testing, I prepared a dataset with less expansion, then training is faster (models are less accurate, but more models can be tried). We have too small data set for positive class, so we can't trust one model, read other discussions, I see somewhere that kfold improved the result by lb. But taking into account all the above, I don't trust public lb in this competition (more precisely, I trust less than usual).",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2134151,
      "postDate": "2023-02-07T19:46:46.073Z",
      "content": "<p>Hello everybody! <br>\nI'm only a few days into the competition, but a quick review showed its \"difficulty\". <br>\nPersonally, I predict a shake-up, but I could be wrong, we'll see. I built an approach that allowed me to get pF1=0.77 on the validation data using one model, on lb I got 0.32, I'm a bit upset about that, but it just says that my approach is wrong, or it didn't work well on the lb part of the data. After looking through the public notebooks and comments, I realized that the gap between val_test and lb is huge. <br>\nSo, who has the biggest gap, and what do you think about it in general. In my opinion, our problem lies precisely in the imbalance of classes. I met with a similar problem for the NLP problem and the bag of words model, then the best solution was to consider only the features (ngrams) that indicate a positive class, such a trick will not work here.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F7ec945d2b2918d7d2ce7ce31ba377417%2F1.PNG?generation=1675799168302642&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello everybody! \nI'm only a few days into the competition, but a quick review showed its \"difficulty\". \nPersonally, I predict a shake-up, but I could be wrong, we'll see. I built an approach that allowed me to get pF1=0.77 on the validation data using one model, on lb I got 0.32, I'm a bit upset about that, but it just says that my approach is wrong, or it didn't work well on the lb part of the data. After looking through the public notebooks and comments, I realized that the gap between val_test and lb is huge. \nSo, who has the biggest gap, and what do you think about it in general. In my opinion, our problem lies precisely in the imbalance of classes. I met with a similar problem for the NLP problem and the bag of words model, then the best solution was to consider only the features (ngrams) that indicate a positive class, such a trick will not work here.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F7ec945d2b2918d7d2ce7ce31ba377417%2F1.PNG?generation=1675799168302642&alt=media)"
    }
  ],
  "comments": [
    {
      "id": 2134251,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2023-02-07T21:15:12.883000",
      "content": "<p>The metric is indeed quite shaky, but your validation score \"looks to good to be true\". Especially, compared to the values reported by other competitors. Most of the time, when this happens, there is a leak in your validation. Did you not split by patient ID maybe? Is this a super small subset?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2134297,
          "author_name": "Andrij",
          "author_url": "",
          "post_date": "2023-02-07T21:49:34.027000",
          "content": "<p>There is such a chance, but what I tried should theoretically eliminate the problem with the leakage of part of the data. In my opinion, a deep neural network (it doesn't matter which one you use) cannot distinguish features that are specific to the positive and negative classes (there is not enough data for the positive class). Most of the features of the negative class found by the network are irrelevant and can be discarded. In general, we do not need features that characterize the negative class, we only need features that characterize the positive class. For me, a black box is a disadvantage of the network in a situation where it is necessary to correctly interpret the results of its work. I tried to work around this limitation.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2134667,
              "author_name": "Andrij",
              "author_url": "",
              "post_date": "2023-02-08T07:10:41.213000",
              "content": "<p>I checked everything again, you are absolutely right, it was a date leak. Good luck</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2137542,
      "author_name": "taruto",
      "author_url": "",
      "post_date": "2023-02-10T05:51:46.597000",
      "content": "<p>Thanks sharing.</p>\n<p>Do  you use some external data?<br>\nOr  pretrain use train data? Then you should take care about  making CV folds.<br>\nAnyway I feel there can be a leak.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2137581,
          "author_name": "Andrij",
          "author_url": "",
          "post_date": "2023-02-10T06:51:23.247000",
          "content": "<p>Hello! I had a data leak. pF1 on the \"correct\" validation sets 0.25) I didn't remove the topic so people wouldn't lose their medals for commenting. Yes, cross-validation helped me. At first, the result on lb was worse compared to the public model, but now everything is fine. I don't trust lb in this contest, it can be accurate, but its less likely than usual to be accurate. In the end, the winners will compete for the item, good luck!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2134684,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-02-08T07:32:48.320000",
      "content": "<p>I just started and made a single submission.</p>\n<pre><code>OOF ROC AUC \nOOF PF1: \nLB PF1: \n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2134534,
      "author_name": "fate",
      "author_url": "",
      "post_date": "2023-02-08T04:10:47.390000",
      "content": "<p>How about your ROC_AUC ??</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2134174,
      "author_name": "A.P.",
      "author_url": "",
      "post_date": "2023-02-07T20:03:33.787000",
      "content": "<p>For me single model with val scores pF1: 0.41, F1: 0.57(t=0.5), ROCAUC: 0.84 gives me LB 0.51(t=0.5)</p>\n<p>you can find more about results here<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/380657\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/380657</a><br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2134176,
          "author_name": "Andrij",
          "author_url": "",
          "post_date": "2023-02-07T20:04:30.750000",
          "content": "<p>Thank you! I'll take a look</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2134840,
      "author_name": "Paul Bacher",
      "author_url": "",
      "post_date": "2023-02-08T09:31:34.990000",
      "content": "<p>0.77 seems too high to be accurate. In my case, the situation is different. I have a validation accuracy of 20% with F1 scores usually between 0.3 and 0.4, which gave me a best score of 0.54. Also, I do not use cross-validation.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2135013,
          "author_name": "Andrij",
          "author_url": "",
          "post_date": "2023-02-08T11:56:53.890000",
          "content": "<p>Excellent result. I checked my pipeline again and realized that the problem was a data leak. Regarding the differences between lb and val_pF1, I use and as I understand many people also clearly overfitting model. Therefore, it is better to use kfold. Good luck!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2135114,
              "author_name": "Paul Bacher",
              "author_url": "",
              "post_date": "2023-02-08T13:10:24.383000",
              "content": "<p>Great to hear that! My model also overfits a lot, but I haven't found a solution that doesn't harm my score yet. I'll try to improve my data aug.</p>\n<p>K-Fold cross-validation would definitely be better. Have you tried it? Does it consume significantly more computational time? In my case, I want to conserve my TPU credits as much as possible to try different ideas. Currently, a single training takes me between 1 to 1h30.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2135141,
              "author_name": "Andrij",
              "author_url": "",
              "post_date": "2023-02-08T13:30:36.723000",
              "content": "<p>For testing, I prepared a dataset with less expansion, then training is faster (models are less accurate, but more models can be tried). We have too small data set for positive class, so we can't trust one model, read other discussions, I see somewhere that kfold improved the result by lb. But taking into account all the above, I don't trust public lb in this competition (more precisely, I trust less than usual).</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2134251": "The metric is indeed quite shaky, but your validation score \"looks to good to be true\". Especially, compared to the values reported by other competitors. Most of the time, when this happens, there is a leak in your validation. Did you not split by patient ID maybe? Is this a super small subset?",
    "2137542": "Thanks sharing.\n\nDo  you use some external data?\nOr  pretrain use train data? Then you should take care about  making CV folds.\nAnyway I feel there can be a leak.",
    "2134684": "I just started and made a single submission.\n\n```python\nOOF ROC AUC 0.8502\nOOF PF1: 0.3171\nLB PF1: 0.39\n```",
    "2134534": "How about your ROC_AUC ??",
    "2134174": "For me single model with val scores pF1: 0.41, F1: 0.57(t=0.5), ROCAUC: 0.84 gives me LB 0.51(t=0.5)\n\nyou can find more about results here\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/380657\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333",
    "2134840": "0.77 seems too high to be accurate. In my case, the situation is different. I have a validation accuracy of 20% with F1 scores usually between 0.3 and 0.4, which gave me a best score of 0.54. Also, I do not use cross-validation.",
    "2134151": "Hello everybody! \nI'm only a few days into the competition, but a quick review showed its \"difficulty\". \nPersonally, I predict a shake-up, but I could be wrong, we'll see. I built an approach that allowed me to get pF1=0.77 on the validation data using one model, on lb I got 0.32, I'm a bit upset about that, but it just says that my approach is wrong, or it didn't work well on the lb part of the data. After looking through the public notebooks and comments, I realized that the gap between val_test and lb is huge. \nSo, who has the biggest gap, and what do you think about it in general. In my opinion, our problem lies precisely in the imbalance of classes. I met with a similar problem for the NLP problem and the bag of words model, then the best solution was to consider only the features (ngrams) that indicate a positive class, such a trick will not work here.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F7ec945d2b2918d7d2ce7ce31ba377417%2F1.PNG?generation=1675799168302642&alt=media)"
  }
}