{
  "id": 434783,
  "title": "CV is inversely proportional to LB",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/434783",
  "author_name": "Harshanand",
  "post_date": "2023-08-26T14:05:20.053000",
  "votes": 1,
  "comment_count": 12,
  "views": 0,
  "content": "<p>CV    -    LB<br>\n0.87  -   3.33<br>\n1.45  -   2.53<br>\n2.50  -   1.99 <br>\n3.50  -   1.56</p>\n<p>I am following a 2.5D approach by sampling 4 slices. For different iterations of training I've tried to tamper with label smoothing, weight decay (l2 penalty) and learning rate. </p>\n<p>From the results I can obviously understand my model is overfitting to the training distribution. Does anyone have better insights about this.</p>",
  "messages": [
    {
      "id": 2435438,
      "postDate": "2023-09-13T01:45:50.427Z",
      "content": "<p>in the worst case, your model should predict \"mean probabiliy value\".<br>\nThis has shown to be around 0.68.</p>\n<p>Hence anything above that is \"meaningless to be discussed\"  in CV or LB.</p>\n<p>if you predict above 0.68 in training, your model is not learning (likely a pipline bug)<br>\nif it is above 0.68 in validation (and not training),  it is generalisation error</p>",
      "rawMarkdown": "in the worst case, your model should predict \"mean probabiliy value\".\nThis has shown to be around 0.68.\n\nHence anything above that is \"meaningless to be discussed\"  in CV or LB.\n\nif you predict above 0.68 in training, your model is not learning (likely a pipline bug)\nif it is above 0.68 in validation (and not training),  it is generalisation error",
      "votes": 3,
      "replies": [
        {
          "id": 2436583,
          "postDate": "2023-09-13T16:44:14.813Z",
          "content": "<p>This is a good tip. You should always start with the baseline and try to make improvements over it.</p>",
          "rawMarkdown": "This is a good tip. You should always start with the baseline and try to make improvements over it."
        },
        {
          "id": 2438264,
          "postDate": "2023-09-14T07:05:20.630Z",
          "content": "<p>Looked at this awhile ago and the notebook is using this reduced dataset with its own issues, including most if not all of the image level labels dcms are not included in it.  <br>\n<a href=\"https://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm\" target=\"_blank\">https://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm</a></p>\n<p>This competition has many difficulties in approaching not the least are all the constant prediction weighted means notebooks, and discussions like \"setting the prediction value for a specific label 'extravasation_injury' to a constant 0.99 yielded outstanding results \"</p>\n<p>In some respects they have prevented meaningful notebooks on data prep, pipelines and modelling making any headway. </p>",
          "rawMarkdown": "Looked at this awhile ago and the notebook is using this reduced dataset with its own issues, including most if not all of the image level labels dcms are not included in it.  \nhttps://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm\n\nThis competition has many difficulties in approaching not the least are all the constant prediction weighted means notebooks, and discussions like \"setting the prediction value for a specific label 'extravasation_injury' to a constant 0.99 yielded outstanding results \"\n\nIn some respects they have prevented meaningful notebooks on data prep, pipelines and modelling making any headway. \n",
          "votes": 4,
          "replies": [
            {
              "id": 2438638,
              "postDate": "2023-09-14T12:02:53.090Z",
              "content": "<p>I agree this competition has a barrier to entry but it's not that high. Someone who has solid fundamentals shouldn't have much problem beating the mean baseline. The main problem is most people depend so much on starter notebooks. When I first started Kaggle 5 years ago, I was always building my pipeline from scratch and it started to pay off immediately. I've never used any public notebook as a starter since then. I suggest you to do the same. You might lose so much time at first but it will pay off eventually.</p>",
              "rawMarkdown": "I agree this competition has a barrier to entry but it's not that high. Someone who has solid fundamentals shouldn't have much problem beating the mean baseline. The main problem is most people depend so much on starter notebooks. When I first started Kaggle 5 years ago, I was always building my pipeline from scratch and it started to pay off immediately. I've never used any public notebook as a starter since then. I suggest you to do the same. You might lose so much time at first but it will pay off eventually.",
              "votes": 3
            },
            {
              "id": 2438669,
              "postDate": "2023-09-14T12:19:03.373Z",
              "content": "<p>I completely agree with you.<br>\nI never used any public notebook - for me, it's part of the competition to build it myself. This costs more time and can lead to some bugs (as seen before) but it's actually fun.</p>",
              "rawMarkdown": "I completely agree with you.\nI never used any public notebook - for me, it's part of the competition to build it myself. This costs more time and can lead to some bugs (as seen before) but it's actually fun.",
              "votes": 1
            },
            {
              "id": 2439955,
              "postDate": "2023-09-15T07:25:06.103Z",
              "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> - am working on my own pipeline and ideas, hopefully will get things working in time.  Just pointing out the issues with some of the reduced datasets and lack of many model notebooks for those trying to look at ideas on how to approach, what is possible on Kaggle resources, how to overcome submission issues, etc.,  e.g., someone asked awhile back about Pytorch XLA and have been able to get that working after much research and attempts with the environment here. Being able to leverage off another model notebook would be helpful to make that available to others.   Many competitions get bogged down in one relatively high scoring notebook or ensemble of ensembles that tend to discourage innovation or participation even in discussions.  Perhaps more is happening on Discord.  </p>\n<blockquote>\n  <p>Hence anything above that is \"meaningless to be discussed\" in CV or LB.</p>\n</blockquote>\n<p>thought that comment a bit harsh, given the poster actually made the effort to try with a public notebook and in discussions.  Get that RSNA competitions are difficult and possibly out of reach for many here.  But it helps when others can help people to try. </p>",
              "rawMarkdown": "@gunesevitan - am working on my own pipeline and ideas, hopefully will get things working in time.  Just pointing out the issues with some of the reduced datasets and lack of many model notebooks for those trying to look at ideas on how to approach, what is possible on Kaggle resources, how to overcome submission issues, etc.,  e.g., someone asked awhile back about Pytorch XLA and have been able to get that working after much research and attempts with the environment here. Being able to leverage off another model notebook would be helpful to make that available to others.   Many competitions get bogged down in one relatively high scoring notebook or ensemble of ensembles that tend to discourage innovation or participation even in discussions.  Perhaps more is happening on Discord.  \n\n>Hence anything above that is \"meaningless to be discussed\" in CV or LB.\n\nthought that comment a bit harsh, given the poster actually made the effort to try with a public notebook and in discussions.  Get that RSNA competitions are difficult and possibly out of reach for many here.  But it helps when others can help people to try. \n",
              "votes": 1
            },
            {
              "id": 2439970,
              "postDate": "2023-09-15T07:30:48.013Z",
              "content": "<p>\"Hence anything above that is \"meaningless to be discussed\" in CV or LB.\"</p>\n<p>i am sorry if i cause misunderstanding.<br>\ni did NOT mean literally disscusion in the forum.</p>\n<p>maybe a better wording is: </p>\n<pre><code>meaningless   considered  analysed\nbecuase this  more likely   pipline bug.\nhence spend time  fault- bug  maybe better than  spend time  analyse/explain the results.\n</code></pre>",
              "rawMarkdown": "\"Hence anything above that is \"meaningless to be discussed\" in CV or LB.\"\n\ni am sorry if i cause misunderstanding.\ni did NOT mean literally disscusion in the forum.\n\nmaybe a better wording is: \n\n```\n\"Hence anything above that (score of 0.68 in CV or LB)  is \"meaningless to be considered or analysed\".\nbecuase this is more likely to be pipline bug.\nhence spend time to fault-find bug is maybe better than to spend time to analyse/explain the results.\n\n```\n",
              "votes": 1
            },
            {
              "id": 2440673,
              "postDate": "2023-09-15T16:52:12.387Z",
              "content": "<p>This is an interesting discussion. Although I would disagree that anything above the baseline is \"meaningless to be considered or analyzed\" or \"meaningless to be discussed\". Imagine a case where your model is perfectly dissociating your output but in a complete flipped manner. Error is going to be huge but the model is totally worth an investigation, debugging, etc.</p>",
              "rawMarkdown": "This is an interesting discussion. Although I would disagree that anything above the baseline is \"meaningless to be considered or analyzed\" or \"meaningless to be discussed\". Imagine a case where your model is perfectly dissociating your output but in a complete flipped manner. Error is going to be huge but the model is totally worth an investigation, debugging, etc.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2435456,
      "postDate": "2023-09-13T02:10:39.037Z",
      "content": "<p>When embarking on a difficult task, it is advisable to monitor the AUC at the beginning. I have a feeling that the AUC for your model is around 0.5. And it would also be good to see if the loss of train continues to drop steadily.</p>",
      "rawMarkdown": "When embarking on a difficult task, it is advisable to monitor the AUC at the beginning. I have a feeling that the AUC for your model is around 0.5. And it would also be good to see if the loss of train continues to drop steadily.",
      "votes": 4
    },
    {
      "id": 2409910,
      "postDate": "2023-08-26T14:05:20.053Z",
      "content": "<p>CV    -    LB<br>\n0.87  -   3.33<br>\n1.45  -   2.53<br>\n2.50  -   1.99 <br>\n3.50  -   1.56</p>\n<p>I am following a 2.5D approach by sampling 4 slices. For different iterations of training I've tried to tamper with label smoothing, weight decay (l2 penalty) and learning rate. </p>\n<p>From the results I can obviously understand my model is overfitting to the training distribution. Does anyone have better insights about this.</p>",
      "rawMarkdown": "CV    -    LB\n0.87  -   3.33\n1.45  -   2.53\n2.50  -   1.99 \n3.50  -   1.56\n\nI am following a 2.5D approach by sampling 4 slices. For different iterations of training I've tried to tamper with label smoothing, weight decay (l2 penalty) and learning rate. \n\nFrom the results I can obviously understand my model is overfitting to the training distribution. Does anyone have better insights about this.",
      "votes": 1
    },
    {
      "id": 2436278,
      "postDate": "2023-09-13T13:23:14.507Z",
      "content": "<p>Thank you both!<br>\nI tested my whole data generation process and there was indeed a bug.<br>\nMaybe the explanation can help someone who's facing similar issues:<br>\nIf you use Pandas with its function to_numpy() it is important to mention that the return value is a shallow copy! I used it to shuffle the patient IDs so the column itself was shuffled and the targets didn't make any sense.</p>\n<pre><code>\nPatients = trainListShuffle[].to_numpy(copy=True)\nsplit = int(Patients.shape[]*TRAIN_TEST_SPLIT)\nnp..shuffle(Patients)\n</code></pre>",
      "rawMarkdown": "Thank you both!\nI tested my whole data generation process and there was indeed a bug.\nMaybe the explanation can help someone who's facing similar issues:\nIf you use Pandas with its function to_numpy() it is important to mention that the return value is a shallow copy! I used it to shuffle the patient IDs so the column itself was shuffled and the targets didn't make any sense.\n```\n# Without the copy=True a reference will be returned\nallPatients = trainListShuffle[\"patient_id\"].to_numpy(copy=True)\nsplit = int(allPatients.shape[0]*TRAIN_TEST_SPLIT)\nnp.random.shuffle(allPatients)\n```",
      "votes": 2
    },
    {
      "id": 2435297,
      "postDate": "2023-09-12T20:58:19.670Z",
      "content": "<p>Same for me here.<br>\nI actually tried the same model after epochs 1, 2 and 3. The validation cross entropy was decreasing every epoch but in the leaderboard I got <br>\nEpoch 1 : 0.85<br>\nEpoch 2: 0.89<br>\nEpoch 3: 1.05</p>\n<p>Is there any abnormality in the test dataset we don't know about? Maybe the voxel values are way higher or lower? Or are there more longitudinal than transversal slices?</p>",
      "rawMarkdown": "Same for me here.\nI actually tried the same model after epochs 1, 2 and 3. The validation cross entropy was decreasing every epoch but in the leaderboard I got \nEpoch 1 : 0.85\nEpoch 2: 0.89\nEpoch 3: 1.05\n\nIs there any abnormality in the test dataset we don't know about? Maybe the voxel values are way higher or lower? Or are there more longitudinal than transversal slices?"
    },
    {
      "id": 2410523,
      "postDate": "2023-08-27T03:46:12.223Z",
      "content": "<p>I have the same issue too. I get an ok score on my validation data but when I submit the same code I get crazy high log loss score..</p>",
      "rawMarkdown": "I have the same issue too. I get an ok score on my validation data but when I submit the same code I get crazy high log loss score.."
    }
  ],
  "comments": [
    {
      "id": 2435438,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-09-13T01:45:50.427000",
      "content": "<p>in the worst case, your model should predict \"mean probabiliy value\".<br>\nThis has shown to be around 0.68.</p>\n<p>Hence anything above that is \"meaningless to be discussed\"  in CV or LB.</p>\n<p>if you predict above 0.68 in training, your model is not learning (likely a pipline bug)<br>\nif it is above 0.68 in validation (and not training),  it is generalisation error</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2436583,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-09-13T16:44:14.813000",
          "content": "<p>This is a good tip. You should always start with the baseline and try to make improvements over it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2438264,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2023-09-14T07:05:20.630000",
          "content": "<p>Looked at this awhile ago and the notebook is using this reduced dataset with its own issues, including most if not all of the image level labels dcms are not included in it.  <br>\n<a href=\"https://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm\" target=\"_blank\">https://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm</a></p>\n<p>This competition has many difficulties in approaching not the least are all the constant prediction weighted means notebooks, and discussions like \"setting the prediction value for a specific label 'extravasation_injury' to a constant 0.99 yielded outstanding results \"</p>\n<p>In some respects they have prevented meaningful notebooks on data prep, pipelines and modelling making any headway. </p>",
          "votes": 4,
          "replies": [
            {
              "id": 2438638,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2023-09-14T12:02:53.090000",
              "content": "<p>I agree this competition has a barrier to entry but it's not that high. Someone who has solid fundamentals shouldn't have much problem beating the mean baseline. The main problem is most people depend so much on starter notebooks. When I first started Kaggle 5 years ago, I was always building my pipeline from scratch and it started to pay off immediately. I've never used any public notebook as a starter since then. I suggest you to do the same. You might lose so much time at first but it will pay off eventually.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2438669,
              "author_name": "Manuel K",
              "author_url": "",
              "post_date": "2023-09-14T12:19:03.373000",
              "content": "<p>I completely agree with you.<br>\nI never used any public notebook - for me, it's part of the competition to build it myself. This costs more time and can lead to some bugs (as seen before) but it's actually fun.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2439955,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2023-09-15T07:25:06.103000",
              "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> - am working on my own pipeline and ideas, hopefully will get things working in time.  Just pointing out the issues with some of the reduced datasets and lack of many model notebooks for those trying to look at ideas on how to approach, what is possible on Kaggle resources, how to overcome submission issues, etc.,  e.g., someone asked awhile back about Pytorch XLA and have been able to get that working after much research and attempts with the environment here. Being able to leverage off another model notebook would be helpful to make that available to others.   Many competitions get bogged down in one relatively high scoring notebook or ensemble of ensembles that tend to discourage innovation or participation even in discussions.  Perhaps more is happening on Discord.  </p>\n<blockquote>\n  <p>Hence anything above that is \"meaningless to be discussed\" in CV or LB.</p>\n</blockquote>\n<p>thought that comment a bit harsh, given the poster actually made the effort to try with a public notebook and in discussions.  Get that RSNA competitions are difficult and possibly out of reach for many here.  But it helps when others can help people to try. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2439970,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-09-15T07:30:48.013000",
              "content": "<p>\"Hence anything above that is \"meaningless to be discussed\" in CV or LB.\"</p>\n<p>i am sorry if i cause misunderstanding.<br>\ni did NOT mean literally disscusion in the forum.</p>\n<p>maybe a better wording is: </p>\n<pre><code>meaningless   considered  analysed\nbecuase this  more likely   pipline bug.\nhence spend time  fault- bug  maybe better than  spend time  analyse/explain the results.\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2440673,
              "author_name": "Parham Mostame",
              "author_url": "",
              "post_date": "2023-09-15T16:52:12.387000",
              "content": "<p>This is an interesting discussion. Although I would disagree that anything above the baseline is \"meaningless to be considered or analyzed\" or \"meaningless to be discussed\". Imagine a case where your model is perfectly dissociating your output but in a complete flipped manner. Error is going to be huge but the model is totally worth an investigation, debugging, etc.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2435456,
      "author_name": "YujiAriyasu",
      "author_url": "",
      "post_date": "2023-09-13T02:10:39.037000",
      "content": "<p>When embarking on a difficult task, it is advisable to monitor the AUC at the beginning. I have a feeling that the AUC for your model is around 0.5. And it would also be good to see if the loss of train continues to drop steadily.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2436278,
      "author_name": "Manuel K",
      "author_url": "",
      "post_date": "2023-09-13T13:23:14.507000",
      "content": "<p>Thank you both!<br>\nI tested my whole data generation process and there was indeed a bug.<br>\nMaybe the explanation can help someone who's facing similar issues:<br>\nIf you use Pandas with its function to_numpy() it is important to mention that the return value is a shallow copy! I used it to shuffle the patient IDs so the column itself was shuffled and the targets didn't make any sense.</p>\n<pre><code>\nPatients = trainListShuffle[].to_numpy(copy=True)\nsplit = int(Patients.shape[]*TRAIN_TEST_SPLIT)\nnp..shuffle(Patients)\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2435297,
      "author_name": "Manuel K",
      "author_url": "",
      "post_date": "2023-09-12T20:58:19.670000",
      "content": "<p>Same for me here.<br>\nI actually tried the same model after epochs 1, 2 and 3. The validation cross entropy was decreasing every epoch but in the leaderboard I got <br>\nEpoch 1 : 0.85<br>\nEpoch 2: 0.89<br>\nEpoch 3: 1.05</p>\n<p>Is there any abnormality in the test dataset we don't know about? Maybe the voxel values are way higher or lower? Or are there more longitudinal than transversal slices?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2410523,
      "author_name": "Parham Mostame",
      "author_url": "",
      "post_date": "2023-08-27T03:46:12.223000",
      "content": "<p>I have the same issue too. I get an ok score on my validation data but when I submit the same code I get crazy high log loss score..</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2435438": "in the worst case, your model should predict \"mean probabiliy value\".\nThis has shown to be around 0.68.\n\nHence anything above that is \"meaningless to be discussed\"  in CV or LB.\n\nif you predict above 0.68 in training, your model is not learning (likely a pipline bug)\nif it is above 0.68 in validation (and not training),  it is generalisation error",
    "2435456": "When embarking on a difficult task, it is advisable to monitor the AUC at the beginning. I have a feeling that the AUC for your model is around 0.5. And it would also be good to see if the loss of train continues to drop steadily.",
    "2409910": "CV    -    LB\n0.87  -   3.33\n1.45  -   2.53\n2.50  -   1.99 \n3.50  -   1.56\n\nI am following a 2.5D approach by sampling 4 slices. For different iterations of training I've tried to tamper with label smoothing, weight decay (l2 penalty) and learning rate. \n\nFrom the results I can obviously understand my model is overfitting to the training distribution. Does anyone have better insights about this.",
    "2436278": "Thank you both!\nI tested my whole data generation process and there was indeed a bug.\nMaybe the explanation can help someone who's facing similar issues:\nIf you use Pandas with its function to_numpy() it is important to mention that the return value is a shallow copy! I used it to shuffle the patient IDs so the column itself was shuffled and the targets didn't make any sense.\n```\n# Without the copy=True a reference will be returned\nallPatients = trainListShuffle[\"patient_id\"].to_numpy(copy=True)\nsplit = int(allPatients.shape[0]*TRAIN_TEST_SPLIT)\nnp.random.shuffle(allPatients)\n```",
    "2435297": "Same for me here.\nI actually tried the same model after epochs 1, 2 and 3. The validation cross entropy was decreasing every epoch but in the leaderboard I got \nEpoch 1 : 0.85\nEpoch 2: 0.89\nEpoch 3: 1.05\n\nIs there any abnormality in the test dataset we don't know about? Maybe the voxel values are way higher or lower? Or are there more longitudinal than transversal slices?",
    "2410523": "I have the same issue too. I get an ok score on my validation data but when I submit the same code I get crazy high log loss score.."
  }
}