{
  "id": 434890,
  "title": "Submission score",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/434890",
  "author_name": "Parham Mostame",
  "post_date": "2023-08-27T03:44:39.064000",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hey folks!</p>\n<p>When I run my notebook on train images (the validation set), I get a somewhat ok accuracy. But when I submit my notebook my score is really bad on the test images even though it's the same exact code. This is my first competition and I'm wondering if there is any difference between train and test images in nature? Or has anyone encountered this problem before?</p>",
  "messages": [
    {
      "id": 2410522,
      "postDate": "2023-08-27T03:44:39.063Z",
      "content": "<p>Hey folks!</p>\n<p>When I run my notebook on train images (the validation set), I get a somewhat ok accuracy. But when I submit my notebook my score is really bad on the test images even though it's the same exact code. This is my first competition and I'm wondering if there is any difference between train and test images in nature? Or has anyone encountered this problem before?</p>",
      "rawMarkdown": "Hey folks!\n\nWhen I run my notebook on train images (the validation set), I get a somewhat ok accuracy. But when I submit my notebook my score is really bad on the test images even though it's the same exact code. This is my first competition and I'm wondering if there is any difference between train and test images in nature? Or has anyone encountered this problem before?",
      "votes": 3
    },
    {
      "id": 2411203,
      "postDate": "2023-08-27T13:16:46.870Z",
      "content": "<p>Hello, I have the same issue.</p>\n<p>Given the \"weighted mean average\" notebooks give almost a top 10 score, I wonder if there isn't some kind of line order problem with the hidden sample_submission.csv, I mean compared to the true solution. At least that would explain the disrepancies… Moreover the kaggle published script for scoring doesn't care about patient_id at all, just line order.</p>\n<p>I think maybe an admin should check this…</p>",
      "rawMarkdown": "Hello, I have the same issue.\n\nGiven the \"weighted mean average\" notebooks give almost a top 10 score, I wonder if there isn't some kind of line order problem with the hidden sample_submission.csv, I mean compared to the true solution. At least that would explain the disrepancies... Moreover the kaggle published script for scoring doesn't care about patient_id at all, just line order.\n\nI think maybe an admin should check this...",
      "votes": 2,
      "replies": [
        {
          "id": 2411299,
          "postDate": "2023-08-27T14:30:07.917Z",
          "content": "<p>This simple code, when scored, throws an exception:</p>\n<p><code>submission = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')</code><br>\n<code>sub_sorted = submission.sort_values('patient_id')</code><br>\n<code>assert(sub_sorted['patient_id'].tolist() == submission['patient_id'].tolist())</code></p>\n<p>This means the lines aren't sorted in  numerical order, like it seems usual…<br>\nIf you see me in the top 10 soon, this would prove my prior assumption! 😅</p>\n<p>Edit:</p>\n<p>I did a few more tests, I'm now sure the hidden sample_submission.csv is sorted in alphabetical order on patient_id.<br>\nBUT when I tried submitting a proper model with numerical sort on patient_id, my score didn't change at all, <em>probably</em> meaning this isn't either the right order for lines, and that no sorting does garble the score.</p>\n<p>Long story short, I think there's a bug with the scoring on this kaggle.</p>",
          "rawMarkdown": "This simple code, when scored, throws an exception:\n\n`submission = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')`\n`sub_sorted = submission.sort_values('patient_id')`\n`assert(sub_sorted['patient_id'].tolist() == submission['patient_id'].tolist())`\n\nThis means the lines aren't sorted in ~~alphabetical~~ numerical order, like it seems usual...\nIf you see me in the top 10 soon, this would prove my prior assumption! 😅\n\nEdit:\n\nI did a few more tests, I'm now sure the hidden sample_submission.csv is sorted in alphabetical order on patient_id.\nBUT when I tried submitting a proper model with numerical sort on patient_id, my score didn't change at all, *probably* meaning this isn't either the right order for lines, and that no sorting does garble the score.\n\nLong story short, I think there's a bug with the scoring on this kaggle.",
          "votes": 2,
          "replies": [
            {
              "id": 2411804,
              "postDate": "2023-08-27T21:40:50.340Z",
              "content": "<p>Can you try this on your model: estimate log loss from sklearn on your predictions. then do the same while use .flatten() on the data. I do this on a perfect prediction and get ~0 for .flatten version but around 10 for the non flattened version.. so weird! </p>",
              "rawMarkdown": "Can you try this on your model: estimate log loss from sklearn on your predictions. then do the same while use .flatten() on the data. I do this on a perfect prediction and get ~0 for .flatten version but around 10 for the non flattened version.. so weird! "
            }
          ]
        },
        {
          "id": 2411800,
          "postDate": "2023-08-27T21:38:18.117Z",
          "content": "<p>Exactly.. I get a good score on the validation set too but also get like 15 on the actual submission. If it was overfitting it wouldn't have been 15 or so.. there is some systematic difference between the datasets.. maybe as you say the order of the patients or so!<br>\nLet me know how it goes and I will update you as well!</p>",
          "rawMarkdown": "Exactly.. I get a good score on the validation set too but also get like 15 on the actual submission. If it was overfitting it wouldn't have been 15 or so.. there is some systematic difference between the datasets.. maybe as you say the order of the patients or so!\nLet me know how it goes and I will update you as well!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2411203,
      "author_name": "GliGli",
      "author_url": "",
      "post_date": "2023-08-27T13:16:46.870000",
      "content": "<p>Hello, I have the same issue.</p>\n<p>Given the \"weighted mean average\" notebooks give almost a top 10 score, I wonder if there isn't some kind of line order problem with the hidden sample_submission.csv, I mean compared to the true solution. At least that would explain the disrepancies… Moreover the kaggle published script for scoring doesn't care about patient_id at all, just line order.</p>\n<p>I think maybe an admin should check this…</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2411299,
          "author_name": "GliGli",
          "author_url": "",
          "post_date": "2023-08-27T14:30:07.917000",
          "content": "<p>This simple code, when scored, throws an exception:</p>\n<p><code>submission = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')</code><br>\n<code>sub_sorted = submission.sort_values('patient_id')</code><br>\n<code>assert(sub_sorted['patient_id'].tolist() == submission['patient_id'].tolist())</code></p>\n<p>This means the lines aren't sorted in  numerical order, like it seems usual…<br>\nIf you see me in the top 10 soon, this would prove my prior assumption! 😅</p>\n<p>Edit:</p>\n<p>I did a few more tests, I'm now sure the hidden sample_submission.csv is sorted in alphabetical order on patient_id.<br>\nBUT when I tried submitting a proper model with numerical sort on patient_id, my score didn't change at all, <em>probably</em> meaning this isn't either the right order for lines, and that no sorting does garble the score.</p>\n<p>Long story short, I think there's a bug with the scoring on this kaggle.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2411804,
              "author_name": "Parham Mostame",
              "author_url": "",
              "post_date": "2023-08-27T21:40:50.340000",
              "content": "<p>Can you try this on your model: estimate log loss from sklearn on your predictions. then do the same while use .flatten() on the data. I do this on a perfect prediction and get ~0 for .flatten version but around 10 for the non flattened version.. so weird! </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2411800,
          "author_name": "Parham Mostame",
          "author_url": "",
          "post_date": "2023-08-27T21:38:18.117000",
          "content": "<p>Exactly.. I get a good score on the validation set too but also get like 15 on the actual submission. If it was overfitting it wouldn't have been 15 or so.. there is some systematic difference between the datasets.. maybe as you say the order of the patients or so!<br>\nLet me know how it goes and I will update you as well!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2410522": "Hey folks!\n\nWhen I run my notebook on train images (the validation set), I get a somewhat ok accuracy. But when I submit my notebook my score is really bad on the test images even though it's the same exact code. This is my first competition and I'm wondering if there is any difference between train and test images in nature? Or has anyone encountered this problem before?",
    "2411203": "Hello, I have the same issue.\n\nGiven the \"weighted mean average\" notebooks give almost a top 10 score, I wonder if there isn't some kind of line order problem with the hidden sample_submission.csv, I mean compared to the true solution. At least that would explain the disrepancies... Moreover the kaggle published script for scoring doesn't care about patient_id at all, just line order.\n\nI think maybe an admin should check this..."
  }
}