{
  "id": 437814,
  "title": "Post-processing problem, why does the bug fix improve CV from 1.00 to 0.70 but degrade lb from 0.69 to 0.97?",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/437814",
  "author_name": "Andrij",
  "post_date": "2023-09-08T08:25:13.649000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello everybody!<br>\nAs a basis, I use this notebook branch (thanks to the authors): <a href=\"https://www.kaggle.com/code/coderrkj/rsna-atd-cnn-tpu-infer-corrections\" target=\"_blank\">https://www.kaggle.com/code/coderrkj/rsna-atd-cnn-tpu-infer-corrections</a>.<br>\nThis notebook has the following post-processing:</p>\n<pre><code> post_proc_v2(pred):\n     = np.empty((pred.shape[], * + *), dtype='float32')\n\n    \n    [:, ] = pred[:, ]\n    [:, ] =  - proc_pred[:, ]\n    [:, ] = pred[:, ]\n    [:, ] =  - proc_pred[:, ]\n\n    \n    [:, :] = pred[:, :]\n    [:, :] = pred[:, :]\n    [:, :] = pred[:, :]\n\n     proc_pred\n</code></pre>\n<p>​As far as I understand it has a bug (proc_pred[:, 0] and proc_pred[:, 1]; proc_pred[:, 2] and proc_pred[:, 3] are selected incorrectly).</p>\n<p>This is what I think this post-processing should look like:</p>\n<pre><code> post_proc_v2(pred):\n     = np.empty((pred.shape[], * + *), dtype='float32')\n\n    \n    [:, ] =  - pred[:, ]\n    [:, ] = pred[:, ]\n    [:, ] =  - pred[:, ]\n    [:, ] = pred[:, ]\n\n    \n    [:, :] = pred[:, :]\n    [:, :] = pred[:, :]\n    [:, :] = pred[:, :]\n\n     proc_pred\n</code></pre>\n<p>Simply put, we assign the model prediction for \"bowel_injury\" &amp; \"extravasation_injury\" to \"bowel_healthy\" &amp; \"extravasation_healthy\". <br>\nThis gives us 0.69 lb and 1.00 CV. If we assign correctly, we estimate 0.97 lb and 0.7 CV.<br>\nThe authors wrote about this in the comments.</p>\n<p>Does anyone know what the problem is?</p>",
  "messages": [
    {
      "id": 2430554,
      "postDate": "2023-09-09T11:07:23.800Z",
      "content": "<p>actually don't mind that much</p>\n<ol>\n<li>first just need to check there is no bug in your code.</li>\n<li>then ensure that results is producible.</li>\n<li>If there is a trend, capture it (if i imporve xxx, yyy is +ve or -ve correlated)</li>\n<li>the rest are just data … everything can be explained by data distribution of the train/test, etc …</li>\n</ol>\n<p>when the sample size is tool small, anything can happens, so don't mind that much.</p>\n<p>(espeically here, the high grade sample has low number and test case and high weighing).</p>\n<hr>\n<p>as an example, i find that for bowel submitting mean values is better than  a trained model prediction</p>",
      "rawMarkdown": "actually don't mind that much\n\n1. first just need to check there is no bug in your code.\n2. then ensure that results is producible.\n3. If there is a trend, capture it (if i imporve xxx, yyy is +ve or -ve correlated)\n4. the rest are just data ... everything can be explained by data distribution of the train/test, etc ...\n\nwhen the sample size is tool small, anything can happens, so don't mind that much.\n\n(espeically here, the high grade sample has low number and test case and high weighing).\n\n---\n\nas an example, i find that for bowel submitting mean values is better than  a trained model prediction",
      "votes": 1,
      "replies": [
        {
          "id": 2430615,
          "postDate": "2023-09-09T12:11:26.430Z",
          "content": "<p>Thanks, I think I found the problem. <br>\nThis is a problem with my training data. I think there will be a shake-up in this competition. <br>\nTo be honest, when I figured out what it was, I was shocked that I had such a good result. <br>\nThis is where I could be wrong, but it seems to me that my model misinterpreted some of the images, which is why I got this wrong result. <br>\nThanks for the important information about bowel. <br>\nBut now I feel that the result can be seriously improved. <br>\nBut where is the limit is another question… <br>\nThe first result of 0.42 may already be somewhere close, taking into account the specifics of the data and possible inaccuracies in the data themselves</p>",
          "rawMarkdown": "Thanks, I think I found the problem. \nThis is a problem with my training data. I think there will be a shake-up in this competition. \nTo be honest, when I figured out what it was, I was shocked that I had such a good result. \nThis is where I could be wrong, but it seems to me that my model misinterpreted some of the images, which is why I got this wrong result. \nThanks for the important information about bowel. \nBut now I feel that the result can be seriously improved. \nBut where is the limit is another question... \nThe first result of 0.42 may already be somewhere close, taking into account the specifics of the data and possible inaccuracies in the data themselves",
          "replies": [
            {
              "id": 2430998,
              "postDate": "2023-09-09T17:56:32.340Z",
              "content": "<p>Yup, I checked after making changes to the data on the validation and public sets I got the same 0.83.</p>",
              "rawMarkdown": "Yup, I checked after making changes to the data on the validation and public sets I got the same 0.83."
            }
          ]
        }
      ]
    },
    {
      "id": 2429137,
      "postDate": "2023-09-08T11:58:00.117Z",
      "content": "<p>I wouldn't trust public notebooks on a competition like this.</p>",
      "rawMarkdown": "I wouldn't trust public notebooks on a competition like this.",
      "votes": 1,
      "replies": [
        {
          "id": 2429194,
          "postDate": "2023-09-08T12:43:10.863Z",
          "content": "<p>I'm more interested in where the problem is: <br>\nis the problem in my code or in the test data. If the problem is in my code, that's bad. <br>\nMaybe the participants fixed some error that I didn't see. <br>\nIn addition, incorrect data processing can produce such results. <br>\nIf the problem is precisely in the public LB, it means that the shake-up will be unpleasant (According to my personal classification: minimal, normal, unpleasant, large, epic)</p>",
          "rawMarkdown": "I'm more interested in where the problem is: \nis the problem in my code or in the test data. If the problem is in my code, that's bad. \nMaybe the participants fixed some error that I didn't see. \nIn addition, incorrect data processing can produce such results. \nIf the problem is precisely in the public LB, it means that the shake-up will be unpleasant (According to my personal classification: minimal, normal, unpleasant, large, epic)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2428892,
      "postDate": "2023-09-08T08:25:13.650Z",
      "content": "<p>Hello everybody!<br>\nAs a basis, I use this notebook branch (thanks to the authors): <a href=\"https://www.kaggle.com/code/coderrkj/rsna-atd-cnn-tpu-infer-corrections\" target=\"_blank\">https://www.kaggle.com/code/coderrkj/rsna-atd-cnn-tpu-infer-corrections</a>.<br>\nThis notebook has the following post-processing:</p>\n<pre><code> post_proc_v2(pred):\n     = np.empty((pred.shape[], * + *), dtype='float32')\n\n    \n    [:, ] = pred[:, ]\n    [:, ] =  - proc_pred[:, ]\n    [:, ] = pred[:, ]\n    [:, ] =  - proc_pred[:, ]\n\n    \n    [:, :] = pred[:, :]\n    [:, :] = pred[:, :]\n    [:, :] = pred[:, :]\n\n     proc_pred\n</code></pre>\n<p>​As far as I understand it has a bug (proc_pred[:, 0] and proc_pred[:, 1]; proc_pred[:, 2] and proc_pred[:, 3] are selected incorrectly).</p>\n<p>This is what I think this post-processing should look like:</p>\n<pre><code> post_proc_v2(pred):\n     = np.empty((pred.shape[], * + *), dtype='float32')\n\n    \n    [:, ] =  - pred[:, ]\n    [:, ] = pred[:, ]\n    [:, ] =  - pred[:, ]\n    [:, ] = pred[:, ]\n\n    \n    [:, :] = pred[:, :]\n    [:, :] = pred[:, :]\n    [:, :] = pred[:, :]\n\n     proc_pred\n</code></pre>\n<p>Simply put, we assign the model prediction for \"bowel_injury\" &amp; \"extravasation_injury\" to \"bowel_healthy\" &amp; \"extravasation_healthy\". <br>\nThis gives us 0.69 lb and 1.00 CV. If we assign correctly, we estimate 0.97 lb and 0.7 CV.<br>\nThe authors wrote about this in the comments.</p>\n<p>Does anyone know what the problem is?</p>",
      "rawMarkdown": "Hello everybody!\nAs a basis, I use this notebook branch (thanks to the authors): https://www.kaggle.com/code/coderrkj/rsna-atd-cnn-tpu-infer-corrections.\nThis notebook has the following post-processing:\n\n    def post_proc_v2(pred):\n        proc_pred = np.empty((pred.shape[0], 2*2 + 3*3), dtype='float32')\n\n        # bowel, extravasation\n        proc_pred[:, 0] = pred[:, 0]\n        proc_pred[:, 1] = 1 - proc_pred[:, 0]\n        proc_pred[:, 2] = pred[:, 1]\n        proc_pred[:, 3] = 1 - proc_pred[:, 1]\n\n        # liver, kidney, sneel\n        proc_pred[:, 4:7] = pred[:, 2:5]\n        proc_pred[:, 7:10] = pred[:, 5:8]\n        proc_pred[:, 10:13] = pred[:, 8:11]\n\n        return proc_pred\n\n\n​As far as I understand it has a bug (proc_pred[:, 0] and proc_pred[:, 1]; proc_pred[:, 2] and proc_pred[:, 3] are selected incorrectly).\n\nThis is what I think this post-processing should look like:\n\n\n    def post_proc_v2(pred):\n        proc_pred = np.empty((pred.shape[0], 2*2 + 3*3), dtype='float32')\n\n        # bowel, extravasation\n        proc_pred[:, 0] = 1 - pred[:, 0]\n        proc_pred[:, 1] = pred[:, 0]\n        proc_pred[:, 2] = 1 - pred[:, 1]\n        proc_pred[:, 3] = pred[:, 1]\n\n        # liver, kidney, sneel\n        proc_pred[:, 4:7] = pred[:, 2:5]\n        proc_pred[:, 7:10] = pred[:, 5:8]\n        proc_pred[:, 10:13] = pred[:, 8:11]\n\n        return proc_pred\n\nSimply put, we assign the model prediction for \"bowel_injury\" & \"extravasation_injury\" to \"bowel_healthy\" & \"extravasation_healthy\". \nThis gives us 0.69 lb and 1.00 CV. If we assign correctly, we estimate 0.97 lb and 0.7 CV.\nThe authors wrote about this in the comments.\n\nDoes anyone know what the problem is?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2430554,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-09-09T11:07:23.800000",
      "content": "<p>actually don't mind that much</p>\n<ol>\n<li>first just need to check there is no bug in your code.</li>\n<li>then ensure that results is producible.</li>\n<li>If there is a trend, capture it (if i imporve xxx, yyy is +ve or -ve correlated)</li>\n<li>the rest are just data … everything can be explained by data distribution of the train/test, etc …</li>\n</ol>\n<p>when the sample size is tool small, anything can happens, so don't mind that much.</p>\n<p>(espeically here, the high grade sample has low number and test case and high weighing).</p>\n<hr>\n<p>as an example, i find that for bowel submitting mean values is better than  a trained model prediction</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2430615,
          "author_name": "Andrij",
          "author_url": "",
          "post_date": "2023-09-09T12:11:26.430000",
          "content": "<p>Thanks, I think I found the problem. <br>\nThis is a problem with my training data. I think there will be a shake-up in this competition. <br>\nTo be honest, when I figured out what it was, I was shocked that I had such a good result. <br>\nThis is where I could be wrong, but it seems to me that my model misinterpreted some of the images, which is why I got this wrong result. <br>\nThanks for the important information about bowel. <br>\nBut now I feel that the result can be seriously improved. <br>\nBut where is the limit is another question… <br>\nThe first result of 0.42 may already be somewhere close, taking into account the specifics of the data and possible inaccuracies in the data themselves</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2430998,
              "author_name": "Andrij",
              "author_url": "",
              "post_date": "2023-09-09T17:56:32.340000",
              "content": "<p>Yup, I checked after making changes to the data on the validation and public sets I got the same 0.83.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2429137,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-09-08T11:58:00.117000",
      "content": "<p>I wouldn't trust public notebooks on a competition like this.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2429194,
          "author_name": "Andrij",
          "author_url": "",
          "post_date": "2023-09-08T12:43:10.863000",
          "content": "<p>I'm more interested in where the problem is: <br>\nis the problem in my code or in the test data. If the problem is in my code, that's bad. <br>\nMaybe the participants fixed some error that I didn't see. <br>\nIn addition, incorrect data processing can produce such results. <br>\nIf the problem is precisely in the public LB, it means that the shake-up will be unpleasant (According to my personal classification: minimal, normal, unpleasant, large, epic)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2430554": "actually don't mind that much\n\n1. first just need to check there is no bug in your code.\n2. then ensure that results is producible.\n3. If there is a trend, capture it (if i imporve xxx, yyy is +ve or -ve correlated)\n4. the rest are just data ... everything can be explained by data distribution of the train/test, etc ...\n\nwhen the sample size is tool small, anything can happens, so don't mind that much.\n\n(espeically here, the high grade sample has low number and test case and high weighing).\n\n---\n\nas an example, i find that for bowel submitting mean values is better than  a trained model prediction",
    "2429137": "I wouldn't trust public notebooks on a competition like this.",
    "2428892": "Hello everybody!\nAs a basis, I use this notebook branch (thanks to the authors): https://www.kaggle.com/code/coderrkj/rsna-atd-cnn-tpu-infer-corrections.\nThis notebook has the following post-processing:\n\n    def post_proc_v2(pred):\n        proc_pred = np.empty((pred.shape[0], 2*2 + 3*3), dtype='float32')\n\n        # bowel, extravasation\n        proc_pred[:, 0] = pred[:, 0]\n        proc_pred[:, 1] = 1 - proc_pred[:, 0]\n        proc_pred[:, 2] = pred[:, 1]\n        proc_pred[:, 3] = 1 - proc_pred[:, 1]\n\n        # liver, kidney, sneel\n        proc_pred[:, 4:7] = pred[:, 2:5]\n        proc_pred[:, 7:10] = pred[:, 5:8]\n        proc_pred[:, 10:13] = pred[:, 8:11]\n\n        return proc_pred\n\n\n​As far as I understand it has a bug (proc_pred[:, 0] and proc_pred[:, 1]; proc_pred[:, 2] and proc_pred[:, 3] are selected incorrectly).\n\nThis is what I think this post-processing should look like:\n\n\n    def post_proc_v2(pred):\n        proc_pred = np.empty((pred.shape[0], 2*2 + 3*3), dtype='float32')\n\n        # bowel, extravasation\n        proc_pred[:, 0] = 1 - pred[:, 0]\n        proc_pred[:, 1] = pred[:, 0]\n        proc_pred[:, 2] = 1 - pred[:, 1]\n        proc_pred[:, 3] = pred[:, 1]\n\n        # liver, kidney, sneel\n        proc_pred[:, 4:7] = pred[:, 2:5]\n        proc_pred[:, 7:10] = pred[:, 5:8]\n        proc_pred[:, 10:13] = pred[:, 8:11]\n\n        return proc_pred\n\nSimply put, we assign the model prediction for \"bowel_injury\" & \"extravasation_injury\" to \"bowel_healthy\" & \"extravasation_healthy\". \nThis gives us 0.69 lb and 1.00 CV. If we assign correctly, we estimate 0.97 lb and 0.7 CV.\nThe authors wrote about this in the comments.\n\nDoes anyone know what the problem is?"
  }
}