{
  "id": 607503,
  "title": "Localizers Coordinates Check",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/607503",
  "author_name": "k2-gc",
  "post_date": "2025-09-14T13:50:25.541000",
  "votes": 8,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I was checking the annotation data in <code>train_localizers.csv</code> and found something strange.</p>\n<p>For the following data:</p>\n<table>\n<thead>\n<tr>\n<th>idx</th>\n<th>SeriesInstanceUID</th>\n<th>SOPInstanceUID</th>\n<th>coordinates</th>\n<th>location</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>146</td>\n<td>1.2.826.0.1.3680043.8.498.10843288560910004558081082597234683103</td>\n<td>1.2.826.0.1.3680043.8.498.52735654348729179289737731303223326692</td>\n<td>{'x': 281.2280438311688, 'y': 167.23323863636364}</td>\n<td>Left Anterior Cerebral Artery</td>\n</tr>\n</tbody>\n</table>\n<p>The coordinates in <code>train_localizers.csv</code> seem to point to a location where nothing is visually apparent, and the corresponding pixel value is around unphysical -20000 HU (Red point below).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7024429%2Fed339dbe014ce58c100f33c9ceba0f4d%2Fclean.png?generation=1757857773864870&amp;alt=media\" alt=\"\"></p>\n<p>I am not sure if this is caused by a bug in my processing pipeline, or if it could be a mistake in the dataset itself.<br>\nCould anyone confirm whether this is expected behavior (e.g., a sentinel value for missing data) or an actual error in the annotations?</p>\n<p>If this is indeed a dataset issue, I would appreciate it if the data could be updated.</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 3288671,
      "postDate": "2025-09-14T13:50:25.540Z",
      "content": "<p>I was checking the annotation data in <code>train_localizers.csv</code> and found something strange.</p>\n<p>For the following data:</p>\n<table>\n<thead>\n<tr>\n<th>idx</th>\n<th>SeriesInstanceUID</th>\n<th>SOPInstanceUID</th>\n<th>coordinates</th>\n<th>location</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>146</td>\n<td>1.2.826.0.1.3680043.8.498.10843288560910004558081082597234683103</td>\n<td>1.2.826.0.1.3680043.8.498.52735654348729179289737731303223326692</td>\n<td>{'x': 281.2280438311688, 'y': 167.23323863636364}</td>\n<td>Left Anterior Cerebral Artery</td>\n</tr>\n</tbody>\n</table>\n<p>The coordinates in <code>train_localizers.csv</code> seem to point to a location where nothing is visually apparent, and the corresponding pixel value is around unphysical -20000 HU (Red point below).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7024429%2Fed339dbe014ce58c100f33c9ceba0f4d%2Fclean.png?generation=1757857773864870&amp;alt=media\" alt=\"\"></p>\n<p>I am not sure if this is caused by a bug in my processing pipeline, or if it could be a mistake in the dataset itself.<br>\nCould anyone confirm whether this is expected behavior (e.g., a sentinel value for missing data) or an actual error in the annotations?</p>\n<p>If this is indeed a dataset issue, I would appreciate it if the data could be updated.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "I was checking the annotation data in `train_localizers.csv` and found something strange.\n\nFor the following data:\n|idx | SeriesInstanceUID |  SOPInstanceUID|coordinates| location |\n| ---| ---| ---| --- | --- |\n| 146 | 1.2.826.0.1.3680043.8.498.10843288560910004558081082597234683103 | 1.2.826.0.1.3680043.8.498.52735654348729179289737731303223326692| {'x': 281.2280438311688, 'y': 167.23323863636364} | Left Anterior Cerebral Artery | \n\n\nThe coordinates in `train_localizers.csv` seem to point to a location where nothing is visually apparent, and the corresponding pixel value is around unphysical -20000 HU (Red point below).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7024429%2Fed339dbe014ce58c100f33c9ceba0f4d%2Fclean.png?generation=1757857773864870&alt=media)\n\nI am not sure if this is caused by a bug in my processing pipeline, or if it could be a mistake in the dataset itself.\nCould anyone confirm whether this is expected behavior (e.g., a sentinel value for missing data) or an actual error in the annotations?\n\nIf this is indeed a dataset issue, I would appreciate it if the data could be updated.\n\nThanks!",
      "votes": 8
    },
    {
      "id": 3288841,
      "postDate": "2025-09-15T00:07:41.147Z",
      "content": "<p>Thanks for identifying this issue. We are obviously very disappointed that this update has introduced a new issue, particularly given all of the additional expert person-hours it took to re-annotate these cases. We will get it solved soon!</p>",
      "rawMarkdown": "Thanks for identifying this issue. We are obviously very disappointed that this update has introduced a new issue, particularly given all of the additional expert person-hours it took to re-annotate these cases. We will get it solved soon!",
      "votes": 2,
      "replies": [
        {
          "id": 3288857,
          "postDate": "2025-09-15T01:47:02.393Z",
          "content": "<p>Thank you very much for your comment.</p>\n<p>For reference, here is how I checked this:</p>\n<ul>\n<li>Loaded DICOM data with pydicom</li>\n<li>Parsed Rescale Intercept and Rescale Slope</li>\n<li>Calculated real values using pixel_array * slope + intercept</li>\n<li>Plotted them with matplotlib</li>\n</ul>\n<p>I also checked all the CTA localizer points within a 2x2 pixel region.  <br>\nThe resulting histogram is shown below, and based on this I suspect that more series may have been annotated incorrectly.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7024429%2F3da77686527061ec8c12364a3caa0706%2Fclean.png?generation=1757900680630406&amp;alt=media\" alt=\"\"></p>\n<p>Another participant also reported a similar issue on this discussion, so it might be related to the updated DICOM files.</p>\n<blockquote>\n  <p>Hi,</p>\n  <p>I think it's not just you. I'm seeing the same after update 2. I visualized a sample of localizers, and all of them look incorrect. Here's an example for series 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382:<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F6db152d9db67c70f50e2e4b9befcf107%2Fproj_now.png?generation=1757871347201556&amp;alt=media\" alt=\"\"></p>\n  <p>And this is how it looked before update 2:<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F0a341715608a32af35799bd5e69e0908%2Fproj_before.png?generation=1757871385574308&amp;alt=media\" alt=\"\"></p>\n  <p>Interestingly, the records in train_localizers.csv have not changed:</p>\n<pre><code>.....,.....,,Other Posterior Circulation\n</code></pre>\n  <p>But the actual .dcm files have changed. Here's how the above slice looks now:<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2Fbbfd4e221d5556f23afc8cb6cdd7308c%2Fslice_now.png?generation=1757871579099642&amp;alt=media\" alt=\"\"></p>\n  <p>And before update 2, the same slice was:<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F24e43970e02535c74c721c67386cc4cb%2Fslice_before.png?generation=1757871617055706&amp;alt=media\" alt=\"\"></p>\n  <p>I hope this gets resolved.</p>\n</blockquote>\n<p>I’ll wait for the update and re-check the data afterwards.  <br>\nThank you very much.</p>",
          "rawMarkdown": "Thank you very much for your comment.\n\nFor reference, here is how I checked this:\n* Loaded DICOM data with pydicom\n* Parsed Rescale Intercept and Rescale Slope\n* Calculated real values using pixel_array * slope + intercept\n* Plotted them with matplotlib\n\nI also checked all the CTA localizer points within a 2x2 pixel region.  \nThe resulting histogram is shown below, and based on this I suspect that more series may have been annotated incorrectly.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7024429%2F3da77686527061ec8c12364a3caa0706%2Fclean.png?generation=1757900680630406&alt=media)\n\nAnother participant also reported a similar issue on this discussion, so it might be related to the updated DICOM files.\n> Hi,\n> \n> I think it's not just you. I'm seeing the same after update 2. I visualized a sample of localizers, and all of them look incorrect. Here's an example for series 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382:\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F6db152d9db67c70f50e2e4b9befcf107%2Fproj_now.png?generation=1757871347201556&alt=media)\n> \n> And this is how it looked before update 2:\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F0a341715608a32af35799bd5e69e0908%2Fproj_before.png?generation=1757871385574308&alt=media)\n> \n> Interestingly, the records in train_localizers.csv have not changed:\n> ```\n> 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382,1.2.826.0.1.3680043.8.498.26441166830185072464905619055293584091,\"{'x': 240.87196819276753, 'y': 299.0633884651388}\",Other Posterior Circulation\n> ```\n> \n> But the actual .dcm files have changed. Here's how the above slice looks now:\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2Fbbfd4e221d5556f23afc8cb6cdd7308c%2Fslice_now.png?generation=1757871579099642&alt=media)\n> \n> And before update 2, the same slice was:\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F24e43970e02535c74c721c67386cc4cb%2Fslice_before.png?generation=1757871617055706&alt=media)\n> \n> I hope this gets resolved.\n\n\n\nI’ll wait for the update and re-check the data afterwards.  \nThank you very much.",
          "votes": 3,
          "replies": [
            {
              "id": 3290435,
              "postDate": "2025-09-17T19:58:36.820Z",
              "content": "<p>Thanks again for this. While this is a good screening method, HU value at the exact aneurysm coordinate may not be 100% reliable. Particularly for some of the bone removed series like you show in your last few images. I am told that the old version of the dataset is the current one right now, so there should only be a very few localizer issues and most of them should be accurate. We will upload the re-annotated localizers as soon as possible. Just waiting to hear from kaggle what exactly the issue was since it seems like it was on their end.</p>",
              "rawMarkdown": "Thanks again for this. While this is a good screening method, HU value at the exact aneurysm coordinate may not be 100% reliable. Particularly for some of the bone removed series like you show in your last few images. I am told that the old version of the dataset is the current one right now, so there should only be a very few localizer issues and most of them should be accurate. We will upload the re-annotated localizers as soon as possible. Just waiting to hear from kaggle what exactly the issue was since it seems like it was on their end.",
              "votes": 5
            }
          ]
        },
        {
          "id": 3289000,
          "postDate": "2025-09-15T08:39:26.860Z",
          "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> <br>\nAlso this might help.</p>\n<blockquote>\n  <p>The coordinates are ok. The issue is that many images in the series were renamed, but the SOPInstanceUID column in train_localizers.csv was mistakenly left unchanged. So when you try to visualize the annotation you draw the correct x,y coordinates, but you open another image in the series (the image names are mixed up after the last data update)<br>\n  Each image in a series has a unique InstanceNumber, so you can find the new filename if you know its InstanceNumber (if you saved the competition data before the update).</p>\n  <p>SOPInstanceUID -&gt; InstanceNumber -&gt; SOPInstanceUID_new</p>\n</blockquote>\n<p>Thank you,</p>",
          "rawMarkdown": "@evancalabrese \nAlso this might help.\n\n> The coordinates are ok. The issue is that many images in the series were renamed, but the SOPInstanceUID column in train_localizers.csv was mistakenly left unchanged. So when you try to visualize the annotation you draw the correct x,y coordinates, but you open another image in the series (the image names are mixed up after the last data update)\n> Each image in a series has a unique InstanceNumber, so you can find the new filename if you know its InstanceNumber (if you saved the competition data before the update).\n> \n> SOPInstanceUID -> InstanceNumber -> SOPInstanceUID_new\n\n\n\nThank you,",
          "votes": 1,
          "replies": [
            {
              "id": 3289291,
              "postDate": "2025-09-15T19:50:13.873Z",
              "content": "<p>Thanks for pointing this out! I also just ran into this problem. I wanted to be thorough and went through a lot of annotated locations and most of them seemed really strange. I checked my pipeline like 100 times but this explains it. <br>\nIs the old dataset still available somewhere or can we hope that this will be fixed soon <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>? </p>",
              "rawMarkdown": "Thanks for pointing this out! I also just ran into this problem. I wanted to be thorough and went through a lot of annotated locations and most of them seemed really strange. I checked my pipeline like 100 times but this explains it. \nIs the old dataset still available somewhere or can we hope that this will be fixed soon @evancalabrese? "
            },
            {
              "id": 3289316,
              "postDate": "2025-09-15T21:05:56.940Z",
              "content": "<p>I'm sorry about the confusion. I thought I had already downloaded the latest data. But I just saw that the train_localizers.csv I downloaded on the 11th of Septembre has a different row count than the one that is available now. Have the file names of the images changed? Is it enough to just get the latest csvs or will I need to download the whole dataset again? Thanks for your help.</p>",
              "rawMarkdown": "I'm sorry about the confusion. I thought I had already downloaded the latest data. But I just saw that the train_localizers.csv I downloaded on the 11th of Septembre has a different row count than the one that is available now. Have the file names of the images changed? Is it enough to just get the latest csvs or will I need to download the whole dataset again? Thanks for your help."
            },
            {
              "id": 3289556,
              "postDate": "2025-09-16T08:25:38.420Z",
              "content": "<p>I made a minimal notebook to check the localizers before downloading the dataset again:</p>\n<p><a href=\"https://www.kaggle.com/code/thalro/check-localizers\" target=\"_blank\">https://www.kaggle.com/code/thalro/check-localizers</a></p>\n<p>The online data now seems to correct as opposed to my local version. I will download it again and see how it goes.</p>",
              "rawMarkdown": "I made a minimal notebook to check the localizers before downloading the dataset again:\n\nhttps://www.kaggle.com/code/thalro/check-localizers\n\nThe online data now seems to correct as opposed to my local version. I will download it again and see how it goes.",
              "votes": 1
            },
            {
              "id": 3289584,
              "postDate": "2025-09-16T09:12:53.330Z",
              "content": "<p>Also, I just inserted the id from the top of this thread and the localizer seems to be showing something more reasonable now. (<a href=\"https://www.kaggle.com/code/thalro/check-localizers\" target=\"_blank\">https://www.kaggle.com/code/thalro/check-localizers</a>)</p>",
              "rawMarkdown": "Also, I just inserted the id from the top of this thread and the localizer seems to be showing something more reasonable now. (https://www.kaggle.com/code/thalro/check-localizers)",
              "votes": 1
            },
            {
              "id": 3290377,
              "postDate": "2025-09-17T17:49:06.173Z",
              "content": "<p>We have reverted to the older dataset for now which does not have this issue.</p>",
              "rawMarkdown": "We have reverted to the older dataset for now which does not have this issue.",
              "votes": 2
            },
            {
              "id": 3290392,
              "postDate": "2025-09-17T18:19:18.183Z",
              "content": "<p>Hi. Thanks. But what issue exactly? I'm not confident if there it was actually any issue rather than may be preprocessing bugs.</p>\n<p>EDIT: I've checked and I agree 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382 case.</p>",
              "rawMarkdown": "Hi. Thanks. But what issue exactly? I'm not confident if there it was actually any issue rather than may be preprocessing bugs.\n\nEDIT: I've checked and I agree 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382 case."
            },
            {
              "id": 3291125,
              "postDate": "2025-09-19T03:23:15.857Z",
              "content": "<blockquote>\n  <p>We have reverted to the older dataset for now which does not have this issue.</p>\n</blockquote>\n<p>Hi Evan, thanks for the confirmation. I’m a bit unclear about the dataset version we’re using now. As I understand it:</p>\n<p>We started with dataset v1 - Then Update 1 → dataset v2 - Then Update 2 → dataset v3 (which has the SOPInstanceUID ordering issue)</p>\n<p>So is the current version (as of 09/18) the dataset before Update 2 (i.e. version 2)? Please correct me if I’ve got this wrong.</p>",
              "rawMarkdown": "> We have reverted to the older dataset for now which does not have this issue.\n\nHi Evan, thanks for the confirmation. I’m a bit unclear about the dataset version we’re using now. As I understand it:\n\nWe started with dataset v1 - Then Update 1 → dataset v2 - Then Update 2 → dataset v3 (which has the SOPInstanceUID ordering issue)\n\nSo is the current version (as of 09/18) the dataset before Update 2 (i.e. version 2)? Please correct me if I’ve got this wrong.",
              "votes": 2
            },
            {
              "id": 3291213,
              "postDate": "2025-09-19T07:41:03.117Z",
              "content": "<p>I want to know about it, too.</p>",
              "rawMarkdown": "I want to know about it, too."
            }
          ]
        }
      ]
    },
    {
      "id": 3288761,
      "postDate": "2025-09-14T17:47:11.857Z",
      "content": "<p>Hi,</p>\n<p>I think it's not just you. I'm seeing the same after update 2. I visualized a sample of localizers, and all of them look incorrect. Here's an example for series 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F6db152d9db67c70f50e2e4b9befcf107%2Fproj_now.png?generation=1757871347201556&amp;alt=media\" alt=\"\"></p>\n<p>And this is how it looked before update 2:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F0a341715608a32af35799bd5e69e0908%2Fproj_before.png?generation=1757871385574308&amp;alt=media\" alt=\"\"></p>\n<p>Interestingly, the records in train_localizers.csv have not changed:</p>\n<pre><code>.....,.....,,Other Posterior Circulation\n</code></pre>\n<p>But the actual .dcm files have changed. Here's how the above slice looks now:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2Fbbfd4e221d5556f23afc8cb6cdd7308c%2Fslice_now.png?generation=1757871579099642&amp;alt=media\" alt=\"\"></p>\n<p>And before update 2, the same slice was:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F24e43970e02535c74c721c67386cc4cb%2Fslice_before.png?generation=1757871617055706&amp;alt=media\" alt=\"\"></p>\n<p>I hope this gets resolved.</p>",
      "rawMarkdown": "Hi,\n\nI think it's not just you. I'm seeing the same after update 2. I visualized a sample of localizers, and all of them look incorrect. Here's an example for series 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F6db152d9db67c70f50e2e4b9befcf107%2Fproj_now.png?generation=1757871347201556&alt=media)\n\nAnd this is how it looked before update 2:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F0a341715608a32af35799bd5e69e0908%2Fproj_before.png?generation=1757871385574308&alt=media)\n\nInterestingly, the records in train_localizers.csv have not changed:\n```\n1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382,1.2.826.0.1.3680043.8.498.26441166830185072464905619055293584091,\"{'x': 240.87196819276753, 'y': 299.0633884651388}\",Other Posterior Circulation\n```\n\nBut the actual .dcm files have changed. Here's how the above slice looks now:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2Fbbfd4e221d5556f23afc8cb6cdd7308c%2Fslice_now.png?generation=1757871579099642&alt=media)\n\nAnd before update 2, the same slice was:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F24e43970e02535c74c721c67386cc4cb%2Fslice_before.png?generation=1757871617055706&alt=media)\n\nI hope this gets resolved.",
      "votes": 2,
      "replies": [
        {
          "id": 3288854,
          "postDate": "2025-09-15T01:30:38.443Z",
          "content": "<p>Thank you for your reply.  <br>\nThe images you posted also clearly show incorrect annotations.  <br>\nSo I think this is not an issue on my side, but rather a dataset problem.  <br>\nSince the competition host also replied to this discussion, I will follow up with them.</p>",
          "rawMarkdown": "Thank you for your reply.  \nThe images you posted also clearly show incorrect annotations.  \nSo I think this is not an issue on my side, but rather a dataset problem.  \nSince the competition host also replied to this discussion, I will follow up with them.",
          "votes": 1,
          "replies": [
            {
              "id": 3288956,
              "postDate": "2025-09-15T07:05:25.583Z",
              "content": "<p>Hi. Try my fixed file train_localizers.csv, where I added two new columns: Number (true InstanceNumber in the series) and fixed column SOPInstanceUID_new. Original columns unchanged.<br>\n<a href=\"https://www.kaggle.com/datasets/baurzhanurazalinov/rsna-detection-localizers/\" target=\"_blank\">https://www.kaggle.com/datasets/baurzhanurazalinov/rsna-detection-localizers/</a></p>\n<pre><code>train_locs = pd.read_csv()\ntrain_locs[] = train_locs[]\n</code></pre>",
              "rawMarkdown": "Hi. Try my fixed file train_localizers.csv, where I added two new columns: Number (true InstanceNumber in the series) and fixed column SOPInstanceUID_new. Original columns unchanged.\nhttps://www.kaggle.com/datasets/baurzhanurazalinov/rsna-detection-localizers/\n\n```python\ntrain_locs = pd.read_csv('/kaggle/input/rsna-detection-localizers/train_localizers.csv')\ntrain_locs['SOPInstanceUID'] = train_locs['SOPInstanceUID_new']\n```",
              "votes": 1
            },
            {
              "id": 3288981,
              "postDate": "2025-09-15T07:48:41.507Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3288983,
              "postDate": "2025-09-15T07:48:57.607Z",
              "content": "<p>Thank you very much.<br>\nI will give it a try later!</p>\n<p>By the way, how did you fix the annotation?<br>\nAnd as far as I can see your fixed file, the mistakes are caused by invalid SOPInstanceUID, not invalid  coordinates?</p>",
              "rawMarkdown": "Thank you very much.\nI will give it a try later!\n\nBy the way, how did you fix the annotation?\nAnd as far as I can see your fixed file, the mistakes are caused by invalid SOPInstanceUID, not invalid  coordinates?",
              "votes": 1
            },
            {
              "id": 3288995,
              "postDate": "2025-09-15T08:31:17.810Z",
              "content": "<p>The coordinates are ok. The issue is that many images in series were renamed, but the SOPInstanceUID column in train_localizers.csv was mistakenly left unchanged. So when you try to visualize the annotation you draw the correct x,y coordinates, but you open another image in the series (the image names are mixed up after the last data update)<br>\nEach image in a series has a unique InstanceNumber, so you can find the new filename if you know its InstanceNumber (if you saved the competition data before the update).</p>\n<p>SOPInstanceUID -&gt; InstanceNumber -&gt; SOPInstanceUID_new</p>",
              "rawMarkdown": "The coordinates are ok. The issue is that many images in series were renamed, but the SOPInstanceUID column in train_localizers.csv was mistakenly left unchanged. So when you try to visualize the annotation you draw the correct x,y coordinates, but you open another image in the series (the image names are mixed up after the last data update)\nEach image in a series has a unique InstanceNumber, so you can find the new filename if you know its InstanceNumber (if you saved the competition data before the update).\n\nSOPInstanceUID -> InstanceNumber -> SOPInstanceUID_new",
              "votes": 1
            },
            {
              "id": 3288999,
              "postDate": "2025-09-15T08:38:15.777Z",
              "content": "<p>Thank you very much!<br>\nI will check for it!</p>\n<p>And I will share this with the competition host on this discussion.</p>",
              "rawMarkdown": "Thank you very much!\nI will check for it!\n\nAnd I will share this with the competition host on this discussion.",
              "votes": 1
            },
            {
              "id": 3289048,
              "postDate": "2025-09-15T10:36:37.790Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F61b153383bcdcad036fa34b29a3410db%2FSin%20ttulo.jpg?generation=1757932595901351&amp;alt=media\" alt=\"\"></p>\n<p>Thanks for share. But are you sure there is 1886 wrong SOP_UID? I've been training with supposed wrong coordinates and for my results is hard to belive.</p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F61b153383bcdcad036fa34b29a3410db%2FSin%20ttulo.jpg?generation=1757932595901351&alt=media)\n\nThanks for share. But are you sure there is 1886 wrong SOP_UID? I've been training with supposed wrong coordinates and for my results is hard to belive."
            }
          ]
        }
      ]
    },
    {
      "id": 3288768,
      "postDate": "2025-09-14T18:00:40.963Z",
      "content": "<p>I guess it depends on whether it's a medical monitor or not.</p>",
      "rawMarkdown": "I guess it depends on whether it's a medical monitor or not.",
      "replies": [
        {
          "id": 3288853,
          "postDate": "2025-09-15T01:25:07.117Z",
          "content": "<p>Thank you very much.<br>\nInteresting point.  <br>\nHowever, I believe that a value around -20,000 HU does not correspond to any physical quantity in CT and might indicate invalid or missing data.</p>",
          "rawMarkdown": "Thank you very much.\nInteresting point.  \nHowever, I believe that a value around -20,000 HU does not correspond to any physical quantity in CT and might indicate invalid or missing data.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3288841,
      "author_name": "Evan Calabrese",
      "author_url": "",
      "post_date": "2025-09-15T00:07:41.147000",
      "content": "<p>Thanks for identifying this issue. We are obviously very disappointed that this update has introduced a new issue, particularly given all of the additional expert person-hours it took to re-annotate these cases. We will get it solved soon!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3288857,
          "author_name": "k2-gc",
          "author_url": "",
          "post_date": "2025-09-15T01:47:02.393000",
          "content": "<p>Thank you very much for your comment.</p>\n<p>For reference, here is how I checked this:</p>\n<ul>\n<li>Loaded DICOM data with pydicom</li>\n<li>Parsed Rescale Intercept and Rescale Slope</li>\n<li>Calculated real values using pixel_array * slope + intercept</li>\n<li>Plotted them with matplotlib</li>\n</ul>\n<p>I also checked all the CTA localizer points within a 2x2 pixel region.  <br>\nThe resulting histogram is shown below, and based on this I suspect that more series may have been annotated incorrectly.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7024429%2F3da77686527061ec8c12364a3caa0706%2Fclean.png?generation=1757900680630406&amp;alt=media\" alt=\"\"></p>\n<p>Another participant also reported a similar issue on this discussion, so it might be related to the updated DICOM files.</p>\n<blockquote>\n  <p>Hi,</p>\n  <p>I think it's not just you. I'm seeing the same after update 2. I visualized a sample of localizers, and all of them look incorrect. Here's an example for series 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382:<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F6db152d9db67c70f50e2e4b9befcf107%2Fproj_now.png?generation=1757871347201556&amp;alt=media\" alt=\"\"></p>\n  <p>And this is how it looked before update 2:<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F0a341715608a32af35799bd5e69e0908%2Fproj_before.png?generation=1757871385574308&amp;alt=media\" alt=\"\"></p>\n  <p>Interestingly, the records in train_localizers.csv have not changed:</p>\n<pre><code>.....,.....,,Other Posterior Circulation\n</code></pre>\n  <p>But the actual .dcm files have changed. Here's how the above slice looks now:<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2Fbbfd4e221d5556f23afc8cb6cdd7308c%2Fslice_now.png?generation=1757871579099642&amp;alt=media\" alt=\"\"></p>\n  <p>And before update 2, the same slice was:<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F24e43970e02535c74c721c67386cc4cb%2Fslice_before.png?generation=1757871617055706&amp;alt=media\" alt=\"\"></p>\n  <p>I hope this gets resolved.</p>\n</blockquote>\n<p>I’ll wait for the update and re-check the data afterwards.  <br>\nThank you very much.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3290435,
              "author_name": "Evan Calabrese",
              "author_url": "",
              "post_date": "2025-09-17T19:58:36.820000",
              "content": "<p>Thanks again for this. While this is a good screening method, HU value at the exact aneurysm coordinate may not be 100% reliable. Particularly for some of the bone removed series like you show in your last few images. I am told that the old version of the dataset is the current one right now, so there should only be a very few localizer issues and most of them should be accurate. We will upload the re-annotated localizers as soon as possible. Just waiting to hear from kaggle what exactly the issue was since it seems like it was on their end.</p>",
              "votes": 5,
              "replies": []
            }
          ]
        },
        {
          "id": 3289000,
          "author_name": "k2-gc",
          "author_url": "",
          "post_date": "2025-09-15T08:39:26.860000",
          "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> <br>\nAlso this might help.</p>\n<blockquote>\n  <p>The coordinates are ok. The issue is that many images in the series were renamed, but the SOPInstanceUID column in train_localizers.csv was mistakenly left unchanged. So when you try to visualize the annotation you draw the correct x,y coordinates, but you open another image in the series (the image names are mixed up after the last data update)<br>\n  Each image in a series has a unique InstanceNumber, so you can find the new filename if you know its InstanceNumber (if you saved the competition data before the update).</p>\n  <p>SOPInstanceUID -&gt; InstanceNumber -&gt; SOPInstanceUID_new</p>\n</blockquote>\n<p>Thank you,</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3289291,
              "author_name": "thomas rost",
              "author_url": "",
              "post_date": "2025-09-15T19:50:13.873000",
              "content": "<p>Thanks for pointing this out! I also just ran into this problem. I wanted to be thorough and went through a lot of annotated locations and most of them seemed really strange. I checked my pipeline like 100 times but this explains it. <br>\nIs the old dataset still available somewhere or can we hope that this will be fixed soon <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3289316,
              "author_name": "thomas rost",
              "author_url": "",
              "post_date": "2025-09-15T21:05:56.940000",
              "content": "<p>I'm sorry about the confusion. I thought I had already downloaded the latest data. But I just saw that the train_localizers.csv I downloaded on the 11th of Septembre has a different row count than the one that is available now. Have the file names of the images changed? Is it enough to just get the latest csvs or will I need to download the whole dataset again? Thanks for your help.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3289556,
              "author_name": "thomas rost",
              "author_url": "",
              "post_date": "2025-09-16T08:25:38.420000",
              "content": "<p>I made a minimal notebook to check the localizers before downloading the dataset again:</p>\n<p><a href=\"https://www.kaggle.com/code/thalro/check-localizers\" target=\"_blank\">https://www.kaggle.com/code/thalro/check-localizers</a></p>\n<p>The online data now seems to correct as opposed to my local version. I will download it again and see how it goes.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3289584,
              "author_name": "thomas rost",
              "author_url": "",
              "post_date": "2025-09-16T09:12:53.330000",
              "content": "<p>Also, I just inserted the id from the top of this thread and the localizer seems to be showing something more reasonable now. (<a href=\"https://www.kaggle.com/code/thalro/check-localizers\" target=\"_blank\">https://www.kaggle.com/code/thalro/check-localizers</a>)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3290377,
              "author_name": "Evan Calabrese",
              "author_url": "",
              "post_date": "2025-09-17T17:49:06.173000",
              "content": "<p>We have reverted to the older dataset for now which does not have this issue.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3290392,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2025-09-17T18:19:18.183000",
              "content": "<p>Hi. Thanks. But what issue exactly? I'm not confident if there it was actually any issue rather than may be preprocessing bugs.</p>\n<p>EDIT: I've checked and I agree 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382 case.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3291125,
              "author_name": "Luca",
              "author_url": "",
              "post_date": "2025-09-19T03:23:15.857000",
              "content": "<blockquote>\n  <p>We have reverted to the older dataset for now which does not have this issue.</p>\n</blockquote>\n<p>Hi Evan, thanks for the confirmation. I’m a bit unclear about the dataset version we’re using now. As I understand it:</p>\n<p>We started with dataset v1 - Then Update 1 → dataset v2 - Then Update 2 → dataset v3 (which has the SOPInstanceUID ordering issue)</p>\n<p>So is the current version (as of 09/18) the dataset before Update 2 (i.e. version 2)? Please correct me if I’ve got this wrong.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3291213,
              "author_name": "k2-gc",
              "author_url": "",
              "post_date": "2025-09-19T07:41:03.117000",
              "content": "<p>I want to know about it, too.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3288761,
      "author_name": "vsemionov",
      "author_url": "",
      "post_date": "2025-09-14T17:47:11.857000",
      "content": "<p>Hi,</p>\n<p>I think it's not just you. I'm seeing the same after update 2. I visualized a sample of localizers, and all of them look incorrect. Here's an example for series 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F6db152d9db67c70f50e2e4b9befcf107%2Fproj_now.png?generation=1757871347201556&amp;alt=media\" alt=\"\"></p>\n<p>And this is how it looked before update 2:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F0a341715608a32af35799bd5e69e0908%2Fproj_before.png?generation=1757871385574308&amp;alt=media\" alt=\"\"></p>\n<p>Interestingly, the records in train_localizers.csv have not changed:</p>\n<pre><code>.....,.....,,Other Posterior Circulation\n</code></pre>\n<p>But the actual .dcm files have changed. Here's how the above slice looks now:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2Fbbfd4e221d5556f23afc8cb6cdd7308c%2Fslice_now.png?generation=1757871579099642&amp;alt=media\" alt=\"\"></p>\n<p>And before update 2, the same slice was:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F24e43970e02535c74c721c67386cc4cb%2Fslice_before.png?generation=1757871617055706&amp;alt=media\" alt=\"\"></p>\n<p>I hope this gets resolved.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3288854,
          "author_name": "k2-gc",
          "author_url": "",
          "post_date": "2025-09-15T01:30:38.443000",
          "content": "<p>Thank you for your reply.  <br>\nThe images you posted also clearly show incorrect annotations.  <br>\nSo I think this is not an issue on my side, but rather a dataset problem.  <br>\nSince the competition host also replied to this discussion, I will follow up with them.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3288956,
              "author_name": "Urazalinov Baurzhan",
              "author_url": "",
              "post_date": "2025-09-15T07:05:25.583000",
              "content": "<p>Hi. Try my fixed file train_localizers.csv, where I added two new columns: Number (true InstanceNumber in the series) and fixed column SOPInstanceUID_new. Original columns unchanged.<br>\n<a href=\"https://www.kaggle.com/datasets/baurzhanurazalinov/rsna-detection-localizers/\" target=\"_blank\">https://www.kaggle.com/datasets/baurzhanurazalinov/rsna-detection-localizers/</a></p>\n<pre><code>train_locs = pd.read_csv()\ntrain_locs[] = train_locs[]\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3288981,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-09-15T07:48:41.507000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3288983,
              "author_name": "k2-gc",
              "author_url": "",
              "post_date": "2025-09-15T07:48:57.607000",
              "content": "<p>Thank you very much.<br>\nI will give it a try later!</p>\n<p>By the way, how did you fix the annotation?<br>\nAnd as far as I can see your fixed file, the mistakes are caused by invalid SOPInstanceUID, not invalid  coordinates?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3288995,
              "author_name": "Urazalinov Baurzhan",
              "author_url": "",
              "post_date": "2025-09-15T08:31:17.810000",
              "content": "<p>The coordinates are ok. The issue is that many images in series were renamed, but the SOPInstanceUID column in train_localizers.csv was mistakenly left unchanged. So when you try to visualize the annotation you draw the correct x,y coordinates, but you open another image in the series (the image names are mixed up after the last data update)<br>\nEach image in a series has a unique InstanceNumber, so you can find the new filename if you know its InstanceNumber (if you saved the competition data before the update).</p>\n<p>SOPInstanceUID -&gt; InstanceNumber -&gt; SOPInstanceUID_new</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3288999,
              "author_name": "k2-gc",
              "author_url": "",
              "post_date": "2025-09-15T08:38:15.777000",
              "content": "<p>Thank you very much!<br>\nI will check for it!</p>\n<p>And I will share this with the competition host on this discussion.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3289048,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2025-09-15T10:36:37.790000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F61b153383bcdcad036fa34b29a3410db%2FSin%20ttulo.jpg?generation=1757932595901351&amp;alt=media\" alt=\"\"></p>\n<p>Thanks for share. But are you sure there is 1886 wrong SOP_UID? I've been training with supposed wrong coordinates and for my results is hard to belive.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3288768,
      "author_name": "ktt proc",
      "author_url": "",
      "post_date": "2025-09-14T18:00:40.963000",
      "content": "<p>I guess it depends on whether it's a medical monitor or not.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3288853,
          "author_name": "k2-gc",
          "author_url": "",
          "post_date": "2025-09-15T01:25:07.117000",
          "content": "<p>Thank you very much.<br>\nInteresting point.  <br>\nHowever, I believe that a value around -20,000 HU does not correspond to any physical quantity in CT and might indicate invalid or missing data.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3288671": "I was checking the annotation data in `train_localizers.csv` and found something strange.\n\nFor the following data:\n|idx | SeriesInstanceUID |  SOPInstanceUID|coordinates| location |\n| ---| ---| ---| --- | --- |\n| 146 | 1.2.826.0.1.3680043.8.498.10843288560910004558081082597234683103 | 1.2.826.0.1.3680043.8.498.52735654348729179289737731303223326692| {'x': 281.2280438311688, 'y': 167.23323863636364} | Left Anterior Cerebral Artery | \n\n\nThe coordinates in `train_localizers.csv` seem to point to a location where nothing is visually apparent, and the corresponding pixel value is around unphysical -20000 HU (Red point below).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7024429%2Fed339dbe014ce58c100f33c9ceba0f4d%2Fclean.png?generation=1757857773864870&alt=media)\n\nI am not sure if this is caused by a bug in my processing pipeline, or if it could be a mistake in the dataset itself.\nCould anyone confirm whether this is expected behavior (e.g., a sentinel value for missing data) or an actual error in the annotations?\n\nIf this is indeed a dataset issue, I would appreciate it if the data could be updated.\n\nThanks!",
    "3288841": "Thanks for identifying this issue. We are obviously very disappointed that this update has introduced a new issue, particularly given all of the additional expert person-hours it took to re-annotate these cases. We will get it solved soon!",
    "3288761": "Hi,\n\nI think it's not just you. I'm seeing the same after update 2. I visualized a sample of localizers, and all of them look incorrect. Here's an example for series 1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F6db152d9db67c70f50e2e4b9befcf107%2Fproj_now.png?generation=1757871347201556&alt=media)\n\nAnd this is how it looked before update 2:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F0a341715608a32af35799bd5e69e0908%2Fproj_before.png?generation=1757871385574308&alt=media)\n\nInterestingly, the records in train_localizers.csv have not changed:\n```\n1.2.826.0.1.3680043.8.498.10935907012185032169927418164924236382,1.2.826.0.1.3680043.8.498.26441166830185072464905619055293584091,\"{'x': 240.87196819276753, 'y': 299.0633884651388}\",Other Posterior Circulation\n```\n\nBut the actual .dcm files have changed. Here's how the above slice looks now:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2Fbbfd4e221d5556f23afc8cb6cdd7308c%2Fslice_now.png?generation=1757871579099642&alt=media)\n\nAnd before update 2, the same slice was:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16527972%2F24e43970e02535c74c721c67386cc4cb%2Fslice_before.png?generation=1757871617055706&alt=media)\n\nI hope this gets resolved.",
    "3288768": "I guess it depends on whether it's a medical monitor or not."
  }
}