{
  "id": 610101,
  "title": "Regarding the remaining issues with competition data",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/610101",
  "author_name": "Evan Calabrese",
  "post_date": "2025-10-01T16:33:45.765000",
  "votes": 23,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>As many of you have pointed out, there are still a few remaining issues with the competition data. We have been working hard behind the scenes for the past few months to correct a variety of issues that have come up including several that required reprocessing and expert manual review of the data. We are happy that we fixed several of the bigger issues, but we know we haven't been able to fix them all. </p>\n<p>Unfortunately, kaggle's process for updating the competition dataset is extremely slow, taking ~1 week for minor updates such as modifying a single CSV. They have their reasons for this, and I can't comment on the merit of this approach. Given this, and the &lt;2 week timeframe remaining in the challenge, <strong>we will NOT be making any additional changes to the dataset</strong>.</p>\n<p>This project was very ambitious from the start. Not only is it one of the largest ever medical imaging competition datasets ever released, but it also has several other complexities including inhomogeneous multi-site data, nuances in DICOM format, and several different data types (image, categorical, location, segmentation etc). It also required hundreds of hours of expert manual annotation. We are proud of the dataset we have created, but we know it still has many flaws. We sincerely hope you understand.</p>\n<p><strong>I am pinning this thread for visibility and in the hopes that we can use it to:</strong><br>\n<strong>1) Have a single place for aggregating known remaining issues with the dataset.</strong><br>\n<strong>2) Sharing any code (if you are willing) that competitors can be use to patch, fix, or avoid any remaining dataset issues.</strong></p>\n<p>Thanks, and good luck in the final 2 weeks!</p>",
  "messages": [
    {
      "id": 3296805,
      "postDate": "2025-10-01T16:33:45.767Z",
      "content": "<p>Hi all,</p>\n<p>As many of you have pointed out, there are still a few remaining issues with the competition data. We have been working hard behind the scenes for the past few months to correct a variety of issues that have come up including several that required reprocessing and expert manual review of the data. We are happy that we fixed several of the bigger issues, but we know we haven't been able to fix them all. </p>\n<p>Unfortunately, kaggle's process for updating the competition dataset is extremely slow, taking ~1 week for minor updates such as modifying a single CSV. They have their reasons for this, and I can't comment on the merit of this approach. Given this, and the &lt;2 week timeframe remaining in the challenge, <strong>we will NOT be making any additional changes to the dataset</strong>.</p>\n<p>This project was very ambitious from the start. Not only is it one of the largest ever medical imaging competition datasets ever released, but it also has several other complexities including inhomogeneous multi-site data, nuances in DICOM format, and several different data types (image, categorical, location, segmentation etc). It also required hundreds of hours of expert manual annotation. We are proud of the dataset we have created, but we know it still has many flaws. We sincerely hope you understand.</p>\n<p><strong>I am pinning this thread for visibility and in the hopes that we can use it to:</strong><br>\n<strong>1) Have a single place for aggregating known remaining issues with the dataset.</strong><br>\n<strong>2) Sharing any code (if you are willing) that competitors can be use to patch, fix, or avoid any remaining dataset issues.</strong></p>\n<p>Thanks, and good luck in the final 2 weeks!</p>",
      "rawMarkdown": "Hi all,\n\nAs many of you have pointed out, there are still a few remaining issues with the competition data. We have been working hard behind the scenes for the past few months to correct a variety of issues that have come up including several that required reprocessing and expert manual review of the data. We are happy that we fixed several of the bigger issues, but we know we haven't been able to fix them all. \n\nUnfortunately, kaggle's process for updating the competition dataset is extremely slow, taking ~1 week for minor updates such as modifying a single CSV. They have their reasons for this, and I can't comment on the merit of this approach. Given this, and the <2 week timeframe remaining in the challenge, **we will NOT be making any additional changes to the dataset**.\n\nThis project was very ambitious from the start. Not only is it one of the largest ever medical imaging competition datasets ever released, but it also has several other complexities including inhomogeneous multi-site data, nuances in DICOM format, and several different data types (image, categorical, location, segmentation etc). It also required hundreds of hours of expert manual annotation. We are proud of the dataset we have created, but we know it still has many flaws. We sincerely hope you understand.\n\n**I am pinning this thread for visibility and in the hopes that we can use it to:**\n**1) Have a single place for aggregating known remaining issues with the dataset.**\n**2) Sharing any code (if you are willing) that competitors can be use to patch, fix, or avoid any remaining dataset issues.**\n\nThanks, and good luck in the final 2 weeks!",
      "votes": 23
    },
    {
      "id": 3299025,
      "postDate": "2025-10-07T01:58:50.393Z",
      "content": "<p>It appears that some participants rely on DICOM attributes during inference, and for certain Multi-Frame DICOMs, these attributes may be missing, leading to invalid or failed inferences. In such cases, you may consider using the following default parameters as a shared function group:<br>\nspacing=(0.5, 0.5), origin=(0, 0, 0), orientation=[1, 0, 0, 0, 1, 0], slice_thickness = 5.0</p>\n<p>While I cannot confirm that this will improve performance, it may be worth trying.</p>",
      "rawMarkdown": "It appears that some participants rely on DICOM attributes during inference, and for certain Multi-Frame DICOMs, these attributes may be missing, leading to invalid or failed inferences. In such cases, you may consider using the following default parameters as a shared function group:\nspacing=(0.5, 0.5), origin=(0, 0, 0), orientation=[1, 0, 0, 0, 1, 0], slice_thickness = 5.0\n\nWhile I cannot confirm that this will improve performance, it may be worth trying.",
      "votes": 3,
      "replies": [
        {
          "id": 3299031,
          "postDate": "2025-10-07T02:51:30.237Z",
          "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> . <a href=\"https://www.kaggle.com/rachitsaluja\" target=\"_blank\">@rachitsaluja</a> <br>\nYes, as you say, it's important to specify it by default.<br>\nOn the other hand, I have one request.<br>\nIf possible, could you provide a few examples of multi-frame DICOM data?<br>\nOf course, the pixel data can be provided as dummy data.</p>\n<p>Simply by sharing such data, I believe we can significantly reduce our debugging time and number of submissions, allowing us to dedicate more time toward improving accuracy. I think this would benefit not only us competition participants but also you as well.</p>",
          "rawMarkdown": "@evancalabrese . @rachitsaluja \nYes, as you say, it's important to specify it by default.\nOn the other hand, I have one request.\nIf possible, could you provide a few examples of multi-frame DICOM data?\nOf course, the pixel data can be provided as dummy data.\n\nSimply by sharing such data, I believe we can significantly reduce our debugging time and number of submissions, allowing us to dedicate more time toward improving accuracy. I think this would benefit not only us competition participants but also you as well.",
          "replies": [
            {
              "id": 3299382,
              "postDate": "2025-10-08T00:10:20.767Z",
              "content": "<p>There is multi-frame DICOM data in the training. Which you can find by search using the following- </p>\n<pre><code>ds.SOPClassUID.name  \n</code></pre>",
              "rawMarkdown": "There is multi-frame DICOM data in the training. Which you can find by search using the following- \n```\nds.SOPClassUID.name == \"Enhanced MR Image Storage\"\n```",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3301971,
      "postDate": "2025-10-14T17:02:01.433Z",
      "content": "<p>Hi there, I am trying to submit my notebook, but is not allowing me to d o it. </p>",
      "rawMarkdown": "Hi there, I am trying to submit my notebook, but is not allowing me to d o it. "
    },
    {
      "id": 3298393,
      "postDate": "2025-10-05T12:36:29.193Z",
      "content": "<p>It has truly been very useful to try to unify working methods to obtain images and that is already a great contribution to the work being carried out. Thank you very much.</p>",
      "rawMarkdown": "It has truly been very useful to try to unify working methods to obtain images and that is already a great contribution to the work being carried out. Thank you very much."
    },
    {
      "id": 3296901,
      "postDate": "2025-10-01T22:26:43.360Z",
      "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>  Thanks for your hard work</p>",
      "rawMarkdown": "@evancalabrese  Thanks for your hard work"
    }
  ],
  "comments": [
    {
      "id": 3299025,
      "author_name": "rsjx",
      "author_url": "",
      "post_date": "2025-10-07T01:58:50.393000",
      "content": "<p>It appears that some participants rely on DICOM attributes during inference, and for certain Multi-Frame DICOMs, these attributes may be missing, leading to invalid or failed inferences. In such cases, you may consider using the following default parameters as a shared function group:<br>\nspacing=(0.5, 0.5), origin=(0, 0, 0), orientation=[1, 0, 0, 0, 1, 0], slice_thickness = 5.0</p>\n<p>While I cannot confirm that this will improve performance, it may be worth trying.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3299031,
          "author_name": "shiba-inu",
          "author_url": "",
          "post_date": "2025-10-07T02:51:30.237000",
          "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> . <a href=\"https://www.kaggle.com/rachitsaluja\" target=\"_blank\">@rachitsaluja</a> <br>\nYes, as you say, it's important to specify it by default.<br>\nOn the other hand, I have one request.<br>\nIf possible, could you provide a few examples of multi-frame DICOM data?<br>\nOf course, the pixel data can be provided as dummy data.</p>\n<p>Simply by sharing such data, I believe we can significantly reduce our debugging time and number of submissions, allowing us to dedicate more time toward improving accuracy. I think this would benefit not only us competition participants but also you as well.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3299382,
              "author_name": "rsjx",
              "author_url": "",
              "post_date": "2025-10-08T00:10:20.767000",
              "content": "<p>There is multi-frame DICOM data in the training. Which you can find by search using the following- </p>\n<pre><code>ds.SOPClassUID.name  \n</code></pre>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3301971,
      "author_name": "AldoCamargo",
      "author_url": "",
      "post_date": "2025-10-14T17:02:01.433000",
      "content": "<p>Hi there, I am trying to submit my notebook, but is not allowing me to d o it. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3298393,
      "author_name": "Jose Luis Sabogal",
      "author_url": "",
      "post_date": "2025-10-05T12:36:29.193000",
      "content": "<p>It has truly been very useful to try to unify working methods to obtain images and that is already a great contribution to the work being carried out. Thank you very much.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3296901,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-10-01T22:26:43.360000",
      "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>  Thanks for your hard work</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3296805": "Hi all,\n\nAs many of you have pointed out, there are still a few remaining issues with the competition data. We have been working hard behind the scenes for the past few months to correct a variety of issues that have come up including several that required reprocessing and expert manual review of the data. We are happy that we fixed several of the bigger issues, but we know we haven't been able to fix them all. \n\nUnfortunately, kaggle's process for updating the competition dataset is extremely slow, taking ~1 week for minor updates such as modifying a single CSV. They have their reasons for this, and I can't comment on the merit of this approach. Given this, and the <2 week timeframe remaining in the challenge, **we will NOT be making any additional changes to the dataset**.\n\nThis project was very ambitious from the start. Not only is it one of the largest ever medical imaging competition datasets ever released, but it also has several other complexities including inhomogeneous multi-site data, nuances in DICOM format, and several different data types (image, categorical, location, segmentation etc). It also required hundreds of hours of expert manual annotation. We are proud of the dataset we have created, but we know it still has many flaws. We sincerely hope you understand.\n\n**I am pinning this thread for visibility and in the hopes that we can use it to:**\n**1) Have a single place for aggregating known remaining issues with the dataset.**\n**2) Sharing any code (if you are willing) that competitors can be use to patch, fix, or avoid any remaining dataset issues.**\n\nThanks, and good luck in the final 2 weeks!",
    "3299025": "It appears that some participants rely on DICOM attributes during inference, and for certain Multi-Frame DICOMs, these attributes may be missing, leading to invalid or failed inferences. In such cases, you may consider using the following default parameters as a shared function group:\nspacing=(0.5, 0.5), origin=(0, 0, 0), orientation=[1, 0, 0, 0, 1, 0], slice_thickness = 5.0\n\nWhile I cannot confirm that this will improve performance, it may be worth trying.",
    "3301971": "Hi there, I am trying to submit my notebook, but is not allowing me to d o it. ",
    "3298393": "It has truly been very useful to try to unify working methods to obtain images and that is already a great contribution to the work being carried out. Thank you very much.",
    "3296901": "@evancalabrese  Thanks for your hard work"
  }
}