{
  "id": 344342,
  "title": "Identifying vertebrae C1-C7 in each image",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/344342",
  "author_name": "Samuel Cortinhas",
  "post_date": "2022-08-14T22:53:17.416000",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I haven't seen much progress on <strong>finding which vertebrae is in each image</strong> yet so I have just released my <a href=\"https://www.kaggle.com/code/samuelcortinhas/extracting-vertebrae-c1-c7\" target=\"_blank\">NOTEBOOK</a>. The idea is that we can identify the targets from the <strong>unique values in the segmentation masks</strong>. I have collected this info and stored it in this <a href=\"https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata/code\" target=\"_blank\">DATASET</a>. </p>\n<p><img src=\"https://i.postimg.cc/sD4ZyMGN/sample-df.png\"></p>\n<p>The <strong>next step</strong> would be to <strong>build a supervised model</strong> (using metadata/images) to predict which vertebrae is in each image for <strong>all the other patients</strong> in the train (&amp; test) set which <strong>don't have segmentation masks</strong>. The <strong>challenging</strong> part is <strong>preserving the monotonicity</strong> (C1-&gt;C2-&gt;etc). I haven't figured this out yet so I'm sharing this resource to hopefully kickstart some progress. </p>",
  "messages": [
    {
      "id": 1898870,
      "postDate": "2022-08-14T22:53:17.417Z",
      "content": "<p>I haven't seen much progress on <strong>finding which vertebrae is in each image</strong> yet so I have just released my <a href=\"https://www.kaggle.com/code/samuelcortinhas/extracting-vertebrae-c1-c7\" target=\"_blank\">NOTEBOOK</a>. The idea is that we can identify the targets from the <strong>unique values in the segmentation masks</strong>. I have collected this info and stored it in this <a href=\"https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata/code\" target=\"_blank\">DATASET</a>. </p>\n<p><img src=\"https://i.postimg.cc/sD4ZyMGN/sample-df.png\"></p>\n<p>The <strong>next step</strong> would be to <strong>build a supervised model</strong> (using metadata/images) to predict which vertebrae is in each image for <strong>all the other patients</strong> in the train (&amp; test) set which <strong>don't have segmentation masks</strong>. The <strong>challenging</strong> part is <strong>preserving the monotonicity</strong> (C1-&gt;C2-&gt;etc). I haven't figured this out yet so I'm sharing this resource to hopefully kickstart some progress. </p>",
      "rawMarkdown": "I haven't seen much progress on **finding which vertebrae is in each image** yet so I have just released my [NOTEBOOK](https://www.kaggle.com/code/samuelcortinhas/extracting-vertebrae-c1-c7). The idea is that we can identify the targets from the **unique values in the segmentation masks**. I have collected this info and stored it in this [DATASET](https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata/code). \n\n<img src='https://i.postimg.cc/sD4ZyMGN/sample-df.png' width=400>\n\nThe **next step** would be to **build a supervised model** (using metadata/images) to predict which vertebrae is in each image for **all the other patients** in the train (& test) set which **don't have segmentation masks**. The **challenging** part is **preserving the monotonicity** (C1->C2->etc). I haven't figured this out yet so I'm sharing this resource to hopefully kickstart some progress. ",
      "votes": 7
    },
    {
      "id": 1900969,
      "postDate": "2022-08-16T11:42:32.963Z",
      "content": "<p>Excellent work and great initiative in sharing your progress!<br>\nDid you try to train a model end2end with this and submit it? Is it improving it's performance? <br>\nAgain, nice work! </p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Excellent work and great initiative in sharing your progress!\nDid you try to train a model end2end with this and submit it? Is it improving it's performance? \nAgain, nice work! \n\nThe Devastator.\n",
      "votes": 1,
      "replies": [
        {
          "id": 1901595,
          "postDate": "2022-08-16T19:16:39.467Z",
          "content": "<p>I am in the process of doing this. Will let you how it goes.</p>",
          "rawMarkdown": "I am in the process of doing this. Will let you how it goes."
        }
      ]
    },
    {
      "id": 1898913,
      "postDate": "2022-08-14T23:57:09.450Z",
      "content": "<p>A simple enough supervised model gets like 85% accuracy (in my tests), won't making it 2.5d solve your problem with monotonicity most of the time, if that does not work, you can always just make a 3d model.</p>",
      "rawMarkdown": "A simple enough supervised model gets like 85% accuracy (in my tests), won't making it 2.5d solve your problem with monotonicity most of the time, if that does not work, you can always just make a 3d model.",
      "votes": 1,
      "replies": [
        {
          "id": 1899530,
          "postDate": "2022-08-15T10:27:22.507Z",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> thanks for the suggestion. I trained a random forest classifier and got 95% accuracy. The code has been added it to my notebook linked above. </p>",
          "rawMarkdown": "@harshitsheoran thanks for the suggestion. I trained a random forest classifier and got 95% accuracy. The code has been added it to my notebook linked above. "
        }
      ]
    },
    {
      "id": 1904524,
      "postDate": "2022-08-18T08:52:19.253Z",
      "content": "<p>Big Data Leak! Not good result if the leak is solved, leak is in train_test_split, you are not splitting it by StudyInstanceUID, instead taking really similar values in train and validation</p>",
      "rawMarkdown": "Big Data Leak! Not good result if the leak is solved, leak is in train_test_split, you are not splitting it by StudyInstanceUID, instead taking really similar values in train and validation",
      "votes": 2,
      "replies": [
        {
          "id": 1913262,
          "postDate": "2022-08-25T09:03:43.850Z",
          "content": "<p>Thank you for spotting this! I've fixed it now.</p>",
          "rawMarkdown": "Thank you for spotting this! I've fixed it now."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1900969,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-08-16T11:42:32.963000",
      "content": "<p>Excellent work and great initiative in sharing your progress!<br>\nDid you try to train a model end2end with this and submit it? Is it improving it's performance? <br>\nAgain, nice work! </p>\n<p>The Devastator.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1901595,
          "author_name": "Samuel Cortinhas",
          "author_url": "",
          "post_date": "2022-08-16T19:16:39.467000",
          "content": "<p>I am in the process of doing this. Will let you how it goes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1898913,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-08-14T23:57:09.450000",
      "content": "<p>A simple enough supervised model gets like 85% accuracy (in my tests), won't making it 2.5d solve your problem with monotonicity most of the time, if that does not work, you can always just make a 3d model.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1899530,
          "author_name": "Samuel Cortinhas",
          "author_url": "",
          "post_date": "2022-08-15T10:27:22.507000",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> thanks for the suggestion. I trained a random forest classifier and got 95% accuracy. The code has been added it to my notebook linked above. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1904524,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-08-18T08:52:19.253000",
      "content": "<p>Big Data Leak! Not good result if the leak is solved, leak is in train_test_split, you are not splitting it by StudyInstanceUID, instead taking really similar values in train and validation</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1913262,
          "author_name": "Samuel Cortinhas",
          "author_url": "",
          "post_date": "2022-08-25T09:03:43.850000",
          "content": "<p>Thank you for spotting this! I've fixed it now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1898870": "I haven't seen much progress on **finding which vertebrae is in each image** yet so I have just released my [NOTEBOOK](https://www.kaggle.com/code/samuelcortinhas/extracting-vertebrae-c1-c7). The idea is that we can identify the targets from the **unique values in the segmentation masks**. I have collected this info and stored it in this [DATASET](https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata/code). \n\n<img src='https://i.postimg.cc/sD4ZyMGN/sample-df.png' width=400>\n\nThe **next step** would be to **build a supervised model** (using metadata/images) to predict which vertebrae is in each image for **all the other patients** in the train (& test) set which **don't have segmentation masks**. The **challenging** part is **preserving the monotonicity** (C1->C2->etc). I haven't figured this out yet so I'm sharing this resource to hopefully kickstart some progress. ",
    "1900969": "Excellent work and great initiative in sharing your progress!\nDid you try to train a model end2end with this and submit it? Is it improving it's performance? \nAgain, nice work! \n\nThe Devastator.\n",
    "1898913": "A simple enough supervised model gets like 85% accuracy (in my tests), won't making it 2.5d solve your problem with monotonicity most of the time, if that does not work, you can always just make a 3d model.",
    "1904524": "Big Data Leak! Not good result if the leak is solved, leak is in train_test_split, you are not splitting it by StudyInstanceUID, instead taking really similar values in train and validation"
  }
}