{
  "id": 354451,
  "title": "patient_overall estimates",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/354451",
  "author_name": "Samuel Cortinhas",
  "post_date": "2022-09-22T10:30:51.450000",
  "votes": 33,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I wanted to open a discussion on how to best deal with predicting the most important component of the competition loss, <strong>patient_overall</strong>.</p>\n<p>A big advantage of <strong>3D models</strong> is that you can make this a training label and let the model figure out the best predictions. I think most people are moving towards <strong>2D models</strong> though as they are performing better, which makes this not possible anymore.</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> showed us how to start; by <strong>assuming independence</strong> of the vertebrae C1, …, C7. </p>\n<p><img src=\"https://i.postimg.cc/0NqQyvTK/indep.png\"></p>\n<p>This works surprisingly (at least to me) well, although I find that it tends to <strong>slightly overestimate</strong> the probability of patient_overall. We can see this by applying a 'squishing' function, which either pushes probabilities closer to the tails 0, 1 or to the middle 0.5, by varying the parameter beta.</p>\n<p><img src=\"https://i.postimg.cc/W3YSv5dz/equ.png\"></p>\n<p><img src=\"https://i.postimg.cc/rF7fYZ4w/beta-new.png\"></p>\n<p>And if you plot beta against a local CV, you get a curve looking like this. (note: this is a different 2D model to Vladimir's)</p>\n<p><img src=\"https://i.postimg.cc/FH7YCmS2/comp.png\"></p>\n<p>This shows you that you can get a slight improvement by squishing your probabilities closer to 0.5 (i.e. making them less extreme)</p>\n<hr>\n<p>This had me thinking of another way to estimate patient_overall. The exact formula including dependence is quite messy, it looks like this </p>\n<p><img src=\"https://i.postimg.cc/qqD3g7nx/incexc.png\"></p>\n<p>I tried some different approximations, e.g. truncating this series to only include interactions of up to 2 sets and estimating conditional probabilities from the training set. Long story short it doesn't work because you end up with lots of negative probabilities and it is just a bad approximation.</p>\n<hr>\n<p>So my question is has anybody else thought about this? Is there a way to improve on the Vladimir's formula by not assuming independence? </p>",
  "messages": [
    {
      "id": 1950433,
      "postDate": "2022-09-22T10:30:51.450Z",
      "content": "<p>I wanted to open a discussion on how to best deal with predicting the most important component of the competition loss, <strong>patient_overall</strong>.</p>\n<p>A big advantage of <strong>3D models</strong> is that you can make this a training label and let the model figure out the best predictions. I think most people are moving towards <strong>2D models</strong> though as they are performing better, which makes this not possible anymore.</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> showed us how to start; by <strong>assuming independence</strong> of the vertebrae C1, …, C7. </p>\n<p><img src=\"https://i.postimg.cc/0NqQyvTK/indep.png\"></p>\n<p>This works surprisingly (at least to me) well, although I find that it tends to <strong>slightly overestimate</strong> the probability of patient_overall. We can see this by applying a 'squishing' function, which either pushes probabilities closer to the tails 0, 1 or to the middle 0.5, by varying the parameter beta.</p>\n<p><img src=\"https://i.postimg.cc/W3YSv5dz/equ.png\"></p>\n<p><img src=\"https://i.postimg.cc/rF7fYZ4w/beta-new.png\"></p>\n<p>And if you plot beta against a local CV, you get a curve looking like this. (note: this is a different 2D model to Vladimir's)</p>\n<p><img src=\"https://i.postimg.cc/FH7YCmS2/comp.png\"></p>\n<p>This shows you that you can get a slight improvement by squishing your probabilities closer to 0.5 (i.e. making them less extreme)</p>\n<hr>\n<p>This had me thinking of another way to estimate patient_overall. The exact formula including dependence is quite messy, it looks like this </p>\n<p><img src=\"https://i.postimg.cc/qqD3g7nx/incexc.png\"></p>\n<p>I tried some different approximations, e.g. truncating this series to only include interactions of up to 2 sets and estimating conditional probabilities from the training set. Long story short it doesn't work because you end up with lots of negative probabilities and it is just a bad approximation.</p>\n<hr>\n<p>So my question is has anybody else thought about this? Is there a way to improve on the Vladimir's formula by not assuming independence? </p>",
      "rawMarkdown": "I wanted to open a discussion on how to best deal with predicting the most important component of the competition loss, **patient_overall**.\n\nA big advantage of **3D models** is that you can make this a training label and let the model figure out the best predictions. I think most people are moving towards **2D models** though as they are performing better, which makes this not possible anymore.\n\n<hr>\n\n@vslaykovsky showed us how to start; by **assuming independence** of the vertebrae C1, ..., C7. \n\n<img src='https://i.postimg.cc/0NqQyvTK/indep.png' width=250>\n\nThis works surprisingly (at least to me) well, although I find that it tends to **slightly overestimate** the probability of patient_overall. We can see this by applying a 'squishing' function, which either pushes probabilities closer to the tails 0, 1 or to the middle 0.5, by varying the parameter beta.\n\n<img src='https://i.postimg.cc/W3YSv5dz/equ.png' width=125>\n\n<img src='https://i.postimg.cc/rF7fYZ4w/beta-new.png' width=600>\n\nAnd if you plot beta against a local CV, you get a curve looking like this. (note: this is a different 2D model to Vladimir's)\n\n<img src='https://i.postimg.cc/FH7YCmS2/comp.png' width=600>\n\nThis shows you that you can get a slight improvement by squishing your probabilities closer to 0.5 (i.e. making them less extreme)\n\n<hr>\n\nThis had me thinking of another way to estimate patient_overall. The exact formula including dependence is quite messy, it looks like this \n\n<img src='https://i.postimg.cc/qqD3g7nx/incexc.png' width=700>\n\nI tried some different approximations, e.g. truncating this series to only include interactions of up to 2 sets and estimating conditional probabilities from the training set. Long story short it doesn't work because you end up with lots of negative probabilities and it is just a bad approximation.\n\n<hr>\n\nSo my question is has anybody else thought about this? Is there a way to improve on the Vladimir's formula by not assuming independence? ",
      "votes": 33
    },
    {
      "id": 1958427,
      "postDate": "2022-09-27T12:41:57.507Z",
      "content": "<p>I think 2D based models are effective because there is an implicit effect of increasing the number of samples. </p>\n<p>I'm working on a hybrid approcha to use CNNs trained using <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> 's approach as a feature extractor to pre-compute features(1280 for each image) and train a Transformer-based model to aggregate information from the 2D features of all the slices when predicting fractures for each patient. </p>",
      "rawMarkdown": "I think 2D based models are effective because there is an implicit effect of increasing the number of samples. \n\nI'm working on a hybrid approcha to use CNNs trained using @vslaykovsky 's approach as a feature extractor to pre-compute features(1280 for each image) and train a Transformer-based model to aggregate information from the 2D features of all the slices when predicting fractures for each patient. ",
      "votes": 2
    },
    {
      "id": 1989602,
      "postDate": "2022-10-16T05:05:44.430Z",
      "content": "<p>Interesting topic, I like probabilities. I think 3D models are not so bad as they do not use pretrained models. They have to learn from scratch. But 3D models do not see all vertebrae and could be better to recognize one or two vertebrae but not 7. It depends on the volume shape that must not be a cube. Recognize the vertebra is not only useful for patient_overall, but for all the predictions.</p>",
      "rawMarkdown": "Interesting topic, I like probabilities. I think 3D models are not so bad as they do not use pretrained models. They have to learn from scratch. But 3D models do not see all vertebrae and could be better to recognize one or two vertebrae but not 7. It depends on the volume shape that must not be a cube. Recognize the vertebra is not only useful for patient_overall, but for all the predictions."
    }
  ],
  "comments": [
    {
      "id": 1958427,
      "author_name": "SieunPark",
      "author_url": "",
      "post_date": "2022-09-27T12:41:57.507000",
      "content": "<p>I think 2D based models are effective because there is an implicit effect of increasing the number of samples. </p>\n<p>I'm working on a hybrid approcha to use CNNs trained using <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> 's approach as a feature extractor to pre-compute features(1280 for each image) and train a Transformer-based model to aggregate information from the 2D features of all the slices when predicting fractures for each patient. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1989602,
      "author_name": "Pierre Tisseur",
      "author_url": "",
      "post_date": "2022-10-16T05:05:44.430000",
      "content": "<p>Interesting topic, I like probabilities. I think 3D models are not so bad as they do not use pretrained models. They have to learn from scratch. But 3D models do not see all vertebrae and could be better to recognize one or two vertebrae but not 7. It depends on the volume shape that must not be a cube. Recognize the vertebra is not only useful for patient_overall, but for all the predictions.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1950433": "I wanted to open a discussion on how to best deal with predicting the most important component of the competition loss, **patient_overall**.\n\nA big advantage of **3D models** is that you can make this a training label and let the model figure out the best predictions. I think most people are moving towards **2D models** though as they are performing better, which makes this not possible anymore.\n\n<hr>\n\n@vslaykovsky showed us how to start; by **assuming independence** of the vertebrae C1, ..., C7. \n\n<img src='https://i.postimg.cc/0NqQyvTK/indep.png' width=250>\n\nThis works surprisingly (at least to me) well, although I find that it tends to **slightly overestimate** the probability of patient_overall. We can see this by applying a 'squishing' function, which either pushes probabilities closer to the tails 0, 1 or to the middle 0.5, by varying the parameter beta.\n\n<img src='https://i.postimg.cc/W3YSv5dz/equ.png' width=125>\n\n<img src='https://i.postimg.cc/rF7fYZ4w/beta-new.png' width=600>\n\nAnd if you plot beta against a local CV, you get a curve looking like this. (note: this is a different 2D model to Vladimir's)\n\n<img src='https://i.postimg.cc/FH7YCmS2/comp.png' width=600>\n\nThis shows you that you can get a slight improvement by squishing your probabilities closer to 0.5 (i.e. making them less extreme)\n\n<hr>\n\nThis had me thinking of another way to estimate patient_overall. The exact formula including dependence is quite messy, it looks like this \n\n<img src='https://i.postimg.cc/qqD3g7nx/incexc.png' width=700>\n\nI tried some different approximations, e.g. truncating this series to only include interactions of up to 2 sets and estimating conditional probabilities from the training set. Long story short it doesn't work because you end up with lots of negative probabilities and it is just a bad approximation.\n\n<hr>\n\nSo my question is has anybody else thought about this? Is there a way to improve on the Vladimir's formula by not assuming independence? ",
    "1958427": "I think 2D based models are effective because there is an implicit effect of increasing the number of samples. \n\nI'm working on a hybrid approcha to use CNNs trained using @vslaykovsky 's approach as a feature extractor to pre-compute features(1280 for each image) and train a Transformer-based model to aggregate information from the 2D features of all the slices when predicting fractures for each patient. ",
    "1989602": "Interesting topic, I like probabilities. I think 3D models are not so bad as they do not use pretrained models. They have to learn from scratch. But 3D models do not see all vertebrae and could be better to recognize one or two vertebrae but not 7. It depends on the volume shape that must not be a cube. Recognize the vertebra is not only useful for patient_overall, but for all the predictions."
  }
}