{
  "id": 186846,
  "title": "2D vs 3D ConvNet?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/186846",
  "author_name": "Hoang Nguyen",
  "post_date": "2020-09-26T09:30:37.009000",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Hope you all are enjoying this competition! </p>\n<p>After viewing a couple of notebooks and discussions, I have some questions about model choices that I hope someone can give me suggestions - of course it's a competition so I'm not expecting a detailed configuration/architecture, but a little general guidance will be much appreciated!</p>\n<p>Particularly, I'm seeing some people preprocess their images as 2D (examples: <a href=\"https://www.kaggle.com/teeyee314/pulmonary-embolism-create-tfrecords\" target=\"_blank\">TFRecords preparation by </a><a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a>, <a href=\"https://www.kaggle.com/orkatz2/resnext-pulmonary-embolism-inference\" target=\"_blank\">RestNet-based inference by </a><a href=\"https://www.kaggle.com/orkatz2\" target=\"_blank\">@orkatz2</a>), while others use 3D (examples: <a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">this amazing preprocessing notebooks by </a><a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> ). If indeed there are options for both 2D and 3D images, my questions are:</p>\n<ol>\n<li>Because all labels are at the study-level, not image-level, I'm thinking a 3D-based ConvNet that uses 3D images (which is combined from all scans at study level, like shown in <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a>'s notebook above) will be better. However most of what I, as a newbie, have come across are presented on 2D images (AlexNet, GoogLeNet, VGG, ResNet, Xception, etc.). Are there similarly big names for 3D models, or are 2D models just very generalizable to 3D ones that those famous architecture can be applied directly on 3D images after some simple re-configuration?</li>\n<li>If it's the latter, What are the things that one should pay attention to? For example in this competition's dataset, I noticed that many scan in the same study doesn't show much information, such as one below. Feeding to the model this image, which has the same label with an image with lots of information, can create the kind of noise we don't want, I think? Is there a way to workaround that?</li>\n</ol>\n<p>Thanks everyone!</p>\n<p>P/s: Sorry, somehow I couldn't any images so I've attached my example image with the post.</p>",
  "messages": [
    {
      "id": 1027652,
      "postDate": "2020-09-26T09:30:37.010Z",
      "content": "<p>Hi everyone,</p>\n<p>Hope you all are enjoying this competition! </p>\n<p>After viewing a couple of notebooks and discussions, I have some questions about model choices that I hope someone can give me suggestions - of course it's a competition so I'm not expecting a detailed configuration/architecture, but a little general guidance will be much appreciated!</p>\n<p>Particularly, I'm seeing some people preprocess their images as 2D (examples: <a href=\"https://www.kaggle.com/teeyee314/pulmonary-embolism-create-tfrecords\" target=\"_blank\">TFRecords preparation by </a><a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a>, <a href=\"https://www.kaggle.com/orkatz2/resnext-pulmonary-embolism-inference\" target=\"_blank\">RestNet-based inference by </a><a href=\"https://www.kaggle.com/orkatz2\" target=\"_blank\">@orkatz2</a>), while others use 3D (examples: <a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">this amazing preprocessing notebooks by </a><a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> ). If indeed there are options for both 2D and 3D images, my questions are:</p>\n<ol>\n<li>Because all labels are at the study-level, not image-level, I'm thinking a 3D-based ConvNet that uses 3D images (which is combined from all scans at study level, like shown in <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a>'s notebook above) will be better. However most of what I, as a newbie, have come across are presented on 2D images (AlexNet, GoogLeNet, VGG, ResNet, Xception, etc.). Are there similarly big names for 3D models, or are 2D models just very generalizable to 3D ones that those famous architecture can be applied directly on 3D images after some simple re-configuration?</li>\n<li>If it's the latter, What are the things that one should pay attention to? For example in this competition's dataset, I noticed that many scan in the same study doesn't show much information, such as one below. Feeding to the model this image, which has the same label with an image with lots of information, can create the kind of noise we don't want, I think? Is there a way to workaround that?</li>\n</ol>\n<p>Thanks everyone!</p>\n<p>P/s: Sorry, somehow I couldn't any images so I've attached my example image with the post.</p>",
      "rawMarkdown": "Hi everyone,\n\nHope you all are enjoying this competition! \n\nAfter viewing a couple of notebooks and discussions, I have some questions about model choices that I hope someone can give me suggestions - of course it's a competition so I'm not expecting a detailed configuration/architecture, but a little general guidance will be much appreciated!\n\nParticularly, I'm seeing some people preprocess their images as 2D (examples: [TFRecords preparation by @teeyee314](https://www.kaggle.com/teeyee314/pulmonary-embolism-create-tfrecords), [RestNet-based inference by @orkatz2](https://www.kaggle.com/orkatz2/resnext-pulmonary-embolism-inference)), while others use 3D (examples: [this amazing preprocessing notebooks by @allunia ](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing)). If indeed there are options for both 2D and 3D images, my questions are:\n1. Because all labels are at the study-level, not image-level, I'm thinking a 3D-based ConvNet that uses 3D images (which is combined from all scans at study level, like shown in @allunia's notebook above) will be better. However most of what I, as a newbie, have come across are presented on 2D images (AlexNet, GoogLeNet, VGG, ResNet, Xception, etc.). Are there similarly big names for 3D models, or are 2D models just very generalizable to 3D ones that those famous architecture can be applied directly on 3D images after some simple re-configuration?\n2. If it's the latter, What are the things that one should pay attention to? For example in this competition's dataset, I noticed that many scan in the same study doesn't show much information, such as one below. Feeding to the model this image, which has the same label with an image with lots of information, can create the kind of noise we don't want, I think? Is there a way to workaround that?\n\n\nThanks everyone!\n\nP/s: Sorry, somehow I couldn't any images so I've attached my example image with the post.",
      "votes": 6
    },
    {
      "id": 1034447,
      "postDate": "2020-10-01T18:43:14.063Z",
      "content": "<p>Something you'll probably be interested in is an Inflated 3D CNN. This paper has something on it: <a href=\"https://arxiv.org/pdf/1705.07750.pdf\" target=\"_blank\">https://arxiv.org/pdf/1705.07750.pdf</a> and here's the related code: <a href=\"https://github.com/deepmind/kinetics-i3d\" target=\"_blank\">https://github.com/deepmind/kinetics-i3d</a>.</p>\n<p>While I'm not an expert, I remember reading in some video processing paper (not the above) that the authors took the weight kernels of a ResNet and \"inflated\" them to 3D by repeating the matrix along the new axis (e.g. 2D kernel of 3x3x1 to 3D kernel of 3x3x1n) with the weights scaled by 1/n. This seems to be something that the folks in video-processing research play with. My memory might be wrong though, I didn't look too much into it as it didn't pertain much to what I was looking for.</p>",
      "rawMarkdown": "Something you'll probably be interested in is an Inflated 3D CNN. This paper has something on it: https://arxiv.org/pdf/1705.07750.pdf and here's the related code: https://github.com/deepmind/kinetics-i3d.\n\nWhile I'm not an expert, I remember reading in some video processing paper (not the above) that the authors took the weight kernels of a ResNet and \"inflated\" them to 3D by repeating the matrix along the new axis (e.g. 2D kernel of 3x3x1 to 3D kernel of 3x3x1n) with the weights scaled by 1/n. This seems to be something that the folks in video-processing research play with. My memory might be wrong though, I didn't look too much into it as it didn't pertain much to what I was looking for.",
      "votes": 2
    },
    {
      "id": 1165409,
      "postDate": "2021-01-23T00:22:14.677Z",
      "content": "<p>Is there any expert working on video prediction here?<br>\nI am new and need some help with I3D model.</p>",
      "rawMarkdown": "Is there any expert working on video prediction here?\nI am new and need some help with I3D model."
    },
    {
      "id": 1033115,
      "postDate": "2020-09-30T17:11:41.367Z",
      "content": "<p>Bumping this as I'd be interested in that too :)</p>",
      "rawMarkdown": "Bumping this as I'd be interested in that too :)"
    },
    {
      "id": 1027664,
      "postDate": "2020-09-26T09:35:14.537Z",
      "content": "<p>Sorry, somehow I couldn't any images so I've attached my example image with the post.</p>",
      "rawMarkdown": "Sorry, somehow I couldn't any images so I've attached my example image with the post."
    },
    {
      "id": 1027660,
      "postDate": "2020-09-26T09:33:34.650Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1034447,
      "author_name": "Joseph Tan",
      "author_url": "",
      "post_date": "2020-10-01T18:43:14.063000",
      "content": "<p>Something you'll probably be interested in is an Inflated 3D CNN. This paper has something on it: <a href=\"https://arxiv.org/pdf/1705.07750.pdf\" target=\"_blank\">https://arxiv.org/pdf/1705.07750.pdf</a> and here's the related code: <a href=\"https://github.com/deepmind/kinetics-i3d\" target=\"_blank\">https://github.com/deepmind/kinetics-i3d</a>.</p>\n<p>While I'm not an expert, I remember reading in some video processing paper (not the above) that the authors took the weight kernels of a ResNet and \"inflated\" them to 3D by repeating the matrix along the new axis (e.g. 2D kernel of 3x3x1 to 3D kernel of 3x3x1n) with the weights scaled by 1/n. This seems to be something that the folks in video-processing research play with. My memory might be wrong though, I didn't look too much into it as it didn't pertain much to what I was looking for.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1165409,
      "author_name": "Victor Adewopo",
      "author_url": "",
      "post_date": "2021-01-23T00:22:14.677000",
      "content": "<p>Is there any expert working on video prediction here?<br>\nI am new and need some help with I3D model.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1033115,
      "author_name": "Alex Bader",
      "author_url": "",
      "post_date": "2020-09-30T17:11:41.367000",
      "content": "<p>Bumping this as I'd be interested in that too :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1027664,
      "author_name": "Hoang Nguyen",
      "author_url": "",
      "post_date": "2020-09-26T09:35:14.537000",
      "content": "<p>Sorry, somehow I couldn't any images so I've attached my example image with the post.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1027660,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-26T09:33:34.650000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1027652": "Hi everyone,\n\nHope you all are enjoying this competition! \n\nAfter viewing a couple of notebooks and discussions, I have some questions about model choices that I hope someone can give me suggestions - of course it's a competition so I'm not expecting a detailed configuration/architecture, but a little general guidance will be much appreciated!\n\nParticularly, I'm seeing some people preprocess their images as 2D (examples: [TFRecords preparation by @teeyee314](https://www.kaggle.com/teeyee314/pulmonary-embolism-create-tfrecords), [RestNet-based inference by @orkatz2](https://www.kaggle.com/orkatz2/resnext-pulmonary-embolism-inference)), while others use 3D (examples: [this amazing preprocessing notebooks by @allunia ](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing)). If indeed there are options for both 2D and 3D images, my questions are:\n1. Because all labels are at the study-level, not image-level, I'm thinking a 3D-based ConvNet that uses 3D images (which is combined from all scans at study level, like shown in @allunia's notebook above) will be better. However most of what I, as a newbie, have come across are presented on 2D images (AlexNet, GoogLeNet, VGG, ResNet, Xception, etc.). Are there similarly big names for 3D models, or are 2D models just very generalizable to 3D ones that those famous architecture can be applied directly on 3D images after some simple re-configuration?\n2. If it's the latter, What are the things that one should pay attention to? For example in this competition's dataset, I noticed that many scan in the same study doesn't show much information, such as one below. Feeding to the model this image, which has the same label with an image with lots of information, can create the kind of noise we don't want, I think? Is there a way to workaround that?\n\n\nThanks everyone!\n\nP/s: Sorry, somehow I couldn't any images so I've attached my example image with the post.",
    "1034447": "Something you'll probably be interested in is an Inflated 3D CNN. This paper has something on it: https://arxiv.org/pdf/1705.07750.pdf and here's the related code: https://github.com/deepmind/kinetics-i3d.\n\nWhile I'm not an expert, I remember reading in some video processing paper (not the above) that the authors took the weight kernels of a ResNet and \"inflated\" them to 3D by repeating the matrix along the new axis (e.g. 2D kernel of 3x3x1 to 3D kernel of 3x3x1n) with the weights scaled by 1/n. This seems to be something that the folks in video-processing research play with. My memory might be wrong though, I didn't look too much into it as it didn't pertain much to what I was looking for.",
    "1165409": "Is there any expert working on video prediction here?\nI am new and need some help with I3D model.",
    "1033115": "Bumping this as I'd be interested in that too :)",
    "1027664": "Sorry, somehow I couldn't any images so I've attached my example image with the post.",
    "1027660": ""
  }
}