{
  "id": 190879,
  "title": "MONAI 3D CNN Baseline",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/190879",
  "author_name": "Bo",
  "post_date": "2020-10-13T17:44:20.215000",
  "votes": 82,
  "comment_count": 30,
  "views": 0,
  "content": "<p>A challenge in this competition is how to handle the ~200 images in the same study, and make the 9 study-level predictions. So far, baselines have been using target means.</p>\n<p>One way to tackle this is to build 3-dimensional CNN models. For this purpose, I find <a href=\"https://github.com/Project-MONAI/MONAI\" target=\"_blank\">MONAI</a> library's 3D CNN models and the 3D augmentations easy to use.</p>\n<p>I'm sharing a MONAI 3D CNN baseline, hoping this can help more people get started in this competition.</p>\n<p>Training notebook (just a demo, since Kaggle notebook's hardware is limited):<br>\n<a href=\"https://www.kaggle.com/boliu0/monai-3d-cnn-training\" target=\"_blank\">https://www.kaggle.com/boliu0/monai-3d-cnn-training</a></p>\n<p>Inference notebook (with weights trained locally, LB = 0.295):<br>\n<a href=\"https://www.kaggle.com/boliu0/monai-3d-cnn-inference\" target=\"_blank\">https://www.kaggle.com/boliu0/monai-3d-cnn-inference</a></p>\n<p>Ackowlegements:</p>\n<ul>\n<li>I'm using Dr. Ian Pan's preprocessed <a href=\"https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256\" target=\"_blank\">dataset</a> </li>\n<li>I'm using OsciiArt's great <a href=\"https://www.kaggle.com/osciiart/baseline-with-no-image\" target=\"_blank\">baseline</a> for image-level predictions </li>\n</ul>",
  "messages": [
    {
      "id": 1048673,
      "postDate": "2020-10-13T17:44:20.217Z",
      "content": "<p>A challenge in this competition is how to handle the ~200 images in the same study, and make the 9 study-level predictions. So far, baselines have been using target means.</p>\n<p>One way to tackle this is to build 3-dimensional CNN models. For this purpose, I find <a href=\"https://github.com/Project-MONAI/MONAI\" target=\"_blank\">MONAI</a> library's 3D CNN models and the 3D augmentations easy to use.</p>\n<p>I'm sharing a MONAI 3D CNN baseline, hoping this can help more people get started in this competition.</p>\n<p>Training notebook (just a demo, since Kaggle notebook's hardware is limited):<br>\n<a href=\"https://www.kaggle.com/boliu0/monai-3d-cnn-training\" target=\"_blank\">https://www.kaggle.com/boliu0/monai-3d-cnn-training</a></p>\n<p>Inference notebook (with weights trained locally, LB = 0.295):<br>\n<a href=\"https://www.kaggle.com/boliu0/monai-3d-cnn-inference\" target=\"_blank\">https://www.kaggle.com/boliu0/monai-3d-cnn-inference</a></p>\n<p>Ackowlegements:</p>\n<ul>\n<li>I'm using Dr. Ian Pan's preprocessed <a href=\"https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256\" target=\"_blank\">dataset</a> </li>\n<li>I'm using OsciiArt's great <a href=\"https://www.kaggle.com/osciiart/baseline-with-no-image\" target=\"_blank\">baseline</a> for image-level predictions </li>\n</ul>",
      "rawMarkdown": "A challenge in this competition is how to handle the ~200 images in the same study, and make the 9 study-level predictions. So far, baselines have been using target means.\n\nOne way to tackle this is to build 3-dimensional CNN models. For this purpose, I find [MONAI](https://github.com/Project-MONAI/MONAI) library's 3D CNN models and the 3D augmentations easy to use.\n\nI'm sharing a MONAI 3D CNN baseline, hoping this can help more people get started in this competition.\n\nTraining notebook (just a demo, since Kaggle notebook's hardware is limited):\nhttps://www.kaggle.com/boliu0/monai-3d-cnn-training\n\nInference notebook (with weights trained locally, LB = 0.295):\nhttps://www.kaggle.com/boliu0/monai-3d-cnn-inference\n\nAckowlegements:\n- I'm using Dr. Ian Pan's preprocessed [dataset](https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256) \n- I'm using OsciiArt's great [baseline](https://www.kaggle.com/osciiart/baseline-with-no-image) for image-level predictions ",
      "votes": 82
    },
    {
      "id": 1048933,
      "postDate": "2020-10-14T01:04:24.277Z",
      "content": "<p>Extremely good notebook, well done! May I ask why you don't use image level predictions from the model, instead using osci's notebook? </p>",
      "rawMarkdown": "Extremely good notebook, well done! May I ask why you don't use image level predictions from the model, instead using osci's notebook? ",
      "votes": 1,
      "replies": [
        {
          "id": 1048995,
          "postDate": "2020-10-14T02:59:06.320Z",
          "content": "<p>I think if you stack 2d images to 3d images by patient, you cannot predict each image's image-level labels.</p>",
          "rawMarkdown": "I think if you stack 2d images to 3d images by patient, you cannot predict each image's image-level labels.",
          "votes": 3
        },
        {
          "id": 1048999,
          "postDate": "2020-10-14T03:07:18.503Z",
          "content": "<p>Yeah, sin is right. </p>\n<p>The 3D model is at study level. Each data point is a study, which includes 100 to 1000 images. The model predictions are 9 study-level labels, not at image level.</p>",
          "rawMarkdown": "Yeah, sin is right. \n\nThe 3D model is at study level. Each data point is a study, which includes 100 to 1000 images. The model predictions are 9 study-level labels, not at image level.",
          "votes": 5
        },
        {
          "id": 1049010,
          "postDate": "2020-10-14T03:12:53.270Z",
          "content": "<p>I'm working on Resnet3D. The goal is to predict image and study level together.</p>",
          "rawMarkdown": "I'm working on Resnet3D. The goal is to predict image and study level together.",
          "votes": 7
        },
        {
          "id": 1049041,
          "postDate": "2020-10-14T03:44:22.267Z",
          "content": "<p>Alright, thank you guys for your responses. I just realized I'm doing everything completely wrong then… Was doing 2d convolutional on image level features and getting scores of 0.31ish. Thanks again, saved me a ton of time. </p>",
          "rawMarkdown": "Alright, thank you guys for your responses. I just realized I'm doing everything completely wrong then... Was doing 2d convolutional on image level features and getting scores of 0.31ish. Thanks again, saved me a ton of time. ",
          "votes": 4
        },
        {
          "id": 1052706,
          "postDate": "2020-10-18T06:57:52.753Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1048809,
      "postDate": "2020-10-13T19:49:04.190Z",
      "content": "<p>Thanks for your contribution, well done 🙏 What hw is needed for training, \"since Kaggle notebook's hardware is limited)\" and what were the results for training for the checkpoint used here, epochs, loss, training time etc?</p>",
      "rawMarkdown": "Thanks for your contribution, well done 🙏 What hw is needed for training, \"since Kaggle notebook's hardware is limited)\" and what were the results for training for the checkpoint used here, epochs, loss, training time etc?",
      "votes": 1,
      "replies": [
        {
          "id": 1048812,
          "postDate": "2020-10-13T19:52:19.723Z",
          "content": "<blockquote>\n  <p>I use apex for mixed precision locally, but I don't use it here as I cannot install apex<br>\n  I use 2 GPUs (32G each) and 32 CPU cores locally with batch size 48, here I can only fit in batch size 8<br>\n  I use DEBUG mode in the notebook for illustration. In DEBUG mode I only train for 3 epochs using first 600 studies. Locally I train 20 epochs using all 7000 studies by setting DEBUG = False in the beginning<br>\n  In my local setup, valid loss can go down to 0.306</p>\n</blockquote>\n<p>Via Bo's notebook.</p>",
          "rawMarkdown": "> I use apex for mixed precision locally, but I don't use it here as I cannot install apex\nI use 2 GPUs (32G each) and 32 CPU cores locally with batch size 48, here I can only fit in batch size 8\nI use DEBUG mode in the notebook for illustration. In DEBUG mode I only train for 3 epochs using first 600 studies. Locally I train 20 epochs using all 7000 studies by setting DEBUG = False in the beginning\nIn my local setup, valid loss can go down to 0.306\n\nVia Bo's notebook."
        },
        {
          "id": 1048815,
          "postDate": "2020-10-13T19:59:11.733Z",
          "content": "<p>I used 2x32G GPUs for training, but I think a single card is also ok. You just need to reduce the batch size.</p>\n<p>The biggest bottle neck in Kaggle GPU docker is CPUs. There are only 2 cores. Data loading is too slow. </p>\n<p>I talked about the hardware, epochs, local valid loss in the training notebook. It took me about 6.5 hours to train 20 epochs locally.</p>",
          "rawMarkdown": "I used 2x32G GPUs for training, but I think a single card is also ok. You just need to reduce the batch size.\n\nThe biggest bottle neck in Kaggle GPU docker is CPUs. There are only 2 cores. Data loading is too slow. \n\nI talked about the hardware, epochs, local valid loss in the training notebook. It took me about 6.5 hours to train 20 epochs locally.",
          "votes": 2
        },
        {
          "id": 1048824,
          "postDate": "2020-10-13T20:11:33.663Z",
          "content": "<p>Sorry didn't saw the comments between the codes just looked after comments in the beginning. Btw, mixed precision is now included in torch 1.6, no need for installation, easy to implement.</p>",
          "rawMarkdown": "Sorry didn't saw the comments between the codes just looked after comments in the beginning. Btw, mixed precision is now included in torch 1.6, no need for installation, easy to implement."
        }
      ]
    },
    {
      "id": 1388447,
      "postDate": "2021-07-15T00:57:15.683Z",
      "content": "<p>Hi there, great notebook<br>\nis it possible to use this on .nii fMRI files or if anyone have any leads on this issue.<br>\ni'm quite new to this and it's for a school project.</p>",
      "rawMarkdown": "Hi there, great notebook\nis it possible to use this on .nii fMRI files or if anyone have any leads on this issue.\ni'm quite new to this and it's for a school project."
    },
    {
      "id": 1051921,
      "postDate": "2020-10-17T04:09:26.750Z",
      "content": "<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a><br>\n1)  in your model are you not predicting the image level pes.. i dint quite understand that<br>\n2) each CT has lot of of slices..  and different ones ,who would they got fit to gpu ?<br>\n3) during training we dont have to use weighted loss ?</p>",
      "rawMarkdown": "@boliu0\n1)  in your model are you not predicting the image level pes.. i dint quite understand that\n2) each CT has lot of of slices..  and different ones ,who would they got fit to gpu ?\n3) during training we dont have to use weighted loss ?",
      "replies": [
        {
          "id": 1051945,
          "postDate": "2020-10-17T05:06:55.213Z",
          "content": "<ol>\n<li>Correct. Each data sample of the 3D model is a study, so I'm predicting study level only.</li>\n<li>\"lots of slices\" are different z levels of the 3D array. They are all read into the 3D array, got resized, then are fed into the model and put on the GPU.</li>\n<li>You can, but you don't have to. Training loss and evaluation metric don't have to be the same.</li>\n</ol>",
          "rawMarkdown": "1. Correct. Each data sample of the 3D model is a study, so I'm predicting study level only.\n2. \"lots of slices\" are different z levels of the 3D array. They are all read into the 3D array, got resized, then are fed into the model and put on the GPU.\n3. You can, but you don't have to. Training loss and evaluation metric don't have to be the same.",
          "votes": 1
        },
        {
          "id": 1053568,
          "postDate": "2020-10-19T05:53:21.970Z",
          "content": "<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a>  thanks..<br>\nare you   using  metric  weighted loss or unweighted loss ?</p>\n<p>if metric loss how we compute qi wrt to batch</p>",
          "rawMarkdown": "@boliu0  thanks..\nare you   using  metric  weighted loss or unweighted loss ?\n\nif metric loss how we compute qi wrt to batch"
        },
        {
          "id": 1053597,
          "postDate": "2020-10-19T06:25:25.083Z",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> you can pre-compute it and then can add it to the Dataset class pretty easily. You may check <a href=\"https://www.kaggle.com/underwearfitting/rsna-pulmonary-meta-full-with-folds\" target=\"_blank\">https://www.kaggle.com/underwearfitting/rsna-pulmonary-meta-full-with-folds</a>, in which I pre-computed the q_i for each image. You may use the following function to calculate the competition metrics:<br>\n<code>def custom_loss(y_true_imag, y_pred_imag, q_image): return np.sum(- q_image * (y_true_imag*np.log(y_pred_imag) + (1-y_true_imag)*np.log(1-y_pred_imag))) / np.sum(q_image)</code></p>",
          "rawMarkdown": "@jaideepvalani you can pre-compute it and then can add it to the Dataset class pretty easily. You may check https://www.kaggle.com/underwearfitting/rsna-pulmonary-meta-full-with-folds, in which I pre-computed the q_i for each image. You may use the following function to calculate the competition metrics:\n`def custom_loss(y_true_imag, y_pred_imag, q_image): return np.sum(- q_image * (y_true_imag*np.log(y_pred_imag) + (1-y_true_imag)*np.log(1-y_pred_imag))) / np.sum(q_image)`"
        },
        {
          "id": 1053601,
          "postDate": "2020-10-19T06:48:51.563Z",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> thanks..<br>\nshould we not do logits here before  log ?<br>\nm still confused why competition page shows bce log loss insteadof with logits, during training if we use only bce log loss will model converge ?</p>",
          "rawMarkdown": "@underwearfitting thanks..\nshould we not do logits here before  log ?\nm still confused why competition page shows bce log loss insteadof with logits, during training if we use only bce log loss will model converge ?"
        },
        {
          "id": 1053607,
          "postDate": "2020-10-19T06:56:34.033Z",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> depends on your need. I am only tracking this loss but take grads based on BCELoss.</p>",
          "rawMarkdown": "@jaideepvalani depends on your need. I am only tracking this loss but take grads based on BCELoss."
        },
        {
          "id": 1053610,
          "postDate": "2020-10-19T07:01:07.647Z",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> if we use BCELoss without doing sigmoid any where like in model forward,we will get outputs as negatives value or close to zero values  in absence of Logits.. so our loss would always be coming high..  let me know if  i have not understood you.. in official documentation of bce loss they do sigmoid first</p>\n<pre><code>m = nn.Sigmoid()\n&gt;&gt;&gt; loss = nn.BCELoss()\n&gt;&gt;&gt; input = torch.randn(3, requires_grad=True)\n&gt;&gt;&gt; target = torch.empty(3).random_(2)\n&gt;&gt;&gt; output = loss(m(input), target)\n&gt;&gt;&gt; output.backward()\n</code></pre>",
          "rawMarkdown": "@underwearfitting if we use BCELoss without doing sigmoid any where like in model forward,we will get outputs as negatives value or close to zero values  in absence of Logits.. so our loss would always be coming high..  let me know if  i have not understood you.. in official documentation of bce loss they do sigmoid first\n\n```\nm = nn.Sigmoid()\n>>> loss = nn.BCELoss()\n>>> input = torch.randn(3, requires_grad=True)\n>>> target = torch.empty(3).random_(2)\n>>> output = loss(m(input), target)\n>>> output.backward()\n```"
        },
        {
          "id": 1053615,
          "postDate": "2020-10-19T07:06:07.807Z",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> I used <code>nn.BCEWithLogitsLoss()</code>……</p>",
          "rawMarkdown": "@jaideepvalani I used `nn.BCEWithLogitsLoss()`......",
          "votes": 1
        }
      ]
    },
    {
      "id": 1051547,
      "postDate": "2020-10-16T16:10:02.180Z",
      "content": "<p>Thank you for sharing your work with us. I just have a couple of questions.</p>\n<ol>\n<li>why did you choose 3D CNN instead of a sequential model as in last year's RSNA competition 1st solution?</li>\n<li>you used all 3 windows that Dr. Ian provided in his dataset. I was under the impression that we use windows to speed up convergence. And that the 3 windows have a lot of redundant information between them. So why didn't you just use the largest window (i.e. retaining the most information)? Do you think that using all 3 windows provided enough information to account for slowing down the convergence?</li>\n</ol>",
      "rawMarkdown": "Thank you for sharing your work with us. I just have a couple of questions.\n1. why did you choose 3D CNN instead of a sequential model as in last year's RSNA competition 1st solution?\n2. you used all 3 windows that Dr. Ian provided in his dataset. I was under the impression that we use windows to speed up convergence. And that the 3 windows have a lot of redundant information between them. So why didn't you just use the largest window (i.e. retaining the most information)? Do you think that using all 3 windows provided enough information to account for slowing down the convergence?",
      "replies": [
        {
          "id": 1051644,
          "postDate": "2020-10-16T17:41:08.383Z",
          "content": "<ol>\n<li>(1) Since RNN method was already shared, I wanted to share some new method. (2) I'm not too familiar with last yeas's data, but this year's data is perfect for 3D model since there are about 200 images per study, enough for the z axis. (3) 3D model is suitable for study-level predictions: one data sample in 3D model is one study</li>\n<li>I'm not a M.D., so I rely on Dr Ian Pan's window strategy without questioning him 😂  It's ok for the 3 channels to have redundant information, like your Red channel and Green channel in RGB images may have redundancy too. Also this is just a starter, feel free to experiment with other windowing strategies and improve on the baseline</li>\n</ol>",
          "rawMarkdown": "1. (1) Since RNN method was already shared, I wanted to share some new method. (2) I'm not too familiar with last yeas's data, but this year's data is perfect for 3D model since there are about 200 images per study, enough for the z axis. (3) 3D model is suitable for study-level predictions: one data sample in 3D model is one study\n2. I'm not a M.D., so I rely on Dr Ian Pan's window strategy without questioning him 😂  It's ok for the 3 channels to have redundant information, like your Red channel and Green channel in RGB images may have redundancy too. Also this is just a starter, feel free to experiment with other windowing strategies and improve on the baseline",
          "votes": 2
        },
        {
          "id": 1051694,
          "postDate": "2020-10-16T18:30:15.620Z",
          "content": "<p>Oh, I don't intend on improving the baseline. I'm not much familiar with pytorch. And my hardware is really really limited; since your baseline is so heavy for kaggle to train I have to find a more efficient method. Besides, I think that training two separate models for study-level and image-level prediction is overkill. Much of the information extracted with the image-level model can be reused for study level prediction.</p>\n<p>Thank you so much for replying btw.</p>",
          "rawMarkdown": "Oh, I don't intend on improving the baseline. I'm not much familiar with pytorch. And my hardware is really really limited; since your baseline is so heavy for kaggle to train I have to find a more efficient method. Besides, I think that training two separate models for study-level and image-level prediction is overkill. Much of the information extracted with the image-level model can be reused for study level prediction.\n\nThank you so much for replying btw."
        }
      ]
    },
    {
      "id": 1050338,
      "postDate": "2020-10-15T09:39:58.293Z",
      "content": "<p>that gives a basic clarity over the running the model</p>",
      "rawMarkdown": "that gives a basic clarity over the running the model"
    },
    {
      "id": 1049926,
      "postDate": "2020-10-14T22:24:32.743Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a>. Thank you for this nice baseline. I have a question regarding DataParallel mode. Do you use DistributedDataParallel or just DataParallel? I am having troubles running DataParallel with MONAI. Thank you.</p>",
      "rawMarkdown": "Hi, @boliu0. Thank you for this nice baseline. I have a question regarding DataParallel mode. Do you use DistributedDataParallel or just DataParallel? I am having troubles running DataParallel with MONAI. Thank you.",
      "replies": [
        {
          "id": 1049936,
          "postDate": "2020-10-14T22:39:31.943Z",
          "content": "<p>I used <code>DataParallel</code></p>\n<p>Just uncomment these 3 lines in training notebook then it will work</p>\n<pre><code>#    os.environ['CUDA_VISIBLE_DEVICES'] = '0,1' # specify GPUs locally\n...\n#     if len(os.environ['CUDA_VISIBLE_DEVICES'].split(',')) &gt; 1:\n#         model = nn.DataParallel(model)  \n</code></pre>",
          "rawMarkdown": "I used `DataParallel`\n\nJust uncomment these 3 lines in training notebook then it will work\n```\n#    os.environ['CUDA_VISIBLE_DEVICES'] = '0,1' # specify GPUs locally\n...\n#     if len(os.environ['CUDA_VISIBLE_DEVICES'].split(',')) > 1:\n#         model = nn.DataParallel(model)  \n```",
          "votes": 2
        },
        {
          "id": 1050288,
          "postDate": "2020-10-15T08:27:40.803Z",
          "content": "<p>Yeah, I am actually using these lines, but my machine does not want to run with multi-gpus. Thank you!</p>",
          "rawMarkdown": "Yeah, I am actually using these lines, but my machine does not want to run with multi-gpus. Thank you!"
        }
      ]
    },
    {
      "id": 1048738,
      "postDate": "2020-10-13T18:41:47.480Z",
      "content": "<p>Actually, I think that it's possible to train your notebook on Kaggle if you switch to TPU.</p>",
      "rawMarkdown": "Actually, I think that it's possible to train your notebook on Kaggle if you switch to TPU.",
      "replies": [
        {
          "id": 1048761,
          "postDate": "2020-10-13T19:16:57.587Z",
          "content": "<p>But unfortunately, it seems that MONAI augmentations are not supported by TPU.</p>",
          "rawMarkdown": "But unfortunately, it seems that MONAI augmentations are not supported by TPU.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1053468,
      "postDate": "2020-10-19T02:56:52.760Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1050983,
      "postDate": "2020-10-16T02:28:05.743Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1048933,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2020-10-14T01:04:24.277000",
      "content": "<p>Extremely good notebook, well done! May I ask why you don't use image level predictions from the model, instead using osci's notebook? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1048995,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-10-14T02:59:06.320000",
          "content": "<p>I think if you stack 2d images to 3d images by patient, you cannot predict each image's image-level labels.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1048999,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-10-14T03:07:18.503000",
          "content": "<p>Yeah, sin is right. </p>\n<p>The 3D model is at study level. Each data point is a study, which includes 100 to 1000 images. The model predictions are 9 study-level labels, not at image level.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1049010,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2020-10-14T03:12:53.270000",
          "content": "<p>I'm working on Resnet3D. The goal is to predict image and study level together.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1049041,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-14T03:44:22.267000",
          "content": "<p>Alright, thank you guys for your responses. I just realized I'm doing everything completely wrong then… Was doing 2d convolutional on image level features and getting scores of 0.31ish. Thanks again, saved me a ton of time. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1052706,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-18T06:57:52.753000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1048809,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-10-13T19:49:04.190000",
      "content": "<p>Thanks for your contribution, well done 🙏 What hw is needed for training, \"since Kaggle notebook's hardware is limited)\" and what were the results for training for the checkpoint used here, epochs, loss, training time etc?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1048812,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2020-10-13T19:52:19.723000",
          "content": "<blockquote>\n  <p>I use apex for mixed precision locally, but I don't use it here as I cannot install apex<br>\n  I use 2 GPUs (32G each) and 32 CPU cores locally with batch size 48, here I can only fit in batch size 8<br>\n  I use DEBUG mode in the notebook for illustration. In DEBUG mode I only train for 3 epochs using first 600 studies. Locally I train 20 epochs using all 7000 studies by setting DEBUG = False in the beginning<br>\n  In my local setup, valid loss can go down to 0.306</p>\n</blockquote>\n<p>Via Bo's notebook.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1048815,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-10-13T19:59:11.733000",
          "content": "<p>I used 2x32G GPUs for training, but I think a single card is also ok. You just need to reduce the batch size.</p>\n<p>The biggest bottle neck in Kaggle GPU docker is CPUs. There are only 2 cores. Data loading is too slow. </p>\n<p>I talked about the hardware, epochs, local valid loss in the training notebook. It took me about 6.5 hours to train 20 epochs locally.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1048824,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-10-13T20:11:33.663000",
          "content": "<p>Sorry didn't saw the comments between the codes just looked after comments in the beginning. Btw, mixed precision is now included in torch 1.6, no need for installation, easy to implement.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1388447,
      "author_name": "Mustapha KERROU",
      "author_url": "",
      "post_date": "2021-07-15T00:57:15.683000",
      "content": "<p>Hi there, great notebook<br>\nis it possible to use this on .nii fMRI files or if anyone have any leads on this issue.<br>\ni'm quite new to this and it's for a school project.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1051921,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-10-17T04:09:26.750000",
      "content": "<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a><br>\n1)  in your model are you not predicting the image level pes.. i dint quite understand that<br>\n2) each CT has lot of of slices..  and different ones ,who would they got fit to gpu ?<br>\n3) during training we dont have to use weighted loss ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1051945,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-10-17T05:06:55.213000",
          "content": "<ol>\n<li>Correct. Each data sample of the 3D model is a study, so I'm predicting study level only.</li>\n<li>\"lots of slices\" are different z levels of the 3D array. They are all read into the 3D array, got resized, then are fed into the model and put on the GPU.</li>\n<li>You can, but you don't have to. Training loss and evaluation metric don't have to be the same.</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1053568,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-19T05:53:21.970000",
          "content": "<p><a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a>  thanks..<br>\nare you   using  metric  weighted loss or unweighted loss ?</p>\n<p>if metric loss how we compute qi wrt to batch</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1053597,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-10-19T06:25:25.083000",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> you can pre-compute it and then can add it to the Dataset class pretty easily. You may check <a href=\"https://www.kaggle.com/underwearfitting/rsna-pulmonary-meta-full-with-folds\" target=\"_blank\">https://www.kaggle.com/underwearfitting/rsna-pulmonary-meta-full-with-folds</a>, in which I pre-computed the q_i for each image. You may use the following function to calculate the competition metrics:<br>\n<code>def custom_loss(y_true_imag, y_pred_imag, q_image): return np.sum(- q_image * (y_true_imag*np.log(y_pred_imag) + (1-y_true_imag)*np.log(1-y_pred_imag))) / np.sum(q_image)</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1053601,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-19T06:48:51.563000",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> thanks..<br>\nshould we not do logits here before  log ?<br>\nm still confused why competition page shows bce log loss insteadof with logits, during training if we use only bce log loss will model converge ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1053607,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-10-19T06:56:34.033000",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> depends on your need. I am only tracking this loss but take grads based on BCELoss.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1053610,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-19T07:01:07.647000",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> if we use BCELoss without doing sigmoid any where like in model forward,we will get outputs as negatives value or close to zero values  in absence of Logits.. so our loss would always be coming high..  let me know if  i have not understood you.. in official documentation of bce loss they do sigmoid first</p>\n<pre><code>m = nn.Sigmoid()\n&gt;&gt;&gt; loss = nn.BCELoss()\n&gt;&gt;&gt; input = torch.randn(3, requires_grad=True)\n&gt;&gt;&gt; target = torch.empty(3).random_(2)\n&gt;&gt;&gt; output = loss(m(input), target)\n&gt;&gt;&gt; output.backward()\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1053615,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-10-19T07:06:07.807000",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> I used <code>nn.BCEWithLogitsLoss()</code>……</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1051547,
      "author_name": "DarkCube",
      "author_url": "",
      "post_date": "2020-10-16T16:10:02.180000",
      "content": "<p>Thank you for sharing your work with us. I just have a couple of questions.</p>\n<ol>\n<li>why did you choose 3D CNN instead of a sequential model as in last year's RSNA competition 1st solution?</li>\n<li>you used all 3 windows that Dr. Ian provided in his dataset. I was under the impression that we use windows to speed up convergence. And that the 3 windows have a lot of redundant information between them. So why didn't you just use the largest window (i.e. retaining the most information)? Do you think that using all 3 windows provided enough information to account for slowing down the convergence?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 1051644,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-10-16T17:41:08.383000",
          "content": "<ol>\n<li>(1) Since RNN method was already shared, I wanted to share some new method. (2) I'm not too familiar with last yeas's data, but this year's data is perfect for 3D model since there are about 200 images per study, enough for the z axis. (3) 3D model is suitable for study-level predictions: one data sample in 3D model is one study</li>\n<li>I'm not a M.D., so I rely on Dr Ian Pan's window strategy without questioning him 😂  It's ok for the 3 channels to have redundant information, like your Red channel and Green channel in RGB images may have redundancy too. Also this is just a starter, feel free to experiment with other windowing strategies and improve on the baseline</li>\n</ol>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1051694,
          "author_name": "DarkCube",
          "author_url": "",
          "post_date": "2020-10-16T18:30:15.620000",
          "content": "<p>Oh, I don't intend on improving the baseline. I'm not much familiar with pytorch. And my hardware is really really limited; since your baseline is so heavy for kaggle to train I have to find a more efficient method. Besides, I think that training two separate models for study-level and image-level prediction is overkill. Much of the information extracted with the image-level model can be reused for study level prediction.</p>\n<p>Thank you so much for replying btw.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1050338,
      "author_name": "ayusharma18",
      "author_url": "",
      "post_date": "2020-10-15T09:39:58.293000",
      "content": "<p>that gives a basic clarity over the running the model</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1049926,
      "author_name": "john doea",
      "author_url": "",
      "post_date": "2020-10-14T22:24:32.743000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a>. Thank you for this nice baseline. I have a question regarding DataParallel mode. Do you use DistributedDataParallel or just DataParallel? I am having troubles running DataParallel with MONAI. Thank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1049936,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-10-14T22:39:31.943000",
          "content": "<p>I used <code>DataParallel</code></p>\n<p>Just uncomment these 3 lines in training notebook then it will work</p>\n<pre><code>#    os.environ['CUDA_VISIBLE_DEVICES'] = '0,1' # specify GPUs locally\n...\n#     if len(os.environ['CUDA_VISIBLE_DEVICES'].split(',')) &gt; 1:\n#         model = nn.DataParallel(model)  \n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1050288,
          "author_name": "john doea",
          "author_url": "",
          "post_date": "2020-10-15T08:27:40.803000",
          "content": "<p>Yeah, I am actually using these lines, but my machine does not want to run with multi-gpus. Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1048738,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2020-10-13T18:41:47.480000",
      "content": "<p>Actually, I think that it's possible to train your notebook on Kaggle if you switch to TPU.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1048761,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2020-10-13T19:16:57.587000",
          "content": "<p>But unfortunately, it seems that MONAI augmentations are not supported by TPU.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1053468,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-19T02:56:52.760000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1050983,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-16T02:28:05.743000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1048673": "A challenge in this competition is how to handle the ~200 images in the same study, and make the 9 study-level predictions. So far, baselines have been using target means.\n\nOne way to tackle this is to build 3-dimensional CNN models. For this purpose, I find [MONAI](https://github.com/Project-MONAI/MONAI) library's 3D CNN models and the 3D augmentations easy to use.\n\nI'm sharing a MONAI 3D CNN baseline, hoping this can help more people get started in this competition.\n\nTraining notebook (just a demo, since Kaggle notebook's hardware is limited):\nhttps://www.kaggle.com/boliu0/monai-3d-cnn-training\n\nInference notebook (with weights trained locally, LB = 0.295):\nhttps://www.kaggle.com/boliu0/monai-3d-cnn-inference\n\nAckowlegements:\n- I'm using Dr. Ian Pan's preprocessed [dataset](https://www.kaggle.com/vaillant/rsna-str-pe-detection-jpeg-256) \n- I'm using OsciiArt's great [baseline](https://www.kaggle.com/osciiart/baseline-with-no-image) for image-level predictions ",
    "1048933": "Extremely good notebook, well done! May I ask why you don't use image level predictions from the model, instead using osci's notebook? ",
    "1048809": "Thanks for your contribution, well done 🙏 What hw is needed for training, \"since Kaggle notebook's hardware is limited)\" and what were the results for training for the checkpoint used here, epochs, loss, training time etc?",
    "1388447": "Hi there, great notebook\nis it possible to use this on .nii fMRI files or if anyone have any leads on this issue.\ni'm quite new to this and it's for a school project.",
    "1051921": "@boliu0\n1)  in your model are you not predicting the image level pes.. i dint quite understand that\n2) each CT has lot of of slices..  and different ones ,who would they got fit to gpu ?\n3) during training we dont have to use weighted loss ?",
    "1051547": "Thank you for sharing your work with us. I just have a couple of questions.\n1. why did you choose 3D CNN instead of a sequential model as in last year's RSNA competition 1st solution?\n2. you used all 3 windows that Dr. Ian provided in his dataset. I was under the impression that we use windows to speed up convergence. And that the 3 windows have a lot of redundant information between them. So why didn't you just use the largest window (i.e. retaining the most information)? Do you think that using all 3 windows provided enough information to account for slowing down the convergence?",
    "1050338": "that gives a basic clarity over the running the model",
    "1049926": "Hi, @boliu0. Thank you for this nice baseline. I have a question regarding DataParallel mode. Do you use DistributedDataParallel or just DataParallel? I am having troubles running DataParallel with MONAI. Thank you.",
    "1048738": "Actually, I think that it's possible to train your notebook on Kaggle if you switch to TPU.",
    "1053468": "",
    "1050983": ""
  }
}