{
  "id": 416746,
  "title": "Pytorch Lightning baseline",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/416746",
  "author_name": "Egor Trushin",
  "post_date": "2023-06-12T22:55:48.427000",
  "votes": 27,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I am a <a href=\"https://www.pytorchlightning.ai/index.html\" target=\"_blank\">Pytorch Lightning</a> (PL) user and have created a PL baseline. Check <a href=\"https://www.kaggle.com/code/egortrushin/gr-icrgw-pytorch-lightning-baseline-unet-resnest\" target=\"_blank\">corresponding notebook</a>. LB can be improved by increasing the number of iterations, encoder size, image size, etc. Especially if you have the possibility to train locally. My current LB score was achieved with this pipeline.</p>\n<p>Good luck in the competition.</p>\n<p>Notebooks with updates:<br>\n<a href=\"https://www.kaggle.com/code/egortrushin/gr-icrgw-pl-pipeline-improved\" target=\"_blank\">[GR-ICRGW] PL Pipeline Improved</a><br>\n<a href=\"https://www.kaggle.com/code/egortrushin/gr-icrgw-training-with-4-folds?scriptVersionId=134148499\" target=\"_blank\">[GR-ICRGW] Training with 4 folds</a></p>",
  "messages": [
    {
      "id": 2299876,
      "postDate": "2023-06-12T22:55:48.427Z",
      "content": "<p>Hello everyone,</p>\n<p>I am a <a href=\"https://www.pytorchlightning.ai/index.html\" target=\"_blank\">Pytorch Lightning</a> (PL) user and have created a PL baseline. Check <a href=\"https://www.kaggle.com/code/egortrushin/gr-icrgw-pytorch-lightning-baseline-unet-resnest\" target=\"_blank\">corresponding notebook</a>. LB can be improved by increasing the number of iterations, encoder size, image size, etc. Especially if you have the possibility to train locally. My current LB score was achieved with this pipeline.</p>\n<p>Good luck in the competition.</p>\n<p>Notebooks with updates:<br>\n<a href=\"https://www.kaggle.com/code/egortrushin/gr-icrgw-pl-pipeline-improved\" target=\"_blank\">[GR-ICRGW] PL Pipeline Improved</a><br>\n<a href=\"https://www.kaggle.com/code/egortrushin/gr-icrgw-training-with-4-folds?scriptVersionId=134148499\" target=\"_blank\">[GR-ICRGW] Training with 4 folds</a></p>",
      "rawMarkdown": "Hello everyone,\n\nI am a [Pytorch Lightning](https://www.pytorchlightning.ai/index.html) (PL) user and have created a PL baseline. Check [corresponding notebook](https://www.kaggle.com/code/egortrushin/gr-icrgw-pytorch-lightning-baseline-unet-resnest). LB can be improved by increasing the number of iterations, encoder size, image size, etc. Especially if you have the possibility to train locally. My current LB score was achieved with this pipeline.\n\nGood luck in the competition.\n\nNotebooks with updates:\n[[GR-ICRGW] PL Pipeline Improved](https://www.kaggle.com/code/egortrushin/gr-icrgw-pl-pipeline-improved)\n[[GR-ICRGW] Training with 4 folds](https://www.kaggle.com/code/egortrushin/gr-icrgw-training-with-4-folds?scriptVersionId=134148499)",
      "votes": 26
    },
    {
      "id": 2313325,
      "postDate": "2023-06-22T15:30:41.383Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> <br>\nThank you for great notebook.</p>\n<p>I am using your notebook, but when I use \"efficientnet-b0\" I get the following error Do you know any solution?<br>\nThsnks.</p>\n<p>\"It looks like your LightningModule has parameters that were not used in producing the loss returned by training_step. If this is intentional, you must enable the detection of unused parameters in DDP, either by setting the string value <code>strategy='ddp_find_unused_parameters_true'</code> or by setting the flag in the strategy with <code>strategy=DDPStrategy(find_unused_parameters=True)</code>\"</p>",
      "rawMarkdown": "Hi @egortrushin \nThank you for great notebook.\n\nI am using your notebook, but when I use \"efficientnet-b0\" I get the following error Do you know any solution?\nThsnks.\n\n\"It looks like your LightningModule has parameters that were not used in producing the loss returned by training_step. If this is intentional, you must enable the detection of unused parameters in DDP, either by setting the string value `strategy='ddp_find_unused_parameters_true'` or by setting the flag in the strategy with `strategy=DDPStrategy(find_unused_parameters=True)`\"",
      "votes": 3,
      "replies": [
        {
          "id": 2313461,
          "postDate": "2023-06-22T17:33:29.840Z",
          "content": "<p>Same here, I couldn't find a solution until now, nothing relevant on PL docs or Google so far. </p>",
          "rawMarkdown": "Same here, I couldn't find a solution until now, nothing relevant on PL docs or Google so far. ",
          "votes": 3
        },
        {
          "id": 2314304,
          "postDate": "2023-06-23T10:33:40.933Z",
          "content": "<p>Hi. I have looked at this problem a bit and could not find a solution at the moment.</p>",
          "rawMarkdown": "Hi. I have looked at this problem a bit and could not find a solution at the moment.",
          "votes": 4
        },
        {
          "id": 2314513,
          "postDate": "2023-06-23T13:03:41.287Z",
          "content": "<p>It seems like a problem with ddp+efficientnet. If you use only one cuda device pipeline will work. </p>",
          "rawMarkdown": "It seems like a problem with ddp+efficientnet. If you use only one cuda device pipeline will work. ",
          "votes": 5,
          "replies": [
            {
              "id": 2314541,
              "postDate": "2023-06-23T13:18:51.227Z",
              "content": "<p>Yes Aleksandr, it's just a DDP+eff issue, I wonder if it works with just Pytorch, sometimes PL it's a bit buggy. Apparently, there're some unused parameters on the encoder, and DDP can't handle that and crash, <a href=\"https://github.com/Lightning-AI/lightning/issues/17212\" target=\"_blank\">https://github.com/Lightning-AI/lightning/issues/17212</a>. </p>",
              "rawMarkdown": "Yes Aleksandr, it's just a DDP+eff issue, I wonder if it works with just Pytorch, sometimes PL it's a bit buggy. Apparently, there're some unused parameters on the encoder, and DDP can't handle that and crash, https://github.com/Lightning-AI/lightning/issues/17212. ",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 2323710,
      "postDate": "2023-06-30T06:20:37.097Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> <br>\nthanks for great notebook, i checked the size of train dataset and find that the label seems not resized as same size as image.  would you please explain why label not resize as the image size, thanks advance</p>\n<p>img, label = next(iter(data_loader_train))<br>\nimg.shape, label.shape</p>\n<p>(torch.Size([32, 3, 384, 384]), torch.Size([32, 256, 256]))</p>",
      "rawMarkdown": "Hi @egortrushin \nthanks for great notebook, i checked the size of train dataset and find that the label seems not resized as same size as image.  would you please explain why label not resize as the image size, thanks advance\n\nimg, label = next(iter(data_loader_train))\nimg.shape, label.shape\n\n(torch.Size([32, 3, 384, 384]), torch.Size([32, 256, 256]))",
      "votes": 1,
      "replies": [
        {
          "id": 2324038,
          "postDate": "2023-06-30T10:40:46.887Z",
          "content": "<p>Hi. In this approach, a model predicts a mask of the size of the resized image, but then this mask is rescaled to 256x256 to be consistent with the 256x256 ground truth mask.</p>",
          "rawMarkdown": "Hi. In this approach, a model predicts a mask of the size of the resized image, but then this mask is rescaled to 256x256 to be consistent with the 256x256 ground truth mask.",
          "votes": 1,
          "replies": [
            {
              "id": 2324130,
              "postDate": "2023-06-30T12:04:20.947Z",
              "content": "<p>I saw  the code in training_step, thanks for your reply</p>\n<pre><code>     self != :\n        preds = torch(preds, size=, mode=)\n</code></pre>",
              "rawMarkdown": "I saw  the code in training_step, thanks for your reply\n\n        if self.config[\"image_size\"] != 256:\n            preds = torch.nn.functional.interpolate(preds, size=256, mode='bilinear')"
            }
          ]
        }
      ]
    },
    {
      "id": 2320710,
      "postDate": "2023-06-28T03:28:41.570Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> <br>\nThank you for this great notebook! What is the reason for changing the image size from the original 256 to other numbers like 384? Generally speaking, when and why increasing the image resolution may help?<br>\nThanks!</p>",
      "rawMarkdown": "Hi @egortrushin \nThank you for this great notebook! What is the reason for changing the image size from the original 256 to other numbers like 384? Generally speaking, when and why increasing the image resolution may help?\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 2321030,
          "postDate": "2023-06-28T08:32:51.453Z",
          "content": "<p>In the manuscript written by the organizers (see <a href=\"https://arxiv.org/pdf/2304.02122.pdf\" target=\"_blank\">OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI</a>) there is a stable improvement of the results with increasing image size. So it's worth trying to increase the image size.</p>",
          "rawMarkdown": "In the manuscript written by the organizers (see [OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI](https://arxiv.org/pdf/2304.02122.pdf)) there is a stable improvement of the results with increasing image size. So it's worth trying to increase the image size.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2304169,
      "postDate": "2023-06-15T18:21:22.543Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a>,<br>\nThank you for this amazing notebook.<br>\nHave you tried using TPUs with Pytorch and/or Pytorch Lightning? <br>\nI want to use TPUs in this competition but I'm unsure if they work well with Pytorch.<br>\nThanks.</p>",
      "rawMarkdown": "Hi @egortrushin,\nThank you for this amazing notebook.\nHave you tried using TPUs with Pytorch and/or Pytorch Lightning? \nI want to use TPUs in this competition but I'm unsure if they work well with Pytorch.\nThanks.",
      "votes": 2,
      "replies": [
        {
          "id": 2304256,
          "postDate": "2023-06-15T21:13:20.493Z",
          "content": "<p>Hi. Yes, I was trying to run Lightning code on TPU. Looks like it is not even possible to install torch-xla on the current Kaggle TPU environment.</p>",
          "rawMarkdown": "Hi. Yes, I was trying to run Lightning code on TPU. Looks like it is not even possible to install torch-xla on the current Kaggle TPU environment.",
          "replies": [
            {
              "id": 2304552,
              "postDate": "2023-06-16T04:45:05.807Z",
              "content": "<p>Oh 🙃. What to do then? <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> </p>",
              "rawMarkdown": "Oh 🙃. What to do then? @egortrushin ",
              "votes": 1
            },
            {
              "id": 2305345,
              "postDate": "2023-06-16T16:04:37.953Z",
              "content": "<p>Pytorch is bad on TPU anyway. It makes more sense to write Tensorflow code for TPU training. We already have some <a href=\"https://www.kaggle.com/code/egorfokin/contrails-tensorflow-train-submission-public\" target=\"_blank\">GPU Tensorflow baseline</a>. </p>",
              "rawMarkdown": "Pytorch is bad on TPU anyway. It makes more sense to write Tensorflow code for TPU training. We already have some [GPU Tensorflow baseline](https://www.kaggle.com/code/egorfokin/contrails-tensorflow-train-submission-public). "
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2313325,
      "author_name": "Never$",
      "author_url": "",
      "post_date": "2023-06-22T15:30:41.383000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> <br>\nThank you for great notebook.</p>\n<p>I am using your notebook, but when I use \"efficientnet-b0\" I get the following error Do you know any solution?<br>\nThsnks.</p>\n<p>\"It looks like your LightningModule has parameters that were not used in producing the loss returned by training_step. If this is intentional, you must enable the detection of unused parameters in DDP, either by setting the string value <code>strategy='ddp_find_unused_parameters_true'</code> or by setting the flag in the strategy with <code>strategy=DDPStrategy(find_unused_parameters=True)</code>\"</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2313461,
          "author_name": "Maximiliano Diaz Battan",
          "author_url": "",
          "post_date": "2023-06-22T17:33:29.840000",
          "content": "<p>Same here, I couldn't find a solution until now, nothing relevant on PL docs or Google so far. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2314304,
          "author_name": "Egor Trushin",
          "author_url": "",
          "post_date": "2023-06-23T10:33:40.933000",
          "content": "<p>Hi. I have looked at this problem a bit and could not find a solution at the moment.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2314513,
          "author_name": "Aleksandr Lavrikov",
          "author_url": "",
          "post_date": "2023-06-23T13:03:41.287000",
          "content": "<p>It seems like a problem with ddp+efficientnet. If you use only one cuda device pipeline will work. </p>",
          "votes": 5,
          "replies": [
            {
              "id": 2314541,
              "author_name": "Maximiliano Diaz Battan",
              "author_url": "",
              "post_date": "2023-06-23T13:18:51.227000",
              "content": "<p>Yes Aleksandr, it's just a DDP+eff issue, I wonder if it works with just Pytorch, sometimes PL it's a bit buggy. Apparently, there're some unused parameters on the encoder, and DDP can't handle that and crash, <a href=\"https://github.com/Lightning-AI/lightning/issues/17212\" target=\"_blank\">https://github.com/Lightning-AI/lightning/issues/17212</a>. </p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2323710,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2023-06-30T06:20:37.097000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> <br>\nthanks for great notebook, i checked the size of train dataset and find that the label seems not resized as same size as image.  would you please explain why label not resize as the image size, thanks advance</p>\n<p>img, label = next(iter(data_loader_train))<br>\nimg.shape, label.shape</p>\n<p>(torch.Size([32, 3, 384, 384]), torch.Size([32, 256, 256]))</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2324038,
          "author_name": "Egor Trushin",
          "author_url": "",
          "post_date": "2023-06-30T10:40:46.887000",
          "content": "<p>Hi. In this approach, a model predicts a mask of the size of the resized image, but then this mask is rescaled to 256x256 to be consistent with the 256x256 ground truth mask.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2324130,
              "author_name": "liuzhangzhen",
              "author_url": "",
              "post_date": "2023-06-30T12:04:20.947000",
              "content": "<p>I saw  the code in training_step, thanks for your reply</p>\n<pre><code>     self != :\n        preds = torch(preds, size=, mode=)\n</code></pre>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2320710,
      "author_name": "Zeyuan Hu",
      "author_url": "",
      "post_date": "2023-06-28T03:28:41.570000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> <br>\nThank you for this great notebook! What is the reason for changing the image size from the original 256 to other numbers like 384? Generally speaking, when and why increasing the image resolution may help?<br>\nThanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2321030,
          "author_name": "Egor Trushin",
          "author_url": "",
          "post_date": "2023-06-28T08:32:51.453000",
          "content": "<p>In the manuscript written by the organizers (see <a href=\"https://arxiv.org/pdf/2304.02122.pdf\" target=\"_blank\">OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI</a>) there is a stable improvement of the results with increasing image size. So it's worth trying to increase the image size.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2304169,
      "author_name": "Shashwat Raman",
      "author_url": "",
      "post_date": "2023-06-15T18:21:22.543000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a>,<br>\nThank you for this amazing notebook.<br>\nHave you tried using TPUs with Pytorch and/or Pytorch Lightning? <br>\nI want to use TPUs in this competition but I'm unsure if they work well with Pytorch.<br>\nThanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2304256,
          "author_name": "Egor Trushin",
          "author_url": "",
          "post_date": "2023-06-15T21:13:20.493000",
          "content": "<p>Hi. Yes, I was trying to run Lightning code on TPU. Looks like it is not even possible to install torch-xla on the current Kaggle TPU environment.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2304552,
              "author_name": "Shashwat Raman",
              "author_url": "",
              "post_date": "2023-06-16T04:45:05.807000",
              "content": "<p>Oh 🙃. What to do then? <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305345,
              "author_name": "Egor Trushin",
              "author_url": "",
              "post_date": "2023-06-16T16:04:37.953000",
              "content": "<p>Pytorch is bad on TPU anyway. It makes more sense to write Tensorflow code for TPU training. We already have some <a href=\"https://www.kaggle.com/code/egorfokin/contrails-tensorflow-train-submission-public\" target=\"_blank\">GPU Tensorflow baseline</a>. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2299876": "Hello everyone,\n\nI am a [Pytorch Lightning](https://www.pytorchlightning.ai/index.html) (PL) user and have created a PL baseline. Check [corresponding notebook](https://www.kaggle.com/code/egortrushin/gr-icrgw-pytorch-lightning-baseline-unet-resnest). LB can be improved by increasing the number of iterations, encoder size, image size, etc. Especially if you have the possibility to train locally. My current LB score was achieved with this pipeline.\n\nGood luck in the competition.\n\nNotebooks with updates:\n[[GR-ICRGW] PL Pipeline Improved](https://www.kaggle.com/code/egortrushin/gr-icrgw-pl-pipeline-improved)\n[[GR-ICRGW] Training with 4 folds](https://www.kaggle.com/code/egortrushin/gr-icrgw-training-with-4-folds?scriptVersionId=134148499)",
    "2313325": "Hi @egortrushin \nThank you for great notebook.\n\nI am using your notebook, but when I use \"efficientnet-b0\" I get the following error Do you know any solution?\nThsnks.\n\n\"It looks like your LightningModule has parameters that were not used in producing the loss returned by training_step. If this is intentional, you must enable the detection of unused parameters in DDP, either by setting the string value `strategy='ddp_find_unused_parameters_true'` or by setting the flag in the strategy with `strategy=DDPStrategy(find_unused_parameters=True)`\"",
    "2323710": "Hi @egortrushin \nthanks for great notebook, i checked the size of train dataset and find that the label seems not resized as same size as image.  would you please explain why label not resize as the image size, thanks advance\n\nimg, label = next(iter(data_loader_train))\nimg.shape, label.shape\n\n(torch.Size([32, 3, 384, 384]), torch.Size([32, 256, 256]))",
    "2320710": "Hi @egortrushin \nThank you for this great notebook! What is the reason for changing the image size from the original 256 to other numbers like 384? Generally speaking, when and why increasing the image resolution may help?\nThanks!",
    "2304169": "Hi @egortrushin,\nThank you for this amazing notebook.\nHave you tried using TPUs with Pytorch and/or Pytorch Lightning? \nI want to use TPUs in this competition but I'm unsure if they work well with Pytorch.\nThanks."
  }
}