{
  "id": 370134,
  "title": "How to train faster?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/370134",
  "author_name": "Myo Min Htet(wnp)🇲🇲",
  "post_date": "2022-12-03T09:00:48.847000",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Training is taking too much time. It takes 1 hour to get a 15% of one epoch and got 13% accuracy and train loss is 0.2.  I don't convert the image into png format. I use dicom images.</p>\n<p>I don't have external GPU. I use GPU from Kaggle.</p>",
  "messages": [
    {
      "id": 2053424,
      "postDate": "2022-12-03T09:00:48.847Z",
      "content": "<p>Training is taking too much time. It takes 1 hour to get a 15% of one epoch and got 13% accuracy and train loss is 0.2.  I don't convert the image into png format. I use dicom images.</p>\n<p>I don't have external GPU. I use GPU from Kaggle.</p>",
      "rawMarkdown": "Training is taking too much time. It takes 1 hour to get a 15% of one epoch and got 13% accuracy and train loss is 0.2.  I don't convert the image into png format. I use dicom images.\n\nI don't have external GPU. I use GPU from Kaggle.",
      "votes": 5
    },
    {
      "id": 2053480,
      "postDate": "2022-12-03T10:19:44.317Z",
      "content": "<p>Now is a beggining of cometition - time to establish fast experimenting pipeline. My current attitude:</p>\n<ul>\n<li>make images smaller (256x256)</li>\n<li>chose lighter/smaller model </li>\n<li>decrease amout of data (now I experimenting using only 500 images)</li>\n<li>will use Colab Pro+ and A100 to increase batch size </li>\n</ul>\n<p>I do everything to speed up experimenting. Then (after strong pipeline discovery) will be time to train on multi gpu configuration (Lambda Cloud, Kaggle 2xGPU … or my personal machine).</p>",
      "rawMarkdown": "Now is a beggining of cometition - time to establish fast experimenting pipeline. My current attitude:\n- make images smaller (256x256)\n- chose lighter/smaller model \n- decrease amout of data (now I experimenting using only 500 images)\n- will use Colab Pro+ and A100 to increase batch size \n\nI do everything to speed up experimenting. Then (after strong pipeline discovery) will be time to train on multi gpu configuration (Lambda Cloud, Kaggle 2xGPU ... or my personal machine).",
      "votes": 3,
      "replies": [
        {
          "id": 2053484,
          "postDate": "2022-12-03T10:27:01.140Z",
          "content": "<p>Thanks a lot.</p>",
          "rawMarkdown": "Thanks a lot."
        }
      ]
    },
    {
      "id": 2053676,
      "postDate": "2022-12-03T13:41:58.533Z",
      "content": "<p>Plus you can train with fp16, that should speed up your training quite a bit and also help you train with larger batch sizes!</p>",
      "rawMarkdown": "Plus you can train with fp16, that should speed up your training quite a bit and also help you train with larger batch sizes!",
      "votes": 1
    },
    {
      "id": 2053675,
      "postDate": "2022-12-03T13:41:13.810Z",
      "content": "<p>Based on your numbers it seems that you are training on DICOM files and are converting them on the fly.</p>\n<p>The conversion is computationally extremely expensive and should be something offloaded to before train. I would suggest you change your pipeline to train on one of the datasets that have been made publicly available (I shared one myself with resolutions of 256px, 512px, 768px and 1024px so there is plenty of sizes to chose from!)</p>",
      "rawMarkdown": "Based on your numbers it seems that you are training on DICOM files and are converting them on the fly.\n\nThe conversion is computationally extremely expensive and should be something offloaded to before train. I would suggest you change your pipeline to train on one of the datasets that have been made publicly available (I shared one myself with resolutions of 256px, 512px, 768px and 1024px so there is plenty of sizes to chose from!)",
      "votes": 1,
      "replies": [
        {
          "id": 2053699,
          "postDate": "2022-12-03T14:22:04.593Z",
          "content": "<p>love it. I'll try a better training pipeline.</p>",
          "rawMarkdown": "love it. I'll try a better training pipeline.",
          "votes": 1
        },
        {
          "id": 2080860,
          "postDate": "2022-12-30T15:17:38.243Z",
          "content": "<p>does GPU or TPU accelerate the conversion of images to png format?</p>",
          "rawMarkdown": "does GPU or TPU accelerate the conversion of images to png format?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2053480,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-03T10:19:44.317000",
      "content": "<p>Now is a beggining of cometition - time to establish fast experimenting pipeline. My current attitude:</p>\n<ul>\n<li>make images smaller (256x256)</li>\n<li>chose lighter/smaller model </li>\n<li>decrease amout of data (now I experimenting using only 500 images)</li>\n<li>will use Colab Pro+ and A100 to increase batch size </li>\n</ul>\n<p>I do everything to speed up experimenting. Then (after strong pipeline discovery) will be time to train on multi gpu configuration (Lambda Cloud, Kaggle 2xGPU … or my personal machine).</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2053484,
          "author_name": "Myo Min Htet(wnp)🇲🇲",
          "author_url": "",
          "post_date": "2022-12-03T10:27:01.140000",
          "content": "<p>Thanks a lot.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2053676,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-03T13:41:58.533000",
      "content": "<p>Plus you can train with fp16, that should speed up your training quite a bit and also help you train with larger batch sizes!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2053675,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-03T13:41:13.810000",
      "content": "<p>Based on your numbers it seems that you are training on DICOM files and are converting them on the fly.</p>\n<p>The conversion is computationally extremely expensive and should be something offloaded to before train. I would suggest you change your pipeline to train on one of the datasets that have been made publicly available (I shared one myself with resolutions of 256px, 512px, 768px and 1024px so there is plenty of sizes to chose from!)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2053699,
          "author_name": "Myo Min Htet(wnp)🇲🇲",
          "author_url": "",
          "post_date": "2022-12-03T14:22:04.593000",
          "content": "<p>love it. I'll try a better training pipeline.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2080860,
          "author_name": "Mohammad Zahrawi",
          "author_url": "",
          "post_date": "2022-12-30T15:17:38.243000",
          "content": "<p>does GPU or TPU accelerate the conversion of images to png format?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2053424": "Training is taking too much time. It takes 1 hour to get a 15% of one epoch and got 13% accuracy and train loss is 0.2.  I don't convert the image into png format. I use dicom images.\n\nI don't have external GPU. I use GPU from Kaggle.",
    "2053480": "Now is a beggining of cometition - time to establish fast experimenting pipeline. My current attitude:\n- make images smaller (256x256)\n- chose lighter/smaller model \n- decrease amout of data (now I experimenting using only 500 images)\n- will use Colab Pro+ and A100 to increase batch size \n\nI do everything to speed up experimenting. Then (after strong pipeline discovery) will be time to train on multi gpu configuration (Lambda Cloud, Kaggle 2xGPU ... or my personal machine).",
    "2053676": "Plus you can train with fp16, that should speed up your training quite a bit and also help you train with larger batch sizes!",
    "2053675": "Based on your numbers it seems that you are training on DICOM files and are converting them on the fly.\n\nThe conversion is computationally extremely expensive and should be something offloaded to before train. I would suggest you change your pipeline to train on one of the datasets that have been made publicly available (I shared one myself with resolutions of 256px, 512px, 768px and 1024px so there is plenty of sizes to chose from!)"
  }
}