{
  "id": 382598,
  "title": "batch size",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/382598",
  "author_name": "Igor Litvin",
  "post_date": "2023-01-31T14:43:53.418000",
  "votes": 0,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Good day my friends.</p>\n<p>I have one problem and misunderstanding. I see that lots people use batch size 16-32 for 1024 resolution.</p>\n<p>How you managed it? I am able to reach batch size 4 maximum for kaggle GPU. I use Eff  B2 v1 and freeze almost everything till block 7 to increase the batch size. I use TensorFlow with ImageDataGenerator and class_weight for training as well.</p>\n<p>For such small batches I have convergence problem (i think) because for 512 and batch 8 I have nice convergence.</p>\n<p>Possible it is not batch size problem as well. Let me know if you know it.</p>\n<p>I know that for most of you it will be stupid question because most of you can use large batches. </p>\n<p>Thanks for help. </p>\n<p>Igor. </p>",
  "messages": [
    {
      "id": 2127154,
      "postDate": "2023-02-02T17:54:33.773Z",
      "content": "<blockquote>\n  <p>For such small batches I have convergence problem (i think) because for 512 and batch 8 I have nice convergence.</p>\n</blockquote>\n<p>Adapt learning rate when you change other hyperparameters, especially when changing the batch size.</p>",
      "rawMarkdown": ">For such small batches I have convergence problem (i think) because for 512 and batch 8 I have nice convergence.\n\nAdapt learning rate when you change other hyperparameters, especially when changing the batch size."
    },
    {
      "id": 2123685,
      "postDate": "2023-01-31T16:25:04.040Z",
      "content": "<p>You need to use TPU for large batch size.</p>",
      "rawMarkdown": "You need to use TPU for large batch size.",
      "replies": [
        {
          "id": 2123800,
          "postDate": "2023-01-31T17:06:41.640Z",
          "content": "<p>Thanks. It slow down a lot. What about keras-gradient-accumulation. I see it stops the NN coefficients correction for the required amount of batches. I think it can be quite cool. </p>",
          "rawMarkdown": "Thanks. It slow down a lot. What about keras-gradient-accumulation. I see it stops the NN coefficients correction for the required amount of batches. I think it can be quite cool. ",
          "replies": [
            {
              "id": 2127588,
              "postDate": "2023-02-03T01:33:28.280Z",
              "content": "<p>It's a good option for people working with limited hardware, don't think there is any downside to doing gradient accumulation</p>",
              "rawMarkdown": "It's a good option for people working with limited hardware, don't think there is any downside to doing gradient accumulation"
            },
            {
              "id": 2127604,
              "postDate": "2023-02-03T01:49:49.030Z",
              "content": "<p>When you mentioned it slow down a lot, did you mean that it's slower when training with TPU? I think you didn't set the configs properly, like the batch size or steps per epoch if you are using tensorflow. In my notebook I train the model with a input size of 1024 * 512, and each epoch takes around 1000s, which is not slower than GPU.</p>",
              "rawMarkdown": "When you mentioned it slow down a lot, did you mean that it's slower when training with TPU? I think you didn't set the configs properly, like the batch size or steps per epoch if you are using tensorflow. In my notebook I train the model with a input size of 1024 * 512, and each epoch takes around 1000s, which is not slower than GPU."
            }
          ]
        }
      ]
    },
    {
      "id": 2123481,
      "postDate": "2023-01-31T14:43:53.417Z",
      "content": "<p>Good day my friends.</p>\n<p>I have one problem and misunderstanding. I see that lots people use batch size 16-32 for 1024 resolution.</p>\n<p>How you managed it? I am able to reach batch size 4 maximum for kaggle GPU. I use Eff  B2 v1 and freeze almost everything till block 7 to increase the batch size. I use TensorFlow with ImageDataGenerator and class_weight for training as well.</p>\n<p>For such small batches I have convergence problem (i think) because for 512 and batch 8 I have nice convergence.</p>\n<p>Possible it is not batch size problem as well. Let me know if you know it.</p>\n<p>I know that for most of you it will be stupid question because most of you can use large batches. </p>\n<p>Thanks for help. </p>\n<p>Igor. </p>",
      "rawMarkdown": "Good day my friends.\n\nI have one problem and misunderstanding. I see that lots people use batch size 16-32 for 1024 resolution.\n\nHow you managed it? I am able to reach batch size 4 maximum for kaggle GPU. I use Eff  B2 v1 and freeze almost everything till block 7 to increase the batch size. I use TensorFlow with ImageDataGenerator and class_weight for training as well.\n\nFor such small batches I have convergence problem (i think) because for 512 and batch 8 I have nice convergence.\n\nPossible it is not batch size problem as well. Let me know if you know it.\n\nI know that for most of you it will be stupid question because most of you can use large batches. \n\nThanks for help. \n\nIgor. "
    }
  ],
  "comments": [
    {
      "id": 2127154,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2023-02-02T17:54:33.773000",
      "content": "<blockquote>\n  <p>For such small batches I have convergence problem (i think) because for 512 and batch 8 I have nice convergence.</p>\n</blockquote>\n<p>Adapt learning rate when you change other hyperparameters, especially when changing the batch size.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2123685,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2023-01-31T16:25:04.040000",
      "content": "<p>You need to use TPU for large batch size.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2123800,
          "author_name": "Igor Litvin",
          "author_url": "",
          "post_date": "2023-01-31T17:06:41.640000",
          "content": "<p>Thanks. It slow down a lot. What about keras-gradient-accumulation. I see it stops the NN coefficients correction for the required amount of batches. I think it can be quite cool. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2127588,
              "author_name": "Cozy Doomer",
              "author_url": "",
              "post_date": "2023-02-03T01:33:28.280000",
              "content": "<p>It's a good option for people working with limited hardware, don't think there is any downside to doing gradient accumulation</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2127604,
              "author_name": "Lau2664",
              "author_url": "",
              "post_date": "2023-02-03T01:49:49.030000",
              "content": "<p>When you mentioned it slow down a lot, did you mean that it's slower when training with TPU? I think you didn't set the configs properly, like the batch size or steps per epoch if you are using tensorflow. In my notebook I train the model with a input size of 1024 * 512, and each epoch takes around 1000s, which is not slower than GPU.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2127154": ">For such small batches I have convergence problem (i think) because for 512 and batch 8 I have nice convergence.\n\nAdapt learning rate when you change other hyperparameters, especially when changing the batch size.",
    "2123685": "You need to use TPU for large batch size.",
    "2123481": "Good day my friends.\n\nI have one problem and misunderstanding. I see that lots people use batch size 16-32 for 1024 resolution.\n\nHow you managed it? I am able to reach batch size 4 maximum for kaggle GPU. I use Eff  B2 v1 and freeze almost everything till block 7 to increase the batch size. I use TensorFlow with ImageDataGenerator and class_weight for training as well.\n\nFor such small batches I have convergence problem (i think) because for 512 and batch 8 I have nice convergence.\n\nPossible it is not batch size problem as well. Let me know if you know it.\n\nI know that for most of you it will be stupid question because most of you can use large batches. \n\nThanks for help. \n\nIgor. "
  }
}