{
  "id": 191310,
  "title": "Understanding of Batchsize in this competition",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/191310",
  "author_name": "Dewei Chen",
  "post_date": "2020-10-15T19:01:09.313000",
  "votes": -4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, the definition of Batch Size: <strong>The number of samples selected for one batch training</strong>.</p>\n<p>For a small sample dataset training purpose, full batch training is the best idea.<br>\nWhen we training a big dataset, we cannot send all data into the network, so we need a training batch by batch. The main reason we have to use batch learning is that <strong>GPU RAM is not enough</strong> to loading the full dataset. Hence, we need to select a suitable batch size number.<br>\nThe <strong>default value</strong> of batch size is <strong>32</strong>, and it is accustomed to x2 or x0.5 to adjust this parameter.</p>\n<p>A figure to understanding the principle of batch size: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1860617%2Fb618981705fd63eb2e912919609c3151%2Fv2-7ffb5d7ee7b6d582a2be9768dc6fa4b1_720w.jpg?generation=1602786940737855&amp;alt=media\" alt=\"picture\"><br>\nWhen we did Gradient Descent, <strong>green line is a batch size 2</strong>; <strong>yellow line is full batch</strong>, and <strong>black line is a batch size 3</strong>; Stochastic Gradient Descent can be considered as <strong>batch size equal to 1</strong></p>\n<p>Therefore, the most suitable value of batch size depends on the dataset, it is not always the same.  </p>\n<p>The dataset of this <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection\" target=\"_blank\">PE Detection</a> competition, dataset is near 1TB data. This is a huge dataset, we cannot using full batch learning; So we need to determine a suitable batch size, I start with 32. Kaggle gives us 16GB RAM Tesla GPU for training.  I try to training a <strong>DenseNet201</strong>, with <strong>batch size 4</strong>; and also <strong>ResNet50</strong> with <strong>batch size 8</strong><br>\nThe large batch size will give faster training. Small batch sizes always training slow. So we always making batch size as much as possible. (But it will be limited by GPU RAM and numbers)</p>\n<p>Reference: <br>\n[1] <a href=\"https://www.zhihu.com/question/61607442\" target=\"_blank\">https://www.zhihu.com/question/61607442</a><br>\n[2] <a href=\"https://stats.stackexchange.com/questions/153531/what-is-batch-size-in-neural-network\" target=\"_blank\">https://stats.stackexchange.com/questions/153531/what-is-batch-size-in-neural-network</a></p>",
  "messages": [
    {
      "id": 1050827,
      "postDate": "2020-10-15T19:01:09.313Z",
      "content": "<p>First of all, the definition of Batch Size: <strong>The number of samples selected for one batch training</strong>.</p>\n<p>For a small sample dataset training purpose, full batch training is the best idea.<br>\nWhen we training a big dataset, we cannot send all data into the network, so we need a training batch by batch. The main reason we have to use batch learning is that <strong>GPU RAM is not enough</strong> to loading the full dataset. Hence, we need to select a suitable batch size number.<br>\nThe <strong>default value</strong> of batch size is <strong>32</strong>, and it is accustomed to x2 or x0.5 to adjust this parameter.</p>\n<p>A figure to understanding the principle of batch size: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1860617%2Fb618981705fd63eb2e912919609c3151%2Fv2-7ffb5d7ee7b6d582a2be9768dc6fa4b1_720w.jpg?generation=1602786940737855&amp;alt=media\" alt=\"picture\"><br>\nWhen we did Gradient Descent, <strong>green line is a batch size 2</strong>; <strong>yellow line is full batch</strong>, and <strong>black line is a batch size 3</strong>; Stochastic Gradient Descent can be considered as <strong>batch size equal to 1</strong></p>\n<p>Therefore, the most suitable value of batch size depends on the dataset, it is not always the same.  </p>\n<p>The dataset of this <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection\" target=\"_blank\">PE Detection</a> competition, dataset is near 1TB data. This is a huge dataset, we cannot using full batch learning; So we need to determine a suitable batch size, I start with 32. Kaggle gives us 16GB RAM Tesla GPU for training.  I try to training a <strong>DenseNet201</strong>, with <strong>batch size 4</strong>; and also <strong>ResNet50</strong> with <strong>batch size 8</strong><br>\nThe large batch size will give faster training. Small batch sizes always training slow. So we always making batch size as much as possible. (But it will be limited by GPU RAM and numbers)</p>\n<p>Reference: <br>\n[1] <a href=\"https://www.zhihu.com/question/61607442\" target=\"_blank\">https://www.zhihu.com/question/61607442</a><br>\n[2] <a href=\"https://stats.stackexchange.com/questions/153531/what-is-batch-size-in-neural-network\" target=\"_blank\">https://stats.stackexchange.com/questions/153531/what-is-batch-size-in-neural-network</a></p>",
      "rawMarkdown": "First of all, the definition of Batch Size: **The number of samples selected for one batch training**.\n\nFor a small sample dataset training purpose, full batch training is the best idea.\nWhen we training a big dataset, we cannot send all data into the network, so we need a training batch by batch. The main reason we have to use batch learning is that **GPU RAM is not enough** to loading the full dataset. Hence, we need to select a suitable batch size number.\nThe **default value** of batch size is **32**, and it is accustomed to x2 or x0.5 to adjust this parameter.\n\nA figure to understanding the principle of batch size: \n![picture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1860617%2Fb618981705fd63eb2e912919609c3151%2Fv2-7ffb5d7ee7b6d582a2be9768dc6fa4b1_720w.jpg?generation=1602786940737855&alt=media)\nWhen we did Gradient Descent, **green line is a batch size 2**; **yellow line is full batch**, and **black line is a batch size 3**; Stochastic Gradient Descent can be considered as **batch size equal to 1**\n\nTherefore, the most suitable value of batch size depends on the dataset, it is not always the same.  \n\nThe dataset of this [PE Detection](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection) competition, dataset is near 1TB data. This is a huge dataset, we cannot using full batch learning; So we need to determine a suitable batch size, I start with 32. Kaggle gives us 16GB RAM Tesla GPU for training.  I try to training a **DenseNet201**, with **batch size 4**; and also **ResNet50** with **batch size 8**\nThe large batch size will give faster training. Small batch sizes always training slow. So we always making batch size as much as possible. (But it will be limited by GPU RAM and numbers)\n\n\n\n\nReference: \n[1] https://www.zhihu.com/question/61607442\n[2] https://stats.stackexchange.com/questions/153531/what-is-batch-size-in-neural-network",
      "votes": -4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1050827": "First of all, the definition of Batch Size: **The number of samples selected for one batch training**.\n\nFor a small sample dataset training purpose, full batch training is the best idea.\nWhen we training a big dataset, we cannot send all data into the network, so we need a training batch by batch. The main reason we have to use batch learning is that **GPU RAM is not enough** to loading the full dataset. Hence, we need to select a suitable batch size number.\nThe **default value** of batch size is **32**, and it is accustomed to x2 or x0.5 to adjust this parameter.\n\nA figure to understanding the principle of batch size: \n![picture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1860617%2Fb618981705fd63eb2e912919609c3151%2Fv2-7ffb5d7ee7b6d582a2be9768dc6fa4b1_720w.jpg?generation=1602786940737855&alt=media)\nWhen we did Gradient Descent, **green line is a batch size 2**; **yellow line is full batch**, and **black line is a batch size 3**; Stochastic Gradient Descent can be considered as **batch size equal to 1**\n\nTherefore, the most suitable value of batch size depends on the dataset, it is not always the same.  \n\nThe dataset of this [PE Detection](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection) competition, dataset is near 1TB data. This is a huge dataset, we cannot using full batch learning; So we need to determine a suitable batch size, I start with 32. Kaggle gives us 16GB RAM Tesla GPU for training.  I try to training a **DenseNet201**, with **batch size 4**; and also **ResNet50** with **batch size 8**\nThe large batch size will give faster training. Small batch sizes always training slow. So we always making batch size as much as possible. (But it will be limited by GPU RAM and numbers)\n\n\n\n\nReference: \n[1] https://www.zhihu.com/question/61607442\n[2] https://stats.stackexchange.com/questions/153531/what-is-batch-size-in-neural-network"
  }
}