{
  "id": 457418,
  "title": "Training of large datasets?",
  "url": "/competitions/UBC-OCEAN/discussion/457418",
  "author_name": "v_parth7",
  "post_date": "2023-11-24T16:18:41.620000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Since the dataset size too large, how you guys perform model training, Google Collab or Amazon Sagemaker does not support this much heavy usage. What would be the best advise - buy GPU's or rent online or something else to train the model???</p>",
  "messages": [
    {
      "id": 2536960,
      "postDate": "2023-11-24T16:18:41.620Z",
      "content": "<p>Since the dataset size too large, how you guys perform model training, Google Collab or Amazon Sagemaker does not support this much heavy usage. What would be the best advise - buy GPU's or rent online or something else to train the model???</p>",
      "rawMarkdown": "Since the dataset size too large, how you guys perform model training, Google Collab or Amazon Sagemaker does not support this much heavy usage. What would be the best advise - buy GPU's or rent online or something else to train the model???",
      "votes": 2
    },
    {
      "id": 2552105,
      "postDate": "2023-12-07T07:38:55.193Z",
      "content": "<p>you can make multiples files of tfr record format  and load all files  to single  dataset  and train using gpu or tpu .the fastest I could reach was probably around 100000 images per 3 min in tpu and 23 minutes in gpu .</p>",
      "rawMarkdown": "you can make multiples files of tfr record format  and load all files  to single  dataset  and train using gpu or tpu .the fastest I could reach was probably around 100000 images per 3 min in tpu and 23 minutes in gpu ."
    },
    {
      "id": 2550227,
      "postDate": "2023-12-05T23:15:47.503Z",
      "content": "<p>You can try subsampling from the dataset based on features distribution (histograms, intensity, etc)<br>\nI took this approach and started from dummy random sub-sampling</p>",
      "rawMarkdown": "You can try subsampling from the dataset based on features distribution (histograms, intensity, etc)\nI took this approach and started from dummy random sub-sampling"
    }
  ],
  "comments": [
    {
      "id": 2552105,
      "author_name": "pratham_adhikari",
      "author_url": "",
      "post_date": "2023-12-07T07:38:55.193000",
      "content": "<p>you can make multiples files of tfr record format  and load all files  to single  dataset  and train using gpu or tpu .the fastest I could reach was probably around 100000 images per 3 min in tpu and 23 minutes in gpu .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2550227,
      "author_name": "Pavel Solomein",
      "author_url": "",
      "post_date": "2023-12-05T23:15:47.503000",
      "content": "<p>You can try subsampling from the dataset based on features distribution (histograms, intensity, etc)<br>\nI took this approach and started from dummy random sub-sampling</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2536960": "Since the dataset size too large, how you guys perform model training, Google Collab or Amazon Sagemaker does not support this much heavy usage. What would be the best advise - buy GPU's or rent online or something else to train the model???",
    "2552105": "you can make multiples files of tfr record format  and load all files  to single  dataset  and train using gpu or tpu .the fastest I could reach was probably around 100000 images per 3 min in tpu and 23 minutes in gpu .",
    "2550227": "You can try subsampling from the dataset based on features distribution (histograms, intensity, etc)\nI took this approach and started from dummy random sub-sampling"
  }
}