{
  "id": 513588,
  "title": "Amount of training data and its effect on the result",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/513588",
  "author_name": "Dmitry Uarov",
  "post_date": "2024-06-20T18:57:16.229000",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>All this time I have been training my model using only 1.87M samples. It is important for me to know how much the result can change if I use 5 times more data. It is important because either I will continue to experiment with model's architecture, or I will already start thinking about computing resources.</p>",
  "messages": [
    {
      "id": 2881551,
      "postDate": "2024-06-20T18:57:16.230Z",
      "content": "<p>All this time I have been training my model using only 1.87M samples. It is important for me to know how much the result can change if I use 5 times more data. It is important because either I will continue to experiment with model's architecture, or I will already start thinking about computing resources.</p>",
      "rawMarkdown": "All this time I have been training my model using only 1.87M samples. It is important for me to know how much the result can change if I use 5 times more data. It is important because either I will continue to experiment with model's architecture, or I will already start thinking about computing resources.",
      "votes": 2
    },
    {
      "id": 2894485,
      "postDate": "2024-06-28T14:07:22.327Z",
      "content": "<p>hi, I am confused. How are the 1.87M samples calculated</p>",
      "rawMarkdown": "hi, I am confused. How are the 1.87M samples calculated"
    },
    {
      "id": 2887098,
      "postDate": "2024-06-24T03:13:59.330Z",
      "content": "<p>hi ,kaggler!  I define a very deep convolutional neural network (CNN) model,just like this <br>\n<code>def build_cnn(activation='relu'):    \n    return keras.Sequential([\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.LSTM(64, return_sequences=True),\n        keras.layers.BatchNormalization(),\n        keras.layers.GRU(64, return_sequences=True),\n        keras.layers.BatchNormalization(),\n        keras.layers.Dense(64, activation=activation),\n    ])</code><br>\nBut it runs too slow, I try to make his Trainable param close to 2M.But it's running too slow, what's the best way to fix it?</p>",
      "rawMarkdown": "hi ,kaggler!  I define a very deep convolutional neural network (CNN) model,just like this \n`def build_cnn(activation='relu'):    \n    return keras.Sequential([\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.LSTM(64, return_sequences=True),\n        keras.layers.BatchNormalization(),\n        keras.layers.GRU(64, return_sequences=True),\n        keras.layers.BatchNormalization(),\n        keras.layers.Dense(64, activation=activation),\n    ])`\nBut it runs too slow, I try to make his Trainable param close to 2M.But it's running too slow, what's the best way to fix it?\n\n"
    },
    {
      "id": 2881729,
      "postDate": "2024-06-20T23:39:22.707Z",
      "content": "<p>The train data is time ordered - you probably want a full year, and I would suggest it be the last year rather than the first on the assumption that the test year is later.</p>",
      "rawMarkdown": "The train data is time ordered - you probably want a full year, and I would suggest it be the last year rather than the first on the assumption that the test year is later.",
      "replies": [
        {
          "id": 2882266,
          "postDate": "2024-06-21T08:59:56.540Z",
          "content": "<p>My question is very simple - how much does the result of the model increase with an increase in the amount of data? I'm waiting for at least someone to answer that \"I trained on <strong>n</strong> samples, got … score, and on full data I got … improvement\". This is extremely important to me because I don't want to spend money on cloud computing for the sake of a result that will not reach at least silver. 😁</p>",
          "rawMarkdown": "My question is very simple - how much does the result of the model increase with an increase in the amount of data? I'm waiting for at least someone to answer that \"I trained on **n** samples, got ... score, and on full data I got ... improvement\". This is extremely important to me because I don't want to spend money on cloud computing for the sake of a result that will not reach at least silver. 😁",
          "replies": [
            {
              "id": 2882320,
              "postDate": "2024-06-21T09:46:48.427Z",
              "content": "<p>This is for old LB but should be similar for new too.<br>\nI trained on ~0.9M for ~0.75<br>\nTrain the same model on ~9M got ~0.78.<br>\nSo *10 times data (also more epochs but I don't remember if 10 times too) was +0.03.<br>\nFor you it's *5 times so worst case scenario you will get less than +0.015 on LB. Although you might get more since some models get more boost from more data (usually those with lower score)</p>",
              "rawMarkdown": "This is for old LB but should be similar for new too.\nI trained on ~0.9M for ~0.75\nTrain the same model on ~9M got ~0.78.\nSo *10 times data (also more epochs but I don't remember if 10 times too) was +0.03.\nFor you it's *5 times so worst case scenario you will get less than +0.015 on LB. Although you might get more since some models get more boost from more data (usually those with lower score)\n",
              "votes": 5
            },
            {
              "id": 2882366,
              "postDate": "2024-06-21T10:09:29.463Z",
              "content": "<p>Hmmm.. Thank you! A little strange of course, because when I increased the amount of data from 0.625M to 1.875M, I got an improvement +0.02 🤔. It is also interesting that you got 0.75 with training only on ~0.9M samples - great result!</p>",
              "rawMarkdown": "Hmmm.. Thank you! A little strange of course, because when I increased the amount of data from 0.625M to 1.875M, I got an improvement +0.02 🤔. It is also interesting that you got 0.75 with training only on ~0.9M samples - great result!"
            },
            {
              "id": 2882414,
              "postDate": "2024-06-21T10:34:59.890Z",
              "content": "<p>Not so strange. As I said, bad models can get a greater boost. Based on you boost from 0.625M, the maximum boost you will get from going to *5 is +0.033, but probably will be closer to +0.02, +0.025 max.</p>",
              "rawMarkdown": "Not so strange. As I said, bad models can get a greater boost. Based on you boost from 0.625M, the maximum boost you will get from going to *5 is +0.033, but probably will be closer to +0.02, +0.025 max.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2881636,
      "postDate": "2024-06-20T21:24:22.380Z",
      "content": "<p>I've noticed that my CV scales with the amount of training data I use. Could be that my approach is not great though….</p>",
      "rawMarkdown": "I've noticed that my CV scales with the amount of training data I use. Could be that my approach is not great though....\n",
      "replies": [
        {
          "id": 2882274,
          "postDate": "2024-06-21T09:02:36.437Z",
          "content": "<p>No, it is quite logical that with an increase in the amount of training data, the result improves. My question is, <strong>how much</strong> of an improvement is this?</p>",
          "rawMarkdown": "No, it is quite logical that with an increase in the amount of training data, the result improves. My question is, **how much** of an improvement is this?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2894485,
      "author_name": "Baiph",
      "author_url": "",
      "post_date": "2024-06-28T14:07:22.327000",
      "content": "<p>hi, I am confused. How are the 1.87M samples calculated</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2887098,
      "author_name": "Aristotle.Chen",
      "author_url": "",
      "post_date": "2024-06-24T03:13:59.330000",
      "content": "<p>hi ,kaggler!  I define a very deep convolutional neural network (CNN) model,just like this <br>\n<code>def build_cnn(activation='relu'):    \n    return keras.Sequential([\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.LSTM(64, return_sequences=True),\n        keras.layers.BatchNormalization(),\n        keras.layers.GRU(64, return_sequences=True),\n        keras.layers.BatchNormalization(),\n        keras.layers.Dense(64, activation=activation),\n    ])</code><br>\nBut it runs too slow, I try to make his Trainable param close to 2M.But it's running too slow, what's the best way to fix it?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2881729,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2024-06-20T23:39:22.707000",
      "content": "<p>The train data is time ordered - you probably want a full year, and I would suggest it be the last year rather than the first on the assumption that the test year is later.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2882266,
          "author_name": "Dmitry Uarov",
          "author_url": "",
          "post_date": "2024-06-21T08:59:56.540000",
          "content": "<p>My question is very simple - how much does the result of the model increase with an increase in the amount of data? I'm waiting for at least someone to answer that \"I trained on <strong>n</strong> samples, got … score, and on full data I got … improvement\". This is extremely important to me because I don't want to spend money on cloud computing for the sake of a result that will not reach at least silver. 😁</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2882320,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-21T09:46:48.427000",
              "content": "<p>This is for old LB but should be similar for new too.<br>\nI trained on ~0.9M for ~0.75<br>\nTrain the same model on ~9M got ~0.78.<br>\nSo *10 times data (also more epochs but I don't remember if 10 times too) was +0.03.<br>\nFor you it's *5 times so worst case scenario you will get less than +0.015 on LB. Although you might get more since some models get more boost from more data (usually those with lower score)</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2882366,
              "author_name": "Dmitry Uarov",
              "author_url": "",
              "post_date": "2024-06-21T10:09:29.463000",
              "content": "<p>Hmmm.. Thank you! A little strange of course, because when I increased the amount of data from 0.625M to 1.875M, I got an improvement +0.02 🤔. It is also interesting that you got 0.75 with training only on ~0.9M samples - great result!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2882414,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-21T10:34:59.890000",
              "content": "<p>Not so strange. As I said, bad models can get a greater boost. Based on you boost from 0.625M, the maximum boost you will get from going to *5 is +0.033, but probably will be closer to +0.02, +0.025 max.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2881636,
      "author_name": "Anthony Meza",
      "author_url": "",
      "post_date": "2024-06-20T21:24:22.380000",
      "content": "<p>I've noticed that my CV scales with the amount of training data I use. Could be that my approach is not great though….</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2882274,
          "author_name": "Dmitry Uarov",
          "author_url": "",
          "post_date": "2024-06-21T09:02:36.437000",
          "content": "<p>No, it is quite logical that with an increase in the amount of training data, the result improves. My question is, <strong>how much</strong> of an improvement is this?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2881551": "All this time I have been training my model using only 1.87M samples. It is important for me to know how much the result can change if I use 5 times more data. It is important because either I will continue to experiment with model's architecture, or I will already start thinking about computing resources.",
    "2894485": "hi, I am confused. How are the 1.87M samples calculated",
    "2887098": "hi ,kaggler!  I define a very deep convolutional neural network (CNN) model,just like this \n`def build_cnn(activation='relu'):    \n    return keras.Sequential([\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(256, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(128, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.Conv1D(64, 3, padding='same', activation=activation),\n        keras.layers.BatchNormalization(),\n        keras.layers.LSTM(64, return_sequences=True),\n        keras.layers.BatchNormalization(),\n        keras.layers.GRU(64, return_sequences=True),\n        keras.layers.BatchNormalization(),\n        keras.layers.Dense(64, activation=activation),\n    ])`\nBut it runs too slow, I try to make his Trainable param close to 2M.But it's running too slow, what's the best way to fix it?\n\n",
    "2881729": "The train data is time ordered - you probably want a full year, and I would suggest it be the last year rather than the first on the assumption that the test year is later.",
    "2881636": "I've noticed that my CV scales with the amount of training data I use. Could be that my approach is not great though....\n"
  }
}