{
  "id": 515017,
  "title": "Ensemble of models improvement",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/515017",
  "author_name": "Vasilis",
  "post_date": "2024-06-26T15:00:38.301000",
  "votes": 9,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Hello guys, my single model trained on 9m data validated at 1m data achives around 0.74 LB score. If i train another one with a totally new validation set of 1m samples is it expected that taking the average result of the two models will perform significantly better? Like at least 0.01+? The train takes too long and i could spend the time in other experiments if the ensemble is not very promising.</p>",
  "messages": [
    {
      "id": 2898653,
      "postDate": "2024-07-01T08:59:26.113Z",
      "content": "<p>Okay the models are finally ready, I took the avg of a 0.74984 and a 0.75317 and i got 0.76092, all LBs. I guess if i keep adding models and take their avg the extra value will be less and less right?</p>",
      "rawMarkdown": "Okay the models are finally ready, I took the avg of a 0.74984 and a 0.75317 and i got 0.76092, all LBs. I guess if i keep adding models and take their avg the extra value will be less and less right?",
      "votes": 7,
      "replies": [
        {
          "id": 2898678,
          "postDate": "2024-07-01T09:24:49.777Z",
          "content": "<p>Yes, but you will still get a boost, up to a certain point when it will become negligible. As long as you have compute- the more the merrier (esp. as you got a large initial boost, much better than I anticipated lol)</p>",
          "rawMarkdown": "Yes, but you will still get a boost, up to a certain point when it will become negligible. As long as you have compute- the more the merrier (esp. as you got a large initial boost, much better than I anticipated lol)",
          "replies": [
            {
              "id": 2898694,
              "postDate": "2024-07-01T09:35:05.863Z",
              "content": "<p>Is it a good idea instead of avg to train a metamodel? But then i guess i have to reserve extra data. What boost should i expect then if you compare with just avg</p>",
              "rawMarkdown": "Is it a good idea instead of avg to train a metamodel? But then i guess i have to reserve extra data. What boost should i expect then if you compare with just avg"
            }
          ]
        },
        {
          "id": 2902930,
          "postDate": "2024-07-03T14:41:17.447Z",
          "content": "<p>I have added a third model with LB 0.74878. The average of the three models achieved went up tp 0.76277, up only a bit from 0.76092 that was with the avg of two models</p>",
          "rawMarkdown": "I have added a third model with LB 0.74878. The average of the three models achieved went up tp 0.76277, up only a bit from 0.76092 that was with the avg of two models",
          "votes": 1,
          "replies": [
            {
              "id": 2903143,
              "postDate": "2024-07-03T16:17:20.647Z",
              "content": "<p>That uplift is still higher than I'm seeing so far</p>",
              "rawMarkdown": "That uplift is still higher than I'm seeing so far"
            },
            {
              "id": 2903277,
              "postDate": "2024-07-03T17:36:52.757Z",
              "content": "<p>Not related to the topic, i just discovered by pure luck that reducing massively the LR to train the transformer gave a huge bump to the validation loss, so now i keep training all three of them with tiny LRs. I have put a linear scheduler with a big warmup with a 2 to 4 e-6 LR. So far i saw an improvement from 0.753 to 0.76 in val set. </p>\n<p>Update: I saw all of the improvement reflected in LB and even more. So the model that was at 0.74878 went up to 0.75633 after got refined with small LRs</p>",
              "rawMarkdown": "Not related to the topic, i just discovered by pure luck that reducing massively the LR to train the transformer gave a huge bump to the validation loss, so now i keep training all three of them with tiny LRs. I have put a linear scheduler with a big warmup with a 2 to 4 e-6 LR. So far i saw an improvement from 0.753 to 0.76 in val set. \n\nUpdate: I saw all of the improvement reflected in LB and even more. So the model that was at 0.74878 went up to 0.75633 after got refined with small LRs",
              "votes": 4
            },
            {
              "id": 2905299,
              "postDate": "2024-07-04T19:53:25.987Z",
              "content": "<p>Nice! I also saw benefit from tinkering with the LR, though I think I'm running out of runway.</p>",
              "rawMarkdown": "Nice! I also saw benefit from tinkering with the LR, though I think I'm running out of runway.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2891276,
      "postDate": "2024-06-26T15:00:38.300Z",
      "content": "<p>Hello guys, my single model trained on 9m data validated at 1m data achives around 0.74 LB score. If i train another one with a totally new validation set of 1m samples is it expected that taking the average result of the two models will perform significantly better? Like at least 0.01+? The train takes too long and i could spend the time in other experiments if the ensemble is not very promising.</p>",
      "rawMarkdown": "Hello guys, my single model trained on 9m data validated at 1m data achives around 0.74 LB score. If i train another one with a totally new validation set of 1m samples is it expected that taking the average result of the two models will perform significantly better? Like at least 0.01+? The train takes too long and i could spend the time in other experiments if the ensemble is not very promising.",
      "votes": 8
    },
    {
      "id": 2891294,
      "postDate": "2024-06-26T15:16:17.717Z",
      "content": "<p>Yes, ensemble helps (well, from my Kaggle experience, ensembling ALWAYS helps). But 0.01 is a bit too optimistic for an ensemble of two. Expect to get ~0.003. This is at least my experience for ~0.78LB same models. Maybe for 0.74 it would be more.</p>",
      "rawMarkdown": "Yes, ensemble helps (well, from my Kaggle experience, ensembling ALWAYS helps). But 0.01 is a bit too optimistic for an ensemble of two. Expect to get ~0.003. This is at least my experience for ~0.78LB same models. Maybe for 0.74 it would be more.",
      "votes": 6
    },
    {
      "id": 2893274,
      "postDate": "2024-06-27T18:01:41.947Z",
      "content": "<p>The main idea is training multiple models on different subsets of data which is called bagging. Decision tree based models do that inherently but you can do it with any kind of model.</p>",
      "rawMarkdown": "The main idea is training multiple models on different subsets of data which is called bagging. Decision tree based models do that inherently but you can do it with any kind of model.",
      "votes": 2
    },
    {
      "id": 2911728,
      "postDate": "2024-07-08T14:06:58.260Z",
      "content": "<p>Hey guys, in the 4th train-val pair the loss is considerably higher in the val set than the other train-val pairs so far. As a matter of fact adding the 4th model to the average it dropped my score from 0.76534 to 0.76412, does this happen often? it it possible that the second lower score generalizes better and achieves a higher score when the whole test set will be used in the end?</p>",
      "rawMarkdown": "Hey guys, in the 4th train-val pair the loss is considerably higher in the val set than the other train-val pairs so far. As a matter of fact adding the 4th model to the average it dropped my score from 0.76534 to 0.76412, does this happen often? it it possible that the second lower score generalizes better and achieves a higher score when the whole test set will be used in the end?",
      "replies": [
        {
          "id": 2913579,
          "postDate": "2024-07-09T15:00:13.987Z",
          "content": "<p>It doesn't seem that weird to me. It also seems plausible that your largest gains from the ensemble may come from the hard to predict targets (12-15). I suspect that there is something enabling reliable prediction of these targets for the highest LB scores. That said, I don't think the trick for q0002 is stable across months/years in the training data so that may be a source of bad times on the private lb. My assumption is that the baseline U-Net implementation has masking for these targets which enables the consistency on the public/private lb. Maybe it is as simple as KNN to guess whether or not to use the mean or use the trick.</p>",
          "rawMarkdown": "It doesn't seem that weird to me. It also seems plausible that your largest gains from the ensemble may come from the hard to predict targets (12-15). I suspect that there is something enabling reliable prediction of these targets for the highest LB scores. That said, I don't think the trick for q0002 is stable across months/years in the training data so that may be a source of bad times on the private lb. My assumption is that the baseline U-Net implementation has masking for these targets which enables the consistency on the public/private lb. Maybe it is as simple as KNN to guess whether or not to use the mean or use the trick.",
          "replies": [
            {
              "id": 2913784,
              "postDate": "2024-07-09T16:22:59.723Z",
              "content": "<blockquote>\n  <p>I don't think the trick for q0002 is stable across months/years in the training data</p>\n</blockquote>\n<p>It's an easy thing to check. No need to 'think' here, you can know.<br>\nAlso easy to check for public LB with some probing.</p>",
              "rawMarkdown": "> I don't think the trick for q0002 is stable across months/years in the training data\n\nIt's an easy thing to check. No need to 'think' here, you can know.\nAlso easy to check for public LB with some probing."
            },
            {
              "id": 2914120,
              "postDate": "2024-07-09T18:24:56.700Z",
              "content": "<p>replacing the ptend_q0002_12 upt to 27 works well. These are the numbers in the validation set</p>\n<p>R2 for ptend_q0002_12: 1.0<br>\nR2 for ptend_q0002_13: 1.0<br>\nR2 for ptend_q0002_14: 1.0<br>\nR2 for ptend_q0002_15: 1.0<br>\nR2 for ptend_q0002_16: 1.0<br>\nR2 for ptend_q0002_17: 1.0<br>\nR2 for ptend_q0002_18: 1.0<br>\nR2 for ptend_q0002_19: 1.0<br>\nR2 for ptend_q0002_20: 1.0<br>\nR2 for ptend_q0002_21: 1.0<br>\nR2 for ptend_q0002_22: 1.0<br>\nR2 for ptend_q0002_23: 1.0<br>\nR2 for ptend_q0002_24: 1.0<br>\nR2 for ptend_q0002_25: 1.0<br>\nR2 for ptend_q0002_26: 1.0<br>\nR2 for ptend_q0002_27: 1.0</p>",
              "rawMarkdown": "replacing the ptend_q0002_12 upt to 27 works well. These are the numbers in the validation set\n\nR2 for ptend_q0002_12: 1.0\nR2 for ptend_q0002_13: 1.0\nR2 for ptend_q0002_14: 1.0\nR2 for ptend_q0002_15: 1.0\nR2 for ptend_q0002_16: 1.0\nR2 for ptend_q0002_17: 1.0\nR2 for ptend_q0002_18: 1.0\nR2 for ptend_q0002_19: 1.0\nR2 for ptend_q0002_20: 1.0\nR2 for ptend_q0002_21: 1.0\nR2 for ptend_q0002_22: 1.0\nR2 for ptend_q0002_23: 1.0\nR2 for ptend_q0002_24: 1.0\nR2 for ptend_q0002_25: 1.0\nR2 for ptend_q0002_26: 1.0\nR2 for ptend_q0002_27: 1.0"
            },
            {
              "id": 2914275,
              "postDate": "2024-07-09T20:01:44.770Z",
              "content": "<p>thanks, i found where my error was</p>",
              "rawMarkdown": "thanks, i found where my error was"
            }
          ]
        },
        {
          "id": 2914140,
          "postDate": "2024-07-09T18:28:14.797Z",
          "content": "<p>the thing is for the 4th test-val pair i collected randomly ids (that i had not in other val sets) every 100k rows. For the other val sets i collected randomly every 1000 rows, so maybe there are highly correlated data across the dataset or we just cover all long lats if we sample every 1000 rows. I am just a bit worried that if there are highly correlated data and then there are not in the hidden test set then my results will get down in the hidden test set</p>",
          "rawMarkdown": "the thing is for the 4th test-val pair i collected randomly ids (that i had not in other val sets) every 100k rows. For the other val sets i collected randomly every 1000 rows, so maybe there are highly correlated data across the dataset or we just cover all long lats if we sample every 1000 rows. I am just a bit worried that if there are highly correlated data and then there are not in the hidden test set then my results will get down in the hidden test set"
        }
      ]
    },
    {
      "id": 2891877,
      "postDate": "2024-06-27T00:56:15.193Z",
      "content": "<p>I was wondering how you broke through 0.7, but then I realized something. 😅😅</p>",
      "rawMarkdown": "I was wondering how you broke through 0.7, but then I realized something. 😅😅",
      "replies": [
        {
          "id": 2892586,
          "postDate": "2024-06-27T10:32:58.917Z",
          "content": "<p>big transformer model trained for more than 10 epochs, and i replace the values of the problematic columns using the state columns as have been discussed in other threads</p>",
          "rawMarkdown": "big transformer model trained for more than 10 epochs, and i replace the values of the problematic columns using the state columns as have been discussed in other threads",
          "votes": 3,
          "replies": [
            {
              "id": 2892858,
              "postDate": "2024-06-27T13:44:20.443Z",
              "content": "<p>Gotcha, yes, the latter is something I missed in my own notebooks and only realized before posting. Also, the transformer encoder models I've created are exceptionally slow to train; I haven't even run a full epoch. Do you think it was worth it versus adjusting model parameters including dropout and running longer with another model?</p>",
              "rawMarkdown": "Gotcha, yes, the latter is something I missed in my own notebooks and only realized before posting. Also, the transformer encoder models I've created are exceptionally slow to train; I haven't even run a full epoch. Do you think it was worth it versus adjusting model parameters including dropout and running longer with another model?"
            },
            {
              "id": 2893564,
              "postDate": "2024-06-27T21:49:00.850Z",
              "content": "<p>for me a big mind-blowing revelation was that if i process the data batched using polars and also with smaller batch size the model was trained much faster. I suspect because it did not have to move data in and out in cuda all the time. Before the competition i was under the impression that whatever do not blow my cuda memory is good but i notice that with smaller batch size and less stuff in cuda the train can be much faster. Considering dropout i tried without it as suggested in other discussion and also more transformer layers helped, i saw difference from 3 to 6 for example</p>",
              "rawMarkdown": "for me a big mind-blowing revelation was that if i process the data batched using polars and also with smaller batch size the model was trained much faster. I suspect because it did not have to move data in and out in cuda all the time. Before the competition i was under the impression that whatever do not blow my cuda memory is good but i notice that with smaller batch size and less stuff in cuda the train can be much faster. Considering dropout i tried without it as suggested in other discussion and also more transformer layers helped, i saw difference from 3 to 6 for example",
              "votes": 4
            },
            {
              "id": 2893646,
              "postDate": "2024-06-28T00:21:32.843Z",
              "content": "<p>Thanks! I appreciate your feedback. I might see what happens with dramatically smaller batches.</p>",
              "rawMarkdown": "Thanks! I appreciate your feedback. I might see what happens with dramatically smaller batches."
            },
            {
              "id": 2894472,
              "postDate": "2024-06-28T13:56:38.723Z",
              "content": "<p>no need to go very low, i think i use a batch of 256, i read every time 1/100th of the train data using polars, make it into tensor, then start feeding the model, after that data are done i go to the next hundredth. Same for validation</p>\n<pre><code>     :\n        reader = pl.read_csv_batched(train_file, =chunk_size)\n        counter = 0\n\n         :\n\n            prep_chunk_time_start = time.time()\n            try:\n                batches = reader.next_batches(20)\n                 batches[0] is None:\n                    break  #  more data  read\n            except:\n                break\n\n            end_batches = time.time()\n            (, end_batches - prep_chunk_time_start)\n\n             df  batches:\n\n                start_chunk_time = time.time()\n                train_dataset, _ = # make it into tensor here\n\n                train_loader = DataLoader(train_dataset,\n                                          =BATCH_SIZE,\n                                          =,\n                                          =collate_fn,\n                                          )\n</code></pre>\n<p>that way only a small piece is in cuda, and that makes it faster. I train a 15 million parameter transformer using all the train data (9m) and 1m validation in a laptop with an NVIDIA GeForce RTX 2070 Super Max-Q</p>",
              "rawMarkdown": "no need to go very low, i think i use a batch of 256, i read every time 1/100th of the train data using polars, make it into tensor, then start feeding the model, after that data are done i go to the next hundredth. Same for validation\n\n```\n    while True:\n        reader = pl.read_csv_batched(train_file, batch_size=chunk_size)\n        counter = 0\n\n        while True:\n\n            prep_chunk_time_start = time.time()\n            try:\n                batches = reader.next_batches(20)\n                if batches[0] is None:\n                    break  # No more data to read\n            except:\n                break\n\n            end_batches = time.time()\n            print(\"end batches time:\", end_batches - prep_chunk_time_start)\n\n            for df in batches:\n\n                start_chunk_time = time.time()\n                train_dataset, _ = # make it into tensor here\n\n                train_loader = DataLoader(train_dataset,\n                                          batch_size=BATCH_SIZE,\n                                          shuffle=True,\n                                          collate_fn=collate_fn,\n                                          )\n```\n\nthat way only a small piece is in cuda, and that makes it faster. I train a 15 million parameter transformer using all the train data (9m) and 1m validation in a laptop with an NVIDIA GeForce RTX 2070 Super Max-Q",
              "votes": 2
            },
            {
              "id": 2894643,
              "postDate": "2024-06-28T15:57:01.170Z",
              "content": "<p>It's cool that you have already achieved such a result! Let me ask you, how do you prepare data before training? Only normalizing like in public notebooks or something else? </p>",
              "rawMarkdown": "It's cool that you have already achieved such a result! Let me ask you, how do you prepare data before training? Only normalizing like in public notebooks or something else? "
            },
            {
              "id": 2894796,
              "postDate": "2024-06-28T17:29:56.497Z",
              "content": "<p>Thank you! yes nothing special, just (x-mean_x) / std_x and (y -mean_y)/std_y like has mentioned in many other discussions.</p>",
              "rawMarkdown": "Thank you! yes nothing special, just (x-mean_x) / std_x and (y -mean_y)/std_y like has mentioned in many other discussions.",
              "votes": 2
            },
            {
              "id": 2895864,
              "postDate": "2024-06-29T12:18:29.540Z",
              "content": "<p>Thanks! I still had cuda memory / rematerialization errors except for smaller transformer encoder networks and switched to an LSTM-based model instead.</p>",
              "rawMarkdown": "Thanks! I still had cuda memory / rematerialization errors except for smaller transformer encoder networks and switched to an LSTM-based model instead.",
              "votes": 2
            },
            {
              "id": 2896017,
              "postDate": "2024-06-29T14:33:32.657Z",
              "content": "<p>Interesting that LSTM-based models achieve similar result</p>",
              "rawMarkdown": "Interesting that LSTM-based models achieve similar result",
              "votes": 2
            },
            {
              "id": 2896284,
              "postDate": "2024-06-29T18:05:24.377Z",
              "content": "<p>I trained for six epochs and it looks like I ran out of steam at epoch 4, though I suppose I could tinker with the learning rate, etc. I'm suspicious that multiple architectures are sufficient and its likely that our R2 plots are very similar. I've seen higher performing model plots (R2) with different architectures (eg. transformer encoder vs cnn1d), but have the same performance boost in targets my models are struggling with.</p>",
              "rawMarkdown": "I trained for six epochs and it looks like I ran out of steam at epoch 4, though I suppose I could tinker with the learning rate, etc. I'm suspicious that multiple architectures are sufficient and its likely that our R2 plots are very similar. I've seen higher performing model plots (R2) with different architectures (eg. transformer encoder vs cnn1d), but have the same performance boost in targets my models are struggling with."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2898653,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-07-01T08:59:26.113000",
      "content": "<p>Okay the models are finally ready, I took the avg of a 0.74984 and a 0.75317 and i got 0.76092, all LBs. I guess if i keep adding models and take their avg the extra value will be less and less right?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2898678,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-07-01T09:24:49.777000",
          "content": "<p>Yes, but you will still get a boost, up to a certain point when it will become negligible. As long as you have compute- the more the merrier (esp. as you got a large initial boost, much better than I anticipated lol)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2898694,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-07-01T09:35:05.863000",
              "content": "<p>Is it a good idea instead of avg to train a metamodel? But then i guess i have to reserve extra data. What boost should i expect then if you compare with just avg</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2902930,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-07-03T14:41:17.447000",
          "content": "<p>I have added a third model with LB 0.74878. The average of the three models achieved went up tp 0.76277, up only a bit from 0.76092 that was with the avg of two models</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2903143,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2024-07-03T16:17:20.647000",
              "content": "<p>That uplift is still higher than I'm seeing so far</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2903277,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-07-03T17:36:52.757000",
              "content": "<p>Not related to the topic, i just discovered by pure luck that reducing massively the LR to train the transformer gave a huge bump to the validation loss, so now i keep training all three of them with tiny LRs. I have put a linear scheduler with a big warmup with a 2 to 4 e-6 LR. So far i saw an improvement from 0.753 to 0.76 in val set. </p>\n<p>Update: I saw all of the improvement reflected in LB and even more. So the model that was at 0.74878 went up to 0.75633 after got refined with small LRs</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2905299,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2024-07-04T19:53:25.987000",
              "content": "<p>Nice! I also saw benefit from tinkering with the LR, though I think I'm running out of runway.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2891294,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-06-26T15:16:17.717000",
      "content": "<p>Yes, ensemble helps (well, from my Kaggle experience, ensembling ALWAYS helps). But 0.01 is a bit too optimistic for an ensemble of two. Expect to get ~0.003. This is at least my experience for ~0.78LB same models. Maybe for 0.74 it would be more.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2893274,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-06-27T18:01:41.947000",
      "content": "<p>The main idea is training multiple models on different subsets of data which is called bagging. Decision tree based models do that inherently but you can do it with any kind of model.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2911728,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-07-08T14:06:58.260000",
      "content": "<p>Hey guys, in the 4th train-val pair the loss is considerably higher in the val set than the other train-val pairs so far. As a matter of fact adding the 4th model to the average it dropped my score from 0.76534 to 0.76412, does this happen often? it it possible that the second lower score generalizes better and achieves a higher score when the whole test set will be used in the end?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2913579,
          "author_name": "Rob Freeman",
          "author_url": "",
          "post_date": "2024-07-09T15:00:13.987000",
          "content": "<p>It doesn't seem that weird to me. It also seems plausible that your largest gains from the ensemble may come from the hard to predict targets (12-15). I suspect that there is something enabling reliable prediction of these targets for the highest LB scores. That said, I don't think the trick for q0002 is stable across months/years in the training data so that may be a source of bad times on the private lb. My assumption is that the baseline U-Net implementation has masking for these targets which enables the consistency on the public/private lb. Maybe it is as simple as KNN to guess whether or not to use the mean or use the trick.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2913784,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-07-09T16:22:59.723000",
              "content": "<blockquote>\n  <p>I don't think the trick for q0002 is stable across months/years in the training data</p>\n</blockquote>\n<p>It's an easy thing to check. No need to 'think' here, you can know.<br>\nAlso easy to check for public LB with some probing.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2914120,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-07-09T18:24:56.700000",
              "content": "<p>replacing the ptend_q0002_12 upt to 27 works well. These are the numbers in the validation set</p>\n<p>R2 for ptend_q0002_12: 1.0<br>\nR2 for ptend_q0002_13: 1.0<br>\nR2 for ptend_q0002_14: 1.0<br>\nR2 for ptend_q0002_15: 1.0<br>\nR2 for ptend_q0002_16: 1.0<br>\nR2 for ptend_q0002_17: 1.0<br>\nR2 for ptend_q0002_18: 1.0<br>\nR2 for ptend_q0002_19: 1.0<br>\nR2 for ptend_q0002_20: 1.0<br>\nR2 for ptend_q0002_21: 1.0<br>\nR2 for ptend_q0002_22: 1.0<br>\nR2 for ptend_q0002_23: 1.0<br>\nR2 for ptend_q0002_24: 1.0<br>\nR2 for ptend_q0002_25: 1.0<br>\nR2 for ptend_q0002_26: 1.0<br>\nR2 for ptend_q0002_27: 1.0</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2914275,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2024-07-09T20:01:44.770000",
              "content": "<p>thanks, i found where my error was</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2914140,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-07-09T18:28:14.797000",
          "content": "<p>the thing is for the 4th test-val pair i collected randomly ids (that i had not in other val sets) every 100k rows. For the other val sets i collected randomly every 1000 rows, so maybe there are highly correlated data across the dataset or we just cover all long lats if we sample every 1000 rows. I am just a bit worried that if there are highly correlated data and then there are not in the hidden test set then my results will get down in the hidden test set</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2891877,
      "author_name": "Rob Freeman",
      "author_url": "",
      "post_date": "2024-06-27T00:56:15.193000",
      "content": "<p>I was wondering how you broke through 0.7, but then I realized something. 😅😅</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2892586,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-06-27T10:32:58.917000",
          "content": "<p>big transformer model trained for more than 10 epochs, and i replace the values of the problematic columns using the state columns as have been discussed in other threads</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2892858,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2024-06-27T13:44:20.443000",
              "content": "<p>Gotcha, yes, the latter is something I missed in my own notebooks and only realized before posting. Also, the transformer encoder models I've created are exceptionally slow to train; I haven't even run a full epoch. Do you think it was worth it versus adjusting model parameters including dropout and running longer with another model?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2893564,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-27T21:49:00.850000",
              "content": "<p>for me a big mind-blowing revelation was that if i process the data batched using polars and also with smaller batch size the model was trained much faster. I suspect because it did not have to move data in and out in cuda all the time. Before the competition i was under the impression that whatever do not blow my cuda memory is good but i notice that with smaller batch size and less stuff in cuda the train can be much faster. Considering dropout i tried without it as suggested in other discussion and also more transformer layers helped, i saw difference from 3 to 6 for example</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2893646,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2024-06-28T00:21:32.843000",
              "content": "<p>Thanks! I appreciate your feedback. I might see what happens with dramatically smaller batches.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2894472,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-28T13:56:38.723000",
              "content": "<p>no need to go very low, i think i use a batch of 256, i read every time 1/100th of the train data using polars, make it into tensor, then start feeding the model, after that data are done i go to the next hundredth. Same for validation</p>\n<pre><code>     :\n        reader = pl.read_csv_batched(train_file, =chunk_size)\n        counter = 0\n\n         :\n\n            prep_chunk_time_start = time.time()\n            try:\n                batches = reader.next_batches(20)\n                 batches[0] is None:\n                    break  #  more data  read\n            except:\n                break\n\n            end_batches = time.time()\n            (, end_batches - prep_chunk_time_start)\n\n             df  batches:\n\n                start_chunk_time = time.time()\n                train_dataset, _ = # make it into tensor here\n\n                train_loader = DataLoader(train_dataset,\n                                          =BATCH_SIZE,\n                                          =,\n                                          =collate_fn,\n                                          )\n</code></pre>\n<p>that way only a small piece is in cuda, and that makes it faster. I train a 15 million parameter transformer using all the train data (9m) and 1m validation in a laptop with an NVIDIA GeForce RTX 2070 Super Max-Q</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2894643,
              "author_name": "Dmitry Uarov",
              "author_url": "",
              "post_date": "2024-06-28T15:57:01.170000",
              "content": "<p>It's cool that you have already achieved such a result! Let me ask you, how do you prepare data before training? Only normalizing like in public notebooks or something else? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2894796,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-28T17:29:56.497000",
              "content": "<p>Thank you! yes nothing special, just (x-mean_x) / std_x and (y -mean_y)/std_y like has mentioned in many other discussions.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2895864,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2024-06-29T12:18:29.540000",
              "content": "<p>Thanks! I still had cuda memory / rematerialization errors except for smaller transformer encoder networks and switched to an LSTM-based model instead.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2896017,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-29T14:33:32.657000",
              "content": "<p>Interesting that LSTM-based models achieve similar result</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2896284,
              "author_name": "Rob Freeman",
              "author_url": "",
              "post_date": "2024-06-29T18:05:24.377000",
              "content": "<p>I trained for six epochs and it looks like I ran out of steam at epoch 4, though I suppose I could tinker with the learning rate, etc. I'm suspicious that multiple architectures are sufficient and its likely that our R2 plots are very similar. I've seen higher performing model plots (R2) with different architectures (eg. transformer encoder vs cnn1d), but have the same performance boost in targets my models are struggling with.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2898653": "Okay the models are finally ready, I took the avg of a 0.74984 and a 0.75317 and i got 0.76092, all LBs. I guess if i keep adding models and take their avg the extra value will be less and less right?",
    "2891276": "Hello guys, my single model trained on 9m data validated at 1m data achives around 0.74 LB score. If i train another one with a totally new validation set of 1m samples is it expected that taking the average result of the two models will perform significantly better? Like at least 0.01+? The train takes too long and i could spend the time in other experiments if the ensemble is not very promising.",
    "2891294": "Yes, ensemble helps (well, from my Kaggle experience, ensembling ALWAYS helps). But 0.01 is a bit too optimistic for an ensemble of two. Expect to get ~0.003. This is at least my experience for ~0.78LB same models. Maybe for 0.74 it would be more.",
    "2893274": "The main idea is training multiple models on different subsets of data which is called bagging. Decision tree based models do that inherently but you can do it with any kind of model.",
    "2911728": "Hey guys, in the 4th train-val pair the loss is considerably higher in the val set than the other train-val pairs so far. As a matter of fact adding the 4th model to the average it dropped my score from 0.76534 to 0.76412, does this happen often? it it possible that the second lower score generalizes better and achieves a higher score when the whole test set will be used in the end?",
    "2891877": "I was wondering how you broke through 0.7, but then I realized something. 😅😅"
  }
}