{
  "id": 501829,
  "title": "The secret for beating baseline",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/501829",
  "author_name": "greySnow",
  "post_date": "2024-05-10T23:06:07.625000",
  "votes": 46,
  "comment_count": 43,
  "views": 0,
  "content": "<p>I struggled to get my seq-to-seq i.e. [60,dim]-&gt;[60, cols_target_dim=14] models to converge. Judging by the LB, I believe I'm not the only one to experience this struggle. Be it transformer, 1dconv, or Unet, they will get, at best, to ~0.3 val. While the dense models i.e. [dim]-&gt;[targets_dim = 368] easily converge to ~500-600 val. Then I had the idea…to treat this as seq-to-multiple targets instead of seq-to-seq, i.e. [60,dim]-&gt;[pool on col_axis]-&gt;[targets_dim = 368]. And suddenly all my models converge easily, transformers Unet etc. Well, it still requires some work to beat the baseline. You need a good model, etc., but this is the main point. With this I believe it is not too hard to beat the baseline (getting to 0.7 range is a different story though lol). Anyway, good luck.</p>\n<h1>EDIT:</h1>\n<p>After <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> said that seq2seq worked for him, I suspected I had a bug. So I revisited my code, and indeed, I found a bug exactly where I suspected. It turns out I missed a transposed somewhere. So, seq2seq works, too…if you don't have a bug in your code 🤣. My current training with seq2seq shows powerful results, better than my seq-to-multiple targets. I am sorry about the misleading above. Well, this is why I love sharing on Kaggle. If you share, it eventually helps you even more than you helped the others; I swear it's true haha.</p>",
  "messages": [
    {
      "id": 2806144,
      "postDate": "2024-05-10T23:06:07.627Z",
      "content": "<p>I struggled to get my seq-to-seq i.e. [60,dim]-&gt;[60, cols_target_dim=14] models to converge. Judging by the LB, I believe I'm not the only one to experience this struggle. Be it transformer, 1dconv, or Unet, they will get, at best, to ~0.3 val. While the dense models i.e. [dim]-&gt;[targets_dim = 368] easily converge to ~500-600 val. Then I had the idea…to treat this as seq-to-multiple targets instead of seq-to-seq, i.e. [60,dim]-&gt;[pool on col_axis]-&gt;[targets_dim = 368]. And suddenly all my models converge easily, transformers Unet etc. Well, it still requires some work to beat the baseline. You need a good model, etc., but this is the main point. With this I believe it is not too hard to beat the baseline (getting to 0.7 range is a different story though lol). Anyway, good luck.</p>\n<h1>EDIT:</h1>\n<p>After <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> said that seq2seq worked for him, I suspected I had a bug. So I revisited my code, and indeed, I found a bug exactly where I suspected. It turns out I missed a transposed somewhere. So, seq2seq works, too…if you don't have a bug in your code 🤣. My current training with seq2seq shows powerful results, better than my seq-to-multiple targets. I am sorry about the misleading above. Well, this is why I love sharing on Kaggle. If you share, it eventually helps you even more than you helped the others; I swear it's true haha.</p>",
      "rawMarkdown": "I struggled to get my seq-to-seq i.e. [60,dim]->[60, cols_target_dim=14] models to converge. Judging by the LB, I believe I'm not the only one to experience this struggle. Be it transformer, 1dconv, or Unet, they will get, at best, to ~0.3 val. While the dense models i.e. [dim]->[targets_dim = 368] easily converge to ~500-600 val. Then I had the idea...to treat this as seq-to-multiple targets instead of seq-to-seq, i.e. [60,dim]->[pool on col_axis]->[targets_dim = 368]. And suddenly all my models converge easily, transformers Unet etc. Well, it still requires some work to beat the baseline. You need a good model, etc., but this is the main point. With this I believe it is not too hard to beat the baseline (getting to 0.7 range is a different story though lol). Anyway, good luck.\n\n#EDIT:\nAfter @ulrich07 said that seq2seq worked for him, I suspected I had a bug. So I revisited my code, and indeed, I found a bug exactly where I suspected. It turns out I missed a transposed somewhere. So, seq2seq works, too...if you don't have a bug in your code 🤣. My current training with seq2seq shows powerful results, better than my seq-to-multiple targets. I am sorry about the misleading above. Well, this is why I love sharing on Kaggle. If you share, it eventually helps you even more than you helped the others; I swear it's true haha.",
      "votes": 46
    },
    {
      "id": 3034424,
      "postDate": "2024-11-02T06:04:48.747Z",
      "content": "<p>transformer, 1dconv, or Unet :) You are rieght!</p>",
      "rawMarkdown": "transformer, 1dconv, or Unet :) You are rieght!",
      "votes": 1
    },
    {
      "id": 2806793,
      "postDate": "2024-05-11T09:49:01.567Z",
      "content": "<p>seq to seq worked for me.</p>",
      "rawMarkdown": "seq to seq worked for me.",
      "votes": 5,
      "replies": [
        {
          "id": 2806814,
          "postDate": "2024-05-11T10:09:37.097Z",
          "content": "<p>Interesting, good to know</p>",
          "rawMarkdown": "Interesting, good to know"
        },
        {
          "id": 2807915,
          "postDate": "2024-05-12T01:08:09.017Z",
          "content": "<p>You are right. See my edit and thank you.</p>",
          "rawMarkdown": "You are right. See my edit and thank you."
        }
      ]
    },
    {
      "id": 2823653,
      "postDate": "2024-05-19T10:43:35.313Z",
      "content": "<p>So following the advice i made my data as a sequence and since i do not have experience in implementing transformers i started with a very simple Transformer that maps sequence to scalars and it does not do anything with position (i think not needed in this problem). The thing is that i let it train for many hours and i notice the train and val R2 to drop very very slowly. My Transformer has only like 4 million params does this make sense? the error started 0.667 and after almost 20 hours it is 0.524. Actually i just noticed that it has only done 4k iterations, that is too slow. Are transformers much slower to train than MLPs?</p>",
      "rawMarkdown": "So following the advice i made my data as a sequence and since i do not have experience in implementing transformers i started with a very simple Transformer that maps sequence to scalars and it does not do anything with position (i think not needed in this problem). The thing is that i let it train for many hours and i notice the train and val R2 to drop very very slowly. My Transformer has only like 4 million params does this make sense? the error started 0.667 and after almost 20 hours it is 0.524. Actually i just noticed that it has only done 4k iterations, that is too slow. Are transformers much slower to train than MLPs?",
      "votes": 3,
      "replies": [
        {
          "id": 2824035,
          "postDate": "2024-05-19T14:48:39.213Z",
          "content": "<p>loss.backward() and optimizer.step() takes on average 9-10 secs does this make sense?</p>",
          "rawMarkdown": "loss.backward() and optimizer.step() takes on average 9-10 secs does this make sense?",
          "votes": 1
        },
        {
          "id": 2824072,
          "postDate": "2024-05-19T15:09:37.837Z",
          "content": "<p>'Are transformers much slower to train than MLP' depend on the parameters, number of layers hidden dimension etc, but generally yes, transformer takes more time. Do you train on GPU? Also remember that for vanilla transformer you have to include some sort of positional encoding. Getting started with transformers can be hard, so find some good example for transformer seq-to-seq and make sure you have a good grasp over it before trying to apply to this problem.</p>",
          "rawMarkdown": "'Are transformers much slower to train than MLP' depend on the parameters, number of layers hidden dimension etc, but generally yes, transformer takes more time. Do you train on GPU? Also remember that for vanilla transformer you have to include some sort of positional encoding. Getting started with transformers can be hard, so find some good example for transformer seq-to-seq and make sure you have a good grasp over it before trying to apply to this problem.",
          "votes": 1,
          "replies": [
            {
              "id": 2824257,
              "postDate": "2024-05-19T16:55:10.067Z",
              "content": "<p>I train on a GPU, but does it make sense the prediction time to be like 2-3 secs and then the backward to be 10 secs? </p>",
              "rawMarkdown": "I train on a GPU, but does it make sense the prediction time to be like 2-3 secs and then the backward to be 10 secs? "
            },
            {
              "id": 2824284,
              "postDate": "2024-05-19T17:24:49.390Z",
              "content": "<p>IDK how to answer that since I work with tensorflow model.fit :/</p>",
              "rawMarkdown": "IDK how to answer that since I work with tensorflow model.fit :/"
            }
          ]
        },
        {
          "id": 2824591,
          "postDate": "2024-05-19T21:41:14.193Z",
          "content": "<p>I am trying to learn seq2seq, what do you mean by \"I made my data as a sequence\"?</p>\n<p>Considering the default (rows, columns), What do you need to change in the shape of the data?</p>",
          "rawMarkdown": "I am trying to learn seq2seq, what do you mean by \"I made my data as a sequence\"?\n\nConsidering the default (rows, columns), What do you need to change in the shape of the data?",
          "replies": [
            {
              "id": 2825365,
              "postDate": "2024-05-20T10:46:58.333Z",
              "content": "<p>so i use polars to load the dataframe and then i use this code</p>\n<pre><code> = [, , ,\n                                 , , ,\n                                 , , ]  \n\n = [, , , ,\n                    , , , , ,\n                    , , , , ,\n                    , ]  \n\n = [, , , , , ]\n = [, , , , ,\n                      , , ]\n</code></pre>",
              "rawMarkdown": "so i use polars to load the dataframe and then i use this code\n\n```\nseq_variables_x = ['state_t', 'state_q0001', 'state_q0002',\n                                 'state_q0003', 'state_u', 'state_v',\n                                 'pbuf_ozone', 'pbuf_CH4', 'pbuf_N2O']  # Example vertically resolved variables\n\nscalar_variables_x = ['state_ps', 'pbuf_SOLIN', 'pbuf_LHFLX', 'pbuf_SHFLX',\n                    'pbuf_TAUX', 'pbuf_TAUY', 'pbuf_COSZRS', 'cam_in_ALDIF', 'cam_in_ALDIR',\n                    'cam_in_ASDIF', 'cam_in_ASDIR', 'cam_in_LWUP', 'cam_in_ICEFRAC', 'cam_in_LANDFRAC',\n                    'cam_in_OCNFRAC', 'cam_in_SNOWHLAND']  # Example scalar variables\n\nseq_variables_y = ['ptend_t', 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', 'ptend_v']\nscalar_variables_y = ['cam_out_NETSW', 'cam_out_FLWDS', 'cam_out_PRECSC', 'cam_out_PRECC', 'cam_out_SOLS',\n                      'cam_out_SOLL', 'cam_out_SOLSD', 'cam_out_SOLLD']\n\n```",
              "votes": 2
            },
            {
              "id": 2825377,
              "postDate": "2024-05-20T10:51:55.037Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2825397,
              "postDate": "2024-05-20T11:28:39.273Z",
              "content": "<pre><code> ():\n    num_variables = (seq_variables_x) + (scalar_variables_x)\n\n    data_X = np.zeros((batch_size, sequence_length, num_variables))\n\n    \n     i, var  (seq_variables_x):\n        columns = [  level  (sequence_length)]\n        data_X[:, :, i] = df[columns].to_numpy()\n\n    \n     j, var  (scalar_variables_x):\n        data_X[:, :, (seq_variables_x) + j] = df[var].to_numpy()[:, np.newaxis]\n\n    \n    tensor_data = torch.tensor(data_X, dtype=torch.float32).cuda()\n\n     tensor_data\n</code></pre>",
              "rawMarkdown": "```\ndef to_tensor(df, batch_size, sequence_length, seq_variables_x, scalar_variables_x):\n    num_variables = len(seq_variables_x) + len(scalar_variables_x)\n\n    data_X = np.zeros((batch_size, sequence_length, num_variables))\n\n    # Process vertically resolved variables\n    for i, var in enumerate(seq_variables_x):\n        columns = [f\"{var}_{level}\" for level in range(sequence_length)]\n        data_X[:, :, i] = df[columns].to_numpy()\n\n    # Process scalar variables\n    for j, var in enumerate(scalar_variables_x):\n        data_X[:, :, len(seq_variables_x) + j] = df[var].to_numpy()[:, np.newaxis]\n\n    # Convert the numpy array to a PyTorch tensor\n    tensor_data = torch.tensor(data_X, dtype=torch.float32).cuda()\n\n    return tensor_data\n\n```",
              "votes": 1
            },
            {
              "id": 2825398,
              "postDate": "2024-05-20T11:29:27.250Z",
              "content": "<p>So then i feed this data to a pytorch nn.Transformer. The batch size in this case is the length of all my data</p>",
              "rawMarkdown": "So then i feed this data to a pytorch nn.Transformer. The batch size in this case is the length of all my data",
              "votes": 1
            },
            {
              "id": 2825407,
              "postDate": "2024-05-20T11:35:12.557Z",
              "content": "<p>i dont know if i do something wrong but 100 iterations take 1350 seconds so its not practical for me to train that transformer. By the way its not seq2seq but seq to scalar</p>",
              "rawMarkdown": "i dont know if i do something wrong but 100 iterations take 1350 seconds so its not practical for me to train that transformer. By the way its not seq2seq but seq to scalar",
              "votes": 1
            },
            {
              "id": 2845428,
              "postDate": "2024-05-30T14:58:58.267Z",
              "content": "<p>How are you converting the targets back to submission file?</p>",
              "rawMarkdown": "How are you converting the targets back to submission file?"
            }
          ]
        }
      ]
    },
    {
      "id": 2806222,
      "postDate": "2024-05-11T01:12:06.110Z",
      "content": "<p>I have been struggling to get 0.7+ for 5 days after I got 0.68, however others who got 0.68 at the same time with me have already got 0.72+… So hard 🤣</p>",
      "rawMarkdown": "I have been struggling to get 0.7+ for 5 days after I got 0.68, however others who got 0.68 at the same time with me have already got 0.72+... So hard 🤣",
      "votes": 4
    },
    {
      "id": 2843127,
      "postDate": "2024-05-29T12:27:41.940Z",
      "content": "<p>Thank you. Can you tell me the number of parameters in your model?</p>",
      "rawMarkdown": "Thank you. Can you tell me the number of parameters in your model?",
      "votes": 1,
      "replies": [
        {
          "id": 2843174,
          "postDate": "2024-05-29T12:48:10.427Z",
          "content": "<p>For beating the baseline 5.8M</p>",
          "rawMarkdown": "For beating the baseline 5.8M",
          "votes": 2
        }
      ]
    },
    {
      "id": 2835810,
      "postDate": "2024-05-25T14:34:16.160Z",
      "content": "<p>seq to seq worked for me. thanks for your share.  i have a question。  the some weights of sample_sub are zero，also the labels of zero weight are not  used。  why not delete those labels in seq?   I don't know if these labels are useful for the net</p>",
      "rawMarkdown": "seq to seq worked for me. thanks for your share.  i have a question。  the some weights of sample_sub are zero，also the labels of zero weight are not  used。  why not delete those labels in seq?   I don't know if these labels are useful for the net",
      "votes": 1
    },
    {
      "id": 2807773,
      "postDate": "2024-05-11T20:31:42.453Z",
      "content": "<p>In this competition, seq2seq is like in NLP where you have an input of size N (number of features in this case) and the output is the number of features (368)?<br>\nLike \"translating\" features to the outputs?</p>",
      "rawMarkdown": "In this competition, seq2seq is like in NLP where you have an input of size N (number of features in this case) and the output is the number of features (368)?\nLike \"translating\" features to the outputs?",
      "votes": 2,
      "replies": [
        {
          "id": 2807781,
          "postDate": "2024-05-11T20:47:34.557Z",
          "content": "<p>No. Seq to seq is when we look at the features and targets along the height axis i.e. [60,dim]</p>",
          "rawMarkdown": "No. Seq to seq is when we look at the features and targets along the height axis i.e. [60,dim]",
          "votes": 2
        }
      ]
    },
    {
      "id": 2806838,
      "postDate": "2024-05-11T10:43:53.973Z",
      "content": "<p>What do you mean by seq2seq? Transformer encoder and rnn head?</p>",
      "rawMarkdown": "What do you mean by seq2seq? Transformer encoder and rnn head?",
      "replies": [
        {
          "id": 2806849,
          "postDate": "2024-05-11T10:56:41.690Z",
          "content": "<p>Exactly as I wrote,  [60,dim]-&gt;[60, cols_target_dim=14]. Like in semantic segmentation.</p>",
          "rawMarkdown": "Exactly as I wrote,  [60,dim]->[60, cols_target_dim=14]. Like in semantic segmentation.",
          "votes": 4
        }
      ]
    },
    {
      "id": 2807300,
      "postDate": "2024-05-11T15:50:48.653Z",
      "content": "<p>What do you mean by seq2seq? Transformer encoder and rnn head?😃</p>",
      "rawMarkdown": "What do you mean by seq2seq? Transformer encoder and rnn head?😃",
      "votes": -10,
      "replies": [
        {
          "id": 2807321,
          "postDate": "2024-05-11T16:03:34.247Z",
          "content": "<p>[60,dim]-&gt;[60, cols_target_dim=14] = 1d semantic segmentation</p>",
          "rawMarkdown": "[60,dim]->[60, cols_target_dim=14] = 1d semantic segmentation",
          "votes": 1,
          "replies": [
            {
              "id": 2822026,
              "postDate": "2024-05-18T12:02:05.823Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2821910,
      "postDate": "2024-05-18T10:28:28.940Z",
      "content": "<p>Hello, thanks for this helpful post, is seq 2 seq better than just a transformer to map the input to 368 targets?</p>",
      "rawMarkdown": "Hello, thanks for this helpful post, is seq 2 seq better than just a transformer to map the input to 368 targets?",
      "votes": -1,
      "replies": [
        {
          "id": 2821991,
          "postDate": "2024-05-18T11:08:17.550Z",
          "content": "<p>Yes, mapping to 60*14 give me better results</p>",
          "rawMarkdown": "Yes, mapping to 60*14 give me better results",
          "votes": 1,
          "replies": [
            {
              "id": 2822027,
              "postDate": "2024-05-18T12:02:48.393Z",
              "content": "<p>a bit counter intuitive since the sequence length is always fixed at 60 and the positions are fixed always, so i dont get why seq 2 seq is better</p>",
              "rawMarkdown": "a bit counter intuitive since the sequence length is always fixed at 60 and the positions are fixed always, so i dont get why seq 2 seq is better"
            },
            {
              "id": 2822056,
              "postDate": "2024-05-18T12:22:18.150Z",
              "content": "<p>I think it's because the data can flow more easily. In [dim]-&gt;[368], the model needs to compress all the data to [dim] and then predict. Whereas in [60, dim]-&gt;[60, 14], the model has more 'space'. Well, anyway, seq-to-seq is usually used in these kinds of problems with the greatest success; if it works, it works.</p>",
              "rawMarkdown": "I think it's because the data can flow more easily. In [dim]->[368], the model needs to compress all the data to [dim] and then predict. Whereas in [60, dim]->[60, 14], the model has more 'space'. Well, anyway, seq-to-seq is usually used in these kinds of problems with the greatest success; if it works, it works.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2842982,
      "postDate": "2024-05-29T11:27:20.680Z",
      "content": "<p>Sorry to bother you, I am a beginner. May I ask if you mean the input shape is [batch_size, 60, dim]?</p>",
      "rawMarkdown": "Sorry to bother you, I am a beginner. May I ask if you mean the input shape is [batch_size, 60, dim]?",
      "replies": [
        {
          "id": 2843020,
          "postDate": "2024-05-29T11:44:08.137Z",
          "content": "<p>The input shape is [batch_size, 60, input_dim=9+16=25] (look at features list to understand why)<br>\nThen there is the 'dim' a.k.e. hidden_dim which is the inner hidden dimension of the networks, usually 128/192/256 etc.<br>\nAnd finally there is the 'output_dim'=14</p>",
          "rawMarkdown": "The input shape is [batch_size, 60, input_dim=9+16=25] (look at features list to understand why)\nThen there is the 'dim' a.k.e. hidden_dim which is the inner hidden dimension of the networks, usually 128/192/256 etc.\nAnd finally there is the 'output_dim'=14",
          "votes": 3,
          "replies": [
            {
              "id": 2843065,
              "postDate": "2024-05-29T11:55:32.213Z",
              "content": "<p>Your response is very detailed. I am starting to try this approach now. Thank you so much.</p>",
              "rawMarkdown": "Your response is very detailed. I am starting to try this approach now. Thank you so much."
            }
          ]
        }
      ]
    },
    {
      "id": 2834126,
      "postDate": "2024-05-24T15:02:02.080Z",
      "content": "<p>How did you handle the 16 global variables in the input? Did you copy them 60 times and then concatenate them to form [60, 9 + 16], or did you use a different approach?</p>",
      "rawMarkdown": "How did you handle the 16 global variables in the input? Did you copy them 60 times and then concatenate them to form [60, 9 + 16], or did you use a different approach?",
      "replies": [
        {
          "id": 2834130,
          "postDate": "2024-05-24T15:04:12.947Z",
          "content": "<p>Yes.[ 60, 9 + 16]</p>",
          "rawMarkdown": "Yes.[ 60, 9 + 16]",
          "votes": 2,
          "replies": [
            {
              "id": 2834752,
              "postDate": "2024-05-25T02:11:12.057Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!"
            }
          ]
        }
      ]
    },
    {
      "id": 2814118,
      "postDate": "2024-05-15T07:06:12.940Z",
      "content": "<p>Did you implement a Bidirectional LSTM? I managed to get a lstm working but the results are not really good. </p>\n<pre><code>input_shape = (, dim)\noutput_shape = (, cols_target_dim)\n\n ():\n    \n    encoder_inputs = Input(shape=input_shape)\n    encoder = Bidirectional(LSTM(units=latent_dim, return_sequences=))\n    encoder_outputs = encoder(encoder_inputs)\n\n    \n    decoder_lstm = LSTM(units=latent_dim, return_sequences=)\n    decoder_outputs = decoder_lstm(encoder_outputs)\n\n    \n    output_layer = Dense(cols_target_dim, activation=)\n    outputs = output_layer(decoder_outputs)\n\n     Model(encoder_inputs, outputs)\n\nmodel = seq2seq_model(latent_dim=)  \n</code></pre>",
      "rawMarkdown": "Did you implement a Bidirectional LSTM? I managed to get a lstm working but the results are not really good. \n```python\ninput_shape = (60, dim)\noutput_shape = (60, cols_target_dim)\n\ndef seq2seq_model(latent_dim):\n    # Encoder\n    encoder_inputs = Input(shape=input_shape)\n    encoder = Bidirectional(LSTM(units=latent_dim, return_sequences=True))\n    encoder_outputs = encoder(encoder_inputs)\n    \n    # Decoder\n    decoder_lstm = LSTM(units=latent_dim, return_sequences=True)\n    decoder_outputs = decoder_lstm(encoder_outputs)\n    \n    # Output layer\n    output_layer = Dense(cols_target_dim, activation='linear')\n    outputs = output_layer(decoder_outputs)\n\n    return Model(encoder_inputs, outputs)\n\nmodel = seq2seq_model(latent_dim=64)  # Choose the number of units for the LSTM layer\n```",
      "replies": [
        {
          "id": 2814129,
          "postDate": "2024-05-15T07:27:01.310Z",
          "content": "<p>Try 1D Unet or transformers. </p>",
          "rawMarkdown": "Try 1D Unet or transformers. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 2813007,
      "postDate": "2024-05-14T14:26:44.243Z",
      "content": "<p>Another question:<br>\nIs it possible to beat the market with a single RTX 4090?</p>",
      "rawMarkdown": "Another question:\nIs it possible to beat the market with a single RTX 4090?\n"
    },
    {
      "id": 2810909,
      "postDate": "2024-05-13T13:59:16.487Z",
      "content": "<p>I'm trying to learn from your post and understand how seq2seq and seq2multipletargets work on this dataset, but I'm struggling to understand it:</p>\n<ol>\n<li>In seq2seq you have 14 target features as said in the Data section, but how are 368 target features converted to 14? I don't understand this completely.</li>\n<li>When we send 368 output targets, how are they converted to 14?</li>\n</ol>",
      "rawMarkdown": "I'm trying to learn from your post and understand how seq2seq and seq2multipletargets work on this dataset, but I'm struggling to understand it:\n1. In seq2seq you have 14 target features as said in the Data section, but how are 368 target features converted to 14? I don't understand this completely.\n2. When we send 368 output targets, how are they converted to 14?",
      "replies": [
        {
          "id": 2810931,
          "postDate": "2024-05-13T14:12:58.157Z",
          "content": "<p>The sequence length is 60. So 14*60=840. But 8 out of the 14 are global. So we average them and then 60*6+8 = 368 </p>",
          "rawMarkdown": "The sequence length is 60. So 14\\*60=840. But 8 out of the 14 are global. So we average them and then 60\\*6+8 = 368 ",
          "votes": 2,
          "replies": [
            {
              "id": 2810949,
              "postDate": "2024-05-13T14:32:37.103Z",
              "content": "<p>Thank you for your reply. </p>\n<p>If I understand correctly, the output of the seq2seq model has a shape of ((60, 14) . For the global targets (i.e., cam_out_NETSW, cam_out_FLWDS, cam_out_PRECSC…), you average them to have one value for each, and for the others, you maintain the regular 60.</p>\n<p>Your post is enlightening.</p>",
              "rawMarkdown": "Thank you for your reply. \n\nIf I understand correctly, the output of the seq2seq model has a shape of ((60, 14) . For the global targets (i.e., cam_out_NETSW, cam_out_FLWDS, cam_out_PRECSC…), you average them to have one value for each, and for the others, you maintain the regular 60.\n\nYour post is enlightening.",
              "votes": 4
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3034424,
      "author_name": "Nandan Nilekani",
      "author_url": "",
      "post_date": "2024-11-02T06:04:48.747000",
      "content": "<p>transformer, 1dconv, or Unet :) You are rieght!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2806793,
      "author_name": "Ulrich G.",
      "author_url": "",
      "post_date": "2024-05-11T09:49:01.567000",
      "content": "<p>seq to seq worked for me.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2806814,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-11T10:09:37.097000",
          "content": "<p>Interesting, good to know</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2807915,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-12T01:08:09.017000",
          "content": "<p>You are right. See my edit and thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2823653,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-05-19T10:43:35.313000",
      "content": "<p>So following the advice i made my data as a sequence and since i do not have experience in implementing transformers i started with a very simple Transformer that maps sequence to scalars and it does not do anything with position (i think not needed in this problem). The thing is that i let it train for many hours and i notice the train and val R2 to drop very very slowly. My Transformer has only like 4 million params does this make sense? the error started 0.667 and after almost 20 hours it is 0.524. Actually i just noticed that it has only done 4k iterations, that is too slow. Are transformers much slower to train than MLPs?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2824035,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-05-19T14:48:39.213000",
          "content": "<p>loss.backward() and optimizer.step() takes on average 9-10 secs does this make sense?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2824072,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-19T15:09:37.837000",
          "content": "<p>'Are transformers much slower to train than MLP' depend on the parameters, number of layers hidden dimension etc, but generally yes, transformer takes more time. Do you train on GPU? Also remember that for vanilla transformer you have to include some sort of positional encoding. Getting started with transformers can be hard, so find some good example for transformer seq-to-seq and make sure you have a good grasp over it before trying to apply to this problem.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2824257,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-05-19T16:55:10.067000",
              "content": "<p>I train on a GPU, but does it make sense the prediction time to be like 2-3 secs and then the backward to be 10 secs? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2824284,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-05-19T17:24:49.390000",
              "content": "<p>IDK how to answer that since I work with tensorflow model.fit :/</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2824591,
          "author_name": "Fernando Melo",
          "author_url": "",
          "post_date": "2024-05-19T21:41:14.193000",
          "content": "<p>I am trying to learn seq2seq, what do you mean by \"I made my data as a sequence\"?</p>\n<p>Considering the default (rows, columns), What do you need to change in the shape of the data?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2825365,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-05-20T10:46:58.333000",
              "content": "<p>so i use polars to load the dataframe and then i use this code</p>\n<pre><code> = [, , ,\n                                 , , ,\n                                 , , ]  \n\n = [, , , ,\n                    , , , , ,\n                    , , , , ,\n                    , ]  \n\n = [, , , , , ]\n = [, , , , ,\n                      , , ]\n</code></pre>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2825377,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-05-20T10:51:55.037000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2825397,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-05-20T11:28:39.273000",
              "content": "<pre><code> ():\n    num_variables = (seq_variables_x) + (scalar_variables_x)\n\n    data_X = np.zeros((batch_size, sequence_length, num_variables))\n\n    \n     i, var  (seq_variables_x):\n        columns = [  level  (sequence_length)]\n        data_X[:, :, i] = df[columns].to_numpy()\n\n    \n     j, var  (scalar_variables_x):\n        data_X[:, :, (seq_variables_x) + j] = df[var].to_numpy()[:, np.newaxis]\n\n    \n    tensor_data = torch.tensor(data_X, dtype=torch.float32).cuda()\n\n     tensor_data\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2825398,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-05-20T11:29:27.250000",
              "content": "<p>So then i feed this data to a pytorch nn.Transformer. The batch size in this case is the length of all my data</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2825407,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-05-20T11:35:12.557000",
              "content": "<p>i dont know if i do something wrong but 100 iterations take 1350 seconds so its not practical for me to train that transformer. By the way its not seq2seq but seq to scalar</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2845428,
              "author_name": "Fernando Melo",
              "author_url": "",
              "post_date": "2024-05-30T14:58:58.267000",
              "content": "<p>How are you converting the targets back to submission file?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2806222,
      "author_name": "嘴爷",
      "author_url": "",
      "post_date": "2024-05-11T01:12:06.110000",
      "content": "<p>I have been struggling to get 0.7+ for 5 days after I got 0.68, however others who got 0.68 at the same time with me have already got 0.72+… So hard 🤣</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2843127,
      "author_name": "Zhuoqun Li",
      "author_url": "",
      "post_date": "2024-05-29T12:27:41.940000",
      "content": "<p>Thank you. Can you tell me the number of parameters in your model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2843174,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-29T12:48:10.427000",
          "content": "<p>For beating the baseline 5.8M</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2835810,
      "author_name": "Yousef",
      "author_url": "",
      "post_date": "2024-05-25T14:34:16.160000",
      "content": "<p>seq to seq worked for me. thanks for your share.  i have a question。  the some weights of sample_sub are zero，also the labels of zero weight are not  used。  why not delete those labels in seq?   I don't know if these labels are useful for the net</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2807773,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-05-11T20:31:42.453000",
      "content": "<p>In this competition, seq2seq is like in NLP where you have an input of size N (number of features in this case) and the output is the number of features (368)?<br>\nLike \"translating\" features to the outputs?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2807781,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-11T20:47:34.557000",
          "content": "<p>No. Seq to seq is when we look at the features and targets along the height axis i.e. [60,dim]</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2806838,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-05-11T10:43:53.973000",
      "content": "<p>What do you mean by seq2seq? Transformer encoder and rnn head?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2806849,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-11T10:56:41.690000",
          "content": "<p>Exactly as I wrote,  [60,dim]-&gt;[60, cols_target_dim=14]. Like in semantic segmentation.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2807300,
      "author_name": "Sheema Zain",
      "author_url": "",
      "post_date": "2024-05-11T15:50:48.653000",
      "content": "<p>What do you mean by seq2seq? Transformer encoder and rnn head?😃</p>",
      "votes": -10,
      "replies": [
        {
          "id": 2807321,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-11T16:03:34.247000",
          "content": "<p>[60,dim]-&gt;[60, cols_target_dim=14] = 1d semantic segmentation</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2822026,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-05-18T12:02:05.823000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2821910,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-05-18T10:28:28.940000",
      "content": "<p>Hello, thanks for this helpful post, is seq 2 seq better than just a transformer to map the input to 368 targets?</p>",
      "votes": -1,
      "replies": [
        {
          "id": 2821991,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-18T11:08:17.550000",
          "content": "<p>Yes, mapping to 60*14 give me better results</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2822027,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-05-18T12:02:48.393000",
              "content": "<p>a bit counter intuitive since the sequence length is always fixed at 60 and the positions are fixed always, so i dont get why seq 2 seq is better</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2822056,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-05-18T12:22:18.150000",
              "content": "<p>I think it's because the data can flow more easily. In [dim]-&gt;[368], the model needs to compress all the data to [dim] and then predict. Whereas in [60, dim]-&gt;[60, 14], the model has more 'space'. Well, anyway, seq-to-seq is usually used in these kinds of problems with the greatest success; if it works, it works.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2842982,
      "author_name": "Leon",
      "author_url": "",
      "post_date": "2024-05-29T11:27:20.680000",
      "content": "<p>Sorry to bother you, I am a beginner. May I ask if you mean the input shape is [batch_size, 60, dim]?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2843020,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-29T11:44:08.137000",
          "content": "<p>The input shape is [batch_size, 60, input_dim=9+16=25] (look at features list to understand why)<br>\nThen there is the 'dim' a.k.e. hidden_dim which is the inner hidden dimension of the networks, usually 128/192/256 etc.<br>\nAnd finally there is the 'output_dim'=14</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2843065,
              "author_name": "Leon",
              "author_url": "",
              "post_date": "2024-05-29T11:55:32.213000",
              "content": "<p>Your response is very detailed. I am starting to try this approach now. Thank you so much.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2834126,
      "author_name": "Zhuoqun Li",
      "author_url": "",
      "post_date": "2024-05-24T15:02:02.080000",
      "content": "<p>How did you handle the 16 global variables in the input? Did you copy them 60 times and then concatenate them to form [60, 9 + 16], or did you use a different approach?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2834130,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-24T15:04:12.947000",
          "content": "<p>Yes.[ 60, 9 + 16]</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2834752,
              "author_name": "Zhuoqun Li",
              "author_url": "",
              "post_date": "2024-05-25T02:11:12.057000",
              "content": "<p>Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2814118,
      "author_name": "Carlos Pérez Ricardo",
      "author_url": "",
      "post_date": "2024-05-15T07:06:12.940000",
      "content": "<p>Did you implement a Bidirectional LSTM? I managed to get a lstm working but the results are not really good. </p>\n<pre><code>input_shape = (, dim)\noutput_shape = (, cols_target_dim)\n\n ():\n    \n    encoder_inputs = Input(shape=input_shape)\n    encoder = Bidirectional(LSTM(units=latent_dim, return_sequences=))\n    encoder_outputs = encoder(encoder_inputs)\n\n    \n    decoder_lstm = LSTM(units=latent_dim, return_sequences=)\n    decoder_outputs = decoder_lstm(encoder_outputs)\n\n    \n    output_layer = Dense(cols_target_dim, activation=)\n    outputs = output_layer(decoder_outputs)\n\n     Model(encoder_inputs, outputs)\n\nmodel = seq2seq_model(latent_dim=)  \n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 2814129,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-15T07:27:01.310000",
          "content": "<p>Try 1D Unet or transformers. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2813007,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-05-14T14:26:44.243000",
      "content": "<p>Another question:<br>\nIs it possible to beat the market with a single RTX 4090?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2810909,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-05-13T13:59:16.487000",
      "content": "<p>I'm trying to learn from your post and understand how seq2seq and seq2multipletargets work on this dataset, but I'm struggling to understand it:</p>\n<ol>\n<li>In seq2seq you have 14 target features as said in the Data section, but how are 368 target features converted to 14? I don't understand this completely.</li>\n<li>When we send 368 output targets, how are they converted to 14?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 2810931,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-13T14:12:58.157000",
          "content": "<p>The sequence length is 60. So 14*60=840. But 8 out of the 14 are global. So we average them and then 60*6+8 = 368 </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2810949,
              "author_name": "Fernando Melo",
              "author_url": "",
              "post_date": "2024-05-13T14:32:37.103000",
              "content": "<p>Thank you for your reply. </p>\n<p>If I understand correctly, the output of the seq2seq model has a shape of ((60, 14) . For the global targets (i.e., cam_out_NETSW, cam_out_FLWDS, cam_out_PRECSC…), you average them to have one value for each, and for the others, you maintain the regular 60.</p>\n<p>Your post is enlightening.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2806144": "I struggled to get my seq-to-seq i.e. [60,dim]->[60, cols_target_dim=14] models to converge. Judging by the LB, I believe I'm not the only one to experience this struggle. Be it transformer, 1dconv, or Unet, they will get, at best, to ~0.3 val. While the dense models i.e. [dim]->[targets_dim = 368] easily converge to ~500-600 val. Then I had the idea...to treat this as seq-to-multiple targets instead of seq-to-seq, i.e. [60,dim]->[pool on col_axis]->[targets_dim = 368]. And suddenly all my models converge easily, transformers Unet etc. Well, it still requires some work to beat the baseline. You need a good model, etc., but this is the main point. With this I believe it is not too hard to beat the baseline (getting to 0.7 range is a different story though lol). Anyway, good luck.\n\n#EDIT:\nAfter @ulrich07 said that seq2seq worked for him, I suspected I had a bug. So I revisited my code, and indeed, I found a bug exactly where I suspected. It turns out I missed a transposed somewhere. So, seq2seq works, too...if you don't have a bug in your code 🤣. My current training with seq2seq shows powerful results, better than my seq-to-multiple targets. I am sorry about the misleading above. Well, this is why I love sharing on Kaggle. If you share, it eventually helps you even more than you helped the others; I swear it's true haha.",
    "3034424": "transformer, 1dconv, or Unet :) You are rieght!",
    "2806793": "seq to seq worked for me.",
    "2823653": "So following the advice i made my data as a sequence and since i do not have experience in implementing transformers i started with a very simple Transformer that maps sequence to scalars and it does not do anything with position (i think not needed in this problem). The thing is that i let it train for many hours and i notice the train and val R2 to drop very very slowly. My Transformer has only like 4 million params does this make sense? the error started 0.667 and after almost 20 hours it is 0.524. Actually i just noticed that it has only done 4k iterations, that is too slow. Are transformers much slower to train than MLPs?",
    "2806222": "I have been struggling to get 0.7+ for 5 days after I got 0.68, however others who got 0.68 at the same time with me have already got 0.72+... So hard 🤣",
    "2843127": "Thank you. Can you tell me the number of parameters in your model?",
    "2835810": "seq to seq worked for me. thanks for your share.  i have a question。  the some weights of sample_sub are zero，also the labels of zero weight are not  used。  why not delete those labels in seq?   I don't know if these labels are useful for the net",
    "2807773": "In this competition, seq2seq is like in NLP where you have an input of size N (number of features in this case) and the output is the number of features (368)?\nLike \"translating\" features to the outputs?",
    "2806838": "What do you mean by seq2seq? Transformer encoder and rnn head?",
    "2807300": "What do you mean by seq2seq? Transformer encoder and rnn head?😃",
    "2821910": "Hello, thanks for this helpful post, is seq 2 seq better than just a transformer to map the input to 368 targets?",
    "2842982": "Sorry to bother you, I am a beginner. May I ask if you mean the input shape is [batch_size, 60, dim]?",
    "2834126": "How did you handle the 16 global variables in the input? Did you copy them 60 times and then concatenate them to form [60, 9 + 16], or did you use a different approach?",
    "2814118": "Did you implement a Bidirectional LSTM? I managed to get a lstm working but the results are not really good. \n```python\ninput_shape = (60, dim)\noutput_shape = (60, cols_target_dim)\n\ndef seq2seq_model(latent_dim):\n    # Encoder\n    encoder_inputs = Input(shape=input_shape)\n    encoder = Bidirectional(LSTM(units=latent_dim, return_sequences=True))\n    encoder_outputs = encoder(encoder_inputs)\n    \n    # Decoder\n    decoder_lstm = LSTM(units=latent_dim, return_sequences=True)\n    decoder_outputs = decoder_lstm(encoder_outputs)\n    \n    # Output layer\n    output_layer = Dense(cols_target_dim, activation='linear')\n    outputs = output_layer(decoder_outputs)\n\n    return Model(encoder_inputs, outputs)\n\nmodel = seq2seq_model(latent_dim=64)  # Choose the number of units for the LSTM layer\n```",
    "2813007": "Another question:\nIs it possible to beat the market with a single RTX 4090?\n",
    "2810909": "I'm trying to learn from your post and understand how seq2seq and seq2multipletargets work on this dataset, but I'm struggling to understand it:\n1. In seq2seq you have 14 target features as said in the Data section, but how are 368 target features converted to 14? I don't understand this completely.\n2. When we send 368 output targets, how are they converted to 14?"
  }
}