{
  "id": 509725,
  "title": "Transformer sequence to sequence training",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/509725",
  "author_name": "Vasilis",
  "post_date": "2024-06-03T15:39:03.127000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello everybody, after succesfuly trained transformers for sequence to scalars i try to train transformers sequence to sequence but the only way to not do teacher forcing (so depend on the previous target to predict the next one) is through a for loop inside forward method. That makes it extremely slow. Is there any other way to train a sequence to sequence transformer and force the model to depend on its own predictions gradually that is faster than a for loop inside forward? or maybe any other suggested sequence to sequence architectures?</p>",
  "messages": [
    {
      "id": 2853101,
      "postDate": "2024-06-03T15:39:03.127Z",
      "content": "<p>Hello everybody, after succesfuly trained transformers for sequence to scalars i try to train transformers sequence to sequence but the only way to not do teacher forcing (so depend on the previous target to predict the next one) is through a for loop inside forward method. That makes it extremely slow. Is there any other way to train a sequence to sequence transformer and force the model to depend on its own predictions gradually that is faster than a for loop inside forward? or maybe any other suggested sequence to sequence architectures?</p>",
      "rawMarkdown": "Hello everybody, after succesfuly trained transformers for sequence to scalars i try to train transformers sequence to sequence but the only way to not do teacher forcing (so depend on the previous target to predict the next one) is through a for loop inside forward method. That makes it extremely slow. Is there any other way to train a sequence to sequence transformer and force the model to depend on its own predictions gradually that is faster than a for loop inside forward? or maybe any other suggested sequence to sequence architectures?",
      "votes": 4
    },
    {
      "id": 2853109,
      "postDate": "2024-06-03T15:48:01.700Z",
      "content": "<p>try encoder only?</p>",
      "rawMarkdown": "try encoder only?",
      "replies": [
        {
          "id": 2853145,
          "postDate": "2024-06-03T16:08:13.077Z",
          "content": "<p>The targets have some positional dependence with each other. With encoder only we do not learn anything about it right? i have tried encoder (i called it above sequence to scalars).</p>",
          "rawMarkdown": "The targets have some positional dependence with each other. With encoder only we do not learn anything about it right? i have tried encoder (i called it above sequence to scalars).",
          "replies": [
            {
              "id": 2853207,
              "postDate": "2024-06-03T16:51:00.347Z",
              "content": "<p>add position embedding?</p>",
              "rawMarkdown": "add position embedding?"
            },
            {
              "id": 2853213,
              "postDate": "2024-06-03T16:56:28.777Z",
              "content": "<p>yes i have positional embeddings but when in encoder you only use them in the input variables, with just an encoder you do not learn anything of the relation between target variables, or at least i dont know how to do it.</p>",
              "rawMarkdown": "yes i have positional embeddings but when in encoder you only use them in the input variables, with just an encoder you do not learn anything of the relation between target variables, or at least i dont know how to do it."
            },
            {
              "id": 2853234,
              "postDate": "2024-06-03T17:13:54.433Z",
              "content": "<p>See <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">here</a> an example for encoder-only seq2seq transformer. It's from the Ribonanza competition so you will have to do some digging in order to understand it. Good luck.</p>",
              "rawMarkdown": "See [here](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb) an example for encoder-only seq2seq transformer. It's from the Ribonanza competition so you will have to do some digging in order to understand it. Good luck.",
              "votes": 6
            },
            {
              "id": 2853310,
              "postDate": "2024-06-03T17:55:53.990Z",
              "content": "<p>Thanks for the suggestion, this is more or less what i have and call sequence to scalars. So the encoder outputs something and then after a linear i end up with N, 368. I thought all these scores that you guys have would require an encoder-decoder architecture instead but maybe you just configure the encoder better or do other engineering tricks that i miss.</p>",
              "rawMarkdown": "Thanks for the suggestion, this is more or less what i have and call sequence to scalars. So the encoder outputs something and then after a linear i end up with N, 368. I thought all these scores that you guys have would require an encoder-decoder architecture instead but maybe you just configure the encoder better or do other engineering tricks that i miss."
            },
            {
              "id": 2853312,
              "postDate": "2024-06-03T17:58:57.080Z",
              "content": "<p>Yeah my score is encoder only. This is not the kind of problem that call for a decoder.</p>",
              "rawMarkdown": "Yeah my score is encoder only. This is not the kind of problem that call for a decoder.",
              "votes": 6
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2853109,
      "author_name": "Zhuoqun Li",
      "author_url": "",
      "post_date": "2024-06-03T15:48:01.700000",
      "content": "<p>try encoder only?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2853145,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-06-03T16:08:13.077000",
          "content": "<p>The targets have some positional dependence with each other. With encoder only we do not learn anything about it right? i have tried encoder (i called it above sequence to scalars).</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2853207,
              "author_name": "Zhuoqun Li",
              "author_url": "",
              "post_date": "2024-06-03T16:51:00.347000",
              "content": "<p>add position embedding?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2853213,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-03T16:56:28.777000",
              "content": "<p>yes i have positional embeddings but when in encoder you only use them in the input variables, with just an encoder you do not learn anything of the relation between target variables, or at least i dont know how to do it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2853234,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-03T17:13:54.433000",
              "content": "<p>See <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">here</a> an example for encoder-only seq2seq transformer. It's from the Ribonanza competition so you will have to do some digging in order to understand it. Good luck.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2853310,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-03T17:55:53.990000",
              "content": "<p>Thanks for the suggestion, this is more or less what i have and call sequence to scalars. So the encoder outputs something and then after a linear i end up with N, 368. I thought all these scores that you guys have would require an encoder-decoder architecture instead but maybe you just configure the encoder better or do other engineering tricks that i miss.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2853312,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-03T17:58:57.080000",
              "content": "<p>Yeah my score is encoder only. This is not the kind of problem that call for a decoder.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2853101": "Hello everybody, after succesfuly trained transformers for sequence to scalars i try to train transformers sequence to sequence but the only way to not do teacher forcing (so depend on the previous target to predict the next one) is through a for loop inside forward method. That makes it extremely slow. Is there any other way to train a sequence to sequence transformer and force the model to depend on its own predictions gradually that is faster than a for loop inside forward? or maybe any other suggested sequence to sequence architectures?",
    "2853109": "try encoder only?"
  }
}