{
  "id": 506081,
  "title": "Transformers train very slowly ",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/506081",
  "author_name": "Vasilis",
  "post_date": "2024-05-20T12:26:19.916000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello i am new to sequence to sequence/scalars models. I tried to use pytorch Transformers to model the data. My Transformer is very simple</p>\n<pre><code> (nn.Module):\n     ():\n        (SequenceToScalarTransformer, ).__init__()\n\n        \n        .input_linear = nn.Linear(input_dim, d_model)\n\n        \n        .transformer_encoder = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=d_model, nhead=nhead, dim_feedforward=dim_feedforward, dropout=dropout),\n            num_layers=num_encoder_layers\n        )\n\n        \n        .output_linear = nn.Linear(d_model, output_dim)\n\n     ():\n        \n\n        \n        src = .input_linear(src).permute(, , )\n\n        \n        src = .transformer_encoder(src)  \n\n        \n        src = src.mean(dim=)  \n\n        \n        output = .output_linear(src)  \n\n         output\n</code></pre>\n<p>I use an nvidia 4070ti GPU and a batch size of 4096. The prediction time takes around 2.5 seconds and the loss.backward() around 9 seconds. This seems to slow for me, even after 24 hours very few iterations happened and the R2 loss is very high. Does it make sense to be so slow?</p>",
  "messages": [
    {
      "id": 2825533,
      "postDate": "2024-05-20T13:07:08.960Z",
      "content": "<p>If I had to guess, you're passing the tensor with the wrong axis positions, causing the transformer encoder to treat your tensor as one with a sequence length of 4096 :D</p>",
      "rawMarkdown": "If I had to guess, you're passing the tensor with the wrong axis positions, causing the transformer encoder to treat your tensor as one with a sequence length of 4096 :D",
      "votes": 2,
      "replies": [
        {
          "id": 2825888,
          "postDate": "2024-05-20T15:55:34.873Z",
          "content": "<p>Thanks! it seems that i was feeding the data in the wrong shape indeed. Now 100 iterations happen 3x faster and the error seems to drop much faster than before! </p>",
          "rawMarkdown": "Thanks! it seems that i was feeding the data in the wrong shape indeed. Now 100 iterations happen 3x faster and the error seems to drop much faster than before! ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2825472,
      "postDate": "2024-05-20T12:26:19.917Z",
      "content": "<p>Hello i am new to sequence to sequence/scalars models. I tried to use pytorch Transformers to model the data. My Transformer is very simple</p>\n<pre><code> (nn.Module):\n     ():\n        (SequenceToScalarTransformer, ).__init__()\n\n        \n        .input_linear = nn.Linear(input_dim, d_model)\n\n        \n        .transformer_encoder = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=d_model, nhead=nhead, dim_feedforward=dim_feedforward, dropout=dropout),\n            num_layers=num_encoder_layers\n        )\n\n        \n        .output_linear = nn.Linear(d_model, output_dim)\n\n     ():\n        \n\n        \n        src = .input_linear(src).permute(, , )\n\n        \n        src = .transformer_encoder(src)  \n\n        \n        src = src.mean(dim=)  \n\n        \n        output = .output_linear(src)  \n\n         output\n</code></pre>\n<p>I use an nvidia 4070ti GPU and a batch size of 4096. The prediction time takes around 2.5 seconds and the loss.backward() around 9 seconds. This seems to slow for me, even after 24 hours very few iterations happened and the R2 loss is very high. Does it make sense to be so slow?</p>",
      "rawMarkdown": "Hello i am new to sequence to sequence/scalars models. I tried to use pytorch Transformers to model the data. My Transformer is very simple\n\n```\nclass SequenceToScalarTransformer(nn.Module):\n    def __init__(self, input_dim, output_dim, d_model, nhead, num_encoder_layers, dim_feedforward, dropout=0.1):\n        super(SequenceToScalarTransformer, self).__init__()\n\n        # Linear layer to project the input features to the model dimension\n        self.input_linear = nn.Linear(input_dim, d_model)\n\n        # Transformer Encoder to process the sequence\n        self.transformer_encoder = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=d_model, nhead=nhead, dim_feedforward=dim_feedforward, dropout=dropout),\n            num_layers=num_encoder_layers\n        )\n\n        # Output linear layer that maps from the model dimension to the desired output dimension\n        self.output_linear = nn.Linear(d_model, output_dim)\n\n    def forward(self, src):\n        # src shape: [seq_len, batch_size, input_dim]\n\n        # Project input to model dimension\n        src = self.input_linear(src).permute(1, 0, 2)\n\n        # Process sequence with the transformer encoder\n        src = self.transformer_encoder(src)  # shape: [seq_len, batch_size, d_model]\n\n        # Pooling over the sequence dimension, aggregate information\n        src = src.mean(dim=0)  # shape: [batch_size, d_model]\n\n        # Map to the desired output dimension\n        output = self.output_linear(src)  # shape: [batch_size, output_dim]\n\n        return output\n``` \n\n I use an nvidia 4070ti GPU and a batch size of 4096. The prediction time takes around 2.5 seconds and the loss.backward() around 9 seconds. This seems to slow for me, even after 24 hours very few iterations happened and the R2 loss is very high. Does it make sense to be so slow?",
      "votes": 2
    },
    {
      "id": 2825488,
      "postDate": "2024-05-20T12:40:13.470Z",
      "content": "<p>This is my very simple train loop</p>\n<pre><code>for batch_idx, (src, tgt) in (train_loader):\n      optimizer.()\n      preds = (src)\n\n      loss = (preds, tgt)\n      loss.()\n      optimizer.()\n</code></pre>",
      "rawMarkdown": "This is my very simple train loop\n\n```\nfor batch_idx, (src, tgt) in enumerate(train_loader):\n      optimizer.zero_grad()\n      preds = model(src)\n\n      loss = r2_score(preds, tgt)\n      loss.backward()\n      optimizer.step()\n```\n"
    }
  ],
  "comments": [
    {
      "id": 2825533,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-05-20T13:07:08.960000",
      "content": "<p>If I had to guess, you're passing the tensor with the wrong axis positions, causing the transformer encoder to treat your tensor as one with a sequence length of 4096 :D</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2825888,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-05-20T15:55:34.873000",
          "content": "<p>Thanks! it seems that i was feeding the data in the wrong shape indeed. Now 100 iterations happen 3x faster and the error seems to drop much faster than before! </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2825488,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-05-20T12:40:13.470000",
      "content": "<p>This is my very simple train loop</p>\n<pre><code>for batch_idx, (src, tgt) in (train_loader):\n      optimizer.()\n      preds = (src)\n\n      loss = (preds, tgt)\n      loss.()\n      optimizer.()\n</code></pre>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2825533": "If I had to guess, you're passing the tensor with the wrong axis positions, causing the transformer encoder to treat your tensor as one with a sequence length of 4096 :D",
    "2825472": "Hello i am new to sequence to sequence/scalars models. I tried to use pytorch Transformers to model the data. My Transformer is very simple\n\n```\nclass SequenceToScalarTransformer(nn.Module):\n    def __init__(self, input_dim, output_dim, d_model, nhead, num_encoder_layers, dim_feedforward, dropout=0.1):\n        super(SequenceToScalarTransformer, self).__init__()\n\n        # Linear layer to project the input features to the model dimension\n        self.input_linear = nn.Linear(input_dim, d_model)\n\n        # Transformer Encoder to process the sequence\n        self.transformer_encoder = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=d_model, nhead=nhead, dim_feedforward=dim_feedforward, dropout=dropout),\n            num_layers=num_encoder_layers\n        )\n\n        # Output linear layer that maps from the model dimension to the desired output dimension\n        self.output_linear = nn.Linear(d_model, output_dim)\n\n    def forward(self, src):\n        # src shape: [seq_len, batch_size, input_dim]\n\n        # Project input to model dimension\n        src = self.input_linear(src).permute(1, 0, 2)\n\n        # Process sequence with the transformer encoder\n        src = self.transformer_encoder(src)  # shape: [seq_len, batch_size, d_model]\n\n        # Pooling over the sequence dimension, aggregate information\n        src = src.mean(dim=0)  # shape: [batch_size, d_model]\n\n        # Map to the desired output dimension\n        output = self.output_linear(src)  # shape: [batch_size, output_dim]\n\n        return output\n``` \n\n I use an nvidia 4070ti GPU and a batch size of 4096. The prediction time takes around 2.5 seconds and the loss.backward() around 9 seconds. This seems to slow for me, even after 24 hours very few iterations happened and the R2 loss is very high. Does it make sense to be so slow?",
    "2825488": "This is my very simple train loop\n\n```\nfor batch_idx, (src, tgt) in enumerate(train_loader):\n      optimizer.zero_grad()\n      preds = model(src)\n\n      loss = r2_score(preds, tgt)\n      loss.backward()\n      optimizer.step()\n```\n"
  }
}