{
  "id": 194168,
  "title": "Dealing with Transformers OOM",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/194168",
  "author_name": "Kerem Turgutlu",
  "post_date": "2020-10-31T04:41:03.634000",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The competition is over now and we are able to read and learn from many great solutions. I will try to incorporate all these great tricks into our solution one by one to see how far we could go as a team, and of course to learn. </p>\n<p>I started off by using BERT in-place of LSTM, in our original pipeline we padded batch inputs to the longest sequence length of that batch. In this case some batches might end up having sequence lengths which are 1000+ and it looks like BERT consumes lots memory due to its attention head as sequence length grows. At the moment I am able to fit batch size of 16 (with fp16) but it's not ideal as our solution with LSTM used 128.</p>\n<p>Bert config:</p>\n<pre><code>config = BertConfig(vocab_size=10000, # dummy vocabsize - we will feed embeds directly\n                    hidden_size=combined_embeddings.shape[1],\n                    max_position_embeddings=1256,\n                    num_hidden_layers=4,\n                    num_attention_heads=6, # hidden_size should be divisible by num attention heads, concat of multihead attn\n                    output_hidden_states=False)\nbert_model = BertModel(config)\n</code></pre>\n<p>What are some of the best options to allow larger batch size without compromising from data quality/performance?</p>\n<ul>\n<li>Sample sequences to reduce the sequence length?</li>\n<li>something else?</li>\n</ul>",
  "messages": [
    {
      "id": 1065249,
      "postDate": "2020-10-31T04:41:03.633Z",
      "content": "<p>The competition is over now and we are able to read and learn from many great solutions. I will try to incorporate all these great tricks into our solution one by one to see how far we could go as a team, and of course to learn. </p>\n<p>I started off by using BERT in-place of LSTM, in our original pipeline we padded batch inputs to the longest sequence length of that batch. In this case some batches might end up having sequence lengths which are 1000+ and it looks like BERT consumes lots memory due to its attention head as sequence length grows. At the moment I am able to fit batch size of 16 (with fp16) but it's not ideal as our solution with LSTM used 128.</p>\n<p>Bert config:</p>\n<pre><code>config = BertConfig(vocab_size=10000, # dummy vocabsize - we will feed embeds directly\n                    hidden_size=combined_embeddings.shape[1],\n                    max_position_embeddings=1256,\n                    num_hidden_layers=4,\n                    num_attention_heads=6, # hidden_size should be divisible by num attention heads, concat of multihead attn\n                    output_hidden_states=False)\nbert_model = BertModel(config)\n</code></pre>\n<p>What are some of the best options to allow larger batch size without compromising from data quality/performance?</p>\n<ul>\n<li>Sample sequences to reduce the sequence length?</li>\n<li>something else?</li>\n</ul>",
      "rawMarkdown": "The competition is over now and we are able to read and learn from many great solutions. I will try to incorporate all these great tricks into our solution one by one to see how far we could go as a team, and of course to learn. \n\nI started off by using BERT in-place of LSTM, in our original pipeline we padded batch inputs to the longest sequence length of that batch. In this case some batches might end up having sequence lengths which are 1000+ and it looks like BERT consumes lots memory due to its attention head as sequence length grows. At the moment I am able to fit batch size of 16 (with fp16) but it's not ideal as our solution with LSTM used 128.\n\nBert config:\n```\nconfig = BertConfig(vocab_size=10000, # dummy vocabsize - we will feed embeds directly\n                    hidden_size=combined_embeddings.shape[1],\n                    max_position_embeddings=1256,\n                    num_hidden_layers=4,\n                    num_attention_heads=6, # hidden_size should be divisible by num attention heads, concat of multihead attn\n                    output_hidden_states=False)\nbert_model = BertModel(config)\n```\n\nWhat are some of the best options to allow larger batch size without compromising from data quality/performance?\n\n- Sample sequences to reduce the sequence length?\n- something else?",
      "votes": 4
    },
    {
      "id": 1066253,
      "postDate": "2020-11-01T14:45:38.293Z",
      "content": "<p>I also used a transformer, but I split long sequences to sequences of length 128 (selecting random slices for each sub-sequence), the final series class for series with more then 128 slices was the average of all it's sub-sequences classes.<br>\nThis method gave better CV and LB then using the full sequence. <br>\n(I didn't use BERT but a transformer encoder with 4 encoder layers and 2-4 attention heads)  </p>",
      "rawMarkdown": "I also used a transformer, but I split long sequences to sequences of length 128 (selecting random slices for each sub-sequence), the final series class for series with more then 128 slices was the average of all it's sub-sequences classes.\nThis method gave better CV and LB then using the full sequence. \n(I didn't use BERT but a transformer encoder with 4 encoder layers and 2-4 attention heads)  ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1066253,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2020-11-01T14:45:38.293000",
      "content": "<p>I also used a transformer, but I split long sequences to sequences of length 128 (selecting random slices for each sub-sequence), the final series class for series with more then 128 slices was the average of all it's sub-sequences classes.<br>\nThis method gave better CV and LB then using the full sequence. <br>\n(I didn't use BERT but a transformer encoder with 4 encoder layers and 2-4 attention heads)  </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1065249": "The competition is over now and we are able to read and learn from many great solutions. I will try to incorporate all these great tricks into our solution one by one to see how far we could go as a team, and of course to learn. \n\nI started off by using BERT in-place of LSTM, in our original pipeline we padded batch inputs to the longest sequence length of that batch. In this case some batches might end up having sequence lengths which are 1000+ and it looks like BERT consumes lots memory due to its attention head as sequence length grows. At the moment I am able to fit batch size of 16 (with fp16) but it's not ideal as our solution with LSTM used 128.\n\nBert config:\n```\nconfig = BertConfig(vocab_size=10000, # dummy vocabsize - we will feed embeds directly\n                    hidden_size=combined_embeddings.shape[1],\n                    max_position_embeddings=1256,\n                    num_hidden_layers=4,\n                    num_attention_heads=6, # hidden_size should be divisible by num attention heads, concat of multihead attn\n                    output_hidden_states=False)\nbert_model = BertModel(config)\n```\n\nWhat are some of the best options to allow larger batch size without compromising from data quality/performance?\n\n- Sample sequences to reduce the sequence length?\n- something else?",
    "1066253": "I also used a transformer, but I split long sequences to sequences of length 128 (selecting random slices for each sub-sequence), the final series class for series with more then 128 slices was the average of all it's sub-sequences classes.\nThis method gave better CV and LB then using the full sequence. \n(I didn't use BERT but a transformer encoder with 4 encoder layers and 2-4 attention heads)  "
  }
}