{
  "id": 155967,
  "title": "Tensorflow Convergence",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/155967",
  "author_name": "Richard Xiao",
  "post_date": "2020-06-03T19:41:34.298000",
  "votes": 0,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all! For those using Tensorflow/Keras, I was wondering if you have tried using Max Pooling, GeM pooling, or Concat Pooling. It seems that all my experiments with those pooling layers result in either diverging loss or no convergence. However, when I add batch normalization, all the pooling methods can converge. Despite this, Average Pooling gives the best results for me without batch normalization. In fact, I find that Average Pooling converges much faster without batch normalization. Has anyone else experienced this? If so, I would love to hear some new insights!</p>",
  "messages": [
    {
      "id": 874405,
      "postDate": "2020-06-05T00:35:13.753Z",
      "content": "<p>I've been using Max pooling and Global average pooling, which both give me around the same results.\nWhen I try GeM pooling using my own custom layer, it won't converge. I have no idea why, are you using a custom layer for GeM pooling as well, since I have not found one for Keras/Tensorflow.</p>\n\n<p>On the matter of batchnormalization, I have no problems with convergence with or without it. In fact, I think it should converge (with max pooling) in any case, since batchnormalization does not change the position of the maximum.</p>",
      "rawMarkdown": "I've been using Max pooling and Global average pooling, which both give me around the same results.\nWhen I try GeM pooling using my own custom layer, it won't converge. I have no idea why, are you using a custom layer for GeM pooling as well, since I have not found one for Keras/Tensorflow.\n\nOn the matter of batchnormalization, I have no problems with convergence with or without it. In fact, I think it should converge (with max pooling) in any case, since batchnormalization does not change the position of the maximum.",
      "replies": [
        {
          "id": 875520,
          "postDate": "2020-06-05T21:43:55.007Z",
          "content": "<p>My GeM layer is defined as the following\n<code>\nclass GeM(tf.keras.layers.Layer):\n    def __init__(self, **kwargs):\n        super(GeM, self).__init__(**kwargs)\n        self.p = tf.Variable(3.)\n        self.eps = 1e-6\n    def call(self, inputs):\n        x = tf.math.pow(tf.math.maximum(self.eps, inputs), self.p)\n        size = (x.shape[1], x.shape[2])\n        return tf.math.pow(tf.nn.avg_pool2d(x, size, 1, padding='VALID'), 1./self.p)\n</code>\n Another question, are you using tiling or are you concatenating the tiles into one large image?</p>",
          "rawMarkdown": "My GeM layer is defined as the following\n```\nclass GeM(tf.keras.layers.Layer):\n    def __init__(self, **kwargs):\n        super(GeM, self).__init__(**kwargs)\n        self.p = tf.Variable(3.)\n        self.eps = 1e-6\n    def call(self, inputs):\n        x = tf.math.pow(tf.math.maximum(self.eps, inputs), self.p)\n        size = (x.shape[1], x.shape[2])\n        return tf.math.pow(tf.nn.avg_pool2d(x, size, 1, padding='VALID'), 1./self.p)\n```\n Another question, are you using tiling or are you concatenating the tiles into one large image?",
          "votes": 1
        },
        {
          "id": 875540,
          "postDate": "2020-06-05T22:14:50.917Z",
          "content": "<p>Thank you for this code, I will see if this will make any difference.\nI'm using a list of tiles (16 at level 1 at this moment) via a TimeDistributed setup defined in this kernel:\n<a href=\"https://www.kaggle.com/vgarshin/panda-keras-timedistributed\">https://www.kaggle.com/vgarshin/panda-keras-timedistributed</a></p>\n\n<p>Using a EfficientNetB1 backend and globalmaxpooling2D, I get a score of roughly 0.81 on LB.\nIn my next experiment I am going to try using a Average pooling layer and then I'll report on what difference it made.</p>",
          "rawMarkdown": "Thank you for this code, I will see if this will make any difference.\nI'm using a list of tiles (16 at level 1 at this moment) via a TimeDistributed setup defined in this kernel:\nhttps://www.kaggle.com/vgarshin/panda-keras-timedistributed\n\nUsing a EfficientNetB1 backend and globalmaxpooling2D, I get a score of roughly 0.81 on LB.\nIn my next experiment I am going to try using a Average pooling layer and then I'll report on what difference it made.\n\n"
        },
        {
          "id": 875541,
          "postDate": "2020-06-05T22:17:14.167Z",
          "content": "<p>It's odd. For me, I find that concatenating the tiles boosts my score tremendously. Using a list of tiles, I can only go up to about 0.82 on LB but concatenating boosts me up to around 0.86-0.87. </p>",
          "rawMarkdown": "It's odd. For me, I find that concatenating the tiles boosts my score tremendously. Using a list of tiles, I can only go up to about 0.82 on LB but concatenating boosts me up to around 0.86-0.87. "
        },
        {
          "id": 875579,
          "postDate": "2020-06-06T00:08:44.380Z",
          "content": "<p>That adds up with my max pooling approach up until now. I have not tried concatenating it into a large image yet, since I was still getting to that. </p>\n\n<p>However, the model is trained differently when using  big images instead of tiles, due to the way the convolutional kernels (for instance) slide over the images. So differences can be expected.</p>\n\n<p>I just turned on my network with GlobalAveragePooling 2D, so I'll probably report back tommorow when it's done.</p>",
          "rawMarkdown": "That adds up with my max pooling approach up until now. I have not tried concatenating it into a large image yet, since I was still getting to that. \n\nHowever, the model is trained differently when using  big images instead of tiles, due to the way the convolutional kernels (for instance) slide over the images. So differences can be expected.\n\nI just turned on my network with GlobalAveragePooling 2D, so I'll probably report back tommorow when it's done."
        }
      ]
    },
    {
      "id": 873099,
      "postDate": "2020-06-03T19:41:34.297Z",
      "content": "<p>Hi all! For those using Tensorflow/Keras, I was wondering if you have tried using Max Pooling, GeM pooling, or Concat Pooling. It seems that all my experiments with those pooling layers result in either diverging loss or no convergence. However, when I add batch normalization, all the pooling methods can converge. Despite this, Average Pooling gives the best results for me without batch normalization. In fact, I find that Average Pooling converges much faster without batch normalization. Has anyone else experienced this? If so, I would love to hear some new insights!</p>",
      "rawMarkdown": "Hi all! For those using Tensorflow/Keras, I was wondering if you have tried using Max Pooling, GeM pooling, or Concat Pooling. It seems that all my experiments with those pooling layers result in either diverging loss or no convergence. However, when I add batch normalization, all the pooling methods can converge. Despite this, Average Pooling gives the best results for me without batch normalization. In fact, I find that Average Pooling converges much faster without batch normalization. Has anyone else experienced this? If so, I would love to hear some new insights!"
    }
  ],
  "comments": [
    {
      "id": 874405,
      "author_name": "Stephan",
      "author_url": "",
      "post_date": "2020-06-05T00:35:13.753000",
      "content": "<p>I've been using Max pooling and Global average pooling, which both give me around the same results.\nWhen I try GeM pooling using my own custom layer, it won't converge. I have no idea why, are you using a custom layer for GeM pooling as well, since I have not found one for Keras/Tensorflow.</p>\n\n<p>On the matter of batchnormalization, I have no problems with convergence with or without it. In fact, I think it should converge (with max pooling) in any case, since batchnormalization does not change the position of the maximum.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 875520,
          "author_name": "Richard Xiao",
          "author_url": "",
          "post_date": "2020-06-05T21:43:55.007000",
          "content": "<p>My GeM layer is defined as the following\n<code>\nclass GeM(tf.keras.layers.Layer):\n    def __init__(self, **kwargs):\n        super(GeM, self).__init__(**kwargs)\n        self.p = tf.Variable(3.)\n        self.eps = 1e-6\n    def call(self, inputs):\n        x = tf.math.pow(tf.math.maximum(self.eps, inputs), self.p)\n        size = (x.shape[1], x.shape[2])\n        return tf.math.pow(tf.nn.avg_pool2d(x, size, 1, padding='VALID'), 1./self.p)\n</code>\n Another question, are you using tiling or are you concatenating the tiles into one large image?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 875540,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-06-05T22:14:50.917000",
          "content": "<p>Thank you for this code, I will see if this will make any difference.\nI'm using a list of tiles (16 at level 1 at this moment) via a TimeDistributed setup defined in this kernel:\n<a href=\"https://www.kaggle.com/vgarshin/panda-keras-timedistributed\">https://www.kaggle.com/vgarshin/panda-keras-timedistributed</a></p>\n\n<p>Using a EfficientNetB1 backend and globalmaxpooling2D, I get a score of roughly 0.81 on LB.\nIn my next experiment I am going to try using a Average pooling layer and then I'll report on what difference it made.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 875541,
          "author_name": "Richard Xiao",
          "author_url": "",
          "post_date": "2020-06-05T22:17:14.167000",
          "content": "<p>It's odd. For me, I find that concatenating the tiles boosts my score tremendously. Using a list of tiles, I can only go up to about 0.82 on LB but concatenating boosts me up to around 0.86-0.87. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 875579,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-06-06T00:08:44.380000",
          "content": "<p>That adds up with my max pooling approach up until now. I have not tried concatenating it into a large image yet, since I was still getting to that. </p>\n\n<p>However, the model is trained differently when using  big images instead of tiles, due to the way the convolutional kernels (for instance) slide over the images. So differences can be expected.</p>\n\n<p>I just turned on my network with GlobalAveragePooling 2D, so I'll probably report back tommorow when it's done.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "874405": "I've been using Max pooling and Global average pooling, which both give me around the same results.\nWhen I try GeM pooling using my own custom layer, it won't converge. I have no idea why, are you using a custom layer for GeM pooling as well, since I have not found one for Keras/Tensorflow.\n\nOn the matter of batchnormalization, I have no problems with convergence with or without it. In fact, I think it should converge (with max pooling) in any case, since batchnormalization does not change the position of the maximum.",
    "873099": "Hi all! For those using Tensorflow/Keras, I was wondering if you have tried using Max Pooling, GeM pooling, or Concat Pooling. It seems that all my experiments with those pooling layers result in either diverging loss or no convergence. However, when I add batch normalization, all the pooling methods can converge. Despite this, Average Pooling gives the best results for me without batch normalization. In fact, I find that Average Pooling converges much faster without batch normalization. Has anyone else experienced this? If so, I would love to hear some new insights!"
  }
}