{
  "id": 388465,
  "title": "Input layer (1024,1024,3) or (1024,1024,1)",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/388465",
  "author_name": "Igor Litvin",
  "post_date": "2023-02-17T15:44:54.794000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Good day.</p>\n<p>I use EffNetB3 with input layer (1024,1024,3). I have added one more layer in front of it to use single channel input (BW, png) Conv2D(3, (1,1), padding='same'):</p>\n<p>input_V=Input(shape=(1024,1024,1))<br>\n  L41=Conv2D(3, (1,1), padding='same')(input_V)</p>\n<p>base_model = keras_efficientnet_v2.EfficientNetV2S(<br>\n       pretrained=\"imagenet\",<br>\n        num_classes=0,<br>\n        input_shape=(1024,1024,3),<br>\n    )(L41)</p>\n<p>D5=tf.keras.layers.GlobalAveragePooling2D()(base_model)</p>\n<p>I did it to save memory from overloading in TPU training. </p>\n<p>I am able to input normal BW image into this NN. </p>\n<p>However I have some problem. My NN trains not so good. I have no idea why. Or this Conv2D(3, (1,1)) layer create problem or the problem somewhere else?</p>\n<p>What do you think. If I add such Conv2D(3, (1,1) in front of EffNetB3 to increase the number of channels will it destroy my training?</p>\n<p>Thanks for your answer.</p>",
  "messages": [
    {
      "id": 2149389,
      "postDate": "2023-02-18T08:45:56.327Z",
      "content": "<p>For comparison, timm library has a 3x3 convolution layer called stem at the start of models. It makes the models compatible with different input channels.</p>\n<p><a href=\"https://github.com/rwightman/pytorch-image-models/blob/main/timm/models/efficientnet.py#L104-L108\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/main/timm/models/efficientnet.py#L104-L108</a> </p>\n<p>I don't use keras so I don't know if that functionality exists in what you are using right now. If it doesn't exist, then you have two options.</p>\n<ul>\n<li>You either stack same grayscale image 3 times on channel axis</li>\n<li>You add a convolution layer that outputs 3 channels</li>\n</ul>\n<p>Both of them are hacky and we can't tell which one is better without trying. A single layer that makes your input compatible with rest of the network won't have any drastic effect so it won't destroy anything.</p>",
      "rawMarkdown": "For comparison, timm library has a 3x3 convolution layer called stem at the start of models. It makes the models compatible with different input channels.\n\nhttps://github.com/rwightman/pytorch-image-models/blob/main/timm/models/efficientnet.py#L104-L108 \n\nI don't use keras so I don't know if that functionality exists in what you are using right now. If it doesn't exist, then you have two options.\n\n* You either stack same grayscale image 3 times on channel axis\n* You add a convolution layer that outputs 3 channels\n\nBoth of them are hacky and we can't tell which one is better without trying. A single layer that makes your input compatible with rest of the network won't have any drastic effect so it won't destroy anything."
    },
    {
      "id": 2149164,
      "postDate": "2023-02-18T02:59:35.507Z",
      "content": "<p>In your case, you did it because you wanted to reduce the computation of your process. However, 1 by 1 Conv is made to create and decrease the dimension of a model. Therefore, it must be added to the layer that has valuable information, which means adding to the next layer of the first reduce the spatial, valuable information. (input layer is valuable information, but it has only a few channels.)</p>\n<p>If you want to add the layer in order to get some better computation than before, you have to delete a bigger layer than others, then add the layer in front of some large layers.</p>",
      "rawMarkdown": "In your case, you did it because you wanted to reduce the computation of your process. However, 1 by 1 Conv is made to create and decrease the dimension of a model. Therefore, it must be added to the layer that has valuable information, which means adding to the next layer of the first reduce the spatial, valuable information. (input layer is valuable information, but it has only a few channels.)\n\nIf you want to add the layer in order to get some better computation than before, you have to delete a bigger layer than others, then add the layer in front of some large layers."
    },
    {
      "id": 2148674,
      "postDate": "2023-02-17T16:00:37.510Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/leighplt/pytorch-pre-trained-from-3-channels-to-1/notebook\" target=\"_blank\">This</a> notebook (PyTorch framework was used) might help to get confirmation on your approach. Having one more layer in front of the network might not help if your aim is to make the model more lightweight. Moreover, the first layer is probably initialized with random weights, so that can affect your training--I am not sure if freezing the pre-trained part for few epochs will help. <a href=\"https://stackoverflow.com/a/54777347\" target=\"_blank\">This</a> answer proposes a solution for similar question.</p>",
      "rawMarkdown": "[This](https://www.kaggle.com/code/leighplt/pytorch-pre-trained-from-3-channels-to-1/notebook) notebook (PyTorch framework was used) might help to get confirmation on your approach. Having one more layer in front of the network might not help if your aim is to make the model more lightweight. Moreover, the first layer is probably initialized with random weights, so that can affect your training--I am not sure if freezing the pre-trained part for few epochs will help. [This](https://stackoverflow.com/a/54777347) answer proposes a solution for similar question.",
      "replies": [
        {
          "id": 2148759,
          "postDate": "2023-02-17T16:59:16.470Z",
          "content": "<p>I unfreeze first block of EffNet to fit the new resolution and match the layers. I use it to decrease the amout of memory. 3 channels image needs 3 times more memory than 1 channel. Neural Network dont overload my memory. Images are too big.</p>",
          "rawMarkdown": "I unfreeze first block of EffNet to fit the new resolution and match the layers. I use it to decrease the amout of memory. 3 channels image needs 3 times more memory than 1 channel. Neural Network dont overload my memory. Images are too big.",
          "votes": 1,
          "replies": [
            {
              "id": 2148838,
              "postDate": "2023-02-17T17:49:32.490Z",
              "content": "<p>Images are indeed very large. I am training a 4 view model with augmentations and I keep the images in their original dimension and as 16-bit gray-scale images. The augmentations almost tipped the memory over.</p>",
              "rawMarkdown": "Images are indeed very large. I am training a 4 view model with augmentations and I keep the images in their original dimension and as 16-bit gray-scale images. The augmentations almost tipped the memory over."
            },
            {
              "id": 2149915,
              "postDate": "2023-02-18T19:26:24.257Z",
              "content": "<p>Out of interest, do you do the layer adding like this:</p>\n<pre><code>inp_shape = (img_height, img_width, )\nimg_inp = Input(shape=inp_shape, name=)\nx = Conv2D(, (,),  padding=, name=)(img_inp)\nmodel = Net(img_inp, x, name=)\n</code></pre>\n<p>I am not familiar with Keras, so I might have written something funny. 😁</p>",
              "rawMarkdown": "Out of interest, do you do the layer adding like this:\n```python\ninp_shape = (img_height, img_width, 1)\nimg_inp = Input(shape=inp_shape, name='grayscale_inp_layer')\nx = Conv2D(3, (3,3),  padding='same', name='grayscale_to_RGB_layer')(img_inp)\nmodel = Net(img_inp, x, name='mdl')\n```\nI am not familiar with Keras, so I might have written something funny. 😁"
            }
          ]
        }
      ]
    },
    {
      "id": 2148654,
      "postDate": "2023-02-17T15:44:54.793Z",
      "content": "<p>Good day.</p>\n<p>I use EffNetB3 with input layer (1024,1024,3). I have added one more layer in front of it to use single channel input (BW, png) Conv2D(3, (1,1), padding='same'):</p>\n<p>input_V=Input(shape=(1024,1024,1))<br>\n  L41=Conv2D(3, (1,1), padding='same')(input_V)</p>\n<p>base_model = keras_efficientnet_v2.EfficientNetV2S(<br>\n       pretrained=\"imagenet\",<br>\n        num_classes=0,<br>\n        input_shape=(1024,1024,3),<br>\n    )(L41)</p>\n<p>D5=tf.keras.layers.GlobalAveragePooling2D()(base_model)</p>\n<p>I did it to save memory from overloading in TPU training. </p>\n<p>I am able to input normal BW image into this NN. </p>\n<p>However I have some problem. My NN trains not so good. I have no idea why. Or this Conv2D(3, (1,1)) layer create problem or the problem somewhere else?</p>\n<p>What do you think. If I add such Conv2D(3, (1,1) in front of EffNetB3 to increase the number of channels will it destroy my training?</p>\n<p>Thanks for your answer.</p>",
      "rawMarkdown": "Good day.\n\nI use EffNetB3 with input layer (1024,1024,3). I have added one more layer in front of it to use single channel input (BW, png) Conv2D(3, (1,1), padding='same'):\n\n  input_V=Input(shape=(1024,1024,1))\n  L41=Conv2D(3, (1,1), padding='same')(input_V)\n\n  base_model = keras_efficientnet_v2.EfficientNetV2S(\n       pretrained=\"imagenet\",\n        num_classes=0,\n        input_shape=(1024,1024,3),\n    )(L41)\n\n  D5=tf.keras.layers.GlobalAveragePooling2D()(base_model)\n\nI did it to save memory from overloading in TPU training. \n\nI am able to input normal BW image into this NN. \n\nHowever I have some problem. My NN trains not so good. I have no idea why. Or this Conv2D(3, (1,1)) layer create problem or the problem somewhere else?\n\nWhat do you think. If I add such Conv2D(3, (1,1) in front of EffNetB3 to increase the number of channels will it destroy my training?\n\nThanks for your answer."
    }
  ],
  "comments": [
    {
      "id": 2149389,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-02-18T08:45:56.327000",
      "content": "<p>For comparison, timm library has a 3x3 convolution layer called stem at the start of models. It makes the models compatible with different input channels.</p>\n<p><a href=\"https://github.com/rwightman/pytorch-image-models/blob/main/timm/models/efficientnet.py#L104-L108\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/main/timm/models/efficientnet.py#L104-L108</a> </p>\n<p>I don't use keras so I don't know if that functionality exists in what you are using right now. If it doesn't exist, then you have two options.</p>\n<ul>\n<li>You either stack same grayscale image 3 times on channel axis</li>\n<li>You add a convolution layer that outputs 3 channels</li>\n</ul>\n<p>Both of them are hacky and we can't tell which one is better without trying. A single layer that makes your input compatible with rest of the network won't have any drastic effect so it won't destroy anything.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2149164,
      "author_name": "Hyunsoo Lee 1010",
      "author_url": "",
      "post_date": "2023-02-18T02:59:35.507000",
      "content": "<p>In your case, you did it because you wanted to reduce the computation of your process. However, 1 by 1 Conv is made to create and decrease the dimension of a model. Therefore, it must be added to the layer that has valuable information, which means adding to the next layer of the first reduce the spatial, valuable information. (input layer is valuable information, but it has only a few channels.)</p>\n<p>If you want to add the layer in order to get some better computation than before, you have to delete a bigger layer than others, then add the layer in front of some large layers.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2148674,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-17T16:00:37.510000",
      "content": "<p><a href=\"https://www.kaggle.com/code/leighplt/pytorch-pre-trained-from-3-channels-to-1/notebook\" target=\"_blank\">This</a> notebook (PyTorch framework was used) might help to get confirmation on your approach. Having one more layer in front of the network might not help if your aim is to make the model more lightweight. Moreover, the first layer is probably initialized with random weights, so that can affect your training--I am not sure if freezing the pre-trained part for few epochs will help. <a href=\"https://stackoverflow.com/a/54777347\" target=\"_blank\">This</a> answer proposes a solution for similar question.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2148759,
          "author_name": "Igor Litvin",
          "author_url": "",
          "post_date": "2023-02-17T16:59:16.470000",
          "content": "<p>I unfreeze first block of EffNet to fit the new resolution and match the layers. I use it to decrease the amout of memory. 3 channels image needs 3 times more memory than 1 channel. Neural Network dont overload my memory. Images are too big.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2148838,
              "author_name": "Antti Isosalo",
              "author_url": "",
              "post_date": "2023-02-17T17:49:32.490000",
              "content": "<p>Images are indeed very large. I am training a 4 view model with augmentations and I keep the images in their original dimension and as 16-bit gray-scale images. The augmentations almost tipped the memory over.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2149915,
              "author_name": "Antti Isosalo",
              "author_url": "",
              "post_date": "2023-02-18T19:26:24.257000",
              "content": "<p>Out of interest, do you do the layer adding like this:</p>\n<pre><code>inp_shape = (img_height, img_width, )\nimg_inp = Input(shape=inp_shape, name=)\nx = Conv2D(, (,),  padding=, name=)(img_inp)\nmodel = Net(img_inp, x, name=)\n</code></pre>\n<p>I am not familiar with Keras, so I might have written something funny. 😁</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2149389": "For comparison, timm library has a 3x3 convolution layer called stem at the start of models. It makes the models compatible with different input channels.\n\nhttps://github.com/rwightman/pytorch-image-models/blob/main/timm/models/efficientnet.py#L104-L108 \n\nI don't use keras so I don't know if that functionality exists in what you are using right now. If it doesn't exist, then you have two options.\n\n* You either stack same grayscale image 3 times on channel axis\n* You add a convolution layer that outputs 3 channels\n\nBoth of them are hacky and we can't tell which one is better without trying. A single layer that makes your input compatible with rest of the network won't have any drastic effect so it won't destroy anything.",
    "2149164": "In your case, you did it because you wanted to reduce the computation of your process. However, 1 by 1 Conv is made to create and decrease the dimension of a model. Therefore, it must be added to the layer that has valuable information, which means adding to the next layer of the first reduce the spatial, valuable information. (input layer is valuable information, but it has only a few channels.)\n\nIf you want to add the layer in order to get some better computation than before, you have to delete a bigger layer than others, then add the layer in front of some large layers.",
    "2148674": "[This](https://www.kaggle.com/code/leighplt/pytorch-pre-trained-from-3-channels-to-1/notebook) notebook (PyTorch framework was used) might help to get confirmation on your approach. Having one more layer in front of the network might not help if your aim is to make the model more lightweight. Moreover, the first layer is probably initialized with random weights, so that can affect your training--I am not sure if freezing the pre-trained part for few epochs will help. [This](https://stackoverflow.com/a/54777347) answer proposes a solution for similar question.",
    "2148654": "Good day.\n\nI use EffNetB3 with input layer (1024,1024,3). I have added one more layer in front of it to use single channel input (BW, png) Conv2D(3, (1,1), padding='same'):\n\n  input_V=Input(shape=(1024,1024,1))\n  L41=Conv2D(3, (1,1), padding='same')(input_V)\n\n  base_model = keras_efficientnet_v2.EfficientNetV2S(\n       pretrained=\"imagenet\",\n        num_classes=0,\n        input_shape=(1024,1024,3),\n    )(L41)\n\n  D5=tf.keras.layers.GlobalAveragePooling2D()(base_model)\n\nI did it to save memory from overloading in TPU training. \n\nI am able to input normal BW image into this NN. \n\nHowever I have some problem. My NN trains not so good. I have no idea why. Or this Conv2D(3, (1,1)) layer create problem or the problem somewhere else?\n\nWhat do you think. If I add such Conv2D(3, (1,1) in front of EffNetB3 to increase the number of channels will it destroy my training?\n\nThanks for your answer."
  }
}