{
  "id": 348605,
  "title": "Shape Mismatch and Reshape in model building phase !",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/348605",
  "author_name": "Collide_Conquer_19",
  "post_date": "2022-08-29T10:19:44.250000",
  "votes": 0,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I'm trying to build an model which takes the input of shape (None,None,None,1) and this has been passed to another conv2d layer which results in the shape of (None,None,None,3). it's then passed into a efficientnet-b5 which results in shape of (None,NoneNone,2048). After this the inputs are needed to be passed to Conv3d. Conv3d accepts the 5d tensors and i have tried of reshaping the tensor into 5d but everything resulted me an error. So, is there any method you can help me to continue the work.</p>\n<p>Thanks in advance,<br>\nConquer_19.</p>",
  "messages": [
    {
      "id": 1918505,
      "postDate": "2022-08-29T16:12:09.240Z",
      "content": "<p>If the Conv3D accepts 5d tensor, I assume it is [batch, y, x, z, channel] or some such thing.  What is your [None,None,None,2048]?  Sounds like maybe the embeddings (the output with the top removed from EN) of 2048, so it is something like  [batch, slice, ?, 2048]?  Please clarify.  Without knowing what the dimensions are, it is hard to recommend a solution.  </p>\n<p>Also, if those are the embeddings, I am not sure you want to do a convolution along the embedding dimension.</p>",
      "rawMarkdown": "If the Conv3D accepts 5d tensor, I assume it is [batch, y, x, z, channel] or some such thing.  What is your [None,None,None,2048]?  Sounds like maybe the embeddings (the output with the top removed from EN) of 2048, so it is something like  [batch, slice, ?, 2048]?  Please clarify.  Without knowing what the dimensions are, it is hard to recommend a solution.  \n\nAlso, if those are the embeddings, I am not sure you want to do a convolution along the embedding dimension.\n",
      "replies": [
        {
          "id": 1918567,
          "postDate": "2022-08-29T17:10:37.390Z",
          "content": "<p><a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> we can think those as like embeddings which are got after passing through the EN and the included _top = False which means the classification layers are freezed in the pre-trained network.<br>\nSo after that (batch_size, h,w,channels) are to be the initial inputs given to network but batch_size, h,w,c are not mentioned in inputs so they became (None, None, None, 2048) after passing through EN. So this shape is needed to be passed through the conv3d for another network input which is a 3d model. So it accepts 5d array. <br>\nHope the explanation is clear, but still some parts maybe missing <br>\nCan you tell me a solution for this to move forward with my work.</p>\n<p>Thanks,<br>\nConquer _19</p>",
          "rawMarkdown": "@solverworld we can think those as like embeddings which are got after passing through the EN and the included _top = False which means the classification layers are freezed in the pre-trained network.\nSo after that (batch_size, h,w,channels) are to be the initial inputs given to network but batch_size, h,w,c are not mentioned in inputs so they became (None, None, None, 2048) after passing through EN. So this shape is needed to be passed through the conv3d for another network input which is a 3d model. So it accepts 5d array. \nHope the explanation is clear, but still some parts maybe missing \nCan you tell me a solution for this to move forward with my work.\n\nThanks,\nConquer _19"
        },
        {
          "id": 1918581,
          "postDate": "2022-08-29T17:19:12.893Z",
          "content": "<p>Is the output of the EN something like [batch, 10, 10, 2048]?  Normally you would do a something like a GlobalPoolingAverage to get the 2048 embeddings.  If you wanted to pass this to a 3D convolution as is you need to add a channel.  Something like</p>\n<pre><code># convert to 300x300x1\nx = tf.reshape(inp, ((-1, 10, 10, 2048,1))) \n#repeat to get 3 channels\ntf.repeat(inp2, repeats=[3] , axis=4, name='repeat') \n</code></pre>\n<p>would work in TensorFlow.</p>",
          "rawMarkdown": "Is the output of the EN something like [batch, 10, 10, 2048]?  Normally you would do a something like a GlobalPoolingAverage to get the 2048 embeddings.  If you wanted to pass this to a 3D convolution as is you need to add a channel.  Something like\n```\n# convert to 300x300x1\nx = tf.reshape(inp, ((-1, 10, 10, 2048,1))) \n#repeat to get 3 channels\ntf.repeat(inp2, repeats=[3] , axis=4, name='repeat') \n```\nwould work in TensorFlow."
        },
        {
          "id": 1934107,
          "postDate": "2022-09-11T06:42:38.463Z",
          "content": "<p>What would be the inp2 in 4th line. </p>",
          "rawMarkdown": "What would be the inp2 in 4th line. "
        },
        {
          "id": 1936005,
          "postDate": "2022-09-12T13:09:39.163Z",
          "content": "<p>Sorry, cut-and-paste error from code.  It should be the output of the reshape, so:</p>\n<pre><code># convert to 10x10x2048x1\nx = tf.reshape(inp, ((-1, 10, 10, 2048,1))) \n#repeat to get 3 channels 10x10x2048x3\nx = tf.repeat(x, repeats=[3] , axis=4, name='repeat') \n</code></pre>\n<p>Note that the first dimension is the batch size</p>",
          "rawMarkdown": "Sorry, cut-and-paste error from code.  It should be the output of the reshape, so:\n```\n# convert to 10x10x2048x1\nx = tf.reshape(inp, ((-1, 10, 10, 2048,1))) \n#repeat to get 3 channels 10x10x2048x3\nx = tf.repeat(x, repeats=[3] , axis=4, name='repeat') \n```\nNote that the first dimension is the batch size"
        },
        {
          "id": 1936133,
          "postDate": "2022-09-12T14:26:40.210Z",
          "content": "<p><a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> I got that!<br>\nBut, I have another question regarding the ensembling of the models.<br>\nI have an Inception-Resnet3d model and this output is needed to be fed into a 2d efficientnet model. As we know that our images are of shape 512,512,3 shape and I have searched over the internet but didn't find the correct resource. can you help me with doing this?</p>\n<p>Thanking you in Advance,<br>\nConuer_19.</p>",
          "rawMarkdown": "@solverworld I got that!\nBut, I have another question regarding the ensembling of the models.\nI have an Inception-Resnet3d model and this output is needed to be fed into a 2d efficientnet model. As we know that our images are of shape 512,512,3 shape and I have searched over the internet but didn't find the correct resource. can you help me with doing this?\n\nThanking you in Advance,\nConuer_19."
        },
        {
          "id": 1936596,
          "postDate": "2022-09-12T21:48:10.507Z",
          "content": "<p>It is not clear what you are trying to do.<br>\nA 2d (ed) EffNet would want an input of 512x512x3 as you say.<br>\nWhat is the output of your Resnet3d model?  presumably it is some embedding layer, like 10x10x10x2048 maybe because it is 3D, or is it Globally Pooled into something like 2048 features, which you would mix into some number of predictions (8 maybe?).</p>\n<p>What Inception-Resnet3d model are you using?</p>",
          "rawMarkdown": "It is not clear what you are trying to do.\nA ~~3d ~~2d (ed) EffNet would want an input of 512x512x3 as you say.\nWhat is the output of your Resnet3d model?  presumably it is some embedding layer, like 10x10x10x2048 maybe because it is 3D, or is it Globally Pooled into something like 2048 features, which you would mix into some number of predictions (8 maybe?).\n\nWhat Inception-Resnet3d model are you using?\n"
        },
        {
          "id": 1936708,
          "postDate": "2022-09-13T02:07:27.097Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> <br>\nWe know that commonly a 3d network learn in 3 dimensions and those feature are sent to 2d network and finally flattened for classification. This is what I know about 3d and 2d network.</p>\n<p>So as you said that a 3d model output is 4 dimensional like an embedding. So this need to be reshaped in a way that it can be used to feed into a 2d pre-trained model and then finally a classification layer is used for classification.</p>\n<p>Why 3d to 2d?<br>\n Features which are learnt by 3d network is more when compared to a 2d model so those features are sent to a 2d model for extraction of more significant features in the 3d model output.</p>\n<p>3d model which I'm using is <a href=\"https://github.com/ZFTurbo/classification_models_3D\" target=\"_blank\">efficientnet B5</a> and 2d model is <a href=\"https://keras.io/api/applications/inceptionresnetv2/\" target=\"_blank\">inception ResNet-50 </a>model. So can you explain me how can this can be achieved in implementing this ?</p>\n<p>Thanking you in advance,<br>\nConquer_19.</p>",
          "rawMarkdown": "Hello @solverworld \nWe know that commonly a 3d network learn in 3 dimensions and those feature are sent to 2d network and finally flattened for classification. This is what I know about 3d and 2d network.\n\nSo as you said that a 3d model output is 4 dimensional like an embedding. So this need to be reshaped in a way that it can be used to feed into a 2d pre-trained model and then finally a classification layer is used for classification.\n\nWhy 3d to 2d?\n Features which are learnt by 3d network is more when compared to a 2d model so those features are sent to a 2d model for extraction of more significant features in the 3d model output.\n\n3d model which I'm using is [efficientnet B5](https://github.com/ZFTurbo/classification_models_3D) and 2d model is [inception ResNet-50 ](https://keras.io/api/applications/inceptionresnetv2/)model. So can you explain me how can this can be achieved in implementing this ?\n\nThanking you in advance,\nConquer_19."
        },
        {
          "id": 1937491,
          "postDate": "2022-09-13T14:04:05.053Z",
          "content": "<p>Sorry, I am not exactly understanding what you are trying to do.  Perhaps someone else can help.  I am curious about it, though.<br>\nDo you have a reference for an example of a 3D network output being sent to a 2D network?</p>",
          "rawMarkdown": "Sorry, I am not exactly understanding what you are trying to do.  Perhaps someone else can help.  I am curious about it, though.\nDo you have a reference for an example of a 3D network output being sent to a 2D network?"
        },
        {
          "id": 1937567,
          "postDate": "2022-09-13T15:01:04.360Z",
          "content": "<p>I don't have any examples to do so I'm experimenting on my own, that's why I sought help from you!</p>",
          "rawMarkdown": "I don't have any examples to do so I'm experimenting on my own, that's why I sought help from you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1918100,
      "postDate": "2022-08-29T10:19:44.250Z",
      "content": "<p>Hello everyone,</p>\n<p>I'm trying to build an model which takes the input of shape (None,None,None,1) and this has been passed to another conv2d layer which results in the shape of (None,None,None,3). it's then passed into a efficientnet-b5 which results in shape of (None,NoneNone,2048). After this the inputs are needed to be passed to Conv3d. Conv3d accepts the 5d tensors and i have tried of reshaping the tensor into 5d but everything resulted me an error. So, is there any method you can help me to continue the work.</p>\n<p>Thanks in advance,<br>\nConquer_19.</p>",
      "rawMarkdown": "Hello everyone,\n\nI'm trying to build an model which takes the input of shape (None,None,None,1) and this has been passed to another conv2d layer which results in the shape of (None,None,None,3). it's then passed into a efficientnet-b5 which results in shape of (None,NoneNone,2048). After this the inputs are needed to be passed to Conv3d. Conv3d accepts the 5d tensors and i have tried of reshaping the tensor into 5d but everything resulted me an error. So, is there any method you can help me to continue the work.\n\nThanks in advance,\nConquer_19."
    }
  ],
  "comments": [
    {
      "id": 1918505,
      "author_name": "SolverWorld",
      "author_url": "",
      "post_date": "2022-08-29T16:12:09.240000",
      "content": "<p>If the Conv3D accepts 5d tensor, I assume it is [batch, y, x, z, channel] or some such thing.  What is your [None,None,None,2048]?  Sounds like maybe the embeddings (the output with the top removed from EN) of 2048, so it is something like  [batch, slice, ?, 2048]?  Please clarify.  Without knowing what the dimensions are, it is hard to recommend a solution.  </p>\n<p>Also, if those are the embeddings, I am not sure you want to do a convolution along the embedding dimension.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1918567,
          "author_name": "Collide_Conquer_19",
          "author_url": "",
          "post_date": "2022-08-29T17:10:37.390000",
          "content": "<p><a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> we can think those as like embeddings which are got after passing through the EN and the included _top = False which means the classification layers are freezed in the pre-trained network.<br>\nSo after that (batch_size, h,w,channels) are to be the initial inputs given to network but batch_size, h,w,c are not mentioned in inputs so they became (None, None, None, 2048) after passing through EN. So this shape is needed to be passed through the conv3d for another network input which is a 3d model. So it accepts 5d array. <br>\nHope the explanation is clear, but still some parts maybe missing <br>\nCan you tell me a solution for this to move forward with my work.</p>\n<p>Thanks,<br>\nConquer _19</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1918581,
          "author_name": "SolverWorld",
          "author_url": "",
          "post_date": "2022-08-29T17:19:12.893000",
          "content": "<p>Is the output of the EN something like [batch, 10, 10, 2048]?  Normally you would do a something like a GlobalPoolingAverage to get the 2048 embeddings.  If you wanted to pass this to a 3D convolution as is you need to add a channel.  Something like</p>\n<pre><code># convert to 300x300x1\nx = tf.reshape(inp, ((-1, 10, 10, 2048,1))) \n#repeat to get 3 channels\ntf.repeat(inp2, repeats=[3] , axis=4, name='repeat') \n</code></pre>\n<p>would work in TensorFlow.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1934107,
          "author_name": "Collide_Conquer_19",
          "author_url": "",
          "post_date": "2022-09-11T06:42:38.463000",
          "content": "<p>What would be the inp2 in 4th line. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1936005,
          "author_name": "SolverWorld",
          "author_url": "",
          "post_date": "2022-09-12T13:09:39.163000",
          "content": "<p>Sorry, cut-and-paste error from code.  It should be the output of the reshape, so:</p>\n<pre><code># convert to 10x10x2048x1\nx = tf.reshape(inp, ((-1, 10, 10, 2048,1))) \n#repeat to get 3 channels 10x10x2048x3\nx = tf.repeat(x, repeats=[3] , axis=4, name='repeat') \n</code></pre>\n<p>Note that the first dimension is the batch size</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1936133,
          "author_name": "Collide_Conquer_19",
          "author_url": "",
          "post_date": "2022-09-12T14:26:40.210000",
          "content": "<p><a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> I got that!<br>\nBut, I have another question regarding the ensembling of the models.<br>\nI have an Inception-Resnet3d model and this output is needed to be fed into a 2d efficientnet model. As we know that our images are of shape 512,512,3 shape and I have searched over the internet but didn't find the correct resource. can you help me with doing this?</p>\n<p>Thanking you in Advance,<br>\nConuer_19.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1936596,
          "author_name": "SolverWorld",
          "author_url": "",
          "post_date": "2022-09-12T21:48:10.507000",
          "content": "<p>It is not clear what you are trying to do.<br>\nA 2d (ed) EffNet would want an input of 512x512x3 as you say.<br>\nWhat is the output of your Resnet3d model?  presumably it is some embedding layer, like 10x10x10x2048 maybe because it is 3D, or is it Globally Pooled into something like 2048 features, which you would mix into some number of predictions (8 maybe?).</p>\n<p>What Inception-Resnet3d model are you using?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1936708,
          "author_name": "Collide_Conquer_19",
          "author_url": "",
          "post_date": "2022-09-13T02:07:27.097000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> <br>\nWe know that commonly a 3d network learn in 3 dimensions and those feature are sent to 2d network and finally flattened for classification. This is what I know about 3d and 2d network.</p>\n<p>So as you said that a 3d model output is 4 dimensional like an embedding. So this need to be reshaped in a way that it can be used to feed into a 2d pre-trained model and then finally a classification layer is used for classification.</p>\n<p>Why 3d to 2d?<br>\n Features which are learnt by 3d network is more when compared to a 2d model so those features are sent to a 2d model for extraction of more significant features in the 3d model output.</p>\n<p>3d model which I'm using is <a href=\"https://github.com/ZFTurbo/classification_models_3D\" target=\"_blank\">efficientnet B5</a> and 2d model is <a href=\"https://keras.io/api/applications/inceptionresnetv2/\" target=\"_blank\">inception ResNet-50 </a>model. So can you explain me how can this can be achieved in implementing this ?</p>\n<p>Thanking you in advance,<br>\nConquer_19.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1937491,
          "author_name": "SolverWorld",
          "author_url": "",
          "post_date": "2022-09-13T14:04:05.053000",
          "content": "<p>Sorry, I am not exactly understanding what you are trying to do.  Perhaps someone else can help.  I am curious about it, though.<br>\nDo you have a reference for an example of a 3D network output being sent to a 2D network?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1937567,
          "author_name": "Collide_Conquer_19",
          "author_url": "",
          "post_date": "2022-09-13T15:01:04.360000",
          "content": "<p>I don't have any examples to do so I'm experimenting on my own, that's why I sought help from you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1918505": "If the Conv3D accepts 5d tensor, I assume it is [batch, y, x, z, channel] or some such thing.  What is your [None,None,None,2048]?  Sounds like maybe the embeddings (the output with the top removed from EN) of 2048, so it is something like  [batch, slice, ?, 2048]?  Please clarify.  Without knowing what the dimensions are, it is hard to recommend a solution.  \n\nAlso, if those are the embeddings, I am not sure you want to do a convolution along the embedding dimension.\n",
    "1918100": "Hello everyone,\n\nI'm trying to build an model which takes the input of shape (None,None,None,1) and this has been passed to another conv2d layer which results in the shape of (None,None,None,3). it's then passed into a efficientnet-b5 which results in shape of (None,NoneNone,2048). After this the inputs are needed to be passed to Conv3d. Conv3d accepts the 5d tensors and i have tried of reshaping the tensor into 5d but everything resulted me an error. So, is there any method you can help me to continue the work.\n\nThanks in advance,\nConquer_19."
  }
}