{
  "id": 152229,
  "title": "Different approaches to model input",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/152229",
  "author_name": "sroger",
  "post_date": "2020-05-19T01:00:27.482000",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Iafoss' concat tiles approach seems to be the defacto method of model input in notebooks and discussions, have anyone experimented with other input methods?</p>\n\n<p>I have tried a stacking approach (batch_size, n_tiles, H, W, C) with each tile being encoded by a small network (b0) followed by 3d pooling layers. This did not work that well. Thoughts on this approach or alternative methods?</p>",
  "messages": [
    {
      "id": 853105,
      "postDate": "2020-05-19T01:00:27.483Z",
      "content": "<p>Iafoss' concat tiles approach seems to be the defacto method of model input in notebooks and discussions, have anyone experimented with other input methods?</p>\n\n<p>I have tried a stacking approach (batch_size, n_tiles, H, W, C) with each tile being encoded by a small network (b0) followed by 3d pooling layers. This did not work that well. Thoughts on this approach or alternative methods?</p>",
      "rawMarkdown": "Iafoss' concat tiles approach seems to be the defacto method of model input in notebooks and discussions, have anyone experimented with other input methods?\n\nI have tried a stacking approach (batch_size, n_tiles, H, W, C) with each tile being encoded by a small network (b0) followed by 3d pooling layers. This did not work that well. Thoughts on this approach or alternative methods?",
      "votes": 1
    },
    {
      "id": 853240,
      "postDate": "2020-05-19T03:51:43.973Z",
      "content": "<p>The problem with your approach using Pool(Conv?) 3D is that you are forcing some spatial dependence across the 3rd axis. But your tiles do not share any spatial dependency like this so your network is likely learning noise.</p>",
      "rawMarkdown": "The problem with your approach using Pool(Conv?) 3D is that you are forcing some spatial dependence across the 3rd axis. But your tiles do not share any spatial dependency like this so your network is likely learning noise.",
      "replies": [
        {
          "id": 853357,
          "postDate": "2020-05-19T05:58:58.643Z",
          "content": "<p>Pooling removes spatial information, convolution layers would force spatial dependence.</p>",
          "rawMarkdown": "Pooling removes spatial information, convolution layers would force spatial dependence."
        },
        {
          "id": 853798,
          "postDate": "2020-05-19T13:47:22.010Z",
          "content": "<p>If you just do pooling without convolution how is that different than global pooling then ?</p>",
          "rawMarkdown": "If you just do pooling without convolution how is that different than global pooling then ?"
        },
        {
          "id": 853925,
          "postDate": "2020-05-19T15:59:21.613Z",
          "content": "<p>I think you misunderstood, it is global pooling, just 3d instead of 2d.</p>",
          "rawMarkdown": "I think you misunderstood, it is global pooling, just 3d instead of 2d."
        },
        {
          "id": 853929,
          "postDate": "2020-05-19T16:03:22.163Z",
          "content": "<p>which should be the same than 2d over concatenated tiles. The max (or average) of a volume should be the same as the 2d view (pooling over the spatial dimensions). This is why I'm surprised you would get different results.</p>",
          "rawMarkdown": "which should be the same than 2d over concatenated tiles. The max (or average) of a volume should be the same as the 2d view (pooling over the spatial dimensions). This is why I'm surprised you would get different results."
        },
        {
          "id": 854307,
          "postDate": "2020-05-20T00:25:15.653Z",
          "content": "<p>the results are different because in my case each tile is \"individually\" fed into a cnn via stacking and not together via concat</p>",
          "rawMarkdown": "the results are different because in my case each tile is \"individually\" fed into a cnn via stacking and not together via concat"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 853240,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-05-19T03:51:43.973000",
      "content": "<p>The problem with your approach using Pool(Conv?) 3D is that you are forcing some spatial dependence across the 3rd axis. But your tiles do not share any spatial dependency like this so your network is likely learning noise.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 853357,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "2020-05-19T05:58:58.643000",
          "content": "<p>Pooling removes spatial information, convolution layers would force spatial dependence.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 853798,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-19T13:47:22.010000",
          "content": "<p>If you just do pooling without convolution how is that different than global pooling then ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 853925,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "2020-05-19T15:59:21.613000",
          "content": "<p>I think you misunderstood, it is global pooling, just 3d instead of 2d.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 853929,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-19T16:03:22.163000",
          "content": "<p>which should be the same than 2d over concatenated tiles. The max (or average) of a volume should be the same as the 2d view (pooling over the spatial dimensions). This is why I'm surprised you would get different results.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 854307,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "2020-05-20T00:25:15.653000",
          "content": "<p>the results are different because in my case each tile is \"individually\" fed into a cnn via stacking and not together via concat</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "853105": "Iafoss' concat tiles approach seems to be the defacto method of model input in notebooks and discussions, have anyone experimented with other input methods?\n\nI have tried a stacking approach (batch_size, n_tiles, H, W, C) with each tile being encoded by a small network (b0) followed by 3d pooling layers. This did not work that well. Thoughts on this approach or alternative methods?",
    "853240": "The problem with your approach using Pool(Conv?) 3D is that you are forcing some spatial dependence across the 3rd axis. But your tiles do not share any spatial dependency like this so your network is likely learning noise."
  }
}