{
  "id": 377835,
  "title": "Reconstruction by using ViT",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/377835",
  "author_name": "Lau2664",
  "post_date": "2023-01-13T04:28:41.423000",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Recently I used ViT to reconstruct image, with followed parameters:</p>\n<ul>\n<li>image size: 224x224</li>\n<li>batch size: 64</li>\n<li>patch size: 32x32</li>\n<li>number of patches: 49</li>\n<li>number of masked patches: 10</li>\n<li>number of transformer block: 16</li>\n<li>image size of reconstructed image: 64<br>\nAnd here is the result:</li>\n</ul>",
  "messages": [
    {
      "id": 2097912,
      "postDate": "2023-01-13T04:28:41.423Z",
      "content": "<p>Recently I used ViT to reconstruct image, with followed parameters:</p>\n<ul>\n<li>image size: 224x224</li>\n<li>batch size: 64</li>\n<li>patch size: 32x32</li>\n<li>number of patches: 49</li>\n<li>number of masked patches: 10</li>\n<li>number of transformer block: 16</li>\n<li>image size of reconstructed image: 64<br>\nAnd here is the result:</li>\n</ul>",
      "rawMarkdown": "Recently I used ViT to reconstruct image, with followed parameters:\n- image size: 224x224\n- batch size: 64\n- patch size: 32x32\n- number of patches: 49\n- number of masked patches: 10\n- number of transformer block: 16\n- image size of reconstructed image: 64\nAnd here is the result:",
      "votes": 1
    },
    {
      "id": 2098249,
      "postDate": "2023-01-13T11:09:29.167Z",
      "content": "<p><a href=\"https://www.kaggle.com/jay2333\" target=\"_blank\">@jay2333</a> try to focus on cropped region of image not on whole resized image, by cropping random sample from image you able to produce much more variative dataset and your network might be able to find more critical features and learn represent them. Then trained path feature extractor might be used as a part of multi instance learning pipeline </p>",
      "rawMarkdown": "@jay2333 try to focus on cropped region of image not on whole resized image, by cropping random sample from image you able to produce much more variative dataset and your network might be able to find more critical features and learn represent them. Then trained path feature extractor might be used as a part of multi instance learning pipeline "
    },
    {
      "id": 2098188,
      "postDate": "2023-01-13T09:52:23.373Z",
      "content": "<p>I noticed that the breast in target image has some texture, yet the reconstructed one is very smooth. Maybe the texture is the critical feature to the classification?</p>",
      "rawMarkdown": "I noticed that the breast in target image has some texture, yet the reconstructed one is very smooth. Maybe the texture is the critical feature to the classification?"
    },
    {
      "id": 2097940,
      "postDate": "2023-01-13T04:55:11.180Z",
      "content": "<p>And here's another demonstration:<br>\n<img src=\"https://user-images.githubusercontent.com/89793933/212237716-ef189862-ed2d-4c7d-a219-5cfeb7a9ba68.png\" alt=\"ssd\"></p>",
      "rawMarkdown": "And here's another demonstration:\n![ssd](https://user-images.githubusercontent.com/89793933/212237716-ef189862-ed2d-4c7d-a219-5cfeb7a9ba68.png)"
    },
    {
      "id": 2097934,
      "postDate": "2023-01-13T04:52:06.310Z",
      "content": "<p>I think we need some carefully-designed pre-text task to help the model to extract the critical feature. </p>",
      "rawMarkdown": "I think we need some carefully-designed pre-text task to help the model to extract the critical feature. "
    },
    {
      "id": 2097929,
      "postDate": "2023-01-13T04:47:49.740Z",
      "content": "<p>I saw a lot of people having the overfitting problem, and I think that's also because the model can't extract the really needed features to classify negative and positive image. When the model extracts some low-level features and use them to classify, it can easily overfit on training dataset.</p>",
      "rawMarkdown": "I saw a lot of people having the overfitting problem, and I think that's also because the model can't extract the really needed features to classify negative and positive image. When the model extracts some low-level features and use them to classify, it can easily overfit on training dataset."
    },
    {
      "id": 2097921,
      "postDate": "2023-01-13T04:40:02.340Z",
      "content": "<p>But then I build a fine-tune model by freezing ViT part and add a MLP head to it as a classifier. The result was no very promising as expected… There are many false positive and little true positive. I supposed that's because the model did learn many information during reconstruction training, but it didn't learn enough information that is critical to distinguish positive and negative.</p>",
      "rawMarkdown": "But then I build a fine-tune model by freezing ViT part and add a MLP head to it as a classifier. The result was no very promising as expected... There are many false positive and little true positive. I supposed that's because the model did learn many information during reconstruction training, but it didn't learn enough information that is critical to distinguish positive and negative."
    },
    {
      "id": 2097916,
      "postDate": "2023-01-13T04:33:32.147Z",
      "content": "<p>I didn't use any augmentation or upsample which can deal with the imbalanced distribution of data. Yet the model still learned enough information from the data to reconstruct image.</p>",
      "rawMarkdown": "I didn't use any augmentation or upsample which can deal with the imbalanced distribution of data. Yet the model still learned enough information from the data to reconstruct image."
    },
    {
      "id": 2097913,
      "postDate": "2023-01-13T04:29:19.963Z",
      "content": "<p><img src=\"https://user-images.githubusercontent.com/89793933/212237348-9381a8ea-dd14-4f84-a863-122d130c43a5.png\" alt=\"result\"></p>",
      "rawMarkdown": "![result](https://user-images.githubusercontent.com/89793933/212237348-9381a8ea-dd14-4f84-a863-122d130c43a5.png)"
    }
  ],
  "comments": [
    {
      "id": 2098249,
      "author_name": "A.P.",
      "author_url": "",
      "post_date": "2023-01-13T11:09:29.167000",
      "content": "<p><a href=\"https://www.kaggle.com/jay2333\" target=\"_blank\">@jay2333</a> try to focus on cropped region of image not on whole resized image, by cropping random sample from image you able to produce much more variative dataset and your network might be able to find more critical features and learn represent them. Then trained path feature extractor might be used as a part of multi instance learning pipeline </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2098188,
      "author_name": "Lau2664",
      "author_url": "",
      "post_date": "2023-01-13T09:52:23.373000",
      "content": "<p>I noticed that the breast in target image has some texture, yet the reconstructed one is very smooth. Maybe the texture is the critical feature to the classification?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2097940,
      "author_name": "Lau2664",
      "author_url": "",
      "post_date": "2023-01-13T04:55:11.180000",
      "content": "<p>And here's another demonstration:<br>\n<img src=\"https://user-images.githubusercontent.com/89793933/212237716-ef189862-ed2d-4c7d-a219-5cfeb7a9ba68.png\" alt=\"ssd\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2097934,
      "author_name": "Lau2664",
      "author_url": "",
      "post_date": "2023-01-13T04:52:06.310000",
      "content": "<p>I think we need some carefully-designed pre-text task to help the model to extract the critical feature. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2097929,
      "author_name": "Lau2664",
      "author_url": "",
      "post_date": "2023-01-13T04:47:49.740000",
      "content": "<p>I saw a lot of people having the overfitting problem, and I think that's also because the model can't extract the really needed features to classify negative and positive image. When the model extracts some low-level features and use them to classify, it can easily overfit on training dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2097921,
      "author_name": "Lau2664",
      "author_url": "",
      "post_date": "2023-01-13T04:40:02.340000",
      "content": "<p>But then I build a fine-tune model by freezing ViT part and add a MLP head to it as a classifier. The result was no very promising as expected… There are many false positive and little true positive. I supposed that's because the model did learn many information during reconstruction training, but it didn't learn enough information that is critical to distinguish positive and negative.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2097916,
      "author_name": "Lau2664",
      "author_url": "",
      "post_date": "2023-01-13T04:33:32.147000",
      "content": "<p>I didn't use any augmentation or upsample which can deal with the imbalanced distribution of data. Yet the model still learned enough information from the data to reconstruct image.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2097913,
      "author_name": "Lau2664",
      "author_url": "",
      "post_date": "2023-01-13T04:29:19.963000",
      "content": "<p><img src=\"https://user-images.githubusercontent.com/89793933/212237348-9381a8ea-dd14-4f84-a863-122d130c43a5.png\" alt=\"result\"></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2097912": "Recently I used ViT to reconstruct image, with followed parameters:\n- image size: 224x224\n- batch size: 64\n- patch size: 32x32\n- number of patches: 49\n- number of masked patches: 10\n- number of transformer block: 16\n- image size of reconstructed image: 64\nAnd here is the result:",
    "2098249": "@jay2333 try to focus on cropped region of image not on whole resized image, by cropping random sample from image you able to produce much more variative dataset and your network might be able to find more critical features and learn represent them. Then trained path feature extractor might be used as a part of multi instance learning pipeline ",
    "2098188": "I noticed that the breast in target image has some texture, yet the reconstructed one is very smooth. Maybe the texture is the critical feature to the classification?",
    "2097940": "And here's another demonstration:\n![ssd](https://user-images.githubusercontent.com/89793933/212237716-ef189862-ed2d-4c7d-a219-5cfeb7a9ba68.png)",
    "2097934": "I think we need some carefully-designed pre-text task to help the model to extract the critical feature. ",
    "2097929": "I saw a lot of people having the overfitting problem, and I think that's also because the model can't extract the really needed features to classify negative and positive image. When the model extracts some low-level features and use them to classify, it can easily overfit on training dataset.",
    "2097921": "But then I build a fine-tune model by freezing ViT part and add a MLP head to it as a classifier. The result was no very promising as expected... There are many false positive and little true positive. I supposed that's because the model did learn many information during reconstruction training, but it didn't learn enough information that is critical to distinguish positive and negative.",
    "2097916": "I didn't use any augmentation or upsample which can deal with the imbalanced distribution of data. Yet the model still learned enough information from the data to reconstruct image.",
    "2097913": "![result](https://user-images.githubusercontent.com/89793933/212237348-9381a8ea-dd14-4f84-a863-122d130c43a5.png)"
  }
}