{
  "id": 336022,
  "title": "Zoom-In Network - Efficient Classification of Very Large Images with Tiny Objects",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/336022",
  "author_name": "Darien Schettler",
  "post_date": "2022-07-09T00:56:16.466000",
  "votes": 26,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I've seen a couple of papers like this floating around… I'm not sure this one is the best one but I thought I would share. If you want to find similar papers you can also look at <a href=\"https://scholar.google.com/scholar_lookup?arxiv_id=2106.02694\" target=\"_blank\">this paper's Google Scholar</a> to see papers that cite/were-cited. This will also help you get an idea of the terminology used around this problem.</p>\n<p><br></p>\n<p><strong>Paper Title</strong> - Efficient Classification of Very Large Images with Tiny Objects<br>\n<strong>Paper Link</strong> - <a href=\"https://arxiv.org/abs/2106.02694\" target=\"_blank\">https://arxiv.org/abs/2106.02694</a><br>\n<strong>Paper Authors</strong> - Fanjie Kong, Ricardo Henao</p>\n<p><br></p>\n<p><strong>Paper Abstract</strong></p>\n<blockquote>\n  <p>An increasing number of applications in computer vision, especially, in medical imaging and remote sensing, become challenging when the goal is to classify very large images with tiny informative objects. Specifically, these classification tasks face two key challenges: i) the size of the input image is usually in the order of mega- or giga-pixels, however, existing deep architectures do not easily operate on such big images due to memory constraints, consequently, we seek a memory-efficient method to process these images; and ii) only a very small fraction of the input images are informative of the label of interest, resulting in low region of interest (ROI) to image ratio. However, most of the current convolutional neural networks (CNNs) are designed for image classification datasets that have relatively large ROIs and small image sizes (sub-megapixel). Existing approaches have addressed these two challenges in isolation. We present an end-to-end CNN model termed Zoom-In network that leverages hierarchical attention sampling for classification of large images with tiny objects using a single GPU. We evaluate our method on four large-image histopathology, road-scene and satellite imaging datasets, and one gigapixel pathology dataset. Experimental results show that our model achieves higher accuracy than existing methods while requiring less memory resources.</p>\n</blockquote>\n<p><br></p>\n<p><strong>Architecture Figure</strong><br>\n<img src=\"https://i.ibb.co/C1w405x/Screen-Shot-2022-07-08-at-8-57-12-PM.png\"></p>",
  "messages": [
    {
      "id": 1848847,
      "postDate": "2022-07-09T00:56:16.467Z",
      "content": "<p>I've seen a couple of papers like this floating around… I'm not sure this one is the best one but I thought I would share. If you want to find similar papers you can also look at <a href=\"https://scholar.google.com/scholar_lookup?arxiv_id=2106.02694\" target=\"_blank\">this paper's Google Scholar</a> to see papers that cite/were-cited. This will also help you get an idea of the terminology used around this problem.</p>\n<p><br></p>\n<p><strong>Paper Title</strong> - Efficient Classification of Very Large Images with Tiny Objects<br>\n<strong>Paper Link</strong> - <a href=\"https://arxiv.org/abs/2106.02694\" target=\"_blank\">https://arxiv.org/abs/2106.02694</a><br>\n<strong>Paper Authors</strong> - Fanjie Kong, Ricardo Henao</p>\n<p><br></p>\n<p><strong>Paper Abstract</strong></p>\n<blockquote>\n  <p>An increasing number of applications in computer vision, especially, in medical imaging and remote sensing, become challenging when the goal is to classify very large images with tiny informative objects. Specifically, these classification tasks face two key challenges: i) the size of the input image is usually in the order of mega- or giga-pixels, however, existing deep architectures do not easily operate on such big images due to memory constraints, consequently, we seek a memory-efficient method to process these images; and ii) only a very small fraction of the input images are informative of the label of interest, resulting in low region of interest (ROI) to image ratio. However, most of the current convolutional neural networks (CNNs) are designed for image classification datasets that have relatively large ROIs and small image sizes (sub-megapixel). Existing approaches have addressed these two challenges in isolation. We present an end-to-end CNN model termed Zoom-In network that leverages hierarchical attention sampling for classification of large images with tiny objects using a single GPU. We evaluate our method on four large-image histopathology, road-scene and satellite imaging datasets, and one gigapixel pathology dataset. Experimental results show that our model achieves higher accuracy than existing methods while requiring less memory resources.</p>\n</blockquote>\n<p><br></p>\n<p><strong>Architecture Figure</strong><br>\n<img src=\"https://i.ibb.co/C1w405x/Screen-Shot-2022-07-08-at-8-57-12-PM.png\"></p>",
      "rawMarkdown": "I've seen a couple of papers like this floating around... I'm not sure this one is the best one but I thought I would share. If you want to find similar papers you can also look at [this paper's Google Scholar](https://scholar.google.com/scholar_lookup?arxiv_id=2106.02694) to see papers that cite/were-cited. This will also help you get an idea of the terminology used around this problem.\n\n<br>\n\n**Paper Title** - Efficient Classification of Very Large Images with Tiny Objects\n**Paper Link** - https://arxiv.org/abs/2106.02694\n**Paper Authors** - Fanjie Kong, Ricardo Henao\n\n<br>\n\n**Paper Abstract**\n> An increasing number of applications in computer vision, especially, in medical imaging and remote sensing, become challenging when the goal is to classify very large images with tiny informative objects. Specifically, these classification tasks face two key challenges: i) the size of the input image is usually in the order of mega- or giga-pixels, however, existing deep architectures do not easily operate on such big images due to memory constraints, consequently, we seek a memory-efficient method to process these images; and ii) only a very small fraction of the input images are informative of the label of interest, resulting in low region of interest (ROI) to image ratio. However, most of the current convolutional neural networks (CNNs) are designed for image classification datasets that have relatively large ROIs and small image sizes (sub-megapixel). Existing approaches have addressed these two challenges in isolation. We present an end-to-end CNN model termed Zoom-In network that leverages hierarchical attention sampling for classification of large images with tiny objects using a single GPU. We evaluate our method on four large-image histopathology, road-scene and satellite imaging datasets, and one gigapixel pathology dataset. Experimental results show that our model achieves higher accuracy than existing methods while requiring less memory resources.\n\n<br>\n\n**Architecture Figure**\n<center><img src=\"https://i.ibb.co/C1w405x/Screen-Shot-2022-07-08-at-8-57-12-PM.png\" width=100%></center>",
      "votes": 25
    },
    {
      "id": 1850858,
      "postDate": "2022-07-10T20:37:08.370Z",
      "content": "<p>seems it has implementation over here: <a href=\"https://github.com/timqqt/pytorch-zoom-in-network\" target=\"_blank\">https://github.com/timqqt/pytorch-zoom-in-network</a></p>",
      "rawMarkdown": "seems it has implementation over here: https://github.com/timqqt/pytorch-zoom-in-network",
      "votes": 4
    },
    {
      "id": 1850925,
      "postDate": "2022-07-10T23:39:10.647Z",
      "content": "<p>Thanks for sharing this paper.<br>\nDo you have a link for the code?</p>",
      "rawMarkdown": "Thanks for sharing this paper.\nDo you have a link for the code?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1850940,
          "postDate": "2022-07-11T00:04:59.233Z",
          "content": "<p><a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> posted a link to the PyTorch implementation below! :)</p>",
          "rawMarkdown": "@jirkaborovec posted a link to the PyTorch implementation below! :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1906093,
      "postDate": "2022-08-19T15:07:37.653Z",
      "content": "<p>Thanks!! It's Useful!</p>",
      "rawMarkdown": "Thanks!! It's Useful!"
    }
  ],
  "comments": [
    {
      "id": 1850858,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2022-07-10T20:37:08.370000",
      "content": "<p>seems it has implementation over here: <a href=\"https://github.com/timqqt/pytorch-zoom-in-network\" target=\"_blank\">https://github.com/timqqt/pytorch-zoom-in-network</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1850925,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-10T23:39:10.647000",
      "content": "<p>Thanks for sharing this paper.<br>\nDo you have a link for the code?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1850940,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-07-11T00:04:59.233000",
          "content": "<p><a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> posted a link to the PyTorch implementation below! :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1906093,
      "author_name": "Wongi Park",
      "author_url": "",
      "post_date": "2022-08-19T15:07:37.653000",
      "content": "<p>Thanks!! It's Useful!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1848847": "I've seen a couple of papers like this floating around... I'm not sure this one is the best one but I thought I would share. If you want to find similar papers you can also look at [this paper's Google Scholar](https://scholar.google.com/scholar_lookup?arxiv_id=2106.02694) to see papers that cite/were-cited. This will also help you get an idea of the terminology used around this problem.\n\n<br>\n\n**Paper Title** - Efficient Classification of Very Large Images with Tiny Objects\n**Paper Link** - https://arxiv.org/abs/2106.02694\n**Paper Authors** - Fanjie Kong, Ricardo Henao\n\n<br>\n\n**Paper Abstract**\n> An increasing number of applications in computer vision, especially, in medical imaging and remote sensing, become challenging when the goal is to classify very large images with tiny informative objects. Specifically, these classification tasks face two key challenges: i) the size of the input image is usually in the order of mega- or giga-pixels, however, existing deep architectures do not easily operate on such big images due to memory constraints, consequently, we seek a memory-efficient method to process these images; and ii) only a very small fraction of the input images are informative of the label of interest, resulting in low region of interest (ROI) to image ratio. However, most of the current convolutional neural networks (CNNs) are designed for image classification datasets that have relatively large ROIs and small image sizes (sub-megapixel). Existing approaches have addressed these two challenges in isolation. We present an end-to-end CNN model termed Zoom-In network that leverages hierarchical attention sampling for classification of large images with tiny objects using a single GPU. We evaluate our method on four large-image histopathology, road-scene and satellite imaging datasets, and one gigapixel pathology dataset. Experimental results show that our model achieves higher accuracy than existing methods while requiring less memory resources.\n\n<br>\n\n**Architecture Figure**\n<center><img src=\"https://i.ibb.co/C1w405x/Screen-Shot-2022-07-08-at-8-57-12-PM.png\" width=100%></center>",
    "1850858": "seems it has implementation over here: https://github.com/timqqt/pytorch-zoom-in-network",
    "1850925": "Thanks for sharing this paper.\nDo you have a link for the code?\n",
    "1906093": "Thanks!! It's Useful!"
  }
}