{
  "id": 386213,
  "title": "Learning from Frequency Domain for Image classification",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/386213",
  "author_name": "Youness EL BRAG",
  "post_date": "2023-02-11T19:49:52.514000",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>During my journey in this competition, i tried to re-implement some of SOTA architecture models , and one of them took my interest Called <strong><a href=\"https://openreview.net/forum?id=K_Mnsw5VoOW\" target=\"_blank\">Global Filter network</a></strong> <br>\nnow let me give an overview of this method quickly :</p>\n<ul>\n<li><strong>Learning from the Frequency Domain:</strong></li>\n</ul>\n<p>To overcome the limitations of the spatial domain, learning from the frequency domain has become a popular approach. The frequency domain is obtained by transforming an image from the spatial domain to the frequency domain using the Fast Fourier Transform (FFT). In this domain, each pixel in the image is represented by its frequency components, which capture the patterns and relationships between the pixels. By learning in the frequency domain, it becomes possible to adjust the importance of each frequency component during training, leading to improved performance.To implement learning from the frequency domain, we first perform the FFT on the input image to transform it from the spatial domain to the frequency domain. Next, we multiply the frequency components with learnable weights, which are parameters that can be adjusted during training. Finally, we perform the Inverse Fast Fourier Transform (IFFT) to transform the image back to the spatial domain.</p>\n<ul>\n<li><strong>Benefits of Learning from the Frequency Domain:</strong></li>\n</ul>\n<p>The frequency domain provides a more suitable representation for capturing the spatial relationships between different regions of the image. This results in a more effective representation of the underlying patterns in medical images, leading to improved performance in biomedical image segmentation.</p>\n<p>here Network Layer i re-Designed based on the task of in this competition  : <br>\n<strong><em>Notation:</em></strong> <strong>The Full implementation coming soon</strong> <br>\ni wrote an article about this topic and i explained in depth this approach here <strong><a href=\"https://younsess-elbrag.medium.com/the-benefit-of-learning-from-the-frequency-domain-in-segmentation-biomedical-images-6da0274a2d0c\" target=\"_blank\">the Benefit of Learning from the Frequency Domain in Segmentation Biomedical Images</a></strong></p>\n<ul>\n<li>this implementation layer i re-designed based on GFNET </li>\n</ul>\n<pre><code> torch\n torch.nn  nn\n torch.nn.functional  F\n\n (nn.Module):\n     ():\n        (LearnableWeightsFFT, self).__init__()\n        self.conv = nn.Conv2d(in_channels, out_channels, kernel_size, stride, padding, dilation, groups, bias)\n        self.weights = nn.Parameter(torch.Tensor(out_channels, in_channels, *kernel_size))\n        self.weights.data.fill_()\n\n     ():\n        x_f = torch.rfft(x, , onesided=)\n        weights_f = torch.rfft(self.weights, , onesided=)\n        y_f = x_f * weights_f\n        y = torch.irfft(y_f, , onesided=)\n         self.conv(y)\n</code></pre>",
  "messages": [
    {
      "id": 2140528,
      "postDate": "2023-02-11T19:49:52.513Z",
      "content": "<p>During my journey in this competition, i tried to re-implement some of SOTA architecture models , and one of them took my interest Called <strong><a href=\"https://openreview.net/forum?id=K_Mnsw5VoOW\" target=\"_blank\">Global Filter network</a></strong> <br>\nnow let me give an overview of this method quickly :</p>\n<ul>\n<li><strong>Learning from the Frequency Domain:</strong></li>\n</ul>\n<p>To overcome the limitations of the spatial domain, learning from the frequency domain has become a popular approach. The frequency domain is obtained by transforming an image from the spatial domain to the frequency domain using the Fast Fourier Transform (FFT). In this domain, each pixel in the image is represented by its frequency components, which capture the patterns and relationships between the pixels. By learning in the frequency domain, it becomes possible to adjust the importance of each frequency component during training, leading to improved performance.To implement learning from the frequency domain, we first perform the FFT on the input image to transform it from the spatial domain to the frequency domain. Next, we multiply the frequency components with learnable weights, which are parameters that can be adjusted during training. Finally, we perform the Inverse Fast Fourier Transform (IFFT) to transform the image back to the spatial domain.</p>\n<ul>\n<li><strong>Benefits of Learning from the Frequency Domain:</strong></li>\n</ul>\n<p>The frequency domain provides a more suitable representation for capturing the spatial relationships between different regions of the image. This results in a more effective representation of the underlying patterns in medical images, leading to improved performance in biomedical image segmentation.</p>\n<p>here Network Layer i re-Designed based on the task of in this competition  : <br>\n<strong><em>Notation:</em></strong> <strong>The Full implementation coming soon</strong> <br>\ni wrote an article about this topic and i explained in depth this approach here <strong><a href=\"https://younsess-elbrag.medium.com/the-benefit-of-learning-from-the-frequency-domain-in-segmentation-biomedical-images-6da0274a2d0c\" target=\"_blank\">the Benefit of Learning from the Frequency Domain in Segmentation Biomedical Images</a></strong></p>\n<ul>\n<li>this implementation layer i re-designed based on GFNET </li>\n</ul>\n<pre><code> torch\n torch.nn  nn\n torch.nn.functional  F\n\n (nn.Module):\n     ():\n        (LearnableWeightsFFT, self).__init__()\n        self.conv = nn.Conv2d(in_channels, out_channels, kernel_size, stride, padding, dilation, groups, bias)\n        self.weights = nn.Parameter(torch.Tensor(out_channels, in_channels, *kernel_size))\n        self.weights.data.fill_()\n\n     ():\n        x_f = torch.rfft(x, , onesided=)\n        weights_f = torch.rfft(self.weights, , onesided=)\n        y_f = x_f * weights_f\n        y = torch.irfft(y_f, , onesided=)\n         self.conv(y)\n</code></pre>",
      "rawMarkdown": "During my journey in this competition, i tried to re-implement some of SOTA architecture models , and one of them took my interest Called **[Global Filter network](https://openreview.net/forum?id=K_Mnsw5VoOW)** \nnow let me give an overview of this method quickly :\n*  **Learning from the Frequency Domain:**\n\nTo overcome the limitations of the spatial domain, learning from the frequency domain has become a popular approach. The frequency domain is obtained by transforming an image from the spatial domain to the frequency domain using the Fast Fourier Transform (FFT). In this domain, each pixel in the image is represented by its frequency components, which capture the patterns and relationships between the pixels. By learning in the frequency domain, it becomes possible to adjust the importance of each frequency component during training, leading to improved performance.To implement learning from the frequency domain, we first perform the FFT on the input image to transform it from the spatial domain to the frequency domain. Next, we multiply the frequency components with learnable weights, which are parameters that can be adjusted during training. Finally, we perform the Inverse Fast Fourier Transform (IFFT) to transform the image back to the spatial domain.\n\n* **Benefits of Learning from the Frequency Domain:**\n\nThe frequency domain provides a more suitable representation for capturing the spatial relationships between different regions of the image. This results in a more effective representation of the underlying patterns in medical images, leading to improved performance in biomedical image segmentation.\n\nhere Network Layer i re-Designed based on the task of in this competition  : \n***Notation:*** **The Full implementation coming soon** \ni wrote an article about this topic and i explained in depth this approach here **[the Benefit of Learning from the Frequency Domain in Segmentation Biomedical Images](https://younsess-elbrag.medium.com/the-benefit-of-learning-from-the-frequency-domain-in-segmentation-biomedical-images-6da0274a2d0c)**\n\n * this implementation layer i re-designed based on GFNET \n```python\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nclass LearnableWeightsFFT(nn.Module):\n    def __init__(self, in_channels, out_channels, kernel_size, stride=1, padding=0, dilation=1, groups=1, bias=True):\n        super(LearnableWeightsFFT, self).__init__()\n        self.conv = nn.Conv2d(in_channels, out_channels, kernel_size, stride, padding, dilation, groups, bias)\n        self.weights = nn.Parameter(torch.Tensor(out_channels, in_channels, *kernel_size))\n        self.weights.data.fill_(1)\n\n    def forward(self, x):\n        x_f = torch.rfft(x, 2, onesided=False)\n        weights_f = torch.rfft(self.weights, 2, onesided=False)\n        y_f = x_f * weights_f\n        y = torch.irfft(y_f, 2, onesided=False)\n        return self.conv(y)\n``` ",
      "votes": 5
    },
    {
      "id": 2145905,
      "postDate": "2023-02-15T12:54:53.780Z",
      "content": "<p>actually if you use FFT vision transform<br>\n<a href=\"https://github.com/raoyongming/GFNet\" target=\"_blank\">https://github.com/raoyongming/GFNet</a></p>\n<p>the result is reasonable. (not too bad)</p>\n<p>but there is one issue. we are dealing with image greater 1024 x1024.<br>\nthere is no fast way to execute fft and ifft, unless you code in cuda.<br>\n(i was looking at availble functions in tensorrt)</p>\n<p>don't forget that decoding the dicom jpeg images eats up 2hr of our time </p>",
      "rawMarkdown": "actually if you use FFT vision transform\nhttps://github.com/raoyongming/GFNet\n\nthe result is reasonable. (not too bad)\n\nbut there is one issue. we are dealing with image greater 1024 x1024.\nthere is no fast way to execute fft and ifft, unless you code in cuda.\n(i was looking at availble functions in tensorrt)\n\ndon't forget that decoding the dicom jpeg images eats up 2hr of our time ",
      "votes": 2,
      "replies": [
        {
          "id": 2146062,
          "postDate": "2023-02-15T15:20:15.207Z",
          "content": "<p>yes you are right the dimensionality of the image is high which will make the latency of the algorithm increase the consumption energy , only if we try three approach </p>\n<ul>\n<li>reduce the parameters of neurons </li>\n<li>using model Paralisim strategies </li>\n<li>resize the image ( may lead to loss of information in the spatial domain also</li>\n</ul>",
          "rawMarkdown": "yes you are right the dimensionality of the image is high which will make the latency of the algorithm increase the consumption energy , only if we try three approach \n \n* reduce the parameters of neurons \n* using model Paralisim strategies \n* resize the image ( may lead to loss of information in the spatial domain also"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2145905,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-15T12:54:53.780000",
      "content": "<p>actually if you use FFT vision transform<br>\n<a href=\"https://github.com/raoyongming/GFNet\" target=\"_blank\">https://github.com/raoyongming/GFNet</a></p>\n<p>the result is reasonable. (not too bad)</p>\n<p>but there is one issue. we are dealing with image greater 1024 x1024.<br>\nthere is no fast way to execute fft and ifft, unless you code in cuda.<br>\n(i was looking at availble functions in tensorrt)</p>\n<p>don't forget that decoding the dicom jpeg images eats up 2hr of our time </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2146062,
          "author_name": "Youness EL BRAG",
          "author_url": "",
          "post_date": "2023-02-15T15:20:15.207000",
          "content": "<p>yes you are right the dimensionality of the image is high which will make the latency of the algorithm increase the consumption energy , only if we try three approach </p>\n<ul>\n<li>reduce the parameters of neurons </li>\n<li>using model Paralisim strategies </li>\n<li>resize the image ( may lead to loss of information in the spatial domain also</li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2140528": "During my journey in this competition, i tried to re-implement some of SOTA architecture models , and one of them took my interest Called **[Global Filter network](https://openreview.net/forum?id=K_Mnsw5VoOW)** \nnow let me give an overview of this method quickly :\n*  **Learning from the Frequency Domain:**\n\nTo overcome the limitations of the spatial domain, learning from the frequency domain has become a popular approach. The frequency domain is obtained by transforming an image from the spatial domain to the frequency domain using the Fast Fourier Transform (FFT). In this domain, each pixel in the image is represented by its frequency components, which capture the patterns and relationships between the pixels. By learning in the frequency domain, it becomes possible to adjust the importance of each frequency component during training, leading to improved performance.To implement learning from the frequency domain, we first perform the FFT on the input image to transform it from the spatial domain to the frequency domain. Next, we multiply the frequency components with learnable weights, which are parameters that can be adjusted during training. Finally, we perform the Inverse Fast Fourier Transform (IFFT) to transform the image back to the spatial domain.\n\n* **Benefits of Learning from the Frequency Domain:**\n\nThe frequency domain provides a more suitable representation for capturing the spatial relationships between different regions of the image. This results in a more effective representation of the underlying patterns in medical images, leading to improved performance in biomedical image segmentation.\n\nhere Network Layer i re-Designed based on the task of in this competition  : \n***Notation:*** **The Full implementation coming soon** \ni wrote an article about this topic and i explained in depth this approach here **[the Benefit of Learning from the Frequency Domain in Segmentation Biomedical Images](https://younsess-elbrag.medium.com/the-benefit-of-learning-from-the-frequency-domain-in-segmentation-biomedical-images-6da0274a2d0c)**\n\n * this implementation layer i re-designed based on GFNET \n```python\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nclass LearnableWeightsFFT(nn.Module):\n    def __init__(self, in_channels, out_channels, kernel_size, stride=1, padding=0, dilation=1, groups=1, bias=True):\n        super(LearnableWeightsFFT, self).__init__()\n        self.conv = nn.Conv2d(in_channels, out_channels, kernel_size, stride, padding, dilation, groups, bias)\n        self.weights = nn.Parameter(torch.Tensor(out_channels, in_channels, *kernel_size))\n        self.weights.data.fill_(1)\n\n    def forward(self, x):\n        x_f = torch.rfft(x, 2, onesided=False)\n        weights_f = torch.rfft(self.weights, 2, onesided=False)\n        y_f = x_f * weights_f\n        y = torch.irfft(y_f, 2, onesided=False)\n        return self.conv(y)\n``` ",
    "2145905": "actually if you use FFT vision transform\nhttps://github.com/raoyongming/GFNet\n\nthe result is reasonable. (not too bad)\n\nbut there is one issue. we are dealing with image greater 1024 x1024.\nthere is no fast way to execute fft and ifft, unless you code in cuda.\n(i was looking at availble functions in tensorrt)\n\ndon't forget that decoding the dicom jpeg images eats up 2hr of our time "
  }
}