{
  "id": 499278,
  "title": "3rd Place Solution - a novel approach from scratch",
  "url": "/competitions/spr-head-ct-age-prediction-challenge/discussion/499278",
  "author_name": "Skelp",
  "post_date": "2024-05-01T09:06:29.483000",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5958987%2F274fde73db96188f794f53dac42a0a6f%2Fdiagram.png?generation=1714628266825394&amp;alt=media\"><br>\nI'm thrilled to share the approach that led to my model securing the first place on the public leaderboard and third on the private leaderboard in this challenging competition. Unlike the first and second-place solutions that utilized existing EfficientNet architectures (EfficientNetV2S and EfficientNetV2L respectively), I developed a custom 3D CNN that diverges significantly from these popular choices.</p>\n<p>Model definition and training notebook can be found here: <a href=\"https://gitlab.com/Skelp/headct3dtinynet\" target=\"_blank\">https://gitlab.com/Skelp/headct3dtinynet</a><br>\nThe inference notebook will be pushed later, I will update this post once this is done.</p>\n<p>In addition to the provided dataset, I also made use of the external, public dataset I mentioned here: <a href=\"https://www.kaggle.com/competitions/spr-head-ct-age-prediction-challenge/discussion/486285\" target=\"_blank\">https://www.kaggle.com/competitions/spr-head-ct-age-prediction-challenge/discussion/486285</a></p>\n<p><strong>Model Architecture and Features:</strong></p>\n<ul>\n<li><strong>Parameters</strong>: The model is relatively lightweight with approximately 2.3 million parameters.</li>\n<li><strong>Attention Mechanisms</strong>: It incorporates both channel and spatial attention mechanisms, enhancing its ability to focus on relevant features within the CT scans.</li>\n<li><strong>Activation Function</strong>: A key innovation in my model is the novel activation function, APTx. The exact function used is a per-channel-parameterized version of APTx, designed to emulate the performance of the MISH activation function but with reduced computational demand.</li>\n<li><strong>Dimensionality</strong>: The architecture is a full 3D CNN, leveraging the spatial context better than the commonly used 2.5D or 2D approaches.</li>\n<li><strong>Imaging Parameters</strong>: For preprocessing, I used a single windowing parameter set at 40, 80 to standardize the input data.</li>\n<li><strong>Data Augmentation</strong>: Utilized TorchIO for extensive data augmentation, which was critical in enhancing the model's robustness and generalizability.</li>\n<li><strong>Optimization</strong>: Initially, I used the AdaBelief optimizer with a learning rate of 3e-4. Upon observing plateaus in the leaderboard scores, I adjusted the learning rate to 3e-5 and introduced a weight decay of 0.5 for more stringent regularization, which significantly improved the model's performance.</li>\n</ul>\n<p>These choices, particularly the use of a full 3D approach and the innovative APTx activation function, were pivotal in differentiating my model from others and achieving the high accuracy necessary to excel in this competition.</p>\n<p><strong>Detailed Model Architecture:</strong></p>\n<p>My model's architecture is designed to effectively process 3D CT scan data for age prediction. Below are the specifics of the architecture, visualized in the attached diagram:</p>\n<ul>\n<li><p><strong>Input Layer</strong>: Accepts 3D data of size 128x128x128.</p></li>\n<li><p><strong>Convolutional Blocks</strong>: The model consists of a series of convolutional blocks (ConvBlock 1 to 6), each followed by 3D max pooling to reduce dimensionality while capturing the most salient features.</p>\n<ul>\n<li><strong>Block 1</strong>: Outputs 32 channels, uses Group Normalization (groups=32, effectively instance norm), and includes both Channel Attention 3D and Spatial Attention 3D mechanisms. The APTx Activation, a novel activation function I developed, is used here.</li>\n<li><strong>Subsequent Blocks</strong>: Increase in channel output progressively (64, 96, 128, 160, and finally 192).</li></ul></li>\n<li><p><strong>Flattening Layer</strong>: After the convolutional blocks, the data is flattened to feed into the fully connected layers.</p></li>\n<li><p><strong>Fully Connected Layer</strong>: A dense layer that reduces the high-level features into a single prediction output.</p></li>\n<li><p><strong>Output</strong>: The mean of predictions is calculated to determine the final age estimate. Specifically, this refers to the mean of predictions from one or more head CT scans per patient. To handle multiple scans efficiently, I parallelized the processing by placing each scan into the batch dimension. This approach ensures that the model can manage additional data without compromising on processing speed or accuracy.</p></li>\n</ul>\n<p>This architecture leverages attention mechanisms to focus on important spatial and channel-specific features, which is crucial for accurately interpreting medical images like CT scans. The use of APTx activation across the model helps to maintain computational efficiency while achieving performance comparable to more computationally intensive activations like MISH and also learn the optimal parameters per channel.</p>\n<p>I look forward to feedback and discussions on the approach and am happy to dive deeper into any aspect of the model that might interest fellow Kagglers!</p>",
  "messages": [
    {
      "id": 2786321,
      "postDate": "2024-05-01T09:06:29.483Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5958987%2F274fde73db96188f794f53dac42a0a6f%2Fdiagram.png?generation=1714628266825394&amp;alt=media\"><br>\nI'm thrilled to share the approach that led to my model securing the first place on the public leaderboard and third on the private leaderboard in this challenging competition. Unlike the first and second-place solutions that utilized existing EfficientNet architectures (EfficientNetV2S and EfficientNetV2L respectively), I developed a custom 3D CNN that diverges significantly from these popular choices.</p>\n<p>Model definition and training notebook can be found here: <a href=\"https://gitlab.com/Skelp/headct3dtinynet\" target=\"_blank\">https://gitlab.com/Skelp/headct3dtinynet</a><br>\nThe inference notebook will be pushed later, I will update this post once this is done.</p>\n<p>In addition to the provided dataset, I also made use of the external, public dataset I mentioned here: <a href=\"https://www.kaggle.com/competitions/spr-head-ct-age-prediction-challenge/discussion/486285\" target=\"_blank\">https://www.kaggle.com/competitions/spr-head-ct-age-prediction-challenge/discussion/486285</a></p>\n<p><strong>Model Architecture and Features:</strong></p>\n<ul>\n<li><strong>Parameters</strong>: The model is relatively lightweight with approximately 2.3 million parameters.</li>\n<li><strong>Attention Mechanisms</strong>: It incorporates both channel and spatial attention mechanisms, enhancing its ability to focus on relevant features within the CT scans.</li>\n<li><strong>Activation Function</strong>: A key innovation in my model is the novel activation function, APTx. The exact function used is a per-channel-parameterized version of APTx, designed to emulate the performance of the MISH activation function but with reduced computational demand.</li>\n<li><strong>Dimensionality</strong>: The architecture is a full 3D CNN, leveraging the spatial context better than the commonly used 2.5D or 2D approaches.</li>\n<li><strong>Imaging Parameters</strong>: For preprocessing, I used a single windowing parameter set at 40, 80 to standardize the input data.</li>\n<li><strong>Data Augmentation</strong>: Utilized TorchIO for extensive data augmentation, which was critical in enhancing the model's robustness and generalizability.</li>\n<li><strong>Optimization</strong>: Initially, I used the AdaBelief optimizer with a learning rate of 3e-4. Upon observing plateaus in the leaderboard scores, I adjusted the learning rate to 3e-5 and introduced a weight decay of 0.5 for more stringent regularization, which significantly improved the model's performance.</li>\n</ul>\n<p>These choices, particularly the use of a full 3D approach and the innovative APTx activation function, were pivotal in differentiating my model from others and achieving the high accuracy necessary to excel in this competition.</p>\n<p><strong>Detailed Model Architecture:</strong></p>\n<p>My model's architecture is designed to effectively process 3D CT scan data for age prediction. Below are the specifics of the architecture, visualized in the attached diagram:</p>\n<ul>\n<li><p><strong>Input Layer</strong>: Accepts 3D data of size 128x128x128.</p></li>\n<li><p><strong>Convolutional Blocks</strong>: The model consists of a series of convolutional blocks (ConvBlock 1 to 6), each followed by 3D max pooling to reduce dimensionality while capturing the most salient features.</p>\n<ul>\n<li><strong>Block 1</strong>: Outputs 32 channels, uses Group Normalization (groups=32, effectively instance norm), and includes both Channel Attention 3D and Spatial Attention 3D mechanisms. The APTx Activation, a novel activation function I developed, is used here.</li>\n<li><strong>Subsequent Blocks</strong>: Increase in channel output progressively (64, 96, 128, 160, and finally 192).</li></ul></li>\n<li><p><strong>Flattening Layer</strong>: After the convolutional blocks, the data is flattened to feed into the fully connected layers.</p></li>\n<li><p><strong>Fully Connected Layer</strong>: A dense layer that reduces the high-level features into a single prediction output.</p></li>\n<li><p><strong>Output</strong>: The mean of predictions is calculated to determine the final age estimate. Specifically, this refers to the mean of predictions from one or more head CT scans per patient. To handle multiple scans efficiently, I parallelized the processing by placing each scan into the batch dimension. This approach ensures that the model can manage additional data without compromising on processing speed or accuracy.</p></li>\n</ul>\n<p>This architecture leverages attention mechanisms to focus on important spatial and channel-specific features, which is crucial for accurately interpreting medical images like CT scans. The use of APTx activation across the model helps to maintain computational efficiency while achieving performance comparable to more computationally intensive activations like MISH and also learn the optimal parameters per channel.</p>\n<p>I look forward to feedback and discussions on the approach and am happy to dive deeper into any aspect of the model that might interest fellow Kagglers!</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5958987%2F274fde73db96188f794f53dac42a0a6f%2Fdiagram.png?generation=1714628266825394&alt=media)\nI'm thrilled to share the approach that led to my model securing the first place on the public leaderboard and third on the private leaderboard in this challenging competition. Unlike the first and second-place solutions that utilized existing EfficientNet architectures (EfficientNetV2S and EfficientNetV2L respectively), I developed a custom 3D CNN that diverges significantly from these popular choices.\n\nModel definition and training notebook can be found here: https://gitlab.com/Skelp/headct3dtinynet\nThe inference notebook will be pushed later, I will update this post once this is done.\n\nIn addition to the provided dataset, I also made use of the external, public dataset I mentioned here: https://www.kaggle.com/competitions/spr-head-ct-age-prediction-challenge/discussion/486285\n\n**Model Architecture and Features:**\n- **Parameters**: The model is relatively lightweight with approximately 2.3 million parameters.\n- **Attention Mechanisms**: It incorporates both channel and spatial attention mechanisms, enhancing its ability to focus on relevant features within the CT scans.\n- **Activation Function**: A key innovation in my model is the novel activation function, APTx. The exact function used is a per-channel-parameterized version of APTx, designed to emulate the performance of the MISH activation function but with reduced computational demand.\n- **Dimensionality**: The architecture is a full 3D CNN, leveraging the spatial context better than the commonly used 2.5D or 2D approaches.\n- **Imaging Parameters**: For preprocessing, I used a single windowing parameter set at 40, 80 to standardize the input data.\n- **Data Augmentation**: Utilized TorchIO for extensive data augmentation, which was critical in enhancing the model's robustness and generalizability.\n- **Optimization**: Initially, I used the AdaBelief optimizer with a learning rate of 3e-4. Upon observing plateaus in the leaderboard scores, I adjusted the learning rate to 3e-5 and introduced a weight decay of 0.5 for more stringent regularization, which significantly improved the model's performance.\n\nThese choices, particularly the use of a full 3D approach and the innovative APTx activation function, were pivotal in differentiating my model from others and achieving the high accuracy necessary to excel in this competition.\n\n**Detailed Model Architecture:**\n\nMy model's architecture is designed to effectively process 3D CT scan data for age prediction. Below are the specifics of the architecture, visualized in the attached diagram:\n\n- **Input Layer**: Accepts 3D data of size 128x128x128.\n- **Convolutional Blocks**: The model consists of a series of convolutional blocks (ConvBlock 1 to 6), each followed by 3D max pooling to reduce dimensionality while capturing the most salient features.\n  - **Block 1**: Outputs 32 channels, uses Group Normalization (groups=32, effectively instance norm), and includes both Channel Attention 3D and Spatial Attention 3D mechanisms. The APTx Activation, a novel activation function I developed, is used here.\n  - **Subsequent Blocks**: Increase in channel output progressively (64, 96, 128, 160, and finally 192).\n\n- **Flattening Layer**: After the convolutional blocks, the data is flattened to feed into the fully connected layers.\n\n- **Fully Connected Layer**: A dense layer that reduces the high-level features into a single prediction output.\n\n- **Output**: The mean of predictions is calculated to determine the final age estimate. Specifically, this refers to the mean of predictions from one or more head CT scans per patient. To handle multiple scans efficiently, I parallelized the processing by placing each scan into the batch dimension. This approach ensures that the model can manage additional data without compromising on processing speed or accuracy.\n\nThis architecture leverages attention mechanisms to focus on important spatial and channel-specific features, which is crucial for accurately interpreting medical images like CT scans. The use of APTx activation across the model helps to maintain computational efficiency while achieving performance comparable to more computationally intensive activations like MISH and also learn the optimal parameters per channel.\n\nI look forward to feedback and discussions on the approach and am happy to dive deeper into any aspect of the model that might interest fellow Kagglers!",
      "votes": 4
    },
    {
      "id": 2788282,
      "postDate": "2024-05-02T06:59:06.703Z",
      "content": "<p>Really neat! Congratulations! I had a standard 3D CNN approach as well but the performance wasn't great. I think the CBAM inspired attention modules and the activation function you chose really did it's job well. Thanks for sharing</p>",
      "rawMarkdown": "Really neat! Congratulations! I had a standard 3D CNN approach as well but the performance wasn't great. I think the CBAM inspired attention modules and the activation function you chose really did it's job well. Thanks for sharing",
      "votes": 1,
      "replies": [
        {
          "id": 2791124,
          "postDate": "2024-05-03T13:44:45.220Z",
          "content": "<p>Thank you very much! What was your total parameter count? I am curious to see if that might also have played a role. But I agree that the combination of CBAM and a very flexible activation function made this model perform the way it did. I will release the source code and model weights after some cleaning on GitLab, editing this post to reflect that.</p>",
          "rawMarkdown": "Thank you very much! What was your total parameter count? I am curious to see if that might also have played a role. But I agree that the combination of CBAM and a very flexible activation function made this model perform the way it did. I will release the source code and model weights after some cleaning on GitLab, editing this post to reflect that."
        }
      ]
    },
    {
      "id": 2786426,
      "postDate": "2024-05-01T09:54:10.347Z",
      "content": "<p>Thanks for sharing! <br>\nIf you have time, I would like to see a performance comparison with the model implemented <a href=\"https://github.com/ZFTurbo/timm_3d\" target=\"_blank\">here</a> and <a href=\"https://pytorchvideo.readthedocs.io/en/latest/model_zoo.html\" target=\"_blank\">here</a>! ! <br>\nAlso, how much epoch did you learn?</p>",
      "rawMarkdown": "\nThanks for sharing! \nIf you have time, I would like to see a performance comparison with the model implemented [here](https://github.com/ZFTurbo/timm_3d) and [here](https://pytorchvideo.readthedocs.io/en/latest/model_zoo.html)! ! \n\n\nAlso, how much epoch did you learn?",
      "votes": 1,
      "replies": [
        {
          "id": 2786462,
          "postDate": "2024-05-01T10:06:51.667Z",
          "content": "<p>Hello patriot,<br>\nI may be a bit dense right now due to lack of sleep, but I don't perfectly understand what you mean - do you wish to see the model train on benchmark datasets and see the results, or have the architecture implemented in the timm_3d library? Please help me understand, thanks!<br>\nAnd I have trained for hundreds of epochs. With a model this small, it really takes some time. Also, I could've be more precise with the learning rate(s), speeding up convergence. But this was of no concern for me, as I had other things to do and just kept it running in the background.</p>",
          "rawMarkdown": "Hello patriot,\nI may be a bit dense right now due to lack of sleep, but I don't perfectly understand what you mean - do you wish to see the model train on benchmark datasets and see the results, or have the architecture implemented in the timm_3d library? Please help me understand, thanks!\nAnd I have trained for hundreds of epochs. With a model this small, it really takes some time. Also, I could've be more precise with the learning rate(s), speeding up convergence. But this was of no concern for me, as I had other things to do and just kept it running in the background.",
          "votes": 1,
          "replies": [
            {
              "id": 2786483,
              "postDate": "2024-05-01T10:16:33.153Z",
              "content": "<p>I would like to see which is better in your learning script, your original model or the existing model implemented in timm_3d or so.</p>",
              "rawMarkdown": "I would like to see which is better in your learning script, your original model or the existing model implemented in timm_3d or so."
            }
          ]
        }
      ]
    },
    {
      "id": 2786360,
      "postDate": "2024-05-01T09:26:48.980Z",
      "content": "<p>Thank you for the excellent solution!<br>\nI had no success with the 3D model approach, so I find yours particularly intriguing. </p>\n<p>I have two questions:</p>\n<ul>\n<li>My 3D experiments consistently showed predictions on the test set to be about 1 point worse than my validation scores. Did you observe a similar discrepancy between your validation and test scores?</li>\n<li>As I'm not well-versed in 3D models, I'm curious whether an architecture combining 3D convolution and attention is common. Was there a specific architecture that inspired your approach?</li>\n</ul>\n<p>Achieving such performance with compact dimensions like 128 x 128 x 128 is astonishing and truly fascinating.</p>",
      "rawMarkdown": "Thank you for the excellent solution!\nI had no success with the 3D model approach, so I find yours particularly intriguing. \n\nI have two questions:\n\n- My 3D experiments consistently showed predictions on the test set to be about 1 point worse than my validation scores. Did you observe a similar discrepancy between your validation and test scores?\n- As I'm not well-versed in 3D models, I'm curious whether an architecture combining 3D convolution and attention is common. Was there a specific architecture that inspired your approach?\n\nAchieving such performance with compact dimensions like 128 x 128 x 128 is astonishing and truly fascinating.",
      "votes": 1,
      "replies": [
        {
          "id": 2786402,
          "postDate": "2024-05-01T09:43:35.673Z",
          "content": "<p>Hello YYama,<br>\nof yes, in fact my best scoring submission had a train loss of 0.29 - a big discrepancy.<br>\nAnd I am not sure if 3D attention is common, but I just adapted existing attention ideas from the 2D realm to 3D, since it was pretty easy to do. The attention modules were inspired by <a href=\"https://paperswithcode.com/paper/cbam-convolutional-block-attention-module\" target=\"_blank\">CBAM</a>.<br>\nI was aiming for a relatively simple and low-parameter solution, and I am happy with the results. Scaling up the model and / or increasing the input dimensions and introducing 3D-head-cropping would probably lead to even better results.</p>",
          "rawMarkdown": "Hello YYama,\nof yes, in fact my best scoring submission had a train loss of 0.29 - a big discrepancy.\nAnd I am not sure if 3D attention is common, but I just adapted existing attention ideas from the 2D realm to 3D, since it was pretty easy to do. The attention modules were inspired by [CBAM](https://paperswithcode.com/paper/cbam-convolutional-block-attention-module).\nI was aiming for a relatively simple and low-parameter solution, and I am happy with the results. Scaling up the model and / or increasing the input dimensions and introducing 3D-head-cropping would probably lead to even better results.",
          "votes": 1,
          "replies": [
            {
              "id": 2786548,
              "postDate": "2024-05-01T10:53:20.777Z",
              "content": "<p>I have heard of CBAM but haven’t looked into it closely. I’ll give it a try!</p>\n<p>I agree that scaling up can improve scores. In my 2D approach, I tried sizes 160, 256, 384, and 512, and the larger sizes consistently led to better scores.</p>\n<p>Another thing I find impressive is that both you and Patriot used only the brain windowing for the image preprocessing. Of course, brain atrophy and chronic ischemic changes are important clues for predicting age, but I think bone condition might be quite crucial too. Because bone density decreases with age.</p>\n<p>Your solution seems to have a lot of potential for further improvement. It’s very impressive, thank you!</p>",
              "rawMarkdown": "I have heard of CBAM but haven’t looked into it closely. I’ll give it a try!\n\nI agree that scaling up can improve scores. In my 2D approach, I tried sizes 160, 256, 384, and 512, and the larger sizes consistently led to better scores.\n\nAnother thing I find impressive is that both you and Patriot used only the brain windowing for the image preprocessing. Of course, brain atrophy and chronic ischemic changes are important clues for predicting age, but I think bone condition might be quite crucial too. Because bone density decreases with age.\n\nYour solution seems to have a lot of potential for further improvement. It’s very impressive, thank you!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2789238,
      "postDate": "2024-05-02T16:07:16.650Z",
      "content": "<p>Hey! I was thinking about your solution, did you not think of using any residual connection in your architecture? </p>",
      "rawMarkdown": "Hey! I was thinking about your solution, did you not think of using any residual connection in your architecture? ",
      "replies": [
        {
          "id": 2791129,
          "postDate": "2024-05-03T13:47:20.337Z",
          "content": "<p>Hey there.<br>\nPersonally, I did consider it briefly, but then removed that idea from my head. Reason being two-fold: First, it would add a lot more VRAM requirements due to the added gradients. But more importantly, since this is such a shallow, simple model, I don't see how the model would benefit from the residual connections - I would have to downsample the input to the next ConvBlock before processing and add it to the output (or concatenate). That seems like a lot more processing for probably just a marginal gain in validation score. Once the source code is released on GitLab, you can try added residual connections to see if it helps.</p>",
          "rawMarkdown": "Hey there.\nPersonally, I did consider it briefly, but then removed that idea from my head. Reason being two-fold: First, it would add a lot more VRAM requirements due to the added gradients. But more importantly, since this is such a shallow, simple model, I don't see how the model would benefit from the residual connections - I would have to downsample the input to the next ConvBlock before processing and add it to the output (or concatenate). That seems like a lot more processing for probably just a marginal gain in validation score. Once the source code is released on GitLab, you can try added residual connections to see if it helps."
        },
        {
          "id": 2793563,
          "postDate": "2024-05-04T19:34:11.507Z",
          "content": "<p>The model class definition and also the training notebook are now accessible via a link in the original post. If you'd like you can try out adding residual connections now and let us know what you found.</p>",
          "rawMarkdown": "The model class definition and also the training notebook are now accessible via a link in the original post. If you'd like you can try out adding residual connections now and let us know what you found."
        }
      ]
    },
    {
      "id": 2789229,
      "postDate": "2024-05-02T15:59:17.423Z",
      "content": "<p>Thanks for sharing your solution! You  solution was very interesting.  I tried to use a transformer based network (SWin and TimeSformer) for a single imaging approach but both were stagnating. Great to see that you managed to use Attention.</p>\n<p>Also, I saw in preprocessing you used a window of 40 and 80 in the channel axis. How did you establish the cut points for the different patients?</p>",
      "rawMarkdown": "Thanks for sharing your solution! You  solution was very interesting.  I tried to use a transformer based network (SWin and TimeSformer) for a single imaging approach but both were stagnating. Great to see that you managed to use Attention.\n\nAlso, I saw in preprocessing you used a window of 40 and 80 in the channel axis. How did you establish the cut points for the different patients?",
      "replies": [
        {
          "id": 2791134,
          "postDate": "2024-05-03T13:50:05.003Z",
          "content": "<p>I also adore the concept of attention, I was first introduced to it when GPT-2 came out. So having a way to apply a similar concept in the visual realm is just neat - especially because it allows for creating visuals of the attention-maps, being able to visualize what part of the image(s) the model is focusing on.<br>\nI am not sure what you mean with the second paragraph - I used a constant windowing scheme of 40, 80 to focus on brain tissue. I am not sure what you mean with your question, though.</p>",
          "rawMarkdown": "I also adore the concept of attention, I was first introduced to it when GPT-2 came out. So having a way to apply a similar concept in the visual realm is just neat - especially because it allows for creating visuals of the attention-maps, being able to visualize what part of the image(s) the model is focusing on.\nI am not sure what you mean with the second paragraph - I used a constant windowing scheme of 40, 80 to focus on brain tissue. I am not sure what you mean with your question, though."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2788282,
      "author_name": "Shreyas Daniel Gaddam",
      "author_url": "",
      "post_date": "2024-05-02T06:59:06.703000",
      "content": "<p>Really neat! Congratulations! I had a standard 3D CNN approach as well but the performance wasn't great. I think the CBAM inspired attention modules and the activation function you chose really did it's job well. Thanks for sharing</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2791124,
          "author_name": "Skelp",
          "author_url": "",
          "post_date": "2024-05-03T13:44:45.220000",
          "content": "<p>Thank you very much! What was your total parameter count? I am curious to see if that might also have played a role. But I agree that the combination of CBAM and a very flexible activation function made this model perform the way it did. I will release the source code and model weights after some cleaning on GitLab, editing this post to reflect that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2786426,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2024-05-01T09:54:10.347000",
      "content": "<p>Thanks for sharing! <br>\nIf you have time, I would like to see a performance comparison with the model implemented <a href=\"https://github.com/ZFTurbo/timm_3d\" target=\"_blank\">here</a> and <a href=\"https://pytorchvideo.readthedocs.io/en/latest/model_zoo.html\" target=\"_blank\">here</a>! ! <br>\nAlso, how much epoch did you learn?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2786462,
          "author_name": "Skelp",
          "author_url": "",
          "post_date": "2024-05-01T10:06:51.667000",
          "content": "<p>Hello patriot,<br>\nI may be a bit dense right now due to lack of sleep, but I don't perfectly understand what you mean - do you wish to see the model train on benchmark datasets and see the results, or have the architecture implemented in the timm_3d library? Please help me understand, thanks!<br>\nAnd I have trained for hundreds of epochs. With a model this small, it really takes some time. Also, I could've be more precise with the learning rate(s), speeding up convergence. But this was of no concern for me, as I had other things to do and just kept it running in the background.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2786483,
              "author_name": "patriot",
              "author_url": "",
              "post_date": "2024-05-01T10:16:33.153000",
              "content": "<p>I would like to see which is better in your learning script, your original model or the existing model implemented in timm_3d or so.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2786360,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2024-05-01T09:26:48.980000",
      "content": "<p>Thank you for the excellent solution!<br>\nI had no success with the 3D model approach, so I find yours particularly intriguing. </p>\n<p>I have two questions:</p>\n<ul>\n<li>My 3D experiments consistently showed predictions on the test set to be about 1 point worse than my validation scores. Did you observe a similar discrepancy between your validation and test scores?</li>\n<li>As I'm not well-versed in 3D models, I'm curious whether an architecture combining 3D convolution and attention is common. Was there a specific architecture that inspired your approach?</li>\n</ul>\n<p>Achieving such performance with compact dimensions like 128 x 128 x 128 is astonishing and truly fascinating.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2786402,
          "author_name": "Skelp",
          "author_url": "",
          "post_date": "2024-05-01T09:43:35.673000",
          "content": "<p>Hello YYama,<br>\nof yes, in fact my best scoring submission had a train loss of 0.29 - a big discrepancy.<br>\nAnd I am not sure if 3D attention is common, but I just adapted existing attention ideas from the 2D realm to 3D, since it was pretty easy to do. The attention modules were inspired by <a href=\"https://paperswithcode.com/paper/cbam-convolutional-block-attention-module\" target=\"_blank\">CBAM</a>.<br>\nI was aiming for a relatively simple and low-parameter solution, and I am happy with the results. Scaling up the model and / or increasing the input dimensions and introducing 3D-head-cropping would probably lead to even better results.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2786548,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2024-05-01T10:53:20.777000",
              "content": "<p>I have heard of CBAM but haven’t looked into it closely. I’ll give it a try!</p>\n<p>I agree that scaling up can improve scores. In my 2D approach, I tried sizes 160, 256, 384, and 512, and the larger sizes consistently led to better scores.</p>\n<p>Another thing I find impressive is that both you and Patriot used only the brain windowing for the image preprocessing. Of course, brain atrophy and chronic ischemic changes are important clues for predicting age, but I think bone condition might be quite crucial too. Because bone density decreases with age.</p>\n<p>Your solution seems to have a lot of potential for further improvement. It’s very impressive, thank you!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2789238,
      "author_name": "Thiago Matheus",
      "author_url": "",
      "post_date": "2024-05-02T16:07:16.650000",
      "content": "<p>Hey! I was thinking about your solution, did you not think of using any residual connection in your architecture? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2791129,
          "author_name": "Skelp",
          "author_url": "",
          "post_date": "2024-05-03T13:47:20.337000",
          "content": "<p>Hey there.<br>\nPersonally, I did consider it briefly, but then removed that idea from my head. Reason being two-fold: First, it would add a lot more VRAM requirements due to the added gradients. But more importantly, since this is such a shallow, simple model, I don't see how the model would benefit from the residual connections - I would have to downsample the input to the next ConvBlock before processing and add it to the output (or concatenate). That seems like a lot more processing for probably just a marginal gain in validation score. Once the source code is released on GitLab, you can try added residual connections to see if it helps.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2793563,
          "author_name": "Skelp",
          "author_url": "",
          "post_date": "2024-05-04T19:34:11.507000",
          "content": "<p>The model class definition and also the training notebook are now accessible via a link in the original post. If you'd like you can try out adding residual connections now and let us know what you found.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2789229,
      "author_name": "Thiago Matheus",
      "author_url": "",
      "post_date": "2024-05-02T15:59:17.423000",
      "content": "<p>Thanks for sharing your solution! You  solution was very interesting.  I tried to use a transformer based network (SWin and TimeSformer) for a single imaging approach but both were stagnating. Great to see that you managed to use Attention.</p>\n<p>Also, I saw in preprocessing you used a window of 40 and 80 in the channel axis. How did you establish the cut points for the different patients?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2791134,
          "author_name": "Skelp",
          "author_url": "",
          "post_date": "2024-05-03T13:50:05.003000",
          "content": "<p>I also adore the concept of attention, I was first introduced to it when GPT-2 came out. So having a way to apply a similar concept in the visual realm is just neat - especially because it allows for creating visuals of the attention-maps, being able to visualize what part of the image(s) the model is focusing on.<br>\nI am not sure what you mean with the second paragraph - I used a constant windowing scheme of 40, 80 to focus on brain tissue. I am not sure what you mean with your question, though.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2786321": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5958987%2F274fde73db96188f794f53dac42a0a6f%2Fdiagram.png?generation=1714628266825394&alt=media)\nI'm thrilled to share the approach that led to my model securing the first place on the public leaderboard and third on the private leaderboard in this challenging competition. Unlike the first and second-place solutions that utilized existing EfficientNet architectures (EfficientNetV2S and EfficientNetV2L respectively), I developed a custom 3D CNN that diverges significantly from these popular choices.\n\nModel definition and training notebook can be found here: https://gitlab.com/Skelp/headct3dtinynet\nThe inference notebook will be pushed later, I will update this post once this is done.\n\nIn addition to the provided dataset, I also made use of the external, public dataset I mentioned here: https://www.kaggle.com/competitions/spr-head-ct-age-prediction-challenge/discussion/486285\n\n**Model Architecture and Features:**\n- **Parameters**: The model is relatively lightweight with approximately 2.3 million parameters.\n- **Attention Mechanisms**: It incorporates both channel and spatial attention mechanisms, enhancing its ability to focus on relevant features within the CT scans.\n- **Activation Function**: A key innovation in my model is the novel activation function, APTx. The exact function used is a per-channel-parameterized version of APTx, designed to emulate the performance of the MISH activation function but with reduced computational demand.\n- **Dimensionality**: The architecture is a full 3D CNN, leveraging the spatial context better than the commonly used 2.5D or 2D approaches.\n- **Imaging Parameters**: For preprocessing, I used a single windowing parameter set at 40, 80 to standardize the input data.\n- **Data Augmentation**: Utilized TorchIO for extensive data augmentation, which was critical in enhancing the model's robustness and generalizability.\n- **Optimization**: Initially, I used the AdaBelief optimizer with a learning rate of 3e-4. Upon observing plateaus in the leaderboard scores, I adjusted the learning rate to 3e-5 and introduced a weight decay of 0.5 for more stringent regularization, which significantly improved the model's performance.\n\nThese choices, particularly the use of a full 3D approach and the innovative APTx activation function, were pivotal in differentiating my model from others and achieving the high accuracy necessary to excel in this competition.\n\n**Detailed Model Architecture:**\n\nMy model's architecture is designed to effectively process 3D CT scan data for age prediction. Below are the specifics of the architecture, visualized in the attached diagram:\n\n- **Input Layer**: Accepts 3D data of size 128x128x128.\n- **Convolutional Blocks**: The model consists of a series of convolutional blocks (ConvBlock 1 to 6), each followed by 3D max pooling to reduce dimensionality while capturing the most salient features.\n  - **Block 1**: Outputs 32 channels, uses Group Normalization (groups=32, effectively instance norm), and includes both Channel Attention 3D and Spatial Attention 3D mechanisms. The APTx Activation, a novel activation function I developed, is used here.\n  - **Subsequent Blocks**: Increase in channel output progressively (64, 96, 128, 160, and finally 192).\n\n- **Flattening Layer**: After the convolutional blocks, the data is flattened to feed into the fully connected layers.\n\n- **Fully Connected Layer**: A dense layer that reduces the high-level features into a single prediction output.\n\n- **Output**: The mean of predictions is calculated to determine the final age estimate. Specifically, this refers to the mean of predictions from one or more head CT scans per patient. To handle multiple scans efficiently, I parallelized the processing by placing each scan into the batch dimension. This approach ensures that the model can manage additional data without compromising on processing speed or accuracy.\n\nThis architecture leverages attention mechanisms to focus on important spatial and channel-specific features, which is crucial for accurately interpreting medical images like CT scans. The use of APTx activation across the model helps to maintain computational efficiency while achieving performance comparable to more computationally intensive activations like MISH and also learn the optimal parameters per channel.\n\nI look forward to feedback and discussions on the approach and am happy to dive deeper into any aspect of the model that might interest fellow Kagglers!",
    "2788282": "Really neat! Congratulations! I had a standard 3D CNN approach as well but the performance wasn't great. I think the CBAM inspired attention modules and the activation function you chose really did it's job well. Thanks for sharing",
    "2786426": "\nThanks for sharing! \nIf you have time, I would like to see a performance comparison with the model implemented [here](https://github.com/ZFTurbo/timm_3d) and [here](https://pytorchvideo.readthedocs.io/en/latest/model_zoo.html)! ! \n\n\nAlso, how much epoch did you learn?",
    "2786360": "Thank you for the excellent solution!\nI had no success with the 3D model approach, so I find yours particularly intriguing. \n\nI have two questions:\n\n- My 3D experiments consistently showed predictions on the test set to be about 1 point worse than my validation scores. Did you observe a similar discrepancy between your validation and test scores?\n- As I'm not well-versed in 3D models, I'm curious whether an architecture combining 3D convolution and attention is common. Was there a specific architecture that inspired your approach?\n\nAchieving such performance with compact dimensions like 128 x 128 x 128 is astonishing and truly fascinating.",
    "2789238": "Hey! I was thinking about your solution, did you not think of using any residual connection in your architecture? ",
    "2789229": "Thanks for sharing your solution! You  solution was very interesting.  I tried to use a transformer based network (SWin and TimeSformer) for a single imaging approach but both were stagnating. Great to see that you managed to use Attention.\n\nAlso, I saw in preprocessing you used a window of 40 and 80 in the channel axis. How did you establish the cut points for the different patients?"
  }
}