{
  "id": 523105,
  "title": " [47th place solution] Pytorch Lightning Framework + Column-wise Ensemble",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523105",
  "author_name": "Lingji Kong",
  "post_date": "2024-07-30T09:04:09.013000",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Congratulations to all participants! I learned a lot in the past few months and want to thank Kaggle and the hosts for organizing this competition. It has been a truly rewarding experience, and I am excited to share my solution with you all.</p>\n<h2>Context</h2>\n<ul>\n<li>Bussiness context: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data</a></li>\n</ul>\n<h2>Code</h2>\n<p>Open source code is available on <a href=\"https://github.com/nemonemonee/leap-silver-medal-solution\" target=\"_blank\">github repo</a>.</p>\n<h2>TLDR</h2>\n<p>Our solution utilizes the PyTorch Lightning framework for training. All model architectures are based on the transformer encoder. Rotary positional encoding proved effective. Data preparation involved normalization, log transformation, min-max scaling, and clamping. Post-processing employed a column-wise ensemble method. While KAN linear embedding did not improve the score, it contributed as one of the models in the ensemble. We used 90% of the data for training and the remaining 10% for validation.</p>\n<h2>Training - Pytorch Lightning Framework</h2>\n<p>Our solution provides a convenient and neat approach for training machine learning models using the <code>PTLit</code> (PyTorch Lightning module) class. The training and validation loops are fully encapsulated within the class, simplifying the process. When multiple models are used, the class automatically averages their outputs to create an ensemble.</p>\n<ul>\n<li>To train a specific model:</li>\n</ul>\n<pre><code>mdlit = PTLit(mask, learning_rate, step_size, gamma, [YOUR_MODEL()])\n</code></pre>\n<ul>\n<li>To train multiple models as an ensemble:</li>\n</ul>\n<pre><code>mdlit = PTLit(mask, learning_rate, step_size, gamma, [YOUR_1ST_MODEL(), YOUR_2ND_MODEL()])\n</code></pre>\n<ul>\n<li>To load models from checkpoints and train them as an ensemble:</li>\n</ul>\n<pre><code>base_models = PTLit.load_from_checkpoint().models\nbase_models += PTLit.load_from_checkpoint().models\nmdlit = PTLit(mask, learning_rate, step_size, gamma, base_models)\n</code></pre>\n<h2>Models</h2>\n<h3>Rotary positional encoding</h3>\n<p>In our experiment, switching from the original positional encoding to <a href=\"https://arxiv.org/abs/2104.09864\" target=\"_blank\">Rotary positional encoding (RoPE)</a> resulted in an improvement of approximately 0.005 in the public score. Our best single model combined RoPE with a basic transformer encoder, achieving a score of 0.751 on the public leaderboard and 0.744 on the private leaderboard.</p>\n<h3>Model Architectures - What we also tried</h3>\n<p>Our JNet model architecture processes the input sequence of shape (B, 60, 25) through a series of Conv1d layers, each followed by GELU activation functions, to extract hierarchical features. The outputs of these layers are concatenated into a combined feature representation, an embedding. This embedding is enriched with Rotary Positional Encoding (RoPE) to provide positional information and is subsequently processed by a transformer encoder. Finally, a linear layer generates the output sequence of shape (B, 60, 14). </p>\n<p>\n    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17111680%2F15261fc27912dd8005aaaf6deeaf7101%2Fjnet.png?generation=1722330153587269&amp;alt=media\" alt=\"JNet Arch\">\n</p>\n<p>Our KANformer model architecture begins with an input sequence of shape (B, 60, 25) and applies a series of 60 <a href=\"https://arxiv.org/abs/2404.19756\" target=\"_blank\">KAN</a> linear layers (25, 256) to each token of the sequence. These layers transform the input into an embedding of shape (B, 60, 256). This embedding is then processed by a transformer encoder to capture dependencies within the sequence. Finally, a linear layer generates the output sequence of shape (B, 60, 14). While KAN linear embedding did not improve the score, it contributed as one of the models in the ensemble.</p>\n<p>\n    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17111680%2F3786e3b064eab9366cb451b3b4308cc8%2Fjnet.png?generation=1722329905316227&amp;alt=media\" alt=\"KAN former Arch\">\n</p>\n<h2>Data Preprocessing</h2>\n<p>Without domain knowledge, we analyzed the distribution histograms of all feature columns one by one to determine the need for log transformation to address skewness. Afterward, we manually decided whether to apply normalization or min-max scaling to each feature column.</p>\n<p>For the label columns, direct normalization caused extreme values for the minimum and maximum. To mitigate this, we first clamped the label columns and then applied normalization. While clamping made it easier for the model to learn initially, it did not result in a better score after convergence. However, this approach can still be useful for curriculum learning. We included models trained on clamped data as part of the ensemble in the final submission.</p>\n<h2>Data Postprocessing - Column-wise Ensemble</h2>\n<p>We used <code>optuna</code> to search for the best weights to ensemble models for each column. The code is as follows:</p>\n<pre><code>num_targets = label.size()\npreds_em = torch.zeros_like(preds[])\nalphas = torch.zeros(num_targets, num_models)\n\n ():\n    weights = []\n    remaining_sum = \n     j  (num_models - ):\n        w_i = trial.suggest_float(, , remaining_sum, step=)\n        w = w_i / \n        weights.append(w)\n        remaining_sum -= w_i\n    col = (weight * preds[m][:, i]  m, weight  (weights))\n     r2_score(col, labeln[:, i])\n\n i  tqdm((num_targets)):\n     mask[i]:\n        study = optuna.create_study(direction=)\n        study.optimize( trial: objective(trial, i), n_trials=)\n        best_weights = []\n        remaining_sum = \n         j  (num_models - ):\n            w = study.best_params[]\n            best_weights.append(w / )\n            remaining_sum -= w\n        best_weights.append(remaining_sum / ) \n        alphas[i] = torch.tensor(best_weights)\n        preds_em[:, i] = (weight * preds[m][:, i]  m, weight  (best_weights))\n</code></pre>\n<h2>What didn't work</h2>\n<p>We experimented with changing the linear output layer to a convolutional output layer and a KAN linear output layer. However, these modifications increased the likelihood of overfitting. Due to time constraints, we were unable to implement adequate regularization techniques to mitigate this issue effectively.</p>\n<h2>References</h2>\n<p>[1] Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., &amp; Liu, Y. (2021, April 20). RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv.org. <a href=\"https://arxiv.org/abs/2104.09864\" target=\"_blank\">https://arxiv.org/abs/2104.09864</a></p>\n<p>[2] Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljačić, M., Hou, T. Y., &amp; Tegmark, M. (2024, April 30). KAN: Kolmogorov-Arnold Networks. arXiv.org. <a href=\"https://arxiv.org/abs/2404.19756\" target=\"_blank\">https://arxiv.org/abs/2404.19756</a></p>",
  "messages": [
    {
      "id": 2940627,
      "postDate": "2024-07-30T09:04:09.013Z",
      "content": "<p>Congratulations to all participants! I learned a lot in the past few months and want to thank Kaggle and the hosts for organizing this competition. It has been a truly rewarding experience, and I am excited to share my solution with you all.</p>\n<h2>Context</h2>\n<ul>\n<li>Bussiness context: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data</a></li>\n</ul>\n<h2>Code</h2>\n<p>Open source code is available on <a href=\"https://github.com/nemonemonee/leap-silver-medal-solution\" target=\"_blank\">github repo</a>.</p>\n<h2>TLDR</h2>\n<p>Our solution utilizes the PyTorch Lightning framework for training. All model architectures are based on the transformer encoder. Rotary positional encoding proved effective. Data preparation involved normalization, log transformation, min-max scaling, and clamping. Post-processing employed a column-wise ensemble method. While KAN linear embedding did not improve the score, it contributed as one of the models in the ensemble. We used 90% of the data for training and the remaining 10% for validation.</p>\n<h2>Training - Pytorch Lightning Framework</h2>\n<p>Our solution provides a convenient and neat approach for training machine learning models using the <code>PTLit</code> (PyTorch Lightning module) class. The training and validation loops are fully encapsulated within the class, simplifying the process. When multiple models are used, the class automatically averages their outputs to create an ensemble.</p>\n<ul>\n<li>To train a specific model:</li>\n</ul>\n<pre><code>mdlit = PTLit(mask, learning_rate, step_size, gamma, [YOUR_MODEL()])\n</code></pre>\n<ul>\n<li>To train multiple models as an ensemble:</li>\n</ul>\n<pre><code>mdlit = PTLit(mask, learning_rate, step_size, gamma, [YOUR_1ST_MODEL(), YOUR_2ND_MODEL()])\n</code></pre>\n<ul>\n<li>To load models from checkpoints and train them as an ensemble:</li>\n</ul>\n<pre><code>base_models = PTLit.load_from_checkpoint().models\nbase_models += PTLit.load_from_checkpoint().models\nmdlit = PTLit(mask, learning_rate, step_size, gamma, base_models)\n</code></pre>\n<h2>Models</h2>\n<h3>Rotary positional encoding</h3>\n<p>In our experiment, switching from the original positional encoding to <a href=\"https://arxiv.org/abs/2104.09864\" target=\"_blank\">Rotary positional encoding (RoPE)</a> resulted in an improvement of approximately 0.005 in the public score. Our best single model combined RoPE with a basic transformer encoder, achieving a score of 0.751 on the public leaderboard and 0.744 on the private leaderboard.</p>\n<h3>Model Architectures - What we also tried</h3>\n<p>Our JNet model architecture processes the input sequence of shape (B, 60, 25) through a series of Conv1d layers, each followed by GELU activation functions, to extract hierarchical features. The outputs of these layers are concatenated into a combined feature representation, an embedding. This embedding is enriched with Rotary Positional Encoding (RoPE) to provide positional information and is subsequently processed by a transformer encoder. Finally, a linear layer generates the output sequence of shape (B, 60, 14). </p>\n<p>\n    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17111680%2F15261fc27912dd8005aaaf6deeaf7101%2Fjnet.png?generation=1722330153587269&amp;alt=media\" alt=\"JNet Arch\">\n</p>\n<p>Our KANformer model architecture begins with an input sequence of shape (B, 60, 25) and applies a series of 60 <a href=\"https://arxiv.org/abs/2404.19756\" target=\"_blank\">KAN</a> linear layers (25, 256) to each token of the sequence. These layers transform the input into an embedding of shape (B, 60, 256). This embedding is then processed by a transformer encoder to capture dependencies within the sequence. Finally, a linear layer generates the output sequence of shape (B, 60, 14). While KAN linear embedding did not improve the score, it contributed as one of the models in the ensemble.</p>\n<p>\n    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17111680%2F3786e3b064eab9366cb451b3b4308cc8%2Fjnet.png?generation=1722329905316227&amp;alt=media\" alt=\"KAN former Arch\">\n</p>\n<h2>Data Preprocessing</h2>\n<p>Without domain knowledge, we analyzed the distribution histograms of all feature columns one by one to determine the need for log transformation to address skewness. Afterward, we manually decided whether to apply normalization or min-max scaling to each feature column.</p>\n<p>For the label columns, direct normalization caused extreme values for the minimum and maximum. To mitigate this, we first clamped the label columns and then applied normalization. While clamping made it easier for the model to learn initially, it did not result in a better score after convergence. However, this approach can still be useful for curriculum learning. We included models trained on clamped data as part of the ensemble in the final submission.</p>\n<h2>Data Postprocessing - Column-wise Ensemble</h2>\n<p>We used <code>optuna</code> to search for the best weights to ensemble models for each column. The code is as follows:</p>\n<pre><code>num_targets = label.size()\npreds_em = torch.zeros_like(preds[])\nalphas = torch.zeros(num_targets, num_models)\n\n ():\n    weights = []\n    remaining_sum = \n     j  (num_models - ):\n        w_i = trial.suggest_float(, , remaining_sum, step=)\n        w = w_i / \n        weights.append(w)\n        remaining_sum -= w_i\n    col = (weight * preds[m][:, i]  m, weight  (weights))\n     r2_score(col, labeln[:, i])\n\n i  tqdm((num_targets)):\n     mask[i]:\n        study = optuna.create_study(direction=)\n        study.optimize( trial: objective(trial, i), n_trials=)\n        best_weights = []\n        remaining_sum = \n         j  (num_models - ):\n            w = study.best_params[]\n            best_weights.append(w / )\n            remaining_sum -= w\n        best_weights.append(remaining_sum / ) \n        alphas[i] = torch.tensor(best_weights)\n        preds_em[:, i] = (weight * preds[m][:, i]  m, weight  (best_weights))\n</code></pre>\n<h2>What didn't work</h2>\n<p>We experimented with changing the linear output layer to a convolutional output layer and a KAN linear output layer. However, these modifications increased the likelihood of overfitting. Due to time constraints, we were unable to implement adequate regularization techniques to mitigate this issue effectively.</p>\n<h2>References</h2>\n<p>[1] Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., &amp; Liu, Y. (2021, April 20). RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv.org. <a href=\"https://arxiv.org/abs/2104.09864\" target=\"_blank\">https://arxiv.org/abs/2104.09864</a></p>\n<p>[2] Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljačić, M., Hou, T. Y., &amp; Tegmark, M. (2024, April 30). KAN: Kolmogorov-Arnold Networks. arXiv.org. <a href=\"https://arxiv.org/abs/2404.19756\" target=\"_blank\">https://arxiv.org/abs/2404.19756</a></p>",
      "rawMarkdown": "Congratulations to all participants! I learned a lot in the past few months and want to thank Kaggle and the hosts for organizing this competition. It has been a truly rewarding experience, and I am excited to share my solution with you all.\n\n## Context\n\n* Bussiness context: <https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview>\n* Data context: <https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data>\n\n## Code\n\nOpen source code is available on [github repo](https://github.com/nemonemonee/leap-silver-medal-solution).\n\n## TLDR\n\nOur solution utilizes the PyTorch Lightning framework for training. All model architectures are based on the transformer encoder. Rotary positional encoding proved effective. Data preparation involved normalization, log transformation, min-max scaling, and clamping. Post-processing employed a column-wise ensemble method. While KAN linear embedding did not improve the score, it contributed as one of the models in the ensemble. We used 90% of the data for training and the remaining 10% for validation.\n\n## Training - Pytorch Lightning Framework\n\nOur solution provides a convenient and neat approach for training machine learning models using the `PTLit` (PyTorch Lightning module) class. The training and validation loops are fully encapsulated within the class, simplifying the process. When multiple models are used, the class automatically averages their outputs to create an ensemble.\n\n* To train a specific model:\n``` python\nmdlit = PTLit(mask, learning_rate, step_size, gamma, [YOUR_MODEL()])\n```\n* To train multiple models as an ensemble:\n``` python\nmdlit = PTLit(mask, learning_rate, step_size, gamma, [YOUR_1ST_MODEL(), YOUR_2ND_MODEL()])\n```\n* To load models from checkpoints and train them as an ensemble:\n``` python\nbase_models = PTLit.load_from_checkpoint(\"PATH_TO_THE_1ST_SAVED_CHECKPOINT\").models\nbase_models += PTLit.load_from_checkpoint(\"PATH_TO_THE_2ND_SAVED_CHECKPOINT\").models\nmdlit = PTLit(mask, learning_rate, step_size, gamma, base_models)\n```\n\n## Models\n### Rotary positional encoding\n\nIn our experiment, switching from the original positional encoding to [Rotary positional encoding (RoPE)](https://arxiv.org/abs/2104.09864) resulted in an improvement of approximately 0.005 in the public score. Our best single model combined RoPE with a basic transformer encoder, achieving a score of 0.751 on the public leaderboard and 0.744 on the private leaderboard.\n\n### Model Architectures - What we also tried\n\nOur JNet model architecture processes the input sequence of shape (B, 60, 25) through a series of Conv1d layers, each followed by GELU activation functions, to extract hierarchical features. The outputs of these layers are concatenated into a combined feature representation, an embedding. This embedding is enriched with Rotary Positional Encoding (RoPE) to provide positional information and is subsequently processed by a transformer encoder. Finally, a linear layer generates the output sequence of shape (B, 60, 14). \n\n<p align=\"center\">\n    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17111680%2F15261fc27912dd8005aaaf6deeaf7101%2Fjnet.png?generation=1722330153587269&alt=media\" alt=\"JNet Arch\" width=\"500\"/>\n</p>\n\nOur KANformer model architecture begins with an input sequence of shape (B, 60, 25) and applies a series of 60 [KAN](https://arxiv.org/abs/2404.19756) linear layers (25, 256) to each token of the sequence. These layers transform the input into an embedding of shape (B, 60, 256). This embedding is then processed by a transformer encoder to capture dependencies within the sequence. Finally, a linear layer generates the output sequence of shape (B, 60, 14). While KAN linear embedding did not improve the score, it contributed as one of the models in the ensemble.\n\n<p align=\"center\">\n    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17111680%2F3786e3b064eab9366cb451b3b4308cc8%2Fjnet.png?generation=1722329905316227&alt=media\" alt=\"KAN former Arch\" width=\"500\"/>\n</p>\n\n## Data Preprocessing\n\nWithout domain knowledge, we analyzed the distribution histograms of all feature columns one by one to determine the need for log transformation to address skewness. Afterward, we manually decided whether to apply normalization or min-max scaling to each feature column.\n\nFor the label columns, direct normalization caused extreme values for the minimum and maximum. To mitigate this, we first clamped the label columns and then applied normalization. While clamping made it easier for the model to learn initially, it did not result in a better score after convergence. However, this approach can still be useful for curriculum learning. We included models trained on clamped data as part of the ensemble in the final submission.\n\n## Data Postprocessing - Column-wise Ensemble\nWe used `optuna` to search for the best weights to ensemble models for each column. The code is as follows:\n\n```python\nnum_targets = label.size(1)\npreds_em = torch.zeros_like(preds[0])\nalphas = torch.zeros(num_targets, num_models)\n\ndef objective(trial, i):\n    weights = []\n    remaining_sum = 10\n    for j in range(num_models - 1):\n        w_i = trial.suggest_float(f'weight_{j}', 0, remaining_sum, step=.5)\n        w = w_i / 10.\n        weights.append(w)\n        remaining_sum -= w_i\n    col = sum(weight * preds[m][:, i] for m, weight in enumerate(weights))\n    return r2_score(col, labeln[:, i])\n\nfor i in tqdm(range(num_targets)):\n    if mask[i]:\n        study = optuna.create_study(direction='maximize')\n        study.optimize(lambda trial: objective(trial, i), n_trials=50)\n        best_weights = []\n        remaining_sum = 10.0\n        for j in range(num_models - 1):\n            w = study.best_params[f'weight_{j}']\n            best_weights.append(w / 10.)\n            remaining_sum -= w\n        best_weights.append(remaining_sum / 10.) \n        alphas[i] = torch.tensor(best_weights)\n        preds_em[:, i] = sum(weight * preds[m][:, i] for m, weight in enumerate(best_weights))\n```\n\n## What didn't work\n\nWe experimented with changing the linear output layer to a convolutional output layer and a KAN linear output layer. However, these modifications increased the likelihood of overfitting. Due to time constraints, we were unable to implement adequate regularization techniques to mitigate this issue effectively.\n\n## References\n\n[1] Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., & Liu, Y. (2021, April 20). RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv.org. https://arxiv.org/abs/2104.09864\n\n[2] Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljačić, M., Hou, T. Y., & Tegmark, M. (2024, April 30). KAN: Kolmogorov-Arnold Networks. arXiv.org. https://arxiv.org/abs/2404.19756",
      "votes": 11
    },
    {
      "id": 2948459,
      "postDate": "2024-08-05T23:52:16.337Z",
      "content": "<p>Nice work! Thanks for sharing code to use PyTorch lightning! Very helpful!</p>",
      "rawMarkdown": "Nice work! Thanks for sharing code to use PyTorch lightning! Very helpful!",
      "votes": 1
    },
    {
      "id": 2940895,
      "postDate": "2024-07-30T14:31:09.687Z",
      "content": "<p>Thank you for sharing with us. I definitely try this code in this competition for learning.</p>",
      "rawMarkdown": "Thank you for sharing with us. I definitely try this code in this competition for learning.",
      "votes": 1
    },
    {
      "id": 2940758,
      "postDate": "2024-07-30T12:26:46.613Z",
      "content": "<p>Impressive work, <a href=\"https://www.kaggle.com/lingjikong\" target=\"_blank\">@lingjikong</a>!<br>\nYour solution, with its use of PyTorch Lightning and column-wise ensemble, shows great attention to detail. The integration of Rotary Positional Encoding and your various model architectures, like JNet and KANformer, highlight a thoughtful approach to improving performance. I appreciate the detailed breakdown of your preprocessing and postprocessing techniques—it’s clear a lot of effort went into optimizing every step. Sharing your code and insights is incredibly valuable for the community. <br>\nKeep up the great work, and thanks for sharing your learnings.</p>",
      "rawMarkdown": "Impressive work, @lingjikong!\nYour solution, with its use of PyTorch Lightning and column-wise ensemble, shows great attention to detail. The integration of Rotary Positional Encoding and your various model architectures, like JNet and KANformer, highlight a thoughtful approach to improving performance. I appreciate the detailed breakdown of your preprocessing and postprocessing techniques—it’s clear a lot of effort went into optimizing every step. Sharing your code and insights is incredibly valuable for the community. \nKeep up the great work, and thanks for sharing your learnings.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2948459,
      "author_name": "Jiayi Ding",
      "author_url": "",
      "post_date": "2024-08-05T23:52:16.337000",
      "content": "<p>Nice work! Thanks for sharing code to use PyTorch lightning! Very helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2940895,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T14:31:09.687000",
      "content": "<p>Thank you for sharing with us. I definitely try this code in this competition for learning.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2940758,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T12:26:46.613000",
      "content": "<p>Impressive work, <a href=\"https://www.kaggle.com/lingjikong\" target=\"_blank\">@lingjikong</a>!<br>\nYour solution, with its use of PyTorch Lightning and column-wise ensemble, shows great attention to detail. The integration of Rotary Positional Encoding and your various model architectures, like JNet and KANformer, highlight a thoughtful approach to improving performance. I appreciate the detailed breakdown of your preprocessing and postprocessing techniques—it’s clear a lot of effort went into optimizing every step. Sharing your code and insights is incredibly valuable for the community. <br>\nKeep up the great work, and thanks for sharing your learnings.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2940627": "Congratulations to all participants! I learned a lot in the past few months and want to thank Kaggle and the hosts for organizing this competition. It has been a truly rewarding experience, and I am excited to share my solution with you all.\n\n## Context\n\n* Bussiness context: <https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview>\n* Data context: <https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data>\n\n## Code\n\nOpen source code is available on [github repo](https://github.com/nemonemonee/leap-silver-medal-solution).\n\n## TLDR\n\nOur solution utilizes the PyTorch Lightning framework for training. All model architectures are based on the transformer encoder. Rotary positional encoding proved effective. Data preparation involved normalization, log transformation, min-max scaling, and clamping. Post-processing employed a column-wise ensemble method. While KAN linear embedding did not improve the score, it contributed as one of the models in the ensemble. We used 90% of the data for training and the remaining 10% for validation.\n\n## Training - Pytorch Lightning Framework\n\nOur solution provides a convenient and neat approach for training machine learning models using the `PTLit` (PyTorch Lightning module) class. The training and validation loops are fully encapsulated within the class, simplifying the process. When multiple models are used, the class automatically averages their outputs to create an ensemble.\n\n* To train a specific model:\n``` python\nmdlit = PTLit(mask, learning_rate, step_size, gamma, [YOUR_MODEL()])\n```\n* To train multiple models as an ensemble:\n``` python\nmdlit = PTLit(mask, learning_rate, step_size, gamma, [YOUR_1ST_MODEL(), YOUR_2ND_MODEL()])\n```\n* To load models from checkpoints and train them as an ensemble:\n``` python\nbase_models = PTLit.load_from_checkpoint(\"PATH_TO_THE_1ST_SAVED_CHECKPOINT\").models\nbase_models += PTLit.load_from_checkpoint(\"PATH_TO_THE_2ND_SAVED_CHECKPOINT\").models\nmdlit = PTLit(mask, learning_rate, step_size, gamma, base_models)\n```\n\n## Models\n### Rotary positional encoding\n\nIn our experiment, switching from the original positional encoding to [Rotary positional encoding (RoPE)](https://arxiv.org/abs/2104.09864) resulted in an improvement of approximately 0.005 in the public score. Our best single model combined RoPE with a basic transformer encoder, achieving a score of 0.751 on the public leaderboard and 0.744 on the private leaderboard.\n\n### Model Architectures - What we also tried\n\nOur JNet model architecture processes the input sequence of shape (B, 60, 25) through a series of Conv1d layers, each followed by GELU activation functions, to extract hierarchical features. The outputs of these layers are concatenated into a combined feature representation, an embedding. This embedding is enriched with Rotary Positional Encoding (RoPE) to provide positional information and is subsequently processed by a transformer encoder. Finally, a linear layer generates the output sequence of shape (B, 60, 14). \n\n<p align=\"center\">\n    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17111680%2F15261fc27912dd8005aaaf6deeaf7101%2Fjnet.png?generation=1722330153587269&alt=media\" alt=\"JNet Arch\" width=\"500\"/>\n</p>\n\nOur KANformer model architecture begins with an input sequence of shape (B, 60, 25) and applies a series of 60 [KAN](https://arxiv.org/abs/2404.19756) linear layers (25, 256) to each token of the sequence. These layers transform the input into an embedding of shape (B, 60, 256). This embedding is then processed by a transformer encoder to capture dependencies within the sequence. Finally, a linear layer generates the output sequence of shape (B, 60, 14). While KAN linear embedding did not improve the score, it contributed as one of the models in the ensemble.\n\n<p align=\"center\">\n    <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17111680%2F3786e3b064eab9366cb451b3b4308cc8%2Fjnet.png?generation=1722329905316227&alt=media\" alt=\"KAN former Arch\" width=\"500\"/>\n</p>\n\n## Data Preprocessing\n\nWithout domain knowledge, we analyzed the distribution histograms of all feature columns one by one to determine the need for log transformation to address skewness. Afterward, we manually decided whether to apply normalization or min-max scaling to each feature column.\n\nFor the label columns, direct normalization caused extreme values for the minimum and maximum. To mitigate this, we first clamped the label columns and then applied normalization. While clamping made it easier for the model to learn initially, it did not result in a better score after convergence. However, this approach can still be useful for curriculum learning. We included models trained on clamped data as part of the ensemble in the final submission.\n\n## Data Postprocessing - Column-wise Ensemble\nWe used `optuna` to search for the best weights to ensemble models for each column. The code is as follows:\n\n```python\nnum_targets = label.size(1)\npreds_em = torch.zeros_like(preds[0])\nalphas = torch.zeros(num_targets, num_models)\n\ndef objective(trial, i):\n    weights = []\n    remaining_sum = 10\n    for j in range(num_models - 1):\n        w_i = trial.suggest_float(f'weight_{j}', 0, remaining_sum, step=.5)\n        w = w_i / 10.\n        weights.append(w)\n        remaining_sum -= w_i\n    col = sum(weight * preds[m][:, i] for m, weight in enumerate(weights))\n    return r2_score(col, labeln[:, i])\n\nfor i in tqdm(range(num_targets)):\n    if mask[i]:\n        study = optuna.create_study(direction='maximize')\n        study.optimize(lambda trial: objective(trial, i), n_trials=50)\n        best_weights = []\n        remaining_sum = 10.0\n        for j in range(num_models - 1):\n            w = study.best_params[f'weight_{j}']\n            best_weights.append(w / 10.)\n            remaining_sum -= w\n        best_weights.append(remaining_sum / 10.) \n        alphas[i] = torch.tensor(best_weights)\n        preds_em[:, i] = sum(weight * preds[m][:, i] for m, weight in enumerate(best_weights))\n```\n\n## What didn't work\n\nWe experimented with changing the linear output layer to a convolutional output layer and a KAN linear output layer. However, these modifications increased the likelihood of overfitting. Due to time constraints, we were unable to implement adequate regularization techniques to mitigate this issue effectively.\n\n## References\n\n[1] Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., & Liu, Y. (2021, April 20). RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv.org. https://arxiv.org/abs/2104.09864\n\n[2] Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljačić, M., Hou, T. Y., & Tegmark, M. (2024, April 30). KAN: Kolmogorov-Arnold Networks. arXiv.org. https://arxiv.org/abs/2404.19756",
    "2948459": "Nice work! Thanks for sharing code to use PyTorch lightning! Very helpful!",
    "2940895": "Thank you for sharing with us. I definitely try this code in this competition for learning.",
    "2940758": "Impressive work, @lingjikong!\nYour solution, with its use of PyTorch Lightning and column-wise ensemble, shows great attention to detail. The integration of Rotary Positional Encoding and your various model architectures, like JNet and KANformer, highlight a thoughtful approach to improving performance. I appreciate the detailed breakdown of your preprocessing and postprocessing techniques—it’s clear a lot of effort went into optimizing every step. Sharing your code and insights is incredibly valuable for the community. \nKeep up the great work, and thanks for sharing your learnings."
  }
}