{
  "id": 608878,
  "title": "Meaningful(*less?) scoring function",
  "url": "/competitions/ariel-data-challenge-2025/discussion/608878",
  "author_name": "Aleksandr Zakuskin",
  "post_date": "2025-09-22T16:31:50.508000",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'd like to raise some concerns about the scoring function used in the competition. And, hopefully, resolve them by filling the gaps in my understanding.</p>\n<p>I will give three examples that raise some questions in my head when I'm looking at them.<br>\nFirst of all, in these examples I deal with only AIRS or synthetic AIRS-like data, so that there is no need for weighting and the <em>score</em> function becomes simpler and more readable.</p>\n<p>A notebook that makes figures and canculates scores discussed further is available <a href=\"https://www.kaggle.com/code/aleksandrzakuskin/gll-score-illustrations\" target=\"_blank\">here</a></p>\n<h2>Example 1. Ideal spectrum vs. straight line prediction.</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2Fc3e4c4938c85f312218e421fc474a7b7%2FExample%201.PNG?generation=1758539960614763&amp;alt=media\" alt=\"\"></p>\n<p>The blue line is a \"true\" spectrum made up as a sum of two random Gaussian profiles. Vertical bars at the edges are uncertainties (0.00055 in all cases in this Example). Green prediction (mean of \"true\") is obviously a relatively good one and gives 0.99904. The orange and the red are \"true\" and \"mean\" shifted by the same value (0.005). <strong>Here is the first question: why is shifted \"mean\" is better that shifted \"true\"?</strong></p>\n<p>Yes, the difference is very small: 0.98942 vs 0.98939, but still. If we are dealing with <strong>spectra</strong> in this competition, aren't their shapes important?<br>\nSome more observations without particular questions: (1) such behavior is observed only when the prediction lies between true and naive; (2) the difference in score between shifted \"mean\" and shifted \"true\" grows as the predicted uncertainties decrease.</p>\n<h2>Example 2. Good and bad predictions.</h2>\n<p>Here I'd like to compare a couple of true Ariel spectra and our predictions for them.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2Fc7c67353b48677ab4ec1c9d42abb595e%2FExample%202.PNG?generation=1758558322499414&amp;alt=media\" alt=\"\"></p>\n<p>Blue areas are our predictions with uncertainties. Uncertainties are constant across all planets and all spectral points. On the right-hand side not even all \"true\" datapoints lie wihtin the estimated uncertainty, but the score is almost perfect (0.99960). In the case of the left plot, it is true everywhere, but score is 0.32894.<br>\nMathematically, the answer is simple: the left spectrum lies very close to the naive prediction (0.015), while the right is far from it. So, if naive prediction is good enough, the scoring function becomes very strict and penalizes any little deviation in actual prediction. But if target is far from naive, straight line prediction becomes just fine.<br>\nIsn't some physical meaning lost in evaluation of these two spectra in the figure?</p>\n<h2>Example 3. Scoring sensitivity.</h2>\n<p>Now we know, that scoring has different sinsitivity depending on distance between true and naive. But what is the sensitivity?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2F44a7a29684a772c41acd5d1e85f5d1df%2FExample%203.PNG?generation=1758555079070701&amp;alt=media\" alt=\"\"></p>\n<p>Here is one arbitrary prediction (blue) for one arbirary planet (orange) that gives a score of 0.99917. Both are far from naive, so we do not need to worry about shape. But what if we intentionally make prediction worse?</p>\n<p>We can add a small constant (0.004, red) and get 0.99796. Let's add much greater constant (0.013, cyan), so that our prediction is further from \"true\" than \"naive\". And we still have 0.98726. Even predicting zeros with the same constant sigma (1e-15 * pred, magenta) results in 0.998…</p>\n<p>If anyone has any thoughts on phyical meanigfulness of using such score fore spectra calculation/prediction, I'm curious to hear your opinions.</p>",
  "messages": [
    {
      "id": 3292839,
      "postDate": "2025-09-22T16:31:50.510Z",
      "content": "<p>I'd like to raise some concerns about the scoring function used in the competition. And, hopefully, resolve them by filling the gaps in my understanding.</p>\n<p>I will give three examples that raise some questions in my head when I'm looking at them.<br>\nFirst of all, in these examples I deal with only AIRS or synthetic AIRS-like data, so that there is no need for weighting and the <em>score</em> function becomes simpler and more readable.</p>\n<p>A notebook that makes figures and canculates scores discussed further is available <a href=\"https://www.kaggle.com/code/aleksandrzakuskin/gll-score-illustrations\" target=\"_blank\">here</a></p>\n<h2>Example 1. Ideal spectrum vs. straight line prediction.</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2Fc3e4c4938c85f312218e421fc474a7b7%2FExample%201.PNG?generation=1758539960614763&amp;alt=media\" alt=\"\"></p>\n<p>The blue line is a \"true\" spectrum made up as a sum of two random Gaussian profiles. Vertical bars at the edges are uncertainties (0.00055 in all cases in this Example). Green prediction (mean of \"true\") is obviously a relatively good one and gives 0.99904. The orange and the red are \"true\" and \"mean\" shifted by the same value (0.005). <strong>Here is the first question: why is shifted \"mean\" is better that shifted \"true\"?</strong></p>\n<p>Yes, the difference is very small: 0.98942 vs 0.98939, but still. If we are dealing with <strong>spectra</strong> in this competition, aren't their shapes important?<br>\nSome more observations without particular questions: (1) such behavior is observed only when the prediction lies between true and naive; (2) the difference in score between shifted \"mean\" and shifted \"true\" grows as the predicted uncertainties decrease.</p>\n<h2>Example 2. Good and bad predictions.</h2>\n<p>Here I'd like to compare a couple of true Ariel spectra and our predictions for them.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2Fc7c67353b48677ab4ec1c9d42abb595e%2FExample%202.PNG?generation=1758558322499414&amp;alt=media\" alt=\"\"></p>\n<p>Blue areas are our predictions with uncertainties. Uncertainties are constant across all planets and all spectral points. On the right-hand side not even all \"true\" datapoints lie wihtin the estimated uncertainty, but the score is almost perfect (0.99960). In the case of the left plot, it is true everywhere, but score is 0.32894.<br>\nMathematically, the answer is simple: the left spectrum lies very close to the naive prediction (0.015), while the right is far from it. So, if naive prediction is good enough, the scoring function becomes very strict and penalizes any little deviation in actual prediction. But if target is far from naive, straight line prediction becomes just fine.<br>\nIsn't some physical meaning lost in evaluation of these two spectra in the figure?</p>\n<h2>Example 3. Scoring sensitivity.</h2>\n<p>Now we know, that scoring has different sinsitivity depending on distance between true and naive. But what is the sensitivity?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2F44a7a29684a772c41acd5d1e85f5d1df%2FExample%203.PNG?generation=1758555079070701&amp;alt=media\" alt=\"\"></p>\n<p>Here is one arbitrary prediction (blue) for one arbirary planet (orange) that gives a score of 0.99917. Both are far from naive, so we do not need to worry about shape. But what if we intentionally make prediction worse?</p>\n<p>We can add a small constant (0.004, red) and get 0.99796. Let's add much greater constant (0.013, cyan), so that our prediction is further from \"true\" than \"naive\". And we still have 0.98726. Even predicting zeros with the same constant sigma (1e-15 * pred, magenta) results in 0.998…</p>\n<p>If anyone has any thoughts on phyical meanigfulness of using such score fore spectra calculation/prediction, I'm curious to hear your opinions.</p>",
      "rawMarkdown": "I'd like to raise some concerns about the scoring function used in the competition. And, hopefully, resolve them by filling the gaps in my understanding.\n\nI will give three examples that raise some questions in my head when I'm looking at them.\nFirst of all, in these examples I deal with only AIRS or synthetic AIRS-like data, so that there is no need for weighting and the _score_ function becomes simpler and more readable.\n\nA notebook that makes figures and canculates scores discussed further is available [here](https://www.kaggle.com/code/aleksandrzakuskin/gll-score-illustrations)\n\n## Example 1. Ideal spectrum vs. straight line prediction.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2Fc3e4c4938c85f312218e421fc474a7b7%2FExample%201.PNG?generation=1758539960614763&alt=media)\n\nThe blue line is a \"true\" spectrum made up as a sum of two random Gaussian profiles. Vertical bars at the edges are uncertainties (0.00055 in all cases in this Example). Green prediction (mean of \"true\") is obviously a relatively good one and gives 0.99904. The orange and the red are \"true\" and \"mean\" shifted by the same value (0.005). **Here is the first question: why is shifted \"mean\" is better that shifted \"true\"?**\n\nYes, the difference is very small: 0.98942 vs 0.98939, but still. If we are dealing with **spectra** in this competition, aren't their shapes important?\nSome more observations without particular questions: (1) such behavior is observed only when the prediction lies between true and naive; (2) the difference in score between shifted \"mean\" and shifted \"true\" grows as the predicted uncertainties decrease.\n\n## Example 2. Good and bad predictions.\n\nHere I'd like to compare a couple of true Ariel spectra and our predictions for them.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2Fc7c67353b48677ab4ec1c9d42abb595e%2FExample%202.PNG?generation=1758558322499414&alt=media)\n\nBlue areas are our predictions with uncertainties. Uncertainties are constant across all planets and all spectral points. On the right-hand side not even all \"true\" datapoints lie wihtin the estimated uncertainty, but the score is almost perfect (0.99960). In the case of the left plot, it is true everywhere, but score is 0.32894.\nMathematically, the answer is simple: the left spectrum lies very close to the naive prediction (0.015), while the right is far from it. So, if naive prediction is good enough, the scoring function becomes very strict and penalizes any little deviation in actual prediction. But if target is far from naive, straight line prediction becomes just fine.\nIsn't some physical meaning lost in evaluation of these two spectra in the figure?\n\n## Example 3. Scoring sensitivity.\n\nNow we know, that scoring has different sinsitivity depending on distance between true and naive. But what is the sensitivity?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2F44a7a29684a772c41acd5d1e85f5d1df%2FExample%203.PNG?generation=1758555079070701&alt=media)\n\nHere is one arbitrary prediction (blue) for one arbirary planet (orange) that gives a score of 0.99917. Both are far from naive, so we do not need to worry about shape. But what if we intentionally make prediction worse?\n\nWe can add a small constant (0.004, red) and get 0.99796. Let's add much greater constant (0.013, cyan), so that our prediction is further from \"true\" than \"naive\". And we still have 0.98726. Even predicting zeros with the same constant sigma (1e-15 * pred, magenta) results in 0.998...\n\nIf anyone has any thoughts on phyical meanigfulness of using such score fore spectra calculation/prediction, I'm curious to hear your opinions.\n\n",
      "votes": 6
    },
    {
      "id": 3293676,
      "postDate": "2025-09-24T12:04:33.870Z",
      "content": "<p>I’ve been trying to build some intuition around this function beyond the usual “make predictions as close as possible to the true values, and sigmas as close as possible to the absolute deviation,” but haven’t gotten very far.<br>\nWhat really bothers me is the way the scoring function works: it averages the individual scores first and only clips them to the 0…1 range afterward. This means that incorrect predictions (like the ones with clipped start and end) can drag the score down massively, sometimes to something like -13, which is equivalent to the penalty of 13 perfect predictions! That feels excessive.<br>\nI don’t really see the logic in this design. In practice, it leads to situations where one bad prediction outweighs a bunch of good ones. For example, 1 bad prediction and 13 perfect ones still score worse than 14 poor predictions (say, all around 0.1).</p>",
      "rawMarkdown": "I’ve been trying to build some intuition around this function beyond the usual “make predictions as close as possible to the true values, and sigmas as close as possible to the absolute deviation,” but haven’t gotten very far.\nWhat really bothers me is the way the scoring function works: it averages the individual scores first and only clips them to the 0…1 range afterward. This means that incorrect predictions (like the ones with clipped start and end) can drag the score down massively, sometimes to something like -13, which is equivalent to the penalty of 13 perfect predictions! That feels excessive.\nI don’t really see the logic in this design. In practice, it leads to situations where one bad prediction outweighs a bunch of good ones. For example, 1 bad prediction and 13 perfect ones still score worse than 14 poor predictions (say, all around 0.1).",
      "votes": 2
    },
    {
      "id": 3292924,
      "postDate": "2025-09-22T19:02:45.693Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3293676,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2025-09-24T12:04:33.870000",
      "content": "<p>I’ve been trying to build some intuition around this function beyond the usual “make predictions as close as possible to the true values, and sigmas as close as possible to the absolute deviation,” but haven’t gotten very far.<br>\nWhat really bothers me is the way the scoring function works: it averages the individual scores first and only clips them to the 0…1 range afterward. This means that incorrect predictions (like the ones with clipped start and end) can drag the score down massively, sometimes to something like -13, which is equivalent to the penalty of 13 perfect predictions! That feels excessive.<br>\nI don’t really see the logic in this design. In practice, it leads to situations where one bad prediction outweighs a bunch of good ones. For example, 1 bad prediction and 13 perfect ones still score worse than 14 poor predictions (say, all around 0.1).</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3292924,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-09-22T19:02:45.693000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3292839": "I'd like to raise some concerns about the scoring function used in the competition. And, hopefully, resolve them by filling the gaps in my understanding.\n\nI will give three examples that raise some questions in my head when I'm looking at them.\nFirst of all, in these examples I deal with only AIRS or synthetic AIRS-like data, so that there is no need for weighting and the _score_ function becomes simpler and more readable.\n\nA notebook that makes figures and canculates scores discussed further is available [here](https://www.kaggle.com/code/aleksandrzakuskin/gll-score-illustrations)\n\n## Example 1. Ideal spectrum vs. straight line prediction.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2Fc3e4c4938c85f312218e421fc474a7b7%2FExample%201.PNG?generation=1758539960614763&alt=media)\n\nThe blue line is a \"true\" spectrum made up as a sum of two random Gaussian profiles. Vertical bars at the edges are uncertainties (0.00055 in all cases in this Example). Green prediction (mean of \"true\") is obviously a relatively good one and gives 0.99904. The orange and the red are \"true\" and \"mean\" shifted by the same value (0.005). **Here is the first question: why is shifted \"mean\" is better that shifted \"true\"?**\n\nYes, the difference is very small: 0.98942 vs 0.98939, but still. If we are dealing with **spectra** in this competition, aren't their shapes important?\nSome more observations without particular questions: (1) such behavior is observed only when the prediction lies between true and naive; (2) the difference in score between shifted \"mean\" and shifted \"true\" grows as the predicted uncertainties decrease.\n\n## Example 2. Good and bad predictions.\n\nHere I'd like to compare a couple of true Ariel spectra and our predictions for them.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2Fc7c67353b48677ab4ec1c9d42abb595e%2FExample%202.PNG?generation=1758558322499414&alt=media)\n\nBlue areas are our predictions with uncertainties. Uncertainties are constant across all planets and all spectral points. On the right-hand side not even all \"true\" datapoints lie wihtin the estimated uncertainty, but the score is almost perfect (0.99960). In the case of the left plot, it is true everywhere, but score is 0.32894.\nMathematically, the answer is simple: the left spectrum lies very close to the naive prediction (0.015), while the right is far from it. So, if naive prediction is good enough, the scoring function becomes very strict and penalizes any little deviation in actual prediction. But if target is far from naive, straight line prediction becomes just fine.\nIsn't some physical meaning lost in evaluation of these two spectra in the figure?\n\n## Example 3. Scoring sensitivity.\n\nNow we know, that scoring has different sinsitivity depending on distance between true and naive. But what is the sensitivity?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14048299%2F44a7a29684a772c41acd5d1e85f5d1df%2FExample%203.PNG?generation=1758555079070701&alt=media)\n\nHere is one arbitrary prediction (blue) for one arbirary planet (orange) that gives a score of 0.99917. Both are far from naive, so we do not need to worry about shape. But what if we intentionally make prediction worse?\n\nWe can add a small constant (0.004, red) and get 0.99796. Let's add much greater constant (0.013, cyan), so that our prediction is further from \"true\" than \"naive\". And we still have 0.98726. Even predicting zeros with the same constant sigma (1e-15 * pred, magenta) results in 0.998...\n\nIf anyone has any thoughts on phyical meanigfulness of using such score fore spectra calculation/prediction, I'm curious to hear your opinions.\n\n",
    "3293676": "I’ve been trying to build some intuition around this function beyond the usual “make predictions as close as possible to the true values, and sigmas as close as possible to the absolute deviation,” but haven’t gotten very far.\nWhat really bothers me is the way the scoring function works: it averages the individual scores first and only clips them to the 0…1 range afterward. This means that incorrect predictions (like the ones with clipped start and end) can drag the score down massively, sometimes to something like -13, which is equivalent to the penalty of 13 perfect predictions! That feels excessive.\nI don’t really see the logic in this design. In practice, it leads to situations where one bad prediction outweighs a bunch of good ones. For example, 1 bad prediction and 13 perfect ones still score worse than 14 poor predictions (say, all around 0.1).",
    "3292924": ""
  }
}