{
  "id": 588832,
  "title": "Multiply or divide by gain? Add or subtract offset?",
  "url": "/competitions/ariel-data-challenge-2025/discussion/588832",
  "author_name": "CPMP",
  "post_date": "2025-07-08T15:06:53.728000",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>In the data page we can read :</p>\n<blockquote>\n  <p>To restore the full dynamic range you must multiply the data by the matching gain value from adc_info.csv and then add the offset value, also from adc_info.csv. </p>\n</blockquote>\n<p>Yet in some public notebook I see a division by the gain instead of a multiply. </p>\n<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> which one is right? Both yield negative numbers, whcih is surprising.</p>\n<p>if input must be divided by gain then maybe offset should be subtracted as well.</p>\n<p>Offset sign is less of an issue as it cancels out when subtracting the dark frame. But still, I find it disturbing to see negative numbers.</p>\n<p>If this has already been discussed then please point me to the relevant message and I'll update my post.</p>",
  "messages": [
    {
      "id": 3244771,
      "postDate": "2025-07-08T15:06:53.727Z",
      "content": "<p>In the data page we can read :</p>\n<blockquote>\n  <p>To restore the full dynamic range you must multiply the data by the matching gain value from adc_info.csv and then add the offset value, also from adc_info.csv. </p>\n</blockquote>\n<p>Yet in some public notebook I see a division by the gain instead of a multiply. </p>\n<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> which one is right? Both yield negative numbers, whcih is surprising.</p>\n<p>if input must be divided by gain then maybe offset should be subtracted as well.</p>\n<p>Offset sign is less of an issue as it cancels out when subtracting the dark frame. But still, I find it disturbing to see negative numbers.</p>\n<p>If this has already been discussed then please point me to the relevant message and I'll update my post.</p>",
      "rawMarkdown": "In the data page we can read :\n\n> To restore the full dynamic range you must multiply the data by the matching gain value from adc_info.csv and then add the offset value, also from adc_info.csv. \n\nYet in some public notebook I see a division by the gain instead of a multiply. \n\n@gordonyip which one is right? Both yield negative numbers, whcih is surprising.\n\nif input must be divided by gain then maybe offset should be subtracted as well.\n\nOffset sign is less of an issue as it cancels out when subtracting the dark frame. But still, I find it disturbing to see negative numbers.\n\nIf this has already been discussed then please point me to the relevant message and I'll update my post.",
      "votes": 13
    },
    {
      "id": 3245575,
      "postDate": "2025-07-09T15:00:35.933Z",
      "content": "<p>Hi all, thank you for the discussion. According to the manual of exosim2: link here:<a href=\"https://exosim2-public.readthedocs.io/en/latest/user/ndrs/analogtodigtital.html\" target=\"_blank\">https://exosim2-public.readthedocs.io/en/latest/user/ndrs/analogtodigtital.html</a></p>\n<p>it should be DIVIDE by the gain and ADD the offset to get the measured signal back. </p>\n<p>I will look into the notebook now to make sure it is implemented this way!</p>",
      "rawMarkdown": "Hi all, thank you for the discussion. According to the manual of exosim2: link here:https://exosim2-public.readthedocs.io/en/latest/user/ndrs/analogtodigtital.html\n\nit should be DIVIDE by the gain and ADD the offset to get the measured signal back. \n\nI will look into the notebook now to make sure it is implemented this way!",
      "votes": 6,
      "replies": [
        {
          "id": 3247273,
          "postDate": "2025-07-12T12:50:46.913Z",
          "content": "<p>Good, I read the same manual and was wondering.</p>\n<p>The notebook is probably fine, but the data description page is wrong.</p>",
          "rawMarkdown": "Good, I read the same manual and was wondering.\n\nThe notebook is probably fine, but the data description page is wrong."
        }
      ]
    },
    {
      "id": 3249578,
      "postDate": "2025-07-16T17:45:36.197Z",
      "content": "<p>We've updated the data description. It looks like the other point of confusion was that when the dynamic range is restored you'll be left with a float64 despite the raw data being uint16.</p>",
      "rawMarkdown": "We've updated the data description. It looks like the other point of confusion was that when the dynamic range is restored you'll be left with a float64 despite the raw data being uint16.",
      "votes": 1,
      "replies": [
        {
          "id": 3249613,
          "postDate": "2025-07-16T19:00:00.090Z",
          "content": "<p>Thank you.</p>",
          "rawMarkdown": "Thank you.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3244933,
      "postDate": "2025-07-08T17:24:11.213Z",
      "content": "<p>It's not discussed this year, but there was a confusion last year:<br>\n<a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/530247\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/530247</a></p>\n<p>It looks like occasional negative exists due to noise.</p>",
      "rawMarkdown": "It's not discussed this year, but there was a confusion last year:\nhttps://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/530247\n\nIt looks like occasional negative exists due to noise.",
      "votes": 2,
      "replies": [
        {
          "id": 3245142,
          "postDate": "2025-07-09T01:20:15.077Z",
          "content": "<p>It is not occasional negatives. It is frequent negatives.</p>\n<p>Thanks for the pointer. I see I am not the only one who wondered about the offset sign.</p>",
          "rawMarkdown": "It is not occasional negatives. It is frequent negatives.\n\nThanks for the pointer. I see I am not the only one who wondered about the offset sign.",
          "votes": 1
        },
        {
          "id": 3245800,
          "postDate": "2025-07-09T21:37:41.753Z",
          "content": "<p>Indeed ~20% of the signal is negative after the transformation! 🙇‍♂️</p>\n<p>But after Correlated Double Sampling (CDS) and time average =&gt; cds_mean shape (32, 356),<br>\nthe negatives are either outside the target range (282 spectra 39:321) or dead pixels. The gain/offset combination seems correct to me.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fe3a73bec1b2ca5ddfd9b7838f91c52a4%2Fdownload%20(2).png?generation=1752096965756370&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F9a5a62c0fa419e494fc7ae5e9b585400%2Fdownload%20(3).png?generation=1752096988412557&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Indeed ~20% of the signal is negative after the transformation! 🙇‍♂️\n\nBut after Correlated Double Sampling (CDS) and time average => cds_mean shape (32, 356),\nthe negatives are either outside the target range (282 spectra 39:321) or dead pixels. The gain/offset combination seems correct to me.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fe3a73bec1b2ca5ddfd9b7838f91c52a4%2Fdownload%20(2).png?generation=1752096965756370&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F9a5a62c0fa419e494fc7ae5e9b585400%2Fdownload%20(3).png?generation=1752096988412557&alt=media)\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 3245533,
      "postDate": "2025-07-09T14:14:03.230Z",
      "content": "<p>division by the gain leads to a range that's beyond what uint16 can hold, so the dataset description formula is correct.</p>\n<p>ADC conversion: Electron count -&gt; digital number (unit16) </p>\n<p>negative values mean the electron count is below the detector baseline count which is -1000 for both instruments.</p>\n<p>edit:</p>\n<p>From the ExoSim2 framework code, the formula used to convert physical detector values to uint16 is:</p>\n<pre><code>uint16_value = rounder((physical_value - offset) * gain_factor)\n</code></pre>\n<p>Where:</p>\n<ul>\n<li><code>physical_value</code> is the original detector measurement  </li>\n<li><code>offset</code> is the ADC offset (-1000.0)  </li>\n<li><code>gain_factor</code> is the ADC gain (0.4369)  </li>\n<li><code>rounder</code> is a rounding function (floor, ceil, or round)  </li>\n</ul>\n<p>This is visible in the <code>analogToDigital.py</code> file:  </p>\n<pre><code>data = (deepcopy(subexposures.dataset[chunk]) - offset) * gain_factor\nndrs.dataset[chunk] = rounder(data).astype(int_type)\n</code></pre>\n<p>So when the telescope compresses data:</p>\n<ol>\n<li>Subtract offset (-1000.0) from physical value  </li>\n<li>Multiply by gain (0.4369)  </li>\n<li>Round to integer  </li>\n<li>Store as uint16</li>\n</ol>\n<p>Therefore, to recover physical values, we must reverse this process:  </p>\n<pre><code>physical_value = (uint16_value / gain_factor) + offset\n</code></pre>\n<p>From the data description page:</p>\n<blockquote>\n  <p>To restore its original dynamic range you must multiply the data by the matching <code>gain</code> value from adc_info.csv and then add the <code>offset</code> value, also from adc_info.csv.</p>\n</blockquote>\n<p>Since it mentions about restoring original dynamic range, the formula is incorrect. I initially thought they got the direction wrong, which added to the confusion. </p>\n<p>Either way it was time well spent digging into the simulation framework. </p>",
      "rawMarkdown": "division by the gain leads to a range that's beyond what uint16 can hold, so the dataset description formula is correct.\n\nADC conversion: Electron count -> digital number (unit16) \n\nnegative values mean the electron count is below the detector baseline count which is -1000 for both instruments.\n\nedit:\n\nFrom the ExoSim2 framework code, the formula used to convert physical detector values to uint16 is:\n\n```python\nuint16_value = rounder((physical_value - offset) * gain_factor)\n```\n\nWhere:\n\n- `physical_value` is the original detector measurement  \n- `offset` is the ADC offset (-1000.0)  \n- `gain_factor` is the ADC gain (0.4369)  \n- `rounder` is a rounding function (floor, ceil, or round)  \n\nThis is visible in the `analogToDigital.py` file:  \n\n```python\n\ndata = (deepcopy(subexposures.dataset[chunk]) - offset) * gain_factor\nndrs.dataset[chunk] = rounder(data).astype(int_type)\n\n```\n\nSo when the telescope compresses data:\n\n1. Subtract offset (-1000.0) from physical value  \n2. Multiply by gain (0.4369)  \n3. Round to integer  \n4. Store as uint16\n\nTherefore, to recover physical values, we must reverse this process:  \n\n```python\nphysical_value = (uint16_value / gain_factor) + offset\n```\n\nFrom the data description page:\n\n> To restore its original dynamic range you must multiply the data by the matching `gain` value from adc_info.csv and then add the `offset` value, also from adc_info.csv.\n\nSince it mentions about restoring original dynamic range, the formula is incorrect. I initially thought they got the direction wrong, which added to the confusion. \n\nEither way it was time well spent digging into the simulation framework. ",
      "votes": -1
    },
    {
      "id": 3244806,
      "postDate": "2025-07-08T15:49:10.570Z",
      "content": "<p>I think there's a mistake in the quote you mentioned. You need to divide by the gain and then add the offset. At least, that's what I do.</p>",
      "rawMarkdown": "I think there's a mistake in the quote you mentioned. You need to divide by the gain and then add the offset. At least, that's what I do.",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3245575,
      "author_name": "Gordon Yip",
      "author_url": "",
      "post_date": "2025-07-09T15:00:35.933000",
      "content": "<p>Hi all, thank you for the discussion. According to the manual of exosim2: link here:<a href=\"https://exosim2-public.readthedocs.io/en/latest/user/ndrs/analogtodigtital.html\" target=\"_blank\">https://exosim2-public.readthedocs.io/en/latest/user/ndrs/analogtodigtital.html</a></p>\n<p>it should be DIVIDE by the gain and ADD the offset to get the measured signal back. </p>\n<p>I will look into the notebook now to make sure it is implemented this way!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3247273,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-07-12T12:50:46.913000",
          "content": "<p>Good, I read the same manual and was wondering.</p>\n<p>The notebook is probably fine, but the data description page is wrong.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3249578,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2025-07-16T17:45:36.197000",
      "content": "<p>We've updated the data description. It looks like the other point of confusion was that when the dynamic range is restored you'll be left with a float64 despite the raw data being uint16.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3249613,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-07-16T19:00:00.090000",
          "content": "<p>Thank you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3244933,
      "author_name": "🐢 Jun Koda",
      "author_url": "",
      "post_date": "2025-07-08T17:24:11.213000",
      "content": "<p>It's not discussed this year, but there was a confusion last year:<br>\n<a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/530247\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/530247</a></p>\n<p>It looks like occasional negative exists due to noise.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3245142,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-07-09T01:20:15.077000",
          "content": "<p>It is not occasional negatives. It is frequent negatives.</p>\n<p>Thanks for the pointer. I see I am not the only one who wondered about the offset sign.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3245800,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2025-07-09T21:37:41.753000",
          "content": "<p>Indeed ~20% of the signal is negative after the transformation! 🙇‍♂️</p>\n<p>But after Correlated Double Sampling (CDS) and time average =&gt; cds_mean shape (32, 356),<br>\nthe negatives are either outside the target range (282 spectra 39:321) or dead pixels. The gain/offset combination seems correct to me.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fe3a73bec1b2ca5ddfd9b7838f91c52a4%2Fdownload%20(2).png?generation=1752096965756370&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F9a5a62c0fa419e494fc7ae5e9b585400%2Fdownload%20(3).png?generation=1752096988412557&amp;alt=media\" alt=\"\"></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3245533,
      "author_name": "g john rao",
      "author_url": "",
      "post_date": "2025-07-09T14:14:03.230000",
      "content": "<p>division by the gain leads to a range that's beyond what uint16 can hold, so the dataset description formula is correct.</p>\n<p>ADC conversion: Electron count -&gt; digital number (unit16) </p>\n<p>negative values mean the electron count is below the detector baseline count which is -1000 for both instruments.</p>\n<p>edit:</p>\n<p>From the ExoSim2 framework code, the formula used to convert physical detector values to uint16 is:</p>\n<pre><code>uint16_value = rounder((physical_value - offset) * gain_factor)\n</code></pre>\n<p>Where:</p>\n<ul>\n<li><code>physical_value</code> is the original detector measurement  </li>\n<li><code>offset</code> is the ADC offset (-1000.0)  </li>\n<li><code>gain_factor</code> is the ADC gain (0.4369)  </li>\n<li><code>rounder</code> is a rounding function (floor, ceil, or round)  </li>\n</ul>\n<p>This is visible in the <code>analogToDigital.py</code> file:  </p>\n<pre><code>data = (deepcopy(subexposures.dataset[chunk]) - offset) * gain_factor\nndrs.dataset[chunk] = rounder(data).astype(int_type)\n</code></pre>\n<p>So when the telescope compresses data:</p>\n<ol>\n<li>Subtract offset (-1000.0) from physical value  </li>\n<li>Multiply by gain (0.4369)  </li>\n<li>Round to integer  </li>\n<li>Store as uint16</li>\n</ol>\n<p>Therefore, to recover physical values, we must reverse this process:  </p>\n<pre><code>physical_value = (uint16_value / gain_factor) + offset\n</code></pre>\n<p>From the data description page:</p>\n<blockquote>\n  <p>To restore its original dynamic range you must multiply the data by the matching <code>gain</code> value from adc_info.csv and then add the <code>offset</code> value, also from adc_info.csv.</p>\n</blockquote>\n<p>Since it mentions about restoring original dynamic range, the formula is incorrect. I initially thought they got the direction wrong, which added to the confusion. </p>\n<p>Either way it was time well spent digging into the simulation framework. </p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3244806,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-08T15:49:10.570000",
      "content": "<p>I think there's a mistake in the quote you mentioned. You need to divide by the gain and then add the offset. At least, that's what I do.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3244771": "In the data page we can read :\n\n> To restore the full dynamic range you must multiply the data by the matching gain value from adc_info.csv and then add the offset value, also from adc_info.csv. \n\nYet in some public notebook I see a division by the gain instead of a multiply. \n\n@gordonyip which one is right? Both yield negative numbers, whcih is surprising.\n\nif input must be divided by gain then maybe offset should be subtracted as well.\n\nOffset sign is less of an issue as it cancels out when subtracting the dark frame. But still, I find it disturbing to see negative numbers.\n\nIf this has already been discussed then please point me to the relevant message and I'll update my post.",
    "3245575": "Hi all, thank you for the discussion. According to the manual of exosim2: link here:https://exosim2-public.readthedocs.io/en/latest/user/ndrs/analogtodigtital.html\n\nit should be DIVIDE by the gain and ADD the offset to get the measured signal back. \n\nI will look into the notebook now to make sure it is implemented this way!",
    "3249578": "We've updated the data description. It looks like the other point of confusion was that when the dynamic range is restored you'll be left with a float64 despite the raw data being uint16.",
    "3244933": "It's not discussed this year, but there was a confusion last year:\nhttps://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/530247\n\nIt looks like occasional negative exists due to noise.",
    "3245533": "division by the gain leads to a range that's beyond what uint16 can hold, so the dataset description formula is correct.\n\nADC conversion: Electron count -> digital number (unit16) \n\nnegative values mean the electron count is below the detector baseline count which is -1000 for both instruments.\n\nedit:\n\nFrom the ExoSim2 framework code, the formula used to convert physical detector values to uint16 is:\n\n```python\nuint16_value = rounder((physical_value - offset) * gain_factor)\n```\n\nWhere:\n\n- `physical_value` is the original detector measurement  \n- `offset` is the ADC offset (-1000.0)  \n- `gain_factor` is the ADC gain (0.4369)  \n- `rounder` is a rounding function (floor, ceil, or round)  \n\nThis is visible in the `analogToDigital.py` file:  \n\n```python\n\ndata = (deepcopy(subexposures.dataset[chunk]) - offset) * gain_factor\nndrs.dataset[chunk] = rounder(data).astype(int_type)\n\n```\n\nSo when the telescope compresses data:\n\n1. Subtract offset (-1000.0) from physical value  \n2. Multiply by gain (0.4369)  \n3. Round to integer  \n4. Store as uint16\n\nTherefore, to recover physical values, we must reverse this process:  \n\n```python\nphysical_value = (uint16_value / gain_factor) + offset\n```\n\nFrom the data description page:\n\n> To restore its original dynamic range you must multiply the data by the matching `gain` value from adc_info.csv and then add the `offset` value, also from adc_info.csv.\n\nSince it mentions about restoring original dynamic range, the formula is incorrect. I initially thought they got the direction wrong, which added to the confusion. \n\nEither way it was time well spent digging into the simulation framework. ",
    "3244806": "I think there's a mistake in the quote you mentioned. You need to divide by the gain and then add the offset. At least, that's what I do."
  }
}