{
  "id": 424365,
  "title": "Scoring issue!",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/424365",
  "author_name": "Amr Muhammed",
  "post_date": "2023-07-13T16:29:56.997000",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I do not know if this problem happens to me alone or other fellows suffers from it. despite having a acceptabble model outcome in traing and validation, the public score of the test set is consistently zero. Tried almost 6-7 different models and approaches. Still having ths issue. THis is a copy of the submission coding part:</p>\n<pre><code>\ntest_id_list = []\n i  test_folders:\n     = i.split()[-]\n    test_id_list.append(())\n\n\npredictions = model.predict(test)\nbinary_predictions = np.squeeze(np.where(predictions &gt;= , , ), axis =-)\n\n\n\n ():\n    dots = np.where(\n        x.T.flatten() == fg_val)[]  \n    run_lengths = []\n    prev = -\n     b  dots:\n         b &gt; prev + :\n            run_lengths.extend((b + , ))\n        run_lengths[-] += \n        prev = b\n     run_lengths\n\n\n ():\n     x: \n        s = (x).replace(, ).replace(, ).replace(, )\n    :\n        s = \n     s\n\n\n ():\n    img = np.zeros(shape[]*shape[], dtype=np.uint8)\n     mask_rle != : \n        s = mask_rle.split()\n        starts, lengths = [np.asarray(x, dtype=)  x  (s[:][::], s[:][::])]\n        starts -= \n        ends = starts + lengths\n         lo, hi  (starts, ends):\n            img[lo:hi] = \n     img.reshape(shape, order=)  \n\n\n ():\n    submission = pd.read_csv(, index_col=)\n\n     i, rec  (test_id_list):\n        mask = binary_predictions[i]\n        submission.loc[(rec), ] = list_to_string(rle_encode(mask))\n\n    submission.to_csv()\n     submission\n</code></pre>\n<p>`</p>",
  "messages": [
    {
      "id": 2343405,
      "postDate": "2023-07-13T16:29:56.997Z",
      "content": "<p>I do not know if this problem happens to me alone or other fellows suffers from it. despite having a acceptabble model outcome in traing and validation, the public score of the test set is consistently zero. Tried almost 6-7 different models and approaches. Still having ths issue. THis is a copy of the submission coding part:</p>\n<pre><code>\ntest_id_list = []\n i  test_folders:\n     = i.split()[-]\n    test_id_list.append(())\n\n\npredictions = model.predict(test)\nbinary_predictions = np.squeeze(np.where(predictions &gt;= , , ), axis =-)\n\n\n\n ():\n    dots = np.where(\n        x.T.flatten() == fg_val)[]  \n    run_lengths = []\n    prev = -\n     b  dots:\n         b &gt; prev + :\n            run_lengths.extend((b + , ))\n        run_lengths[-] += \n        prev = b\n     run_lengths\n\n\n ():\n     x: \n        s = (x).replace(, ).replace(, ).replace(, )\n    :\n        s = \n     s\n\n\n ():\n    img = np.zeros(shape[]*shape[], dtype=np.uint8)\n     mask_rle != : \n        s = mask_rle.split()\n        starts, lengths = [np.asarray(x, dtype=)  x  (s[:][::], s[:][::])]\n        starts -= \n        ends = starts + lengths\n         lo, hi  (starts, ends):\n            img[lo:hi] = \n     img.reshape(shape, order=)  \n\n\n ():\n    submission = pd.read_csv(, index_col=)\n\n     i, rec  (test_id_list):\n        mask = binary_predictions[i]\n        submission.loc[(rec), ] = list_to_string(rle_encode(mask))\n\n    submission.to_csv()\n     submission\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "I do not know if this problem happens to me alone or other fellows suffers from it. despite having a acceptabble model outcome in traing and validation, the public score of the test set is consistently zero. Tried almost 6-7 different models and approaches. Still having ths issue. THis is a copy of the submission coding part:\n\n\n```python\n\n#### extract folders ids\ntest_id_list = []\nfor i in test_folders:\n    id = i.split('/')[-1]\n    test_id_list.append(int(id))\n\n#### predict and doing binary mask\npredictions = model.predict(test)\nbinary_predictions = np.squeeze(np.where(predictions >= 0.5, 1, 0), axis =-1)\n\n\n### necessary functions for submission\ndef rle_encode(x, fg_val=1):\n    dots = np.where(\n        x.T.flatten() == fg_val)[0]  # .T sets Fortran order down-then-right\n    run_lengths = []\n    prev = -2\n    for b in dots:\n        if b > prev + 1:\n            run_lengths.extend((b + 1, 0))\n        run_lengths[-1] += 1\n        prev = b\n    return run_lengths\n\n\ndef list_to_string(x):\n    if x: # non-empty list\n        s = str(x).replace(\"[\", \"\").replace(\"]\", \"\").replace(\",\", \"\")\n    else:\n        s = '-'\n    return s\n\n\ndef rle_decode(mask_rle, shape=(256, 256)):\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    if mask_rle != '-': \n        s = mask_rle.split()\n        starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n        starts -= 1\n        ends = starts + lengths\n        for lo, hi in zip(starts, ends):\n            img[lo:hi] = 1\n    return img.reshape(shape, order='F')  # Needed to align to RLE direction\n\n\ndef create_submission(binary_predictions):\n    submission = pd.read_csv(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/sample_submission.csv\", index_col='record_id')\n\n    for i, rec in enumerate(test_id_list):\n        mask = binary_predictions[i]\n        submission.loc[int(rec), 'encoded_pixels'] = list_to_string(rle_encode(mask))\n    \n    submission.to_csv('submission.csv')\n    return submission\n````",
      "votes": 4
    },
    {
      "id": 2343505,
      "postDate": "2023-07-13T18:03:32.977Z",
      "content": "<p>Does your model output logits or probabilities? Maybe you should apply sigmoid to the outputs</p>\n<p>Generally to avoid mistakes when running in kaggle environment you could upload all you development code as private dataset and run validation / prediction in the notebook in exactly the same way you do locally (train and validation data is available in competition dataset) e. g. as <code>!python validation.py</code> and check the scores.</p>",
      "rawMarkdown": "Does your model output logits or probabilities? Maybe you should apply sigmoid to the outputs\n\nGenerally to avoid mistakes when running in kaggle environment you could upload all you development code as private dataset and run validation / prediction in the notebook in exactly the same way you do locally (train and validation data is available in competition dataset) e. g. as `!python validation.py` and check the scores.",
      "votes": 1
    },
    {
      "id": 2343498,
      "postDate": "2023-07-13T17:55:25.520Z",
      "content": "<p>I would recommend running the pipeline against the validation folder first. This way you can double check that the inference pipeline is set up correctly.</p>\n<p>Something like this.</p>\n<pre><code>sub = pd.read_csv()\nmetric = torchmetrics.Dice(average = , threshold = )\nm_plots = []\n\n\n record_id, encoded_pixels  tqdm(sub[[, ]].values):\n\n    \n    pred_mask = rle_decode(encoded_pixels).astype()\n    true_mask = np.load(os.path.join(, (record_id), ))[..., -]\n\n    pred_mask = torch.tensor(pred_mask)\n    true_mask = torch.tensor(true_mask)\n\n    \n    metric.update(pred_mask, true_mask)\n\nscore = metric.compute().item()\n(.(score))\n</code></pre>",
      "rawMarkdown": "I would recommend running the pipeline against the validation folder first. This way you can double check that the inference pipeline is set up correctly.\n\nSomething like this.\n```python\nsub = pd.read_csv(\"/kaggle/working/submission.csv\")\nmetric = torchmetrics.Dice(average = 'micro', threshold = 0.5)\nm_plots = []\n\n# Iterate predictions\nfor record_id, encoded_pixels in tqdm(sub[[\"record_id\", \"encoded_pixels\"]].values):\n\n    # Load pred + truth mask\n    pred_mask = rle_decode(encoded_pixels).astype(float)\n    true_mask = np.load(os.path.join(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/validation/\", str(record_id), 'human_pixel_masks.npy'))[..., -1]\n\n    pred_mask = torch.tensor(pred_mask)\n    true_mask = torch.tensor(true_mask)\n\n    # Update metric\n    metric.update(pred_mask, true_mask)\n\nscore = metric.compute().item()\nprint(\"Score: {:.6f}\".format(score))\n```",
      "votes": 1
    },
    {
      "id": 2343466,
      "postDate": "2023-07-13T17:21:50.143Z",
      "content": "<p><code>predictions = model.predict(test)</code><br>\nWhats is test ? it is not shown in your code.</p>\n<p>Did you forget to use normalisation by any chance ? Did you call model.eval() when loading the state dict ? <br>\nYour LB score should be close to CV at least a little, a 0.00 score indicates a mistake somewhere.</p>",
      "rawMarkdown": "`predictions = model.predict(test)`\nWhats is test ? it is not shown in your code.\n\nDid you forget to use normalisation by any chance ? Did you call model.eval() when loading the state dict ? \nYour LB score should be close to CV at least a little, a 0.00 score indicates a mistake somewhere.",
      "replies": [
        {
          "id": 2343470,
          "postDate": "2023-07-13T17:24:16.553Z",
          "content": "<p>One suggestion I have is that you load an image in your training environnement, look at the output of your model, then import the same  image on the inference notebook, and look at the output again. You'll be able to troubleshoot the issue faster.</p>",
          "rawMarkdown": "One suggestion I have is that you load an image in your training environnement, look at the output of your model, then import the same  image on the inference notebook, and look at the output again. You'll be able to troubleshoot the issue faster."
        },
        {
          "id": 2343485,
          "postDate": "2023-07-13T17:39:56.523Z",
          "content": "<p>I am using the follwing loop to load the test set</p>\n<p>The validations et showed the following values in scoring: loss: 0.3820 - binary_accuracy: 0.9964 - iou_score: 0.4492</p>\n<p>This is acheived by a UNET </p>\n<pre><code>\ntest_folders = glob.glob()\n\n i  tqdm.tqdm((test_folders), total = (test_folders)):\n\n    \n    _path = (glob.glob(i + ))\n\n    \n    img = np.zeros((, , , ), np.float32)\n\n     ii  _path:\n           ii:\n            img_band_11 = np.load(ii)\n            img_band_11 = img_band_11[:,:,]\n\n\n           ii:\n            img_band_14 = np.load(ii)\n            img_band_14 = img_band_14[:,:,]\n\n\n           ii:\n            img_band_15 = np.load(ii)\n            img_band_15 = img_band_15[:,:,]       \n\n           ii:\n            lab_ii = np.expand_dims(np.load(ii), )\n\n\n    r = normalize_range(img_band_15 - img_band_14, _TDIFF_BOUNDS)\n    g = normalize_range(img_band_14 - img_band_11, _CLOUD_TOP_TDIFF_BOUNDS)\n    b = normalize_range(img_band_14, _T11_BOUNDS)\n\n    false_color = np.clip(np.stack([r, g, b], axis=), , )\n\n    img[] = false_color\n\n    : \n        test = np.concatenate([test, img], axis = )\n    : \n        test = img\n</code></pre>",
          "rawMarkdown": "I am using the follwing loop to load the test set\n\nThe validations et showed the following values in scoring: loss: 0.3820 - binary_accuracy: 0.9964 - iou_score: 0.4492\n\nThis is acheived by a UNET \n```python\n####### load test set and submission process\ntest_folders = glob.glob('/kaggle/input/google-research-identify-contrails-reduce-global-warming/test/*')\n\nfor i in tqdm.tqdm(sorted(test_folders), total = len(test_folders)):\n    \n    ## path for each subfolder\n    _path = sorted(glob.glob(i + '/*'))\n    \n    ### generate empty 9 ch image, each will carry one band\n    img = np.zeros((1, 256, 256, 3), np.float32)\n    \n    for ii in _path:\n        if 'band_11' in ii:\n            img_band_11 = np.load(ii)\n            img_band_11 = img_band_11[:,:,4]\n            \n        \n        elif 'band_14' in ii:\n            img_band_14 = np.load(ii)\n            img_band_14 = img_band_14[:,:,4]\n            \n            \n        elif 'band_15' in ii:\n            img_band_15 = np.load(ii)\n            img_band_15 = img_band_15[:,:,4]       \n        \n        elif 'human_pixel_' in ii:\n            lab_ii = np.expand_dims(np.load(ii), 0)\n            \n        \n    r = normalize_range(img_band_15 - img_band_14, _TDIFF_BOUNDS)\n    g = normalize_range(img_band_14 - img_band_11, _CLOUD_TOP_TDIFF_BOUNDS)\n    b = normalize_range(img_band_14, _T11_BOUNDS)\n        \n    false_color = np.clip(np.stack([r, g, b], axis=2), 0, 1)\n        \n    img[0] = false_color\n       \n    try: ## concatenate resulting image to X array\n        test = np.concatenate([test, img], axis = 0)\n    except: ## If x is not in memory as var (for first image), build new X from initial image\n        test = img\n```"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2343505,
      "author_name": "Mikhail Kotyushev",
      "author_url": "",
      "post_date": "2023-07-13T18:03:32.977000",
      "content": "<p>Does your model output logits or probabilities? Maybe you should apply sigmoid to the outputs</p>\n<p>Generally to avoid mistakes when running in kaggle environment you could upload all you development code as private dataset and run validation / prediction in the notebook in exactly the same way you do locally (train and validation data is available in competition dataset) e. g. as <code>!python validation.py</code> and check the scores.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2343498,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2023-07-13T17:55:25.520000",
      "content": "<p>I would recommend running the pipeline against the validation folder first. This way you can double check that the inference pipeline is set up correctly.</p>\n<p>Something like this.</p>\n<pre><code>sub = pd.read_csv()\nmetric = torchmetrics.Dice(average = , threshold = )\nm_plots = []\n\n\n record_id, encoded_pixels  tqdm(sub[[, ]].values):\n\n    \n    pred_mask = rle_decode(encoded_pixels).astype()\n    true_mask = np.load(os.path.join(, (record_id), ))[..., -]\n\n    pred_mask = torch.tensor(pred_mask)\n    true_mask = torch.tensor(true_mask)\n\n    \n    metric.update(pred_mask, true_mask)\n\nscore = metric.compute().item()\n(.(score))\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2343466,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-07-13T17:21:50.143000",
      "content": "<p><code>predictions = model.predict(test)</code><br>\nWhats is test ? it is not shown in your code.</p>\n<p>Did you forget to use normalisation by any chance ? Did you call model.eval() when loading the state dict ? <br>\nYour LB score should be close to CV at least a little, a 0.00 score indicates a mistake somewhere.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2343470,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-07-13T17:24:16.553000",
          "content": "<p>One suggestion I have is that you load an image in your training environnement, look at the output of your model, then import the same  image on the inference notebook, and look at the output again. You'll be able to troubleshoot the issue faster.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2343485,
          "author_name": "Amr Muhammed",
          "author_url": "",
          "post_date": "2023-07-13T17:39:56.523000",
          "content": "<p>I am using the follwing loop to load the test set</p>\n<p>The validations et showed the following values in scoring: loss: 0.3820 - binary_accuracy: 0.9964 - iou_score: 0.4492</p>\n<p>This is acheived by a UNET </p>\n<pre><code>\ntest_folders = glob.glob()\n\n i  tqdm.tqdm((test_folders), total = (test_folders)):\n\n    \n    _path = (glob.glob(i + ))\n\n    \n    img = np.zeros((, , , ), np.float32)\n\n     ii  _path:\n           ii:\n            img_band_11 = np.load(ii)\n            img_band_11 = img_band_11[:,:,]\n\n\n           ii:\n            img_band_14 = np.load(ii)\n            img_band_14 = img_band_14[:,:,]\n\n\n           ii:\n            img_band_15 = np.load(ii)\n            img_band_15 = img_band_15[:,:,]       \n\n           ii:\n            lab_ii = np.expand_dims(np.load(ii), )\n\n\n    r = normalize_range(img_band_15 - img_band_14, _TDIFF_BOUNDS)\n    g = normalize_range(img_band_14 - img_band_11, _CLOUD_TOP_TDIFF_BOUNDS)\n    b = normalize_range(img_band_14, _T11_BOUNDS)\n\n    false_color = np.clip(np.stack([r, g, b], axis=), , )\n\n    img[] = false_color\n\n    : \n        test = np.concatenate([test, img], axis = )\n    : \n        test = img\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2343405": "I do not know if this problem happens to me alone or other fellows suffers from it. despite having a acceptabble model outcome in traing and validation, the public score of the test set is consistently zero. Tried almost 6-7 different models and approaches. Still having ths issue. THis is a copy of the submission coding part:\n\n\n```python\n\n#### extract folders ids\ntest_id_list = []\nfor i in test_folders:\n    id = i.split('/')[-1]\n    test_id_list.append(int(id))\n\n#### predict and doing binary mask\npredictions = model.predict(test)\nbinary_predictions = np.squeeze(np.where(predictions >= 0.5, 1, 0), axis =-1)\n\n\n### necessary functions for submission\ndef rle_encode(x, fg_val=1):\n    dots = np.where(\n        x.T.flatten() == fg_val)[0]  # .T sets Fortran order down-then-right\n    run_lengths = []\n    prev = -2\n    for b in dots:\n        if b > prev + 1:\n            run_lengths.extend((b + 1, 0))\n        run_lengths[-1] += 1\n        prev = b\n    return run_lengths\n\n\ndef list_to_string(x):\n    if x: # non-empty list\n        s = str(x).replace(\"[\", \"\").replace(\"]\", \"\").replace(\",\", \"\")\n    else:\n        s = '-'\n    return s\n\n\ndef rle_decode(mask_rle, shape=(256, 256)):\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    if mask_rle != '-': \n        s = mask_rle.split()\n        starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n        starts -= 1\n        ends = starts + lengths\n        for lo, hi in zip(starts, ends):\n            img[lo:hi] = 1\n    return img.reshape(shape, order='F')  # Needed to align to RLE direction\n\n\ndef create_submission(binary_predictions):\n    submission = pd.read_csv(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/sample_submission.csv\", index_col='record_id')\n\n    for i, rec in enumerate(test_id_list):\n        mask = binary_predictions[i]\n        submission.loc[int(rec), 'encoded_pixels'] = list_to_string(rle_encode(mask))\n    \n    submission.to_csv('submission.csv')\n    return submission\n````",
    "2343505": "Does your model output logits or probabilities? Maybe you should apply sigmoid to the outputs\n\nGenerally to avoid mistakes when running in kaggle environment you could upload all you development code as private dataset and run validation / prediction in the notebook in exactly the same way you do locally (train and validation data is available in competition dataset) e. g. as `!python validation.py` and check the scores.",
    "2343498": "I would recommend running the pipeline against the validation folder first. This way you can double check that the inference pipeline is set up correctly.\n\nSomething like this.\n```python\nsub = pd.read_csv(\"/kaggle/working/submission.csv\")\nmetric = torchmetrics.Dice(average = 'micro', threshold = 0.5)\nm_plots = []\n\n# Iterate predictions\nfor record_id, encoded_pixels in tqdm(sub[[\"record_id\", \"encoded_pixels\"]].values):\n\n    # Load pred + truth mask\n    pred_mask = rle_decode(encoded_pixels).astype(float)\n    true_mask = np.load(os.path.join(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/validation/\", str(record_id), 'human_pixel_masks.npy'))[..., -1]\n\n    pred_mask = torch.tensor(pred_mask)\n    true_mask = torch.tensor(true_mask)\n\n    # Update metric\n    metric.update(pred_mask, true_mask)\n\nscore = metric.compute().item()\nprint(\"Score: {:.6f}\".format(score))\n```",
    "2343466": "`predictions = model.predict(test)`\nWhats is test ? it is not shown in your code.\n\nDid you forget to use normalisation by any chance ? Did you call model.eval() when loading the state dict ? \nYour LB score should be close to CV at least a little, a 0.00 score indicates a mistake somewhere."
  }
}