{
  "id": 416630,
  "title": "Dual Thresholds are Slightly Better than Single Threshold",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/416630",
  "author_name": "LUPIN11",
  "post_date": "2023-06-12T12:08:32.313000",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, Kagglers.<br>\nAs we have seen, finding the optimal threshold is a crucial step in improving the score. But I feel that using a fixed threshold globally is somewhat simplistic. So I tried to find a more flexible approach. I attempted to implement thresholds in a trainable approach, but it failed to converge. Later, inspired by the Canny algorithm, I tried the dual thresholds. <br>\nReplacing the single threshold with dual thresholds improved my model's score from 0.55 to 0.552. This model is a U-Net without pre-trained parameters from ImageNet. The single threshold is 0.26 and the dual thresholds are 0.5(high threshold) and 0.25(low threshold).<br>\n Here is an example code and the resulting image using dual thresholds: <a href=\"https://www.kaggle.com/code/lupin11/doubleshreshold/notebook\" target=\"_blank\">dual thresholds</a><br>\nAdditionally, due to limited quota, I did not extensively search for the optimal dual thresholds for my model.<br>\nHope my findings could be helpful to you and further improvements can be shared.😊</p>",
  "messages": [
    {
      "id": 2297186,
      "postDate": "2023-06-12T12:08:32.313Z",
      "content": "<p>Hi, Kagglers.<br>\nAs we have seen, finding the optimal threshold is a crucial step in improving the score. But I feel that using a fixed threshold globally is somewhat simplistic. So I tried to find a more flexible approach. I attempted to implement thresholds in a trainable approach, but it failed to converge. Later, inspired by the Canny algorithm, I tried the dual thresholds. <br>\nReplacing the single threshold with dual thresholds improved my model's score from 0.55 to 0.552. This model is a U-Net without pre-trained parameters from ImageNet. The single threshold is 0.26 and the dual thresholds are 0.5(high threshold) and 0.25(low threshold).<br>\n Here is an example code and the resulting image using dual thresholds: <a href=\"https://www.kaggle.com/code/lupin11/doubleshreshold/notebook\" target=\"_blank\">dual thresholds</a><br>\nAdditionally, due to limited quota, I did not extensively search for the optimal dual thresholds for my model.<br>\nHope my findings could be helpful to you and further improvements can be shared.😊</p>",
      "rawMarkdown": "Hi, Kagglers.\nAs we have seen, finding the optimal threshold is a crucial step in improving the score. But I feel that using a fixed threshold globally is somewhat simplistic. So I tried to find a more flexible approach. I attempted to implement thresholds in a trainable approach, but it failed to converge. Later, inspired by the Canny algorithm, I tried the dual thresholds. \nReplacing the single threshold with dual thresholds improved my model's score from 0.55 to 0.552. This model is a U-Net without pre-trained parameters from ImageNet. The single threshold is 0.26 and the dual thresholds are 0.5(high threshold) and 0.25(low threshold).\n Here is an example code and the resulting image using dual thresholds: [dual thresholds](https://www.kaggle.com/code/lupin11/doubleshreshold/notebook)\nAdditionally, due to limited quota, I did not extensively search for the optimal dual thresholds for my model.\nHope my findings could be helpful to you and further improvements can be shared.😊",
      "votes": 7
    },
    {
      "id": 2297227,
      "postDate": "2023-06-12T12:34:19.353Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/lupin11\" target=\"_blank\">@lupin11</a>, that sounds really interesting, <br>\nyou mentionned that it improved your LB score from 0.55 to 0.552, tbh I don't really trust such small improvements on this small public data. What about your CV ? </p>",
      "rawMarkdown": "Hey @lupin11, that sounds really interesting, \nyou mentionned that it improved your LB score from 0.55 to 0.552, tbh I don't really trust such small improvements on this small public data. What about your CV ? ",
      "votes": 2,
      "replies": [
        {
          "id": 2297288,
          "postDate": "2023-06-12T13:39:13.953Z",
          "content": "<p>The thresholds (0.5, 0.25) are determined by searching on my local validation set(only 2000 samples). It improved my validation score from 0.6124 to 0.6134. This improvement is indeed much smaller than I expected. It seems to suggest that the inflexible threshold is not the bottleneck.</p>",
          "rawMarkdown": "The thresholds (0.5, 0.25) are determined by searching on my local validation set(only 2000 samples). It improved my validation score from 0.6124 to 0.6134. This improvement is indeed much smaller than I expected. It seems to suggest that the inflexible threshold is not the bottleneck.",
          "votes": 1,
          "replies": [
            {
              "id": 2297389,
              "postDate": "2023-06-12T14:42:15.503Z",
              "content": "<blockquote>\n  <p>my local validation set(only 2000 samples)</p>\n</blockquote>\n<p><strong>I assume with your statement that you use the given validation set for CV. You should look at this:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Febc9b77f9b8ce5cb5483b8089683160d%2FScreenshot_108.jpg?generation=1686580779914716&amp;alt=media\" alt=\"\"><br>\n<code>contrail</code> is a boolean that indicates the precense of at least one positive pixel in the image.<br>\nYou can see that the train and valid set are not distributed equally in regards of their presence of contrails. A stratified split would be much better in my opinion.</p>",
              "rawMarkdown": ">my local validation set(only 2000 samples)\n\n**I assume with your statement that you use the given validation set for CV. You should look at this:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Febc9b77f9b8ce5cb5483b8089683160d%2FScreenshot_108.jpg?generation=1686580779914716&alt=media)\n`contrail` is a boolean that indicates the precense of at least one positive pixel in the image.\nYou can see that the train and valid set are not distributed equally in regards of their presence of contrails. A stratified split would be much better in my opinion.",
              "votes": 1
            },
            {
              "id": 2297415,
              "postDate": "2023-06-12T15:05:52.790Z",
              "content": "<p>Thank you for your valuable suggestion! I haven't used the data from the validation folder yet. Instead, I extracted 10% of the training set as my local validation set. The statistical data you provided is very helpful for my further experiments!</p>",
              "rawMarkdown": "Thank you for your valuable suggestion! I haven't used the data from the validation folder yet. Instead, I extracted 10% of the training set as my local validation set. The statistical data you provided is very helpful for my further experiments!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2297227,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-06-12T12:34:19.353000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/lupin11\" target=\"_blank\">@lupin11</a>, that sounds really interesting, <br>\nyou mentionned that it improved your LB score from 0.55 to 0.552, tbh I don't really trust such small improvements on this small public data. What about your CV ? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2297288,
          "author_name": "LUPIN11",
          "author_url": "",
          "post_date": "2023-06-12T13:39:13.953000",
          "content": "<p>The thresholds (0.5, 0.25) are determined by searching on my local validation set(only 2000 samples). It improved my validation score from 0.6124 to 0.6134. This improvement is indeed much smaller than I expected. It seems to suggest that the inflexible threshold is not the bottleneck.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2297389,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-12T14:42:15.503000",
              "content": "<blockquote>\n  <p>my local validation set(only 2000 samples)</p>\n</blockquote>\n<p><strong>I assume with your statement that you use the given validation set for CV. You should look at this:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Febc9b77f9b8ce5cb5483b8089683160d%2FScreenshot_108.jpg?generation=1686580779914716&amp;alt=media\" alt=\"\"><br>\n<code>contrail</code> is a boolean that indicates the precense of at least one positive pixel in the image.<br>\nYou can see that the train and valid set are not distributed equally in regards of their presence of contrails. A stratified split would be much better in my opinion.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2297415,
              "author_name": "LUPIN11",
              "author_url": "",
              "post_date": "2023-06-12T15:05:52.790000",
              "content": "<p>Thank you for your valuable suggestion! I haven't used the data from the validation folder yet. Instead, I extracted 10% of the training set as my local validation set. The statistical data you provided is very helpful for my further experiments!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2297186": "Hi, Kagglers.\nAs we have seen, finding the optimal threshold is a crucial step in improving the score. But I feel that using a fixed threshold globally is somewhat simplistic. So I tried to find a more flexible approach. I attempted to implement thresholds in a trainable approach, but it failed to converge. Later, inspired by the Canny algorithm, I tried the dual thresholds. \nReplacing the single threshold with dual thresholds improved my model's score from 0.55 to 0.552. This model is a U-Net without pre-trained parameters from ImageNet. The single threshold is 0.26 and the dual thresholds are 0.5(high threshold) and 0.25(low threshold).\n Here is an example code and the resulting image using dual thresholds: [dual thresholds](https://www.kaggle.com/code/lupin11/doubleshreshold/notebook)\nAdditionally, due to limited quota, I did not extensively search for the optimal dual thresholds for my model.\nHope my findings could be helpful to you and further improvements can be shared.😊",
    "2297227": "Hey @lupin11, that sounds really interesting, \nyou mentionned that it improved your LB score from 0.55 to 0.552, tbh I don't really trust such small improvements on this small public data. What about your CV ? "
  }
}