{
  "id": 186531,
  "title": "[0929 Updated] More about Competition Metric Implementation",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/186531",
  "author_name": "khyeh",
  "post_date": "2020-09-24T18:49:35.343000",
  "votes": 24,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to the <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924\" target=\"_blank\">post</a> from <a href=\"https://www.kaggle.com/yuval\" target=\"_blank\">@yuval</a>, it makes me less confused about the competition metric. However, when I tried to implement myself, I found someplace still inconsistent with the description. I decided to share my implementation, and hope it could help Kagglers spend less time confirming the competition metric. Please correct me if I'm not doing it correctly.</p>\n<p>This is my implementation and testing:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/khyeh0719/rsna-competition-metric\" target=\"_blank\">https://www.kaggle.com/khyeh0719/rsna-competition-metric</a></p>\n</blockquote>\n<p><strong>====== Steps in the kernel ======</strong></p>\n<ol>\n<li>Make dummy predictions by taking an average for each label</li>\n<li>Split the predictions into chunks</li>\n<li>Calculate the loss and weight for each chunk</li>\n<li>Divide the total loss by the total weight</li>\n</ol>\n<p><strong>====== Update on 0929 ========</strong><br>\nI found my previous testing by taking an average of labels is not a good method to judge the correctness of the implementation of competition metrics. I checked with my cv, <strong>taking the sum</strong> in exam-level matches my cv 0.38x and lb: 0.362, and also just found the [reply](<a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924#1030203\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924#1030203</a><br>\n) from <a href=\"https://www.kaggle.com/yuval\" target=\"_blank\">@yuval</a>. </p>\n<p>I guess the description from the description page means taking the weighted mean; given the sum of weights itself is equal to 1 already.</p>\n<blockquote>\n  <p>Kaggle uses a binary log loss equation for each label and then <strong>takes the mean</strong> of the log loss over all labels.</p>\n</blockquote>\n<p>Just want to update here as well as the competition metric kernel to reduce confusion of Kagglers who are struggling with the metric like me :)</p>\n<p><strong>====== Update on 0926 ========</strong><br>\nThanks to <a href=\"https://www.kaggle.com/pgeiger\" target=\"_blank\">@pgeiger</a> <a href=\"https://www.kaggle.com/sudhiriitb\" target=\"_blank\">@sudhiriitb</a> for pointing out that help me resolve my confusion about the competition metric. I should do <code>train.groupby('StudyInstanceUID', sort=False)</code> in my code implementation instead. After fixing this, <strong>taking the mean</strong> of the log loss over all labels in <strong>exam level</strong> makes the number in the kernel (0.4555) close to the mean baseline (0.434) on the leaderboard now.</p>\n<p><strong>====== Before 0926 Post =======</strong><br>\nIn step 2., according to the description of <strong>exam level</strong> loss from the evaluation page</p>\n<blockquote>\n  <p>Kaggle uses a binary log loss equation for each label and then <strong>takes the mean</strong> of the log loss over all labels.</p>\n</blockquote>\n<p>What I did instead is to <strong>take the sum</strong> of the log loss over all labels<br>\n<code>exam_loss = torch.sum(exam_loss*label_w, 1)[0]</code></p>\n<p>The loss is 0.4827 in my Pytorch implementation, which is close to the 0.434 mean baseline. Also, the sum of the weights for exam labels is 1., should be a normalized weighting already, taking average would make the exam level loss 9 times smaller.</p>\n<p>Welcome for suggestions to the correctness of the competition metric implementation.=</p>",
  "messages": [
    {
      "id": 1025773,
      "postDate": "2020-09-24T18:49:35.343Z",
      "content": "<p>Thanks to the <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924\" target=\"_blank\">post</a> from <a href=\"https://www.kaggle.com/yuval\" target=\"_blank\">@yuval</a>, it makes me less confused about the competition metric. However, when I tried to implement myself, I found someplace still inconsistent with the description. I decided to share my implementation, and hope it could help Kagglers spend less time confirming the competition metric. Please correct me if I'm not doing it correctly.</p>\n<p>This is my implementation and testing:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/khyeh0719/rsna-competition-metric\" target=\"_blank\">https://www.kaggle.com/khyeh0719/rsna-competition-metric</a></p>\n</blockquote>\n<p><strong>====== Steps in the kernel ======</strong></p>\n<ol>\n<li>Make dummy predictions by taking an average for each label</li>\n<li>Split the predictions into chunks</li>\n<li>Calculate the loss and weight for each chunk</li>\n<li>Divide the total loss by the total weight</li>\n</ol>\n<p><strong>====== Update on 0929 ========</strong><br>\nI found my previous testing by taking an average of labels is not a good method to judge the correctness of the implementation of competition metrics. I checked with my cv, <strong>taking the sum</strong> in exam-level matches my cv 0.38x and lb: 0.362, and also just found the [reply](<a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924#1030203\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924#1030203</a><br>\n) from <a href=\"https://www.kaggle.com/yuval\" target=\"_blank\">@yuval</a>. </p>\n<p>I guess the description from the description page means taking the weighted mean; given the sum of weights itself is equal to 1 already.</p>\n<blockquote>\n  <p>Kaggle uses a binary log loss equation for each label and then <strong>takes the mean</strong> of the log loss over all labels.</p>\n</blockquote>\n<p>Just want to update here as well as the competition metric kernel to reduce confusion of Kagglers who are struggling with the metric like me :)</p>\n<p><strong>====== Update on 0926 ========</strong><br>\nThanks to <a href=\"https://www.kaggle.com/pgeiger\" target=\"_blank\">@pgeiger</a> <a href=\"https://www.kaggle.com/sudhiriitb\" target=\"_blank\">@sudhiriitb</a> for pointing out that help me resolve my confusion about the competition metric. I should do <code>train.groupby('StudyInstanceUID', sort=False)</code> in my code implementation instead. After fixing this, <strong>taking the mean</strong> of the log loss over all labels in <strong>exam level</strong> makes the number in the kernel (0.4555) close to the mean baseline (0.434) on the leaderboard now.</p>\n<p><strong>====== Before 0926 Post =======</strong><br>\nIn step 2., according to the description of <strong>exam level</strong> loss from the evaluation page</p>\n<blockquote>\n  <p>Kaggle uses a binary log loss equation for each label and then <strong>takes the mean</strong> of the log loss over all labels.</p>\n</blockquote>\n<p>What I did instead is to <strong>take the sum</strong> of the log loss over all labels<br>\n<code>exam_loss = torch.sum(exam_loss*label_w, 1)[0]</code></p>\n<p>The loss is 0.4827 in my Pytorch implementation, which is close to the 0.434 mean baseline. Also, the sum of the weights for exam labels is 1., should be a normalized weighting already, taking average would make the exam level loss 9 times smaller.</p>\n<p>Welcome for suggestions to the correctness of the competition metric implementation.=</p>",
      "rawMarkdown": "Thanks to the [post](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924) from @yuval, it makes me less confused about the competition metric. However, when I tried to implement myself, I found someplace still inconsistent with the description. I decided to share my implementation, and hope it could help Kagglers spend less time confirming the competition metric. Please correct me if I'm not doing it correctly.\n\nThis is my implementation and testing:\n> https://www.kaggle.com/khyeh0719/rsna-competition-metric\n\n**====== Steps in the kernel ======**\n1. Make dummy predictions by taking an average for each label\n2. Split the predictions into chunks\n3. Calculate the loss and weight for each chunk\n4. Divide the total loss by the total weight\n\n**====== Update on 0929 ========**\nI found my previous testing by taking an average of labels is not a good method to judge the correctness of the implementation of competition metrics. I checked with my cv, **taking the sum** in exam-level matches my cv 0.38x and lb: 0.362, and also just found the [reply](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924#1030203\n) from @yuval. \n\nI guess the description from the description page means taking the weighted mean; given the sum of weights itself is equal to 1 already.\n> Kaggle uses a binary log loss equation for each label and then **takes the mean** of the log loss over all labels.\n\nJust want to update here as well as the competition metric kernel to reduce confusion of Kagglers who are struggling with the metric like me :)\n\n**====== Update on 0926 ========**\nThanks to @pgeiger @sudhiriitb for pointing out that help me resolve my confusion about the competition metric. I should do `train.groupby('StudyInstanceUID', sort=False)` in my code implementation instead. After fixing this, **taking the mean** of the log loss over all labels in **exam level** makes the number in the kernel (0.4555) close to the mean baseline (0.434) on the leaderboard now.\n\n**====== Before 0926 Post =======**\nIn step 2., according to the description of **exam level** loss from the evaluation page\n> Kaggle uses a binary log loss equation for each label and then **takes the mean** of the log loss over all labels.\n\nWhat I did instead is to **take the sum** of the log loss over all labels\n`exam_loss = torch.sum(exam_loss*label_w, 1)[0] `\n\nThe loss is 0.4827 in my Pytorch implementation, which is close to the 0.434 mean baseline. Also, the sum of the weights for exam labels is 1., should be a normalized weighting already, taking average would make the exam level loss 9 times smaller.\n\nWelcome for suggestions to the correctness of the competition metric implementation.=",
      "votes": 24
    },
    {
      "id": 1026601,
      "postDate": "2020-09-25T12:44:35.143Z",
      "content": "<p>Your implementation looks ok to me. I have a similar implementation for the CV and the correlation between CV and LB is fine (taking into account the dependence of the score on the amount of studies with PE and the number of PE images in this study).</p>\n<p>(FOR CV=0.216-0.222 I get LB = 0.187-0.183)</p>",
      "rawMarkdown": "Your implementation looks ok to me. I have a similar implementation for the CV and the correlation between CV and LB is fine (taking into account the dependence of the score on the amount of studies with PE and the number of PE images in this study).\n\n(FOR CV=0.216-0.222 I get LB = 0.187-0.183)",
      "votes": 2,
      "replies": [
        {
          "id": 1026610,
          "postDate": "2020-09-25T12:53:42.443Z",
          "content": "<p>Thanks for confirming</p>",
          "rawMarkdown": "Thanks for confirming"
        },
        {
          "id": 1026747,
          "postDate": "2020-09-25T14:44:45.547Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> did you use a similar loss function as above for training your model or just checked your model performance with this?</p>",
          "rawMarkdown": "@yuval6967 did you use a similar loss function as above for training your model or just checked your model performance with this?"
        },
        {
          "id": 1026819,
          "postDate": "2020-09-25T15:16:27.110Z",
          "content": "<p>I always try to train my models with a loss function which is as close as possible to the competition's metric.</p>",
          "rawMarkdown": "I always try to train my models with a loss function which is as close as possible to the competition's metric.",
          "votes": 5
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1026601,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2020-09-25T12:44:35.143000",
      "content": "<p>Your implementation looks ok to me. I have a similar implementation for the CV and the correlation between CV and LB is fine (taking into account the dependence of the score on the amount of studies with PE and the number of PE images in this study).</p>\n<p>(FOR CV=0.216-0.222 I get LB = 0.187-0.183)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1026610,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-09-25T12:53:42.443000",
          "content": "<p>Thanks for confirming</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1026747,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-09-25T14:44:45.547000",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> did you use a similar loss function as above for training your model or just checked your model performance with this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1026819,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-25T15:16:27.110000",
          "content": "<p>I always try to train my models with a loss function which is as close as possible to the competition's metric.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1025773": "Thanks to the [post](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924) from @yuval, it makes me less confused about the competition metric. However, when I tried to implement myself, I found someplace still inconsistent with the description. I decided to share my implementation, and hope it could help Kagglers spend less time confirming the competition metric. Please correct me if I'm not doing it correctly.\n\nThis is my implementation and testing:\n> https://www.kaggle.com/khyeh0719/rsna-competition-metric\n\n**====== Steps in the kernel ======**\n1. Make dummy predictions by taking an average for each label\n2. Split the predictions into chunks\n3. Calculate the loss and weight for each chunk\n4. Divide the total loss by the total weight\n\n**====== Update on 0929 ========**\nI found my previous testing by taking an average of labels is not a good method to judge the correctness of the implementation of competition metrics. I checked with my cv, **taking the sum** in exam-level matches my cv 0.38x and lb: 0.362, and also just found the [reply](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924#1030203\n) from @yuval. \n\nI guess the description from the description page means taking the weighted mean; given the sum of weights itself is equal to 1 already.\n> Kaggle uses a binary log loss equation for each label and then **takes the mean** of the log loss over all labels.\n\nJust want to update here as well as the competition metric kernel to reduce confusion of Kagglers who are struggling with the metric like me :)\n\n**====== Update on 0926 ========**\nThanks to @pgeiger @sudhiriitb for pointing out that help me resolve my confusion about the competition metric. I should do `train.groupby('StudyInstanceUID', sort=False)` in my code implementation instead. After fixing this, **taking the mean** of the log loss over all labels in **exam level** makes the number in the kernel (0.4555) close to the mean baseline (0.434) on the leaderboard now.\n\n**====== Before 0926 Post =======**\nIn step 2., according to the description of **exam level** loss from the evaluation page\n> Kaggle uses a binary log loss equation for each label and then **takes the mean** of the log loss over all labels.\n\nWhat I did instead is to **take the sum** of the log loss over all labels\n`exam_loss = torch.sum(exam_loss*label_w, 1)[0] `\n\nThe loss is 0.4827 in my Pytorch implementation, which is close to the 0.434 mean baseline. Also, the sum of the weights for exam labels is 1., should be a normalized weighting already, taking average would make the exam level loss 9 times smaller.\n\nWelcome for suggestions to the correctness of the competition metric implementation.=",
    "1026601": "Your implementation looks ok to me. I have a similar implementation for the CV and the correlation between CV and LB is fine (taking into account the dependence of the score on the amount of studies with PE and the number of PE images in this study).\n\n(FOR CV=0.216-0.222 I get LB = 0.187-0.183)"
  }
}