{
  "id": 183924,
  "title": "Clarifying the competition's metric",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/183924",
  "author_name": "yuval reina",
  "post_date": "2020-09-18T15:05:38.763000",
  "votes": 33,
  "comment_count": 23,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <br>\nI have some questions about the competition's evaluation matric:</p>\n<ol>\n<li>in the evaluation tab you state: <code>The total loss is the average of all image- and exam-level loss, divided by the sum of the weights.</code><br>\nFor the images loss do you mean  the summed weight should be  <code>w=0.07361963</code> or <code>q_i*w=m_i/n_i*w</code> </li>\n<li>What about images from non PE exams? <code>m_i=0</code> which means the [q_i*w]=0 for all the images in the exam, which means they don't contribute to the loss?</li>\n<li>Is it correct to say that in exams where a small percentage of the images is positive the weight of all images is smaller? (It is a bit counter intuitive for me).</li>\n</ol>",
  "messages": [
    {
      "id": 1015957,
      "postDate": "2020-09-18T15:05:38.763Z",
      "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <br>\nI have some questions about the competition's evaluation matric:</p>\n<ol>\n<li>in the evaluation tab you state: <code>The total loss is the average of all image- and exam-level loss, divided by the sum of the weights.</code><br>\nFor the images loss do you mean  the summed weight should be  <code>w=0.07361963</code> or <code>q_i*w=m_i/n_i*w</code> </li>\n<li>What about images from non PE exams? <code>m_i=0</code> which means the [q_i*w]=0 for all the images in the exam, which means they don't contribute to the loss?</li>\n<li>Is it correct to say that in exams where a small percentage of the images is positive the weight of all images is smaller? (It is a bit counter intuitive for me).</li>\n</ol>",
      "rawMarkdown": "@anthracene @juliaelliott \nI have some questions about the competition's evaluation matric:\n1. in the evaluation tab you state: `The total loss is the average of all image- and exam-level loss, divided by the sum of the weights.`\nFor the images loss do you mean  the summed weight should be  `w=0.07361963` or `q_i*w=m_i/n_i*w` \n2. What about images from non PE exams? `m_i=0` which means the [q_i*w]=0 for all the images in the exam, which means they don't contribute to the loss?\n3. Is it correct to say that in exams where a small percentage of the images is positive the weight of all images is smaller? (It is a bit counter intuitive for me).",
      "votes": 32
    },
    {
      "id": 1015990,
      "postDate": "2020-09-18T15:29:49.580Z",
      "content": "<p>Hi! Thanks for the questions.</p>\n<ol>\n<li>The image loss weight would be <code>q_i*w</code> for each exam <code>i</code>, correct, so the sum of the image-level weights would be the sum of <code>q_i*w</code> across all exams.</li>\n<li>Correct, images from non-PE exams do not contribute to loss.</li>\n<li>That's correct. The clinical intuition would be something the host should probably weigh in on (I don't want to misstate it), but that is the how the math is set up.</li>\n</ol>",
      "rawMarkdown": "Hi! Thanks for the questions.\n\n1. The image loss weight would be `q_i*w` for each exam `i`, correct, so the sum of the image-level weights would be the sum of `q_i*w` across all exams.\n2. Correct, images from non-PE exams do not contribute to loss.\n3. That's correct. The clinical intuition would be something the host should probably weigh in on (I don't want to misstate it), but that is the how the math is set up.",
      "votes": 4,
      "replies": [
        {
          "id": 1016021,
          "postDate": "2020-09-18T15:49:49.960Z",
          "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Thanks</p>",
          "rawMarkdown": "@philculliton Thanks"
        },
        {
          "id": 1018452,
          "postDate": "2020-09-19T17:29:36.227Z",
          "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> something is still strange. you write:<br>\n<code>The total loss is the average of all image- and exam-level loss, divided by the sum of the weights</code><br>\nWhich means for the images part:</p>\n<p><code>mean(q_i * w * bce(Y_ik,P_ik)) / sum(q_i * w)</code></p>\n<p>But in that case the loss takes the mean of a lot of zeros (which come from the non PE exams) I would have thought the question should be:</p>\n<p><code>sum[q_i * w * bce(Y_ik,P_ik)] / sum(q_i * w * n_i)</code><br>\n where <code>n_i</code> is the number of images in the exam</p>\n<p>This is the weighted average.</p>",
          "rawMarkdown": "@philculliton something is still strange. you write:\n`The total loss is the average of all image- and exam-level loss, divided by the sum of the weights`\nWhich means for the images part:\n\n`mean(q_i * w * bce(Y_ik,P_ik)) / sum(q_i * w)`\n\nBut in that case the loss takes the mean of a lot of zeros (which come from the non PE exams) I would have thought the question should be:\n\n`sum[q_i * w * bce(Y_ik,P_ik)] / sum(q_i * w * n_i) `\n where `n_i` is the number of images in the exam\n\nThis is the weighted average.\n\n\n\n"
        },
        {
          "id": 1018666,
          "postDate": "2020-09-19T20:19:20.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuval\" target=\"_blank\">@yuval</a> Now I am confused, because it says <code>The total loss is the average of all image- and exam-level loss</code>. Which means for the images part, I interpret it to mean that you should not take the average loss across all the images within an exam. I.e. I interpret it to mean:</p>\n<p><code>Loss = mean(concat(q_i * w * bce(Y_ik,P_ik), avg_weighted_exam_loss_i)) / (sum(q_i * w) + #exams)</code></p>\n<p>This is because The weight for all exam labels sum to 1, so we dividing by #exams in addition to the weight from each image</p>\n<p>It would be helpful if <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> can release some code for the equation, even if it is in C# I am sure it would be helpful to understand the equation.</p>",
          "rawMarkdown": "@yuval Now I am confused, because it says `The total loss is the average of all image- and exam-level loss`. Which means for the images part, I interpret it to mean that you should not take the average loss across all the images within an exam. I.e. I interpret it to mean:\n\n`Loss = mean(concat(q_i * w * bce(Y_ik,P_ik), avg_weighted_exam_loss_i)) / (sum(q_i * w) + #exams)`\n\nThis is because The weight for all exam labels sum to 1, so we dividing by #exams in addition to the weight from each image\n\nIt would be helpful if @philculliton can release some code for the equation, even if it is in C# I am sure it would be helpful to understand the equation.",
          "votes": 2
        },
        {
          "id": 1018707,
          "postDate": "2020-09-19T21:51:35.357Z",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> where you successful in reproducing the score the Mean Baseline submission got (i.e. to get the same CV as the LB is got)?</p>",
          "rawMarkdown": "@returnofsputnik where you successful in reproducing the score the Mean Baseline submission got (i.e. to get the same CV as the LB is got)?"
        },
        {
          "id": 1018709,
          "postDate": "2020-09-19T22:02:42.320Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuval\" target=\"_blank\">@yuval</a>, No I did not try that, but that would be a good way to verify if we have the correct equation. I will try it</p>",
          "rawMarkdown": "@yuval, No I did not try that, but that would be a good way to verify if we have the correct equation. I will try it",
          "votes": 1
        },
        {
          "id": 1019082,
          "postDate": "2020-09-20T07:20:36.950Z",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> Looking at your interpretation again, I think it can't be the right one:<br>\nmean(….) must be smaller then 0.1 (bce(0.5,1) &lt; 0.7 and the wights are small), and you divide it by a number greater then 600 =&gt; the loss will be really small.<br>\nIf you replace the <code>mean</code> by  <code>sum</code> it makes sense, but it isn't what's written in the Evaluation section. </p>\n<p>When I checked your interpretation with mean instead of sum against the 'Mean Baseline' I got a close CV - 0.5285 compared to 0.55 LB.</p>",
          "rawMarkdown": "@returnofsputnik Looking at your interpretation again, I think it can't be the right one:\nmean(....) must be smaller then 0.1 (bce(0.5,1) < 0.7 and the wights are small), and you divide it by a number greater then 600 => the loss will be really small.\nIf you replace the `mean` by  `sum` it makes sense, but it isn't what's written in the Evaluation section. \n\nWhen I checked your interpretation with mean instead of sum against the 'Mean Baseline' I got a close CV - 0.5285 compared to 0.55 LB.",
          "votes": 1
        },
        {
          "id": 1019416,
          "postDate": "2020-09-20T12:15:26.407Z",
          "content": "<blockquote>\n  <p>It would be helpful if <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> can release some code for the equation, even if it is in C# I am sure it would be helpful to understand the equation.</p>\n</blockquote>\n<p>That would be the best way to clarify this metric. We have limited time for this competition with such huge dataset :(</p>",
          "rawMarkdown": "> It would be helpful if @philculliton can release some code for the equation, even if it is in C# I am sure it would be helpful to understand the equation.\n\nThat would be the best way to clarify this metric. We have limited time for this competition with such huge dataset :(",
          "votes": 3
        },
        {
          "id": 1020805,
          "postDate": "2020-09-21T12:49:57.480Z",
          "content": "<p>Hi all - sorry for the confusion. There are multiple implementations of weighted loss on our end and they all have their idiosyncrasies. The version used for this competition uses the <em>average</em> of all weights, not the sum. I have updated the Evaluation page. My apologies again.</p>\n<p>I will run the posting of the C# code past the rest of the team and see what they think. I'm not sure how helpful it would be (the code would need to be simplified to match the math being done here, as we are using only one of the code paths), but I'm happy to post it if it'll help clarify things.</p>",
          "rawMarkdown": "Hi all - sorry for the confusion. There are multiple implementations of weighted loss on our end and they all have their idiosyncrasies. The version used for this competition uses the _average_ of all weights, not the sum. I have updated the Evaluation page. My apologies again.\n\nI will run the posting of the C# code past the rest of the team and see what they think. I'm not sure how helpful it would be (the code would need to be simplified to match the math being done here, as we are using only one of the code paths), but I'm happy to post it if it'll help clarify things.",
          "votes": 3
        },
        {
          "id": 1021059,
          "postDate": "2020-09-21T15:35:16.973Z",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> now it makes sense.<br>\nactually taking the average of the losses and dividing it by the average of the weights is the same as taking the sum of the losses and dividing it by the sum of the weights, as the number of losses and the number of weights is the same :) </p>",
          "rawMarkdown": "thanks @philculliton now it makes sense.\nactually taking the average of the losses and dividing it by the average of the weights is the same as taking the sum of the losses and dividing it by the sum of the weights, as the number of losses and the number of weights is the same :) "
        },
        {
          "id": 1030061,
          "postDate": "2020-09-28T11:40:50.900Z",
          "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nI'm still confused, so let me ask you a question.</p>\n<pre><code>total loss is the average of all image- and exam-level loss, divided by the average of all row (both image- and exam-level) weights. \n</code></pre>\n<p>What is \"the average of all image- and exam-level losses\" here.<br>\nFor train data, an average of (1790594 + 65511 = 1856105 losses)? Or is it an average of (1790594 + 65511/9 [take the average of 9 labels] = 1797873 losses)?</p>",
          "rawMarkdown": "@philculliton \nI'm still confused, so let me ask you a question.\n```\ntotal loss is the average of all image- and exam-level loss, divided by the average of all row (both image- and exam-level) weights. \n```\nWhat is \"the average of all image- and exam-level losses\" here.\nFor train data, an average of (1790594 + 65511 = 1856105 losses)? Or is it an average of (1790594 + 65511/9 [take the average of 9 labels] = 1797873 losses)?"
        },
        {
          "id": 1030117,
          "postDate": "2020-09-28T12:37:40.093Z",
          "content": "<p><a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> it's the first  (1790594 + 65511 = 1856105 losses)</p>",
          "rawMarkdown": "@yujiariyasu it's the first  (1790594 + 65511 = 1856105 losses)"
        },
        {
          "id": 1030183,
          "postDate": "2020-09-28T13:14:03.217Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> <br>\nThanks!</p>\n<pre><code>Kaggle uses a binary log loss equation for each label and then takes the mean of the log loss over all labels.\n</code></pre>\n<p>Do you think this \"mean\" is meant to take the average at the end?<br>\nI thought I could take it as saying to take the average of the 9 losses, since it is written in part of the exam level description. (Probably not.)<br>\nWhat do you think about this?</p>",
          "rawMarkdown": "@yuval6967 \nThanks!\n```\nKaggle uses a binary log loss equation for each label and then takes the mean of the log loss over all labels.\n```\nDo you think this \"mean\" is meant to take the average at the end?\nI thought I could take it as saying to take the average of the 9 losses, since it is written in part of the exam level description. (Probably not.)\nWhat do you think about this?"
        },
        {
          "id": 1030203,
          "postDate": "2020-09-28T13:27:01.403Z",
          "content": "<p><a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> <br>\nmean is the average of all the losses (each multiplied by it's weight).<br>\nThe mean of the weights is the average of all weights.<br>\nFor my CV I take the sum for both, as they have the same number of elements which cancels itself at the division. </p>",
          "rawMarkdown": "@yujiariyasu \nmean is the average of all the losses (each multiplied by it's weight).\nThe mean of the weights is the average of all weights.\nFor my CV I take the sum for both, as they have the same number of elements which cancels itself at the division. \n",
          "votes": 2
        },
        {
          "id": 1030231,
          "postDate": "2020-09-28T13:43:38.373Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> <br>\nYes, I was overthinking it. Even in my model, the CV and LB values are close in that calculation method. Thank you!</p>",
          "rawMarkdown": "@yuval6967 \nYes, I was overthinking it. Even in my model, the CV and LB values are close in that calculation method. Thank you!"
        },
        {
          "id": 1030560,
          "postDate": "2020-09-28T18:17:25.647Z",
          "content": "<p>FYI Yuval's response is on point here.</p>",
          "rawMarkdown": "FYI Yuval's response is on point here.",
          "votes": 1
        },
        {
          "id": 1031002,
          "postDate": "2020-09-29T06:44:34.767Z",
          "content": "<p><a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <br>\nThey say it's sum, not mean!<br>\nI hope you find it helpful.</p>",
          "rawMarkdown": "@khyeh0719 \nThey say it's sum, not mean!\nI hope you find it helpful.",
          "votes": 1
        },
        {
          "id": 1031038,
          "postDate": "2020-09-29T07:49:55.433Z",
          "content": "<p>Yes, thank you! I updated my kernel and post as well, including referencing this post. Hope it helps others to experiment with less struggle.</p>",
          "rawMarkdown": "Yes, thank you! I updated my kernel and post as well, including referencing this post. Hope it helps others to experiment with less struggle.",
          "votes": 1
        },
        {
          "id": 1053637,
          "postDate": "2020-10-19T07:28:16.380Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> <br>\nthanks for input above<br>\none doubt if you can help clarify ,competition page mentions BCElogloss , but for training which loss is useful bce with logits  or bce log loss,</p>",
          "rawMarkdown": "@yuval6967 \nthanks for input above\none doubt if you can help clarify ,competition page mentions BCElogloss , but for training which loss is useful bce with logits  or bce log loss,\n"
        }
      ]
    },
    {
      "id": 1031133,
      "postDate": "2020-09-29T09:06:14.677Z",
      "content": "<p>good job ! good</p>",
      "rawMarkdown": "good job ! good",
      "votes": -7
    },
    {
      "id": 1041922,
      "postDate": "2020-10-08T01:14:18.513Z",
      "content": "<p>Hello everyone, I am new to kaggle competition, I am a little confused on how do I start?</p>",
      "rawMarkdown": "Hello everyone, I am new to kaggle competition, I am a little confused on how do I start?"
    },
    {
      "id": 1033542,
      "postDate": "2020-10-01T04:45:35.807Z",
      "content": "<p>This is my implementation of metric function:  <a href=\"https://www.kaggle.com/kingstying/rsna-ped-check-metric\" target=\"_blank\">https://www.kaggle.com/kingstying/rsna-ped-check-metric</a> <br>\nI have checked it on train set and public leardboard, got same score.</p>",
      "rawMarkdown": "This is my implementation of metric function:  https://www.kaggle.com/kingstying/rsna-ped-check-metric \nI have checked it on train set and public leardboard, got same score.\n"
    },
    {
      "id": 1041920,
      "postDate": "2020-10-08T01:13:07.797Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1015990,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2020-09-18T15:29:49.580000",
      "content": "<p>Hi! Thanks for the questions.</p>\n<ol>\n<li>The image loss weight would be <code>q_i*w</code> for each exam <code>i</code>, correct, so the sum of the image-level weights would be the sum of <code>q_i*w</code> across all exams.</li>\n<li>Correct, images from non-PE exams do not contribute to loss.</li>\n<li>That's correct. The clinical intuition would be something the host should probably weigh in on (I don't want to misstate it), but that is the how the math is set up.</li>\n</ol>",
      "votes": 4,
      "replies": [
        {
          "id": 1016021,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-18T15:49:49.960000",
          "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1018452,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-19T17:29:36.227000",
          "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> something is still strange. you write:<br>\n<code>The total loss is the average of all image- and exam-level loss, divided by the sum of the weights</code><br>\nWhich means for the images part:</p>\n<p><code>mean(q_i * w * bce(Y_ik,P_ik)) / sum(q_i * w)</code></p>\n<p>But in that case the loss takes the mean of a lot of zeros (which come from the non PE exams) I would have thought the question should be:</p>\n<p><code>sum[q_i * w * bce(Y_ik,P_ik)] / sum(q_i * w * n_i)</code><br>\n where <code>n_i</code> is the number of images in the exam</p>\n<p>This is the weighted average.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1018666,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-09-19T20:19:20.137000",
          "content": "<p><a href=\"https://www.kaggle.com/yuval\" target=\"_blank\">@yuval</a> Now I am confused, because it says <code>The total loss is the average of all image- and exam-level loss</code>. Which means for the images part, I interpret it to mean that you should not take the average loss across all the images within an exam. I.e. I interpret it to mean:</p>\n<p><code>Loss = mean(concat(q_i * w * bce(Y_ik,P_ik), avg_weighted_exam_loss_i)) / (sum(q_i * w) + #exams)</code></p>\n<p>This is because The weight for all exam labels sum to 1, so we dividing by #exams in addition to the weight from each image</p>\n<p>It would be helpful if <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> can release some code for the equation, even if it is in C# I am sure it would be helpful to understand the equation.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1018707,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-19T21:51:35.357000",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> where you successful in reproducing the score the Mean Baseline submission got (i.e. to get the same CV as the LB is got)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1018709,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-09-19T22:02:42.320000",
          "content": "<p><a href=\"https://www.kaggle.com/yuval\" target=\"_blank\">@yuval</a>, No I did not try that, but that would be a good way to verify if we have the correct equation. I will try it</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1019082,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-20T07:20:36.950000",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> Looking at your interpretation again, I think it can't be the right one:<br>\nmean(….) must be smaller then 0.1 (bce(0.5,1) &lt; 0.7 and the wights are small), and you divide it by a number greater then 600 =&gt; the loss will be really small.<br>\nIf you replace the <code>mean</code> by  <code>sum</code> it makes sense, but it isn't what's written in the Evaluation section. </p>\n<p>When I checked your interpretation with mean instead of sum against the 'Mean Baseline' I got a close CV - 0.5285 compared to 0.55 LB.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1019416,
          "author_name": "SeuTao",
          "author_url": "",
          "post_date": "2020-09-20T12:15:26.407000",
          "content": "<blockquote>\n  <p>It would be helpful if <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> can release some code for the equation, even if it is in C# I am sure it would be helpful to understand the equation.</p>\n</blockquote>\n<p>That would be the best way to clarify this metric. We have limited time for this competition with such huge dataset :(</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1020805,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2020-09-21T12:49:57.480000",
          "content": "<p>Hi all - sorry for the confusion. There are multiple implementations of weighted loss on our end and they all have their idiosyncrasies. The version used for this competition uses the <em>average</em> of all weights, not the sum. I have updated the Evaluation page. My apologies again.</p>\n<p>I will run the posting of the C# code past the rest of the team and see what they think. I'm not sure how helpful it would be (the code would need to be simplified to match the math being done here, as we are using only one of the code paths), but I'm happy to post it if it'll help clarify things.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1021059,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-21T15:35:16.973000",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> now it makes sense.<br>\nactually taking the average of the losses and dividing it by the average of the weights is the same as taking the sum of the losses and dividing it by the sum of the weights, as the number of losses and the number of weights is the same :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1030061,
          "author_name": "YujiAriyasu",
          "author_url": "",
          "post_date": "2020-09-28T11:40:50.900000",
          "content": "<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <br>\nI'm still confused, so let me ask you a question.</p>\n<pre><code>total loss is the average of all image- and exam-level loss, divided by the average of all row (both image- and exam-level) weights. \n</code></pre>\n<p>What is \"the average of all image- and exam-level losses\" here.<br>\nFor train data, an average of (1790594 + 65511 = 1856105 losses)? Or is it an average of (1790594 + 65511/9 [take the average of 9 labels] = 1797873 losses)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1030117,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-28T12:37:40.093000",
          "content": "<p><a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> it's the first  (1790594 + 65511 = 1856105 losses)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1030183,
          "author_name": "YujiAriyasu",
          "author_url": "",
          "post_date": "2020-09-28T13:14:03.217000",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> <br>\nThanks!</p>\n<pre><code>Kaggle uses a binary log loss equation for each label and then takes the mean of the log loss over all labels.\n</code></pre>\n<p>Do you think this \"mean\" is meant to take the average at the end?<br>\nI thought I could take it as saying to take the average of the 9 losses, since it is written in part of the exam level description. (Probably not.)<br>\nWhat do you think about this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1030203,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-28T13:27:01.403000",
          "content": "<p><a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> <br>\nmean is the average of all the losses (each multiplied by it's weight).<br>\nThe mean of the weights is the average of all weights.<br>\nFor my CV I take the sum for both, as they have the same number of elements which cancels itself at the division. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1030231,
          "author_name": "YujiAriyasu",
          "author_url": "",
          "post_date": "2020-09-28T13:43:38.373000",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> <br>\nYes, I was overthinking it. Even in my model, the CV and LB values are close in that calculation method. Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1030560,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-09-28T18:17:25.647000",
          "content": "<p>FYI Yuval's response is on point here.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1031002,
          "author_name": "YujiAriyasu",
          "author_url": "",
          "post_date": "2020-09-29T06:44:34.767000",
          "content": "<p><a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> <br>\nThey say it's sum, not mean!<br>\nI hope you find it helpful.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1031038,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-09-29T07:49:55.433000",
          "content": "<p>Yes, thank you! I updated my kernel and post as well, including referencing this post. Hope it helps others to experiment with less struggle.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1053637,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-19T07:28:16.380000",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> <br>\nthanks for input above<br>\none doubt if you can help clarify ,competition page mentions BCElogloss , but for training which loss is useful bce with logits  or bce log loss,</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1031133,
      "author_name": "Naim Mhedhbi",
      "author_url": "",
      "post_date": "2020-09-29T09:06:14.677000",
      "content": "<p>good job ! good</p>",
      "votes": -7,
      "replies": []
    },
    {
      "id": 1041922,
      "author_name": "DANIEL LIN 88",
      "author_url": "",
      "post_date": "2020-10-08T01:14:18.513000",
      "content": "<p>Hello everyone, I am new to kaggle competition, I am a little confused on how do I start?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1033542,
      "author_name": "HoneJie Li",
      "author_url": "",
      "post_date": "2020-10-01T04:45:35.807000",
      "content": "<p>This is my implementation of metric function:  <a href=\"https://www.kaggle.com/kingstying/rsna-ped-check-metric\" target=\"_blank\">https://www.kaggle.com/kingstying/rsna-ped-check-metric</a> <br>\nI have checked it on train set and public leardboard, got same score.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1041920,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-08T01:13:07.797000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1015957": "@anthracene @juliaelliott \nI have some questions about the competition's evaluation matric:\n1. in the evaluation tab you state: `The total loss is the average of all image- and exam-level loss, divided by the sum of the weights.`\nFor the images loss do you mean  the summed weight should be  `w=0.07361963` or `q_i*w=m_i/n_i*w` \n2. What about images from non PE exams? `m_i=0` which means the [q_i*w]=0 for all the images in the exam, which means they don't contribute to the loss?\n3. Is it correct to say that in exams where a small percentage of the images is positive the weight of all images is smaller? (It is a bit counter intuitive for me).",
    "1015990": "Hi! Thanks for the questions.\n\n1. The image loss weight would be `q_i*w` for each exam `i`, correct, so the sum of the image-level weights would be the sum of `q_i*w` across all exams.\n2. Correct, images from non-PE exams do not contribute to loss.\n3. That's correct. The clinical intuition would be something the host should probably weigh in on (I don't want to misstate it), but that is the how the math is set up.",
    "1031133": "good job ! good",
    "1041922": "Hello everyone, I am new to kaggle competition, I am a little confused on how do I start?",
    "1033542": "This is my implementation of metric function:  https://www.kaggle.com/kingstying/rsna-ped-check-metric \nI have checked it on train set and public leardboard, got same score.\n",
    "1041920": ""
  }
}