{
  "id": 116308,
  "title": "Is public LB score calculation broken?",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/116308",
  "author_name": "nosound",
  "post_date": "2019-11-08T09:04:55.023000",
  "votes": 18,
  "comment_count": 22,
  "views": 0,
  "content": "<p>It seems that there is a bug in public LB calculation. Below is a list of observations that support it.</p>\n\n<ol>\n<li><p>One can get score 0.000 by submitting 1 for all <code>any</code> and 0 for all other classes. If there is no bug, that means that all <code>any</code> that the public test contains are 1, and all other classes are 0. </p></li>\n<li><p>The above observation also means that the public test does not contain all 6 rows per image, as 1 for <code>any</code> necessarily means 1 in one of the classes.</p></li>\n<li><p>By changing threshold of binarization and submitting I observed, and many people reported as well, that the score changed in one step, with intermediate magic score of 13.837. No other scores were observed by submitting only ones and zeros (besides all zeros 16.353, and all ones 18.185). This could indicate that there are just a few rows in the public test, but the weights do not match any pattern.</p></li>\n</ol>\n\n<p><a href=\"/philculliton\">@philculliton</a> , <a href=\"/juliaelliott\">@juliaelliott</a> , can you please confirm that the behavior is as expected and was designed that way.</p>",
  "messages": [
    {
      "id": 668315,
      "postDate": "2019-11-08T09:04:55.023Z",
      "content": "<p>It seems that there is a bug in public LB calculation. Below is a list of observations that support it.</p>\n\n<ol>\n<li><p>One can get score 0.000 by submitting 1 for all <code>any</code> and 0 for all other classes. If there is no bug, that means that all <code>any</code> that the public test contains are 1, and all other classes are 0. </p></li>\n<li><p>The above observation also means that the public test does not contain all 6 rows per image, as 1 for <code>any</code> necessarily means 1 in one of the classes.</p></li>\n<li><p>By changing threshold of binarization and submitting I observed, and many people reported as well, that the score changed in one step, with intermediate magic score of 13.837. No other scores were observed by submitting only ones and zeros (besides all zeros 16.353, and all ones 18.185). This could indicate that there are just a few rows in the public test, but the weights do not match any pattern.</p></li>\n</ol>\n\n<p><a href=\"/philculliton\">@philculliton</a> , <a href=\"/juliaelliott\">@juliaelliott</a> , can you please confirm that the behavior is as expected and was designed that way.</p>",
      "rawMarkdown": "It seems that there is a bug in public LB calculation. Below is a list of observations that support it.\n\n1. One can get score 0.000 by submitting 1 for all `any` and 0 for all other classes. If there is no bug, that means that all `any` that the public test contains are 1, and all other classes are 0. \n\n2. The above observation also means that the public test does not contain all 6 rows per image, as 1 for `any` necessarily means 1 in one of the classes.\n\n3. By changing threshold of binarization and submitting I observed, and many people reported as well, that the score changed in one step, with intermediate magic score of 13.837. No other scores were observed by submitting only ones and zeros (besides all zeros 16.353, and all ones 18.185). This could indicate that there are just a few rows in the public test, but the weights do not match any pattern.\n\n@philculliton , @juliaelliott , can you please confirm that the behavior is as expected and was designed that way.",
      "votes": 18
    },
    {
      "id": 668413,
      "postDate": "2019-11-08T12:04:08.130Z",
      "content": "<p>The issue is the Public LB doesn't do what is is supposed to do - help us fined bugs in our code. The major, and most common but is - an error in the images order. If this LB doesn't help us find this bug (or any other bug).\n<a href=\"/juliaelliott\">@juliaelliott</a> we need help here.</p>",
      "rawMarkdown": "The issue is the Public LB doesn't do what is is supposed to do - help us fined bugs in our code. The major, and most common but is - an error in the images order. If this LB doesn't help us find this bug (or any other bug).\n@juliaelliott we need help here.",
      "votes": 6,
      "replies": [
        {
          "id": 668458,
          "postDate": "2019-11-08T13:00:13.783Z",
          "content": "<p>Hi <a href=\"/yuval6967\">@yuval6967</a>! The public leaderboard was intended to provide the smallest possible amount of feedback for people to debug their submissions.</p>\n\n<p>What do you mean by an error in image order? (The order of the rows in the submission is not important.)</p>",
          "rawMarkdown": "Hi @yuval6967! The public leaderboard was intended to provide the smallest possible amount of feedback for people to debug their submissions.\n\nWhat do you mean by an error in image order? (The order of the rows in the submission is not important.)",
          "votes": 1
        },
        {
          "id": 668462,
          "postDate": "2019-11-08T13:03:33.227Z",
          "content": "<blockquote>\n  <p>What do you mean by an error in image order?</p>\n</blockquote>\n\n<p>I guess he means that current public LB is so non-representative, that you can shuffle labels and IDs in submission but still achieve comparable score (so you cannot even debug your pipeline properly).</p>",
          "rawMarkdown": "&gt; What do you mean by an error in image order?\n\nI guess he means that current public LB is so non-representative, that you can shuffle labels and IDs in submission but still achieve comparable score (so you cannot even debug your pipeline properly).",
          "votes": 4
        },
        {
          "id": 668589,
          "postDate": "2019-11-08T15:39:01.870Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Currently the only indication we get from the LB is: the submission has all the IDs and only them and the Label is in the range 0-1, otherwise the submission will fail - so the number itself gives no indication and would better be replaced with a simple successful/fail indicator.</p>\n\n<p>One of the common disastrous bugs is inferencing on the wrong data, or losing the correct reference between the image examined and it's ID. This error is even more common when changing datasets and when the pipeline isn't a simple model. - This is what I mean by image order.\nAnd there are a lot of common bugs that could have been discovered if the LB would have been calculated on a random 1% of the data.\nAnd if the decision was only to indicate that the submission was successful and you don't have a mechanism to do it without giving a score.  It could have been written explicitly \"The number is not an indication - every submission that didn't fail, is OK'   </p>\n\n<p><strong>Just look at how many discussion are where initiated because of this issue</strong>  </p>",
          "rawMarkdown": "@philculliton Currently the only indication we get from the LB is: the submission has all the IDs and only them and the Label is in the range 0-1, otherwise the submission will fail - so the number itself gives no indication and would better be replaced with a simple successful/fail indicator.\n\nOne of the common disastrous bugs is inferencing on the wrong data, or losing the correct reference between the image examined and it's ID. This error is even more common when changing datasets and when the pipeline isn't a simple model. - This is what I mean by image order.\nAnd there are a lot of common bugs that could have been discovered if the LB would have been calculated on a random 1% of the data.\nAnd if the decision was only to indicate that the submission was successful and you don't have a mechanism to do it without giving a score.  It could have been written explicitly \"The number is not an indication - every submission that didn't fail, is OK'   \n\n**Just look at how many discussion are where initiated because of this issue**  ",
          "votes": 2
        },
        {
          "id": 668594,
          "postDate": "2019-11-08T15:46:58.917Z",
          "content": "<p>This is not first RSNA competition as I know. Stage 2 is not supposed to help you with debugging code - the idea is on contrary to build a strategy during Stage 1 that generalizes well on new data. And if we had f.e. 10% of data - that would be enough for decision making of what strategy works better for Stage 2 (pseudo labeling etc). </p>",
          "rawMarkdown": "This is not first RSNA competition as I know. Stage 2 is not supposed to help you with debugging code - the idea is on contrary to build a strategy during Stage 1 that generalizes well on new data. And if we had f.e. 10% of data - that would be enough for decision making of what strategy works better for Stage 2 (pseudo labeling etc). ",
          "votes": 2
        },
        {
          "id": 668599,
          "postDate": "2019-11-08T15:53:56.110Z",
          "content": "<blockquote>\n  <p>the Label is in the range 0-1, otherwise the submission will fail</p>\n</blockquote>\n\n<p>Actually, even that is not true. I submitted logits instead of sigmoids accidentally and got ~13 LB score :)</p>",
          "rawMarkdown": "&gt; the Label is in the range 0-1, otherwise the submission will fail\n\nActually, even that is not true. I submitted logits instead of sigmoids accidentally and got ~13 LB score :)",
          "votes": 3
        },
        {
          "id": 668601,
          "postDate": "2019-11-08T15:57:16.810Z",
          "content": "<p>No need for 10%. even 1% or 0.5%  is enough for better debugging as long as it is random. </p>",
          "rawMarkdown": "No need for 10%. even 1% or 0.5%  is enough for better debugging as long as it is random. ",
          "votes": 2
        },
        {
          "id": 668638,
          "postDate": "2019-11-08T16:44:12.520Z",
          "content": "<p>But this is not about debugging :)</p>",
          "rawMarkdown": "But this is not about debugging :)",
          "votes": 1
        },
        {
          "id": 668641,
          "postDate": "2019-11-08T16:54:00.043Z",
          "content": "<p>Thanks for all the feedback!</p>\n\n<blockquote>\n  <p>Actually, even that is not true. I submitted logits instead of sigmoids accidentally and got ~13 LB score :)</p>\n</blockquote>\n\n<p>The above is the sort of information you should be gleaning from the public LB. It is not intended for debugging code changes - in Stage 2 you should only be changing the code required to read in new input data. It's intended to give <em>some</em> (as minimal as possible) amount of information about what you're submitting.</p>\n\n<blockquote>\n  <p>This is not first RSNA competition as I know. Stage 2 is not supposed to help you with debugging code - the idea is on contrary to build a strategy during Stage 1 that generalizes well on new data. And if we had f.e. 10% of data - that would be enough for decision making of what strategy works better for Stage 2 (pseudo labeling etc).</p>\n</blockquote>\n\n<p>This is correct. The intention is not for Stage 2 to be a new round of decision-making. It's to run your Stage 1 code on new data.</p>",
          "rawMarkdown": "Thanks for all the feedback!\n\n&gt; Actually, even that is not true. I submitted logits instead of sigmoids accidentally and got ~13 LB score :)\n\nThe above is the sort of information you should be gleaning from the public LB. It is not intended for debugging code changes - in Stage 2 you should only be changing the code required to read in new input data. It's intended to give _some_ (as minimal as possible) amount of information about what you're submitting.\n\n&gt; This is not first RSNA competition as I know. Stage 2 is not supposed to help you with debugging code - the idea is on contrary to build a strategy during Stage 1 that generalizes well on new data. And if we had f.e. 10% of data - that would be enough for decision making of what strategy works better for Stage 2 (pseudo labeling etc).\n\nThis is correct. The intention is not for Stage 2 to be a new round of decision-making. It's to run your Stage 1 code on new data.",
          "votes": 2
        },
        {
          "id": 668669,
          "postDate": "2019-11-08T17:30:57.047Z",
          "content": "<p>&gt; It is not intended for debugging code changes</p>\n\n<p>I understand that it shouldn't allow me to decide that \"model A is better than model B\". But we hoped that it at least should allow us to determine that \"my overly complex pipeline comprised of 7 Jupyter Notebooks generated submit that was intended (not some gibberish)\" 😧 </p>",
          "rawMarkdown": "&gt; It is not intended for debugging code changes\n\nI understand that it shouldn't allow me to decide that \"model A is better than model B\". But we hoped that it at least should allow us to determine that \"my overly complex pipeline comprised of 7 Jupyter Notebooks generated submit that was intended (not some gibberish)\" 😧 ",
          "votes": 3
        },
        {
          "id": 668718,
          "postDate": "2019-11-08T19:10:36.033Z",
          "content": "<p>Just for fun I made 2 tests:\n1) Reorder rows of the submission. Nothing is changed, as expected.\n2) Shuffle 'Label' values. That's what some people afraid of. The score changed from 0.583 to 2.235. So at least this LB can show some errors!</p>",
          "rawMarkdown": "Just for fun I made 2 tests:\n1) Reorder rows of the submission. Nothing is changed, as expected.\n2) Shuffle 'Label' values. That's what some people afraid of. The score changed from 0.583 to 2.235. So at least this LB can show some errors!",
          "votes": 6
        }
      ]
    },
    {
      "id": 668329,
      "postDate": "2019-11-08T09:25:59.713Z",
      "content": "<p>That has gone too far...</p>",
      "rawMarkdown": "That has gone too far...",
      "votes": 5
    },
    {
      "id": 668334,
      "postDate": "2019-11-08T09:35:51.140Z",
      "content": "<p>The right thing to do would be to fix the bug if it exists (sounds like it) and recompute the LB.</p>",
      "rawMarkdown": "The right thing to do would be to fix the bug if it exists (sounds like it) and recompute the LB.",
      "votes": 1
    },
    {
      "id": 668332,
      "postDate": "2019-11-08T09:33:45.113Z",
      "content": "<p>This is bad. Hopefully this bug will not be on private LB.</p>",
      "rawMarkdown": "This is bad. Hopefully this bug will not be on private LB.",
      "votes": 1
    },
    {
      "id": 668356,
      "postDate": "2019-11-08T10:11:09.583Z",
      "content": "<p>But who told that public should include all 6 rows per image?</p>",
      "rawMarkdown": "But who told that public should include all 6 rows per image?",
      "votes": 2,
      "replies": [
        {
          "id": 668360,
          "postDate": "2019-11-08T10:19:28.027Z",
          "content": "<p>Yes, it can be a conscious decision. But I think it is unlikely, and given the other points even more unlikely. If there is a bug, it is better be fixed.</p>",
          "rawMarkdown": "Yes, it can be a conscious decision. But I think it is unlikely, and given the other points even more unlikely. If there is a bug, it is better be fixed.",
          "votes": 3
        },
        {
          "id": 668444,
          "postDate": "2019-11-08T12:47:05.807Z",
          "content": "<p>Hi! Thanks for your questions. It was a conscious decision. The public leaderboard does not contain all six rows from a particular image. It is vanishingly small.</p>",
          "rawMarkdown": "Hi! Thanks for your questions. It was a conscious decision. The public leaderboard does not contain all six rows from a particular image. It is vanishingly small.",
          "votes": 4
        }
      ]
    },
    {
      "id": 668520,
      "postDate": "2019-11-08T14:17:36.830Z",
      "content": "<p>Some strong memes required for such a stress :)\nSorry guys))</p>\n\n<p>As for debug, no changes of the code allowed. So, we just re-run the code to know that it runs, that's all, and wait for private LB, few days remaining. </p>",
      "rawMarkdown": "Some strong memes required for such a stress :)\nSorry guys))\n\nAs for debug, no changes of the code allowed. So, we just re-run the code to know that it runs, that's all, and wait for private LB, few days remaining. \n\n",
      "votes": 1
    },
    {
      "id": 668935,
      "postDate": "2019-11-09T06:11:09.757Z",
      "content": "<p>It's probably best to forget probing Public LB and solely trust local CV...</p>",
      "rawMarkdown": "It's probably best to forget probing Public LB and solely trust local CV..."
    },
    {
      "id": 668446,
      "postDate": "2019-11-08T12:48:55.493Z",
      "rawMarkdown": ""
    },
    {
      "id": 668355,
      "postDate": "2019-11-08T10:06:58.900Z",
      "content": "<p>It could also mean that the selected rows to compare to do not contain all the rows for the images.\ne.g. only ID_xxx_any rows ?</p>",
      "rawMarkdown": "It could also mean that the selected rows to compare to do not contain all the rows for the images.\ne.g. only ID_xxx_any rows ?"
    },
    {
      "id": 668341,
      "postDate": "2019-11-08T09:53:59.257Z",
      "content": "<p>&gt;  If there is no bug, that means that all any that the public test contains are 1, and all other classes are 0. </p>\n\n<p>This also can mean that true labels are not so good.</p>",
      "rawMarkdown": "&gt;  If there is no bug, that means that all any that the public test contains are 1, and all other classes are 0. \n\nThis also can mean that true labels are not so good."
    }
  ],
  "comments": [
    {
      "id": 668413,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2019-11-08T12:04:08.130000",
      "content": "<p>The issue is the Public LB doesn't do what is is supposed to do - help us fined bugs in our code. The major, and most common but is - an error in the images order. If this LB doesn't help us find this bug (or any other bug).\n<a href=\"/juliaelliott\">@juliaelliott</a> we need help here.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 668458,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-11-08T13:00:13.783000",
          "content": "<p>Hi <a href=\"/yuval6967\">@yuval6967</a>! The public leaderboard was intended to provide the smallest possible amount of feedback for people to debug their submissions.</p>\n\n<p>What do you mean by an error in image order? (The order of the rows in the submission is not important.)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 668462,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2019-11-08T13:03:33.227000",
          "content": "<blockquote>\n  <p>What do you mean by an error in image order?</p>\n</blockquote>\n\n<p>I guess he means that current public LB is so non-representative, that you can shuffle labels and IDs in submission but still achieve comparable score (so you cannot even debug your pipeline properly).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 668589,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2019-11-08T15:39:01.870000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Currently the only indication we get from the LB is: the submission has all the IDs and only them and the Label is in the range 0-1, otherwise the submission will fail - so the number itself gives no indication and would better be replaced with a simple successful/fail indicator.</p>\n\n<p>One of the common disastrous bugs is inferencing on the wrong data, or losing the correct reference between the image examined and it's ID. This error is even more common when changing datasets and when the pipeline isn't a simple model. - This is what I mean by image order.\nAnd there are a lot of common bugs that could have been discovered if the LB would have been calculated on a random 1% of the data.\nAnd if the decision was only to indicate that the submission was successful and you don't have a mechanism to do it without giving a score.  It could have been written explicitly \"The number is not an indication - every submission that didn't fail, is OK'   </p>\n\n<p><strong>Just look at how many discussion are where initiated because of this issue</strong>  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 668594,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2019-11-08T15:46:58.917000",
          "content": "<p>This is not first RSNA competition as I know. Stage 2 is not supposed to help you with debugging code - the idea is on contrary to build a strategy during Stage 1 that generalizes well on new data. And if we had f.e. 10% of data - that would be enough for decision making of what strategy works better for Stage 2 (pseudo labeling etc). </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 668599,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2019-11-08T15:53:56.110000",
          "content": "<blockquote>\n  <p>the Label is in the range 0-1, otherwise the submission will fail</p>\n</blockquote>\n\n<p>Actually, even that is not true. I submitted logits instead of sigmoids accidentally and got ~13 LB score :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 668601,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2019-11-08T15:57:16.810000",
          "content": "<p>No need for 10%. even 1% or 0.5%  is enough for better debugging as long as it is random. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 668638,
          "author_name": "Oleg Yaroshevskiy",
          "author_url": "",
          "post_date": "2019-11-08T16:44:12.520000",
          "content": "<p>But this is not about debugging :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 668641,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-11-08T16:54:00.043000",
          "content": "<p>Thanks for all the feedback!</p>\n\n<blockquote>\n  <p>Actually, even that is not true. I submitted logits instead of sigmoids accidentally and got ~13 LB score :)</p>\n</blockquote>\n\n<p>The above is the sort of information you should be gleaning from the public LB. It is not intended for debugging code changes - in Stage 2 you should only be changing the code required to read in new input data. It's intended to give <em>some</em> (as minimal as possible) amount of information about what you're submitting.</p>\n\n<blockquote>\n  <p>This is not first RSNA competition as I know. Stage 2 is not supposed to help you with debugging code - the idea is on contrary to build a strategy during Stage 1 that generalizes well on new data. And if we had f.e. 10% of data - that would be enough for decision making of what strategy works better for Stage 2 (pseudo labeling etc).</p>\n</blockquote>\n\n<p>This is correct. The intention is not for Stage 2 to be a new round of decision-making. It's to run your Stage 1 code on new data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 668669,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2019-11-08T17:30:57.047000",
          "content": "<p>&gt; It is not intended for debugging code changes</p>\n\n<p>I understand that it shouldn't allow me to decide that \"model A is better than model B\". But we hoped that it at least should allow us to determine that \"my overly complex pipeline comprised of 7 Jupyter Notebooks generated submit that was intended (not some gibberish)\" 😧 </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 668718,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2019-11-08T19:10:36.033000",
          "content": "<p>Just for fun I made 2 tests:\n1) Reorder rows of the submission. Nothing is changed, as expected.\n2) Shuffle 'Label' values. That's what some people afraid of. The score changed from 0.583 to 2.235. So at least this LB can show some errors!</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 668329,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-11-08T09:25:59.713000",
      "content": "<p>That has gone too far...</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 668334,
      "author_name": "Cogitae _ Thomas Soumarmon",
      "author_url": "",
      "post_date": "2019-11-08T09:35:51.140000",
      "content": "<p>The right thing to do would be to fix the bug if it exists (sounds like it) and recompute the LB.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 668332,
      "author_name": "Cogitae _ Thomas Soumarmon",
      "author_url": "",
      "post_date": "2019-11-08T09:33:45.113000",
      "content": "<p>This is bad. Hopefully this bug will not be on private LB.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 668356,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-11-08T10:11:09.583000",
      "content": "<p>But who told that public should include all 6 rows per image?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 668360,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-11-08T10:19:28.027000",
          "content": "<p>Yes, it can be a conscious decision. But I think it is unlikely, and given the other points even more unlikely. If there is a bug, it is better be fixed.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 668444,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-11-08T12:47:05.807000",
          "content": "<p>Hi! Thanks for your questions. It was a conscious decision. The public leaderboard does not contain all six rows from a particular image. It is vanishingly small.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 668520,
      "author_name": "Mukharbek Organokov",
      "author_url": "",
      "post_date": "2019-11-08T14:17:36.830000",
      "content": "<p>Some strong memes required for such a stress :)\nSorry guys))</p>\n\n<p>As for debug, no changes of the code allowed. So, we just re-run the code to know that it runs, that's all, and wait for private LB, few days remaining. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 668935,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2019-11-09T06:11:09.757000",
      "content": "<p>It's probably best to forget probing Public LB and solely trust local CV...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 668446,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2019-11-08T12:48:55.493000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 668355,
      "author_name": "Cogitae _ Thomas Soumarmon",
      "author_url": "",
      "post_date": "2019-11-08T10:06:58.900000",
      "content": "<p>It could also mean that the selected rows to compare to do not contain all the rows for the images.\ne.g. only ID_xxx_any rows ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 668341,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2019-11-08T09:53:59.257000",
      "content": "<p>&gt;  If there is no bug, that means that all any that the public test contains are 1, and all other classes are 0. </p>\n\n<p>This also can mean that true labels are not so good.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "668315": "It seems that there is a bug in public LB calculation. Below is a list of observations that support it.\n\n1. One can get score 0.000 by submitting 1 for all `any` and 0 for all other classes. If there is no bug, that means that all `any` that the public test contains are 1, and all other classes are 0. \n\n2. The above observation also means that the public test does not contain all 6 rows per image, as 1 for `any` necessarily means 1 in one of the classes.\n\n3. By changing threshold of binarization and submitting I observed, and many people reported as well, that the score changed in one step, with intermediate magic score of 13.837. No other scores were observed by submitting only ones and zeros (besides all zeros 16.353, and all ones 18.185). This could indicate that there are just a few rows in the public test, but the weights do not match any pattern.\n\n@philculliton , @juliaelliott , can you please confirm that the behavior is as expected and was designed that way.",
    "668413": "The issue is the Public LB doesn't do what is is supposed to do - help us fined bugs in our code. The major, and most common but is - an error in the images order. If this LB doesn't help us find this bug (or any other bug).\n@juliaelliott we need help here.",
    "668329": "That has gone too far...",
    "668334": "The right thing to do would be to fix the bug if it exists (sounds like it) and recompute the LB.",
    "668332": "This is bad. Hopefully this bug will not be on private LB.",
    "668356": "But who told that public should include all 6 rows per image?",
    "668520": "Some strong memes required for such a stress :)\nSorry guys))\n\nAs for debug, no changes of the code allowed. So, we just re-run the code to know that it runs, that's all, and wait for private LB, few days remaining. \n\n",
    "668935": "It's probably best to forget probing Public LB and solely trust local CV...",
    "668446": "",
    "668355": "It could also mean that the selected rows to compare to do not contain all the rows for the images.\ne.g. only ID_xxx_any rows ?",
    "668341": "&gt;  If there is no bug, that means that all any that the public test contains are 1, and all other classes are 0. \n\nThis also can mean that true labels are not so good."
  }
}