{
  "id": 193624,
  "title": "Wow - Baseline Scores 38th Place Silver Medal - Private LB 0.208",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193624",
  "author_name": "Chris Deotte",
  "post_date": "2020-10-27T23:45:10.402000",
  "votes": 29,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Out of curiosity, I just submitted Kun Hao Yeh's baseline <a href=\"https://www.kaggle.com/khyeh0719/cnn-gru-baseline-stage2-train-inference\" target=\"_blank\">here</a> without using his post process. Without post process this notebook would score public LB 0.214 and private LB 0.208 and achieve 38th place and silver medal. Wow!!</p>\n<h1>Baseline with PP - Private LB 0.232</h1>\n<h2>(notebook as is without changes)</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcb2ee239782b3060abcbbc901f7b526f%2FyesPP.png?generation=1603842039101311&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F0a3926001d1032c9ea29ab0474dfa432%2FyesPP2.png?generation=1603842519769998&amp;alt=media\" alt=\"\"></p>\n<h1>Baseline without PP - Private LB 0.208</h1>\n<h2>(notebook with commented out post process)</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F046f24bc9035e67b1ebc29141d393ed8%2FnoPP.png?generation=1603842060148010&amp;alt=media\" alt=\"\"></p>\n<h1>UPDATE - Oct 28th</h1>\n<p>Using a more efficient post process, it is at least possible to achieve private LB 0.209 and conform to label consistency. Details are in comments below. Theoretically post process should improve LB, so it is possible to design a great PP that achieves private LB 0.208 or better and conforms to label consistency.</p>",
  "messages": [
    {
      "id": 1062522,
      "postDate": "2020-10-27T23:45:10.403Z",
      "content": "<p>Out of curiosity, I just submitted Kun Hao Yeh's baseline <a href=\"https://www.kaggle.com/khyeh0719/cnn-gru-baseline-stage2-train-inference\" target=\"_blank\">here</a> without using his post process. Without post process this notebook would score public LB 0.214 and private LB 0.208 and achieve 38th place and silver medal. Wow!!</p>\n<h1>Baseline with PP - Private LB 0.232</h1>\n<h2>(notebook as is without changes)</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcb2ee239782b3060abcbbc901f7b526f%2FyesPP.png?generation=1603842039101311&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F0a3926001d1032c9ea29ab0474dfa432%2FyesPP2.png?generation=1603842519769998&amp;alt=media\" alt=\"\"></p>\n<h1>Baseline without PP - Private LB 0.208</h1>\n<h2>(notebook with commented out post process)</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F046f24bc9035e67b1ebc29141d393ed8%2FnoPP.png?generation=1603842060148010&amp;alt=media\" alt=\"\"></p>\n<h1>UPDATE - Oct 28th</h1>\n<p>Using a more efficient post process, it is at least possible to achieve private LB 0.209 and conform to label consistency. Details are in comments below. Theoretically post process should improve LB, so it is possible to design a great PP that achieves private LB 0.208 or better and conforms to label consistency.</p>",
      "rawMarkdown": "Out of curiosity, I just submitted Kun Hao Yeh's baseline [here][1] without using his post process. Without post process this notebook would score public LB 0.214 and private LB 0.208 and achieve 38th place and silver medal. Wow!!\n\n# Baseline with PP - Private LB 0.232\n## (notebook as is without changes)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcb2ee239782b3060abcbbc901f7b526f%2FyesPP.png?generation=1603842039101311&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F0a3926001d1032c9ea29ab0474dfa432%2FyesPP2.png?generation=1603842519769998&alt=media)\n\n# Baseline without PP - Private LB 0.208\n## (notebook with commented out post process)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F046f24bc9035e67b1ebc29141d393ed8%2FnoPP.png?generation=1603842060148010&alt=media)\n\n# UPDATE - Oct 28th\nUsing a more efficient post process, it is at least possible to achieve private LB 0.209 and conform to label consistency. Details are in comments below. Theoretically post process should improve LB, so it is possible to design a great PP that achieves private LB 0.208 or better and conforms to label consistency.\n\n[1]: https://www.kaggle.com/khyeh0719/cnn-gru-baseline-stage2-train-inference\n",
      "votes": 29
    },
    {
      "id": 1062796,
      "postDate": "2020-10-28T07:31:12.653Z",
      "content": "<p>It's really a shame for people who worked hard and chose a compliant submission to lose 20~ranks because of this. I would love the Kaggle team to look into it but I doubt it will happen.</p>",
      "rawMarkdown": "It's really a shame for people who worked hard and chose a compliant submission to lose 20~ranks because of this. I would love the Kaggle team to look into it but I doubt it will happen.\n",
      "votes": 12
    },
    {
      "id": 1062524,
      "postDate": "2020-10-27T23:51:39.807Z",
      "content": "<p>Yeah, that's why there are so many 0.208 on the private Lb.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2006644%2F61fe87bf22feb2fc3765042f48f77214%2F2020-10-28%208.48.42.png?generation=1603842599326467&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Yeah, that's why there are so many 0.208 on the private Lb.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2006644%2F61fe87bf22feb2fc3765042f48f77214%2F2020-10-28%208.48.42.png?generation=1603842599326467&alt=media)",
      "votes": 8
    },
    {
      "id": 1062734,
      "postDate": "2020-10-28T06:26:55.967Z",
      "content": "<p>Even if it is a rule violation: \"Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page\", <strong>unfortunately kaggle may just ignore this cheating</strong> despite many teams tried to be honest with the rules(  <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193372\" target=\"_blank\">wrote</a> that they are going to check only the first 10 teams… </p>",
      "rawMarkdown": "Even if it is a rule violation: \"Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page\", **unfortunately kaggle may just ignore this cheating** despite many teams tried to be honest with the rules(  @juliaelliott [wrote](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193372) that they are going to check only the first 10 teams... ",
      "votes": 6,
      "replies": [
        {
          "id": 1062740,
          "postDate": "2020-10-28T06:34:00.997Z",
          "content": "<p>I'm running some experiments. It appears that that notebook's post process is very inefficient. I have checked public LB so far and I am checking private LB now. Here are my results so far</p>\n<ul>\n<li>Notebook with notebook PP (conforms to consistency) - Public LB 0.233</li>\n<li>Notebook without notebook PP (does not conform to consistency) - Public LB 0.214</li>\n<li>Notebook with my PP (conforms to consistency) - Public LB 0.215</li>\n</ul>\n<p>When i get results for private LB, i will post. But it may be possible to post process (with an efficient PP) that public notebook and conform to consistency and score private LB 0.208</p>",
          "rawMarkdown": "I'm running some experiments. It appears that that notebook's post process is very inefficient. I have checked public LB so far and I am checking private LB now. Here are my results so far\n* Notebook with notebook PP (conforms to consistency) - Public LB 0.233\n* Notebook without notebook PP (does not conform to consistency) - Public LB 0.214\n* Notebook with my PP (conforms to consistency) - Public LB 0.215\n\nWhen i get results for private LB, i will post. But it may be possible to post process (with an efficient PP) that public notebook and conform to consistency and score private LB 0.208",
          "votes": 4
        },
        {
          "id": 1062757,
          "postDate": "2020-10-28T06:49:32.277Z",
          "content": "<p>Thank you for checking it. The PP posted in that notebook gave quite a huge performance drop when I checked it with one of my models. It indeed may be possible that a better approach could give 0.208 private LB. In this case I apologize.</p>",
          "rawMarkdown": "Thank you for checking it. The PP posted in that notebook gave quite a huge performance drop when I checked it with one of my models. It indeed may be possible that a better approach could give 0.208 private LB. In this case I apologize.",
          "votes": 3
        },
        {
          "id": 1062862,
          "postDate": "2020-10-28T09:10:56.060Z",
          "content": "<p>It's also possible that everyone came up with the same post-processing, hence 0.214's and 0.208's. The chances for this to happen are indeed slim and if i recall correctly in probability theory this is called <strong>almost never</strong></p>",
          "rawMarkdown": "It's also possible that everyone came up with the same post-processing, hence 0.214's and 0.208's. The chances for this to happen are indeed slim and if i recall correctly in probability theory this is called **almost never**",
          "votes": 3
        },
        {
          "id": 1063469,
          "postDate": "2020-10-29T00:47:24.423Z",
          "content": "<p>Using my more efficient post process, the public notebook can conform to label consistency and score private LB 0.209. My PP is not the best, so it is possible to develop a PP that achieves private 0.208 or lower because theoretically enforcing label consistency can improve a model's LB</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F6c93a63070d413094f52ab4e012e5567%2Fmy-pp.png?generation=1603932385470426&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Notebook with notebook PP (conforms to consistency) - Private LB 0.232</li>\n<li>Notebook without notebook PP (does not conform to consistency) - Private LB 0.208</li>\n<li>Notebook with my PP (conforms to consistency) - Private LB 0.209</li>\n<li>Notebook with better PP (conforms to consistency) - Private LB 0.208 or less</li>\n</ul>",
          "rawMarkdown": "Using my more efficient post process, the public notebook can conform to label consistency and score private LB 0.209. My PP is not the best, so it is possible to develop a PP that achieves private 0.208 or lower because theoretically enforcing label consistency can improve a model's LB\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F6c93a63070d413094f52ab4e012e5567%2Fmy-pp.png?generation=1603932385470426&alt=media)\n\n* Notebook with notebook PP (conforms to consistency) - Private LB 0.232\n* Notebook without notebook PP (does not conform to consistency) - Private LB 0.208\n* Notebook with my PP (conforms to consistency) - Private LB 0.209\n* Notebook with better PP (conforms to consistency) - Private LB 0.208 or less",
          "votes": 3
        },
        {
          "id": 1063884,
          "postDate": "2020-10-29T13:05:33.180Z",
          "content": "<p>To be fair to all host should have applied label consistency check to all or to none of the teams. What is the point of checking only the top 10 and how one will know that he going to finish in the top 10 and therefore also have to meet some additional requirements?  </p>\n<p>Btw  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> can you share you PP using which you got 0.209 on Private LB on the above baseline notebook?</p>",
          "rawMarkdown": "To be fair to all host should have applied label consistency check to all or to none of the teams. What is the point of checking only the top 10 and how one will know that he going to finish in the top 10 and therefore also have to meet some additional requirements?  \n\nBtw  @cdeotte can you share you PP using which you got 0.209 on Private LB on the above baseline notebook?",
          "votes": 2
        },
        {
          "id": 1064140,
          "postDate": "2020-10-29T18:54:26.610Z",
          "content": "<p>Sure, I will clean up the code, add comments, and publish.</p>",
          "rawMarkdown": "Sure, I will clean up the code, add comments, and publish."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1062796,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-10-28T07:31:12.653000",
      "content": "<p>It's really a shame for people who worked hard and chose a compliant submission to lose 20~ranks because of this. I would love the Kaggle team to look into it but I doubt it will happen.</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 1062524,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2020-10-27T23:51:39.807000",
      "content": "<p>Yeah, that's why there are so many 0.208 on the private Lb.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2006644%2F61fe87bf22feb2fc3765042f48f77214%2F2020-10-28%208.48.42.png?generation=1603842599326467&amp;alt=media\" alt=\"\"></p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1062734,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-10-28T06:26:55.967000",
      "content": "<p>Even if it is a rule violation: \"Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page\", <strong>unfortunately kaggle may just ignore this cheating</strong> despite many teams tried to be honest with the rules(  <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193372\" target=\"_blank\">wrote</a> that they are going to check only the first 10 teams… </p>",
      "votes": 6,
      "replies": [
        {
          "id": 1062740,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-28T06:34:00.997000",
          "content": "<p>I'm running some experiments. It appears that that notebook's post process is very inefficient. I have checked public LB so far and I am checking private LB now. Here are my results so far</p>\n<ul>\n<li>Notebook with notebook PP (conforms to consistency) - Public LB 0.233</li>\n<li>Notebook without notebook PP (does not conform to consistency) - Public LB 0.214</li>\n<li>Notebook with my PP (conforms to consistency) - Public LB 0.215</li>\n</ul>\n<p>When i get results for private LB, i will post. But it may be possible to post process (with an efficient PP) that public notebook and conform to consistency and score private LB 0.208</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1062757,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-10-28T06:49:32.277000",
          "content": "<p>Thank you for checking it. The PP posted in that notebook gave quite a huge performance drop when I checked it with one of my models. It indeed may be possible that a better approach could give 0.208 private LB. In this case I apologize.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1062862,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-10-28T09:10:56.060000",
          "content": "<p>It's also possible that everyone came up with the same post-processing, hence 0.214's and 0.208's. The chances for this to happen are indeed slim and if i recall correctly in probability theory this is called <strong>almost never</strong></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1063469,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-29T00:47:24.423000",
          "content": "<p>Using my more efficient post process, the public notebook can conform to label consistency and score private LB 0.209. My PP is not the best, so it is possible to develop a PP that achieves private 0.208 or lower because theoretically enforcing label consistency can improve a model's LB</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F6c93a63070d413094f52ab4e012e5567%2Fmy-pp.png?generation=1603932385470426&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Notebook with notebook PP (conforms to consistency) - Private LB 0.232</li>\n<li>Notebook without notebook PP (does not conform to consistency) - Private LB 0.208</li>\n<li>Notebook with my PP (conforms to consistency) - Private LB 0.209</li>\n<li>Notebook with better PP (conforms to consistency) - Private LB 0.208 or less</li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1063884,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-10-29T13:05:33.180000",
          "content": "<p>To be fair to all host should have applied label consistency check to all or to none of the teams. What is the point of checking only the top 10 and how one will know that he going to finish in the top 10 and therefore also have to meet some additional requirements?  </p>\n<p>Btw  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> can you share you PP using which you got 0.209 on Private LB on the above baseline notebook?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1064140,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-29T18:54:26.610000",
          "content": "<p>Sure, I will clean up the code, add comments, and publish.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1062522": "Out of curiosity, I just submitted Kun Hao Yeh's baseline [here][1] without using his post process. Without post process this notebook would score public LB 0.214 and private LB 0.208 and achieve 38th place and silver medal. Wow!!\n\n# Baseline with PP - Private LB 0.232\n## (notebook as is without changes)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcb2ee239782b3060abcbbc901f7b526f%2FyesPP.png?generation=1603842039101311&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F0a3926001d1032c9ea29ab0474dfa432%2FyesPP2.png?generation=1603842519769998&alt=media)\n\n# Baseline without PP - Private LB 0.208\n## (notebook with commented out post process)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F046f24bc9035e67b1ebc29141d393ed8%2FnoPP.png?generation=1603842060148010&alt=media)\n\n# UPDATE - Oct 28th\nUsing a more efficient post process, it is at least possible to achieve private LB 0.209 and conform to label consistency. Details are in comments below. Theoretically post process should improve LB, so it is possible to design a great PP that achieves private LB 0.208 or better and conforms to label consistency.\n\n[1]: https://www.kaggle.com/khyeh0719/cnn-gru-baseline-stage2-train-inference\n",
    "1062796": "It's really a shame for people who worked hard and chose a compliant submission to lose 20~ranks because of this. I would love the Kaggle team to look into it but I doubt it will happen.\n",
    "1062524": "Yeah, that's why there are so many 0.208 on the private Lb.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2006644%2F61fe87bf22feb2fc3765042f48f77214%2F2020-10-28%208.48.42.png?generation=1603842599326467&alt=media)",
    "1062734": "Even if it is a rule violation: \"Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page\", **unfortunately kaggle may just ignore this cheating** despite many teams tried to be honest with the rules(  @juliaelliott [wrote](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193372) that they are going to check only the first 10 teams... "
  }
}