{
  "id": 190411,
  "title": "Only 39 teams have beaten public  baseline",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/190411",
  "author_name": "Kamal Das",
  "post_date": "2020-10-11T17:20:07.073000",
  "votes": 0,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi All,<br>\nLooks like  a difficult competition… Only 39 participants/teams have beaten public  baseline of 0.325</p>\n<p>Those who have, looking forward to your suggestions on approach and next steps to beat the baseline</p>\n<p>As last 2 weeks, not looking for code but suggestions on what would help to improve and also retain the place post shake up!! <br>\nAs private dataset is 3x of public LB dataset, don't want to just be up on public one!</p>\n<p>All suggestions and approaches are welcome and highly appreciated!!</p>",
  "messages": [
    {
      "id": 1046806,
      "postDate": "2020-10-12T02:54:02.390Z",
      "content": "<p>To be honest, it's just the scope of the competition that is difficult. The 920gb of data and insane compute needed, as well as the annoying submission format which has caused numerous errors. My team isn't doing anything spectacular or different to beat the baseline; just a working model with a working submission pipeline. Our current best 0.323 submission doesn't even use TTA yet, just one single solid model.</p>",
      "rawMarkdown": "To be honest, it's just the scope of the competition that is difficult. The 920gb of data and insane compute needed, as well as the annoying submission format which has caused numerous errors. My team isn't doing anything spectacular or different to beat the baseline; just a working model with a working submission pipeline. Our current best 0.323 submission doesn't even use TTA yet, just one single solid model.",
      "votes": 6,
      "replies": [
        {
          "id": 1046826,
          "postDate": "2020-10-12T03:14:42.443Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>\n<p>Agree! A lot of the submissions I have tried are also throwing errors; and its hours for the solution to run. </p>",
          "rawMarkdown": "Thanks @stanleyjzheng \n\nAgree! A lot of the submissions I have tried are also throwing errors; and its hours for the solution to run. ",
          "votes": 2
        },
        {
          "id": 1046838,
          "postDate": "2020-10-12T03:21:36.100Z",
          "content": "<p>Yep! High barrier to entry means that even a baseline model like ours can get Silver. Best of luck, and stay persistent through those errors. </p>",
          "rawMarkdown": "Yep! High barrier to entry means that even a baseline model like ours can get Silver. Best of luck, and stay persistent through those errors. ",
          "votes": 4
        },
        {
          "id": 1046855,
          "postDate": "2020-10-12T03:53:29.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> can you tell us how many days or time it took to train a model on the full data? This will give us all idea of how much time it can take. One more question is that how many epochs are required from your point of view to get a stable score? I hope you do not mind sharing this. Thank you.</p>",
          "rawMarkdown": "@stanleyjzheng can you tell us how many days or time it took to train a model on the full data? This will give us all idea of how much time it can take. One more question is that how many epochs are required from your point of view to get a stable score? I hope you do not mind sharing this. Thank you.",
          "votes": 2
        },
        {
          "id": 1046860,
          "postDate": "2020-10-12T04:02:01.600Z",
          "content": "<p>I use tfrecords I made from Ian Pan's 256x256, which was quite painful. I think my team is a bit unique in that we are using tensorflow, everyone else seems like they're using pytorch.</p>\n<p>Training done fully in Colab Pro TPU. Not looking to share too much about epochs, but I can say that we can train a 5 fold model in less than 16 hours on Colab's TPUv2. Existing public submission notebooks have been helpful, but I wouldn't use a ton of code from any of them; many have some glaring mistakes or oversights. Our model is absolutely nothing special; no hparam tuning or TTA or anything, just a baseline, and it gets 0.323. It feels a bit like we just got out of the hard part, and now we can try our ideas and start iterating, so I look forwards to seeing our score increase in the coming week. </p>\n<p>Best of luck.</p>",
          "rawMarkdown": "I use tfrecords I made from Ian Pan's 256x256, which was quite painful. I think my team is a bit unique in that we are using tensorflow, everyone else seems like they're using pytorch.\n\nTraining done fully in Colab Pro TPU. Not looking to share too much about epochs, but I can say that we can train a 5 fold model in less than 16 hours on Colab's TPUv2. Existing public submission notebooks have been helpful, but I wouldn't use a ton of code from any of them; many have some glaring mistakes or oversights. Our model is absolutely nothing special; no hparam tuning or TTA or anything, just a baseline, and it gets 0.323. It feels a bit like we just got out of the hard part, and now we can try our ideas and start iterating, so I look forwards to seeing our score increase in the coming week. \n\nBest of luck.",
          "votes": 5
        },
        {
          "id": 1046862,
          "postDate": "2020-10-12T04:04:11.700Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a>, it gave some idea that 16 hours on TPUv2 is enough to get some good score. Though, it is a hard competition because of a huge dataset. </p>",
          "rawMarkdown": "Thanks @stanleyjzheng, it gave some idea that 16 hours on TPUv2 is enough to get some good score. Though, it is a hard competition because of a huge dataset. ",
          "votes": 2
        },
        {
          "id": 1046864,
          "postDate": "2020-10-12T04:07:23.243Z",
          "content": "<p>No problem, definitely hard due to the massive data and difficult submission/image format, but it's a double edged sword; at the moment, any model can get a medal. </p>",
          "rawMarkdown": "No problem, definitely hard due to the massive data and difficult submission/image format, but it's a double edged sword; at the moment, any model can get a medal. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 1047499,
      "postDate": "2020-10-12T17:02:14.313Z",
      "content": "<p>Hallo, there is an explanation for this issue. Concerning the 0.325 score: there is a public notebooks which provides this score. I mean, <a href=\"https://www.kaggle.com/osciiart/baseline-with-no-image\" target=\"_blank\">this one</a>. It seems a lot of people were inspired by it, forked and tried to improve it. The notebook is brilliant as it gets a very good score using a basic linear algorithm.</p>",
      "rawMarkdown": "Hallo, there is an explanation for this issue. Concerning the 0.325 score: there is a public notebooks which provides this score. I mean, [this one](https://www.kaggle.com/osciiart/baseline-with-no-image). It seems a lot of people were inspired by it, forked and tried to improve it. The notebook is brilliant as it gets a very good score using a basic linear algorithm.",
      "votes": 1,
      "replies": [
        {
          "id": 1048358,
          "postDate": "2020-10-13T12:10:13.650Z",
          "content": "<p>agree great work by <a href=\"https://www.kaggle.com/OsciiArt\" target=\"_blank\">@OsciiArt</a><br>\nI tried some improvements and many did not work… very good and clean work.. impressive use of image level evaluation metric!!<br>\nHoping to improve on this, and get a medal…</p>",
          "rawMarkdown": "agree great work by @OsciiArt\nI tried some improvements and many did not work... very good and clean work.. impressive use of image level evaluation metric!!\nHoping to improve on this, and get a medal...",
          "votes": -1
        },
        {
          "id": 1048379,
          "postDate": "2020-10-13T12:39:37.060Z",
          "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> : Indeed. I did not choose to fork and compete with it, even if I took some ideas from it. I tried some improvements on it and I shall make a submission only if I get a better score than the original one 0.325. PS: You shall have a medal with that score, most probably.</p>",
          "rawMarkdown": "@kmldas : Indeed. I did not choose to fork and compete with it, even if I took some ideas from it. I tried some improvements on it and I shall make a submission only if I get a better score than the original one 0.325. PS: You shall have a medal with that score, most probably.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1046466,
      "postDate": "2020-10-11T17:20:07.073Z",
      "content": "<p>Hi All,<br>\nLooks like  a difficult competition… Only 39 participants/teams have beaten public  baseline of 0.325</p>\n<p>Those who have, looking forward to your suggestions on approach and next steps to beat the baseline</p>\n<p>As last 2 weeks, not looking for code but suggestions on what would help to improve and also retain the place post shake up!! <br>\nAs private dataset is 3x of public LB dataset, don't want to just be up on public one!</p>\n<p>All suggestions and approaches are welcome and highly appreciated!!</p>",
      "rawMarkdown": "Hi All,\nLooks like  a difficult competition... Only 39 participants/teams have beaten public  baseline of 0.325\n\nThose who have, looking forward to your suggestions on approach and next steps to beat the baseline\n\nAs last 2 weeks, not looking for code but suggestions on what would help to improve and also retain the place post shake up!! \nAs private dataset is 3x of public LB dataset, don't want to just be up on public one!\n\nAll suggestions and approaches are welcome and highly appreciated!!\n\n"
    }
  ],
  "comments": [
    {
      "id": 1046806,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2020-10-12T02:54:02.390000",
      "content": "<p>To be honest, it's just the scope of the competition that is difficult. The 920gb of data and insane compute needed, as well as the annoying submission format which has caused numerous errors. My team isn't doing anything spectacular or different to beat the baseline; just a working model with a working submission pipeline. Our current best 0.323 submission doesn't even use TTA yet, just one single solid model.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1046826,
          "author_name": "Kamal Das",
          "author_url": "",
          "post_date": "2020-10-12T03:14:42.443000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>\n<p>Agree! A lot of the submissions I have tried are also throwing errors; and its hours for the solution to run. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1046838,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-12T03:21:36.100000",
          "content": "<p>Yep! High barrier to entry means that even a baseline model like ours can get Silver. Best of luck, and stay persistent through those errors. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1046855,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2020-10-12T03:53:29.100000",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> can you tell us how many days or time it took to train a model on the full data? This will give us all idea of how much time it can take. One more question is that how many epochs are required from your point of view to get a stable score? I hope you do not mind sharing this. Thank you.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1046860,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-12T04:02:01.600000",
          "content": "<p>I use tfrecords I made from Ian Pan's 256x256, which was quite painful. I think my team is a bit unique in that we are using tensorflow, everyone else seems like they're using pytorch.</p>\n<p>Training done fully in Colab Pro TPU. Not looking to share too much about epochs, but I can say that we can train a 5 fold model in less than 16 hours on Colab's TPUv2. Existing public submission notebooks have been helpful, but I wouldn't use a ton of code from any of them; many have some glaring mistakes or oversights. Our model is absolutely nothing special; no hparam tuning or TTA or anything, just a baseline, and it gets 0.323. It feels a bit like we just got out of the hard part, and now we can try our ideas and start iterating, so I look forwards to seeing our score increase in the coming week. </p>\n<p>Best of luck.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1046862,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2020-10-12T04:04:11.700000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a>, it gave some idea that 16 hours on TPUv2 is enough to get some good score. Though, it is a hard competition because of a huge dataset. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1046864,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-12T04:07:23.243000",
          "content": "<p>No problem, definitely hard due to the massive data and difficult submission/image format, but it's a double edged sword; at the moment, any model can get a medal. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1047499,
      "author_name": "Catadanna",
      "author_url": "",
      "post_date": "2020-10-12T17:02:14.313000",
      "content": "<p>Hallo, there is an explanation for this issue. Concerning the 0.325 score: there is a public notebooks which provides this score. I mean, <a href=\"https://www.kaggle.com/osciiart/baseline-with-no-image\" target=\"_blank\">this one</a>. It seems a lot of people were inspired by it, forked and tried to improve it. The notebook is brilliant as it gets a very good score using a basic linear algorithm.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1048358,
          "author_name": "Kamal Das",
          "author_url": "",
          "post_date": "2020-10-13T12:10:13.650000",
          "content": "<p>agree great work by <a href=\"https://www.kaggle.com/OsciiArt\" target=\"_blank\">@OsciiArt</a><br>\nI tried some improvements and many did not work… very good and clean work.. impressive use of image level evaluation metric!!<br>\nHoping to improve on this, and get a medal…</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1048379,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2020-10-13T12:39:37.060000",
          "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> : Indeed. I did not choose to fork and compete with it, even if I took some ideas from it. I tried some improvements on it and I shall make a submission only if I get a better score than the original one 0.325. PS: You shall have a medal with that score, most probably.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1046806": "To be honest, it's just the scope of the competition that is difficult. The 920gb of data and insane compute needed, as well as the annoying submission format which has caused numerous errors. My team isn't doing anything spectacular or different to beat the baseline; just a working model with a working submission pipeline. Our current best 0.323 submission doesn't even use TTA yet, just one single solid model.",
    "1047499": "Hallo, there is an explanation for this issue. Concerning the 0.325 score: there is a public notebooks which provides this score. I mean, [this one](https://www.kaggle.com/osciiart/baseline-with-no-image). It seems a lot of people were inspired by it, forked and tried to improve it. The notebook is brilliant as it gets a very good score using a basic linear algorithm.",
    "1046466": "Hi All,\nLooks like  a difficult competition... Only 39 participants/teams have beaten public  baseline of 0.325\n\nThose who have, looking forward to your suggestions on approach and next steps to beat the baseline\n\nAs last 2 weeks, not looking for code but suggestions on what would help to improve and also retain the place post shake up!! \nAs private dataset is 3x of public LB dataset, don't want to just be up on public one!\n\nAll suggestions and approaches are welcome and highly appreciated!!\n\n"
  }
}