{
  "id": 192607,
  "title": "Is it Necessary to Use All train data",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/192607",
  "author_name": "Jaideep",
  "post_date": "2020-10-22T10:58:05.034000",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I think data is huge  than any other competition. </p>\n<p>Considering the Available resources/storage  for many it would be difficult to train with whole data. <br>\n,is it possible to get good lb score with subset of data ?</p>\n<p>I tried using subset but loss remain miserably poor at 0.400 range..<br>\nPlease let me know if we need to make use of whole data to get reduce the loss down further</p>",
  "messages": [
    {
      "id": 1057086,
      "postDate": "2020-10-22T10:58:05.033Z",
      "content": "<p>I think data is huge  than any other competition. </p>\n<p>Considering the Available resources/storage  for many it would be difficult to train with whole data. <br>\n,is it possible to get good lb score with subset of data ?</p>\n<p>I tried using subset but loss remain miserably poor at 0.400 range..<br>\nPlease let me know if we need to make use of whole data to get reduce the loss down further</p>",
      "rawMarkdown": "I think data is huge  than any other competition. \n\nConsidering the Available resources/storage  for many it would be difficult to train with whole data. \n,is it possible to get good lb score with subset of data ?\n\nI tried using subset but loss remain miserably poor at 0.400 range..\nPlease let me know if we need to make use of whole data to get reduce the loss down further",
      "votes": 3
    },
    {
      "id": 1057167,
      "postDate": "2020-10-22T12:40:59.103Z",
      "content": "<p>You can get a sub 0.3 score with about 1/2 the train data, although I don't meet all the consistency requirements with that score. </p>",
      "rawMarkdown": "You can get a sub 0.3 score with about 1/2 the train data, although I don't meet all the consistency requirements with that score. ",
      "replies": [
        {
          "id": 1059925,
          "postDate": "2020-10-25T15:38:20.603Z",
          "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> </p>\n<p>Exam loss mentioned in competition is weighted sum , when we take mean of it then should it be for a sum of weighted loss / number of elements in batch ( which is what i think pytorch's bce with logits does ,when reduction='mean' )</p>\n<p>or</p>\n<p>should it be sum (weighted loss) of all labels in exam / number of exams in batch ..</p>\n<p>When i do first one i get too less Exam loss.. will it be  right mean.. does competition metric does it that way it uses second method for exam loss</p>",
          "rawMarkdown": "@richardepstein \n\nExam loss mentioned in competition is weighted sum , when we take mean of it then should it be for a sum of weighted loss / number of elements in batch ( which is what i think pytorch's bce with logits does ,when reduction='mean' )\n\nor\n\nshould it be sum (weighted loss) of all labels in exam / number of exams in batch ..\n\nWhen i do first one i get too less Exam loss.. will it be  right mean.. does competition metric does it that way it uses second method for exam loss",
          "replies": [
            {
              "id": 1060027,
              "postDate": "2020-10-25T17:36:23.573Z",
              "content": "<p>Jaideep,</p>\n<p>Sorry, I cannot help. I don't 100% understand the metric.</p>\n<p>-Rich</p>",
              "rawMarkdown": "Jaideep,\n\nSorry, I cannot help. I don't 100% understand the metric.\n\n-Rich"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1057167,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-10-22T12:40:59.103000",
      "content": "<p>You can get a sub 0.3 score with about 1/2 the train data, although I don't meet all the consistency requirements with that score. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1059925,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-25T15:38:20.603000",
          "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> </p>\n<p>Exam loss mentioned in competition is weighted sum , when we take mean of it then should it be for a sum of weighted loss / number of elements in batch ( which is what i think pytorch's bce with logits does ,when reduction='mean' )</p>\n<p>or</p>\n<p>should it be sum (weighted loss) of all labels in exam / number of exams in batch ..</p>\n<p>When i do first one i get too less Exam loss.. will it be  right mean.. does competition metric does it that way it uses second method for exam loss</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1060027,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-10-25T17:36:23.573000",
              "content": "<p>Jaideep,</p>\n<p>Sorry, I cannot help. I don't 100% understand the metric.</p>\n<p>-Rich</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1057086": "I think data is huge  than any other competition. \n\nConsidering the Available resources/storage  for many it would be difficult to train with whole data. \n,is it possible to get good lb score with subset of data ?\n\nI tried using subset but loss remain miserably poor at 0.400 range..\nPlease let me know if we need to make use of whole data to get reduce the loss down further",
    "1057167": "You can get a sub 0.3 score with about 1/2 the train data, although I don't meet all the consistency requirements with that score. "
  }
}