{
  "id": 438451,
  "title": "Is submission =...detection/sample_submission.csv' necessary for scoring?",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/438451",
  "author_name": "Michalina Hulak",
  "post_date": "2023-09-11T08:59:54.135000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm stuck during submission and realy can't understand what's going on :P<br>\nThere is sample notebook (i belive that most of us saw it): <a href=\"https://www.kaggle.com/code/mirenaborisova/rsna-0-66-lb\" target=\"_blank\">https://www.kaggle.com/code/mirenaborisova/rsna-0-66-lb</a></p>\n<p>and I don't understand why while i change submision path to my I have a scoring error?(my submission looks exactly like sample submission: the same colmn's names, id_patients). Is it the goal to find a properly weight based on train[features].mean? I'm really confused :P </p>\n<pre><code>\n\n\n\n\n\n\n\n\n\n\n\n\n\nsubmission = pd.read_csv()\nsubmission[features] = train[features].mean().tolist()\n\n feature  features:\n     feature.split()[] ==   feature == :\n        submission[feature] *= \n     feature.split()[] == :\n        submission[feature] *= \n     feature.split()[] ==   feature != :\n        submission[feature] *= \n\n\nsubmission.to_csv(, index=)\n</code></pre>",
  "messages": [
    {
      "id": 2433938,
      "postDate": "2023-09-12T01:43:34.877Z",
      "content": "<p>If I had to guess, it probably has something to do with whatever you have \"my_path\" set to in the pd.read_csv line at the top of your code. In any case, the notebook uses clever mathematics, and as the author said, some experimentation, to get a number that can be inputted into all the submission columns as more of a placeholder. So, instead of having a model look at the dicom files in the 'train' dicom file folders, it just takes a calculation of the mean of all values in that column, for bowel_injury maybe it's .22, and since the calculation is *=3, .22 * 3 = .66. So the submission file has .66 for e v e r y test patient_id, which isn't really helpful for real data and real calculations. The notebook is more of a baseline, you can use the numbers listed to make calculations about what the public test set values are like, and perhaps create a model keeping those numbers in mind, but you definitely should not try to have the notebook be your actual submission notebook. For that, refer to the pinned KerasCV starter notebook under the 'code' tab in this competition. Hope that helps, good luck!</p>",
      "rawMarkdown": "If I had to guess, it probably has something to do with whatever you have \"my_path\" set to in the pd.read_csv line at the top of your code. In any case, the notebook uses clever mathematics, and as the author said, some experimentation, to get a number that can be inputted into all the submission columns as more of a placeholder. So, instead of having a model look at the dicom files in the 'train' dicom file folders, it just takes a calculation of the mean of all values in that column, for bowel_injury maybe it's .22, and since the calculation is *=3, .22 * 3 = .66. So the submission file has .66 for e v e r y test patient_id, which isn't really helpful for real data and real calculations. The notebook is more of a baseline, you can use the numbers listed to make calculations about what the public test set values are like, and perhaps create a model keeping those numbers in mind, but you definitely should not try to have the notebook be your actual submission notebook. For that, refer to the pinned KerasCV starter notebook under the 'code' tab in this competition. Hope that helps, good luck!",
      "votes": 1,
      "replies": [
        {
          "id": 2436785,
          "postDate": "2023-09-13T18:52:58.590Z",
          "content": "<p>good, thanks a lot! :)</p>",
          "rawMarkdown": "good, thanks a lot! :)"
        }
      ]
    },
    {
      "id": 2432987,
      "postDate": "2023-09-11T08:59:54.137Z",
      "content": "<p>I'm stuck during submission and realy can't understand what's going on :P<br>\nThere is sample notebook (i belive that most of us saw it): <a href=\"https://www.kaggle.com/code/mirenaborisova/rsna-0-66-lb\" target=\"_blank\">https://www.kaggle.com/code/mirenaborisova/rsna-0-66-lb</a></p>\n<p>and I don't understand why while i change submision path to my I have a scoring error?(my submission looks exactly like sample submission: the same colmn's names, id_patients). Is it the goal to find a properly weight based on train[features].mean? I'm really confused :P </p>\n<pre><code>\n\n\n\n\n\n\n\n\n\n\n\n\n\nsubmission = pd.read_csv()\nsubmission[features] = train[features].mean().tolist()\n\n feature  features:\n     feature.split()[] ==   feature == :\n        submission[feature] *= \n     feature.split()[] == :\n        submission[feature] *= \n     feature.split()[] ==   feature != :\n        submission[feature] *= \n\n\nsubmission.to_csv(, index=)\n</code></pre>",
      "rawMarkdown": "I'm stuck during submission and realy can't understand what's going on :P\nThere is sample notebook (i belive that most of us saw it): https://www.kaggle.com/code/mirenaborisova/rsna-0-66-lb\n\nand I don't understand why while i change submision path to my I have a scoring error?(my submission looks exactly like sample submission: the same colmn's names, id_patients). Is it the goal to find a properly weight based on train[features].mean? I'm really confused :P \n\n```python\n# submission = pd.read_csv(my_path)\n# submission[features] = train[features].mean().tolist()\n\n# for feature in features:\n#     if feature.split('_')[1] == 'low' or feature == 'bowel_injury':\n#         submission[feature] *= 4\n#     elif feature.split('_')[1] == 'high':\n#         submission[feature] *= 6\n#     elif feature.split('_')[1] == 'injury' and feature != 'bowel_injury':\n#         submission[feature] *= 28\n        \n\n# submission.to_csv('submission.csv', index=False)\n\nsubmission = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\nsubmission[features] = train[features].mean().tolist()\n\nfor feature in features:\n    if feature.split('_')[1] == 'low' or feature == 'bowel_injury':\n        submission[feature] *= 3\n    elif feature.split('_')[1] == 'high':\n        submission[feature] *= 8\n    elif feature.split('_')[1] == 'injury' and feature != 'bowel_injury':\n        submission[feature] *= 27\n        \n\nsubmission.to_csv('submission.csv', index=False)\n```",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2433938,
      "author_name": "Art. Berz.",
      "author_url": "",
      "post_date": "2023-09-12T01:43:34.877000",
      "content": "<p>If I had to guess, it probably has something to do with whatever you have \"my_path\" set to in the pd.read_csv line at the top of your code. In any case, the notebook uses clever mathematics, and as the author said, some experimentation, to get a number that can be inputted into all the submission columns as more of a placeholder. So, instead of having a model look at the dicom files in the 'train' dicom file folders, it just takes a calculation of the mean of all values in that column, for bowel_injury maybe it's .22, and since the calculation is *=3, .22 * 3 = .66. So the submission file has .66 for e v e r y test patient_id, which isn't really helpful for real data and real calculations. The notebook is more of a baseline, you can use the numbers listed to make calculations about what the public test set values are like, and perhaps create a model keeping those numbers in mind, but you definitely should not try to have the notebook be your actual submission notebook. For that, refer to the pinned KerasCV starter notebook under the 'code' tab in this competition. Hope that helps, good luck!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2436785,
          "author_name": "Michalina Hulak",
          "author_url": "",
          "post_date": "2023-09-13T18:52:58.590000",
          "content": "<p>good, thanks a lot! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2433938": "If I had to guess, it probably has something to do with whatever you have \"my_path\" set to in the pd.read_csv line at the top of your code. In any case, the notebook uses clever mathematics, and as the author said, some experimentation, to get a number that can be inputted into all the submission columns as more of a placeholder. So, instead of having a model look at the dicom files in the 'train' dicom file folders, it just takes a calculation of the mean of all values in that column, for bowel_injury maybe it's .22, and since the calculation is *=3, .22 * 3 = .66. So the submission file has .66 for e v e r y test patient_id, which isn't really helpful for real data and real calculations. The notebook is more of a baseline, you can use the numbers listed to make calculations about what the public test set values are like, and perhaps create a model keeping those numbers in mind, but you definitely should not try to have the notebook be your actual submission notebook. For that, refer to the pinned KerasCV starter notebook under the 'code' tab in this competition. Hope that helps, good luck!",
    "2432987": "I'm stuck during submission and realy can't understand what's going on :P\nThere is sample notebook (i belive that most of us saw it): https://www.kaggle.com/code/mirenaborisova/rsna-0-66-lb\n\nand I don't understand why while i change submision path to my I have a scoring error?(my submission looks exactly like sample submission: the same colmn's names, id_patients). Is it the goal to find a properly weight based on train[features].mean? I'm really confused :P \n\n```python\n# submission = pd.read_csv(my_path)\n# submission[features] = train[features].mean().tolist()\n\n# for feature in features:\n#     if feature.split('_')[1] == 'low' or feature == 'bowel_injury':\n#         submission[feature] *= 4\n#     elif feature.split('_')[1] == 'high':\n#         submission[feature] *= 6\n#     elif feature.split('_')[1] == 'injury' and feature != 'bowel_injury':\n#         submission[feature] *= 28\n        \n\n# submission.to_csv('submission.csv', index=False)\n\nsubmission = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\nsubmission[features] = train[features].mean().tolist()\n\nfor feature in features:\n    if feature.split('_')[1] == 'low' or feature == 'bowel_injury':\n        submission[feature] *= 3\n    elif feature.split('_')[1] == 'high':\n        submission[feature] *= 8\n    elif feature.split('_')[1] == 'injury' and feature != 'bowel_injury':\n        submission[feature] *= 27\n        \n\nsubmission.to_csv('submission.csv', index=False)\n```"
  }
}