{
  "id": 346201,
  "title": "implications of output for each patient_id and not per image",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/346201",
  "author_name": "Abhishek Vijayan",
  "post_date": "2022-08-18T11:51:31.225000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The data page mentions that <code>Note in particular that you should make one prediction per patient_id, not per image_id.</code></p>\n<ol>\n<li><p>So if we were to create a submission file from the train data, which has multiple images per patient, we are expected to output prob(CE) = (number of CE images) / (total number of images for one patient) ?</p></li>\n<li><p>The evaluation page while explaining the loss function describes in terms of image_id. <code>pij is the predicted probability that image i belongs to class j.</code> and also divides by number of images in the class set. This seems to be a metric suitable for per image_id probability. </p></li>\n</ol>\n<p>Can someone please shed some light on these 2 aspects ?</p>",
  "messages": [
    {
      "id": 1904681,
      "postDate": "2022-08-18T11:51:31.227Z",
      "content": "<p>The data page mentions that <code>Note in particular that you should make one prediction per patient_id, not per image_id.</code></p>\n<ol>\n<li><p>So if we were to create a submission file from the train data, which has multiple images per patient, we are expected to output prob(CE) = (number of CE images) / (total number of images for one patient) ?</p></li>\n<li><p>The evaluation page while explaining the loss function describes in terms of image_id. <code>pij is the predicted probability that image i belongs to class j.</code> and also divides by number of images in the class set. This seems to be a metric suitable for per image_id probability. </p></li>\n</ol>\n<p>Can someone please shed some light on these 2 aspects ?</p>",
      "rawMarkdown": "The data page mentions that `Note in particular that you should make one prediction per patient_id, not per image_id.`\n\n1. So if we were to create a submission file from the train data, which has multiple images per patient, we are expected to output prob(CE) = (number of CE images) / (total number of images for one patient) ?\n\n2. The evaluation page while explaining the loss function describes in terms of image_id. `pij is the predicted probability that image i belongs to class j.` and also divides by number of images in the class set. This seems to be a metric suitable for per image_id probability. \n\nCan someone please shed some light on these 2 aspects ?",
      "votes": 1
    },
    {
      "id": 1906472,
      "postDate": "2022-08-19T23:02:49.770Z",
      "content": "<p>I think for #2, the probabilities for the other classes don't matter, since the y_ij will zero-out those probabilities.</p>\n<p>Regarding #1, I haven't done submissions myself, but just a guess is that we need to generate a single probability per class somehow. The methodology I think is open to competitors. A simple example would be to just average the probabilities across images.</p>\n<p>Edit: I gave a longer explanation of my understanding of the loss function after realizing some things later in another <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/346480\" target=\"_blank\">discussion post</a></p>\n<p>Edit 2: After looking at the loss function, it seems you have a valid point. I think that the probabilities for each image are important for the loss function. I haven't submitted anything yet, but it seems to run counter to what other competitors have discussed in forums. I think I've heard them say that we need a single probability per patient. However, the loss function I think is in fact looking at each probability per image of the true class.</p>",
      "rawMarkdown": "I think for #2, the probabilities for the other classes don't matter, since the y_ij will zero-out those probabilities.\n\nRegarding #1, I haven't done submissions myself, but just a guess is that we need to generate a single probability per class somehow. The methodology I think is open to competitors. A simple example would be to just average the probabilities across images.\n\nEdit: I gave a longer explanation of my understanding of the loss function after realizing some things later in another [discussion post](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/346480)\n\nEdit 2: After looking at the loss function, it seems you have a valid point. I think that the probabilities for each image are important for the loss function. I haven't submitted anything yet, but it seems to run counter to what other competitors have discussed in forums. I think I've heard them say that we need a single probability per patient. However, the loss function I think is in fact looking at each probability per image of the true class.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1906472,
      "author_name": "yqz",
      "author_url": "",
      "post_date": "2022-08-19T23:02:49.770000",
      "content": "<p>I think for #2, the probabilities for the other classes don't matter, since the y_ij will zero-out those probabilities.</p>\n<p>Regarding #1, I haven't done submissions myself, but just a guess is that we need to generate a single probability per class somehow. The methodology I think is open to competitors. A simple example would be to just average the probabilities across images.</p>\n<p>Edit: I gave a longer explanation of my understanding of the loss function after realizing some things later in another <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/346480\" target=\"_blank\">discussion post</a></p>\n<p>Edit 2: After looking at the loss function, it seems you have a valid point. I think that the probabilities for each image are important for the loss function. I haven't submitted anything yet, but it seems to run counter to what other competitors have discussed in forums. I think I've heard them say that we need a single probability per patient. However, the loss function I think is in fact looking at each probability per image of the true class.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1904681": "The data page mentions that `Note in particular that you should make one prediction per patient_id, not per image_id.`\n\n1. So if we were to create a submission file from the train data, which has multiple images per patient, we are expected to output prob(CE) = (number of CE images) / (total number of images for one patient) ?\n\n2. The evaluation page while explaining the loss function describes in terms of image_id. `pij is the predicted probability that image i belongs to class j.` and also divides by number of images in the class set. This seems to be a metric suitable for per image_id probability. \n\nCan someone please shed some light on these 2 aspects ?",
    "1906472": "I think for #2, the probabilities for the other classes don't matter, since the y_ij will zero-out those probabilities.\n\nRegarding #1, I haven't done submissions myself, but just a guess is that we need to generate a single probability per class somehow. The methodology I think is open to competitors. A simple example would be to just average the probabilities across images.\n\nEdit: I gave a longer explanation of my understanding of the loss function after realizing some things later in another [discussion post](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/346480)\n\nEdit 2: After looking at the loss function, it seems you have a valid point. I think that the probabilities for each image are important for the loss function. I haven't submitted anything yet, but it seems to run counter to what other competitors have discussed in forums. I think I've heard them say that we need a single probability per patient. However, the loss function I think is in fact looking at each probability per image of the true class."
  }
}