{
  "id": 375498,
  "title": "Loss Function Idea",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/375498",
  "author_name": "AleNic",
  "post_date": "2023-01-02T01:28:30.050000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm guessing if a change of loss can be useful for this problem. If the cancer outcome is given by breasts, then for every tuple (patient_id, laterality) we have the value of cancer 1 or 0.<br>\nFor every group there are from 2 to 8 images of different views ('CC', 'MLO', 'ML', 'LM', 'AT', 'LMO').<br>\nIf the biopsy revelas the cancer, maybe could be difficult to discover cancer from images (this problem is anti-causal ?).<br>\nMy questions is:</p>\n<p><strong>Q: Having a value of cancer=1 means that all the images reveals the presence of a cancer ?</strong><br>\nIf the answer is NO it means that at least 1 image could reveals the cancer.<br>\nIf this is the case, maybe can be useful to modify the loss as follow:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1532087%2F204c1acc116bfcd12ced36a8fa49c8ee%2Fequation.png?generation=1672622358143783&amp;alt=media\" alt=\"\"></p>\n<p>where:</p>\n<ul>\n<li>w_1 is the positive class weight</li>\n<li>g is the number of groups (patient_id, laterality) inside the batch</li>\n<li>G_j is the set of indices of the j-th group</li>\n<li>y_i are the labels (cancer=1 or 0)</li>\n<li>p_i are the output probability of the model:   p_i = sigmoid(model(x_i))</li>\n<li>B is the sizeo of batch</li>\n</ul>\n<p>What do you think ?</p>",
  "messages": [
    {
      "id": 2082819,
      "postDate": "2023-01-02T01:28:30.050Z",
      "content": "<p>I'm guessing if a change of loss can be useful for this problem. If the cancer outcome is given by breasts, then for every tuple (patient_id, laterality) we have the value of cancer 1 or 0.<br>\nFor every group there are from 2 to 8 images of different views ('CC', 'MLO', 'ML', 'LM', 'AT', 'LMO').<br>\nIf the biopsy revelas the cancer, maybe could be difficult to discover cancer from images (this problem is anti-causal ?).<br>\nMy questions is:</p>\n<p><strong>Q: Having a value of cancer=1 means that all the images reveals the presence of a cancer ?</strong><br>\nIf the answer is NO it means that at least 1 image could reveals the cancer.<br>\nIf this is the case, maybe can be useful to modify the loss as follow:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1532087%2F204c1acc116bfcd12ced36a8fa49c8ee%2Fequation.png?generation=1672622358143783&amp;alt=media\" alt=\"\"></p>\n<p>where:</p>\n<ul>\n<li>w_1 is the positive class weight</li>\n<li>g is the number of groups (patient_id, laterality) inside the batch</li>\n<li>G_j is the set of indices of the j-th group</li>\n<li>y_i are the labels (cancer=1 or 0)</li>\n<li>p_i are the output probability of the model:   p_i = sigmoid(model(x_i))</li>\n<li>B is the sizeo of batch</li>\n</ul>\n<p>What do you think ?</p>",
      "rawMarkdown": "I'm guessing if a change of loss can be useful for this problem. If the cancer outcome is given by breasts, then for every tuple (patient_id, laterality) we have the value of cancer 1 or 0.\nFor every group there are from 2 to 8 images of different views ('CC', 'MLO', 'ML', 'LM', 'AT', 'LMO').\nIf the biopsy revelas the cancer, maybe could be difficult to discover cancer from images (this problem is anti-causal ?).\nMy questions is:\n\n**Q: Having a value of cancer=1 means that all the images reveals the presence of a cancer ?**\nIf the answer is NO it means that at least 1 image could reveals the cancer.\nIf this is the case, maybe can be useful to modify the loss as follow:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1532087%2F204c1acc116bfcd12ced36a8fa49c8ee%2Fequation.png?generation=1672622358143783&alt=media)\n\nwhere:\n* w_1 is the positive class weight\n* g is the number of groups (patient_id, laterality) inside the batch\n* G_j is the set of indices of the j-th group\n* y_i are the labels (cancer=1 or 0)\n* p_i are the output probability of the model:   p_i = sigmoid(model(x_i))\n* B is the sizeo of batch\n\nWhat do you think ?",
      "votes": 3
    },
    {
      "id": 2096826,
      "postDate": "2023-01-12T09:28:10.427Z",
      "content": "<p>Well.. there is only one way to find out..</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Well.. there is only one way to find out..\n\nThe Devastator.\n",
      "votes": 1
    },
    {
      "id": 2084396,
      "postDate": "2023-01-03T13:14:08.323Z",
      "content": "<p>\"Q: Having a value of cancer=1 means that all the images reveals the presence of a cancer ?\"</p>\n<p>Experts should answer here, but my understanding is this is more or less not the case.  Or at least, all views in theory reveal presence, but some certainly more than others.</p>\n<p>\"If the biopsy reveals the cancer, maybe could be difficult to discover cancer from images (this problem is anti-causal ?).\"</p>\n<p>That's an interesting idea, however I don't think they'd do biopsies randomly, even to generate this data set.  I could be wrong of course.</p>\n<p>If I'm right, it means there must have been something suspicious in the images to provoke a BIRADS score and biopsy, which is what we're trying to detect.</p>\n<p><em>Hmmm… unless, are you saying that it was later scans (not present in this set) which recommended biopsies and found cancer and these are the original scans which showed no signs of cancer?</em>  </p>\n<p>In which case… yikes!  No wonder accuracy is struggling here.  Jeez, that's mean if that's the case.</p>\n<p>It would explain the contest formation however, at least for site 2, kinda.  Site 1 has birads scores though they don't seem standard.</p>\n<blockquote>\n  <p>BIRADS - 0 if the breast required follow-up, 1 if the breast was rated as negative for cancer, and 2 if the breast was rated as normal. Only provided for train.</p>\n</blockquote>\n<p>BIRADS, to my understanding, is split into 6 levels, with 6 being the most likely for cancer.  </p>\n<p>If it is true (at least for site 2), I do have to admit, not a big fan of these sort of things not being explained up front.  Just breeds mistrust and suspicion.  </p>\n<p>It could even still be true for site 1, in that we don't know if biopsies were recommended immediately because of the non standard birads scores.  Or frankly if those birads scores are equated with same images (that seems like a stretch though, no way they'd do that, right?)</p>\n<p>The novo enzymes contest is an example of hosts/kaggle playing cute, which had a non random split on a 50/50 leaderboard.  I mean, gimme a break.</p>\n<p>That all said, I remain hopeful nothing gimmicky is happening here.  Perhaps some radiologists / experts could comment by looking at a few of the labelled cancer images in the train set, in particular for site 2 which doesn't have the birads scores.</p>",
      "rawMarkdown": "\"Q: Having a value of cancer=1 means that all the images reveals the presence of a cancer ?\"\n\nExperts should answer here, but my understanding is this is more or less not the case.  Or at least, all views in theory reveal presence, but some certainly more than others.\n\n\"If the biopsy reveals the cancer, maybe could be difficult to discover cancer from images (this problem is anti-causal ?).\"\n\nThat's an interesting idea, however I don't think they'd do biopsies randomly, even to generate this data set.  I could be wrong of course.\n\nIf I'm right, it means there must have been something suspicious in the images to provoke a BIRADS score and biopsy, which is what we're trying to detect.\n\n*Hmmm... unless, are you saying that it was later scans (not present in this set) which recommended biopsies and found cancer and these are the original scans which showed no signs of cancer?*  \n \nIn which case... yikes!  No wonder accuracy is struggling here.  Jeez, that's mean if that's the case.\n\nIt would explain the contest formation however, at least for site 2, kinda.  Site 1 has birads scores though they don't seem standard.\n\n>BIRADS - 0 if the breast required follow-up, 1 if the breast was rated as negative for cancer, and 2 if the breast was rated as normal. Only provided for train.\n\nBIRADS, to my understanding, is split into 6 levels, with 6 being the most likely for cancer.  \n\nIf it is true (at least for site 2), I do have to admit, not a big fan of these sort of things not being explained up front.  Just breeds mistrust and suspicion.  \n\nIt could even still be true for site 1, in that we don't know if biopsies were recommended immediately because of the non standard birads scores.  Or frankly if those birads scores are equated with same images (that seems like a stretch though, no way they'd do that, right?)\n\nThe novo enzymes contest is an example of hosts/kaggle playing cute, which had a non random split on a 50/50 leaderboard.  I mean, gimme a break.\n\nThat all said, I remain hopeful nothing gimmicky is happening here.  Perhaps some radiologists / experts could comment by looking at a few of the labelled cancer images in the train set, in particular for site 2 which doesn't have the birads scores.\n\n\n\n\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2096826,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2023-01-12T09:28:10.427000",
      "content": "<p>Well.. there is only one way to find out..</p>\n<p>The Devastator.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2084396,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2023-01-03T13:14:08.323000",
      "content": "<p>\"Q: Having a value of cancer=1 means that all the images reveals the presence of a cancer ?\"</p>\n<p>Experts should answer here, but my understanding is this is more or less not the case.  Or at least, all views in theory reveal presence, but some certainly more than others.</p>\n<p>\"If the biopsy reveals the cancer, maybe could be difficult to discover cancer from images (this problem is anti-causal ?).\"</p>\n<p>That's an interesting idea, however I don't think they'd do biopsies randomly, even to generate this data set.  I could be wrong of course.</p>\n<p>If I'm right, it means there must have been something suspicious in the images to provoke a BIRADS score and biopsy, which is what we're trying to detect.</p>\n<p><em>Hmmm… unless, are you saying that it was later scans (not present in this set) which recommended biopsies and found cancer and these are the original scans which showed no signs of cancer?</em>  </p>\n<p>In which case… yikes!  No wonder accuracy is struggling here.  Jeez, that's mean if that's the case.</p>\n<p>It would explain the contest formation however, at least for site 2, kinda.  Site 1 has birads scores though they don't seem standard.</p>\n<blockquote>\n  <p>BIRADS - 0 if the breast required follow-up, 1 if the breast was rated as negative for cancer, and 2 if the breast was rated as normal. Only provided for train.</p>\n</blockquote>\n<p>BIRADS, to my understanding, is split into 6 levels, with 6 being the most likely for cancer.  </p>\n<p>If it is true (at least for site 2), I do have to admit, not a big fan of these sort of things not being explained up front.  Just breeds mistrust and suspicion.  </p>\n<p>It could even still be true for site 1, in that we don't know if biopsies were recommended immediately because of the non standard birads scores.  Or frankly if those birads scores are equated with same images (that seems like a stretch though, no way they'd do that, right?)</p>\n<p>The novo enzymes contest is an example of hosts/kaggle playing cute, which had a non random split on a 50/50 leaderboard.  I mean, gimme a break.</p>\n<p>That all said, I remain hopeful nothing gimmicky is happening here.  Perhaps some radiologists / experts could comment by looking at a few of the labelled cancer images in the train set, in particular for site 2 which doesn't have the birads scores.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2082819": "I'm guessing if a change of loss can be useful for this problem. If the cancer outcome is given by breasts, then for every tuple (patient_id, laterality) we have the value of cancer 1 or 0.\nFor every group there are from 2 to 8 images of different views ('CC', 'MLO', 'ML', 'LM', 'AT', 'LMO').\nIf the biopsy revelas the cancer, maybe could be difficult to discover cancer from images (this problem is anti-causal ?).\nMy questions is:\n\n**Q: Having a value of cancer=1 means that all the images reveals the presence of a cancer ?**\nIf the answer is NO it means that at least 1 image could reveals the cancer.\nIf this is the case, maybe can be useful to modify the loss as follow:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1532087%2F204c1acc116bfcd12ced36a8fa49c8ee%2Fequation.png?generation=1672622358143783&alt=media)\n\nwhere:\n* w_1 is the positive class weight\n* g is the number of groups (patient_id, laterality) inside the batch\n* G_j is the set of indices of the j-th group\n* y_i are the labels (cancer=1 or 0)\n* p_i are the output probability of the model:   p_i = sigmoid(model(x_i))\n* B is the sizeo of batch\n\nWhat do you think ?",
    "2096826": "Well.. there is only one way to find out..\n\nThe Devastator.\n",
    "2084396": "\"Q: Having a value of cancer=1 means that all the images reveals the presence of a cancer ?\"\n\nExperts should answer here, but my understanding is this is more or less not the case.  Or at least, all views in theory reveal presence, but some certainly more than others.\n\n\"If the biopsy reveals the cancer, maybe could be difficult to discover cancer from images (this problem is anti-causal ?).\"\n\nThat's an interesting idea, however I don't think they'd do biopsies randomly, even to generate this data set.  I could be wrong of course.\n\nIf I'm right, it means there must have been something suspicious in the images to provoke a BIRADS score and biopsy, which is what we're trying to detect.\n\n*Hmmm... unless, are you saying that it was later scans (not present in this set) which recommended biopsies and found cancer and these are the original scans which showed no signs of cancer?*  \n \nIn which case... yikes!  No wonder accuracy is struggling here.  Jeez, that's mean if that's the case.\n\nIt would explain the contest formation however, at least for site 2, kinda.  Site 1 has birads scores though they don't seem standard.\n\n>BIRADS - 0 if the breast required follow-up, 1 if the breast was rated as negative for cancer, and 2 if the breast was rated as normal. Only provided for train.\n\nBIRADS, to my understanding, is split into 6 levels, with 6 being the most likely for cancer.  \n\nIf it is true (at least for site 2), I do have to admit, not a big fan of these sort of things not being explained up front.  Just breeds mistrust and suspicion.  \n\nIt could even still be true for site 1, in that we don't know if biopsies were recommended immediately because of the non standard birads scores.  Or frankly if those birads scores are equated with same images (that seems like a stretch though, no way they'd do that, right?)\n\nThe novo enzymes contest is an example of hosts/kaggle playing cute, which had a non random split on a 50/50 leaderboard.  I mean, gimme a break.\n\nThat all said, I remain hopeful nothing gimmicky is happening here.  Perhaps some radiologists / experts could comment by looking at a few of the labelled cancer images in the train set, in particular for site 2 which doesn't have the birads scores.\n\n\n\n\n"
  }
}