{
  "id": 371057,
  "title": "difficult_negative_case and BIRADS",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371057",
  "author_name": "delai50",
  "post_date": "2022-12-07T19:03:50.235000",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/hrlmmm\" target=\"_blank\">@hrlmmm</a> mentioned in a comment that when <code>difficult_negative_case=False</code> and <code>BIRADS=0</code> always <code>cancer=1</code>. The complete table is the following:</p>\n<p><code>df_train.groupby(\"patient_id\").tail(1).groupby([\"difficult_negative_case\", \"BIRADS\"])[\"cancer\"].agg([\"mean\",\"size\"])</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n<th>avg cancer</th>\n<th>number of patients</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>difficult_negative_case</td>\n<td>BIRADS</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>False</td>\n<td>0</td>\n<td>1</td>\n<td>125</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>0</td>\n<td>3140</td>\n</tr>\n<tr>\n<td></td>\n<td>2</td>\n<td>0</td>\n<td>424</td>\n</tr>\n<tr>\n<td>True</td>\n<td>0</td>\n<td>0</td>\n<td>1575</td>\n</tr>\n</tbody>\n</table>\n<p>I would say that we must be cautious if we use <code>difficult_negative_case</code> and <code>BIRADS</code> features because our model could learn that relationship (<code>difficult_negative_case=False</code> and <code>BIRADS=0</code> then <code>cancer=1</code>) which doesn't need to be True in the test set. Any thoughts on that?</p>",
  "messages": [
    {
      "id": 2058276,
      "postDate": "2022-12-07T19:03:50.237Z",
      "content": "<p><a href=\"https://www.kaggle.com/hrlmmm\" target=\"_blank\">@hrlmmm</a> mentioned in a comment that when <code>difficult_negative_case=False</code> and <code>BIRADS=0</code> always <code>cancer=1</code>. The complete table is the following:</p>\n<p><code>df_train.groupby(\"patient_id\").tail(1).groupby([\"difficult_negative_case\", \"BIRADS\"])[\"cancer\"].agg([\"mean\",\"size\"])</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n<th>avg cancer</th>\n<th>number of patients</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>difficult_negative_case</td>\n<td>BIRADS</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>False</td>\n<td>0</td>\n<td>1</td>\n<td>125</td>\n</tr>\n<tr>\n<td></td>\n<td>1</td>\n<td>0</td>\n<td>3140</td>\n</tr>\n<tr>\n<td></td>\n<td>2</td>\n<td>0</td>\n<td>424</td>\n</tr>\n<tr>\n<td>True</td>\n<td>0</td>\n<td>0</td>\n<td>1575</td>\n</tr>\n</tbody>\n</table>\n<p>I would say that we must be cautious if we use <code>difficult_negative_case</code> and <code>BIRADS</code> features because our model could learn that relationship (<code>difficult_negative_case=False</code> and <code>BIRADS=0</code> then <code>cancer=1</code>) which doesn't need to be True in the test set. Any thoughts on that?</p>",
      "rawMarkdown": "@hrlmmm mentioned in a comment that when `difficult_negative_case=False` and `BIRADS=0` always `cancer=1`. The complete table is the following:\n\n`df_train.groupby(\"patient_id\").tail(1).groupby([\"difficult_negative_case\", \"BIRADS\"])[\"cancer\"].agg([\"mean\",\"size\"])`\n\n||| avg cancer | number of patients |\n| -- | -- |\n| difficult_negative_case | BIRADS |\n| False | 0 | 1 | 125 |\n|| 1 | 0 | 3140 |\n|| 2 | 0 | 424 |\n| True | 0 | 0 | 1575 |\n\nI would say that we must be cautious if we use `difficult_negative_case` and `BIRADS` features because our model could learn that relationship (`difficult_negative_case=False` and `BIRADS=0` then `cancer=1`) which doesn't need to be True in the test set. Any thoughts on that?",
      "votes": 6
    },
    {
      "id": 2058290,
      "postDate": "2022-12-07T19:17:21.037Z",
      "content": "<p>Some metadata isn't provided for test, check data page (so you can't use difficult_negative_case for example during inference)</p>",
      "rawMarkdown": "Some metadata isn't provided for test, check data page (so you can't use difficult_negative_case for example during inference)",
      "votes": 1,
      "replies": [
        {
          "id": 2058396,
          "postDate": "2022-12-07T22:18:51.513Z",
          "content": "<p>This is the correct answer. We don't provide either <code>difficult_negative_case</code> or <code>BIRADS</code> for the test set, so you shouldn't use either column as a feature.</p>",
          "rawMarkdown": "This is the correct answer. We don't provide either `difficult_negative_case` or `BIRADS` for the test set, so you shouldn't use either column as a feature."
        },
        {
          "id": 2058398,
          "postDate": "2022-12-07T22:32:43.223Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Not sure this is correct, or at least potentially misleading.</p>\n<p>You shouldn't directly use it as a feature, of course, but that doesn't mean you shouldn't use it.  For example, density is a recognized signal for cancer and if one of your models can predict density it might be helpful in the competition - even though density isn't in the test set.</p>\n<p>The same may (or may not, requires experimenting) go for the other features not in the test set such as BIRADS and/or difficult_negative_case.</p>",
          "rawMarkdown": "@sohier \n\nNot sure this is correct, or at least potentially misleading.\n\nYou shouldn't directly use it as a feature, of course, but that doesn't mean you shouldn't use it.  For example, density is a recognized signal for cancer and if one of your models can predict density it might be helpful in the competition - even though density isn't in the test set.\n\nThe same may (or may not, requires experimenting) go for the other features not in the test set such as BIRADS and/or difficult_negative_case.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2058322,
      "postDate": "2022-12-07T19:57:46.480Z",
      "content": "<p>What I said was when difficult_negative_case=False and  BIRADS=0 than cancer=1.😳May be you write a wrong label </p>",
      "rawMarkdown": "What I said was when difficult_negative_case=False and  BIRADS=0 than cancer=1.😳May be you write a wrong label ",
      "replies": [
        {
          "id": 2058325,
          "postDate": "2022-12-07T20:06:23.487Z",
          "content": "<p>Totally right, the table was right but for some reason I write it wrongly, edited!</p>",
          "rawMarkdown": "Totally right, the table was right but for some reason I write it wrongly, edited!"
        }
      ]
    },
    {
      "id": 2058287,
      "postDate": "2022-12-07T19:16:54.300Z",
      "content": "<p>It's worth experimenting I think.  Fine tuning a model such that it can predict BIRADS / difficult negative case, and then seeing if that helps via hill climbing.  Also, fine tuning by predicting density (which is also a signal for cancer).  No promises of course, just ideas for potential diverse models.</p>",
      "rawMarkdown": "It's worth experimenting I think.  Fine tuning a model such that it can predict BIRADS / difficult negative case, and then seeing if that helps via hill climbing.  Also, fine tuning by predicting density (which is also a signal for cancer).  No promises of course, just ideas for potential diverse models.\n\n",
      "replies": [
        {
          "id": 2058299,
          "postDate": "2022-12-07T19:28:04.650Z",
          "content": "<p>Yep, it's worth to experiment. I was thinking about their potential as auxiliary targets (or even metafeatures?)</p>",
          "rawMarkdown": "Yep, it's worth to experiment. I was thinking about their potential as auxiliary targets (or even metafeatures?)",
          "votes": 1
        },
        {
          "id": 2058306,
          "postDate": "2022-12-07T19:31:53.533Z",
          "content": "<p>Do you have a link?  I'm unfamiliar with auxiliary targets / metafeatures in this context.</p>\n<p>Btw, difficult_negative_case=False and BIRADS=0 results in 254 unique patients/breasts, all with cancer. .. all difficult_negative_case==True are negative by definition.</p>\n<pre><code>display(origtrain.query().drop_duplicates([, ]))\ndisplay(origtrain.query().drop_duplicates([, ]).mean())\n</code></pre>\n<p>versus</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/hrlmmm\" target=\"_blank\">@hrlmmm</a> mentioned in a comment that when difficult_negative_case=True and BIRADS=0 always cancer=1. The complete table is the following:</p>\n</blockquote>\n<p>I have to say, it's a very fascinating call out.  Might be worth asking the contest hosters about it.  </p>",
          "rawMarkdown": "Do you have a link?  I'm unfamiliar with auxiliary targets / metafeatures in this context.\n\nBtw, difficult_negative_case=False and BIRADS=0 results in 254 unique patients/breasts, all with cancer. .. all difficult_negative_case==True are negative by definition.\n\n```python\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]))\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]).mean())\n\n```\nversus\n\n> @hrlmmm mentioned in a comment that when difficult_negative_case=True and BIRADS=0 always cancer=1. The complete table is the following:\n\nI have to say, it's a very fascinating call out.  Might be worth asking the contest hosters about it.  "
        },
        {
          "id": 2058351,
          "postDate": "2022-12-07T20:42:46.170Z",
          "content": "<p>For auxiliary targets you can check this <a href=\"https://www.kaggle.com/code/vslaykovsky/train-effnetv2-aux-targets-weighted-loss-thres\" target=\"_blank\">kernel</a> for example</p>",
          "rawMarkdown": "For auxiliary targets you can check this [kernel](https://www.kaggle.com/code/vslaykovsky/train-effnetv2-aux-targets-weighted-loss-thres) for example"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2058290,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2022-12-07T19:17:21.037000",
      "content": "<p>Some metadata isn't provided for test, check data page (so you can't use difficult_negative_case for example during inference)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2058396,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2022-12-07T22:18:51.513000",
          "content": "<p>This is the correct answer. We don't provide either <code>difficult_negative_case</code> or <code>BIRADS</code> for the test set, so you shouldn't use either column as a feature.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2058398,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-07T22:32:43.223000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Not sure this is correct, or at least potentially misleading.</p>\n<p>You shouldn't directly use it as a feature, of course, but that doesn't mean you shouldn't use it.  For example, density is a recognized signal for cancer and if one of your models can predict density it might be helpful in the competition - even though density isn't in the test set.</p>\n<p>The same may (or may not, requires experimenting) go for the other features not in the test set such as BIRADS and/or difficult_negative_case.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2058322,
      "author_name": "hrlmmm",
      "author_url": "",
      "post_date": "2022-12-07T19:57:46.480000",
      "content": "<p>What I said was when difficult_negative_case=False and  BIRADS=0 than cancer=1.😳May be you write a wrong label </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2058325,
          "author_name": "delai50",
          "author_url": "",
          "post_date": "2022-12-07T20:06:23.487000",
          "content": "<p>Totally right, the table was right but for some reason I write it wrongly, edited!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2058287,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-07T19:16:54.300000",
      "content": "<p>It's worth experimenting I think.  Fine tuning a model such that it can predict BIRADS / difficult negative case, and then seeing if that helps via hill climbing.  Also, fine tuning by predicting density (which is also a signal for cancer).  No promises of course, just ideas for potential diverse models.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2058299,
          "author_name": "delai50",
          "author_url": "",
          "post_date": "2022-12-07T19:28:04.650000",
          "content": "<p>Yep, it's worth to experiment. I was thinking about their potential as auxiliary targets (or even metafeatures?)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2058306,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-07T19:31:53.533000",
          "content": "<p>Do you have a link?  I'm unfamiliar with auxiliary targets / metafeatures in this context.</p>\n<p>Btw, difficult_negative_case=False and BIRADS=0 results in 254 unique patients/breasts, all with cancer. .. all difficult_negative_case==True are negative by definition.</p>\n<pre><code>display(origtrain.query().drop_duplicates([, ]))\ndisplay(origtrain.query().drop_duplicates([, ]).mean())\n</code></pre>\n<p>versus</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/hrlmmm\" target=\"_blank\">@hrlmmm</a> mentioned in a comment that when difficult_negative_case=True and BIRADS=0 always cancer=1. The complete table is the following:</p>\n</blockquote>\n<p>I have to say, it's a very fascinating call out.  Might be worth asking the contest hosters about it.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2058351,
          "author_name": "delai50",
          "author_url": "",
          "post_date": "2022-12-07T20:42:46.170000",
          "content": "<p>For auxiliary targets you can check this <a href=\"https://www.kaggle.com/code/vslaykovsky/train-effnetv2-aux-targets-weighted-loss-thres\" target=\"_blank\">kernel</a> for example</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2058276": "@hrlmmm mentioned in a comment that when `difficult_negative_case=False` and `BIRADS=0` always `cancer=1`. The complete table is the following:\n\n`df_train.groupby(\"patient_id\").tail(1).groupby([\"difficult_negative_case\", \"BIRADS\"])[\"cancer\"].agg([\"mean\",\"size\"])`\n\n||| avg cancer | number of patients |\n| -- | -- |\n| difficult_negative_case | BIRADS |\n| False | 0 | 1 | 125 |\n|| 1 | 0 | 3140 |\n|| 2 | 0 | 424 |\n| True | 0 | 0 | 1575 |\n\nI would say that we must be cautious if we use `difficult_negative_case` and `BIRADS` features because our model could learn that relationship (`difficult_negative_case=False` and `BIRADS=0` then `cancer=1`) which doesn't need to be True in the test set. Any thoughts on that?",
    "2058290": "Some metadata isn't provided for test, check data page (so you can't use difficult_negative_case for example during inference)",
    "2058322": "What I said was when difficult_negative_case=False and  BIRADS=0 than cancer=1.😳May be you write a wrong label ",
    "2058287": "It's worth experimenting I think.  Fine tuning a model such that it can predict BIRADS / difficult negative case, and then seeing if that helps via hill climbing.  Also, fine tuning by predicting density (which is also a signal for cancer).  No promises of course, just ideas for potential diverse models.\n\n"
  }
}