{
  "id": 375605,
  "title": "External Data (CMMD)",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/375605",
  "author_name": "Francisco Javier Gallego",
  "post_date": "2023-01-02T13:19:15.507000",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone, </p>\n<p>Reading one paper I came across with the Chinese Mammography Database. This is one the publicly available mammogram dataset. It contains 5,202 mammograms from 1,775 patients with normal, benign or malignant biopsy-confirmed tumours.</p>\n<p>As external data is allowed, this dataset could be useful for several things. We could try to use it as a dev set, with the aim of extending our dataset, to treat the class imbalance we have for cancer, etc.</p>\n<p>Links of interest: </p>\n<ul>\n<li><a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508\" target=\"_blank\">Chinese Mammography Database (CMMD)</a></li>\n<li><a href=\"https://www.kaggle.com/code/javigallego/rsna-complete-eda-external-data#3-External-Data\" target=\"_blank\">RSNA - Complete EDA (+External Data)</a></li>\n</ul>",
  "messages": [
    {
      "id": 2083390,
      "postDate": "2023-01-02T13:19:15.507Z",
      "content": "<p>Hi everyone, </p>\n<p>Reading one paper I came across with the Chinese Mammography Database. This is one the publicly available mammogram dataset. It contains 5,202 mammograms from 1,775 patients with normal, benign or malignant biopsy-confirmed tumours.</p>\n<p>As external data is allowed, this dataset could be useful for several things. We could try to use it as a dev set, with the aim of extending our dataset, to treat the class imbalance we have for cancer, etc.</p>\n<p>Links of interest: </p>\n<ul>\n<li><a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508\" target=\"_blank\">Chinese Mammography Database (CMMD)</a></li>\n<li><a href=\"https://www.kaggle.com/code/javigallego/rsna-complete-eda-external-data#3-External-Data\" target=\"_blank\">RSNA - Complete EDA (+External Data)</a></li>\n</ul>",
      "rawMarkdown": "Hi everyone, \n\nReading one paper I came across with the Chinese Mammography Database. This is one the publicly available mammogram dataset. It contains 5,202 mammograms from 1,775 patients with normal, benign or malignant biopsy-confirmed tumours.\n\nAs external data is allowed, this dataset could be useful for several things. We could try to use it as a dev set, with the aim of extending our dataset, to treat the class imbalance we have for cancer, etc.\n\nLinks of interest: \n\n* [Chinese Mammography Database (CMMD)](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508)\n* [RSNA - Complete EDA (+External Data)](https://www.kaggle.com/code/javigallego/rsna-complete-eda-external-data#3-External-Data)",
      "votes": 10
    },
    {
      "id": 2084865,
      "postDate": "2023-01-03T19:49:58.913Z",
      "content": "<p>I think a big question here, is what exactly cancer means in the RSNA dataset. How long after the scan was taken was cancer confirmed, specifically.</p>\n<p>Discussed more here -<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375498\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375498</a></p>\n<p>That could seriously bring into question what datasets can be used to train with.</p>\n<p>I've scanned this <a href=\"https://pubs.rsna.org/doi/pdf/10.1148/ryai.220072\" target=\"_blank\">https://pubs.rsna.org/doi/pdf/10.1148/ryai.220072</a> and <a href=\"https://www.rsna.org/education/ai-resources-and-training/ai-image-challenge/screening-mammography-breast-cancer-detection-ai-challenge\" target=\"_blank\">https://www.rsna.org/education/ai-resources-and-training/ai-image-challenge/screening-mammography-breast-cancer-detection-ai-challenge</a> which indicates that site 1 and 2 are US / Australia (not necessarily respectively). I didn't see anything that mentioned the time between scans and cancer detection.</p>\n<p>One thing to keep in mind is that cancer is often detected over time. You get the first screen and then a year or more later, you get another one and the delta helps detect. In our case, we only have one screening. Whether it's the initial or not, not sure.</p>",
      "rawMarkdown": "I think a big question here, is what exactly cancer means in the RSNA dataset. How long after the scan was taken was cancer confirmed, specifically.\n\nDiscussed more here -\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375498\n\nThat could seriously bring into question what datasets can be used to train with.\n\nI've scanned this https://pubs.rsna.org/doi/pdf/10.1148/ryai.220072 and https://www.rsna.org/education/ai-resources-and-training/ai-image-challenge/screening-mammography-breast-cancer-detection-ai-challenge which indicates that site 1 and 2 are US / Australia (not necessarily respectively). I didn't see anything that mentioned the time between scans and cancer detection.\n\nOne thing to keep in mind is that cancer is often detected over time. You get the first screen and then a year or more later, you get another one and the delta helps detect. In our case, we only have one screening. Whether it's the initial or not, not sure.",
      "votes": 1,
      "replies": [
        {
          "id": 2085174,
          "postDate": "2023-01-04T00:51:08.290Z",
          "content": "<p>\"is what exactly cancer means in the RSNA dataset. How long after the scan was taken was cancer confirmed, specifically. \"<br>\nrefer to   </p>\n<pre><code>https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2084695\n</code></pre>\n<hr>\n<p>\"site 1 and 2 are US / Australia (not necessarily respectively).\"<br>\nmake a histogram plot of the ages.<br>\ncompare that with the papers</p>\n<pre><code>Name: age, dtype: int64\nd1 = train_df[train_df.site_id==1]\nd1.age.value_counts().sort_index()\nOut[6]: \n26.0    11\n28.0    18\n29.0     7\nd2.age.value_counts().sort_index()\nOut[4]: \n40.0     120  &lt;--- ###### caught the cat!\n41.0     109\n42.0     134\n43.0     169\n44.0     126\n</code></pre>\n<p>further, for site 1:<br>\n<img src=\"https://i.ibb.co/p40pZL7/Selection-461.png\" alt=\"https://i.ibb.co/p40pZL7/Selection-461.png\"></p>",
          "rawMarkdown": "\n\"is what exactly cancer means in the RSNA dataset. How long after the scan was taken was cancer confirmed, specifically. \"\n\nrefer to   \n\n```\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2084695\n\n```\n\n---\n\n\n\"site 1 and 2 are US / Australia (not necessarily respectively).\"\n\nmake a histogram plot of the ages.\ncompare that with the papers\n\n```\nName: age, dtype: int64\nd1 = train_df[train_df.site_id==1]\nd1.age.value_counts().sort_index()\nOut[6]: \n26.0    11\n28.0    18\n29.0     7\n\n\nd2.age.value_counts().sort_index()\nOut[4]: \n40.0     120  <--- ###### caught the cat!\n41.0     109\n42.0     134\n43.0     169\n44.0     126\n```\n\nfurther, for site 1:\n\n![https://i.ibb.co/p40pZL7/Selection-461.png](https://i.ibb.co/p40pZL7/Selection-461.png)\n\n"
        }
      ]
    },
    {
      "id": 2083546,
      "postDate": "2023-01-02T15:47:37.300Z",
      "content": "<p>thanks for sharing, I convert my customize dataset based on your notebook.</p>",
      "rawMarkdown": "thanks for sharing, I convert my customize dataset based on your notebook.",
      "votes": 1
    },
    {
      "id": 2084781,
      "postDate": "2023-01-03T18:30:25.047Z",
      "content": "<p>This dataset appears to be licensed under Creative Commons, but without a non-commercial constraint, so I think we can use it.</p>\n<p>It seems quite useful, because every DICOM shows some abnormality - i.e. are not completely normal. (However, I think technically this is only true to D1-**** IDs - see below). It is therefore very rich for adversarial cases.</p>\n<p>However, the dataset itself is fairly confusing.</p>\n<p>If you look at the labels in CMMD_clinicaldata_revision.xlsx, you see two types of ID.</p>\n<p>The D1-**** IDs I understand are cases where there are 2 images inside the folder of the same name (1-1.png and 1-2.png). These will both be of the same breast, and will both correspond with the label in the excel sheet.</p>\n<p>The The D2-**** IDs are more confusing. Although the Excel sheet lists 2 images per study (as with D1 IDs), the folders often contain 4 images.</p>\n<p>The website, under detailed description, says:</p>\n<p>\"For the D2-XXXX dataset, it is a dataset that only involves malignant tumors. Therefore, only one side of the clinical data is reasonable, such a situation shows that the other side is benign. We provided mammograms from both the left and right breast.\"</p>\n<p>I think we are therefore to assume that the files 1-1.png and 1-2.png in these folders correspond to the labels in the Excel sheet, and that 1-3.png and 1-4.png correspond to the OTHER breast, and are therefore non malignant? These breasts of course, are probably a 'different' benign from the benign ones in the D1 dataset (where there are still lesions, but non-malignant ones).</p>\n<p>Finally, despite the statement above, some cases (e.g. D2-0090) are not malignant, but rather benign. I wonder if these cases were suspected malignant, but ended up being benign on biopsy?</p>\n<p>Has anyone else explored this dataset yet?</p>\n<p>There's also another statement where it says \"Clinical data are saved in .XLSX format. Note that for those rows where there exists BOTH a value for ID1 and ID2, TCIA image database stores ONLY the ID2 value as PatientID.\" - I'm not sure what this means…</p>",
      "rawMarkdown": "This dataset appears to be licensed under Creative Commons, but without a non-commercial constraint, so I think we can use it.\n\nIt seems quite useful, because every DICOM shows some abnormality - i.e. are not completely normal. (However, I think technically this is only true to D1-**** IDs - see below). It is therefore very rich for adversarial cases.\n\nHowever, the dataset itself is fairly confusing.\n\nIf you look at the labels in CMMD_clinicaldata_revision.xlsx, you see two types of ID.\n\nThe D1-**** IDs I understand are cases where there are 2 images inside the folder of the same name (1-1.png and 1-2.png). These will both be of the same breast, and will both correspond with the label in the excel sheet.\n\nThe The D2-**** IDs are more confusing. Although the Excel sheet lists 2 images per study (as with D1 IDs), the folders often contain 4 images.\n\nThe website, under detailed description, says:\n\n\"For the D2-XXXX dataset, it is a dataset that only involves malignant tumors. Therefore, only one side of the clinical data is reasonable, such a situation shows that the other side is benign. We provided mammograms from both the left and right breast.\"\n\nI think we are therefore to assume that the files 1-1.png and 1-2.png in these folders correspond to the labels in the Excel sheet, and that 1-3.png and 1-4.png correspond to the OTHER breast, and are therefore non malignant? These breasts of course, are probably a 'different' benign from the benign ones in the D1 dataset (where there are still lesions, but non-malignant ones).\n\nFinally, despite the statement above, some cases (e.g. D2-0090) are not malignant, but rather benign. I wonder if these cases were suspected malignant, but ended up being benign on biopsy?\n\nHas anyone else explored this dataset yet?\n\nThere's also another statement where it says \"Clinical data are saved in .XLSX format. Note that for those rows where there exists BOTH a value for ID1 and ID2, TCIA image database stores ONLY the ID2 value as PatientID.\" - I'm not sure what this means…",
      "replies": [
        {
          "id": 2084793,
          "postDate": "2023-01-03T18:38:37.177Z",
          "content": "<p>you can either:</p>\n<ol>\n<li><p>mix with kaggle data in common loss function and use the labeling: <br>\neverything that is malignant = kaggle cancer label<br>\neverything that is benign = kaggle non-cancer label</p></li>\n<li><p>have different loss function for each external data, just treat external data label supervision signal as aux loss:<br>\nkaggle loss function = backprop for kaggle data only<br>\nCMMD loss function = backprop for CMMD  data only </p></li>\n</ol>\n<p>of course, use validation cv and lb to see if it work actually works</p>\n<p>check also papers that uses CMMD from google scholar citations.<br>\nI think some papers did experiments on all ADMANI,CMMD,NYU</p>\n<p>i think CMMD is a more difficult data and the results poorer than the rest</p>",
          "rawMarkdown": "you can either:\n\n1. mix with kaggle data in common loss function and use the labeling: \n everything that is malignant = kaggle cancer label\n everything that is benign = kaggle non-cancer label\n\n2. have different loss function for each external data, just treat external data label supervision signal as aux loss:\nkaggle loss function = backprop for kaggle data only\nCMMD loss function = backprop for CMMD  data only \n\nof course, use validation cv and lb to see if it work actually works\n\ncheck also papers that uses CMMD from google scholar citations.\nI think some papers did experiments on all ADMANI,CMMD,NYU\n\ni think CMMD is a more difficult data and the results poorer than the rest"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2084865,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2023-01-03T19:49:58.913000",
      "content": "<p>I think a big question here, is what exactly cancer means in the RSNA dataset. How long after the scan was taken was cancer confirmed, specifically.</p>\n<p>Discussed more here -<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375498\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375498</a></p>\n<p>That could seriously bring into question what datasets can be used to train with.</p>\n<p>I've scanned this <a href=\"https://pubs.rsna.org/doi/pdf/10.1148/ryai.220072\" target=\"_blank\">https://pubs.rsna.org/doi/pdf/10.1148/ryai.220072</a> and <a href=\"https://www.rsna.org/education/ai-resources-and-training/ai-image-challenge/screening-mammography-breast-cancer-detection-ai-challenge\" target=\"_blank\">https://www.rsna.org/education/ai-resources-and-training/ai-image-challenge/screening-mammography-breast-cancer-detection-ai-challenge</a> which indicates that site 1 and 2 are US / Australia (not necessarily respectively). I didn't see anything that mentioned the time between scans and cancer detection.</p>\n<p>One thing to keep in mind is that cancer is often detected over time. You get the first screen and then a year or more later, you get another one and the delta helps detect. In our case, we only have one screening. Whether it's the initial or not, not sure.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2085174,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-04T00:51:08.290000",
          "content": "<p>\"is what exactly cancer means in the RSNA dataset. How long after the scan was taken was cancer confirmed, specifically. \"<br>\nrefer to   </p>\n<pre><code>https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2084695\n</code></pre>\n<hr>\n<p>\"site 1 and 2 are US / Australia (not necessarily respectively).\"<br>\nmake a histogram plot of the ages.<br>\ncompare that with the papers</p>\n<pre><code>Name: age, dtype: int64\nd1 = train_df[train_df.site_id==1]\nd1.age.value_counts().sort_index()\nOut[6]: \n26.0    11\n28.0    18\n29.0     7\nd2.age.value_counts().sort_index()\nOut[4]: \n40.0     120  &lt;--- ###### caught the cat!\n41.0     109\n42.0     134\n43.0     169\n44.0     126\n</code></pre>\n<p>further, for site 1:<br>\n<img src=\"https://i.ibb.co/p40pZL7/Selection-461.png\" alt=\"https://i.ibb.co/p40pZL7/Selection-461.png\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2083546,
      "author_name": "ITOSHIMARO",
      "author_url": "",
      "post_date": "2023-01-02T15:47:37.300000",
      "content": "<p>thanks for sharing, I convert my customize dataset based on your notebook.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2084781,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2023-01-03T18:30:25.047000",
      "content": "<p>This dataset appears to be licensed under Creative Commons, but without a non-commercial constraint, so I think we can use it.</p>\n<p>It seems quite useful, because every DICOM shows some abnormality - i.e. are not completely normal. (However, I think technically this is only true to D1-**** IDs - see below). It is therefore very rich for adversarial cases.</p>\n<p>However, the dataset itself is fairly confusing.</p>\n<p>If you look at the labels in CMMD_clinicaldata_revision.xlsx, you see two types of ID.</p>\n<p>The D1-**** IDs I understand are cases where there are 2 images inside the folder of the same name (1-1.png and 1-2.png). These will both be of the same breast, and will both correspond with the label in the excel sheet.</p>\n<p>The The D2-**** IDs are more confusing. Although the Excel sheet lists 2 images per study (as with D1 IDs), the folders often contain 4 images.</p>\n<p>The website, under detailed description, says:</p>\n<p>\"For the D2-XXXX dataset, it is a dataset that only involves malignant tumors. Therefore, only one side of the clinical data is reasonable, such a situation shows that the other side is benign. We provided mammograms from both the left and right breast.\"</p>\n<p>I think we are therefore to assume that the files 1-1.png and 1-2.png in these folders correspond to the labels in the Excel sheet, and that 1-3.png and 1-4.png correspond to the OTHER breast, and are therefore non malignant? These breasts of course, are probably a 'different' benign from the benign ones in the D1 dataset (where there are still lesions, but non-malignant ones).</p>\n<p>Finally, despite the statement above, some cases (e.g. D2-0090) are not malignant, but rather benign. I wonder if these cases were suspected malignant, but ended up being benign on biopsy?</p>\n<p>Has anyone else explored this dataset yet?</p>\n<p>There's also another statement where it says \"Clinical data are saved in .XLSX format. Note that for those rows where there exists BOTH a value for ID1 and ID2, TCIA image database stores ONLY the ID2 value as PatientID.\" - I'm not sure what this means…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2084793,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-03T18:38:37.177000",
          "content": "<p>you can either:</p>\n<ol>\n<li><p>mix with kaggle data in common loss function and use the labeling: <br>\neverything that is malignant = kaggle cancer label<br>\neverything that is benign = kaggle non-cancer label</p></li>\n<li><p>have different loss function for each external data, just treat external data label supervision signal as aux loss:<br>\nkaggle loss function = backprop for kaggle data only<br>\nCMMD loss function = backprop for CMMD  data only </p></li>\n</ol>\n<p>of course, use validation cv and lb to see if it work actually works</p>\n<p>check also papers that uses CMMD from google scholar citations.<br>\nI think some papers did experiments on all ADMANI,CMMD,NYU</p>\n<p>i think CMMD is a more difficult data and the results poorer than the rest</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2083390": "Hi everyone, \n\nReading one paper I came across with the Chinese Mammography Database. This is one the publicly available mammogram dataset. It contains 5,202 mammograms from 1,775 patients with normal, benign or malignant biopsy-confirmed tumours.\n\nAs external data is allowed, this dataset could be useful for several things. We could try to use it as a dev set, with the aim of extending our dataset, to treat the class imbalance we have for cancer, etc.\n\nLinks of interest: \n\n* [Chinese Mammography Database (CMMD)](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508)\n* [RSNA - Complete EDA (+External Data)](https://www.kaggle.com/code/javigallego/rsna-complete-eda-external-data#3-External-Data)",
    "2084865": "I think a big question here, is what exactly cancer means in the RSNA dataset. How long after the scan was taken was cancer confirmed, specifically.\n\nDiscussed more here -\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375498\n\nThat could seriously bring into question what datasets can be used to train with.\n\nI've scanned this https://pubs.rsna.org/doi/pdf/10.1148/ryai.220072 and https://www.rsna.org/education/ai-resources-and-training/ai-image-challenge/screening-mammography-breast-cancer-detection-ai-challenge which indicates that site 1 and 2 are US / Australia (not necessarily respectively). I didn't see anything that mentioned the time between scans and cancer detection.\n\nOne thing to keep in mind is that cancer is often detected over time. You get the first screen and then a year or more later, you get another one and the delta helps detect. In our case, we only have one screening. Whether it's the initial or not, not sure.",
    "2083546": "thanks for sharing, I convert my customize dataset based on your notebook.",
    "2084781": "This dataset appears to be licensed under Creative Commons, but without a non-commercial constraint, so I think we can use it.\n\nIt seems quite useful, because every DICOM shows some abnormality - i.e. are not completely normal. (However, I think technically this is only true to D1-**** IDs - see below). It is therefore very rich for adversarial cases.\n\nHowever, the dataset itself is fairly confusing.\n\nIf you look at the labels in CMMD_clinicaldata_revision.xlsx, you see two types of ID.\n\nThe D1-**** IDs I understand are cases where there are 2 images inside the folder of the same name (1-1.png and 1-2.png). These will both be of the same breast, and will both correspond with the label in the excel sheet.\n\nThe The D2-**** IDs are more confusing. Although the Excel sheet lists 2 images per study (as with D1 IDs), the folders often contain 4 images.\n\nThe website, under detailed description, says:\n\n\"For the D2-XXXX dataset, it is a dataset that only involves malignant tumors. Therefore, only one side of the clinical data is reasonable, such a situation shows that the other side is benign. We provided mammograms from both the left and right breast.\"\n\nI think we are therefore to assume that the files 1-1.png and 1-2.png in these folders correspond to the labels in the Excel sheet, and that 1-3.png and 1-4.png correspond to the OTHER breast, and are therefore non malignant? These breasts of course, are probably a 'different' benign from the benign ones in the D1 dataset (where there are still lesions, but non-malignant ones).\n\nFinally, despite the statement above, some cases (e.g. D2-0090) are not malignant, but rather benign. I wonder if these cases were suspected malignant, but ended up being benign on biopsy?\n\nHas anyone else explored this dataset yet?\n\nThere's also another statement where it says \"Clinical data are saved in .XLSX format. Note that for those rows where there exists BOTH a value for ID1 and ID2, TCIA image database stores ONLY the ID2 value as PatientID.\" - I'm not sure what this means…"
  }
}