{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [EDA] RSNA Screening Mammography Breast Cancer Detection\n","metadata":{}},{"cell_type":"markdown","source":"This notebook describes the [EDA] for the [RSNA Screening Mammography Breast Cancer Detection](https://www.kaggle.com/competitions/rsna-breast-cancer-detection)  \nこのノートブックは[RSNA Screening Mammography Breast Cancer Detection](https://www.kaggle.com/competitions/rsna-breast-cancer-detection)コンペティションのEDAです。コンペティション参加したての人向けにNotebookを構成しております。","metadata":{}},{"cell_type":"markdown","source":"本notebookのオリジナルは[こちら](https://www.kaggle.com/code/masatakaitakura/eda-for-beginner-rsna-mammography-breast-cancer)です。","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport pandas_profiling\nimport numpy as np\nimport os\n\n\n# visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\n\n# install gdcm to open the DICOM file\n!pip install -qU python-gdcm pydicom pylibjpeg\n\nimport pydicom\nimport pylibjpeg\n","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:43:48.764245Z","iopub.execute_input":"2022-12-17T05:43:48.764727Z","iopub.status.idle":"2022-12-17T05:44:09.529249Z","shell.execute_reply.started":"2022-12-17T05:43:48.764633Z","shell.execute_reply":"2022-12-17T05:44:09.528162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"section-one\"></a>\n# Overview of data / データ概要","metadata":{}},{"cell_type":"markdown","source":"## train data / trainデータ","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:09.531288Z","iopub.execute_input":"2022-12-17T05:44:09.532086Z","iopub.status.idle":"2022-12-17T05:44:09.711594Z","shell.execute_reply.started":"2022-12-17T05:44:09.532044Z","shell.execute_reply":"2022-12-17T05:44:09.710447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following is the explanation of columns in [Kaggle-Data](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/data) \n\n`site_id` - ID code for the source hospital.  \n`patient_id` - ID code for the patient.  \n`image_id` - ID code for the image.  \n`laterality` - Whether the image is of the left or right breast.  \n`view` - The orientation of the image. The default for a screening exam is to capture two views per breast.  \n`age` - The patient's age in years.  \n`implant` - Whether or not the patient had breast implants. Site 1 only provides breast implant information at the patient level, not at the breast level.  \n`density` - A rating for how dense the breast tissue is, with A being the least dense and D being the most dense. Extremely dense tissue can make diagnosis more difficult. <font color=blue>Only provided for train.</font>  \n`machine_id` - An ID code for the imaging device.  \n`cancer` - Whether or not the breast was positive for cancer. The target value. <font color=blue>Only provided for train.  </font>  \n`biopsy` - Whether or not a follow-up biopsy was performed on the breast. <font color=blue>Only provided for train.  </font>  \n`invasive` - If the breast is positive for cancer, whether or not the cancer proved to be invasive. <font color=blue>Only provided for train.  </font>  \n`BIRADS` - 0 if the breast required follow-up, 1 if the breast was rated as negative for cancer, and 2 if the breast was rated as normal. <font color=blue>Only provided for train.  </font>  \n`prediction_id` - The ID for the matching submission row. Multiple images will share the same prediction ID. Test only.  \n`difficult_negative_case` - True if the case was unusually difficult. <font color=blue>Only provided for train.  </font>","metadata":{}},{"cell_type":"markdown","source":"以下はtrainデータの列名と内容の説明を [Kaggle-Data](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/data) から参照しました。\n\n`site_id` - 検査した病院のID  \n`patient_id` - 患者ごとのID <font color=red>[重要]</font>.  </br>\n`image_id` - 画像ごとのID.  \n`laterality` - イメージが「右」か「左」の乳房か.  \n`view` - 画像の傾き. 検査のデフォルトは1患者に対して2画像をとる.  \n`age` - 患者の年齢.  \n`implant` - 患者が乳房にインプラントしているか.  \n`density` - 乳房の乳腺などの器官の密度. A < B < C < Dの順に密度が高くなる. 密度の高い乳房だと検査は難しくなる. <font color=blue>trainデータのみ与えられる.</font>  \n`machine_id` - 検査機器のID.  \n`cancer` - ガン陽性か否か. 本コンペにおける目的変数. <font color=blue>trainデータのみ与えられる.</font>  <font color=red>[重要]</font>.  </br>\n`biopsy` - 追加の生体検査 (biopsy) が実施されたか. <font color=blue>trainデータのみ与えられる.</font>  \n`invasive` - ガンと診断された場合に、そのガンが「浸潤がん」と診断されたか否か. <font color=blue>trainデータのみ与えられる.</font>  \n`BIRADS` - 0 follow-upが必要, 1 ガン陰性 (ガンじゃない), 2 normalと判定. <font color=blue>trainデータのみ与えられる.</font>  \n`prediction_id` - マッチング提出用のID. テストのみ.  \n`difficult_negative_case` - 判定が特別に難しい場合にTRUE. <font color=blue>trainデータのみ与えられる.</font>  ","metadata":{}},{"cell_type":"markdown","source":"Let's see the overview of the train data.  \ntrainデータについて概要を見ていきます。","metadata":{}},{"cell_type":"markdown","source":"## Target columns and features / ガン陽性・陰性列の特徴","metadata":{}},{"cell_type":"code","source":"# 'cancer' and 'age'\nfig, ax1 = plt.subplots()\n\ncolor = 'tab:blue'\nax1.set_xlabel('age')\nax1.set_ylabel('cancer=0 count', color=color)\nax1.hist(train.loc[train['cancer'] == 0, 'age'].dropna(), bins=30, alpha=0.5, label='0', color=color)\nax1.tick_params(axis='y', labelcolor=color)\n\nax2 = ax1.twinx()  # instantiate a second axes that shares the same x-axis\n\ncolor = 'tab:orange'\nax2.set_ylabel('cancer=1 count', color=color)  # we already handled the x-label with ax1\nax2.hist(train.loc[train['cancer'] == 1, 'age'].dropna(), bins=30, alpha=0.8, label='1', color=color)\nax2.tick_params(axis='y', labelcolor=color)\n\nfig.tight_layout()  # otherwise the right y-label is slightly clipped\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:09.713484Z","iopub.execute_input":"2022-12-17T05:44:09.714724Z","iopub.status.idle":"2022-12-17T05:44:10.271333Z","shell.execute_reply.started":"2022-12-17T05:44:09.71467Z","shell.execute_reply":"2022-12-17T05:44:10.26993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Statistical data for 'cancer' and 'age'\ntrain.loc[train['cancer'] == 1, 'age'].describe()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:10.274043Z","iopub.execute_input":"2022-12-17T05:44:10.274471Z","iopub.status.idle":"2022-12-17T05:44:10.28941Z","shell.execute_reply.started":"2022-12-17T05:44:10.274436Z","shell.execute_reply":"2022-12-17T05:44:10.287972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings:\n- cancers count vs age seems normal distribution.\n- the peak is appeared around age 60 to 75.\n- cancers appeared from age 40.\n\n考察:\n- がん数 vs 年齢は正規分布のようです。\n- 60～75歳頃にピークが現れる。\n- がんは40歳から発症。","metadata":{}},{"cell_type":"code","source":"# 'cancer' and 'laterality'\nsns.countplot(x='laterality', hue='cancer', data=train)\nplt.legend(loc='upper right', title='cancer')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:10.290858Z","iopub.execute_input":"2022-12-17T05:44:10.291237Z","iopub.status.idle":"2022-12-17T05:44:10.540528Z","shell.execute_reply.started":"2022-12-17T05:44:10.291205Z","shell.execute_reply":"2022-12-17T05:44:10.539062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings:\n- laterality data seems balanced.\n\n考察：\n- lateralityデータ(乳房の左右)はバランスよく検査されたようです。","metadata":{}},{"cell_type":"code","source":"# 'cancer' and 'view'\nsns.countplot(x='view', hue='cancer', data=train)\nplt.legend(loc='upper right', title='cancer')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:10.542485Z","iopub.execute_input":"2022-12-17T05:44:10.542902Z","iopub.status.idle":"2022-12-17T05:44:10.817461Z","shell.execute_reply.started":"2022-12-17T05:44:10.542864Z","shell.execute_reply":"2022-12-17T05:44:10.816137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings\n- it seems there are no significand difference between 'cancer' and 'non-cancer'.\n- the view of 'CC' and 'MLO' are the most popular.\n  \n考察\n- 結果がガン陽性と陰性の結果で大きな偏りはないようです。\n- viewは'CC'と'MLO'のデータ数が最も多かったです。","metadata":{}},{"cell_type":"code","source":"# 'cancer' and 'density' *ONLY available in TRAIN data\nsns.countplot(x='density', hue='cancer', data=train)\nplt.legend(loc='upper right', title='cancer')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:10.819143Z","iopub.execute_input":"2022-12-17T05:44:10.820051Z","iopub.status.idle":"2022-12-17T05:44:11.036273Z","shell.execute_reply.started":"2022-12-17T05:44:10.82Z","shell.execute_reply":"2022-12-17T05:44:11.035098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"「がん」と「密度」の関係です。\n\n考察\n- 密度は'C'と'B'の中密度が多く、高密度の'D'と低密度の'A'は少ない","metadata":{}},{"cell_type":"code","source":"# 'cancer' and 'BIRADS' *ONLY available in TRAIN data\nsns.countplot(x='BIRADS', hue='cancer', data=train)\nplt.legend(loc='upper right', title='cancer')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:11.03896Z","iopub.execute_input":"2022-12-17T05:44:11.040464Z","iopub.status.idle":"2022-12-17T05:44:11.21024Z","shell.execute_reply.started":"2022-12-17T05:44:11.040405Z","shell.execute_reply":"2022-12-17T05:44:11.209134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"「がん」と「BIRADS」の関係です。\n\n考察\n- がん患者は全員BIRADS=0, つまりフォローアップが必要と分類される","metadata":{}},{"cell_type":"markdown","source":"If you want to see more comprehensive overview, please see bellow.  \nより包括的な概要を見たい場合には以下のレポートを参照ください。","metadata":{}},{"cell_type":"code","source":"train.profile_report()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:11.212385Z","iopub.execute_input":"2022-12-17T05:44:11.21291Z","iopub.status.idle":"2022-12-17T05:44:32.036126Z","shell.execute_reply.started":"2022-12-17T05:44:11.21286Z","shell.execute_reply":"2022-12-17T05:44:32.03518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, see the data of `cancer=1` to understand the target column features.  \n次に`cancer=1` のデータのみを取り出してデータの概要を見ます。","metadata":{}},{"cell_type":"code","source":"train_cancer = train[train['cancer']==1]\ntrain_cancer","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:32.037353Z","iopub.execute_input":"2022-12-17T05:44:32.038146Z","iopub.status.idle":"2022-12-17T05:44:32.065536Z","shell.execute_reply.started":"2022-12-17T05:44:32.038111Z","shell.execute_reply":"2022-12-17T05:44:32.064607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"包括的なレポートを参照する場合には以下を実行してください。","metadata":{}},{"cell_type":"code","source":"# if you want to see the detail [profile_report], please run this cell.\n# train_cancer.profile_report()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:32.069422Z","iopub.execute_input":"2022-12-17T05:44:32.070026Z","iopub.status.idle":"2022-12-17T05:44:32.073881Z","shell.execute_reply.started":"2022-12-17T05:44:32.069988Z","shell.execute_reply":"2022-12-17T05:44:32.07304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## test data / testデータ","metadata":{}},{"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:32.075753Z","iopub.execute_input":"2022-12-17T05:44:32.076422Z","iopub.status.idle":"2022-12-17T05:44:32.103722Z","shell.execute_reply.started":"2022-12-17T05:44:32.076387Z","shell.execute_reply":"2022-12-17T05:44:32.102759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test.csv only contains 4 rows, single person data.\n\"Only the first few rows of the test set are available for download.\"  \ntest.csv には、1 人のデータが 4 行しか含まれていません。\n「テスト セットの最初の数行のみをダウンロードできます。」","metadata":{}},{"cell_type":"markdown","source":"# sample_submission / サンプルサブミッションファイル","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/sample_submission.csv')\nsample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:32.104979Z","iopub.execute_input":"2022-12-17T05:44:32.105798Z","iopub.status.idle":"2022-12-17T05:44:32.123177Z","shell.execute_reply.started":"2022-12-17T05:44:32.105763Z","shell.execute_reply":"2022-12-17T05:44:32.122077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization / データ可視化","metadata":{}},{"cell_type":"markdown","source":"To visualize the DICOM file (.dcm), the following function is defined. Note that the line #12 and #13 is commented out to see the difference between each `machine_id`. The further detail to be discussed later.  \n  \n  DICOM ファイル (.dcm) を可視化するために、次の関数が定義されています。 各「machine_id」の違いを確認するために、行 #12 と #13 がコメントアウトされていることに注意してください。 詳細については後述します。","metadata":{}},{"cell_type":"code","source":"# Function to visualize the DICOM file\n\ndef visualize(patient_id, figsize=(20, 20)):\n    image_num = list(train['patient_id']==patient_id).count(True)\n    fig, ax = plt.subplots(1, image_num, figsize=figsize)\n    ax = ax.flatten()\n    path_to_dcms = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\n    for i, dcm_path in enumerate(os.listdir(os.path.join(path_to_dcms, str(patient_id)))):\n        dcm = pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), dcm_path))\n        dcm = dcm.pixel_array\n        dcm = (dcm - dcm.min()) / (dcm.max() - dcm.min()) * 255\n#         if pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), dcm_path)).PhotometricInterpretation == \"MONOCHROME1\":\n#             dcm = 1 - dcm\n        ax[i].set_title(dcm_path)\n        ax[i].imshow(dcm, cmap=\"bone\")","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:32.124672Z","iopub.execute_input":"2022-12-17T05:44:32.125072Z","iopub.status.idle":"2022-12-17T05:44:32.133809Z","shell.execute_reply.started":"2022-12-17T05:44:32.125031Z","shell.execute_reply":"2022-12-17T05:44:32.132613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# visualize the patient_id images\nvisualize(10006)\nvisualize(10011)\ntrain.head(8)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:32.135047Z","iopub.execute_input":"2022-12-17T05:44:32.135978Z","iopub.status.idle":"2022-12-17T05:44:54.088774Z","shell.execute_reply.started":"2022-12-17T05:44:32.13594Z","shell.execute_reply":"2022-12-17T05:44:54.086965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Basically, each `patient_id` has 4 mammography image, but some patient has more than 4 images. So the `visualization` function consider this by checking an [image_num].","metadata":{}},{"cell_type":"markdown","source":"基本的に各患者 ID には 4 つのマンモグラフィ画像がありますが、一部の患者には 4 つ以上の画像があります。 したがって、可視化関数は [image_num] をチェックしてこれを考慮します。","metadata":{}},{"cell_type":"code","source":"patient_id = 10049\nimage_num = list(train['patient_id']==patient_id).count(True)\nimage_num","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:54.090687Z","iopub.execute_input":"2022-12-17T05:44:54.091315Z","iopub.status.idle":"2022-12-17T05:44:54.106726Z","shell.execute_reply.started":"2022-12-17T05:44:54.091275Z","shell.execute_reply":"2022-12-17T05:44:54.105353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[train['patient_id']==10049]","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:54.108943Z","iopub.execute_input":"2022-12-17T05:44:54.10934Z","iopub.status.idle":"2022-12-17T05:44:54.173765Z","shell.execute_reply.started":"2022-12-17T05:44:54.109307Z","shell.execute_reply":"2022-12-17T05:44:54.172628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings:\n- the mammography images contain some 'blank' area (ex. laterality=Right has blank in left area.)\n- the background color seems different depende on the `machine_id` -> <font color=red>black and white contract should be aligned for each image??</font>","metadata":{}},{"cell_type":"markdown","source":"考察\n- マンモグラフィ画像には「空白」の領域が含まれています (例: laterality = 右の領域は左側の領域に空白があります)。\n- 背景色は machine_id によって異なるようです -> 画像ごとに白黒のコントラクトを揃える必要があるか??\n","metadata":{}},{"cell_type":"markdown","source":"To consider the 2nd bullet of the above findings, we re-defined the function as bellow. The only difference is just removed the comment-out in line #12 and #13.","metadata":{}},{"cell_type":"markdown","source":"上記の調査結果の 2項目を検討するために、関数を次のように再定義しました。 唯一の違いは、12 行目と 13 行目のコメントアウトを削除したことだけです。","metadata":{}},{"cell_type":"code","source":"# Function to visualize the DICOM file\n\ndef visualize(patient_id, figsize=(20, 20)):\n    image_num = list(train['patient_id']==patient_id).count(True)\n    fig, ax = plt.subplots(1, image_num, figsize=figsize)\n    ax = ax.flatten()\n    path_to_dcms = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\n    for i, dcm_path in enumerate(os.listdir(os.path.join(path_to_dcms, str(patient_id)))):\n        dcm = pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), dcm_path))\n        dcm = dcm.pixel_array\n        dcm = (dcm - dcm.min()) / (dcm.max() - dcm.min()) * 255\n        if pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), dcm_path)).PhotometricInterpretation == \"MONOCHROME1\":\n            dcm = 1 - dcm\n        ax[i].set_title(dcm_path)\n        ax[i].imshow(dcm, cmap=\"bone\")","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:54.175301Z","iopub.execute_input":"2022-12-17T05:44:54.175683Z","iopub.status.idle":"2022-12-17T05:44:54.186454Z","shell.execute_reply.started":"2022-12-17T05:44:54.17565Z","shell.execute_reply":"2022-12-17T05:44:54.184924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# visualize the patient_id images\nvisualize(10006)\nvisualize(10011)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:44:54.187498Z","iopub.execute_input":"2022-12-17T05:44:54.187862Z","iopub.status.idle":"2022-12-17T05:45:15.69417Z","shell.execute_reply.started":"2022-12-17T05:44:54.187816Z","shell.execute_reply":"2022-12-17T05:45:15.692985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can obtain the visualization of remaining all DICOM files.","metadata":{}},{"cell_type":"markdown","source":"以前の関数では別の`machine_id`の画像はエラーで表示することができませんでしたが、これで残りの全ての画像を可視化することができるようになりました。","metadata":{}},{"cell_type":"markdown","source":"## cancer mammography / がん結果の画像\nNext we check the cancer mammography images.  \nがんで「陽性」と診断された画像について見ていきます。","metadata":{}},{"cell_type":"code","source":"visualize(10130)\nvisualize(10226)\nvisualize(1025)\nvisualize(10432)\nvisualize(10589)\ntrain_cancer.head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:45:15.695653Z","iopub.execute_input":"2022-12-17T05:45:15.696132Z","iopub.status.idle":"2022-12-17T05:46:18.12914Z","shell.execute_reply.started":"2022-12-17T05:45:15.696096Z","shell.execute_reply":"2022-12-17T05:46:18.127907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is difficult to find out the cancer by visual judgement...however image_id=613462606 has the \"white block area\" so this seems a cancer.\n","metadata":{}},{"cell_type":"markdown","source":"肉眼での判断は難しいですが、image_id=613462606 に「白いブロックの部分」があるので、ガンのようです。","metadata":{}},{"cell_type":"code","source":"def visualize_single(patient_id, image_id, figsize=(20, 20)):\n    fig, ax = plt.subplots(1, 1, figsize=figsize)\n    path_to_dcms = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\n    dcm = pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), str(image_id)+'.dcm'))\n    dcm = dcm.pixel_array\n    ax.imshow(dcm)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:46:18.130744Z","iopub.execute_input":"2022-12-17T05:46:18.131097Z","iopub.status.idle":"2022-12-17T05:46:18.137444Z","shell.execute_reply.started":"2022-12-17T05:46:18.131066Z","shell.execute_reply":"2022-12-17T05:46:18.136153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize_single(patient_id=10130, image_id=613462606)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:46:18.139267Z","iopub.execute_input":"2022-12-17T05:46:18.139904Z","iopub.status.idle":"2022-12-17T05:46:21.364369Z","shell.execute_reply.started":"2022-12-17T05:46:18.139851Z","shell.execute_reply":"2022-12-17T05:46:21.363125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## machine_id","metadata":{}},{"cell_type":"markdown","source":"Now we check the `machine_id` information for more detail.\nThe machine_id has total 10 groups as below.","metadata":{}},{"cell_type":"markdown","source":"次に`machine_id`についてより詳細に見ていきます。machine_idは全部で10 グループあるようです。","metadata":{}},{"cell_type":"code","source":"print(f\"machine_id of all train data: {train['machine_id'].unique()}\")\nprint(f\"machine_id of cancer=1 train data{train_cancer['machine_id'].unique()}\")","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:46:21.366057Z","iopub.execute_input":"2022-12-17T05:46:21.366447Z","iopub.status.idle":"2022-12-17T05:46:21.375123Z","shell.execute_reply.started":"2022-12-17T05:46:21.366404Z","shell.execute_reply":"2022-12-17T05:46:21.373895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`machine_id` is unbalanced. id=49 has the largest data and id=197 has the smallest number.","metadata":{}},{"cell_type":"markdown","source":"`machine_id` は不均衡データです。 id=49 のデータが最大で、id=197 のデータが最小です。","metadata":{}},{"cell_type":"code","source":"# 'machine_id'\nsns.countplot(x='machine_id', data=train)\nplt.legend(loc='upper right', title='machine_id')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:46:21.376541Z","iopub.execute_input":"2022-12-17T05:46:21.376982Z","iopub.status.idle":"2022-12-17T05:46:21.613558Z","shell.execute_reply.started":"2022-12-17T05:46:21.376915Z","shell.execute_reply":"2022-12-17T05:46:21.612233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, let's see the mammography data for each `machine_id`.","metadata":{}},{"cell_type":"markdown","source":"次に`machine_id`ごとの画像データを見ます。","metadata":{}},{"cell_type":"code","source":"train_dropduplicates = train.drop_duplicates(subset='machine_id')\ntrain_dropduplicates","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:46:21.615349Z","iopub.execute_input":"2022-12-17T05:46:21.615688Z","iopub.status.idle":"2022-12-17T05:46:21.638776Z","shell.execute_reply.started":"2022-12-17T05:46:21.615658Z","shell.execute_reply":"2022-12-17T05:46:21.637581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 29\nvisualize(10006)\nvisualize(10025)\nvisualize(10048)\ntrain[train['machine_id']==29].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:46:21.640475Z","iopub.execute_input":"2022-12-17T05:46:21.640889Z","iopub.status.idle":"2022-12-17T05:47:11.346891Z","shell.execute_reply.started":"2022-12-17T05:46:21.640846Z","shell.execute_reply":"2022-12-17T05:47:11.345657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 21\nvisualize(10011)\nvisualize(10106)\nvisualize(10144)\ntrain[train['machine_id']==21].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:47:11.348476Z","iopub.execute_input":"2022-12-17T05:47:11.348974Z","iopub.status.idle":"2022-12-17T05:47:27.71405Z","shell.execute_reply.started":"2022-12-17T05:47:11.348937Z","shell.execute_reply":"2022-12-17T05:47:27.712913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 216\nvisualize(10038)\nvisualize(10285)\nvisualize(10302)\ntrain[train['machine_id']==216].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:47:27.715362Z","iopub.execute_input":"2022-12-17T05:47:27.715692Z","iopub.status.idle":"2022-12-17T05:47:40.877168Z","shell.execute_reply.started":"2022-12-17T05:47:27.715662Z","shell.execute_reply":"2022-12-17T05:47:40.875699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 93\nvisualize(10042)\nvisualize(10215)\nvisualize(10391)\ntrain[train['machine_id']==93].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:47:40.878879Z","iopub.execute_input":"2022-12-17T05:47:40.879312Z","iopub.status.idle":"2022-12-17T05:48:01.194222Z","shell.execute_reply.started":"2022-12-17T05:47:40.87927Z","shell.execute_reply":"2022-12-17T05:48:01.193063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 49\nvisualize(10049)\nvisualize(10095)\nvisualize(10097)\ntrain[train['machine_id']==49].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:48:01.195812Z","iopub.execute_input":"2022-12-17T05:48:01.197112Z","iopub.status.idle":"2022-12-17T05:48:32.14635Z","shell.execute_reply.started":"2022-12-17T05:48:01.197048Z","shell.execute_reply":"2022-12-17T05:48:32.144924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 48\nvisualize(10124)\nvisualize(10126)\nvisualize(10136)\ntrain[train['machine_id']==48].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:48:32.147856Z","iopub.execute_input":"2022-12-17T05:48:32.148667Z","iopub.status.idle":"2022-12-17T05:49:13.965944Z","shell.execute_reply.started":"2022-12-17T05:48:32.148627Z","shell.execute_reply":"2022-12-17T05:49:13.964576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 170\nvisualize(10243)\nvisualize(10589)\nvisualize(10668)\ntrain[train['machine_id']==170].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:49:13.973983Z","iopub.execute_input":"2022-12-17T05:49:13.974517Z","iopub.status.idle":"2022-12-17T05:49:42.972373Z","shell.execute_reply.started":"2022-12-17T05:49:13.974469Z","shell.execute_reply":"2022-12-17T05:49:42.971102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 210\nvisualize(10342)\nvisualize(10438)\nvisualize(10741)\ntrain[train['machine_id']==210].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:49:42.973971Z","iopub.execute_input":"2022-12-17T05:49:42.974347Z","iopub.status.idle":"2022-12-17T05:50:37.156809Z","shell.execute_reply.started":"2022-12-17T05:49:42.974312Z","shell.execute_reply":"2022-12-17T05:50:37.155892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 190\nvisualize(10577)\nvisualize(11664)\nvisualize(21809)\ntrain[train['machine_id']==190].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:50:37.158208Z","iopub.execute_input":"2022-12-17T05:50:37.158968Z","iopub.status.idle":"2022-12-17T05:50:50.270374Z","shell.execute_reply.started":"2022-12-17T05:50:37.158923Z","shell.execute_reply":"2022-12-17T05:50:50.269084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 197\nvisualize(13365)\nvisualize(17095)\ntrain[train['machine_id']==197].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:50:50.272113Z","iopub.execute_input":"2022-12-17T05:50:50.273386Z","iopub.status.idle":"2022-12-17T05:51:06.187266Z","shell.execute_reply.started":"2022-12-17T05:50:50.273334Z","shell.execute_reply":"2022-12-17T05:51:06.185931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The characteristics of the image may appear different for each machine. This may be coming from the setting of each machine, so next we see the <font color=red>DICOM file metadata</font>.\n\n画像の特徴はマシンごとに異なって見える場合があります。これは各マシンの設定に起因する可能性があるため、次に <font color=red>DICOM ファイルのメタデータ</font> を確認します。","metadata":{}},{"cell_type":"markdown","source":"# DICOM file meta information\nDICOM file contain some file information.\nEach image contains a rich information that might help us to understand the image deeply and how we split the data for training. Some data might be provided as a train input directly to our model.\n\nDICOM ファイルには、複数のファイル情報が含まれています。\n各画像には、その画像を深く理解したり、学習のためにデータを分割する方法を理解するのに役立つ豊富な情報が含まれる場合があります。 一部のデータは、トレーニング入力としてモデルに直接提供される場合もあります。","metadata":{}},{"cell_type":"code","source":"dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/10038/1967300488.dcm\"\ndcm_info = pydicom.dcmread(dicom_path)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:51:06.203153Z","iopub.execute_input":"2022-12-17T05:51:06.20407Z","iopub.status.idle":"2022-12-17T05:51:06.212876Z","shell.execute_reply.started":"2022-12-17T05:51:06.204026Z","shell.execute_reply":"2022-12-17T05:51:06.211624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dcm_info","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:51:06.214324Z","iopub.execute_input":"2022-12-17T05:51:06.21465Z","iopub.status.idle":"2022-12-17T05:51:06.228923Z","shell.execute_reply.started":"2022-12-17T05:51:06.21462Z","shell.execute_reply":"2022-12-17T05:51:06.227636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"以下に示すように、「.dir()」によって dicom ファイルの「キーワード」を取得できます。","metadata":{}},{"cell_type":"code","source":"dcm_info.dir()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:51:06.230446Z","iopub.execute_input":"2022-12-17T05:51:06.230813Z","iopub.status.idle":"2022-12-17T05:51:06.243381Z","shell.execute_reply.started":"2022-12-17T05:51:06.230778Z","shell.execute_reply":"2022-12-17T05:51:06.242446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"以下のように行タイトルを空白なしで指定することで、メタ情報を取得することができます。","metadata":{}},{"cell_type":"code","source":"dcm_info_list = [\n    dcm_info.SOPInstanceUID,\n    dcm_info.ContentDate,\n    dcm_info.ContentTime,\n    dcm_info.PatientID, \n    dcm_info.BodyPartThickness,\n    dcm_info.CompressionForce,\n    dcm_info.ExposureControlMode,\n    dcm_info.ExposureControlModeDescription,\n    dcm_info.StudyInstanceUID,\n    dcm_info.SeriesInstanceUID,\n    dcm_info.InstanceNumber,\n    dcm_info.ImageLaterality,\n    dcm_info.SamplesPerPixel,\n    dcm_info.PhotometricInterpretation,\n    dcm_info.Rows,\n    dcm_info.Columns,\n    dcm_info.BitsAllocated,\n    dcm_info.BitsStored,\n    dcm_info.HighBit,\n    dcm_info.PixelRepresentation,\n    dcm_info.PixelIntensityRelationship,\n    dcm_info.PixelIntensityRelationshipSign,\n    dcm_info.WindowCenter,\n    dcm_info.WindowWidth,\n    dcm_info.RescaleIntercept,\n    dcm_info.RescaleSlope,\n    dcm_info.RescaleType,\n    dcm_info.VOILUTFunction,\n    dcm_info.LossyImageCompression,\n]","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:51:06.244976Z","iopub.execute_input":"2022-12-17T05:51:06.24531Z","iopub.status.idle":"2022-12-17T05:51:06.255018Z","shell.execute_reply.started":"2022-12-17T05:51:06.245281Z","shell.execute_reply":"2022-12-17T05:51:06.253428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dcm_info_list","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:51:06.258918Z","iopub.execute_input":"2022-12-17T05:51:06.259372Z","iopub.status.idle":"2022-12-17T05:51:06.275773Z","shell.execute_reply.started":"2022-12-17T05:51:06.259334Z","shell.execute_reply":"2022-12-17T05:51:06.274933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"pydicom の公式ドキュメントには、以下のコード例で示される方法でもメタ情報を取得できます。","metadata":{}},{"cell_type":"code","source":"# source : https://pydicom.github.io/pydicom/stable/auto_examples/input_output/plot_printing_dataset.html#sphx-glr-auto-examples-input-output-plot-printing-dataset-py\n\ndef myprint(dataset, indent=0):\n    \"\"\"Go through all items in the dataset and print them with custom format\n\n    Modelled after Dataset._pretty_str()\n    \"\"\"\n    dont_print = ['Pixel Data', 'File Meta Information Version']\n\n    indent_string = \"   \" * indent\n    next_indent_string = \"   \" * (indent + 1)\n\n    for data_element in dataset:\n        if data_element.VR == \"SQ\":   # a sequence\n            print(indent_string, data_element.name)\n            for sequence_item in data_element.value:\n                myprint(sequence_item, indent + 1)\n                print(next_indent_string + \"---------\")\n        else:\n            if data_element.name in dont_print:\n                print(\"\"\"<item not printed -- in the \"don't print\" list>\"\"\")\n            else:\n                repr_value = repr(data_element.value)\n                if len(repr_value) > 50:\n                    repr_value = repr_value[:50] + \"...\"\n                print(\"{0:s} {1:s} = {2:s}\".format(indent_string,\n                                                   data_element.name,\n                                                   repr_value))\n\n\n# Set the dicom file path\nds = pydicom.dcmread(dicom_path)\n\nmyprint(ds)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:51:30.746397Z","iopub.execute_input":"2022-12-17T05:51:30.746853Z","iopub.status.idle":"2022-12-17T05:51:30.765602Z","shell.execute_reply.started":"2022-12-17T05:51:30.746793Z","shell.execute_reply":"2022-12-17T05:51:30.764194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"次に、machine_id 間のメタ情報タグ名の違いを確認します。","metadata":{}},{"cell_type":"code","source":"# pick-up `image_id`s for each `machine_id`s\ntrain_dropdup = train.drop_duplicates(subset='machine_id')\ntrain_dropdup","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:52:09.003325Z","iopub.execute_input":"2022-12-17T05:52:09.003751Z","iopub.status.idle":"2022-12-17T05:52:09.029773Z","shell.execute_reply.started":"2022-12-17T05:52:09.003716Z","shell.execute_reply":"2022-12-17T05:52:09.028602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# difference of meta information tags name between the `machine_id`\n\n# prepare the blank list\ndcm_info_list = []\nmachine_id_list = []\n\nfor patient_id,image_id, machine_id in zip(train_dropdup['patient_id'], train_dropdup['image_id'], train_dropdup['machine_id']):\n    # obtain the keywords for each `machine_id`\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n    keywords = pydicom.dcmread(dicom_path).dir()\n    dcm_info_list.append(keywords)\n    machine_id_list.append(machine_id)\n\n# convert to df\ndcm_info_df = pd.DataFrame(dcm_info_list, index = machine_id_list)\ndcm_info_df","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:52:24.396766Z","iopub.execute_input":"2022-12-17T05:52:24.397219Z","iopub.status.idle":"2022-12-17T05:52:24.483038Z","shell.execute_reply.started":"2022-12-17T05:52:24.397181Z","shell.execute_reply":"2022-12-17T05:52:24.481751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`machine_id`:170 が最も多くの dicom メタデータ中にキーワードを持っていることがわかります。 正し、これは各 machine_id のスポットチェックであるため、先のdfで抽出されたDICOMファイル以外の他のすべてのデータには適用できない場合があることに注意してください。\nただし、<font color=red>machine_idごとの各データには異なる数の dicom メタ情報がある</font>と言えます。","metadata":{}},{"cell_type":"markdown","source":"## Transform DICOM meta information to DataFrame\n次に、これらのメタ情報を取得して、いくつかのトレーニングデータ用のデータフレームを作成しましょう。","metadata":{}},{"cell_type":"code","source":"# difference of meta information tags name between the `machine_id`\n\n# prepare the blank dataframe\ndcm_df = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(train_dropdup['patient_id'], train_dropdup['image_id'], train_dropdup['machine_id']):\n    # obtain the keywords for each `machine_id`\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df = pd.merge(dcm_df, dcm_df_ind, on='name', how='outer')\n\ndcm_df","metadata":{"execution":{"iopub.status.busy":"2022-12-17T06:00:17.160259Z","iopub.execute_input":"2022-12-17T06:00:17.160668Z","iopub.status.idle":"2022-12-17T06:00:17.297103Z","shell.execute_reply.started":"2022-12-17T06:00:17.160637Z","shell.execute_reply":"2022-12-17T06:00:17.295847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"考察:\n- 各データには異なる種類の dicom メタ情報があります。\n- さらなるレポートについては、この [ノートブック (準備中)](xxx) を参照してください。\n\n\n別途、DICOMファイルのメタ情報をDataFrameに変換しました。 興味のある方は[xxx](xxx)をご覧ください。","metadata":{}},{"cell_type":"markdown","source":"Thank you very much for reading, I hope this notebook helps you!\n\n**If you enjoyed the notebook, please upvote! 🙏 Thank you, appreciate your support!**\n","metadata":{}},{"cell_type":"markdown","source":"最後までご覧いただきありがとうございました。まだこのノートブックは今後更新されていく予定ですが、もしここまでで有用だと思いましたら励みになりますのでぜひUpvoteをお願いします😄","metadata":{}}]}