{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## 背景","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"卵巢癌是女性生殖系统最致命的癌症，包含一系列不同的亚型，每种亚型都有独特的特征。 准确识别这些亚型对于制定有效的治疗计划至关重要。 然而，目前对病理学家诊断的依赖在一致性和可及性方面提出了挑战，特别是在服务不足的社区。 利用数据科学为彻底改变卵巢癌诊断并解决这些关键问题提供了一条有前途的途径。 本简介探讨了数据驱动解决方案在改善亚型识别以及随后推进卵巢癌个性化治疗策略方面的潜力。","metadata":{}},{"cell_type":"markdown","source":"该数据集包含与卵巢癌相关的图像，主要分为两种类型：全幻灯片图像（WSI）和组织微阵列（TMA）。 WSI 图像以 20 倍放大倍率捕获，可能会产生较大的文件大小。 另一方面，TMA 的尺寸较小（大约 4,000x4,000 像素），但放大倍率较高，为 40 倍。\n\n在测试集中，与训练集相比，图像来自不同的来源医院。 值得注意的是，测试集中一些最大的图像非常大，尺寸接近 100,000 x 50,000 像素。 必须为不同的场景做好准备，包括图像尺寸、质量、染色技术等的变化。\n\n测试集包含大约 2,000 张图像，其中大部分是 TMA。 整个数据集大小很大，总计 550 GB。 加载数据需要大量时间和资源。\n\n请注意，一些最大的测试集图像可能无法完全装入配备 GPU 的笔记本电脑的内存中。 解决方案正在研究中，预计将于 10 月 18 日那周左右更新。\n\nCSV 文件：对于训练集和测试集，随附的 CSV 文件提供了重要的标签和信息：\n\nimage_id：每个图像的唯一标识符。 标签：指示卵巢癌亚型的目标类别，例如 CC、EC、HGSC、LGSC、MC 或其他。 值得注意的是，“其他”类是测试集独有的，突出了识别异常值的挑战。 image_width：图像的宽度（以像素为单位）。 image_height：图像的高度（以像素为单位）。 is_tma：一个二进制值，指示载玻片是否是组织微阵列。 此信息仅适用于火车组。\n\n此外，数据集还包括一个名为 [train/test]_thumbnails 的文件夹，其中包含整个幻灯片图像的较小 .png 版本。 但是，不为 TMA 提供缩略图。","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport warnings\nwarnings.filterwarnings(\"ignore\")\nimport math","metadata":{"execution":{"iopub.status.busy":"2023-10-26T14:49:31.773391Z","iopub.execute_input":"2023-10-26T14:49:31.773832Z","iopub.status.idle":"2023-10-26T14:49:31.779565Z","shell.execute_reply.started":"2023-10-26T14:49:31.773802Z","shell.execute_reply":"2023-10-26T14:49:31.778552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\n# train.head().style.set_properties(**{'background-color':'green','color':'white','border-color':'#8b8c8c'})\ntrain","metadata":{"execution":{"iopub.status.busy":"2023-10-26T14:49:50.547246Z","iopub.execute_input":"2023-10-26T14:49:50.547844Z","iopub.status.idle":"2023-10-26T14:49:50.570051Z","shell.execute_reply.started":"2023-10-26T14:49:50.547815Z","shell.execute_reply":"2023-10-26T14:49:50.568944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/UBC-OCEAN/test.csv')\ntest","metadata":{"execution":{"iopub.status.busy":"2023-10-26T14:49:54.535128Z","iopub.execute_input":"2023-10-26T14:49:54.535904Z","iopub.status.idle":"2023-10-26T14:49:54.54824Z","shell.execute_reply.started":"2023-10-26T14:49:54.535864Z","shell.execute_reply":"2023-10-26T14:49:54.547301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2023-10-26T14:49:57.645762Z","iopub.execute_input":"2023-10-26T14:49:57.646864Z","iopub.status.idle":"2023-10-26T14:49:57.659968Z","shell.execute_reply.started":"2023-10-26T14:49:57.646828Z","shell.execute_reply":"2023-10-26T14:49:57.658813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Summary statistics for relevant variables\n# styled_data = train.describe().style.background_gradient(cmap='summer').set_properties(**{'text-align':'center','border':'1px solid black'})\nstyled_data = train.describe()\n# display styled data\ndisplay(styled_data)","metadata":{"execution":{"iopub.status.busy":"2023-10-26T14:50:00.397894Z","iopub.execute_input":"2023-10-26T14:50:00.398271Z","iopub.status.idle":"2023-10-26T14:50:00.420696Z","shell.execute_reply.started":"2023-10-26T14:50:00.398245Z","shell.execute_reply":"2023-10-26T14:50:00.419488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Class Distribution\nclass_distribution = train['label'].value_counts()\nprint(class_distribution)\n\n# TMA Distribution\ntma_distribution = train['is_tma'].value_counts()\nprint(tma_distribution)\n\n# Correlation between Image Dimensions\ncorrelation = train[['image_width', 'image_height']].corr()\nprint(correlation)\n\n# Visualization\nplt.figure(figsize=(10, 6))\nsns.scatterplot(x='image_width', y='image_height', data=train, hue='label')\nplt.title('Scatter plot of Image Dimensions', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\nplt.savefig('Scatter plot of Image Dimensions.png')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-26T14:50:03.334572Z","iopub.execute_input":"2023-10-26T14:50:03.335629Z","iopub.status.idle":"2023-10-26T14:50:04.219352Z","shell.execute_reply.started":"2023-10-26T14:50:03.335586Z","shell.execute_reply":"2023-10-26T14:50:04.218094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HGSC = train[train['label']==\"HGSC\"]\nEC = train[train['label']==\"EC\"]\nCC = train[train['label']==\"CC\"]\nLGSC = train[train['label']==\"LGSC\"]\nMC = train[train['label']==\"MC\"]\n\n# set the figure size and font size\nplt.figure(figsize=(12, 6))\nplt.rcParams['font.size'] = 14\n\n# set the colors (I've selected a nice color scheme)\ncolors = ['#66b3ff','#99ff99','#ffcc99','#c2c2f0', '#ffb3e6']\n\n# plot the pie chart for the training set\nplt.subplot(1, 1, 1)\nplt.pie([len(HGSC), len(EC), len(CC), len(LGSC), len(MC)], labels=['HGSC', 'EC', 'CC', 'LGSC', 'MC'], autopct='%1.1f%%', colors=colors)\nplt.title('Training Set', fontsize = 12, fontweight = 'bold', color = 'darkred')\n\nplt.suptitle('Distribution of Subtypes of Ovarian Cancer', fontsize=14,fontweight = 'bold', y=1.05)\n\nplt.savefig('Distribution of Subtypes of Ovarian Cancer.png')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-26T14:50:12.765756Z","iopub.execute_input":"2023-10-26T14:50:12.766173Z","iopub.status.idle":"2023-10-26T14:50:13.068077Z","shell.execute_reply.started":"2023-10-26T14:50:12.766143Z","shell.execute_reply":"2023-10-26T14:50:13.066472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12, 4))\nsns.barplot(x=class_distribution.index, y=class_distribution.values)\nplt.title('Class Distribution', fontsize=14, fontweight='bold', color='darkgreen')\nplt.xlabel('Class Label', fontsize=12, fontweight='bold', color='darkblue')\nplt.ylabel('Count', fontsize=12, fontweight='bold', color='darkblue')\nplt.savefig('Class Distribution.png')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-26T14:50:16.648278Z","iopub.execute_input":"2023-10-26T14:50:16.649171Z","iopub.status.idle":"2023-10-26T14:50:17.08151Z","shell.execute_reply.started":"2023-10-26T14:50:16.649136Z","shell.execute_reply":"2023-10-26T14:50:17.080697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import glob\nfrom matplotlib import pyplot as plt\nfrom matplotlib.image import imread\n\ntrain_data = glob.glob('/kaggle/input/UBC-OCEAN/train_thumbnails/*.png')\ntest_data = glob.glob('/kaggle/input/UBC-OCEAN/test_thumbnails/*.png')\n\nnum_samples = 5\n\nfig, axes = plt.subplots(1, num_samples, figsize=(15, 5))\n\nfor i, image_path in enumerate(train_data[:num_samples]):\n    img = imread(image_path)\n    axes[i].imshow(img)\n    axes[i].axis('off')\n    axes[i].set_title(f'Train Image {i+1}')\n\nplt.tight_layout()\nplt.savefig('Train Iamge.png')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-24T10:19:45.181832Z","iopub.execute_input":"2023-10-24T10:19:45.182307Z","iopub.status.idle":"2023-10-24T10:20:06.453726Z","shell.execute_reply.started":"2023-10-24T10:19:45.182274Z","shell.execute_reply":"2023-10-24T10:20:06.452363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import glob\nfrom matplotlib import pyplot as plt\nfrom matplotlib.image import imread\n\ntrain_data = glob.glob('/kaggle/input/UBC-OCEAN/train_thumbnails/*.png')\ntest_data = glob.glob('/kaggle/input/UBC-OCEAN/test_thumbnails/*.png')\n\nnum_samples = 10\n\nfig, axes = plt.subplots(1, num_samples, figsize=(15, 5))\n\nfor i, image_path in enumerate(train_data[:num_samples]):\n    img = imread(image_path)\n    axes[i].imshow(img)\n    axes[i].axis('off')\n    axes[i].set_title(f'Train Image {i+1}')\n\nplt.tight_layout()\nplt.savefig('Train Iamge.png')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-26T14:46:35.493558Z","iopub.execute_input":"2023-10-26T14:46:35.493968Z","iopub.status.idle":"2023-10-26T14:47:05.300793Z","shell.execute_reply.started":"2023-10-26T14:46:35.493931Z","shell.execute_reply":"2023-10-26T14:47:05.299969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}