{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":45867,"databundleVersionId":6924515,"sourceType":"competition"}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [Reference](https://www.kaggle.com/code/jefersonpazze/eda-baseline)","metadata":{}},{"cell_type":"markdown","source":"# <mark>UBC 난소암 하위 유형 분류 및 이상값 탐지(UBC-OCEAN) - EDA</mark>\n<span style=\"font-size:22px;color:purple\"> 제 노트북을 살펴봐 주셔서 감사합니다 - 조언과 피드백은 언제나 환영입니다!</span>\n\n\n<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana;\">\n    📌 Dataset Link: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\">https://www.kaggle.com/competitions/UBC-OCEAN/data</a>\n</div>\n\n\n## **탐색적 데이터 분석 (EDA)** \nEDA는 UBC 난소암 아형 분류 및 이상값 탐지(UBC-OCEAN)를 포함한 모든 데이터 분석 또는 머신 러닝 프로젝트를 위해 데이터를 이해하고 준비하는 데 있어 중요한 단계입니다. 다음은 이 데이터 세트에 대해 EDA를 수행하는 방법에 대한 단계별 가이드입니다.:\n\n**개요**\n\nUBC 난소암 아형 분류 및 이상값 검출(UBC-OCEAN) 경진대회의 목표는 난소암 아형을 분류하는 것입니다. 20개 이상의 의료 센터에서 얻은 세계에서 가장 광범위한 조직 병리 이미지의 난소암 데이터 세트를 기반으로 학습된 모델을 구축하게 됩니다.\n\n\n**데이터 수집:**\n\n먼저 난소암 하위 유형에 대한 정보와 이상값 감지 데이터가 포함되어야 하는 UBC-OCEAN 데이터 집합을 가져옵니다. 데이터 세트의 구조와 각 변수의 의미를 명확하게 이해해야 합니다.\n\n**데이터 로딩:**\n\n시각화를 위해 pandas, numpy, matplotlib/seaborn과 같은 라이브러리를 사용하여 Python과 같은 원하는 데이터 분석 환경으로 데이터 집합을 가져옵니다.\n\n\n\n\n### **초기 탐색:**\n\n**1 - 먼저 데이터의 기본 특성을 살펴보는 것부터 시작하세요:**\n\n    데이터 처음 몇 행을 확인합니다. df.head().\n    데이터 유형 및 누락된 값을 확인합니다. df.info().\n    기본 통계를 계산합니다. df.describe().\n    데이터 정리:\n\n**2 - 결측치, 이상치, 중복값 처리하기:**\n\n    imputation과 같은 기술을 결측치에 사용.\n    이상값을 적절히 식별하고 처리.\n    필요한 경우 중복 행 제거.\n    \n**3 - 데이터 시각화:**\n\n    데이터에 대한 인사이트를 얻기 위해 시각화\n    numerical features를 위한 히스토그램 및 박스 플롯.\n    categorical features를 위한 막대형 그래프.\n    변수 간의 관계를 파악하기 위한 상관관계 행렬 및 분산형 차트.\n    \n**4 - 피처 분석:**\n\n    분류 및 이상치 탐지를 위해 features 와 target 간의 관계를 탐색합니다.\n    다양한 하위 유형 또는 클래스에 따라 서로 다른 feature들이 어떻게 달라지는지 시각화하기.\n    박스 플롯, 바이올린 플롯 또는 군집 플롯을 사용하여 특징 분포를 비교합니다.\n\n**5 - 이상치 탐지:**\n\n    데이터 세트에 이상값 탐지와 관련된 정보가 포함되어 있는 경우, 이 부분에 대한 전용 EDA를 수행하세요.:\n    산점도 또는 박스 플롯을 사용하여 이상값 시각화하기.\n    통계적 방법이나 머신러닝 기법을 적용하여 이상값을 식별합니다.\n\n**6 - 차원 축소(선택 사항):**\n\n    데이터 집합에 많은 feature들이 있는 경우, 중요한 정보를 보존하면서 변수 수를 줄이기 위해 주성분 분석(PCA)과 같은 차원 축소 기법을 고려하세요.\n\n**8 - 요약 및 인사이트:**\n\n    관찰된 패턴, 추세 또는 이상 징후를 포함하여 EDA에서 얻은 결과를 요약하세요..\n    적용된 모든 데이터 전처리 단계를 문서화합니다.\n\n**7 - 다음 단계:**\n\n    A 결과를 바탕으로 피처 엔지니어링, 모델 선택, 추가 데이터 전처리 등 다음 단계를 계획합니다.\n\n\nEDA는 반복적인 프로세스이며, 데이터 세트를 더 깊이 파고들어 머신러닝 또는 데이터 분석 모델을 개발할 때 이러한 단계를 다시 살펴봐야 할 수도 있습니다.","metadata":{}},{"cell_type":"code","source":"%%capture \n# 이 Python 3 환경에는 여러 가지 유용한 분석 라이브러리가 설치되어 있습니다.\n# 다음과 같이 정의됩니다. kaggle/python Docker image: https://github.com/kaggle/docker-python\n# 예를 들어 로드할 수 있는 몇 가지 유용한 패키지는 다음과 같습니다.\n\nimport numpy as np # 선형 대수\nimport pandas as pd # 데이터 처리, CSV 파일 I/O(예: pd.read_csv)\n\n# 입력 데이터 파일은 읽기 전용으로 사용할 수 있습니다. \"../input/\" directory\n# 예를 들어, 실행을 클릭하거나 Shift+Enter를 눌러 실행하면 입력 디렉토리 아래에 있는 모든 파일이 나열됩니다.\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# \"모두 저장 및 실행\"을 사용하여 버전을 만들 때 출력으로 보존되는 현재 디렉터리(/kaggle/working/)에 최대 20GB까지 쓸 수 있습니다. \n# 임시 파일을 /kaggle/temp/에 쓸 수도 있지만 현재 세션 외부에 저장되지 않습니다.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-12-03T11:41:48.241228Z","iopub.execute_input":"2023-12-03T11:41:48.241967Z","iopub.status.idle":"2023-12-03T11:41:48.771706Z","shell.execute_reply.started":"2023-12-03T11:41:48.24193Z","shell.execute_reply":"2023-12-03T11:41:48.770896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!pip install scikit-image","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:28:58.718303Z","iopub.execute_input":"2023-10-07T04:28:58.719155Z","iopub.status.idle":"2023-10-07T04:28:58.72344Z","shell.execute_reply.started":"2023-10-07T04:28:58.719122Z","shell.execute_reply":"2023-10-07T04:28:58.722185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nfrom skimage import io\nimport os\nimport seaborn as sns\nimport cv2\nimport random\nimport os\nimport glob","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:41:55.93442Z","iopub.execute_input":"2023-12-03T11:41:55.935488Z","iopub.status.idle":"2023-12-03T11:41:56.910777Z","shell.execute_reply.started":"2023-12-03T11:41:55.935445Z","shell.execute_reply":"2023-12-03T11:41:56.909991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 학습데이터 불러오기\ntrain_df = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:42:13.930029Z","iopub.execute_input":"2023-12-03T11:42:13.930414Z","iopub.status.idle":"2023-12-03T11:42:13.962187Z","shell.execute_reply.started":"2023-12-03T11:42:13.930383Z","shell.execute_reply":"2023-12-03T11:42:13.961374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:42:17.743953Z","iopub.execute_input":"2023-12-03T11:42:17.744323Z","iopub.status.idle":"2023-12-03T11:42:17.758005Z","shell.execute_reply.started":"2023-12-03T11:42:17.744285Z","shell.execute_reply":"2023-12-03T11:42:17.756931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.shape)","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:42:21.07894Z","iopub.execute_input":"2023-12-03T11:42:21.079825Z","iopub.status.idle":"2023-12-03T11:42:21.084535Z","shell.execute_reply.started":"2023-12-03T11:42:21.079791Z","shell.execute_reply":"2023-12-03T11:42:21.083555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isna().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:42:21.962285Z","iopub.execute_input":"2023-12-03T11:42:21.963124Z","iopub.status.idle":"2023-12-03T11:42:21.969726Z","shell.execute_reply.started":"2023-12-03T11:42:21.963095Z","shell.execute_reply":"2023-12-03T11:42:21.968773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:purple\">null 값 없음</span>\n​\n<div class=\"alert alert-block alert-info\">\n<b>좋습니다:</b> 이제 다음 단계로 넘어가겠습니다. </div>","metadata":{}},{"cell_type":"code","source":"# 배포를 원하는 열의 목록입니다.\ncolumns = [col for col in train_df.columns if col!='label']\n\n# 루프를 사용하여 각 열을 반복합니다.\nfor col in columns:\n    \n    # 3개 열에 대한 하위 플롯(3개 플롯)\n    fig, axs = plt.subplots(figsize=(15,5), ncols=3)\n    \n    # 첫 번째 플롯 - 샘플 데이터 세트의 분포\n    sns.histplot(data=train_df, x=col, kde=True, ax=axs[0])\n    axs[0].set_title('Sample Distribution')\n    \n    # 두 번째 플롯 - 결과가 1(당뇨병 있음)인 선택된 열의 분포입니다.\n    sns.histplot(data=train_df[train_df['label']==\"HGSC\"], x=col, kde=True, ax=axs[1], color='orange')\n    axs[1].set_title('label - HGSC')\n    \n    # 세 번째 플롯 - 결과가 0(당뇨병이 없음)인 선택된 열의 분포입니다.\n    sns.histplot(data=train_df[train_df['label']==\"LGSC\"], x=col, kde=True, ax=axs[2], color='green')\n    axs[2].set_title('label - LGSC')\n    \n    # 플롯 보기\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:44:28.765981Z","iopub.execute_input":"2023-12-03T11:44:28.766605Z","iopub.status.idle":"2023-12-03T11:44:32.358138Z","shell.execute_reply.started":"2023-12-03T11:44:28.766569Z","shell.execute_reply":"2023-12-03T11:44:32.357227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:44:34.883863Z","iopub.execute_input":"2023-12-03T11:44:34.884211Z","iopub.status.idle":"2023-12-03T11:44:34.898037Z","shell.execute_reply.started":"2023-12-03T11:44:34.884184Z","shell.execute_reply":"2023-12-03T11:44:34.897041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.image_id.is_unique)\nprint(train_df.label.is_unique)","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:44:35.911586Z","iopub.execute_input":"2023-12-03T11:44:35.912231Z","iopub.status.idle":"2023-12-03T11:44:35.91786Z","shell.execute_reply.started":"2023-12-03T11:44:35.912198Z","shell.execute_reply":"2023-12-03T11:44:35.916923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"이미지 크기를 탐색한다는 것은 이미지의 크기와 해상도의 다양한 측면을 이해하고 작업하는 것을 의미합니다. 디지털 이미지의 맥락에서 고려해야 할 몇 가지 주요 차원과 속성이 있습니다:\n\n1. **해상도(Resolution)**: 해상도는 이미지에 포함된 픽셀(개별 색상 점)의 수를 나타냅니다. 일반적으로 인치당 픽셀 수(PPI) 또는 인치당 도트 수(DPI)로 표현됩니다. 해상도가 높은 이미지는 디테일이 풍부하여 인쇄에 적합하고, 해상도가 낮은 이미지는 웹 디스플레이나 화면 보기에 사용할 수 있습니다.\n\n2. **픽셀 크기(Pixel Dimensions)**: 픽셀 크기는 이미지의 너비와 높이를 픽셀 단위로 지정합니다. 예를 들어 이미지의 크기가 1920x1080픽셀인 경우, 이는 폭이 1920픽셀이고 높이가 1080픽셀임을 의미합니다. 이는 일반적으로 디지털 이미지의 크기를 지정하는 데 사용됩니다.\n\n3. **화면비(Aspect Ratio)**: 화면비(aspect ratio)는 이미지의 가로와 세로의 비율입니다. 일반적인 종횡비에는 4:3(표준 텔레비전), 16:9(와이드스크린 텔레비전), 1:1(정사각형)이 있습니다. 이미지 왜곡을 방지하려면 올바른 화면비를 유지하는 것이 중요합니다.\n\n4. **물리적 크기(Physical Dimensions)**: 물리적 크기는 실제 세계에서 인쇄되거나 표시될 때 이미지의 크기를 나타냅니다. 이는 픽셀 크기와 해상도 모두에 의해 결정됩니다. 예를 들어, 크기가 3000x2000픽셀(300DPI)인 이미지가 인쇄되면 10x6.67인치입니다.\n\n5. **파일 크기(File Size)**: 이미지의 파일 크기는 바이트 또는 킬로바이트(KB), 메가바이트(MB) 등으로 측정됩니다. 색심도, 압축 및 픽셀 크기와 같은 요소에 따라 달라집니다. 세부 사항이 많은 큰 이미지일수록 파일 크기가 커지는 경향이 있습니다.\n\n6. **색 심도(Color Depth)**: 비트 심도라고도 하는 색 심도(Color depth)는 픽셀이 표현할 수 있는 색의 수를 결정합니다. 일반적인 색 심도에는 8비트(256색), 24비트(트루 컬러), 32비트(투명도를 위한 알파 채널이 있는 트루 컬러)가 있습니다.\n\n7. **DPI vs. PPI**: DPI(인치당 도트 수)는 프린터가 선형 인치에서 생성할 수 있는 잉크 도트의 수를 나타내는 것으로, 인쇄와 관련하여 자주 사용됩니다. PPI(인치당 픽셀 수)는 화면 디스플레이에 사용되며, 1인치 화면 공간의 픽셀 수를 나타냅니다.\n\n8. **스케일링(Scaling)**: 이미지 크기를 조정하려면 이미지의 크기를 다른 크기로 조정해야 합니다. 이미지의 크기를 늘리거나(확대) 줄일 수 있습니다. 특히 이미지를 크게 만들 때 크기를 너무 많이 조정하면 이미지 품질이 저하될 수 있다는 점에 유의하세요.\n\n9. **자르기(Cropping)**: 자르기에는 특정 영역에 초점을 맞추기 위해 이미지의 일부를 잘라내는 작업이 포함됩니다. 이렇게 하면 이미지의 픽셀 크기와 종횡비가 변경됩니다.\n\n10. **압축(Compression)**: 이미지 압축은 중복되거나 덜 중요한 데이터를 제거하여 파일 크기를 줄입니다. 무손실(품질 손실 없음) 또는 손실(일부 품질 손실) 압축이 가능합니다. JPEG와 같은 일반적인 이미지 형식은 손실 압축을 사용합니다.\n\n이러한 이미지 크기와 속성을 이해하고 관리하는 것은 그래픽 디자인, 사진, 웹 개발, 인쇄 등 다양한 목적을 위해 필수적입니다. 특정 요구 사항에 따라 원하는 결과를 얻기 위해 이러한 크기와 속성을 적절히 조정해야 할 수도 있습니다.","metadata":{}},{"cell_type":"code","source":"train_df[['image_height', 'image_width']].describe()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:51:21.709915Z","iopub.execute_input":"2023-12-03T11:51:21.710311Z","iopub.status.idle":"2023-12-03T11:51:21.728553Z","shell.execute_reply.started":"2023-12-03T11:51:21.710282Z","shell.execute_reply":"2023-12-03T11:51:21.727588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.image_id.shape[0])\nprint(len(os.listdir('/kaggle/input/UBC-OCEAN/train_images')))\nprint(len(os.listdir('/kaggle/input/UBC-OCEAN/train_thumbnails')))","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:51:23.051948Z","iopub.execute_input":"2023-12-03T11:51:23.052296Z","iopub.status.idle":"2023-12-03T11:51:23.059548Z","shell.execute_reply.started":"2023-12-03T11:51:23.052268Z","shell.execute_reply":"2023-12-03T11:51:23.058582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import imageio\n","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:51:23.777873Z","iopub.execute_input":"2023-12-03T11:51:23.778639Z","iopub.status.idle":"2023-12-03T11:51:23.782575Z","shell.execute_reply.started":"2023-12-03T11:51:23.778605Z","shell.execute_reply":"2023-12-03T11:51:23.781589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom tqdm.notebook import tqdm\n\nimport albumentations as A\nimport matplotlib.image as mpimg\nimport imageio\nimport scipy.ndimage as ndi\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:51:43.467644Z","iopub.execute_input":"2023-12-03T11:51:43.468Z","iopub.status.idle":"2023-12-03T11:51:43.474387Z","shell.execute_reply.started":"2023-12-03T11:51:43.467971Z","shell.execute_reply":"2023-12-03T11:51:43.473407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:51:44.125663Z","iopub.execute_input":"2023-12-03T11:51:44.126018Z","iopub.status.idle":"2023-12-03T11:51:44.130386Z","shell.execute_reply.started":"2023-12-03T11:51:44.125991Z","shell.execute_reply":"2023-12-03T11:51:44.129401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.label.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:51:47.988025Z","iopub.execute_input":"2023-12-03T11:51:47.988407Z","iopub.status.idle":"2023-12-03T11:51:47.996037Z","shell.execute_reply.started":"2023-12-03T11:51:47.988375Z","shell.execute_reply":"2023-12-03T11:51:47.995224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HGSC = train_df[train_df['label']==\"HGSC\"]\nEC = train_df[train_df['label']==\"EC\"]\nCC = train_df[train_df['label']==\"CC\"]\nLGSC = train_df[train_df['label']==\"LGSC\"]\nMC = train_df[train_df['label']==\"MC\"]","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:51:48.943587Z","iopub.execute_input":"2023-12-03T11:51:48.944319Z","iopub.status.idle":"2023-12-03T11:51:48.953644Z","shell.execute_reply.started":"2023-12-03T11:51:48.944289Z","shell.execute_reply":"2023-12-03T11:51:48.95279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 그림 크기 설정\nplt.figure(figsize=(20, 6))\n\n# 글꼴 크기 설정\nplt.rcParams['font.size'] = 14\n\n# 색상 설정\ncolors = ['lightgreen', 'lightblue', 'purple', 'blue', 'yellow']\n\n# training 세트에 대한 파이 차트 그리기\nplt.subplot(1, 1, 1)\nplt.pie([len(HGSC), len(EC), len(CC), len(LGSC), len(MC)], labels=['HGSC', 'EC', 'CC', 'LGSC', 'MC'], autopct='%1.1f%%', colors=colors)\nplt.title('Training Set')\n\n\n# 그림에 기본 제목 추가\nplt.suptitle('Distribution of HGSC, EC, CC, LGSC and MC Images in the Training data', fontsize=20, y=1.05)\n\n# 플롯 보기\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:52:46.930941Z","iopub.execute_input":"2023-12-03T11:52:46.931581Z","iopub.status.idle":"2023-12-03T11:52:47.153004Z","shell.execute_reply.started":"2023-12-03T11:52:46.931551Z","shell.execute_reply":"2023-12-03T11:52:47.15143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train 및 test 데이터의 데이터 분할 살펴보기\ntrain_data = glob.glob('/kaggle/input/UBC-OCEAN/train_images/*.png')\ntest_data = glob.glob('/kaggle/input/UBC-OCEAN/test_images/*.png')\n\nprint(f\"The Training Set contains: {len(train_data)} images\")\nprint(f\"The Testing Set contains: {len(test_data)} images\")","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:53:12.83367Z","iopub.execute_input":"2023-12-03T11:53:12.834426Z","iopub.status.idle":"2023-12-03T11:53:12.844604Z","shell.execute_reply.started":"2023-12-03T11:53:12.834395Z","shell.execute_reply":"2023-12-03T11:53:12.843496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 총 개수 계산\ntotal_train = len(train_data)\ntotal_test = len(test_data)\n\n# 그림 크기 설정\nplt.figure(figsize=(4, 4))\n\n# 글꼴 크기 설정\nplt.rcParams['font.size'] = 12\n\n# 색상 설정\ncolors = ['lightgreen', 'red']\n\n# 전체 집합에 대한 원형 차트 그리기\nplt.pie([total_train, total_test], labels=['Training Set', 'Testing Set'], autopct='%1.1f%%', colors=colors)\nplt.title('Distribution of Images in Training and Testing Sets')\n\n# 플롯 보기\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:53:52.237106Z","iopub.execute_input":"2023-12-03T11:53:52.237484Z","iopub.status.idle":"2023-12-03T11:53:52.354321Z","shell.execute_reply.started":"2023-12-03T11:53:52.237452Z","shell.execute_reply":"2023-12-03T11:53:52.353403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 데이터 시각화 📈 \n\n### 이미지 무작위 시각화","metadata":{}},{"cell_type":"code","source":"# 샘플 train 이미지\n\nio.imshow('/kaggle/input/UBC-OCEAN/train_thumbnails/10642_thumbnail.png')","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:54:22.497517Z","iopub.execute_input":"2023-12-03T11:54:22.497871Z","iopub.status.idle":"2023-12-03T11:54:23.671608Z","shell.execute_reply.started":"2023-12-03T11:54:22.497847Z","shell.execute_reply":"2023-12-03T11:54:23.670615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HGSC = train_df[train_df['label']==\"HGSC\"]\nEC = train_df[train_df['label']==\"EC\"]\nCC = train_df[train_df['label']==\"CC\"]\nLGSC = train_df[train_df['label']==\"LGSC\"]\nMC = train_df[train_df['label']==\"MC\"]","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:54:24.721477Z","iopub.execute_input":"2023-12-03T11:54:24.721835Z","iopub.status.idle":"2023-12-03T11:54:24.734433Z","shell.execute_reply.started":"2023-12-03T11:54:24.721809Z","shell.execute_reply":"2023-12-03T11:54:24.733394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"!pip install torch-summary\n!pip install torch-lr-finder\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:54:28.213884Z","iopub.execute_input":"2023-12-03T11:54:28.214273Z","iopub.status.idle":"2023-12-03T11:54:28.220737Z","shell.execute_reply.started":"2023-12-03T11:54:28.214243Z","shell.execute_reply":"2023-12-03T11:54:28.21953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"import os\nimport numpy as np\nimport cv2\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\n%matplotlib inline\nfrom PIL import Image\nfrom IPython.display import display\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import DataLoader\nimport torch.nn.functional as F\nfrom torchvision import datasets, transforms, models\nfrom torch.optim.lr_scheduler import StepLR\nfrom torchsummary import summary\nfrom tqdm import tqdm\nimport torchvision.models as models\nimport PIL\nfrom torchvision.models import resnet50, ResNet50_Weights\nimport timm\nfrom torch.utils.data import Dataset\nimport glob\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:54:28.60864Z","iopub.execute_input":"2023-12-03T11:54:28.609002Z","iopub.status.idle":"2023-12-03T11:54:28.615744Z","shell.execute_reply.started":"2023-12-03T11:54:28.608972Z","shell.execute_reply":"2023-12-03T11:54:28.614751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"torch.manual_seed(0)\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:54:29.613684Z","iopub.execute_input":"2023-12-03T11:54:29.614173Z","iopub.status.idle":"2023-12-03T11:54:29.621109Z","shell.execute_reply.started":"2023-12-03T11:54:29.614128Z","shell.execute_reply":"2023-12-03T11:54:29.620251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"data_path = '/kaggle/input/UBC-OCEAN/train_images'\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:54:30.268721Z","iopub.execute_input":"2023-12-03T11:54:30.269081Z","iopub.status.idle":"2023-12-03T11:54:30.27555Z","shell.execute_reply.started":"2023-12-03T11:54:30.269052Z","shell.execute_reply":"2023-12-03T11:54:30.274542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"class_name = os.listdir(data_path)\nprint(class_name)\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:54:31.01483Z","iopub.execute_input":"2023-12-03T11:54:31.015165Z","iopub.status.idle":"2023-12-03T11:54:31.020957Z","shell.execute_reply.started":"2023-12-03T11:54:31.015138Z","shell.execute_reply":"2023-12-03T11:54:31.019981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"class_map = {\"HGSC\" : 0, \"EC\": 1, \"CC\": 2, \"LGSC\": 3, \"MC\": 4}\nrev_class_map = {0 : \"HGSC\", 1 : \"EC\", 2 : \"CC\", 3 : \"LGSC\", 4 : \"MC\"}\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:54:34.245211Z","iopub.execute_input":"2023-12-03T11:54:34.246055Z","iopub.status.idle":"2023-12-03T11:54:34.251714Z","shell.execute_reply.started":"2023-12-03T11:54:34.246019Z","shell.execute_reply":"2023-12-03T11:54:34.250623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Activation,Conv2D, Flatten, Dropout, MaxPooling2D, BatchNormalization\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom keras import regularizers, optimizers\nimport os\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:54:37.0663Z","iopub.execute_input":"2023-12-03T11:54:37.067144Z","iopub.status.idle":"2023-12-03T11:54:47.971249Z","shell.execute_reply.started":"2023-12-03T11:54:37.067113Z","shell.execute_reply":"2023-12-03T11:54:47.970396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf=pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv',dtype=str)\ntraindf","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:59:25.377172Z","iopub.execute_input":"2023-12-03T11:59:25.378478Z","iopub.status.idle":"2023-12-03T11:59:25.396534Z","shell.execute_reply.started":"2023-12-03T11:59:25.378441Z","shell.execute_reply":"2023-12-03T11:59:25.395555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testdf=pd.read_csv('/kaggle/input/UBC-OCEAN/test.csv',dtype=str)\ntestdf","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:59:25.651472Z","iopub.execute_input":"2023-12-03T11:59:25.652178Z","iopub.status.idle":"2023-12-03T11:59:25.665564Z","shell.execute_reply.started":"2023-12-03T11:59:25.652145Z","shell.execute_reply":"2023-12-03T11:59:25.664483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"datagen=ImageDataGenerator(rescale=1./255.,validation_split=0.25)\ntrain_generator=datagen.flow_from_dataframe(\ndataframe=traindf,\ndirectory=\"/kaggle/input/UBC-OCEAN/train_images/\",\nx_col=\"image_id\",\ny_col=\"label\",\nsubset=\"training\",\nbatch_size=32,\nseed=42,\nshuffle=True,\nclass_mode=\"categorical\",\ntarget_size=(100,100))","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:59:29.988562Z","iopub.execute_input":"2023-12-03T11:59:29.988908Z","iopub.status.idle":"2023-12-03T11:59:30.001918Z","shell.execute_reply.started":"2023-12-03T11:59:29.988883Z","shell.execute_reply":"2023-12-03T11:59:30.000912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def append_ext(fn):\n    return fn+\".png\"\n\ndef append_ext_thum(fn):\n    return fn+\"_thumbnail.png\"\n\n\ntraindf[\"image_id_path\"]=traindf[\"image_id\"].apply(append_ext)\ntraindf[\"image_id_path_thum\"]=traindf[\"image_id\"].apply(append_ext_thum)\n\n\ntestdf[\"image_id_path\"]=testdf[\"image_id\"].apply(append_ext)\ntestdf[\"image_id_path_thum\"]=testdf[\"image_id\"].apply(append_ext_thum)\n","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:59:31.6693Z","iopub.execute_input":"2023-12-03T11:59:31.670094Z","iopub.status.idle":"2023-12-03T11:59:31.67897Z","shell.execute_reply.started":"2023-12-03T11:59:31.670057Z","shell.execute_reply":"2023-12-03T11:59:31.677902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:59:33.186889Z","iopub.execute_input":"2023-12-03T11:59:33.187657Z","iopub.status.idle":"2023-12-03T11:59:33.202595Z","shell.execute_reply.started":"2023-12-03T11:59:33.187625Z","shell.execute_reply":"2023-12-03T11:59:33.201697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testdf","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:59:33.956913Z","iopub.execute_input":"2023-12-03T11:59:33.957919Z","iopub.status.idle":"2023-12-03T11:59:33.967559Z","shell.execute_reply.started":"2023-12-03T11:59:33.957886Z","shell.execute_reply":"2023-12-03T11:59:33.966593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testdf['Image_path'] = [os.path.join('/kaggle/input/UBC-OCEAN/test_images', image) for image in testdf['image_id_path']]\ntestdf['Image_path_thumbnails'] = [os.path.join('/kaggle/input/UBC-OCEAN/test_thumbnails', image) for image in testdf['image_id_path_thum']]\ntestdf","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:59:34.772441Z","iopub.execute_input":"2023-12-03T11:59:34.772807Z","iopub.status.idle":"2023-12-03T11:59:34.786474Z","shell.execute_reply.started":"2023-12-03T11:59:34.772778Z","shell.execute_reply":"2023-12-03T11:59:34.785519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf['Image_path'] = [os.path.join('/kaggle/input/UBC-OCEAN/train_images', image) for image in traindf['image_id_path']]\ntraindf['Image_path_thumbnails'] = [os.path.join('/kaggle/input/UBC-OCEAN/train_thumbnails', image) for image in traindf['image_id_path_thum']]\ntraindf","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:59:35.381141Z","iopub.execute_input":"2023-12-03T11:59:35.381623Z","iopub.status.idle":"2023-12-03T11:59:35.403558Z","shell.execute_reply.started":"2023-12-03T11:59:35.381589Z","shell.execute_reply":"2023-12-03T11:59:35.402192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_path_random = np.random.choice(traindf['Image_path_thumbnails'],5)\nfull_path_random","metadata":{"execution":{"iopub.status.busy":"2023-12-03T11:59:35.939616Z","iopub.execute_input":"2023-12-03T11:59:35.939954Z","iopub.status.idle":"2023-12-03T11:59:35.947829Z","shell.execute_reply.started":"2023-12-03T11:59:35.939929Z","shell.execute_reply":"2023-12-03T11:59:35.946974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def image_viewer(dataset, index, ax):\n    image_path =  dataset['Image_path_thumbnails'][index]\n    image      =  Image.open(image_path)\n    ax.imshow(image)\n    \ndef plot_some_images(dataset, title):\n    fig, axs = plt.subplots(nrows = 1,ncols = 2,figsize=(20,8))\n    for ind, ax in enumerate(axs.flat):\n            index = random.randrange(len(dataset))\n            image_viewer(dataset, index, ax)\n            ax.set_title(dataset['label'][index], fontsize = 8)\n            ax.axis('off')\n            fig.suptitle(title, fontsize = 15)\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:00:46.357466Z","iopub.execute_input":"2023-12-03T12:00:46.357885Z","iopub.status.idle":"2023-12-03T12:00:46.364935Z","shell.execute_reply.started":"2023-12-03T12:00:46.357855Z","shell.execute_reply":"2023-12-03T12:00:46.363967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_some_images(traindf, 'Trainig Images')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfrom sklearn.metrics import confusion_matrix, classification_report\nfrom sklearn.utils.class_weight import compute_class_weight","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:00:51.709503Z","iopub.execute_input":"2023-12-03T12:00:51.710152Z","iopub.status.idle":"2023-12-03T12:00:51.839119Z","shell.execute_reply.started":"2023-12-03T12:00:51.710122Z","shell.execute_reply":"2023-12-03T12:00:51.83836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_weights = compute_class_weight(class_weight = \"balanced\",\n                                     classes= np.unique(traindf['label']),\n                                     y= traindf['label'])\n\nclasses = (np.unique(traindf['label']))\nclass_weights_forplot = dict(zip(classes, class_weights))","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:00:52.81495Z","iopub.execute_input":"2023-12-03T12:00:52.815589Z","iopub.status.idle":"2023-12-03T12:00:52.822925Z","shell.execute_reply.started":"2023-12-03T12:00:52.815556Z","shell.execute_reply":"2023-12-03T12:00:52.821918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"classes","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:00:53.769829Z","iopub.execute_input":"2023-12-03T12:00:53.770675Z","iopub.status.idle":"2023-12-03T12:00:53.776032Z","shell.execute_reply.started":"2023-12-03T12:00:53.770645Z","shell.execute_reply":"2023-12-03T12:00:53.775145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_weights_forplot","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:00:54.399931Z","iopub.execute_input":"2023-12-03T12:00:54.400599Z","iopub.status.idle":"2023-12-03T12:00:54.407005Z","shell.execute_reply.started":"2023-12-03T12:00:54.400562Z","shell.execute_reply":"2023-12-03T12:00:54.40604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_weights = dict(zip(range(43), class_weights))\n","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:00:55.005161Z","iopub.execute_input":"2023-12-03T12:00:55.006042Z","iopub.status.idle":"2023-12-03T12:00:55.010157Z","shell.execute_reply.started":"2023-12-03T12:00:55.006005Z","shell.execute_reply":"2023-12-03T12:00:55.009148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **모델 구축**\n","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\n","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:27.013186Z","iopub.execute_input":"2023-12-03T12:01:27.014038Z","iopub.status.idle":"2023-12-03T12:01:27.017867Z","shell.execute_reply.started":"2023-12-03T12:01:27.014004Z","shell.execute_reply":"2023-12-03T12:01:27.016993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_generator = tf.keras.preprocessing.image.ImageDataGenerator(\n    preprocessing_function=tf.keras.applications.efficientnet.preprocess_input,\n)\ntest_generator = tf.keras.preprocessing.image.ImageDataGenerator(\n    preprocessing_function=tf.keras.applications.efficientnet.preprocess_input\n)","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:27.738136Z","iopub.execute_input":"2023-12-03T12:01:27.738538Z","iopub.status.idle":"2023-12-03T12:01:27.744072Z","shell.execute_reply.started":"2023-12-03T12:01:27.738501Z","shell.execute_reply":"2023-12-03T12:01:27.742935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_images = train_generator.flow_from_dataframe(\n    dataframe=traindf,\n    x_col='Image_path_thumbnails',\n    y_col= 'label',\n    target_size=(256, 256),\n    color_mode='grayscale',\n    class_mode=\"categorical\",\n    batch_size=64,\n    shuffle=True,\n    seed=210,\n)\n\ntest_images = test_generator.flow_from_dataframe(\n    dataframe=traindf[0:10],\n    x_col='Image_path_thumbnails',\n    y_col= 'label',\n    target_size=(256, 256),\n    class_mode=\"categorical\",\n    color_mode='grayscale',\n    batch_size=64,\n    shuffle=False\n)","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:28.413963Z","iopub.execute_input":"2023-12-03T12:01:28.414353Z","iopub.status.idle":"2023-12-03T12:01:28.678089Z","shell.execute_reply.started":"2023-12-03T12:01:28.414311Z","shell.execute_reply":"2023-12-03T12:01:28.677274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trans_arc = tf.keras.applications.EfficientNetB0(weights = \"imagenet\", include_top = False,\n                         input_shape=(256, 256,3), pooling='max')\nfor l in trans_arc.layers:\n    l.trainable = False\ninputs = trans_arc.input\nflatten = trans_arc.output\n\nx = tf.keras.layers.Dense(256, activation='relu')(flatten)\nx = tf.keras.layers.BatchNormalization()(x)\nx = tf.keras.layers.Dropout(0.3)(x)\n\nx = tf.keras.layers.Dense(128, activation='relu')(x)\nx = tf.keras.layers.BatchNormalization()(x)\n\noutputs = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n\n\nmodel = tf.keras.Model(inputs=inputs, outputs=outputs)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()\n","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:55.900419Z","iopub.status.idle":"2023-12-03T12:01:55.900756Z","shell.execute_reply.started":"2023-12-03T12:01:55.900591Z","shell.execute_reply":"2023-12-03T12:01:55.900607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.mkdir(\"/kaggle/working/checkpoints\")\ncb_csvlogger = tf.keras.callbacks.CSVLogger(\n                                            filename='/content/checkpoints/training_log.csv',\n                                            separator=',',\n                                            append=False)","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:55.902516Z","iopub.status.idle":"2023-12-03T12:01:55.902883Z","shell.execute_reply.started":"2023-12-03T12:01:55.90271Z","shell.execute_reply":"2023-12-03T12:01:55.902728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"loss = [tf.keras.losses.binary_crossentropy]\n\ninitial_learning_rate = 0.005\n\nlr_schedule = tf.keras.optimizers.schedules.ExponentialDecay(\n    initial_learning_rate,\n    decay_steps=82,\n    decay_rate=0.9,\n    staircase=True)\n\noptimizer = tf.keras.optimizers.Adam(\n    learning_rate= lr_schedule,\n    beta_1=0.9,\n    beta_2=0.999,\n    epsilon=1e-07,\n)\nmetrics= ['accuracy']\n\nmodel.compile(\n    optimizer=optimizer,\n    loss= loss,\n    metrics=metrics\n    )","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:55.904313Z","iopub.status.idle":"2023-12-03T12:01:55.904698Z","shell.execute_reply.started":"2023-12-03T12:01:55.904528Z","shell.execute_reply":"2023-12-03T12:01:55.904545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_images","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:55.905589Z","iopub.status.idle":"2023-12-03T12:01:55.905929Z","shell.execute_reply.started":"2023-12-03T12:01:55.905763Z","shell.execute_reply":"2023-12-03T12:01:55.905779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.python.client import device_lib\nprint(device_lib.list_local_devices())","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:55.907263Z","iopub.status.idle":"2023-12-03T12:01:55.907736Z","shell.execute_reply.started":"2023-12-03T12:01:55.907496Z","shell.execute_reply":"2023-12-03T12:01:55.907519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# https://www.tensorflow.org/guide/gpu?hl=pt-br\nwith tf.device(\"/device:GPU:0\"):\n\n    history = model.fit(\n      train_images,\n      validation_data=test_images,\n      verbose = True,\n      epochs=5,\n      class_weight = class_weights,\n      callbacks=[\n          tf.keras.callbacks.LearningRateScheduler(lr_schedule),\n          tf.keras.callbacks.EarlyStopping(\n              monitor='val_loss',\n              patience=7,\n              restore_best_weights=True),\n          #cb_csvlogger\n      ]\n  )","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:55.909417Z","iopub.status.idle":"2023-12-03T12:01:55.909861Z","shell.execute_reply.started":"2023-12-03T12:01:55.909629Z","shell.execute_reply":"2023-12-03T12:01:55.909651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 기록에 있는 모든 데이터 나열\nprint(history.history.keys())\n# summarize history for accuracy\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])  # RAISE ERROR\nplt.title('model accuracy')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validation'], loc='upper left')\nplt.show()\n# summarize history for loss\nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss']) #RAISE ERROR\nplt.title('model loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validation'], loc='upper left')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-03T12:01:55.910954Z","iopub.status.idle":"2023-12-03T12:01:55.911441Z","shell.execute_reply.started":"2023-12-03T12:01:55.91119Z","shell.execute_reply":"2023-12-03T12:01:55.911213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **모델 평가**","metadata":{}},{"cell_type":"code","source":"def plot_model_evaluation(model, test_data, n_classes, target_labels):\n\n    results = model.evaluate(test_data, verbose=0)\n    loss = results[0]\n    acc = results[1]\n\n    print(\"    Test Loss: {:.5f}\".format(loss))\n    print(\"Test Accuracy: {:.2f}%\".format(acc * 100))\n\n    y_pred = np.squeeze((model.predict(test_data) >= 0.5).astype(int))\n    cm = confusion_matrix(test_data.labels, y_pred)\n    clr = classification_report(test_data.labels, y_pred, target_names=target_labels)\n\n    plt.figure(figsize=(15, 15))\n    sns.heatmap(cm, annot=True, fmt='g', vmin=0, cmap='Blues', cbar=False)\n    plt.xticks(ticks=np.arange(n_classes) + 0.5, labels=list(test_data.class_indices.keys()), rotation=90)\n    plt.yticks(ticks=np.arange(n_classes) + 0.5, labels=list(test_data.class_indices.keys()), rotation=0)\n    plt.xlabel(\"Predicted\")\n    plt.ylabel(\"Actual\")\n    plt.title(\"Confusion Matrix\")\n    plt.show()\n\n    print(\"Classification Report:\\n----------------------\\n\", clr)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model_evaluation(model, test_images, 3, traindf[0:10]['label'].unique())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**친애하는 여러분,**\n\n**이 메시지가 여러분에게 잘 전달되기를 바랍니다. UBC 난소암 아형 분류 및 이상치 탐지(UBC-OCEAN) 경연대회에 참가한다는 소식을 전하게 되어 기쁘게 생각하며, 여러분의 지원이 필요합니다.!**\n\n**이 대회의 일환으로, 저는 암 아형 분류 및 이상값 탐지를 위한 효과적인 솔루션 개발의 기본 단계인 데이터 세트에 대한 중요한 통찰력을 얻기 위해 심층적인 탐색적 데이터 분석(EDA)을 수행했습니다. 이제, 제가 제출한 EDA에 대한 여러분의 투표와 지원을 요청하고자 합니다..**\n\n**여러분의 투표는 이 대회에서 큰 변화를 가져올 수 있으며, 제가 다음 단계로 나아가는 데 도움이 될 수 있습니다. 저를 지지하는 방법은 다음과 같습니다.:**\n\n\n\n**1. 투표하기:**\n\n대회 플랫폼을 방문하여 내 EDA 제출물을 찾습니다.\n'vote' 또는 'support' 버튼을 클릭하여 투표합니다..\n\n**2. 네트워크와 공유:**\n\n제 일을 지원하는 데 관심이 있는 친구, 가족, 동료에게 널리 알려주세요.\n\n**3. 피드백 제공:**\n\n제 EDA에 대한 피드백이나 제안 사항이 있으시면 언제든지 공유해 주세요. 여러분의 의견은 소중하며 제가 개선하는 데 도움이 될 수 있습니다.\n저는 암 연구 분야에 긍정적인 영향을 미치기 위해 최선을 다하고 있으며, 여러분의 지원으로 목표 달성에 한 걸음 더 다가갈 수 있을 것입니다.\n\n시간을 내어 이 메시지를 읽어주셔서 감사드리며, 이번 대회에 보내주신 여러분의 성원에 진심으로 감사드립니다. 우리는 함께 난소암 퇴치에 기여하고 데이터 기반 의료 분야를 발전시킬 수 있습니다.\n\n제 EDA에 대해 궁금한 점이 있거나 더 많은 정보가 필요하시면 언제든지 저에게 연락해 주세요. 여러분의 성원은 저에게 큰 힘이 됩니다!\n\nWarm regards,\nJeferson S. Pazze","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}