{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":39272,"databundleVersionId":4629629,"sourceType":"competition"},{"sourceId":4866520,"sourceType":"datasetVersion","datasetId":2820722}],"dockerImageVersionId":31011,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **A Beacon of Hope: Harnessing AI for Breast Cancer Detection**","metadata":{}},{"cell_type":"markdown","source":"### **Chapter 1: The Call to Action**\n**Every year, millions of women undergo mammography screening. For many, the results bring comfort and reassurance; for others, the diagnosis of breast cancer changes their lives. Radiologists like Dr. Amal Ali have long dedicated themselves to early detection, knowing that accurate diagnosis can save lives. Yet, despite her expertise, Dr. Amal knew that even the most skilled eyes can sometimes miss subtle signs—or mistakenly raise false alarms that cause undue stress.**\n\n`What if there were an assistant, a tireless ally, that could help reduce both missed cases and false positives?`\n \n    — Dr. Amal Ali  \n\n**Driven by this vision, Dr. Amal partnered with a team of data scientists to create an AI-powered tool. Their mission was clear: develop a model that could sift through thousands of radiographic breast images and accurately flag potential cases of cancer, while minimizing unnecessary worry.**","metadata":{}},{"cell_type":"markdown","source":"### **Chapter 2: Unveiling the Dataset**\n**The team embarked on their journey by assembling a rich [dataset](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/data) of radiographic breast images. The data was organized as follows:**\n- `Data Source:` train_images/[patient_id]/[image_id].dcm\n- `Subjects:` Female patients undergoing screening exams\n- `Patients:` Approximately 8,000 unique patients\n- `Images per Patient:` Usually, but not always, 4 images\n- `Image Format:` **DICOM** (with many images encoded in JPEG 2000)\n- `train.csv:` a csv file that have corresponding labels (0 for benign, 1 for malignant).\n\n**This diverse and challenging dataset held the promise of training a robust model. However, the team knew that working with DICOM files—and handling JPEG 2000 images in particular—would require special care.**","metadata":{}},{"cell_type":"markdown","source":"### **Chapter 3: Preparing the Data 👨‍💻**\n**Before the AI could learn to detect cancer, the images needed to be carefully processed. Using the pydicom library, the team set out to load and preprocess the data. They also accounted for the fact that some images were stored in JPEG 2000 format, ensuring the proper libraries were in place.**","metadata":{}},{"cell_type":"markdown","source":"**At first he imported needed modules**","metadata":{}},{"cell_type":"code","source":"import os\nfrom PIL import Image\nimport pydicom\n\n# import other libraries\n# import system libs\nimport os\nimport time\nimport shutil\nimport pathlib\nimport itertools\nfrom pathlib import Path\nimport multiprocessing as mp\nfrom tqdm.notebook import tqdm\nfrom joblib import Parallel, delayed\n\n# import data handling tools\nimport cv2\nimport pydicom\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nsns.set_style('darkgrid')\nimport matplotlib.pyplot as plt\n\n# import Deep learning Libraries\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Activation, Dropout, BatchNormalization\nfrom tensorflow.keras.models import Model, load_model, Sequential\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom sklearn.metrics import confusion_matrix, classification_report\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.optimizers import Adam, Adamax\nfrom tensorflow.keras import regularizers\nfrom tensorflow.keras.metrics import categorical_crossentropy\n\n# Ignore Warnings\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nprint ('modules loaded')","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true,"execution":{"iopub.status.busy":"2025-04-28T00:56:30.162555Z","iopub.execute_input":"2025-04-28T00:56:30.16306Z","iopub.status.idle":"2025-04-28T00:56:44.206064Z","shell.execute_reply.started":"2025-04-28T00:56:30.163039Z","shell.execute_reply":"2025-04-28T00:56:44.205263Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\n# المسار لمجلد الصور\nimage_folder = '/kaggle/input/rsna-bcd-1024x512-preprocessed/'\n\n# اعرض أول 5 فايلات عشان نتأكد\nprint(os.listdir(image_folder)[:5])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T23:13:42.635061Z","iopub.execute_input":"2025-04-27T23:13:42.635325Z","iopub.status.idle":"2025-04-27T23:13:42.642527Z","shell.execute_reply.started":"2025-04-27T23:13:42.635306Z","shell.execute_reply":"2025-04-27T23:13:42.641972Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Then the team read data and store it in dataframe. The data was in a folder 📂 `train_images` and there is a `train.csv` file to get the labels, so the team stored images into `df`.**","metadata":{}},{"cell_type":"code","source":"# استيراد المكتبات المطلوبة\nimport pandas as pd\nimport os\n\n# تحديد المسارات\ndataset_path = '/kaggle/input/rsna-bcd-1024x512-preprocessed'  # هنا تحطي المسار الأساسي للداتا عندك\ncsv_path = os.path.join(dataset_path, '/kaggle/input/rsna-breast-cancer-detection/train.csv')\nimages_path = os.path.join(dataset_path, '/kaggle/input/rsna-bcd-1024x512-preprocessed/train_images')\n\n# قراءة ملف train.csv\ndf = pd.read_csv(csv_path)\n\n# إنشاء عمود جديد لمسار كل صورة\ndf['image_path'] = df['image_id'].apply(lambda x: os.path.join(images_path, f\"{x}.png\"))\n\n# عرض أول 5 صفوف للتأكد أن كل حاجة صح\nprint(df.head())\n\n# كمان نطبع عدد الصور واللابيلز لو تحبي\nprint(f\"Number of images: {len(df)}\")\nprint(f\"Labels distribution:\\n{df['cancer'].value_counts()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T00:56:54.609273Z","iopub.execute_input":"2025-04-28T00:56:54.609891Z","iopub.status.idle":"2025-04-28T00:56:54.779157Z","shell.execute_reply.started":"2025-04-28T00:56:54.609858Z","shell.execute_reply":"2025-04-28T00:56:54.778372Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# مسار الداتا الجديدة اللي فيها الصور المعالجة\ndata_path = '/kaggle/input/rsna-bcd-1024x512-preprocessed'\ncsv_path = os.path.join(data_path, '/kaggle/input/rsna-breast-cancer-detection/train.csv')\nimages_folder = os.path.join(data_path, '/kaggle/input/rsna-bcd-1024x512-preprocessed/train_images')\n\n# قراءة ملف اللابلز\ndf = pd.read_csv(csv_path)\n\n# إنشاء عمود مسارات الصور\ndf['image_path'] = df['patient_id'].astype(str) + '/' + df['image_id'].astype(str) + '.png'  # امتداد الصور png مش jpg\ndf['image_path'] = df['image_path'].apply(lambda x: os.path.join(images_folder, x))\n\n# نظرة سريعة\nprint(df[['patient_id', 'image_id', 'image_path']].head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T00:57:09.226841Z","iopub.execute_input":"2025-04-28T00:57:09.227555Z","iopub.status.idle":"2025-04-28T00:57:09.375953Z","shell.execute_reply.started":"2025-04-28T00:57:09.227529Z","shell.execute_reply":"2025-04-28T00:57:09.375339Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# استيراد المكتبات المطلوبة\nimport pandas as pd\nimport os\n\n# تحديد المسارات\nimport os\nimport pandas as pd\n\n# المسار الأساسي للداتا\ndataset_path = '/kaggle/input/rsna-bcd-1024x512-preprocessed'\ncsv_path = os.path.join(dataset_path, '/kaggle/input/rsna-breast-cancer-detection/train.csv')  # تم تعديل المسار هنا\nimages_path = os.path.join(dataset_path, '/kaggle/input/rsna-bcd-1024x512-preprocessed/train_images')  # تم تعديل المسار هنا\n\n# قراءة ملف train.csv\ndf = pd.read_csv(csv_path)\n\n# إنشاء عمود جديد لمسار كل صورة باستخدام \"image_id\" و \"patient_id\"\ndf['image_path'] = df['patient_id'].astype(str) + '/' + df['image_id'].astype(str) + '.png'\ndf['image_path'] = df['image_path'].apply(lambda x: os.path.join(images_path, x))\n# عرض أول 5 صفوف للتأكد من المسارات\nprint(df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T00:57:17.641866Z","iopub.execute_input":"2025-04-28T00:57:17.642135Z","iopub.status.idle":"2025-04-28T00:57:17.785466Z","shell.execute_reply.started":"2025-04-28T00:57:17.642113Z","shell.execute_reply":"2025-04-28T00:57:17.784708Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**After that they split `df` into `train_df`, `valid_df`, and `test_df` with 80% for training and 10% for validation and 10% for testing**","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\n# أول حاجة: نقسم إلى train (80%) و (validation + test) (20%)\ntrain_df, temp_df = train_test_split(df, test_size=0.2, random_state=42, stratify=df['cancer'])\n\n# ثاني حاجة: نقسم الـ temp_df بالتساوي إلى validation و test (كل واحد 10%)\nvalid_df, test_df = train_test_split(temp_df, test_size=0.5, random_state=42, stratify=temp_df['cancer'])\n\n# طباعة أحجام كل مجموعة\nprint(f\"Train size: {len(train_df)}\")\nprint(f\"Validation size: {len(valid_df)}\")\nprint(f\"Test size: {len(test_df)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T00:57:26.362946Z","iopub.execute_input":"2025-04-28T00:57:26.36322Z","iopub.status.idle":"2025-04-28T00:57:26.405566Z","shell.execute_reply.started":"2025-04-28T00:57:26.363201Z","shell.execute_reply":"2025-04-28T00:57:26.404802Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**At this time he was able to convert train_df, valid_df, and test_df into tensors with `batch_size = 32` for train and `batch_size = 16` for valid and test, and `img_size = (512, 512)`**","metadata":{}},{"cell_type":"code","source":"# إعداد التحويلات transformations\nfrom torchvision import transforms\nfrom torch.utils.data import Dataset\nfrom torch.utils.data import DataLoader\nimport torch\nimg_size = (256, 256)\n\ntrain_transform = transforms.Compose([\n    transforms.Resize(img_size),\n    transforms.RandomHorizontalFlip(),\n    transforms.ToTensor(),\n])\n\nvalid_test_transform = transforms.Compose([\n    transforms.Resize(img_size),\n    transforms.ToTensor(),\n])\n\n# كلاس الداتا الخاص بينا\nclass BreastCancerDataset(Dataset):\n    def __init__(self, df, transform=None):\n        self.df = df.reset_index(drop=True)\n        self.transform = transform\n        \n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, idx):\n        img_path = self.df.loc[idx, 'image_path']\n        label = self.df.loc[idx, 'cancer']\n        \n        # فتح الصورة\n        image = Image.open(img_path).convert('RGB')\n        \n        if self.transform:\n            image = self.transform(image)\n        \n        return image, torch.tensor(label, dtype=torch.float32)\n\n# إنشاء الداتاستس\ntrain_dataset = BreastCancerDataset(train_df, transform=train_transform)\nvalid_dataset = BreastCancerDataset(valid_df, transform=valid_test_transform)\ntest_dataset = BreastCancerDataset(test_df, transform=valid_test_transform)\n\n# إنشاء الداتالودرز\ntrain_loader = DataLoader(train_dataset, batch_size=16, shuffle=True, num_workers=2, pin_memory=True)\nvalid_loader = DataLoader(valid_dataset, batch_size=8, shuffle=False, num_workers=2, pin_memory=True)\ntest_loader = DataLoader(test_dataset, batch_size=8, shuffle=False, num_workers=2, pin_memory=True)\n\n# اختبار الشكل\nfor images, labels in train_loader:\n    print(f\"Images batch shape: {images.shape}\")\n    print(f\"Labels batch shape: {labels.shape}\")\n    break","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T00:57:33.713074Z","iopub.execute_input":"2025-04-28T00:57:33.713361Z","iopub.status.idle":"2025-04-28T00:57:41.192785Z","shell.execute_reply.started":"2025-04-28T00:57:33.713339Z","shell.execute_reply":"2025-04-28T00:57:41.191848Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**To show a sample of data, they tried to use a python script to just showing a sample of the data**","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport numpy as np\n\n# Function to show a batch of images\ndef show_sample(data_loader):\n    images, labels = next(iter(data_loader))  # Get one batch\n    images = images.numpy()\n\n    plt.figure(figsize=(12, 6))\n    for i in range(8):  # عرض أول 8 صور\n        plt.subplot(2, 4, i+1)\n        img = images[i].transpose((1, 2, 0))  # رجعنا الشكل (HWC)\n        plt.imshow(img)\n        plt.title(f\"Label: {labels[i].item()}\")\n        plt.axis('off')\n    plt.tight_layout()\n    plt.show()\n\n# Example: Show sample from training data\nshow_sample(train_loader)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T00:58:07.387831Z","iopub.execute_input":"2025-04-28T00:58:07.388384Z","iopub.status.idle":"2025-04-28T00:58:09.440387Z","shell.execute_reply.started":"2025-04-28T00:58:07.388359Z","shell.execute_reply":"2025-04-28T00:58:09.439573Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Chapter 4: Forging the AI Model 🤖**\n**With the data prepped, the team moved on to building the AI model. They chose a Convolutional Neural Network (CNN) architecture specially `EfficientNetB5` as it is a well-suited for image classification tasks like this. Given the critical nature of the diagnosis, the model needed to be both sensitive to true cases of cancer and cautious enough to avoid excessive false positives.**","metadata":{}},{"cell_type":"code","source":"import torch.nn as nn\nimport torchvision.models as models\n\n# Load a pre-trained EfficientNetB5\nmodel = models.efficientnet_b5(weights='IMAGENET1K_V1')\n\n#Modify the classifier to match the number of classes (2 classes: 0 and 1)\nmodel.classifier[1] = nn.Linear(model.classifier[1].in_features, 2)\n\n# Move model to GPU if available\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel = model.to(device)\n\n# Define Loss function and Optimizer\ncriterion = nn.CrossEntropyLoss()\noptimizer = torch.optim.Adam(model.parameters(), lr=1e-4)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T00:58:18.550938Z","iopub.execute_input":"2025-04-28T00:58:18.551591Z","iopub.status.idle":"2025-04-28T00:58:20.742848Z","shell.execute_reply.started":"2025-04-28T00:58:18.551566Z","shell.execute_reply":"2025-04-28T00:58:20.74225Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import torch\nimport torch.optim as optim\nimport torch.nn as nn\nfrom torch.utils.data import DataLoader\nfrom tqdm import tqdm\n\n# تحديد الجهاز (GPU أو CPU)\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# تعريف النموذج (هنا يتم استخدام EfficientNet كمثال)\nmodel = models.efficientnet_b0(pretrained=True)\nmodel.classifier[1] = nn.Linear(model.classifier[1].in_features, 1)  # تصنيف ثنائي (مريض / غير مريض)\nmodel = model.to(device)\n\n# تعريف الـ Loss Function\ncriterion = nn.BCEWithLogitsLoss()\n\n# تعريف الـ Optimizer\noptimizer = optim.Adam(model.parameters(), lr=1e-4)\n\n# عدد الإيبوكس (epochs)\nepochs = 5\n\n# تدريب النموذج\nfor epoch in range(epochs):\n    model.train()  # وضع النموذج في وضع التدريب\n    running_loss = 0.0\n    \n    # استخدام tqdm لعرض شريط التقدم أثناء التدريب\n    pbar = tqdm(train_loader, desc=f\"Epoch {epoch+1}/{epochs}\")\n    \n    for images, labels in pbar:\n        # نقل البيانات إلى الجهاز (GPU أو CPU)\n        images, labels = images.to(device), labels.to(device).float().unsqueeze(1)  # تأكد أن labels في شكل [batch_size, 1]\n        \n        optimizer.zero_grad()  # إعادة تعيين التدرجات\n        outputs = model(images)  # تمرير الصور عبر النموذج\n        \n        loss = criterion(outputs, labels)  # حساب الخسارة\n        loss.backward()  # حساب التدرجات\n        optimizer.step()  # تحديث الأوزان\n        \n        running_loss += loss.item()  # جمع الخسارة\n        \n        # تحديث شريط التقدم بعرض الخسارة\n        pbar.set_postfix({'loss': running_loss / (pbar.n + 1)})\n    \n    # طباعة الخسارة في نهاية كل epoch\n    print(f\"Epoch [{epoch+1}/{epochs}] Loss: {running_loss/len(train_loader):.4f}\")\n\n    # تقييم الأداء على بيانات التحقق بعد كل epoch\n    model.eval()  # وضع النموذج في وضع التحقق\n    correct = 0\n    total = 0\n    with torch.no_grad():  # لا حاجة لحساب التدرجات أثناء التحقق\n        for images, labels in valid_loader:\n            # نقل البيانات إلى الجهاز\n            images, labels = images.to(device), labels.to(device).float().unsqueeze(1)\n            outputs = model(images)\n            \n            # تطبيق دالة sigmoide وتحويل المخرجات إلى تنبؤات ثنائية\n            preds = torch.sigmoid(outputs) > 0.5\n            \n            correct += (preds == labels).sum().item()  # حساب عدد التنبؤات الصحيحة\n            total += labels.size(0)  # إجمالي عدد البيانات\n            \n    # حساب الدقة\n    acc = correct / total\n    print(f\"Validation Accuracy after Epoch {epoch+1}: {acc:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T01:04:22.119056Z","iopub.execute_input":"2025-04-28T01:04:22.119787Z","iopub.status.idle":"2025-04-28T01:35:36.902118Z","shell.execute_reply.started":"2025-04-28T01:04:22.11976Z","shell.execute_reply":"2025-04-28T01:35:36.901227Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# وضع النموذج في وضع التقييم (eval)\nmodel.eval()  \ncorrect = 0\ntotal = 0\n\n# عدم الحاجة لحساب التدرجات أثناء التحقق\nwith torch.no_grad():\n    for images, labels in test_loader:\n        # نقل البيانات إلى الجهاز (GPU أو CPU)\n        images, labels = images.to(device), labels.to(device).float().unsqueeze(1)\n        \n        # تمرير البيانات عبر النموذج\n        outputs = model(images)\n        \n        # تطبيق دالة sigmoide لتحويل المخرجات إلى تنبؤات\n        preds = torch.sigmoid(outputs) > 0.5  # التنبؤ إذا كانت النتيجة 0 أو 1\n        \n        # حساب عدد التنبؤات الصحيحة\n        correct += (preds == labels).sum().item()\n        total += labels.size(0)  # إجمالي عدد البيانات\n\n# حساب دقة النموذج على بيانات الاختبار\nacc = correct / total\nprint(f\"Test Accuracy: {acc:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T01:39:11.1241Z","iopub.execute_input":"2025-04-28T01:39:11.124456Z","iopub.status.idle":"2025-04-28T01:40:05.871254Z","shell.execute_reply.started":"2025-04-28T01:39:11.124431Z","shell.execute_reply":"2025-04-28T01:40:05.870493Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**This neural network was designed to extract hierarchical features from the mammograms, allowing the model to learn the subtle differences that indicate cancerous changes.**","metadata":{}},{"cell_type":"markdown","source":"### **Chapter 5: Training the Champion**\n**The next step was the battle of training. The model was fed batches of preprocessed images along with their labels. With each epoch, the model learned to differentiate between healthy tissue and suspicious findings. Special care was taken to monitor performance, ensuring that false positives were minimized.**","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Chapter 6: The Triumph and Beyond**\n**After rigorous training and validation, the model was put to the test on unseen data. Its performance was promising—a testament to the tireless work of Dr. Amal and her team. The AI assistant demonstrated its potential to serve as a second pair of eyes, supporting radiologists in making more accurate diagnoses.😀**","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Epilogue: A New Dawn of Hope**\n**In the quiet hum of the radiology lab, as mammograms cycled through the imaging machines, a new kind of vigilance took shape. The AI assistant, born from collaboration and fueled by data, was more than just an algorithm—it was a beacon of hope for countless women.**\n\n**Dr. Amal often reflected on the journey:**\n    \n`\"Every image processed, every diagnosis aided, brings us one step closer to a future where early detection is the norm rather than the exception. This is not just technology; it's a promise of better care and a brighter tomorrow.\"`","metadata":{}},{"cell_type":"markdown","source":"#### **And so, the story continues—each breakthrough and every line of code reinforcing the commitment to save lives, one image at a time.❤️**","metadata":{}}]}