{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":39272,"databundleVersionId":4629629,"sourceType":"competition"},{"sourceId":4866520,"sourceType":"datasetVersion","datasetId":2820722}],"dockerImageVersionId":31011,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **A Beacon of Hope: Harnessing AI for Breast Cancer Detection**","metadata":{}},{"cell_type":"markdown","source":"### **Chapter 1: The Call to Action**\n**Every year, millions of women undergo mammography screening. For many, the results bring comfort and reassurance; for others, the diagnosis of breast cancer changes their lives. Radiologists like Dr. Amal Ali have long dedicated themselves to early detection, knowing that accurate diagnosis can save lives. Yet, despite her expertise, Dr. Amal knew that even the most skilled eyes can sometimes miss subtle signs—or mistakenly raise false alarms that cause undue stress.**\n\n`What if there were an assistant, a tireless ally, that could help reduce both missed cases and false positives?`\n \n    — Dr. Amal Ali  \n\n**Driven by this vision, Dr. Amal partnered with a team of data scientists to create an AI-powered tool. Their mission was clear: develop a model that could sift through thousands of radiographic breast images and accurately flag potential cases of cancer, while minimizing unnecessary worry.**","metadata":{}},{"cell_type":"markdown","source":"### **Chapter 2: Unveiling the Dataset**\n**The team embarked on their journey by assembling a rich [dataset](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/data) of radiographic breast images. The data was organized as follows:**\n- `Data Source:` train_images/[patient_id]/[image_id].dcm\n- `Subjects:` Female patients undergoing screening exams\n- `Patients:` Approximately 8,000 unique patients\n- `Images per Patient:` Usually, but not always, 4 images\n- `Image Format:` **DICOM** (with many images encoded in JPEG 2000)\n- `train.csv:` a csv file that have corresponding labels (0 for benign, 1 for malignant).\n\n**This diverse and challenging dataset held the promise of training a robust model. However, the team knew that working with DICOM files—and handling JPEG 2000 images in particular—would require special care.**","metadata":{}},{"cell_type":"markdown","source":"### **Chapter 3: Preparing the Data 👨‍💻**\n**Before the AI could learn to detect cancer, the images needed to be carefully processed. Using the pydicom library, the team set out to load and preprocess the data. They also accounted for the fact that some images were stored in JPEG 2000 format, ensuring the proper libraries were in place.**","metadata":{}},{"cell_type":"markdown","source":"**At first he imported needed modules**","metadata":{}},{"cell_type":"code","source":"# 📌 Step 1: Import needed libraries\nimport os\nimport pandas as pd\nimport numpy as np\nimport pydicom\nfrom PIL import Image\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nimport torch.nn as nn\nimport torchvision.models as models\nimport torch.optim as optim\nfrom tqdm import tqdm\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import train_test_split\nimport timm  # For pre-trained models like EfficientNet\n","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:22:58.331975Z","iopub.execute_input":"2025-04-27T22:22:58.332245Z","iopub.status.idle":"2025-04-27T22:23:06.71324Z","shell.execute_reply.started":"2025-04-27T22:22:58.332225Z","shell.execute_reply":"2025-04-27T22:23:06.712496Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**And here, you can convert `.dcm` files into `.jpg` files, you can skip this step if you need and use this preprocessed [datatset](https://www.kaggle.com/datasets/paulbacher/rsna-bcd-1024x512-preprocessed/data)**","metadata":{}},{"cell_type":"code","source":"\nlabels_df = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:23:28.696199Z","iopub.execute_input":"2025-04-27T22:23:28.697068Z","iopub.status.idle":"2025-04-27T22:23:28.815843Z","shell.execute_reply.started":"2025-04-27T22:23:28.697042Z","shell.execute_reply":"2025-04-27T22:23:28.81493Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"base_image_path = '/kaggle/input/rsna-bcd-1024x512-preprocessed/train_images'\n\nlabels_df['image_path'] = base_image_path + '/' + labels_df['patient_id'].astype(str) + '/' + labels_df['image_id'].astype(str) + '.png'\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:23:31.870037Z","iopub.execute_input":"2025-04-27T22:23:31.870533Z","iopub.status.idle":"2025-04-27T22:23:31.929263Z","shell.execute_reply.started":"2025-04-27T22:23:31.870509Z","shell.execute_reply":"2025-04-27T22:23:31.928756Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nlabels_df['exists'] = labels_df['image_path'].apply(lambda x: os.path.exists(x))\nprint(labels_df['exists'].value_counts())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:23:35.47636Z","iopub.execute_input":"2025-04-27T22:23:35.476921Z","iopub.status.idle":"2025-04-27T22:25:55.80844Z","shell.execute_reply.started":"2025-04-27T22:23:35.476898Z","shell.execute_reply":"2025-04-27T22:25:55.807744Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Then the team read data and store it in dataframe. The data was in a folder 📂 `train_images` and there is a `train.csv` file to get the labels, so the team stored images into `df`.**","metadata":{}},{"cell_type":"code","source":"\n# مسار الداتا الجديدة اللي فيها الصور المعالجة\ndata_path = '/kaggle/input/rsna-bcd-1024x512-preprocessed'\ncsv_path = os.path.join(data_path, '/kaggle/input/rsna-breast-cancer-detection/train.csv')\nimages_folder = os.path.join(data_path, '/kaggle/input/rsna-bcd-1024x512-preprocessed/train_images')\n\n# قراءة ملف اللابلز\ndf = pd.read_csv(csv_path)\n\n# إنشاء عمود مسارات الصور\ndf['image_path'] = df['patient_id'].astype(str) + '/' + df['image_id'].astype(str) + '.png'  # امتداد الصور png مش jpg\ndf['image_path'] = df['image_path'].apply(lambda x: os.path.join(images_folder, x))\n\n# نظرة سريعة\nprint(df[['patient_id', 'image_id', 'image_path']].head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:26:35.40108Z","iopub.execute_input":"2025-04-27T22:26:35.401549Z","iopub.status.idle":"2025-04-27T22:26:35.549737Z","shell.execute_reply.started":"2025-04-27T22:26:35.401523Z","shell.execute_reply":"2025-04-27T22:26:35.549137Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**After that they split `df` into `train_df`, `valid_df`, and `test_df` with 80% for training and 10% for validation and 10% for testing**","metadata":{}},{"cell_type":"code","source":"\n# أولاً: تقسيم 80% تدريب و20% (هقسمهم لاحقًا لـ val/test)\ntrain_df, temp_df = train_test_split(df, test_size=0.2, random_state=42, stratify=df['cancer'])\n\n# ثانياً: تقسيم 20% الباقيين إلى 10% فاليديشن و10% تيست\nvalid_df, test_df = train_test_split(temp_df, test_size=0.5, random_state=42, stratify=temp_df['cancer'])\n\n# طباعة حجم كل مجموعة للتأكد\nprint(f\"Train size: {len(train_df)}\")\nprint(f\"Validation size: {len(valid_df)}\")\nprint(f\"Test size: {len(test_df)}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:26:40.924565Z","iopub.execute_input":"2025-04-27T22:26:40.92518Z","iopub.status.idle":"2025-04-27T22:26:40.966264Z","shell.execute_reply.started":"2025-04-27T22:26:40.925153Z","shell.execute_reply":"2025-04-27T22:26:40.965546Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**At this time he was able to convert train_df, valid_df, and test_df into tensors with `batch_size = 32` for train and `batch_size = 16` for valid and test, and `img_size = (512, 512)`**","metadata":{}},{"cell_type":"code","source":"\n# إعداد التحويلات transformations\nimg_size = (256, 256)\n\ntrain_transform = transforms.Compose([\n    transforms.Resize(img_size),\n    transforms.RandomHorizontalFlip(),\n    transforms.ToTensor(),\n])\n\nvalid_test_transform = transforms.Compose([\n    transforms.Resize(img_size),\n    transforms.ToTensor(),\n])\n\n# كلاس الداتا الخاص بينا\nclass BreastCancerDataset(Dataset):\n    def __init__(self, df, transform=None):\n        self.df = df.reset_index(drop=True)\n        self.transform = transform\n        \n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, idx):\n        img_path = self.df.loc[idx, 'image_path']\n        label = self.df.loc[idx, 'cancer']\n        \n        # فتح الصورة\n        image = Image.open(img_path).convert('RGB')\n        \n        if self.transform:\n            image = self.transform(image)\n        \n        return image, torch.tensor(label, dtype=torch.float32)\n\n# إنشاء الداتاستس\ntrain_dataset = BreastCancerDataset(train_df, transform=train_transform)\nvalid_dataset = BreastCancerDataset(valid_df, transform=valid_test_transform)\ntest_dataset = BreastCancerDataset(test_df, transform=valid_test_transform)\n\n# إنشاء الداتالودرز\ntrain_loader = DataLoader(train_dataset, batch_size=16, shuffle=True, num_workers=2, pin_memory=True)\nvalid_loader = DataLoader(valid_dataset, batch_size=8, shuffle=False, num_workers=2, pin_memory=True)\ntest_loader = DataLoader(test_dataset, batch_size=8, shuffle=False, num_workers=2, pin_memory=True)\n\n# اختبار الشكل\nfor images, labels in train_loader:\n    print(f\"Images batch shape: {images.shape}\")\n    print(f\"Labels batch shape: {labels.shape}\")\n    break\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:26:46.613939Z","iopub.execute_input":"2025-04-27T22:26:46.614201Z","iopub.status.idle":"2025-04-27T22:26:47.432762Z","shell.execute_reply.started":"2025-04-27T22:26:46.614182Z","shell.execute_reply":"2025-04-27T22:26:47.432003Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**To show a sample of data, they tried to use a python script to just showing a sample of the data**","metadata":{}},{"cell_type":"code","source":"\n# دالة لعرض مجموعة صور مع اللابلز\ndef show_batch(loader):\n    images, labels = next(iter(loader))  # نجيب أول batch\n    images = images[:8]  # نعرض أول 8 صور مثلاً\n    labels = labels[:8]\n    \n    plt.figure(figsize=(16, 8))\n    for i in range(len(images)):\n        img = images[i].permute(1, 2, 0).numpy()  # تحويل من (C, H, W) إلى (H, W, C)\n        plt.subplot(2, 4, i + 1)\n        plt.imshow(img)\n        plt.title(f'Label: {int(labels[i].item())}')\n        plt.axis('off')\n    plt.show()\n\n# استدعاء الدالة\nshow_batch(train_loader)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:26:52.926103Z","iopub.execute_input":"2025-04-27T22:26:52.926878Z","iopub.status.idle":"2025-04-27T22:26:54.308166Z","shell.execute_reply.started":"2025-04-27T22:26:52.926849Z","shell.execute_reply":"2025-04-27T22:26:54.307264Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Chapter 4: Forging the AI Model 🤖**\n**With the data prepped, the team moved on to building the AI model. They chose a Convolutional Neural Network (CNN) architecture specially `EfficientNetB5` as it is a well-suited for image classification tasks like this. Given the critical nature of the diagnosis, the model needed to be both sensitive to true cases of cancer and cautious enough to avoid excessive false positives.**","metadata":{}},{"cell_type":"code","source":"\n# التأكد من وجود GPU\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f'Using device: {device}')\n\n# تحميل موديل EfficientNetB5\nfrom torchvision.models import efficientnet_b5, EfficientNet_B5_Weights\n\n# نستخدم الـ Weights المسبقة لو عايزين (Transfer Learning)\nweights = EfficientNet_B5_Weights.IMAGENET1K_V1\nmodel = efficientnet_b5(weights=weights)\n\n# تعديل آخر طبقة ليناسب عدد الكلاسات (هنا 2: Cancer / No Cancer)\nmodel.classifier[1] = nn.Linear(in_features=model.classifier[1].in_features, out_features=1)\n\n# نرسل الموديل للـ device (GPU أو CPU)\nmodel = model.to(device)\n\n# طباعة ملخص للموديل\nprint(model)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:27:04.46409Z","iopub.execute_input":"2025-04-27T22:27:04.4649Z","iopub.status.idle":"2025-04-27T22:27:06.688758Z","shell.execute_reply.started":"2025-04-27T22:27:04.464871Z","shell.execute_reply":"2025-04-27T22:27:06.687917Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**This neural network was designed to extract hierarchical features from the mammograms, allowing the model to learn the subtle differences that indicate cancerous changes.**","metadata":{}},{"cell_type":"code","source":"# تعبير عن تمرير صورة خلال الموديل واستخراج الـ features\n\nsample_batch = next(iter(train_loader))\nimages, labels = sample_batch\nimages = images.to(device)\n\n# مرر صورة خلال الموديل بدون تحديث المعاملات\nwith torch.no_grad():\n    features = model.features(images)  # استخراج الـ features قبل طبقة التصنيف النهائية\n\nprint(f'Extracted feature map shape: {features.shape}')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T21:31:17.128541Z","iopub.execute_input":"2025-04-27T21:31:17.129048Z","iopub.status.idle":"2025-04-27T21:31:18.759952Z","shell.execute_reply.started":"2025-04-27T21:31:17.129025Z","shell.execute_reply":"2025-04-27T21:31:18.758999Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Chapter 5: Training the Champion**\n**The next step was the battle of training. The model was fed batches of preprocessed images along with their labels. With each epoch, the model learned to differentiate between healthy tissue and suspicious findings. Special care was taken to monitor performance, ensuring that false positives were minimized.**","metadata":{}},{"cell_type":"code","source":"\n# تحديد الجهاز\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n# تعريف الفقد\ncriterion = nn.BCEWithLogitsLoss()\n\n# تعريف الأوبتميمايزر\noptimizer = optim.Adam(model.parameters(), lr=1e-4)\n\n# عدد الإيبوكس\nepochs = 5\n\n# التدريب\nfor epoch in range(epochs):\n    model.train()\n    running_loss = 0.0\n    \n    pbar = tqdm(train_loader, desc=f\"Epoch {epoch+1}/{epochs}\")\n    \n    for images, labels in pbar:\n        images, labels = images.to(device), labels.to(device).float().unsqueeze(1)\n        \n        optimizer.zero_grad()\n        \n        outputs = model(images)\n        \n        loss = criterion(outputs, labels)\n        loss.backward()\n        optimizer.step()\n        \n        running_loss += loss.item()\n        pbar.set_postfix({'loss': running_loss / (pbar.n + 1)})\n    \n    print(f\"Epoch [{epoch+1}/{epochs}] Loss: {running_loss/len(train_loader):.4f}\")\n\n    # تقييم بعد كل ايبوك\n    model.eval()\n    correct = 0\n    total = 0\n    with torch.no_grad():\n        for images, labels in valid_loader:\n            images, labels = images.to(device), labels.to(device).float().unsqueeze(1)\n            outputs = model(images)\n            preds = torch.sigmoid(outputs) > 0.5\n            correct += (preds == labels).sum().item()\n            total += labels.size(0)\n    acc = correct / total\n    print(f\"Validation Accuracy after Epoch {epoch+1}: {acc:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T22:30:07.214619Z","iopub.execute_input":"2025-04-27T22:30:07.215384Z","iopub.status.idle":"2025-04-27T23:33:45.367031Z","shell.execute_reply.started":"2025-04-27T22:30:07.215331Z","shell.execute_reply":"2025-04-27T23:33:45.366207Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Chapter 6: The Triumph and Beyond**\n**After rigorous training and validation, the model was put to the test on unseen data. Its performance was promising—a testament to the tireless work of Dr. Amal and her team. The AI assistant demonstrated its potential to serve as a second pair of eyes, supporting radiologists in making more accurate diagnoses.😀**","metadata":{}},{"cell_type":"code","source":"correct = 0\ntotal = 0\n\nwith torch.no_grad():  # We don't need gradients for testing\n    for images, labels in tqdm(test_loader, desc=\"Testing\"):\n        images = images.to(device)\n        labels = labels.to(device)\n\n        outputs = model(images)  # Forward pass\n        _, predicted = torch.max(outputs, 1)  # Get predicted class\n        total += labels.size(0)\n        correct += (predicted == labels).sum().item()\n\naccuracy = 100 * correct / total\nprint(f\"Test Accuracy: {accuracy:.2f}%\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-27T23:34:49.975757Z","iopub.execute_input":"2025-04-27T23:34:49.976023Z","iopub.status.idle":"2025-04-27T23:35:48.576485Z","shell.execute_reply.started":"2025-04-27T23:34:49.976004Z","shell.execute_reply":"2025-04-27T23:35:48.575595Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Epilogue: A New Dawn of Hope**\n**In the quiet hum of the radiology lab, as mammograms cycled through the imaging machines, a new kind of vigilance took shape. The AI assistant, born from collaboration and fueled by data, was more than just an algorithm—it was a beacon of hope for countless women.**\n\n**Dr. Amal often reflected on the journey:**\n    \n`\"Every image processed, every diagnosis aided, brings us one step closer to a future where early detection is the norm rather than the exception. This is not just technology; it's a promise of better care and a brighter tomorrow.\"`","metadata":{}},{"cell_type":"markdown","source":"#### **And so, the story continues—each breakthrough and every line of code reinforcing the commitment to save lives, one image at a time.❤️**","metadata":{}}]}