{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### ROI Extractor \n\nDuring the visual data analysis I noticed that there is a large variation in the arrangement of the object in the images. In addition, some objects occupy only a small part of the image. By converting (resizing) without ROI extraction we have a very inefficient use of the reduced image. Most of the picture is blank.\n\n**Solution:**\n- annotate data - I annotated about 500 images in a human in the loop technique (3 models were created - I started from 300 images and ended up about 500) \n- train object detector - I used yolov5 (small - balance between accuracy and speed) \n\n**Result on train DS:**\n* 54601 images processed successfully \n* 105 images - detection failed\n\nModel performance (on my validation DS):\n* mAP@50 -> 0.995      \n* mAP50-95 -> 0.914\n\n<div class=\"alert alert-warning\">If you are interested in:\n    <ul>\n        <li>Annotation dataset</li>\n        <li>yolov5 training notebook</li>\n    </ul>\n    <p>&nbsp;</p>\nLet me know in comment. I will provide it as well in separate notebooks.</div>\n\n**Next steps:**\n* generate ROI based dataset for training (768 pix) -> is available here: https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images\n* use ROI extractor in inference part ","metadata":{}},{"cell_type":"code","source":"%%capture \n\n# Clone yolov5 repository\n!git clone https://github.com/ultralytics/yolov5","metadata":{"execution":{"iopub.status.busy":"2022-12-01T09:15:09.026021Z","iopub.execute_input":"2022-12-01T09:15:09.026469Z","iopub.status.idle":"2022-12-01T09:15:11.978124Z","shell.execute_reply.started":"2022-12-01T09:15:09.026381Z","shell.execute_reply":"2022-12-01T09:15:11.976864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nimport glob\nimport random\nimport cv2\nimport matplotlib.pyplot as plt","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-01T09:15:11.980626Z","iopub.execute_input":"2022-12-01T09:15:11.980926Z","iopub.status.idle":"2022-12-01T09:15:13.886099Z","shell.execute_reply.started":"2022-12-01T09:15:11.980897Z","shell.execute_reply":"2022-12-01T09:15:13.885107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load trained model\nmodel = torch.hub.load('./yolov5', 'custom', path='/kaggle/input/rsna-breast-cancer-detection-roi-model/rsna-roi-003.pt', source='local')","metadata":{"execution":{"iopub.status.busy":"2022-12-01T09:15:13.888607Z","iopub.execute_input":"2022-12-01T09:15:13.889537Z","iopub.status.idle":"2022-12-01T09:15:19.311387Z","shell.execute_reply.started":"2022-12-01T09:15:13.889498Z","shell.execute_reply":"2022-12-01T09:15:19.310204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create list of input files \n# I use Radek Osmulski dataset -> https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs \n\nfile_list = glob.glob('/kaggle/input/rsna-mammography-images-as-pngs/images_as_pngs_768/train_images_processed_768/*/*.png')","metadata":{"execution":{"iopub.status.busy":"2022-12-01T09:15:19.313994Z","iopub.execute_input":"2022-12-01T09:15:19.314895Z","iopub.status.idle":"2022-12-01T09:16:09.854016Z","shell.execute_reply.started":"2022-12-01T09:15:19.314847Z","shell.execute_reply":"2022-12-01T09:16:09.852888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%matplotlib inline\nimages = []\n\nfor img_file in random.sample(file_list, 25):  # it is fixed to 25 random predictions - if you want to change it remeber to change plot_roi as well\n    \n    # Read file from file\n    frame = cv2.imread(img_file)\n    \n    # Make prediction\n    detections = model(frame)\n    \n    # Convert results to Pandas style\n    results = detections.pandas().xyxy[0].to_dict(orient=\"records\")\n    \n    # Plot result (in 99.99% it predicts only one instance - certainly you can assure that only best prediction is used)\n    for result in results:\n        images.append(cv2.rectangle(frame, (int(result['xmin']), int(result['ymin'])), (int(result['xmax']), int(result['ymax'])), (255,0,0), 4))\n\n# Plot result\nfig, axes = plt.subplots(5, 5, figsize=(20,20))\n    \nfor idx, image in enumerate(images):\n    i = idx % 5 \n    j = idx // 5 \n    axes[i, j].imshow(image)\n\nplt.subplots_adjust(wspace=0, hspace=.2)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T09:16:09.865937Z","iopub.execute_input":"2022-12-01T09:16:09.866308Z","iopub.status.idle":"2022-12-01T09:16:19.966626Z","shell.execute_reply.started":"2022-12-01T09:16:09.866273Z","shell.execute_reply":"2022-12-01T09:16:19.965765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thank you! This is my first contribution to RSNA Screening Mammography Breast Cancer Detection.\n\nHave a nice day and competition.","metadata":{}}]}