{"cells":[{"metadata":{},"cell_type":"markdown","source":"Can we use a model to predict, given an image, which institution it comes from ? \n\nI ask the question because, looking at the pictures, I saw very different looking shapes. I wonder if the two institutions have different way of taking the pictures (or maybe they render them differently ?)\n\nIf the pictures are so different from one institution to another,it could pertubate the model during training... \n\nLet's see if the pictures are so different..."},{"metadata":{},"cell_type":"markdown","source":"- Source of dataset: https://www.kaggle.com/xhlulu/panda-resize-and-save-train-data/\n- Kernel: https://www.kaggle.com/xhlulu/panda-resize-and-save-train-data/output"},{"metadata":{},"cell_type":"markdown","source":"# Install and import fastai2"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"!pip install fastai2 > /dev/null","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import fastai2\nfrom fastai2.vision.all import *","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Get the data"},{"metadata":{"trusted":true},"cell_type":"code","source":"path = Path('../input/prostate-cancer-grade-assessment')\npath.ls()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.read_csv(path/'train.csv')\nimg_path = Path('../input/panda-train-png-images/train/')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head(3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Create DataLoaders"},{"metadata":{"trusted":true},"cell_type":"code","source":"# add .png to filenames\ndf['image_id'] = df['image_id'].apply(lambda x: str(x)+'.png')\ndf.head(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"prostates = DataBlock(blocks=(ImageBlock, CategoryBlock),\n                   splitter=RandomSplitter(),\n                   get_x=ColReader(0, pref=img_path),\n                   get_y=ColReader(1),\n                   item_tfms=Resize(224),\n                   batch_tfms=aug_transforms()\n                     )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dls = prostates.dataloaders(df, bs=16)\ndls.show_batch()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"See they \"grey\" zone around the purple in \"Radboud\" pictures, whereas \"Karolinska\" doesn't seem to have that at all... Also, karolinska pictures seem to be nearly always straight vertical thin lines, whereas radboud look more like masses.\n\nThat's what got me started. Here are more examples:"},{"metadata":{"trusted":true},"cell_type":"code","source":"dls.show_batch()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"learn = cnn_learner(dls, resnet50, metrics=accuracy)\nlearn.fit_one_cycle(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"interp = ClassificationInterpretation.from_learner(learn)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"interp.plot_confusion_matrix()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"interp.plot_top_losses(k = 9)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"A single epoch gives us an accuracy of 100%. I suspect several possibilities\n\n- the imagery process might differ\n- or they have similar process, but then don't encode the data similarly on the computer\n- it's not the same level of zoom\n- the preprocessing of <a href='https://www.kaggle.com/xhlulu/panda-resize-and-save-train-data/output'>this kernel</a> has somethings fishy \n\nIn any case, finding a way for our model to look at the same thing when it sees images from different institution should probably a priority in this competition"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}