{
  "id": 370893,
  "title": "Submission constantly failing",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/370893",
  "author_name": "Giovanni Cavallin",
  "post_date": "2022-12-07T00:36:05.736000",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi all,<br>\nI'm trying to submit my code, but I constantly having a notebook exception while inferring. I've already checked the question asked before, but it seems not to be a correlated problem.</p>\n<p>In particular, I'd like to directly read <code>.dcm</code> into a (imsize, imsize, 3) image, ready for my dataloader.</p>\n<pre><code>try:\n    import pylibjpeg\nexcept:\n    !pip install /kaggle/input/rsna-bce-download-wheels/pydicom-2.3.1-py3-none-any.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg-1.4.0-py3-none-any.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg_libjpeg-1.3.2-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg_openjpeg-1.2.1-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n\nimport pandas as pd\nimport numpy as np\nfrom typing import List\nfrom pathlib import Path\nimport random\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nimport numpy as np\nfrom PIL import Image\nfrom albumentations import Compose, Resize\nimport os\nimport pydicom\nfrom PIL import Image\nimport cv2\n\n\ndef read_xray(path, fix_monochrome = True):\n    dicom = pydicom.dcmread(path)\n    max_px = 2 ** 16 - 1\n#     data = apply_voi_lut(dicom.pixel_array, dicom)\n    data = dicom.pixel_array\n\n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = max_px - data\n\n    data = data / max_px\n    data = (data * 255).astype(np.uint8)\n\n    return data\n\ndef crop_roi_from_image(img: np.ndarray):\n    # Otsu's thresholding after Gaussian filtering\n    blur = cv2.GaussianBlur(img, (5, 5), 3)\n    # _, breast_mask = cv2.threshold(blur,0,255,cv2.THRESH_BINARY+cv2.THRESH_OTSU)\n    _, breast_mask = cv2.threshold(blur,0,255, 16)\n\n    cnts, _ = cv2.findContours(breast_mask.astype(np.uint8), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)\n    cnt = max(cnts, key = cv2.contourArea)\n    x, y, w, h = cv2.boundingRect(cnt)\n    return img[y:y+h, x:x+w, ...]\n</code></pre>\n<p>And the dataloader:</p>\n<pre><code>random.seed(42)\n\nclass RSNA_BCD_Dataset(Dataset):\n    def __init__(\n            self, \n            dataset_path: Path, \n            csv_path: Path,\n            transform = None\n        ):\n        super().__init__()\n\n        df = pd.read_csv(csv_path)\n        patient_ids = list(df.get('patient_id'))\n        lateralities = list(df.get('laterality'))\n        image_ids = list(df.get('image_id'))\n\n        self.patient_ids = {}\n        for patient_id, lat, img_id in zip(patient_ids, lateralities, image_ids):\n            key = f\"{patient_id}_{lat}\"\n            v = dataset_path / Path(str(patient_id)) / Path(f\"{img_id}.dcm\")\n            try:\n                self.patient_ids[key].append(v)\n            except KeyError:\n                self.patient_ids[key] = [v]\n\n        print(\"Ids loaded!\")\n        self.idx_to_key = {i:k for i, k in enumerate(self.patient_ids.keys())}\n        self.transform = transform\n\n    def __len__(self):\n        return len(list(self.patient_ids.keys()))\n\n    def __getitem__(self, idx):\n        img_paths = self.patient_ids[self.idx_to_key[idx]]\n        imgs = []\n        for img_path in img_paths:\n            img = crop_roi_from_image(read_xray(img_path))\n            imgs.append(np.array(Image.fromarray(img).convert(\"RGB\")))\n\n        # Apply transform\n        if self.transform is not None:\n            imgs = np.array([self.transform(image=img)['image'] for img in imgs])\n\n        return imgs, self.idx_to_key[idx]\n</code></pre>\n<p>It really fails as soon as the scoring starts - the commit finishes without a problem. I think there is some problems with the imports… But I'm finding it difficult to debug.<br>\nAny ideas?</p>\n<p>Thank you in advance!</p>",
  "messages": [
    {
      "id": 2057299,
      "postDate": "2022-12-07T00:36:05.737Z",
      "content": "<p>Hi all,<br>\nI'm trying to submit my code, but I constantly having a notebook exception while inferring. I've already checked the question asked before, but it seems not to be a correlated problem.</p>\n<p>In particular, I'd like to directly read <code>.dcm</code> into a (imsize, imsize, 3) image, ready for my dataloader.</p>\n<pre><code>try:\n    import pylibjpeg\nexcept:\n    !pip install /kaggle/input/rsna-bce-download-wheels/pydicom-2.3.1-py3-none-any.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg-1.4.0-py3-none-any.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg_libjpeg-1.3.2-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg_openjpeg-1.2.1-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n\nimport pandas as pd\nimport numpy as np\nfrom typing import List\nfrom pathlib import Path\nimport random\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nimport numpy as np\nfrom PIL import Image\nfrom albumentations import Compose, Resize\nimport os\nimport pydicom\nfrom PIL import Image\nimport cv2\n\n\ndef read_xray(path, fix_monochrome = True):\n    dicom = pydicom.dcmread(path)\n    max_px = 2 ** 16 - 1\n#     data = apply_voi_lut(dicom.pixel_array, dicom)\n    data = dicom.pixel_array\n\n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = max_px - data\n\n    data = data / max_px\n    data = (data * 255).astype(np.uint8)\n\n    return data\n\ndef crop_roi_from_image(img: np.ndarray):\n    # Otsu's thresholding after Gaussian filtering\n    blur = cv2.GaussianBlur(img, (5, 5), 3)\n    # _, breast_mask = cv2.threshold(blur,0,255,cv2.THRESH_BINARY+cv2.THRESH_OTSU)\n    _, breast_mask = cv2.threshold(blur,0,255, 16)\n\n    cnts, _ = cv2.findContours(breast_mask.astype(np.uint8), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)\n    cnt = max(cnts, key = cv2.contourArea)\n    x, y, w, h = cv2.boundingRect(cnt)\n    return img[y:y+h, x:x+w, ...]\n</code></pre>\n<p>And the dataloader:</p>\n<pre><code>random.seed(42)\n\nclass RSNA_BCD_Dataset(Dataset):\n    def __init__(\n            self, \n            dataset_path: Path, \n            csv_path: Path,\n            transform = None\n        ):\n        super().__init__()\n\n        df = pd.read_csv(csv_path)\n        patient_ids = list(df.get('patient_id'))\n        lateralities = list(df.get('laterality'))\n        image_ids = list(df.get('image_id'))\n\n        self.patient_ids = {}\n        for patient_id, lat, img_id in zip(patient_ids, lateralities, image_ids):\n            key = f\"{patient_id}_{lat}\"\n            v = dataset_path / Path(str(patient_id)) / Path(f\"{img_id}.dcm\")\n            try:\n                self.patient_ids[key].append(v)\n            except KeyError:\n                self.patient_ids[key] = [v]\n\n        print(\"Ids loaded!\")\n        self.idx_to_key = {i:k for i, k in enumerate(self.patient_ids.keys())}\n        self.transform = transform\n\n    def __len__(self):\n        return len(list(self.patient_ids.keys()))\n\n    def __getitem__(self, idx):\n        img_paths = self.patient_ids[self.idx_to_key[idx]]\n        imgs = []\n        for img_path in img_paths:\n            img = crop_roi_from_image(read_xray(img_path))\n            imgs.append(np.array(Image.fromarray(img).convert(\"RGB\")))\n\n        # Apply transform\n        if self.transform is not None:\n            imgs = np.array([self.transform(image=img)['image'] for img in imgs])\n\n        return imgs, self.idx_to_key[idx]\n</code></pre>\n<p>It really fails as soon as the scoring starts - the commit finishes without a problem. I think there is some problems with the imports… But I'm finding it difficult to debug.<br>\nAny ideas?</p>\n<p>Thank you in advance!</p>",
      "rawMarkdown": "Hi all,\nI'm trying to submit my code, but I constantly having a notebook exception while inferring. I've already checked the question asked before, but it seems not to be a correlated problem.\n\nIn particular, I'd like to directly read `.dcm` into a (imsize, imsize, 3) image, ready for my dataloader.\n```\ntry:\n    import pylibjpeg\nexcept:\n    !pip install /kaggle/input/rsna-bce-download-wheels/pydicom-2.3.1-py3-none-any.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg-1.4.0-py3-none-any.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg_libjpeg-1.3.2-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg_openjpeg-1.2.1-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n\nimport pandas as pd\nimport numpy as np\nfrom typing import List\nfrom pathlib import Path\nimport random\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nimport numpy as np\nfrom PIL import Image\nfrom albumentations import Compose, Resize\nimport os\nimport pydicom\nfrom PIL import Image\nimport cv2\n\n\ndef read_xray(path, fix_monochrome = True):\n    dicom = pydicom.dcmread(path)\n    max_px = 2 ** 16 - 1\n#     data = apply_voi_lut(dicom.pixel_array, dicom)\n    data = dicom.pixel_array\n               \n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = max_px - data\n        \n    data = data / max_px\n    data = (data * 255).astype(np.uint8)\n        \n    return data\n\ndef crop_roi_from_image(img: np.ndarray):\n    # Otsu's thresholding after Gaussian filtering\n    blur = cv2.GaussianBlur(img, (5, 5), 3)\n    # _, breast_mask = cv2.threshold(blur,0,255,cv2.THRESH_BINARY+cv2.THRESH_OTSU)\n    _, breast_mask = cv2.threshold(blur,0,255, 16)\n    \n    cnts, _ = cv2.findContours(breast_mask.astype(np.uint8), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)\n    cnt = max(cnts, key = cv2.contourArea)\n    x, y, w, h = cv2.boundingRect(cnt)\n    return img[y:y+h, x:x+w, ...]\n```\n\nAnd the dataloader:\n```\nrandom.seed(42)\n\nclass RSNA_BCD_Dataset(Dataset):\n    def __init__(\n            self, \n            dataset_path: Path, \n            csv_path: Path,\n            transform = None\n        ):\n        super().__init__()\n        \n        df = pd.read_csv(csv_path)\n        patient_ids = list(df.get('patient_id'))\n        lateralities = list(df.get('laterality'))\n        image_ids = list(df.get('image_id'))\n        \n        self.patient_ids = {}\n        for patient_id, lat, img_id in zip(patient_ids, lateralities, image_ids):\n            key = f\"{patient_id}_{lat}\"\n            v = dataset_path / Path(str(patient_id)) / Path(f\"{img_id}.dcm\")\n            try:\n                self.patient_ids[key].append(v)\n            except KeyError:\n                self.patient_ids[key] = [v]\n        \n        print(\"Ids loaded!\")\n        self.idx_to_key = {i:k for i, k in enumerate(self.patient_ids.keys())}\n        self.transform = transform\n    \n    def __len__(self):\n        return len(list(self.patient_ids.keys()))\n    \n    def __getitem__(self, idx):\n        img_paths = self.patient_ids[self.idx_to_key[idx]]\n        imgs = []\n        for img_path in img_paths:\n            img = crop_roi_from_image(read_xray(img_path))\n            imgs.append(np.array(Image.fromarray(img).convert(\"RGB\")))\n\n        # Apply transform\n        if self.transform is not None:\n            imgs = np.array([self.transform(image=img)['image'] for img in imgs])\n        \n        return imgs, self.idx_to_key[idx]\n```\nIt really fails as soon as the scoring starts - the commit finishes without a problem. I think there is some problems with the imports... But I'm finding it difficult to debug.\nAny ideas?\n\nThank you in advance!",
      "votes": 3
    },
    {
      "id": 2057388,
      "postDate": "2022-12-07T04:17:00.317Z",
      "content": "<p>I had submission errors like notebook threw exception / timed out in the past. Try running prediction on train data in your notebook and check if it runs fine. </p>",
      "rawMarkdown": "I had submission errors like notebook threw exception / timed out in the past. Try running prediction on train data in your notebook and check if it runs fine. ",
      "votes": 2,
      "replies": [
        {
          "id": 2057593,
          "postDate": "2022-12-07T08:33:05.280Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/dhruvkhatri\" target=\"_blank\">@dhruvkhatri</a> . Once you fixed all your problems in the train set, were you able to infer without a problem then?</p>",
          "rawMarkdown": "Thank you @dhruvkhatri . Once you fixed all your problems in the train set, were you able to infer without a problem then?",
          "votes": 1
        },
        {
          "id": 2057640,
          "postDate": "2022-12-07T09:18:47.820Z",
          "content": "<p>Yes indeed, But I had to add a code segment (shown here <a href=\"https://www.kaggle.com/code/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres/notebook\" target=\"_blank\">preprocess function</a>) to process all the hidden test dicom images to png first (otherwise the notebook will take a long time to process 8000 dicom files). Then I ran the inference which worked fine. Also make sure to handle the duplicate entries, this will give you a 'scoring error' instead. </p>",
          "rawMarkdown": "Yes indeed, But I had to add a code segment (shown here [preprocess function](https://www.kaggle.com/code/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres/notebook)) to process all the hidden test dicom images to png first (otherwise the notebook will take a long time to process 8000 dicom files). Then I ran the inference which worked fine. Also make sure to handle the duplicate entries, this will give you a 'scoring error' instead. ",
          "votes": 1
        },
        {
          "id": 2057653,
          "postDate": "2022-12-07T09:32:28.557Z",
          "content": "<p>Thank you for your suggestions. As it turned out, the packages requirements for the environment were different in this platform wrt my own - so I wrongly assumed I needed to install <code>pylibjpeg_libjpeg</code> and <code>pylibjpeg_openjpeg</code> along with <code>pylibjpeg</code>. Indeed, I only needed to install <code>pylibjpeg</code> and <code>python-gdcm</code> in order to make the thing run! <br>\nThank you so much for pointing out trying my method with train set. I'm not so new to Kaggle, but I've never had to do so. D:</p>\n<p>As for the loading suggestion, I've already taken a look at your great work :) But I've decided to go with a PyTorch custom DataLoader, where I can exploit the online multiprocessing. As for now, I processed (that is, photo is loaded, then ROI is cut, then it is offered to the model for inferring) 300 images in ~10mins, that is ~5h for the entire test set (assuming the same images distribution). Is it similar to your timings? <br>\nIn my case I won't have repetitions since the network takes all the images for each <code>&lt;patient_id&gt;_&lt;laterality&gt;</code> at once! :D</p>",
          "rawMarkdown": "Thank you for your suggestions. As it turned out, the packages requirements for the environment were different in this platform wrt my own - so I wrongly assumed I needed to install `pylibjpeg_libjpeg` and `pylibjpeg_openjpeg` along with `pylibjpeg`. Indeed, I only needed to install `pylibjpeg` and `python-gdcm` in order to make the thing run! \nThank you so much for pointing out trying my method with train set. I'm not so new to Kaggle, but I've never had to do so. D:\n\nAs for the loading suggestion, I've already taken a look at your great work :) But I've decided to go with a PyTorch custom DataLoader, where I can exploit the online multiprocessing. As for now, I processed (that is, photo is loaded, then ROI is cut, then it is offered to the model for inferring) 300 images in ~10mins, that is ~5h for the entire test set (assuming the same images distribution). Is it similar to your timings? \nIn my case I won't have repetitions since the network takes all the images for each `<patient_id>_<laterality>` at once! :D"
        },
        {
          "id": 2057714,
          "postDate": "2022-12-07T10:07:25.613Z",
          "content": "<blockquote>\n  <p>As for now, I processed (that is, photo is loaded, then ROI is cut, then it is offered to the model for inferring) 300 images in ~10mins, that is ~5h for the entire test set (assuming the same images distribution). Is it similar to your timings?</p>\n</blockquote>\n<p>yes - about 6h</p>",
          "rawMarkdown": "> As for now, I processed (that is, photo is loaded, then ROI is cut, then it is offered to the model for inferring) 300 images in ~10mins, that is ~5h for the entire test set (assuming the same images distribution). Is it similar to your timings?\n\nyes - about 6h",
          "votes": 2
        },
        {
          "id": 2057762,
          "postDate": "2022-12-07T10:23:35.157Z",
          "content": "<p>Great! I am glad i could help. Just one thing, the notebook I shared is not mine :) I was going through it to find a solution to  my \" Notebook timed out\" errors</p>",
          "rawMarkdown": "Great! I am glad i could help. Just one thing, the notebook I shared is not mine :) I was going through it to find a solution to  my \" Notebook timed out\" errors"
        },
        {
          "id": 2058379,
          "postDate": "2022-12-07T21:31:09.930Z",
          "content": "<p>Ah! I didn't notice, thanks for the note. :) Did you manage to find a good solution for yourself? In order to take less than ~6h, I mean…</p>",
          "rawMarkdown": "Ah! I didn't notice, thanks for the note. :) Did you manage to find a good solution for yourself? In order to take less than ~6h, I mean..."
        },
        {
          "id": 2059088,
          "postDate": "2022-12-08T13:26:55.340Z",
          "content": "<p>No luck there! Will update you if I find something faster. Best </p>",
          "rawMarkdown": "No luck there! Will update you if I find something faster. Best ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2057388,
      "author_name": "Dhruv_K",
      "author_url": "",
      "post_date": "2022-12-07T04:17:00.317000",
      "content": "<p>I had submission errors like notebook threw exception / timed out in the past. Try running prediction on train data in your notebook and check if it runs fine. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2057593,
          "author_name": "Giovanni Cavallin",
          "author_url": "",
          "post_date": "2022-12-07T08:33:05.280000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/dhruvkhatri\" target=\"_blank\">@dhruvkhatri</a> . Once you fixed all your problems in the train set, were you able to infer without a problem then?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2057640,
          "author_name": "Dhruv_K",
          "author_url": "",
          "post_date": "2022-12-07T09:18:47.820000",
          "content": "<p>Yes indeed, But I had to add a code segment (shown here <a href=\"https://www.kaggle.com/code/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres/notebook\" target=\"_blank\">preprocess function</a>) to process all the hidden test dicom images to png first (otherwise the notebook will take a long time to process 8000 dicom files). Then I ran the inference which worked fine. Also make sure to handle the duplicate entries, this will give you a 'scoring error' instead. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2057653,
          "author_name": "Giovanni Cavallin",
          "author_url": "",
          "post_date": "2022-12-07T09:32:28.557000",
          "content": "<p>Thank you for your suggestions. As it turned out, the packages requirements for the environment were different in this platform wrt my own - so I wrongly assumed I needed to install <code>pylibjpeg_libjpeg</code> and <code>pylibjpeg_openjpeg</code> along with <code>pylibjpeg</code>. Indeed, I only needed to install <code>pylibjpeg</code> and <code>python-gdcm</code> in order to make the thing run! <br>\nThank you so much for pointing out trying my method with train set. I'm not so new to Kaggle, but I've never had to do so. D:</p>\n<p>As for the loading suggestion, I've already taken a look at your great work :) But I've decided to go with a PyTorch custom DataLoader, where I can exploit the online multiprocessing. As for now, I processed (that is, photo is loaded, then ROI is cut, then it is offered to the model for inferring) 300 images in ~10mins, that is ~5h for the entire test set (assuming the same images distribution). Is it similar to your timings? <br>\nIn my case I won't have repetitions since the network takes all the images for each <code>&lt;patient_id&gt;_&lt;laterality&gt;</code> at once! :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2057714,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-07T10:07:25.613000",
          "content": "<blockquote>\n  <p>As for now, I processed (that is, photo is loaded, then ROI is cut, then it is offered to the model for inferring) 300 images in ~10mins, that is ~5h for the entire test set (assuming the same images distribution). Is it similar to your timings?</p>\n</blockquote>\n<p>yes - about 6h</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2057762,
          "author_name": "Dhruv_K",
          "author_url": "",
          "post_date": "2022-12-07T10:23:35.157000",
          "content": "<p>Great! I am glad i could help. Just one thing, the notebook I shared is not mine :) I was going through it to find a solution to  my \" Notebook timed out\" errors</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2058379,
          "author_name": "Giovanni Cavallin",
          "author_url": "",
          "post_date": "2022-12-07T21:31:09.930000",
          "content": "<p>Ah! I didn't notice, thanks for the note. :) Did you manage to find a good solution for yourself? In order to take less than ~6h, I mean…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2059088,
          "author_name": "Dhruv_K",
          "author_url": "",
          "post_date": "2022-12-08T13:26:55.340000",
          "content": "<p>No luck there! Will update you if I find something faster. Best </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2057299": "Hi all,\nI'm trying to submit my code, but I constantly having a notebook exception while inferring. I've already checked the question asked before, but it seems not to be a correlated problem.\n\nIn particular, I'd like to directly read `.dcm` into a (imsize, imsize, 3) image, ready for my dataloader.\n```\ntry:\n    import pylibjpeg\nexcept:\n    !pip install /kaggle/input/rsna-bce-download-wheels/pydicom-2.3.1-py3-none-any.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg-1.4.0-py3-none-any.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg_libjpeg-1.3.2-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n    !pip install /kaggle/input/rsna-bce-download-wheels/pylibjpeg_openjpeg-1.2.1-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n\nimport pandas as pd\nimport numpy as np\nfrom typing import List\nfrom pathlib import Path\nimport random\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nimport numpy as np\nfrom PIL import Image\nfrom albumentations import Compose, Resize\nimport os\nimport pydicom\nfrom PIL import Image\nimport cv2\n\n\ndef read_xray(path, fix_monochrome = True):\n    dicom = pydicom.dcmread(path)\n    max_px = 2 ** 16 - 1\n#     data = apply_voi_lut(dicom.pixel_array, dicom)\n    data = dicom.pixel_array\n               \n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = max_px - data\n        \n    data = data / max_px\n    data = (data * 255).astype(np.uint8)\n        \n    return data\n\ndef crop_roi_from_image(img: np.ndarray):\n    # Otsu's thresholding after Gaussian filtering\n    blur = cv2.GaussianBlur(img, (5, 5), 3)\n    # _, breast_mask = cv2.threshold(blur,0,255,cv2.THRESH_BINARY+cv2.THRESH_OTSU)\n    _, breast_mask = cv2.threshold(blur,0,255, 16)\n    \n    cnts, _ = cv2.findContours(breast_mask.astype(np.uint8), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)\n    cnt = max(cnts, key = cv2.contourArea)\n    x, y, w, h = cv2.boundingRect(cnt)\n    return img[y:y+h, x:x+w, ...]\n```\n\nAnd the dataloader:\n```\nrandom.seed(42)\n\nclass RSNA_BCD_Dataset(Dataset):\n    def __init__(\n            self, \n            dataset_path: Path, \n            csv_path: Path,\n            transform = None\n        ):\n        super().__init__()\n        \n        df = pd.read_csv(csv_path)\n        patient_ids = list(df.get('patient_id'))\n        lateralities = list(df.get('laterality'))\n        image_ids = list(df.get('image_id'))\n        \n        self.patient_ids = {}\n        for patient_id, lat, img_id in zip(patient_ids, lateralities, image_ids):\n            key = f\"{patient_id}_{lat}\"\n            v = dataset_path / Path(str(patient_id)) / Path(f\"{img_id}.dcm\")\n            try:\n                self.patient_ids[key].append(v)\n            except KeyError:\n                self.patient_ids[key] = [v]\n        \n        print(\"Ids loaded!\")\n        self.idx_to_key = {i:k for i, k in enumerate(self.patient_ids.keys())}\n        self.transform = transform\n    \n    def __len__(self):\n        return len(list(self.patient_ids.keys()))\n    \n    def __getitem__(self, idx):\n        img_paths = self.patient_ids[self.idx_to_key[idx]]\n        imgs = []\n        for img_path in img_paths:\n            img = crop_roi_from_image(read_xray(img_path))\n            imgs.append(np.array(Image.fromarray(img).convert(\"RGB\")))\n\n        # Apply transform\n        if self.transform is not None:\n            imgs = np.array([self.transform(image=img)['image'] for img in imgs])\n        \n        return imgs, self.idx_to_key[idx]\n```\nIt really fails as soon as the scoring starts - the commit finishes without a problem. I think there is some problems with the imports... But I'm finding it difficult to debug.\nAny ideas?\n\nThank you in advance!",
    "2057388": "I had submission errors like notebook threw exception / timed out in the past. Try running prediction on train data in your notebook and check if it runs fine. "
  }
}