{
  "id": 428962,
  "title": "Extra Files in train_images ? ",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/428962",
  "author_name": "AyushS9020",
  "post_date": "2023-08-03T14:12:23.725000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I was merging the <code>dicom paths</code> and <code>patient_id</code>, but when I started converting the <code>dicom - numpy</code>. I got this </p>\n<pre><code> invalid value encountered  divide\n</code></pre>\n<p>When I traversed for the <code>[508]</code> file ---&gt; <code>train['train_paths'][508]</code>. I got the following error. </p>\n<pre><code>---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nFile condapython3.pandasndexes/base.py:,  Index.get_loc(self, key, method, tolerance)\n    try:\n-&gt;      return self._engine.get_loc(casted_key)\n    except KeyError as err:\n\nFile condapython3.pandasndex.pyx:,  pandas._libs.index.IndexEngine.get_loc()\n\nFile condapython3.pandasndex.pyx:,  pandas._libs.index.IndexEngine.get_loc()\n\nFile pandashashtable_class_helper.pxi:,  pandas._libs.hashtable.Int64HashTable.get_item()\n\nFile pandashashtable_class_helper.pxi:,  pandas._libs.hashtable.Int64HashTable.get_item()\n\nKeyError: \n\nThe above exception was the direct cause of the following exception:\n\nKeyError                                  Traceback (most recent call last)\nCell In[], line \n----&gt;  train[][]\n\nFile condapython3.pandasseries.py:,  Series.__getitem__(self, key)\n         return self._values[key]\n     elif key_is_scalar:\n--&gt;      return self._get_value(key)\n      is_hashable(key):\n         \n         try:\n             \n\nFile condapython3.pandasseries.py:,  Series._get_value(self, label, takeable)\n        return self._values[label]\n    \n-&gt;  loc = self.index.get_loc(label)\n    return self.index._get_values_for_loc(self, loc, label)\n\nFile condapython3.pandasndexes/base.py:,  Index.get_loc(self, key, method, tolerance)\n        return self._engine.get_loc(casted_key)\n    except KeyError as err:\n-&gt;      raise KeyError(key) from err\n    except TypeError:\n        \n        \n        \n        self._check_indexing_error(key)\n\nKeyError: \n</code></pre>\n<p>When I used <code>iloc</code> ---&gt; <code>train.iloc[508]</code>, I got this</p>\n<pre><code>patient_id                                                           \nbowel_healthy                                                          \nbowel_injury                                                           \nextravasation_healthy                                                  \nextravasation_injury                                                   \nkidney_healthy                                                         \nkidney_low                                                             \nkidney_high                                                            \nliver_healthy                                                          \nliver_low                                                              \nliver_high                                                             \nspleen_healthy                                                         \nspleen_low                                                             \nspleen_high                                                            \nany_injury                                                             \nmaked_files                                                            \ntrain_paths              /kaggle/input/rsna-abdominal-trauma-detec...\nName: , : object\n</code></pre>\n<p>which showed that the <code>12580</code> only exists in <code>train_images</code> and we do not have <code>annotations</code>. I further rechecked the anomly with </p>\n<pre><code>In : pd.read\n\nOut : ---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nCell In, line \n----&gt;  pd.read\n\nFile /opt/conda/lib/python3./site-packages/pandas/core/series.py:,  self, key)\n         return self._values\n     elif key_is_scalar:\n--&gt;      return self.\n      is:\n         # Otherwise index.get_value will raise InvalidIndexError\n         :\n             # For labels that don't resolve  scalars like tuples  frozensets\n\nFile /opt/conda/lib/python3./site-packages/pandas/core/series.py:,  self, label, takeable)\n        return self._values\n    # Similar  get_value, but we  not fall back  positional\n-&gt;  loc = self.index.get\n    return self.index.\n\nFile /opt/conda/lib/python3./site-packages/pandas/core/indexes/range.py:,  get\n                 raise  from err\n         self.\n--&gt;      raise \n     return super.get\n\nKeyError: '' \n</code></pre>\n<pre><code>In :  value  (os()):\n     value  (train) : (value)\n\nOut : \n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n</code></pre>\n<pre><code>In : train = train(int)\n\n  train\n\nOut : False\n</code></pre>",
  "messages": [
    {
      "id": 2372148,
      "postDate": "2023-08-03T14:12:23.727Z",
      "content": "<p>I was merging the <code>dicom paths</code> and <code>patient_id</code>, but when I started converting the <code>dicom - numpy</code>. I got this </p>\n<pre><code> invalid value encountered  divide\n</code></pre>\n<p>When I traversed for the <code>[508]</code> file ---&gt; <code>train['train_paths'][508]</code>. I got the following error. </p>\n<pre><code>---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nFile condapython3.pandasndexes/base.py:,  Index.get_loc(self, key, method, tolerance)\n    try:\n-&gt;      return self._engine.get_loc(casted_key)\n    except KeyError as err:\n\nFile condapython3.pandasndex.pyx:,  pandas._libs.index.IndexEngine.get_loc()\n\nFile condapython3.pandasndex.pyx:,  pandas._libs.index.IndexEngine.get_loc()\n\nFile pandashashtable_class_helper.pxi:,  pandas._libs.hashtable.Int64HashTable.get_item()\n\nFile pandashashtable_class_helper.pxi:,  pandas._libs.hashtable.Int64HashTable.get_item()\n\nKeyError: \n\nThe above exception was the direct cause of the following exception:\n\nKeyError                                  Traceback (most recent call last)\nCell In[], line \n----&gt;  train[][]\n\nFile condapython3.pandasseries.py:,  Series.__getitem__(self, key)\n         return self._values[key]\n     elif key_is_scalar:\n--&gt;      return self._get_value(key)\n      is_hashable(key):\n         \n         try:\n             \n\nFile condapython3.pandasseries.py:,  Series._get_value(self, label, takeable)\n        return self._values[label]\n    \n-&gt;  loc = self.index.get_loc(label)\n    return self.index._get_values_for_loc(self, loc, label)\n\nFile condapython3.pandasndexes/base.py:,  Index.get_loc(self, key, method, tolerance)\n        return self._engine.get_loc(casted_key)\n    except KeyError as err:\n-&gt;      raise KeyError(key) from err\n    except TypeError:\n        \n        \n        \n        self._check_indexing_error(key)\n\nKeyError: \n</code></pre>\n<p>When I used <code>iloc</code> ---&gt; <code>train.iloc[508]</code>, I got this</p>\n<pre><code>patient_id                                                           \nbowel_healthy                                                          \nbowel_injury                                                           \nextravasation_healthy                                                  \nextravasation_injury                                                   \nkidney_healthy                                                         \nkidney_low                                                             \nkidney_high                                                            \nliver_healthy                                                          \nliver_low                                                              \nliver_high                                                             \nspleen_healthy                                                         \nspleen_low                                                             \nspleen_high                                                            \nany_injury                                                             \nmaked_files                                                            \ntrain_paths              /kaggle/input/rsna-abdominal-trauma-detec...\nName: , : object\n</code></pre>\n<p>which showed that the <code>12580</code> only exists in <code>train_images</code> and we do not have <code>annotations</code>. I further rechecked the anomly with </p>\n<pre><code>In : pd.read\n\nOut : ---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nCell In, line \n----&gt;  pd.read\n\nFile /opt/conda/lib/python3./site-packages/pandas/core/series.py:,  self, key)\n         return self._values\n     elif key_is_scalar:\n--&gt;      return self.\n      is:\n         # Otherwise index.get_value will raise InvalidIndexError\n         :\n             # For labels that don't resolve  scalars like tuples  frozensets\n\nFile /opt/conda/lib/python3./site-packages/pandas/core/series.py:,  self, label, takeable)\n        return self._values\n    # Similar  get_value, but we  not fall back  positional\n-&gt;  loc = self.index.get\n    return self.index.\n\nFile /opt/conda/lib/python3./site-packages/pandas/core/indexes/range.py:,  get\n                 raise  from err\n         self.\n--&gt;      raise \n     return super.get\n\nKeyError: '' \n</code></pre>\n<pre><code>In :  value  (os()):\n     value  (train) : (value)\n\nOut : \n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n</code></pre>\n<pre><code>In : train = train(int)\n\n  train\n\nOut : False\n</code></pre>",
      "rawMarkdown": "I was merging the `dicom paths` and `patient_id`, but when I started converting the `dicom - numpy`. I got this \n```\nRuntimeWarning: invalid value encountered in divide\n```\nWhen I traversed for the `[508]` file ---> `train['train_paths'][508]`. I got the following error. \n```\n---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance)\n   3801 try:\n-> 3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()\n\nFile pandas/_libs/hashtable_class_helper.pxi:2263, in pandas._libs.hashtable.Int64HashTable.get_item()\n\nFile pandas/_libs/hashtable_class_helper.pxi:2273, in pandas._libs.hashtable.Int64HashTable.get_item()\n\nKeyError: 508\n\nThe above exception was the direct cause of the following exception:\n\nKeyError                                  Traceback (most recent call last)\nCell In[16], line 1\n----> 1 train['train_paths'][508]\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/series.py:981, in Series.__getitem__(self, key)\n    978     return self._values[key]\n    980 elif key_is_scalar:\n--> 981     return self._get_value(key)\n    983 if is_hashable(key):\n    984     # Otherwise index.get_value will raise InvalidIndexError\n    985     try:\n    986         # For labels that don't resolve as scalars like tuples and frozensets\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/series.py:1089, in Series._get_value(self, label, takeable)\n   1086     return self._values[label]\n   1088 # Similar to Index.get_value, but we do not fall back to positional\n-> 1089 loc = self.index.get_loc(label)\n   1090 return self.index._get_values_for_loc(self, loc, label)\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance)\n   3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n-> 3804     raise KeyError(key) from err\n   3805 except TypeError:\n   3806     # If we have a listlike key, _check_indexing_error will raise\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise\n   3808     #  the TypeError.\n   3809     self._check_indexing_error(key)\n\nKeyError: 508\n\n```\nWhen I used `iloc` ---> `train.iloc[508]`, I got this\n```\npatient_id                                                           12580\nbowel_healthy                                                          NaN\nbowel_injury                                                           NaN\nextravasation_healthy                                                  NaN\nextravasation_injury                                                   NaN\nkidney_healthy                                                         NaN\nkidney_low                                                             NaN\nkidney_high                                                            NaN\nliver_healthy                                                          NaN\nliver_low                                                              NaN\nliver_high                                                             NaN\nspleen_healthy                                                         NaN\nspleen_low                                                             NaN\nspleen_high                                                            NaN\nany_injury                                                             NaN\nmaked_files                                                            NaN\ntrain_paths              /kaggle/input/rsna-2023-abdominal-trauma-detec...\nName: 68949, dtype: object\n```\nwhich showed that the `12580` only exists in `train_images` and we do not have `annotations`. I further rechecked the anomly with \n```\nIn : pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')['patient_id']['12580']\n\nOut : ---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nCell In[20], line 1\n----> 1 pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')['patient_id']['12580']\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/series.py:981, in Series.__getitem__(self, key)\n    978     return self._values[key]\n    980 elif key_is_scalar:\n--> 981     return self._get_value(key)\n    983 if is_hashable(key):\n    984     # Otherwise index.get_value will raise InvalidIndexError\n    985     try:\n    986         # For labels that don't resolve as scalars like tuples and frozensets\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/series.py:1089, in Series._get_value(self, label, takeable)\n   1086     return self._values[label]\n   1088 # Similar to Index.get_value, but we do not fall back to positional\n-> 1089 loc = self.index.get_loc(label)\n   1090 return self.index._get_values_for_loc(self, loc, label)\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/range.py:395, in RangeIndex.get_loc(self, key, method, tolerance)\n    393             raise KeyError(key) from err\n    394     self._check_indexing_error(key)\n--> 395     raise KeyError(key)\n    396 return super().get_loc(key, method=method, tolerance=tolerance)\n\nKeyError: '12580' \n```\n```\nIn : for value in sorted(os.listdir('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images')):\n    if value in str(train['patient_id']) : print(value)\n\nOut : \n\n10004\n10005\n10007\n10026\n10051\n26\n43\n951\n96\n980\n9951\n9960\n9961\n9980\n9983\n```\n```\nIn : train['patient_id'] = train['patient_id'].astype(int)\n\n10082 in train['patient_id']\n\nOut : False\n```",
      "votes": 3
    },
    {
      "id": 2372395,
      "postDate": "2023-08-03T16:32:52.360Z",
      "content": "<p>Yes, the train images file includes some series that were dropped from the final dataset after QC checks. I should have dropped them from <code>train_images.csv</code> as well but overlooked it. Sorry for the confusion!</p>",
      "rawMarkdown": "Yes, the train images file includes some series that were dropped from the final dataset after QC checks. I should have dropped them from `train_images.csv` as well but overlooked it. Sorry for the confusion!",
      "replies": [
        {
          "id": 2372415,
          "postDate": "2023-08-03T16:54:05.013Z",
          "content": "<p>Just confirming, when I iterated through the all the images in the directory there were 1,500,653 images separated into 4,711 <code>series_id</code> folders. But there were 60 <code>series_id</code> missing than what was given in the <code>train_dicom_tags.csv</code>. So is this also because of the QC resulting in dropping these series? (<code>train_dicom_tags.csv</code> had 1,510,373 images, a difference of 9,720 from <code>train_images</code> folder)</p>",
          "rawMarkdown": "Just confirming, when I iterated through the all the images in the directory there were 1,500,653 images separated into 4,711 `series_id` folders. But there were 60 `series_id` missing than what was given in the `train_dicom_tags.csv`. So is this also because of the QC resulting in dropping these series? (`train_dicom_tags.csv` had 1,510,373 images, a difference of 9,720 from `train_images` folder)"
        },
        {
          "id": 2372876,
          "postDate": "2023-08-04T03:11:41.700Z",
          "content": "<p>Thanks Sohier for the clarification :)</p>",
          "rawMarkdown": "Thanks Sohier for the clarification :)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2372395,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-08-03T16:32:52.360000",
      "content": "<p>Yes, the train images file includes some series that were dropped from the final dataset after QC checks. I should have dropped them from <code>train_images.csv</code> as well but overlooked it. Sorry for the confusion!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2372415,
          "author_name": "coderRKJ",
          "author_url": "",
          "post_date": "2023-08-03T16:54:05.013000",
          "content": "<p>Just confirming, when I iterated through the all the images in the directory there were 1,500,653 images separated into 4,711 <code>series_id</code> folders. But there were 60 <code>series_id</code> missing than what was given in the <code>train_dicom_tags.csv</code>. So is this also because of the QC resulting in dropping these series? (<code>train_dicom_tags.csv</code> had 1,510,373 images, a difference of 9,720 from <code>train_images</code> folder)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2372876,
          "author_name": "AyushS9020",
          "author_url": "",
          "post_date": "2023-08-04T03:11:41.700000",
          "content": "<p>Thanks Sohier for the clarification :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2372148": "I was merging the `dicom paths` and `patient_id`, but when I started converting the `dicom - numpy`. I got this \n```\nRuntimeWarning: invalid value encountered in divide\n```\nWhen I traversed for the `[508]` file ---> `train['train_paths'][508]`. I got the following error. \n```\n---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance)\n   3801 try:\n-> 3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()\n\nFile pandas/_libs/hashtable_class_helper.pxi:2263, in pandas._libs.hashtable.Int64HashTable.get_item()\n\nFile pandas/_libs/hashtable_class_helper.pxi:2273, in pandas._libs.hashtable.Int64HashTable.get_item()\n\nKeyError: 508\n\nThe above exception was the direct cause of the following exception:\n\nKeyError                                  Traceback (most recent call last)\nCell In[16], line 1\n----> 1 train['train_paths'][508]\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/series.py:981, in Series.__getitem__(self, key)\n    978     return self._values[key]\n    980 elif key_is_scalar:\n--> 981     return self._get_value(key)\n    983 if is_hashable(key):\n    984     # Otherwise index.get_value will raise InvalidIndexError\n    985     try:\n    986         # For labels that don't resolve as scalars like tuples and frozensets\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/series.py:1089, in Series._get_value(self, label, takeable)\n   1086     return self._values[label]\n   1088 # Similar to Index.get_value, but we do not fall back to positional\n-> 1089 loc = self.index.get_loc(label)\n   1090 return self.index._get_values_for_loc(self, loc, label)\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance)\n   3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n-> 3804     raise KeyError(key) from err\n   3805 except TypeError:\n   3806     # If we have a listlike key, _check_indexing_error will raise\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise\n   3808     #  the TypeError.\n   3809     self._check_indexing_error(key)\n\nKeyError: 508\n\n```\nWhen I used `iloc` ---> `train.iloc[508]`, I got this\n```\npatient_id                                                           12580\nbowel_healthy                                                          NaN\nbowel_injury                                                           NaN\nextravasation_healthy                                                  NaN\nextravasation_injury                                                   NaN\nkidney_healthy                                                         NaN\nkidney_low                                                             NaN\nkidney_high                                                            NaN\nliver_healthy                                                          NaN\nliver_low                                                              NaN\nliver_high                                                             NaN\nspleen_healthy                                                         NaN\nspleen_low                                                             NaN\nspleen_high                                                            NaN\nany_injury                                                             NaN\nmaked_files                                                            NaN\ntrain_paths              /kaggle/input/rsna-2023-abdominal-trauma-detec...\nName: 68949, dtype: object\n```\nwhich showed that the `12580` only exists in `train_images` and we do not have `annotations`. I further rechecked the anomly with \n```\nIn : pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')['patient_id']['12580']\n\nOut : ---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nCell In[20], line 1\n----> 1 pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')['patient_id']['12580']\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/series.py:981, in Series.__getitem__(self, key)\n    978     return self._values[key]\n    980 elif key_is_scalar:\n--> 981     return self._get_value(key)\n    983 if is_hashable(key):\n    984     # Otherwise index.get_value will raise InvalidIndexError\n    985     try:\n    986         # For labels that don't resolve as scalars like tuples and frozensets\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/series.py:1089, in Series._get_value(self, label, takeable)\n   1086     return self._values[label]\n   1088 # Similar to Index.get_value, but we do not fall back to positional\n-> 1089 loc = self.index.get_loc(label)\n   1090 return self.index._get_values_for_loc(self, loc, label)\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/range.py:395, in RangeIndex.get_loc(self, key, method, tolerance)\n    393             raise KeyError(key) from err\n    394     self._check_indexing_error(key)\n--> 395     raise KeyError(key)\n    396 return super().get_loc(key, method=method, tolerance=tolerance)\n\nKeyError: '12580' \n```\n```\nIn : for value in sorted(os.listdir('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images')):\n    if value in str(train['patient_id']) : print(value)\n\nOut : \n\n10004\n10005\n10007\n10026\n10051\n26\n43\n951\n96\n980\n9951\n9960\n9961\n9980\n9983\n```\n```\nIn : train['patient_id'] = train['patient_id'].astype(int)\n\n10082 in train['patient_id']\n\nOut : False\n```",
    "2372395": "Yes, the train images file includes some series that were dropped from the final dataset after QC checks. I should have dropped them from `train_images.csv` as well but overlooked it. Sorry for the confusion!"
  }
}