{
  "id": 440269,
  "title": "Don't trust test_dicom_tags.parquet",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/440269",
  "author_name": "Gunes Evitan",
  "post_date": "2023-09-14T11:10:00.515000",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I was using train_dicom_tags.parquet since all DICOM attributes were extracted perfectly but it's not the case for test. I found out submissions fail when I use test_dicom_tags.parquet the same way I use train_dicom_tags.parquet.</p>\n<p>I made a dummy submission with only these lines added and it failed.</p>\n<pre><code>df_dicom_tags[] = df_dicom_tags[].apply( x: (x).split()[]).astype(np.uint64)\ndf_dicom_tags[] = df_dicom_tags[].apply( x: (x).split()[]).astype(np.uint64)\ndf_dicom_tags[] = df_dicom_tags[].apply( x: (x).split()[].split()[]).astype(np.uint64)\ndf_dicom_tags[] = df_dicom_tags[].apply( x: (x)[-]).astype(np.float32)\ndf_dicom_tags.sort_values(by=[, , ], ascending=[, , ], inplace=)\ndf_dicom_tags.reset_index(drop=, inplace=)\n</code></pre>\n<p>It's probably failing because of the line that creates z_position column so some or at least 1 row has missing ImagePositionPatient value. We can't risk it and go along with that parquet file. Z position should be extracted on the fly while creating 3d data. </p>",
  "messages": [
    {
      "id": 2438575,
      "postDate": "2023-09-14T11:10:00.517Z",
      "content": "<p>I was using train_dicom_tags.parquet since all DICOM attributes were extracted perfectly but it's not the case for test. I found out submissions fail when I use test_dicom_tags.parquet the same way I use train_dicom_tags.parquet.</p>\n<p>I made a dummy submission with only these lines added and it failed.</p>\n<pre><code>df_dicom_tags[] = df_dicom_tags[].apply( x: (x).split()[]).astype(np.uint64)\ndf_dicom_tags[] = df_dicom_tags[].apply( x: (x).split()[]).astype(np.uint64)\ndf_dicom_tags[] = df_dicom_tags[].apply( x: (x).split()[].split()[]).astype(np.uint64)\ndf_dicom_tags[] = df_dicom_tags[].apply( x: (x)[-]).astype(np.float32)\ndf_dicom_tags.sort_values(by=[, , ], ascending=[, , ], inplace=)\ndf_dicom_tags.reset_index(drop=, inplace=)\n</code></pre>\n<p>It's probably failing because of the line that creates z_position column so some or at least 1 row has missing ImagePositionPatient value. We can't risk it and go along with that parquet file. Z position should be extracted on the fly while creating 3d data. </p>",
      "rawMarkdown": "I was using train_dicom_tags.parquet since all DICOM attributes were extracted perfectly but it's not the case for test. I found out submissions fail when I use test_dicom_tags.parquet the same way I use train_dicom_tags.parquet.\n\nI made a dummy submission with only these lines added and it failed.\n\n```python\ndf_dicom_tags['scan_id'] = df_dicom_tags['path'].apply(lambda x: str(x).split('/')[2]).astype(np.uint64)\ndf_dicom_tags['patient_id'] = df_dicom_tags['path'].apply(lambda x: str(x).split('/')[1]).astype(np.uint64)\ndf_dicom_tags['slice_id'] = df_dicom_tags['path'].apply(lambda x: str(x).split('/')[3].split('.')[0]).astype(np.uint64)\ndf_dicom_tags['z_position'] = df_dicom_tags['ImagePositionPatient'].apply(lambda x: eval(x)[-1]).astype(np.float32)\ndf_dicom_tags.sort_values(by=['patient_id', 'scan_id', 'z_position'], ascending=[True, True, False], inplace=True)\ndf_dicom_tags.reset_index(drop=True, inplace=True)\n```\n  \nIt's probably failing because of the line that creates z_position column so some or at least 1 row has missing ImagePositionPatient value. We can't risk it and go along with that parquet file. Z position should be extracted on the fly while creating 3d data. ",
      "votes": 9
    },
    {
      "id": 2445163,
      "postDate": "2023-09-18T16:32:07.290Z",
      "content": "<p>Acutally, when  i try to use the index <code>(patient, series)</code>  to get all the attributes of corresponding seriese,  it returns with <code>Series.sort_values() got an unexpected keyword argument 'by'</code>. It seems like in the test_dicom_tags.parquet, there is only one one image for every series. That's why it returns with a  <code>pd.Series</code> instead of a <code>pd.DataFrame</code>. I'm wondering dose the hosts really change the  test_dicom_tags.parquet file in their private test set ? Here is the code i used:</p>\n<pre><code>\ndf_dicom = pd.read_parquet()\n\ndf_dicom[] = df_dicom[].astype()\ndf_dicom[] = df_dicom[].apply( x: x.split()[-])\ndf_dicom = df_dicom.set_index([, ]).sort_index()\n\np_ids = os.listdir()\n p  tqdm(p_ids):\n    p_folder = J(, p)\n    s_ids = os.listdir(p_folder)\n     s  s_ids:\n        curr_df=df_dicom.loc[(p, s)]\n        curr_df=curr_df.sort_values(by=)\n</code></pre>",
      "rawMarkdown": "Acutally, when  i try to use the index ``(patient, series)``  to get all the attributes of corresponding seriese,  it returns with ``Series.sort_values() got an unexpected keyword argument 'by'``. It seems like in the test_dicom_tags.parquet, there is only one one image for every series. That's why it returns with a  ``pd.Series`` instead of a ``pd.DataFrame``. I'm wondering dose the hosts really change the  test_dicom_tags.parquet file in their private test set ? Here is the code i used:\n\n```python\n\n# Get DICOM meta-data\ndf_dicom = pd.read_parquet(f\"/kaggle/input/rsna-2023-abdominal-trauma-detection/{Config.MODE}_dicom_tags.parquet\")\n# Gest series folder\ndf_dicom[\"PatientID\"] = df_dicom[\"PatientID\"].astype(str)\ndf_dicom[\"serie\"] = df_dicom[\"SeriesInstanceUID\"].apply(lambda x: x.split(\".\")[-1])\ndf_dicom = df_dicom.set_index([\"PatientID\", \"serie\"]).sort_index()\n\np_ids = os.listdir(f\"{BASE_PATH}/{Config.MODE}_images\")\nfor p in tqdm(p_ids):\n    p_folder = J(f\"{BASE_PATH}/{Config.MODE}_images\", p)\n    s_ids = os.listdir(p_folder)\n    for s in s_ids:\n        curr_df=df_dicom.loc[(p, s)]\n        curr_df=curr_df.sort_values(by='InstanceNumber')\n```",
      "votes": 1
    },
    {
      "id": 2445449,
      "postDate": "2023-09-18T20:46:18.760Z",
      "content": "<p>On my end, I am also debugging my submission. I also have problems and spent the entire week trying to spot the problem. Please, share your findings. Thanks! </p>",
      "rawMarkdown": "On my end, I am also debugging my submission. I also have problems and spent the entire week trying to spot the problem. Please, share your findings. Thanks! ",
      "votes": 2
    },
    {
      "id": 2442902,
      "postDate": "2023-09-17T11:38:19.790Z",
      "content": "<p>I am currently also debugging my submission, and everything is pointing to some operations failing on the data in test_dicom_tags.parquet</p>\n<p>Will post here once I figure out why, but it might take a while because of the 5 submission per day limit.</p>",
      "rawMarkdown": "I am currently also debugging my submission, and everything is pointing to some operations failing on the data in test_dicom_tags.parquet\n\nWill post here once I figure out why, but it might take a while because of the 5 submission per day limit.",
      "votes": 2
    },
    {
      "id": 2439283,
      "postDate": "2023-09-14T18:32:48.410Z",
      "content": "<p>I couldn't get a lambda function to work properly here either, so I hackishly treat them as strings and split them. I'm sure there's a better method, but this doesn't cause subs to fail for me.</p>\n<pre><code>tags = pd.read_parquet()\ntags[] = tags[]..replace(,, regex=).replace(,, regex=)\ntags[[,,]] = tags[]..split(, expand=).astype()\n</code></pre>",
      "rawMarkdown": "I couldn't get a lambda function to work properly here either, so I hackishly treat them as strings and split them. I'm sure there's a better method, but this doesn't cause subs to fail for me.\n\n```python\ntags = pd.read_parquet('/kaggle/input/rsna-2023-abdominal-trauma-detection/test_dicom_tags.parquet')\ntags['ImagePositionPatient'] = tags['ImagePositionPatient'].str.replace(\"\\[\",\"\", regex=True).replace(\"\\]\",\"\", regex=True)\ntags[['x','y','z']] = tags['ImagePositionPatient'].str.split(', ', expand=True).astype('float')\n```",
      "votes": 2
    },
    {
      "id": 2456101,
      "postDate": "2023-09-26T02:03:09.633Z",
      "content": "<p>I faced the same problem… what should we do?</p>",
      "rawMarkdown": "I faced the same problem... what should we do?"
    }
  ],
  "comments": [
    {
      "id": 2445163,
      "author_name": "Gooring",
      "author_url": "",
      "post_date": "2023-09-18T16:32:07.290000",
      "content": "<p>Acutally, when  i try to use the index <code>(patient, series)</code>  to get all the attributes of corresponding seriese,  it returns with <code>Series.sort_values() got an unexpected keyword argument 'by'</code>. It seems like in the test_dicom_tags.parquet, there is only one one image for every series. That's why it returns with a  <code>pd.Series</code> instead of a <code>pd.DataFrame</code>. I'm wondering dose the hosts really change the  test_dicom_tags.parquet file in their private test set ? Here is the code i used:</p>\n<pre><code>\ndf_dicom = pd.read_parquet()\n\ndf_dicom[] = df_dicom[].astype()\ndf_dicom[] = df_dicom[].apply( x: x.split()[-])\ndf_dicom = df_dicom.set_index([, ]).sort_index()\n\np_ids = os.listdir()\n p  tqdm(p_ids):\n    p_folder = J(, p)\n    s_ids = os.listdir(p_folder)\n     s  s_ids:\n        curr_df=df_dicom.loc[(p, s)]\n        curr_df=curr_df.sort_values(by=)\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2445449,
      "author_name": "Byzantine Monk",
      "author_url": "",
      "post_date": "2023-09-18T20:46:18.760000",
      "content": "<p>On my end, I am also debugging my submission. I also have problems and spent the entire week trying to spot the problem. Please, share your findings. Thanks! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2442902,
      "author_name": "amobular",
      "author_url": "",
      "post_date": "2023-09-17T11:38:19.790000",
      "content": "<p>I am currently also debugging my submission, and everything is pointing to some operations failing on the data in test_dicom_tags.parquet</p>\n<p>Will post here once I figure out why, but it might take a while because of the 5 submission per day limit.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2439283,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2023-09-14T18:32:48.410000",
      "content": "<p>I couldn't get a lambda function to work properly here either, so I hackishly treat them as strings and split them. I'm sure there's a better method, but this doesn't cause subs to fail for me.</p>\n<pre><code>tags = pd.read_parquet()\ntags[] = tags[]..replace(,, regex=).replace(,, regex=)\ntags[[,,]] = tags[]..split(, expand=).astype()\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2456101,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "2023-09-26T02:03:09.633000",
      "content": "<p>I faced the same problem… what should we do?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2438575": "I was using train_dicom_tags.parquet since all DICOM attributes were extracted perfectly but it's not the case for test. I found out submissions fail when I use test_dicom_tags.parquet the same way I use train_dicom_tags.parquet.\n\nI made a dummy submission with only these lines added and it failed.\n\n```python\ndf_dicom_tags['scan_id'] = df_dicom_tags['path'].apply(lambda x: str(x).split('/')[2]).astype(np.uint64)\ndf_dicom_tags['patient_id'] = df_dicom_tags['path'].apply(lambda x: str(x).split('/')[1]).astype(np.uint64)\ndf_dicom_tags['slice_id'] = df_dicom_tags['path'].apply(lambda x: str(x).split('/')[3].split('.')[0]).astype(np.uint64)\ndf_dicom_tags['z_position'] = df_dicom_tags['ImagePositionPatient'].apply(lambda x: eval(x)[-1]).astype(np.float32)\ndf_dicom_tags.sort_values(by=['patient_id', 'scan_id', 'z_position'], ascending=[True, True, False], inplace=True)\ndf_dicom_tags.reset_index(drop=True, inplace=True)\n```\n  \nIt's probably failing because of the line that creates z_position column so some or at least 1 row has missing ImagePositionPatient value. We can't risk it and go along with that parquet file. Z position should be extracted on the fly while creating 3d data. ",
    "2445163": "Acutally, when  i try to use the index ``(patient, series)``  to get all the attributes of corresponding seriese,  it returns with ``Series.sort_values() got an unexpected keyword argument 'by'``. It seems like in the test_dicom_tags.parquet, there is only one one image for every series. That's why it returns with a  ``pd.Series`` instead of a ``pd.DataFrame``. I'm wondering dose the hosts really change the  test_dicom_tags.parquet file in their private test set ? Here is the code i used:\n\n```python\n\n# Get DICOM meta-data\ndf_dicom = pd.read_parquet(f\"/kaggle/input/rsna-2023-abdominal-trauma-detection/{Config.MODE}_dicom_tags.parquet\")\n# Gest series folder\ndf_dicom[\"PatientID\"] = df_dicom[\"PatientID\"].astype(str)\ndf_dicom[\"serie\"] = df_dicom[\"SeriesInstanceUID\"].apply(lambda x: x.split(\".\")[-1])\ndf_dicom = df_dicom.set_index([\"PatientID\", \"serie\"]).sort_index()\n\np_ids = os.listdir(f\"{BASE_PATH}/{Config.MODE}_images\")\nfor p in tqdm(p_ids):\n    p_folder = J(f\"{BASE_PATH}/{Config.MODE}_images\", p)\n    s_ids = os.listdir(p_folder)\n    for s in s_ids:\n        curr_df=df_dicom.loc[(p, s)]\n        curr_df=curr_df.sort_values(by='InstanceNumber')\n```",
    "2445449": "On my end, I am also debugging my submission. I also have problems and spent the entire week trying to spot the problem. Please, share your findings. Thanks! ",
    "2442902": "I am currently also debugging my submission, and everything is pointing to some operations failing on the data in test_dicom_tags.parquet\n\nWill post here once I figure out why, but it might take a while because of the 5 submission per day limit.",
    "2439283": "I couldn't get a lambda function to work properly here either, so I hackishly treat them as strings and split them. I'm sure there's a better method, but this doesn't cause subs to fail for me.\n\n```python\ntags = pd.read_parquet('/kaggle/input/rsna-2023-abdominal-trauma-detection/test_dicom_tags.parquet')\ntags['ImagePositionPatient'] = tags['ImagePositionPatient'].str.replace(\"\\[\",\"\", regex=True).replace(\"\\]\",\"\", regex=True)\ntags[['x','y','z']] = tags['ImagePositionPatient'].str.split(', ', expand=True).astype('float')\n```",
    "2456101": "I faced the same problem... what should we do?"
  }
}