{
  "id": 433004,
  "title": "Reduced Dataset from 400 GB to 7.5 GB",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/433004",
  "author_name": "AleNic",
  "post_date": "2023-08-20T00:09:39.356000",
  "votes": 15,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, I added this dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm\" target=\"_blank\">https://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm</a></p>\n<p>generated with this notebook:<br>\n<a href=\"https://www.kaggle.com/code/alenic/dataset-size-reduction\" target=\"_blank\">https://www.kaggle.com/code/alenic/dataset-size-reduction</a></p>\n<p>Hope that this can help to start with a baseline.</p>",
  "messages": [
    {
      "id": 2398787,
      "postDate": "2023-08-20T00:09:39.357Z",
      "content": "<p>Hi, I added this dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm\" target=\"_blank\">https://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm</a></p>\n<p>generated with this notebook:<br>\n<a href=\"https://www.kaggle.com/code/alenic/dataset-size-reduction\" target=\"_blank\">https://www.kaggle.com/code/alenic/dataset-size-reduction</a></p>\n<p>Hope that this can help to start with a baseline.</p>",
      "rawMarkdown": "Hi, I added this dataset:\nhttps://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm\n\ngenerated with this notebook:\nhttps://www.kaggle.com/code/alenic/dataset-size-reduction\n\nHope that this can help to start with a baseline.",
      "votes": 15
    },
    {
      "id": 2421959,
      "postDate": "2023-09-03T16:01:09.693Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc86860d3a7acc5ab0eb32e180a617ff8%2FSelection_999(3033).png?generation=1693756839048817&amp;alt=media\" alt=\"\"></p>\n<p>you intensity is wrong. some series end up with almost black images</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc86860d3a7acc5ab0eb32e180a617ff8%2FSelection_999(3033).png?generation=1693756839048817&alt=media)\n\nyou intensity is wrong. some series end up with almost black images",
      "votes": 2,
      "replies": [
        {
          "id": 2435458,
          "postDate": "2023-09-13T02:14:53.553Z",
          "content": "<p>He is using the widowing so some images become dark. Here is my data with basic conversion, windowing, and high-contrast. Please upvote it if you find it useful! Thank you!<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/438690\" target=\"_blank\">Data DICOM to 3 different types of Image (Grayscale)</a></p>",
          "rawMarkdown": "He is using the widowing so some images become dark. Here is my data with basic conversion, windowing, and high-contrast. Please upvote it if you find it useful! Thank you!\n[Data DICOM to 3 different types of Image (Grayscale)](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/438690)"
        }
      ]
    },
    {
      "id": 2419978,
      "postDate": "2023-09-02T09:57:42.240Z",
      "content": "<p>Thanks! To me this looks like a meaningful loss of image quality.</p>",
      "rawMarkdown": "Thanks! To me this looks like a meaningful loss of image quality.",
      "votes": -1,
      "replies": [
        {
          "id": 2746866,
          "postDate": "2024-04-11T14:17:06.400Z",
          "content": "<p>at least he tried to help, why don't you generate better data 🤗?!</p>",
          "rawMarkdown": "at least he tried to help, why don't you generate better data 🤗?!"
        }
      ]
    },
    {
      "id": 2746850,
      "postDate": "2024-04-11T14:09:50.770Z",
      "content": "<p>I'm sorry but I think <strong>42%</strong> data is corrupted🫠. <br>\nI designed code to check scan with almost black frames with two thresholds one for black frames 98.0 and other for whole scan 78 </p>\n<p><em>to try code yourself just add data path in classify_scans() function</em></p>\n<pre><code>import os\nimport cv2\nimport numpy as np\nfrom tqdm import tqdm\nimport concurrent\n\ndef (img_path):\n    img = cv2.(img_path, cv2.IMREAD_GRAYSCALE)\n    black_pixels = np.(img == )\n    total_pixels = img.shape[] * img.shape[]\n    return (black_pixels / total_pixels) * \n\ndef (folder_path, threshold):\n    img_paths = [os.path.(folder_path, img_name) for img_name in os.(folder_path)]\n    with concurrent.futures.() as executor:\n        black_percentages = (executor.(calculate_black_percentage, img_paths))\n    n_black_images = ( for pct in black_percentages if pct &gt;= threshold)\n    return (n_black_images / (black_percentages)) * \n\ndef (data_path=, threshold=, scan_thresh=):\n    patient_ids = os.(data_path)\n    corrupted_scans = []\n\n    total_patients = (patient_ids)\n    progress_bar = (total=total_patients, desc=)\n\n    for patient_id in patient_ids:\n        patient_path = os.path.(data_path, patient_id)\n        scan_ids = os.(patient_path)\n        for scan_id in scan_ids:\n            scan_path = os.path.(patient_path, scan_id)\n\n            if (scan_path, threshold) &gt;= scan_thresh:\n                corrupted_scans.(os.path.(data_path, patient_id, scan_id))\n        progress_bar.()\n\n    progress_bar.()\n    return corrupted_scans\n\ndebug_lst =  ()\n(((debug_lst) / (os.())) * )\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8229747%2F60885e453abc03ee290618fcc1ba6e95%2FScreenshot%202024-04-11%20161057.png?generation=1712844710488920&amp;alt=media\"><br>\nI checked 10 scans from list and all of them are truly black.🙂</p>\n<p>BUT the good news that after removing corrupted scans(if you do this) there will be 905 empty patient folder not all scans for corrupted patient is wrong</p>",
      "rawMarkdown": "I'm sorry but I think **42%** data is corrupted🫠. \nI designed code to check scan with almost black frames with two thresholds one for black frames 98.0 and other for whole scan 78 \n\n*to try code yourself just add data path in classify_scans() function*\n```\nimport os\nimport cv2\nimport numpy as np\nfrom tqdm import tqdm\nimport concurrent.futures\n\ndef calculate_black_percentage(img_path):\n    img = cv2.imread(img_path, cv2.IMREAD_GRAYSCALE)\n    black_pixels = np.sum(img == 0)\n    total_pixels = img.shape[0] * img.shape[1]\n    return (black_pixels / total_pixels) * 100\n\ndef get_scan_pct(folder_path, threshold):\n    img_paths = [os.path.join(folder_path, img_name) for img_name in os.listdir(folder_path)]\n    with concurrent.futures.ThreadPoolExecutor() as executor:\n        black_percentages = list(executor.map(calculate_black_percentage, img_paths))\n    n_black_images = sum(1 for pct in black_percentages if pct >= threshold)\n    return (n_black_images / len(black_percentages)) * 100\n\ndef classify_scans(data_path=\"C:/codingWorkspace/reduced_256_tickness_5\", threshold=98.0, scan_thresh=78):\n    patient_ids = os.listdir(data_path)\n    corrupted_scans = []\n\n    total_patients = len(patient_ids)\n    progress_bar = tqdm(total=total_patients, desc=\"Processing Patients\")\n\n    for patient_id in patient_ids:\n        patient_path = os.path.join(data_path, patient_id)\n        scan_ids = os.listdir(patient_path)\n        for scan_id in scan_ids:\n            scan_path = os.path.join(patient_path, scan_id)\n\n            if get_scan_pct(scan_path, threshold) >= scan_thresh:\n                corrupted_scans.append(os.path.join(data_path, patient_id, scan_id))\n        progress_bar.update(1)\n\n    progress_bar.close()\n    return corrupted_scans\n\ndebug_lst =  classify_scans()\nprint((len(debug_lst) / len(os.listdir(\"C:/codingWorkspace/reduced_256_tickness_5\"))) * 100)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8229747%2F60885e453abc03ee290618fcc1ba6e95%2FScreenshot%202024-04-11%20161057.png?generation=1712844710488920&alt=media)\nI checked 10 scans from list and all of them are truly black.🙂\n\nBUT the good news that after removing corrupted scans(if you do this) there will be 905 empty patient folder not all scans for corrupted patient is wrong\n"
    }
  ],
  "comments": [
    {
      "id": 2421959,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-09-03T16:01:09.693000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc86860d3a7acc5ab0eb32e180a617ff8%2FSelection_999(3033).png?generation=1693756839048817&amp;alt=media\" alt=\"\"></p>\n<p>you intensity is wrong. some series end up with almost black images</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2435458,
          "author_name": "Minh-Long Pham",
          "author_url": "",
          "post_date": "2023-09-13T02:14:53.553000",
          "content": "<p>He is using the widowing so some images become dark. Here is my data with basic conversion, windowing, and high-contrast. Please upvote it if you find it useful! Thank you!<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/438690\" target=\"_blank\">Data DICOM to 3 different types of Image (Grayscale)</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2419978,
      "author_name": "Gyula Maloveczky4",
      "author_url": "",
      "post_date": "2023-09-02T09:57:42.240000",
      "content": "<p>Thanks! To me this looks like a meaningful loss of image quality.</p>",
      "votes": -1,
      "replies": [
        {
          "id": 2746866,
          "author_name": "Ahmed Kamal El-Senussi",
          "author_url": "",
          "post_date": "2024-04-11T14:17:06.400000",
          "content": "<p>at least he tried to help, why don't you generate better data 🤗?!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2746850,
      "author_name": "Ahmed Kamal El-Senussi",
      "author_url": "",
      "post_date": "2024-04-11T14:09:50.770000",
      "content": "<p>I'm sorry but I think <strong>42%</strong> data is corrupted🫠. <br>\nI designed code to check scan with almost black frames with two thresholds one for black frames 98.0 and other for whole scan 78 </p>\n<p><em>to try code yourself just add data path in classify_scans() function</em></p>\n<pre><code>import os\nimport cv2\nimport numpy as np\nfrom tqdm import tqdm\nimport concurrent\n\ndef (img_path):\n    img = cv2.(img_path, cv2.IMREAD_GRAYSCALE)\n    black_pixels = np.(img == )\n    total_pixels = img.shape[] * img.shape[]\n    return (black_pixels / total_pixels) * \n\ndef (folder_path, threshold):\n    img_paths = [os.path.(folder_path, img_name) for img_name in os.(folder_path)]\n    with concurrent.futures.() as executor:\n        black_percentages = (executor.(calculate_black_percentage, img_paths))\n    n_black_images = ( for pct in black_percentages if pct &gt;= threshold)\n    return (n_black_images / (black_percentages)) * \n\ndef (data_path=, threshold=, scan_thresh=):\n    patient_ids = os.(data_path)\n    corrupted_scans = []\n\n    total_patients = (patient_ids)\n    progress_bar = (total=total_patients, desc=)\n\n    for patient_id in patient_ids:\n        patient_path = os.path.(data_path, patient_id)\n        scan_ids = os.(patient_path)\n        for scan_id in scan_ids:\n            scan_path = os.path.(patient_path, scan_id)\n\n            if (scan_path, threshold) &gt;= scan_thresh:\n                corrupted_scans.(os.path.(data_path, patient_id, scan_id))\n        progress_bar.()\n\n    progress_bar.()\n    return corrupted_scans\n\ndebug_lst =  ()\n(((debug_lst) / (os.())) * )\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8229747%2F60885e453abc03ee290618fcc1ba6e95%2FScreenshot%202024-04-11%20161057.png?generation=1712844710488920&amp;alt=media\"><br>\nI checked 10 scans from list and all of them are truly black.🙂</p>\n<p>BUT the good news that after removing corrupted scans(if you do this) there will be 905 empty patient folder not all scans for corrupted patient is wrong</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2398787": "Hi, I added this dataset:\nhttps://www.kaggle.com/datasets/alenic/rsna-2023-atd-reduced-256-5mm\n\ngenerated with this notebook:\nhttps://www.kaggle.com/code/alenic/dataset-size-reduction\n\nHope that this can help to start with a baseline.",
    "2421959": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc86860d3a7acc5ab0eb32e180a617ff8%2FSelection_999(3033).png?generation=1693756839048817&alt=media)\n\nyou intensity is wrong. some series end up with almost black images",
    "2419978": "Thanks! To me this looks like a meaningful loss of image quality.",
    "2746850": "I'm sorry but I think **42%** data is corrupted🫠. \nI designed code to check scan with almost black frames with two thresholds one for black frames 98.0 and other for whole scan 78 \n\n*to try code yourself just add data path in classify_scans() function*\n```\nimport os\nimport cv2\nimport numpy as np\nfrom tqdm import tqdm\nimport concurrent.futures\n\ndef calculate_black_percentage(img_path):\n    img = cv2.imread(img_path, cv2.IMREAD_GRAYSCALE)\n    black_pixels = np.sum(img == 0)\n    total_pixels = img.shape[0] * img.shape[1]\n    return (black_pixels / total_pixels) * 100\n\ndef get_scan_pct(folder_path, threshold):\n    img_paths = [os.path.join(folder_path, img_name) for img_name in os.listdir(folder_path)]\n    with concurrent.futures.ThreadPoolExecutor() as executor:\n        black_percentages = list(executor.map(calculate_black_percentage, img_paths))\n    n_black_images = sum(1 for pct in black_percentages if pct >= threshold)\n    return (n_black_images / len(black_percentages)) * 100\n\ndef classify_scans(data_path=\"C:/codingWorkspace/reduced_256_tickness_5\", threshold=98.0, scan_thresh=78):\n    patient_ids = os.listdir(data_path)\n    corrupted_scans = []\n\n    total_patients = len(patient_ids)\n    progress_bar = tqdm(total=total_patients, desc=\"Processing Patients\")\n\n    for patient_id in patient_ids:\n        patient_path = os.path.join(data_path, patient_id)\n        scan_ids = os.listdir(patient_path)\n        for scan_id in scan_ids:\n            scan_path = os.path.join(patient_path, scan_id)\n\n            if get_scan_pct(scan_path, threshold) >= scan_thresh:\n                corrupted_scans.append(os.path.join(data_path, patient_id, scan_id))\n        progress_bar.update(1)\n\n    progress_bar.close()\n    return corrupted_scans\n\ndebug_lst =  classify_scans()\nprint((len(debug_lst) / len(os.listdir(\"C:/codingWorkspace/reduced_256_tickness_5\"))) * 100)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8229747%2F60885e453abc03ee290618fcc1ba6e95%2FScreenshot%202024-04-11%20161057.png?generation=1712844710488920&alt=media)\nI checked 10 scans from list and all of them are truly black.🙂\n\nBUT the good news that after removing corrupted scans(if you do this) there will be 905 empty patient folder not all scans for corrupted patient is wrong\n"
  }
}