{
  "id": 431480,
  "title": "Question about the class imbalance",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/431480",
  "author_name": "Jinyoung Seo",
  "post_date": "2023-08-13T17:41:32.615000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Isn't the dataset a little too imbalanced? <br>\nThe number of data with any_injury seems a lot, but there are too many labels and injury for each organ is only a small portion. Does anyone know how to address this problem?</p>",
  "messages": [
    {
      "id": 2388952,
      "postDate": "2023-08-13T17:41:32.617Z",
      "content": "<p>Isn't the dataset a little too imbalanced? <br>\nThe number of data with any_injury seems a lot, but there are too many labels and injury for each organ is only a small portion. Does anyone know how to address this problem?</p>",
      "rawMarkdown": "Isn't the dataset a little too imbalanced? \nThe number of data with any_injury seems a lot, but there are too many labels and injury for each organ is only a small portion. Does anyone know how to address this problem?",
      "votes": 5
    },
    {
      "id": 2390978,
      "postDate": "2023-08-14T21:40:40.597Z",
      "content": "<p>This data brings a lot of interesting issues to address.</p>\n<p>Large file sizes, weak labels and class imbalance.  With only a single leader board score that's a lot better than the mean solution it does not appear that many of us have found how to handle these.</p>\n<p>I started with a fork of the Tensorflow <a href=\"https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-train\" target=\"_blank\">model</a> shared by <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">AWSAF</a> where the data set contains a fraction of the full data.  That certainly helped keep training time down to address the large file size issue, but I found the noise on 5 fold training to be much larger than any signal from my experiments on model parameters, loss types, etc.  </p>\n<p>Training with the full data is a real compute time monster - on my dual GPU system a single fold with 20 epochs took over 24 hours.  I think the model is pretty solid but I believe you do need full 4/5 folds with augmentation to have any handle on the weak labels issue.  4 days to run a single experiment is not going to cut it (I have 3 more machines I can bring into the project but…).</p>\n<p>So my current attempt is to start to address the class imbalance (and the above) by down sampling the full set of images based on the 'any_injury'.  Using the code below the number of rows in my full data goes from 1,694,669 to 925,698.  Still pretty long run times for a 5 fold training and not sure how well it will help with the class imbalance.  If you see me at any point with a LB score better than 0.66 you will know it helped :)   </p>\n<pre><code> CFG.downsample == :\n    \n    df_minority = df[df[] == ]  \n    df_majority = df[df[] == ]\n\n\n    \n    groups_minority = (df_minority.groupby([, ]).groups.keys())\n    groups_majority = (df_majority.groupby([, ]).groups.keys())\n\n    \n     random\n    selected_groups_keys = random.sample(groups_majority, (groups_minority))\n\n    \n    selected_majority_rows = df_majority[df_majority.set_index([, ]).index.isin(selected_groups_keys)]\n\n    \n    balanced_df = pd.concat([df_minority, selected_majority_rows])\n\n    \n    balanced_df = balanced_df.sample(frac=).reset_index(drop=)\n\n    df = balanced_df\n\n    \n    df.sort_values([, , ], inplace=)\n     balanced_df\n    gc.collect()\n    display()\n</code></pre>",
      "rawMarkdown": "This data brings a lot of interesting issues to address.\n\nLarge file sizes, weak labels and class imbalance.  With only a single leader board score that's a lot better than the mean solution it does not appear that many of us have found how to handle these.\n\nI started with a fork of the Tensorflow [model](https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-train) shared by [AWSAF](https://www.kaggle.com/awsaf49) where the data set contains a fraction of the full data.  That certainly helped keep training time down to address the large file size issue, but I found the noise on 5 fold training to be much larger than any signal from my experiments on model parameters, loss types, etc.  \n\nTraining with the full data is a real compute time monster - on my dual GPU system a single fold with 20 epochs took over 24 hours.  I think the model is pretty solid but I believe you do need full 4/5 folds with augmentation to have any handle on the weak labels issue.  4 days to run a single experiment is not going to cut it (I have 3 more machines I can bring into the project but...).\n\nSo my current attempt is to start to address the class imbalance (and the above) by down sampling the full set of images based on the 'any_injury'.  Using the code below the number of rows in my full data goes from 1,694,669 to 925,698.  Still pretty long run times for a 5 fold training and not sure how well it will help with the class imbalance.  If you see me at any point with a LB score better than 0.66 you will know it helped :)   \n\n\n\n```python\nif CFG.downsample == True:\n    # Split the data into minority and majority\n    df_minority = df[df['any_injury'] == 1]  # 1 is the minority class\n    df_majority = df[df['any_injury'] == 0]\n\n\n    # Group by 'patient_id' and 'series_id' for both classes\n    groups_minority = list(df_minority.groupby(['patient_id', 'series_id']).groups.keys())\n    groups_majority = list(df_majority.groupby(['patient_id', 'series_id']).groups.keys())\n\n    # Randomly sample group keys from the majority set\n    import random\n    selected_groups_keys = random.sample(groups_majority, len(groups_minority))\n\n    # Get the rows corresponding to the selected groups\n    selected_majority_rows = df_majority[df_majority.set_index(['patient_id', 'series_id']).index.isin(selected_groups_keys)]\n\n    # Concatenate the minority class data with the selected majority class data\n    balanced_df = pd.concat([df_minority, selected_majority_rows])\n\n    # Shuffle the rows to randomize the data order\n    balanced_df = balanced_df.sample(frac=1).reset_index(drop=True)\n\n    df = balanced_df\n\n    # Sort values in dataframe by 'patient_id', 'series_id', and 'instance_number'\n    df.sort_values(['patient_id', 'series_id', 'instance_number'], inplace=True)\n    del balanced_df\n    gc.collect()\n    display('downsampled data frame')\n```",
      "votes": 4,
      "replies": [
        {
          "id": 2392115,
          "postDate": "2023-08-15T13:30:23.620Z",
          "content": "<p>Thank you for sharing this valuable information! Maybe I should also consider down sampling…</p>",
          "rawMarkdown": "Thank you for sharing this valuable information! Maybe I should also consider down sampling..."
        }
      ]
    },
    {
      "id": 2393192,
      "postDate": "2023-08-16T07:15:42.180Z",
      "content": "<p>This competition's evaluation metric already addresses this issue.</p>",
      "rawMarkdown": "This competition's evaluation metric already addresses this issue.",
      "replies": [
        {
          "id": 2393995,
          "postDate": "2023-08-16T16:38:54.300Z",
          "content": "<p>I am not using the competition metric for training my model, only a simple sum loss.  But perhaps we are answering different questions (not sure which was Jinyoug's).   I don't think that using the metric addresses the issue of training a 'not balanced' data set.</p>\n<p>The competition metric seems more aimed at balancing clinical importance.</p>",
          "rawMarkdown": "I am not using the competition metric for training my model, only a simple sum loss.  But perhaps we are answering different questions (not sure which was Jinyoug's).   I don't think that using the metric addresses the issue of training a 'not balanced' data set.\n\nThe competition metric seems more aimed at balancing clinical importance.\n\n"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2390978,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2023-08-14T21:40:40.597000",
      "content": "<p>This data brings a lot of interesting issues to address.</p>\n<p>Large file sizes, weak labels and class imbalance.  With only a single leader board score that's a lot better than the mean solution it does not appear that many of us have found how to handle these.</p>\n<p>I started with a fork of the Tensorflow <a href=\"https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-train\" target=\"_blank\">model</a> shared by <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">AWSAF</a> where the data set contains a fraction of the full data.  That certainly helped keep training time down to address the large file size issue, but I found the noise on 5 fold training to be much larger than any signal from my experiments on model parameters, loss types, etc.  </p>\n<p>Training with the full data is a real compute time monster - on my dual GPU system a single fold with 20 epochs took over 24 hours.  I think the model is pretty solid but I believe you do need full 4/5 folds with augmentation to have any handle on the weak labels issue.  4 days to run a single experiment is not going to cut it (I have 3 more machines I can bring into the project but…).</p>\n<p>So my current attempt is to start to address the class imbalance (and the above) by down sampling the full set of images based on the 'any_injury'.  Using the code below the number of rows in my full data goes from 1,694,669 to 925,698.  Still pretty long run times for a 5 fold training and not sure how well it will help with the class imbalance.  If you see me at any point with a LB score better than 0.66 you will know it helped :)   </p>\n<pre><code> CFG.downsample == :\n    \n    df_minority = df[df[] == ]  \n    df_majority = df[df[] == ]\n\n\n    \n    groups_minority = (df_minority.groupby([, ]).groups.keys())\n    groups_majority = (df_majority.groupby([, ]).groups.keys())\n\n    \n     random\n    selected_groups_keys = random.sample(groups_majority, (groups_minority))\n\n    \n    selected_majority_rows = df_majority[df_majority.set_index([, ]).index.isin(selected_groups_keys)]\n\n    \n    balanced_df = pd.concat([df_minority, selected_majority_rows])\n\n    \n    balanced_df = balanced_df.sample(frac=).reset_index(drop=)\n\n    df = balanced_df\n\n    \n    df.sort_values([, , ], inplace=)\n     balanced_df\n    gc.collect()\n    display()\n</code></pre>",
      "votes": 4,
      "replies": [
        {
          "id": 2392115,
          "author_name": "Jinyoung Seo",
          "author_url": "",
          "post_date": "2023-08-15T13:30:23.620000",
          "content": "<p>Thank you for sharing this valuable information! Maybe I should also consider down sampling…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2393192,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-08-16T07:15:42.180000",
      "content": "<p>This competition's evaluation metric already addresses this issue.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2393995,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2023-08-16T16:38:54.300000",
          "content": "<p>I am not using the competition metric for training my model, only a simple sum loss.  But perhaps we are answering different questions (not sure which was Jinyoug's).   I don't think that using the metric addresses the issue of training a 'not balanced' data set.</p>\n<p>The competition metric seems more aimed at balancing clinical importance.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2388952": "Isn't the dataset a little too imbalanced? \nThe number of data with any_injury seems a lot, but there are too many labels and injury for each organ is only a small portion. Does anyone know how to address this problem?",
    "2390978": "This data brings a lot of interesting issues to address.\n\nLarge file sizes, weak labels and class imbalance.  With only a single leader board score that's a lot better than the mean solution it does not appear that many of us have found how to handle these.\n\nI started with a fork of the Tensorflow [model](https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-train) shared by [AWSAF](https://www.kaggle.com/awsaf49) where the data set contains a fraction of the full data.  That certainly helped keep training time down to address the large file size issue, but I found the noise on 5 fold training to be much larger than any signal from my experiments on model parameters, loss types, etc.  \n\nTraining with the full data is a real compute time monster - on my dual GPU system a single fold with 20 epochs took over 24 hours.  I think the model is pretty solid but I believe you do need full 4/5 folds with augmentation to have any handle on the weak labels issue.  4 days to run a single experiment is not going to cut it (I have 3 more machines I can bring into the project but...).\n\nSo my current attempt is to start to address the class imbalance (and the above) by down sampling the full set of images based on the 'any_injury'.  Using the code below the number of rows in my full data goes from 1,694,669 to 925,698.  Still pretty long run times for a 5 fold training and not sure how well it will help with the class imbalance.  If you see me at any point with a LB score better than 0.66 you will know it helped :)   \n\n\n\n```python\nif CFG.downsample == True:\n    # Split the data into minority and majority\n    df_minority = df[df['any_injury'] == 1]  # 1 is the minority class\n    df_majority = df[df['any_injury'] == 0]\n\n\n    # Group by 'patient_id' and 'series_id' for both classes\n    groups_minority = list(df_minority.groupby(['patient_id', 'series_id']).groups.keys())\n    groups_majority = list(df_majority.groupby(['patient_id', 'series_id']).groups.keys())\n\n    # Randomly sample group keys from the majority set\n    import random\n    selected_groups_keys = random.sample(groups_majority, len(groups_minority))\n\n    # Get the rows corresponding to the selected groups\n    selected_majority_rows = df_majority[df_majority.set_index(['patient_id', 'series_id']).index.isin(selected_groups_keys)]\n\n    # Concatenate the minority class data with the selected majority class data\n    balanced_df = pd.concat([df_minority, selected_majority_rows])\n\n    # Shuffle the rows to randomize the data order\n    balanced_df = balanced_df.sample(frac=1).reset_index(drop=True)\n\n    df = balanced_df\n\n    # Sort values in dataframe by 'patient_id', 'series_id', and 'instance_number'\n    df.sort_values(['patient_id', 'series_id', 'instance_number'], inplace=True)\n    del balanced_df\n    gc.collect()\n    display('downsampled data frame')\n```",
    "2393192": "This competition's evaluation metric already addresses this issue."
  }
}