{
  "id": 444889,
  "title": "The CategoricalCrossentropy Problem",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/444889",
  "author_name": "Andrij",
  "post_date": "2023-10-04T05:54:37.864000",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi to all!<br>\nI don't understand what I'm missing. I use this code as a base from <a href=\"https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-infer\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-infer</a> (with major modifications). Now I'm testing spleen classification, but it seems that tf.keras.losses.CategoricalCrossentropy(label_smoothing=0.05) instead of classifying objects just dances around their frequency:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F9adfa6782e56688e5b0b4a45dd620794%2F999.PNG?generation=1696398858857774&amp;alt=media\" alt=\"\"></p>\n<p>Here, the frequencies differ slightly from the real ones due to the weighting of the classes. I tried Huge loss but it only made my result worse. At the same time, the AUC at this fold is stable at 0.92-0.93. This model performs slightly better than the better public model 0.65, but essentially <br>\nmodel trained using tf.keras.losses.CategoricalCrossentropy does not actually classify the spleen, it only returns the average frequencies of the classes.<br>\nWhat could be the reason? The model cannot distinguish features that separate ct, or did I miss something in the implementation? Considering that 95% of participants could not cross the 0.65 barrier, this problem is not only mine. Important I transfer the whole ct to the model as one object. For the spleen I form a slice. Thank you all for your comments!</p>",
  "messages": [
    {
      "id": 2466862,
      "postDate": "2023-10-04T05:54:37.863Z",
      "content": "<p>Hi to all!<br>\nI don't understand what I'm missing. I use this code as a base from <a href=\"https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-infer\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-infer</a> (with major modifications). Now I'm testing spleen classification, but it seems that tf.keras.losses.CategoricalCrossentropy(label_smoothing=0.05) instead of classifying objects just dances around their frequency:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F9adfa6782e56688e5b0b4a45dd620794%2F999.PNG?generation=1696398858857774&amp;alt=media\" alt=\"\"></p>\n<p>Here, the frequencies differ slightly from the real ones due to the weighting of the classes. I tried Huge loss but it only made my result worse. At the same time, the AUC at this fold is stable at 0.92-0.93. This model performs slightly better than the better public model 0.65, but essentially <br>\nmodel trained using tf.keras.losses.CategoricalCrossentropy does not actually classify the spleen, it only returns the average frequencies of the classes.<br>\nWhat could be the reason? The model cannot distinguish features that separate ct, or did I miss something in the implementation? Considering that 95% of participants could not cross the 0.65 barrier, this problem is not only mine. Important I transfer the whole ct to the model as one object. For the spleen I form a slice. Thank you all for your comments!</p>",
      "rawMarkdown": "Hi to all!\nI don't understand what I'm missing. I use this code as a base from https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-infer (with major modifications). Now I'm testing spleen classification, but it seems that tf.keras.losses.CategoricalCrossentropy(label_smoothing=0.05) instead of classifying objects just dances around their frequency:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F9adfa6782e56688e5b0b4a45dd620794%2F999.PNG?generation=1696398858857774&alt=media)\n\n\nHere, the frequencies differ slightly from the real ones due to the weighting of the classes. I tried Huge loss but it only made my result worse. At the same time, the AUC at this fold is stable at 0.92-0.93. This model performs slightly better than the better public model 0.65, but essentially \nmodel trained using tf.keras.losses.CategoricalCrossentropy does not actually classify the spleen, it only returns the average frequencies of the classes.\nWhat could be the reason? The model cannot distinguish features that separate ct, or did I miss something in the implementation? Considering that 95% of participants could not cross the 0.65 barrier, this problem is not only mine. Important I transfer the whole ct to the model as one object. For the spleen I form a slice. Thank you all for your comments!",
      "votes": 6
    },
    {
      "id": 2468294,
      "postDate": "2023-10-05T11:30:08.690Z",
      "content": "<p>I think that the problem that we are facing might not be related to the choice of model or loss function, rather it is with the data that we are feeding this model. There is a very high probability that the model is not able to learn anything meaningful from the data, and it just learns the frequency of the predictions.</p>",
      "rawMarkdown": "I think that the problem that we are facing might not be related to the choice of model or loss function, rather it is with the data that we are feeding this model. There is a very high probability that the model is not able to learn anything meaningful from the data, and it just learns the frequency of the predictions.",
      "votes": 3,
      "replies": [
        {
          "id": 2472010,
          "postDate": "2023-10-06T21:03:17.143Z",
          "content": "<p>Agree, in my experience full size image doesn't works. To many noise for a weak label. You should realize how to cease area of interests for an every organ. For example it will work if you try to classify images inside bboxes objects like spleen kidney liver. </p>",
          "rawMarkdown": "Agree, in my experience full size image doesn't works. To many noise for a weak label. You should realize how to cease area of interests for an every organ. For example it will work if you try to classify images inside bboxes objects like spleen kidney liver. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 2467290,
      "postDate": "2023-10-04T12:19:38.887Z",
      "content": "<p>I think the worst part is most of the submissions were just weighted scaled mean values of particular target category.</p>",
      "rawMarkdown": "I think the worst part is most of the submissions were just weighted scaled mean values of particular target category.",
      "votes": 3,
      "replies": [
        {
          "id": 2467296,
          "postDate": "2023-10-04T12:24:22.120Z",
          "content": "<p>That's right, these models don't solve the real problem, simple statistics work better.</p>",
          "rawMarkdown": "That's right, these models don't solve the real problem, simple statistics work better.",
          "votes": 2,
          "replies": [
            {
              "id": 2467370,
              "postDate": "2023-10-04T13:47:19.380Z",
              "content": "<p>literally today my youtube homepage has this video : <a href=\"https://youtube.com/shorts/h21Hc8O6nhY?feature=shared\" target=\"_blank\">https://youtube.com/shorts/h21Hc8O6nhY?feature=shared</a></p>",
              "rawMarkdown": "literally today my youtube homepage has this video : https://youtube.com/shorts/h21Hc8O6nhY?feature=shared",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2472473,
      "postDate": "2023-10-07T11:03:41.753Z",
      "content": "<p>I don't know if you noticed but some people instantly broke that barrier you just mentioned (including me). The reason is they have experience of dealing with similar data. If it was early stages, I would share my approach but it's too late to share anything now. </p>",
      "rawMarkdown": "I don't know if you noticed but some people instantly broke that barrier you just mentioned (including me). The reason is they have experience of dealing with similar data. If it was early stages, I would share my approach but it's too late to share anything now. ",
      "votes": 4
    },
    {
      "id": 2472305,
      "postDate": "2023-10-07T07:57:52.490Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a>, I think it is better if you don't mind about </p>\n<blockquote>\n  <p>Considering that 95% of participants could not cross the 0.65 barrier, this problem is not only mine.</p>\n</blockquote>\n<p>I am also one of these 95% you have mentioned.<br>\nI assume most of the people who are not crossing 0.65 barrier are the ones who make the simple modifications to the scales.<br>\nAs I mentioned in one of my notebooks (focusing on this scale modification), there will be big changes to the leaderboard for those who are focusing on the scales (Like myself).<br>\nThat is also why I am working on vision techniques (but it is tough to break that 0.65 as well xD)<br>\nI will share my solution with you once I find one!</p>",
      "rawMarkdown": "Hi @aikhmelnytskyy, I think it is better if you don't mind about \n> Considering that 95% of participants could not cross the 0.65 barrier, this problem is not only mine.\n\nI am also one of these 95% you have mentioned.\nI assume most of the people who are not crossing 0.65 barrier are the ones who make the simple modifications to the scales.\nAs I mentioned in one of my notebooks (focusing on this scale modification), there will be big changes to the leaderboard for those who are focusing on the scales (Like myself).\nThat is also why I am working on vision techniques (but it is tough to break that 0.65 as well xD)\nI will share my solution with you once I find one!",
      "votes": 1
    },
    {
      "id": 2475504,
      "postDate": "2023-10-10T01:03:45.017Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2468294,
      "author_name": "Aakash Rana",
      "author_url": "",
      "post_date": "2023-10-05T11:30:08.690000",
      "content": "<p>I think that the problem that we are facing might not be related to the choice of model or loss function, rather it is with the data that we are feeding this model. There is a very high probability that the model is not able to learn anything meaningful from the data, and it just learns the frequency of the predictions.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2472010,
          "author_name": "Aleksandr Lavrikov",
          "author_url": "",
          "post_date": "2023-10-06T21:03:17.143000",
          "content": "<p>Agree, in my experience full size image doesn't works. To many noise for a weak label. You should realize how to cease area of interests for an every organ. For example it will work if you try to classify images inside bboxes objects like spleen kidney liver. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2467290,
      "author_name": "naidu_tmp",
      "author_url": "",
      "post_date": "2023-10-04T12:19:38.887000",
      "content": "<p>I think the worst part is most of the submissions were just weighted scaled mean values of particular target category.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2467296,
          "author_name": "Andrij",
          "author_url": "",
          "post_date": "2023-10-04T12:24:22.120000",
          "content": "<p>That's right, these models don't solve the real problem, simple statistics work better.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2467370,
              "author_name": "naidu_tmp",
              "author_url": "",
              "post_date": "2023-10-04T13:47:19.380000",
              "content": "<p>literally today my youtube homepage has this video : <a href=\"https://youtube.com/shorts/h21Hc8O6nhY?feature=shared\" target=\"_blank\">https://youtube.com/shorts/h21Hc8O6nhY?feature=shared</a></p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2472473,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-10-07T11:03:41.753000",
      "content": "<p>I don't know if you noticed but some people instantly broke that barrier you just mentioned (including me). The reason is they have experience of dealing with similar data. If it was early stages, I would share my approach but it's too late to share anything now. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2472305,
      "author_name": "Jason Heesang Lee",
      "author_url": "",
      "post_date": "2023-10-07T07:57:52.490000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a>, I think it is better if you don't mind about </p>\n<blockquote>\n  <p>Considering that 95% of participants could not cross the 0.65 barrier, this problem is not only mine.</p>\n</blockquote>\n<p>I am also one of these 95% you have mentioned.<br>\nI assume most of the people who are not crossing 0.65 barrier are the ones who make the simple modifications to the scales.<br>\nAs I mentioned in one of my notebooks (focusing on this scale modification), there will be big changes to the leaderboard for those who are focusing on the scales (Like myself).<br>\nThat is also why I am working on vision techniques (but it is tough to break that 0.65 as well xD)<br>\nI will share my solution with you once I find one!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2475504,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-10T01:03:45.017000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2466862": "Hi to all!\nI don't understand what I'm missing. I use this code as a base from https://www.kaggle.com/code/awsaf49/rsna-atd-cnn-tpu-infer (with major modifications). Now I'm testing spleen classification, but it seems that tf.keras.losses.CategoricalCrossentropy(label_smoothing=0.05) instead of classifying objects just dances around their frequency:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2626211%2F9adfa6782e56688e5b0b4a45dd620794%2F999.PNG?generation=1696398858857774&alt=media)\n\n\nHere, the frequencies differ slightly from the real ones due to the weighting of the classes. I tried Huge loss but it only made my result worse. At the same time, the AUC at this fold is stable at 0.92-0.93. This model performs slightly better than the better public model 0.65, but essentially \nmodel trained using tf.keras.losses.CategoricalCrossentropy does not actually classify the spleen, it only returns the average frequencies of the classes.\nWhat could be the reason? The model cannot distinguish features that separate ct, or did I miss something in the implementation? Considering that 95% of participants could not cross the 0.65 barrier, this problem is not only mine. Important I transfer the whole ct to the model as one object. For the spleen I form a slice. Thank you all for your comments!",
    "2468294": "I think that the problem that we are facing might not be related to the choice of model or loss function, rather it is with the data that we are feeding this model. There is a very high probability that the model is not able to learn anything meaningful from the data, and it just learns the frequency of the predictions.",
    "2467290": "I think the worst part is most of the submissions were just weighted scaled mean values of particular target category.",
    "2472473": "I don't know if you noticed but some people instantly broke that barrier you just mentioned (including me). The reason is they have experience of dealing with similar data. If it was early stages, I would share my approach but it's too late to share anything now. ",
    "2472305": "Hi @aikhmelnytskyy, I think it is better if you don't mind about \n> Considering that 95% of participants could not cross the 0.65 barrier, this problem is not only mine.\n\nI am also one of these 95% you have mentioned.\nI assume most of the people who are not crossing 0.65 barrier are the ones who make the simple modifications to the scales.\nAs I mentioned in one of my notebooks (focusing on this scale modification), there will be big changes to the leaderboard for those who are focusing on the scales (Like myself).\nThat is also why I am working on vision techniques (but it is tough to break that 0.65 as well xD)\nI will share my solution with you once I find one!",
    "2475504": ""
  }
}