{
  "id": 382056,
  "title": "pF1 score not improving at all",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/382056",
  "author_name": "Aryamaan Thakur",
  "post_date": "2023-01-29T12:02:36.242000",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am using EfficientNetB3 with AdamW optimizer and OneCycleLR scheduler. Training on ROI images of size 1024x512. Using BCEWithLogitsLoss with pos_weight=20 (also tried focal loss).<br>\nBut I am still stuck on pF1 score of 0.03-0.04. Is this normal or am I missing something? I would be highly grateful if someone can help me out.<br>\nThanks in advance!</p>",
  "messages": [
    {
      "id": 2120171,
      "postDate": "2023-01-29T12:02:36.243Z",
      "content": "<p>I am using EfficientNetB3 with AdamW optimizer and OneCycleLR scheduler. Training on ROI images of size 1024x512. Using BCEWithLogitsLoss with pos_weight=20 (also tried focal loss).<br>\nBut I am still stuck on pF1 score of 0.03-0.04. Is this normal or am I missing something? I would be highly grateful if someone can help me out.<br>\nThanks in advance!</p>",
      "rawMarkdown": "I am using EfficientNetB3 with AdamW optimizer and OneCycleLR scheduler. Training on ROI images of size 1024x512. Using BCEWithLogitsLoss with pos_weight=20 (also tried focal loss).\nBut I am still stuck on pF1 score of 0.03-0.04. Is this normal or am I missing something? I would be highly grateful if someone can help me out.\nThanks in advance!",
      "votes": 5
    },
    {
      "id": 2120611,
      "postDate": "2023-01-29T17:57:32.483Z",
      "content": "<p>assume there is no bug in your code or pippline, your problem is \"not able to learn\"</p>\n<p>pF1 score of 0.03-0.04 means the model is predicting all zeros (the majority negative class), i.e. it is not learning anything.<br>\nin machine learning terms, it cannot perform gradient or perform gradient goes to nowhere.</p>\n<p>a few possibility:<br>\n1) poor initialisation : did you forget to load your pretrain model?<br>\n2) the input to the pretrain model is not the same: did you forget to to do std and mean normalisation? are the std, mean values correct?<br>\n3) wrong sampling : you did not feed all train data, you only feed one class, etc <br>\n4) wrong loss, wrong model configuration: e.g. you forget to do relu before last global pooling, you feed prob in bce loss function which actually accept logit<br>\n5) wrong hyper paramters : your learning rate is wrong etc ….<br>\n6) wrong use of software framework : forget to set to train mode? your model  tensor graph is not connected, etc …</p>",
      "rawMarkdown": "assume there is no bug in your code or pippline, your problem is \"not able to learn\"\n\npF1 score of 0.03-0.04 means the model is predicting all zeros (the majority negative class), i.e. it is not learning anything.\nin machine learning terms, it cannot perform gradient or perform gradient goes to nowhere.\n\na few possibility:\n1) poor initialisation : did you forget to load your pretrain model?\n2) the input to the pretrain model is not the same: did you forget to to do std and mean normalisation? are the std, mean values correct?\n3) wrong sampling : you did not feed all train data, you only feed one class, etc \n4) wrong loss, wrong model configuration: e.g. you forget to do relu before last global pooling, you feed prob in bce loss function which actually accept logit\n5) wrong hyper paramters : your learning rate is wrong etc ....\n6) wrong use of software framework : forget to set to train mode? your model  tensor graph is not connected, etc ...",
      "votes": 6
    },
    {
      "id": 2120359,
      "postDate": "2023-01-29T14:47:51.320Z",
      "content": "<p>Hello,<br>\nYou should resample the data's pos_neg_rate to a normal number or delete some negative samples, which is helpful to model and debug. You can easily check the accuracy and training loss. But that is a shortcoming that is easier to overfit. </p>",
      "rawMarkdown": "Hello,\nYou should resample the data's pos_neg_rate to a normal number or delete some negative samples, which is helpful to model and debug. You can easily check the accuracy and training loss. But that is a shortcoming that is easier to overfit. ",
      "votes": 2,
      "replies": [
        {
          "id": 2120512,
          "postDate": "2023-01-29T16:46:19.950Z",
          "content": "<p>Thanks, I will try WeightedRandomSampler to get more positive samples. Will update you once its done :)</p>",
          "rawMarkdown": "Thanks, I will try WeightedRandomSampler to get more positive samples. Will update you once its done :)",
          "votes": 1
        },
        {
          "id": 2121104,
          "postDate": "2023-01-30T04:18:38.280Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2121795,
          "postDate": "2023-01-30T14:56:04.837Z",
          "content": "<p>Using a sampler worked well. pF1 straightaway increased to 0.11 on validation set. And upon using a threshold, it gave 0.24. Thanks a lot!</p>",
          "rawMarkdown": "Using a sampler worked well. pF1 straightaway increased to 0.11 on validation set. And upon using a threshold, it gave 0.24. Thanks a lot!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2121012,
      "postDate": "2023-01-30T01:16:24.007Z",
      "content": "<p>I am having the same problem. I will try the tips provided by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> . Thank both of you! :)</p>\n<p>I recently published a notebook with the code I am using just in case someone wants to point out what is more likely to be wrong there. </p>",
      "rawMarkdown": "I am having the same problem. I will try the tips provided by @hengck23 . Thank both of you! :)\n\nI recently published a notebook with the code I am using just in case someone wants to point out what is more likely to be wrong there. "
    },
    {
      "id": 2120742,
      "postDate": "2023-01-29T19:35:24.683Z",
      "content": "<p>my guess, big image size -&gt; small batch, combined with the imbalanced dataset -&gt; the model doesn't learn. I have a similar problem with big resolution, but I am able to train … something, on a smaller resolution</p>",
      "rawMarkdown": "my guess, big image size -> small batch, combined with the imbalanced dataset -> the model doesn't learn. I have a similar problem with big resolution, but I am able to train ... something, on a smaller resolution\n"
    }
  ],
  "comments": [
    {
      "id": 2120611,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-29T17:57:32.483000",
      "content": "<p>assume there is no bug in your code or pippline, your problem is \"not able to learn\"</p>\n<p>pF1 score of 0.03-0.04 means the model is predicting all zeros (the majority negative class), i.e. it is not learning anything.<br>\nin machine learning terms, it cannot perform gradient or perform gradient goes to nowhere.</p>\n<p>a few possibility:<br>\n1) poor initialisation : did you forget to load your pretrain model?<br>\n2) the input to the pretrain model is not the same: did you forget to to do std and mean normalisation? are the std, mean values correct?<br>\n3) wrong sampling : you did not feed all train data, you only feed one class, etc <br>\n4) wrong loss, wrong model configuration: e.g. you forget to do relu before last global pooling, you feed prob in bce loss function which actually accept logit<br>\n5) wrong hyper paramters : your learning rate is wrong etc ….<br>\n6) wrong use of software framework : forget to set to train mode? your model  tensor graph is not connected, etc …</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2120359,
      "author_name": "yoyobar",
      "author_url": "",
      "post_date": "2023-01-29T14:47:51.320000",
      "content": "<p>Hello,<br>\nYou should resample the data's pos_neg_rate to a normal number or delete some negative samples, which is helpful to model and debug. You can easily check the accuracy and training loss. But that is a shortcoming that is easier to overfit. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2120512,
          "author_name": "Aryamaan Thakur",
          "author_url": "",
          "post_date": "2023-01-29T16:46:19.950000",
          "content": "<p>Thanks, I will try WeightedRandomSampler to get more positive samples. Will update you once its done :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2121104,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-30T04:18:38.280000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2121795,
          "author_name": "Aryamaan Thakur",
          "author_url": "",
          "post_date": "2023-01-30T14:56:04.837000",
          "content": "<p>Using a sampler worked well. pF1 straightaway increased to 0.11 on validation set. And upon using a threshold, it gave 0.24. Thanks a lot!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2121012,
      "author_name": "Enrique Alanís",
      "author_url": "",
      "post_date": "2023-01-30T01:16:24.007000",
      "content": "<p>I am having the same problem. I will try the tips provided by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> . Thank both of you! :)</p>\n<p>I recently published a notebook with the code I am using just in case someone wants to point out what is more likely to be wrong there. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2120742,
      "author_name": "FlaviuPaul",
      "author_url": "",
      "post_date": "2023-01-29T19:35:24.683000",
      "content": "<p>my guess, big image size -&gt; small batch, combined with the imbalanced dataset -&gt; the model doesn't learn. I have a similar problem with big resolution, but I am able to train … something, on a smaller resolution</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2120171": "I am using EfficientNetB3 with AdamW optimizer and OneCycleLR scheduler. Training on ROI images of size 1024x512. Using BCEWithLogitsLoss with pos_weight=20 (also tried focal loss).\nBut I am still stuck on pF1 score of 0.03-0.04. Is this normal or am I missing something? I would be highly grateful if someone can help me out.\nThanks in advance!",
    "2120611": "assume there is no bug in your code or pippline, your problem is \"not able to learn\"\n\npF1 score of 0.03-0.04 means the model is predicting all zeros (the majority negative class), i.e. it is not learning anything.\nin machine learning terms, it cannot perform gradient or perform gradient goes to nowhere.\n\na few possibility:\n1) poor initialisation : did you forget to load your pretrain model?\n2) the input to the pretrain model is not the same: did you forget to to do std and mean normalisation? are the std, mean values correct?\n3) wrong sampling : you did not feed all train data, you only feed one class, etc \n4) wrong loss, wrong model configuration: e.g. you forget to do relu before last global pooling, you feed prob in bce loss function which actually accept logit\n5) wrong hyper paramters : your learning rate is wrong etc ....\n6) wrong use of software framework : forget to set to train mode? your model  tensor graph is not connected, etc ...",
    "2120359": "Hello,\nYou should resample the data's pos_neg_rate to a normal number or delete some negative samples, which is helpful to model and debug. You can easily check the accuracy and training loss. But that is a shortcoming that is easier to overfit. ",
    "2121012": "I am having the same problem. I will try the tips provided by @hengck23 . Thank both of you! :)\n\nI recently published a notebook with the code I am using just in case someone wants to point out what is more likely to be wrong there. ",
    "2120742": "my guess, big image size -> small batch, combined with the imbalanced dataset -> the model doesn't learn. I have a similar problem with big resolution, but I am able to train ... something, on a smaller resolution\n"
  }
}