{
  "id": 375018,
  "title": "Limited in LB 0.13, No improvement anymore.",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/375018",
  "author_name": "ChenxiangSun@NJU",
  "post_date": "2022-12-29T23:47:18.682000",
  "votes": 7,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I used efficientnetb4, weighted loss, balanced sampler, and augmentation. But I only could get around 0.1 on LB best. Do you have some suggestions? Do I miss something?</p>\n<p>You could check here. Thank you for all your suggestions.😭</p>\n<p><a href=\"https://www.kaggle.com/code/fanyang99/train\" target=\"_blank\">https://www.kaggle.com/code/fanyang99/train</a></p>",
  "messages": [
    {
      "id": 2080183,
      "postDate": "2022-12-29T23:47:18.683Z",
      "content": "<p>Hi all,</p>\n<p>I used efficientnetb4, weighted loss, balanced sampler, and augmentation. But I only could get around 0.1 on LB best. Do you have some suggestions? Do I miss something?</p>\n<p>You could check here. Thank you for all your suggestions.😭</p>\n<p><a href=\"https://www.kaggle.com/code/fanyang99/train\" target=\"_blank\">https://www.kaggle.com/code/fanyang99/train</a></p>",
      "rawMarkdown": "Hi all,\n\nI used efficientnetb4, weighted loss, balanced sampler, and augmentation. But I only could get around 0.1 on LB best. Do you have some suggestions? Do I miss something?\n\nYou could check here. Thank you for all your suggestions.😭\n\nhttps://www.kaggle.com/code/fanyang99/train\n\n",
      "votes": 6
    },
    {
      "id": 2081074,
      "postDate": "2022-12-30T19:36:00.560Z",
      "content": "<p>Hi Nick, I’m a beginner in Image ML so sorry for any stupid comment but some ideas I can see in your nice notebook :</p>\n<ul>\n<li>Try StratifiedGroupKFold ( by patient_id)</li>\n<li>I don’t understand the loss in your train, maybe I’m wrong but it seems you lost the pos_weight when you write : <br>\nloss = nn.BCEWithLogitsLoss(pos_weight=torch.tensor([7])).cuda()<br>\nloss = nn.BCEWithLogitsLoss().cuda()</li>\n<li>try blending models , it is boring but often useful for competition </li>\n<li>I’m not sure about your resize (512,1024) not (1024,512) ? </li>\n</ul>",
      "rawMarkdown": "Hi Nick, I’m a beginner in Image ML so sorry for any stupid comment but some ideas I can see in your nice notebook :\n- Try StratifiedGroupKFold ( by patient_id)\n- I don’t understand the loss in your train, maybe I’m wrong but it seems you lost the pos_weight when you write : \nloss = nn.BCEWithLogitsLoss(pos_weight=torch.tensor([7])).cuda()\nloss = nn.BCEWithLogitsLoss().cuda()\n- try blending models , it is boring but often useful for competition \n- I’m not sure about your resize (512,1024) not (1024,512) ? ",
      "votes": 1,
      "replies": [
        {
          "id": 2081082,
          "postDate": "2022-12-30T19:48:20.627Z",
          "content": "<p>Thank you very much, sir.</p>\n<ol>\n<li>Yes, I just split it by the label of images. What number of splits is better for this task… I just use 4 or 5 I remember.</li>\n<li>I also found that I am stupid about my loss function. I will correct it. Do you think If I use some other loss function like, BCE+Dice loss would be better? like the combination of two loss functions for classification.</li>\n<li>This suggestion is just like, hmmm, I have two models, and I use them to do prediction, and I could just use the mean of these two predictions as my final result to improve. Am I right?</li>\n<li>Yes, I will try this. And In the past, I also tried no ROI images 512*512, I get 0.17 on the LB. Do you think is there any improvement space for this dataset or module? <br>\nA very interesting thing is that I found someone who is very good at ML on the Kaggle, They just said that they use very basic methods and could get an unbelievable baseline… Like LB 0.4. When I follow their suggestions, I found I could get only 0.17😂 I think I must miss some things… lol</li>\n</ol>",
          "rawMarkdown": "Thank you very much, sir.\n1. Yes, I just split it by the label of images. What number of splits is better for this task... I just use 4 or 5 I remember.\n2. I also found that I am stupid about my loss function. I will correct it. Do you think If I use some other loss function like, BCE+Dice loss would be better? like the combination of two loss functions for classification.\n3. This suggestion is just like, hmmm, I have two models, and I use them to do prediction, and I could just use the mean of these two predictions as my final result to improve. Am I right?\n4. Yes, I will try this. And In the past, I also tried no ROI images 512*512, I get 0.17 on the LB. Do you think is there any improvement space for this dataset or module? \nA very interesting thing is that I found someone who is very good at ML on the Kaggle, They just said that they use very basic methods and could get an unbelievable baseline... Like LB 0.4. When I follow their suggestions, I found I could get only 0.17😂 I think I must miss some things... lol\n",
          "replies": [
            {
              "id": 2081130,
              "postDate": "2022-12-30T20:34:40.623Z",
              "content": "<p>Hi, again I'm only beginner in image ML but I would suggest :<br>\n1- yes 5 or 4 should be correct<br>\n2- BCE should be sufficient (I presume)<br>\n3- Yes you make a loop to get predictions for each saved model for each fold and you stack the predictions for mean or max later for the submission<br>\n4- My guess is that 1024x512 (y,x) it a correct size, the problem is the performance and the machine you have… to go further you need a datacenter, and some lucky competitors have one 😁<br>\nI'm sure the secret is in boring blending…and the good threshold. </p>",
              "rawMarkdown": "Hi, again I'm only beginner in image ML but I would suggest :\n1- yes 5 or 4 should be correct\n2- BCE should be sufficient (I presume)\n3- Yes you make a loop to get predictions for each saved model for each fold and you stack the predictions for mean or max later for the submission\n4- My guess is that 1024x512 (y,x) it a correct size, the problem is the performance and the machine you have... to go further you need a datacenter, and some lucky competitors have one 😁\nI'm sure the secret is in boring blending...and the good threshold. ",
              "votes": 1
            },
            {
              "id": 2081136,
              "postDate": "2022-12-30T20:41:35.297Z",
              "content": "<p>thank you so much!🤜</p>",
              "rawMarkdown": "thank you so much!🤜"
            }
          ]
        }
      ]
    },
    {
      "id": 2094399,
      "postDate": "2023-01-10T18:25:42.130Z",
      "content": "<p>Try playing with the hyperparameters, there are usually some low hanging fruits there.. <br>\nGood luck!</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Try playing with the hyperparameters, there are usually some low hanging fruits there.. \nGood luck!\n\nThe Devastator.\n",
      "replies": [
        {
          "id": 2094456,
          "postDate": "2023-01-10T19:04:06.253Z",
          "content": "<p>Thank you. I will try it.</p>\n<p>What I recently found is that the local best threshold is not equal to the LB best threshold. The mode, weighted BCE, and sampler could decide the threshold.</p>\n<p>And some models cloud not work very well on this task. Not 100% sure, maybe my hyperparameters have something wrong or unsuitable.</p>\n<p>I will continue to check.🤜</p>",
          "rawMarkdown": "Thank you. I will try it.\n\nWhat I recently found is that the local best threshold is not equal to the LB best threshold. The mode, weighted BCE, and sampler could decide the threshold.\n\nAnd some models cloud not work very well on this task. Not 100% sure, maybe my hyperparameters have something wrong or unsuitable.\n\nI will continue to check.🤜"
        }
      ]
    },
    {
      "id": 2081483,
      "postDate": "2022-12-31T08:17:30.750Z",
      "content": "<p>Hi, i'm also stuck at 0.18. Have you solve this problem? Could you please give me some advice?</p>",
      "rawMarkdown": "Hi, i'm also stuck at 0.18. Have you solve this problem? Could you please give me some advice?",
      "replies": [
        {
          "id": 2081490,
          "postDate": "2022-12-31T08:29:36.927Z",
          "content": "<p>I plan to work on the no ROI first, but now I still get 0.17. I guess it is because I do not try using the ensemble method.  I found some other peoples SOTA results using this method on LB. You can check this on other people's inference notebooks, they load some models instead of one model. Not sure. I have just submitted an ensemble method, I hope I can get a good result tomorrow morning. That all I know, if you have some ideas you can also share them.💪</p>",
          "rawMarkdown": "I plan to work on the no ROI first, but now I still get 0.17. I guess it is because I do not try using the ensemble method.  I found some other peoples SOTA results using this method on LB. You can check this on other people's inference notebooks, they load some models instead of one model. Not sure. I have just submitted an ensemble method, I hope I can get a good result tomorrow morning. That all I know, if you have some ideas you can also share them.💪"
        },
        {
          "id": 2081773,
          "postDate": "2022-12-31T15:51:57.160Z",
          "content": "<p>Ok, now my best score is 0.21. I found using the ensemble method could get improvement. However, I do not know how to get the best threshold to submit. I just use 0.5 without trying other thresholds. I'm 100% sure that It could have some improvement. (at least I have🙉)</p>",
          "rawMarkdown": "Ok, now my best score is 0.21. I found using the ensemble method could get improvement. However, I do not know how to get the best threshold to submit. I just use 0.5 without trying other thresholds. I'm 100% sure that It could have some improvement. (at least I have🙉)",
          "replies": [
            {
              "id": 2082016,
              "postDate": "2023-01-01T01:44:10.020Z",
              "content": "<p>Now is 0.24. Using the ensemble method.</p>",
              "rawMarkdown": "Now is 0.24. Using the ensemble method."
            },
            {
              "id": 2082112,
              "postDate": "2023-01-01T05:44:09.097Z",
              "content": "<p>thanks for your sharing, but i saw your lb is 0.42. I will try your advice to see if it can improve.</p>",
              "rawMarkdown": "thanks for your sharing, but i saw your lb is 0.42. I will try your advice to see if it can improve."
            },
            {
              "id": 2082115,
              "postDate": "2023-01-01T05:50:11.020Z",
              "content": "<p>Oh, 0.42 is copied from another kaggler's notebook to submit. So it's not mine. I will plan to share the code recently. lol, very basic/easy notebook.</p>",
              "rawMarkdown": "Oh, 0.42 is copied from another kaggler's notebook to submit. So it's not mine. I will plan to share the code recently. lol, very basic/easy notebook."
            },
            {
              "id": 2082117,
              "postDate": "2023-01-01T05:55:59.510Z",
              "content": "<p>ok thanks. could you please share the nb, I will check if i have something wrong in my infer nb.</p>",
              "rawMarkdown": "ok thanks. could you please share the nb, I will check if i have something wrong in my infer nb."
            },
            {
              "id": 2082120,
              "postDate": "2023-01-01T06:03:21.483Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2082127,
              "postDate": "2023-01-01T06:18:16.133Z",
              "content": "<p>Training notebook and inference here. <br>\n<a href=\"https://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24\" target=\"_blank\">https://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24</a><br>\n<a href=\"https://www.kaggle.com/code/fanyang99/inference-no-roi-512-lb0-24/\" target=\"_blank\">https://www.kaggle.com/code/fanyang99/inference-no-roi-512-lb0-24/</a></p>",
              "rawMarkdown": "Training notebook and inference here. \nhttps://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24\nhttps://www.kaggle.com/code/fanyang99/inference-no-roi-512-lb0-24/"
            }
          ]
        }
      ]
    },
    {
      "id": 2081048,
      "postDate": "2022-12-30T19:02:33.053Z",
      "content": "<p>Have you tried experimenting with lower learning rates?</p>",
      "rawMarkdown": "Have you tried experimenting with lower learning rates?",
      "replies": [
        {
          "id": 2081057,
          "postDate": "2022-12-30T19:15:00.760Z",
          "content": "<p>Not yet, I will try this. Thank you very much.</p>",
          "rawMarkdown": "Not yet, I will try this. Thank you very much."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2081074,
      "author_name": "Laurent Pourchot",
      "author_url": "",
      "post_date": "2022-12-30T19:36:00.560000",
      "content": "<p>Hi Nick, I’m a beginner in Image ML so sorry for any stupid comment but some ideas I can see in your nice notebook :</p>\n<ul>\n<li>Try StratifiedGroupKFold ( by patient_id)</li>\n<li>I don’t understand the loss in your train, maybe I’m wrong but it seems you lost the pos_weight when you write : <br>\nloss = nn.BCEWithLogitsLoss(pos_weight=torch.tensor([7])).cuda()<br>\nloss = nn.BCEWithLogitsLoss().cuda()</li>\n<li>try blending models , it is boring but often useful for competition </li>\n<li>I’m not sure about your resize (512,1024) not (1024,512) ? </li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 2081082,
          "author_name": "ChenxiangSun@NJU",
          "author_url": "",
          "post_date": "2022-12-30T19:48:20.627000",
          "content": "<p>Thank you very much, sir.</p>\n<ol>\n<li>Yes, I just split it by the label of images. What number of splits is better for this task… I just use 4 or 5 I remember.</li>\n<li>I also found that I am stupid about my loss function. I will correct it. Do you think If I use some other loss function like, BCE+Dice loss would be better? like the combination of two loss functions for classification.</li>\n<li>This suggestion is just like, hmmm, I have two models, and I use them to do prediction, and I could just use the mean of these two predictions as my final result to improve. Am I right?</li>\n<li>Yes, I will try this. And In the past, I also tried no ROI images 512*512, I get 0.17 on the LB. Do you think is there any improvement space for this dataset or module? <br>\nA very interesting thing is that I found someone who is very good at ML on the Kaggle, They just said that they use very basic methods and could get an unbelievable baseline… Like LB 0.4. When I follow their suggestions, I found I could get only 0.17😂 I think I must miss some things… lol</li>\n</ol>",
          "votes": 0,
          "replies": [
            {
              "id": 2081130,
              "author_name": "Laurent Pourchot",
              "author_url": "",
              "post_date": "2022-12-30T20:34:40.623000",
              "content": "<p>Hi, again I'm only beginner in image ML but I would suggest :<br>\n1- yes 5 or 4 should be correct<br>\n2- BCE should be sufficient (I presume)<br>\n3- Yes you make a loop to get predictions for each saved model for each fold and you stack the predictions for mean or max later for the submission<br>\n4- My guess is that 1024x512 (y,x) it a correct size, the problem is the performance and the machine you have… to go further you need a datacenter, and some lucky competitors have one 😁<br>\nI'm sure the secret is in boring blending…and the good threshold. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2081136,
              "author_name": "ChenxiangSun@NJU",
              "author_url": "",
              "post_date": "2022-12-30T20:41:35.297000",
              "content": "<p>thank you so much!🤜</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2094399,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2023-01-10T18:25:42.130000",
      "content": "<p>Try playing with the hyperparameters, there are usually some low hanging fruits there.. <br>\nGood luck!</p>\n<p>The Devastator.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2094456,
          "author_name": "ChenxiangSun@NJU",
          "author_url": "",
          "post_date": "2023-01-10T19:04:06.253000",
          "content": "<p>Thank you. I will try it.</p>\n<p>What I recently found is that the local best threshold is not equal to the LB best threshold. The mode, weighted BCE, and sampler could decide the threshold.</p>\n<p>And some models cloud not work very well on this task. Not 100% sure, maybe my hyperparameters have something wrong or unsuitable.</p>\n<p>I will continue to check.🤜</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2081483,
      "author_name": "Yingshu Li",
      "author_url": "",
      "post_date": "2022-12-31T08:17:30.750000",
      "content": "<p>Hi, i'm also stuck at 0.18. Have you solve this problem? Could you please give me some advice?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2081490,
          "author_name": "ChenxiangSun@NJU",
          "author_url": "",
          "post_date": "2022-12-31T08:29:36.927000",
          "content": "<p>I plan to work on the no ROI first, but now I still get 0.17. I guess it is because I do not try using the ensemble method.  I found some other peoples SOTA results using this method on LB. You can check this on other people's inference notebooks, they load some models instead of one model. Not sure. I have just submitted an ensemble method, I hope I can get a good result tomorrow morning. That all I know, if you have some ideas you can also share them.💪</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2081773,
          "author_name": "ChenxiangSun@NJU",
          "author_url": "",
          "post_date": "2022-12-31T15:51:57.160000",
          "content": "<p>Ok, now my best score is 0.21. I found using the ensemble method could get improvement. However, I do not know how to get the best threshold to submit. I just use 0.5 without trying other thresholds. I'm 100% sure that It could have some improvement. (at least I have🙉)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2082016,
              "author_name": "ChenxiangSun@NJU",
              "author_url": "",
              "post_date": "2023-01-01T01:44:10.020000",
              "content": "<p>Now is 0.24. Using the ensemble method.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082112,
              "author_name": "Yingshu Li",
              "author_url": "",
              "post_date": "2023-01-01T05:44:09.097000",
              "content": "<p>thanks for your sharing, but i saw your lb is 0.42. I will try your advice to see if it can improve.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082115,
              "author_name": "ChenxiangSun@NJU",
              "author_url": "",
              "post_date": "2023-01-01T05:50:11.020000",
              "content": "<p>Oh, 0.42 is copied from another kaggler's notebook to submit. So it's not mine. I will plan to share the code recently. lol, very basic/easy notebook.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082117,
              "author_name": "Yingshu Li",
              "author_url": "",
              "post_date": "2023-01-01T05:55:59.510000",
              "content": "<p>ok thanks. could you please share the nb, I will check if i have something wrong in my infer nb.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082120,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-01T06:03:21.483000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082127,
              "author_name": "ChenxiangSun@NJU",
              "author_url": "",
              "post_date": "2023-01-01T06:18:16.133000",
              "content": "<p>Training notebook and inference here. <br>\n<a href=\"https://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24\" target=\"_blank\">https://www.kaggle.com/code/fanyang99/train-no-roi-512-lb0-24</a><br>\n<a href=\"https://www.kaggle.com/code/fanyang99/inference-no-roi-512-lb0-24/\" target=\"_blank\">https://www.kaggle.com/code/fanyang99/inference-no-roi-512-lb0-24/</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2081048,
      "author_name": "Nick Navid Yazdani",
      "author_url": "",
      "post_date": "2022-12-30T19:02:33.053000",
      "content": "<p>Have you tried experimenting with lower learning rates?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2081057,
          "author_name": "ChenxiangSun@NJU",
          "author_url": "",
          "post_date": "2022-12-30T19:15:00.760000",
          "content": "<p>Not yet, I will try this. Thank you very much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2080183": "Hi all,\n\nI used efficientnetb4, weighted loss, balanced sampler, and augmentation. But I only could get around 0.1 on LB best. Do you have some suggestions? Do I miss something?\n\nYou could check here. Thank you for all your suggestions.😭\n\nhttps://www.kaggle.com/code/fanyang99/train\n\n",
    "2081074": "Hi Nick, I’m a beginner in Image ML so sorry for any stupid comment but some ideas I can see in your nice notebook :\n- Try StratifiedGroupKFold ( by patient_id)\n- I don’t understand the loss in your train, maybe I’m wrong but it seems you lost the pos_weight when you write : \nloss = nn.BCEWithLogitsLoss(pos_weight=torch.tensor([7])).cuda()\nloss = nn.BCEWithLogitsLoss().cuda()\n- try blending models , it is boring but often useful for competition \n- I’m not sure about your resize (512,1024) not (1024,512) ? ",
    "2094399": "Try playing with the hyperparameters, there are usually some low hanging fruits there.. \nGood luck!\n\nThe Devastator.\n",
    "2081483": "Hi, i'm also stuck at 0.18. Have you solve this problem? Could you please give me some advice?",
    "2081048": "Have you tried experimenting with lower learning rates?"
  }
}