{
  "id": 111198,
  "title": "Correlation between LB and loss",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/111198",
  "author_name": "Alimbekov Renat [dsmlkz]",
  "post_date": "2019-10-04T04:52:04.642000",
  "votes": 6,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi.  Using BCE loss I have different between LB and loss equal 0,015 in few experiments. Do you also?</p>",
  "messages": [
    {
      "id": 640537,
      "postDate": "2019-10-04T04:52:04.643Z",
      "content": "<p>Hi.  Using BCE loss I have different between LB and loss equal 0,015 in few experiments. Do you also?</p>",
      "rawMarkdown": "Hi.  Using BCE loss I have different between LB and loss equal 0,015 in few experiments. Do you also?",
      "votes": 6
    },
    {
      "id": 641613,
      "postDate": "2019-10-04T20:58:13.297Z",
      "content": "<p>Do you use the correct weights? \nIt might be a statistical effect, but I usually see 0-0.005 difference (I use 10% of the data for validation) and your difference is x3. </p>",
      "rawMarkdown": "Do you use the correct weights? \nIt might be a statistical effect, but I usually see 0-0.005 difference (I use 10% of the data for validation) and your difference is x3. ",
      "votes": 2,
      "replies": [
        {
          "id": 641844,
          "postDate": "2019-10-05T08:18:43.547Z",
          "content": "<p>What model architecture did u use*? I use se-resnext50.  Could it be the network architecture? But there is definitely a dependency</p>",
          "rawMarkdown": "What model architecture did u use*? I use se-resnext50.  Could it be the network architecture? But there is definitely a dependency"
        },
        {
          "id": 641870,
          "postDate": "2019-10-05T08:52:18.843Z",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> sorry for my stupid question.\nAre you guys talking about LB score and validation loss?\nI though everyone's validation set is different so cant guarantee their better validation score = better LB score.  However in competitions many top performers can kind of get the relation between local validation score with LB score. May I know how to do that?</p>\n\n<p>Also what do you mean by using the correct weight?</p>\n\n<p>Sorry for the above questions, I am new in Kaggle.\nThanks so much</p>",
          "rawMarkdown": "@yuval6967 sorry for my stupid question.\nAre you guys talking about LB score and validation loss?\nI though everyone's validation set is different so cant guarantee their better validation score = better LB score.  However in competitions many top performers can kind of get the relation between local validation score with LB score. May I know how to do that?\n\nAlso what do you mean by using the correct weight?\n\nSorry for the above questions, I am new in Kaggle.\nThanks so much",
          "votes": 1
        },
        {
          "id": 641978,
          "postDate": "2019-10-05T11:56:26.310Z",
          "content": "<p><a href=\"/alimbekovkz\">@alimbekovkz</a> I used several Architectures, mostly Densenet, I don't think the dependency on the architecture should be that big. It could depend on the score itself, as it gets better the difference get smaller (up to nearly 0 near 0.068)</p>\n\n<p><a href=\"/fiyeroleung\">@fiyeroleung</a> Don't be sorry, your question is not stupid.\n1. Yes we are. If you use a weighted BCE you should get similar CV (validation score/loss) and LB score. You need to use the correct weights for the different classes. it was already established that all classes except 'any' get a weight of 1, and 'any' gets 2.\n2. The difference between CV and LB, can come from few sources, the main ones are:\na. A fundamental difference between the test and the training data\nb. A statistical difference - The test data is statistically chosen and the same goes for the validation set. </p>\n\n<p>In this competition b. is the case (for the first stage at least), and if I randomly choose different validation sets, I will get different CVs, I estimate this range to be +-0.003 depending on the score (look above). When I test different architectures and parameters, I make sure I use the same seeds (for random) in order to get the same validation set, and after I already submitted once I can estimate the difference between the CV and LB.   </p>",
          "rawMarkdown": "@alimbekovkz I used several Architectures, mostly Densenet, I don't think the dependency on the architecture should be that big. It could depend on the score itself, as it gets better the difference get smaller (up to nearly 0 near 0.068)\n\n@fiyeroleung Don't be sorry, your question is not stupid.\n1. Yes we are. If you use a weighted BCE you should get similar CV (validation score/loss) and LB score. You need to use the correct weights for the different classes. it was already established that all classes except 'any' get a weight of 1, and 'any' gets 2.\n2. The difference between CV and LB, can come from few sources, the main ones are:\na. A fundamental difference between the test and the training data\nb. A statistical difference - The test data is statistically chosen and the same goes for the validation set. \n\nIn this competition b. is the case (for the first stage at least), and if I randomly choose different validation sets, I will get different CVs, I estimate this range to be +-0.003 depending on the score (look above). When I test different architectures and parameters, I make sure I use the same seeds (for random) in order to get the same validation set, and after I already submitted once I can estimate the difference between the CV and LB.   ",
          "votes": 8
        },
        {
          "id": 642142,
          "postDate": "2019-10-05T15:47:57.297Z",
          "content": "<p>Thank you for answer.  I didn't implement weighted BCE Loss yet. Just use BCE. Therefore, I still can not try use the correct weights for the different classes. If you tell me the finished implementation, I will be grateful</p>",
          "rawMarkdown": "Thank you for answer.  I didn't implement weighted BCE Loss yet. Just use BCE. Therefore, I still can not try use the correct weights for the different classes. If you tell me the finished implementation, I will be grateful"
        },
        {
          "id": 642157,
          "postDate": "2019-10-05T16:08:54.600Z",
          "content": "<p>pytotch BCE has option to pass weights, you don’t have to implement anything :) </p>",
          "rawMarkdown": "pytotch BCE has option to pass weights, you don’t have to implement anything :) ",
          "votes": 2
        },
        {
          "id": 642196,
          "postDate": "2019-10-05T17:18:13.660Z",
          "content": "<p>here is an example</p>\n\n<p>```\ndef my_loss(y_pred,y_true,weights):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))</p>\n\n<p>```</p>",
          "rawMarkdown": "here is an example\n\n\n\n```\ndef my_loss(y_pred,y_true,weights):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))\n\n```",
          "votes": 6
        },
        {
          "id": 642202,
          "postDate": "2019-10-05T17:24:56.433Z",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> Thanks so much for your kind reply , learn a lots from you 🙌  </p>",
          "rawMarkdown": "@yuval6967 Thanks so much for your kind reply , learn a lots from you 🙌  ",
          "votes": 1
        },
        {
          "id": 642452,
          "postDate": "2019-10-06T05:10:41.273Z",
          "content": "<p>Wow, thanks. But how we calc weight?</p>",
          "rawMarkdown": "Wow, thanks. But how we calc weight?"
        },
        {
          "id": 642464,
          "postDate": "2019-10-06T05:28:32.330Z",
          "content": "<p>weights = torch.tensor([1.0,1.0,1.0,1.0,1.0,2.0])</p>\n\n<p>If  'any' is the 6th column in y_true </p>",
          "rawMarkdown": "weights = torch.tensor([1.0,1.0,1.0,1.0,1.0,2.0])\n\nIf  'any' is the 6th column in y_true ",
          "votes": 1
        },
        {
          "id": 642490,
          "postDate": "2019-10-06T06:50:36.850Z",
          "content": "<p>Thank you, but I think coefficients must be dependent on class balances or I'm wrong?</p>",
          "rawMarkdown": "Thank you, but I think coefficients must be dependent on class balances or I'm wrong?"
        },
        {
          "id": 642620,
          "postDate": "2019-10-06T11:58:27.470Z",
          "content": "<p>Usually it is best to use a loss function that mimics the competition's metric, that way, when you minimize the loss you also get the best score you can.</p>",
          "rawMarkdown": "Usually it is best to use a loss function that mimics the competition's metric, that way, when you minimize the loss you also get the best score you can.",
          "votes": 2
        },
        {
          "id": 643120,
          "postDate": "2019-10-07T06:19:18.473Z",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> Just to make sure, because I also use weighted BCE as mentioned above and efficientnet-b0 architecture. I observe as much as .007 difference between local val and LB using .15 val data which is always higher). I suppose this loss implementation is correct?</p>\n\n<p>weights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))</p>",
          "rawMarkdown": "@yuval6967 Just to make sure, because I also use weighted BCE as mentioned above and efficientnet-b0 architecture. I observe as much as .007 difference between local val and LB using .15 val data which is always higher). I suppose this loss implementation is correct?\n\n\nweights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))",
          "votes": 1
        },
        {
          "id": 643190,
          "postDate": "2019-10-07T08:17:09.170Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> You can look at the explanation above about the sources for the difference.\nAnother issue to keep in mind is that there are series of images and close images from the same series might be similar. Hance if you just choose images randomly, you might have a small leak between train and validation.</p>",
          "rawMarkdown": "@roguekk007 You can look at the explanation above about the sources for the difference.\nAnother issue to keep in mind is that there are series of images and close images from the same series might be similar. Hance if you just choose images randomly, you might have a small leak between train and validation.",
          "votes": 1
        },
        {
          "id": 643237,
          "postDate": "2019-10-07T09:46:36.610Z",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> OK understood! Validation strategies are important for this competition, I guess. An adversarial validation kernel can be very welcome</p>",
          "rawMarkdown": "@yuval6967 OK understood! Validation strategies are important for this competition, I guess. An adversarial validation kernel can be very welcome"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 641613,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2019-10-04T20:58:13.297000",
      "content": "<p>Do you use the correct weights? \nIt might be a statistical effect, but I usually see 0-0.005 difference (I use 10% of the data for validation) and your difference is x3. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 641844,
          "author_name": "Alimbekov Renat [dsmlkz]",
          "author_url": "",
          "post_date": "2019-10-05T08:18:43.547000",
          "content": "<p>What model architecture did u use*? I use se-resnext50.  Could it be the network architecture? But there is definitely a dependency</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 641870,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2019-10-05T08:52:18.843000",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> sorry for my stupid question.\nAre you guys talking about LB score and validation loss?\nI though everyone's validation set is different so cant guarantee their better validation score = better LB score.  However in competitions many top performers can kind of get the relation between local validation score with LB score. May I know how to do that?</p>\n\n<p>Also what do you mean by using the correct weight?</p>\n\n<p>Sorry for the above questions, I am new in Kaggle.\nThanks so much</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 641978,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2019-10-05T11:56:26.310000",
          "content": "<p><a href=\"/alimbekovkz\">@alimbekovkz</a> I used several Architectures, mostly Densenet, I don't think the dependency on the architecture should be that big. It could depend on the score itself, as it gets better the difference get smaller (up to nearly 0 near 0.068)</p>\n\n<p><a href=\"/fiyeroleung\">@fiyeroleung</a> Don't be sorry, your question is not stupid.\n1. Yes we are. If you use a weighted BCE you should get similar CV (validation score/loss) and LB score. You need to use the correct weights for the different classes. it was already established that all classes except 'any' get a weight of 1, and 'any' gets 2.\n2. The difference between CV and LB, can come from few sources, the main ones are:\na. A fundamental difference between the test and the training data\nb. A statistical difference - The test data is statistically chosen and the same goes for the validation set. </p>\n\n<p>In this competition b. is the case (for the first stage at least), and if I randomly choose different validation sets, I will get different CVs, I estimate this range to be +-0.003 depending on the score (look above). When I test different architectures and parameters, I make sure I use the same seeds (for random) in order to get the same validation set, and after I already submitted once I can estimate the difference between the CV and LB.   </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 642142,
          "author_name": "Alimbekov Renat [dsmlkz]",
          "author_url": "",
          "post_date": "2019-10-05T15:47:57.297000",
          "content": "<p>Thank you for answer.  I didn't implement weighted BCE Loss yet. Just use BCE. Therefore, I still can not try use the correct weights for the different classes. If you tell me the finished implementation, I will be grateful</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 642157,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-10-05T16:08:54.600000",
          "content": "<p>pytotch BCE has option to pass weights, you don’t have to implement anything :) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 642196,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2019-10-05T17:18:13.660000",
          "content": "<p>here is an example</p>\n\n<p>```\ndef my_loss(y_pred,y_true,weights):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))</p>\n\n<p>```</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 642202,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2019-10-05T17:24:56.433000",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> Thanks so much for your kind reply , learn a lots from you 🙌  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 642452,
          "author_name": "Alimbekov Renat [dsmlkz]",
          "author_url": "",
          "post_date": "2019-10-06T05:10:41.273000",
          "content": "<p>Wow, thanks. But how we calc weight?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 642464,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2019-10-06T05:28:32.330000",
          "content": "<p>weights = torch.tensor([1.0,1.0,1.0,1.0,1.0,2.0])</p>\n\n<p>If  'any' is the 6th column in y_true </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 642490,
          "author_name": "Alimbekov Renat [dsmlkz]",
          "author_url": "",
          "post_date": "2019-10-06T06:50:36.850000",
          "content": "<p>Thank you, but I think coefficients must be dependent on class balances or I'm wrong?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 642620,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2019-10-06T11:58:27.470000",
          "content": "<p>Usually it is best to use a loss function that mimics the competition's metric, that way, when you minimize the loss you also get the best score you can.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 643120,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-07T06:19:18.473000",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> Just to make sure, because I also use weighted BCE as mentioned above and efficientnet-b0 architecture. I observe as much as .007 difference between local val and LB using .15 val data which is always higher). I suppose this loss implementation is correct?</p>\n\n<p>weights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 643190,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2019-10-07T08:17:09.170000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> You can look at the explanation above about the sources for the difference.\nAnother issue to keep in mind is that there are series of images and close images from the same series might be similar. Hance if you just choose images randomly, you might have a small leak between train and validation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 643237,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-07T09:46:36.610000",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> OK understood! Validation strategies are important for this competition, I guess. An adversarial validation kernel can be very welcome</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "640537": "Hi.  Using BCE loss I have different between LB and loss equal 0,015 in few experiments. Do you also?",
    "641613": "Do you use the correct weights? \nIt might be a statistical effect, but I usually see 0-0.005 difference (I use 10% of the data for validation) and your difference is x3. "
  }
}