{
  "id": 111729,
  "title": "Similarity between train and test images",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/111729",
  "author_name": "Nicholas Lyu",
  "post_date": "2019-10-08T12:14:38.470000",
  "votes": 18,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I have been frustrated over difference between local validation and LB scores (as much as .01+) when using random selection of validation data. Though many has reported little difference between local validation and LB, this is not my case. Discussions hinted that there may exist significant difference between training and test data. </p>\n\n<p>I am doing an adversarial validation test (train a model to try to classify whether an image belongs to training or test data), and preliminary results show that classifier can achieve as much as 77% validation accuracy for adversarial validation (I checked so training and testing data are balanced). </p>\n\n<p>Has anyone done similar analysis to quantize the difference between training / test set? How is you local validation comparing to LB? I think this is very important for establishing useful local validation scores to avoid overfitting.</p>",
  "messages": [
    {
      "id": 644146,
      "postDate": "2019-10-08T12:14:38.470Z",
      "content": "<p>I have been frustrated over difference between local validation and LB scores (as much as .01+) when using random selection of validation data. Though many has reported little difference between local validation and LB, this is not my case. Discussions hinted that there may exist significant difference between training and test data. </p>\n\n<p>I am doing an adversarial validation test (train a model to try to classify whether an image belongs to training or test data), and preliminary results show that classifier can achieve as much as 77% validation accuracy for adversarial validation (I checked so training and testing data are balanced). </p>\n\n<p>Has anyone done similar analysis to quantize the difference between training / test set? How is you local validation comparing to LB? I think this is very important for establishing useful local validation scores to avoid overfitting.</p>",
      "rawMarkdown": "I have been frustrated over difference between local validation and LB scores (as much as .01+) when using random selection of validation data. Though many has reported little difference between local validation and LB, this is not my case. Discussions hinted that there may exist significant difference between training and test data. \n\nI am doing an adversarial validation test (train a model to try to classify whether an image belongs to training or test data), and preliminary results show that classifier can achieve as much as 77% validation accuracy for adversarial validation (I checked so training and testing data are balanced). \n\nHas anyone done similar analysis to quantize the difference between training / test set? How is you local validation comparing to LB? I think this is very important for establishing useful local validation scores to avoid overfitting.",
      "votes": 18
    },
    {
      "id": 644685,
      "postDate": "2019-10-09T07:27:18.213Z",
      "content": "<p>I am using <code>GroupKFold</code> on Patient ID and am seeing good CV/LB correlation (LB is 0.03 higher)</p>",
      "rawMarkdown": "I am using `GroupKFold` on Patient ID and am seeing good CV/LB correlation (LB is 0.03 higher)",
      "votes": 4
    },
    {
      "id": 647268,
      "postDate": "2019-10-12T10:11:34.527Z",
      "content": "<p>When I was randomly splitting, I found that while CV/LB scores initially appeared closer, after X epochs they would diverge and LB would actually start increasing as CV decreased. I think this over fitting is likely due to same patient being in train/validation. </p>\n\n<p>Right now I train using a PatientID split, and find that CV loss stops decreasing after so many epoch, and this correlates to LB performance. However, scores are a little different between CV/LB, differing as much as 0.01. </p>",
      "rawMarkdown": "When I was randomly splitting, I found that while CV/LB scores initially appeared closer, after X epochs they would diverge and LB would actually start increasing as CV decreased. I think this over fitting is likely due to same patient being in train/validation. \n\nRight now I train using a PatientID split, and find that CV loss stops decreasing after so many epoch, and this correlates to LB performance. However, scores are a little different between CV/LB, differing as much as 0.01. ",
      "votes": 3,
      "replies": [
        {
          "id": 647370,
          "postDate": "2019-10-12T14:17:00.357Z",
          "content": "<p><a href=\"/taindow\">@taindow</a> This confirms my observation, thanks!</p>",
          "rawMarkdown": "@taindow This confirms my observation, thanks!"
        },
        {
          "id": 647373,
          "postDate": "2019-10-12T14:19:56.550Z",
          "rawMarkdown": ""
        },
        {
          "id": 647909,
          "postDate": "2019-10-13T13:22:11.043Z",
          "content": "<p>Excuse me,  have you try other splitting?  such as according to studyInstanceID....</p>",
          "rawMarkdown": "Excuse me,  have you try other splitting?  such as according to studyInstanceID...."
        }
      ]
    },
    {
      "id": 644299,
      "postDate": "2019-10-08T15:46:27.657Z",
      "content": "<p>Update: I am using efficientnet-b0 trained for 2 epochs, and predictions are made on training data that was totally unseen. I observe a bimodal distribution of predicted confidence that training data is actually training data. Anyone has idea why this happens? I am going to investigate its relationship with labels tomorrow.</p>\n\n<p>Probability distribution of predicted confidence on unseen training data\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F111494fb2bcd616b6e7279d01eb5f7f8%2Fdownload%20(1\" alt=\"\">.png?generation=1570549578416793&amp;alt=media)\n:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2Fa7db8624e70d3b94bd663ee2c2cc9b6f%2Fdownload.png?generation=1570549575213198&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Update: I am using efficientnet-b0 trained for 2 epochs, and predictions are made on training data that was totally unseen. I observe a bimodal distribution of predicted confidence that training data is actually training data. Anyone has idea why this happens? I am going to investigate its relationship with labels tomorrow.\n\nProbability distribution of predicted confidence on unseen training data\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F111494fb2bcd616b6e7279d01eb5f7f8%2Fdownload%20(1).png?generation=1570549578416793&amp;alt=media)\n:![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2Fa7db8624e70d3b94bd663ee2c2cc9b6f%2Fdownload.png?generation=1570549575213198&amp;alt=media)",
      "votes": 3
    },
    {
      "id": 644647,
      "postDate": "2019-10-09T06:13:14.380Z",
      "content": "<p>If you end up with images from the same study in both the train and validation sets, it's easy to over fit without realizing it. </p>\n\n<p>I'd guess this is also why your adversarial validation accuracy is so high. </p>\n\n<p>Randomly splitting for train/validation is probably not a good idea in this competition. I've been using the metadata and splitting by study and my validation/lb scores are typically under 0.005</p>\n\n<p>Check out how similar some of these examples are: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1202828%2Fd76ab2c2d3e5a85835115917f0cc1ee1%2Fsimilar.png?generation=1570601171279933&amp;alt=media\" alt=\"\">\n(Image borrowed from Bruno Aquino's excellent <a href=\"https://www.kaggle.com/braquino/correct-images-sequece\">Correct Images Sequece</a> Kernel)</p>",
      "rawMarkdown": "If you end up with images from the same study in both the train and validation sets, it's easy to over fit without realizing it. \n\nI'd guess this is also why your adversarial validation accuracy is so high. \n\nRandomly splitting for train/validation is probably not a good idea in this competition. I've been using the metadata and splitting by study and my validation/lb scores are typically under 0.005\n\nCheck out how similar some of these examples are: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1202828%2Fd76ab2c2d3e5a85835115917f0cc1ee1%2Fsimilar.png?generation=1570601171279933&amp;alt=media)\n(Image borrowed from Bruno Aquino's excellent [Correct Images Sequece](https://www.kaggle.com/braquino/correct-images-sequece) Kernel)",
      "votes": 3,
      "replies": [
        {
          "id": 646329,
          "postDate": "2019-10-11T06:13:07.960Z",
          "content": "<p>Great insight. If not random splitting then what is your validation scheme?</p>",
          "rawMarkdown": "Great insight. If not random splitting then what is your validation scheme?",
          "votes": 1
        },
        {
          "id": 646982,
          "postDate": "2019-10-11T23:46:37.753Z",
          "content": "<p>Thanks Neerja. I’ve been splitting in StudyId but I think datasaurus’s suggestion of using PatientId is probably even better.</p>",
          "rawMarkdown": "Thanks Neerja. I’ve been splitting in StudyId but I think datasaurus’s suggestion of using PatientId is probably even better."
        }
      ]
    },
    {
      "id": 644783,
      "postDate": "2019-10-09T10:36:17.740Z",
      "content": "<p>Ok, Great! Love the insight:). I will try out your methods</p>",
      "rawMarkdown": "Ok, Great! Love the insight:). I will try out your methods",
      "votes": 1
    },
    {
      "id": 647074,
      "postDate": "2019-10-12T03:18:33.977Z",
      "content": "<p>I grouped on patients and saw an immediate increase in validation loss. Right now the problem seems to be that my validation loss is much larger than my LB (.0814 / .073). Haven't studied the correlation, though.</p>",
      "rawMarkdown": "I grouped on patients and saw an immediate increase in validation loss. Right now the problem seems to be that my validation loss is much larger than my LB (.0814 / .073). Haven't studied the correlation, though."
    },
    {
      "id": 644909,
      "postDate": "2019-10-09T14:04:32.217Z",
      "content": "<p><a href=\"/anjum48\">@anjum48</a> <a href=\"/reppic\">@reppic</a> Are you using this loss, where [any] is the 6th class? How about val/LB correlation without grouping, is overfitting and leak serious?\n<code>\nweights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))\n</code></p>",
      "rawMarkdown": "@anjum48 @reppic Are you using this loss, where [any] is the 6th class? How about val/LB correlation without grouping, is overfitting and leak serious?\n```\nweights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))\n```",
      "replies": [
        {
          "id": 645003,
          "postDate": "2019-10-09T16:38:21.623Z",
          "content": "<p>Yea, essentially. My codes a little different since I’m using Keras and I have <em>any</em> as my first class, but I’m also using weighted bce.</p>\n\n<p>I believe <em>with_logits</em> means the outputs are expected to be unbounded and since my outputs are all between 0 and 1. I’m not using the logits option. It also seemed to take longer to converge when I used <em>with_logits</em></p>",
          "rawMarkdown": "Yea, essentially. My codes a little different since I’m using Keras and I have *any* as my first class, but I’m also using weighted bce.\n\nI believe *with\\_logits* means the outputs are expected to be unbounded and since my outputs are all between 0 and 1. I’m not using the logits option. It also seemed to take longer to converge when I used *with\\_logits*"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 644685,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2019-10-09T07:27:18.213000",
      "content": "<p>I am using <code>GroupKFold</code> on Patient ID and am seeing good CV/LB correlation (LB is 0.03 higher)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 647268,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2019-10-12T10:11:34.527000",
      "content": "<p>When I was randomly splitting, I found that while CV/LB scores initially appeared closer, after X epochs they would diverge and LB would actually start increasing as CV decreased. I think this over fitting is likely due to same patient being in train/validation. </p>\n\n<p>Right now I train using a PatientID split, and find that CV loss stops decreasing after so many epoch, and this correlates to LB performance. However, scores are a little different between CV/LB, differing as much as 0.01. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 647370,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-12T14:17:00.357000",
          "content": "<p><a href=\"/taindow\">@taindow</a> This confirms my observation, thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 647373,
          "author_name": "baigou",
          "author_url": "",
          "post_date": "2019-10-12T14:19:56.550000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 647909,
          "author_name": "tik_boa",
          "author_url": "",
          "post_date": "2019-10-13T13:22:11.043000",
          "content": "<p>Excuse me,  have you try other splitting?  such as according to studyInstanceID....</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 644299,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2019-10-08T15:46:27.657000",
      "content": "<p>Update: I am using efficientnet-b0 trained for 2 epochs, and predictions are made on training data that was totally unseen. I observe a bimodal distribution of predicted confidence that training data is actually training data. Anyone has idea why this happens? I am going to investigate its relationship with labels tomorrow.</p>\n\n<p>Probability distribution of predicted confidence on unseen training data\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F111494fb2bcd616b6e7279d01eb5f7f8%2Fdownload%20(1\" alt=\"\">.png?generation=1570549578416793&amp;alt=media)\n:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2Fa7db8624e70d3b94bd663ee2c2cc9b6f%2Fdownload.png?generation=1570549575213198&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 644647,
      "author_name": "Ryan Epp",
      "author_url": "",
      "post_date": "2019-10-09T06:13:14.380000",
      "content": "<p>If you end up with images from the same study in both the train and validation sets, it's easy to over fit without realizing it. </p>\n\n<p>I'd guess this is also why your adversarial validation accuracy is so high. </p>\n\n<p>Randomly splitting for train/validation is probably not a good idea in this competition. I've been using the metadata and splitting by study and my validation/lb scores are typically under 0.005</p>\n\n<p>Check out how similar some of these examples are: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1202828%2Fd76ab2c2d3e5a85835115917f0cc1ee1%2Fsimilar.png?generation=1570601171279933&amp;alt=media\" alt=\"\">\n(Image borrowed from Bruno Aquino's excellent <a href=\"https://www.kaggle.com/braquino/correct-images-sequece\">Correct Images Sequece</a> Kernel)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 646329,
          "author_name": "NeerajSharma",
          "author_url": "",
          "post_date": "2019-10-11T06:13:07.960000",
          "content": "<p>Great insight. If not random splitting then what is your validation scheme?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 646982,
          "author_name": "Ryan Epp",
          "author_url": "",
          "post_date": "2019-10-11T23:46:37.753000",
          "content": "<p>Thanks Neerja. I’ve been splitting in StudyId but I think datasaurus’s suggestion of using PatientId is probably even better.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 644783,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2019-10-09T10:36:17.740000",
      "content": "<p>Ok, Great! Love the insight:). I will try out your methods</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 647074,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2019-10-12T03:18:33.977000",
      "content": "<p>I grouped on patients and saw an immediate increase in validation loss. Right now the problem seems to be that my validation loss is much larger than my LB (.0814 / .073). Haven't studied the correlation, though.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 644909,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2019-10-09T14:04:32.217000",
      "content": "<p><a href=\"/anjum48\">@anjum48</a> <a href=\"/reppic\">@reppic</a> Are you using this loss, where [any] is the 6th class? How about val/LB correlation without grouping, is overfitting and leak serious?\n<code>\nweights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))\n</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 645003,
          "author_name": "Ryan Epp",
          "author_url": "",
          "post_date": "2019-10-09T16:38:21.623000",
          "content": "<p>Yea, essentially. My codes a little different since I’m using Keras and I have <em>any</em> as my first class, but I’m also using weighted bce.</p>\n\n<p>I believe <em>with_logits</em> means the outputs are expected to be unbounded and since my outputs are all between 0 and 1. I’m not using the logits option. It also seemed to take longer to converge when I used <em>with_logits</em></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "644146": "I have been frustrated over difference between local validation and LB scores (as much as .01+) when using random selection of validation data. Though many has reported little difference between local validation and LB, this is not my case. Discussions hinted that there may exist significant difference between training and test data. \n\nI am doing an adversarial validation test (train a model to try to classify whether an image belongs to training or test data), and preliminary results show that classifier can achieve as much as 77% validation accuracy for adversarial validation (I checked so training and testing data are balanced). \n\nHas anyone done similar analysis to quantize the difference between training / test set? How is you local validation comparing to LB? I think this is very important for establishing useful local validation scores to avoid overfitting.",
    "644685": "I am using `GroupKFold` on Patient ID and am seeing good CV/LB correlation (LB is 0.03 higher)",
    "647268": "When I was randomly splitting, I found that while CV/LB scores initially appeared closer, after X epochs they would diverge and LB would actually start increasing as CV decreased. I think this over fitting is likely due to same patient being in train/validation. \n\nRight now I train using a PatientID split, and find that CV loss stops decreasing after so many epoch, and this correlates to LB performance. However, scores are a little different between CV/LB, differing as much as 0.01. ",
    "644299": "Update: I am using efficientnet-b0 trained for 2 epochs, and predictions are made on training data that was totally unseen. I observe a bimodal distribution of predicted confidence that training data is actually training data. Anyone has idea why this happens? I am going to investigate its relationship with labels tomorrow.\n\nProbability distribution of predicted confidence on unseen training data\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F111494fb2bcd616b6e7279d01eb5f7f8%2Fdownload%20(1).png?generation=1570549578416793&amp;alt=media)\n:![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2Fa7db8624e70d3b94bd663ee2c2cc9b6f%2Fdownload.png?generation=1570549575213198&amp;alt=media)",
    "644647": "If you end up with images from the same study in both the train and validation sets, it's easy to over fit without realizing it. \n\nI'd guess this is also why your adversarial validation accuracy is so high. \n\nRandomly splitting for train/validation is probably not a good idea in this competition. I've been using the metadata and splitting by study and my validation/lb scores are typically under 0.005\n\nCheck out how similar some of these examples are: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1202828%2Fd76ab2c2d3e5a85835115917f0cc1ee1%2Fsimilar.png?generation=1570601171279933&amp;alt=media)\n(Image borrowed from Bruno Aquino's excellent [Correct Images Sequece](https://www.kaggle.com/braquino/correct-images-sequece) Kernel)",
    "644783": "Ok, Great! Love the insight:). I will try out your methods",
    "647074": "I grouped on patients and saw an immediate increase in validation loss. Right now the problem seems to be that my validation loss is much larger than my LB (.0814 / .073). Haven't studied the correlation, though.",
    "644909": "@anjum48 @reppic Are you using this loss, where [any] is the 6th class? How about val/LB correlation without grouping, is overfitting and leak serious?\n```\nweights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return F.binary_cross_entropy_with_logits(y_pred,\n                                  y_true,\n                                  weights.repeat(y_pred.shape[0],1))\n```"
  }
}