{
  "id": 164167,
  "title": "Any possibility of changing public dataset",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/164167",
  "author_name": "Jaideep",
  "post_date": "2020-07-05T03:15:00.504000",
  "votes": -1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Just want to check with organizers if public dataset getting randomised while calculating  getting  public lb score so it can be reason for inconsistency of scores for many .</p>",
  "messages": [
    {
      "id": 916480,
      "postDate": "2020-07-05T17:36:27.343Z",
      "content": "<p>What ? Just run twice your submission and your score won't change... People are getting so much variation because their CV set is not such a good representation of the test set. It's even mentioned in the overview that the test and train set have had different processes so it's not exactly surprising that the CV LB correlation is not great.</p>",
      "rawMarkdown": "What ? Just run twice your submission and your score won't change... People are getting so much variation because their CV set is not such a good representation of the test set. It's even mentioned in the overview that the test and train set have had different processes so it's not exactly surprising that the CV LB correlation is not great.",
      "replies": [
        {
          "id": 916487,
          "postDate": "2020-07-05T17:41:47.083Z",
          "content": "<p>im unable to produce my current score with same set of model and params so it is making me think that. </p>",
          "rawMarkdown": "im unable to produce my current score with same set of model and params so it is making me think that. "
        },
        {
          "id": 916502,
          "postDate": "2020-07-05T17:49:45.620Z",
          "content": "<p>Are you sure you are setting all seeds correctly and running torch.backends.cudnn.deterministic = True (if you're using pytorch)?</p>\n\n<p>Something like:</p>\n\n<p><code>\nos.environ['PYTHONHASHSEED'] = str(SEED)\nrandom.seed(SEED)\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\ntorch.cuda.manual_seed(SEED)\ntorch.backends.cudnn.deterministic = True\n</code>\nResults should be reproducible in this case (at least I'm finding that they are).</p>",
          "rawMarkdown": "Are you sure you are setting all seeds correctly and running torch.backends.cudnn.deterministic = True (if you're using pytorch)?\n\nSomething like:\n\n```\nos.environ['PYTHONHASHSEED'] = str(SEED)\nrandom.seed(SEED)\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\ntorch.cuda.manual_seed(SEED)\ntorch.backends.cudnn.deterministic = True\n```\nResults should be reproducible in this case (at least I'm finding that they are)."
        }
      ]
    },
    {
      "id": 915796,
      "postDate": "2020-07-05T05:19:35.077Z",
      "content": "<p>Even if you just randomize the public test set, why would it change the LB score? I think in evaluation, randomizing the test set has no effect cause all ur model weights including the batchnorm are fixed.</p>",
      "rawMarkdown": "Even if you just randomize the public test set, why would it change the LB score? I think in evaluation, randomizing the test set has no effect cause all ur model weights including the batchnorm are fixed.",
      "replies": [
        {
          "id": 916053,
          "postDate": "2020-07-05T10:08:17.953Z",
          "content": "<p>Here test set is half or less ,some part of test set would give x qwk ,some would give Y qwk .if we were to get same metric value irrespective of size of test set being evaluated then public lb score would have been 98 percent reflecting final standing </p>",
          "rawMarkdown": "Here test set is half or less ,some part of test set would give x qwk ,some would give Y qwk .if we were to get same metric value irrespective of size of test set being evaluated then public lb score would have been 98 percent reflecting final standing "
        },
        {
          "id": 916068,
          "postDate": "2020-07-05T10:27:18.773Z",
          "content": "<p>So, were u asking whether the public test set is created (again from scratch) by sampling images from the whole (private + public) test set <strong>every-time</strong> we make a submission?</p>",
          "rawMarkdown": "So, were u asking whether the public test set is created (again from scratch) by sampling images from the whole (private + public) test set **every-time** we make a submission?"
        },
        {
          "id": 916074,
          "postDate": "2020-07-05T10:34:17.563Z",
          "content": "<p>yes.. some thing like Public Test set=random.sample(Private_set,% of Private)</p>\n\n<p>otherwise imagine why should there be so much oscillations in the scores of many individuals except those who have pretty much same Ra and Ka CV score.</p>",
          "rawMarkdown": "yes.. some thing like Public Test set=random.sample(Private_set,% of Private)\n\notherwise imagine why should there be so much oscillations in the scores of many individuals except those who have pretty much same Ra and Ka CV score."
        }
      ]
    },
    {
      "id": 915702,
      "postDate": "2020-07-05T03:15:00.503Z",
      "content": "<p>Just want to check with organizers if public dataset getting randomised while calculating  getting  public lb score so it can be reason for inconsistency of scores for many .</p>",
      "rawMarkdown": "Just want to check with organizers if public dataset getting randomised while calculating  getting  public lb score so it can be reason for inconsistency of scores for many .",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 916480,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-07-05T17:36:27.343000",
      "content": "<p>What ? Just run twice your submission and your score won't change... People are getting so much variation because their CV set is not such a good representation of the test set. It's even mentioned in the overview that the test and train set have had different processes so it's not exactly surprising that the CV LB correlation is not great.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 916487,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-05T17:41:47.083000",
          "content": "<p>im unable to produce my current score with same set of model and params so it is making me think that. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916502,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-07-05T17:49:45.620000",
          "content": "<p>Are you sure you are setting all seeds correctly and running torch.backends.cudnn.deterministic = True (if you're using pytorch)?</p>\n\n<p>Something like:</p>\n\n<p><code>\nos.environ['PYTHONHASHSEED'] = str(SEED)\nrandom.seed(SEED)\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\ntorch.cuda.manual_seed(SEED)\ntorch.backends.cudnn.deterministic = True\n</code>\nResults should be reproducible in this case (at least I'm finding that they are).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 915796,
      "author_name": "Viraj Bagal",
      "author_url": "",
      "post_date": "2020-07-05T05:19:35.077000",
      "content": "<p>Even if you just randomize the public test set, why would it change the LB score? I think in evaluation, randomizing the test set has no effect cause all ur model weights including the batchnorm are fixed.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 916053,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-05T10:08:17.953000",
          "content": "<p>Here test set is half or less ,some part of test set would give x qwk ,some would give Y qwk .if we were to get same metric value irrespective of size of test set being evaluated then public lb score would have been 98 percent reflecting final standing </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916068,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2020-07-05T10:27:18.773000",
          "content": "<p>So, were u asking whether the public test set is created (again from scratch) by sampling images from the whole (private + public) test set <strong>every-time</strong> we make a submission?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916074,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-05T10:34:17.563000",
          "content": "<p>yes.. some thing like Public Test set=random.sample(Private_set,% of Private)</p>\n\n<p>otherwise imagine why should there be so much oscillations in the scores of many individuals except those who have pretty much same Ra and Ka CV score.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "916480": "What ? Just run twice your submission and your score won't change... People are getting so much variation because their CV set is not such a good representation of the test set. It's even mentioned in the overview that the test and train set have had different processes so it's not exactly surprising that the CV LB correlation is not great.",
    "915796": "Even if you just randomize the public test set, why would it change the LB score? I think in evaluation, randomizing the test set has no effect cause all ur model weights including the batchnorm are fixed.",
    "915702": "Just want to check with organizers if public dataset getting randomised while calculating  getting  public lb score so it can be reason for inconsistency of scores for many ."
  }
}