{
  "id": 336514,
  "title": "The Weirdness of Randomness",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/336514",
  "author_name": "Harshit Sheoran",
  "post_date": "2022-07-11T15:17:24.268000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>So, I was testing out models on my crop-tiles-1024-balanced dataset, and I do my CV in a format which is exactly how the LB does it, scoring it by patient_id (patient id is the first subpart of 'id' in train.csv), and everything was going altright, I was seeing subtle improvements from small changes in models, like from resnet18 to 34 to 50, but when I changed it from 50 to 101, I saw a major boost, the score (accuracy) went from .68 to .73, which was a huge difference, and first thing I did was not to accept it but to check other big models, from other big models I have tried, they were not manage to replicate the results, so I did a batch-size stability check, meaning, I changed my per-GPU batch-size from 32 to 8, and score of resnet101 also dropped to a bit above that of resnet50, one explaination is my validation set of this fold is only ~150 patients, it could be really high roll  on the randomness dice.<br>\nLet me know what conclusion would you make from this study?</p>",
  "messages": [
    {
      "id": 1851849,
      "postDate": "2022-07-11T15:17:24.270Z",
      "content": "<p>So, I was testing out models on my crop-tiles-1024-balanced dataset, and I do my CV in a format which is exactly how the LB does it, scoring it by patient_id (patient id is the first subpart of 'id' in train.csv), and everything was going altright, I was seeing subtle improvements from small changes in models, like from resnet18 to 34 to 50, but when I changed it from 50 to 101, I saw a major boost, the score (accuracy) went from .68 to .73, which was a huge difference, and first thing I did was not to accept it but to check other big models, from other big models I have tried, they were not manage to replicate the results, so I did a batch-size stability check, meaning, I changed my per-GPU batch-size from 32 to 8, and score of resnet101 also dropped to a bit above that of resnet50, one explaination is my validation set of this fold is only ~150 patients, it could be really high roll  on the randomness dice.<br>\nLet me know what conclusion would you make from this study?</p>",
      "rawMarkdown": "So, I was testing out models on my crop-tiles-1024-balanced dataset, and I do my CV in a format which is exactly how the LB does it, scoring it by patient_id (patient id is the first subpart of 'id' in train.csv), and everything was going altright, I was seeing subtle improvements from small changes in models, like from resnet18 to 34 to 50, but when I changed it from 50 to 101, I saw a major boost, the score (accuracy) went from .68 to .73, which was a huge difference, and first thing I did was not to accept it but to check other big models, from other big models I have tried, they were not manage to replicate the results, so I did a batch-size stability check, meaning, I changed my per-GPU batch-size from 32 to 8, and score of resnet101 also dropped to a bit above that of resnet50, one explaination is my validation set of this fold is only ~150 patients, it could be really high roll  on the randomness dice.\nLet me know what conclusion would you make from this study?",
      "votes": 4
    },
    {
      "id": 1851992,
      "postDate": "2022-07-11T17:29:29.767Z",
      "content": "<p>That's a nice insight. My guess is that since the competition's datasets are quite small then randomness plays a huge role here. Have you seed everything in your notebooks to ensure that when you change the CNN architecture data is split the same way? </p>",
      "rawMarkdown": "That's a nice insight. My guess is that since the competition's datasets are quite small then randomness plays a huge role here. Have you seed everything in your notebooks to ensure that when you change the CNN architecture data is split the same way? ",
      "replies": [
        {
          "id": 1851994,
          "postDate": "2022-07-11T17:31:27.197Z",
          "content": "<p>Ofc, absolutely everything is seeded, I can reproduce results to the last decimal digit (means, completely same results every time)</p>",
          "rawMarkdown": "Ofc, absolutely everything is seeded, I can reproduce results to the last decimal digit (means, completely same results every time)"
        },
        {
          "id": 1852007,
          "postDate": "2022-07-11T17:37:20.087Z",
          "content": "<p>I will let you know if I achieve higher performance when switching to larger CNN architectures 👌🏽</p>",
          "rawMarkdown": "I will let you know if I achieve higher performance when switching to larger CNN architectures 👌🏽",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1851992,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2022-07-11T17:29:29.767000",
      "content": "<p>That's a nice insight. My guess is that since the competition's datasets are quite small then randomness plays a huge role here. Have you seed everything in your notebooks to ensure that when you change the CNN architecture data is split the same way? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1851994,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-07-11T17:31:27.197000",
          "content": "<p>Ofc, absolutely everything is seeded, I can reproduce results to the last decimal digit (means, completely same results every time)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1852007,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-07-11T17:37:20.087000",
          "content": "<p>I will let you know if I achieve higher performance when switching to larger CNN architectures 👌🏽</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1851849": "So, I was testing out models on my crop-tiles-1024-balanced dataset, and I do my CV in a format which is exactly how the LB does it, scoring it by patient_id (patient id is the first subpart of 'id' in train.csv), and everything was going altright, I was seeing subtle improvements from small changes in models, like from resnet18 to 34 to 50, but when I changed it from 50 to 101, I saw a major boost, the score (accuracy) went from .68 to .73, which was a huge difference, and first thing I did was not to accept it but to check other big models, from other big models I have tried, they were not manage to replicate the results, so I did a batch-size stability check, meaning, I changed my per-GPU batch-size from 32 to 8, and score of resnet101 also dropped to a bit above that of resnet50, one explaination is my validation set of this fold is only ~150 patients, it could be really high roll  on the randomness dice.\nLet me know what conclusion would you make from this study?",
    "1851992": "That's a nice insight. My guess is that since the competition's datasets are quite small then randomness plays a huge role here. Have you seed everything in your notebooks to ensure that when you change the CNN architecture data is split the same way? "
  }
}