{
  "id": 340926,
  "title": "Why does shuffle causes accuracy drop?",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/340926",
  "author_name": "Chet",
  "post_date": "2022-07-31T13:58:42.687000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I followed this notebook<a href=\"https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-training\" target=\"_blank\">KABIR IVAN</a>.<br>\nOnce turn on shuffle for the trainloader, it causes the validation set accuracy to drop from 50% to less than 20%.</p>",
  "messages": [
    {
      "id": 1878624,
      "postDate": "2022-07-31T13:58:42.687Z",
      "content": "<p>I followed this notebook<a href=\"https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-training\" target=\"_blank\">KABIR IVAN</a>.<br>\nOnce turn on shuffle for the trainloader, it causes the validation set accuracy to drop from 50% to less than 20%.</p>",
      "rawMarkdown": "I followed this notebook[KABIR IVAN](https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-training).\nOnce turn on shuffle for the trainloader, it causes the validation set accuracy to drop from 50% to less than 20%.",
      "votes": 4
    },
    {
      "id": 1929823,
      "postDate": "2022-09-07T11:44:47.130Z",
      "content": "<p>True, I just joined and had the same observation/question - the same holds even even when you split by GroupKFold by patient id  </p>",
      "rawMarkdown": "True, I just joined and had the same observation/question - the same holds even even when you split by GroupKFold by patient id  ",
      "votes": 1
    },
    {
      "id": 1879200,
      "postDate": "2022-07-31T22:46:20.500Z",
      "content": "<p>Maybe you have something wrong with arrays of labels and images themselves when loading a dataset. So your labels shouldn't be 'static' but attached to the images even when they are shuffled.</p>\n<p>paths_for_labels = [path[:-4] for path in self.files]<br>\nself.labels = []<br>\nfor i in range(len(paths_for_labels)):<br>\n    self.labels.append(dataset_for_training.loc[dataset_for_training['image_id'] == paths_for_labels[i], 'label'].item())</p>\n<p>That worked for me.<br>\nMy code isn't nice but I think you'll understand.<br>\nMy solution may help in case the problem was in labels.</p>",
      "rawMarkdown": "Maybe you have something wrong with arrays of labels and images themselves when loading a dataset. So your labels shouldn't be 'static' but attached to the images even when they are shuffled.\n\npaths_for_labels = [path[:-4] for path in self.files]\nself.labels = []\nfor i in range(len(paths_for_labels)):\n    self.labels.append(dataset_for_training.loc[dataset_for_training['image_id'] == paths_for_labels[i], 'label'].item())\n\nThat worked for me.\nMy code isn't nice but I think you'll understand.\nMy solution may help in case the problem was in labels.",
      "votes": 1,
      "replies": [
        {
          "id": 1890140,
          "postDate": "2022-08-08T15:10:12.617Z",
          "content": "<p>Sorry, actually I can't understand your code. I follow the <a href=\"https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-training\" target=\"_blank\">code</a>, and I think this code does not have the problem that you called 'static'.<br>\nI think this problem may come from the continuity of images of the same patient. <strong>Shuffing</strong> damages the continuity and may cause a shock of optimization.</p>",
          "rawMarkdown": "Sorry, actually I can't understand your code. I follow the [code](https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-training), and I think this code does not have the problem that you called 'static'.\nI think this problem may come from the continuity of images of the same patient. **Shuffing** damages the continuity and may cause a shock of optimization.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1929823,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2022-09-07T11:44:47.130000",
      "content": "<p>True, I just joined and had the same observation/question - the same holds even even when you split by GroupKFold by patient id  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1879200,
      "author_name": "pastafarian",
      "author_url": "",
      "post_date": "2022-07-31T22:46:20.500000",
      "content": "<p>Maybe you have something wrong with arrays of labels and images themselves when loading a dataset. So your labels shouldn't be 'static' but attached to the images even when they are shuffled.</p>\n<p>paths_for_labels = [path[:-4] for path in self.files]<br>\nself.labels = []<br>\nfor i in range(len(paths_for_labels)):<br>\n    self.labels.append(dataset_for_training.loc[dataset_for_training['image_id'] == paths_for_labels[i], 'label'].item())</p>\n<p>That worked for me.<br>\nMy code isn't nice but I think you'll understand.<br>\nMy solution may help in case the problem was in labels.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1890140,
          "author_name": "Chet",
          "author_url": "",
          "post_date": "2022-08-08T15:10:12.617000",
          "content": "<p>Sorry, actually I can't understand your code. I follow the <a href=\"https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-training\" target=\"_blank\">code</a>, and I think this code does not have the problem that you called 'static'.<br>\nI think this problem may come from the continuity of images of the same patient. <strong>Shuffing</strong> damages the continuity and may cause a shock of optimization.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1878624": "I followed this notebook[KABIR IVAN](https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-training).\nOnce turn on shuffle for the trainloader, it causes the validation set accuracy to drop from 50% to less than 20%.",
    "1929823": "True, I just joined and had the same observation/question - the same holds even even when you split by GroupKFold by patient id  ",
    "1879200": "Maybe you have something wrong with arrays of labels and images themselves when loading a dataset. So your labels shouldn't be 'static' but attached to the images even when they are shuffled.\n\npaths_for_labels = [path[:-4] for path in self.files]\nself.labels = []\nfor i in range(len(paths_for_labels)):\n    self.labels.append(dataset_for_training.loc[dataset_for_training['image_id'] == paths_for_labels[i], 'label'].item())\n\nThat worked for me.\nMy code isn't nice but I think you'll understand.\nMy solution may help in case the problem was in labels."
  }
}