{
  "id": 416395,
  "title": "cross validation",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/416395",
  "author_name": "ma_si_620",
  "post_date": "2023-06-11T07:16:44.175000",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am new to kaggle and have a question about cross-validation. cv(k-fold) divides the data into test and train (in deep learning, train is further divided into train and valid), evaluates k times and calculates the result.</p>\n<p>Q1. I have then obtained 5 trained models after cv. Which one should I submit in the end?</p>\n<p>Q2. In this competition, we already have data divided into train and validation. How should we perform cross-validation? (Do I do cv after merging train and validation?)</p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": 2295731,
      "postDate": "2023-06-11T08:08:53.680Z",
      "content": "<p>Q1- You can either pick one randomly or pick the 5, and <a href=\"https://arxiv.org/abs/2111.13280\" target=\"_blank\">ensemble </a>their predictions <br>\nQ2- You can and I think you should. the original split is a 90-10 split, which is a little small on the validation part in my opinion.<br>\nif you are computationnaly limited, peform the 80-20 5 split (stratified on the presence or not of a contrail) and train/valid on 2-3 splits, that should be robust enough and not that expensive.</p>\n<p>Here is how I do things so you get an idea of how it is done:</p>\n<pre><code> ():\n\n    skfold = StratifiedKFold(n_splits=num_splits, shuffle=, random_state=random_state)\n    train_loaders = []\n    valid_loaders = []\n\n     train_indices, valid_indices  skfold.split(metadata, metadata[]):\n        train_metadata = metadata.iloc[train_indices].reset_index(drop=)\n        valid_metadata = metadata.iloc[valid_indices].reset_index(drop=)\n\n        train_dataset = FakeColorDataset(train_metadata, transform_train, train=)\n        valid_dataset = FakeColorDataset(valid_metadata, transform_test, train=)\n\n        train_loader = DataLoader(train_dataset, batch_size=batch_size_train, shuffle=)\n        valid_loader = DataLoader(valid_dataset, batch_size=batch_size_valid, shuffle=)\n\n        train_loaders.append(train_loader)\n        valid_loaders.append(valid_loader)\n\n     train_loaders, valid_loaders\n</code></pre>\n<p><em>Note: metadata is the concatenation of the train and valid dataframes. I added the path to the labels and images before which I use in the <code>FakeColorDataset</code></em></p>",
      "rawMarkdown": "Q1- You can either pick one randomly or pick the 5, and [ensemble ](https://arxiv.org/abs/2111.13280)their predictions \nQ2- You can and I think you should. the original split is a 90-10 split, which is a little small on the validation part in my opinion.\nif you are computationnaly limited, peform the 80-20 5 split (stratified on the presence or not of a contrail) and train/valid on 2-3 splits, that should be robust enough and not that expensive.\n\nHere is how I do things so you get an idea of how it is done:\n```python\ndef stratified_kfold_loaders(metadata, transform_train = None, transform_test = None, batch_size_train = 64, batch_size_valid = 64, num_splits=5, random_state=42):\n\n    skfold = StratifiedKFold(n_splits=num_splits, shuffle=True, random_state=random_state)\n    train_loaders = []\n    valid_loaders = []\n\n    for train_indices, valid_indices in skfold.split(metadata, metadata[\"contrail\"]):\n        train_metadata = metadata.iloc[train_indices].reset_index(drop=True)\n        valid_metadata = metadata.iloc[valid_indices].reset_index(drop=True)\n\n        train_dataset = FakeColorDataset(train_metadata, transform_train, train=True)\n        valid_dataset = FakeColorDataset(valid_metadata, transform_test, train=False)\n\n        train_loader = DataLoader(train_dataset, batch_size=batch_size_train, shuffle=True)\n        valid_loader = DataLoader(valid_dataset, batch_size=batch_size_valid, shuffle=False)\n\n        train_loaders.append(train_loader)\n        valid_loaders.append(valid_loader)\n    \n    return train_loaders, valid_loaders\n```\n\n*Note: metadata is the concatenation of the train and valid dataframes. I added the path to the labels and images before which I use in the `FakeColorDataset`*",
      "votes": 2,
      "replies": [
        {
          "id": 2295829,
          "postDate": "2023-06-11T09:42:07.147Z",
          "content": "<p>Thank you for providing an answer to my question. I really appreciate it.</p>",
          "rawMarkdown": "Thank you for providing an answer to my question. I really appreciate it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2295650,
      "postDate": "2023-06-11T07:16:44.177Z",
      "content": "<p>I am new to kaggle and have a question about cross-validation. cv(k-fold) divides the data into test and train (in deep learning, train is further divided into train and valid), evaluates k times and calculates the result.</p>\n<p>Q1. I have then obtained 5 trained models after cv. Which one should I submit in the end?</p>\n<p>Q2. In this competition, we already have data divided into train and validation. How should we perform cross-validation? (Do I do cv after merging train and validation?)</p>\n<p>Thank you.</p>",
      "rawMarkdown": "I am new to kaggle and have a question about cross-validation. cv(k-fold) divides the data into test and train (in deep learning, train is further divided into train and valid), evaluates k times and calculates the result.\n\nQ1. I have then obtained 5 trained models after cv. Which one should I submit in the end?\n\nQ2. In this competition, we already have data divided into train and validation. How should we perform cross-validation? (Do I do cv after merging train and validation?)\n\nThank you."
    }
  ],
  "comments": [
    {
      "id": 2295731,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-06-11T08:08:53.680000",
      "content": "<p>Q1- You can either pick one randomly or pick the 5, and <a href=\"https://arxiv.org/abs/2111.13280\" target=\"_blank\">ensemble </a>their predictions <br>\nQ2- You can and I think you should. the original split is a 90-10 split, which is a little small on the validation part in my opinion.<br>\nif you are computationnaly limited, peform the 80-20 5 split (stratified on the presence or not of a contrail) and train/valid on 2-3 splits, that should be robust enough and not that expensive.</p>\n<p>Here is how I do things so you get an idea of how it is done:</p>\n<pre><code> ():\n\n    skfold = StratifiedKFold(n_splits=num_splits, shuffle=, random_state=random_state)\n    train_loaders = []\n    valid_loaders = []\n\n     train_indices, valid_indices  skfold.split(metadata, metadata[]):\n        train_metadata = metadata.iloc[train_indices].reset_index(drop=)\n        valid_metadata = metadata.iloc[valid_indices].reset_index(drop=)\n\n        train_dataset = FakeColorDataset(train_metadata, transform_train, train=)\n        valid_dataset = FakeColorDataset(valid_metadata, transform_test, train=)\n\n        train_loader = DataLoader(train_dataset, batch_size=batch_size_train, shuffle=)\n        valid_loader = DataLoader(valid_dataset, batch_size=batch_size_valid, shuffle=)\n\n        train_loaders.append(train_loader)\n        valid_loaders.append(valid_loader)\n\n     train_loaders, valid_loaders\n</code></pre>\n<p><em>Note: metadata is the concatenation of the train and valid dataframes. I added the path to the labels and images before which I use in the <code>FakeColorDataset</code></em></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2295829,
          "author_name": "ma_si_620",
          "author_url": "",
          "post_date": "2023-06-11T09:42:07.147000",
          "content": "<p>Thank you for providing an answer to my question. I really appreciate it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2295731": "Q1- You can either pick one randomly or pick the 5, and [ensemble ](https://arxiv.org/abs/2111.13280)their predictions \nQ2- You can and I think you should. the original split is a 90-10 split, which is a little small on the validation part in my opinion.\nif you are computationnaly limited, peform the 80-20 5 split (stratified on the presence or not of a contrail) and train/valid on 2-3 splits, that should be robust enough and not that expensive.\n\nHere is how I do things so you get an idea of how it is done:\n```python\ndef stratified_kfold_loaders(metadata, transform_train = None, transform_test = None, batch_size_train = 64, batch_size_valid = 64, num_splits=5, random_state=42):\n\n    skfold = StratifiedKFold(n_splits=num_splits, shuffle=True, random_state=random_state)\n    train_loaders = []\n    valid_loaders = []\n\n    for train_indices, valid_indices in skfold.split(metadata, metadata[\"contrail\"]):\n        train_metadata = metadata.iloc[train_indices].reset_index(drop=True)\n        valid_metadata = metadata.iloc[valid_indices].reset_index(drop=True)\n\n        train_dataset = FakeColorDataset(train_metadata, transform_train, train=True)\n        valid_dataset = FakeColorDataset(valid_metadata, transform_test, train=False)\n\n        train_loader = DataLoader(train_dataset, batch_size=batch_size_train, shuffle=True)\n        valid_loader = DataLoader(valid_dataset, batch_size=batch_size_valid, shuffle=False)\n\n        train_loaders.append(train_loader)\n        valid_loaders.append(valid_loader)\n    \n    return train_loaders, valid_loaders\n```\n\n*Note: metadata is the concatenation of the train and valid dataframes. I added the path to the labels and images before which I use in the `FakeColorDataset`*",
    "2295650": "I am new to kaggle and have a question about cross-validation. cv(k-fold) divides the data into test and train (in deep learning, train is further divided into train and valid), evaluates k times and calculates the result.\n\nQ1. I have then obtained 5 trained models after cv. Which one should I submit in the end?\n\nQ2. In this competition, we already have data divided into train and validation. How should we perform cross-validation? (Do I do cv after merging train and validation?)\n\nThank you."
  }
}