{
  "id": 451424,
  "title": "How to know when we stop training a DL model while training with all data? ",
  "url": "/competitions/UBC-OCEAN/discussion/451424",
  "author_name": "turkenm",
  "post_date": "2023-10-28T17:43:06.197000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>My question is not specific to this competition but rather related to the general domain of DL model training. Let’s say we have performed a 5-fold cross-validation for a DL model, and based on the training and validation metrics, we have selected the model architecture, optimizer, and hyperparameters to use. Later, we want to train the model on the entire dataset for the final model. At this point, when do we need to stop training the model that we have trained with all available data, how can we make this decision? I have researched model selection and model validation topics, but I couldn’t find a specific answer to the question I mentioned.</p>",
  "messages": [
    {
      "id": 2504733,
      "postDate": "2023-10-30T06:00:50.567Z",
      "content": "<p>There is no way to know when to stop the training without validation data. There are various ways to estimate or guess, but there is no established way that is guaranteed to work. I think using N folds and creating N models, and averaging predictions from those N models, is a more general way to solve this problem. More importantly, it is the way where there is no need to guess when to stop the training.</p>\n<p>To put it simply: we understand how to make 5 DL models each trained on 80% of data and average their predictions. We don't know for sure how to use the information from those 5 models to stop the training on full data, nor do we know with certainty that a single model trained on 100% of data would be better than an average of 5 models trained on 80% of data. And that goes for any reasonably large value of N - I used 5 only as an example.</p>",
      "rawMarkdown": "There is no way to know when to stop the training without validation data. There are various ways to estimate or guess, but there is no established way that is guaranteed to work. I think using N folds and creating N models, and averaging predictions from those N models, is a more general way to solve this problem. More importantly, it is the way where there is no need to guess when to stop the training.\n\nTo put it simply: we understand how to make 5 DL models each trained on 80% of data and average their predictions. We don't know for sure how to use the information from those 5 models to stop the training on full data, nor do we know with certainty that a single model trained on 100% of data would be better than an average of 5 models trained on 80% of data. And that goes for any reasonably large value of N - I used 5 only as an example.",
      "votes": 4,
      "replies": [
        {
          "id": 2504939,
          "postDate": "2023-10-30T09:07:01.727Z",
          "content": "<p>Thank you for the answer. In that case, the safest approach would be split the data into train, test and validation, if we have enough data, and ignore the validation data as if it doesn't exist. </p>",
          "rawMarkdown": "Thank you for the answer. In that case, the safest approach would be split the data into train, test and validation, if we have enough data, and ignore the validation data as if it doesn't exist. ",
          "votes": -2,
          "replies": [
            {
              "id": 2505848,
              "postDate": "2023-10-30T21:13:22.613Z",
              "content": "<blockquote>\n  <p>In that case, the safest approach would be split the data into train, test and validation, if we have enough data, and ignore the validation data as if it doesn't exist.</p>\n</blockquote>\n<p>Not sure how you made that conclusion from my post. The best approach I endorsed was to use the N-fold validation, where one uses all the training data in one fold or another, and makes N predictions which are then averaged.</p>",
              "rawMarkdown": "> In that case, the safest approach would be split the data into train, test and validation, if we have enough data, and ignore the validation data as if it doesn't exist.\n\nNot sure how you made that conclusion from my post. The best approach I endorsed was to use the N-fold validation, where one uses all the training data in one fold or another, and makes N predictions which are then averaged.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2503044,
      "postDate": "2023-10-28T17:43:06.197Z",
      "content": "<p>My question is not specific to this competition but rather related to the general domain of DL model training. Let’s say we have performed a 5-fold cross-validation for a DL model, and based on the training and validation metrics, we have selected the model architecture, optimizer, and hyperparameters to use. Later, we want to train the model on the entire dataset for the final model. At this point, when do we need to stop training the model that we have trained with all available data, how can we make this decision? I have researched model selection and model validation topics, but I couldn’t find a specific answer to the question I mentioned.</p>",
      "rawMarkdown": "My question is not specific to this competition but rather related to the general domain of DL model training. Let’s say we have performed a 5-fold cross-validation for a DL model, and based on the training and validation metrics, we have selected the model architecture, optimizer, and hyperparameters to use. Later, we want to train the model on the entire dataset for the final model. At this point, when do we need to stop training the model that we have trained with all available data, how can we make this decision? I have researched model selection and model validation topics, but I couldn’t find a specific answer to the question I mentioned.",
      "votes": 1
    },
    {
      "id": 2554685,
      "postDate": "2023-12-09T10:05:46.163Z",
      "content": "<p>I think there's no certain way to determine that. One possibility is to use average parameters obtained from CV training (like average number of epochs with best val score). But without the validation set it's not possible to determine if selected parameters are optimal.</p>",
      "rawMarkdown": "I think there's no certain way to determine that. One possibility is to use average parameters obtained from CV training (like average number of epochs with best val score). But without the validation set it's not possible to determine if selected parameters are optimal."
    },
    {
      "id": 2503384,
      "postDate": "2023-10-29T03:45:57.117Z",
      "content": "<p>Hmm good question. Probably just stick to the number of epochs that worked best in CV?</p>\n<p>Alternatively, try 1 epoch more &amp; 1 epoch less (compared to what worked best in CV), submit each of these, and then select the checkpoint that gives the highest public LB score.</p>",
      "rawMarkdown": "Hmm good question. Probably just stick to the number of epochs that worked best in CV?\n\nAlternatively, try 1 epoch more & 1 epoch less (compared to what worked best in CV), submit each of these, and then select the checkpoint that gives the highest public LB score.",
      "replies": [
        {
          "id": 2503840,
          "postDate": "2023-10-29T13:53:24.323Z",
          "content": "<p>Thank you for the answer. It can be good approach but selecting checkpoint wrt public LB can mislead. I think there should be a more robust approach for stop training. </p>",
          "rawMarkdown": "Thank you for the answer. It can be good approach but selecting checkpoint wrt public LB can mislead. I think there should be a more robust approach for stop training. "
        },
        {
          "id": 2554662,
          "postDate": "2023-12-09T09:35:05.417Z",
          "content": "<p>Setting up any hyperparameter based on leaderboard score is information leak from the test set. This can help improve the score on the public leaderboard, but the private leaderboard is still hidden and contains 78% of the data. Hence, it is not advisable.</p>",
          "rawMarkdown": "Setting up any hyperparameter based on leaderboard score is information leak from the test set. This can help improve the score on the public leaderboard, but the private leaderboard is still hidden and contains 78% of the data. Hence, it is not advisable."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2504733,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2023-10-30T06:00:50.567000",
      "content": "<p>There is no way to know when to stop the training without validation data. There are various ways to estimate or guess, but there is no established way that is guaranteed to work. I think using N folds and creating N models, and averaging predictions from those N models, is a more general way to solve this problem. More importantly, it is the way where there is no need to guess when to stop the training.</p>\n<p>To put it simply: we understand how to make 5 DL models each trained on 80% of data and average their predictions. We don't know for sure how to use the information from those 5 models to stop the training on full data, nor do we know with certainty that a single model trained on 100% of data would be better than an average of 5 models trained on 80% of data. And that goes for any reasonably large value of N - I used 5 only as an example.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2504939,
          "author_name": "turkenm",
          "author_url": "",
          "post_date": "2023-10-30T09:07:01.727000",
          "content": "<p>Thank you for the answer. In that case, the safest approach would be split the data into train, test and validation, if we have enough data, and ignore the validation data as if it doesn't exist. </p>",
          "votes": -2,
          "replies": [
            {
              "id": 2505848,
              "author_name": "Tilii",
              "author_url": "",
              "post_date": "2023-10-30T21:13:22.613000",
              "content": "<blockquote>\n  <p>In that case, the safest approach would be split the data into train, test and validation, if we have enough data, and ignore the validation data as if it doesn't exist.</p>\n</blockquote>\n<p>Not sure how you made that conclusion from my post. The best approach I endorsed was to use the N-fold validation, where one uses all the training data in one fold or another, and makes N predictions which are then averaged.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2554685,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2023-12-09T10:05:46.163000",
      "content": "<p>I think there's no certain way to determine that. One possibility is to use average parameters obtained from CV training (like average number of epochs with best val score). But without the validation set it's not possible to determine if selected parameters are optimal.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2503384,
      "author_name": "Sadhaklal",
      "author_url": "",
      "post_date": "2023-10-29T03:45:57.117000",
      "content": "<p>Hmm good question. Probably just stick to the number of epochs that worked best in CV?</p>\n<p>Alternatively, try 1 epoch more &amp; 1 epoch less (compared to what worked best in CV), submit each of these, and then select the checkpoint that gives the highest public LB score.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2503840,
          "author_name": "turkenm",
          "author_url": "",
          "post_date": "2023-10-29T13:53:24.323000",
          "content": "<p>Thank you for the answer. It can be good approach but selecting checkpoint wrt public LB can mislead. I think there should be a more robust approach for stop training. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2554662,
          "author_name": "Akash Gupta",
          "author_url": "",
          "post_date": "2023-12-09T09:35:05.417000",
          "content": "<p>Setting up any hyperparameter based on leaderboard score is information leak from the test set. This can help improve the score on the public leaderboard, but the private leaderboard is still hidden and contains 78% of the data. Hence, it is not advisable.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2504733": "There is no way to know when to stop the training without validation data. There are various ways to estimate or guess, but there is no established way that is guaranteed to work. I think using N folds and creating N models, and averaging predictions from those N models, is a more general way to solve this problem. More importantly, it is the way where there is no need to guess when to stop the training.\n\nTo put it simply: we understand how to make 5 DL models each trained on 80% of data and average their predictions. We don't know for sure how to use the information from those 5 models to stop the training on full data, nor do we know with certainty that a single model trained on 100% of data would be better than an average of 5 models trained on 80% of data. And that goes for any reasonably large value of N - I used 5 only as an example.",
    "2503044": "My question is not specific to this competition but rather related to the general domain of DL model training. Let’s say we have performed a 5-fold cross-validation for a DL model, and based on the training and validation metrics, we have selected the model architecture, optimizer, and hyperparameters to use. Later, we want to train the model on the entire dataset for the final model. At this point, when do we need to stop training the model that we have trained with all available data, how can we make this decision? I have researched model selection and model validation topics, but I couldn’t find a specific answer to the question I mentioned.",
    "2554685": "I think there's no certain way to determine that. One possibility is to use average parameters obtained from CV training (like average number of epochs with best val score). But without the validation set it's not possible to determine if selected parameters are optimal.",
    "2503384": "Hmm good question. Probably just stick to the number of epochs that worked best in CV?\n\nAlternatively, try 1 epoch more & 1 epoch less (compared to what worked best in CV), submit each of these, and then select the checkpoint that gives the highest public LB score."
  }
}