{
  "id": 377831,
  "title": "My VAL pF1 score is around 0.04, however my VAL is around 0.66 to 0.69. Unable to improve score",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/377831",
  "author_name": "Pranay Barkataki",
  "post_date": "2023-01-13T04:09:27.663000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am using the EfficientNetB4 model, and training PNG images are of dimension 1024x512. I have tried upsampling, weighted loss, low learning rates, etc. Also, my PNG images are scaled between [0, 1]. I have uploaded the training screenshot <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3791541%2Fda6c0e20ef49e202c0419e0252be8e34%2Fkaggle_problem.png?generation=1673582928792333&amp;alt=media\" alt=\"Training\"></p>\n<p>Please provide me with some tips.<br>\nThank you in advance</p>",
  "messages": [
    {
      "id": 2105447,
      "postDate": "2023-01-18T14:09:54.157Z",
      "content": "<p>On increasing the batch size the aforementioned problem got fixed</p>",
      "rawMarkdown": "On increasing the batch size the aforementioned problem got fixed",
      "votes": 1
    },
    {
      "id": 2097906,
      "postDate": "2023-01-13T04:09:27.663Z",
      "content": "<p>I am using the EfficientNetB4 model, and training PNG images are of dimension 1024x512. I have tried upsampling, weighted loss, low learning rates, etc. Also, my PNG images are scaled between [0, 1]. I have uploaded the training screenshot <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3791541%2Fda6c0e20ef49e202c0419e0252be8e34%2Fkaggle_problem.png?generation=1673582928792333&amp;alt=media\" alt=\"Training\"></p>\n<p>Please provide me with some tips.<br>\nThank you in advance</p>",
      "rawMarkdown": "I am using the EfficientNetB4 model, and training PNG images are of dimension 1024x512. I have tried upsampling, weighted loss, low learning rates, etc. Also, my PNG images are scaled between [0, 1]. I have uploaded the training screenshot ![Training](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3791541%2Fda6c0e20ef49e202c0419e0252be8e34%2Fkaggle_problem.png?generation=1673582928792333&alt=media)\n\nPlease provide me with some tips.\nThank you in advance",
      "votes": 2
    },
    {
      "id": 2100081,
      "postDate": "2023-01-14T23:30:14.420Z",
      "content": "<p>It looks to me like your model is overfitting.. What kind of validation strategy do you use?</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "It looks to me like your model is overfitting.. What kind of validation strategy do you use?\n\nThe Devastator.\n",
      "replies": [
        {
          "id": 2100917,
          "postDate": "2023-01-15T13:47:00.793Z",
          "content": "<p>Created 5 folds of the dataset by using Stratified Group K-Fold, and thereafter 4 folds are taken into the training dataset and the remaining fold into the validation dataset. I used the same strategy as defined in the following notebook, <br>\n<a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train</a></p>",
          "rawMarkdown": "Created 5 folds of the dataset by using Stratified Group K-Fold, and thereafter 4 folds are taken into the training dataset and the remaining fold into the validation dataset. I used the same strategy as defined in the following notebook, \n[https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train)"
        }
      ]
    },
    {
      "id": 2097979,
      "postDate": "2023-01-13T05:55:13.343Z",
      "content": "<p>Leak (trianing auc is very high for this type of problem - I have not seen such great score in research papers)?<br>\nCheck your DS for sure.</p>",
      "rawMarkdown": "Leak (trianing auc is very high for this type of problem - I have not seen such great score in research papers)?\nCheck your DS for sure.",
      "replies": [
        {
          "id": 2097992,
          "postDate": "2023-01-13T06:07:36.440Z",
          "content": "<p>Does it mean I should try a different dataset to train than the one on which I am currently training?</p>",
          "rawMarkdown": "Does it mean I should try a different dataset to train than the one on which I am currently training?",
          "replies": [
            {
              "id": 2098003,
              "postDate": "2023-01-13T06:18:37.963Z",
              "content": "<p>I don't know how you split the dataset between test and validation. There may be a problem here that part of the data for one patient is both in training and validation.</p>",
              "rawMarkdown": "I don't know how you split the dataset between test and validation. There may be a problem here that part of the data for one patient is both in training and validation."
            },
            {
              "id": 2098013,
              "postDate": "2023-01-13T06:30:54.210Z",
              "content": "<p>I have split my dataset by using Stratified Group K-Fold and used the concept from the following Kaggle Notebook.<br>\n<a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train</a></p>",
              "rawMarkdown": "I have split my dataset by using Stratified Group K-Fold and used the concept from the following Kaggle Notebook.\n[https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train)"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2105447,
      "author_name": "Pranay Barkataki",
      "author_url": "",
      "post_date": "2023-01-18T14:09:54.157000",
      "content": "<p>On increasing the batch size the aforementioned problem got fixed</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2100081,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2023-01-14T23:30:14.420000",
      "content": "<p>It looks to me like your model is overfitting.. What kind of validation strategy do you use?</p>\n<p>The Devastator.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2100917,
          "author_name": "Pranay Barkataki",
          "author_url": "",
          "post_date": "2023-01-15T13:47:00.793000",
          "content": "<p>Created 5 folds of the dataset by using Stratified Group K-Fold, and thereafter 4 folds are taken into the training dataset and the remaining fold into the validation dataset. I used the same strategy as defined in the following notebook, <br>\n<a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2097979,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-01-13T05:55:13.343000",
      "content": "<p>Leak (trianing auc is very high for this type of problem - I have not seen such great score in research papers)?<br>\nCheck your DS for sure.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2097992,
          "author_name": "Pranay Barkataki",
          "author_url": "",
          "post_date": "2023-01-13T06:07:36.440000",
          "content": "<p>Does it mean I should try a different dataset to train than the one on which I am currently training?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2098003,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-13T06:18:37.963000",
              "content": "<p>I don't know how you split the dataset between test and validation. There may be a problem here that part of the data for one patient is both in training and validation.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2098013,
              "author_name": "Pranay Barkataki",
              "author_url": "",
              "post_date": "2023-01-13T06:30:54.210000",
              "content": "<p>I have split my dataset by using Stratified Group K-Fold and used the concept from the following Kaggle Notebook.<br>\n<a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2105447": "On increasing the batch size the aforementioned problem got fixed",
    "2097906": "I am using the EfficientNetB4 model, and training PNG images are of dimension 1024x512. I have tried upsampling, weighted loss, low learning rates, etc. Also, my PNG images are scaled between [0, 1]. I have uploaded the training screenshot ![Training](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3791541%2Fda6c0e20ef49e202c0419e0252be8e34%2Fkaggle_problem.png?generation=1673582928792333&alt=media)\n\nPlease provide me with some tips.\nThank you in advance",
    "2100081": "It looks to me like your model is overfitting.. What kind of validation strategy do you use?\n\nThe Devastator.\n",
    "2097979": "Leak (trianing auc is very high for this type of problem - I have not seen such great score in research papers)?\nCheck your DS for sure."
  }
}