{
  "id": 374504,
  "title": "Could I use 2 normalize in my training pipeline? ",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/374504",
  "author_name": "happymentee",
  "post_date": "2022-12-27T16:22:41.998000",
  "votes": 0,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi! I'm a complete beginner at the Kaggle competition, and I'm confused about choosing the right transformation method:</p>\n<p>I have seen 2 types of normalization:</p>\n<ol>\n<li>to subtract the mean of the training set and divide by the standard deviation</li>\n<li>scales the data to the range [0, 1] by subtracting min and then dividing by (max - min) <br>\nIn this convert dataset from theoviel <a href=\"https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\" target=\"_blank\">Dicom -&gt; Resized PNG/JPG\n</a> , I noticed that the original image had been normalized by type 2 (I listed above).</li>\n</ol>\n<p>So my question is: If I use this converted dataset, could I use normalization type 1(I listed above) in my custom transform function in my training code?</p>\n<p>I guess there is no robust answer, but I hope to hear about others' experiences.<br>\nI really appreciate any help you can provide. </p>",
  "messages": [
    {
      "id": 2077903,
      "postDate": "2022-12-28T00:06:32.257Z",
      "content": "<p>Hey! You can normalize data however you see fit (in fact, that is something worth experimenting with! 🙂)</p>\n<p>You just need to make sure that you normalize the data in the same way in training as in test. As in, you can read in the data from whatever format, you just need to make sure that once the data reaches your model, that it has been preprocessed in the same way (that is, you can construct a custom preprocessing pipeline that sits between reading the data and it being shown to your model for training/inference)</p>\n<p>All the best, hope this helps! 🙌</p>",
      "rawMarkdown": "Hey! You can normalize data however you see fit (in fact, that is something worth experimenting with! 🙂)\n\nYou just need to make sure that you normalize the data in the same way in training as in test. As in, you can read in the data from whatever format, you just need to make sure that once the data reaches your model, that it has been preprocessed in the same way (that is, you can construct a custom preprocessing pipeline that sits between reading the data and it being shown to your model for training/inference)\n\nAll the best, hope this helps! 🙌",
      "votes": 1
    },
    {
      "id": 2077574,
      "postDate": "2022-12-27T17:55:15.327Z",
      "content": "<p>What you refer to as Type1 is data standardization. You can rescale images by  Type2 first, then calculate mean and std of the normalized dataset, and finally standardize data. It might help the model to generalize better.</p>",
      "rawMarkdown": "What you refer to as Type1 is data standardization. You can rescale images by  Type2 first, then calculate mean and std of the normalized dataset, and finally standardize data. It might help the model to generalize better.",
      "votes": 1
    },
    {
      "id": 2077489,
      "postDate": "2022-12-27T16:22:41.997Z",
      "content": "<p>Hi! I'm a complete beginner at the Kaggle competition, and I'm confused about choosing the right transformation method:</p>\n<p>I have seen 2 types of normalization:</p>\n<ol>\n<li>to subtract the mean of the training set and divide by the standard deviation</li>\n<li>scales the data to the range [0, 1] by subtracting min and then dividing by (max - min) <br>\nIn this convert dataset from theoviel <a href=\"https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\" target=\"_blank\">Dicom -&gt; Resized PNG/JPG\n</a> , I noticed that the original image had been normalized by type 2 (I listed above).</li>\n</ol>\n<p>So my question is: If I use this converted dataset, could I use normalization type 1(I listed above) in my custom transform function in my training code?</p>\n<p>I guess there is no robust answer, but I hope to hear about others' experiences.<br>\nI really appreciate any help you can provide. </p>",
      "rawMarkdown": "Hi! I'm a complete beginner at the Kaggle competition, and I'm confused about choosing the right transformation method:\n\nI have seen 2 types of normalization:\n\n1. to subtract the mean of the training set and divide by the standard deviation\n2. scales the data to the range [0, 1] by subtracting min and then dividing by (max - min) \nIn this convert dataset from theoviel [Dicom -> Resized PNG/JPG\n](https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg) , I noticed that the original image had been normalized by type 2 (I listed above).\n\nSo my question is: If I use this converted dataset, could I use normalization type 1(I listed above) in my custom transform function in my training code?\n\nI guess there is no robust answer, but I hope to hear about others' experiences.\nI really appreciate any help you can provide. "
    }
  ],
  "comments": [
    {
      "id": 2077903,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-28T00:06:32.257000",
      "content": "<p>Hey! You can normalize data however you see fit (in fact, that is something worth experimenting with! 🙂)</p>\n<p>You just need to make sure that you normalize the data in the same way in training as in test. As in, you can read in the data from whatever format, you just need to make sure that once the data reaches your model, that it has been preprocessed in the same way (that is, you can construct a custom preprocessing pipeline that sits between reading the data and it being shown to your model for training/inference)</p>\n<p>All the best, hope this helps! 🙌</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2077574,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2022-12-27T17:55:15.327000",
      "content": "<p>What you refer to as Type1 is data standardization. You can rescale images by  Type2 first, then calculate mean and std of the normalized dataset, and finally standardize data. It might help the model to generalize better.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2077903": "Hey! You can normalize data however you see fit (in fact, that is something worth experimenting with! 🙂)\n\nYou just need to make sure that you normalize the data in the same way in training as in test. As in, you can read in the data from whatever format, you just need to make sure that once the data reaches your model, that it has been preprocessed in the same way (that is, you can construct a custom preprocessing pipeline that sits between reading the data and it being shown to your model for training/inference)\n\nAll the best, hope this helps! 🙌",
    "2077574": "What you refer to as Type1 is data standardization. You can rescale images by  Type2 first, then calculate mean and std of the normalized dataset, and finally standardize data. It might help the model to generalize better.",
    "2077489": "Hi! I'm a complete beginner at the Kaggle competition, and I'm confused about choosing the right transformation method:\n\nI have seen 2 types of normalization:\n\n1. to subtract the mean of the training set and divide by the standard deviation\n2. scales the data to the range [0, 1] by subtracting min and then dividing by (max - min) \nIn this convert dataset from theoviel [Dicom -> Resized PNG/JPG\n](https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg) , I noticed that the original image had been normalized by type 2 (I listed above).\n\nSo my question is: If I use this converted dataset, could I use normalization type 1(I listed above) in my custom transform function in my training code?\n\nI guess there is no robust answer, but I hope to hear about others' experiences.\nI really appreciate any help you can provide. "
  }
}