{
  "id": 341936,
  "title": "Question about image augmentation",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/341936",
  "author_name": "yqz",
  "post_date": "2022-08-04T23:29:25.488000",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This is my first computer vision competition, and I've taken some courses on it and have a basic idea of how CNN's work. However, I'm currently not quite sure what's considered best practices when it comes to image augmentation. I know that one of the most popular libraries is Albumentations, and I've even met one of the co-founders in person. My difficulty though is when it comes to how to implement it within a typical machine learning development/deployment pipeline. I posted this on their <a href=\"https://github.com/albumentations-team/albumentations/issues/1249\" target=\"_blank\">GitHub</a>, hopefully someone here can answer this question for me also.</p>\n<blockquote>\n  <p>I have some basic understanding of machine learning, but when I saw how Albumentations works I got kind of confused. I noticed that with the A.Compose() function, we can augment an image with for example some given probability. Then for image_1, we can get either image_1a, image_1b, or image_1c. However, this leaves our training dataset having the exact same size. I imagined before using Albumentations that we would instead want to have all of images 1a, 1b, and 1c in our training set to help expand the size of our training data.</p>\n  <p>Also, I would imagine that we wouldn't want to have any of the same images 1a, 1b, or 1c in both the training and validation dataset.</p>\n  <p>I feel like the current system as is doesn't well suit the my idea of having all the 1a, 1b, and 1c in the training set. Given that the A.Compose() is utilized within the PyTorch Dataset class, then it really only has the chance to happen once per image. I would imagine that we would need to call different A.Compose() multiple times per image to get the end result. Is there some workaround for this?</p>\n</blockquote>\n<p>Any help would be greatly appreciated.</p>",
  "messages": [
    {
      "id": 1885134,
      "postDate": "2022-08-04T23:29:25.490Z",
      "content": "<p>This is my first computer vision competition, and I've taken some courses on it and have a basic idea of how CNN's work. However, I'm currently not quite sure what's considered best practices when it comes to image augmentation. I know that one of the most popular libraries is Albumentations, and I've even met one of the co-founders in person. My difficulty though is when it comes to how to implement it within a typical machine learning development/deployment pipeline. I posted this on their <a href=\"https://github.com/albumentations-team/albumentations/issues/1249\" target=\"_blank\">GitHub</a>, hopefully someone here can answer this question for me also.</p>\n<blockquote>\n  <p>I have some basic understanding of machine learning, but when I saw how Albumentations works I got kind of confused. I noticed that with the A.Compose() function, we can augment an image with for example some given probability. Then for image_1, we can get either image_1a, image_1b, or image_1c. However, this leaves our training dataset having the exact same size. I imagined before using Albumentations that we would instead want to have all of images 1a, 1b, and 1c in our training set to help expand the size of our training data.</p>\n  <p>Also, I would imagine that we wouldn't want to have any of the same images 1a, 1b, or 1c in both the training and validation dataset.</p>\n  <p>I feel like the current system as is doesn't well suit the my idea of having all the 1a, 1b, and 1c in the training set. Given that the A.Compose() is utilized within the PyTorch Dataset class, then it really only has the chance to happen once per image. I would imagine that we would need to call different A.Compose() multiple times per image to get the end result. Is there some workaround for this?</p>\n</blockquote>\n<p>Any help would be greatly appreciated.</p>",
      "rawMarkdown": "This is my first computer vision competition, and I've taken some courses on it and have a basic idea of how CNN's work. However, I'm currently not quite sure what's considered best practices when it comes to image augmentation. I know that one of the most popular libraries is Albumentations, and I've even met one of the co-founders in person. My difficulty though is when it comes to how to implement it within a typical machine learning development/deployment pipeline. I posted this on their [GitHub](https://github.com/albumentations-team/albumentations/issues/1249), hopefully someone here can answer this question for me also.\n\n> I have some basic understanding of machine learning, but when I saw how Albumentations works I got kind of confused. I noticed that with the A.Compose() function, we can augment an image with for example some given probability. Then for image_1, we can get either image_1a, image_1b, or image_1c. However, this leaves our training dataset having the exact same size. I imagined before using Albumentations that we would instead want to have all of images 1a, 1b, and 1c in our training set to help expand the size of our training data.\n\n> Also, I would imagine that we wouldn't want to have any of the same images 1a, 1b, or 1c in both the training and validation dataset.\n\n> I feel like the current system as is doesn't well suit the my idea of having all the 1a, 1b, and 1c in the training set. Given that the A.Compose() is utilized within the PyTorch Dataset class, then it really only has the chance to happen once per image. I would imagine that we would need to call different A.Compose() multiple times per image to get the end result. Is there some workaround for this?\n\nAny help would be greatly appreciated.",
      "votes": 4
    },
    {
      "id": 1887776,
      "postDate": "2022-08-07T03:00:06.047Z",
      "content": "<p>I see your point. To start with, augmentations, as well as any methods that aims to increase the amount of data that we have, also comes with their own bias-variance tradeoff. Swing too far in either direction in this tradeoff will give you poor results on the modelling end. Augmentations generally result in <em>less</em> variance and more bias. If you take a pair of random images and compare the pair's variance with another pair of images, but this one contains one original image and a rotated version of that image, it is very likely that the variance of the second pair will be smaller. To relate this to your question, you should (in normal settings, but you can of course change the pipeline to accommodate otherwise, but it's not recommended) only have 1 transformation per pass through the entire dataset, so that the extra bias that you're introducing does not influence the model training process too much.</p>",
      "rawMarkdown": "I see your point. To start with, augmentations, as well as any methods that aims to increase the amount of data that we have, also comes with their own bias-variance tradeoff. Swing too far in either direction in this tradeoff will give you poor results on the modelling end. Augmentations generally result in *less* variance and more bias. If you take a pair of random images and compare the pair's variance with another pair of images, but this one contains one original image and a rotated version of that image, it is very likely that the variance of the second pair will be smaller. To relate this to your question, you should (in normal settings, but you can of course change the pipeline to accommodate otherwise, but it's not recommended) only have 1 transformation per pass through the entire dataset, so that the extra bias that you're introducing does not influence the model training process too much.",
      "votes": 2,
      "replies": [
        {
          "id": 1888598,
          "postDate": "2022-08-07T17:25:29.343Z",
          "content": "<p>Thanks for the explanation in terms of bias/variance tradeoff, it makes sense to me now.</p>",
          "rawMarkdown": "Thanks for the explanation in terms of bias/variance tradeoff, it makes sense to me now."
        }
      ]
    },
    {
      "id": 1886998,
      "postDate": "2022-08-06T10:54:45.450Z",
      "content": "<p>Compose is a sequence of transforms, each transform is randomly executed when image is passed through. If you pass same image multiple times through transform, you'll get different output. For example random flip with p=.5 means output of transform half of the times will produce flipped image. So, when you run multiple epochs during training, you augment your data every time before passing it into model.</p>",
      "rawMarkdown": "Compose is a sequence of transforms, each transform is randomly executed when image is passed through. If you pass same image multiple times through transform, you'll get different output. For example random flip with p=.5 means output of transform half of the times will produce flipped image. So, when you run multiple epochs during training, you augment your data every time before passing it into model.",
      "votes": 2,
      "replies": [
        {
          "id": 1888600,
          "postDate": "2022-08-07T17:25:57.860Z",
          "content": "<p>Thanks for highlighting that the augmentation happens each epoch, so for each time it happens we can get a slightly different image depending on the number of epochs.</p>",
          "rawMarkdown": "Thanks for highlighting that the augmentation happens each epoch, so for each time it happens we can get a slightly different image depending on the number of epochs."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1887776,
      "author_name": "Phan Nguyen",
      "author_url": "",
      "post_date": "2022-08-07T03:00:06.047000",
      "content": "<p>I see your point. To start with, augmentations, as well as any methods that aims to increase the amount of data that we have, also comes with their own bias-variance tradeoff. Swing too far in either direction in this tradeoff will give you poor results on the modelling end. Augmentations generally result in <em>less</em> variance and more bias. If you take a pair of random images and compare the pair's variance with another pair of images, but this one contains one original image and a rotated version of that image, it is very likely that the variance of the second pair will be smaller. To relate this to your question, you should (in normal settings, but you can of course change the pipeline to accommodate otherwise, but it's not recommended) only have 1 transformation per pass through the entire dataset, so that the extra bias that you're introducing does not influence the model training process too much.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1888598,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-07T17:25:29.343000",
          "content": "<p>Thanks for the explanation in terms of bias/variance tradeoff, it makes sense to me now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1886998,
      "author_name": "Sergey",
      "author_url": "",
      "post_date": "2022-08-06T10:54:45.450000",
      "content": "<p>Compose is a sequence of transforms, each transform is randomly executed when image is passed through. If you pass same image multiple times through transform, you'll get different output. For example random flip with p=.5 means output of transform half of the times will produce flipped image. So, when you run multiple epochs during training, you augment your data every time before passing it into model.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1888600,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-07T17:25:57.860000",
          "content": "<p>Thanks for highlighting that the augmentation happens each epoch, so for each time it happens we can get a slightly different image depending on the number of epochs.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1885134": "This is my first computer vision competition, and I've taken some courses on it and have a basic idea of how CNN's work. However, I'm currently not quite sure what's considered best practices when it comes to image augmentation. I know that one of the most popular libraries is Albumentations, and I've even met one of the co-founders in person. My difficulty though is when it comes to how to implement it within a typical machine learning development/deployment pipeline. I posted this on their [GitHub](https://github.com/albumentations-team/albumentations/issues/1249), hopefully someone here can answer this question for me also.\n\n> I have some basic understanding of machine learning, but when I saw how Albumentations works I got kind of confused. I noticed that with the A.Compose() function, we can augment an image with for example some given probability. Then for image_1, we can get either image_1a, image_1b, or image_1c. However, this leaves our training dataset having the exact same size. I imagined before using Albumentations that we would instead want to have all of images 1a, 1b, and 1c in our training set to help expand the size of our training data.\n\n> Also, I would imagine that we wouldn't want to have any of the same images 1a, 1b, or 1c in both the training and validation dataset.\n\n> I feel like the current system as is doesn't well suit the my idea of having all the 1a, 1b, and 1c in the training set. Given that the A.Compose() is utilized within the PyTorch Dataset class, then it really only has the chance to happen once per image. I would imagine that we would need to call different A.Compose() multiple times per image to get the end result. Is there some workaround for this?\n\nAny help would be greatly appreciated.",
    "1887776": "I see your point. To start with, augmentations, as well as any methods that aims to increase the amount of data that we have, also comes with their own bias-variance tradeoff. Swing too far in either direction in this tradeoff will give you poor results on the modelling end. Augmentations generally result in *less* variance and more bias. If you take a pair of random images and compare the pair's variance with another pair of images, but this one contains one original image and a rotated version of that image, it is very likely that the variance of the second pair will be smaller. To relate this to your question, you should (in normal settings, but you can of course change the pipeline to accommodate otherwise, but it's not recommended) only have 1 transformation per pass through the entire dataset, so that the extra bias that you're introducing does not influence the model training process too much.",
    "1886998": "Compose is a sequence of transforms, each transform is randomly executed when image is passed through. If you pass same image multiple times through transform, you'll get different output. For example random flip with p=.5 means output of transform half of the times will produce flipped image. So, when you run multiple epochs during training, you augment your data every time before passing it into model."
  }
}