{
  "id": 464900,
  "title": "Learning Recap",
  "url": "/competitions/UBC-OCEAN/discussion/464900",
  "author_name": "chemdatafarmer",
  "post_date": "2024-01-02T05:11:51.817000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n<p>I'm writing this mostly as a way to recap the things I've learned during this competition. In the end, I'm not particularly happy with the models I came up with, but I'm pleased with the things I learned this time around and that's really why I'm here. WSIs brought a bunch of challenges I had never encountered before and I'm really glad I participated in this competition and had the chance to learn about them. Learning about pyvips and how useful it is for these high res images was pretty neat and I'm grateful for all of the awesome notebooks that used this library. While it was a little frustrating figuring out how to get the tiling done quickly enough to make viable submissions, it was also a lot of fun learning about the different multiprocessing libraries in Python and the differences between threading and multiprocessing. It took some experimentation, but I also got much more comfortable using a tf.data.Dataset workflow and writing mapping functions that could handle tiling. Thanks to the folks who put together the keras-cv starter notebook, this was my first experience using keras-cv and it was really enjoyable to use. I burned through a lot of GPU hours early on experimenting with things, which led me to learn about docker containers and how to use docker in WSL2 to get the docker images Kaggle provides to us up and running as containers on my computer so next time I'll be a lot less compute constrained for experimentation (this was a bit of a rabbit hole, but it was an fun one). That said, I also learned some tips and tricks for being more efficient with my GPU time such as pre-calculating the tensors output from frozen pretrained models and training the classifier layers on them (though I did find this reduced the amount of runtime augmentation I could do and would love to hear if people have any tips for this). </p>\n<p>I think the main thing that I struggled with in this competition, besides running out of time after figuring out the things mentioned above, was overfitting. While the WSIs each contained a lot of data, I wasn't able to find a way to tighten the gap between my training loss and my validation loss, particularly when splitting on image_id. I suspect there are some normalization tricks I'm not using that can help here. Random splits of tiles weren't particularly diagnostic, even with really heavy augmentations during training, as I noticed the models tended to fit to something about the image and not about the phenotype displayed by the particular cancer strain on the slide. Most of the work I was able to do focused on training a tile level classifier and I think it would have been really interesting to spend more time working on training a slide level classifier that uses tiles, rather than a tile level classifier where I aggregate votes. It would have also been interesting to get to spend more time on the outlier detection portion of this problem, but without relatively good classifiers I didn't end up spending much time there. I stayed away from thumbnail based approaches as there were already really good notebooks out there for them when I started the competition and I was curious how tiling approaches would do in comparison.</p>\n<p>I'm really looking forward to seeing posts from folks that detail how they approached this competition and what their final solutions ended up looking like. It's one of the best parts of competing in Kaggle competitions and I always get an interesting perspective on the competition going through how others approached the problem.</p>\n<p>Feel free to use this thread to share things you learned as well :).</p>",
  "messages": [
    {
      "id": 2583153,
      "postDate": "2024-01-02T05:11:51.817Z",
      "content": "<p>Hi Everyone,</p>\n<p>I'm writing this mostly as a way to recap the things I've learned during this competition. In the end, I'm not particularly happy with the models I came up with, but I'm pleased with the things I learned this time around and that's really why I'm here. WSIs brought a bunch of challenges I had never encountered before and I'm really glad I participated in this competition and had the chance to learn about them. Learning about pyvips and how useful it is for these high res images was pretty neat and I'm grateful for all of the awesome notebooks that used this library. While it was a little frustrating figuring out how to get the tiling done quickly enough to make viable submissions, it was also a lot of fun learning about the different multiprocessing libraries in Python and the differences between threading and multiprocessing. It took some experimentation, but I also got much more comfortable using a tf.data.Dataset workflow and writing mapping functions that could handle tiling. Thanks to the folks who put together the keras-cv starter notebook, this was my first experience using keras-cv and it was really enjoyable to use. I burned through a lot of GPU hours early on experimenting with things, which led me to learn about docker containers and how to use docker in WSL2 to get the docker images Kaggle provides to us up and running as containers on my computer so next time I'll be a lot less compute constrained for experimentation (this was a bit of a rabbit hole, but it was an fun one). That said, I also learned some tips and tricks for being more efficient with my GPU time such as pre-calculating the tensors output from frozen pretrained models and training the classifier layers on them (though I did find this reduced the amount of runtime augmentation I could do and would love to hear if people have any tips for this). </p>\n<p>I think the main thing that I struggled with in this competition, besides running out of time after figuring out the things mentioned above, was overfitting. While the WSIs each contained a lot of data, I wasn't able to find a way to tighten the gap between my training loss and my validation loss, particularly when splitting on image_id. I suspect there are some normalization tricks I'm not using that can help here. Random splits of tiles weren't particularly diagnostic, even with really heavy augmentations during training, as I noticed the models tended to fit to something about the image and not about the phenotype displayed by the particular cancer strain on the slide. Most of the work I was able to do focused on training a tile level classifier and I think it would have been really interesting to spend more time working on training a slide level classifier that uses tiles, rather than a tile level classifier where I aggregate votes. It would have also been interesting to get to spend more time on the outlier detection portion of this problem, but without relatively good classifiers I didn't end up spending much time there. I stayed away from thumbnail based approaches as there were already really good notebooks out there for them when I started the competition and I was curious how tiling approaches would do in comparison.</p>\n<p>I'm really looking forward to seeing posts from folks that detail how they approached this competition and what their final solutions ended up looking like. It's one of the best parts of competing in Kaggle competitions and I always get an interesting perspective on the competition going through how others approached the problem.</p>\n<p>Feel free to use this thread to share things you learned as well :).</p>",
      "rawMarkdown": "Hi Everyone,\n\nI'm writing this mostly as a way to recap the things I've learned during this competition. In the end, I'm not particularly happy with the models I came up with, but I'm pleased with the things I learned this time around and that's really why I'm here. WSIs brought a bunch of challenges I had never encountered before and I'm really glad I participated in this competition and had the chance to learn about them. Learning about pyvips and how useful it is for these high res images was pretty neat and I'm grateful for all of the awesome notebooks that used this library. While it was a little frustrating figuring out how to get the tiling done quickly enough to make viable submissions, it was also a lot of fun learning about the different multiprocessing libraries in Python and the differences between threading and multiprocessing. It took some experimentation, but I also got much more comfortable using a tf.data.Dataset workflow and writing mapping functions that could handle tiling. Thanks to the folks who put together the keras-cv starter notebook, this was my first experience using keras-cv and it was really enjoyable to use. I burned through a lot of GPU hours early on experimenting with things, which led me to learn about docker containers and how to use docker in WSL2 to get the docker images Kaggle provides to us up and running as containers on my computer so next time I'll be a lot less compute constrained for experimentation (this was a bit of a rabbit hole, but it was an fun one). That said, I also learned some tips and tricks for being more efficient with my GPU time such as pre-calculating the tensors output from frozen pretrained models and training the classifier layers on them (though I did find this reduced the amount of runtime augmentation I could do and would love to hear if people have any tips for this). \n\nI think the main thing that I struggled with in this competition, besides running out of time after figuring out the things mentioned above, was overfitting. While the WSIs each contained a lot of data, I wasn't able to find a way to tighten the gap between my training loss and my validation loss, particularly when splitting on image_id. I suspect there are some normalization tricks I'm not using that can help here. Random splits of tiles weren't particularly diagnostic, even with really heavy augmentations during training, as I noticed the models tended to fit to something about the image and not about the phenotype displayed by the particular cancer strain on the slide. Most of the work I was able to do focused on training a tile level classifier and I think it would have been really interesting to spend more time working on training a slide level classifier that uses tiles, rather than a tile level classifier where I aggregate votes. It would have also been interesting to get to spend more time on the outlier detection portion of this problem, but without relatively good classifiers I didn't end up spending much time there. I stayed away from thumbnail based approaches as there were already really good notebooks out there for them when I started the competition and I was curious how tiling approaches would do in comparison.\n\nI'm really looking forward to seeing posts from folks that detail how they approached this competition and what their final solutions ended up looking like. It's one of the best parts of competing in Kaggle competitions and I always get an interesting perspective on the competition going through how others approached the problem.\n\nFeel free to use this thread to share things you learned as well :).",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2583153": "Hi Everyone,\n\nI'm writing this mostly as a way to recap the things I've learned during this competition. In the end, I'm not particularly happy with the models I came up with, but I'm pleased with the things I learned this time around and that's really why I'm here. WSIs brought a bunch of challenges I had never encountered before and I'm really glad I participated in this competition and had the chance to learn about them. Learning about pyvips and how useful it is for these high res images was pretty neat and I'm grateful for all of the awesome notebooks that used this library. While it was a little frustrating figuring out how to get the tiling done quickly enough to make viable submissions, it was also a lot of fun learning about the different multiprocessing libraries in Python and the differences between threading and multiprocessing. It took some experimentation, but I also got much more comfortable using a tf.data.Dataset workflow and writing mapping functions that could handle tiling. Thanks to the folks who put together the keras-cv starter notebook, this was my first experience using keras-cv and it was really enjoyable to use. I burned through a lot of GPU hours early on experimenting with things, which led me to learn about docker containers and how to use docker in WSL2 to get the docker images Kaggle provides to us up and running as containers on my computer so next time I'll be a lot less compute constrained for experimentation (this was a bit of a rabbit hole, but it was an fun one). That said, I also learned some tips and tricks for being more efficient with my GPU time such as pre-calculating the tensors output from frozen pretrained models and training the classifier layers on them (though I did find this reduced the amount of runtime augmentation I could do and would love to hear if people have any tips for this). \n\nI think the main thing that I struggled with in this competition, besides running out of time after figuring out the things mentioned above, was overfitting. While the WSIs each contained a lot of data, I wasn't able to find a way to tighten the gap between my training loss and my validation loss, particularly when splitting on image_id. I suspect there are some normalization tricks I'm not using that can help here. Random splits of tiles weren't particularly diagnostic, even with really heavy augmentations during training, as I noticed the models tended to fit to something about the image and not about the phenotype displayed by the particular cancer strain on the slide. Most of the work I was able to do focused on training a tile level classifier and I think it would have been really interesting to spend more time working on training a slide level classifier that uses tiles, rather than a tile level classifier where I aggregate votes. It would have also been interesting to get to spend more time on the outlier detection portion of this problem, but without relatively good classifiers I didn't end up spending much time there. I stayed away from thumbnail based approaches as there were already really good notebooks out there for them when I started the competition and I was curious how tiling approaches would do in comparison.\n\nI'm really looking forward to seeing posts from folks that detail how they approached this competition and what their final solutions ended up looking like. It's one of the best parts of competing in Kaggle competitions and I always get an interesting perspective on the competition going through how others approached the problem.\n\nFeel free to use this thread to share things you learned as well :)."
  }
}