{
  "id": 465030,
  "title": "My Competition Recap",
  "url": "/competitions/UBC-OCEAN/discussion/465030",
  "author_name": "Connor",
  "post_date": "2024-01-02T16:23:24.547000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This competition took my interest from the beginning and although I faced many challenges when building a solution, I took a lot away from this experience.</p>\n<p>The first of many challenges was processing the full WSI.  While many solutions used the thumbnails, I opted to attempt to efficiently tile the WSI and train the TransMIL transformer architecture on the patches as well as a CNN that aggregates the predictions.  This brought the first challenge, tiling the image within the given time and only tiling important information.  I took a two step approach to this challenge.  The first step was to utilize some of the functions in the CV2 library to find an array of coordinates that represent an outline of the non-background portion of the WSI.  From these coordinates, patch coordinates are created.  The second step involves actually extracting the patches from the WSI.  For the TransMIL model I created the feature embeddings using a pretrained ResNet-50 model and created an (B, N, 1028) tensor where B is the batch size, N is the number of patches, and 1028 is the feature embeddings size.  For the CNN solution, I got the prediction once the tile was created and then removed the tile from memory.</p>\n<p>The next challenge was the memory constraint.  Both of the solutions mentioned in the previous paragraph avoid writing images to memory due to the large number of patches.  In addition to the memory challenge, there was a computational time challenge due to the 12 hr maximum runtime.  Processing these images takes a long time so I had to find a way to be optimal when tiling the image.  I started this by running step 1 of finding the outline of the important information using the thumbnail, and then transforming the pixel coordinates based on the downscaling ratio between the WSI and the thumbnail. This alone significantly sped up the runtime, allowing for the coordinates to be calculated for the entire training set in 20 minutes.  Next, I utilized the GPU pyvips library and the python-openslide library to efficiently read, process, and tile the WSIs alongwith some parallel programming.  Using this, I was able to get the runtime for inference on the TRAINING set done in 12.5 hours and because the test set is smaller in size (more images but ~200 less GB), I was able to get my baseline model (not fine-tuned) to run within the allowed time.</p>\n<p>Finally, while I believe the TransMIL model has more potential for high results, I worry that it will have a difficult time generalizing to data from other hospitals.  Because the data loading time takes very long (12 hours to get the feature embeddings for every train WSI), augmenting the data to expand the size of the training set would have taken many days or even weeks which I was unable to commit the time to do.  For the CNN model I was able to augment each 512x512 patch to allow the model to generalize better and not overfit to this specific hospital's data.  </p>\n<p>My biggest contribution to this competition was my algorithm that optimally tiled each WSI in the test set within the allowed computation time.  I saw early on that there were many notebooks that used tiling solutions but were only able to retrieve ~30-50 patches.  I decided to go deeper and try to find a way to get every patch possible from every sample within the allowed time.  I believe that there must have been other competitors that were able to do the same thing but to me this was my biggest achievement during this competition and I will be able to see later tonight if it paid off.</p>\n<p>At the time I am making this post my leaderboard submission is still my baseline submission and I am hopeful that either my TransMIL fine-tuned model or CNN solution will provide strong results.  Regardless of where I end up I am very interested to see other competitor's solutions and I plan to share my highly optimized code for tiling the entire image for each WSI after the competition.  Good luck to all competitors and I am very excited to see the final results!  Thank you to the competition owners for putting this competition out and for every code and discussion post along the way that has taught me so much more about data science.</p>",
  "messages": [
    {
      "id": 2584007,
      "postDate": "2024-01-02T16:23:24.547Z",
      "content": "<p>This competition took my interest from the beginning and although I faced many challenges when building a solution, I took a lot away from this experience.</p>\n<p>The first of many challenges was processing the full WSI.  While many solutions used the thumbnails, I opted to attempt to efficiently tile the WSI and train the TransMIL transformer architecture on the patches as well as a CNN that aggregates the predictions.  This brought the first challenge, tiling the image within the given time and only tiling important information.  I took a two step approach to this challenge.  The first step was to utilize some of the functions in the CV2 library to find an array of coordinates that represent an outline of the non-background portion of the WSI.  From these coordinates, patch coordinates are created.  The second step involves actually extracting the patches from the WSI.  For the TransMIL model I created the feature embeddings using a pretrained ResNet-50 model and created an (B, N, 1028) tensor where B is the batch size, N is the number of patches, and 1028 is the feature embeddings size.  For the CNN solution, I got the prediction once the tile was created and then removed the tile from memory.</p>\n<p>The next challenge was the memory constraint.  Both of the solutions mentioned in the previous paragraph avoid writing images to memory due to the large number of patches.  In addition to the memory challenge, there was a computational time challenge due to the 12 hr maximum runtime.  Processing these images takes a long time so I had to find a way to be optimal when tiling the image.  I started this by running step 1 of finding the outline of the important information using the thumbnail, and then transforming the pixel coordinates based on the downscaling ratio between the WSI and the thumbnail. This alone significantly sped up the runtime, allowing for the coordinates to be calculated for the entire training set in 20 minutes.  Next, I utilized the GPU pyvips library and the python-openslide library to efficiently read, process, and tile the WSIs alongwith some parallel programming.  Using this, I was able to get the runtime for inference on the TRAINING set done in 12.5 hours and because the test set is smaller in size (more images but ~200 less GB), I was able to get my baseline model (not fine-tuned) to run within the allowed time.</p>\n<p>Finally, while I believe the TransMIL model has more potential for high results, I worry that it will have a difficult time generalizing to data from other hospitals.  Because the data loading time takes very long (12 hours to get the feature embeddings for every train WSI), augmenting the data to expand the size of the training set would have taken many days or even weeks which I was unable to commit the time to do.  For the CNN model I was able to augment each 512x512 patch to allow the model to generalize better and not overfit to this specific hospital's data.  </p>\n<p>My biggest contribution to this competition was my algorithm that optimally tiled each WSI in the test set within the allowed computation time.  I saw early on that there were many notebooks that used tiling solutions but were only able to retrieve ~30-50 patches.  I decided to go deeper and try to find a way to get every patch possible from every sample within the allowed time.  I believe that there must have been other competitors that were able to do the same thing but to me this was my biggest achievement during this competition and I will be able to see later tonight if it paid off.</p>\n<p>At the time I am making this post my leaderboard submission is still my baseline submission and I am hopeful that either my TransMIL fine-tuned model or CNN solution will provide strong results.  Regardless of where I end up I am very interested to see other competitor's solutions and I plan to share my highly optimized code for tiling the entire image for each WSI after the competition.  Good luck to all competitors and I am very excited to see the final results!  Thank you to the competition owners for putting this competition out and for every code and discussion post along the way that has taught me so much more about data science.</p>",
      "rawMarkdown": "This competition took my interest from the beginning and although I faced many challenges when building a solution, I took a lot away from this experience.\n\nThe first of many challenges was processing the full WSI.  While many solutions used the thumbnails, I opted to attempt to efficiently tile the WSI and train the TransMIL transformer architecture on the patches as well as a CNN that aggregates the predictions.  This brought the first challenge, tiling the image within the given time and only tiling important information.  I took a two step approach to this challenge.  The first step was to utilize some of the functions in the CV2 library to find an array of coordinates that represent an outline of the non-background portion of the WSI.  From these coordinates, patch coordinates are created.  The second step involves actually extracting the patches from the WSI.  For the TransMIL model I created the feature embeddings using a pretrained ResNet-50 model and created an (B, N, 1028) tensor where B is the batch size, N is the number of patches, and 1028 is the feature embeddings size.  For the CNN solution, I got the prediction once the tile was created and then removed the tile from memory.\n\nThe next challenge was the memory constraint.  Both of the solutions mentioned in the previous paragraph avoid writing images to memory due to the large number of patches.  In addition to the memory challenge, there was a computational time challenge due to the 12 hr maximum runtime.  Processing these images takes a long time so I had to find a way to be optimal when tiling the image.  I started this by running step 1 of finding the outline of the important information using the thumbnail, and then transforming the pixel coordinates based on the downscaling ratio between the WSI and the thumbnail. This alone significantly sped up the runtime, allowing for the coordinates to be calculated for the entire training set in 20 minutes.  Next, I utilized the GPU pyvips library and the python-openslide library to efficiently read, process, and tile the WSIs alongwith some parallel programming.  Using this, I was able to get the runtime for inference on the TRAINING set done in 12.5 hours and because the test set is smaller in size (more images but ~200 less GB), I was able to get my baseline model (not fine-tuned) to run within the allowed time.\n\nFinally, while I believe the TransMIL model has more potential for high results, I worry that it will have a difficult time generalizing to data from other hospitals.  Because the data loading time takes very long (12 hours to get the feature embeddings for every train WSI), augmenting the data to expand the size of the training set would have taken many days or even weeks which I was unable to commit the time to do.  For the CNN model I was able to augment each 512x512 patch to allow the model to generalize better and not overfit to this specific hospital's data.  \n\nMy biggest contribution to this competition was my algorithm that optimally tiled each WSI in the test set within the allowed computation time.  I saw early on that there were many notebooks that used tiling solutions but were only able to retrieve ~30-50 patches.  I decided to go deeper and try to find a way to get every patch possible from every sample within the allowed time.  I believe that there must have been other competitors that were able to do the same thing but to me this was my biggest achievement during this competition and I will be able to see later tonight if it paid off.\n\nAt the time I am making this post my leaderboard submission is still my baseline submission and I am hopeful that either my TransMIL fine-tuned model or CNN solution will provide strong results.  Regardless of where I end up I am very interested to see other competitor's solutions and I plan to share my highly optimized code for tiling the entire image for each WSI after the competition.  Good luck to all competitors and I am very excited to see the final results!  Thank you to the competition owners for putting this competition out and for every code and discussion post along the way that has taught me so much more about data science.",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2584007": "This competition took my interest from the beginning and although I faced many challenges when building a solution, I took a lot away from this experience.\n\nThe first of many challenges was processing the full WSI.  While many solutions used the thumbnails, I opted to attempt to efficiently tile the WSI and train the TransMIL transformer architecture on the patches as well as a CNN that aggregates the predictions.  This brought the first challenge, tiling the image within the given time and only tiling important information.  I took a two step approach to this challenge.  The first step was to utilize some of the functions in the CV2 library to find an array of coordinates that represent an outline of the non-background portion of the WSI.  From these coordinates, patch coordinates are created.  The second step involves actually extracting the patches from the WSI.  For the TransMIL model I created the feature embeddings using a pretrained ResNet-50 model and created an (B, N, 1028) tensor where B is the batch size, N is the number of patches, and 1028 is the feature embeddings size.  For the CNN solution, I got the prediction once the tile was created and then removed the tile from memory.\n\nThe next challenge was the memory constraint.  Both of the solutions mentioned in the previous paragraph avoid writing images to memory due to the large number of patches.  In addition to the memory challenge, there was a computational time challenge due to the 12 hr maximum runtime.  Processing these images takes a long time so I had to find a way to be optimal when tiling the image.  I started this by running step 1 of finding the outline of the important information using the thumbnail, and then transforming the pixel coordinates based on the downscaling ratio between the WSI and the thumbnail. This alone significantly sped up the runtime, allowing for the coordinates to be calculated for the entire training set in 20 minutes.  Next, I utilized the GPU pyvips library and the python-openslide library to efficiently read, process, and tile the WSIs alongwith some parallel programming.  Using this, I was able to get the runtime for inference on the TRAINING set done in 12.5 hours and because the test set is smaller in size (more images but ~200 less GB), I was able to get my baseline model (not fine-tuned) to run within the allowed time.\n\nFinally, while I believe the TransMIL model has more potential for high results, I worry that it will have a difficult time generalizing to data from other hospitals.  Because the data loading time takes very long (12 hours to get the feature embeddings for every train WSI), augmenting the data to expand the size of the training set would have taken many days or even weeks which I was unable to commit the time to do.  For the CNN model I was able to augment each 512x512 patch to allow the model to generalize better and not overfit to this specific hospital's data.  \n\nMy biggest contribution to this competition was my algorithm that optimally tiled each WSI in the test set within the allowed computation time.  I saw early on that there were many notebooks that used tiling solutions but were only able to retrieve ~30-50 patches.  I decided to go deeper and try to find a way to get every patch possible from every sample within the allowed time.  I believe that there must have been other competitors that were able to do the same thing but to me this was my biggest achievement during this competition and I will be able to see later tonight if it paid off.\n\nAt the time I am making this post my leaderboard submission is still my baseline submission and I am hopeful that either my TransMIL fine-tuned model or CNN solution will provide strong results.  Regardless of where I end up I am very interested to see other competitor's solutions and I plan to share my highly optimized code for tiling the entire image for each WSI after the competition.  Good luck to all competitors and I am very excited to see the final results!  Thank you to the competition owners for putting this competition out and for every code and discussion post along the way that has taught me so much more about data science."
  }
}