{
  "id": 344617,
  "title": "Divide images into tiles for submission - running for many hours and error",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/344617",
  "author_name": "Rodrigo M Carrillo Larco",
  "post_date": "2022-08-16T00:21:31.222000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi all, </p>\n<p>Has anyone tried to read the original images and divide into tiles (e.g., <a href=\"https://www.youtube.com/watch?v=tNfcvgPKgyU&amp;t=1148s\" target=\"_blank\">https://www.youtube.com/watch?v=tNfcvgPKgyU&amp;t=1148s</a>) for the submission? </p>\n<p>Here's my notebook:<br>\n<a href=\"https://www.kaggle.com/code/rodrigocarrillo/import-model-trained-locally-with-tiles-20220814/notebook\" target=\"_blank\">https://www.kaggle.com/code/rodrigocarrillo/import-model-trained-locally-with-tiles-20220814/notebook</a></p>\n<ol>\n<li>Load the original images (the ones they'll use for testing).</li>\n<li>Split each original image into tiles and drop those tiles with a lot of white space (i.e., without much information and thought it may save resources). </li>\n<li>Because you will have many many tiles per patient (as the original image has been divided), I even randomly selected 100 images per patient. I thought that with fewer tiles, it should work just fine. </li>\n<li>Load a simple pre-trained model.</li>\n<li>Apply the model to each tile.</li>\n<li>Save the predictions.</li>\n<li>Take the average by patient, so that there will only be one row per patient. </li>\n</ol>\n<p>Yesterday I did not select the random 100 tiles and the notebook ran for ~13 hours until I got an error \"Kaggle Error\". <br>\nToday, because I thought it was going to run faster, I randomly selected 100 tiles per patient; the notebook has been running for ~8 hours.</p>\n<p>Any suggestions?</p>\n<p>Thank you all<br>\nBest,</p>",
  "messages": [
    {
      "id": 1901684,
      "postDate": "2022-08-16T20:07:20.557Z",
      "content": "<p>I think your memory usage is intense, it's pretty amazing that your notebook continued for so long without OOM error  :)</p>\n<p>A few things to consider:</p>\n<p>1) You can downscale the image to a lower size before tiling, I have tried downscale factors of 10 and 5, both worked and created tiles within 3-4 hours for the entire training set - I use <a href=\"https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1\" target=\"_blank\">this</a> technique with downsized image tile creation, it has been used in such previous competitions successfully.</p>\n<p>2) You are also converting your RGB image from uint8 to a numpy array. Check if that is converting your array into a float because that will increase the size of your array, which means longer writing/reading/processing times. Best to write the image tile as is (i.e. in uint8 format) for fastest processing. You can process the array while you read the tile for training etc.</p>\n<p>3) Consider limiting the image size that you convert to tiles for the first round - I left out images above 1.2 Gb and overall left out only 34 images from the 754 images. So that's another way to save some time initially.</p>\n<p>All the best for the process.</p>\n<p><strong>PS.</strong> you can write the tiles in one notebook and commit it as a version that saves the output, then use that output using 'add data' tab from the top right corner of the kernel and add it to a new notebook that reads the tile for training. Kaggle allows making notebook chains like this, no need to do everything in one go.</p>",
      "rawMarkdown": "I think your memory usage is intense, it's pretty amazing that your notebook continued for so long without OOM error  :)\n\nA few things to consider:\n\n1) You can downscale the image to a lower size before tiling, I have tried downscale factors of 10 and 5, both worked and created tiles within 3-4 hours for the entire training set - I use [this](https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1) technique with downsized image tile creation, it has been used in such previous competitions successfully.\n\n2) You are also converting your RGB image from uint8 to a numpy array. Check if that is converting your array into a float because that will increase the size of your array, which means longer writing/reading/processing times. Best to write the image tile as is (i.e. in uint8 format) for fastest processing. You can process the array while you read the tile for training etc.\n\n3) Consider limiting the image size that you convert to tiles for the first round - I left out images above 1.2 Gb and overall left out only 34 images from the 754 images. So that's another way to save some time initially.\n\nAll the best for the process.\n\n**PS.** you can write the tiles in one notebook and commit it as a version that saves the output, then use that output using 'add data' tab from the top right corner of the kernel and add it to a new notebook that reads the tile for training. Kaggle allows making notebook chains like this, no need to do everything in one go.\n\n\n",
      "votes": 3
    },
    {
      "id": 1901020,
      "postDate": "2022-08-16T12:25:12.757Z",
      "content": "<p>Does this process of tiling have to be part of the submission notebook (the one having to run in 9h)? Or can we do it separately, save the tiles and then load those in another notebook?</p>",
      "rawMarkdown": "Does this process of tiling have to be part of the submission notebook (the one having to run in 9h)? Or can we do it separately, save the tiles and then load those in another notebook?",
      "votes": 3,
      "replies": [
        {
          "id": 1901151,
          "postDate": "2022-08-16T13:39:52.923Z",
          "content": "<p>Great question I have't thought of that. How can you do that? I thought you can only submit one notebook. Even if you separate the process in two notebooks (tile + model inference), both notebooks would count towards the total time, would it not? So in the end, it will still take many many hours?</p>",
          "rawMarkdown": "Great question I have't thought of that. How can you do that? I thought you can only submit one notebook. Even if you separate the process in two notebooks (tile + model inference), both notebooks would count towards the total time, would it not? So in the end, it will still take many many hours?",
          "votes": 1
        },
        {
          "id": 1905038,
          "postDate": "2022-08-18T17:41:33.210Z",
          "content": "<p>My thinking right now is that the since the test data only includes 4 data points, the competition will release some new test images after the competition ends. Therefore, we're supposed to also do data preprocessing on these unseen test images. My guess then is that we would actually need to do data preprocessing on these unseen test images and also do inference on them under the 9 hour time limit.</p>\n<p>I feel that this system for this competition makes it quite difficult for competitors to approach the competition with certainty. There are just so many issues with preprocessing the data that having to try and do it blindly in the end, with the hopes that our code will manage the unseen data correctly and under time limit is a bit difficult.</p>",
          "rawMarkdown": "My thinking right now is that the since the test data only includes 4 data points, the competition will release some new test images after the competition ends. Therefore, we're supposed to also do data preprocessing on these unseen test images. My guess then is that we would actually need to do data preprocessing on these unseen test images and also do inference on them under the 9 hour time limit.\n\nI feel that this system for this competition makes it quite difficult for competitors to approach the competition with certainty. There are just so many issues with preprocessing the data that having to try and do it blindly in the end, with the hopes that our code will manage the unseen data correctly and under time limit is a bit difficult.",
          "votes": 1
        },
        {
          "id": 1905057,
          "postDate": "2022-08-18T17:56:56.200Z",
          "content": "<p>Yes, that is true. Final submission must handle image processing and inference together, only training can be done separately and models saved and loaded later. I'm also confused on the notebook run times for the final case - as of now, are all 280 unseen test images running inference on our submission? Or just about 20 images are running inference? That is key to know if your notebook will complete 280 images within 9 hours. My guess is all 280 images are running submission inference, but score on LB are for 7% of that inference, based on my training set data inference speeds. But waiting for confirmation on this as well. Too many layers of uncertainty to handle here - at least for beginners like me :)</p>",
          "rawMarkdown": "Yes, that is true. Final submission must handle image processing and inference together, only training can be done separately and models saved and loaded later. I'm also confused on the notebook run times for the final case - as of now, are all 280 unseen test images running inference on our submission? Or just about 20 images are running inference? That is key to know if your notebook will complete 280 images within 9 hours. My guess is all 280 images are running submission inference, but score on LB are for 7% of that inference, based on my training set data inference speeds. But waiting for confirmation on this as well. Too many layers of uncertainty to handle here - at least for beginners like me :)",
          "votes": 1
        },
        {
          "id": 1905088,
          "postDate": "2022-08-18T18:14:50.617Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> are either of you able to help give us some clarity on the situation with the hidden data? Thank you for your time.</p>",
          "rawMarkdown": "Hi @barbaroserdal @ashleychow are either of you able to help give us some clarity on the situation with the hidden data? Thank you for your time.",
          "votes": 1
        },
        {
          "id": 1906193,
          "postDate": "2022-08-19T16:57:34.717Z",
          "content": "<p>Update: I tried a no inference submission just to have an idea of the notebook run times… Did the following in the notebook:</p>\n<ul>\n<li>Read image if it is upto 2 Gb in size</li>\n<li>Perform image resize operation to downscale by a factor of 5</li>\n<li>delete image and perform garbage collection</li>\n<li>assign random 0.5 probability to the two classes</li>\n<li>finally create submission.csv</li>\n</ul>\n<p>This notebook, without model loading or any inference operation, still ran for over an hour. I assume at this point that the submission notebook runs on all 280 images but scores for LB only take 7% of that inference data - over an hour is too long for 20 odd images.</p>\n<p>So I guess if your submission currently runs within 9 hours, final scoring should be fine. Would be great to have a confirmation though :) </p>\n<p>Anyone find any relevant information, please do share. Thanks and all the best.</p>",
          "rawMarkdown": "Update: I tried a no inference submission just to have an idea of the notebook run times... Did the following in the notebook:\n\n- Read image if it is upto 2 Gb in size\n- Perform image resize operation to downscale by a factor of 5\n- delete image and perform garbage collection\n- assign random 0.5 probability to the two classes\n- finally create submission.csv\n\nThis notebook, without model loading or any inference operation, still ran for over an hour. I assume at this point that the submission notebook runs on all 280 images but scores for LB only take 7% of that inference data - over an hour is too long for 20 odd images.\n\nSo I guess if your submission currently runs within 9 hours, final scoring should be fine. Would be great to have a confirmation though :) \n\nAnyone find any relevant information, please do share. Thanks and all the best.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1900363,
      "postDate": "2022-08-16T00:21:31.223Z",
      "content": "<p>Hi all, </p>\n<p>Has anyone tried to read the original images and divide into tiles (e.g., <a href=\"https://www.youtube.com/watch?v=tNfcvgPKgyU&amp;t=1148s\" target=\"_blank\">https://www.youtube.com/watch?v=tNfcvgPKgyU&amp;t=1148s</a>) for the submission? </p>\n<p>Here's my notebook:<br>\n<a href=\"https://www.kaggle.com/code/rodrigocarrillo/import-model-trained-locally-with-tiles-20220814/notebook\" target=\"_blank\">https://www.kaggle.com/code/rodrigocarrillo/import-model-trained-locally-with-tiles-20220814/notebook</a></p>\n<ol>\n<li>Load the original images (the ones they'll use for testing).</li>\n<li>Split each original image into tiles and drop those tiles with a lot of white space (i.e., without much information and thought it may save resources). </li>\n<li>Because you will have many many tiles per patient (as the original image has been divided), I even randomly selected 100 images per patient. I thought that with fewer tiles, it should work just fine. </li>\n<li>Load a simple pre-trained model.</li>\n<li>Apply the model to each tile.</li>\n<li>Save the predictions.</li>\n<li>Take the average by patient, so that there will only be one row per patient. </li>\n</ol>\n<p>Yesterday I did not select the random 100 tiles and the notebook ran for ~13 hours until I got an error \"Kaggle Error\". <br>\nToday, because I thought it was going to run faster, I randomly selected 100 tiles per patient; the notebook has been running for ~8 hours.</p>\n<p>Any suggestions?</p>\n<p>Thank you all<br>\nBest,</p>",
      "rawMarkdown": "Hi all, \n\nHas anyone tried to read the original images and divide into tiles (e.g., https://www.youtube.com/watch?v=tNfcvgPKgyU&t=1148s) for the submission? \n\n\nHere's my notebook:\nhttps://www.kaggle.com/code/rodrigocarrillo/import-model-trained-locally-with-tiles-20220814/notebook\n1. Load the original images (the ones they'll use for testing).\n2. Split each original image into tiles and drop those tiles with a lot of white space (i.e., without much information and thought it may save resources). \n3. Because you will have many many tiles per patient (as the original image has been divided), I even randomly selected 100 images per patient. I thought that with fewer tiles, it should work just fine. \n4. Load a simple pre-trained model.\n5. Apply the model to each tile.\n6. Save the predictions.\n7. Take the average by patient, so that there will only be one row per patient. \n\n\nYesterday I did not select the random 100 tiles and the notebook ran for ~13 hours until I got an error \"Kaggle Error\". \nToday, because I thought it was going to run faster, I randomly selected 100 tiles per patient; the notebook has been running for ~8 hours.\n\nAny suggestions?\n\n\nThank you all\nBest,\n\n\n\n\n\n\n",
      "votes": 4
    }
  ],
  "comments": [
    {
      "id": 1901684,
      "author_name": "tdiceman",
      "author_url": "",
      "post_date": "2022-08-16T20:07:20.557000",
      "content": "<p>I think your memory usage is intense, it's pretty amazing that your notebook continued for so long without OOM error  :)</p>\n<p>A few things to consider:</p>\n<p>1) You can downscale the image to a lower size before tiling, I have tried downscale factors of 10 and 5, both worked and created tiles within 3-4 hours for the entire training set - I use <a href=\"https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1\" target=\"_blank\">this</a> technique with downsized image tile creation, it has been used in such previous competitions successfully.</p>\n<p>2) You are also converting your RGB image from uint8 to a numpy array. Check if that is converting your array into a float because that will increase the size of your array, which means longer writing/reading/processing times. Best to write the image tile as is (i.e. in uint8 format) for fastest processing. You can process the array while you read the tile for training etc.</p>\n<p>3) Consider limiting the image size that you convert to tiles for the first round - I left out images above 1.2 Gb and overall left out only 34 images from the 754 images. So that's another way to save some time initially.</p>\n<p>All the best for the process.</p>\n<p><strong>PS.</strong> you can write the tiles in one notebook and commit it as a version that saves the output, then use that output using 'add data' tab from the top right corner of the kernel and add it to a new notebook that reads the tile for training. Kaggle allows making notebook chains like this, no need to do everything in one go.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1901020,
      "author_name": "Daniel Nicolae",
      "author_url": "",
      "post_date": "2022-08-16T12:25:12.757000",
      "content": "<p>Does this process of tiling have to be part of the submission notebook (the one having to run in 9h)? Or can we do it separately, save the tiles and then load those in another notebook?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1901151,
          "author_name": "Rodrigo M Carrillo Larco",
          "author_url": "",
          "post_date": "2022-08-16T13:39:52.923000",
          "content": "<p>Great question I have't thought of that. How can you do that? I thought you can only submit one notebook. Even if you separate the process in two notebooks (tile + model inference), both notebooks would count towards the total time, would it not? So in the end, it will still take many many hours?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1905038,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-18T17:41:33.210000",
          "content": "<p>My thinking right now is that the since the test data only includes 4 data points, the competition will release some new test images after the competition ends. Therefore, we're supposed to also do data preprocessing on these unseen test images. My guess then is that we would actually need to do data preprocessing on these unseen test images and also do inference on them under the 9 hour time limit.</p>\n<p>I feel that this system for this competition makes it quite difficult for competitors to approach the competition with certainty. There are just so many issues with preprocessing the data that having to try and do it blindly in the end, with the hopes that our code will manage the unseen data correctly and under time limit is a bit difficult.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1905057,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-08-18T17:56:56.200000",
          "content": "<p>Yes, that is true. Final submission must handle image processing and inference together, only training can be done separately and models saved and loaded later. I'm also confused on the notebook run times for the final case - as of now, are all 280 unseen test images running inference on our submission? Or just about 20 images are running inference? That is key to know if your notebook will complete 280 images within 9 hours. My guess is all 280 images are running submission inference, but score on LB are for 7% of that inference, based on my training set data inference speeds. But waiting for confirmation on this as well. Too many layers of uncertainty to handle here - at least for beginners like me :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1905088,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-18T18:14:50.617000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> are either of you able to help give us some clarity on the situation with the hidden data? Thank you for your time.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1906193,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-08-19T16:57:34.717000",
          "content": "<p>Update: I tried a no inference submission just to have an idea of the notebook run times… Did the following in the notebook:</p>\n<ul>\n<li>Read image if it is upto 2 Gb in size</li>\n<li>Perform image resize operation to downscale by a factor of 5</li>\n<li>delete image and perform garbage collection</li>\n<li>assign random 0.5 probability to the two classes</li>\n<li>finally create submission.csv</li>\n</ul>\n<p>This notebook, without model loading or any inference operation, still ran for over an hour. I assume at this point that the submission notebook runs on all 280 images but scores for LB only take 7% of that inference data - over an hour is too long for 20 odd images.</p>\n<p>So I guess if your submission currently runs within 9 hours, final scoring should be fine. Would be great to have a confirmation though :) </p>\n<p>Anyone find any relevant information, please do share. Thanks and all the best.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1901684": "I think your memory usage is intense, it's pretty amazing that your notebook continued for so long without OOM error  :)\n\nA few things to consider:\n\n1) You can downscale the image to a lower size before tiling, I have tried downscale factors of 10 and 5, both worked and created tiles within 3-4 hours for the entire training set - I use [this](https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1) technique with downsized image tile creation, it has been used in such previous competitions successfully.\n\n2) You are also converting your RGB image from uint8 to a numpy array. Check if that is converting your array into a float because that will increase the size of your array, which means longer writing/reading/processing times. Best to write the image tile as is (i.e. in uint8 format) for fastest processing. You can process the array while you read the tile for training etc.\n\n3) Consider limiting the image size that you convert to tiles for the first round - I left out images above 1.2 Gb and overall left out only 34 images from the 754 images. So that's another way to save some time initially.\n\nAll the best for the process.\n\n**PS.** you can write the tiles in one notebook and commit it as a version that saves the output, then use that output using 'add data' tab from the top right corner of the kernel and add it to a new notebook that reads the tile for training. Kaggle allows making notebook chains like this, no need to do everything in one go.\n\n\n",
    "1901020": "Does this process of tiling have to be part of the submission notebook (the one having to run in 9h)? Or can we do it separately, save the tiles and then load those in another notebook?",
    "1900363": "Hi all, \n\nHas anyone tried to read the original images and divide into tiles (e.g., https://www.youtube.com/watch?v=tNfcvgPKgyU&t=1148s) for the submission? \n\n\nHere's my notebook:\nhttps://www.kaggle.com/code/rodrigocarrillo/import-model-trained-locally-with-tiles-20220814/notebook\n1. Load the original images (the ones they'll use for testing).\n2. Split each original image into tiles and drop those tiles with a lot of white space (i.e., without much information and thought it may save resources). \n3. Because you will have many many tiles per patient (as the original image has been divided), I even randomly selected 100 images per patient. I thought that with fewer tiles, it should work just fine. \n4. Load a simple pre-trained model.\n5. Apply the model to each tile.\n6. Save the predictions.\n7. Take the average by patient, so that there will only be one row per patient. \n\n\nYesterday I did not select the random 100 tiles and the notebook ran for ~13 hours until I got an error \"Kaggle Error\". \nToday, because I thought it was going to run faster, I randomly selected 100 tiles per patient; the notebook has been running for ~8 hours.\n\nAny suggestions?\n\n\nThank you all\nBest,\n\n\n\n\n\n\n"
  }
}