{
  "id": 440463,
  "title": "What to do for reducing the gb of test dataset?",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/440463",
  "author_name": "UEIGHT8",
  "post_date": "2023-09-15T02:38:17.943000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, I'm UEIGHT8 and a Kaggle beginner.<br>\nI'm poor at English, so if there is anything wrong, let me know.</p>\n<p>Then, I made a simple model and completed the notebook.<br>\nHowever, when I submitted it, I got \"Notebook out of memory\".<br>\nIt seems that the test set has approximately 1100(patients) × around 50.5 (the average from 1 to 100) images, so just downloading the test dataset, I'll spend about 30 to 40 gb.<br>\nBecause the limit of memory of RAM is 30gb, so I has not succeeded in submitting.</p>\n<p>Including the memory of the notebook, I have already spent 10 to 16 gb.<br>\nSo, if you know the way to reduce the size of the test set, tell me!</p>\n<p>My notebook link<br>\n→simple CNN road <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ueight8/simple-cnn-road-map/edit/run/142962939</a></p>",
  "messages": [
    {
      "id": 2440721,
      "postDate": "2023-09-15T17:26:14.353Z",
      "content": "<p>I think that looking into batch-wise prediction would be very helpful. It takes a subset of the data, loads to memory, predicts on it, and then releases the memory so that it isn't overloaded. There are some built-in as a part of some modeling types, but you may need to code a custom way to do that. I've used for-loops in the past, as that doesn't use hard-coded variables in direct memory, but temporary variables that clear when the loop finishes. Hope that helps, good luck!</p>",
      "rawMarkdown": "I think that looking into batch-wise prediction would be very helpful. It takes a subset of the data, loads to memory, predicts on it, and then releases the memory so that it isn't overloaded. There are some built-in as a part of some modeling types, but you may need to code a custom way to do that. I've used for-loops in the past, as that doesn't use hard-coded variables in direct memory, but temporary variables that clear when the loop finishes. Hope that helps, good luck!",
      "votes": 1,
      "replies": [
        {
          "id": 2447745,
          "postDate": "2023-09-20T08:25:22.720Z",
          "content": "<p>Thanks a lot !<br>\nThen, what is 'batch-wise prediction'?<br>\nI would appreciate it if you share any websites describing it very well !!</p>",
          "rawMarkdown": "Thanks a lot !\nThen, what is 'batch-wise prediction'?\nI would appreciate it if you share any websites describing it very well !!",
          "votes": 1,
          "replies": [
            {
              "id": 2448570,
              "postDate": "2023-09-20T17:12:58.287Z",
              "content": "<p>When I originally responded, I actually tried to look up some websites / references but I couldn't really find any that explained it well, or at all. It's probably easier for me to try explaining it myself, so here goes. <br>\nImagine that your RAM limit, 30gb, is a basket. The basket has a limit for the number of dicom files that it can hold. My assumption is that your current setup tries to put all of the test images into this basket, and of course it gets overloaded, you're looking at probably 50000-100000 images in the hidden test set, when the most that the basket can hold is probably somewhere around 2000 at one time. So, how do you fix this? Well, the thing about the basket is that it doesn't have limits if you keep emptying it. <br>\nInstead of <br>\n<code>basket = []</code><br>\n<code>for i in test_data:</code><br>\n<code>basket.append(i)</code><br>\n<code>model.predict(basket)</code><br>\nIn this case you have no way of emptying the basket, as we can see. It will become full after a certain point, and then your submission will likely fail.<br>\nBut, if you do this:<br>\n<code>def predictor(test_data[0:25]):</code><br>\n<code>basket = []</code><br>\n<code>for i in test_data:</code><br>\n<code>basket.append(i)</code><br>\n<code>predictions = model.predict(basket)</code><br>\n<code>del basket</code><br>\n<code>return predictions</code><br>\nYou do lots of important things, now your basket isn't hardcoded, so it only exists while the function is running, you can take a smaller section of your data, fill the basket, make predictions on it, delete it (which removes the memory usage of it, by the way), and then return that predictions array for you to use in your final submission file. I hope that makes sense, if not I can try to explain a little bit more.</p>",
              "rawMarkdown": "When I originally responded, I actually tried to look up some websites / references but I couldn't really find any that explained it well, or at all. It's probably easier for me to try explaining it myself, so here goes. \nImagine that your RAM limit, 30gb, is a basket. The basket has a limit for the number of dicom files that it can hold. My assumption is that your current setup tries to put all of the test images into this basket, and of course it gets overloaded, you're looking at probably 50000-100000 images in the hidden test set, when the most that the basket can hold is probably somewhere around 2000 at one time. So, how do you fix this? Well, the thing about the basket is that it doesn't have limits if you keep emptying it. \nInstead of \n`basket = [] `\n`for i in test_data:`\n` basket.append(i)`\n`model.predict(basket)`\nIn this case you have no way of emptying the basket, as we can see. It will become full after a certain point, and then your submission will likely fail.\nBut, if you do this:\n`def predictor(test_data[0:25]):`\n`basket = []`\n`for i in test_data:`\n`basket.append(i)`\n`predictions = model.predict(basket)`\n`del basket`\n`return predictions`\nYou do lots of important things, now your basket isn't hardcoded, so it only exists while the function is running, you can take a smaller section of your data, fill the basket, make predictions on it, delete it (which removes the memory usage of it, by the way), and then return that predictions array for you to use in your final submission file. I hope that makes sense, if not I can try to explain a little bit more.",
              "votes": 1
            },
            {
              "id": 2454691,
              "postDate": "2023-09-25T03:21:04.657Z",
              "content": "<p>Wonderful !!<br>\nProbably, I understand the point.<br>\nThank you so much.</p>",
              "rawMarkdown": "Wonderful !!\nProbably, I understand the point.\nThank you so much.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2439607,
      "postDate": "2023-09-15T02:38:17.943Z",
      "content": "<p>Hi, I'm UEIGHT8 and a Kaggle beginner.<br>\nI'm poor at English, so if there is anything wrong, let me know.</p>\n<p>Then, I made a simple model and completed the notebook.<br>\nHowever, when I submitted it, I got \"Notebook out of memory\".<br>\nIt seems that the test set has approximately 1100(patients) × around 50.5 (the average from 1 to 100) images, so just downloading the test dataset, I'll spend about 30 to 40 gb.<br>\nBecause the limit of memory of RAM is 30gb, so I has not succeeded in submitting.</p>\n<p>Including the memory of the notebook, I have already spent 10 to 16 gb.<br>\nSo, if you know the way to reduce the size of the test set, tell me!</p>\n<p>My notebook link<br>\n→simple CNN road <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ueight8/simple-cnn-road-map/edit/run/142962939</a></p>",
      "rawMarkdown": "Hi, I'm UEIGHT8 and a Kaggle beginner.\nI'm poor at English, so if there is anything wrong, let me know.\n\nThen, I made a simple model and completed the notebook.\nHowever, when I submitted it, I got \"Notebook out of memory\".\nIt seems that the test set has approximately 1100(patients) × around 50.5 (the average from 1 to 100) images, so just downloading the test dataset, I'll spend about 30 to 40 gb.\nBecause the limit of memory of RAM is 30gb, so I has not succeeded in submitting.\n\nIncluding the memory of the notebook, I have already spent 10 to 16 gb.\nSo, if you know the way to reduce the size of the test set, tell me!\n\n\nMy notebook link\n→simple CNN road [https://www.kaggle.com/code/ueight8/simple-cnn-road-map/edit/run/142962939](url)",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2440721,
      "author_name": "Art. Berz.",
      "author_url": "",
      "post_date": "2023-09-15T17:26:14.353000",
      "content": "<p>I think that looking into batch-wise prediction would be very helpful. It takes a subset of the data, loads to memory, predicts on it, and then releases the memory so that it isn't overloaded. There are some built-in as a part of some modeling types, but you may need to code a custom way to do that. I've used for-loops in the past, as that doesn't use hard-coded variables in direct memory, but temporary variables that clear when the loop finishes. Hope that helps, good luck!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2447745,
          "author_name": "UEIGHT8",
          "author_url": "",
          "post_date": "2023-09-20T08:25:22.720000",
          "content": "<p>Thanks a lot !<br>\nThen, what is 'batch-wise prediction'?<br>\nI would appreciate it if you share any websites describing it very well !!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2448570,
              "author_name": "Art. Berz.",
              "author_url": "",
              "post_date": "2023-09-20T17:12:58.287000",
              "content": "<p>When I originally responded, I actually tried to look up some websites / references but I couldn't really find any that explained it well, or at all. It's probably easier for me to try explaining it myself, so here goes. <br>\nImagine that your RAM limit, 30gb, is a basket. The basket has a limit for the number of dicom files that it can hold. My assumption is that your current setup tries to put all of the test images into this basket, and of course it gets overloaded, you're looking at probably 50000-100000 images in the hidden test set, when the most that the basket can hold is probably somewhere around 2000 at one time. So, how do you fix this? Well, the thing about the basket is that it doesn't have limits if you keep emptying it. <br>\nInstead of <br>\n<code>basket = []</code><br>\n<code>for i in test_data:</code><br>\n<code>basket.append(i)</code><br>\n<code>model.predict(basket)</code><br>\nIn this case you have no way of emptying the basket, as we can see. It will become full after a certain point, and then your submission will likely fail.<br>\nBut, if you do this:<br>\n<code>def predictor(test_data[0:25]):</code><br>\n<code>basket = []</code><br>\n<code>for i in test_data:</code><br>\n<code>basket.append(i)</code><br>\n<code>predictions = model.predict(basket)</code><br>\n<code>del basket</code><br>\n<code>return predictions</code><br>\nYou do lots of important things, now your basket isn't hardcoded, so it only exists while the function is running, you can take a smaller section of your data, fill the basket, make predictions on it, delete it (which removes the memory usage of it, by the way), and then return that predictions array for you to use in your final submission file. I hope that makes sense, if not I can try to explain a little bit more.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2454691,
              "author_name": "UEIGHT8",
              "author_url": "",
              "post_date": "2023-09-25T03:21:04.657000",
              "content": "<p>Wonderful !!<br>\nProbably, I understand the point.<br>\nThank you so much.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2440721": "I think that looking into batch-wise prediction would be very helpful. It takes a subset of the data, loads to memory, predicts on it, and then releases the memory so that it isn't overloaded. There are some built-in as a part of some modeling types, but you may need to code a custom way to do that. I've used for-loops in the past, as that doesn't use hard-coded variables in direct memory, but temporary variables that clear when the loop finishes. Hope that helps, good luck!",
    "2439607": "Hi, I'm UEIGHT8 and a Kaggle beginner.\nI'm poor at English, so if there is anything wrong, let me know.\n\nThen, I made a simple model and completed the notebook.\nHowever, when I submitted it, I got \"Notebook out of memory\".\nIt seems that the test set has approximately 1100(patients) × around 50.5 (the average from 1 to 100) images, so just downloading the test dataset, I'll spend about 30 to 40 gb.\nBecause the limit of memory of RAM is 30gb, so I has not succeeded in submitting.\n\nIncluding the memory of the notebook, I have already spent 10 to 16 gb.\nSo, if you know the way to reduce the size of the test set, tell me!\n\n\nMy notebook link\n→simple CNN road [https://www.kaggle.com/code/ueight8/simple-cnn-road-map/edit/run/142962939](url)"
  }
}