{
  "id": 152732,
  "title": "how to monitor memory usage for \"Notebook Exceeded Allowed Compute\"",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/152732",
  "author_name": "hengck23",
  "post_date": "2020-05-21T06:54:04.243000",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>i have problems with \"Notebook Exceeded Allowed Compute\" (it is surprising that my \"try and except code don't stop this error\")</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc8e9ada936a4ee41dd59fb0818364f6c%2FSelection_162.png?generation=1590043936265640&amp;alt=media\" alt=\"\"></p>\n\n<p>i want to monitor my memory usage for the python script like this:</p>\n\n<p>```\nfor i in range(num_image):\n    ... do some processing ...\n    print( i, time, memory_used)</p>\n\n<p>```</p>\n\n<p>how can i do this? which package is supported by kaggle kernel?</p>",
  "messages": [
    {
      "id": 855740,
      "postDate": "2020-05-21T06:54:04.243Z",
      "content": "<p>i have problems with \"Notebook Exceeded Allowed Compute\" (it is surprising that my \"try and except code don't stop this error\")</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc8e9ada936a4ee41dd59fb0818364f6c%2FSelection_162.png?generation=1590043936265640&amp;alt=media\" alt=\"\"></p>\n\n<p>i want to monitor my memory usage for the python script like this:</p>\n\n<p>```\nfor i in range(num_image):\n    ... do some processing ...\n    print( i, time, memory_used)</p>\n\n<p>```</p>\n\n<p>how can i do this? which package is supported by kaggle kernel?</p>",
      "rawMarkdown": "i have problems with \"Notebook Exceeded Allowed Compute\" (it is surprising that my \"try and except code don't stop this error\")\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc8e9ada936a4ee41dd59fb0818364f6c%2FSelection_162.png?generation=1590043936265640&amp;alt=media)\n\n\ni want to monitor my memory usage for the python script like this:\n\n```\nfor i in range(num_image):\n    ... do some processing ...\n    print( i, time, memory_used)\n\n```\n\nhow can i do this? which package is supported by kaggle kernel?",
      "votes": 6
    },
    {
      "id": 859997,
      "postDate": "2020-05-24T23:40:07.937Z",
      "content": "<p>BTW, for my experiments on high res layer I wrote a function that uses OpenSlide to read only tiles that I want to load rather than entire high res tiff layer. So the memory consumption is greatly reduced, even with two workers.</p>\n\n<p>```\ndef tile(fname):\n    #use layer 2 for tile selection\n    img = skimage.io.MultiImage(fname)[-1]\n    shape = img.shape\n    r = 16 # ratio of layer 0 vs layer 2 res\n    sz16 = sz//r\n    pad0,pad1 = (sz16 - shape[0]%sz16)%sz16, (sz16 - shape[1]%sz16)%sz16\n    img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],constant_values=255)\n    img = img.reshape(img.shape[0]//sz16,sz16,img.shape[1]//sz16,sz16,3)\n    img = img.transpose(0,2,1,3,4).reshape(-1,sz16,sz16,3)\n    idxs = np.argsort(img.reshape(img.shape[0],-1).sum(-1))[:min(N,len(img))]</p>\n\n<pre><code>#read layer 0 tile by tile with use of openslide\nn0,n1 = (pad0+shape[0])//sz16, (pad1+shape[1])//sz16\nimg0 = openslide.OpenSlide(fname)\ntiles = []\nfor idx in idxs:\n    x = (-pad0//2 + sz16*(idx//n1))*r\n    y = (-pad1//2 + sz16*(idx%n1))*r\n    t = np.array(img0.read_region((y,x),0,(sz,sz)))[:,:,:3]\n    tiles.append(t)\nfor i in range(N - len(tiles)): tiles.append(np.full((sz,sz,3), 255, dtype=np.uint8))\n\nreturn np.stack(tiles)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "BTW, for my experiments on high res layer I wrote a function that uses OpenSlide to read only tiles that I want to load rather than entire high res tiff layer. So the memory consumption is greatly reduced, even with two workers.\n\n```\ndef tile(fname):\n    #use layer 2 for tile selection\n    img = skimage.io.MultiImage(fname)[-1]\n    shape = img.shape\n    r = 16 # ratio of layer 0 vs layer 2 res\n    sz16 = sz//r\n    pad0,pad1 = (sz16 - shape[0]%sz16)%sz16, (sz16 - shape[1]%sz16)%sz16\n    img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],constant_values=255)\n    img = img.reshape(img.shape[0]//sz16,sz16,img.shape[1]//sz16,sz16,3)\n    img = img.transpose(0,2,1,3,4).reshape(-1,sz16,sz16,3)\n    idxs = np.argsort(img.reshape(img.shape[0],-1).sum(-1))[:min(N,len(img))]\n\n    #read layer 0 tile by tile with use of openslide\n    n0,n1 = (pad0+shape[0])//sz16, (pad1+shape[1])//sz16\n    img0 = openslide.OpenSlide(fname)\n    tiles = []\n    for idx in idxs:\n        x = (-pad0//2 + sz16*(idx//n1))*r\n        y = (-pad1//2 + sz16*(idx%n1))*r\n        t = np.array(img0.read_region((y,x),0,(sz,sz)))[:,:,:3]\n        tiles.append(t)\n    for i in range(N - len(tiles)): tiles.append(np.full((sz,sz,3), 255, dtype=np.uint8))\n    \n    return np.stack(tiles)\n```",
      "votes": 4
    },
    {
      "id": 856405,
      "postDate": "2020-05-21T17:27:33.263Z",
      "content": "<p>The thing that may help with loading large images during inference is using only a single worker to ensure that only one tiff is loaded at the time. </p>\n\n<p>Your suggestion about monitoring the inference kernel makes sense since it's quite vital for debugging, but opens a huge gap for data probing: ppl would print out not only memory usage but test data stats as well.</p>",
      "rawMarkdown": "The thing that may help with loading large images during inference is using only a single worker to ensure that only one tiff is loaded at the time. \n\nYour suggestion about monitoring the inference kernel makes sense since it's quite vital for debugging, but opens a huge gap for data probing: ppl would print out not only memory usage but test data stats as well.",
      "votes": 1
    },
    {
      "id": 855808,
      "postDate": "2020-05-21T08:36:45.337Z",
      "content": "<p>Did you try to examine the read in tiff image size? I had similar issue when I read the image and mask of level 1 tiff resolution in the same time, then the notebook just ran out of memory. Since some of the level 1 images/masks are really huge, and the memory just running out before I do any processing on them. </p>",
      "rawMarkdown": "Did you try to examine the read in tiff image size? I had similar issue when I read the image and mask of level 1 tiff resolution in the same time, then the notebook just ran out of memory. Since some of the level 1 images/masks are really huge, and the memory just running out before I do any processing on them. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 859997,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-05-24T23:40:07.937000",
      "content": "<p>BTW, for my experiments on high res layer I wrote a function that uses OpenSlide to read only tiles that I want to load rather than entire high res tiff layer. So the memory consumption is greatly reduced, even with two workers.</p>\n\n<p>```\ndef tile(fname):\n    #use layer 2 for tile selection\n    img = skimage.io.MultiImage(fname)[-1]\n    shape = img.shape\n    r = 16 # ratio of layer 0 vs layer 2 res\n    sz16 = sz//r\n    pad0,pad1 = (sz16 - shape[0]%sz16)%sz16, (sz16 - shape[1]%sz16)%sz16\n    img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],constant_values=255)\n    img = img.reshape(img.shape[0]//sz16,sz16,img.shape[1]//sz16,sz16,3)\n    img = img.transpose(0,2,1,3,4).reshape(-1,sz16,sz16,3)\n    idxs = np.argsort(img.reshape(img.shape[0],-1).sum(-1))[:min(N,len(img))]</p>\n\n<pre><code>#read layer 0 tile by tile with use of openslide\nn0,n1 = (pad0+shape[0])//sz16, (pad1+shape[1])//sz16\nimg0 = openslide.OpenSlide(fname)\ntiles = []\nfor idx in idxs:\n    x = (-pad0//2 + sz16*(idx//n1))*r\n    y = (-pad1//2 + sz16*(idx%n1))*r\n    t = np.array(img0.read_region((y,x),0,(sz,sz)))[:,:,:3]\n    tiles.append(t)\nfor i in range(N - len(tiles)): tiles.append(np.full((sz,sz,3), 255, dtype=np.uint8))\n\nreturn np.stack(tiles)\n</code></pre>\n\n<p>```</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 856405,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-05-21T17:27:33.263000",
      "content": "<p>The thing that may help with loading large images during inference is using only a single worker to ensure that only one tiff is loaded at the time. </p>\n\n<p>Your suggestion about monitoring the inference kernel makes sense since it's quite vital for debugging, but opens a huge gap for data probing: ppl would print out not only memory usage but test data stats as well.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 855808,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2020-05-21T08:36:45.337000",
      "content": "<p>Did you try to examine the read in tiff image size? I had similar issue when I read the image and mask of level 1 tiff resolution in the same time, then the notebook just ran out of memory. Since some of the level 1 images/masks are really huge, and the memory just running out before I do any processing on them. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "855740": "i have problems with \"Notebook Exceeded Allowed Compute\" (it is surprising that my \"try and except code don't stop this error\")\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc8e9ada936a4ee41dd59fb0818364f6c%2FSelection_162.png?generation=1590043936265640&amp;alt=media)\n\n\ni want to monitor my memory usage for the python script like this:\n\n```\nfor i in range(num_image):\n    ... do some processing ...\n    print( i, time, memory_used)\n\n```\n\nhow can i do this? which package is supported by kaggle kernel?",
    "859997": "BTW, for my experiments on high res layer I wrote a function that uses OpenSlide to read only tiles that I want to load rather than entire high res tiff layer. So the memory consumption is greatly reduced, even with two workers.\n\n```\ndef tile(fname):\n    #use layer 2 for tile selection\n    img = skimage.io.MultiImage(fname)[-1]\n    shape = img.shape\n    r = 16 # ratio of layer 0 vs layer 2 res\n    sz16 = sz//r\n    pad0,pad1 = (sz16 - shape[0]%sz16)%sz16, (sz16 - shape[1]%sz16)%sz16\n    img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],constant_values=255)\n    img = img.reshape(img.shape[0]//sz16,sz16,img.shape[1]//sz16,sz16,3)\n    img = img.transpose(0,2,1,3,4).reshape(-1,sz16,sz16,3)\n    idxs = np.argsort(img.reshape(img.shape[0],-1).sum(-1))[:min(N,len(img))]\n\n    #read layer 0 tile by tile with use of openslide\n    n0,n1 = (pad0+shape[0])//sz16, (pad1+shape[1])//sz16\n    img0 = openslide.OpenSlide(fname)\n    tiles = []\n    for idx in idxs:\n        x = (-pad0//2 + sz16*(idx//n1))*r\n        y = (-pad1//2 + sz16*(idx%n1))*r\n        t = np.array(img0.read_region((y,x),0,(sz,sz)))[:,:,:3]\n        tiles.append(t)\n    for i in range(N - len(tiles)): tiles.append(np.full((sz,sz,3), 255, dtype=np.uint8))\n    \n    return np.stack(tiles)\n```",
    "856405": "The thing that may help with loading large images during inference is using only a single worker to ensure that only one tiff is loaded at the time. \n\nYour suggestion about monitoring the inference kernel makes sense since it's quite vital for debugging, but opens a huge gap for data probing: ppl would print out not only memory usage but test data stats as well.",
    "855808": "Did you try to examine the read in tiff image size? I had similar issue when I read the image and mask of level 1 tiff resolution in the same time, then the notebook just ran out of memory. Since some of the level 1 images/masks are really huge, and the memory just running out before I do any processing on them. "
  }
}