{
  "id": 154875,
  "title": "Some ways that boost my score",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/154875",
  "author_name": "Shiyuan Zeng",
  "post_date": "2020-05-30T07:06:22.744000",
  "votes": 27,
  "comment_count": 29,
  "views": 0,
  "content": "<p>As experiment some thoughts. I boost my score to cv/lb 0.807475/0.81. Here I wanna show a part of ways that boost my score.(not all of course) And the main reason why I post this topic is 抛砖引玉(I don't know how to translate it as many tools' translations are not so accurate)\nMy model forward is: backbone.features -&gt; pool -&gt; dropout(0.2) -&gt; linear(flatten(), 6).\n- Model\n    - efficientnet b0\n- Pool layer\n    - GeM &gt; concat(max+avg) &gt; avg &gt; max\n- Data source\n    - <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">panda-2020-level-1-2</a>(cv2.imread from this dataset is faster than skimage.io.MultiImage read from original competition data if my code isn't wrong, but just a little faster)\n- Data process\n    - level 1 resolution\n    - tiles, like what iafoss did(the idea is so great and code is so tidy!)\n    - transform each tile then combine into one image &gt; tiles &gt; combine tiles into one image then transform it(the whole image), theoretical analysis which I accepted is <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/153872\">List of tiles or large single image made of tiles?</a> on comment by iafoss\n    - augmentations\n        - flip + oneof(rotate, randomrotate90)\n        - brightness, contrast, ssr, channelshuffle aren't work for me\n- Training details\n    - random split, 80:20\n    - 16 tiles(transform each then combine to one image(4tilesx4tiles))\n    - bs: 4\n    - epochs: 5(more cause cv/lb decrease for me, maybe b0 is too samll?)\n    - lr-scheduler: none, lr: 5e-04, no tta\n    - save model weights by score(the competition metric)\n    - hardware: kaggle kernel(plan to rent a 2080ti, or 40min/epoch for b0 is too slow...)</p>\n\n<p>Some questions\n- Has anyone try tpu? I try tpu by torch-xla(xm.spawn), but it crash after committing it ~10k secs. And not so fast as I except(I sample 1k image, and it run ~4min/epoch, but my model on gpu just cost 45min/epoch while it runs 10k image...) I use 1 core.\n- Is there any way to speed up doing tiles/transform with np/torch/albumentations? TF has, mentioned <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">rotation-augmentation-gpu-tpu-0-96</a> by Chris\n- proper way to train multi folds?(I failed in using parallel on tpu by modifing the <a href=\"https://www.kaggle.com/tarunpaparaju/panda-challenge-resnet-multitask-8-fold-on-tpu\">panda-challenge-resnet-multitask-8-fold-on-tpu</a> )\n- Which model do you choose for testing augmentations/ideas(it seems res18 faster than b0, I plan to turn to res18)? And the result is consistent or not while you trun to deeper model?\n- What's your training time(model? hardware?)?</p>\n\n<p>Feel free to discuss anything that properly, and thank you for your coming to see my topic.\nWhat's your ways to boost score?(not all of course, because it's a competition)</p>",
  "messages": [
    {
      "id": 867317,
      "postDate": "2020-05-30T07:06:22.743Z",
      "content": "<p>As experiment some thoughts. I boost my score to cv/lb 0.807475/0.81. Here I wanna show a part of ways that boost my score.(not all of course) And the main reason why I post this topic is 抛砖引玉(I don't know how to translate it as many tools' translations are not so accurate)\nMy model forward is: backbone.features -&gt; pool -&gt; dropout(0.2) -&gt; linear(flatten(), 6).\n- Model\n    - efficientnet b0\n- Pool layer\n    - GeM &gt; concat(max+avg) &gt; avg &gt; max\n- Data source\n    - <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">panda-2020-level-1-2</a>(cv2.imread from this dataset is faster than skimage.io.MultiImage read from original competition data if my code isn't wrong, but just a little faster)\n- Data process\n    - level 1 resolution\n    - tiles, like what iafoss did(the idea is so great and code is so tidy!)\n    - transform each tile then combine into one image &gt; tiles &gt; combine tiles into one image then transform it(the whole image), theoretical analysis which I accepted is <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/153872\">List of tiles or large single image made of tiles?</a> on comment by iafoss\n    - augmentations\n        - flip + oneof(rotate, randomrotate90)\n        - brightness, contrast, ssr, channelshuffle aren't work for me\n- Training details\n    - random split, 80:20\n    - 16 tiles(transform each then combine to one image(4tilesx4tiles))\n    - bs: 4\n    - epochs: 5(more cause cv/lb decrease for me, maybe b0 is too samll?)\n    - lr-scheduler: none, lr: 5e-04, no tta\n    - save model weights by score(the competition metric)\n    - hardware: kaggle kernel(plan to rent a 2080ti, or 40min/epoch for b0 is too slow...)</p>\n\n<p>Some questions\n- Has anyone try tpu? I try tpu by torch-xla(xm.spawn), but it crash after committing it ~10k secs. And not so fast as I except(I sample 1k image, and it run ~4min/epoch, but my model on gpu just cost 45min/epoch while it runs 10k image...) I use 1 core.\n- Is there any way to speed up doing tiles/transform with np/torch/albumentations? TF has, mentioned <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">rotation-augmentation-gpu-tpu-0-96</a> by Chris\n- proper way to train multi folds?(I failed in using parallel on tpu by modifing the <a href=\"https://www.kaggle.com/tarunpaparaju/panda-challenge-resnet-multitask-8-fold-on-tpu\">panda-challenge-resnet-multitask-8-fold-on-tpu</a> )\n- Which model do you choose for testing augmentations/ideas(it seems res18 faster than b0, I plan to turn to res18)? And the result is consistent or not while you trun to deeper model?\n- What's your training time(model? hardware?)?</p>\n\n<p>Feel free to discuss anything that properly, and thank you for your coming to see my topic.\nWhat's your ways to boost score?(not all of course, because it's a competition)</p>",
      "rawMarkdown": "As experiment some thoughts. I boost my score to cv/lb 0.807475/0.81. Here I wanna show a part of ways that boost my score.(not all of course) And the main reason why I post this topic is 抛砖引玉(I don't know how to translate it as many tools' translations are not so accurate)\nMy model forward is: backbone.features -&gt; pool -&gt; dropout(0.2) -&gt; linear(flatten(), 6).\n- Model\n    - efficientnet b0\n- Pool layer\n    - GeM &gt; concat(max+avg) &gt; avg &gt; max\n- Data source\n    - [panda-2020-level-1-2](https://www.kaggle.com/lopuhin/panda-2020-level-1-2)(cv2.imread from this dataset is faster than skimage.io.MultiImage read from original competition data if my code isn't wrong, but just a little faster)\n- Data process\n    - level 1 resolution\n    - tiles, like what iafoss did(the idea is so great and code is so tidy!)\n    - transform each tile then combine into one image &gt; tiles &gt; combine tiles into one image then transform it(the whole image), theoretical analysis which I accepted is [List of tiles or large single image made of tiles?](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/153872) on comment by iafoss\n    - augmentations\n        - flip + oneof(rotate, randomrotate90)\n        - brightness, contrast, ssr, channelshuffle aren't work for me\n- Training details\n    - random split, 80:20\n    - 16 tiles(transform each then combine to one image(4tilesx4tiles))\n    - bs: 4\n    - epochs: 5(more cause cv/lb decrease for me, maybe b0 is too samll?)\n    - lr-scheduler: none, lr: 5e-04, no tta\n    - save model weights by score(the competition metric)\n    - hardware: kaggle kernel(plan to rent a 2080ti, or 40min/epoch for b0 is too slow...)\n\nSome questions\n- Has anyone try tpu? I try tpu by torch-xla(xm.spawn), but it crash after committing it ~10k secs. And not so fast as I except(I sample 1k image, and it run ~4min/epoch, but my model on gpu just cost 45min/epoch while it runs 10k image...) I use 1 core.\n- Is there any way to speed up doing tiles/transform with np/torch/albumentations? TF has, mentioned [rotation-augmentation-gpu-tpu-0-96](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96) by Chris\n- proper way to train multi folds?(I failed in using parallel on tpu by modifing the [panda-challenge-resnet-multitask-8-fold-on-tpu](https://www.kaggle.com/tarunpaparaju/panda-challenge-resnet-multitask-8-fold-on-tpu) )\n- Which model do you choose for testing augmentations/ideas(it seems res18 faster than b0, I plan to turn to res18)? And the result is consistent or not while you trun to deeper model?\n- What's your training time(model? hardware?)?\n\nFeel free to discuss anything that properly, and thank you for your coming to see my topic.\nWhat's your ways to boost score?(not all of course, because it's a competition)",
      "votes": 25
    },
    {
      "id": 867375,
      "postDate": "2020-05-30T08:21:29.007Z",
      "content": "<p>Hello <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>, I'm using a B0 (0.87 LB). </p>\n\n<p>About # epochs: Perhaps increase the number of epochs you train (you will see fluctuations in the CV/LB score, but on average it increases (at least for me)). 5 epochs seems very little.\nAbout training run time: I think it could be a good idea to precompute all the tile/patch-coordinates for all images (and save it to file). Then feed the precomputed coordinates to/through your pipeline together with the corresponding images.</p>\n\n<p>Further, for me max pooling isn't working well. For some reason my loss get really high. I need to use only the average pooling layer.</p>",
      "rawMarkdown": "Hello @cnzengshiyuan, I'm using a B0 (0.87 LB). \n\nAbout # epochs: Perhaps increase the number of epochs you train (you will see fluctuations in the CV/LB score, but on average it increases (at least for me)). 5 epochs seems very little.\nAbout training run time: I think it could be a good idea to precompute all the tile/patch-coordinates for all images (and save it to file). Then feed the precomputed coordinates to/through your pipeline together with the corresponding images.\n\nFurther, for me max pooling isn't working well. For some reason my loss get really high. I need to use only the average pooling layer.",
      "votes": 3,
      "replies": [
        {
          "id": 867410,
          "postDate": "2020-05-30T08:58:49.020Z",
          "content": "<p><a href=\"/akensert\">@akensert</a> Thanks for your advice! :) I will train more epochs and precompute(it seems should be done locally and upload to kaggle? I try in kaggle kernel to save as .npy, but get disk error and kernel crashed)\nMaxpool doesn't perform well for me, too.\nWhat's more, b0 get 0.87 is so impressive, what's the key to this if you don'y mind disclosure?(about network architecture/augmentations/post-process?) No need to specific details, just a direction is ok.😄 </p>",
          "rawMarkdown": "@akensert Thanks for your advice! :) I will train more epochs and precompute(it seems should be done locally and upload to kaggle? I try in kaggle kernel to save as .npy, but get disk error and kernel crashed)\nMaxpool doesn't perform well for me, too.\nWhat's more, b0 get 0.87 is so impressive, what's the key to this if you don'y mind disclosure?(about network architecture/augmentations/post-process?) No need to specific details, just a direction is ok.😄 \n"
        },
        {
          "id": 867471,
          "postDate": "2020-05-30T10:28:12.500Z",
          "content": "<p>No problem :-)</p>\n\n<p>I save it as .npy file, and this file should not be that big (few MB). You only need to save the y[_start] and x[_start] coordninates (so e.g. (10616, 16, 2) shaped array)</p>\n\n<p>About your last question: I think by simply train for many more epochs will help you get much better results (you just have to suffer through ~24 hours of training :-D). Other adjustments could be to try more tiles/patches, somewhat lower learning rate, and image augmentation. Remember these are just some suggestions, that may, or may not, improve the performance of your models.</p>",
          "rawMarkdown": "No problem :-)\n\nI save it as .npy file, and this file should not be that big (few MB). You only need to save the y[\\_start\\] and x[\\_start] coordninates (so e.g. (10616, 16, 2) shaped array)\n\nAbout your last question: I think by simply train for many more epochs will help you get much better results (you just have to suffer through ~24 hours of training :-D). Other adjustments could be to try more tiles/patches, somewhat lower learning rate, and image augmentation. Remember these are just some suggestions, that may, or may not, improve the performance of your models.",
          "votes": 1
        },
        {
          "id": 867512,
          "postDate": "2020-05-30T11:31:47.053Z",
          "content": "<p>&gt; (10616, 16, 2) shaped array</p>\n\n<p>! Great! I get the meaning, puzzle with the \"coordinates\" first time, but now understand. This is really a smart approach💯💯💯 However, it should also be padded, right? And just jump the argsort part to save time?(correct me if I miss sth important)\nThanks for your sharing👍👍👍 I will update the improvement about more epochs or others after dayssssss</p>\n\n<p>&gt; you just have to suffer through ~24 hours of training :-D</p>\n\n<p>suffer soon😆 </p>",
          "rawMarkdown": "&gt; (10616, 16, 2) shaped array\n\n! Great! I get the meaning, puzzle with the \"coordinates\" first time, but now understand. This is really a smart approach💯💯💯 However, it should also be padded, right? And just jump the argsort part to save time?(correct me if I miss sth important)\nThanks for your sharing👍👍👍 I will update the improvement about more epochs or others after dayssssss\n\n&gt; you just have to suffer through ~24 hours of training :-D\n\nsuffer soon😆 "
        },
        {
          "id": 867961,
          "postDate": "2020-05-30T19:20:58.370Z",
          "content": "<p>Great :-) Yes you can pad it (I guess it depends on your pipeline etc, if you need/should do it or not). Good luck!</p>",
          "rawMarkdown": "Great :-) Yes you can pad it (I guess it depends on your pipeline etc, if you need/should do it or not). Good luck!",
          "votes": 1
        },
        {
          "id": 868124,
          "postDate": "2020-05-31T00:28:42.273Z",
          "content": "<p>Thanks for your kindness and patience! I did it yesterday and save as .npy as expected, can't wait to see the difference!</p>",
          "rawMarkdown": "Thanks for your kindness and patience! I did it yesterday and save as .npy as expected, can't wait to see the difference!"
        }
      ]
    },
    {
      "id": 867546,
      "postDate": "2020-05-30T11:59:23.257Z",
      "content": "<p>hey 曾，\nWe are offering free GPU for Chinese, feel free to join our wechat group, our product is at <a href=\"https://featurize.cn\">https://featurize.cn</a> . You can find the QRCode of our WeChat after you signing in.\nBesides, we are just launched this competition too.</p>",
      "rawMarkdown": "hey 曾，\nWe are offering free GPU for Chinese, feel free to join our wechat group, our product is at https://featurize.cn . You can find the QRCode of our WeChat after you signing in.\nBesides, we are just launched this competition too.",
      "votes": 1,
      "replies": [
        {
          "id": 867589,
          "postDate": "2020-05-30T12:42:02.553Z",
          "content": "<p>Sounds attractive! I will check it later :)</p>",
          "rawMarkdown": "Sounds attractive! I will check it later :)"
        }
      ]
    },
    {
      "id": 867430,
      "postDate": "2020-05-30T09:26:56.160Z",
      "content": "<p>Interesting</p>",
      "rawMarkdown": "Interesting",
      "votes": 1
    },
    {
      "id": 874253,
      "postDate": "2020-06-04T18:46:37.630Z",
      "content": "<p>what you can do is to get little more boost is that dont use full dataset to train .\nUse just top 7K in one fold and  then bottom 7K records .\nDo train in 2 folds and blend.</p>",
      "rawMarkdown": "what you can do is to get little more boost is that dont use full dataset to train .\nUse just top 7K in one fold and  then bottom 7K records .\nDo train in 2 folds and blend.\n",
      "votes": 2,
      "replies": [
        {
          "id": 874407,
          "postDate": "2020-06-05T00:38:50.813Z",
          "content": "<p>What's the meaning of top 7k and bottom 7k? Is that you mean use like the pseudo code shows?\n<code>\nfold_1 = dataset[:7k], fold_2 = dataset[-7k:]\ntrain each fold until converged\n</code></p>",
          "rawMarkdown": "What's the meaning of top 7k and bottom 7k? Is that you mean use like the pseudo code shows?\n```\nfold_1 = dataset[:7k], fold_2 = dataset[-7k:]\ntrain each fold until converged\n```"
        },
        {
          "id": 875389,
          "postDate": "2020-06-05T18:32:39.367Z",
          "content": "<p>yes , as above </p>",
          "rawMarkdown": "yes , as above "
        },
        {
          "id": 875578,
          "postDate": "2020-06-06T00:05:50.453Z",
          "content": "<p>Thanks for your reply :D</p>",
          "rawMarkdown": "Thanks for your reply :D"
        }
      ]
    },
    {
      "id": 867404,
      "postDate": "2020-05-30T08:53:30.907Z",
      "content": "<p>Hi !\nI've tried so many things (augmentation, models, losses, learning rates, even some new implementation using LSTM or bagging) and I was still at 0.79. \nThen increasing the number of tiles significantly boosted my score.  </p>",
      "rawMarkdown": "Hi !\nI've tried so many things (augmentation, models, losses, learning rates, even some new implementation using LSTM or bagging) and I was still at 0.79. \nThen increasing the number of tiles significantly boosted my score.  ",
      "votes": 2,
      "replies": [
        {
          "id": 867414,
          "postDate": "2020-05-30T09:04:35.353Z",
          "content": "<p>It makes sense! More tiles contains more information. As you mentioned, your network architecture seems contains many customized layers(except the backbone)?\nThe number of tiles I choose(16) seems a bit small as I check some other number.</p>",
          "rawMarkdown": "It makes sense! More tiles contains more information. As you mentioned, your network architecture seems contains many customized layers(except the backbone)?\nThe number of tiles I choose(16) seems a bit small as I check some other number."
        },
        {
          "id": 867977,
          "postDate": "2020-05-30T19:38:43.627Z",
          "content": "<p>I'm using a standard resnext, no custom layers. Increasing the number of tiles could help you</p>",
          "rawMarkdown": "I'm using a standard resnext, no custom layers. Increasing the number of tiles could help you",
          "votes": 2
        },
        {
          "id": 868121,
          "postDate": "2020-05-31T00:25:15.477Z",
          "content": "<p>Thanks for your kindness!💯  I plan to train more epochs and then increase the number of tiles. =)</p>",
          "rawMarkdown": "Thanks for your kindness!💯  I plan to train more epochs and then increase the number of tiles. =)"
        }
      ]
    },
    {
      "id": 874722,
      "postDate": "2020-06-05T08:54:01.913Z",
      "content": "<p>Hi thanks for sharing. \nDid GEM pooling significantly boost your score over other pooling methods ?</p>",
      "rawMarkdown": "Hi thanks for sharing. \nDid GEM pooling significantly boost your score over other pooling methods ?",
      "replies": [
        {
          "id": 874753,
          "postDate": "2020-06-05T09:18:14.140Z",
          "content": "<p>This topic is just based on 5~10 epochs, so I can't said what GeM will perform further.\nBut GeM boost ~1e-04 on my local score(over others based on 5~10 epochs), and it give me the most consistent trend between train loss and valid loss(even I train more epochs), so I choose it :) Because I choose GeM in 5~10 epochs, so I did't try other pool layers for more epochs.😄  Hope it's useful for you.</p>",
          "rawMarkdown": "This topic is just based on 5~10 epochs, so I can't said what GeM will perform further.\nBut GeM boost ~1e-04 on my local score(over others based on 5~10 epochs), and it give me the most consistent trend between train loss and valid loss(even I train more epochs), so I choose it :) Because I choose GeM in 5~10 epochs, so I did't try other pool layers for more epochs.😄  Hope it's useful for you."
        },
        {
          "id": 874775,
          "postDate": "2020-06-05T09:38:56.517Z",
          "content": "<p>good to know, thanks 👍 </p>",
          "rawMarkdown": "good to know, thanks 👍 "
        }
      ]
    },
    {
      "id": 868899,
      "postDate": "2020-05-31T15:01:48.747Z",
      "content": "<p><a href=\"/akensert\">@akensert</a> how we can input asymmetric image created form 128x128x36, by cocatenation image will be of size 512 x 1152 (approx),  Thanks for help in advance.</p>",
      "rawMarkdown": "@akensert how we can input asymmetric image created form 128x128x36, by cocatenation image will be of size 512 x 1152 (approx),  Thanks for help in advance.\n",
      "replies": [
        {
          "id": 872142,
          "postDate": "2020-06-03T00:12:25.383Z",
          "content": "<p><a href=\"/rajnishe\">@rajnishe</a> I don't think asymmetric image is a problem, as the CNN can also accept the different size of the image what feeds in it(you can just fees them to your CNN model, but all the data should have the same size, or it will cause error when concatenate some images into a batch). Can my answer solve your problem?</p>",
          "rawMarkdown": "@rajnishe I don't think asymmetric image is a problem, as the CNN can also accept the different size of the image what feeds in it(you can just fees them to your CNN model, but all the data should have the same size, or it will cause error when concatenate some images into a batch). Can my answer solve your problem?"
        },
        {
          "id": 872580,
          "postDate": "2020-06-03T10:22:51.700Z",
          "content": "<p><a href=\"/rajnishe\">@rajnishe</a> I make sure it's symmetric. And 128x128x3x36 would be 768x768x3 no?</p>",
          "rawMarkdown": "@rajnishe I make sure it's symmetric. And 128x128x3x36 would be 768x768x3 no?"
        },
        {
          "id": 873780,
          "postDate": "2020-06-04T12:25:20.917Z",
          "content": "<p>yes , thanks <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> </p>",
          "rawMarkdown": "yes , thanks @cnzengshiyuan "
        },
        {
          "id": 873783,
          "postDate": "2020-06-04T12:26:10.743Z",
          "content": "<p>thanks <a href=\"/akensert\">@akensert</a> </p>",
          "rawMarkdown": "thanks @akensert "
        },
        {
          "id": 873831,
          "postDate": "2020-06-04T13:11:15.790Z",
          "content": "<p>😄 You are welcome.</p>",
          "rawMarkdown": "😄 You are welcome."
        }
      ]
    },
    {
      "id": 867359,
      "postDate": "2020-05-30T07:51:11.127Z",
      "content": "<p>Did you experiment with EfficientNetB7? How long does your current approach take to train?</p>",
      "rawMarkdown": "Did you experiment with EfficientNetB7? How long does your current approach take to train?",
      "replies": [
        {
          "id": 867369,
          "postDate": "2020-05-30T08:12:31.400Z",
          "content": "<p>I never dream about b7😂 , now my model takes ~45min/epoch on kaggle kernel gpu. I use fast learn.to_fp16() and then learn.fit</p>",
          "rawMarkdown": "I never dream about b7😂 , now my model takes ~45min/epoch on kaggle kernel gpu. I use fast learn.to_fp16() and then learn.fit"
        }
      ]
    },
    {
      "id": 872539,
      "postDate": "2020-06-03T09:36:25.890Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 867375,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-05-30T08:21:29.007000",
      "content": "<p>Hello <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a>, I'm using a B0 (0.87 LB). </p>\n\n<p>About # epochs: Perhaps increase the number of epochs you train (you will see fluctuations in the CV/LB score, but on average it increases (at least for me)). 5 epochs seems very little.\nAbout training run time: I think it could be a good idea to precompute all the tile/patch-coordinates for all images (and save it to file). Then feed the precomputed coordinates to/through your pipeline together with the corresponding images.</p>\n\n<p>Further, for me max pooling isn't working well. For some reason my loss get really high. I need to use only the average pooling layer.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 867410,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-05-30T08:58:49.020000",
          "content": "<p><a href=\"/akensert\">@akensert</a> Thanks for your advice! :) I will train more epochs and precompute(it seems should be done locally and upload to kaggle? I try in kaggle kernel to save as .npy, but get disk error and kernel crashed)\nMaxpool doesn't perform well for me, too.\nWhat's more, b0 get 0.87 is so impressive, what's the key to this if you don'y mind disclosure?(about network architecture/augmentations/post-process?) No need to specific details, just a direction is ok.😄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 867471,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-05-30T10:28:12.500000",
          "content": "<p>No problem :-)</p>\n\n<p>I save it as .npy file, and this file should not be that big (few MB). You only need to save the y[_start] and x[_start] coordninates (so e.g. (10616, 16, 2) shaped array)</p>\n\n<p>About your last question: I think by simply train for many more epochs will help you get much better results (you just have to suffer through ~24 hours of training :-D). Other adjustments could be to try more tiles/patches, somewhat lower learning rate, and image augmentation. Remember these are just some suggestions, that may, or may not, improve the performance of your models.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 867512,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-05-30T11:31:47.053000",
          "content": "<p>&gt; (10616, 16, 2) shaped array</p>\n\n<p>! Great! I get the meaning, puzzle with the \"coordinates\" first time, but now understand. This is really a smart approach💯💯💯 However, it should also be padded, right? And just jump the argsort part to save time?(correct me if I miss sth important)\nThanks for your sharing👍👍👍 I will update the improvement about more epochs or others after dayssssss</p>\n\n<p>&gt; you just have to suffer through ~24 hours of training :-D</p>\n\n<p>suffer soon😆 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 867961,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-05-30T19:20:58.370000",
          "content": "<p>Great :-) Yes you can pad it (I guess it depends on your pipeline etc, if you need/should do it or not). Good luck!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 868124,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-05-31T00:28:42.273000",
          "content": "<p>Thanks for your kindness and patience! I did it yesterday and save as .npy as expected, can't wait to see the difference!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 867546,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2020-05-30T11:59:23.257000",
      "content": "<p>hey 曾，\nWe are offering free GPU for Chinese, feel free to join our wechat group, our product is at <a href=\"https://featurize.cn\">https://featurize.cn</a> . You can find the QRCode of our WeChat after you signing in.\nBesides, we are just launched this competition too.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 867589,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-05-30T12:42:02.553000",
          "content": "<p>Sounds attractive! I will check it later :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 867430,
      "author_name": "Andreas Akarepis",
      "author_url": "",
      "post_date": "2020-05-30T09:26:56.160000",
      "content": "<p>Interesting</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 874253,
      "author_name": "Rajnish Chauhan",
      "author_url": "",
      "post_date": "2020-06-04T18:46:37.630000",
      "content": "<p>what you can do is to get little more boost is that dont use full dataset to train .\nUse just top 7K in one fold and  then bottom 7K records .\nDo train in 2 folds and blend.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 874407,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-05T00:38:50.813000",
          "content": "<p>What's the meaning of top 7k and bottom 7k? Is that you mean use like the pseudo code shows?\n<code>\nfold_1 = dataset[:7k], fold_2 = dataset[-7k:]\ntrain each fold until converged\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 875389,
          "author_name": "Rajnish Chauhan",
          "author_url": "",
          "post_date": "2020-06-05T18:32:39.367000",
          "content": "<p>yes , as above </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 875578,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-06T00:05:50.453000",
          "content": "<p>Thanks for your reply :D</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 867404,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-05-30T08:53:30.907000",
      "content": "<p>Hi !\nI've tried so many things (augmentation, models, losses, learning rates, even some new implementation using LSTM or bagging) and I was still at 0.79. \nThen increasing the number of tiles significantly boosted my score.  </p>",
      "votes": 2,
      "replies": [
        {
          "id": 867414,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-05-30T09:04:35.353000",
          "content": "<p>It makes sense! More tiles contains more information. As you mentioned, your network architecture seems contains many customized layers(except the backbone)?\nThe number of tiles I choose(16) seems a bit small as I check some other number.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 867977,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-05-30T19:38:43.627000",
          "content": "<p>I'm using a standard resnext, no custom layers. Increasing the number of tiles could help you</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 868121,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-05-31T00:25:15.477000",
          "content": "<p>Thanks for your kindness!💯  I plan to train more epochs and then increase the number of tiles. =)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 874722,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-06-05T08:54:01.913000",
      "content": "<p>Hi thanks for sharing. \nDid GEM pooling significantly boost your score over other pooling methods ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 874753,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-05T09:18:14.140000",
          "content": "<p>This topic is just based on 5~10 epochs, so I can't said what GeM will perform further.\nBut GeM boost ~1e-04 on my local score(over others based on 5~10 epochs), and it give me the most consistent trend between train loss and valid loss(even I train more epochs), so I choose it :) Because I choose GeM in 5~10 epochs, so I did't try other pool layers for more epochs.😄  Hope it's useful for you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 874775,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-06-05T09:38:56.517000",
          "content": "<p>good to know, thanks 👍 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 868899,
      "author_name": "Rajnish Chauhan",
      "author_url": "",
      "post_date": "2020-05-31T15:01:48.747000",
      "content": "<p><a href=\"/akensert\">@akensert</a> how we can input asymmetric image created form 128x128x36, by cocatenation image will be of size 512 x 1152 (approx),  Thanks for help in advance.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 872142,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-03T00:12:25.383000",
          "content": "<p><a href=\"/rajnishe\">@rajnishe</a> I don't think asymmetric image is a problem, as the CNN can also accept the different size of the image what feeds in it(you can just fees them to your CNN model, but all the data should have the same size, or it will cause error when concatenate some images into a batch). Can my answer solve your problem?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 872580,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-06-03T10:22:51.700000",
          "content": "<p><a href=\"/rajnishe\">@rajnishe</a> I make sure it's symmetric. And 128x128x3x36 would be 768x768x3 no?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873780,
          "author_name": "Rajnish Chauhan",
          "author_url": "",
          "post_date": "2020-06-04T12:25:20.917000",
          "content": "<p>yes , thanks <a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873783,
          "author_name": "Rajnish Chauhan",
          "author_url": "",
          "post_date": "2020-06-04T12:26:10.743000",
          "content": "<p>thanks <a href=\"/akensert\">@akensert</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873831,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-04T13:11:15.790000",
          "content": "<p>😄 You are welcome.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 867359,
      "author_name": "Kurian Benoy",
      "author_url": "",
      "post_date": "2020-05-30T07:51:11.127000",
      "content": "<p>Did you experiment with EfficientNetB7? How long does your current approach take to train?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 867369,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-05-30T08:12:31.400000",
          "content": "<p>I never dream about b7😂 , now my model takes ~45min/epoch on kaggle kernel gpu. I use fast learn.to_fp16() and then learn.fit</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 872539,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-03T09:36:25.890000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "867317": "As experiment some thoughts. I boost my score to cv/lb 0.807475/0.81. Here I wanna show a part of ways that boost my score.(not all of course) And the main reason why I post this topic is 抛砖引玉(I don't know how to translate it as many tools' translations are not so accurate)\nMy model forward is: backbone.features -&gt; pool -&gt; dropout(0.2) -&gt; linear(flatten(), 6).\n- Model\n    - efficientnet b0\n- Pool layer\n    - GeM &gt; concat(max+avg) &gt; avg &gt; max\n- Data source\n    - [panda-2020-level-1-2](https://www.kaggle.com/lopuhin/panda-2020-level-1-2)(cv2.imread from this dataset is faster than skimage.io.MultiImage read from original competition data if my code isn't wrong, but just a little faster)\n- Data process\n    - level 1 resolution\n    - tiles, like what iafoss did(the idea is so great and code is so tidy!)\n    - transform each tile then combine into one image &gt; tiles &gt; combine tiles into one image then transform it(the whole image), theoretical analysis which I accepted is [List of tiles or large single image made of tiles?](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/153872) on comment by iafoss\n    - augmentations\n        - flip + oneof(rotate, randomrotate90)\n        - brightness, contrast, ssr, channelshuffle aren't work for me\n- Training details\n    - random split, 80:20\n    - 16 tiles(transform each then combine to one image(4tilesx4tiles))\n    - bs: 4\n    - epochs: 5(more cause cv/lb decrease for me, maybe b0 is too samll?)\n    - lr-scheduler: none, lr: 5e-04, no tta\n    - save model weights by score(the competition metric)\n    - hardware: kaggle kernel(plan to rent a 2080ti, or 40min/epoch for b0 is too slow...)\n\nSome questions\n- Has anyone try tpu? I try tpu by torch-xla(xm.spawn), but it crash after committing it ~10k secs. And not so fast as I except(I sample 1k image, and it run ~4min/epoch, but my model on gpu just cost 45min/epoch while it runs 10k image...) I use 1 core.\n- Is there any way to speed up doing tiles/transform with np/torch/albumentations? TF has, mentioned [rotation-augmentation-gpu-tpu-0-96](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96) by Chris\n- proper way to train multi folds?(I failed in using parallel on tpu by modifing the [panda-challenge-resnet-multitask-8-fold-on-tpu](https://www.kaggle.com/tarunpaparaju/panda-challenge-resnet-multitask-8-fold-on-tpu) )\n- Which model do you choose for testing augmentations/ideas(it seems res18 faster than b0, I plan to turn to res18)? And the result is consistent or not while you trun to deeper model?\n- What's your training time(model? hardware?)?\n\nFeel free to discuss anything that properly, and thank you for your coming to see my topic.\nWhat's your ways to boost score?(not all of course, because it's a competition)",
    "867375": "Hello @cnzengshiyuan, I'm using a B0 (0.87 LB). \n\nAbout # epochs: Perhaps increase the number of epochs you train (you will see fluctuations in the CV/LB score, but on average it increases (at least for me)). 5 epochs seems very little.\nAbout training run time: I think it could be a good idea to precompute all the tile/patch-coordinates for all images (and save it to file). Then feed the precomputed coordinates to/through your pipeline together with the corresponding images.\n\nFurther, for me max pooling isn't working well. For some reason my loss get really high. I need to use only the average pooling layer.",
    "867546": "hey 曾，\nWe are offering free GPU for Chinese, feel free to join our wechat group, our product is at https://featurize.cn . You can find the QRCode of our WeChat after you signing in.\nBesides, we are just launched this competition too.",
    "867430": "Interesting",
    "874253": "what you can do is to get little more boost is that dont use full dataset to train .\nUse just top 7K in one fold and  then bottom 7K records .\nDo train in 2 folds and blend.\n",
    "867404": "Hi !\nI've tried so many things (augmentation, models, losses, learning rates, even some new implementation using LSTM or bagging) and I was still at 0.79. \nThen increasing the number of tiles significantly boosted my score.  ",
    "874722": "Hi thanks for sharing. \nDid GEM pooling significantly boost your score over other pooling methods ?",
    "868899": "@akensert how we can input asymmetric image created form 128x128x36, by cocatenation image will be of size 512 x 1152 (approx),  Thanks for help in advance.\n",
    "867359": "Did you experiment with EfficientNetB7? How long does your current approach take to train?",
    "872539": ""
  }
}