{
  "id": 170979,
  "title": "Questions from a newbie - Score and process to submission",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/170979",
  "author_name": "Hernoo",
  "post_date": "2020-07-29T21:05:03.856000",
  "votes": 0,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>as said, I'm a newbie, but I'd like to learn a lot from this competition.</p>\n\n<p>I tried by using a dataset with tiles (512x512, thanks to the author !) and a ResNet50 to train my model. I finish with GPU after 100 epochs around 0.97 accuracy. Most likely I overfit there, but to know if I do i'd like to submit.\n1) First question : do you agree that with such a simple model I most likely overfit ? Any other reasons for such scores otherwise?</p>\n\n<p>I'd like to submit, but here I'm lost. I read loads of topics, and it seems to be a bit more cumbersome than in others in this competition. What I understood is : when submitting (!= comitting) the secret folder with test_images will appear. Great. But It also means that I have tio pre-process all the images ?\n2) if so, should I just pre process and write them in a folder that I'll access afterwards?\n3) some authors showed a \"minimal notebook\" to submit, and save their model to be loaded in this specific notebook. From what I understood it's neater, easier, but not necessary right ? I can train and \"on-the-go\" predict after having preprocessed the test images, right? </p>\n\n<p>I'm sorry to ask silly questions, this is a lot of discoveries to me : deep learning, kaggle, image processing and this submission system.</p>\n\n<p>Thank you for your help ++ \nHernoo</p>",
  "messages": [
    {
      "id": 981238,
      "postDate": "2020-08-22T10:03:13.067Z",
      "content": "<p>Hello all,</p>\n<p>I'm now a bit more confident in what I have done. I have a training model which on training data reaches 0.55 QWK, so I'm expecting more than 0 in the test set !</p>\n<p>However, I'm having a issue now about submitting : it goes all the time in Timeout.</p>\n<p>If I understood well, there's about 1000 tiff images roughly the same as the ones we got for training. The preprocessing algorithm (tiles) roughly takes 2 minutes for 50 images tiff to be preprocessed and predicted by the algorithm. Then my submitting process should take ~40 minutes to do the 1000 images ? It's not. <br>\nI have used what you gave me <a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> : a if/else with only the directory changing, so it should work smoothly, but it's not.</p>\n<p>Do you guys have a clue of what could happen behind the scene? </p>\n<p>thanks ++ </p>",
      "rawMarkdown": "Hello all,\n\nI'm now a bit more confident in what I have done. I have a training model which on training data reaches 0.55 QWK, so I'm expecting more than 0 in the test set !\n\nHowever, I'm having a issue now about submitting : it goes all the time in Timeout.\n\nIf I understood well, there's about 1000 tiff images roughly the same as the ones we got for training. The preprocessing algorithm (tiles) roughly takes 2 minutes for 50 images tiff to be preprocessed and predicted by the algorithm. Then my submitting process should take ~40 minutes to do the 1000 images ? It's not. \nI have used what you gave me @arroqc : a if/else with only the directory changing, so it should work smoothly, but it's not.\n\nDo you guys have a clue of what could happen behind the scene? \n\nthanks ++ "
    },
    {
      "id": 956527,
      "postDate": "2020-08-03T15:25:10.007Z",
      "content": "<p>Hi,\nI'm sorry but I'm still stuck in the process of submitting.\nI have read a lot of different discussions in here, and still after ~30 minutes of checking, I have a CSV not found.\nI tried with a random value csv which works, (although I logically get a score of 0), so the process in itself is fine.</p>\n\n<p>I think I have an issue in my code, and I suspect the preprocessing to be the reason for this issue.\nMy code has the <code>if os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):</code>\nand thanks to it I process either the train or the test set.</p>\n\n<p>Can you create a directory and save your preprocessed images to be used by the algorithm for prediction afterwards?\nCan one juste have a glance to my notebook and tell me where could the my issue? <a href=\"https://www.kaggle.com/hernoo/submission-panda\">https://www.kaggle.com/hernoo/submission-panda</a></p>\n\n<p>would be of a great help !</p>\n\n<p>Thanks !</p>",
      "rawMarkdown": "Hi,\nI'm sorry but I'm still stuck in the process of submitting.\nI have read a lot of different discussions in here, and still after ~30 minutes of checking, I have a CSV not found.\nI tried with a random value csv which works, (although I logically get a score of 0), so the process in itself is fine.\n\nI think I have an issue in my code, and I suspect the preprocessing to be the reason for this issue.\nMy code has the `if os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):`\nand thanks to it I process either the train or the test set.\n\nCan you create a directory and save your preprocessed images to be used by the algorithm for prediction afterwards?\nCan one juste have a glance to my notebook and tell me where could the my issue? [https://www.kaggle.com/hernoo/submission-panda](https://www.kaggle.com/hernoo/submission-panda)\n\nwould be of a great help !\n\nThanks !\n\n",
      "replies": [
        {
          "id": 962139,
          "postDate": "2020-08-07T20:59:23.340Z",
          "content": "<p>Your pipeline should be able to handle the train set. If it works perfectly for train set with no bug then just changing the dataframe and the image directory should work.</p>\n<p>My suggestion is to change this:<br>\n<code>if os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):</code><br>\nto this<br>\n<code>if os.path.exists(IMAGE_FOLDER):</code></p>\n<p>Remove this:<br>\n<code>testcsv=pd.read_csv('../input/prostate-cancer-grade-assessment/test.csv')</code></p>\n<p>and finaly at the start of the kernel use:</p>\n<pre><code>if os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):\n    testcsv=pd.read_csv('../input/prostate-cancer-grade-assessment/test.csv')\n    IMAGE_FOLDER = '../input/prostate-cancer-grade-assessment/test_images'\nelse:\n    testcsv=pd.read_csv('../input/prostate-cancer-grade-assessment/train.csv')[:16]\n    IMAGE_FOLDER = '../input/prostate-cancer-grade-assessment/test_images'\n</code></pre>\n<p>Then when commiting it should run with the trainset (16 images of it) so that when submitting it does the transition to test set seemlessly. If there is a bug you'll know at commit stage (or just by running the kernel in edit mode).</p>",
          "rawMarkdown": "Your pipeline should be able to handle the train set. If it works perfectly for train set with no bug then just changing the dataframe and the image directory should work.\n\nMy suggestion is to change this:\n`if os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):`\nto this\n`if os.path.exists(IMAGE_FOLDER):`\n\nRemove this:\n`testcsv=pd.read_csv('../input/prostate-cancer-grade-assessment/test.csv')`\n\nand finaly at the start of the kernel use:\n```\nif os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):\n    testcsv=pd.read_csv('../input/prostate-cancer-grade-assessment/test.csv')\n    IMAGE_FOLDER = '../input/prostate-cancer-grade-assessment/test_images'\nelse:\n    testcsv=pd.read_csv('../input/prostate-cancer-grade-assessment/train.csv')[:16]\n    IMAGE_FOLDER = '../input/prostate-cancer-grade-assessment/test_images'\n```\n\nThen when commiting it should run with the trainset (16 images of it) so that when submitting it does the transition to test set seemlessly. If there is a bug you'll know at commit stage (or just by running the kernel in edit mode)."
        },
        {
          "id": 962740,
          "postDate": "2020-08-08T11:53:22.373Z",
          "content": "<p>Thanks Arnaud, very appreciated !\nIt seems that I managed to do something, and even though my pipeline was rather simple (tile processing + training with a ResNet50) I'm a bit disappointed by the scores I got :/\nPrivate Score\n-0.00961\nPublic Score\n-0.00945</p>\n\n<p>Do you think it's more a question or inefficient training or a bug in my notebook that can cause that? </p>\n\n<p>Best</p>",
          "rawMarkdown": "Thanks Arnaud, very appreciated !\nIt seems that I managed to do something, and even though my pipeline was rather simple (tile processing + training with a ResNet50) I'm a bit disappointed by the scores I got :/\nPrivate Score\n-0.00961\nPublic Score\n-0.00945\n\nDo you think it's more a question or inefficient training or a bug in my notebook that can cause that? \n\nBest\n\n"
        },
        {
          "id": 962992,
          "postDate": "2020-08-08T15:31:31.243Z",
          "content": "<p>Could be a bunch of reasons:</p>\n<ul>\n<li>Bug in training </li>\n<li>Bug in inference<br>\n=&gt; Try to verify your model is doing good on a validation set using your inference pipeline</li>\n<li>Some case of data leakage or terrible overfitting</li>\n</ul>",
          "rawMarkdown": "Could be a bunch of reasons:\n* Bug in training \n* Bug in inference\n=&gt; Try to verify your model is doing good on a validation set using your inference pipeline\n* Some case of data leakage or terrible overfitting",
          "votes": 1
        },
        {
          "id": 967821,
          "postDate": "2020-08-12T14:24:55.463Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> I'll try that </p>",
          "rawMarkdown": "Thanks @arroqc I'll try that "
        }
      ]
    },
    {
      "id": 951799,
      "postDate": "2020-07-30T12:22:52.440Z",
      "content": "<p>Thanks <a href=\"/yukkyo\">@yukkyo</a> for your answer ! \nI'm trying to submit indeed but I'm having issues the last attempt got a \" Submission Scoring Error\nError\nError \" \nwithout more precise messages to help me finding the error ! Any idea how one could troubleshoot that easily? :)</p>",
      "rawMarkdown": "Thanks @yukkyo for your answer ! \nI'm trying to submit indeed but I'm having issues the last attempt got a \" Submission Scoring Error\nError\nError \" \nwithout more precise messages to help me finding the error ! Any idea how one could troubleshoot that easily? :)"
    },
    {
      "id": 951784,
      "postDate": "2020-07-30T12:05:36.130Z",
      "content": "<p>Hi <a href=\"/hernoo\">@hernoo</a> . Welcome to PANDA competition.\nThis is just my personal opinion.</p>\n\n<blockquote>\n  <p>1) First question : do you agree that with such a simple model I most likely overfit ? Any other reasons for such scores otherwise?</p>\n</blockquote>\n\n<p>No. In general, a simple model is harder to overfit than a complex model.\nI suspect two things\n- The split in CV is not good (not considering duplicate images)\n- The metrics implementation is wrong (try using QWK as well as accuracy)</p>\n\n<p>Of course, the quickest way to try is to SUBMIT.</p>\n\n<blockquote>\n  <p>2) if so, should I just pre process and write them in a folder that I'll access afterwards?</p>\n</blockquote>\n\n<p>Test data cannot be written to a place that we can access later.</p>\n\n<blockquote>\n  <p>3) some authors showed a \"minimal notebook\" to submit, and save their model to be loaded in this specific notebook. From what I understood it's neater, easier, but not necessary right ? I can train and \"on-the-go\" predict after having preprocessed the test images, right?</p>\n</blockquote>\n\n<p>It is easier to exceed the time limit when training and inference are done together. This is especially noticeable if you're doing an ensemble.</p>\n\n<p>You'll probably learn far more by running it yourself than by reading this discussion. <br>\nI wish you good luck.</p>",
      "rawMarkdown": "Hi @hernoo . Welcome to PANDA competition.\nThis is just my personal opinion.\n\n&gt; 1) First question : do you agree that with such a simple model I most likely overfit ? Any other reasons for such scores otherwise?\n\nNo. In general, a simple model is harder to overfit than a complex model.\nI suspect two things\n- The split in CV is not good (not considering duplicate images)\n- The metrics implementation is wrong (try using QWK as well as accuracy)\n\nOf course, the quickest way to try is to SUBMIT.\n\n&gt; 2) if so, should I just pre process and write them in a folder that I'll access afterwards?\n\nTest data cannot be written to a place that we can access later.\n\n&gt; 3) some authors showed a \"minimal notebook\" to submit, and save their model to be loaded in this specific notebook. From what I understood it's neater, easier, but not necessary right ? I can train and \"on-the-go\" predict after having preprocessed the test images, right?\n\nIt is easier to exceed the time limit when training and inference are done together. This is especially noticeable if you're doing an ensemble.\n\nYou'll probably learn far more by running it yourself than by reading this discussion.  \nI wish you good luck."
    },
    {
      "id": 951065,
      "postDate": "2020-07-29T21:05:03.857Z",
      "content": "<p>Hi,</p>\n\n<p>as said, I'm a newbie, but I'd like to learn a lot from this competition.</p>\n\n<p>I tried by using a dataset with tiles (512x512, thanks to the author !) and a ResNet50 to train my model. I finish with GPU after 100 epochs around 0.97 accuracy. Most likely I overfit there, but to know if I do i'd like to submit.\n1) First question : do you agree that with such a simple model I most likely overfit ? Any other reasons for such scores otherwise?</p>\n\n<p>I'd like to submit, but here I'm lost. I read loads of topics, and it seems to be a bit more cumbersome than in others in this competition. What I understood is : when submitting (!= comitting) the secret folder with test_images will appear. Great. But It also means that I have tio pre-process all the images ?\n2) if so, should I just pre process and write them in a folder that I'll access afterwards?\n3) some authors showed a \"minimal notebook\" to submit, and save their model to be loaded in this specific notebook. From what I understood it's neater, easier, but not necessary right ? I can train and \"on-the-go\" predict after having preprocessed the test images, right? </p>\n\n<p>I'm sorry to ask silly questions, this is a lot of discoveries to me : deep learning, kaggle, image processing and this submission system.</p>\n\n<p>Thank you for your help ++ \nHernoo</p>",
      "rawMarkdown": "Hi,\n\nas said, I'm a newbie, but I'd like to learn a lot from this competition.\n\nI tried by using a dataset with tiles (512x512, thanks to the author !) and a ResNet50 to train my model. I finish with GPU after 100 epochs around 0.97 accuracy. Most likely I overfit there, but to know if I do i'd like to submit.\n1) First question : do you agree that with such a simple model I most likely overfit ? Any other reasons for such scores otherwise?\n\nI'd like to submit, but here I'm lost. I read loads of topics, and it seems to be a bit more cumbersome than in others in this competition. What I understood is : when submitting (!= comitting) the secret folder with test_images will appear. Great. But It also means that I have tio pre-process all the images ?\n2) if so, should I just pre process and write them in a folder that I'll access afterwards?\n3) some authors showed a \"minimal notebook\" to submit, and save their model to be loaded in this specific notebook. From what I understood it's neater, easier, but not necessary right ? I can train and \"on-the-go\" predict after having preprocessed the test images, right? \n\nI'm sorry to ask silly questions, this is a lot of discoveries to me : deep learning, kaggle, image processing and this submission system.\n\nThank you for your help ++ \nHernoo"
    }
  ],
  "comments": [
    {
      "id": 981238,
      "author_name": "Hernoo",
      "author_url": "",
      "post_date": "2020-08-22T10:03:13.067000",
      "content": "<p>Hello all,</p>\n<p>I'm now a bit more confident in what I have done. I have a training model which on training data reaches 0.55 QWK, so I'm expecting more than 0 in the test set !</p>\n<p>However, I'm having a issue now about submitting : it goes all the time in Timeout.</p>\n<p>If I understood well, there's about 1000 tiff images roughly the same as the ones we got for training. The preprocessing algorithm (tiles) roughly takes 2 minutes for 50 images tiff to be preprocessed and predicted by the algorithm. Then my submitting process should take ~40 minutes to do the 1000 images ? It's not. <br>\nI have used what you gave me <a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> : a if/else with only the directory changing, so it should work smoothly, but it's not.</p>\n<p>Do you guys have a clue of what could happen behind the scene? </p>\n<p>thanks ++ </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 956527,
      "author_name": "Hernoo",
      "author_url": "",
      "post_date": "2020-08-03T15:25:10.007000",
      "content": "<p>Hi,\nI'm sorry but I'm still stuck in the process of submitting.\nI have read a lot of different discussions in here, and still after ~30 minutes of checking, I have a CSV not found.\nI tried with a random value csv which works, (although I logically get a score of 0), so the process in itself is fine.</p>\n\n<p>I think I have an issue in my code, and I suspect the preprocessing to be the reason for this issue.\nMy code has the <code>if os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):</code>\nand thanks to it I process either the train or the test set.</p>\n\n<p>Can you create a directory and save your preprocessed images to be used by the algorithm for prediction afterwards?\nCan one juste have a glance to my notebook and tell me where could the my issue? <a href=\"https://www.kaggle.com/hernoo/submission-panda\">https://www.kaggle.com/hernoo/submission-panda</a></p>\n\n<p>would be of a great help !</p>\n\n<p>Thanks !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 962139,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-07T20:59:23.340000",
          "content": "<p>Your pipeline should be able to handle the train set. If it works perfectly for train set with no bug then just changing the dataframe and the image directory should work.</p>\n<p>My suggestion is to change this:<br>\n<code>if os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):</code><br>\nto this<br>\n<code>if os.path.exists(IMAGE_FOLDER):</code></p>\n<p>Remove this:<br>\n<code>testcsv=pd.read_csv('../input/prostate-cancer-grade-assessment/test.csv')</code></p>\n<p>and finaly at the start of the kernel use:</p>\n<pre><code>if os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):\n    testcsv=pd.read_csv('../input/prostate-cancer-grade-assessment/test.csv')\n    IMAGE_FOLDER = '../input/prostate-cancer-grade-assessment/test_images'\nelse:\n    testcsv=pd.read_csv('../input/prostate-cancer-grade-assessment/train.csv')[:16]\n    IMAGE_FOLDER = '../input/prostate-cancer-grade-assessment/test_images'\n</code></pre>\n<p>Then when commiting it should run with the trainset (16 images of it) so that when submitting it does the transition to test set seemlessly. If there is a bug you'll know at commit stage (or just by running the kernel in edit mode).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 962740,
          "author_name": "Hernoo",
          "author_url": "",
          "post_date": "2020-08-08T11:53:22.373000",
          "content": "<p>Thanks Arnaud, very appreciated !\nIt seems that I managed to do something, and even though my pipeline was rather simple (tile processing + training with a ResNet50) I'm a bit disappointed by the scores I got :/\nPrivate Score\n-0.00961\nPublic Score\n-0.00945</p>\n\n<p>Do you think it's more a question or inefficient training or a bug in my notebook that can cause that? </p>\n\n<p>Best</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 962992,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-08T15:31:31.243000",
          "content": "<p>Could be a bunch of reasons:</p>\n<ul>\n<li>Bug in training </li>\n<li>Bug in inference<br>\n=&gt; Try to verify your model is doing good on a validation set using your inference pipeline</li>\n<li>Some case of data leakage or terrible overfitting</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 967821,
          "author_name": "Hernoo",
          "author_url": "",
          "post_date": "2020-08-12T14:24:55.463000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> I'll try that </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 951799,
      "author_name": "Hernoo",
      "author_url": "",
      "post_date": "2020-07-30T12:22:52.440000",
      "content": "<p>Thanks <a href=\"/yukkyo\">@yukkyo</a> for your answer ! \nI'm trying to submit indeed but I'm having issues the last attempt got a \" Submission Scoring Error\nError\nError \" \nwithout more precise messages to help me finding the error ! Any idea how one could troubleshoot that easily? :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 951784,
      "author_name": "fam_taro",
      "author_url": "",
      "post_date": "2020-07-30T12:05:36.130000",
      "content": "<p>Hi <a href=\"/hernoo\">@hernoo</a> . Welcome to PANDA competition.\nThis is just my personal opinion.</p>\n\n<blockquote>\n  <p>1) First question : do you agree that with such a simple model I most likely overfit ? Any other reasons for such scores otherwise?</p>\n</blockquote>\n\n<p>No. In general, a simple model is harder to overfit than a complex model.\nI suspect two things\n- The split in CV is not good (not considering duplicate images)\n- The metrics implementation is wrong (try using QWK as well as accuracy)</p>\n\n<p>Of course, the quickest way to try is to SUBMIT.</p>\n\n<blockquote>\n  <p>2) if so, should I just pre process and write them in a folder that I'll access afterwards?</p>\n</blockquote>\n\n<p>Test data cannot be written to a place that we can access later.</p>\n\n<blockquote>\n  <p>3) some authors showed a \"minimal notebook\" to submit, and save their model to be loaded in this specific notebook. From what I understood it's neater, easier, but not necessary right ? I can train and \"on-the-go\" predict after having preprocessed the test images, right?</p>\n</blockquote>\n\n<p>It is easier to exceed the time limit when training and inference are done together. This is especially noticeable if you're doing an ensemble.</p>\n\n<p>You'll probably learn far more by running it yourself than by reading this discussion. <br>\nI wish you good luck.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "981238": "Hello all,\n\nI'm now a bit more confident in what I have done. I have a training model which on training data reaches 0.55 QWK, so I'm expecting more than 0 in the test set !\n\nHowever, I'm having a issue now about submitting : it goes all the time in Timeout.\n\nIf I understood well, there's about 1000 tiff images roughly the same as the ones we got for training. The preprocessing algorithm (tiles) roughly takes 2 minutes for 50 images tiff to be preprocessed and predicted by the algorithm. Then my submitting process should take ~40 minutes to do the 1000 images ? It's not. \nI have used what you gave me @arroqc : a if/else with only the directory changing, so it should work smoothly, but it's not.\n\nDo you guys have a clue of what could happen behind the scene? \n\nthanks ++ ",
    "956527": "Hi,\nI'm sorry but I'm still stuck in the process of submitting.\nI have read a lot of different discussions in here, and still after ~30 minutes of checking, I have a CSV not found.\nI tried with a random value csv which works, (although I logically get a score of 0), so the process in itself is fine.\n\nI think I have an issue in my code, and I suspect the preprocessing to be the reason for this issue.\nMy code has the `if os.path.exists('../input/prostate-cancer-grade-assessment/test_images'):`\nand thanks to it I process either the train or the test set.\n\nCan you create a directory and save your preprocessed images to be used by the algorithm for prediction afterwards?\nCan one juste have a glance to my notebook and tell me where could the my issue? [https://www.kaggle.com/hernoo/submission-panda](https://www.kaggle.com/hernoo/submission-panda)\n\nwould be of a great help !\n\nThanks !\n\n",
    "951799": "Thanks @yukkyo for your answer ! \nI'm trying to submit indeed but I'm having issues the last attempt got a \" Submission Scoring Error\nError\nError \" \nwithout more precise messages to help me finding the error ! Any idea how one could troubleshoot that easily? :)",
    "951784": "Hi @hernoo . Welcome to PANDA competition.\nThis is just my personal opinion.\n\n&gt; 1) First question : do you agree that with such a simple model I most likely overfit ? Any other reasons for such scores otherwise?\n\nNo. In general, a simple model is harder to overfit than a complex model.\nI suspect two things\n- The split in CV is not good (not considering duplicate images)\n- The metrics implementation is wrong (try using QWK as well as accuracy)\n\nOf course, the quickest way to try is to SUBMIT.\n\n&gt; 2) if so, should I just pre process and write them in a folder that I'll access afterwards?\n\nTest data cannot be written to a place that we can access later.\n\n&gt; 3) some authors showed a \"minimal notebook\" to submit, and save their model to be loaded in this specific notebook. From what I understood it's neater, easier, but not necessary right ? I can train and \"on-the-go\" predict after having preprocessed the test images, right?\n\nIt is easier to exceed the time limit when training and inference are done together. This is especially noticeable if you're doing an ensemble.\n\nYou'll probably learn far more by running it yourself than by reading this discussion.  \nI wish you good luck.",
    "951065": "Hi,\n\nas said, I'm a newbie, but I'd like to learn a lot from this competition.\n\nI tried by using a dataset with tiles (512x512, thanks to the author !) and a ResNet50 to train my model. I finish with GPU after 100 epochs around 0.97 accuracy. Most likely I overfit there, but to know if I do i'd like to submit.\n1) First question : do you agree that with such a simple model I most likely overfit ? Any other reasons for such scores otherwise?\n\nI'd like to submit, but here I'm lost. I read loads of topics, and it seems to be a bit more cumbersome than in others in this competition. What I understood is : when submitting (!= comitting) the secret folder with test_images will appear. Great. But It also means that I have tio pre-process all the images ?\n2) if so, should I just pre process and write them in a folder that I'll access afterwards?\n3) some authors showed a \"minimal notebook\" to submit, and save their model to be loaded in this specific notebook. From what I understood it's neater, easier, but not necessary right ? I can train and \"on-the-go\" predict after having preprocessed the test images, right? \n\nI'm sorry to ask silly questions, this is a lot of discoveries to me : deep learning, kaggle, image processing and this submission system.\n\nThank you for your help ++ \nHernoo"
  }
}