{
  "id": 420251,
  "title": "Code submission requirements and how to run sessions for a long time",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/420251",
  "author_name": "Sashi",
  "post_date": "2023-06-29T22:56:30.108000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello, <br>\nA few newbie questions:</p>\n<ol>\n<li><p>The code submission rules states that internet access is disabled but also allows pre-trained models? Is the expectation that we download the pre-trained models and add it as a \"data source\" so that we can use it without internet when they run the notebook?</p></li>\n<li><p>I see some public notebooks use <code>pip install segmentation_models_pytorch</code>. Are these notebooks only for illustration - i.e. how does <code>pip install</code> work without network access.</p></li>\n<li><p>I am not able run the session in the Kaggle notebook beyond the duration of when I am logged into the laptop. If I am idle, I get a prompt that asks me whether I want to continue editing and if I respond after a long time, the session is dead and has to be restarted. Any suggestions on how to keep the Kaggle notebook running?</p></li>\n</ol>\n<p>Thank you,</p>",
  "messages": [
    {
      "id": 2323795,
      "postDate": "2023-06-30T06:55:46.307Z",
      "content": "<p>Hi! For questions 1 and 2 my general approach is </p>\n<ol>\n<li>Create inference kaggle notebook with no internet access and pinned environment version (so it will not break at some point with version update from kaggle)</li>\n<li>Download all the packages in *.whl (or *.tar.gz) format from your python environment via <code>pip list --format=freeze &gt; requirements.txt &amp;&amp; pip download -r requirements.txt --no-deps --no-binary=:all: -d .</code>, add it to zip archive and upload to kaggle dataset</li>\n<li>Upload your model checkpoints / state dicts to separate kaggle dataset</li>\n<li>Add the software dataset and model dataset to your inference notebook and use them without internet: <code>!pip install --no-index --find-links ../input/your-software-dataset-name/ -r requirements.txt</code> and load weights from your model dataset with your favorite ML framework</li>\n</ol>\n<p>If some additional system packages are required, similar steps could be performed with <code>apt download package_name</code>, uploading to dataset and installing with <code>!dpkg -i ../input/your-software-dataset-name/package_name.deb</code>.</p>\n<p>You could refer to <a href=\"https://www.kaggle.com/code/mkotyushev/scrolls-inference-agg\" target=\"_blank\">my inference notebook</a> from vesuvius challenge to see this approach in action.</p>\n<p><strong>Some side notes</strong>:</p>\n<ul>\n<li>Kaggle's docker images used in competitions are pretty much packed with almost all the required packages and you could avoid uploading and installing <em>all</em> the packages and e. g. some heavy torch distribution. So just you should try to run your code and see which dependencies are missing.</li>\n<li>There are many public datasets with preinstalled packages, you could search for them in notebook's data tab.</li>\n<li>You need to verify that you model output in kaggle environment is the same (or at least similar) as in your local environment because of different versions of packages. Pay attention to CUDA version in local &amp; kaggle environments.</li>\n<li>kaggle unpacks .zip and .tar.gz files on uploading to dataset, so try to avoid using nested archives (if it is absolutely necessary, you could rename it to .zip.nounzip and rename back in in notebook)</li>\n<li>In addition to all above, I prefer to pack all the development source code to separate dataset to avoid writing / copy-pasting separate inference code in notebook: pack it with <code>git archive --format=zip --output /full/path/to/zipfile.zip master</code> and use in notebook like this <code>!python inference.py --model ../input/your-model-dataset-name/model.pth --data ../input/your-data-dataset-name/ --output ../working/</code></li>\n</ul>\n<p>Good luck in kaggling!</p>",
      "rawMarkdown": "Hi! For questions 1 and 2 my general approach is \n\n1. Create inference kaggle notebook with no internet access and pinned environment version (so it will not break at some point with version update from kaggle)\n2. Download all the packages in *.whl (or *.tar.gz) format from your python environment via `pip list --format=freeze > requirements.txt && pip download -r requirements.txt --no-deps --no-binary=:all: -d .`, add it to zip archive and upload to kaggle dataset\n3. Upload your model checkpoints / state dicts to separate kaggle dataset\n4. Add the software dataset and model dataset to your inference notebook and use them without internet: `!pip install --no-index --find-links ../input/your-software-dataset-name/ -r requirements.txt` and load weights from your model dataset with your favorite ML framework\n\nIf some additional system packages are required, similar steps could be performed with `apt download package_name`, uploading to dataset and installing with `!dpkg -i ../input/your-software-dataset-name/package_name.deb`.\n\nYou could refer to [my inference notebook](https://www.kaggle.com/code/mkotyushev/scrolls-inference-agg) from vesuvius challenge to see this approach in action.\n\n**Some side notes**:\n\n- Kaggle's docker images used in competitions are pretty much packed with almost all the required packages and you could avoid uploading and installing *all* the packages and e. g. some heavy torch distribution. So just you should try to run your code and see which dependencies are missing.\n- There are many public datasets with preinstalled packages, you could search for them in notebook's data tab.\n- You need to verify that you model output in kaggle environment is the same (or at least similar) as in your local environment because of different versions of packages. Pay attention to CUDA version in local & kaggle environments.\n- kaggle unpacks .zip and .tar.gz files on uploading to dataset, so try to avoid using nested archives (if it is absolutely necessary, you could rename it to .zip.nounzip and rename back in in notebook)\n- In addition to all above, I prefer to pack all the development source code to separate dataset to avoid writing / copy-pasting separate inference code in notebook: pack it with `git archive --format=zip --output /full/path/to/zipfile.zip master` and use in notebook like this `!python inference.py --model ../input/your-model-dataset-name/model.pth --data ../input/your-data-dataset-name/ --output ../working/`\n\nGood luck in kaggling!",
      "votes": 3
    },
    {
      "id": 2323387,
      "postDate": "2023-06-29T22:56:30.110Z",
      "content": "<p>Hello, <br>\nA few newbie questions:</p>\n<ol>\n<li><p>The code submission rules states that internet access is disabled but also allows pre-trained models? Is the expectation that we download the pre-trained models and add it as a \"data source\" so that we can use it without internet when they run the notebook?</p></li>\n<li><p>I see some public notebooks use <code>pip install segmentation_models_pytorch</code>. Are these notebooks only for illustration - i.e. how does <code>pip install</code> work without network access.</p></li>\n<li><p>I am not able run the session in the Kaggle notebook beyond the duration of when I am logged into the laptop. If I am idle, I get a prompt that asks me whether I want to continue editing and if I respond after a long time, the session is dead and has to be restarted. Any suggestions on how to keep the Kaggle notebook running?</p></li>\n</ol>\n<p>Thank you,</p>",
      "rawMarkdown": "Hello, \nA few newbie questions:\n\n1. The code submission rules states that internet access is disabled but also allows pre-trained models? Is the expectation that we download the pre-trained models and add it as a \"data source\" so that we can use it without internet when they run the notebook?\n\n2. I see some public notebooks use `pip install segmentation_models_pytorch`. Are these notebooks only for illustration - i.e. how does `pip install` work without network access.\n\n3. I am not able run the session in the Kaggle notebook beyond the duration of when I am logged into the laptop. If I am idle, I get a prompt that asks me whether I want to continue editing and if I respond after a long time, the session is dead and has to be restarted. Any suggestions on how to keep the Kaggle notebook running?\n\nThank you,",
      "votes": 1
    },
    {
      "id": 2323707,
      "postDate": "2023-06-30T06:19:34.823Z",
      "content": "<p>Hello, I've been challenged to Kaggle recently, so I'd like to make some comments so that I can have a better Kaggle life.</p>\n<p>About question 2.<br>\nSince Submit doesn't work over the internet, there are generally two approach.</p>\n<ul>\n<li><strong>A.sys.path.append</strong></li>\n<li><strong>B.whl using Installation using</strong></li>\n</ul>\n<p>refer to the code that uses <strong>smp</strong> with the Internet off.<br>\n<a href=\"https://www.kaggle.com/code/hideyukizushi/contrails-sample-libinstall-internet-off/notebook\" target=\"_blank\">https://www.kaggle.com/code/hideyukizushi/contrails-sample-libinstall-internet-off/notebook</a></p>",
      "rawMarkdown": "Hello, I've been challenged to Kaggle recently, so I'd like to make some comments so that I can have a better Kaggle life.\n\nAbout question 2.\nSince Submit doesn't work over the internet, there are generally two approach.\n* **A.sys.path.append**\n* **B.whl using Installation using**\n\nrefer to the code that uses **smp** with the Internet off.\nhttps://www.kaggle.com/code/hideyukizushi/contrails-sample-libinstall-internet-off/notebook",
      "replies": [
        {
          "id": 2324478,
          "postDate": "2023-06-30T16:10:50.937Z",
          "content": "<p>And if you want to grab the whl files by yourself, you can go to the library's pypi page and go under Download Files (ex.<a href=\"https://pypi.org/project/segmentation-models-pytorch/#files\" target=\"_blank\"> segmentation models pytorch</a>)</p>",
          "rawMarkdown": "And if you want to grab the whl files by yourself, you can go to the library's pypi page and go under Download Files (ex.[ segmentation models pytorch](https://pypi.org/project/segmentation-models-pytorch/#files))"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2323795,
      "author_name": "Mikhail Kotyushev",
      "author_url": "",
      "post_date": "2023-06-30T06:55:46.307000",
      "content": "<p>Hi! For questions 1 and 2 my general approach is </p>\n<ol>\n<li>Create inference kaggle notebook with no internet access and pinned environment version (so it will not break at some point with version update from kaggle)</li>\n<li>Download all the packages in *.whl (or *.tar.gz) format from your python environment via <code>pip list --format=freeze &gt; requirements.txt &amp;&amp; pip download -r requirements.txt --no-deps --no-binary=:all: -d .</code>, add it to zip archive and upload to kaggle dataset</li>\n<li>Upload your model checkpoints / state dicts to separate kaggle dataset</li>\n<li>Add the software dataset and model dataset to your inference notebook and use them without internet: <code>!pip install --no-index --find-links ../input/your-software-dataset-name/ -r requirements.txt</code> and load weights from your model dataset with your favorite ML framework</li>\n</ol>\n<p>If some additional system packages are required, similar steps could be performed with <code>apt download package_name</code>, uploading to dataset and installing with <code>!dpkg -i ../input/your-software-dataset-name/package_name.deb</code>.</p>\n<p>You could refer to <a href=\"https://www.kaggle.com/code/mkotyushev/scrolls-inference-agg\" target=\"_blank\">my inference notebook</a> from vesuvius challenge to see this approach in action.</p>\n<p><strong>Some side notes</strong>:</p>\n<ul>\n<li>Kaggle's docker images used in competitions are pretty much packed with almost all the required packages and you could avoid uploading and installing <em>all</em> the packages and e. g. some heavy torch distribution. So just you should try to run your code and see which dependencies are missing.</li>\n<li>There are many public datasets with preinstalled packages, you could search for them in notebook's data tab.</li>\n<li>You need to verify that you model output in kaggle environment is the same (or at least similar) as in your local environment because of different versions of packages. Pay attention to CUDA version in local &amp; kaggle environments.</li>\n<li>kaggle unpacks .zip and .tar.gz files on uploading to dataset, so try to avoid using nested archives (if it is absolutely necessary, you could rename it to .zip.nounzip and rename back in in notebook)</li>\n<li>In addition to all above, I prefer to pack all the development source code to separate dataset to avoid writing / copy-pasting separate inference code in notebook: pack it with <code>git archive --format=zip --output /full/path/to/zipfile.zip master</code> and use in notebook like this <code>!python inference.py --model ../input/your-model-dataset-name/model.pth --data ../input/your-data-dataset-name/ --output ../working/</code></li>\n</ul>\n<p>Good luck in kaggling!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2323707,
      "author_name": "yukiZ",
      "author_url": "",
      "post_date": "2023-06-30T06:19:34.823000",
      "content": "<p>Hello, I've been challenged to Kaggle recently, so I'd like to make some comments so that I can have a better Kaggle life.</p>\n<p>About question 2.<br>\nSince Submit doesn't work over the internet, there are generally two approach.</p>\n<ul>\n<li><strong>A.sys.path.append</strong></li>\n<li><strong>B.whl using Installation using</strong></li>\n</ul>\n<p>refer to the code that uses <strong>smp</strong> with the Internet off.<br>\n<a href=\"https://www.kaggle.com/code/hideyukizushi/contrails-sample-libinstall-internet-off/notebook\" target=\"_blank\">https://www.kaggle.com/code/hideyukizushi/contrails-sample-libinstall-internet-off/notebook</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2324478,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2023-06-30T16:10:50.937000",
          "content": "<p>And if you want to grab the whl files by yourself, you can go to the library's pypi page and go under Download Files (ex.<a href=\"https://pypi.org/project/segmentation-models-pytorch/#files\" target=\"_blank\"> segmentation models pytorch</a>)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2323795": "Hi! For questions 1 and 2 my general approach is \n\n1. Create inference kaggle notebook with no internet access and pinned environment version (so it will not break at some point with version update from kaggle)\n2. Download all the packages in *.whl (or *.tar.gz) format from your python environment via `pip list --format=freeze > requirements.txt && pip download -r requirements.txt --no-deps --no-binary=:all: -d .`, add it to zip archive and upload to kaggle dataset\n3. Upload your model checkpoints / state dicts to separate kaggle dataset\n4. Add the software dataset and model dataset to your inference notebook and use them without internet: `!pip install --no-index --find-links ../input/your-software-dataset-name/ -r requirements.txt` and load weights from your model dataset with your favorite ML framework\n\nIf some additional system packages are required, similar steps could be performed with `apt download package_name`, uploading to dataset and installing with `!dpkg -i ../input/your-software-dataset-name/package_name.deb`.\n\nYou could refer to [my inference notebook](https://www.kaggle.com/code/mkotyushev/scrolls-inference-agg) from vesuvius challenge to see this approach in action.\n\n**Some side notes**:\n\n- Kaggle's docker images used in competitions are pretty much packed with almost all the required packages and you could avoid uploading and installing *all* the packages and e. g. some heavy torch distribution. So just you should try to run your code and see which dependencies are missing.\n- There are many public datasets with preinstalled packages, you could search for them in notebook's data tab.\n- You need to verify that you model output in kaggle environment is the same (or at least similar) as in your local environment because of different versions of packages. Pay attention to CUDA version in local & kaggle environments.\n- kaggle unpacks .zip and .tar.gz files on uploading to dataset, so try to avoid using nested archives (if it is absolutely necessary, you could rename it to .zip.nounzip and rename back in in notebook)\n- In addition to all above, I prefer to pack all the development source code to separate dataset to avoid writing / copy-pasting separate inference code in notebook: pack it with `git archive --format=zip --output /full/path/to/zipfile.zip master` and use in notebook like this `!python inference.py --model ../input/your-model-dataset-name/model.pth --data ../input/your-data-dataset-name/ --output ../working/`\n\nGood luck in kaggling!",
    "2323387": "Hello, \nA few newbie questions:\n\n1. The code submission rules states that internet access is disabled but also allows pre-trained models? Is the expectation that we download the pre-trained models and add it as a \"data source\" so that we can use it without internet when they run the notebook?\n\n2. I see some public notebooks use `pip install segmentation_models_pytorch`. Are these notebooks only for illustration - i.e. how does `pip install` work without network access.\n\n3. I am not able run the session in the Kaggle notebook beyond the duration of when I am logged into the laptop. If I am idle, I get a prompt that asks me whether I want to continue editing and if I respond after a long time, the session is dead and has to be restarted. Any suggestions on how to keep the Kaggle notebook running?\n\nThank you,",
    "2323707": "Hello, I've been challenged to Kaggle recently, so I'd like to make some comments so that I can have a better Kaggle life.\n\nAbout question 2.\nSince Submit doesn't work over the internet, there are generally two approach.\n* **A.sys.path.append**\n* **B.whl using Installation using**\n\nrefer to the code that uses **smp** with the Internet off.\nhttps://www.kaggle.com/code/hideyukizushi/contrails-sample-libinstall-internet-off/notebook"
  }
}