{
  "id": 377143,
  "title": "Submit time is unstable",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/377143",
  "author_name": "ForcewithMe",
  "post_date": "2023-01-10T01:57:16.894000",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I submitted the same model with different thresholds, while the rest of the settings were exactly the same. However, the fastest one takes less than 7.5 hours, while the slowest one ran out of the limit of 9 hours. 👀</p>",
  "messages": [
    {
      "id": 2093416,
      "postDate": "2023-01-10T01:57:16.893Z",
      "content": "<p>I submitted the same model with different thresholds, while the rest of the settings were exactly the same. However, the fastest one takes less than 7.5 hours, while the slowest one ran out of the limit of 9 hours. 👀</p>",
      "rawMarkdown": "I submitted the same model with different thresholds, while the rest of the settings were exactly the same. However, the fastest one takes less than 7.5 hours, while the slowest one ran out of the limit of 9 hours. 👀",
      "votes": 6
    },
    {
      "id": 2093875,
      "postDate": "2023-01-10T11:46:49.120Z",
      "content": "<p>The randomness in the runtime, in my opinion has something to do with the saving the converted and preprocessed DCM files on disk. Since we save about 34K images on disk each time, and the fact that each submission notebook uses network storage space, depending on the network traffic the read/write operations could take longer than normal.</p>",
      "rawMarkdown": "The randomness in the runtime, in my opinion has something to do with the saving the converted and preprocessed DCM files on disk. Since we save about 34K images on disk each time, and the fact that each submission notebook uses network storage space, depending on the network traffic the read/write operations could take longer than normal.",
      "votes": 1,
      "replies": [
        {
          "id": 2093884,
          "postDate": "2023-01-10T11:54:46.690Z",
          "content": "<p>I suspect you are right, as I don't save anything to disk and my runtimes are very stable.</p>",
          "rawMarkdown": "I suspect you are right, as I don't save anything to disk and my runtimes are very stable.",
          "votes": 1
        },
        {
          "id": 2094040,
          "postDate": "2023-01-10T14:58:13.123Z",
          "content": "<p>It really make sense! 👍</p>",
          "rawMarkdown": "It really make sense! 👍"
        }
      ]
    },
    {
      "id": 2093488,
      "postDate": "2023-01-10T03:39:45.430Z",
      "content": "<p>I noticed the same thing so I wonder if it is a problem with Kaggle kernels.</p>",
      "rawMarkdown": "I noticed the same thing so I wonder if it is a problem with Kaggle kernels.",
      "votes": 1
    },
    {
      "id": 2094058,
      "postDate": "2023-01-10T15:24:07.477Z",
      "content": "<p>for me,  some submissions scored, some OOM.  same kernel</p>",
      "rawMarkdown": "for me,  some submissions scored, some OOM.  same kernel"
    },
    {
      "id": 2094045,
      "postDate": "2023-01-10T15:05:41.900Z",
      "content": "<p>I am also wondering if the submission time is related to the number of submissions at the same time. As <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a> said, the randomness in the runtime is caused by network traffic, I wonder if kaggle will cut the network resources of users who frequently take submissions at the same time, and offer resources first to users who previously used fewer submission resources. I have this guess because my submission can be finished in 7 hours if I only make 2 submissions at one time. But if I take 5 submissions at one time, some of them takes more than 8 hours, and one of them takes more than 9 hours.</p>",
      "rawMarkdown": "I am also wondering if the submission time is related to the number of submissions at the same time. As @rasoulmojtahedzadeh said, the randomness in the runtime is caused by network traffic, I wonder if kaggle will cut the network resources of users who frequently take submissions at the same time, and offer resources first to users who previously used fewer submission resources. I have this guess because my submission can be finished in 7 hours if I only make 2 submissions at one time. But if I take 5 submissions at one time, some of them takes more than 8 hours, and one of them takes more than 9 hours."
    },
    {
      "id": 2093537,
      "postDate": "2023-01-10T04:30:33.320Z",
      "content": "<p>In previous competitions, the variance of time between different submissions wasn't as large as this one. In this pipeline, I use <code>Dali</code> to generate PNG files. For Inference, I don't apply any TTAs, with only <code>normalize</code>,  <code>resize</code>, and <code>ToTensor</code>. So I can't identify any randomness in my pipeline for now. Maybe I missed something.</p>",
      "rawMarkdown": "In previous competitions, the variance of time between different submissions wasn't as large as this one. In this pipeline, I use `Dali` to generate PNG files. For Inference, I don't apply any TTAs, with only `normalize`,  `resize`, and `ToTensor`. So I can't identify any randomness in my pipeline for now. Maybe I missed something."
    }
  ],
  "comments": [
    {
      "id": 2093875,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2023-01-10T11:46:49.120000",
      "content": "<p>The randomness in the runtime, in my opinion has something to do with the saving the converted and preprocessed DCM files on disk. Since we save about 34K images on disk each time, and the fact that each submission notebook uses network storage space, depending on the network traffic the read/write operations could take longer than normal.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2093884,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2023-01-10T11:54:46.690000",
          "content": "<p>I suspect you are right, as I don't save anything to disk and my runtimes are very stable.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2094040,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2023-01-10T14:58:13.123000",
          "content": "<p>It really make sense! 👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2093488,
      "author_name": "outwrest",
      "author_url": "",
      "post_date": "2023-01-10T03:39:45.430000",
      "content": "<p>I noticed the same thing so I wonder if it is a problem with Kaggle kernels.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2094058,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-01-10T15:24:07.477000",
      "content": "<p>for me,  some submissions scored, some OOM.  same kernel</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2094045,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2023-01-10T15:05:41.900000",
      "content": "<p>I am also wondering if the submission time is related to the number of submissions at the same time. As <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a> said, the randomness in the runtime is caused by network traffic, I wonder if kaggle will cut the network resources of users who frequently take submissions at the same time, and offer resources first to users who previously used fewer submission resources. I have this guess because my submission can be finished in 7 hours if I only make 2 submissions at one time. But if I take 5 submissions at one time, some of them takes more than 8 hours, and one of them takes more than 9 hours.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2093537,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2023-01-10T04:30:33.320000",
      "content": "<p>In previous competitions, the variance of time between different submissions wasn't as large as this one. In this pipeline, I use <code>Dali</code> to generate PNG files. For Inference, I don't apply any TTAs, with only <code>normalize</code>,  <code>resize</code>, and <code>ToTensor</code>. So I can't identify any randomness in my pipeline for now. Maybe I missed something.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2093416": "I submitted the same model with different thresholds, while the rest of the settings were exactly the same. However, the fastest one takes less than 7.5 hours, while the slowest one ran out of the limit of 9 hours. 👀",
    "2093875": "The randomness in the runtime, in my opinion has something to do with the saving the converted and preprocessed DCM files on disk. Since we save about 34K images on disk each time, and the fact that each submission notebook uses network storage space, depending on the network traffic the read/write operations could take longer than normal.",
    "2093488": "I noticed the same thing so I wonder if it is a problem with Kaggle kernels.",
    "2094058": "for me,  some submissions scored, some OOM.  same kernel",
    "2094045": "I am also wondering if the submission time is related to the number of submissions at the same time. As @rasoulmojtahedzadeh said, the randomness in the runtime is caused by network traffic, I wonder if kaggle will cut the network resources of users who frequently take submissions at the same time, and offer resources first to users who previously used fewer submission resources. I have this guess because my submission can be finished in 7 hours if I only make 2 submissions at one time. But if I take 5 submissions at one time, some of them takes more than 8 hours, and one of them takes more than 9 hours.",
    "2093537": "In previous competitions, the variance of time between different submissions wasn't as large as this one. In this pipeline, I use `Dali` to generate PNG files. For Inference, I don't apply any TTAs, with only `normalize`,  `resize`, and `ToTensor`. So I can't identify any randomness in my pipeline for now. Maybe I missed something."
  }
}