{
  "id": 372228,
  "title": "Data and training flow... basic beginner question...help appreciated!",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372228",
  "author_name": "Ben Lebovitz",
  "post_date": "2022-12-14T22:38:07.293000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi! I'm a complete beginner at anything dealing with image, and I'm feeling confused how to set up a good flow to do experiments with especially in regards to this being a code competition. So I was hoping that some of the higher ranked folks here (shout out to <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> your shared notebooks and datasets have been amazing…) could let me know if my basic idea on how to structure things are completely dumb or not.</p>\n<p>So… yeah, let me know if I'm going wrong here…<br>\n<strong>Train</strong></p>\n<ol>\n<li>Use a processed data set <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images\" target=\"_blank\">like this one</a> to train on.</li>\n<li>Train model offline so can get more power than on Kaggle<br>\n 2.a. Use CV/hill climbing to determine best models/params/etc… </li>\n<li>Upload saved model(s) to kaggle notebook (or should I upload to hugging face?)</li>\n</ol>\n<p><strong>Test/submit</strong><br>\nAll the below is done in a kaggle notebook to satisfy code contest rules:</p>\n<ol>\n<li>Upload pre-trained model(s)</li>\n<li>Process test images to result in same format as the image dataset chosen in #1</li>\n<li>Predict from saved models, create CSV, profit</li>\n</ol>",
  "messages": [
    {
      "id": 2065890,
      "postDate": "2022-12-15T07:02:17.627Z",
      "content": "<p>Thank you for mentioning my work. I'm still working locally to improve the processing pipeline. I spent the first days in the competition improving the speed of processing dicom files. This point took the most time - there was no time for inference. Now looking for training improvement - still not happy with results (I still can't get satisfactory results).</p>\n<p>I train models locally:</p>\n<ul>\n<li>Pytorch + timm</li>\n<li>models: efficientnet_b2/b3/resnet</li>\n<li>data: use this <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images\" target=\"_blank\">https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images</a> -&gt; resized to 1024 (but most experimentd I do on 512 to speed up process)</li>\n<li>cv: patient_id + laterality -&gt; 4 folds</li>\n<li>augumentation: only few operations </li>\n<li>data sampling (6x more cancer images)</li>\n<li>f1prob optimization </li>\n</ul>\n<p>Then have inference notebook (I upload models to Dataset):</p>\n<ul>\n<li>process file using library from this notebook <a href=\"https://www.kaggle.com/code/remekkinas/fast-dicom-processing-1-6-2x-faster\" target=\"_blank\">https://www.kaggle.com/code/remekkinas/fast-dicom-processing-1-6-2x-faster</a></li>\n<li>blend of models (2 models so far) with TTA</li>\n</ul>",
      "rawMarkdown": "Thank you for mentioning my work. I'm still working locally to improve the processing pipeline. I spent the first days in the competition improving the speed of processing dicom files. This point took the most time - there was no time for inference. Now looking for training improvement - still not happy with results (I still can't get satisfactory results).\n\n I train models locally:\n- Pytorch + timm\n- models: efficientnet_b2/b3/resnet\n- data: use this https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images -> resized to 1024 (but most experimentd I do on 512 to speed up process)\n- cv: patient_id + laterality -> 4 folds\n- augumentation: only few operations \n- data sampling (6x more cancer images)\n- f1prob optimization \n\n\nThen have inference notebook (I upload models to Dataset):\n- process file using library from this notebook https://www.kaggle.com/code/remekkinas/fast-dicom-processing-1-6-2x-faster\n- blend of models (2 models so far) with TTA\n\n",
      "votes": 1
    },
    {
      "id": 2065644,
      "postDate": "2022-12-14T23:18:26.577Z",
      "content": "<p>Yes, I believe that's exactly the idea 🙂 you might want to start with one of the publically shared notebooks and build on that (many have a train/inference part, mine wraps it all into a single one and one of the earlier version had training as submission in a single notebook)</p>\n<p>But yeah, you are in the right track 🙂</p>",
      "rawMarkdown": "Yes, I believe that's exactly the idea 🙂 you might want to start with one of the publically shared notebooks and build on that (many have a train/inference part, mine wraps it all into a single one and one of the earlier version had training as submission in a single notebook)\n\nBut yeah, you are in the right track 🙂",
      "votes": 2,
      "replies": [
        {
          "id": 2065686,
          "postDate": "2022-12-15T00:51:50.477Z",
          "content": "<p>Thanks! </p>\n<p>I think writing out the post actually gave me the answer. I think I just needed to organize my thoughts a bit. Thanks so much for taking the time to read and confirm!</p>",
          "rawMarkdown": "Thanks! \n\nI think writing out the post actually gave me the answer. I think I just needed to organize my thoughts a bit. Thanks so much for taking the time to read and confirm!",
          "votes": 1
        },
        {
          "id": 2065716,
          "postDate": "2022-12-15T01:49:58.557Z",
          "content": "<p>No worries! 🙂 Glad you've found what you were looking for! Writing often feels like magic to me for the reason that you mention.</p>",
          "rawMarkdown": "No worries! 🙂 Glad you've found what you were looking for! Writing often feels like magic to me for the reason that you mention.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2065619,
      "postDate": "2022-12-14T22:38:07.293Z",
      "content": "<p>Hi! I'm a complete beginner at anything dealing with image, and I'm feeling confused how to set up a good flow to do experiments with especially in regards to this being a code competition. So I was hoping that some of the higher ranked folks here (shout out to <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> your shared notebooks and datasets have been amazing…) could let me know if my basic idea on how to structure things are completely dumb or not.</p>\n<p>So… yeah, let me know if I'm going wrong here…<br>\n<strong>Train</strong></p>\n<ol>\n<li>Use a processed data set <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images\" target=\"_blank\">like this one</a> to train on.</li>\n<li>Train model offline so can get more power than on Kaggle<br>\n 2.a. Use CV/hill climbing to determine best models/params/etc… </li>\n<li>Upload saved model(s) to kaggle notebook (or should I upload to hugging face?)</li>\n</ol>\n<p><strong>Test/submit</strong><br>\nAll the below is done in a kaggle notebook to satisfy code contest rules:</p>\n<ol>\n<li>Upload pre-trained model(s)</li>\n<li>Process test images to result in same format as the image dataset chosen in #1</li>\n<li>Predict from saved models, create CSV, profit</li>\n</ol>",
      "rawMarkdown": "Hi! I'm a complete beginner at anything dealing with image, and I'm feeling confused how to set up a good flow to do experiments with especially in regards to this being a code competition. So I was hoping that some of the higher ranked folks here (shout out to @radek1, @remekkinas your shared notebooks and datasets have been amazing...) could let me know if my basic idea on how to structure things are completely dumb or not.\n\nSo... yeah, let me know if I'm going wrong here...\n**Train**\n1. Use a processed data set [like this one](https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images) to train on.\n2. Train model offline so can get more power than on Kaggle\n     2.a. Use CV/hill climbing to determine best models/params/etc... \n3. Upload saved model(s) to kaggle notebook (or should I upload to hugging face?)\n\n**Test/submit**\nAll the below is done in a kaggle notebook to satisfy code contest rules:\n4. Upload pre-trained model(s)\n5. Process test images to result in same format as the image dataset chosen in #1\n6. Predict from saved models, create CSV, profit\n"
    }
  ],
  "comments": [
    {
      "id": 2065890,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-15T07:02:17.627000",
      "content": "<p>Thank you for mentioning my work. I'm still working locally to improve the processing pipeline. I spent the first days in the competition improving the speed of processing dicom files. This point took the most time - there was no time for inference. Now looking for training improvement - still not happy with results (I still can't get satisfactory results).</p>\n<p>I train models locally:</p>\n<ul>\n<li>Pytorch + timm</li>\n<li>models: efficientnet_b2/b3/resnet</li>\n<li>data: use this <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images\" target=\"_blank\">https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images</a> -&gt; resized to 1024 (but most experimentd I do on 512 to speed up process)</li>\n<li>cv: patient_id + laterality -&gt; 4 folds</li>\n<li>augumentation: only few operations </li>\n<li>data sampling (6x more cancer images)</li>\n<li>f1prob optimization </li>\n</ul>\n<p>Then have inference notebook (I upload models to Dataset):</p>\n<ul>\n<li>process file using library from this notebook <a href=\"https://www.kaggle.com/code/remekkinas/fast-dicom-processing-1-6-2x-faster\" target=\"_blank\">https://www.kaggle.com/code/remekkinas/fast-dicom-processing-1-6-2x-faster</a></li>\n<li>blend of models (2 models so far) with TTA</li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2065644,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-14T23:18:26.577000",
      "content": "<p>Yes, I believe that's exactly the idea 🙂 you might want to start with one of the publically shared notebooks and build on that (many have a train/inference part, mine wraps it all into a single one and one of the earlier version had training as submission in a single notebook)</p>\n<p>But yeah, you are in the right track 🙂</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2065686,
          "author_name": "Ben Lebovitz",
          "author_url": "",
          "post_date": "2022-12-15T00:51:50.477000",
          "content": "<p>Thanks! </p>\n<p>I think writing out the post actually gave me the answer. I think I just needed to organize my thoughts a bit. Thanks so much for taking the time to read and confirm!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2065716,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-15T01:49:58.557000",
          "content": "<p>No worries! 🙂 Glad you've found what you were looking for! Writing often feels like magic to me for the reason that you mention.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2065890": "Thank you for mentioning my work. I'm still working locally to improve the processing pipeline. I spent the first days in the competition improving the speed of processing dicom files. This point took the most time - there was no time for inference. Now looking for training improvement - still not happy with results (I still can't get satisfactory results).\n\n I train models locally:\n- Pytorch + timm\n- models: efficientnet_b2/b3/resnet\n- data: use this https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images -> resized to 1024 (but most experimentd I do on 512 to speed up process)\n- cv: patient_id + laterality -> 4 folds\n- augumentation: only few operations \n- data sampling (6x more cancer images)\n- f1prob optimization \n\n\nThen have inference notebook (I upload models to Dataset):\n- process file using library from this notebook https://www.kaggle.com/code/remekkinas/fast-dicom-processing-1-6-2x-faster\n- blend of models (2 models so far) with TTA\n\n",
    "2065644": "Yes, I believe that's exactly the idea 🙂 you might want to start with one of the publically shared notebooks and build on that (many have a train/inference part, mine wraps it all into a single one and one of the earlier version had training as submission in a single notebook)\n\nBut yeah, you are in the right track 🙂",
    "2065619": "Hi! I'm a complete beginner at anything dealing with image, and I'm feeling confused how to set up a good flow to do experiments with especially in regards to this being a code competition. So I was hoping that some of the higher ranked folks here (shout out to @radek1, @remekkinas your shared notebooks and datasets have been amazing...) could let me know if my basic idea on how to structure things are completely dumb or not.\n\nSo... yeah, let me know if I'm going wrong here...\n**Train**\n1. Use a processed data set [like this one](https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images) to train on.\n2. Train model offline so can get more power than on Kaggle\n     2.a. Use CV/hill climbing to determine best models/params/etc... \n3. Upload saved model(s) to kaggle notebook (or should I upload to hugging face?)\n\n**Test/submit**\nAll the below is done in a kaggle notebook to satisfy code contest rules:\n4. Upload pre-trained model(s)\n5. Process test images to result in same format as the image dataset chosen in #1\n6. Predict from saved models, create CSV, profit\n"
  }
}