{
  "id": 379084,
  "title": "non public data and kaggle product idea",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/379084",
  "author_name": "hengck23",
  "post_date": "2023-01-18T01:31:50.228000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>this is idea for the future. </p>\n<p>as an example, you will have find that there are large scale mammography image datasets (millions in size) like NYU, OPTIMAM, CSAW. these are open to public (academic and some are even commercial) upon request.</p>\n<p>actually kaggle can talk to these institutes  and host them as hidden data for long term.<br>\nkagglers cannot download these data, but kaggle can train and evaluate models as what they would do in code competition.</p>\n<ol>\n<li><p>the objective is to let kaggle be a place to evaluate model on real industrial data and not the academic dummy data.</p></li>\n<li><p>it may also be easier to handle external data at competition, i.e. the instruction page can say which hidden external data can be used.</p></li>\n<li><p>it will further improve quality of kaggle solution as kaggle now has access to large scale data.</p></li>\n<li><p>it will encourage the use of TPU (and maybe google cloud compute) as these data are not downloadable and you need power machine to train them.</p></li>\n<li><p>when reading resume from applicants, it is hard to jugde if new comers really have ML training skills. Today codes are easily downloadable. public data are not large and not real enough to test the skill of job applicants. maybe kaggle can fill this gap?</p></li>\n</ol>",
  "messages": [
    {
      "id": 2104707,
      "postDate": "2023-01-18T01:31:50.230Z",
      "content": "<p>this is idea for the future. </p>\n<p>as an example, you will have find that there are large scale mammography image datasets (millions in size) like NYU, OPTIMAM, CSAW. these are open to public (academic and some are even commercial) upon request.</p>\n<p>actually kaggle can talk to these institutes  and host them as hidden data for long term.<br>\nkagglers cannot download these data, but kaggle can train and evaluate models as what they would do in code competition.</p>\n<ol>\n<li><p>the objective is to let kaggle be a place to evaluate model on real industrial data and not the academic dummy data.</p></li>\n<li><p>it may also be easier to handle external data at competition, i.e. the instruction page can say which hidden external data can be used.</p></li>\n<li><p>it will further improve quality of kaggle solution as kaggle now has access to large scale data.</p></li>\n<li><p>it will encourage the use of TPU (and maybe google cloud compute) as these data are not downloadable and you need power machine to train them.</p></li>\n<li><p>when reading resume from applicants, it is hard to jugde if new comers really have ML training skills. Today codes are easily downloadable. public data are not large and not real enough to test the skill of job applicants. maybe kaggle can fill this gap?</p></li>\n</ol>",
      "rawMarkdown": "this is idea for the future. \n\nas an example, you will have find that there are large scale mammography image datasets (millions in size) like NYU, OPTIMAM, CSAW. these are open to public (academic and some are even commercial) upon request.\n\nactually kaggle can talk to these institutes  and host them as hidden data for long term.\nkagglers cannot download these data, but kaggle can train and evaluate models as what they would do in code competition.\n\n1. the objective is to let kaggle be a place to evaluate model on real industrial data and not the academic dummy data.\n\n2. it may also be easier to handle external data at competition, i.e. the instruction page can say which hidden external data can be used.\n\n3. it will further improve quality of kaggle solution as kaggle now has access to large scale data.\n\n4. it will encourage the use of TPU (and maybe google cloud compute) as these data are not downloadable and you need power machine to train them.\n\n5. when reading resume from applicants, it is hard to jugde if new comers really have ML training skills. Today codes are easily downloadable. public data are not large and not real enough to test the skill of job applicants. maybe kaggle can fill this gap?\n",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2104707": "this is idea for the future. \n\nas an example, you will have find that there are large scale mammography image datasets (millions in size) like NYU, OPTIMAM, CSAW. these are open to public (academic and some are even commercial) upon request.\n\nactually kaggle can talk to these institutes  and host them as hidden data for long term.\nkagglers cannot download these data, but kaggle can train and evaluate models as what they would do in code competition.\n\n1. the objective is to let kaggle be a place to evaluate model on real industrial data and not the academic dummy data.\n\n2. it may also be easier to handle external data at competition, i.e. the instruction page can say which hidden external data can be used.\n\n3. it will further improve quality of kaggle solution as kaggle now has access to large scale data.\n\n4. it will encourage the use of TPU (and maybe google cloud compute) as these data are not downloadable and you need power machine to train them.\n\n5. when reading resume from applicants, it is hard to jugde if new comers really have ML training skills. Today codes are easily downloadable. public data are not large and not real enough to test the skill of job applicants. maybe kaggle can fill this gap?\n"
  }
}