{
  "id": 379133,
  "title": "Why is External Data allowed?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/379133",
  "author_name": "beluga",
  "post_date": "2023-01-18T08:41:26.245000",
  "votes": 18,
  "comment_count": 16,
  "views": 0,
  "content": "<p>AFAIK the most relevant and useful datasets ADMANI, NYU with millions of images are not publicly available. (They were probably used to create the train/test set.)</p>\n<p>Pretrained models on these datsets might be available although they have weird licences. (e.g. <a href=\"https://github.com/nyukat/mammography_metarepository)\" target=\"_blank\">https://github.com/nyukat/mammography_metarepository)</a>.</p>\n<p>Other datasets have quality concerns (DDSM) or licence issues (VinDr) and generally even smaller than the current RSNA training set. </p>\n<p>IMHO sharing more data from ADMANI/NYU and disallowing external data would have been better.<br>\nAnd maybe using larger or a bit more balanced test set :D</p>",
  "messages": [
    {
      "id": 2158394,
      "postDate": "2023-02-24T20:16:34.950Z",
      "content": "<p>Hello all and apologies for the very late reply. While the use of publicly available datasets is allowed in the competition, in the sponsor's interpretation of the contest rules, entries employing models based on the NYU dataset would not be eligible as contest winners as that dataset is not made publicly available. Moreover, the models referred to are released under the BSD2 \"viral\" license that, in our understanding, fails to meet the clause in the competition rules that winning models must have a license \"that in no event limits commercial use of such code or models containing or depending on such code.\"</p>",
      "rawMarkdown": "Hello all and apologies for the very late reply. While the use of publicly available datasets is allowed in the competition, in the sponsor's interpretation of the contest rules, entries employing models based on the NYU dataset would not be eligible as contest winners as that dataset is not made publicly available. Moreover, the models referred to are released under the BSD2 \"viral\" license that, in our understanding, fails to meet the clause in the competition rules that winning models must have a license \"that in no event limits commercial use of such code or models containing or depending on such code.\"",
      "votes": 20,
      "replies": [
        {
          "id": 2158891,
          "postDate": "2023-02-25T08:58:46.940Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2159625,
          "postDate": "2023-02-25T22:05:12.470Z",
          "content": "<p>There is no chance that everyone sees this now. You should have created a pinned topic weeks ago to clarify the rules. I guess lots of people will be in a difficult situation in the end, hosts included…</p>",
          "rawMarkdown": "There is no chance that everyone sees this now. You should have created a pinned topic weeks ago to clarify the rules. I guess lots of people will be in a difficult situation in the end, hosts included...",
          "votes": 3,
          "replies": [
            {
              "id": 2159640,
              "postDate": "2023-02-25T22:33:53.173Z",
              "content": "<p>I would argue that anyone who uses this model/data is aware of the risk involved and will monitor these and other threads. It is now clearly stated that it is not allowed to use it. </p>\n<p>That said, I would also wish these decisions would be communicated way earlier.</p>",
              "rawMarkdown": "I would argue that anyone who uses this model/data is aware of the risk involved and will monitor these and other threads. It is now clearly stated that it is not allowed to use it. \n\nThat said, I would also wish these decisions would be communicated way earlier.",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 2105062,
      "postDate": "2023-01-18T08:41:26.247Z",
      "content": "<p>AFAIK the most relevant and useful datasets ADMANI, NYU with millions of images are not publicly available. (They were probably used to create the train/test set.)</p>\n<p>Pretrained models on these datsets might be available although they have weird licences. (e.g. <a href=\"https://github.com/nyukat/mammography_metarepository)\" target=\"_blank\">https://github.com/nyukat/mammography_metarepository)</a>.</p>\n<p>Other datasets have quality concerns (DDSM) or licence issues (VinDr) and generally even smaller than the current RSNA training set. </p>\n<p>IMHO sharing more data from ADMANI/NYU and disallowing external data would have been better.<br>\nAnd maybe using larger or a bit more balanced test set :D</p>",
      "rawMarkdown": "AFAIK the most relevant and useful datasets ADMANI, NYU with millions of images are not publicly available. (They were probably used to create the train/test set.)\n\nPretrained models on these datsets might be available although they have weird licences. (e.g. https://github.com/nyukat/mammography_metarepository).\n\nOther datasets have quality concerns (DDSM) or licence issues (VinDr) and generally even smaller than the current RSNA training set. \n\nIMHO sharing more data from ADMANI/NYU and disallowing external data would have been better.\nAnd maybe using larger or a bit more balanced test set :D",
      "votes": 18
    },
    {
      "id": 2105091,
      "postDate": "2023-01-18T08:58:13.337Z",
      "content": "<p>I am personally not aware of any external data that is allowed according to license, apart from maybe DDSM.</p>\n<p>I am suggesting for a long time to just disallow any external data in competitions, to not always be confronted with these discussions about what is allowed and what is not allowed. From experience, hosts and Kaggle will not give you precise answers early enough. So I am fully on board with what you say.</p>\n<p>Disallowing external data would solve so many issues for competitors, but also Kaggle and hosts. If the models are promising, hosts can always just collect more data post-hoc and it will improve results in most cases.</p>",
      "rawMarkdown": "I am personally not aware of any external data that is allowed according to license, apart from maybe DDSM.\n\nI am suggesting for a long time to just disallow any external data in competitions, to not always be confronted with these discussions about what is allowed and what is not allowed. From experience, hosts and Kaggle will not give you precise answers early enough. So I am fully on board with what you say.\n\nDisallowing external data would solve so many issues for competitors, but also Kaggle and hosts. If the models are promising, hosts can always just collect more data post-hoc and it will improve results in most cases.",
      "votes": 15,
      "replies": [
        {
          "id": 2106066,
          "postDate": "2023-01-18T23:49:16.430Z",
          "content": "<p>Another possibility could also be having a thread which say which dataset are allowed or not, with a specific timeline (maybe one month after the beggining of the competition) so that everyone know what is or is not possible to use?</p>",
          "rawMarkdown": "Another possibility could also be having a thread which say which dataset are allowed or not, with a specific timeline (maybe one month after the beggining of the competition) so that everyone know what is or is not possible to use?",
          "votes": 3,
          "replies": [
            {
              "id": 2106959,
              "postDate": "2023-01-19T13:27:16.670Z",
              "content": "<p>Something like this has not worked too well across competitions in the past. The issue is that kaggle and hosts need to make a decision / legal judgement upfront. I think just disallowing is easier and also has the benefit of people playing the game with the same conditions.</p>",
              "rawMarkdown": "Something like this has not worked too well across competitions in the past. The issue is that kaggle and hosts need to make a decision / legal judgement upfront. I think just disallowing is easier and also has the benefit of people playing the game with the same conditions.",
              "votes": 2
            }
          ]
        },
        {
          "id": 2109179,
          "postDate": "2023-01-21T07:36:19.793Z",
          "rawMarkdown": "",
          "votes": -9,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2143083,
      "postDate": "2023-02-14T02:41:21.563Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> For example, this repository (<a href=\"https://github.com/nyukat/BIRADS_classifier\" target=\"_blank\">https://github.com/nyukat/BIRADS_classifier</a>) includes model weights trained with NYU dataset (not publicly available), but published under BSD 2-Clause \"Simplified\" License. Could you please clarify whether it is acceptable to use those trained weights?</p>",
      "rawMarkdown": "@cdcarr For example, this repository (https://github.com/nyukat/BIRADS_classifier) includes model weights trained with NYU dataset (not publicly available), but published under BSD 2-Clause \"Simplified\" License. Could you please clarify whether it is acceptable to use those trained weights?",
      "votes": 2,
      "replies": [
        {
          "id": 2143093,
          "postDate": "2023-02-14T03:07:18.647Z",
          "content": "<p>It has very minute differences from MIT license, it is most likely allowed.</p>",
          "rawMarkdown": "It has very minute differences from MIT license, it is most likely allowed."
        },
        {
          "id": 2143343,
          "postDate": "2023-02-14T07:52:50.950Z",
          "content": "<p><a href=\"https://www.kaggle.com/cacarr\" target=\"_blank\">@cacarr</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Given that </p>\n<ul>\n<li>such pretrained models are trained on not publicly available dataset</li>\n<li>the possibility of overlap with competition dataset cannot be ruled out (see discussion <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2076911\" target=\"_blank\">here</a> )</li>\n</ul>\n<p>, I strongly suggest <strong>any model weights trained on NYU/ADMANI should not be allowed</strong>. </p>",
          "rawMarkdown": "@cacarr @sohier \n\nGiven that \n- such pretrained models are trained on not publicly available dataset\n- the possibility of overlap with competition dataset cannot be ruled out (see discussion [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2076911) )\n\n, I strongly suggest **any model weights trained on NYU/ADMANI should not be allowed**. ",
          "votes": 7,
          "replies": [
            {
              "id": 2156405,
              "postDate": "2023-02-23T10:15:43.440Z",
              "content": "<blockquote>\n  <p>It has very minute differences from MIT license, it is most likely allowed.</p>\n</blockquote>\n<p>If I train using unauthorized data and publish the weights on my github with an MIT license, it would mean that we could use the weights. This is of course incorrect. </p>",
              "rawMarkdown": "> It has very minute differences from MIT license, it is most likely allowed.\n\nIf I train using unauthorized data and publish the weights on my github with an MIT license, it would mean that we could use the weights. This is of course incorrect. ",
              "votes": 2
            },
            {
              "id": 2156465,
              "postDate": "2023-02-23T11:12:43.983Z",
              "content": "<p>I get the obvious point on why it should be that way but one could argue the same with imagenet weights, not all the images are licensed to be used for commercial in the imagenet dataset, but so many weights trained on imagenet has MIT or other licenses which do not restrict commercial usage..</p>",
              "rawMarkdown": "I get the obvious point on why it should be that way but one could argue the same with imagenet weights, not all the images are licensed to be used for commercial in the imagenet dataset, but so many weights trained on imagenet has MIT or other licenses which do not restrict commercial usage..",
              "votes": -3
            },
            {
              "id": 2156475,
              "postDate": "2023-02-23T11:22:20.570Z",
              "content": "<p>Apart from license issues, another important aspect is the possibility of data leaks. It is about the reliability / external validity of models. I believe only the competition host can give an clear answer to this.</p>",
              "rawMarkdown": "Apart from license issues, another important aspect is the possibility of data leaks. It is about the reliability / external validity of models. I believe only the competition host can give an clear answer to this.",
              "votes": 6
            },
            {
              "id": 2156518,
              "postDate": "2023-02-23T11:43:26.640Z",
              "content": "<p><a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> <a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> <br>\nCould you please make a statement on this? Only 5 days left.</p>\n<p>I personally think such pretrained models should not be allowed as they are trained on non-public data and might even leak towards this dataset.</p>",
              "rawMarkdown": "@maggiemd @cdcarr \nCould you please make a statement on this? Only 5 days left.\n\nI personally think such pretrained models should not be allowed as they are trained on non-public data and might even leak towards this dataset.",
              "votes": 6
            }
          ]
        },
        {
          "id": 2157640,
          "postDate": "2023-02-24T07:35:51.257Z",
          "content": "<p>in theory, from my understanding, training dataset and trained weights are two different entities, each of them is self contained, then this leads to the 1st conclusion, trained-weights on its own license make total sense. i believe in the up coming of AI world, we can have open-sourced (mit, bsd, apache 2.0 etc) trained-weights, it uses private data-set, or sell proprietary trained-weights based on open data-set ( open.ai just did this ). then 2nd conclusion, i think using BSD 2-Clause trained weights does not conflict with the competition rule.</p>",
          "rawMarkdown": "in theory, from my understanding, training dataset and trained weights are two different entities, each of them is self contained, then this leads to the 1st conclusion, trained-weights on its own license make total sense. i believe in the up coming of AI world, we can have open-sourced (mit, bsd, apache 2.0 etc) trained-weights, it uses private data-set, or sell proprietary trained-weights based on open data-set ( open.ai just did this ). then 2nd conclusion, i think using BSD 2-Clause trained weights does not conflict with the competition rule.",
          "votes": -6
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2158394,
      "author_name": "Chris Carr",
      "author_url": "",
      "post_date": "2023-02-24T20:16:34.950000",
      "content": "<p>Hello all and apologies for the very late reply. While the use of publicly available datasets is allowed in the competition, in the sponsor's interpretation of the contest rules, entries employing models based on the NYU dataset would not be eligible as contest winners as that dataset is not made publicly available. Moreover, the models referred to are released under the BSD2 \"viral\" license that, in our understanding, fails to meet the clause in the competition rules that winning models must have a license \"that in no event limits commercial use of such code or models containing or depending on such code.\"</p>",
      "votes": 20,
      "replies": [
        {
          "id": 2158891,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-02-25T08:58:46.940000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2159625,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-02-25T22:05:12.470000",
          "content": "<p>There is no chance that everyone sees this now. You should have created a pinned topic weeks ago to clarify the rules. I guess lots of people will be in a difficult situation in the end, hosts included…</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2159640,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2023-02-25T22:33:53.173000",
              "content": "<p>I would argue that anyone who uses this model/data is aware of the risk involved and will monitor these and other threads. It is now clearly stated that it is not allowed to use it. </p>\n<p>That said, I would also wish these decisions would be communicated way earlier.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2105091,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2023-01-18T08:58:13.337000",
      "content": "<p>I am personally not aware of any external data that is allowed according to license, apart from maybe DDSM.</p>\n<p>I am suggesting for a long time to just disallow any external data in competitions, to not always be confronted with these discussions about what is allowed and what is not allowed. From experience, hosts and Kaggle will not give you precise answers early enough. So I am fully on board with what you say.</p>\n<p>Disallowing external data would solve so many issues for competitors, but also Kaggle and hosts. If the models are promising, hosts can always just collect more data post-hoc and it will improve results in most cases.</p>",
      "votes": 15,
      "replies": [
        {
          "id": 2106066,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2023-01-18T23:49:16.430000",
          "content": "<p>Another possibility could also be having a thread which say which dataset are allowed or not, with a specific timeline (maybe one month after the beggining of the competition) so that everyone know what is or is not possible to use?</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2106959,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2023-01-19T13:27:16.670000",
              "content": "<p>Something like this has not worked too well across competitions in the past. The issue is that kaggle and hosts need to make a decision / legal judgement upfront. I think just disallowing is easier and also has the benefit of people playing the game with the same conditions.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2109179,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-21T07:36:19.793000",
          "content": "",
          "votes": -9,
          "replies": []
        }
      ]
    },
    {
      "id": 2143083,
      "author_name": "RabotniKuma",
      "author_url": "",
      "post_date": "2023-02-14T02:41:21.563000",
      "content": "<p><a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> For example, this repository (<a href=\"https://github.com/nyukat/BIRADS_classifier\" target=\"_blank\">https://github.com/nyukat/BIRADS_classifier</a>) includes model weights trained with NYU dataset (not publicly available), but published under BSD 2-Clause \"Simplified\" License. Could you please clarify whether it is acceptable to use those trained weights?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2143093,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2023-02-14T03:07:18.647000",
          "content": "<p>It has very minute differences from MIT license, it is most likely allowed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2143343,
          "author_name": "RabotniKuma",
          "author_url": "",
          "post_date": "2023-02-14T07:52:50.950000",
          "content": "<p><a href=\"https://www.kaggle.com/cacarr\" target=\"_blank\">@cacarr</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Given that </p>\n<ul>\n<li>such pretrained models are trained on not publicly available dataset</li>\n<li>the possibility of overlap with competition dataset cannot be ruled out (see discussion <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2076911\" target=\"_blank\">here</a> )</li>\n</ul>\n<p>, I strongly suggest <strong>any model weights trained on NYU/ADMANI should not be allowed</strong>. </p>",
          "votes": 7,
          "replies": [
            {
              "id": 2156405,
              "author_name": "YujiAriyasu",
              "author_url": "",
              "post_date": "2023-02-23T10:15:43.440000",
              "content": "<blockquote>\n  <p>It has very minute differences from MIT license, it is most likely allowed.</p>\n</blockquote>\n<p>If I train using unauthorized data and publish the weights on my github with an MIT license, it would mean that we could use the weights. This is of course incorrect. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2156465,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-02-23T11:12:43.983000",
              "content": "<p>I get the obvious point on why it should be that way but one could argue the same with imagenet weights, not all the images are licensed to be used for commercial in the imagenet dataset, but so many weights trained on imagenet has MIT or other licenses which do not restrict commercial usage..</p>",
              "votes": -3,
              "replies": []
            },
            {
              "id": 2156475,
              "author_name": "RabotniKuma",
              "author_url": "",
              "post_date": "2023-02-23T11:22:20.570000",
              "content": "<p>Apart from license issues, another important aspect is the possibility of data leaks. It is about the reliability / external validity of models. I believe only the competition host can give an clear answer to this.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2156518,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2023-02-23T11:43:26.640000",
              "content": "<p><a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> <a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> <br>\nCould you please make a statement on this? Only 5 days left.</p>\n<p>I personally think such pretrained models should not be allowed as they are trained on non-public data and might even leak towards this dataset.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        },
        {
          "id": 2157640,
          "author_name": "Kefan Xu",
          "author_url": "",
          "post_date": "2023-02-24T07:35:51.257000",
          "content": "<p>in theory, from my understanding, training dataset and trained weights are two different entities, each of them is self contained, then this leads to the 1st conclusion, trained-weights on its own license make total sense. i believe in the up coming of AI world, we can have open-sourced (mit, bsd, apache 2.0 etc) trained-weights, it uses private data-set, or sell proprietary trained-weights based on open data-set ( open.ai just did this ). then 2nd conclusion, i think using BSD 2-Clause trained weights does not conflict with the competition rule.</p>",
          "votes": -6,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2158394": "Hello all and apologies for the very late reply. While the use of publicly available datasets is allowed in the competition, in the sponsor's interpretation of the contest rules, entries employing models based on the NYU dataset would not be eligible as contest winners as that dataset is not made publicly available. Moreover, the models referred to are released under the BSD2 \"viral\" license that, in our understanding, fails to meet the clause in the competition rules that winning models must have a license \"that in no event limits commercial use of such code or models containing or depending on such code.\"",
    "2105062": "AFAIK the most relevant and useful datasets ADMANI, NYU with millions of images are not publicly available. (They were probably used to create the train/test set.)\n\nPretrained models on these datsets might be available although they have weird licences. (e.g. https://github.com/nyukat/mammography_metarepository).\n\nOther datasets have quality concerns (DDSM) or licence issues (VinDr) and generally even smaller than the current RSNA training set. \n\nIMHO sharing more data from ADMANI/NYU and disallowing external data would have been better.\nAnd maybe using larger or a bit more balanced test set :D",
    "2105091": "I am personally not aware of any external data that is allowed according to license, apart from maybe DDSM.\n\nI am suggesting for a long time to just disallow any external data in competitions, to not always be confronted with these discussions about what is allowed and what is not allowed. From experience, hosts and Kaggle will not give you precise answers early enough. So I am fully on board with what you say.\n\nDisallowing external data would solve so many issues for competitors, but also Kaggle and hosts. If the models are promising, hosts can always just collect more data post-hoc and it will improve results in most cases.",
    "2143083": "@cdcarr For example, this repository (https://github.com/nyukat/BIRADS_classifier) includes model weights trained with NYU dataset (not publicly available), but published under BSD 2-Clause \"Simplified\" License. Could you please clarify whether it is acceptable to use those trained weights?"
  }
}