{
  "id": 391568,
  "title": "Did Kaggle require compliance with the same set of rules for all submitted solutions?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/391568",
  "author_name": "Allie K.",
  "post_date": "2023-03-01T20:05:03.524000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>A lot was said about which external datasets and models trained on them were allowed in this competition, especially in this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/379133\">discussion</a>. As usual Kaggle team kept silent and left all the responsibility on the host. <br>\nThe host's reaction came unfortunately very late by <a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> 's comment prohibiting the use of the most profound ADMANI/NYU dataset and models trained on it. <br>\nWhat struct me most in his comment was that the author spoke <strong>only about requirements for winning solutions</strong>.</p>\n<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a>  does it mean that if somebody aspires \"only\" to lower gold medal position or to even less glittering medal, then he or she needn't obey any datasets' rules and can use anything what he or she finds anywhere no matter what the license or origin is?<br>\nProbably yes, because the chance to be caught is extremely low.</p>\n<p>This competition was closed in one day, definitely without Kaggle team's checking anything but double accounts.<br>\nA brief look at the leaderboard shows a typical picture - some high ranked silver medals going to teams of novices/contributors with hardly any or no activity on Kaggle. The only difference from a usual leaderboard is that this time there was another possibility how to get a medal without any significant effort - simply by using the prohibited models.</p>\n<p>Was this an equal opportunity competition - yes, because in fact everybody has the same opportunity (if character) to cheat without being caught.<br>\nWas this competition set to be fair - not at all.<br>\nKaggle has not only the right to check any submitted solution but also the duty to keep competitions fair. <br>\nIn a case like this it would be easy to check at least all unpublished gold solutions and \"sudden geniuses' \" medal solutions for datasets/models compliance. Of course it would require more active anti-cheating approach by the organizer. </p>",
  "messages": [
    {
      "id": 2165861,
      "postDate": "2023-03-02T14:09:29.010Z",
      "content": "<p>I agree with you, but the issue is that it is just impossible to check the training code of all solutions and only the winners have to license their solution. I thought a few times about this issue, and also I cannot find a good solution. The only way I would see is that all training needs to be done in the kaggle solution kernel, combined with inference, and only pre-approved datasets / models can be attached. While technically possible, and probably interesting, this would seriously limit the quality of produced results.</p>\n<p>One step in the right direction, which I repeatedly suggest, is to <strong>completely disallow any external data</strong> from the start of the competition. This would discourage people from exploring extra data from the start, and the temptation to use it would be lower.</p>\n<p>Or at least, clearly state in the beginning which data is allowed, and all other is disallowed.</p>\n<p>But regardless of what you do, there will always be cheaters in life. It can be incredibly frustrating, but I still believe that long-term honesty will be rewarded. That still means, that certain steps can and should be done to make cheating much more risky.</p>",
      "rawMarkdown": "I agree with you, but the issue is that it is just impossible to check the training code of all solutions and only the winners have to license their solution. I thought a few times about this issue, and also I cannot find a good solution. The only way I would see is that all training needs to be done in the kaggle solution kernel, combined with inference, and only pre-approved datasets / models can be attached. While technically possible, and probably interesting, this would seriously limit the quality of produced results.\n\nOne step in the right direction, which I repeatedly suggest, is to **completely disallow any external data** from the start of the competition. This would discourage people from exploring extra data from the start, and the temptation to use it would be lower.\n\nOr at least, clearly state in the beginning which data is allowed, and all other is disallowed.\n\nBut regardless of what you do, there will always be cheaters in life. It can be incredibly frustrating, but I still believe that long-term honesty will be rewarded. That still means, that certain steps can and should be done to make cheating much more risky.",
      "votes": 6,
      "replies": [
        {
          "id": 2166392,
          "postDate": "2023-03-02T19:02:41.370Z",
          "content": "<p>Thank you for your reaction and congrats to your team's 2 golds in 2 days!</p>\n<p>You are completely right that all this issue is very complex and difficult to handle by the organizer. This is also why I suggested at least the very limited number of solutions checking.<br>\nBut what is probably most frustrating is that despite many repeated suggestions coming even from the top Kagglers including you, there are no visible steps made to address the external data problem.</p>\n<p>And here every step would count because it would demonstrate that Kaggle team does care both about the fairness of competitions and about the opinions of those, who create the real Kaggle value - top Kagglers.</p>",
          "rawMarkdown": "Thank you for your reaction and congrats to your team's 2 golds in 2 days!\n\nYou are completely right that all this issue is very complex and difficult to handle by the organizer. This is also why I suggested at least the very limited number of solutions checking.\nBut what is probably most frustrating is that despite many repeated suggestions coming even from the top Kagglers including you, there are no visible steps made to address the external data problem.\n\nAnd here every step would count because it would demonstrate that Kaggle team does care both about the fairness of competitions and about the opinions of those, who create the real Kaggle value - top Kagglers.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2164834,
      "postDate": "2023-03-01T20:26:49.970Z",
      "content": "<p>I don’t think that publicly available models pretrained on NYU brought any significant advantage. So IMO there was no easy cheating method in this competition.</p>\n<p>However I do agree that following the changing rules during the completion is an extra disadvantage for honest people and completely useless if you don’t aim at a prize pool solution.</p>\n<p>In the end, except for money positions, this is all just about learning things and getting better at what you like, so don’t worry too much about people bending the rules. Overall I still think this was a fair competition.</p>",
      "rawMarkdown": "I don’t think that publicly available models pretrained on NYU brought any significant advantage. So IMO there was no easy cheating method in this competition.\n\nHowever I do agree that following the changing rules during the completion is an extra disadvantage for honest people and completely useless if you don’t aim at a prize pool solution.\n\nIn the end, except for money positions, this is all just about learning things and getting better at what you like, so don’t worry too much about people bending the rules. Overall I still think this was a fair competition.",
      "votes": 4,
      "replies": [
        {
          "id": 2165630,
          "postDate": "2023-03-02T10:13:18.260Z",
          "content": "<p>Thank you for expressing your opinion and congrats to your team's results.<br>\nI used the ADMANI/NYU dataset only as an example but <strong>the problem is much more general</strong> and didn't appear only in this competition.</p>\n<p>When the host or Kaggle team prohibits something (dataset, model, algorithm etc.) in the rules but then doesn't enforce compliance with this rule (except for marginal number of winners) then such a <strong>competition setting is not fair from the basis</strong>.</p>\n<p>I agree with you that at some point of our career the most we can get from a competition is the knowledge. For this I can participate but needn't submit to be listed in the leaderboard.</p>\n<p>But most participants aspiring to at least silver medal use leaderboard (rankings) to enhance their career. So this is why I repeatedly ask the Kaggle team and its head <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> why cheaters should be favoured at the expense of honest participants.</p>",
          "rawMarkdown": "Thank you for expressing your opinion and congrats to your team's results.\nI used the ADMANI/NYU dataset only as an example but **the problem is much more general** and didn't appear only in this competition.\n\nWhen the host or Kaggle team prohibits something (dataset, model, algorithm etc.) in the rules but then doesn't enforce compliance with this rule (except for marginal number of winners) then such a **competition setting is not fair from the basis**.\n\nI agree with you that at some point of our career the most we can get from a competition is the knowledge. For this I can participate but needn't submit to be listed in the leaderboard.\n\nBut most participants aspiring to at least silver medal use leaderboard (rankings) to enhance their career. So this is why I repeatedly ask the Kaggle team and its head @wcukierski why cheaters should be favoured at the expense of honest participants.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2164806,
      "postDate": "2023-03-01T20:05:03.523Z",
      "content": "<p>A lot was said about which external datasets and models trained on them were allowed in this competition, especially in this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/379133\">discussion</a>. As usual Kaggle team kept silent and left all the responsibility on the host. <br>\nThe host's reaction came unfortunately very late by <a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> 's comment prohibiting the use of the most profound ADMANI/NYU dataset and models trained on it. <br>\nWhat struct me most in his comment was that the author spoke <strong>only about requirements for winning solutions</strong>.</p>\n<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a>  does it mean that if somebody aspires \"only\" to lower gold medal position or to even less glittering medal, then he or she needn't obey any datasets' rules and can use anything what he or she finds anywhere no matter what the license or origin is?<br>\nProbably yes, because the chance to be caught is extremely low.</p>\n<p>This competition was closed in one day, definitely without Kaggle team's checking anything but double accounts.<br>\nA brief look at the leaderboard shows a typical picture - some high ranked silver medals going to teams of novices/contributors with hardly any or no activity on Kaggle. The only difference from a usual leaderboard is that this time there was another possibility how to get a medal without any significant effort - simply by using the prohibited models.</p>\n<p>Was this an equal opportunity competition - yes, because in fact everybody has the same opportunity (if character) to cheat without being caught.<br>\nWas this competition set to be fair - not at all.<br>\nKaggle has not only the right to check any submitted solution but also the duty to keep competitions fair. <br>\nIn a case like this it would be easy to check at least all unpublished gold solutions and \"sudden geniuses' \" medal solutions for datasets/models compliance. Of course it would require more active anti-cheating approach by the organizer. </p>",
      "rawMarkdown": "A lot was said about which external datasets and models trained on them were allowed in this competition, especially in this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/379133\">discussion</a>. As usual Kaggle team kept silent and left all the responsibility on the host. \nThe host's reaction came unfortunately very late by @cdcarr 's comment prohibiting the use of the most profound ADMANI/NYU dataset and models trained on it. \nWhat struct me most in his comment was that the author spoke **only about requirements for winning solutions**.\n\n@wcukierski  does it mean that if somebody aspires \"only\" to lower gold medal position or to even less glittering medal, then he or she needn't obey any datasets' rules and can use anything what he or she finds anywhere no matter what the license or origin is?\nProbably yes, because the chance to be caught is extremely low.\n\nThis competition was closed in one day, definitely without Kaggle team's checking anything but double accounts.\nA brief look at the leaderboard shows a typical picture - some high ranked silver medals going to teams of novices/contributors with hardly any or no activity on Kaggle. The only difference from a usual leaderboard is that this time there was another possibility how to get a medal without any significant effort - simply by using the prohibited models.\n\nWas this an equal opportunity competition - yes, because in fact everybody has the same opportunity (if character) to cheat without being caught.\nWas this competition set to be fair - not at all.\nKaggle has not only the right to check any submitted solution but also the duty to keep competitions fair. \nIn a case like this it would be easy to check at least all unpublished gold solutions and \"sudden geniuses' \" medal solutions for datasets/models compliance. Of course it would require more active anti-cheating approach by the organizer. ",
      "votes": 1
    },
    {
      "id": 2165425,
      "postDate": "2023-03-02T07:47:50.387Z",
      "content": "<p>I see your point and hope that the authorities answer your questions.<br>\nI personally tried hard not to use any external dataset and pretrained model that either not allowed for commercial use or not publicly available. But I am quite happy since without using any such prohibited material I managed to get a good model and on top of that I learned a lot of new things 🙂</p>",
      "rawMarkdown": "I see your point and hope that the authorities answer your questions.\nI personally tried hard not to use any external dataset and pretrained model that either not allowed for commercial use or not publicly available. But I am quite happy since without using any such prohibited material I managed to get a good model and on top of that I learned a lot of new things 🙂",
      "votes": 2
    },
    {
      "id": 2178293,
      "postDate": "2023-03-12T09:47:26.080Z",
      "content": "<p>External data is one thing, another is the models allowed. I kept seeing in discussions that ConvNeXt is not allowed so I didn't even try it… In the next competition, I will simply not read the rules… </p>",
      "rawMarkdown": "External data is one thing, another is the models allowed. I kept seeing in discussions that ConvNeXt is not allowed so I didn't even try it... In the next competition, I will simply not read the rules... "
    }
  ],
  "comments": [
    {
      "id": 2165861,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2023-03-02T14:09:29.010000",
      "content": "<p>I agree with you, but the issue is that it is just impossible to check the training code of all solutions and only the winners have to license their solution. I thought a few times about this issue, and also I cannot find a good solution. The only way I would see is that all training needs to be done in the kaggle solution kernel, combined with inference, and only pre-approved datasets / models can be attached. While technically possible, and probably interesting, this would seriously limit the quality of produced results.</p>\n<p>One step in the right direction, which I repeatedly suggest, is to <strong>completely disallow any external data</strong> from the start of the competition. This would discourage people from exploring extra data from the start, and the temptation to use it would be lower.</p>\n<p>Or at least, clearly state in the beginning which data is allowed, and all other is disallowed.</p>\n<p>But regardless of what you do, there will always be cheaters in life. It can be incredibly frustrating, but I still believe that long-term honesty will be rewarded. That still means, that certain steps can and should be done to make cheating much more risky.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2166392,
          "author_name": "Allie K.",
          "author_url": "",
          "post_date": "2023-03-02T19:02:41.370000",
          "content": "<p>Thank you for your reaction and congrats to your team's 2 golds in 2 days!</p>\n<p>You are completely right that all this issue is very complex and difficult to handle by the organizer. This is also why I suggested at least the very limited number of solutions checking.<br>\nBut what is probably most frustrating is that despite many repeated suggestions coming even from the top Kagglers including you, there are no visible steps made to address the external data problem.</p>\n<p>And here every step would count because it would demonstrate that Kaggle team does care both about the fairness of competitions and about the opinions of those, who create the real Kaggle value - top Kagglers.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2164834,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-03-01T20:26:49.970000",
      "content": "<p>I don’t think that publicly available models pretrained on NYU brought any significant advantage. So IMO there was no easy cheating method in this competition.</p>\n<p>However I do agree that following the changing rules during the completion is an extra disadvantage for honest people and completely useless if you don’t aim at a prize pool solution.</p>\n<p>In the end, except for money positions, this is all just about learning things and getting better at what you like, so don’t worry too much about people bending the rules. Overall I still think this was a fair competition.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2165630,
          "author_name": "Allie K.",
          "author_url": "",
          "post_date": "2023-03-02T10:13:18.260000",
          "content": "<p>Thank you for expressing your opinion and congrats to your team's results.<br>\nI used the ADMANI/NYU dataset only as an example but <strong>the problem is much more general</strong> and didn't appear only in this competition.</p>\n<p>When the host or Kaggle team prohibits something (dataset, model, algorithm etc.) in the rules but then doesn't enforce compliance with this rule (except for marginal number of winners) then such a <strong>competition setting is not fair from the basis</strong>.</p>\n<p>I agree with you that at some point of our career the most we can get from a competition is the knowledge. For this I can participate but needn't submit to be listed in the leaderboard.</p>\n<p>But most participants aspiring to at least silver medal use leaderboard (rankings) to enhance their career. So this is why I repeatedly ask the Kaggle team and its head <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> why cheaters should be favoured at the expense of honest participants.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2165425,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2023-03-02T07:47:50.387000",
      "content": "<p>I see your point and hope that the authorities answer your questions.<br>\nI personally tried hard not to use any external dataset and pretrained model that either not allowed for commercial use or not publicly available. But I am quite happy since without using any such prohibited material I managed to get a good model and on top of that I learned a lot of new things 🙂</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2178293,
      "author_name": "FlaviuPaul",
      "author_url": "",
      "post_date": "2023-03-12T09:47:26.080000",
      "content": "<p>External data is one thing, another is the models allowed. I kept seeing in discussions that ConvNeXt is not allowed so I didn't even try it… In the next competition, I will simply not read the rules… </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2165861": "I agree with you, but the issue is that it is just impossible to check the training code of all solutions and only the winners have to license their solution. I thought a few times about this issue, and also I cannot find a good solution. The only way I would see is that all training needs to be done in the kaggle solution kernel, combined with inference, and only pre-approved datasets / models can be attached. While technically possible, and probably interesting, this would seriously limit the quality of produced results.\n\nOne step in the right direction, which I repeatedly suggest, is to **completely disallow any external data** from the start of the competition. This would discourage people from exploring extra data from the start, and the temptation to use it would be lower.\n\nOr at least, clearly state in the beginning which data is allowed, and all other is disallowed.\n\nBut regardless of what you do, there will always be cheaters in life. It can be incredibly frustrating, but I still believe that long-term honesty will be rewarded. That still means, that certain steps can and should be done to make cheating much more risky.",
    "2164834": "I don’t think that publicly available models pretrained on NYU brought any significant advantage. So IMO there was no easy cheating method in this competition.\n\nHowever I do agree that following the changing rules during the completion is an extra disadvantage for honest people and completely useless if you don’t aim at a prize pool solution.\n\nIn the end, except for money positions, this is all just about learning things and getting better at what you like, so don’t worry too much about people bending the rules. Overall I still think this was a fair competition.",
    "2164806": "A lot was said about which external datasets and models trained on them were allowed in this competition, especially in this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/379133\">discussion</a>. As usual Kaggle team kept silent and left all the responsibility on the host. \nThe host's reaction came unfortunately very late by @cdcarr 's comment prohibiting the use of the most profound ADMANI/NYU dataset and models trained on it. \nWhat struct me most in his comment was that the author spoke **only about requirements for winning solutions**.\n\n@wcukierski  does it mean that if somebody aspires \"only\" to lower gold medal position or to even less glittering medal, then he or she needn't obey any datasets' rules and can use anything what he or she finds anywhere no matter what the license or origin is?\nProbably yes, because the chance to be caught is extremely low.\n\nThis competition was closed in one day, definitely without Kaggle team's checking anything but double accounts.\nA brief look at the leaderboard shows a typical picture - some high ranked silver medals going to teams of novices/contributors with hardly any or no activity on Kaggle. The only difference from a usual leaderboard is that this time there was another possibility how to get a medal without any significant effort - simply by using the prohibited models.\n\nWas this an equal opportunity competition - yes, because in fact everybody has the same opportunity (if character) to cheat without being caught.\nWas this competition set to be fair - not at all.\nKaggle has not only the right to check any submitted solution but also the duty to keep competitions fair. \nIn a case like this it would be easy to check at least all unpublished gold solutions and \"sudden geniuses' \" medal solutions for datasets/models compliance. Of course it would require more active anti-cheating approach by the organizer. ",
    "2165425": "I see your point and hope that the authorities answer your questions.\nI personally tried hard not to use any external dataset and pretrained model that either not allowed for commercial use or not publicly available. But I am quite happy since without using any such prohibited material I managed to get a good model and on top of that I learned a lot of new things 🙂",
    "2178293": "External data is one thing, another is the models allowed. I kept seeing in discussions that ConvNeXt is not allowed so I didn't even try it... In the next competition, I will simply not read the rules... "
  }
}