{
  "topic": {
    "id": 433184,
    "title": "CAFA5: last day competition stats",
    "authorName": "Oleksiy Kononenko",
    "commentCount": 20,
    "votes": 10,
    "postDate": "2023-08-20T16:49:15.619000"
  },
  "comments": [
    {
      "id": 2401400,
      "authorName": "John Mitchell",
      "votes": 4,
      "postDate": "2023-08-21T15:22:45.733000",
      "content": "<p>Note that this is substantially similar to my answer <a href=\"https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/431181#2396668\" target=\"_blank\">here</a>:</p>\n<p>Broadly, teams near the top of the LB are there for one of two reasons:</p>\n<p>[1] They are bioinformatics professionals (academic or biotech) who know what they're doing and have excellent scores because they have strong models. These will remain near the top.</p>\n<p>Based on previous CAFAs, I'd expect 50-100 'professional' teams to end up in the medals. Most of those will have decent public scores too come submission time, even if they post their models only today (the last day). Noted that this is essentially the same point as <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>'s comment that <em>I think there will be a shakeup, but maybe for half the competitors in top 50. I would assume that many of the quiet guys in top 20 that have been in this competition before will stay as they are or move upward</em>.</p>\n<p>[2] They are have overfitted the public LB and are heading for an almighty fall when the results come out.</p>\n<p>My team has dropped ~200 places on the public LB in the last week or so, at the same time as higher-scoring public codes appeared. My judgement is that there's little new innovation in those codes beyond what has been shared earlier. rather a tweaking of parameters. Given how small a set the public LB is probably based on, fitting multiple parameters to maximise public score seems foolhardy.</p>\n<p>I'd also expect that any genuinely good model would do reasonably well on the public LB too. I previously wrote that <em>I would be surprised if many teams come 'from nowhere' (outside the medals, ~top 150) into the gold zone, unless they're deliberately sandbagging their scores on the public LB.</em> Given that I think there are 150-200 seriously overfitted models in the top 250 now, I'd probably be less surprised now to see risers from further down.  After all, we will be coming from 236th place or worse.</p>"
    },
    {
      "id": 2401625,
      "authorName": "Alexander Chervov",
      "votes": 1,
      "postDate": "2023-08-21T17:35:57.353000",
      "content": "<p>By the way , I guess if you \"merge\" your solution with quick-go taking same params as jn top public notebook.<br>\nThen, I guess your LB score would jump back on top. I guess many people doing that. I had wrong impression it will not work on private at all. Looking more carefully I guess there's some chance for at least not a complete shake down for that quick-go thing. It is also unexpected for me that small changes they do in params change LB score that much. Any way you may choose different solution for final private judgment.</p>"
    },
    {
      "id": 2401824,
      "authorName": "Oleksiy Kononenko",
      "votes": 1,
      "postDate": "2023-08-21T20:12:31.883000",
      "content": "<blockquote>\n  <p>My team has dropped ~200 places on the public LB in the last week or so, at the same time as higher-scoring public codes appeared</p>\n</blockquote>\n<p>The interesting thing here is that there are some identical scores (for instance, <code>0.58938</code> or <code>0.59144</code>) demonstrated by ~10 teams with no corresponding public notebooks.</p>"
    },
    {
      "id": 2401839,
      "authorName": "Tilii",
      "votes": 1,
      "postDate": "2023-08-21T20:32:53.220000",
      "content": "<blockquote>\n  <p>The interesting thing here is that there are some identical scores (for instance, 0.58938 or 0.59144) demonstrated by ~10 teams with no corresponding public notebooks.</p>\n</blockquote>\n<p>This would have been near-impossible if the public LB had a very large number of data points, or if the number of hyperparameters in those public notebooks was large. For 100 or so data points, it is possible that several groups made identical changes to public notebooks that got them identical scores.</p>"
    },
    {
      "id": 2401854,
      "authorName": "Oleksiy Kononenko",
      "votes": 0,
      "postDate": "2023-08-21T20:50:52.793000",
      "content": "<p>Yes, however, those high scoring public notebooks don’t have any hyperparameters to tune, they only combine several submissions into one. The number of combinations is pretty large, so I find it highly unlikely that ~10 teams would end up with identical scores. </p>\n<p>May be a submission was shared somewhere as a dataset, or some public notebook got deleted, not sure.</p>"
    },
    {
      "id": 2401904,
      "authorName": "Alexander Chervov",
      "votes": 2,
      "postDate": "2023-08-21T21:55:16.213000",
      "content": "<p>Public notebooks have simple integer hyperparameter.<br>\nWhich everybody tunes -  \"top\" - how many labels allowed for protein</p>\n<p>Originally MT  put 35,later others used  45.<br>\nNow they checked more options.<br>\nSee different versions of the top public notebook.<br>\nI guess these teams with same score put the other integer values. </p>"
    },
    {
      "id": 2401911,
      "authorName": "Tilii",
      "votes": 1,
      "postDate": "2023-08-21T21:59:38.653000",
      "content": "<blockquote>\n  <p>I guess these teams with same score put other integer values.</p>\n</blockquote>\n<p>That was also my thinking, and perhaps they used a slightly different combination of public models.</p>"
    },
    {
      "id": 2401917,
      "authorName": "Alexander Chervov",
      "votes": 0,
      "postDate": "2023-08-21T22:04:58.370000",
      "content": "<p>I think just changing \"top\" otherwise score might change more.<br>\nBut the newest version of public contains an option to have different \"top\" for different ontologies.<br>\nSo now we have 3 integers and appropriate choice is less obvious.<br>\nSo those who have more luck chosen better integers.</p>\n<p>PS<br>\nIn my mind it is overfishing to LB.</p>"
    },
    {
      "id": 2401934,
      "authorName": "Oleksiy Kononenko",
      "votes": 1,
      "postDate": "2023-08-21T22:33:15.530000",
      "content": "<p>Ah, I thought it was the model hyperparameter that was meant by <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>. Then, there is also a <code>bonus</code> that could be adjusted. Still, I'm not sure why ten teams would change these in exactly the same way.</p>"
    },
    {
      "id": 2401938,
      "authorName": "Tilii",
      "votes": 1,
      "postDate": "2023-08-21T22:38:16.637000",
      "content": "<blockquote>\n  <p>Still, I'm not sure why ten teams would change these in exactly the same way.</p>\n</blockquote>\n<p>A relatively limited number of choices, ~1500 teams, 5 submissions per day. I have no trouble accepting that 10 teams would stumble upon the same combination of parameters.</p>"
    },
    {
      "id": 2401946,
      "authorName": "Alexander Chervov",
      "votes": 1,
      "postDate": "2023-08-21T22:57:47.277000",
      "content": "<p>People always look on what is working well, and try push it further.<br>\nThat top public changed \"top\" from 42 to 39 and got uplift - compare early version with the recent version.</p>\n<p>So it is natural to try 38 37 36 <br>\nNot so many options <br>\nThus we see several clusters.</p>\n<p>PS<br>\n\"Bonus\" does not seen to be used by any one. Probably it is not working well</p>"
    },
    {
      "id": 2401947,
      "authorName": "Oleksiy Kononenko",
      "votes": 1,
      "postDate": "2023-08-21T23:01:20.797000",
      "content": "<p><a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> I was curious to try it some time ago and even <code>bonus == 0.001</code> makes things worse. </p>"
    },
    {
      "id": 2404377,
      "authorName": "hoho",
      "votes": 2,
      "postDate": "2023-08-23T07:59:19.720000",
      "content": "<p>I used <code>bonus=0.05</code>, but I'm not just kept the top score when <code>add_prediction</code>, I think it can increased the public LB.</p>"
    },
    {
      "id": 2401987,
      "authorName": "Oleksiy Kononenko",
      "votes": 1,
      "postDate": "2023-08-22T00:19:53.280000",
      "content": "<p>Just updated <a href=\"https://www.kaggle.com/kononenko/cafa5-competition-statistics/edit\" target=\"_blank\">cafa5-competition-statistics</a> notebook and here is the final public LB / medal thresholds plot:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F2ea73c8d7116dc9d88f1f9526c6b34c2%2FImage%208-21-23%20at%205.17%20PM.jpeg?generation=1692663529867804&amp;alt=media\" alt=\"\"></p>"
    },
    {
      "id": 2400185,
      "authorName": "Samuel Cortinhas",
      "votes": 2,
      "postDate": "2023-08-20T21:47:16.360000",
      "content": "<p>What's the general feeling of a shakeup happening?</p>"
    },
    {
      "id": 2400301,
      "authorName": "Oleksiy Kononenko",
      "votes": 0,
      "postDate": "2023-08-21T01:07:30.330000",
      "content": "<p>I’ve joined very late, so not sure. May be those who have a better feeling of the competition could comment.</p>"
    },
    {
      "id": 2400329,
      "authorName": "",
      "votes": 0,
      "postDate": "2023-08-21T02:08:40.037000",
      "content": ""
    },
    {
      "id": 2400341,
      "authorName": "Oleksiy Kononenko",
      "votes": 1,
      "postDate": "2023-08-21T02:29:48.003000",
      "content": "<p>For ICR the shake-up was caused by a significant difference between the public and private test data distributions. Public/private scores discrepancy was major even for those who used their own models. The key was to use minimal feature engineering and very basic classifier parameters, sometimes even just defaults.</p>"
    },
    {
      "id": 2400351,
      "authorName": "Tilii",
      "votes": 7,
      "postDate": "2023-08-21T02:50:14.597000",
      "content": "<p>It is impossible to tell how big of a shakeup will happen, as not even the organizers know the private test at this point. It will depend to some degree on luck. My understanding is that an objectively better model could end up looking worse if the other model better predicts the private set of proteins that acquire new annotations by the end of this year. To a degree, we are all at the mercy of those providing solid annotations.</p>\n<p>I think there will be a shakeup, but maybe for half the competitors in top 50. I would assume that many of the quiet guys in top 20 that have been in this competition before will stay as they are or move upward.</p>"
    },
    {
      "id": 2400117,
      "authorName": "Michalina Hulak",
      "votes": 1,
      "postDate": "2023-08-20T19:33:34.790000",
      "content": "<p>Nice :) thanks </p>"
    }
  ],
  "index": {
    "id": "433184",
    "title": "CAFA5: last day competition stats",
    "authorName": "",
    "commentCount": "20",
    "votes": "10",
    "postDate": "2023-08-20 16:49:15.619000"
  },
  "competition": "cafa-5-protein-function-prediction"
}