{
  "competition": "cafa-5-protein-function-prediction",
  "topic_id": "433091",
  "comments": [],
  "messages": [],
  "raw_show": {
    "topic": {
      "id": 433091,
      "title": "Ensembling of public test set submissions ?",
      "authorName": "Henri Upton",
      "commentCount": 7,
      "votes": 1,
      "postDate": "2023-08-20T08:58:55.430000"
    },
    "comments": [
      {
        "id": 2399378,
        "authorName": "Alexander Chervov",
        "votes": 5,
        "postDate": "2023-08-20T10:00:38.507000",
        "content": "<p>Top public notebook are overfitting to public LB.<br>\nIn particular gradually people understood that parameter \"35\" from original MT's notebook \"merge_datatasets\", should be increased - see recent public notebooks.  And a couple of other changes.</p>\n<p>PS<br>\nI thought it is explained in the post here<br>\n<a href=\"https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/432907\" target=\"_blank\">https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/432907</a><br>\nbut that was misunderstanding, sorry</p>"
      },
      {
        "id": 2400347,
        "authorName": "Tilii",
        "votes": 3,
        "postDate": "2023-08-21T02:42:44.490000",
        "content": "<p>I show below cumulative distribution functions (CDFs) for a \"regular\" submission and for one of those public merge notebooks. Not only does public CDF predict a huge proportion of points at &gt;0.5 probability, but you can also observe those horizontal parts where unnatural rounding or clipping was done to the data. Not saying it is impossible that these merged predictions will work on private data, but I would not put my money on that horse.</p>\n<p><img src=\"https://i.ibb.co/zNHbcFn/distributions-submission-05-public-merge.png\" alt=\"CDF comparison\"></p>"
      },
      {
        "id": 2400361,
        "authorName": "ZMC",
        "votes": 1,
        "postDate": "2023-08-21T03:11:39.597000",
        "content": "<p>Some of the models that end up within the public ensembles do produce non continuous predictions so the flat parts in the CDF is actually unsurprising! I do not completely understand why that might make these predictions less robust but it is an interesting observation.</p>"
      },
      {
        "id": 2403007,
        "authorName": "John Mitchell",
        "votes": 1,
        "postDate": "2023-08-22T13:06:02.010000",
        "content": "<p>Thanks <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, that's a really interesting plot. Noting that this lies below the y=x line, and approximately differentiating in my head. I don't know if your 'regular submission' is a single model or ensemble, but it looks as I'd expect a single model to. One of the interesting things about this competition is that there are many ways to make an ensemble, which treat the 'probability' in different manners. My team have some ensembles designed to manipulate the 'probability' in ways that reflect our estimate of the precision of the scoring (i.e., above best F1 threshold) portion of each model. They will have 'weird' CDFs, I'm sure.</p>\n<p>Some other thoughts on the 'probability'. I believe there's no requirement for it to scale as a probability, just for it to rank the predictions within [0.001, 1.0] as per the order of our genuine confidence (or estimate of probability) for each, and to disperse them sensibly with regard to the trial thresholds (i.e., the multiples of 0.001; namely 0.001, 0.002 … 0.998, 0.999, 1.000) such that not too many predictions get the same score to 3 dp. There may even be a small benefit in morphing the distribution to allow more trial thrsholds to appear in the region where we expect to find the scoring threshold (we didn't do this).</p>"
      },
      {
        "id": 2403904,
        "authorName": "Tilii",
        "votes": 1,
        "postDate": "2023-08-22T23:54:41.617000",
        "content": "<blockquote>\n  <p>I don't know if your 'regular submission' is a single model or ensemble</p>\n</blockquote>\n<p>It is an ensemble of three models (one for each ontology category) but without any funny business - a simple concatenation.</p>\n<p>It is to be seen at a later time whether manipulating probabilities by thresholding and/or weighting is a better approach. Part of me is happy that I didn't have time to dig into that and was doing the simplest ensembling possible, though I may come to regret that come Christmas.</p>"
      },
      {
        "id": 2399553,
        "authorName": "Michalina Hulak",
        "votes": 3,
        "postDate": "2023-08-20T12:26:05.747000",
        "content": "<p>there is no private test dataset. There is only a public test dataset from which proteins will be selected, and the final score will be calculated based on these proteins later on.</p>"
      },
      {
        "id": 2399327,
        "authorName": "John Mitchell",
        "votes": 3,
        "postDate": "2023-08-20T09:13:07.143000",
        "content": "<p>Every protein in the private test set is within the visible test set that our models have already predicted. There is no 'hidden' set.</p>\n<p>This is only nominally a code competition. It is fine to prepare submission.tsv offline and then submit by attaching that file to a notebook.</p>"
      }
    ]
  },
  "topic": null,
  "index": {
    "id": "433091",
    "title": "Ensembling of public test set submissions ?",
    "authorName": "",
    "commentCount": "7",
    "votes": "1",
    "postDate": "2023-08-20 08:58:55.430000"
  }
}