{
  "topic": {
    "id": 434064,
    "title": "Private 2nd/Public 5th solution: Py-Boost and GCN",
    "authorName": "Btbpanda",
    "commentCount": 7,
    "votes": 59,
    "postDate": "2023-08-23T20:42:14.987000"
  },
  "comments": [
    {
      "id": 2405383,
      "authorName": "Marília Prata",
      "votes": 7,
      "postDate": "2023-08-23T20:57:01.880000",
      "content": "<p>Thank you for sharing yours Py-Boost and GCN solution, on such high level competition.<br>\nCongratulations too for the 5th position.</p>"
    },
    {
      "id": 2405585,
      "authorName": "hoho",
      "votes": 6,
      "postDate": "2023-08-24T02:04:51.050000",
      "content": "<p>Thanks for sharing!</p>\n<ul>\n<li>1. I think using taxons is a great idea, but I didn't see anyone using it in public notebook. Selecting the features in taxons is critical, could you please give me more information about it? Did you use term label or ESM features to determine which \"are good enough represented\"? How much improvement when concat these 30 features?</li>\n<li>3 and 6. Great idea. I think they were not wrong assumptions, but some hidden condition in the propagation rules and priori knowledge. But very difficult to understand and implement into the model or pre-/post-processing. Especially the postprocessing part, I think it is the \"simplest\" way to boost the score, but I didn't make it work. I'm still trying to understand it better. </li>\n</ul>\n<p>BTW, I expect a big shake-up, but the shake-up will not \"shake\" you, haha, congratulations in advance!</p>"
    },
    {
      "id": 2412537,
      "authorName": "Btbpanda",
      "votes": 2,
      "postDate": "2023-08-28T10:44:03.967000",
      "content": "<p>Thanks for your feedback. Yes, applying the graph transformations was a little bit tricky and I spent a lot of time to learn it better and write the efficient implementations. So, I am not surpised, it could be difficult to understand, but I have no idea, how to explain it more clear unfortunately.</p>\n<p>I used only taxon frequencies to select. I made one-hot features only for the taxons, that has frequency above the threshold (I don't remember exactly, I just saved the IDs at the begining after selection, but it is around 80) in both train and test set. </p>"
    },
    {
      "id": 2405544,
      "authorName": "z7777",
      "votes": 6,
      "postDate": "2023-08-24T01:15:49.527000",
      "content": "<p>Thank you for sharing yours sharing btbpanda! Congratulations! Your GCN solution is so amazing!</p>"
    },
    {
      "id": 2406898,
      "authorName": "Mystic Shadow",
      "votes": 4,
      "postDate": "2023-08-24T17:41:41.017000",
      "content": "<p>So great! Nice learning for me :)</p>"
    },
    {
      "id": 2405633,
      "authorName": "Muhammad Usman",
      "votes": 4,
      "postDate": "2023-08-24T03:30:46.917000",
      "content": "<p>Congratulations on your 5th position. </p>\n<p>Your solution was amazing.</p>"
    },
    {
      "id": 2414221,
      "authorName": "Henri Upton",
      "votes": 2,
      "postDate": "2023-08-29T13:35:24.443000",
      "content": "<p>so nice approach, i am so shocked of the level of your ideas so far. Will definitely check for <em>py-boost</em> documentation ! Congratulations for your work :) </p>"
    }
  ],
  "index": {
    "id": "434064",
    "title": "Private 2nd/Public 5th solution: Py-Boost and GCN",
    "authorName": "",
    "commentCount": "7",
    "votes": "59",
    "postDate": "2024-01-14 22:18:27.513000"
  },
  "competition": "cafa-5-protein-function-prediction"
}