{
  "topic": {
    "id": 552449,
    "title": "Competition Upgrade Plan Details",
    "authorName": "Sohier Dane",
    "commentCount": 12,
    "votes": 21,
    "postDate": "2024-12-19T19:14:50.817000"
  },
  "comments": [
    {
      "id": 3076699,
      "authorName": "Khoi Nguyen",
      "votes": 18,
      "postDate": "2024-12-20T07:39:37.380000",
      "content": "<p>I second the idea of extending the duration, this is not a 3 month task. I'd also love to see a sample submission that goes beyond 0.00, or to get a confirmation that Kaggle team has managed that when preparing this competition.</p>"
    },
    {
      "id": 3076750,
      "authorName": "Thanh Hau Nguyen",
      "votes": 1,
      "postDate": "2024-12-20T08:25:28.263000",
      "content": "<p>Looking at the current scores on swebench leaderboard, I'm still wondering if the best submission of this competition would be better than skipping everything (0.00) :D And one other option to remove the penalization in the metric.</p>"
    },
    {
      "id": 3076757,
      "authorName": "Khoi Nguyen",
      "votes": 1,
      "postDate": "2024-12-20T08:30:35.850000",
      "content": "<p>The metric is incredibly ambitious indeed</p>"
    },
    {
      "id": 3077030,
      "authorName": "",
      "votes": 0,
      "postDate": "2024-12-20T13:45:31.127000",
      "content": ""
    },
    {
      "id": 3076560,
      "authorName": "Pavel Orlov",
      "votes": 7,
      "postDate": "2024-12-20T04:23:45.267000",
      "content": "<p>Maybe it would be worth extending the competition duration by 2-3 months?</p>"
    },
    {
      "id": 3076696,
      "authorName": "Gabriel Mirea",
      "votes": 6,
      "postDate": "2024-12-20T07:35:57.917000",
      "content": "<blockquote>\n  <p>Additional train data. We're hoping to add another few dozen instances to the train set, with a very rough timeline of early February.</p>\n</blockquote>\n<p>I think releasing a limited number of actual tests (5-10%?) would be overall better for the competition. Since the public LB isn't going to matter anyway, the integrity of the competition is not impacted. On the other hand, having a few \"ground truth\" entries would allow us to validate the tooling, repo id, format and everything else that the test env expects. </p>"
    },
    {
      "id": 3078665,
      "authorName": "Sohier Dane",
      "votes": 2,
      "postDate": "2024-12-22T16:35:42.367000",
      "content": "<p>For anyone following this thread via notifications, please see the metric section that I've added to the main post as an update.</p>"
    },
    {
      "id": 3082759,
      "authorName": "Camaro",
      "votes": 0,
      "postDate": "2024-12-28T15:33:23.853000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<blockquote>\n  <p>We're going to look into providing access to the test set environments to make it possible for your submitted agents to run unit tests on patches as recommended in this post. This would change the predict function signature. We're especially interested in any feedback you might have on this item.</p>\n</blockquote>\n<p>This item is extremely important. While providing a Docker-based test environment similar to SWE-Bench would be desirable, I feel it might not be compatible with Kaggle's private test environment. The most realistic approach would be to include a pre-verified virtual environment (e.g., .venv) in the repository. Ideally, a function would be provided as an argument to the predict function, which takes a patch and returns the test results. It would be even better if there was a way to add arbitrary tests, but that seems a bit too complex. <br>\nIn any case, it is essential to provide information about the test execution environment that is the same as, or at least on par with, SWE-Bench.</p>"
    },
    {
      "id": 3083729,
      "authorName": "Jim White",
      "votes": 0,
      "postDate": "2024-12-30T00:33:10.280000",
      "content": "<p>I disagree, as I commented in the other discussion (<a href=\"https://www.kaggle.com/competitions/konwinski-prize/discussion/551236#3082310)\" target=\"_blank\">https://www.kaggle.com/competitions/konwinski-prize/discussion/551236#3082310)</a>.</p>\n<p>Moving the goal post beyond SWE-bench to evaluating coding agents for real-world usage means not only using unseen test cases but also doing the environment setup.  Environment setup is the required first step for issue resolution.</p>\n<p>I agree that it (env setup) is a difficult problem and could be a shared task in its own right.  I proposed exactly that in 2014.  But I don't think it is feasible to use README-EVAL's metric for the K-prize (it depended on having working solutions to score against, although they were naturally occurring rather than \"gold labels\" per se).</p>\n<p><em>Towards README-EVAL : Interpreting README File Instructions</em><br>\n<a href=\"https://aclanthology.org/W14-2415/\" target=\"_blank\">https://aclanthology.org/W14-2415/</a></p>"
    },
    {
      "id": 3083733,
      "authorName": "Camaro",
      "votes": 1,
      "postDate": "2024-12-30T00:55:18.110000",
      "content": "<p>Thank you for the comment!<br>\nI agree with your opinion for real-world applications, setting up the environment is a tedious task and should be automated.<br>\nHowever, the problem here is that the Kaggle test environment is completely offline. How can we set up an unknown repository without internet access? In my opinion, it must be pre-configured to make the issue solvable.</p>"
    },
    {
      "id": 3083736,
      "authorName": "Jim White",
      "votes": 0,
      "postDate": "2024-12-30T01:10:27.417000",
      "content": "<p>I wonder that myself but I believe the installation issue can be addressed by not requiring package versions that aren't already installed or at least cached.  It is not necessary to have Internet access to run <code>apt install</code> (or <code>apt-get</code>) and <code>pip install</code>.  The lack of access just means <code>apt update</code> doesn't work.  Working with multiple Python versions and environments is also normal.</p>"
    },
    {
      "id": 3077026,
      "authorName": "Quin O’Malley",
      "votes": 0,
      "postDate": "2024-12-20T13:41:14.053000",
      "content": "<p>Do you recommend that we make our own train data? </p>"
    }
  ],
  "index": {
    "id": "552449",
    "title": "Competition Upgrade Plan Details",
    "authorName": "",
    "commentCount": "12",
    "votes": "21",
    "postDate": "2024-12-19 19:14:50.817000"
  },
  "competition": "konwinski-prize"
}