{
  "id": 463042,
  "title": "Percentage of the outliers \"other\"",
  "url": "/competitions/UBC-OCEAN/discussion/463042",
  "author_name": "Cyrus",
  "post_date": "2023-12-23T01:33:02.588000",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>What is the percentage of outliers categorized as \"other\" in the dataset?</p>",
  "messages": [
    {
      "id": 2571182,
      "postDate": "2023-12-23T01:33:02.590Z",
      "content": "<p>What is the percentage of outliers categorized as \"other\" in the dataset?</p>",
      "rawMarkdown": "What is the percentage of outliers categorized as \"other\" in the dataset?",
      "votes": 1
    },
    {
      "id": 2584374,
      "postDate": "2024-01-02T21:41:31.073Z",
      "content": "<p>Let me waste a submission to demonstrate the obvious:</p>\n<pre><code> pathlib  Path\n pandas  pd\ndf_test = pd.read_csv(Path())\ndf_test[] = \ndf_test[[, ]].to_csv(, index=)\n</code></pre>\n<p>Score: 0.16</p>\n<p>== 1/6. Accuracy is balanced.</p>",
      "rawMarkdown": "Let me waste a submission to demonstrate the obvious:\n\n```python\nfrom pathlib import Path\nimport pandas as pd\ndf_test = pd.read_csv(Path(\"/kaggle/input/UBC-OCEAN/test.csv\"))\ndf_test[\"label\"] = \"Other\"\ndf_test[[\"image_id\", \"label\"]].to_csv(\"submission.csv\", index=False)\n```\n\nScore: 0.16\n\n== 1/6. Accuracy is balanced."
    },
    {
      "id": 2584063,
      "postDate": "2024-01-02T16:45:52.673Z",
      "content": "<p>There is no way of knowing this.  Because the test set is hidden and the metric used is balanced accuracy there is no way to work backwards to find the % of \"other\" labels in the test set.  </p>",
      "rawMarkdown": "There is no way of knowing this.  Because the test set is hidden and the metric used is balanced accuracy there is no way to work backwards to find the % of \"other\" labels in the test set.  "
    }
  ],
  "comments": [
    {
      "id": 2584374,
      "author_name": "tinkei",
      "author_url": "",
      "post_date": "2024-01-02T21:41:31.073000",
      "content": "<p>Let me waste a submission to demonstrate the obvious:</p>\n<pre><code> pathlib  Path\n pandas  pd\ndf_test = pd.read_csv(Path())\ndf_test[] = \ndf_test[[, ]].to_csv(, index=)\n</code></pre>\n<p>Score: 0.16</p>\n<p>== 1/6. Accuracy is balanced.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2584063,
      "author_name": "Connor",
      "author_url": "",
      "post_date": "2024-01-02T16:45:52.673000",
      "content": "<p>There is no way of knowing this.  Because the test set is hidden and the metric used is balanced accuracy there is no way to work backwards to find the % of \"other\" labels in the test set.  </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2571182": "What is the percentage of outliers categorized as \"other\" in the dataset?",
    "2584374": "Let me waste a submission to demonstrate the obvious:\n\n```python\nfrom pathlib import Path\nimport pandas as pd\ndf_test = pd.read_csv(Path(\"/kaggle/input/UBC-OCEAN/test.csv\"))\ndf_test[\"label\"] = \"Other\"\ndf_test[[\"image_id\", \"label\"]].to_csv(\"submission.csv\", index=False)\n```\n\nScore: 0.16\n\n== 1/6. Accuracy is balanced.",
    "2584063": "There is no way of knowing this.  Because the test set is hidden and the metric used is balanced accuracy there is no way to work backwards to find the % of \"other\" labels in the test set.  "
  }
}