{
  "id": 446102,
  "title": "Why is there no \"is_tma\" in the test csv?",
  "url": "/competitions/UBC-OCEAN/discussion/446102",
  "author_name": "Jonas Gjerris",
  "post_date": "2023-10-10T09:15:08.668000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Can anyone explain the omission of \"is_tma\" from the test csv? Although the information seems discernible—if the image_id isn't in the test_thumbnails folder, then the image is a tma—it feels unnecessarily complicated. Am I missing something?</p>",
  "messages": [
    {
      "id": 2482591,
      "postDate": "2023-10-15T06:45:58.973Z",
      "content": "<p>Yeah it doesn't make sense. You can basically do</p>\n<pre><code>df_test[] = \ndf_test.loc[df_test[].isin(test_thumbnails), ] = \n</code></pre>",
      "rawMarkdown": "Yeah it doesn't make sense. You can basically do\n\n```python\ndf_test['is_tma'] = True\ndf_test.loc[df_test['image_id'].isin(test_thumbnails), 'is_tma'] = False\n```\n",
      "votes": 3,
      "replies": [
        {
          "id": 2483025,
          "postDate": "2023-10-15T12:31:17.727Z",
          "content": "<p>I was starting to complain but this is a very nice fix, thank you.</p>",
          "rawMarkdown": "I was starting to complain but this is a very nice fix, thank you."
        },
        {
          "id": 2483073,
          "postDate": "2023-10-15T13:03:06.863Z",
          "content": "<p>I had to do smth like this tho</p>\n<pre><code>\ntest_thumbnails_folder = \n\n\nthumbnails_filenames = os.listdir(test_thumbnails_folder)\n\n\nimage_ids_in_thumbnails = [filename.split()[]  filename  thumbnails_filenames]\n\n\ntest_df[] = ~test_df[].isin(image_ids_in_thumbnails)\n</code></pre>",
          "rawMarkdown": "I had to do smth like this tho\n\n```python\n# Replace 'test_thumbnails' with the actual path to your folder\ntest_thumbnails_folder = '/kaggle/input/UBC-OCEAN/test_thumbnails'\n\n# Get a list of filenames in the folder\nthumbnails_filenames = os.listdir(test_thumbnails_folder)\n\n# Extract image IDs from filenames\nimage_ids_in_thumbnails = [filename.split('_')[0] for filename in thumbnails_filenames]\n\n# Set 'is_tma' in test_df based on image IDs\ntest_df['is_tma'] = ~test_df['image_id'].isin(image_ids_in_thumbnails)\n```",
          "votes": 2
        },
        {
          "id": 2546914,
          "postDate": "2023-12-03T01:43:05.963Z",
          "content": "<p>Does this still work? It seems that all test images have thumbnails. Am I wrong?</p>",
          "rawMarkdown": "Does this still work? It seems that all test images have thumbnails. Am I wrong?",
          "replies": [
            {
              "id": 2552639,
              "postDate": "2023-12-07T16:18:55.973Z",
              "content": "<p>Did you figure it out?</p>",
              "rawMarkdown": "Did you figure it out?"
            },
            {
              "id": 2562040,
              "postDate": "2023-12-15T03:59:58.570Z",
              "content": "<p>After some testing, it seem there is no TMA in the public LB test data. Johnny, right?</p>",
              "rawMarkdown": "After some testing, it seem there is no TMA in the public LB test data. Johnny, right?"
            },
            {
              "id": 2574595,
              "postDate": "2023-12-26T05:22:41.900Z",
              "content": "<p>you can judge by the width and height</p>",
              "rawMarkdown": "you can judge by the width and height"
            }
          ]
        }
      ]
    },
    {
      "id": 2476001,
      "postDate": "2023-10-10T09:15:08.667Z",
      "content": "<p>Can anyone explain the omission of \"is_tma\" from the test csv? Although the information seems discernible—if the image_id isn't in the test_thumbnails folder, then the image is a tma—it feels unnecessarily complicated. Am I missing something?</p>",
      "rawMarkdown": "Can anyone explain the omission of \"is_tma\" from the test csv? Although the information seems discernible—if the image_id isn't in the test_thumbnails folder, then the image is a tma—it feels unnecessarily complicated. Am I missing something?",
      "votes": 3
    }
  ],
  "comments": [
    {
      "id": 2482591,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-10-15T06:45:58.973000",
      "content": "<p>Yeah it doesn't make sense. You can basically do</p>\n<pre><code>df_test[] = \ndf_test.loc[df_test[].isin(test_thumbnails), ] = \n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 2483025,
          "author_name": "Parham Gousheh",
          "author_url": "",
          "post_date": "2023-10-15T12:31:17.727000",
          "content": "<p>I was starting to complain but this is a very nice fix, thank you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2483073,
          "author_name": "Parham Gousheh",
          "author_url": "",
          "post_date": "2023-10-15T13:03:06.863000",
          "content": "<p>I had to do smth like this tho</p>\n<pre><code>\ntest_thumbnails_folder = \n\n\nthumbnails_filenames = os.listdir(test_thumbnails_folder)\n\n\nimage_ids_in_thumbnails = [filename.split()[]  filename  thumbnails_filenames]\n\n\ntest_df[] = ~test_df[].isin(image_ids_in_thumbnails)\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2546914,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2023-12-03T01:43:05.963000",
          "content": "<p>Does this still work? It seems that all test images have thumbnails. Am I wrong?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2552639,
              "author_name": "Jonas Gjerris",
              "author_url": "",
              "post_date": "2023-12-07T16:18:55.973000",
              "content": "<p>Did you figure it out?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2562040,
              "author_name": "Yoobin",
              "author_url": "",
              "post_date": "2023-12-15T03:59:58.570000",
              "content": "<p>After some testing, it seem there is no TMA in the public LB test data. Johnny, right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2574595,
              "author_name": "Toby",
              "author_url": "",
              "post_date": "2023-12-26T05:22:41.900000",
              "content": "<p>you can judge by the width and height</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2482591": "Yeah it doesn't make sense. You can basically do\n\n```python\ndf_test['is_tma'] = True\ndf_test.loc[df_test['image_id'].isin(test_thumbnails), 'is_tma'] = False\n```\n",
    "2476001": "Can anyone explain the omission of \"is_tma\" from the test csv? Although the information seems discernible—if the image_id isn't in the test_thumbnails folder, then the image is a tma—it feels unnecessarily complicated. Am I missing something?"
  }
}