{
  "id": 418016,
  "title": "What's the point of {train|validation}_metadata.json ?",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/418016",
  "author_name": "JEANMPIA",
  "post_date": "2023-06-18T11:00:56.681000",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p><strong>I don't know why I didn't ask this before,</strong><br>\nbut why doesn't the provided metadata match what we have as a train and valid set ? It would be interesting and most likely useful to have temporal and spacial context for each image.</p>\n<p><em>I don't know if this is a known issue or if it's not an issue at all and i'm the only one who had to make my own metadata but I'd rather ask than not know.</em></p>",
  "messages": [
    {
      "id": 2307649,
      "postDate": "2023-06-18T11:00:56.683Z",
      "content": "<p><strong>I don't know why I didn't ask this before,</strong><br>\nbut why doesn't the provided metadata match what we have as a train and valid set ? It would be interesting and most likely useful to have temporal and spacial context for each image.</p>\n<p><em>I don't know if this is a known issue or if it's not an issue at all and i'm the only one who had to make my own metadata but I'd rather ask than not know.</em></p>",
      "rawMarkdown": "**I don't know why I didn't ask this before,**\nbut why doesn't the provided metadata match what we have as a train and valid set ? It would be interesting and most likely useful to have temporal and spacial context for each image.\n\n*I don't know if this is a known issue or if it's not an issue at all and i'm the only one who had to make my own metadata but I'd rather ask than not know.*",
      "votes": 6
    },
    {
      "id": 2311796,
      "postDate": "2023-06-21T13:05:35.883Z",
      "content": "<p>The record_id field matches the directories as far as I can tell. The first record_id in train_metadata.json, which is \"3283699311323360356\", is also a directory inside the train directory with the appropriate .npy files .. and so on.</p>",
      "rawMarkdown": "The record_id field matches the directories as far as I can tell. The first record_id in train_metadata.json, which is \"3283699311323360356\", is also a directory inside the train directory with the appropriate .npy files .. and so on.",
      "votes": 1,
      "replies": [
        {
          "id": 2311818,
          "postDate": "2023-06-21T13:15:46.917Z",
          "content": "<p>yes, but if you read it using Pandas record_id is broken - you have to declare data type in Pandas for this column. I published below how to read json file properly. </p>",
          "rawMarkdown": "yes, but if you read it using Pandas record_id is broken - you have to declare data type in Pandas for this column. I published below how to read json file properly. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 2310538,
      "postDate": "2023-06-20T12:45:50.883Z",
      "content": "<p>Can you elaborate on what doesn't match?</p>",
      "rawMarkdown": "Can you elaborate on what doesn't match?",
      "replies": [
        {
          "id": 2310745,
          "postDate": "2023-06-20T15:25:47.323Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/aaronsarna\" target=\"_blank\">@aaronsarna</a>,<br>\nThe record_ids from the metadatas are nowhere to be found in the given folders. I tried many things to see if it was a string/int comparaison problem but it is not. Such notebooks are made in response but I don't think that's really optimal for this comp: <a href=\"https://www.kaggle.com/code/takaito/gr-icrgw-make-table-data\" target=\"_blank\">https://www.kaggle.com/code/takaito/gr-icrgw-make-table-data</a></p>",
          "rawMarkdown": "Hello @aaronsarna,\nThe record_ids from the metadatas are nowhere to be found in the given folders. I tried many things to see if it was a string/int comparaison problem but it is not. Such notebooks are made in response but I don't think that's really optimal for this comp: https://www.kaggle.com/code/takaito/gr-icrgw-make-table-data",
          "votes": 2,
          "replies": [
            {
              "id": 2311331,
              "postDate": "2023-06-21T05:39:48.137Z",
              "content": "<p>I am also interested in this. I just started so maybe I misunderstand, but shouldn't record_ids from train_metadata.json match the subfolder names of train/ ?</p>",
              "rawMarkdown": "I am also interested in this. I just started so maybe I misunderstand, but shouldn't record_ids from train_metadata.json match the subfolder names of train/ ?",
              "votes": 1
            },
            {
              "id": 2311814,
              "postDate": "2023-06-21T13:13:13.877Z",
              "content": "<p>Everything is ok. You have to declare data type for long numbers (record_id) in Pandas reading json file. </p>\n<pre><code> = {: str} \n\n = pd.read_json(, dtype=data_types)\n = pd.read_json(, dtype=data_types)\n</code></pre>",
              "rawMarkdown": "Everything is ok. You have to declare data type for long numbers (record_id) in Pandas reading json file. \n\n```\ndata_types = {'record_id': str} \n\ntr_json = pd.read_json(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/train_metadata.json\", dtype=data_types)\nval_json = pd.read_json(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/validation_metadata.json\", dtype=data_types)\n```",
              "votes": 11
            },
            {
              "id": 2312089,
              "postDate": "2023-06-21T17:00:14.630Z",
              "content": "<p>ty, that helps</p>",
              "rawMarkdown": "ty, that helps",
              "votes": 1
            },
            {
              "id": 2312423,
              "postDate": "2023-06-21T21:55:09.230Z",
              "content": "<p>thx ! that will for sure help</p>",
              "rawMarkdown": "thx ! that will for sure help",
              "votes": 1
            },
            {
              "id": 2343221,
              "postDate": "2023-07-13T13:43:26.560Z",
              "content": "<p>I was having the same problem.<br>\nThanks for sharing.</p>",
              "rawMarkdown": "I was having the same problem.\nThanks for sharing.",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2311796,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2023-06-21T13:05:35.883000",
      "content": "<p>The record_id field matches the directories as far as I can tell. The first record_id in train_metadata.json, which is \"3283699311323360356\", is also a directory inside the train directory with the appropriate .npy files .. and so on.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2311818,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2023-06-21T13:15:46.917000",
          "content": "<p>yes, but if you read it using Pandas record_id is broken - you have to declare data type in Pandas for this column. I published below how to read json file properly. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2310538,
      "author_name": "Aaron Sarna",
      "author_url": "",
      "post_date": "2023-06-20T12:45:50.883000",
      "content": "<p>Can you elaborate on what doesn't match?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2310745,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-06-20T15:25:47.323000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/aaronsarna\" target=\"_blank\">@aaronsarna</a>,<br>\nThe record_ids from the metadatas are nowhere to be found in the given folders. I tried many things to see if it was a string/int comparaison problem but it is not. Such notebooks are made in response but I don't think that's really optimal for this comp: <a href=\"https://www.kaggle.com/code/takaito/gr-icrgw-make-table-data\" target=\"_blank\">https://www.kaggle.com/code/takaito/gr-icrgw-make-table-data</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2311331,
              "author_name": "Dieter",
              "author_url": "",
              "post_date": "2023-06-21T05:39:48.137000",
              "content": "<p>I am also interested in this. I just started so maybe I misunderstand, but shouldn't record_ids from train_metadata.json match the subfolder names of train/ ?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2311814,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-06-21T13:13:13.877000",
              "content": "<p>Everything is ok. You have to declare data type for long numbers (record_id) in Pandas reading json file. </p>\n<pre><code> = {: str} \n\n = pd.read_json(, dtype=data_types)\n = pd.read_json(, dtype=data_types)\n</code></pre>",
              "votes": 11,
              "replies": []
            },
            {
              "id": 2312089,
              "author_name": "Dieter",
              "author_url": "",
              "post_date": "2023-06-21T17:00:14.630000",
              "content": "<p>ty, that helps</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2312423,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-21T21:55:09.230000",
              "content": "<p>thx ! that will for sure help</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2343221,
              "author_name": "moaiinthesky",
              "author_url": "",
              "post_date": "2023-07-13T13:43:26.560000",
              "content": "<p>I was having the same problem.<br>\nThanks for sharing.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2307649": "**I don't know why I didn't ask this before,**\nbut why doesn't the provided metadata match what we have as a train and valid set ? It would be interesting and most likely useful to have temporal and spacial context for each image.\n\n*I don't know if this is a known issue or if it's not an issue at all and i'm the only one who had to make my own metadata but I'd rather ask than not know.*",
    "2311796": "The record_id field matches the directories as far as I can tell. The first record_id in train_metadata.json, which is \"3283699311323360356\", is also a directory inside the train directory with the appropriate .npy files .. and so on.",
    "2310538": "Can you elaborate on what doesn't match?"
  }
}