{
  "id": 189459,
  "title": "How to add tabular_data with your image_model",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/189459",
  "author_name": "Nitin Datta",
  "post_date": "2020-10-07T17:00:06.838000",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi everyone, </p>\n<p>We have a lot of information from train.csv. <br>\nIs anyone using this tabular data along with images?<br>\nIf yes can anyone share their ideas on how to add tabular data with an image model.</p>\n<p>Thanks in advance.</p>",
  "messages": [
    {
      "id": 1041274,
      "postDate": "2020-10-07T17:00:06.837Z",
      "content": "<p>Hi everyone, </p>\n<p>We have a lot of information from train.csv. <br>\nIs anyone using this tabular data along with images?<br>\nIf yes can anyone share their ideas on how to add tabular data with an image model.</p>\n<p>Thanks in advance.</p>",
      "rawMarkdown": "Hi everyone, \n\nWe have a lot of information from train.csv. \nIs anyone using this tabular data along with images?\nIf yes can anyone share their ideas on how to add tabular data with an image model.\n\nThanks in advance.\n",
      "votes": 4
    },
    {
      "id": 1041373,
      "postDate": "2020-10-07T17:49:32.773Z",
      "content": "<p>Try to utilise it in TFRecord format like Chris Deotte has done in SIIM melanoma, <strong>or</strong> you could go and turn images into tabular data using <a href=\"https://www.kaggle.com/christofhenkel/extract-image-features-from-pretrained-nn\" target=\"_blank\">this kernel by Dieter</a>.</p>",
      "rawMarkdown": "Try to utilise it in TFRecord format like Chris Deotte has done in SIIM melanoma, **or** you could go and turn images into tabular data using [this kernel by Dieter](https://www.kaggle.com/christofhenkel/extract-image-features-from-pretrained-nn).",
      "votes": 4,
      "replies": [
        {
          "id": 1041892,
          "postDate": "2020-10-08T00:50:39.080Z",
          "content": "<p>Thanks for the notebook <a href=\"https://www.kaggle.com/nxrprime\" target=\"_blank\">@nxrprime</a> <br>\nIt was helpful </p>",
          "rawMarkdown": "Thanks for the notebook @nxrprime \nIt was helpful "
        }
      ]
    },
    {
      "id": 1041286,
      "postDate": "2020-10-07T17:10:12.643Z",
      "content": "<p>i dont think its a good idea to add features from train.csv into the model. the reason is , while running on test data, we will not be provided with these additional data columns that are provided in train.csv. test.csv has only ID columns.</p>",
      "rawMarkdown": "i dont think its a good idea to add features from train.csv into the model. the reason is , while running on test data, we will not be provided with these additional data columns that are provided in train.csv. test.csv has only ID columns.",
      "votes": 1,
      "replies": [
        {
          "id": 1041306,
          "postDate": "2020-10-07T17:16:57.887Z",
          "content": "<p>I understand that but can you explain how you are dealing with the submission_csv where we have UID+label combination??</p>",
          "rawMarkdown": "I understand that but can you explain how you are dealing with the submission_csv where we have UID+label combination??"
        },
        {
          "id": 1041438,
          "postDate": "2020-10-07T18:33:28.920Z",
          "content": "<p>i am not sure if you are referring to the code or logic behind the UID+label combination. the expectation for the submission is to have StudyInstanceUID + exam level predictions and SOPInstanceUID should be paired with image-level prediction. <br>\nthere could be several reasons for getting submission scoring error. one of them could be your pipeline missing some dicom file as there are few files in private test set that are not readable by pydicom .</p>\n<p>the following discussion might help in the second case. <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/187823\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/187823</a></p>",
          "rawMarkdown": "i am not sure if you are referring to the code or logic behind the UID+label combination. the expectation for the submission is to have StudyInstanceUID + exam level predictions and SOPInstanceUID should be paired with image-level prediction. \nthere could be several reasons for getting submission scoring error. one of them could be your pipeline missing some dicom file as there are few files in private test set that are not readable by pydicom .\n\nthe following discussion might help in the second case. https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/187823"
        },
        {
          "id": 1041893,
          "postDate": "2020-10-08T00:52:21.933Z",
          "content": "<p>I do not get any submission errors, but I have a ton of label inconsistencies…<br>\nThis bothers me the most </p>",
          "rawMarkdown": "I do not get any submission errors, but I have a ton of label inconsistencies...\nThis bothers me the most "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1041373,
      "author_name": "Trigram",
      "author_url": "",
      "post_date": "2020-10-07T17:49:32.773000",
      "content": "<p>Try to utilise it in TFRecord format like Chris Deotte has done in SIIM melanoma, <strong>or</strong> you could go and turn images into tabular data using <a href=\"https://www.kaggle.com/christofhenkel/extract-image-features-from-pretrained-nn\" target=\"_blank\">this kernel by Dieter</a>.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1041892,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2020-10-08T00:50:39.080000",
          "content": "<p>Thanks for the notebook <a href=\"https://www.kaggle.com/nxrprime\" target=\"_blank\">@nxrprime</a> <br>\nIt was helpful </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1041286,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "2020-10-07T17:10:12.643000",
      "content": "<p>i dont think its a good idea to add features from train.csv into the model. the reason is , while running on test data, we will not be provided with these additional data columns that are provided in train.csv. test.csv has only ID columns.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1041306,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2020-10-07T17:16:57.887000",
          "content": "<p>I understand that but can you explain how you are dealing with the submission_csv where we have UID+label combination??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041438,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "2020-10-07T18:33:28.920000",
          "content": "<p>i am not sure if you are referring to the code or logic behind the UID+label combination. the expectation for the submission is to have StudyInstanceUID + exam level predictions and SOPInstanceUID should be paired with image-level prediction. <br>\nthere could be several reasons for getting submission scoring error. one of them could be your pipeline missing some dicom file as there are few files in private test set that are not readable by pydicom .</p>\n<p>the following discussion might help in the second case. <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/187823\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/187823</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1041893,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2020-10-08T00:52:21.933000",
          "content": "<p>I do not get any submission errors, but I have a ton of label inconsistencies…<br>\nThis bothers me the most </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1041274": "Hi everyone, \n\nWe have a lot of information from train.csv. \nIs anyone using this tabular data along with images?\nIf yes can anyone share their ideas on how to add tabular data with an image model.\n\nThanks in advance.\n",
    "1041373": "Try to utilise it in TFRecord format like Chris Deotte has done in SIIM melanoma, **or** you could go and turn images into tabular data using [this kernel by Dieter](https://www.kaggle.com/christofhenkel/extract-image-features-from-pretrained-nn).",
    "1041286": "i dont think its a good idea to add features from train.csv into the model. the reason is , while running on test data, we will not be provided with these additional data columns that are provided in train.csv. test.csv has only ID columns."
  }
}