{
  "id": 445592,
  "title": "'Other' class - outlier identification",
  "url": "/competitions/UBC-OCEAN/discussion/445592",
  "author_name": "Bill Ray",
  "post_date": "2023-10-07T19:00:38.677000",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>The <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\" target=\"_blank\">Data</a> section of the challenge description indicates the label will be:</p>\n<blockquote>\n  <p>one of these subtypes of ovarian cancer: CC, EC, HGSC, LGSC, MC, Other. The Other class is not present in the training set; identifying outliers is one of the challenges of this competition. Only available for the train set.</p>\n</blockquote>\n<p>Will the Other class be included in the test set? If this class is present in the test set but not the training set, under what circumstances should our systems assign an 'Other' label to samples in the test set?</p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": 2477532,
      "postDate": "2023-10-11T11:05:05.713Z",
      "content": "<p>My guess is that outliers/other refers to mixed, uncommon, unusual, or rare tumors; all of which can have varying appearances. <br>\nAlmost all types of tumors have increased DNA content, which makes the cell nucleus bigger and takes up more of the hematoxylin (blue-purple) stain; essentially resulting in areas with many pixels that are dark blue to dark purple. <br>\nI see the handling of the task as pathologist and AI model in a similar way. Since it's already known that there is a tumor, viewing the images and focusing on the generally blue-purple tumor areas (which may be mixed with other colored areas) will push for deciding which class it fits best. If there are multiple classes seen, or doesn't fit any class well, then it would be time to seek help from other colleagues, or in the AI's case send an alert for a case that needs more attention.</p>",
      "rawMarkdown": "My guess is that outliers/other refers to mixed, uncommon, unusual, or rare tumors; all of which can have varying appearances. \nAlmost all types of tumors have increased DNA content, which makes the cell nucleus bigger and takes up more of the hematoxylin (blue-purple) stain; essentially resulting in areas with many pixels that are dark blue to dark purple. \nI see the handling of the task as pathologist and AI model in a similar way. Since it's already known that there is a tumor, viewing the images and focusing on the generally blue-purple tumor areas (which may be mixed with other colored areas) will push for deciding which class it fits best. If there are multiple classes seen, or doesn't fit any class well, then it would be time to seek help from other colleagues, or in the AI's case send an alert for a case that needs more attention.",
      "votes": 13,
      "replies": [
        {
          "id": 2482598,
          "postDate": "2023-10-15T06:59:53.660Z",
          "content": "<p>Thanks for sharing your insights. Is it possible for a sample to have multiple classes?</p>",
          "rawMarkdown": "Thanks for sharing your insights. Is it possible for a sample to have multiple classes?",
          "replies": [
            {
              "id": 2482713,
              "postDate": "2023-10-15T08:15:54.460Z",
              "content": "<p>In the field of medicine, almost anything is possible. </p>\n<p>For example <br>\n<a href=\"url\" target=\"_blank\">https://journals.lww.com/intjgynpathology/Abstract/2021/05000/Ovarian_Mixed_Epithelial_Carcinoma_With_Extensive.16.aspx</a></p>\n<blockquote>\n  <p>Seromucinous carcinoma of the ovary was a newly defined category in the revised 2014 World Health Organization Classification of Tumors of Female Reproductive Organs. It was defined as a carcinoma composed of predominantly of serous and endocervical-type mucinous epithelium. Foci containing clear cells, and areas of endometrioid and squamous differentiation are not uncommon. It is a rare entity with morphologic and immunophenotypic features overlapping other types of ovarian carcinoma.&gt;</p>\n</blockquote>\n<p>It's essentially saying that Endometrioid Carcinoma (EC) has a rare subtype that can have patches containing any class, and the deciding factor is partly based on which class is predominant, and which class is not present. Technically this applies to almost every tumor.</p>",
              "rawMarkdown": "In the field of medicine, almost anything is possible. \n\nFor example \n[https://journals.lww.com/intjgynpathology/Abstract/2021/05000/Ovarian_Mixed_Epithelial_Carcinoma_With_Extensive.16.aspx](url)\n\n>Seromucinous carcinoma of the ovary was a newly defined category in the revised 2014 World Health Organization Classification of Tumors of Female Reproductive Organs. It was defined as a carcinoma composed of predominantly of serous and endocervical-type mucinous epithelium. Foci containing clear cells, and areas of endometrioid and squamous differentiation are not uncommon. It is a rare entity with morphologic and immunophenotypic features overlapping other types of ovarian carcinoma.>\n\nIt's essentially saying that Endometrioid Carcinoma (EC) has a rare subtype that can have patches containing any class, and the deciding factor is partly based on which class is predominant, and which class is not present. Technically this applies to almost every tumor.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2511998,
      "postDate": "2023-11-04T05:39:39.537Z",
      "content": "<p>I'm thinking of two approaches to detect the 'Other' class:</p>\n<ol>\n<li>Train an autoencoder, and use it to find outliers. On the test set, predict 'Other' for outliers.</li>\n<li>Formulate the problem as multi-label classification and use a probability threshold to predict the correct class. By using the multi-label approach, we avoid using the softmax activation function (which 'forces' the highest probability to be too high, even if it's an incorrect class prediction). Softmax is replaced by sigmoid. If for a particular example, ALL the sigmoid probabilities are lower than the probability threshold, then we can infer that it's \"none of the classes in the training set\", and hence likely to be the 'Other' class.</li>\n</ol>",
      "rawMarkdown": "I'm thinking of two approaches to detect the 'Other' class:\n\n1. Train an autoencoder, and use it to find outliers. On the test set, predict 'Other' for outliers.\n2. Formulate the problem as multi-label classification and use a probability threshold to predict the correct class. By using the multi-label approach, we avoid using the softmax activation function (which 'forces' the highest probability to be too high, even if it's an incorrect class prediction). Softmax is replaced by sigmoid. If for a particular example, ALL the sigmoid probabilities are lower than the probability threshold, then we can infer that it's \"none of the classes in the training set\", and hence likely to be the 'Other' class.",
      "votes": 7,
      "replies": [
        {
          "id": 2554941,
          "postDate": "2023-12-09T14:46:40.993Z",
          "content": "<p>Thanks for sharing your idea. </p>\n<p>While using an autoencoder for outlier detection is an interesting idea, its effectiveness might be limited, particularly with high-dimensional data. Autoencoders excel in learning and compressing data representations, but they can struggle when the input data has very high dimensionality. This challenge arises because learning a comprehensive and accurate representation in such cases becomes more complex, potentially leading to issues like overfitting or inadequate feature learning. Additionally, training autoencoders on high-dimensional datasets can be computationally demanding. Therefore, while autoencoders could be part of the solution, relying solely on them for outlier detection in high-dimensional spaces might not yield optimal results. It could be beneficial to explore other methods or a combination of techniques tailored to the specific characteristics of the dataset.</p>",
          "rawMarkdown": "Thanks for sharing your idea. \n\nWhile using an autoencoder for outlier detection is an interesting idea, its effectiveness might be limited, particularly with high-dimensional data. Autoencoders excel in learning and compressing data representations, but they can struggle when the input data has very high dimensionality. This challenge arises because learning a comprehensive and accurate representation in such cases becomes more complex, potentially leading to issues like overfitting or inadequate feature learning. Additionally, training autoencoders on high-dimensional datasets can be computationally demanding. Therefore, while autoencoders could be part of the solution, relying solely on them for outlier detection in high-dimensional spaces might not yield optimal results. It could be beneficial to explore other methods or a combination of techniques tailored to the specific characteristics of the dataset.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2472981,
      "postDate": "2023-10-07T19:00:38.677Z",
      "content": "<p>The <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\" target=\"_blank\">Data</a> section of the challenge description indicates the label will be:</p>\n<blockquote>\n  <p>one of these subtypes of ovarian cancer: CC, EC, HGSC, LGSC, MC, Other. The Other class is not present in the training set; identifying outliers is one of the challenges of this competition. Only available for the train set.</p>\n</blockquote>\n<p>Will the Other class be included in the test set? If this class is present in the test set but not the training set, under what circumstances should our systems assign an 'Other' label to samples in the test set?</p>\n<p>Thank you.</p>",
      "rawMarkdown": "The [Data](https://www.kaggle.com/competitions/UBC-OCEAN/data) section of the challenge description indicates the label will be:\n> one of these subtypes of ovarian cancer: CC, EC, HGSC, LGSC, MC, Other. The Other class is not present in the training set; identifying outliers is one of the challenges of this competition. Only available for the train set.\n\nWill the Other class be included in the test set? If this class is present in the test set but not the training set, under what circumstances should our systems assign an 'Other' label to samples in the test set?\n\nThank you.",
      "votes": 6
    },
    {
      "id": 2482602,
      "postDate": "2023-10-15T07:03:54.777Z",
      "content": "<p>I think it's pretty much self-explanatory. Anything other than CC, EC, HGSC, LGSC, MC is <strong>Other</strong>.</p>",
      "rawMarkdown": "I think it's pretty much self-explanatory. Anything other than CC, EC, HGSC, LGSC, MC is **Other**.",
      "votes": 1
    },
    {
      "id": 2527101,
      "postDate": "2023-11-16T09:46:39.883Z",
      "content": "<p>I was thinking that maybe the test set has normal images, and normal images also belong to the others categories?</p>",
      "rawMarkdown": "I was thinking that maybe the test set has normal images, and normal images also belong to the others categories?"
    },
    {
      "id": 2496283,
      "postDate": "2023-10-23T22:29:15.787Z",
      "content": "<p>I have the same question. In the description is written: \"The Other class is not present in the training set\" and then, in the next sentence it says \"Only available for the train set.\"<br>\nSo what does it mean? \"Other\" is not present in the training set but will be present in the test set?</p>",
      "rawMarkdown": "I have the same question. In the description is written: \"The Other class is not present in the training set\" and then, in the next sentence it says \"Only available for the train set.\"\nSo what does it mean? \"Other\" is not present in the training set but will be present in the test set?",
      "replies": [
        {
          "id": 2499667,
          "postDate": "2023-10-26T07:06:07.937Z",
          "content": "<blockquote>\n  <p>\"Only available for the train set.\"</p>\n</blockquote>\n<p>It means that \"Label\" is only available for train set. This sentence is not the explanation of \"Other class\" but \"Label\"</p>",
          "rawMarkdown": ">\"Only available for the train set.\"\n\nIt means that \"Label\" is only available for train set. This sentence is not the explanation of \"Other class\" but \"Label\""
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2477532,
      "author_name": "Noli Alonso",
      "author_url": "",
      "post_date": "2023-10-11T11:05:05.713000",
      "content": "<p>My guess is that outliers/other refers to mixed, uncommon, unusual, or rare tumors; all of which can have varying appearances. <br>\nAlmost all types of tumors have increased DNA content, which makes the cell nucleus bigger and takes up more of the hematoxylin (blue-purple) stain; essentially resulting in areas with many pixels that are dark blue to dark purple. <br>\nI see the handling of the task as pathologist and AI model in a similar way. Since it's already known that there is a tumor, viewing the images and focusing on the generally blue-purple tumor areas (which may be mixed with other colored areas) will push for deciding which class it fits best. If there are multiple classes seen, or doesn't fit any class well, then it would be time to seek help from other colleagues, or in the AI's case send an alert for a case that needs more attention.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 2482598,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-10-15T06:59:53.660000",
          "content": "<p>Thanks for sharing your insights. Is it possible for a sample to have multiple classes?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2482713,
              "author_name": "Noli Alonso",
              "author_url": "",
              "post_date": "2023-10-15T08:15:54.460000",
              "content": "<p>In the field of medicine, almost anything is possible. </p>\n<p>For example <br>\n<a href=\"url\" target=\"_blank\">https://journals.lww.com/intjgynpathology/Abstract/2021/05000/Ovarian_Mixed_Epithelial_Carcinoma_With_Extensive.16.aspx</a></p>\n<blockquote>\n  <p>Seromucinous carcinoma of the ovary was a newly defined category in the revised 2014 World Health Organization Classification of Tumors of Female Reproductive Organs. It was defined as a carcinoma composed of predominantly of serous and endocervical-type mucinous epithelium. Foci containing clear cells, and areas of endometrioid and squamous differentiation are not uncommon. It is a rare entity with morphologic and immunophenotypic features overlapping other types of ovarian carcinoma.&gt;</p>\n</blockquote>\n<p>It's essentially saying that Endometrioid Carcinoma (EC) has a rare subtype that can have patches containing any class, and the deciding factor is partly based on which class is predominant, and which class is not present. Technically this applies to almost every tumor.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2511998,
      "author_name": "Sadhaklal",
      "author_url": "",
      "post_date": "2023-11-04T05:39:39.537000",
      "content": "<p>I'm thinking of two approaches to detect the 'Other' class:</p>\n<ol>\n<li>Train an autoencoder, and use it to find outliers. On the test set, predict 'Other' for outliers.</li>\n<li>Formulate the problem as multi-label classification and use a probability threshold to predict the correct class. By using the multi-label approach, we avoid using the softmax activation function (which 'forces' the highest probability to be too high, even if it's an incorrect class prediction). Softmax is replaced by sigmoid. If for a particular example, ALL the sigmoid probabilities are lower than the probability threshold, then we can infer that it's \"none of the classes in the training set\", and hence likely to be the 'Other' class.</li>\n</ol>",
      "votes": 7,
      "replies": [
        {
          "id": 2554941,
          "author_name": "Mohammad Dehghanmanshadi",
          "author_url": "",
          "post_date": "2023-12-09T14:46:40.993000",
          "content": "<p>Thanks for sharing your idea. </p>\n<p>While using an autoencoder for outlier detection is an interesting idea, its effectiveness might be limited, particularly with high-dimensional data. Autoencoders excel in learning and compressing data representations, but they can struggle when the input data has very high dimensionality. This challenge arises because learning a comprehensive and accurate representation in such cases becomes more complex, potentially leading to issues like overfitting or inadequate feature learning. Additionally, training autoencoders on high-dimensional datasets can be computationally demanding. Therefore, while autoencoders could be part of the solution, relying solely on them for outlier detection in high-dimensional spaces might not yield optimal results. It could be beneficial to explore other methods or a combination of techniques tailored to the specific characteristics of the dataset.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2482602,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-10-15T07:03:54.777000",
      "content": "<p>I think it's pretty much self-explanatory. Anything other than CC, EC, HGSC, LGSC, MC is <strong>Other</strong>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2527101,
      "author_name": "Metavers",
      "author_url": "",
      "post_date": "2023-11-16T09:46:39.883000",
      "content": "<p>I was thinking that maybe the test set has normal images, and normal images also belong to the others categories?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2496283,
      "author_name": "Ali Shadman",
      "author_url": "",
      "post_date": "2023-10-23T22:29:15.787000",
      "content": "<p>I have the same question. In the description is written: \"The Other class is not present in the training set\" and then, in the next sentence it says \"Only available for the train set.\"<br>\nSo what does it mean? \"Other\" is not present in the training set but will be present in the test set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2499667,
          "author_name": "Aurora_blue",
          "author_url": "",
          "post_date": "2023-10-26T07:06:07.937000",
          "content": "<blockquote>\n  <p>\"Only available for the train set.\"</p>\n</blockquote>\n<p>It means that \"Label\" is only available for train set. This sentence is not the explanation of \"Other class\" but \"Label\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2477532": "My guess is that outliers/other refers to mixed, uncommon, unusual, or rare tumors; all of which can have varying appearances. \nAlmost all types of tumors have increased DNA content, which makes the cell nucleus bigger and takes up more of the hematoxylin (blue-purple) stain; essentially resulting in areas with many pixels that are dark blue to dark purple. \nI see the handling of the task as pathologist and AI model in a similar way. Since it's already known that there is a tumor, viewing the images and focusing on the generally blue-purple tumor areas (which may be mixed with other colored areas) will push for deciding which class it fits best. If there are multiple classes seen, or doesn't fit any class well, then it would be time to seek help from other colleagues, or in the AI's case send an alert for a case that needs more attention.",
    "2511998": "I'm thinking of two approaches to detect the 'Other' class:\n\n1. Train an autoencoder, and use it to find outliers. On the test set, predict 'Other' for outliers.\n2. Formulate the problem as multi-label classification and use a probability threshold to predict the correct class. By using the multi-label approach, we avoid using the softmax activation function (which 'forces' the highest probability to be too high, even if it's an incorrect class prediction). Softmax is replaced by sigmoid. If for a particular example, ALL the sigmoid probabilities are lower than the probability threshold, then we can infer that it's \"none of the classes in the training set\", and hence likely to be the 'Other' class.",
    "2472981": "The [Data](https://www.kaggle.com/competitions/UBC-OCEAN/data) section of the challenge description indicates the label will be:\n> one of these subtypes of ovarian cancer: CC, EC, HGSC, LGSC, MC, Other. The Other class is not present in the training set; identifying outliers is one of the challenges of this competition. Only available for the train set.\n\nWill the Other class be included in the test set? If this class is present in the test set but not the training set, under what circumstances should our systems assign an 'Other' label to samples in the test set?\n\nThank you.",
    "2482602": "I think it's pretty much self-explanatory. Anything other than CC, EC, HGSC, LGSC, MC is **Other**.",
    "2527101": "I was thinking that maybe the test set has normal images, and normal images also belong to the others categories?",
    "2496283": "I have the same question. In the description is written: \"The Other class is not present in the training set\" and then, in the next sentence it says \"Only available for the train set.\"\nSo what does it mean? \"Other\" is not present in the training set but will be present in the test set?"
  }
}