{
  "id": 149302,
  "title": "Distribution of test set",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/149302",
  "author_name": "Ian Pan",
  "post_date": "2020-05-07T18:28:23.891000",
  "votes": 33,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I don't expect to get a concrete answer to this question, so this is just to speculate and to caution people against going purely by LB over CV. </p>\n\n<p>Based on some evidence, it seems like public test set is skewed towards Karolinska data:\n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/148894#836585\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/148894#836585</a></p>\n\n<p>Basically, CV for Karolinska data better correlates with LB than overall CV. @iafoss CV/LB results are more compelling than mine. </p>\n\n<p>Given that the training set is almost equally split between data providers (51.4% Karolinska/48.6% Radboud), there are several possibilities:</p>\n\n<ol>\n<li>Distribution of public and private test sets together are similar to training, and public test set is skewed towards Karolinska whereas private test set is skewed towards Radboud. </li>\n<li>Distribution of public and private test sets are similar to each other but different from training, and Karolinska data predominates the overall test set. </li>\n<li>The distribution of data providers is the same, but the distribution of labels more closely matches the label distribution of Karolinska vs. Radboud. </li>\n<li>The label and data provider distributions of public/private test sets and training set are similar, but because the test set is multi-annotated, this results in CV-LB discrepancy.</li>\n</ol>\n\n<p>There are probably more possibilities, but these are the ones that came to mind. </p>\n\n<p>Does anyone else have CV (by data provider)/LB results that either support/refute this hypothesis?</p>",
  "messages": [
    {
      "id": 837397,
      "postDate": "2020-05-07T18:28:23.890Z",
      "content": "<p>I don't expect to get a concrete answer to this question, so this is just to speculate and to caution people against going purely by LB over CV. </p>\n\n<p>Based on some evidence, it seems like public test set is skewed towards Karolinska data:\n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/148894#836585\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/148894#836585</a></p>\n\n<p>Basically, CV for Karolinska data better correlates with LB than overall CV. @iafoss CV/LB results are more compelling than mine. </p>\n\n<p>Given that the training set is almost equally split between data providers (51.4% Karolinska/48.6% Radboud), there are several possibilities:</p>\n\n<ol>\n<li>Distribution of public and private test sets together are similar to training, and public test set is skewed towards Karolinska whereas private test set is skewed towards Radboud. </li>\n<li>Distribution of public and private test sets are similar to each other but different from training, and Karolinska data predominates the overall test set. </li>\n<li>The distribution of data providers is the same, but the distribution of labels more closely matches the label distribution of Karolinska vs. Radboud. </li>\n<li>The label and data provider distributions of public/private test sets and training set are similar, but because the test set is multi-annotated, this results in CV-LB discrepancy.</li>\n</ol>\n\n<p>There are probably more possibilities, but these are the ones that came to mind. </p>\n\n<p>Does anyone else have CV (by data provider)/LB results that either support/refute this hypothesis?</p>",
      "rawMarkdown": "I don't expect to get a concrete answer to this question, so this is just to speculate and to caution people against going purely by LB over CV. \n\nBased on some evidence, it seems like public test set is skewed towards Karolinska data:\nhttps://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/148894#836585\n\nBasically, CV for Karolinska data better correlates with LB than overall CV. @iafoss CV/LB results are more compelling than mine. \n\nGiven that the training set is almost equally split between data providers (51.4% Karolinska/48.6% Radboud), there are several possibilities:\n\n1. Distribution of public and private test sets together are similar to training, and public test set is skewed towards Karolinska whereas private test set is skewed towards Radboud. \n2. Distribution of public and private test sets are similar to each other but different from training, and Karolinska data predominates the overall test set. \n3. The distribution of data providers is the same, but the distribution of labels more closely matches the label distribution of Karolinska vs. Radboud. \n4. The label and data provider distributions of public/private test sets and training set are similar, but because the test set is multi-annotated, this results in CV-LB discrepancy.\n\nThere are probably more possibilities, but these are the ones that came to mind. \n\nDoes anyone else have CV (by data provider)/LB results that either support/refute this hypothesis?\n",
      "votes": 33
    },
    {
      "id": 839679,
      "postDate": "2020-05-09T14:30:23.550Z",
      "content": "<p>The results from our current best CV model <a href=\"/bdubreu\">@bdubreu</a> . Training on level 1 images, 4 folds CV, regression task, metric MSE loss.</p>\n\n<p>LB 0.87 (one fold scores 0.88)\nOverall CV: 0.8501680964330703</p>\n\n<p>```\narray([[1837,  909,  102,   18,    7,    0],\n       [ 140, 1784,  641,   43,    8,    0],\n       [  12,  267,  749,  297,   16,    0],\n       [   6,   67,  221,  635,  291,    6],\n       [   4,   66,   98,  393,  626,   58],\n       [   3,   38,   38,  164,  675,  297]])</p>\n\n<p>```\nkarolinska: \nCV: 0.8636174720017733</p>\n\n<p>```\narray([[1498,  405,   19,    3,    0,    0],\n       [ 119, 1410,  276,    8,    1,    0],\n       [   8,  198,  371,   88,    3,    0],\n       [   0,   18,   99,  155,   45,    0],\n       [   3,   30,   52,  174,  209,   13],\n       [   3,    8,    8,   39,  123,   70]])</p>\n\n<p>```</p>\n\n<p>radboud\nCV: 0.7951884923518456</p>\n\n<p>```\narray([[339, 504,  83,  15,   7,   0],\n       [ 21, 374, 365,  35,   7,   0],\n       [  4,  69, 378, 209,  13,   0],\n       [  6,  49, 122, 480, 246,   6],\n       [  1,  36,  46, 219, 417,  45],\n       [  0,  30,  30, 125, 552, 227]])</p>\n\n<p>```</p>",
      "rawMarkdown": "The results from our current best CV model @bdubreu . Training on level 1 images, 4 folds CV, regression task, metric MSE loss.\n\nLB 0.87 (one fold scores 0.88)\nOverall CV: 0.8501680964330703\n\n```\narray([[1837,  909,  102,   18,    7,    0],\n       [ 140, 1784,  641,   43,    8,    0],\n       [  12,  267,  749,  297,   16,    0],\n       [   6,   67,  221,  635,  291,    6],\n       [   4,   66,   98,  393,  626,   58],\n       [   3,   38,   38,  164,  675,  297]])\n\n```\nkarolinska: \nCV: 0.8636174720017733\n\n```\narray([[1498,  405,   19,    3,    0,    0],\n       [ 119, 1410,  276,    8,    1,    0],\n       [   8,  198,  371,   88,    3,    0],\n       [   0,   18,   99,  155,   45,    0],\n       [   3,   30,   52,  174,  209,   13],\n       [   3,    8,    8,   39,  123,   70]])\n\n```\n\nradboud\nCV: 0.7951884923518456\n\n```\narray([[339, 504,  83,  15,   7,   0],\n       [ 21, 374, 365,  35,   7,   0],\n       [  4,  69, 378, 209,  13,   0],\n       [  6,  49, 122, 480, 246,   6],\n       [  1,  36,  46, 219, 417,  45],\n       [  0,  30,  30, 125, 552, 227]])\n\n```",
      "votes": 9,
      "replies": [
        {
          "id": 840434,
          "postDate": "2020-05-10T03:01:16.187Z",
          "content": "<p>It's quite interesting that tile based approach gives consistently lower score to radboud. While the approach used by <a href=\"/vaillant\">@vaillant</a> gets an opposite: better score for radboud.</p>",
          "rawMarkdown": "It's quite interesting that tile based approach gives consistently lower score to radboud. While the approach used by @vaillant gets an opposite: better score for radboud."
        },
        {
          "id": 840437,
          "postDate": "2020-05-10T03:08:51.680Z",
          "content": "<p>I am actually also using tiles, but a little differently.</p>",
          "rawMarkdown": "I am actually also using tiles, but a little differently."
        },
        {
          "id": 840766,
          "postDate": "2020-05-10T11:09:36.213Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a>, I think we can expect this, given that Radboud has more hard examples (of isup grade 3-5) in the training set.\nIntuitively, those examples are difficult to learn, as they have the composition of the same patterns (3,4,5) in different amounts. And while using tiles, but not all that have tissue pixels, we are probably doing a little mess with amounts of patterns presented to the network.\nI suppose, those intermediate grades are hard to have a consensus about for human experts as well.</p>\n\n<p>What are your thoughts about this, guys?</p>",
          "rawMarkdown": "@iafoss, I think we can expect this, given that Radboud has more hard examples (of isup grade 3-5) in the training set.\nIntuitively, those examples are difficult to learn, as they have the composition of the same patterns (3,4,5) in different amounts. And while using tiles, but not all that have tissue pixels, we are probably doing a little mess with amounts of patterns presented to the network.\nI suppose, those intermediate grades are hard to have a consensus about for human experts as well.\n\nWhat are your thoughts about this, guys?",
          "votes": 2
        },
        {
          "id": 841111,
          "postDate": "2020-05-10T15:51:14.107Z",
          "content": "<p>A lot to think about indeed. For example, do we train a model by institution? But if we do so we lose the transfer learning from one institution to the other. </p>",
          "rawMarkdown": "A lot to think about indeed. For example, do we train a model by institution? But if we do so we lose the transfer learning from one institution to the other. \n"
        },
        {
          "id": 841130,
          "postDate": "2020-05-10T16:13:21.860Z",
          "content": "<p>I briefly tried adding data provider as a feature in the last layer, and performance was a bit worse. It's also possible to use separate heads per provider, I didn't try this yet.</p>",
          "rawMarkdown": "I briefly tried adding data provider as a feature in the last layer, and performance was a bit worse. It's also possible to use separate heads per provider, I didn't try this yet.",
          "votes": 2
        },
        {
          "id": 841345,
          "postDate": "2020-05-10T18:38:41.633Z",
          "content": "<p><a href=\"/ademyanchuk\">@ademyanchuk</a> Yes, it may be an issue. The thing bothers me is that karolinska train data was labeled by one expert while radboud one by students. The test data is labeled by 3 experts, and the level of label noise in radboud train set is estimated to be 0.85.  Another bad thing is that karolinska training data has images with cancerous regions annotated with a pen, while no pen annotation is used in the test data, if I understood correctly.</p>",
          "rawMarkdown": "@ademyanchuk Yes, it may be an issue. The thing bothers me is that karolinska train data was labeled by one expert while radboud one by students. The test data is labeled by 3 experts, and the level of label noise in radboud train set is estimated to be 0.85.  Another bad thing is that karolinska training data has images with cancerous regions annotated with a pen, while no pen annotation is used in the test data, if I understood correctly."
        },
        {
          "id": 841841,
          "postDate": "2020-05-11T03:57:38.670Z",
          "content": "<p><a href=\"/msmelguizo\">@msmelguizo</a> \n1 Are you training separate model for each provider \n2. What could be contributing more to score,  way of tile extraction ,number of tile ,size of tile </p>",
          "rawMarkdown": "@msmelguizo \n1 Are you training separate model for each provider \n2. What could be contributing more to score,  way of tile extraction ,number of tile ,size of tile \n\n"
        },
        {
          "id": 842475,
          "postDate": "2020-05-11T12:18:55.743Z",
          "content": "<ol>\n<li>No, we are training one model. We are thinking about separate models among other ideas.</li>\n<li>The LB score is way better on level 1 images than level 2 while CV is about the same as others have noted.</li>\n</ol>",
          "rawMarkdown": "1. No, we are training one model. We are thinking about separate models among other ideas.\n2. The LB score is way better on level 1 images than level 2 while CV is about the same as others have noted."
        }
      ]
    },
    {
      "id": 837503,
      "postDate": "2020-05-07T20:32:32.940Z",
      "content": "<p>Great questions - but I think we can quite easily find out provider distribution in test with probing, as the provider is present in test.csv?</p>\n\n<p>EDIT: in more details, we can find out overall distribution of providers by binary tests (kernel fail/not fail), and distribution of providers in public by changing predictions of one provider - and from that deduce distribution of providers in private.</p>",
      "rawMarkdown": "Great questions - but I think we can quite easily find out provider distribution in test with probing, as the provider is present in test.csv?\n\nEDIT: in more details, we can find out overall distribution of providers by binary tests (kernel fail/not fail), and distribution of providers in public by changing predictions of one provider - and from that deduce distribution of providers in private.",
      "votes": 5,
      "replies": [
        {
          "id": 842238,
          "postDate": "2020-05-11T09:15:12.120Z",
          "content": "<p>I did some probing, just one submission so far: it was set to output a sample submission in case of radboud frequency in test being &lt; 0.4, and set to replace prediction for radboud with 5 otherwise (base LB score is 0.85). The score is 0.16 - so overall radboud frequency is not low (I was checking for this as LB score seemed to correlate more with karolinska). But if I do the same replacement in CV, I get 0.47, so there is a huge discrepancy here: does it mean that radboud is more frequent in public LB, or more frequent overall, or that class distribution is different? Still have 2 more submissions to probe today.</p>",
          "rawMarkdown": "I did some probing, just one submission so far: it was set to output a sample submission in case of radboud frequency in test being &lt; 0.4, and set to replace prediction for radboud with 5 otherwise (base LB score is 0.85). The score is 0.16 - so overall radboud frequency is not low (I was checking for this as LB score seemed to correlate more with karolinska). But if I do the same replacement in CV, I get 0.47, so there is a huge discrepancy here: does it mean that radboud is more frequent in public LB, or more frequent overall, or that class distribution is different? Still have 2 more submissions to probe today.",
          "votes": 4
        },
        {
          "id": 842268,
          "postDate": "2020-05-11T09:41:42.657Z",
          "content": "<p>And one more probe: now I know that radboud ratio is between 0.4 and 0.6 (so more or less similar to train), and the LB score is 0.34 when I replace karolinska predictions with 5. When I do the same in CV, my score goes to 0.10.</p>\n\n<p>So I didn't find evidence of any large discrepancy of provider ratio in test. I'm not yet sure how to robustly interpret the change in scores when altering per-provider predictions.</p>",
          "rawMarkdown": "And one more probe: now I know that radboud ratio is between 0.4 and 0.6 (so more or less similar to train), and the LB score is 0.34 when I replace karolinska predictions with 5. When I do the same in CV, my score goes to 0.10.\n\nSo I didn't find evidence of any large discrepancy of provider ratio in test. I'm not yet sure how to robustly interpret the change in scores when altering per-provider predictions.",
          "votes": 3
        },
        {
          "id": 842286,
          "postDate": "2020-05-11T09:57:03.153Z",
          "content": "<p>Actually this information was already present in the report at <a href=\"https://zenodo.org/record/3715938#.XrVzk6gzZPa\">https://zenodo.org/record/3715938#.XrVzk6gzZPa</a> which was cited in another thread:</p>\n\n<blockquote>\n  <p>Karolinska: The test data were additionally graded by 3 experienced pathologists and a consensus grade has been derived (at least 2 of 3 agree in the grade). For these annotations, the approximately 200 cancer cases from STHLM3 were split in 100 + 100 slides. Two pathologists (plus the pathologists from the training/public set) graded one set each. Whenever there was a disagreement in the grade, another pathologist acted as a tiebreaker. If none of the 3 pathologists agreed the case was removed. Finally approximately 50 benign cases were added.</p>\n</blockquote>\n\n<p>So we should expect ~450 minus some removed cases in test set for Karolinska, which would make approximately a half of overall test.</p>",
          "rawMarkdown": "Actually this information was already present in the report at https://zenodo.org/record/3715938#.XrVzk6gzZPa which was cited in another thread:\n\n&gt; Karolinska: The test data were additionally graded by 3 experienced pathologists and a consensus grade has been derived (at least 2 of 3 agree in the grade). For these annotations, the approximately 200 cancer cases from STHLM3 were split in 100 + 100 slides. Two pathologists (plus the pathologists from the training/public set) graded one set each. Whenever there was a disagreement in the grade, another pathologist acted as a tiebreaker. If none of the 3 pathologists agreed the case was removed. Finally approximately 50 benign cases were added.\n\nSo we should expect ~450 minus some removed cases in test set for Karolinska, which would make approximately a half of overall test.",
          "votes": 2
        },
        {
          "id": 842389,
          "postDate": "2020-05-11T11:17:45.137Z",
          "content": "<p>Thanks for info! I think it's 200 (splited into 100 + 100 slides) + 50 benign = ~250. Am I missing something, correct please?</p>",
          "rawMarkdown": "Thanks for info! I think it's 200 (splited into 100 + 100 slides) + 50 benign = ~250. Am I missing something, correct please?",
          "votes": 2
        },
        {
          "id": 842406,
          "postDate": "2020-05-11T11:31:50.050Z",
          "content": "<p><a href=\"/ademyanchuk\">@ademyanchuk</a> oh right, good catch. Maybe there are more cases not accounted for? According to probing there should be at least 320 Karolinska samples and likely more.</p>",
          "rawMarkdown": "@ademyanchuk oh right, good catch. Maybe there are more cases not accounted for? According to probing there should be at least 320 Karolinska samples and likely more.",
          "votes": 1
        },
        {
          "id": 842409,
          "postDate": "2020-05-11T11:34:41.927Z",
          "content": "<p>I would interpret the findings as follows: trainset has more severe (3-5 isup) cases from radboud, that is why changing radboud cases in CV to 5s does not drop the score so dramatically.  Public has more low-moderate cases (0-2) from radboud, so changing them to 5s should drop score more than locally. The same logic applies to karolinska. IMO: public has most of low-moderate cases from radboud and most of severe cases from karolinska. Unfortunately, we can not imply that private would be the same or opposite :\\\nP.S.: another variable to have in mind is the label noise</p>",
          "rawMarkdown": "I would interpret the findings as follows: trainset has more severe (3-5 isup) cases from radboud, that is why changing radboud cases in CV to 5s does not drop the score so dramatically.  Public has more low-moderate cases (0-2) from radboud, so changing them to 5s should drop score more than locally. The same logic applies to karolinska. IMO: public has most of low-moderate cases from radboud and most of severe cases from karolinska. Unfortunately, we can not imply that private would be the same or opposite :\\\nP.S.: another variable to have in mind is the label noise",
          "votes": 1
        },
        {
          "id": 842653,
          "postDate": "2020-05-11T14:29:52.327Z",
          "content": "<p>Based on your results I'm inclined to think that the class distribution is different like <a href=\"/ademyanchuk\">@ademyanchuk</a> said. Probably the public test distribution for Radboud is closer to the train distribution of Karolinska. </p>",
          "rawMarkdown": "Based on your results I'm inclined to think that the class distribution is different like @ademyanchuk said. Probably the public test distribution for Radboud is closer to the train distribution of Karolinska. "
        }
      ]
    },
    {
      "id": 853755,
      "postDate": "2020-05-19T13:08:06.670Z",
      "content": "<p>i \"estimate\" that there are 674 images from Radboud in the test LB set</p>",
      "rawMarkdown": "i \"estimate\" that there are 674 images from Radboud in the test LB set",
      "replies": [
        {
          "id": 853884,
          "postDate": "2020-05-19T15:10:45.527Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> How did you estimate this? According to <a href=\"https://zenodo.org/record/3715938#.XpTU3PJKiUl\">https://zenodo.org/record/3715938#.XpTU3PJKiUl</a> (bottom of page 9), public and private LB each only have about 400 cases.</p>",
          "rawMarkdown": "@hengck23 How did you estimate this? According to https://zenodo.org/record/3715938#.XpTU3PJKiUl (bottom of page 9), public and private LB each only have about 400 cases.",
          "votes": 1
        }
      ]
    },
    {
      "id": 839871,
      "postDate": "2020-05-09T16:28:10.930Z",
      "content": "<p><a href=\"/msmelguizo\">@msmelguizo</a>  are you using tiles ?</p>",
      "rawMarkdown": "@msmelguizo  are you using tiles ?",
      "replies": [
        {
          "id": 839952,
          "postDate": "2020-05-09T17:12:07.010Z",
          "content": "<p>Yes 32 patches of level 1</p>",
          "rawMarkdown": "Yes 32 patches of level 1",
          "votes": 2
        },
        {
          "id": 841246,
          "postDate": "2020-05-10T17:25:50.737Z",
          "content": "<p><a href=\"/msmelguizo\">@msmelguizo</a> \nthanks maria\n1) ,how about size of patches\nWhat i find that if we use less size path we need more tiles to understand the patterns\nand may be less if patch size is more\n2) How loss L1  loss or MSE?</p>",
          "rawMarkdown": "@msmelguizo \nthanks maria\n1) ,how about size of patches\nWhat i find that if we use less size path we need more tiles to understand the patterns\nand may be less if patch size is more\n2) How loss L1  loss or MSE?"
        },
        {
          "id": 842476,
          "postDate": "2020-05-11T12:20:03.070Z",
          "content": "<p>1) Yes its a lot more data to process with level 1.\n2) You need to modify the network to have a continuous output. </p>",
          "rawMarkdown": "1) Yes its a lot more data to process with level 1.\n2) You need to modify the network to have a continuous output. ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 839679,
      "author_name": "Maria Wellen",
      "author_url": "",
      "post_date": "2020-05-09T14:30:23.550000",
      "content": "<p>The results from our current best CV model <a href=\"/bdubreu\">@bdubreu</a> . Training on level 1 images, 4 folds CV, regression task, metric MSE loss.</p>\n\n<p>LB 0.87 (one fold scores 0.88)\nOverall CV: 0.8501680964330703</p>\n\n<p>```\narray([[1837,  909,  102,   18,    7,    0],\n       [ 140, 1784,  641,   43,    8,    0],\n       [  12,  267,  749,  297,   16,    0],\n       [   6,   67,  221,  635,  291,    6],\n       [   4,   66,   98,  393,  626,   58],\n       [   3,   38,   38,  164,  675,  297]])</p>\n\n<p>```\nkarolinska: \nCV: 0.8636174720017733</p>\n\n<p>```\narray([[1498,  405,   19,    3,    0,    0],\n       [ 119, 1410,  276,    8,    1,    0],\n       [   8,  198,  371,   88,    3,    0],\n       [   0,   18,   99,  155,   45,    0],\n       [   3,   30,   52,  174,  209,   13],\n       [   3,    8,    8,   39,  123,   70]])</p>\n\n<p>```</p>\n\n<p>radboud\nCV: 0.7951884923518456</p>\n\n<p>```\narray([[339, 504,  83,  15,   7,   0],\n       [ 21, 374, 365,  35,   7,   0],\n       [  4,  69, 378, 209,  13,   0],\n       [  6,  49, 122, 480, 246,   6],\n       [  1,  36,  46, 219, 417,  45],\n       [  0,  30,  30, 125, 552, 227]])</p>\n\n<p>```</p>",
      "votes": 9,
      "replies": [
        {
          "id": 840434,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-10T03:01:16.187000",
          "content": "<p>It's quite interesting that tile based approach gives consistently lower score to radboud. While the approach used by <a href=\"/vaillant\">@vaillant</a> gets an opposite: better score for radboud.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840437,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-05-10T03:08:51.680000",
          "content": "<p>I am actually also using tiles, but a little differently.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840766,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-10T11:09:36.213000",
          "content": "<p><a href=\"/iafoss\">@iafoss</a>, I think we can expect this, given that Radboud has more hard examples (of isup grade 3-5) in the training set.\nIntuitively, those examples are difficult to learn, as they have the composition of the same patterns (3,4,5) in different amounts. And while using tiles, but not all that have tissue pixels, we are probably doing a little mess with amounts of patterns presented to the network.\nI suppose, those intermediate grades are hard to have a consensus about for human experts as well.</p>\n\n<p>What are your thoughts about this, guys?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 841111,
          "author_name": "Maria Wellen",
          "author_url": "",
          "post_date": "2020-05-10T15:51:14.107000",
          "content": "<p>A lot to think about indeed. For example, do we train a model by institution? But if we do so we lose the transfer learning from one institution to the other. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 841130,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-10T16:13:21.860000",
          "content": "<p>I briefly tried adding data provider as a feature in the last layer, and performance was a bit worse. It's also possible to use separate heads per provider, I didn't try this yet.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 841345,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-10T18:38:41.633000",
          "content": "<p><a href=\"/ademyanchuk\">@ademyanchuk</a> Yes, it may be an issue. The thing bothers me is that karolinska train data was labeled by one expert while radboud one by students. The test data is labeled by 3 experts, and the level of label noise in radboud train set is estimated to be 0.85.  Another bad thing is that karolinska training data has images with cancerous regions annotated with a pen, while no pen annotation is used in the test data, if I understood correctly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 841841,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-05-11T03:57:38.670000",
          "content": "<p><a href=\"/msmelguizo\">@msmelguizo</a> \n1 Are you training separate model for each provider \n2. What could be contributing more to score,  way of tile extraction ,number of tile ,size of tile </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 842475,
          "author_name": "Maria Wellen",
          "author_url": "",
          "post_date": "2020-05-11T12:18:55.743000",
          "content": "<ol>\n<li>No, we are training one model. We are thinking about separate models among other ideas.</li>\n<li>The LB score is way better on level 1 images than level 2 while CV is about the same as others have noted.</li>\n</ol>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 837503,
      "author_name": "Konstantin Lopukhin",
      "author_url": "",
      "post_date": "2020-05-07T20:32:32.940000",
      "content": "<p>Great questions - but I think we can quite easily find out provider distribution in test with probing, as the provider is present in test.csv?</p>\n\n<p>EDIT: in more details, we can find out overall distribution of providers by binary tests (kernel fail/not fail), and distribution of providers in public by changing predictions of one provider - and from that deduce distribution of providers in private.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 842238,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-11T09:15:12.120000",
          "content": "<p>I did some probing, just one submission so far: it was set to output a sample submission in case of radboud frequency in test being &lt; 0.4, and set to replace prediction for radboud with 5 otherwise (base LB score is 0.85). The score is 0.16 - so overall radboud frequency is not low (I was checking for this as LB score seemed to correlate more with karolinska). But if I do the same replacement in CV, I get 0.47, so there is a huge discrepancy here: does it mean that radboud is more frequent in public LB, or more frequent overall, or that class distribution is different? Still have 2 more submissions to probe today.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 842268,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-11T09:41:42.657000",
          "content": "<p>And one more probe: now I know that radboud ratio is between 0.4 and 0.6 (so more or less similar to train), and the LB score is 0.34 when I replace karolinska predictions with 5. When I do the same in CV, my score goes to 0.10.</p>\n\n<p>So I didn't find evidence of any large discrepancy of provider ratio in test. I'm not yet sure how to robustly interpret the change in scores when altering per-provider predictions.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 842286,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-11T09:57:03.153000",
          "content": "<p>Actually this information was already present in the report at <a href=\"https://zenodo.org/record/3715938#.XrVzk6gzZPa\">https://zenodo.org/record/3715938#.XrVzk6gzZPa</a> which was cited in another thread:</p>\n\n<blockquote>\n  <p>Karolinska: The test data were additionally graded by 3 experienced pathologists and a consensus grade has been derived (at least 2 of 3 agree in the grade). For these annotations, the approximately 200 cancer cases from STHLM3 were split in 100 + 100 slides. Two pathologists (plus the pathologists from the training/public set) graded one set each. Whenever there was a disagreement in the grade, another pathologist acted as a tiebreaker. If none of the 3 pathologists agreed the case was removed. Finally approximately 50 benign cases were added.</p>\n</blockquote>\n\n<p>So we should expect ~450 minus some removed cases in test set for Karolinska, which would make approximately a half of overall test.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 842389,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-11T11:17:45.137000",
          "content": "<p>Thanks for info! I think it's 200 (splited into 100 + 100 slides) + 50 benign = ~250. Am I missing something, correct please?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 842406,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-11T11:31:50.050000",
          "content": "<p><a href=\"/ademyanchuk\">@ademyanchuk</a> oh right, good catch. Maybe there are more cases not accounted for? According to probing there should be at least 320 Karolinska samples and likely more.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 842409,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-11T11:34:41.927000",
          "content": "<p>I would interpret the findings as follows: trainset has more severe (3-5 isup) cases from radboud, that is why changing radboud cases in CV to 5s does not drop the score so dramatically.  Public has more low-moderate cases (0-2) from radboud, so changing them to 5s should drop score more than locally. The same logic applies to karolinska. IMO: public has most of low-moderate cases from radboud and most of severe cases from karolinska. Unfortunately, we can not imply that private would be the same or opposite :\\\nP.S.: another variable to have in mind is the label noise</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 842653,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-05-11T14:29:52.327000",
          "content": "<p>Based on your results I'm inclined to think that the class distribution is different like <a href=\"/ademyanchuk\">@ademyanchuk</a> said. Probably the public test distribution for Radboud is closer to the train distribution of Karolinska. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 853755,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-05-19T13:08:06.670000",
      "content": "<p>i \"estimate\" that there are 674 images from Radboud in the test LB set</p>",
      "votes": 0,
      "replies": [
        {
          "id": 853884,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-05-19T15:10:45.527000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> How did you estimate this? According to <a href=\"https://zenodo.org/record/3715938#.XpTU3PJKiUl\">https://zenodo.org/record/3715938#.XpTU3PJKiUl</a> (bottom of page 9), public and private LB each only have about 400 cases.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 839871,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-05-09T16:28:10.930000",
      "content": "<p><a href=\"/msmelguizo\">@msmelguizo</a>  are you using tiles ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 839952,
          "author_name": "Maria Wellen",
          "author_url": "",
          "post_date": "2020-05-09T17:12:07.010000",
          "content": "<p>Yes 32 patches of level 1</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 841246,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-05-10T17:25:50.737000",
          "content": "<p><a href=\"/msmelguizo\">@msmelguizo</a> \nthanks maria\n1) ,how about size of patches\nWhat i find that if we use less size path we need more tiles to understand the patterns\nand may be less if patch size is more\n2) How loss L1  loss or MSE?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 842476,
          "author_name": "Maria Wellen",
          "author_url": "",
          "post_date": "2020-05-11T12:20:03.070000",
          "content": "<p>1) Yes its a lot more data to process with level 1.\n2) You need to modify the network to have a continuous output. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "837397": "I don't expect to get a concrete answer to this question, so this is just to speculate and to caution people against going purely by LB over CV. \n\nBased on some evidence, it seems like public test set is skewed towards Karolinska data:\nhttps://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/148894#836585\n\nBasically, CV for Karolinska data better correlates with LB than overall CV. @iafoss CV/LB results are more compelling than mine. \n\nGiven that the training set is almost equally split between data providers (51.4% Karolinska/48.6% Radboud), there are several possibilities:\n\n1. Distribution of public and private test sets together are similar to training, and public test set is skewed towards Karolinska whereas private test set is skewed towards Radboud. \n2. Distribution of public and private test sets are similar to each other but different from training, and Karolinska data predominates the overall test set. \n3. The distribution of data providers is the same, but the distribution of labels more closely matches the label distribution of Karolinska vs. Radboud. \n4. The label and data provider distributions of public/private test sets and training set are similar, but because the test set is multi-annotated, this results in CV-LB discrepancy.\n\nThere are probably more possibilities, but these are the ones that came to mind. \n\nDoes anyone else have CV (by data provider)/LB results that either support/refute this hypothesis?\n",
    "839679": "The results from our current best CV model @bdubreu . Training on level 1 images, 4 folds CV, regression task, metric MSE loss.\n\nLB 0.87 (one fold scores 0.88)\nOverall CV: 0.8501680964330703\n\n```\narray([[1837,  909,  102,   18,    7,    0],\n       [ 140, 1784,  641,   43,    8,    0],\n       [  12,  267,  749,  297,   16,    0],\n       [   6,   67,  221,  635,  291,    6],\n       [   4,   66,   98,  393,  626,   58],\n       [   3,   38,   38,  164,  675,  297]])\n\n```\nkarolinska: \nCV: 0.8636174720017733\n\n```\narray([[1498,  405,   19,    3,    0,    0],\n       [ 119, 1410,  276,    8,    1,    0],\n       [   8,  198,  371,   88,    3,    0],\n       [   0,   18,   99,  155,   45,    0],\n       [   3,   30,   52,  174,  209,   13],\n       [   3,    8,    8,   39,  123,   70]])\n\n```\n\nradboud\nCV: 0.7951884923518456\n\n```\narray([[339, 504,  83,  15,   7,   0],\n       [ 21, 374, 365,  35,   7,   0],\n       [  4,  69, 378, 209,  13,   0],\n       [  6,  49, 122, 480, 246,   6],\n       [  1,  36,  46, 219, 417,  45],\n       [  0,  30,  30, 125, 552, 227]])\n\n```",
    "837503": "Great questions - but I think we can quite easily find out provider distribution in test with probing, as the provider is present in test.csv?\n\nEDIT: in more details, we can find out overall distribution of providers by binary tests (kernel fail/not fail), and distribution of providers in public by changing predictions of one provider - and from that deduce distribution of providers in private.",
    "853755": "i \"estimate\" that there are 674 images from Radboud in the test LB set",
    "839871": "@msmelguizo  are you using tiles ?"
  }
}