{
  "id": 157997,
  "title": "Which Loss Function is best for Kappa?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/157997",
  "author_name": "Salman",
  "post_date": "2020-06-12T21:51:25.374000",
  "votes": 1,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Kappa is kind of sensitive in this case. Could you please share best loss function for your approach so far? For Classification and Regression?\nFor me Categorical Cross entropy is working best.  </p>",
  "messages": [
    {
      "id": 883807,
      "postDate": "2020-06-13T00:39:31.893Z",
      "content": "<p>Some people reported that both <code>CE</code> and <code>MSE</code> loss works the best. I observed the same.. </p>",
      "rawMarkdown": "Some people reported that both `CE` and `MSE` loss works the best. I observed the same.. ",
      "votes": 1,
      "replies": [
        {
          "id": 884443,
          "postDate": "2020-06-13T11:36:13.030Z",
          "content": "<p>Yeah. I observed the same. But for me MAE works best with my current approach.</p>",
          "rawMarkdown": "Yeah. I observed the same. But for me MAE works best with my current approach.",
          "votes": 1
        },
        {
          "id": 887635,
          "postDate": "2020-06-15T19:25:40.097Z",
          "content": "<p>You round predictions from your network to assign label is MSE or using some kind of thresholds or something?</p>",
          "rawMarkdown": "You round predictions from your network to assign label is MSE or using some kind of thresholds or something?",
          "votes": 1
        },
        {
          "id": 887678,
          "postDate": "2020-06-15T20:00:28.060Z",
          "content": "<p>yes, just using standart  threshold <code>[0.5, 1.5, 2.5, 3.5, 4.5]</code></p>",
          "rawMarkdown": "yes, just using standart  threshold `[0.5, 1.5, 2.5, 3.5, 4.5]`"
        },
        {
          "id": 887727,
          "postDate": "2020-06-15T20:31:37.470Z",
          "content": "<p>Can you please share your validation / CV MSE so that I can make a comparison?\nIt would be a great help. Thanks.</p>",
          "rawMarkdown": "Can you please share your validation / CV MSE so that I can make a comparison?\nIt would be a great help. Thanks."
        },
        {
          "id": 887738,
          "postDate": "2020-06-15T20:45:37.277Z",
          "content": "<p>This is just an example. This scores <code>0.89</code> on LB </p>\n\n<p><code>\nVAL LOSS:\n0.561787 <br>\nSCORE OVERALL:\n0.8920411196145195\n[[492  64   7   2   2   0]\n [ 36 385  71   7   1   0]\n [  4  89 124  25  10   0]\n [  7  13  46  83  70  11]\n [  1   8  16  35 127  40]\n [  1   3   8  16  62 130]]\nSCORE KAROLINSKA:\n0.8835960297315442\n[[358  44   5   1   2   0]\n [ 21 285  45   3   0   0]\n [  3  41  53  17   5   0]\n [  0   4  14  17  17   1]\n [  1   5   8  14  51  12]\n [  0   1   1   2  13  24]]\nSCORE RADBOUND:\n0.8677251938510091\n[[134  20   2   1   0   0]\n [ 15 100  26   4   1   0]\n [  1  48  71   8   5   0]\n [  7   9  32  66  53  10]\n [  0   3   8  21  76  28]\n [  1   2   7  14  49 106]]\n</code></p>\n\n<p>Good luck=)</p>",
          "rawMarkdown": "This is just an example. This scores `0.89` on LB \n\n```\nVAL LOSS:\n0.561787\t\nSCORE OVERALL:\n0.8920411196145195\n[[492  64   7   2   2   0]\n [ 36 385  71   7   1   0]\n [  4  89 124  25  10   0]\n [  7  13  46  83  70  11]\n [  1   8  16  35 127  40]\n [  1   3   8  16  62 130]]\nSCORE KAROLINSKA:\n0.8835960297315442\n[[358  44   5   1   2   0]\n [ 21 285  45   3   0   0]\n [  3  41  53  17   5   0]\n [  0   4  14  17  17   1]\n [  1   5   8  14  51  12]\n [  0   1   1   2  13  24]]\nSCORE RADBOUND:\n0.8677251938510091\n[[134  20   2   1   0   0]\n [ 15 100  26   4   1   0]\n [  1  48  71   8   5   0]\n [  7   9  32  66  53  10]\n [  0   3   8  21  76  28]\n [  1   2   7  14  49 106]]\n```\n\nGood luck=)",
          "votes": 1
        },
        {
          "id": 887749,
          "postDate": "2020-06-15T20:54:38.080Z",
          "content": "<p>Thanks :)</p>",
          "rawMarkdown": "Thanks :)"
        },
        {
          "id": 887882,
          "postDate": "2020-06-16T01:06:52.697Z",
          "content": "<p>Ok. So I have achieved MSE LOSS : 0.90 using simple intermediate tiles in 30 epochs.\nBut my code is implemented in tensorflow.\nI am not familiar with pytorch. Are there any differences in AdaptiveAveragePooling from pytorch and AveragePooling from tensorflow? \nKind of confused with this.\nMoreover, I am following approach to concatenate patches to make a bigger image. \nPassing individual patch from pretrained network and concatenating their feature map for dense layers will work better or this bigger image is enough with augmentation on each patch and over all?\nLet me know. Thanks. :)</p>",
          "rawMarkdown": "Ok. So I have achieved MSE LOSS : 0.90 using simple intermediate tiles in 30 epochs.\nBut my code is implemented in tensorflow.\nI am not familiar with pytorch. Are there any differences in AdaptiveAveragePooling from pytorch and AveragePooling from tensorflow? \nKind of confused with this.\nMoreover, I am following approach to concatenate patches to make a bigger image. \nPassing individual patch from pretrained network and concatenating their feature map for dense layers will work better or this bigger image is enough with augmentation on each patch and over all?\nLet me know. Thanks. :)"
        },
        {
          "id": 887960,
          "postDate": "2020-06-16T03:16:27.970Z",
          "content": "<p>Adaptive average pooling is just average pooling over the whole spatial dimensions regardless of the size. Its keras equivalent would be GlobalAverage I think.</p>",
          "rawMarkdown": "Adaptive average pooling is just average pooling over the whole spatial dimensions regardless of the size. Its keras equivalent would be GlobalAverage I think.",
          "votes": 1
        }
      ]
    },
    {
      "id": 883745,
      "postDate": "2020-06-12T21:51:25.373Z",
      "content": "<p>Kappa is kind of sensitive in this case. Could you please share best loss function for your approach so far? For Classification and Regression?\nFor me Categorical Cross entropy is working best.  </p>",
      "rawMarkdown": "Kappa is kind of sensitive in this case. Could you please share best loss function for your approach so far? For Classification and Regression?\nFor me Categorical Cross entropy is working best.  ",
      "votes": 1
    },
    {
      "id": 884552,
      "postDate": "2020-06-13T12:44:55.043Z",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel/weighted-kappa-loss-for-keras-tensorflow\">kappa-loss</a>\nHas anyone tried this loss function in this competition? How is the result? Can you share it?</p>",
      "rawMarkdown": "[kappa-loss](https://www.kaggle.com/christofhenkel/weighted-kappa-loss-for-keras-tensorflow)\nHas anyone tried this loss function in this competition? How is the result? Can you share it?",
      "replies": [
        {
          "id": 885173,
          "postDate": "2020-06-14T01:35:00.617Z",
          "content": "<p>I also found this article while looking around. <a href=\"https://arxiv.org/pdf/1612.00775.pdf\">https://arxiv.org/pdf/1612.00775.pdf</a></p>",
          "rawMarkdown": "I also found this article while looking around. https://arxiv.org/pdf/1612.00775.pdf"
        },
        {
          "id": 885194,
          "postDate": "2020-06-14T02:12:29.200Z",
          "content": "<p>Thanks for sharing. Let me have a look！</p>",
          "rawMarkdown": "Thanks for sharing. Let me have a look！"
        },
        {
          "id": 886149,
          "postDate": "2020-06-14T18:33:11.053Z",
          "content": "<p>So I looked into that TF code and couldn't really follow so I went back to the paper source: <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0167865517301666\">https://www.sciencedirect.com/science/article/abs/pii/S0167865517301666</a> (paywall) </p>\n\n<p>One strong weakness of that loss (mentioned in paper) is that it uses normalization terms per class. We can try it but it will probably fail because of the low batch size which makes your (soft) kappa loss really unstable. Below batches of 10 they show their loss give bad results compared to standard cross entropy in their experiment.</p>\n\n<p>I think the other paper I linked which uses just a constrained output and MSE is more promising.</p>",
          "rawMarkdown": "So I looked into that TF code and couldn't really follow so I went back to the paper source: https://www.sciencedirect.com/science/article/abs/pii/S0167865517301666 (paywall) \n\nOne strong weakness of that loss (mentioned in paper) is that it uses normalization terms per class. We can try it but it will probably fail because of the low batch size which makes your (soft) kappa loss really unstable. Below batches of 10 they show their loss give bad results compared to standard cross entropy in their experiment.\n\nI think the other paper I linked which uses just a constrained output and MSE is more promising.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 883807,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-06-13T00:39:31.893000",
      "content": "<p>Some people reported that both <code>CE</code> and <code>MSE</code> loss works the best. I observed the same.. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 884443,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2020-06-13T11:36:13.030000",
          "content": "<p>Yeah. I observed the same. But for me MAE works best with my current approach.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 887635,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2020-06-15T19:25:40.097000",
          "content": "<p>You round predictions from your network to assign label is MSE or using some kind of thresholds or something?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 887678,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-06-15T20:00:28.060000",
          "content": "<p>yes, just using standart  threshold <code>[0.5, 1.5, 2.5, 3.5, 4.5]</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 887727,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2020-06-15T20:31:37.470000",
          "content": "<p>Can you please share your validation / CV MSE so that I can make a comparison?\nIt would be a great help. Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 887738,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-06-15T20:45:37.277000",
          "content": "<p>This is just an example. This scores <code>0.89</code> on LB </p>\n\n<p><code>\nVAL LOSS:\n0.561787 <br>\nSCORE OVERALL:\n0.8920411196145195\n[[492  64   7   2   2   0]\n [ 36 385  71   7   1   0]\n [  4  89 124  25  10   0]\n [  7  13  46  83  70  11]\n [  1   8  16  35 127  40]\n [  1   3   8  16  62 130]]\nSCORE KAROLINSKA:\n0.8835960297315442\n[[358  44   5   1   2   0]\n [ 21 285  45   3   0   0]\n [  3  41  53  17   5   0]\n [  0   4  14  17  17   1]\n [  1   5   8  14  51  12]\n [  0   1   1   2  13  24]]\nSCORE RADBOUND:\n0.8677251938510091\n[[134  20   2   1   0   0]\n [ 15 100  26   4   1   0]\n [  1  48  71   8   5   0]\n [  7   9  32  66  53  10]\n [  0   3   8  21  76  28]\n [  1   2   7  14  49 106]]\n</code></p>\n\n<p>Good luck=)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 887749,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2020-06-15T20:54:38.080000",
          "content": "<p>Thanks :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 887882,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2020-06-16T01:06:52.697000",
          "content": "<p>Ok. So I have achieved MSE LOSS : 0.90 using simple intermediate tiles in 30 epochs.\nBut my code is implemented in tensorflow.\nI am not familiar with pytorch. Are there any differences in AdaptiveAveragePooling from pytorch and AveragePooling from tensorflow? \nKind of confused with this.\nMoreover, I am following approach to concatenate patches to make a bigger image. \nPassing individual patch from pretrained network and concatenating their feature map for dense layers will work better or this bigger image is enough with augmentation on each patch and over all?\nLet me know. Thanks. :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 887960,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-16T03:16:27.970000",
          "content": "<p>Adaptive average pooling is just average pooling over the whole spatial dimensions regardless of the size. Its keras equivalent would be GlobalAverage I think.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 884552,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2020-06-13T12:44:55.043000",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel/weighted-kappa-loss-for-keras-tensorflow\">kappa-loss</a>\nHas anyone tried this loss function in this competition? How is the result? Can you share it?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 885173,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-14T01:35:00.617000",
          "content": "<p>I also found this article while looking around. <a href=\"https://arxiv.org/pdf/1612.00775.pdf\">https://arxiv.org/pdf/1612.00775.pdf</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 885194,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2020-06-14T02:12:29.200000",
          "content": "<p>Thanks for sharing. Let me have a look！</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 886149,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-14T18:33:11.053000",
          "content": "<p>So I looked into that TF code and couldn't really follow so I went back to the paper source: <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0167865517301666\">https://www.sciencedirect.com/science/article/abs/pii/S0167865517301666</a> (paywall) </p>\n\n<p>One strong weakness of that loss (mentioned in paper) is that it uses normalization terms per class. We can try it but it will probably fail because of the low batch size which makes your (soft) kappa loss really unstable. Below batches of 10 they show their loss give bad results compared to standard cross entropy in their experiment.</p>\n\n<p>I think the other paper I linked which uses just a constrained output and MSE is more promising.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "883807": "Some people reported that both `CE` and `MSE` loss works the best. I observed the same.. ",
    "883745": "Kappa is kind of sensitive in this case. Could you please share best loss function for your approach so far? For Classification and Regression?\nFor me Categorical Cross entropy is working best.  ",
    "884552": "[kappa-loss](https://www.kaggle.com/christofhenkel/weighted-kappa-loss-for-keras-tensorflow)\nHas anyone tried this loss function in this competition? How is the result? Can you share it?"
  }
}