{
  "id": 346480,
  "title": "Technical Flaws in Error Evaluation Function",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/346480",
  "author_name": "Syed Muhammad Atif",
  "post_date": "2022-08-19T18:13:08.844000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Submissions are evaluated using a weighted multi-class logarithmic loss.<br>\nThe formula is then:</p>\n<p>\\[ \\text{Log Loss} = - \\left( \\frac{\\sum_{i=1}^{M} w_{i} \\cdot \\sum_{j=1}^{N_{i}} \\frac{y_{ij}}{N_{i}} \\cdot \\ln  p_{ij} }{\\sum_{i=1}^{M} w_{i}} \\right) \\]</p>\n<p>where \\(N\\) is the number of images in the class set, \\(M\\) is the number of classes, \\(ln\\) is the natural logarithm, \\( y_{i,j} \\) is 1 if observation \\(i\\) belongs to class \\(j\\) and 0 otherwise, \\(p_{i,j}\\) is the predicted probability that image \\(i\\) belongs to class \\(j\\).</p>\n<p>The above formula has following technical flaws:<br>\n1) \\( N_{i} \\)  is not defined.<br>\n2) The index of class is \\(i\\) running from 1 to \\(M\\), then how \\(y_{i,j}\\) belongs to class \\(j\\) (see. \\( y_{i,j} \\) is \\(1\\) if observation \\(i\\) belongs to class \\(j\\) and \\(0\\) otherwise).<br>\n3) Same problem is with \\( p_{i,j} \\), how \\( p_{i,j} \\) belongs to class \\(j\\).</p>",
  "messages": [
    {
      "id": 1906501,
      "postDate": "2022-08-19T23:37:30.983Z",
      "content": "<p>I think that for this loss function it's really the Log Loss for a single observation. My guess for #1 then is that \\(N_i\\) is the number of images per class of an observation. However, an observation can only belong to one class, so if we have \\(i=1\\) representing CA and \\(i=2\\) representing LAA, then if the example belongs to CA, \\(N_2=0.\\)</p>\n<p>For #2 and # 3, I believe \\(y_{i.j}\\) is an <a href=\"https://en.wikipedia.org/wiki/Indicator_function\" target=\"_blank\">indicator function</a>, which is 1 is the observation belongs to the \\(i\\)'th class and zero otherwise. So following our previous example of the single observation being CA with \\(i=1\\), then the indicator function would appear as follows, \\(y_{1,j}=1,\\) and \\(y_{2,j}=0\\).</p>\n<p>The effect then is that the \\(p_{2,j}\\) values are zeroed out from the loss function. They're not really required then for the resulting Log Loss. The calculation really then is just the negative natural log of the weighted probabilities for the true class. I think a typo they have though is that instead of </p>\n<blockquote>\n  <p>ln is the natural logarithm, \\(y_{ij}\\) is 1 if observation \\(i\\) belongs to class \\(j\\) and 0 otherwise, \\(p_{ij}\\) is the predicted probability that image \\(i\\) belongs to class \\(j\\).</p>\n</blockquote>\n<p>it should read</p>\n<blockquote>\n  <p>ln is the natural logarithm, \\(y_{ij}\\) is 1 if observation \\(j\\) belongs to class \\(i\\) and 0 otherwise, \\(p_{ij}\\) is the predicted probability that image \\(j\\) belongs to class \\(i\\).</p>\n</blockquote>\n<p>I have a statistics background (not saying that I'm sure I understand this correctly though) so I'm pretty used to seeing indices like this. I believe an issue though that the writer made is by simplifying the loss function to that of a single example (rather than using 3 indices, \\(i\\), \\(j\\), and \\(k\\)), they got a bit mixed up by what they mean by class and observation. What I wrote in the end is still a bit off in my mind. I've seen indices like this (e.g., <a href=\"https://www.microsoft.com/en-us/research/uploads/prod/2006/01/Bishop-Pattern-Recognition-and-Machine-Learning-2006.pdf\" target=\"_blank\">Bishop</a>) where it feels like the \\(i\\) and \\(j\\) are inverted. Generally, we like to think of i as belonging to the observation, and j belonging to the feature. So having \\(i\\) as class and \\(j\\) as image number for the \\(i\\)'th class runs a bit counterintuitive to how we normally think about it.</p>",
      "rawMarkdown": "I think that for this loss function it's really the Log Loss for a single observation. My guess for #1 then is that \\\\(N_i\\\\) is the number of images per class of an observation. However, an observation can only belong to one class, so if we have \\\\(i=1\\\\) representing CA and \\\\(i=2\\\\) representing LAA, then if the example belongs to CA, \\\\(N_2=0.\\\\)\n\nFor #2 and # 3, I believe \\\\(y_{i.j}\\\\) is an [indicator function](https://en.wikipedia.org/wiki/Indicator_function), which is 1 is the observation belongs to the \\\\(i\\\\)'th class and zero otherwise. So following our previous example of the single observation being CA with \\\\(i=1\\\\), then the indicator function would appear as follows, \\\\(y_{1,j}=1,\\\\) and \\\\(y_{2,j}=0\\\\).\n\nThe effect then is that the \\\\(p_{2,j}\\\\) values are zeroed out from the loss function. They're not really required then for the resulting Log Loss. The calculation really then is just the negative natural log of the weighted probabilities for the true class. I think a typo they have though is that instead of \n> ln is the natural logarithm, \\\\(y_{ij}\\\\) is 1 if observation \\\\(i\\\\) belongs to class \\\\(j\\\\) and 0 otherwise, \\\\(p_{ij}\\\\) is the predicted probability that image \\\\(i\\\\) belongs to class \\\\(j\\\\).\n\nit should read\n\n> ln is the natural logarithm, \\\\(y_{ij}\\\\) is 1 if observation \\\\(j\\\\) belongs to class \\\\(i\\\\) and 0 otherwise, \\\\(p_{ij}\\\\) is the predicted probability that image \\\\(j\\\\) belongs to class \\\\(i\\\\).\n\nI have a statistics background (not saying that I'm sure I understand this correctly though) so I'm pretty used to seeing indices like this. I believe an issue though that the writer made is by simplifying the loss function to that of a single example (rather than using 3 indices, \\\\(i\\\\), \\\\(j\\\\), and \\\\(k\\\\)), they got a bit mixed up by what they mean by class and observation. What I wrote in the end is still a bit off in my mind. I've seen indices like this (e.g., [Bishop](https://www.microsoft.com/en-us/research/uploads/prod/2006/01/Bishop-Pattern-Recognition-and-Machine-Learning-2006.pdf)) where it feels like the \\\\(i\\\\) and \\\\(j\\\\) are inverted. Generally, we like to think of i as belonging to the observation, and j belonging to the feature. So having \\\\(i\\\\) as class and \\\\(j\\\\) as image number for the \\\\(i\\\\)'th class runs a bit counterintuitive to how we normally think about it.",
      "votes": 3,
      "replies": [
        {
          "id": 1908338,
          "postDate": "2022-08-21T14:58:00.087Z",
          "content": "<p>Thanks for your valuable comments. But I believe the loss function is for all observations and all classes.<br>\nA multi-class log loss function is usually defined as:</p>\n<p>\\[  \\text{Log Loss} = -\\frac{1}{N} \\sum_{j=1}^{M} \\sum_{i=1}^{N} y_{j,i} \\ln p_{j,i} \\]</p>\n<p>where \\( N \\) is the total number of observations (images), \\( \\ln \\) is the natural logarithm, \\( y_{j,i} \\) is \\( 1 \\) if the \\( ith \\) observation (image) belongs to class \\( j \\) and \\( 0 \\) otherwise, \\( p_{j,i} \\) is the predicted probability that image \\( i \\) belongs class \\(  j \\).</p>\n<p>The primary disadvantage of this multi-class loss function is that if number of observations in each class are not equal (or not almost equal), then the loss function value is biased toward the class having the most number of observations. </p>\n<p>A simple and straight forward extension of this multi-class log loss is weighted multi-class log loss function i.e.,:</p>\n<p>\\[  \\text{Log Loss} = -\\frac{1}{N} \\sum_{j=1}^{M} w_{j} \\sum_{i=1}^{N} y_{j,i} \\ln p_{j,i} \\] such that \\[  \\sum_{j=1}^{M} w_{j} = 1 \\]</p>\n<p>where, \\( w_{j} \\) is the assigned weight for class \\( j \\). </p>\n<p>Further, if we want to give equal importance to all classes then higher weights \\( w_{j} \\)s should be assigned to classes that have less number of observations.</p>\n<p>Valuable input from all the Kagglers specially the competition organizers and Kaggle staff will be highly appreciated.</p>",
          "rawMarkdown": "Thanks for your valuable comments. But I believe the loss function is for all observations and all classes.\nA multi-class log loss function is usually defined as:\n\n\\\\[  \\text{Log Loss} = -\\frac{1}{N} \\sum_{j=1}^{M} \\sum_{i=1}^{N} y_{j,i} \\ln p_{j,i} \\\\]\n\nwhere \\\\( N \\\\) is the total number of observations (images), \\\\( \\ln \\\\) is the natural logarithm, \\\\( y_{j,i} \\\\) is \\\\( 1 \\\\) if the \\\\( ith \\\\) observation (image) belongs to class \\\\( j \\\\) and \\\\( 0 \\\\) otherwise, \\\\( p_{j,i} \\\\) is the predicted probability that image \\\\( i \\\\) belongs class \\\\(  j \\\\).\n\nThe primary disadvantage of this multi-class loss function is that if number of observations in each class are not equal (or not almost equal), then the loss function value is biased toward the class having the most number of observations. \n\nA simple and straight forward extension of this multi-class log loss is weighted multi-class log loss function i.e.,:\n\n\\\\[  \\text{Log Loss} = -\\frac{1}{N} \\sum_{j=1}^{M} w_{j} \\sum_{i=1}^{N} y_{j,i} \\ln p_{j,i} \\\\] such that \\\\[  \\sum_{j=1}^{M} w_{j} = 1 \\\\]\n\nwhere, \\\\( w_{j} \\\\) is the assigned weight for class \\\\( j \\\\). \n\nFurther, if we want to give equal importance to all classes then higher weights \\\\( w_{j} \\\\)s should be assigned to classes that have less number of observations.\n\nValuable input from all the Kagglers specially the competition organizers and Kaggle staff will be highly appreciated.\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 1906242,
      "postDate": "2022-08-19T18:13:08.843Z",
      "content": "<p>Submissions are evaluated using a weighted multi-class logarithmic loss.<br>\nThe formula is then:</p>\n<p>\\[ \\text{Log Loss} = - \\left( \\frac{\\sum_{i=1}^{M} w_{i} \\cdot \\sum_{j=1}^{N_{i}} \\frac{y_{ij}}{N_{i}} \\cdot \\ln  p_{ij} }{\\sum_{i=1}^{M} w_{i}} \\right) \\]</p>\n<p>where \\(N\\) is the number of images in the class set, \\(M\\) is the number of classes, \\(ln\\) is the natural logarithm, \\( y_{i,j} \\) is 1 if observation \\(i\\) belongs to class \\(j\\) and 0 otherwise, \\(p_{i,j}\\) is the predicted probability that image \\(i\\) belongs to class \\(j\\).</p>\n<p>The above formula has following technical flaws:<br>\n1) \\( N_{i} \\)  is not defined.<br>\n2) The index of class is \\(i\\) running from 1 to \\(M\\), then how \\(y_{i,j}\\) belongs to class \\(j\\) (see. \\( y_{i,j} \\) is \\(1\\) if observation \\(i\\) belongs to class \\(j\\) and \\(0\\) otherwise).<br>\n3) Same problem is with \\( p_{i,j} \\), how \\( p_{i,j} \\) belongs to class \\(j\\).</p>",
      "rawMarkdown": "Submissions are evaluated using a weighted multi-class logarithmic loss.\nThe formula is then:\n\n\\\\[ \\text{Log Loss} = - \\left( \\frac{\\sum_{i=1}^{M} w_{i} \\cdot \\sum_{j=1}^{N_{i}} \\frac{y_{ij}}{N_{i}} \\cdot \\ln  p_{ij} }{\\sum_{i=1}^{M} w_{i}} \\right) \\\\]\n\n\nwhere \\\\(N\\\\) is the number of images in the class set, \\\\(M\\\\) is the number of classes, \\\\(ln\\\\) is the natural logarithm, \\\\( y_{i,j} \\\\) is 1 if observation \\\\(i\\\\) belongs to class \\\\(j\\\\) and 0 otherwise, \\\\(p_{i,j}\\\\) is the predicted probability that image \\\\(i\\\\) belongs to class \\\\(j\\\\).\n\nThe above formula has following technical flaws:\n1) \\\\( N_{i} \\\\)  is not defined.\n2) The index of class is \\\\(i\\\\) running from 1 to \\\\(M\\\\), then how \\\\(y_{i,j}\\\\) belongs to class \\\\(j\\\\) (see. \\\\( y_{i,j} \\\\) is \\\\(1\\\\) if observation \\\\(i\\\\) belongs to class \\\\(j\\\\) and \\\\(0\\\\) otherwise).\n3) Same problem is with \\\\( p_{i,j} \\\\), how \\\\( p_{i,j} \\\\) belongs to class \\\\(j\\\\).\n\n",
      "votes": 4
    },
    {
      "id": 1907315,
      "postDate": "2022-08-20T16:58:57.513Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1906501,
      "author_name": "yqz",
      "author_url": "",
      "post_date": "2022-08-19T23:37:30.983000",
      "content": "<p>I think that for this loss function it's really the Log Loss for a single observation. My guess for #1 then is that \\(N_i\\) is the number of images per class of an observation. However, an observation can only belong to one class, so if we have \\(i=1\\) representing CA and \\(i=2\\) representing LAA, then if the example belongs to CA, \\(N_2=0.\\)</p>\n<p>For #2 and # 3, I believe \\(y_{i.j}\\) is an <a href=\"https://en.wikipedia.org/wiki/Indicator_function\" target=\"_blank\">indicator function</a>, which is 1 is the observation belongs to the \\(i\\)'th class and zero otherwise. So following our previous example of the single observation being CA with \\(i=1\\), then the indicator function would appear as follows, \\(y_{1,j}=1,\\) and \\(y_{2,j}=0\\).</p>\n<p>The effect then is that the \\(p_{2,j}\\) values are zeroed out from the loss function. They're not really required then for the resulting Log Loss. The calculation really then is just the negative natural log of the weighted probabilities for the true class. I think a typo they have though is that instead of </p>\n<blockquote>\n  <p>ln is the natural logarithm, \\(y_{ij}\\) is 1 if observation \\(i\\) belongs to class \\(j\\) and 0 otherwise, \\(p_{ij}\\) is the predicted probability that image \\(i\\) belongs to class \\(j\\).</p>\n</blockquote>\n<p>it should read</p>\n<blockquote>\n  <p>ln is the natural logarithm, \\(y_{ij}\\) is 1 if observation \\(j\\) belongs to class \\(i\\) and 0 otherwise, \\(p_{ij}\\) is the predicted probability that image \\(j\\) belongs to class \\(i\\).</p>\n</blockquote>\n<p>I have a statistics background (not saying that I'm sure I understand this correctly though) so I'm pretty used to seeing indices like this. I believe an issue though that the writer made is by simplifying the loss function to that of a single example (rather than using 3 indices, \\(i\\), \\(j\\), and \\(k\\)), they got a bit mixed up by what they mean by class and observation. What I wrote in the end is still a bit off in my mind. I've seen indices like this (e.g., <a href=\"https://www.microsoft.com/en-us/research/uploads/prod/2006/01/Bishop-Pattern-Recognition-and-Machine-Learning-2006.pdf\" target=\"_blank\">Bishop</a>) where it feels like the \\(i\\) and \\(j\\) are inverted. Generally, we like to think of i as belonging to the observation, and j belonging to the feature. So having \\(i\\) as class and \\(j\\) as image number for the \\(i\\)'th class runs a bit counterintuitive to how we normally think about it.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1908338,
          "author_name": "Syed Muhammad Atif",
          "author_url": "",
          "post_date": "2022-08-21T14:58:00.087000",
          "content": "<p>Thanks for your valuable comments. But I believe the loss function is for all observations and all classes.<br>\nA multi-class log loss function is usually defined as:</p>\n<p>\\[  \\text{Log Loss} = -\\frac{1}{N} \\sum_{j=1}^{M} \\sum_{i=1}^{N} y_{j,i} \\ln p_{j,i} \\]</p>\n<p>where \\( N \\) is the total number of observations (images), \\( \\ln \\) is the natural logarithm, \\( y_{j,i} \\) is \\( 1 \\) if the \\( ith \\) observation (image) belongs to class \\( j \\) and \\( 0 \\) otherwise, \\( p_{j,i} \\) is the predicted probability that image \\( i \\) belongs class \\(  j \\).</p>\n<p>The primary disadvantage of this multi-class loss function is that if number of observations in each class are not equal (or not almost equal), then the loss function value is biased toward the class having the most number of observations. </p>\n<p>A simple and straight forward extension of this multi-class log loss is weighted multi-class log loss function i.e.,:</p>\n<p>\\[  \\text{Log Loss} = -\\frac{1}{N} \\sum_{j=1}^{M} w_{j} \\sum_{i=1}^{N} y_{j,i} \\ln p_{j,i} \\] such that \\[  \\sum_{j=1}^{M} w_{j} = 1 \\]</p>\n<p>where, \\( w_{j} \\) is the assigned weight for class \\( j \\). </p>\n<p>Further, if we want to give equal importance to all classes then higher weights \\( w_{j} \\)s should be assigned to classes that have less number of observations.</p>\n<p>Valuable input from all the Kagglers specially the competition organizers and Kaggle staff will be highly appreciated.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1907315,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-20T16:58:57.513000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1906501": "I think that for this loss function it's really the Log Loss for a single observation. My guess for #1 then is that \\\\(N_i\\\\) is the number of images per class of an observation. However, an observation can only belong to one class, so if we have \\\\(i=1\\\\) representing CA and \\\\(i=2\\\\) representing LAA, then if the example belongs to CA, \\\\(N_2=0.\\\\)\n\nFor #2 and # 3, I believe \\\\(y_{i.j}\\\\) is an [indicator function](https://en.wikipedia.org/wiki/Indicator_function), which is 1 is the observation belongs to the \\\\(i\\\\)'th class and zero otherwise. So following our previous example of the single observation being CA with \\\\(i=1\\\\), then the indicator function would appear as follows, \\\\(y_{1,j}=1,\\\\) and \\\\(y_{2,j}=0\\\\).\n\nThe effect then is that the \\\\(p_{2,j}\\\\) values are zeroed out from the loss function. They're not really required then for the resulting Log Loss. The calculation really then is just the negative natural log of the weighted probabilities for the true class. I think a typo they have though is that instead of \n> ln is the natural logarithm, \\\\(y_{ij}\\\\) is 1 if observation \\\\(i\\\\) belongs to class \\\\(j\\\\) and 0 otherwise, \\\\(p_{ij}\\\\) is the predicted probability that image \\\\(i\\\\) belongs to class \\\\(j\\\\).\n\nit should read\n\n> ln is the natural logarithm, \\\\(y_{ij}\\\\) is 1 if observation \\\\(j\\\\) belongs to class \\\\(i\\\\) and 0 otherwise, \\\\(p_{ij}\\\\) is the predicted probability that image \\\\(j\\\\) belongs to class \\\\(i\\\\).\n\nI have a statistics background (not saying that I'm sure I understand this correctly though) so I'm pretty used to seeing indices like this. I believe an issue though that the writer made is by simplifying the loss function to that of a single example (rather than using 3 indices, \\\\(i\\\\), \\\\(j\\\\), and \\\\(k\\\\)), they got a bit mixed up by what they mean by class and observation. What I wrote in the end is still a bit off in my mind. I've seen indices like this (e.g., [Bishop](https://www.microsoft.com/en-us/research/uploads/prod/2006/01/Bishop-Pattern-Recognition-and-Machine-Learning-2006.pdf)) where it feels like the \\\\(i\\\\) and \\\\(j\\\\) are inverted. Generally, we like to think of i as belonging to the observation, and j belonging to the feature. So having \\\\(i\\\\) as class and \\\\(j\\\\) as image number for the \\\\(i\\\\)'th class runs a bit counterintuitive to how we normally think about it.",
    "1906242": "Submissions are evaluated using a weighted multi-class logarithmic loss.\nThe formula is then:\n\n\\\\[ \\text{Log Loss} = - \\left( \\frac{\\sum_{i=1}^{M} w_{i} \\cdot \\sum_{j=1}^{N_{i}} \\frac{y_{ij}}{N_{i}} \\cdot \\ln  p_{ij} }{\\sum_{i=1}^{M} w_{i}} \\right) \\\\]\n\n\nwhere \\\\(N\\\\) is the number of images in the class set, \\\\(M\\\\) is the number of classes, \\\\(ln\\\\) is the natural logarithm, \\\\( y_{i,j} \\\\) is 1 if observation \\\\(i\\\\) belongs to class \\\\(j\\\\) and 0 otherwise, \\\\(p_{i,j}\\\\) is the predicted probability that image \\\\(i\\\\) belongs to class \\\\(j\\\\).\n\nThe above formula has following technical flaws:\n1) \\\\( N_{i} \\\\)  is not defined.\n2) The index of class is \\\\(i\\\\) running from 1 to \\\\(M\\\\), then how \\\\(y_{i,j}\\\\) belongs to class \\\\(j\\\\) (see. \\\\( y_{i,j} \\\\) is \\\\(1\\\\) if observation \\\\(i\\\\) belongs to class \\\\(j\\\\) and \\\\(0\\\\) otherwise).\n3) Same problem is with \\\\( p_{i,j} \\\\), how \\\\( p_{i,j} \\\\) belongs to class \\\\(j\\\\).\n\n",
    "1907315": ""
  }
}