{
  "id": 369879,
  "title": "Competition Metric Explained - probabilistic F1 score ",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369879",
  "author_name": "Sandy",
  "post_date": "2022-12-01T17:06:39.448000",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi Kagglers! I found it a bit difficult to grasp the evaluation metric as cited in the paper under the evaluation tab. So I decided to probe further with some simplified examples. I am just sharing my understanding here to gauge the findings &amp; help others, though most of you guys might already understood it clearly. Do let me know if I have gone wrong somewhere.</p>\n<p><strong>Assumptions/Simplifications:</strong></p>\n<ol>\n<li>We are dealing with a binary classification problem (which we really are in this competition).</li>\n<li>An imaginary dataset of 10 examples with their true labels.</li>\n<li>An imaginary model with a sigmoid head which outputs the probability/confidence(p) of a dataset example belonging to class 0 or probability (1-p) belonging to class 1. This can also be represented as model confidence score for a particular dataset (say x1) as M(x1, 0) for class 0 i.e (p) and M(x1, 1) for class 1 i.e (1-p).</li>\n</ol>\n<p><strong>Example Dataset, Model Confidence &amp; Predictions:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F6744945526c9df2e39e9fdaf4ab59010%2FCapture.PNG?generation=1669910536161507&amp;alt=media\" alt=\"\"></p>\n<p><strong>Building the probabilistic confusion matrix (pCM):</strong></p>\n<p>The probabilistic confusion matrix is calculated from the formula:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2Fef0a11bbe6d46be929ff515a97c29121%2FCapture1.PNG?generation=1669911539957936&amp;alt=media\" alt=\"\"></p>\n<p>where, jref  is the true label &amp; jhyp is the label predicted by model.</p>\n<p>We will begin with datasets whose True label is 0 first -- i.e jref = 0:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F0eb96f3f009e9878e4499476c116cdb9%2FCapture2.PNG?generation=1669911518835995&amp;alt=media\" alt=\"\"></p>\n<p>Out of this filtered dataset, we can calculate** pCM(jref=0 , jhyp=0)** as:</p>\n<p>M(x1, 0) + M(x2, 0),+M(x4, 0)+M(x8, 0)+M(x9, 0)= 0.735+0.381+0.689+0.423+0.821 = <strong>3.049</strong></p>\n<p>and <strong>pCM(jref=0 , jhyp=1)</strong> as:</p>\n<p>M(x1, 1) + M(x2, 1),+M(x4, 1)+M(x8, 1)+M(x9, 1)= 0.265+0.619+0.311+0.577+0.179 = <strong>1.951</strong> (alternatively 5 - 3.049)</p>\n<p>Similarly, <strong>pCM(jref=1 , jhyp=0) &amp; pCM(jref=1 , jhyp=0)</strong> can be calculated to be <strong>2.585</strong> &amp; <strong>2.415</strong> respectively.</p>\n<p>Then probabilistic confusion matrix can be populated as below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F62d21a2ddfe7f161fafcf61b0db4af74%2FCapture3.PNG?generation=1669913937698645&amp;alt=media\" alt=\"\"></p>\n<p>pPrecision, pRecall &amp; pF1 score can then be deduced.</p>\n<p>Please forgive the naive approach or erros. Once again request all your valuable inputs. Best wishes for the competition!!!</p>",
  "messages": [
    {
      "id": 2051877,
      "postDate": "2022-12-01T17:06:39.450Z",
      "content": "<p>Hi Kagglers! I found it a bit difficult to grasp the evaluation metric as cited in the paper under the evaluation tab. So I decided to probe further with some simplified examples. I am just sharing my understanding here to gauge the findings &amp; help others, though most of you guys might already understood it clearly. Do let me know if I have gone wrong somewhere.</p>\n<p><strong>Assumptions/Simplifications:</strong></p>\n<ol>\n<li>We are dealing with a binary classification problem (which we really are in this competition).</li>\n<li>An imaginary dataset of 10 examples with their true labels.</li>\n<li>An imaginary model with a sigmoid head which outputs the probability/confidence(p) of a dataset example belonging to class 0 or probability (1-p) belonging to class 1. This can also be represented as model confidence score for a particular dataset (say x1) as M(x1, 0) for class 0 i.e (p) and M(x1, 1) for class 1 i.e (1-p).</li>\n</ol>\n<p><strong>Example Dataset, Model Confidence &amp; Predictions:</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F6744945526c9df2e39e9fdaf4ab59010%2FCapture.PNG?generation=1669910536161507&amp;alt=media\" alt=\"\"></p>\n<p><strong>Building the probabilistic confusion matrix (pCM):</strong></p>\n<p>The probabilistic confusion matrix is calculated from the formula:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2Fef0a11bbe6d46be929ff515a97c29121%2FCapture1.PNG?generation=1669911539957936&amp;alt=media\" alt=\"\"></p>\n<p>where, jref  is the true label &amp; jhyp is the label predicted by model.</p>\n<p>We will begin with datasets whose True label is 0 first -- i.e jref = 0:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F0eb96f3f009e9878e4499476c116cdb9%2FCapture2.PNG?generation=1669911518835995&amp;alt=media\" alt=\"\"></p>\n<p>Out of this filtered dataset, we can calculate** pCM(jref=0 , jhyp=0)** as:</p>\n<p>M(x1, 0) + M(x2, 0),+M(x4, 0)+M(x8, 0)+M(x9, 0)= 0.735+0.381+0.689+0.423+0.821 = <strong>3.049</strong></p>\n<p>and <strong>pCM(jref=0 , jhyp=1)</strong> as:</p>\n<p>M(x1, 1) + M(x2, 1),+M(x4, 1)+M(x8, 1)+M(x9, 1)= 0.265+0.619+0.311+0.577+0.179 = <strong>1.951</strong> (alternatively 5 - 3.049)</p>\n<p>Similarly, <strong>pCM(jref=1 , jhyp=0) &amp; pCM(jref=1 , jhyp=0)</strong> can be calculated to be <strong>2.585</strong> &amp; <strong>2.415</strong> respectively.</p>\n<p>Then probabilistic confusion matrix can be populated as below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F62d21a2ddfe7f161fafcf61b0db4af74%2FCapture3.PNG?generation=1669913937698645&amp;alt=media\" alt=\"\"></p>\n<p>pPrecision, pRecall &amp; pF1 score can then be deduced.</p>\n<p>Please forgive the naive approach or erros. Once again request all your valuable inputs. Best wishes for the competition!!!</p>",
      "rawMarkdown": "Hi Kagglers! I found it a bit difficult to grasp the evaluation metric as cited in the paper under the evaluation tab. So I decided to probe further with some simplified examples. I am just sharing my understanding here to gauge the findings & help others, though most of you guys might already understood it clearly. Do let me know if I have gone wrong somewhere.\n\n**Assumptions/Simplifications:**\n1. We are dealing with a binary classification problem (which we really are in this competition).\n2. An imaginary dataset of 10 examples with their true labels.\n3. An imaginary model with a sigmoid head which outputs the probability/confidence(p) of a dataset example belonging to class 0 or probability (1-p) belonging to class 1. This can also be represented as model confidence score for a particular dataset (say x1) as M(x1, 0) for class 0 i.e (p) and M(x1, 1) for class 1 i.e (1-p).\n\n**Example Dataset, Model Confidence & Predictions:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F6744945526c9df2e39e9fdaf4ab59010%2FCapture.PNG?generation=1669910536161507&alt=media)\n\n**Building the probabilistic confusion matrix (pCM):**\n\nThe probabilistic confusion matrix is calculated from the formula:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2Fef0a11bbe6d46be929ff515a97c29121%2FCapture1.PNG?generation=1669911539957936&alt=media)\n\nwhere, jref  is the true label & jhyp is the label predicted by model.\n\nWe will begin with datasets whose True label is 0 first -- i.e jref = 0:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F0eb96f3f009e9878e4499476c116cdb9%2FCapture2.PNG?generation=1669911518835995&alt=media)\n\nOut of this filtered dataset, we can calculate** pCM(jref=0 , jhyp=0)** as:\n\nM(x1, 0) + M(x2, 0),+M(x4, 0)+M(x8, 0)+M(x9, 0)= 0.735+0.381+0.689+0.423+0.821 = **3.049**\n\nand **pCM(jref=0 , jhyp=1)** as:\n\nM(x1, 1) + M(x2, 1),+M(x4, 1)+M(x8, 1)+M(x9, 1)= 0.265+0.619+0.311+0.577+0.179 = **1.951** (alternatively 5 - 3.049)\n\nSimilarly, **pCM(jref=1 , jhyp=0) & pCM(jref=1 , jhyp=0)** can be calculated to be **2.585** & **2.415** respectively.\n\nThen probabilistic confusion matrix can be populated as below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F62d21a2ddfe7f161fafcf61b0db4af74%2FCapture3.PNG?generation=1669913937698645&alt=media)\n\npPrecision, pRecall & pF1 score can then be deduced.\n\nPlease forgive the naive approach or erros. Once again request all your valuable inputs. Best wishes for the competition!!!\n",
      "votes": 9
    },
    {
      "id": 2113685,
      "postDate": "2023-01-24T12:22:37.137Z",
      "content": "<p>Though this topic has already 2 months: worthy contribution explaining:<br>\nprobabilistic confusion matrix (pPrecision, pRecall &amp; pF1) competition metric.</p>",
      "rawMarkdown": "Though this topic has already 2 months: worthy contribution explaining:\nprobabilistic confusion matrix (pPrecision, pRecall & pF1) competition metric."
    }
  ],
  "comments": [
    {
      "id": 2113685,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2023-01-24T12:22:37.137000",
      "content": "<p>Though this topic has already 2 months: worthy contribution explaining:<br>\nprobabilistic confusion matrix (pPrecision, pRecall &amp; pF1) competition metric.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2051877": "Hi Kagglers! I found it a bit difficult to grasp the evaluation metric as cited in the paper under the evaluation tab. So I decided to probe further with some simplified examples. I am just sharing my understanding here to gauge the findings & help others, though most of you guys might already understood it clearly. Do let me know if I have gone wrong somewhere.\n\n**Assumptions/Simplifications:**\n1. We are dealing with a binary classification problem (which we really are in this competition).\n2. An imaginary dataset of 10 examples with their true labels.\n3. An imaginary model with a sigmoid head which outputs the probability/confidence(p) of a dataset example belonging to class 0 or probability (1-p) belonging to class 1. This can also be represented as model confidence score for a particular dataset (say x1) as M(x1, 0) for class 0 i.e (p) and M(x1, 1) for class 1 i.e (1-p).\n\n**Example Dataset, Model Confidence & Predictions:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F6744945526c9df2e39e9fdaf4ab59010%2FCapture.PNG?generation=1669910536161507&alt=media)\n\n**Building the probabilistic confusion matrix (pCM):**\n\nThe probabilistic confusion matrix is calculated from the formula:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2Fef0a11bbe6d46be929ff515a97c29121%2FCapture1.PNG?generation=1669911539957936&alt=media)\n\nwhere, jref  is the true label & jhyp is the label predicted by model.\n\nWe will begin with datasets whose True label is 0 first -- i.e jref = 0:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F0eb96f3f009e9878e4499476c116cdb9%2FCapture2.PNG?generation=1669911518835995&alt=media)\n\nOut of this filtered dataset, we can calculate** pCM(jref=0 , jhyp=0)** as:\n\nM(x1, 0) + M(x2, 0),+M(x4, 0)+M(x8, 0)+M(x9, 0)= 0.735+0.381+0.689+0.423+0.821 = **3.049**\n\nand **pCM(jref=0 , jhyp=1)** as:\n\nM(x1, 1) + M(x2, 1),+M(x4, 1)+M(x8, 1)+M(x9, 1)= 0.265+0.619+0.311+0.577+0.179 = **1.951** (alternatively 5 - 3.049)\n\nSimilarly, **pCM(jref=1 , jhyp=0) & pCM(jref=1 , jhyp=0)** can be calculated to be **2.585** & **2.415** respectively.\n\nThen probabilistic confusion matrix can be populated as below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9335656%2F62d21a2ddfe7f161fafcf61b0db4af74%2FCapture3.PNG?generation=1669913937698645&alt=media)\n\npPrecision, pRecall & pF1 score can then be deduced.\n\nPlease forgive the naive approach or erros. Once again request all your valuable inputs. Best wishes for the competition!!!\n",
    "2113685": "Though this topic has already 2 months: worthy contribution explaining:\nprobabilistic confusion matrix (pPrecision, pRecall & pF1) competition metric."
  }
}