{
  "id": 371583,
  "title": "Is the evaluation metric pF1 score appropriate for screening mammography?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371583",
  "author_name": "Taiji",
  "post_date": "2022-12-11T05:13:27.888000",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I think the pF1 score is not very appropriate metric for evaluating the effectiveness of a screening mammogram system. Please correct me if I was wrong …</p>\n<p>The pF1 score is a probabilistic version of F1, which is 2*Precision *Recall/(Precision + Recall). The problem is that it weights the precision and recall the same importance and is maximized only when  P and R are closed. </p>\n<p>In screening mammography however, the <strong>Recall is much more important than Precision</strong>, and <a href=\"https://en.wikipedia.org/wiki/Sensitivity_and_specificity\" target=\"_blank\">Sensitvity (same as recall) and Specifity (different from precision)</a> are normally used for evaluation. </p>\n<p>If the pF1 is used to evaluate the performance of a screening mammography system in real-life scenario , the precision will be extremely low and so will the pF1 score. </p>\n<p>I refer to this paper:<br>\n<em>Lehman, Constance D., et al. \"National performance benchmarks for modern screening digital mammography: update from the Breast Cancer Surveillance Consortium.\" Radiology 283.1 (2017): 49.</em><br>\nand cite several facts:</p>\n<ol>\n<li>The US national benchmark performance of screening mammography is sensitivity=0.87 and specificity=0.89.</li>\n<li>On average for every 1000 women screened, 116 will be called-back for further examination. (Rated as Positive by a mammogram reader)</li>\n<li>Among the 116 women called-back,  only 5.1 women will be finally diagnoised with cancer. <br>\nAs a result, if evaluated with the F1 score, recall=sensitivity=0.87, precision=5.1/116=0.044, so<br>\nF1 = 2<em>0.044</em>0.87/(0.87+0.044)=<strong>0.084</strong><br>\nwhich is** very low**. Changing F1 to pF1 may not help since pF1 still weight the importance of precision and recall the same. </li>\n</ol>\n<p>In my opinion, it is better to use something like 2*Sensitivity *Sepcificity/(sensitivity+specificity), or use AUROC or pAUROC as evaluation metric, otherwise the evaluation will diverage too much from the real life scenario and have limited value for clinical references. </p>",
  "messages": [
    {
      "id": 2061406,
      "postDate": "2022-12-11T05:13:27.890Z",
      "content": "<p>I think the pF1 score is not very appropriate metric for evaluating the effectiveness of a screening mammogram system. Please correct me if I was wrong …</p>\n<p>The pF1 score is a probabilistic version of F1, which is 2*Precision *Recall/(Precision + Recall). The problem is that it weights the precision and recall the same importance and is maximized only when  P and R are closed. </p>\n<p>In screening mammography however, the <strong>Recall is much more important than Precision</strong>, and <a href=\"https://en.wikipedia.org/wiki/Sensitivity_and_specificity\" target=\"_blank\">Sensitvity (same as recall) and Specifity (different from precision)</a> are normally used for evaluation. </p>\n<p>If the pF1 is used to evaluate the performance of a screening mammography system in real-life scenario , the precision will be extremely low and so will the pF1 score. </p>\n<p>I refer to this paper:<br>\n<em>Lehman, Constance D., et al. \"National performance benchmarks for modern screening digital mammography: update from the Breast Cancer Surveillance Consortium.\" Radiology 283.1 (2017): 49.</em><br>\nand cite several facts:</p>\n<ol>\n<li>The US national benchmark performance of screening mammography is sensitivity=0.87 and specificity=0.89.</li>\n<li>On average for every 1000 women screened, 116 will be called-back for further examination. (Rated as Positive by a mammogram reader)</li>\n<li>Among the 116 women called-back,  only 5.1 women will be finally diagnoised with cancer. <br>\nAs a result, if evaluated with the F1 score, recall=sensitivity=0.87, precision=5.1/116=0.044, so<br>\nF1 = 2<em>0.044</em>0.87/(0.87+0.044)=<strong>0.084</strong><br>\nwhich is** very low**. Changing F1 to pF1 may not help since pF1 still weight the importance of precision and recall the same. </li>\n</ol>\n<p>In my opinion, it is better to use something like 2*Sensitivity *Sepcificity/(sensitivity+specificity), or use AUROC or pAUROC as evaluation metric, otherwise the evaluation will diverage too much from the real life scenario and have limited value for clinical references. </p>",
      "rawMarkdown": "I think the pF1 score is not very appropriate metric for evaluating the effectiveness of a screening mammogram system. Please correct me if I was wrong ...\n\nThe pF1 score is a probabilistic version of F1, which is 2*Precision *Recall/(Precision + Recall). The problem is that it weights the precision and recall the same importance and is maximized only when  P and R are closed. \n\nIn screening mammography however, the **Recall is much more important than Precision**, and [Sensitvity (same as recall) and Specifity (different from precision)](https://en.wikipedia.org/wiki/Sensitivity_and_specificity) are normally used for evaluation. \n\nIf the pF1 is used to evaluate the performance of a screening mammography system in real-life scenario , the precision will be extremely low and so will the pF1 score. \n\nI refer to this paper:\n*Lehman, Constance D., et al. \"National performance benchmarks for modern screening digital mammography: update from the Breast Cancer Surveillance Consortium.\" Radiology 283.1 (2017): 49.*\nand cite several facts:\n1. The US national benchmark performance of screening mammography is sensitivity=0.87 and specificity=0.89.\n2.  On average for every 1000 women screened, 116 will be called-back for further examination. (Rated as Positive by a mammogram reader)\n3. Among the 116 women called-back,  only 5.1 women will be finally diagnoised with cancer. \n As a result, if evaluated with the F1 score, recall=sensitivity=0.87, precision=5.1/116=0.044, so\nF1 = 2*0.044*0.87/(0.87+0.044)=**0.084**\nwhich is** very low**. Changing F1 to pF1 may not help since pF1 still weight the importance of precision and recall the same. \n\nIn my opinion, it is better to use something like 2*Sensitivity *Sepcificity/(sensitivity+specificity), or use AUROC or pAUROC as evaluation metric, otherwise the evaluation will diverage too much from the real life scenario and have limited value for clinical references. \n\n \n\n",
      "votes": 6
    },
    {
      "id": 2061415,
      "postDate": "2022-12-11T05:50:13.407Z",
      "content": "<p>FDA usually reported sensitivity, specificity, AUROC for Computer-Aided Detection (CAD) software. <br>\ni would think a score like the below is good.<br>\nscore = w1 * sensitivity + w2 * specificity +w3 * AUROC </p>\n<p>w1,w2,w3 are importance weights</p>",
      "rawMarkdown": "FDA usually reported sensitivity, specificity, AUROC for Computer-Aided Detection (CAD) software. \ni would think a score like the below is good.\nscore = w1 * sensitivity + w2 * specificity +w3 * AUROC \n\nw1,w2,w3 are importance weights\n",
      "replies": [
        {
          "id": 2061436,
          "postDate": "2022-12-11T06:13:17.143Z",
          "content": "<p>yes, the pF1 evaluation troubles me a lot. To get a 50% precision rate, the specificity should be at least 99% according to US national statistics. <br>\nTo meet this metric, I have to tune the model output score in a very strange way, even modifying the loss function to let the prediction value be heavily screwed towards 0. I think it may heavily diverge from real clinical practices. </p>",
          "rawMarkdown": "yes, the pF1 evaluation troubles me a lot. To get a 50% precision rate, the specificity should be at least 99% according to US national statistics. \nTo meet this metric, I have to tune the model output score in a very strange way, even modifying the loss function to let the prediction value be heavily screwed towards 0. I think it may heavily diverge from real clinical practices. \n"
        },
        {
          "id": 2061455,
          "postDate": "2022-12-11T06:34:58.357Z",
          "content": "<p>\" I have to tune the model output score in a very strange way,\"</p>\n<p>competition and practical applications are two different things.</p>",
          "rawMarkdown": "\" I have to tune the model output score in a very strange way,\"\n\ncompetition and practical applications are two different things.\n ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2061415,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-11T05:50:13.407000",
      "content": "<p>FDA usually reported sensitivity, specificity, AUROC for Computer-Aided Detection (CAD) software. <br>\ni would think a score like the below is good.<br>\nscore = w1 * sensitivity + w2 * specificity +w3 * AUROC </p>\n<p>w1,w2,w3 are importance weights</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2061436,
          "author_name": "Taiji",
          "author_url": "",
          "post_date": "2022-12-11T06:13:17.143000",
          "content": "<p>yes, the pF1 evaluation troubles me a lot. To get a 50% precision rate, the specificity should be at least 99% according to US national statistics. <br>\nTo meet this metric, I have to tune the model output score in a very strange way, even modifying the loss function to let the prediction value be heavily screwed towards 0. I think it may heavily diverge from real clinical practices. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2061455,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-11T06:34:58.357000",
          "content": "<p>\" I have to tune the model output score in a very strange way,\"</p>\n<p>competition and practical applications are two different things.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2061406": "I think the pF1 score is not very appropriate metric for evaluating the effectiveness of a screening mammogram system. Please correct me if I was wrong ...\n\nThe pF1 score is a probabilistic version of F1, which is 2*Precision *Recall/(Precision + Recall). The problem is that it weights the precision and recall the same importance and is maximized only when  P and R are closed. \n\nIn screening mammography however, the **Recall is much more important than Precision**, and [Sensitvity (same as recall) and Specifity (different from precision)](https://en.wikipedia.org/wiki/Sensitivity_and_specificity) are normally used for evaluation. \n\nIf the pF1 is used to evaluate the performance of a screening mammography system in real-life scenario , the precision will be extremely low and so will the pF1 score. \n\nI refer to this paper:\n*Lehman, Constance D., et al. \"National performance benchmarks for modern screening digital mammography: update from the Breast Cancer Surveillance Consortium.\" Radiology 283.1 (2017): 49.*\nand cite several facts:\n1. The US national benchmark performance of screening mammography is sensitivity=0.87 and specificity=0.89.\n2.  On average for every 1000 women screened, 116 will be called-back for further examination. (Rated as Positive by a mammogram reader)\n3. Among the 116 women called-back,  only 5.1 women will be finally diagnoised with cancer. \n As a result, if evaluated with the F1 score, recall=sensitivity=0.87, precision=5.1/116=0.044, so\nF1 = 2*0.044*0.87/(0.87+0.044)=**0.084**\nwhich is** very low**. Changing F1 to pF1 may not help since pF1 still weight the importance of precision and recall the same. \n\nIn my opinion, it is better to use something like 2*Sensitivity *Sepcificity/(sensitivity+specificity), or use AUROC or pAUROC as evaluation metric, otherwise the evaluation will diverage too much from the real life scenario and have limited value for clinical references. \n\n \n\n",
    "2061415": "FDA usually reported sensitivity, specificity, AUROC for Computer-Aided Detection (CAD) software. \ni would think a score like the below is good.\nscore = w1 * sensitivity + w2 * specificity +w3 * AUROC \n\nw1,w2,w3 are importance weights\n"
  }
}