{
  "id": 186680,
  "title": "Metric Clarification, Python Code - 2",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/186680",
  "author_name": "Sergey Bryansky",
  "post_date": "2020-09-25T13:48:54.221000",
  "votes": 18,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Due to Evaluation Page and <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924\" target=\"_blank\">discussion</a> started by <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> I found mistake in my own implementation of competition metric, please, review this :) </p>\n<pre><code>from sklearn.metrics import log_loss\n\n\ndef competition_score(submission, ground_truth):\n    exam_level = [\n        'negative_exam_for_pe',\n        'rv_lv_ratio_gte_1',\n        'rv_lv_ratio_lt_1',\n        'leftsided_pe',\n        'chronic_pe',\n        'rightsided_pe',\n        'acute_and_chronic_pe', \n        'central_pe',\n        'indeterminate'\n    ]\n\n    image_level = 'pe_present_on_image'\n\n    exam_weights = {\n        \"negative_exam_for_pe\": 0.0736196319,\n        \"indeterminate\": 0.09202453988,\n        \"chronic_pe\": 0.1042944785,\n        \"acute_and_chronic_pe\": 0.1042944785,\n        \"central_pe\": 0.1877300613,\n        \"leftsided_pe\": 0.06257668712,\n        \"rightsided_pe\": 0.06257668712,\n        \"rv_lv_ratio_gte_1\": 0.2346625767,\n        \"rv_lv_ratio_lt_1\": 0.0782208589\n    }\n\n    # EXAM LEVEL\n    exam_loss = []\n    n_exam_samples = ground_truth.StudyInstanceUID.nunique()\n    for target in exam_level:\n        exam_loss.append(\n            log_loss(\n                ground_truth.groupby(\"StudyInstanceUID\")[target].mean(),\n                submission.groupby(\"StudyInstanceUID\")[target].mean(),\n                sample_weight=[exam_weights[target]] * n_exam_samples\n            )\n        )\n\n    # IMAGE LEVEL\n    weight = 0.07361963\n    image_weights = weight * ground_truth.groupby(\"StudyInstanceUID\")[image_level].transform(\"mean\")\n    image_loss = log_loss(\n        ground_truth[image_level],\n        submission[image_level],\n        sample_weight=image_weights\n    )\n\n    # TOTAL LOSS\n    exam_loss.append(image_loss)\n    total_loss = exam_loss\n    total_loss = (np.mean(total_loss)) / (np.mean(image_weights) + np.sum(list(exam_weights.values())))\n\n    return total_loss\n</code></pre>\n<p>Using <code>scipy.optimize.minimize</code> routine I found best constant predictions on train part</p>\n<pre><code>np.array(\n    [\n        0.67498484, 0.12917683, 0.17459449, \n        0.21217964, 0.04017838, 0.25771374, \n        0.01987881, 0.05501893, 0.02157019,\n        0.28973937\n    ]\n)\n</code></pre>\n<p>for labels</p>\n<pre><code>[\n    'negative_exam_for_pe', 'rv_lv_ratio_gte_1', 'rv_lv_ratio_lt_1',\n    'leftsided_pe', 'chronic_pe', 'rightsided_pe',\n    'acute_and_chronic_pe', 'central_pe', 'indeterminate',\n    'pe_present_on_image'\n]\n</code></pre>\n<p>It is similar to means baseline but slightly better.</p>\n<p>And it seems it would be good to add reviewed version of metric implementation to check consistency <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">notebook</a>.</p>",
  "messages": [
    {
      "id": 1026688,
      "postDate": "2020-09-25T13:48:54.220Z",
      "content": "<p>Due to Evaluation Page and <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924\" target=\"_blank\">discussion</a> started by <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> I found mistake in my own implementation of competition metric, please, review this :) </p>\n<pre><code>from sklearn.metrics import log_loss\n\n\ndef competition_score(submission, ground_truth):\n    exam_level = [\n        'negative_exam_for_pe',\n        'rv_lv_ratio_gte_1',\n        'rv_lv_ratio_lt_1',\n        'leftsided_pe',\n        'chronic_pe',\n        'rightsided_pe',\n        'acute_and_chronic_pe', \n        'central_pe',\n        'indeterminate'\n    ]\n\n    image_level = 'pe_present_on_image'\n\n    exam_weights = {\n        \"negative_exam_for_pe\": 0.0736196319,\n        \"indeterminate\": 0.09202453988,\n        \"chronic_pe\": 0.1042944785,\n        \"acute_and_chronic_pe\": 0.1042944785,\n        \"central_pe\": 0.1877300613,\n        \"leftsided_pe\": 0.06257668712,\n        \"rightsided_pe\": 0.06257668712,\n        \"rv_lv_ratio_gte_1\": 0.2346625767,\n        \"rv_lv_ratio_lt_1\": 0.0782208589\n    }\n\n    # EXAM LEVEL\n    exam_loss = []\n    n_exam_samples = ground_truth.StudyInstanceUID.nunique()\n    for target in exam_level:\n        exam_loss.append(\n            log_loss(\n                ground_truth.groupby(\"StudyInstanceUID\")[target].mean(),\n                submission.groupby(\"StudyInstanceUID\")[target].mean(),\n                sample_weight=[exam_weights[target]] * n_exam_samples\n            )\n        )\n\n    # IMAGE LEVEL\n    weight = 0.07361963\n    image_weights = weight * ground_truth.groupby(\"StudyInstanceUID\")[image_level].transform(\"mean\")\n    image_loss = log_loss(\n        ground_truth[image_level],\n        submission[image_level],\n        sample_weight=image_weights\n    )\n\n    # TOTAL LOSS\n    exam_loss.append(image_loss)\n    total_loss = exam_loss\n    total_loss = (np.mean(total_loss)) / (np.mean(image_weights) + np.sum(list(exam_weights.values())))\n\n    return total_loss\n</code></pre>\n<p>Using <code>scipy.optimize.minimize</code> routine I found best constant predictions on train part</p>\n<pre><code>np.array(\n    [\n        0.67498484, 0.12917683, 0.17459449, \n        0.21217964, 0.04017838, 0.25771374, \n        0.01987881, 0.05501893, 0.02157019,\n        0.28973937\n    ]\n)\n</code></pre>\n<p>for labels</p>\n<pre><code>[\n    'negative_exam_for_pe', 'rv_lv_ratio_gte_1', 'rv_lv_ratio_lt_1',\n    'leftsided_pe', 'chronic_pe', 'rightsided_pe',\n    'acute_and_chronic_pe', 'central_pe', 'indeterminate',\n    'pe_present_on_image'\n]\n</code></pre>\n<p>It is similar to means baseline but slightly better.</p>\n<p>And it seems it would be good to add reviewed version of metric implementation to check consistency <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">notebook</a>.</p>",
      "rawMarkdown": "Due to Evaluation Page and [discussion](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924) started by @yuval6967 I found mistake in my own implementation of competition metric, please, review this :) \n\n```\nfrom sklearn.metrics import log_loss\n\n\ndef competition_score(submission, ground_truth):\n    exam_level = [\n        'negative_exam_for_pe',\n        'rv_lv_ratio_gte_1',\n        'rv_lv_ratio_lt_1',\n        'leftsided_pe',\n        'chronic_pe',\n        'rightsided_pe',\n        'acute_and_chronic_pe', \n        'central_pe',\n        'indeterminate'\n    ]\n    \n    image_level = 'pe_present_on_image'\n    \n    exam_weights = {\n        \"negative_exam_for_pe\": 0.0736196319,\n        \"indeterminate\": 0.09202453988,\n        \"chronic_pe\": 0.1042944785,\n        \"acute_and_chronic_pe\": 0.1042944785,\n        \"central_pe\": 0.1877300613,\n        \"leftsided_pe\": 0.06257668712,\n        \"rightsided_pe\": 0.06257668712,\n        \"rv_lv_ratio_gte_1\": 0.2346625767,\n        \"rv_lv_ratio_lt_1\": 0.0782208589\n    }\n    \n    # EXAM LEVEL\n    exam_loss = []\n    n_exam_samples = ground_truth.StudyInstanceUID.nunique()\n    for target in exam_level:\n        exam_loss.append(\n            log_loss(\n                ground_truth.groupby(\"StudyInstanceUID\")[target].mean(),\n                submission.groupby(\"StudyInstanceUID\")[target].mean(),\n                sample_weight=[exam_weights[target]] * n_exam_samples\n            )\n        )\n        \n    # IMAGE LEVEL\n    weight = 0.07361963\n    image_weights = weight * ground_truth.groupby(\"StudyInstanceUID\")[image_level].transform(\"mean\")\n    image_loss = log_loss(\n        ground_truth[image_level],\n        submission[image_level],\n        sample_weight=image_weights\n    )\n    \n    # TOTAL LOSS\n    exam_loss.append(image_loss)\n    total_loss = exam_loss\n    total_loss = (np.mean(total_loss)) / (np.mean(image_weights) + np.sum(list(exam_weights.values())))\n    \n    return total_loss\n```\n\nUsing `scipy.optimize.minimize` routine I found best constant predictions on train part\n```\nnp.array(\n    [\n        0.67498484, 0.12917683, 0.17459449, \n        0.21217964, 0.04017838, 0.25771374, \n        0.01987881, 0.05501893, 0.02157019,\n        0.28973937\n    ]\n)\n```\nfor labels\n```\n[\n    'negative_exam_for_pe', 'rv_lv_ratio_gte_1', 'rv_lv_ratio_lt_1',\n    'leftsided_pe', 'chronic_pe', 'rightsided_pe',\n    'acute_and_chronic_pe', 'central_pe', 'indeterminate',\n    'pe_present_on_image'\n]\n```\nIt is similar to means baseline but slightly better.\n\nAnd it seems it would be good to add reviewed version of metric implementation to check consistency [notebook](https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check).",
      "votes": 17
    },
    {
      "id": 1027696,
      "postDate": "2020-09-26T09:55:30.497Z",
      "content": "<p>Bravo ! GOOD job </p>",
      "rawMarkdown": "Bravo ! GOOD job ",
      "votes": -3
    },
    {
      "id": 1027677,
      "postDate": "2020-09-26T09:45:05.857Z",
      "content": "<p>good job ! go ahead</p>",
      "rawMarkdown": "good job ! go ahead",
      "votes": -3
    },
    {
      "id": 1040409,
      "postDate": "2020-10-07T05:59:15.930Z",
      "content": "<p>Hello Sergey, </p>\n<p>Thanks for the implementation.</p>\n<p>When would you use the \"constant\" predictions?</p>\n<p>Thanks </p>",
      "rawMarkdown": "Hello Sergey, \n\nThanks for the implementation.\n\nWhen would you use the \"constant\" predictions?\n\nThanks "
    }
  ],
  "comments": [
    {
      "id": 1027696,
      "author_name": "Naim Mhedhbi",
      "author_url": "",
      "post_date": "2020-09-26T09:55:30.497000",
      "content": "<p>Bravo ! GOOD job </p>",
      "votes": -3,
      "replies": []
    },
    {
      "id": 1027677,
      "author_name": "Naim Mhedhbi",
      "author_url": "",
      "post_date": "2020-09-26T09:45:05.857000",
      "content": "<p>good job ! go ahead</p>",
      "votes": -3,
      "replies": []
    },
    {
      "id": 1040409,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-10-07T05:59:15.930000",
      "content": "<p>Hello Sergey, </p>\n<p>Thanks for the implementation.</p>\n<p>When would you use the \"constant\" predictions?</p>\n<p>Thanks </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1026688": "Due to Evaluation Page and [discussion](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183924) started by @yuval6967 I found mistake in my own implementation of competition metric, please, review this :) \n\n```\nfrom sklearn.metrics import log_loss\n\n\ndef competition_score(submission, ground_truth):\n    exam_level = [\n        'negative_exam_for_pe',\n        'rv_lv_ratio_gte_1',\n        'rv_lv_ratio_lt_1',\n        'leftsided_pe',\n        'chronic_pe',\n        'rightsided_pe',\n        'acute_and_chronic_pe', \n        'central_pe',\n        'indeterminate'\n    ]\n    \n    image_level = 'pe_present_on_image'\n    \n    exam_weights = {\n        \"negative_exam_for_pe\": 0.0736196319,\n        \"indeterminate\": 0.09202453988,\n        \"chronic_pe\": 0.1042944785,\n        \"acute_and_chronic_pe\": 0.1042944785,\n        \"central_pe\": 0.1877300613,\n        \"leftsided_pe\": 0.06257668712,\n        \"rightsided_pe\": 0.06257668712,\n        \"rv_lv_ratio_gte_1\": 0.2346625767,\n        \"rv_lv_ratio_lt_1\": 0.0782208589\n    }\n    \n    # EXAM LEVEL\n    exam_loss = []\n    n_exam_samples = ground_truth.StudyInstanceUID.nunique()\n    for target in exam_level:\n        exam_loss.append(\n            log_loss(\n                ground_truth.groupby(\"StudyInstanceUID\")[target].mean(),\n                submission.groupby(\"StudyInstanceUID\")[target].mean(),\n                sample_weight=[exam_weights[target]] * n_exam_samples\n            )\n        )\n        \n    # IMAGE LEVEL\n    weight = 0.07361963\n    image_weights = weight * ground_truth.groupby(\"StudyInstanceUID\")[image_level].transform(\"mean\")\n    image_loss = log_loss(\n        ground_truth[image_level],\n        submission[image_level],\n        sample_weight=image_weights\n    )\n    \n    # TOTAL LOSS\n    exam_loss.append(image_loss)\n    total_loss = exam_loss\n    total_loss = (np.mean(total_loss)) / (np.mean(image_weights) + np.sum(list(exam_weights.values())))\n    \n    return total_loss\n```\n\nUsing `scipy.optimize.minimize` routine I found best constant predictions on train part\n```\nnp.array(\n    [\n        0.67498484, 0.12917683, 0.17459449, \n        0.21217964, 0.04017838, 0.25771374, \n        0.01987881, 0.05501893, 0.02157019,\n        0.28973937\n    ]\n)\n```\nfor labels\n```\n[\n    'negative_exam_for_pe', 'rv_lv_ratio_gte_1', 'rv_lv_ratio_lt_1',\n    'leftsided_pe', 'chronic_pe', 'rightsided_pe',\n    'acute_and_chronic_pe', 'central_pe', 'indeterminate',\n    'pe_present_on_image'\n]\n```\nIt is similar to means baseline but slightly better.\n\nAnd it seems it would be good to add reviewed version of metric implementation to check consistency [notebook](https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check).",
    "1027696": "Bravo ! GOOD job ",
    "1027677": "good job ! go ahead",
    "1040409": "Hello Sergey, \n\nThanks for the implementation.\n\nWhen would you use the \"constant\" predictions?\n\nThanks "
  }
}