{
  "id": 353665,
  "title": "What accuracy should be a good result?",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/353665",
  "author_name": "Pablo Larrosa",
  "post_date": "2022-09-19T13:48:17.559000",
  "votes": 2,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi everyone, until now I got a loss around 0.56 but the accuracy is very low  (0.73).</p>\n<p>Thanks in advance</p>",
  "messages": [
    {
      "id": 1946046,
      "postDate": "2022-09-19T15:21:23.137Z",
      "content": "<p><em>&gt;very low</em></p>\n<p><em>&gt;0.73</em></p>\n<p>Are you even sure that you aren't overfitted yet</p>",
      "rawMarkdown": "*>very low*\n\n*>0.73*\n\nAre you even sure that you aren't overfitted yet",
      "votes": 5
    },
    {
      "id": 1945923,
      "postDate": "2022-09-19T13:48:17.560Z",
      "content": "<p>Hi everyone, until now I got a loss around 0.56 but the accuracy is very low  (0.73).</p>\n<p>Thanks in advance</p>",
      "rawMarkdown": "Hi everyone, until now I got a loss around 0.56 but the accuracy is very low  (0.73).\n\nThanks in advance",
      "votes": 2
    },
    {
      "id": 1954335,
      "postDate": "2022-09-25T06:48:06.270Z",
      "content": "<p>I use 5 fold cross-validation (stratify target and non-overlapping patient_ids). This is my best model. Not sure this is good or not…</p>\n<pre><code>{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6655420602789024,\n      \"log_loss\": 0.5360595129819867\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5798369457148538,\n      \"log_loss\": 0.6151186826747733\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6552579365079365,\n      \"log_loss\": 0.5286322120778464\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6325799955247259,\n      \"log_loss\": 0.5594933565209309\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.7004329004329004,\n      \"log_loss\": 0.5591903770679996\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6483851310176723,\n    \"log_loss\": 0.5599818563222174\n  }\n}\n</code></pre>",
      "rawMarkdown": "I use 5 fold cross-validation (stratify target and non-overlapping patient_ids). This is my best model. Not sure this is good or not...\n\n```\n{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6655420602789024,\n      \"log_loss\": 0.5360595129819867\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5798369457148538,\n      \"log_loss\": 0.6151186826747733\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6552579365079365,\n      \"log_loss\": 0.5286322120778464\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6325799955247259,\n      \"log_loss\": 0.5594933565209309\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.7004329004329004,\n      \"log_loss\": 0.5591903770679996\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6483851310176723,\n    \"log_loss\": 0.5599818563222174\n  }\n}\n```",
      "votes": 2
    },
    {
      "id": 1958871,
      "postDate": "2022-09-27T17:18:44.173Z",
      "content": "<p>Accuracy doesn't mean much in this competition. I have an accuracy of about 0.9. The validation model confidently determines the image class (adjusted for accuracy). But here is another metric, good accuracy only harms it. How it is calculated, I still do not understand.</p>",
      "rawMarkdown": "Accuracy doesn't mean much in this competition. I have an accuracy of about 0.9. The validation model confidently determines the image class (adjusted for accuracy). But here is another metric, good accuracy only harms it. How it is calculated, I still do not understand.",
      "replies": [
        {
          "id": 1958878,
          "postDate": "2022-09-27T17:22:54.553Z",
          "content": "<p>You mean the scoring metric? It's a variant on the weighted log loss (weighted binary crossentropy). The formula is shown on the dataset page </p>",
          "rawMarkdown": "You mean the scoring metric? It's a variant on the weighted log loss (weighted binary crossentropy). The formula is shown on the dataset page ",
          "votes": 1
        },
        {
          "id": 1958914,
          "postDate": "2022-09-27T17:55:01.823Z",
          "content": "<p>Yes, good accuracy worsens the metric of this competition.</p>",
          "rawMarkdown": "Yes, good accuracy worsens the metric of this competition.",
          "votes": 1
        },
        {
          "id": 1959813,
          "postDate": "2022-09-28T10:39:31.477Z",
          "content": "<p>Keep in mind that the test set for the public score is really small and that the current scores on the LB can be totally wrong. You should try to optimize your model with cross-validation for a better idea of how it'll fare in the end. :)</p>",
          "rawMarkdown": "Keep in mind that the test set for the public score is really small and that the current scores on the LB can be totally wrong. You should try to optimize your model with cross-validation for a better idea of how it'll fare in the end. :)",
          "votes": 1
        },
        {
          "id": 1960644,
          "postDate": "2022-09-28T17:33:07.280Z",
          "content": "<p>Thanks for the advice, next time I'll definitely do it that way))</p>",
          "rawMarkdown": "Thanks for the advice, next time I'll definitely do it that way))"
        }
      ]
    },
    {
      "id": 1952138,
      "postDate": "2022-09-23T13:17:58.687Z",
      "content": "<p>I have got 0.68 with a small homemade cnn with 21K parameters, but without taking into account the images of the same person.</p>",
      "rawMarkdown": "I have got 0.68 with a small homemade cnn with 21K parameters, but without taking into account the images of the same person."
    },
    {
      "id": 1947252,
      "postDate": "2022-09-20T11:32:05.610Z",
      "content": "<p>Public test took 9 minutes</p>",
      "rawMarkdown": "Public test took 9 minutes",
      "replies": [
        {
          "id": 1947572,
          "postDate": "2022-09-20T14:34:25.940Z",
          "content": "<p>There's ~70x the images in the hidden set compared to the public set. We don't know the sizes of them, so that affects the public execution time too. You'll likely have to get the inference time below 6min to get it passing, I think. </p>",
          "rawMarkdown": "There's ~70x the images in the hidden set compared to the public set. We don't know the sizes of them, so that affects the public execution time too. You'll likely have to get the inference time below 6min to get it passing, I think. "
        }
      ]
    },
    {
      "id": 1947227,
      "postDate": "2022-09-20T11:05:49.853Z",
      "content": "<p>Anyway my submission failed by timeout ☹️</p>",
      "rawMarkdown": "Anyway my submission failed by timeout ☹️",
      "replies": [
        {
          "id": 1947251,
          "postDate": "2022-09-20T11:28:58.940Z",
          "content": "<p>How long did it take to run on the public test set?</p>",
          "rawMarkdown": "How long did it take to run on the public test set?"
        }
      ]
    },
    {
      "id": 1946597,
      "postDate": "2022-09-19T23:52:18.783Z",
      "content": "<p>Hi, I want to clarify that the value 0.73 is a mean of 5 folders, I have accuracy of 0.77, 0.84, 0.69, 0.63 , 0.72</p>",
      "rawMarkdown": "Hi, I want to clarify that the value 0.73 is a mean of 5 folders, I have accuracy of 0.77, 0.84, 0.69, 0.63 , 0.72",
      "replies": [
        {
          "id": 1947009,
          "postDate": "2022-09-20T07:56:40.350Z",
          "content": "<blockquote>\n  <p>0.84</p>\n</blockquote>\n<p>oh boy</p>",
          "rawMarkdown": ">0.84\n\noh boy",
          "votes": 1
        },
        {
          "id": 1947110,
          "postDate": "2022-09-20T09:08:53.320Z",
          "content": "<p>Weighted or unweighted accuracies and how have you addressed class imbalance? Just a single metric like this doesn't tell us a lot about your workflow, and pretty much every step in the process can affect it. 0.84 is not bad at all :)</p>\n<p>Since there's a fair bit of wiggle room between your fold accuracies, a couple of things that come to mind are:</p>\n<ul>\n<li>Increase your fold sizes (and hence the support for the metrics)</li>\n<li>Lower your learning rates</li>\n<li>Check for overfitting</li>\n</ul>\n<p>Again, there's a million things that can cause this, but these are some common reasons that might be affecting your workflow. If you share your notebook, we can look into it :)</p>",
          "rawMarkdown": "Weighted or unweighted accuracies and how have you addressed class imbalance? Just a single metric like this doesn't tell us a lot about your workflow, and pretty much every step in the process can affect it. 0.84 is not bad at all :)\n\nSince there's a fair bit of wiggle room between your fold accuracies, a couple of things that come to mind are:\n\n- Increase your fold sizes (and hence the support for the metrics)\n- Lower your learning rates\n- Check for overfitting\n\nAgain, there's a million things that can cause this, but these are some common reasons that might be affecting your workflow. If you share your notebook, we can look into it :)",
          "votes": 1
        },
        {
          "id": 1953172,
          "postDate": "2022-09-24T09:20:07.140Z",
          "content": "<p>What is your AUC-ROC score?</p>",
          "rawMarkdown": "What is your AUC-ROC score?"
        },
        {
          "id": 1953753,
          "postDate": "2022-09-24T17:56:48.530Z",
          "content": "<p>0.57 is my AUC ROC</p>",
          "rawMarkdown": "0.57 is my AUC ROC",
          "votes": 1
        },
        {
          "id": 1953812,
          "postDate": "2022-09-24T18:45:11.810Z",
          "content": "<p>I was able to get up to a 0.49(avg) AUC score leading to 0.6 lb. 0.57 is a questionable score.<br>\nAlso, I would like to argue whether accuracy is a good way to judge the performance of our models for this comp? or even AUC-ROC for that matter? </p>",
          "rawMarkdown": "I was able to get up to a 0.49(avg) AUC score leading to 0.6 lb. 0.57 is a questionable score.\nAlso, I would like to argue whether accuracy is a good way to judge the performance of our models for this comp? or even AUC-ROC for that matter? "
        }
      ]
    },
    {
      "id": 1946380,
      "postDate": "2022-09-19T19:32:45.333Z",
      "content": "<p>An accuracy of 0.73 isn't very low, assuming it's weighted, and that you've accounted for the imbalance in the set. Apparently, many people in the forum here struggle to get it past 0.5.<br>\nThat being said, it's unknown how easily discernable these images are (hence, there's a competition to chart the territory), so nobody will be able to tell what a good accuracy for this dataset is. Just keep doing your best and good luck :)</p>",
      "rawMarkdown": "An accuracy of 0.73 isn't very low, assuming it's weighted, and that you've accounted for the imbalance in the set. Apparently, many people in the forum here struggle to get it past 0.5.\nThat being said, it's unknown how easily discernable these images are (hence, there's a competition to chart the territory), so nobody will be able to tell what a good accuracy for this dataset is. Just keep doing your best and good luck :)"
    }
  ],
  "comments": [
    {
      "id": 1946046,
      "author_name": "majoraregalia",
      "author_url": "",
      "post_date": "2022-09-19T15:21:23.137000",
      "content": "<p><em>&gt;very low</em></p>\n<p><em>&gt;0.73</em></p>\n<p>Are you even sure that you aren't overfitted yet</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1954335,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-09-25T06:48:06.270000",
      "content": "<p>I use 5 fold cross-validation (stratify target and non-overlapping patient_ids). This is my best model. Not sure this is good or not…</p>\n<pre><code>{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6655420602789024,\n      \"log_loss\": 0.5360595129819867\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5798369457148538,\n      \"log_loss\": 0.6151186826747733\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6552579365079365,\n      \"log_loss\": 0.5286322120778464\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6325799955247259,\n      \"log_loss\": 0.5594933565209309\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.7004329004329004,\n      \"log_loss\": 0.5591903770679996\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6483851310176723,\n    \"log_loss\": 0.5599818563222174\n  }\n}\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1958871,
      "author_name": "Zaakcii Ru",
      "author_url": "",
      "post_date": "2022-09-27T17:18:44.173000",
      "content": "<p>Accuracy doesn't mean much in this competition. I have an accuracy of about 0.9. The validation model confidently determines the image class (adjusted for accuracy). But here is another metric, good accuracy only harms it. How it is calculated, I still do not understand.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1958878,
          "author_name": "David Landup",
          "author_url": "",
          "post_date": "2022-09-27T17:22:54.553000",
          "content": "<p>You mean the scoring metric? It's a variant on the weighted log loss (weighted binary crossentropy). The formula is shown on the dataset page </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1958914,
          "author_name": "Zaakcii Ru",
          "author_url": "",
          "post_date": "2022-09-27T17:55:01.823000",
          "content": "<p>Yes, good accuracy worsens the metric of this competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1959813,
          "author_name": "David Landup",
          "author_url": "",
          "post_date": "2022-09-28T10:39:31.477000",
          "content": "<p>Keep in mind that the test set for the public score is really small and that the current scores on the LB can be totally wrong. You should try to optimize your model with cross-validation for a better idea of how it'll fare in the end. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1960644,
          "author_name": "Zaakcii Ru",
          "author_url": "",
          "post_date": "2022-09-28T17:33:07.280000",
          "content": "<p>Thanks for the advice, next time I'll definitely do it that way))</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1952138,
      "author_name": "Pierre Tisseur",
      "author_url": "",
      "post_date": "2022-09-23T13:17:58.687000",
      "content": "<p>I have got 0.68 with a small homemade cnn with 21K parameters, but without taking into account the images of the same person.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1947252,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-09-20T11:32:05.610000",
      "content": "<p>Public test took 9 minutes</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1947572,
          "author_name": "David Landup",
          "author_url": "",
          "post_date": "2022-09-20T14:34:25.940000",
          "content": "<p>There's ~70x the images in the hidden set compared to the public set. We don't know the sizes of them, so that affects the public execution time too. You'll likely have to get the inference time below 6min to get it passing, I think. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1947227,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-09-20T11:05:49.853000",
      "content": "<p>Anyway my submission failed by timeout ☹️</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1947251,
          "author_name": "David Landup",
          "author_url": "",
          "post_date": "2022-09-20T11:28:58.940000",
          "content": "<p>How long did it take to run on the public test set?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1946597,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-09-19T23:52:18.783000",
      "content": "<p>Hi, I want to clarify that the value 0.73 is a mean of 5 folders, I have accuracy of 0.77, 0.84, 0.69, 0.63 , 0.72</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1947009,
          "author_name": "majoraregalia",
          "author_url": "",
          "post_date": "2022-09-20T07:56:40.350000",
          "content": "<blockquote>\n  <p>0.84</p>\n</blockquote>\n<p>oh boy</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1947110,
          "author_name": "David Landup",
          "author_url": "",
          "post_date": "2022-09-20T09:08:53.320000",
          "content": "<p>Weighted or unweighted accuracies and how have you addressed class imbalance? Just a single metric like this doesn't tell us a lot about your workflow, and pretty much every step in the process can affect it. 0.84 is not bad at all :)</p>\n<p>Since there's a fair bit of wiggle room between your fold accuracies, a couple of things that come to mind are:</p>\n<ul>\n<li>Increase your fold sizes (and hence the support for the metrics)</li>\n<li>Lower your learning rates</li>\n<li>Check for overfitting</li>\n</ul>\n<p>Again, there's a million things that can cause this, but these are some common reasons that might be affecting your workflow. If you share your notebook, we can look into it :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1953172,
          "author_name": "ANANDIYA SHEEL DIWAN",
          "author_url": "",
          "post_date": "2022-09-24T09:20:07.140000",
          "content": "<p>What is your AUC-ROC score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1953753,
          "author_name": "Pablo Larrosa",
          "author_url": "",
          "post_date": "2022-09-24T17:56:48.530000",
          "content": "<p>0.57 is my AUC ROC</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1953812,
          "author_name": "ANANDIYA SHEEL DIWAN",
          "author_url": "",
          "post_date": "2022-09-24T18:45:11.810000",
          "content": "<p>I was able to get up to a 0.49(avg) AUC score leading to 0.6 lb. 0.57 is a questionable score.<br>\nAlso, I would like to argue whether accuracy is a good way to judge the performance of our models for this comp? or even AUC-ROC for that matter? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1946380,
      "author_name": "David Landup",
      "author_url": "",
      "post_date": "2022-09-19T19:32:45.333000",
      "content": "<p>An accuracy of 0.73 isn't very low, assuming it's weighted, and that you've accounted for the imbalance in the set. Apparently, many people in the forum here struggle to get it past 0.5.<br>\nThat being said, it's unknown how easily discernable these images are (hence, there's a competition to chart the territory), so nobody will be able to tell what a good accuracy for this dataset is. Just keep doing your best and good luck :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1946046": "*>very low*\n\n*>0.73*\n\nAre you even sure that you aren't overfitted yet",
    "1945923": "Hi everyone, until now I got a loss around 0.56 but the accuracy is very low  (0.73).\n\nThanks in advance",
    "1954335": "I use 5 fold cross-validation (stratify target and non-overlapping patient_ids). This is my best model. Not sure this is good or not...\n\n```\n{\n  \"fold_scores\": {\n    \"fold1\": {\n      \"accuracy\": 0.7450980392156863,\n      \"roc_auc\": 0.6655420602789024,\n      \"log_loss\": 0.5360595129819867\n    },\n    \"fold2\": {\n      \"accuracy\": 0.6948051948051948,\n      \"roc_auc\": 0.5798369457148538,\n      \"log_loss\": 0.6151186826747733\n    },\n    \"fold3\": {\n      \"accuracy\": 0.7635135135135135,\n      \"roc_auc\": 0.6552579365079365,\n      \"log_loss\": 0.5286322120778464\n    },\n    \"fold4\": {\n      \"accuracy\": 0.74,\n      \"roc_auc\": 0.6325799955247259,\n      \"log_loss\": 0.5594933565209309\n    },\n    \"fold5\": {\n      \"accuracy\": 0.7181208053691275,\n      \"roc_auc\": 0.7004329004329004,\n      \"log_loss\": 0.5591903770679996\n    }\n  },\n  \"oof_scores\": {\n    \"accuracy\": 0.7320954907161804,\n    \"roc_auc\": 0.6483851310176723,\n    \"log_loss\": 0.5599818563222174\n  }\n}\n```",
    "1958871": "Accuracy doesn't mean much in this competition. I have an accuracy of about 0.9. The validation model confidently determines the image class (adjusted for accuracy). But here is another metric, good accuracy only harms it. How it is calculated, I still do not understand.",
    "1952138": "I have got 0.68 with a small homemade cnn with 21K parameters, but without taking into account the images of the same person.",
    "1947252": "Public test took 9 minutes",
    "1947227": "Anyway my submission failed by timeout ☹️",
    "1946597": "Hi, I want to clarify that the value 0.73 is a mean of 5 folders, I have accuracy of 0.77, 0.84, 0.69, 0.63 , 0.72",
    "1946380": "An accuracy of 0.73 isn't very low, assuming it's weighted, and that you've accounted for the imbalance in the set. Apparently, many people in the forum here struggle to get it past 0.5.\nThat being said, it's unknown how easily discernable these images are (hence, there's a competition to chart the territory), so nobody will be able to tell what a good accuracy for this dataset is. Just keep doing your best and good luck :)"
  }
}