{
  "id": 371493,
  "title": "Only CSV data gives good score? why?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371493",
  "author_name": "VK",
  "post_date": "2022-12-10T14:34:26.661000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I was trying only csv data and apply xgboost model. And the score is 0.02</p>\n<p>But the real time is different instead of competition i now !</p>\n<p>But why we focus on image data?</p>\n<p>And another question in the tabular data input feature \"Machine id\" ?</p>",
  "messages": [
    {
      "id": 2060892,
      "postDate": "2022-12-10T14:34:26.660Z",
      "content": "<p>I was trying only csv data and apply xgboost model. And the score is 0.02</p>\n<p>But the real time is different instead of competition i now !</p>\n<p>But why we focus on image data?</p>\n<p>And another question in the tabular data input feature \"Machine id\" ?</p>",
      "rawMarkdown": "I was trying only csv data and apply xgboost model. And the score is 0.02\n\nBut the real time is different instead of competition i now !\n\nBut why we focus on image data?\n\nAnd another question in the tabular data input feature \"Machine id\" ?",
      "votes": 5
    },
    {
      "id": 2062600,
      "postDate": "2022-12-12T08:43:14.243Z",
      "content": "<p>As far as I remember, competition hosts in an other topic clarified, that not all 'table variables' are available for the test images. </p>\n<blockquote>\n  <p>This is the correct answer. We don't provide either difficult_negative_case or BIRADS for the test set, so you shouldn't use either column as a feature.</p>\n</blockquote>\n<p>Same for <code>density</code> and <code>biopsy</code></p>",
      "rawMarkdown": "As far as I remember, competition hosts in an other topic clarified, that not all 'table variables' are available for the test images. \n> This is the correct answer. We don't provide either difficult_negative_case or BIRADS for the test set, so you shouldn't use either column as a feature.\n\nSame for `density` and `biopsy`",
      "votes": 3
    },
    {
      "id": 2061149,
      "postDate": "2022-12-10T18:22:46.577Z",
      "content": "<p>If you trained an image classifier and a tabular data classifier you can later blend their predictions together. In the melanoma competition it resulted for me in a nice LB boost (private + public)</p>\n<p>I have assigned 70% weight to the image classifier and 30% weight for the tabular classifier by easily summing the predictions up:<br>\n<code>blended_results = img_clf_results * 0.7 + tabular_clf_results * 0.3</code></p>\n<p>Overweighting the image classifier is natural because it does, when correctly trained, perform better, recognizes more patterns and therefore is more accurate then the tabular classifier in this competition which also has way less features to work with.</p>",
      "rawMarkdown": "If you trained an image classifier and a tabular data classifier you can later blend their predictions together. In the melanoma competition it resulted for me in a nice LB boost (private + public)\n\nI have assigned 70% weight to the image classifier and 30% weight for the tabular classifier by easily summing the predictions up:\n`blended_results = img_clf_results * 0.7 + tabular_clf_results * 0.3`\n\nOverweighting the image classifier is natural because it does, when correctly trained, perform better, recognizes more patterns and therefore is more accurate then the tabular classifier in this competition which also has way less features to work with.",
      "votes": 2,
      "replies": [
        {
          "id": 2062396,
          "postDate": "2022-12-12T04:27:41.867Z",
          "content": "<p>It maybe not work on pF1.</p>",
          "rawMarkdown": "It maybe not work on pF1."
        },
        {
          "id": 2063252,
          "postDate": "2022-12-12T19:04:29.617Z",
          "content": "<p>The way to calculate pF1 is affected directly by its model predictions. Thus blending has an impact on the score. Following numpy code excerpt shows the impact on the pF1 score:</p>\n<pre><code>[...]\nctp = preds[labels==].()\ncfp = preds[labels==].()\nbeta_squared = beta * beta\nc_precision = ctp / (ctp + cfp)\nc_recall = ctp / y_true_count\n[...]\nresult = ( + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall)\n</code></pre>\n<p>My previous defined variable <code>blended_results</code> would be the <code>preds</code> in this code and would directly affect the result after. </p>",
          "rawMarkdown": "The way to calculate pF1 is affected directly by its model predictions. Thus blending has an impact on the score. Following numpy code excerpt shows the impact on the pF1 score:\n\n```python\n[...]\nctp = preds[labels==1].sum()\ncfp = preds[labels==0].sum()\nbeta_squared = beta * beta\nc_precision = ctp / (ctp + cfp)\nc_recall = ctp / y_true_count\n[...]\nresult = (1 + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall)\n```\nMy previous defined variable `blended_results` would be the `preds` in this code and would directly affect the result after. ",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2062600,
      "author_name": "JanGlinko2",
      "author_url": "",
      "post_date": "2022-12-12T08:43:14.243000",
      "content": "<p>As far as I remember, competition hosts in an other topic clarified, that not all 'table variables' are available for the test images. </p>\n<blockquote>\n  <p>This is the correct answer. We don't provide either difficult_negative_case or BIRADS for the test set, so you shouldn't use either column as a feature.</p>\n</blockquote>\n<p>Same for <code>density</code> and <code>biopsy</code></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2061149,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2022-12-10T18:22:46.577000",
      "content": "<p>If you trained an image classifier and a tabular data classifier you can later blend their predictions together. In the melanoma competition it resulted for me in a nice LB boost (private + public)</p>\n<p>I have assigned 70% weight to the image classifier and 30% weight for the tabular classifier by easily summing the predictions up:<br>\n<code>blended_results = img_clf_results * 0.7 + tabular_clf_results * 0.3</code></p>\n<p>Overweighting the image classifier is natural because it does, when correctly trained, perform better, recognizes more patterns and therefore is more accurate then the tabular classifier in this competition which also has way less features to work with.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2062396,
          "author_name": "fate",
          "author_url": "",
          "post_date": "2022-12-12T04:27:41.867000",
          "content": "<p>It maybe not work on pF1.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2063252,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2022-12-12T19:04:29.617000",
          "content": "<p>The way to calculate pF1 is affected directly by its model predictions. Thus blending has an impact on the score. Following numpy code excerpt shows the impact on the pF1 score:</p>\n<pre><code>[...]\nctp = preds[labels==].()\ncfp = preds[labels==].()\nbeta_squared = beta * beta\nc_precision = ctp / (ctp + cfp)\nc_recall = ctp / y_true_count\n[...]\nresult = ( + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall)\n</code></pre>\n<p>My previous defined variable <code>blended_results</code> would be the <code>preds</code> in this code and would directly affect the result after. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2060892": "I was trying only csv data and apply xgboost model. And the score is 0.02\n\nBut the real time is different instead of competition i now !\n\nBut why we focus on image data?\n\nAnd another question in the tabular data input feature \"Machine id\" ?",
    "2062600": "As far as I remember, competition hosts in an other topic clarified, that not all 'table variables' are available for the test images. \n> This is the correct answer. We don't provide either difficult_negative_case or BIRADS for the test set, so you shouldn't use either column as a feature.\n\nSame for `density` and `biopsy`",
    "2061149": "If you trained an image classifier and a tabular data classifier you can later blend their predictions together. In the melanoma competition it resulted for me in a nice LB boost (private + public)\n\nI have assigned 70% weight to the image classifier and 30% weight for the tabular classifier by easily summing the predictions up:\n`blended_results = img_clf_results * 0.7 + tabular_clf_results * 0.3`\n\nOverweighting the image classifier is natural because it does, when correctly trained, perform better, recognizes more patterns and therefore is more accurate then the tabular classifier in this competition which also has way less features to work with."
  }
}