{
  "id": 145105,
  "title": "Fast QWK Computation",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/145105",
  "author_name": "CPMP",
  "post_date": "2020-04-21T22:11:20.300000",
  "votes": 56,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Here is a fast implementation of the competition metric.  A comparison with other methods is available at <a href=\"https://www.kaggle.com/cpmpml/ultra-fast-qwk-calc-method\">https://www.kaggle.com/cpmpml/ultra-fast-qwk-calc-method</a></p>\n\n<p>If you use GPU, which is likely here, you also have a very efficient GPU accelerated code from @jiweiliu at  <a href=\"https://www.kaggle.com/jiweiliu/qwk-cupy-vs-numpy-vs-numba\">https://www.kaggle.com/jiweiliu/qwk-cupy-vs-numpy-vs-numba</a></p>\n\n<p>In the code below, <code>max_rat</code> is the maximal target value.</p>\n\n<p>```\nimport numpy as np\nfrom numba import jit </p>\n\n<p>@jit\ndef qwk3(a1, a2, max_rat):\n    assert(len(a1) == len(a2))\n    a1 = np.asarray(a1, dtype=int)\n    a2 = np.asarray(a2, dtype=int)</p>\n\n<pre><code>hist1 = np.zeros((max_rat + 1, ))\nhist2 = np.zeros((max_rat + 1, ))\n\no = 0\nfor k in range(a1.shape[0]):\n    i, j = a1[k], a2[k]\n    hist1[i] += 1\n    hist2[j] += 1\n    o +=  (i - j) * (i - j)\n\ne = 0\nfor i in range(max_rat + 1):\n    for j in range(max_rat + 1):\n        e += hist1[i] * hist2[j] * (i - j) * (i - j)\n\ne = e / a1.shape[0]\n\nreturn 1 - o / e\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": 815852,
      "postDate": "2020-04-21T22:11:20.300Z",
      "content": "<p>Here is a fast implementation of the competition metric.  A comparison with other methods is available at <a href=\"https://www.kaggle.com/cpmpml/ultra-fast-qwk-calc-method\">https://www.kaggle.com/cpmpml/ultra-fast-qwk-calc-method</a></p>\n\n<p>If you use GPU, which is likely here, you also have a very efficient GPU accelerated code from @jiweiliu at  <a href=\"https://www.kaggle.com/jiweiliu/qwk-cupy-vs-numpy-vs-numba\">https://www.kaggle.com/jiweiliu/qwk-cupy-vs-numpy-vs-numba</a></p>\n\n<p>In the code below, <code>max_rat</code> is the maximal target value.</p>\n\n<p>```\nimport numpy as np\nfrom numba import jit </p>\n\n<p>@jit\ndef qwk3(a1, a2, max_rat):\n    assert(len(a1) == len(a2))\n    a1 = np.asarray(a1, dtype=int)\n    a2 = np.asarray(a2, dtype=int)</p>\n\n<pre><code>hist1 = np.zeros((max_rat + 1, ))\nhist2 = np.zeros((max_rat + 1, ))\n\no = 0\nfor k in range(a1.shape[0]):\n    i, j = a1[k], a2[k]\n    hist1[i] += 1\n    hist2[j] += 1\n    o +=  (i - j) * (i - j)\n\ne = 0\nfor i in range(max_rat + 1):\n    for j in range(max_rat + 1):\n        e += hist1[i] * hist2[j] * (i - j) * (i - j)\n\ne = e / a1.shape[0]\n\nreturn 1 - o / e\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Here is a fast implementation of the competition metric.  A comparison with other methods is available at https://www.kaggle.com/cpmpml/ultra-fast-qwk-calc-method\n\nIf you use GPU, which is likely here, you also have a very efficient GPU accelerated code from @jiweiliu at  https://www.kaggle.com/jiweiliu/qwk-cupy-vs-numpy-vs-numba\n\nIn the code below, `max_rat` is the maximal target value.\n\n```\nimport numpy as np\nfrom numba import jit \n\n@jit\ndef qwk3(a1, a2, max_rat):\n    assert(len(a1) == len(a2))\n    a1 = np.asarray(a1, dtype=int)\n    a2 = np.asarray(a2, dtype=int)\n\n    hist1 = np.zeros((max_rat + 1, ))\n    hist2 = np.zeros((max_rat + 1, ))\n\n    o = 0\n    for k in range(a1.shape[0]):\n        i, j = a1[k], a2[k]\n        hist1[i] += 1\n        hist2[j] += 1\n        o +=  (i - j) * (i - j)\n\n    e = 0\n    for i in range(max_rat + 1):\n        for j in range(max_rat + 1):\n            e += hist1[i] * hist2[j] * (i - j) * (i - j)\n\n    e = e / a1.shape[0]\n\n    return 1 - o / e\n```",
      "votes": 55
    },
    {
      "id": 815924,
      "postDate": "2020-04-22T00:17:39.150Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 815861,
      "postDate": "2020-04-21T22:17:39.303Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 815924,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-22T00:17:39.150000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 815861,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-21T22:17:39.303000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "815852": "Here is a fast implementation of the competition metric.  A comparison with other methods is available at https://www.kaggle.com/cpmpml/ultra-fast-qwk-calc-method\n\nIf you use GPU, which is likely here, you also have a very efficient GPU accelerated code from @jiweiliu at  https://www.kaggle.com/jiweiliu/qwk-cupy-vs-numpy-vs-numba\n\nIn the code below, `max_rat` is the maximal target value.\n\n```\nimport numpy as np\nfrom numba import jit \n\n@jit\ndef qwk3(a1, a2, max_rat):\n    assert(len(a1) == len(a2))\n    a1 = np.asarray(a1, dtype=int)\n    a2 = np.asarray(a2, dtype=int)\n\n    hist1 = np.zeros((max_rat + 1, ))\n    hist2 = np.zeros((max_rat + 1, ))\n\n    o = 0\n    for k in range(a1.shape[0]):\n        i, j = a1[k], a2[k]\n        hist1[i] += 1\n        hist2[j] += 1\n        o +=  (i - j) * (i - j)\n\n    e = 0\n    for i in range(max_rat + 1):\n        for j in range(max_rat + 1):\n            e += hist1[i] * hist2[j] * (i - j) * (i - j)\n\n    e = e / a1.shape[0]\n\n    return 1 - o / e\n```",
    "815924": "",
    "815861": ""
  }
}