{"cells":[{"cell_type":"markdown","metadata":{},"source":"# A title score costs RSNA 32% of its views\n\nA well-known Kaggle growth-hack is: put a real, checkable number in your notebook's title (a\nscore, a rank, a metric) and it gets more views. A 1,672-notebook census backs this up in\naggregate -- title-carries-a-score notebooks get 1.48x the pooled-average views.\n\nThat number is pooled across every competition on the platform. Nobody seems to have checked\nwhether it holds inside any one competition. This notebook checks it across five specific\ncompetitions, and the answer is: no. In one of them, putting a score in the title is associated\nwith **32% fewer** views, not more."},{"cell_type":"markdown","metadata":{},"source":"## Load every kernel currently attributed to five competitions"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import glob, time\nt0 = time.time()\n\ndef find(pattern):\n    hits = sorted(glob.glob(f\"/kaggle/input/**/{pattern}\", recursive=True))\n    assert hits, f\"{pattern} not found under /kaggle/input -- is kaggle/meta-kaggle mounted?\"\n    return hits[0]\n\nKERNELS_PATH = find(\"Kernels.csv\")\nKV_PATH = find(\"KernelVersions.csv\")\nKVCS_PATH = find(\"KernelVersionCompetitionSources.csv\")\nCOMPS_PATH = find(\"Competitions.csv\")\nprint(f\"reading {KERNELS_PATH}\\n reading {KV_PATH}\\n reading {KVCS_PATH}\\n reading {COMPS_PATH}\")"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import pandas as pd\nimport numpy as np\nimport re\n\n# Kernels.csv does not carry a Title column -- titles are per-version, in KernelVersions.csv.\nkernels = pd.read_csv(KERNELS_PATH, usecols=[\"Id\", \"CurrentKernelVersionId\", \"TotalViews\", \"TotalVotes\"])\nkernel_versions = pd.read_csv(KV_PATH, usecols=[\"Id\", \"Title\"])\nkvcs = pd.read_csv(KVCS_PATH, usecols=[\"KernelVersionId\", \"SourceCompetitionId\"])\ncomps = pd.read_csv(COMPS_PATH, usecols=[\"Id\", \"Slug\", \"Title\", \"TotalTeams\"])\n\nkernels = kernels.merge(kernel_versions, left_on=\"CurrentKernelVersionId\", right_on=\"Id\",\n                         how=\"left\", suffixes=(\"\", \"_ver\"))\nmerged = kernels.merge(kvcs, left_on=\"CurrentKernelVersionId\", right_on=\"KernelVersionId\", how=\"inner\")\nmerged = merged.merge(comps, left_on=\"SourceCompetitionId\", right_on=\"Id\", how=\"left\", suffixes=(\"\", \"_comp\"))\nprint(f\"{len(merged):,} competition-attributed kernels, {merged['SourceCompetitionId'].nunique():,} \"\n      f\"distinct competitions [{time.time()-t0:.0f}s]\")"},{"cell_type":"markdown","metadata":{},"source":"## Classify \"score in the title\"\n\nReusing the exact regex this account already uses elsewhere to make this call\n(`0\\.\\d{3,}|lb[ _-]?0`, case-insensitive) rather than inventing a friendlier definition after\nseeing the result."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"SCORE_RE = re.compile(r\"0\\.\\d{3,}|lb[ _-]?0\", flags=re.IGNORECASE)\nmerged[\"score_in_title\"] = merged[\"Title\"].fillna(\"\").str.contains(SCORE_RE)\nprint(merged[\"score_in_title\"].value_counts())"},{"cell_type":"markdown","metadata":{},"source":"## Five venues, one question: does the score-in-title multiplier hold across all of them?"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"MY_VENUES = {\n    \"playground-series-s6e8\": \"S6E8\",\n    \"kaggriculture\": \"Kaggriculture\",\n    \"biohub-cell-tracking-during-development\": \"Biohub\",\n    \"rsna-knee-abnormality-detection\": \"RSNA\",\n    \"pokemon-tcg-ai-battle-challenge-strategy\": \"Pokemon\",\n}\n\nrows = []\nfor slug, label in MY_VENUES.items():\n    sub = merged.loc[merged[\"Slug\"] == slug]\n    with_score = sub.loc[sub[\"score_in_title\"], \"TotalViews\"].dropna()\n    without_score = sub.loc[~sub[\"score_in_title\"], \"TotalViews\"].dropna()\n    mv_s = with_score.mean() if len(with_score) else np.nan\n    mv_n = without_score.mean() if len(without_score) else np.nan\n    mult = (mv_s / mv_n) if (mv_n and mv_n > 0 and not np.isnan(mv_s)) else np.nan\n    rows.append({\n        \"venue\": label, \"n_total\": len(sub),\n        \"n_score_in_title\": len(with_score), \"n_no_score\": len(without_score),\n        \"mean_views_score\": round(float(mv_s), 1) if not np.isnan(mv_s) else None,\n        \"mean_views_no_score\": round(float(mv_n), 1) if not np.isnan(mv_n) else None,\n        \"multiplier\": round(float(mult), 2) if mult == mult else None,\n    })\n\nresults = pd.DataFrame(rows)\nprint(results.to_string(index=False))\nresults.to_csv(\"venue_score_title_multipliers.csv\", index=False)\nprint(\"\\nwrote venue_score_title_multipliers.csv\")"},{"cell_type":"markdown","metadata":{},"source":"## What this actually says\n\n| venue | n with score-title | multiplier | reads as |\n|---|---:|---:|---|\n| S6E8 | 48 | **2.85x** | the trick works, hard |\n| Biohub | 25 | **1.70x** | the trick works |\n| RSNA | 10 | **0.68x** | the trick backfires -- 32% *fewer* views |\n| Kaggriculture | 0 | not measured | nobody in this snapshot has tried it |\n| Pokemon | 0 | not measured | nobody in this snapshot has tried it |\n\nThree things, stated plainly rather than smoothed over:\n\n**1. The multiplier is not one number.** Among the three venues where it can even be measured\n(n >= 5 on both sides), it ranges from 0.68x to 2.85x -- a **4.2x** spread between the best and\nworst venue for the exact same title genre. A rule learned from a 1,672-notebook pool does not\ntransfer venue-by-venue at anywhere near that pooled multiplier.\n\n**2. RSNA is a reversal, not just a smaller win.** 0.68x means the average score-titled notebook\nin RSNA gets fewer views than the average non-score-titled one -- the opposite sign from the\npooled census finding. This account has already published an RSNA notebook using exactly this\ntitle genre without ever checking whether the genre helps there. Whether that specific choice hurt\nits reach can't be said from a group average, but the assumption behind making it (that the\n1.48x pooled lift applies inside RSNA) is now falsified.\n\n**3. \"Not measured\" is not the same as \"no effect.\"** Kaggriculture and Pokemon show zero\nscore-titled notebooks in this snapshot -- the genre isn't failing there, it simply hasn't been\ntried by anyone yet. Anyone publishing the first one is running an uncontrolled experiment, not\nfollowing a validated rule.\n\n**Caveats that matter:** this is an observational split, not a randomized one -- a score-titled\nRSNA notebook might differ from a non-score-titled one in ways other than its title (author\nhistory, publish timing, notebook type). n=10 for RSNA's score-in-title cell is thin; it clears a\n\"at least 5 on both sides\" bar set before this was run, but thin is still thin, and a single\noutlier notebook could move it. Treat the direction as a real correction to a previously untested\nassumption, not as a precise causal estimate."},{"cell_type":"markdown","metadata":{},"source":"## Check your own venue\n\nSwap `MY_VENUE` for any competition slug and re-run this cell -- Copy & Edit this notebook to get\nyour own venue's score-in-title multiplier before you assume the pooled 1.48x number applies to\nyou."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"MY_VENUE = \"biohub-cell-tracking-during-development\"   # <-- swap for any competition slug\n\nsub = merged.loc[merged[\"Slug\"] == MY_VENUE]\nif len(sub) == 0:\n    print(f\"{MY_VENUE}: no competition-attributed kernels found in this snapshot\")\nelse:\n    with_score = sub.loc[sub[\"score_in_title\"], \"TotalViews\"].dropna()\n    without_score = sub.loc[~sub[\"score_in_title\"], \"TotalViews\"].dropna()\n    print(f\"{MY_VENUE}: n_score={len(with_score)} n_no_score={len(without_score)}\")\n    if len(with_score) >= 5 and len(without_score) >= 5:\n        print(f\"  multiplier = {with_score.mean() / without_score.mean():.2f}x\")\n    else:\n        print(\"  fewer than 5 notebooks on one side -- not enough to call a multiplier yet\")"},{"cell_type":"markdown","metadata":{},"source":"## The honest bottom line\n\n\"Put a score in your title\" is real advice with real pooled evidence behind it, and it is still\nworth doing in most places. But it is a venue-conditional lever, not a universal one, and this\naccount had been treating it as universal. Check your own competition's split before writing the\ntitle, the same way this notebook just did for five of them."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.0"}},"nbformat":4,"nbformat_minor":5}