{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Check a notebook before you spend GPU on it\n\nForking a strong public notebook is a normal and legitimate way to work on a competition. The\ntrap is that **a notebook being public does not mean it can run for you.**\n\nA notebook can be public, well documented, heavily upvoted, and still depend on an artifact that\nonly its author can mount. If that happens you find out after the queue, after the mount pull, and\nsometimes after hours of accelerator time - for a run that was never going to produce a submission.\n\nThis notebook is a short pre-flight check. It costs seconds, runs on CPU, and answers one question:\n**does everything this notebook needs actually exist for me?**\n","id":"c00"},{"cell_type":"markdown","metadata":{},"source":"## Signal 1 - an empty string in the mount list\n\nKernel metadata lists the datasets a notebook mounts. Pull it with the Kaggle API and you get\nsomething like:\n\n```json\n\"dataset_sources\": [\"\", \"\", \"owner/pack-a\", \"owner/pack-b\"]\n```\n\nThose empty strings are not cosmetic and they are not a bug in your pull. **A dataset you cannot\nsee does not serialise to a name.** Private datasets, and datasets since deleted or made private,\ncome back as empty strings for everyone except the owner.\n\nSo an empty entry means: *this notebook mounts something you do not have access to.* You cannot\nsupply it, you cannot substitute it, and no amount of re-pushing will resolve it.\n\n**This is the cheapest check in competitive machine learning.** One metadata pull, before you build\nanything.\n","id":"c01"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def check_mounts(metadata: dict) -> dict:\n    \"Pre-flight a kernel's metadata. Returns a verdict dict; never raises.\"\n    sources = metadata.get(\"dataset_sources\", []) or []\n    named = [s for s in sources if str(s).strip()]\n    blank = [s for s in sources if not str(s).strip()]\n    verdict = {\n        \"declared_mounts\": len(sources),\n        \"named\": named,\n        \"unresolvable\": len(blank),\n        \"runnable_by_you\": len(blank) == 0,\n    }\n    return verdict\n\n\n# A metadata block shaped like one that cannot run for you.\nexample_blocked = {\n    \"dataset_sources\": [\"\", \"\", \"someowner/public-pack-a\", \"someowner/public-pack-b\"],\n}\n# ... and one that can.\nexample_clear = {\n    \"dataset_sources\": [\"someowner/public-pack-a\", \"someowner/public-pack-b\"],\n}\n\nfor label, meta in ((\"blocked\", example_blocked), (\"clear\", example_clear)):\n    v = check_mounts(meta)\n    print(f\"{label:>8}: declares {v['declared_mounts']} mounts, \"\n          f\"{v['unresolvable']} unresolvable -> runnable_by_you={v['runnable_by_you']}\")\nprint()\nprint(\"If unresolvable > 0, stop. Building the fork cannot fix it.\")\n","id":"c02"},{"cell_type":"markdown","metadata":{},"source":"## Signal 2 - hash-pinned weight loaders\n\nThe second signal is in the code rather than the metadata. Some pipelines authenticate their\nweights: they walk the input tree looking for a filename, then require its SHA256 to equal a\nconstant baked into the source.\n\n```python\nif sha256(path) == expected:\n    candidates.append(path)\nif len(set(candidates)) != 1:\n    raise RuntimeError(f\"expected one authenticated {name}; matches={sorted(set(candidates))}\")\n```\n\nThis is good engineering - it fails fast and loudly instead of silently loading the wrong tensor.\nFor a forker it is also free information:\n\n- The error prints **the exact filename and hash** it wanted.\n- `matches=[]` means the file is not in your input tree at all.\n- Because the check is on the **hash**, a similarly named public file will not satisfy it. You\n  cannot substitute a lookalike.\n\nSo if you grep a notebook for `sha256` and find a table of expected digests, list those filenames\nand confirm every one of them exists in a dataset you can actually mount.\n","id":"c03"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import re\n\ndef scan_source_for_pinned_files(src: str) -> list:\n    \"Find filenames that a notebook authenticates by hash. Heuristic, deliberately loose.\"\n    hits = []\n    for m in re.finditer(r'[\"\\']([\\w\\-.]+\\.(?:pt|pth|bin|safetensors|json))[\"\\']\\s*:\\s*[\"\\']([0-9a-f]{64})[\"\\']', src):\n        hits.append((m.group(1), m.group(2)[:12] + \"...\"))\n    return hits\n\n\ndemo = '''\nEXPECTED_WEIGHTS = {\n    \"model_axial_state.pt\": \"0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef\",\n    \"model_fusion_head.pt\": \"fedcba9876543210fedcba9876543210fedcba9876543210fedcba9876543210\",\n}\n'''\nfound = scan_source_for_pinned_files(demo)\nprint(f\"hash-pinned artifacts found: {len(found)}\")\nfor name, digest in found:\n    print(f\"  {name:<28} sha256 {digest}\")\nprint()\nprint(\"Each of these must exist, byte-for-byte, in a dataset you can mount.\")\nprint(\"A file with the same name but a different hash will be rejected.\")\n","id":"c04"},{"cell_type":"markdown","metadata":{},"source":"## The pre-flight, in order\n\n1. **Pull the metadata**, not the notebook. `kaggle kernels pull <ref> -m` is seconds and costs\n   nothing.\n2. **Count empty `dataset_sources` entries.** Any at all means a dependency you cannot access.\n   Stop here.\n3. **Grep the source for a table of SHA256 digests.** If one exists, list the filenames.\n4. **Confirm each filename is inside a dataset you can list.** `kaggle datasets files <ref>` will\n   tell you. Matching names is not enough when the loader checks hashes, but a *missing* name is\n   already decisive.\n5. **Only then** build the fork.\n\nSteps 1 to 4 take about a minute. Step 5 can take hours of accelerator time, and on a competition\nwith a daily submission cap, a failed run can also cost you a slot if you submit before reading the\nlog.\n","id":"c05"},{"cell_type":"markdown","metadata":{},"source":"## Limits, and one thing this is not\n\n- An empty mount entry is strong evidence of an inaccessible dependency, but the API is the only\n  thing being read here. If a notebook resolves an artifact at runtime by some other route, this\n  check will not see it.\n- A clean pre-flight does not mean a fork will reproduce the author's score. It only means the run\n  can start. Configuration, environment and versioning are all still ahead of you.\n- **This is not a complaint about anyone.** Keeping fine-tuned weights private is entirely within\n  the rules and is a normal competitive choice. No notebook, dataset or competitor is named\n  anywhere in this notebook, and none of this implies anything improper. The point is narrow and\n  practical: *check what you can run before you spend time running it.*\n","id":"c06"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"}},"nbformat":4,"nbformat_minor":5}