The longest task slug in the skillsbench@1.1 registry — 42 characters, three hyphens, full ascender/descender range. If a candidate is going to fringe against Satoshi anywhere, it is here. Amber is Satoshi, blue is the candidate; where the two agree exactly they cancel toward a neutral tone.
SkillsBench — leaderboard
Harness × model matrix, the site's real top rows. The stress case: four right-aligned numeric columns that must hold still while an agent name goes bold on hover.
| Model | Harness | Without | With Skills | Δ | Gain (g) |
|---|---|---|---|---|---|
| GPT-5.5 | OpenHands | 51.5 | 67.3 | 15.8 | 32.6 |
| GPT-5.5 | Codex | 46.8 | 66.5 | 19.7 | 37.0 |
| Opus 4.7 | Claude Code | 43.0 | 61.2 | 18.2 | 31.9 |
| Gemini 3.1 Pro | Gemini CLI | 36.0 | 60.8 | 24.8 | 38.7 |
| GLM 5.1 | OpenHands | 32.7 | 58.4 | 25.7 | 38.1 |
| HY3 skills-only | Claude Code | — | 55.9 | — | — |
| Gemini 3 Flash | Gemini CLI | 34.2 | 54.6 | 20.4 | 31.0 |
| Opus 4.8 | OpenHands | 45.7 | 54.1 | 8.4 | 15.5 |
| Kimi K2.6 | OpenHands | 33.4 | 54.0 | 20.6 | 31.0 |
Capability viewer — run detail
List plus detail — the hardest pairing here: dense rows on the left against running prose and a grader checklist on the right.
| Task | Category | Difficulty |
|---|---|---|
| citation-check | office-white-collar | medium |
| crystallographic-wyckoff-position-analysis | natural-science | medium |
| fix-druid-loophole-cve | cybersecurity | hard |
| energy-ac-optimal-power-flow | industrial-physical-systems | medium |
| debug-trl-grpo | software-engineering | hard |
You are helping a research team verify the integrity of their bibliography before submitting a paper. The team suspects that some citations in their BibTeX file may be fake or hallucinated.
TaskMiner — mining queue
Queue view: two-line rows and a status word that has to read at a glance, without leaning on colour alone.