Roman shipped a skill that turns the GVS5H discovery into a one-word command. If you have not read the paper: five orchestrated open-weight Qwen3.8-27B instances matched and slightly beat Claude Fable 5 on the LiveCodeBench-Hard 100-problem set — 92.4% vs 90.4% pass@1 — without a single line of extra training. The trick was not a bigger model; it was fresh contexts per role, coordinating through a shared filesystem ledger. zcode-smart-skill packages exactly that protocol into a /smart command with auto-triggering, so any ZCode user gets the orchestration without reimplementing the paper.
The GVS5H project (Persis Capital, ICLR 2027 paper) is a ledger-based zero-shot self-orchestration method. It is training-free: fresh instances of the same model decompose a problem and coordinate through a shared filesystem that holds a plan, notes and the current solution. The headline result from the abstract:
locally served, open-weight Qwen3.8-27B rises from 69.2% to 92.4%, slightly exceeding Fable 5 — and orchestrated GPT-5.6-Terra reaches 88.0% against Fable 5's 90.4% at 19% of the cost.
Two things make this notable. First, the gain is architectural, not scale: the same weights, reorganized at inference time, jumped 23 percentage points. Second, the authors are honest that gains are not universal — some models were unchanged or worse. Orchestration is a discipline that works reliably on some backends, not a magic multiplier for every model.
The paper's transcript analysis attributes the improvement to decomposition and persistent context. A single long-context attempt on a hard problem degrades for three compounding reasons:
GVS5H replaces all three with structure: fresh contexts per role eliminate pollution, a notes file that prunes superseded findings prevents the ledger itself from bloating, forced ideation of distinct approaches escapes local optima, and a mandatory verify step makes running the code the only ground truth.
zcode-smart-skill is a single SKILL.md (plain markdown, MIT licensed, no code to trust) plus an optional AGENTS-snippet.md for auto-triggering. When you type /smart — or the snippet fires it — the agent announces "Entering smart mode (ledger orchestration)" and runs this loop:
notes.md, pruning anything superseded.State lives only in a ledger directory: .smart/<hash>/ with task.md, plan.md, notes.md (hard cap 800 words), tasks.json (max 12 tasks), solution.* and verify.log.
The loop is capped and self-correcting, which is what separates it from a naive agent swarm:
done requires a non-empty artifact and green verification. No exceptions.Install is two commands, and the skill works in any agent harness that has skills, fresh-context subagents and file IO (tested on ZCode; adapts trivially to Claude Code, Codex CLI, OpenCode):
git clone https://github.com/romangalaxys10-spec/zcode-smart-skill.git
mkdir -p ~/.agents/skills
cp -R zcode-smart-skill/SKILL.md ~/.agents/skills/smart/SKILL.md
Restart the agent once and /smart is callable. For auto-triggering, append AGENTS-snippet.md to AGENTS.md (or CLAUDE.md); it fires the skill on hard algorithm work, concurrency design, multi-file refactors with subtle invariants, competitive-programming-style tasks, a fix that already failed twice, and the words "smart", "hard" or "properly".
⚡ Cost note: smart mode spawns subagents, so it burns more tokens than a plain answer — by design. The skill says so up front and keeps the ledger lean so the overhead is bounded, not explosive.
If you live in a coding agent daily, yes — and the cost is a single markdown file. The honest caveats: the GVS5H numbers come from pinned backends and the 100 latest hard LiveCodeBench problems, so your mileage depends on the model behind ZCode; and orchestration is a thinking-discipline, so it shines on problems where single-context attempts genuinely degrade — gnarly debugging, concurrency, architecture with subtle invariants. On trivial tasks it is overkill, and the skill itself skips IDEATE when the scope is trivially small. As a pattern, this is the most interesting open technique of the year so far: evidence that engineering can substitute for scale, at least on the problems hardest for current models.
All figures above were verified against the GVS5H repository README and paper abstract (slee-persis/GVS5H, ICLR 2027), and the zcode-smart-skill README and SKILL.md (romangalaxys10-spec/zcode-smart-skill).
Source: GVS5H on GitHub · zcode-smart-skill on GitHub