RELEASE VERIFIED · gitduck-bench-v1.1.0

28 attempts.
0 ducks.

25 geese · 2 chicks · 1 egg

One frozen Git decision. Every model chose an action. None provided exact provenance.

Different ecosystems. Same way to goose. Sometimes byte for byte.

GOOSE PENDING HATCH DUCK
28Attempts
0Ducks
25Geese
2Chicks
1Egg

🔍 6 exact-text collision events across 5 shared answer digests — jump to the convergence ledger.

Unofficial · derived from the published records

Distance to Duck

How many exact-set edits would turn each eligible answer into a duck? Dense ranks; chronology never breaks ties; byte-identical answers share a rank but stay separate attempts.

Rank Model Award Axes Δexact 🔑 Notes
1 Claude Sonnet 5 API 🦢 3/4 1 🔑 1 extra key
2 nemotron-3-nano:30b 🦢 3/4 2 1 extra key, 1 missing key
3 gpt-oss:120b 🦢 3/4 5 2 extra keys, 1 missing key, 2 fabricated refs · exact-byte twin (#5)
3 gpt-oss:20b 🦢 3/4 5 2 extra keys, 1 missing key, 2 fabricated refs
3 kimi-k3 🦢 3/4 5 2 extra keys, 1 missing key, 2 fabricated refs · exact-byte twin (#2)
3 minimax-m3 🦢 3/4 5 2 extra keys, 1 missing key, 2 fabricated refs
3 nemotron-3-ultra 🦢 3/4 5 2 extra keys, 1 missing key, 2 fabricated refs
3 GPT-5.6-sol API 🦢 3/4 5 2 extra keys, 1 missing key, 2 fabricated refs · exact-byte twin (#2)
3 Gemini 3.6 Flash API 🦢 3/4 5 2 extra keys, 1 missing key, 2 fabricated refs · exact-byte twin (#5)
3 Gemini 3.1 Pro Preview API 🦢 3/4 5 2 extra keys, 1 missing key, 2 fabricated refs · exact-byte twin (#6)
3 Gemini 3.5 Flash API 🦢 3/4 5 2 extra keys, 1 missing key, 2 fabricated refs · exact-byte twin (#6)
4 nemotron-3-super 🦢 3/4 6 2 extra keys, 1 missing key, 3 fabricated refs
4 Claude Opus 5 API 🦢 3/4 6 🔑 3 extra keys, 3 fabricated refs
5 kimi-k2.7-code 🦢 3/4 8 3 extra keys, 1 missing key, 4 fabricated refs
6 deepseek-v4-flash 🦢 3/4 10 4 extra keys, 1 missing key, 5 fabricated refs
6 deepseek-v4-flash:0731 🦢 3/4 10 4 extra keys, 1 missing key, 5 fabricated refs
6 deepseek-v4-pro 🦢 3/4 10 4 extra keys, 1 missing key, 5 fabricated refs
6 glm-5.1 🦢 3/4 10 4 extra keys, 1 missing key, 5 fabricated refs · exact-byte twin (#3)
6 glm-5.2 🦢 3/4 10 4 extra keys, 1 missing key, 5 fabricated refs · exact-byte twin (#4)
6 minimax-m2.7 🦢 3/4 10 🔑 5 extra keys, 5 fabricated refs
6 qwen3.5:397b 🦢 3/4 10 4 extra keys, 1 missing key, 5 fabricated refs
6 Claude Fable 5 API 🦢 3/4 10 🔑 5 extra keys, 5 fabricated refs
6 Claude Haiku 4.5 API 🦢 3/4 10 4 extra keys, 1 missing key, 5 fabricated refs · exact-byte twin (#3)
6 Grok 4.5 API 🦢 3/4 10 4 extra keys, 1 missing key, 5 fabricated refs · exact-byte twin (#4)
7 gemma4:31b 🐤 2/4 Wrong authorization
7 mistral-large-3:675b 🐤 2/4 Wrong authorization
8 kimi-k2.6 🥚 0/4 Schema failure

🦢 3/4 axes · 🐤 2/4 axes · 🥚 0/4 axes (never hatched) 🔑 = cited the load-bearing key (badge only, not a rank input) Δexact = edit-distance to an exact match; defined only when schema, authorization and selection all PASS

Full method → Machine-generated report →

Agent Product Lane

codex-cli / gpt-5.6-sol — scored separately. An agent product is never placed on the raw-model board.

Rank Model Award Axes Δexact 🔑 Notes
1 codex-cli 0.144.6 agent · gpt-5.6-sol 🦢 3/4 11 5 extra keys, 1 missing key, 5 fabricated refs

Diagnostic projection, not a ranking

Who actually found the load-bearing key?

Four attempts cited verification.PILOT_BUILD.outcome — the one key the correct action rests on. This table keeps that detection visible. It is not the leaderboard.

Model Found outcome Extra keys Fabricated refs Diagnosis
Claude Sonnet 5 1 0 Precise detection, minimal overreach
Claude Opus 5 3 3 Detection + overpacking
MiniMax M2.7 5 5 Provenance shotgun
Claude Fable 5 5 5 Provenance shotgun
Nemotron 3 Nano 30B (control) 1 0 Clean, but the wrong key

Finding the right key is not the same as proving the decision.

Case-induced exact convergence

Different providers. Identical bytes.

Ranked as ledger events, never as synthetic pairs. Provider-attributed identities, distinct transaction IDs, byte-identical answers.

Digest Provider A Provider B Bytes Class
3ec9f38e… Ollama kimi-k3 Ollama glm-5.2 369 Historical (unattested)
056ff186… GPT-5.6-sol API Ollama kimi-k3 255 Cross-provider
2efda97d… Claude Haiku 4.5 API Ollama glm-5.1 467 Cross-provider
3ec9f38e… Grok 4.5 API Ollama glm-5.2 369 Cross-provider recurrence
687a169a… Gemini 3.6 Flash API Ollama gpt-oss:120b 316 Cross-provider
6e5ca818… Gemini 3.5 Flash API Gemini 3.1 Pro Preview API 328 Intraprovider
6 collision events
5 shared answer digests
4 cross-provider events
1 intraprovider event
1 historical pair remains UNVERIFIED FOREVER

On one frozen case, each of the four major closed-provider API ecosystems produced at least one answer that was byte-identical to an answer previously produced by an Ollama-hosted arm. Exact bytes do not prove copying, distillation, caching, aliasing, or a shared physical backend. Event-level detail: exact-byte-convergence.md.

Reference

Award taxonomy

Descriptive labels for how an attempt fell short. Memorable, not the metric — DUCK 1/1 is the only official score.

🦆
DUCK
0%

All 4 axes PASS — exact and complete.

🦢
GOOSE
89%

3/4 axes — the right action found, provenance fumbled.

🐤
CHICK
7%

2/4 axes — structure and selection partly there.

🥚
EGG
4%

Schema failure, or nothing scorable was produced.

The goose is drawn as a swan (🦢) — the goose glyph U+1FABF is Unicode 15 and renders as a box on Windows 10 and other frozen emoji fonts. The award is still named GOOSE everywhere data speaks. Full taxonomy: award-taxonomy.md.

Latest release: v1.1.0

Immutable evidence in v1.0.0. Reproducible derived projections above.

$ git clone https://github.com/atlas-ocm/gitduck-bench
$ cd gitduck-bench && git checkout gitduck-bench-v1.1.0
$ node tools/verify-manifest.mjs
MANIFEST = VERIFIED