Fabervant Research · Technical report

Six Language Models Authoring Logic Puzzles Through a Game's Production Gates

Passes, hardness, time and a blind read over ten requests per model

These puzzles are live: play GRIDIGMA free →

Version 1.0 · Published 29 September 2026 · First published on 27 September 2026 as Section 10 of the nine-model prose paper [1]

Abstract

Background. A logic-puzzle game needs puzzles whose clues admit exactly one solution at the requested difficulty, and a story that sets them up. We asked which current language models author such puzzles best when every draft must pass the game's own production gates, and what they spend doing it.

Methods. Six models from Anthropic, OpenAI and xAI each authored the logic, clues and story of the same 10 puzzle requests (two themed packs, five difficulty tiers each) through GRIDIGMA's production authoring loop and its gates, at reasoning effort high and with up to 3 attempts per request. In an order fixed before the run, we measured passes through the gates, the hardness of the passing puzzles and model time; a blind read of the 46 passing puzzles by one AI agent is reported separately.

Results. No model won every measure: Grok 4.7 passed the most requests (9 of 10) but was the slowest (27.7 model minutes per pass); Opus 5.5 and Fable 5.1 passed 8 each and read best in the blind read (8.13 and 8.00 of 10); GPT-6 Sol was the fastest (14.1 minutes per pass) and wrote the least demanding logic. Every model passed both very easy requests; failures gathered at the harder tiers, and in 10 of the 14 failures the best draft was sound and missed only the requested tier.

Conclusions. Ten requests per model give a direction, not a choice. The commonest defect, clues built on a sum, came from the instructions rather than from any one model, and the blind read rests on a single AI reader of the same make as two of the authors; players have not rated these puzzles.

Keywords: large language models; procedural content generation; logic puzzles; puzzle difficulty; game design; LLM-as-a-judge

1. Introduction

GRIDIGMA is a logic-grid deduction game played mostly on phones, published in English and Turkish. Its puzzles are authored by a loop in which a language model drafts the logic, the clues and the story, and the game's gates check the draft before it can be published. A companion paper [1] asked which language models best write and translate the text of a finished puzzle. This paper asks the step before it: which model best authors the puzzle itself, a grid whose clues admit exactly one solution at the requested difficulty, and the story that sets them up.

Authoring is judged differently from prose. A draft passes the gates or it does not; a passing puzzle can be easy or demanding; and the reasoning it takes varies widely between models. We asked three questions, in an order fixed with the product owner before the run:

  1. RQ1. Which models best author a puzzle's logic, passing the game's own gates at the requested difficulty?
  2. RQ2. How hard are the puzzles that pass?
  3. RQ3. What does each model spend on them, in model time, calls and tokens?

A blind read of theme, story, clues and appeal is reported alongside and enters none of the three. Section 2 sets out the game, the requests, the loop and the measures; Section 3 gives the results; Sections 4, 5 and 6 discuss them, state their limits and conclude.

2. Method

2.1 The game and its puzzles

GRIDIGMA is a logic-grid deduction game played mostly on phones, published in English and Turkish. Each puzzle has a title, a one-line description, a short story and a set of clues, and the clues together admit exactly one solution. The game sorts its puzzles into five difficulty tiers, from very easy to very hard, and also publishes them in themed packs. New puzzles reach the game through an authoring loop in which a model drafts it and the game's gates check the draft (Section 2.3). The companion paper [1] compared language models as writers and translators of the text of finished puzzles; this study needs nothing from it beyond what is set out below.

2.2 Design

Six models took part (Table 1): Grok 4.6, Grok 4.7, GPT-6 Sol, GPT-6 Astra, Fable 5.1 and Opus 5.5. The requests were July's two themed packs, one on the Wimbledon tennis championship and one on World Heritage sites, one puzzle per difficulty tier per pack (very easy, easy, medium, hard, very hard): ten requests, identical for every model. Each request fixed the setting, a cast of four (English names in the Wimbledon pack, Korean names in the World Heritage pack, the same in every language), the labels of the grid and the tier; the model wrote the logic, the clues and the story. The instructions asked for plain everyday words, no clue told as arithmetic, no real person, organisation or named place, and encouraged the clue types the game's catalogue uses least (Section 5).

Table 1. The six models, the model identifier every call requested (the identifier each answer reported was the same on all 138 answered calls) and the route that reached each.
ModelMakerModel IDRoute
Grok 4.6xAIgrok-4.6xAI interface
Grok 4.7xAIgrok-4.7xAI interface
GPT-6 SolOpenAIgpt-6-solOpenAI's Codex client, Plus subscription
GPT-6 AstraOpenAIgpt-6-astraOpenAI's Codex client, Plus subscription
Fable 5.1Anthropicclaude-fable-5-1Anthropic's Claude command-line client, subscription
Opus 5.5Anthropicclaude-opus-5-5Anthropic's Claude command-line client, subscription

2.3 The production loop

Every request ran through GRIDIGMA's own authoring loop, the one that makes the game's puzzles, with one route added to reach each model. Every call received exactly the prompt the production client builds: the tier brief, the state of the catalogue, the gate contract, the craft guide, the clue-type ceilings, the setting and the place. A draft that failed a gate went back to the same model with the gates' findings, up to three attempts per request (a first draft and two revisions), with no escalation to another model. The gates were those every published puzzle passes: exactly one solution, no clue that can be removed without losing it, the clue-quality checks, and a computed difficulty tier equal to the requested one. The tier is computed from a simulated player's solving time per unit of deduction, so a model cannot reach it by counting clues.

2.4 Settings and interruptions

Settings. Every call ran at reasoning effort high, set and recorded per call with the requested and the reported model identifier, the time and the token usage; the two identifiers agreed on all 138 answered calls. Each call had a 30-minute cap. The Grok models ran through the xAI interface, the GPT-6 models through OpenAI's Codex client on a Plus subscription, and Fable 5.1 and Opus 5.5 through Anthropic's Claude command-line client on a subscription.

Interruptions. The subscriptions' usage limits paused GPT-6 Sol and GPT-6 Astra (Codex) and Fable 5.1 and Opus 5.5 (Claude). The caller waited for each reset; a request whose calls had failed was restarted from scratch (GPT-6 Astra 1, Fable 5.1 7, Opus 5.5 7 requests), and the failed runs' records are kept. Model minutes are the time from sending a call to its answer, summed over the answered calls, with every limit wait removed; failed calls are not counted (one Fable 5.1 call reached the 30-minute cap without an answer). The run went from 11:19 UTC on 25 September to 03:50 UTC on 26 September 2026.

2.5 Measures

The order of the measures was fixed with the product owner before the run: first the gates (RQ1), then how hard the passing puzzles are (RQ2), then time and cost (RQ3). A blind read of theme and story is reported separately (Section 3.4) and enters none of them.

  • Gates. A request passes when a draft clears every gate within the attempts allowed; each model's passes are counted over the ten requests.
  • Hardness. For each pass, the game engine's solve-path walker computes readableShare, the share of the board a solver can fill by reading the clues one at a time (lower = harder; Section 3.2), and speculativeWork, the depth of trial reasoning the solve needed.
  • Time and cost. Model minutes (Section 2.4), answered calls and output tokens, per request and per pass; money only where a route reports it.

Each model ran each request once. The results are counts and means over those runs; the uncertainty of a pass rate is given as a Wilson interval [2] in Section 5.

3. Results

3.1 Gates (RQ1)

Figure 1. The outcome of every request (rows: the two packs' five tiers) for every model (columns, ordered by passes, then by readableShare). A filled cell is a pass through all the production gates, labelled with its readableShare: the share of the board a solver can fill by reading the clues one at a time (lower = harder; darker = harder). A dashed cell is a fail after 3 attempts; a corner mark shows that the best draft of the run failed only the difficulty-tier gate. The right column counts the models that passed each request. One generation per cell. Source: GRIDIGMA out/s3/ (pass, fail and run records; readableShare from the engine's solve-path walker).
Grok 4.7 Fable 5.1 Grok 4.6 Opus 5.5 GPT-6 Sol GPT-6 Astra Models passing Wimbledon pack very easy 1.00 Grok 4.7, Wimbledon, very easy: pass after 1 revision; readableShare 1.00; computed tier very easy 0.69 Fable 5.1, Wimbledon, very easy: pass after 2 revisions; readableShare 0.69; computed tier very easy 0.83 Grok 4.6, Wimbledon, very easy: pass after 0 revisions; readableShare 0.83; computed tier very easy 1.00 Opus 5.5, Wimbledon, very easy: pass after 0 revisions; readableShare 1.00; computed tier very easy 1.00 GPT-6 Sol, Wimbledon, very easy: pass after 1 revision; readableShare 1.00; computed tier very easy 1.00 GPT-6 Astra, Wimbledon, very easy: pass after 1 revision; readableShare 1.00; computed tier very easy 6 of 6 easy 0.71 Grok 4.7, Wimbledon, easy: pass after 0 revisions; readableShare 0.71; computed tier easy 1.00 Fable 5.1, Wimbledon, easy: pass after 1 revision; readableShare 1.00; computed tier easy 0.88 Grok 4.6, Wimbledon, easy: pass after 1 revision; readableShare 0.88; computed tier easy 0.67 Opus 5.5, Wimbledon, easy: pass after 1 revision; readableShare 0.67; computed tier easy 1.00 GPT-6 Sol, Wimbledon, easy: pass after 1 revision; readableShare 1.00; computed tier easy fail GPT-6 Astra, Wimbledon, easy: fail after 3 attempts; the best draft failed the gate name diversity; 0 sound off-tier drafts kept 5 of 6 medium 0.88 Grok 4.7, Wimbledon, medium: pass after 1 revision; readableShare 0.88; computed tier medium 0.25 Fable 5.1, Wimbledon, medium: pass after 0 revisions; readableShare 0.25; computed tier medium 0.25 Grok 4.6, Wimbledon, medium: pass after 1 revision; readableShare 0.25; computed tier medium 0.75 Opus 5.5, Wimbledon, medium: pass after 2 revisions; readableShare 0.75; computed tier medium fail GPT-6 Sol, Wimbledon, medium: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 2 sound off-tier drafts kept fail GPT-6 Astra, Wimbledon, medium: fail after 3 attempts; the best draft failed the gates difficulty target and name diversity; 0 sound off-tier drafts kept 4 of 6 hard fail Grok 4.7, Wimbledon, hard: fail after 3 attempts; the best draft failed the gate clue quality; 0 sound off-tier drafts kept fail Fable 5.1, Wimbledon, hard: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 2 sound off-tier drafts kept 0.20 Grok 4.6, Wimbledon, hard: pass after 2 revisions; readableShare 0.20; computed tier hard 0.19 Opus 5.5, Wimbledon, hard: pass after 1 revision; readableShare 0.19; computed tier hard 0.16 GPT-6 Sol, Wimbledon, hard: pass after 2 revisions; readableShare 0.16; computed tier hard fail GPT-6 Astra, Wimbledon, hard: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 2 sound off-tier drafts kept 3 of 6 very hard 0.06 Grok 4.7, Wimbledon, very hard: pass after 0 revisions; readableShare 0.06; computed tier very hard 0.05 Fable 5.1, Wimbledon, very hard: pass after 2 revisions; readableShare 0.05; computed tier very hard 0.04 Grok 4.6, Wimbledon, very hard: pass after 1 revision; readableShare 0.04; computed tier very hard fail Opus 5.5, Wimbledon, very hard: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 3 sound off-tier drafts kept 0.13 GPT-6 Sol, Wimbledon, very hard: pass after 2 revisions; readableShare 0.13; computed tier very hard 0.02 GPT-6 Astra, Wimbledon, very hard: pass after 2 revisions; readableShare 0.02; computed tier very hard 5 of 6 World Heritage pack very easy 1.00 Grok 4.7, World Heritage, very easy: pass after 0 revisions; readableShare 1.00; computed tier very easy 0.69 Fable 5.1, World Heritage, very easy: pass after 0 revisions; readableShare 0.69; computed tier very easy 0.83 Grok 4.6, World Heritage, very easy: pass after 0 revisions; readableShare 0.83; computed tier very easy 1.00 Opus 5.5, World Heritage, very easy: pass after 0 revisions; readableShare 1.00; computed tier very easy 1.00 GPT-6 Sol, World Heritage, very easy: pass after 0 revisions; readableShare 1.00; computed tier very easy 1.00 GPT-6 Astra, World Heritage, very easy: pass after 1 revision; readableShare 1.00; computed tier very easy 6 of 6 easy 0.88 Grok 4.7, World Heritage, easy: pass after 0 revisions; readableShare 0.88; computed tier easy 1.00 Fable 5.1, World Heritage, easy: pass after 2 revisions; readableShare 1.00; computed tier easy 1.00 Grok 4.6, World Heritage, easy: pass after 1 revision; readableShare 1.00; computed tier easy 0.88 Opus 5.5, World Heritage, easy: pass after 1 revision; readableShare 0.88; computed tier easy 1.00 GPT-6 Sol, World Heritage, easy: pass after 1 revision; readableShare 1.00; computed tier easy 1.00 GPT-6 Astra, World Heritage, easy: pass after 1 revision; readableShare 1.00; computed tier easy 6 of 6 medium 0.32 Grok 4.7, World Heritage, medium: pass after 1 revision; readableShare 0.32; computed tier medium fail Fable 5.1, World Heritage, medium: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 3 sound off-tier drafts kept 0.31 Grok 4.6, World Heritage, medium: pass after 0 revisions; readableShare 0.31; computed tier medium 0.20 Opus 5.5, World Heritage, medium: pass after 1 revision; readableShare 0.20; computed tier medium 0.59 GPT-6 Sol, World Heritage, medium: pass after 2 revisions; readableShare 0.59; computed tier medium 0.29 GPT-6 Astra, World Heritage, medium: pass after 1 revision; readableShare 0.29; computed tier medium 5 of 6 hard 0.14 Grok 4.7, World Heritage, hard: pass after 0 revisions; readableShare 0.14; computed tier hard 0.24 Fable 5.1, World Heritage, hard: pass after 2 revisions; readableShare 0.24; computed tier hard fail Grok 4.6, World Heritage, hard: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 1 sound off-tier draft kept fail Opus 5.5, World Heritage, hard: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 3 sound off-tier drafts kept fail GPT-6 Sol, World Heritage, hard: fail after 3 attempts; the best draft failed the gates difficulty target and locale structure; 0 sound off-tier drafts kept fail GPT-6 Astra, World Heritage, hard: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 2 sound off-tier drafts kept 2 of 6 very hard 0.04 Grok 4.7, World Heritage, very hard: pass after 1 revision; readableShare 0.04; computed tier very hard 0.34 Fable 5.1, World Heritage, very hard: pass after 1 revision; readableShare 0.34; computed tier very hard fail Grok 4.6, World Heritage, very hard: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 1 sound off-tier draft kept 0.10 Opus 5.5, World Heritage, very hard: pass after 1 revision; readableShare 0.10; computed tier very hard fail GPT-6 Sol, World Heritage, very hard: fail after 3 attempts; the best draft failed only the difficulty-tier gate; 2 sound off-tier drafts kept 0.05 GPT-6 Astra, World Heritage, very hard: pass after 1 revision; readableShare 0.05; computed tier very hard 4 of 6 Passed 9 of 10 8 of 10 8 of 10 8 of 10 7 of 10 6 of 10 readableShare, darker = harder 1 0.75 0.5 0.25 0 fail; corner mark: the best draft missed only the requested tier
Table 2. Per model over the 10 requests. readableShare: the mean over the model's passes (lower = harder); "vs same request": the mean difference from the mean of all passes of the same request. Specul. work: the walker's speculativeWork, the depth of trial reasoning the solve needed (higher = more). Calls, model minutes (Min) and output tokens per request are totals over the ten requests divided by ten; per pass, divided by the model's passes.
ModelPassed
of 10
readable-
Share
vs same
request
Specul.
work
Calls /
request
Min /
request
Min /
pass
Tokens /
request
Grok 4.790.558+0.01310921.624.927.7119,174
Fable 5.180.533−0.03813292.417.622.094,074
Grok 4.680.543−0.05315522.021.226.582,850
Opus 5.580.598−0.00713302.213.116.488,527
GPT-6 Sol70.696+0.09214972.79.914.128,326
GPT-6 Astra60.561+0.00416132.913.823.026,985

Grok 4.7 passed 9 of the 10 requests; Fable 5.1, Grok 4.6 and Opus 5.5 passed 8, GPT-6 Sol 7 and GPT-6 Astra 6. Every request was passed by at least two models (the World Heritage hard request, by Grok 4.7 and Fable 5.1). The easy end held: all six models passed both very easy requests (12 of 12), and 11 of 12 easy ones passed. Failures gathered higher up: 9 of 12 medium, 5 of 12 hard and 9 of 12 very hard passed. Most failures were near misses: in 10 of the 14 the best draft of the run was sound and failed only the difficulty-tier gate, computing a different tier from the one requested.

3.2 Hardness (RQ2)

readableShare is the share of the board that a solver can fill by reading the clues one at a time, before combining any two or trying an assumption. It is computed by the game engine's solve-path walker; lower means harder. It falls with the tier as intended: over all passes it averaged 0.92 at very easy and 0.09 at very hard.

Figure 2. How hard each model's passing puzzles are. Each mark is a pass (circles: Wimbledon pack; squares: World Heritage pack) at its readableShare, the share of the board readable one clue at a time as computed by the game engine's solve-path walker; lower = harder. Passes of one pack at the same value share a mark whose area is proportional to their number. Diamonds and the right column: the mean over the model's passes. Each model passed a different set of requests, so a mean also reflects which tiers it passed; Table 2 gives the difference from the mean of all passes of the same request. Source: GRIDIGMA out/s3/.
Wimbledon pack World Heritage pack Mean 0.0 0.2 0.4 0.6 0.8 1.0 Mean Grok 4.7 Grok 4.7, Wimbledon, very hard: readableShare 0.06 Grok 4.7, Wimbledon, easy: readableShare 0.71 Grok 4.7, Wimbledon, medium: readableShare 0.88 Grok 4.7, Wimbledon, very easy: readableShare 1.00 Grok 4.7, World Heritage, very hard: readableShare 0.04 Grok 4.7, World Heritage, hard: readableShare 0.14 Grok 4.7, World Heritage, medium: readableShare 0.32 Grok 4.7, World Heritage, easy: readableShare 0.88 Grok 4.7, World Heritage, very easy: readableShare 1.00 Grok 4.7: mean over its 9 passes 0.558 0.558 Fable 5.1 Fable 5.1, Wimbledon, very hard: readableShare 0.05 Fable 5.1, Wimbledon, medium: readableShare 0.25 Fable 5.1, Wimbledon, very easy: readableShare 0.69 Fable 5.1, Wimbledon, easy: readableShare 1.00 Fable 5.1, World Heritage, hard: readableShare 0.24 Fable 5.1, World Heritage, very hard: readableShare 0.34 Fable 5.1, World Heritage, very easy: readableShare 0.69 Fable 5.1, World Heritage, easy: readableShare 1.00 Fable 5.1: mean over its 8 passes 0.533 0.533 Grok 4.6 Grok 4.6, Wimbledon, very hard: readableShare 0.04 Grok 4.6, Wimbledon, hard: readableShare 0.20 Grok 4.6, Wimbledon, medium: readableShare 0.25 Grok 4.6, Wimbledon, very easy: readableShare 0.83 Grok 4.6, Wimbledon, easy: readableShare 0.88 Grok 4.6, World Heritage, medium: readableShare 0.31 Grok 4.6, World Heritage, very easy: readableShare 0.83 Grok 4.6, World Heritage, easy: readableShare 1.00 Grok 4.6: mean over its 8 passes 0.543 0.543 Opus 5.5 Opus 5.5, Wimbledon, hard: readableShare 0.19 Opus 5.5, Wimbledon, easy: readableShare 0.67 Opus 5.5, Wimbledon, medium: readableShare 0.75 Opus 5.5, Wimbledon, very easy: readableShare 1.00 Opus 5.5, World Heritage, very hard: readableShare 0.10 Opus 5.5, World Heritage, medium: readableShare 0.20 Opus 5.5, World Heritage, easy: readableShare 0.88 Opus 5.5, World Heritage, very easy: readableShare 1.00 Opus 5.5: mean over its 8 passes 0.598 0.598 GPT-6 Sol GPT-6 Sol, Wimbledon, very hard: readableShare 0.13 GPT-6 Sol, Wimbledon, hard: readableShare 0.16 GPT-6 Sol, Wimbledon, very easy; Wimbledon, easy: readableShare 1.00 GPT-6 Sol, World Heritage, medium: readableShare 0.59 GPT-6 Sol, World Heritage, very easy; World Heritage, easy: readableShare 1.00 GPT-6 Sol: mean over its 7 passes 0.696 0.696 GPT-6 Astra GPT-6 Astra, Wimbledon, very hard: readableShare 0.02 GPT-6 Astra, Wimbledon, very easy: readableShare 1.00 GPT-6 Astra, World Heritage, very hard: readableShare 0.05 GPT-6 Astra, World Heritage, medium: readableShare 0.29 GPT-6 Astra, World Heritage, very easy; World Heritage, easy: readableShare 1.00 GPT-6 Astra: mean over its 6 passes 0.561 0.561 readableShare of each pass (lower = harder)

Five of the six models are close: their passes average 0.533 to 0.598 (Fable 5.1 lowest). GPT-6 Sol's passes are the easiest to read clue by clue (0.696). A mean covers only the requests a model passed, so a model that failed the harder requests is left with an easier set. Compared with the mean of all passes of the same request (Table 2), the picture holds: GPT-6 Sol is the most readable (+0.092), Grok 4.6 (−0.053) and Fable 5.1 (−0.038) the least, and the rest lie within a few hundredths of the request's mean.

3.3 Time and cost (RQ3)

Figure 3. Time per model. Model minutes are the time from sending a call to its answer, summed over the model's answered calls for the ten requests, with every usage-limit wait removed. Circles: per request (the total over ten); squares: per pass (the total over the model's passes). The right column is output tokens per request, as each route reports them. Source: GRIDIGMA out/s3/<model>/calls.jsonl.
Model minutes per request Model minutes per pass 0 5 10 15 20 25 30 Tokens out Grok 4.7 Grok 4.7: 24.9 model minutes per request (16 answered calls, 249.2 minutes in all) Grok 4.7: 27.7 model minutes per pass (9 passes) 119,174 Fable 5.1 Fable 5.1: 17.6 model minutes per request (24 answered calls, 175.9 minutes in all) Fable 5.1: 22.0 model minutes per pass (8 passes) 94,074 Grok 4.6 Grok 4.6: 21.2 model minutes per request (20 answered calls, 211.6 minutes in all) Grok 4.6: 26.5 model minutes per pass (8 passes) 82,850 Opus 5.5 Opus 5.5: 13.1 model minutes per request (22 answered calls, 131.2 minutes in all) Opus 5.5: 16.4 model minutes per pass (8 passes) 88,527 GPT-6 Sol GPT-6 Sol: 9.9 model minutes per request (27 answered calls, 98.6 minutes in all) GPT-6 Sol: 14.1 model minutes per pass (7 passes) 28,326 GPT-6 Astra GPT-6 Astra: 13.8 model minutes per request (29 answered calls, 138.0 minutes in all) GPT-6 Astra: 23.0 model minutes per pass (6 passes) 26,985 Model minutes (limit waits excluded)

GPT-6 Sol spent the least model time: 9.9 minutes per request and 14.1 per pass. Among the models that passed eight or more, Opus 5.5 was the fastest per pass (16.4 minutes), then Fable 5.1 (22.0), Grok 4.6 (26.5) and Grok 4.7, the slowest of all (27.7). The GPT-6 models reported about 28,326 and 26,985 output tokens per request, against 82,850 to 119,174 for the others.

The time goes to the model. The two Grok models never waited on a limit, so their calls ran back to back. Between the start of its first call and the end of its last, Grok 4.6 worked 211.6 of 213.7 minutes and Grok 4.7 249.2 of 250.7; everything between the calls (the gates, building the prompt, moving to the next request) took 2.1 and 1.5 minutes in all.

Money. Only two routes report a price. xAI reported USD 2.52 for Grok 4.6's answered calls and USD 3.05 for Grok 4.7's. The Claude client reports USD 75.39 for Fable 5.1 and USD 27.47 for Opus 5.5, although both ran on a flat subscription whose limits paused them; the Codex client reports no price for the GPT-6 models. Cost is therefore compared here by calls, tokens and model minutes.

3.4 The blind read

An AI agent running Opus 5.5 (itself one of the six authors) that had not seen which model wrote what read the English of all 46 passing puzzles as a player would, one request at a time with the versions shuffled, and scored theme (0–2), story (0–3: does it set up everything the clues rely on), clues (0–3: plain words, no arithmetic, one relation per clue, natural) and appeal (0–2). It is a single reader, and it is reported on its own.

Figure 4. The blind read: every passing puzzle's English, read by one AI agent that did not know the author, request by request with the versions shuffled, and scored for theme (0–2), story (0–3), clues (0–3) and appeal (0–2). Each circle's area is proportional to the number of the model's passes at that total; diamonds and the right column give the mean. Reported separately; it enters none of the measures in Figures 1–3. Source: GRIDIGMA out/s3/blindread/ (scores and key).
Passes at that score (area = count) Mean 0 1 2 3 4 5 6 7 8 9 10 Mean Grok 4.7 Grok 4.7: 2 of its 9 passes scored 4 Grok 4.7: 1 of its 9 passes scored 6 Grok 4.7: 2 of its 9 passes scored 7 Grok 4.7: 2 of its 9 passes scored 8 Grok 4.7: 2 of its 9 passes scored 9 Grok 4.7: mean 6.89 6.89 Fable 5.1 Fable 5.1: 1 of its 8 passes scored 4 Fable 5.1: 2 of its 8 passes scored 7 Fable 5.1: 1 of its 8 passes scored 8 Fable 5.1: 2 of its 8 passes scored 9 Fable 5.1: 2 of its 8 passes scored 10 Fable 5.1: mean 8.00 8.00 Grok 4.6 Grok 4.6: 3 of its 8 passes scored 5 Grok 4.6: 2 of its 8 passes scored 6 Grok 4.6: 2 of its 8 passes scored 7 Grok 4.6: 1 of its 8 passes scored 9 Grok 4.6: mean 6.25 6.25 Opus 5.5 Opus 5.5: 1 of its 8 passes scored 5 Opus 5.5: 1 of its 8 passes scored 6 Opus 5.5: 6 of its 8 passes scored 9 Opus 5.5: mean 8.13 8.13 GPT-6 Sol GPT-6 Sol: 2 of its 7 passes scored 5 GPT-6 Sol: 2 of its 7 passes scored 6 GPT-6 Sol: 2 of its 7 passes scored 7 GPT-6 Sol: 1 of its 7 passes scored 8 GPT-6 Sol: mean 6.29 6.29 GPT-6 Astra GPT-6 Astra: 1 of its 6 passes scored 2 GPT-6 Astra: 3 of its 6 passes scored 4 GPT-6 Astra: 2 of its 6 passes scored 5 GPT-6 Astra: mean 4.00 4.00 Blind read of one pass, out of 10 (theme 2, story 3, clues 3, appeal 2)
Table 3. Blind read per model: mean scores over its passes. "Best alone": requests in which its version scored highest outright; "Best shared": requests in which it tied for the highest.
ModelPasses
read
Total
/10
Theme
/2
Story
/3
Clues
/3
Appeal
/2
Best
alone
Best
shared
Opus 5.588.132.002.751.751.6313
Fable 5.188.002.002.751.751.5052
Grok 4.796.892.002.221.780.8911
GPT-6 Sol76.292.001.292.001.0000
Grok 4.686.251.751.631.751.1301
GPT-6 Astra64.001.831.170.670.3300

Opus 5.5 (8.13) and Fable 5.1 (8.00) read best, then Grok 4.7 (6.89), GPT-6 Sol (6.29), Grok 4.6 (6.25) and GPT-6 Astra (4.00). Fable 5.1 wrote the best version of its request outright in 5 of the 10 requests and tied for it in 2 more; Opus 5.5 outright in 1 and tied in 3; GPT-6 Sol and GPT-6 Astra never wrote the best version. The models differ more in the story (means 1.17–2.75 of 3) than in the clues (0.67–2.00 of 3). The faults the reader named most were clues built on a total of neighbouring values, told as arithmetic; textbook templates ("X, Y and Z were three different people", "exactly one of these two claims is true"); stories that never define what a clue relies on; and one sentence carrying several relations. The best versions define every time, place and grouping in the story and keep each clue a short, plain statement.

4. Discussion

No model wins every measure. Grok 4.7 passes the most requests but is the slowest per pass and reads below the two Anthropic models. Opus 5.5 and Fable 5.1 pass 8 each and read best; Opus 5.5 is the fastest of the models with eight or more passes, and Fable 5.1 wrote the most best versions. Grok 4.6 passes 8 and writes the hardest puzzles against the same requests, with a weaker read. GPT-6 Sol is the fastest but passes 7 and writes the most readable, that is the least demanding, logic. GPT-6 Astra passes 6 and reads worst. GRIDIGMA's study agent recommends Opus 5.5 for authoring the logic, or Fable 5.1 if the story is weighed above speed, and Grok 4.7 if the pass rate alone decides. That recommendation is the study agent's; the product owner has not decided.

5. Limitations

  • Ten requests. Each model ran each of the ten requests once, so a difference of one pass is one request. The 95% interval (Wilson [2]) for a pass rate of 9 of 10 runs from 0.60 to 0.98, and for 6 of 10 from 0.31 to 0.83: the ranking is a direction, not a verdict.
  • The commonest defect was the prompt's. The instructions encouraged the clue types the catalogue uses least, among them the window sum, which is a sum by nature. All 9 very hard passes used one, as did 2 medium ones, and in all 11 the blind reader named the sum or total; those passes averaged 0.55 of 3 for their clues, against 2.00 for the rest. Several drafts stated the total as arithmetic in prose, which the product owner had ruled out. The next authoring prompt has to say how a window sum is told without arithmetic, or stop encouraging it.
  • Time goes to reasoning, not to the gates (Section 3.3), so a faster run needs a faster model or fewer revisions, not faster gates.
  • Subscription limits. Four of the six models ran on subscriptions whose limits paused them, and three of them had requests restarted. Model minutes exclude the waits, but the elapsed time of a run on those routes is far longer than its model minutes.
  • One reader. The blind read is one AI agent's judgement: a subagent of the Claude session that ran the study, running Opus 5.5, the same model as one of the six authors and from the same maker as a second, Fable 5.1. Those two received the best reads (8.13 and 8.00); the read was blind, but a preference for its own or its maker's style [3] cannot be ruled out. Players have not rated these puzzles.
  • Model identity. As in the companion paper [1] (Section 7.5), the GPT-6 identifiers are those the Codex client reports; they could not be confirmed independently.
  • Not placed. None of the passing puzzles has been published in the game; they are kept for review.

6. Conclusion

Six models authored the logic of 10 puzzles each through GRIDIGMA's production loop and gates; Grok 4.7 passed the most (9 of 10), Opus 5.5 and Fable 5.1 passed 8 and read best, and GPT-6 Sol was the fastest. No model won every measure, and the product owner has not chosen one. With ten requests per model, a difference of one pass is one request, so these results give a direction for choosing an authoring model, not a choice. The next authoring prompt should say how a window sum is told without arithmetic, or stop encouraging it; and players have not yet rated the passing puzzles.

Declarations

Data availability. Every draft, passing puzzle, gate record, call log and blind-read score, and the analysis scripts, are held by Fabervant and are available to researchers on request at the address above. The puzzles' solutions and stored logic are withheld; none of the puzzles is in the game yet.

Competing interests. Fabervant develops GRIDIGMA and has a commercial interest in how its puzzles are authored. Fabervant is not affiliated with Anthropic, OpenAI or xAI, received no funding or model access from them for this study, and none of them reviewed this report.

Use of AI tools. The puzzles in this study were authored by the models under test, and the blind read was performed by an AI agent of the study running Opus 5.5, one of those models. The pipeline, the analysis and the drafting of this report were carried out with AI coding assistants under the author's direction; the author set the research questions and the instructions, and is responsible for the content.

Ethics. No human participants were involved and no personal data was processed.

References

  1. Güneş, C. (2026). Nine language models rewriting and translating logic-puzzle text: A blind cross-rating study (Version 3.2). Fabervant. https://fabervant.com/research/nine-model-prose-trial/
  2. Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association, 22(158), 209–212.
  3. Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems 37. arXiv:2404.13076.

How to cite

Güneş, C. (2026). Six language models authoring logic puzzles through a game's production gates (Version 1.0). Fabervant. https://fabervant.com/research/six-model-logic-authoring/

Version history: 1.0 (29 September 2026): published as a paper of its own. The study first appeared as Section 10 of the nine-model prose paper [1], version 3.1 (27 September 2026), whose Figures 15–18 and Tables 19 and 20 are Figures 1–4 and Tables 2 and 3 here. Its numbers and findings are unchanged; the method is now self-contained (Section 2), with the models listed in Table 1, and the abstract, introduction, discussion and conclusion are written for a paper of its own.