Fabervant Research · Technical report · Stages 1, 2, 2b and 2c

Nine Language Models Rewriting and Translating Logic-Puzzle Text: A Blind Cross-Rating Study

Stage 1: English and Turkish. Stage 2: twenty-one languages. Stages 2b and 2c: plain words and a naturalness-first rubric

These puzzles are live: play GRIDIGMA free →

Version 3.2 · Published 24 September 2026, Stages 2, 2b and 2c added 25 September 2026 · Native-speaker ratings to follow

Abstract

Background. Text in a logic-puzzle game must read naturally while every clue keeps one exact logical meaning. We asked which current language models do this best, in English and in many languages, and at what cost.

Methods. In Stage 1, nine models from Anthropic, OpenAI and xAI each rewrote the English of three published GRIDIGMA puzzles (27 versions) and translated every model's English into Turkish (243 versions). Every model then rated every version blind on a fixed 10-point rubric (clarity, readability, naturalness), with the text players read today inserted unmarked as a control: 2,430 ratings of the 270 versions, plus 270 of the controls. In Stage 2, the same nine models translated one English text into 21 languages at a matched reasoning effort (189 translations), and each language's nine translations were rated blind by all nine models (1,701 ratings). A separate check judged each clue as keeping, changing or blurring its meaning. Stage 2b repeated Stage 2 on a plain-words edit of the source (189 translations, 1,701 ratings). Stage 2c required plain words with no arithmetic, weighted the rubric towards naturalness (6 of 10 points), had all nine models write the English of the three puzzles and three of them translate it into the 21 languages by two routes (153 versions, 1,377 ratings). We report means with 95% two-way cluster-bootstrap confidence intervals, rater agreement, and cost at list prices.

Results. Raters agreed well in Stages 1 and 2 (Kendall's W 0.67 to 0.70; ICC for the mean of nine raters 0.90 to 0.92), less on the plain source (Stage 2b: W 0.52; ICC 0.86) and least on the Stage 2c translations (W 0.36; ICC 0.78), so per-language picks there are directions only. In Stage 1, GPT-6 Astra wrote the highest-rated English (9.67/10, 95% CI 9.33–9.93) and Opus 5.5 the highest-rated Turkish (9.18, 8.78–9.52); the pair of the two scored 9.70 with no clue flagged. Today's live text scored below the rewrites in 265 of 270 batches, and Turkish naturalness varied far more across translators than across English sources. In Stage 2, across 21 languages, Opus 5.5 (9.28, 8.90–9.57) and GPT-6 Astra (9.20, 8.80–9.54) were statistically tied at the top and clear of the other seven; only 5 of 1,701 translated clues were flagged. On the plain source (Stage 2b) the weaker translators closed much of the gap, and GPT-6 Astra, Opus 5.5 and GPT-6 Sol were tied at the top. Under the naturalness-first rubric (Stage 2c), GPT-6 Astra and Opus 5.5 tied as English writers (8.87 and 8.80) and shared the translation lead (10 and 9 of 21 languages); translating from the stored logic instead of the English made no difference on average (7.80 against 7.86).

Conclusions. For this task the translator matters more than the English writer, and two translators, Opus 5.5 and GPT-6 Astra, stay at the top through every change of language, source and rubric, with notable per-language exceptions; the mechanical text came more from the inputs (a formula clue, technical labels) than from the models. The evidence rests on model raters; native-speaker ratings follow as an addendum.

Keywords: large language models; machine translation; multilingual evaluation; Turkish; LLM-as-a-judge; controlled text generation; game localisation; logic puzzles

1. Introduction

GRIDIGMA is a logic-grid deduction game played mostly on phones, published in English and Turkish. Each puzzle has a title, a one-line description, a short story and a set of clues. The text has two jobs that pull against each other. It should read easily and sound natural, because it is what draws a player in. And each clue is a piece of logic: together the clues admit exactly one answer, so a clue that says a little more, a little less, or can be read two ways breaks the puzzle. Rewriting such text for style, or translating it, is therefore a constrained generation task in which fluency and fidelity can trade off.

Language models are now used both to produce such text and to judge it. Model judges agree with human preferences at levels comparable to agreement between humans on some tasks [1], and have been used successfully to assess translation quality [2]. They also have known biases, among them a preference for their own generations [3]. A study that uses models as judges must therefore measure how much the judges agree and whether they favour themselves.

We asked six questions:

  1. RQ1. Which models write the clearest and most natural English puzzle text, and which translate it into the most natural Turkish?
  2. RQ2. Does the quality of a Turkish translation depend more on the translator or on the English it starts from?
  3. RQ3. Do models keep every clue's logical meaning while rewriting and translating?
  4. RQ4. What does each option cost, in money and in time?
  5. RQ5. Does the translators' ranking hold across languages when every model runs at the same reasoning effort?
  6. RQ6. Do the rankings hold when the source is written in plain words and the rubric weights naturalness first?

Stage 1 (Sections 2–6) answers RQ1–RQ4 for English and Turkish. Stage 2 (Section 7) answers RQ5 across 21 languages, and Stages 2b and 2c (Sections 8 and 9) answer RQ6.

2. Methods

2.1 Design

A fully crossed design: every model wrote English for every puzzle, every model translated every model's English, and every model rated every version. Ratings were blind to authorship. The text players read today was rated alongside, unmarked, as a control.

2.2 Models

Table 1. The nine models, called by the model IDs shown on 24 September 2026.
ModelProviderModel ID
Sonnet 5Anthropicclaude-sonnet-5
Opus 5Anthropicclaude-opus-5
Opus 5.5Anthropicclaude-opus-5-5
Fable 5.1Anthropicclaude-fable-5-1
Grok 4.6xAIgrok-4.6
Grok 4.7xAIgrok-4.7
GPT-6 LunaOpenAIgpt-6-luna
GPT-6 SolOpenAIgpt-6-sol
GPT-6 AstraOpenAIgpt-6-astra

2.3 Materials

Three published GRIDIGMA daily puzzles, all 4 × 4 grids, chosen for different kinds of clue: The Seed-Sorting Trays (8 clues, positional), The Rope-Line Survey (9 clues, exactly one of three lettered field notes is false) and The Ropewalk Ledger (9 clues). The controls were the English and Turkish texts players read in the game on 24 September 2026. The live English had been written with an earlier model-based authoring pipeline and the live Turkish translated from it under the project's Turkish style guide.

2.4 Procedure

English rewrite (9 models × 3 puzzles = 27 versions). Each model rewrote a puzzle's title, description, story and clues. The instruction, in the product owner's words, was to make the text "simple, natural and fluid. The aim is not a piece of literature: the text is there to prepare the players for the puzzle and sweep them along." Models were free to improvise the story and wording, but had to keep the number and order of clues and make each clue "state exactly the same fact as the clue it replaces – no more, no less, and with only one possible reading". To let a model check its own wording, the prompt included the puzzle's answer and each clue's stored logic.

Turkish translation (9 translators × 9 English sources × 3 puzzles = 243 versions). Each model translated each model's English, its own included. The instruction asked for "the way a Turkish writer would tell it to Turkish players, not a word-for-word rendering", under the same fidelity constraint, using the Turkish edition's own cast of names.

Every call used the same one-line system prompt ("You are a writer preparing the text of a puzzle game."), and one user message per call, with no conversation, no tools and no other context. A reply that was not valid JSON with the right number of clues was re-sent once with a one-line reminder.

2.5 Measures

Rubric. Fixed before any rating: clarity 0–4 (one reading per clue; the setting easy to grasp), readability 0–3 (flows; easy on a phone) and naturalness 0–3 (sounds the way a native speaker would say it). The total runs from 0 to 10.

Fidelity. Each English clue was compared with the puzzle's stored logic and answer, and each Turkish clue with the English clue it translated, and judged same, changed (states a different fact) or ambiguous (admits a second reading that is a different fact). Opus 5.5 performed this check; every changed verdict was re-read by hand.

2.6 Rating and blinding

For each puzzle, each rater received one English batch (the 9 English versions plus the live English) and nine Turkish batches (the 9 translations of one English source plus the live Turkish). Within every batch the 10 texts were shuffled with a deterministic per-rater seed and labelled V1–V10. Raters were told that the versions came from different writers in random order, and to judge each on its own text without rewarding length, ornament or similarity. Each returned the three criterion scores and a one-sentence note per version. In total: 270 rating calls, and 2,700 scores, 2,430 of them of rewrites. Each model also rated its own work; we report every mean with and without self-ratings.

2.7 Settings

No temperature, top-p or reasoning effort was set, so every model ran at its provider's defaults. Reasoning effort was therefore not matched across models in Stage 1. The defaults were established afterwards: high for Sonnet 5, Opus 5 and Fable 5.1, and medium for Opus 5.5 (read back from the client); at least high for both Grok models (measured: with no effort set they reason as much as at high or more); and most likely medium for the three GPT-6 models (measured from output length against each level; firm for GPT-6 Luna, not proven for Sol and Astra). Stage 2 set the effort explicitly (Section 7.1).

2.8 Statistical analysis

The unit is one rater's total for one version. A model's English score is the mean over its 3 English versions × 9 raters; its Turkish score is the mean over its 27 translations × 9 raters. The 95% confidence intervals come from a two-way cluster bootstrap [4] that resamples raters and versions independently (4,000 resamples, fixed seed), so each interval carries both rater and item variation. The stability of the best pair was estimated by resampling raters and puzzles 1,000 times and counting how often each pair ranked first. Agreement between raters was measured with Kendall's coefficient of concordance W [5], corrected for ties, within each batch of 10 versions × 9 raters. The intraclass correlation [6] was measured as ICC(2,1) for a single rater and ICC(2,k) for the mean of the nine, two-way random effects with absolute agreement, and interpreted after [7]. The controls were compared with the rewrites by counting the batches in which the control scored below the batch's mean rewrite.

2.9 Cost accounting

Cost is each provider's list price on 24 September 2026 multiplied by the tokens used, the same for every provider: reasoning tokens billed as output, no cache discounts. The prices, in USD per million input/output tokens: Sonnet 5 2/10, Opus 5 5/25, Opus 5.5 4/20, Fable 5.1 10/50, Grok 4.6 2/6, Grok 4.7 2/6, GPT-6 Luna 0.1/0.5, GPT-6 Sol 2/10, GPT-6 Astra 10/50.

3. Results

3.1 Agreement between raters

Kendall's W averaged 0.70 over the 3 English batches (range 0.59–0.79) and 0.69 over the 27 Turkish batch groups (range 0.57–0.82): strong agreement on the ordering of versions. On absolute scores, a single rater's reliability was moderate (ICC(2,1) = 0.51 for English, 0.56 for Turkish), and the mean of nine raters was excellent (ICC(2,k) = 0.90 and 0.92). Every score below is therefore a nine-rater mean.

3.2 Scores by model (RQ1)

Figure 1. Stage 1: mean blind score out of 10 for each model over everything it wrote, with 95% confidence intervals. Circles: English (3 versions × 9 raters = 27 ratings per model); squares: Turkish (27 translations × 9 raters = 243 ratings per model). Intervals: two-way cluster bootstrap resampling raters and versions (4,000 resamples, seed 20260924, as stats.py). The last row is today's live text, a point estimate (English n = 27 ratings, Turkish n = 243). Source: GRIDIGMA out/rate/.
English Turkish 5 6 7 8 9 10 GPT-6 Astra GPT-6 Astra: English 9.67 (95% CI 9.33 to 9.93) GPT-6 Astra: Turkish 9.09 (95% CI 8.77 to 9.37) GPT-6 Sol GPT-6 Sol: English 9.30 (95% CI 8.93 to 9.63) GPT-6 Sol: Turkish 8.93 (95% CI 8.61 to 9.25) Opus 5.5 Opus 5.5: English 9.26 (95% CI 8.56 to 9.85) Opus 5.5: Turkish 9.18 (95% CI 8.78 to 9.52) Grok 4.6 Grok 4.6: English 9.19 (95% CI 8.44 to 9.89) Grok 4.6: Turkish 7.11 (95% CI 6.62 to 7.62) Fable 5.1 Fable 5.1: English 8.89 (95% CI 8.11 to 9.67) Fable 5.1: Turkish 8.69 (95% CI 8.40 to 9.00) Opus 5 Opus 5: English 8.59 (95% CI 7.78 to 9.37) Opus 5: Turkish 8.47 (95% CI 7.97 to 8.92) Grok 4.7 Grok 4.7: English 8.48 (95% CI 7.74 to 9.11) Grok 4.7: Turkish 7.60 (95% CI 7.07 to 8.14) Sonnet 5 Sonnet 5: English 7.52 (95% CI 6.19 to 8.70) Sonnet 5: Turkish 7.48 (95% CI 6.99 to 7.98) GPT-6 Luna GPT-6 Luna: English 7.22 (95% CI 6.26 to 8.11) GPT-6 Luna: Turkish 7.35 (95% CI 6.85 to 7.84) Live text Live text: English 6.33 Live text: Turkish 5.51 Mean blind score out of 10 (Stage 1 rubric)
Table 2. Mean score out of 10 with 95% CI; "no self" leaves out the model's rating of its own work; the last column is the mean score the model gave others.
ModelEnglish [95% CI]No selfTurkish [95% CI]No selfGave
GPT-6 Astra9.67 [9.33, 9.93]9.639.09 [8.77, 9.37]9.018.14
GPT-6 Sol9.30 [8.93, 9.63]9.258.93 [8.61, 9.25]8.917.79
Opus 5.59.26 [8.56, 9.85]9.259.18 [8.78, 9.52]9.177.56
Grok 4.69.19 [8.44, 9.89]9.177.11 [6.62, 7.62]7.107.64
Fable 5.18.89 [8.11, 9.67]8.838.69 [8.40, 9.00]8.638.28
Opus 58.59 [7.78, 9.37]8.508.47 [7.97, 8.92]8.438.16
Grok 4.78.48 [7.74, 9.11]8.467.60 [7.07, 8.14]7.577.85
Sonnet 57.52 [6.19, 8.70]7.427.48 [6.99, 7.98]7.358.64
GPT-6 Luna7.22 [6.26, 8.11]7.217.35 [6.85, 7.84]7.317.86

The English intervals are wide because each rests on only three versions; the top four English writers are not separable at this sample size. The Turkish intervals, each resting on 27 versions, are narrower: the intervals of Opus 5.5, GPT-6 Astra and GPT-6 Sol lie entirely above those of Grok 4.7, Sonnet 5, GPT-6 Luna and Grok 4.6, while Fable 5.1 and Opus 5 overlap both groups.

Self-preference. Removing each model's rating of its own work moved its mean by at most 0.13 points (Sonnet 5, Turkish), well inside every interval. We found no meaningful self-preference in this design.

Table 3. Mean criterion scores over all raters and versions. Clarity is out of 4; readability and naturalness out of 3.
ModelEN clarityEN readabilityEN naturalnessTR clarityTR readabilityTR naturalness
GPT-6 Astra3.892.962.813.822.682.59
GPT-6 Sol3.782.962.563.642.732.56
Opus 5.53.562.932.783.752.812.61
Grok 4.63.412.892.893.002.301.81
Fable 5.13.782.482.633.652.672.38
Opus 53.632.522.443.442.672.35
Grok 4.73.562.672.263.272.401.93
Sonnet 52.742.522.263.192.341.95
GPT-6 Luna2.632.302.303.152.132.06
Table 4. English score per puzzle (mean of nine raters).
ModelSeed-Sorting TraysRope-Line SurveyRopewalk Ledger
GPT-6 Astra9.679.569.78
GPT-6 Sol9.339.229.33
Opus 5.59.788.679.33
Grok 4.69.898.788.89
Fable 5.18.118.899.67
Opus 59.338.567.89
Grok 4.78.678.788.00
Sonnet 58.677.676.22
GPT-6 Luna7.566.337.78

3.3 Comparison with today's text

Today's live text scored below the mean of the rewrites in its batch in 265 of 270 batches (all 27 English batches, 238 of 243 Turkish), and at or below every single rewrite in 200 of 270. Its mean scores were 5.67–6.67 in English and 4.91–5.92 in Turkish (Table 5). The raters' notes point to the same defects again and again: an ornate story, and positional clues ("left", "right", "tray 3") given without saying how the trays are laid out. The example below shows the difference on one clue.

Table 5. Today's live text, rated blind among the rewrites.
PuzzleLive EnglishLive TurkishLive Turkish naturalness (of 3)
The Seed-Sorting Trays6.675.921.83
The Rope-Line Survey6.675.691.63
The Ropewalk Ledger5.674.911.43

Live text, clue 2
The clover sorting was done at the tray immediately to the right of the basil sorting.

Opus 5.5 rewrite
The clover tray number is exactly one higher than the basil tray number.

Story sentence the rewrite added
The trays are numbered 1 to 4 along the bench.

3.4 Translator versus source (RQ2)

Turkish naturalness ranged from 1.81 to 2.61 out of 3 across the nine translators, but only from 2.17 to 2.33 across the nine English sources. The English writer's quality barely carried into the Turkish; the translator decided it (Figure 2). The same holds for the total: Grok 4.6 wrote the fourth-best English (9.19) but produced the weakest Turkish (7.11).

Figure 2. Stage 1: Turkish naturalness (0–3) of the 243 translations, grouped two ways. Squares: the mean over the 27 translations a model made (as translator); circles: the mean over the 27 translations of that model's English (as source). Each mean covers 243 ratings. Across translators naturalness spans 1.81–2.61; across English sources 2.17–2.33. Point estimates, no intervals. Source: GRIDIGMA out/rate/.
Grouped by translator Grouped by English source 1.5 2.0 2.5 3.0 Opus 5.5 Opus 5.5 as English source: Turkish naturalness 2.27 of 3 (27 translations of its English) Opus 5.5 as translator: Turkish naturalness 2.61 of 3 (its 27 translations) GPT-6 Astra GPT-6 Astra as English source: Turkish naturalness 2.33 of 3 (27 translations of its English) GPT-6 Astra as translator: Turkish naturalness 2.59 of 3 (its 27 translations) GPT-6 Sol GPT-6 Sol as English source: Turkish naturalness 2.31 of 3 (27 translations of its English) GPT-6 Sol as translator: Turkish naturalness 2.56 of 3 (its 27 translations) Fable 5.1 Fable 5.1 as English source: Turkish naturalness 2.23 of 3 (27 translations of its English) Fable 5.1 as translator: Turkish naturalness 2.38 of 3 (its 27 translations) Opus 5 Opus 5 as English source: Turkish naturalness 2.23 of 3 (27 translations of its English) Opus 5 as translator: Turkish naturalness 2.35 of 3 (its 27 translations) GPT-6 Luna GPT-6 Luna as English source: Turkish naturalness 2.17 of 3 (27 translations of its English) GPT-6 Luna as translator: Turkish naturalness 2.06 of 3 (its 27 translations) Sonnet 5 Sonnet 5 as English source: Turkish naturalness 2.21 of 3 (27 translations of its English) Sonnet 5 as translator: Turkish naturalness 1.95 of 3 (its 27 translations) Grok 4.7 Grok 4.7 as English source: Turkish naturalness 2.20 of 3 (27 translations of its English) Grok 4.7 as translator: Turkish naturalness 1.93 of 3 (its 27 translations) Grok 4.6 Grok 4.6 as English source: Turkish naturalness 2.29 of 3 (27 translations of its English) Grok 4.6 as translator: Turkish naturalness 1.81 of 3 (its 27 translations) Turkish naturalness, mean out of 3

3.5 Writer and translator pairs

Table 6. The 15 best of 81 pairs. Score: the mean of the English and the Turkish score over three puzzles. Cost: list price per puzzle for both steps. Flagged: clues judged changed or ambiguous.
#English byTurkish byScoreUSD / puzzlePoints / USDFlagged
1GPT-6 AstraOpus 5.59.700.08291170
2GPT-6 SolOpus 5.59.400.05741640
3Opus 5.5Opus 5.59.370.08331130
4Grok 4.6Opus 5.59.350.1170801
5GPT-6 AstraGPT-6 Sol9.350.04981880
6GPT-6 AstraFable 5.19.330.1497620
7GPT-6 SolGPT-6 Astra9.310.05901580
8GPT-6 AstraGPT-6 Astra9.310.08981040
9GPT-6 SolOpus 59.310.06891350
10GPT-6 SolGPT-6 Sol9.290.02294060
11Grok 4.6GPT-6 Sol9.260.08691071
12Grok 4.6GPT-6 Astra9.260.1219761
13Fable 5.1Opus 5.59.200.1387660
14Opus 5.5GPT-6 Astra9.170.09061010
15Opus 5.5Fable 5.19.110.1448630

GPT-6 Astra with Opus 5.5 ranked first in 814 of 1,000 bootstrap resamples of raters and puzzles. The next most frequent leaders were GPT-6 Sol with GPT-6 Astra (55) and GPT-6 Sol with Opus 5.5 (28).

Figure 3. Stage 1: the 81 writer-translator pairs. Each cell is the Turkish score of one writer's English translated by one translator, averaged over the three puzzles and nine raters (27 ratings). A: total out of 10; B: naturalness out of 3. Rows: whose English; columns: who translated it; outlined cells on the diagonal are a model translating its own English; the thick frame marks the highest cell. Darker = higher. Point estimates. The per-puzzle cells are Tables 7–9. Source: GRIDIGMA out/rate/.
A. Turkish total, out of 10 (rows: English by; columns: Turkish by) Sonnet 5 Opus 5 Opus 5.5 Fable 5.1 Grok 4.6 Grok 4.7 GPT-6 Luna GPT-6 Sol GPT-6 Astra Sonnet 5 6.67 English by Sonnet 5, Turkish by Sonnet 5: 6.67 of 10 (27 ratings) 8.85 English by Sonnet 5, Turkish by Opus 5: 8.85 of 10 (27 ratings) 8.30 English by Sonnet 5, Turkish by Opus 5.5: 8.30 of 10 (27 ratings) 8.48 English by Sonnet 5, Turkish by Fable 5.1: 8.48 of 10 (27 ratings) 6.41 English by Sonnet 5, Turkish by Grok 4.6: 6.41 of 10 (27 ratings) 7.78 English by Sonnet 5, Turkish by Grok 4.7: 7.78 of 10 (27 ratings) 7.04 English by Sonnet 5, Turkish by GPT-6 Luna: 7.04 of 10 (27 ratings) 8.96 English by Sonnet 5, Turkish by GPT-6 Sol: 8.96 of 10 (27 ratings) 9.15 English by Sonnet 5, Turkish by GPT-6 Astra: 9.15 of 10 (27 ratings) Opus 5 8.30 English by Opus 5, Turkish by Sonnet 5: 8.30 of 10 (27 ratings) 7.93 English by Opus 5, Turkish by Opus 5: 7.93 of 10 (27 ratings) 9.07 English by Opus 5, Turkish by Opus 5.5: 9.07 of 10 (27 ratings) 8.93 English by Opus 5, Turkish by Fable 5.1: 8.93 of 10 (27 ratings) 6.44 English by Opus 5, Turkish by Grok 4.6: 6.44 of 10 (27 ratings) 7.96 English by Opus 5, Turkish by Grok 4.7: 7.96 of 10 (27 ratings) 7.93 English by Opus 5, Turkish by GPT-6 Luna: 7.93 of 10 (27 ratings) 8.56 English by Opus 5, Turkish by GPT-6 Sol: 8.56 of 10 (27 ratings) 9.30 English by Opus 5, Turkish by GPT-6 Astra: 9.30 of 10 (27 ratings) Opus 5.5 7.37 English by Opus 5.5, Turkish by Sonnet 5: 7.37 of 10 (27 ratings) 8.63 English by Opus 5.5, Turkish by Opus 5: 8.63 of 10 (27 ratings) 9.48 English by Opus 5.5, Turkish by Opus 5.5: 9.48 of 10 (27 ratings) 8.96 English by Opus 5.5, Turkish by Fable 5.1: 8.96 of 10 (27 ratings) 7.33 English by Opus 5.5, Turkish by Grok 4.6: 7.33 of 10 (27 ratings) 8.15 English by Opus 5.5, Turkish by Grok 4.7: 8.15 of 10 (27 ratings) 7.85 English by Opus 5.5, Turkish by GPT-6 Luna: 7.85 of 10 (27 ratings) 8.74 English by Opus 5.5, Turkish by GPT-6 Sol: 8.74 of 10 (27 ratings) 9.07 English by Opus 5.5, Turkish by GPT-6 Astra: 9.07 of 10 (27 ratings) Fable 5.1 7.59 English by Fable 5.1, Turkish by Sonnet 5: 7.59 of 10 (27 ratings) 7.70 English by Fable 5.1, Turkish by Opus 5: 7.70 of 10 (27 ratings) 9.52 English by Fable 5.1, Turkish by Opus 5.5: 9.52 of 10 (27 ratings) 8.26 English by Fable 5.1, Turkish by Fable 5.1: 8.26 of 10 (27 ratings) 6.37 English by Fable 5.1, Turkish by Grok 4.6: 6.37 of 10 (27 ratings) 8.11 English by Fable 5.1, Turkish by Grok 4.7: 8.11 of 10 (27 ratings) 7.22 English by Fable 5.1, Turkish by GPT-6 Luna: 7.22 of 10 (27 ratings) 9.07 English by Fable 5.1, Turkish by GPT-6 Sol: 9.07 of 10 (27 ratings) 8.48 English by Fable 5.1, Turkish by GPT-6 Astra: 8.48 of 10 (27 ratings) Grok 4.6 7.63 English by Grok 4.6, Turkish by Sonnet 5: 7.63 of 10 (27 ratings) 8.48 English by Grok 4.6, Turkish by Opus 5: 8.48 of 10 (27 ratings) 9.52 English by Grok 4.6, Turkish by Opus 5.5: 9.52 of 10 (27 ratings) 8.56 English by Grok 4.6, Turkish by Fable 5.1: 8.56 of 10 (27 ratings) 7.33 English by Grok 4.6, Turkish by Grok 4.6: 7.33 of 10 (27 ratings) 7.22 English by Grok 4.6, Turkish by Grok 4.7: 7.22 of 10 (27 ratings) 6.81 English by Grok 4.6, Turkish by GPT-6 Luna: 6.81 of 10 (27 ratings) 9.33 English by Grok 4.6, Turkish by GPT-6 Sol: 9.33 of 10 (27 ratings) 9.33 English by Grok 4.6, Turkish by GPT-6 Astra: 9.33 of 10 (27 ratings) Grok 4.7 7.15 English by Grok 4.7, Turkish by Sonnet 5: 7.15 of 10 (27 ratings) 8.96 English by Grok 4.7, Turkish by Opus 5: 8.96 of 10 (27 ratings) 8.44 English by Grok 4.7, Turkish by Opus 5.5: 8.44 of 10 (27 ratings) 9.04 English by Grok 4.7, Turkish by Fable 5.1: 9.04 of 10 (27 ratings) 7.33 English by Grok 4.7, Turkish by Grok 4.6: 7.33 of 10 (27 ratings) 7.37 English by Grok 4.7, Turkish by Grok 4.7: 7.37 of 10 (27 ratings) 6.70 English by Grok 4.7, Turkish by GPT-6 Luna: 6.70 of 10 (27 ratings) 9.04 English by Grok 4.7, Turkish by GPT-6 Sol: 9.04 of 10 (27 ratings) 9.11 English by Grok 4.7, Turkish by GPT-6 Astra: 9.11 of 10 (27 ratings) GPT-6 Luna 7.33 English by GPT-6 Luna, Turkish by Sonnet 5: 7.33 of 10 (27 ratings) 7.91 English by GPT-6 Luna, Turkish by Opus 5: 7.91 of 10 (27 ratings) 9.02 English by GPT-6 Luna, Turkish by Opus 5.5: 9.02 of 10 (27 ratings) 8.17 English by GPT-6 Luna, Turkish by Fable 5.1: 8.17 of 10 (27 ratings) 7.00 English by GPT-6 Luna, Turkish by Grok 4.6: 7.00 of 10 (27 ratings) 6.72 English by GPT-6 Luna, Turkish by Grok 4.7: 6.72 of 10 (27 ratings) 7.78 English by GPT-6 Luna, Turkish by GPT-6 Luna: 7.78 of 10 (27 ratings) 8.37 English by GPT-6 Luna, Turkish by GPT-6 Sol: 8.37 of 10 (27 ratings) 9.04 English by GPT-6 Luna, Turkish by GPT-6 Astra: 9.04 of 10 (27 ratings) GPT-6 Sol 7.67 English by GPT-6 Sol, Turkish by Sonnet 5: 7.67 of 10 (27 ratings) 9.31 English by GPT-6 Sol, Turkish by Opus 5: 9.31 of 10 (27 ratings) 9.50 English by GPT-6 Sol, Turkish by Opus 5.5: 9.50 of 10 (27 ratings) 8.85 English by GPT-6 Sol, Turkish by Fable 5.1: 8.85 of 10 (27 ratings) 7.65 English by GPT-6 Sol, Turkish by Grok 4.6: 7.65 of 10 (27 ratings) 7.48 English by GPT-6 Sol, Turkish by Grok 4.7: 7.48 of 10 (27 ratings) 6.67 English by GPT-6 Sol, Turkish by GPT-6 Luna: 6.67 of 10 (27 ratings) 9.28 English by GPT-6 Sol, Turkish by GPT-6 Sol: 9.28 of 10 (27 ratings) 9.33 English by GPT-6 Sol, Turkish by GPT-6 Astra: 9.33 of 10 (27 ratings) GPT-6 Astra 7.59 English by GPT-6 Astra, Turkish by Sonnet 5: 7.59 of 10 (27 ratings) 8.41 English by GPT-6 Astra, Turkish by Opus 5: 8.41 of 10 (27 ratings) 9.74 English by GPT-6 Astra, Turkish by Opus 5.5: 9.74 of 10 (27 ratings) 9.00 English by GPT-6 Astra, Turkish by Fable 5.1: 9.00 of 10 (27 ratings) 8.15 English by GPT-6 Astra, Turkish by Grok 4.6: 8.15 of 10 (27 ratings) 7.63 English by GPT-6 Astra, Turkish by Grok 4.7: 7.63 of 10 (27 ratings) 8.11 English by GPT-6 Astra, Turkish by GPT-6 Luna: 8.11 of 10 (27 ratings) 9.04 English by GPT-6 Astra, Turkish by GPT-6 Sol: 9.04 of 10 (27 ratings) 8.96 English by GPT-6 Astra, Turkish by GPT-6 Astra: 8.96 of 10 (27 ratings) Scale 6 7 8 9 10 B. Turkish naturalness, out of 3 Sonnet 5 Opus 5 Opus 5.5 Fable 5.1 Grok 4.6 Grok 4.7 GPT-6 Luna GPT-6 Sol GPT-6 Astra Sonnet 5 1.78 English by Sonnet 5, Turkish by Sonnet 5: naturalness 1.78 of 3 (27 ratings) 2.59 English by Sonnet 5, Turkish by Opus 5: naturalness 2.59 of 3 (27 ratings) 2.33 English by Sonnet 5, Turkish by Opus 5.5: naturalness 2.33 of 3 (27 ratings) 2.37 English by Sonnet 5, Turkish by Fable 5.1: naturalness 2.37 of 3 (27 ratings) 1.67 English by Sonnet 5, Turkish by Grok 4.6: naturalness 1.67 of 3 (27 ratings) 2.04 English by Sonnet 5, Turkish by Grok 4.7: naturalness 2.04 of 3 (27 ratings) 2.11 English by Sonnet 5, Turkish by GPT-6 Luna: naturalness 2.11 of 3 (27 ratings) 2.56 English by Sonnet 5, Turkish by GPT-6 Sol: naturalness 2.56 of 3 (27 ratings) 2.48 English by Sonnet 5, Turkish by GPT-6 Astra: naturalness 2.48 of 3 (27 ratings) Opus 5 2.22 English by Opus 5, Turkish by Sonnet 5: naturalness 2.22 of 3 (27 ratings) 2.11 English by Opus 5, Turkish by Opus 5: naturalness 2.11 of 3 (27 ratings) 2.44 English by Opus 5, Turkish by Opus 5.5: naturalness 2.44 of 3 (27 ratings) 2.41 English by Opus 5, Turkish by Fable 5.1: naturalness 2.41 of 3 (27 ratings) 1.48 English by Opus 5, Turkish by Grok 4.6: naturalness 1.48 of 3 (27 ratings) 2.07 English by Opus 5, Turkish by Grok 4.7: naturalness 2.07 of 3 (27 ratings) 2.15 English by Opus 5, Turkish by GPT-6 Luna: naturalness 2.15 of 3 (27 ratings) 2.52 English by Opus 5, Turkish by GPT-6 Sol: naturalness 2.52 of 3 (27 ratings) 2.70 English by Opus 5, Turkish by GPT-6 Astra: naturalness 2.70 of 3 (27 ratings) Opus 5.5 1.85 English by Opus 5.5, Turkish by Sonnet 5: naturalness 1.85 of 3 (27 ratings) 2.37 English by Opus 5.5, Turkish by Opus 5: naturalness 2.37 of 3 (27 ratings) 2.70 English by Opus 5.5, Turkish by Opus 5.5: naturalness 2.70 of 3 (27 ratings) 2.48 English by Opus 5.5, Turkish by Fable 5.1: naturalness 2.48 of 3 (27 ratings) 1.78 English by Opus 5.5, Turkish by Grok 4.6: naturalness 1.78 of 3 (27 ratings) 2.04 English by Opus 5.5, Turkish by Grok 4.7: naturalness 2.04 of 3 (27 ratings) 2.07 English by Opus 5.5, Turkish by GPT-6 Luna: naturalness 2.07 of 3 (27 ratings) 2.52 English by Opus 5.5, Turkish by GPT-6 Sol: naturalness 2.52 of 3 (27 ratings) 2.59 English by Opus 5.5, Turkish by GPT-6 Astra: naturalness 2.59 of 3 (27 ratings) Fable 5.1 2.15 English by Fable 5.1, Turkish by Sonnet 5: naturalness 2.15 of 3 (27 ratings) 2.00 English by Fable 5.1, Turkish by Opus 5: naturalness 2.00 of 3 (27 ratings) 2.70 English by Fable 5.1, Turkish by Opus 5.5: naturalness 2.70 of 3 (27 ratings) 2.33 English by Fable 5.1, Turkish by Fable 5.1: naturalness 2.33 of 3 (27 ratings) 1.74 English by Fable 5.1, Turkish by Grok 4.6: naturalness 1.74 of 3 (27 ratings) 2.15 English by Fable 5.1, Turkish by Grok 4.7: naturalness 2.15 of 3 (27 ratings) 1.89 English by Fable 5.1, Turkish by GPT-6 Luna: naturalness 1.89 of 3 (27 ratings) 2.74 English by Fable 5.1, Turkish by GPT-6 Sol: naturalness 2.74 of 3 (27 ratings) 2.37 English by Fable 5.1, Turkish by GPT-6 Astra: naturalness 2.37 of 3 (27 ratings) Grok 4.6 2.00 English by Grok 4.6, Turkish by Sonnet 5: naturalness 2.00 of 3 (27 ratings) 2.44 English by Grok 4.6, Turkish by Opus 5: naturalness 2.44 of 3 (27 ratings) 2.85 English by Grok 4.6, Turkish by Opus 5.5: naturalness 2.85 of 3 (27 ratings) 2.22 English by Grok 4.6, Turkish by Fable 5.1: naturalness 2.22 of 3 (27 ratings) 1.96 English by Grok 4.6, Turkish by Grok 4.6: naturalness 1.96 of 3 (27 ratings) 1.70 English by Grok 4.6, Turkish by Grok 4.7: naturalness 1.70 of 3 (27 ratings) 2.04 English by Grok 4.6, Turkish by GPT-6 Luna: naturalness 2.04 of 3 (27 ratings) 2.63 English by Grok 4.6, Turkish by GPT-6 Sol: naturalness 2.63 of 3 (27 ratings) 2.74 English by Grok 4.6, Turkish by GPT-6 Astra: naturalness 2.74 of 3 (27 ratings) Grok 4.7 1.70 English by Grok 4.7, Turkish by Sonnet 5: naturalness 1.70 of 3 (27 ratings) 2.41 English by Grok 4.7, Turkish by Opus 5: naturalness 2.41 of 3 (27 ratings) 2.30 English by Grok 4.7, Turkish by Opus 5.5: naturalness 2.30 of 3 (27 ratings) 2.52 English by Grok 4.7, Turkish by Fable 5.1: naturalness 2.52 of 3 (27 ratings) 1.85 English by Grok 4.7, Turkish by Grok 4.6: naturalness 1.85 of 3 (27 ratings) 1.93 English by Grok 4.7, Turkish by Grok 4.7: naturalness 1.93 of 3 (27 ratings) 1.89 English by Grok 4.7, Turkish by GPT-6 Luna: naturalness 1.89 of 3 (27 ratings) 2.56 English by Grok 4.7, Turkish by GPT-6 Sol: naturalness 2.56 of 3 (27 ratings) 2.67 English by Grok 4.7, Turkish by GPT-6 Astra: naturalness 2.67 of 3 (27 ratings) GPT-6 Luna 1.93 English by GPT-6 Luna, Turkish by Sonnet 5: naturalness 1.93 of 3 (27 ratings) 2.31 English by GPT-6 Luna, Turkish by Opus 5: naturalness 2.31 of 3 (27 ratings) 2.57 English by GPT-6 Luna, Turkish by Opus 5.5: naturalness 2.57 of 3 (27 ratings) 2.17 English by GPT-6 Luna, Turkish by Fable 5.1: naturalness 2.17 of 3 (27 ratings) 1.78 English by GPT-6 Luna, Turkish by Grok 4.6: naturalness 1.78 of 3 (27 ratings) 1.74 English by GPT-6 Luna, Turkish by Grok 4.7: naturalness 1.74 of 3 (27 ratings) 2.15 English by GPT-6 Luna, Turkish by GPT-6 Luna: naturalness 2.15 of 3 (27 ratings) 2.39 English by GPT-6 Luna, Turkish by GPT-6 Sol: naturalness 2.39 of 3 (27 ratings) 2.54 English by GPT-6 Luna, Turkish by GPT-6 Astra: naturalness 2.54 of 3 (27 ratings) GPT-6 Sol 1.96 English by GPT-6 Sol, Turkish by Sonnet 5: naturalness 1.96 of 3 (27 ratings) 2.61 English by GPT-6 Sol, Turkish by Opus 5: naturalness 2.61 of 3 (27 ratings) 2.72 English by GPT-6 Sol, Turkish by Opus 5.5: naturalness 2.72 of 3 (27 ratings) 2.50 English by GPT-6 Sol, Turkish by Fable 5.1: naturalness 2.50 of 3 (27 ratings) 1.96 English by GPT-6 Sol, Turkish by Grok 4.6: naturalness 1.96 of 3 (27 ratings) 1.85 English by GPT-6 Sol, Turkish by Grok 4.7: naturalness 1.85 of 3 (27 ratings) 1.96 English by GPT-6 Sol, Turkish by GPT-6 Luna: naturalness 1.96 of 3 (27 ratings) 2.50 English by GPT-6 Sol, Turkish by GPT-6 Sol: naturalness 2.50 of 3 (27 ratings) 2.70 English by GPT-6 Sol, Turkish by GPT-6 Astra: naturalness 2.70 of 3 (27 ratings) GPT-6 Astra 1.96 English by GPT-6 Astra, Turkish by Sonnet 5: naturalness 1.96 of 3 (27 ratings) 2.30 English by GPT-6 Astra, Turkish by Opus 5: naturalness 2.30 of 3 (27 ratings) 2.89 English by GPT-6 Astra, Turkish by Opus 5.5: naturalness 2.89 of 3 (27 ratings) 2.41 English by GPT-6 Astra, Turkish by Fable 5.1: naturalness 2.41 of 3 (27 ratings) 2.07 English by GPT-6 Astra, Turkish by Grok 4.6: naturalness 2.07 of 3 (27 ratings) 1.85 English by GPT-6 Astra, Turkish by Grok 4.7: naturalness 1.85 of 3 (27 ratings) 2.30 English by GPT-6 Astra, Turkish by GPT-6 Luna: naturalness 2.30 of 3 (27 ratings) 2.63 English by GPT-6 Astra, Turkish by GPT-6 Sol: naturalness 2.63 of 3 (27 ratings) 2.52 English by GPT-6 Astra, Turkish by GPT-6 Astra: naturalness 2.52 of 3 (27 ratings) Scale 1 1.5 2 2.5 3

Figure 3 shows all 81 pairs averaged over the three puzzles, with Turkish naturalness beside the total. The full 9 × 9 Turkish grids follow (Tables 7–9). Rows give whose English was translated, columns who translated it; each cell is the nine-rater mean out of 10, and the diagonal is a model translating its own English.

Table 7. Turkish grid, The Seed-Sorting Trays
English bySonnet 5Opus 5Opus 5.5Fable 5.1Grok 4.6Grok 4.7GPT-6 LunaGPT-6 SolGPT-6 Astra
Sonnet 57.787.897.008.336.118.898.119.789.11
Opus 59.118.449.569.566.009.679.118.338.89
Opus 5.57.568.789.229.677.118.118.678.338.22
Fable 5.19.337.679.788.566.788.337.448.788.22
Grok 4.68.339.229.678.788.675.787.339.679.33
Grok 4.77.569.009.229.007.337.115.229.338.78
GPT-6 Luna7.898.729.617.947.444.177.339.228.44
GPT-6 Sol8.009.679.897.568.677.786.339.338.33
GPT-6 Astra8.338.8910.009.009.007.786.899.898.22
Table 8. Turkish grid, The Rope-Line Survey
English bySonnet 5Opus 5Opus 5.5Fable 5.1Grok 4.6Grok 4.7GPT-6 LunaGPT-6 SolGPT-6 Astra
Sonnet 55.789.449.118.006.007.335.567.569.44
Opus 57.116.678.678.446.446.447.338.339.56
Opus 5.57.447.899.678.899.118.227.678.449.56
Fable 5.15.786.229.117.785.678.226.449.119.11
Grok 4.66.567.569.008.116.448.005.898.789.44
Grok 4.76.338.448.788.896.676.447.789.009.67
GPT-6 Luna6.566.228.007.675.568.449.008.449.89
GPT-6 Sol6.679.509.509.567.396.896.789.6110.00
GPT-6 Astra6.337.339.448.447.008.007.678.119.44
Table 9. Turkish grid, The Ropewalk Ledger
English bySonnet 5Opus 5Opus 5.5Fable 5.1Grok 4.6Grok 4.7GPT-6 LunaGPT-6 SolGPT-6 Astra
Sonnet 56.449.228.789.117.117.117.449.568.89
Opus 58.678.679.008.786.897.787.339.009.44
Opus 5.57.119.229.568.335.788.117.229.449.44
Fable 5.17.679.229.678.446.677.787.789.338.11
Grok 4.68.008.679.898.786.897.897.229.569.22
Grok 4.77.569.447.339.228.008.567.118.788.89
GPT-6 Luna7.568.789.448.898.007.567.007.448.78
GPT-6 Sol8.338.789.119.446.897.786.898.899.67
GPT-6 Astra8.119.009.789.568.447.119.789.119.22

3.6 Fidelity (RQ3)

Table 10. Clues judged changed or ambiguous, over the 3 English versions each model wrote and the 27 Turkish translations each model made.
ModelEN changedEN ambiguousTR changedTR ambiguous
Opus 50000
Opus 5.50000
Fable 5.10000
GPT-6 Sol0001
Sonnet 50002
GPT-6 Astra0002
Grok 4.70020
Grok 4.60103
GPT-6 Luna0042

No English rewrite changed a clue's meaning. The six changed verdicts, all in Turkish, were confirmed by hand: Grok 4.7 wrote tepki (reaction) for tepsi (tray) in two clues of one translation, and GPT-6 Luna left an English name in four clues where the Turkish edition uses its own.

3.7 Cost and time (RQ4)

Table 11. Per model, means over its 30 writing calls (English and Turkish).
ModelInput tokensOutput tokensMinutes / callUSD / English versionUSD / Turkish version
Sonnet 52,7613,2380.500.03650.0381
Opus 52,6931,9100.400.05280.0621
Opus 5.52,6981,7960.320.03650.0479
Fable 5.12,6981,6500.370.09030.1116
Grok 4.62,53412,5723.210.07010.0817
Grok 4.72,70416,2324.380.10840.1022
GPT-6 Luna1,4087110.300.00040.0005
GPT-6 Sol1,4081,1950.450.00970.0153
GPT-6 Astra1,4087490.420.03680.0531

The Grok models spent most of their output on hidden reasoning, about twenty times the visible answer, and took roughly six to fifteen times longer per call than the other models. For the GPT-6 models the reasoning tokens are counted inside the reported output tokens, which are billed in full here; only their split between reasoning and answer is not reported. At list prices the 270 versions cost USD 15.16 and the rating USD 27.63. On cost alone, GPT-6 Luna in both roles gives the most points per dollar (7.50 out of 10), but it was the weakest model and made the most meaning errors as a translator. Among the high-scoring pairs, GPT-6 Sol in both roles gives 9.29 at 406 points per dollar, and GPT-6 Astra with GPT-6 Sol 9.35 at 188 (Table 6). Figure 4 sets each model's cost against its score.

Figure 4. Stage 1: cost against quality per model. x: list-price cost of one English version (A) or one Turkish translation (B), USD, log scale, from the tokens each call used (Table 11); y: the model's mean blind score out of 10 in that role (Figure 1; intervals omitted here). Nine models per panel. Source: GRIDIGMA out/results.json (cost) and out/rate/ (scores).
A. English rewrite $0.001 $0.01 $0.1 7 8 9 10 GPT-6 Astra GPT-6 Sol Opus 5.5 Grok 4.6 Fable 5.1 Opus 5 Grok 4.7 Sonnet 5 GPT-6 Luna GPT-6 Astra: English 9.67 at USD 0.0368 per version GPT-6 Sol: English 9.30 at USD 0.0097 per version Opus 5.5: English 9.26 at USD 0.0365 per version Grok 4.6: English 9.19 at USD 0.0701 per version Fable 5.1: English 8.89 at USD 0.0903 per version Opus 5: English 8.59 at USD 0.0528 per version Grok 4.7: English 8.48 at USD 0.1084 per version Sonnet 5: English 7.52 at USD 0.0365 per version GPT-6 Luna: English 7.22 at USD 0.0004 per version USD per version, list price (log scale) B. Turkish translation $0.001 $0.01 $0.1 7 8 9 10 Opus 5.5 GPT-6 Astra GPT-6 Sol Fable 5.1 Opus 5 Grok 4.7 Sonnet 5 GPT-6 Luna Grok 4.6 GPT-6 Astra: Turkish 9.09 at USD 0.0531 per version GPT-6 Sol: Turkish 8.93 at USD 0.0153 per version Opus 5.5: Turkish 9.18 at USD 0.0479 per version Grok 4.6: Turkish 7.11 at USD 0.0817 per version Fable 5.1: Turkish 8.69 at USD 0.1116 per version Opus 5: Turkish 8.47 at USD 0.0621 per version Grok 4.7: Turkish 7.60 at USD 0.1022 per version Sonnet 5: Turkish 7.48 at USD 0.0381 per version GPT-6 Luna: Turkish 7.35 at USD 0.0005 per version USD per version, list price (log scale) Mean score out of 10

4. Discussion

Two roles, two different models. English writing and Turkish translation rewarded different models. The best English came from GPT-6 Astra and GPT-6 Sol, the best Turkish from Opus 5.5, and the best pair combined them. A pipeline that uses one model for both roles gives up quality in one of them.

The translator decides naturalness. Naturalness spread about five times wider across translators (0.80 points) than across English sources (0.16 points). A well-written source helps clarity, but it does not make a weak translator sound native. For a product in several languages, the translator is the choice to get right.

Fluency without drift. None of the 27 English rewrites changed a clue's meaning, and the best models were both the most fluent and error-free. Here, readability did not cost correctness. The Turkish errors that did occur were lexical slips (tepki for tepsi) and untranslated names, not reasoning failures, and a mechanical check against the stored logic catches both.

Reasoning volume is not quality. The Grok models reasoned the most, cost the most time, and gave the weakest Turkish. At default settings, more hidden reasoning did not buy better prose.

Model judges. Agreement among the nine raters was high, and self-ratings moved no mean by more than 0.13, so the self-preference reported in other settings [3] did not show up measurably here. We attribute this partly to the blind, shuffled batches and to a rubric fixed in advance. Agreement among models is not agreement with people, however, and the native-speaker ratings planned as an addendum have to test that directly.

5. Limitations

  • Small sample. Three puzzles, all 4 × 4 grids from one game. The English intervals rest on three versions per model and do not separate the top four writers.
  • No human rating. All raters were models, and none is a native Turkish speaker. So far the only native Turkish reader of the outputs is the author. The absolute numbers are machine judgements.
  • The fidelity checker was also a contestant. Opus 5.5 ran the clue check and scored best on it. Every changed verdict was confirmed by hand, but same verdicts were only spot-checked (26 English clues, all agreeing).
  • Unmatched settings in Stage 1. Every model ran at its provider's default reasoning effort, and these differ (Section 2.7). Stage 2 matched them; see Section 7.3 for what that changed.
  • Batch composition. A Turkish batch held one English source's nine translations, so raters never compared translations of different sources side by side.
  • Plain text, not production format. GRIDIGMA's authoring format carries markup that this trial left out.
  • Single generation. Each version was generated once; variation across repeated generations was not measured.

6. Conclusion

For English-and-Turkish text in a logic-puzzle game, the highest-rated combination was GPT-6 Astra writing the English and Opus 5.5 translating it into Turkish (9.70/10, no clue flagged, first in 81% of resamples). Every model's rewrite outscored the text in use. Translation quality depended far more on the translator than on the source. Across 21 languages at a matched effort (Stage 2), Opus 5.5 and GPT-6 Astra were statistically tied as the best translators and led in 17 of the 21 languages between them; neither Grok model, nor GPT-6 Luna, nor Sonnet 5 led in any language. Based on Stage 1, GRIDIGMA decided not to use either Grok model for Turkish translation, and is evaluating a split pipeline with Opus 5.5 as the Turkish translator, piloted on the production format first. Stages 2b and 2c tested what the product owner found in Stage 2's source, technical words and a formula clue. On plain words, and under a rubric that weights naturalness first, Opus 5.5 and GPT-6 Astra remain the two strongest translators (GPT-6 Astra 10, Opus 5.5 9 of 21 languages in Stage 2c), and the route from the stored logic is as good as the route from English on average. Per-language choices from one puzzle and one generation remain fragile.

7. Stage 2: twenty-one languages (RQ5)

7.1 Design and protocol changes

Stage 2 asks whether Stage 1's translators keep their ranking across languages when every model runs at the same reasoning effort. One puzzle, The Rope-Line Survey, was translated from a single English source (GPT-6 Astra's Stage 1 rewrite, unchanged) by all nine models into 21 languages: German, Japanese, French, Korean, Traditional and Simplified Chinese, Dutch, Spanish for Spain and for Latin America, Italian, Swedish, Norwegian (Bokmål), Danish, Finnish, Polish, Czech, Hebrew, Arabic for the Gulf countries, Brazilian Portuguese, Russian, and Turkish again, as the row that links the two stages. That gave 189 translations. For each language, all nine models rated its nine translations blind in one batch, with self-ratings included and the rubric of Stage 1 (1,701 ratings); the Turkish batch also carried the live Turkish text as a control. Opus 5.5 checked every clue of every translation, as in Stage 1. The run took place on 24 September 2026.

The protocol changed deliberately in five respects:

  1. Matched effort. Every call ran at reasoning effort high, set explicitly and recorded per call with the model identifier, the time and the full token usage. Before the run, each provider was shown to act on the setting: an invalid value was rejected by the OpenAI and xAI interfaces, and output grew between low and high for every provider.
  2. Clue order as in the game. The clue "exactly one of notes A, B and C is false" came first, as players see it; Stage 1 had it eighth.
  3. Names. Outside Turkish, each translator named the four people for its own readers; Turkish kept Stage 1's Turkish cast.
  4. One source text instead of nine, and one puzzle instead of three.
  5. Corrected wording in the fidelity prompt, which in Stage 1 stated the number of translations wrongly (eight for nine); the check itself was unchanged.

Because Stage 1's defaults were already high for Sonnet 5, Opus 5, Fable 5.1 and both Grok models (Section 2.7), the effort change between the stages affected only Opus 5.5 and the three GPT-6 models.

7.2 Results across all 21 languages

Raters agreed as in Stage 1: Kendall's W averaged 0.67 over the 21 language batches (range 0.44–0.84), and ICC(2,1) was 0.54 for a single rater and ICC(2,k) 0.91 for the mean of nine, over the 189 translations.

Figure 5. Stage 2: mean blind score out of 10 for each translator over its 21 translations (21 × 9 = 189 ratings), with 95% confidence intervals from a two-way cluster bootstrap resampling raters and languages (4,000 resamples, seed 20260925, as stats2.py). Source: GRIDIGMA out/s2/rate/.
5 6 7 8 9 10 Opus 5.5 Opus 5.5: 9.28 (95% CI 8.90 to 9.57) GPT-6 Astra GPT-6 Astra: 9.20 (95% CI 8.80 to 9.54) GPT-6 Sol GPT-6 Sol: 8.39 (95% CI 7.90 to 8.86) Fable 5.1 Fable 5.1: 8.35 (95% CI 7.88 to 8.77) Opus 5 Opus 5: 8.28 (95% CI 7.77 to 8.79) Grok 4.6 Grok 4.6: 7.63 (95% CI 6.98 to 8.21) Grok 4.7 Grok 4.7: 7.57 (95% CI 6.96 to 8.16) GPT-6 Luna GPT-6 Luna: 7.36 (95% CI 6.79 to 7.93) Sonnet 5 Sonnet 5: 7.10 (95% CI 6.52 to 7.68) Mean blind score out of 10 (Stage 1 rubric)
Table 12. Stage 2 per translator, over 21 languages. "Best in": the number of languages in which it scored highest. "Flagged": clues judged changed / ambiguous. Cost at list price per translation.
TranslatorScore [95% CI]No selfClarityReadabilityNaturalnessBest inFlaggedUSDMinutes
Opus 5.59.28 [8.90, 9.57]9.263.702.852.72100 / 20.0570.4
GPT-6 Astra9.20 [8.80, 9.54]9.133.862.772.5870 / 00.0760.7
GPT-6 Sol8.39 [7.90, 8.86]8.403.402.672.3310 / 10.0220.7
Fable 5.18.35 [7.88, 8.77]8.293.432.642.2720 / 00.1270.4
Opus 58.28 [7.77, 8.79]8.253.392.672.2310 / 00.0670.5
Grok 4.67.63 [6.98, 8.21]7.653.202.471.9600 / 00.0833.5
Grok 4.77.57 [6.96, 8.16]7.513.292.341.9400 / 00.1274.5
GPT-6 Luna7.36 [6.79, 7.93]7.383.142.361.8601 / 10.00090.6
Sonnet 57.10 [6.52, 7.68]6.953.202.161.7400 / 00.0420.5

A tie at the top. Opus 5.5 and GPT-6 Astra cannot be separated: in 1,000 resamples of raters and languages, Opus 5.5 ranked first 612 times and GPT-6 Astra 388 times, and no other model ever ranked first. Both lie clear of the other seven. Removing self-ratings moved no translator's mean by more than 0.15; Figure 9 shows how each rater scored each translator against the other eight. Figure 6 gives every translation's score, Figure 7 each translator's criteria and Figure 8 the spread of single ratings behind each mean.

Figure 6. Stage 2: every translation's score, by language (rows) and translator (columns, ordered by overall mean). Each cell is the mean of the nine raters' totals out of 10 for one translation, shown to one decimal; the frame marks the highest cell in the row (ties framed together). Darker = higher. One translation per cell, generated once. Source: GRIDIGMA out/s2/rate/.
Opus 5.5 GPT-6 Astra GPT-6 Sol Fable 5.1 Opus 5 Grok 4.6 Grok 4.7 GPT-6 Luna Sonnet 5 German 8.9 German, Opus 5.5: 8.89 of 10 (9 raters) 8.3 German, GPT-6 Astra: 8.33 of 10 (9 raters) 9.1 German, GPT-6 Sol: 9.11 of 10 (9 raters) 7.7 German, Fable 5.1: 7.67 of 10 (9 raters) 8.6 German, Opus 5: 8.61 of 10 (9 raters) 6.6 German, Grok 4.6: 6.56 of 10 (9 raters) 6.8 German, Grok 4.7: 6.83 of 10 (9 raters) 6.8 German, GPT-6 Luna: 6.83 of 10 (9 raters) 6.4 German, Sonnet 5: 6.44 of 10 (9 raters) Japanese 9.1 Japanese, Opus 5.5: 9.11 of 10 (9 raters) 9.9 Japanese, GPT-6 Astra: 9.89 of 10 (9 raters) 7.8 Japanese, GPT-6 Sol: 7.78 of 10 (9 raters) 9.6 Japanese, Fable 5.1: 9.56 of 10 (9 raters) 9.4 Japanese, Opus 5: 9.44 of 10 (9 raters) 8.3 Japanese, Grok 4.6: 8.33 of 10 (9 raters) 7.2 Japanese, Grok 4.7: 7.22 of 10 (9 raters) 7.3 Japanese, GPT-6 Luna: 7.33 of 10 (9 raters) 5.9 Japanese, Sonnet 5: 5.89 of 10 (9 raters) French 9.6 French, Opus 5.5: 9.61 of 10 (9 raters) 10.0 French, GPT-6 Astra: 10.00 of 10 (9 raters) 9.6 French, GPT-6 Sol: 9.56 of 10 (9 raters) 8.2 French, Fable 5.1: 8.17 of 10 (9 raters) 9.3 French, Opus 5: 9.28 of 10 (9 raters) 6.8 French, Grok 4.6: 6.78 of 10 (9 raters) 6.7 French, Grok 4.7: 6.67 of 10 (9 raters) 7.2 French, GPT-6 Luna: 7.22 of 10 (9 raters) 7.6 French, Sonnet 5: 7.56 of 10 (9 raters) Korean 9.4 Korean, Opus 5.5: 9.39 of 10 (9 raters) 9.3 Korean, GPT-6 Astra: 9.33 of 10 (9 raters) 6.7 Korean, GPT-6 Sol: 6.72 of 10 (9 raters) 8.6 Korean, Fable 5.1: 8.56 of 10 (9 raters) 7.7 Korean, Opus 5: 7.67 of 10 (9 raters) 5.9 Korean, Grok 4.6: 5.89 of 10 (9 raters) 7.3 Korean, Grok 4.7: 7.28 of 10 (9 raters) 7.7 Korean, GPT-6 Luna: 7.72 of 10 (9 raters) 7.3 Korean, Sonnet 5: 7.28 of 10 (9 raters) Chinese (Traditional) 7.1 Chinese (Traditional), Opus 5.5: 7.11 of 10 (9 raters) 9.1 Chinese (Traditional), GPT-6 Astra: 9.11 of 10 (9 raters) 9.0 Chinese (Traditional), GPT-6 Sol: 9.00 of 10 (9 raters) 8.4 Chinese (Traditional), Fable 5.1: 8.44 of 10 (9 raters) 9.3 Chinese (Traditional), Opus 5: 9.33 of 10 (9 raters) 9.2 Chinese (Traditional), Grok 4.6: 9.22 of 10 (9 raters) 8.1 Chinese (Traditional), Grok 4.7: 8.11 of 10 (9 raters) 7.8 Chinese (Traditional), GPT-6 Luna: 7.78 of 10 (9 raters) 7.7 Chinese (Traditional), Sonnet 5: 7.67 of 10 (9 raters) Chinese (Simplified) 9.5 Chinese (Simplified), Opus 5.5: 9.50 of 10 (9 raters) 9.4 Chinese (Simplified), GPT-6 Astra: 9.38 of 10 (9 raters) 9.0 Chinese (Simplified), GPT-6 Sol: 9.01 of 10 (9 raters) 8.9 Chinese (Simplified), Fable 5.1: 8.90 of 10 (9 raters) 9.0 Chinese (Simplified), Opus 5: 9.00 of 10 (9 raters) 8.9 Chinese (Simplified), Grok 4.6: 8.94 of 10 (9 raters) 9.2 Chinese (Simplified), Grok 4.7: 9.21 of 10 (9 raters) 7.5 Chinese (Simplified), GPT-6 Luna: 7.50 of 10 (9 raters) 7.3 Chinese (Simplified), Sonnet 5: 7.28 of 10 (9 raters) Dutch 9.3 Dutch, Opus 5.5: 9.33 of 10 (9 raters) 9.6 Dutch, GPT-6 Astra: 9.56 of 10 (9 raters) 6.7 Dutch, GPT-6 Sol: 6.67 of 10 (9 raters) 8.4 Dutch, Fable 5.1: 8.44 of 10 (9 raters) 8.3 Dutch, Opus 5: 8.33 of 10 (9 raters) 7.1 Dutch, Grok 4.6: 7.11 of 10 (9 raters) 6.3 Dutch, Grok 4.7: 6.33 of 10 (9 raters) 8.9 Dutch, GPT-6 Luna: 8.89 of 10 (9 raters) 7.6 Dutch, Sonnet 5: 7.56 of 10 (9 raters) Spanish (Spain) 9.2 Spanish (Spain), Opus 5.5: 9.17 of 10 (9 raters) 9.4 Spanish (Spain), GPT-6 Astra: 9.44 of 10 (9 raters) 7.4 Spanish (Spain), GPT-6 Sol: 7.39 of 10 (9 raters) 9.6 Spanish (Spain), Fable 5.1: 9.56 of 10 (9 raters) 7.5 Spanish (Spain), Opus 5: 7.50 of 10 (9 raters) 9.0 Spanish (Spain), Grok 4.6: 9.00 of 10 (9 raters) 7.4 Spanish (Spain), Grok 4.7: 7.39 of 10 (9 raters) 8.4 Spanish (Spain), GPT-6 Luna: 8.39 of 10 (9 raters) 8.6 Spanish (Spain), Sonnet 5: 8.61 of 10 (9 raters) Spanish (Latin America) 9.7 Spanish (Latin America), Opus 5.5: 9.67 of 10 (9 raters) 10.0 Spanish (Latin America), GPT-6 Astra: 10.00 of 10 (9 raters) 8.2 Spanish (Latin America), GPT-6 Sol: 8.17 of 10 (9 raters) 7.1 Spanish (Latin America), Fable 5.1: 7.06 of 10 (9 raters) 7.8 Spanish (Latin America), Opus 5: 7.83 of 10 (9 raters) 6.7 Spanish (Latin America), Grok 4.6: 6.67 of 10 (9 raters) 8.6 Spanish (Latin America), Grok 4.7: 8.61 of 10 (9 raters) 9.1 Spanish (Latin America), GPT-6 Luna: 9.06 of 10 (9 raters) 7.6 Spanish (Latin America), Sonnet 5: 7.56 of 10 (9 raters) Italian 9.4 Italian, Opus 5.5: 9.44 of 10 (9 raters) 9.9 Italian, GPT-6 Astra: 9.89 of 10 (9 raters) 9.4 Italian, GPT-6 Sol: 9.44 of 10 (9 raters) 7.2 Italian, Fable 5.1: 7.22 of 10 (9 raters) 8.4 Italian, Opus 5: 8.39 of 10 (9 raters) 6.1 Italian, Grok 4.6: 6.06 of 10 (9 raters) 5.6 Italian, Grok 4.7: 5.61 of 10 (9 raters) 8.7 Italian, GPT-6 Luna: 8.67 of 10 (9 raters) 8.3 Italian, Sonnet 5: 8.28 of 10 (9 raters) Swedish 8.3 Swedish, Opus 5.5: 8.33 of 10 (9 raters) 9.0 Swedish, GPT-6 Astra: 9.00 of 10 (9 raters) 8.8 Swedish, GPT-6 Sol: 8.78 of 10 (9 raters) 8.8 Swedish, Fable 5.1: 8.78 of 10 (9 raters) 7.2 Swedish, Opus 5: 7.22 of 10 (9 raters) 7.2 Swedish, Grok 4.6: 7.22 of 10 (9 raters) 8.8 Swedish, Grok 4.7: 8.78 of 10 (9 raters) 8.3 Swedish, GPT-6 Luna: 8.33 of 10 (9 raters) 7.3 Swedish, Sonnet 5: 7.33 of 10 (9 raters) Norwegian 10.0 Norwegian, Opus 5.5: 10.00 of 10 (9 raters) 9.6 Norwegian, GPT-6 Astra: 9.56 of 10 (9 raters) 8.2 Norwegian, GPT-6 Sol: 8.22 of 10 (9 raters) 7.8 Norwegian, Fable 5.1: 7.78 of 10 (9 raters) 7.8 Norwegian, Opus 5: 7.78 of 10 (9 raters) 6.6 Norwegian, Grok 4.6: 6.56 of 10 (9 raters) 7.7 Norwegian, Grok 4.7: 7.67 of 10 (9 raters) 6.1 Norwegian, GPT-6 Luna: 6.11 of 10 (9 raters) 6.2 Norwegian, Sonnet 5: 6.22 of 10 (9 raters) Danish 9.1 Danish, Opus 5.5: 9.11 of 10 (9 raters) 9.6 Danish, GPT-6 Astra: 9.56 of 10 (9 raters) 9.3 Danish, GPT-6 Sol: 9.33 of 10 (9 raters) 8.6 Danish, Fable 5.1: 8.56 of 10 (9 raters) 8.9 Danish, Opus 5: 8.89 of 10 (9 raters) 6.0 Danish, Grok 4.6: 6.00 of 10 (9 raters) 6.9 Danish, Grok 4.7: 6.89 of 10 (9 raters) 6.3 Danish, GPT-6 Luna: 6.33 of 10 (9 raters) 5.2 Danish, Sonnet 5: 5.22 of 10 (9 raters) Finnish 9.9 Finnish, Opus 5.5: 9.89 of 10 (9 raters) 8.0 Finnish, GPT-6 Astra: 8.00 of 10 (9 raters) 9.7 Finnish, GPT-6 Sol: 9.67 of 10 (9 raters) 7.9 Finnish, Fable 5.1: 7.89 of 10 (9 raters) 8.3 Finnish, Opus 5: 8.33 of 10 (9 raters) 9.3 Finnish, Grok 4.6: 9.33 of 10 (9 raters) 9.0 Finnish, Grok 4.7: 9.00 of 10 (9 raters) 6.9 Finnish, GPT-6 Luna: 6.89 of 10 (9 raters) 5.9 Finnish, Sonnet 5: 5.89 of 10 (9 raters) Polish 9.7 Polish, Opus 5.5: 9.67 of 10 (9 raters) 8.8 Polish, GPT-6 Astra: 8.78 of 10 (9 raters) 8.7 Polish, GPT-6 Sol: 8.67 of 10 (9 raters) 7.9 Polish, Fable 5.1: 7.89 of 10 (9 raters) 8.2 Polish, Opus 5: 8.22 of 10 (9 raters) 8.2 Polish, Grok 4.6: 8.22 of 10 (9 raters) 7.8 Polish, Grok 4.7: 7.78 of 10 (9 raters) 5.2 Polish, GPT-6 Luna: 5.22 of 10 (9 raters) 6.3 Polish, Sonnet 5: 6.33 of 10 (9 raters) Czech 9.6 Czech, Opus 5.5: 9.56 of 10 (9 raters) 8.3 Czech, GPT-6 Astra: 8.33 of 10 (9 raters) 7.8 Czech, GPT-6 Sol: 7.78 of 10 (9 raters) 9.2 Czech, Fable 5.1: 9.22 of 10 (9 raters) 6.6 Czech, Opus 5: 6.56 of 10 (9 raters) 9.0 Czech, Grok 4.6: 9.00 of 10 (9 raters) 7.3 Czech, Grok 4.7: 7.33 of 10 (9 raters) 6.8 Czech, GPT-6 Luna: 6.78 of 10 (9 raters) 7.3 Czech, Sonnet 5: 7.33 of 10 (9 raters) Hebrew 9.7 Hebrew, Opus 5.5: 9.67 of 10 (9 raters) 8.8 Hebrew, GPT-6 Astra: 8.83 of 10 (9 raters) 7.7 Hebrew, GPT-6 Sol: 7.67 of 10 (9 raters) 6.4 Hebrew, Fable 5.1: 6.44 of 10 (9 raters) 7.9 Hebrew, Opus 5: 7.89 of 10 (9 raters) 7.9 Hebrew, Grok 4.6: 7.89 of 10 (9 raters) 8.6 Hebrew, Grok 4.7: 8.61 of 10 (9 raters) 7.7 Hebrew, GPT-6 Luna: 7.72 of 10 (9 raters) 6.8 Hebrew, Sonnet 5: 6.83 of 10 (9 raters) Arabic (Gulf) 9.6 Arabic (Gulf), Opus 5.5: 9.56 of 10 (9 raters) 8.4 Arabic (Gulf), GPT-6 Astra: 8.44 of 10 (9 raters) 7.6 Arabic (Gulf), GPT-6 Sol: 7.56 of 10 (9 raters) 8.6 Arabic (Gulf), Fable 5.1: 8.56 of 10 (9 raters) 8.2 Arabic (Gulf), Opus 5: 8.22 of 10 (9 raters) 8.2 Arabic (Gulf), Grok 4.6: 8.22 of 10 (9 raters) 7.0 Arabic (Gulf), Grok 4.7: 7.00 of 10 (9 raters) 5.9 Arabic (Gulf), GPT-6 Luna: 5.89 of 10 (9 raters) 7.6 Arabic (Gulf), Sonnet 5: 7.56 of 10 (9 raters) Portuguese (Brazil) 9.3 Portuguese (Brazil), Opus 5.5: 9.33 of 10 (9 raters) 9.1 Portuguese (Brazil), GPT-6 Astra: 9.06 of 10 (9 raters) 8.8 Portuguese (Brazil), GPT-6 Sol: 8.77 of 10 (9 raters) 8.3 Portuguese (Brazil), Fable 5.1: 8.26 of 10 (9 raters) 8.3 Portuguese (Brazil), Opus 5: 8.26 of 10 (9 raters) 9.0 Portuguese (Brazil), Grok 4.6: 8.97 of 10 (9 raters) 7.5 Portuguese (Brazil), Grok 4.7: 7.47 of 10 (9 raters) 6.7 Portuguese (Brazil), GPT-6 Luna: 6.69 of 10 (9 raters) 7.9 Portuguese (Brazil), Sonnet 5: 7.88 of 10 (9 raters) Russian 9.6 Russian, Opus 5.5: 9.56 of 10 (9 raters) 9.4 Russian, GPT-6 Astra: 9.44 of 10 (9 raters) 8.7 Russian, GPT-6 Sol: 8.67 of 10 (9 raters) 8.8 Russian, Fable 5.1: 8.78 of 10 (9 raters) 7.8 Russian, Opus 5: 7.78 of 10 (9 raters) 6.0 Russian, Grok 4.6: 6.00 of 10 (9 raters) 7.4 Russian, Grok 4.7: 7.44 of 10 (9 raters) 7.1 Russian, GPT-6 Luna: 7.11 of 10 (9 raters) 7.1 Russian, Sonnet 5: 7.11 of 10 (9 raters) Turkish 8.9 Turkish, Opus 5.5: 8.89 of 10 (9 raters) 9.3 Turkish, GPT-6 Astra: 9.33 of 10 (9 raters) 8.3 Turkish, GPT-6 Sol: 8.33 of 10 (9 raters) 9.6 Turkish, Fable 5.1: 9.56 of 10 (9 raters) 9.4 Turkish, Opus 5: 9.44 of 10 (9 raters) 8.2 Turkish, Grok 4.6: 8.22 of 10 (9 raters) 7.7 Turkish, Grok 4.7: 7.67 of 10 (9 raters) 8.1 Turkish, GPT-6 Luna: 8.11 of 10 (9 raters) 7.2 Turkish, Sonnet 5: 7.22 of 10 (9 raters) Scale 5 6 7 8 9 10

Rankings travel, with exceptions. A language's own ranking of the nine translators agreed with the overall ranking with a median Spearman correlation of 0.72, highest in Norwegian (0.97) and lowest in Traditional Chinese (0.15), where Opus 5.5 scored its lowest (7.11) because two of its clues admitted a second reading. Table 13 lists the best translator in each language.

Table 13. The highest-scoring translator in each language, and the runner-up. Scores are relative within a language and are not comparable across languages (Section 7.5).
LanguageBestScoreRunner-upScore
GermanGPT-6 Sol9.11Opus 5.58.89
JapaneseGPT-6 Astra9.89Fable 5.19.56
FrenchGPT-6 Astra10.00Opus 5.59.61
KoreanOpus 5.59.39GPT-6 Astra9.33
Chinese (Traditional)Opus 59.33Grok 4.69.22
Chinese (Simplified)Opus 5.59.50GPT-6 Astra9.38
DutchGPT-6 Astra9.56Opus 5.59.33
Spanish (Spain)Fable 5.19.56GPT-6 Astra9.44
Spanish (Latin America)GPT-6 Astra10.00Opus 5.59.67
ItalianGPT-6 Astra9.89Opus 5.59.44
SwedishGPT-6 Astra9.00three tied8.78
NorwegianOpus 5.510.00GPT-6 Astra9.56
DanishGPT-6 Astra9.56GPT-6 Sol9.33
FinnishOpus 5.59.89GPT-6 Sol9.67
PolishOpus 5.59.67GPT-6 Astra8.78
CzechOpus 5.59.56Fable 5.19.22
HebrewOpus 5.59.67GPT-6 Astra8.83
Arabic (Gulf)Opus 5.59.56Fable 5.18.56
Portuguese (Brazil)Opus 5.59.33GPT-6 Astra9.06
RussianOpus 5.59.56GPT-6 Astra9.44
TurkishFable 5.19.56Opus 59.44
Figure 7. Stage 2: each translator's mean score on each criterion as a share of that criterion's maximum (clarity 4, readability 3, naturalness 3), over its 189 ratings. Naturalness separates the translators most. Point estimates; the raw means are in Table 12. Source: GRIDIGMA out/s2/rate/.
Clarity (of 4) Readability (of 3) Naturalness (of 3) 50% 60% 70% 80% 90% 100% Opus 5.5 Opus 5.5: clarity 3.70 of 4 (93%) Opus 5.5: readability 2.85 of 3 (95%) Opus 5.5: naturalness 2.72 of 3 (91%) GPT-6 Astra GPT-6 Astra: clarity 3.86 of 4 (96%) GPT-6 Astra: readability 2.77 of 3 (92%) GPT-6 Astra: naturalness 2.58 of 3 (86%) GPT-6 Sol GPT-6 Sol: clarity 3.40 of 4 (85%) GPT-6 Sol: readability 2.67 of 3 (89%) GPT-6 Sol: naturalness 2.33 of 3 (78%) Fable 5.1 Fable 5.1: clarity 3.43 of 4 (86%) Fable 5.1: readability 2.64 of 3 (88%) Fable 5.1: naturalness 2.27 of 3 (76%) Opus 5 Opus 5: clarity 3.39 of 4 (85%) Opus 5: readability 2.67 of 3 (89%) Opus 5: naturalness 2.23 of 3 (74%) Grok 4.6 Grok 4.6: clarity 3.20 of 4 (80%) Grok 4.6: readability 2.47 of 3 (82%) Grok 4.6: naturalness 1.96 of 3 (65%) Grok 4.7 Grok 4.7: clarity 3.29 of 4 (82%) Grok 4.7: readability 2.34 of 3 (78%) Grok 4.7: naturalness 1.94 of 3 (65%) GPT-6 Luna GPT-6 Luna: clarity 3.14 of 4 (79%) GPT-6 Luna: readability 2.36 of 3 (79%) GPT-6 Luna: naturalness 1.86 of 3 (62%) Sonnet 5 Sonnet 5: clarity 3.20 of 4 (80%) Sonnet 5: readability 2.16 of 3 (72%) Sonnet 5: naturalness 1.74 of 3 (58%) Mean criterion score, % of its maximum
Figure 8. Stage 2: the distribution behind each translator's mean. Each circle's area is proportional to the number of its 189 ratings (21 languages × 9 raters) at that total; totals off the whole numbers come from GPT-6 raters that gave half points or, rarely, tenths. Diamonds and the right-hand column give the mean. Source: GRIDIGMA out/s2/rate/.
Ratings at that score (area = count) Mean 3 4 5 6 7 8 9 10 Mean Opus 5.5 Opus 5.5: 1 of 189 ratings gave 5 Opus 5.5: 2 of 189 ratings gave 6 Opus 5.5: 9 of 189 ratings gave 7 Opus 5.5: 15 of 189 ratings gave 8 Opus 5.5: 64 of 189 ratings gave 9 Opus 5.5: 6 of 189 ratings gave 9.5 Opus 5.5: 92 of 189 ratings gave 10 Opus 5.5: mean 9.28 9.28 GPT-6 Astra GPT-6 Astra: 2 of 189 ratings gave 6 GPT-6 Astra: 11 of 189 ratings gave 7 GPT-6 Astra: 33 of 189 ratings gave 8 GPT-6 Astra: 2 of 189 ratings gave 8.5 GPT-6 Astra: 1 of 189 ratings gave 8.9 GPT-6 Astra: 38 of 189 ratings gave 9 GPT-6 Astra: 3 of 189 ratings gave 9.5 GPT-6 Astra: 99 of 189 ratings gave 10 GPT-6 Astra: mean 9.20 9.20 GPT-6 Sol GPT-6 Sol: 3 of 189 ratings gave 5 GPT-6 Sol: 8 of 189 ratings gave 6 GPT-6 Sol: 34 of 189 ratings gave 7 GPT-6 Sol: 1 of 189 ratings gave 7.5 GPT-6 Sol: 48 of 189 ratings gave 8 GPT-6 Sol: 1 of 189 ratings gave 8.1 GPT-6 Sol: 1 of 189 ratings gave 8.5 GPT-6 Sol: 52 of 189 ratings gave 9 GPT-6 Sol: 1 of 189 ratings gave 9.5 GPT-6 Sol: 1 of 189 ratings gave 9.9 GPT-6 Sol: 39 of 189 ratings gave 10 GPT-6 Sol: mean 8.39 8.39 Fable 5.1 Fable 5.1: 1 of 189 ratings gave 4 Fable 5.1: 14 of 189 ratings gave 6 Fable 5.1: 33 of 189 ratings gave 7 Fable 5.1: 44 of 189 ratings gave 8 Fable 5.1: 4 of 189 ratings gave 8.5 Fable 5.1: 1 of 189 ratings gave 8.6 Fable 5.1: 55 of 189 ratings gave 9 Fable 5.1: 1 of 189 ratings gave 9.3 Fable 5.1: 1 of 189 ratings gave 9.5 Fable 5.1: 35 of 189 ratings gave 10 Fable 5.1: mean 8.35 8.35 Opus 5 Opus 5: 3 of 189 ratings gave 5 Opus 5: 11 of 189 ratings gave 6 Opus 5: 1 of 189 ratings gave 6.5 Opus 5: 39 of 189 ratings gave 7 Opus 5: 1 of 189 ratings gave 7.5 Opus 5: 42 of 189 ratings gave 8 Opus 5: 3 of 189 ratings gave 8.5 Opus 5: 51 of 189 ratings gave 9 Opus 5: 1 of 189 ratings gave 9.3 Opus 5: 4 of 189 ratings gave 9.5 Opus 5: 33 of 189 ratings gave 10 Opus 5: mean 8.28 8.28 Grok 4.6 Grok 4.6: 1 of 189 ratings gave 3 Grok 4.6: 2 of 189 ratings gave 4 Grok 4.6: 18 of 189 ratings gave 5 Grok 4.6: 31 of 189 ratings gave 6 Grok 4.6: 26 of 189 ratings gave 7 Grok 4.6: 2 of 189 ratings gave 7.5 Grok 4.6: 43 of 189 ratings gave 8 Grok 4.6: 2 of 189 ratings gave 8.5 Grok 4.6: 43 of 189 ratings gave 9 Grok 4.6: 1 of 189 ratings gave 9.7 Grok 4.6: 20 of 189 ratings gave 10 Grok 4.6: mean 7.63 7.63 Grok 4.7 Grok 4.7: 1 of 189 ratings gave 3 Grok 4.7: 9 of 189 ratings gave 5 Grok 4.7: 33 of 189 ratings gave 6 Grok 4.7: 2 of 189 ratings gave 6.5 Grok 4.7: 50 of 189 ratings gave 7 Grok 4.7: 2 of 189 ratings gave 7.5 Grok 4.7: 36 of 189 ratings gave 8 Grok 4.7: 4 of 189 ratings gave 8.5 Grok 4.7: 1 of 189 ratings gave 8.9 Grok 4.7: 34 of 189 ratings gave 9 Grok 4.7: 1 of 189 ratings gave 9.2 Grok 4.7: 16 of 189 ratings gave 10 Grok 4.7: mean 7.57 7.57 GPT-6 Luna GPT-6 Luna: 6 of 189 ratings gave 4 GPT-6 Luna: 14 of 189 ratings gave 5 GPT-6 Luna: 2 of 189 ratings gave 5.5 GPT-6 Luna: 30 of 189 ratings gave 6 GPT-6 Luna: 1 of 189 ratings gave 6.5 GPT-6 Luna: 42 of 189 ratings gave 7 GPT-6 Luna: 1 of 189 ratings gave 7.5 GPT-6 Luna: 47 of 189 ratings gave 8 GPT-6 Luna: 1 of 189 ratings gave 8.2 GPT-6 Luna: 2 of 189 ratings gave 8.5 GPT-6 Luna: 32 of 189 ratings gave 9 GPT-6 Luna: 2 of 189 ratings gave 9.5 GPT-6 Luna: 9 of 189 ratings gave 10 GPT-6 Luna: mean 7.36 7.36 Sonnet 5 Sonnet 5: 4 of 189 ratings gave 4 Sonnet 5: 14 of 189 ratings gave 5 Sonnet 5: 45 of 189 ratings gave 6 Sonnet 5: 54 of 189 ratings gave 7 Sonnet 5: 4 of 189 ratings gave 7.5 Sonnet 5: 37 of 189 ratings gave 8 Sonnet 5: 2 of 189 ratings gave 8.5 Sonnet 5: 1 of 189 ratings gave 8.9 Sonnet 5: 24 of 189 ratings gave 9 Sonnet 5: 1 of 189 ratings gave 9.5 Sonnet 5: 3 of 189 ratings gave 10 Sonnet 5: mean 7.10 7.10 One rater's total for one translation, out of 10
Figure 9. Stage 2: how each rater (rows) scored each translator (columns) against the other eight raters. Cell: the rater's mean over the translator's 21 translations minus the other eight raters' mean for the same translations, in points out of 10; blue = more generous, orange = stricter. The outlined diagonal is self-rating. The last column is the rater's mean over all 189 translations (its overall severity). Rows and columns in the order of Figure 5. Source: GRIDIGMA out/s2/rate/.
Translator Opus 5.5 GPT-6 Astra GPT-6 Sol Fable 5.1 Opus 5 Grok 4.6 Grok 4.7 GPT-6 Luna Sonnet 5 Mean given Opus 5.5 +0.1 Rater Opus 5.5 on Opus 5.5: 9.38, against 9.26 from the other eight (+0.12) −0.8 Rater Opus 5.5 on GPT-6 Astra: 8.52, against 9.29 from the other eight (-0.76) −0.6 Rater Opus 5.5 on GPT-6 Sol: 7.90, against 8.46 from the other eight (-0.55) −0.4 Rater Opus 5.5 on Fable 5.1: 8.00, against 8.39 from the other eight (-0.39) +0.1 Rater Opus 5.5 on Opus 5: 8.33, against 8.28 from the other eight (+0.05) −0.8 Rater Opus 5.5 on Grok 4.6: 6.95, against 7.71 from the other eight (-0.76) −0.7 Rater Opus 5.5 on Grok 4.7: 6.90, against 7.65 from the other eight (-0.74) −0.6 Rater Opus 5.5 on GPT-6 Luna: 6.86, against 7.42 from the other eight (-0.57) −0.3 Rater Opus 5.5 on Sonnet 5: 6.81, against 7.13 from the other eight (-0.32) 7.74 GPT-6 Astra +0.1 Rater GPT-6 Astra on Opus 5.5: 9.36, against 9.26 from the other eight (+0.09) +0.7 Rater GPT-6 Astra on GPT-6 Astra: 9.79, against 9.13 from the other eight (+0.66) +0.9 Rater GPT-6 Astra on GPT-6 Sol: 9.16, against 8.30 from the other eight (+0.86) +0.5 Rater GPT-6 Astra on Fable 5.1: 8.82, against 8.29 from the other eight (+0.54) +0.8 Rater GPT-6 Astra on Opus 5: 9.01, against 8.19 from the other eight (+0.82) +0.7 Rater GPT-6 Astra on Grok 4.6: 8.25, against 7.55 from the other eight (+0.70) +0.6 Rater GPT-6 Astra on Grok 4.7: 8.10, against 7.50 from the other eight (+0.61) +0.8 Rater GPT-6 Astra on GPT-6 Luna: 8.08, against 7.27 from the other eight (+0.81) +0.9 Rater GPT-6 Astra on Sonnet 5: 7.92, against 6.99 from the other eight (+0.93) 8.72 GPT-6 Sol −0.3 Rater GPT-6 Sol on Opus 5.5: 9.02, against 9.31 from the other eight (-0.28) +0.2 Rater GPT-6 Sol on GPT-6 Astra: 9.40, against 9.18 from the other eight (+0.23) −0.1 Rater GPT-6 Sol on GPT-6 Sol: 8.33, against 8.40 from the other eight (-0.07) −0.3 Rater GPT-6 Sol on Fable 5.1: 8.07, against 8.38 from the other eight (-0.31) −0.6 Rater GPT-6 Sol on Opus 5: 7.71, against 8.36 from the other eight (-0.64) −0.5 Rater GPT-6 Sol on Grok 4.6: 7.17, against 7.69 from the other eight (-0.52) −0.7 Rater GPT-6 Sol on Grok 4.7: 6.98, against 7.64 from the other eight (-0.66) −0.4 Rater GPT-6 Sol on GPT-6 Luna: 7.00, against 7.41 from the other eight (-0.41) −0.6 Rater GPT-6 Sol on Sonnet 5: 6.60, against 7.16 from the other eight (-0.56) 7.81 Fable 5.1 +0.5 Rater Fable 5.1 on Opus 5.5: 9.71, against 9.22 from the other eight (+0.49) +0.1 Rater Fable 5.1 on GPT-6 Astra: 9.33, against 9.19 from the other eight (+0.15) +0.2 Rater Fable 5.1 on GPT-6 Sol: 8.57, against 8.37 from the other eight (+0.20) +0.5 Rater Fable 5.1 on Fable 5.1: 8.81, against 8.29 from the other eight (+0.52) +0.9 Rater Fable 5.1 on Opus 5: 9.10, against 8.18 from the other eight (+0.91) +0.6 Rater Fable 5.1 on Grok 4.6: 8.14, against 7.56 from the other eight (+0.58) +0.5 Rater Fable 5.1 on Grok 4.7: 8.00, against 7.51 from the other eight (+0.49) +0.3 Rater Fable 5.1 on GPT-6 Luna: 7.67, against 7.32 from the other eight (+0.34) +0.6 Rater Fable 5.1 on Sonnet 5: 7.62, against 7.03 from the other eight (+0.59) 8.55 Opus 5 +0.2 Rater Opus 5 on Opus 5.5: 9.48, against 9.25 from the other eight (+0.23) +0.1 Rater Opus 5 on GPT-6 Astra: 9.29, against 9.19 from the other eight (+0.09) +0.3 Rater Opus 5 on GPT-6 Sol: 8.62, against 8.37 from the other eight (+0.25) +0.3 Rater Opus 5 on Fable 5.1: 8.57, against 8.32 from the other eight (+0.25) +0.3 Rater Opus 5 on Opus 5: 8.52, against 8.25 from the other eight (+0.27) 0.0 Rater Opus 5 on Grok 4.6: 7.67, against 7.62 from the other eight (+0.04) 0.0 Rater Opus 5 on Grok 4.7: 7.57, against 7.57 from the other eight (+0.01) +0.5 Rater Opus 5 on GPT-6 Luna: 7.76, against 7.31 from the other eight (+0.45) +0.2 Rater Opus 5 on Sonnet 5: 7.24, against 7.08 from the other eight (+0.16) 8.30 Grok 4.6 0.0 Rater Grok 4.6 on Opus 5.5: 9.24, against 9.28 from the other eight (-0.04) −0.8 Rater Grok 4.6 on GPT-6 Astra: 8.48, against 9.29 from the other eight (-0.82) −0.9 Rater Grok 4.6 on GPT-6 Sol: 7.57, against 8.50 from the other eight (-0.93) −0.7 Rater Grok 4.6 on Fable 5.1: 7.71, against 8.43 from the other eight (-0.71) −1.1 Rater Grok 4.6 on Opus 5: 7.29, against 8.41 from the other eight (-1.12) −0.2 Rater Grok 4.6 on Grok 4.6: 7.43, against 7.65 from the other eight (-0.22) −0.9 Rater Grok 4.6 on Grok 4.7: 6.81, against 7.66 from the other eight (-0.85) −0.9 Rater Grok 4.6 on GPT-6 Luna: 6.57, against 7.46 from the other eight (-0.89) −1.1 Rater Grok 4.6 on Sonnet 5: 6.10, against 7.22 from the other eight (-1.13) 7.47 Grok 4.7 +0.1 Rater Grok 4.7 on Opus 5.5: 9.33, against 9.27 from the other eight (+0.07) 0.0 Rater Grok 4.7 on GPT-6 Astra: 9.19, against 9.20 from the other eight (-0.01) −0.2 Rater Grok 4.7 on GPT-6 Sol: 8.19, against 8.42 from the other eight (-0.23) −0.2 Rater Grok 4.7 on Fable 5.1: 8.14, against 8.37 from the other eight (-0.23) −0.3 Rater Grok 4.7 on Opus 5: 8.05, against 8.31 from the other eight (-0.27) −0.2 Rater Grok 4.7 on Grok 4.6: 7.43, against 7.65 from the other eight (-0.22) +0.5 Rater Grok 4.7 on Grok 4.7: 8.00, against 7.51 from the other eight (+0.49) −0.2 Rater Grok 4.7 on GPT-6 Luna: 7.14, against 7.39 from the other eight (-0.25) −0.4 Rater Grok 4.7 on Sonnet 5: 6.71, against 7.15 from the other eight (-0.43) 8.02 GPT-6 Luna −0.8 Rater GPT-6 Luna on Opus 5.5: 8.57, against 9.36 from the other eight (-0.79) +0.2 Rater GPT-6 Luna on GPT-6 Astra: 9.40, against 9.18 from the other eight (+0.22) +0.1 Rater GPT-6 Luna on GPT-6 Sol: 8.48, against 8.38 from the other eight (+0.10) −0.5 Rater GPT-6 Luna on Fable 5.1: 7.93, against 8.40 from the other eight (-0.46) −0.4 Rater GPT-6 Luna on Opus 5: 7.93, against 8.33 from the other eight (-0.40) −0.5 Rater GPT-6 Luna on Grok 4.6: 7.14, against 7.69 from the other eight (-0.55) −0.7 Rater GPT-6 Luna on Grok 4.7: 6.92, against 7.65 from the other eight (-0.72) −0.2 Rater GPT-6 Luna on GPT-6 Luna: 7.21, against 7.38 from the other eight (-0.16) −0.6 Rater GPT-6 Luna on Sonnet 5: 6.60, against 7.16 from the other eight (-0.56) 7.80 Sonnet 5 +0.1 Rater Sonnet 5 on Opus 5.5: 9.38, against 9.26 from the other eight (+0.12) +0.3 Rater Sonnet 5 on GPT-6 Astra: 9.43, against 9.18 from the other eight (+0.25) +0.4 Rater Sonnet 5 on GPT-6 Sol: 8.71, against 8.35 from the other eight (+0.36) +0.8 Rater Sonnet 5 on Fable 5.1: 9.05, against 8.26 from the other eight (+0.79) +0.4 Rater Sonnet 5 on Opus 5: 8.62, against 8.24 from the other eight (+0.38) +1.0 Rater Sonnet 5 on Grok 4.6: 8.48, against 7.52 from the other eight (+0.95) +1.4 Rater Sonnet 5 on Grok 4.7: 8.81, against 7.41 from the other eight (+1.40) +0.7 Rater Sonnet 5 on GPT-6 Luna: 7.95, against 7.29 from the other eight (+0.67) +1.3 Rater Sonnet 5 on Sonnet 5: 8.29, against 6.95 from the other eight (+1.34) 8.75 Rater minus the other eight −1.5 −0.75 0 +0.75 +1.5

Fidelity. Of the 189 translations, five had a flagged clue: one changed (GPT-6 Luna's Polish named the wrong pennant colour) and four ambiguous, two of them in Opus 5.5's Traditional Chinese.

Cost and time. At list prices Stage 2 cost USD 33.67: 12.65 for the translations, 19.00 for the rating and 2.02 for the fidelity checks. At effort high the Grok models again spent by far the most output on reasoning and took 3.5–4.5 minutes per translation, against 0.4–0.7 for the others.

7.3 Turkish in both stages

Table 14. The same English (GPT-6 Astra's) translated into Turkish in both stages. Each cell is one translation rated by nine models. Between the stages the clue order changed for all; the effort changed only where marked.
TranslatorStage 1 effortStage 1Stage 2 (high)Change
Sonnet 5high6.337.22+0.89
Opus 5high7.339.44+2.11
Opus 5.5medium → high9.448.89−0.56
Fable 5.1high8.449.56+1.11
Grok 4.6high or more7.008.22+1.22
Grok 4.7high or more8.007.67−0.33
GPT-6 Lunamedium → high7.678.11+0.44
GPT-6 Solmedium? → high8.118.33+0.22
GPT-6 Astramedium? → high9.449.33−0.11
Live text (control)5.445.11

Each cell is a single translation, generated once. Without repeated generations, a change of a point or so cannot be told apart from ordinary variation between runs, so no individual change here is evidence on its own. Two patterns are still worth noting. The largest gains came from models whose effort did not change (Opus 5 +2.11, Grok 4.6 +1.22, Fable 5.1 +1.11), which points to the new clue order or to variation between generations, not to effort. And raising Opus 5.5 and the GPT-6 models to high brought no consistent gain (from −0.56 to +0.44). The live Turkish text again scored below the translations' mean in all nine rater batches.

7.4 What Stage 2 adds

Stage 1's best Turkish translator, Opus 5.5, stays at the top across 21 languages, and GPT-6 Astra, Stage 1's best English writer, turns out to be an equally strong translator. The two are statistically tied and clearly ahead of the rest, and between them they led in 17 of the 21 languages. Choosing a translator by its overall rank is therefore a sound default. Where one language matters most, it pays to check that language: in Traditional Chinese the overall leader scored 2.2 points below that language's best, because of two ambiguous clues that a fidelity check would catch before release.

7.5 Limitations of Stage 2

  • One text. Every language translates the same English puzzle; a ranking can move on other texts.
  • No control outside Turkish. No live text exists in the other languages, so scores are relative within a language and not comparable across languages.
  • Still no human raters. A blind rating sheet for native speakers exists; their scores will be added to this report as an addendum.
  • One fidelity judge, itself a contestant, as in Stage 1.
  • Half points. The GPT-6 raters sometimes gave half points (196 of 5,103 sub-scores in Stage 2 (5,130, controls included), 31 of 8,100 in Stage 1); they were kept as given.
  • Model identity. For the OpenAI models the recorded model identifier is the one requested; it could not be confirmed independently through our access. The Anthropic and xAI identifiers are the providers' own.

8. Stage 2b: the same run on a plain-words source

8.1 Why and what changed

Reading Stage 2, the product owner found two faults in the English it translated (GPT-6 Astra's Stage 1 rewrite) and requested a re-run. The text used technical words that few players know (alidade, theodolite, glacier axe, snow probe, bone white, ochre, plum, and rope positions called lead, upper, lower and anchor), and these reached every language. And it stated the sum clue as a formula. A formula and a dictionary term translate without error, so Stage 2 partly measured the translation of an easy text.

Stage 2b repeated Stage 2 exactly (the nine translators, the 21 languages, the prompts, the rubric, blind rating by all nine models, the fidelity check, effort high on every call) on a source edited by hand from Astra's English. Only the board's item names and one clue changed: the item names became everyday words (compass, telescope, axe, pole; first, second, third, last), and the sum clue lost its number legend. Before any translation, Opus 5.5 judged every clue of the edited source same against the stored logic.

The sum clue was still arithmetic. Stage 2 read "Number the rope positions from front to back: lead = 1, upper = 2, lower = 3, anchor = 4. Bennett's and Gareth's position numbers add up to 6." Stage 2b read "Counting from the front, Bennett's place on the rope and Gareth's add up to 6." The legend was gone; the addition was not, and every Stage 2b translation carried it. Stage 2b therefore tests plain item names, not a formula-free source. Stage 2c (Section 9) removes the arithmetic.

8.2 Results

Raters agreed less than in Stage 2: Kendall's W averaged 0.52 over the 21 language batches (range 0.25–0.76, against 0.67), and ICC(2,1) was 0.42 for a single rater and ICC(2,k) 0.86 for the mean of nine (Stage 2: 0.54 and 0.91). The translations became more alike, which leaves the raters less to order them by.

Figure 10. Stage 2 against Stage 2b, per translator: mean blind score out of 10 over 21 languages × 9 raters (189 ratings per point), with 95% confidence intervals (two-way cluster bootstrap over raters and languages, 4,000 resamples, seed 20260925, run separately per stage as stats2.py). Same rubric, raters, languages and prompts; only the English source differs. The right column is Stage 2b minus Stage 2. Sources: GRIDIGMA out/s2/rate/ and out/s2b/rate/.
Stage 2: Astra's English as written Stage 2b: plain-words edit 6 7 8 9 10 Change Opus 5.5 Opus 5.5, Stage 2: 9.28 (95% CI 8.90 to 9.57) Opus 5.5, Stage 2b: 9.10 (95% CI 8.81 to 9.40) −0.17 GPT-6 Astra GPT-6 Astra, Stage 2: 9.20 (95% CI 8.80 to 9.54) GPT-6 Astra, Stage 2b: 9.19 (95% CI 8.85 to 9.49) −0.02 GPT-6 Sol GPT-6 Sol, Stage 2: 8.39 (95% CI 7.90 to 8.86) GPT-6 Sol, Stage 2b: 9.05 (95% CI 8.73 to 9.37) +0.66 Fable 5.1 Fable 5.1, Stage 2: 8.35 (95% CI 7.88 to 8.77) Fable 5.1, Stage 2b: 8.59 (95% CI 8.25 to 8.93) +0.24 Opus 5 Opus 5, Stage 2: 8.28 (95% CI 7.77 to 8.79) Opus 5, Stage 2b: 8.71 (95% CI 8.41 to 8.99) +0.42 Grok 4.6 Grok 4.6, Stage 2: 7.63 (95% CI 6.98 to 8.21) Grok 4.6, Stage 2b: 8.29 (95% CI 7.84 to 8.70) +0.67 Grok 4.7 Grok 4.7, Stage 2: 7.57 (95% CI 6.96 to 8.16) Grok 4.7, Stage 2b: 8.18 (95% CI 7.77 to 8.52) +0.61 GPT-6 Luna GPT-6 Luna, Stage 2: 7.36 (95% CI 6.79 to 7.93) GPT-6 Luna, Stage 2b: 7.87 (95% CI 7.30 to 8.43) +0.51 Sonnet 5 Sonnet 5, Stage 2: 7.10 (95% CI 6.52 to 7.68) Sonnet 5, Stage 2b: 7.30 (95% CI 6.80 to 7.82) +0.21 Mean blind score out of 10 (Stage 1 rubric, both stages)
Table 15. Stage 2 against Stage 2b per translator, over 21 languages. Score out of 10 with 95% CI for Stage 2b; naturalness and clarity as mean criterion scores (of 3 and of 4); "Best in": languages in which it scored highest among all nine.
TranslatorStage 2Stage 2b [95% CI]ChangeNaturalnessClarityBest in
Opus 5.59.289.10 [8.81, 9.40]−0.172.72 → 2.593.70 → 3.8110 → 5
GPT-6 Astra9.209.19 [8.85, 9.49]−0.022.58 → 2.643.86 → 3.817 → 6
GPT-6 Sol8.399.05 [8.73, 9.37]+0.662.33 → 2.623.40 → 3.631 → 7
Fable 5.18.358.59 [8.25, 8.93]+0.242.27 → 2.323.43 → 3.702 → 1
Opus 58.288.71 [8.41, 8.99]+0.422.23 → 2.453.39 → 3.591 → 2
Grok 4.67.638.29 [7.84, 8.70]+0.671.96 → 2.213.20 → 3.500 → 1
Grok 4.77.578.18 [7.77, 8.52]+0.611.94 → 2.113.29 → 3.630 → 1
GPT-6 Luna7.367.87 [7.30, 8.43]+0.511.86 → 2.123.14 → 3.310 → 0
Sonnet 57.107.30 [6.80, 7.82]+0.211.74 → 1.853.20 → 3.250 → 0

The weaker translators caught up. 7 of the nine translators scored higher on the plain source, by 0.21 to 0.67 points; the two Stage 2 leaders did not (GPT-6 Astra −0.02, Opus 5.5 −0.17). The gap between the best and the weakest translator narrowed from 2.18 to 1.88 points, and naturalness rose for 8 of the nine. The top of the ranking became a three-way tie: GPT-6 Astra 9.19 [8.85, 9.49], Opus 5.5 9.10 [8.81, 9.40] and GPT-6 Sol 9.05 [8.73, 9.37]; in 1,000 resamples of raters and languages they ranked first 569, 309 and 121 times. GPT-6 Sol gained the most among the three (+0.66).

Per-language leaders moved. The best translator changed in 16 of the 21 languages. A language's own ranking agreed with the overall one with a median Spearman correlation of 0.75.

Fidelity and control. The check flagged 1 changed and 2 ambiguous clues, all in one translation (GPT-6 Luna, Korean). The live Turkish text, rated blind in the Turkish batch, scored 4.11 (Stage 2: 5.11) and fell below the batch's translation mean in 9 of 9 rater batches. At list prices Stage 2b cost USD 33.79: 11.72 for the translations, 20.13 for the rating and 1.94 for the checks.

8.3 Limitations of Stage 2b

  • Not a formula-free source. The sum clue still asked the reader to add two places (Section 8.1), so Stage 2b does not show how the translators handle a source with no arithmetic.
  • Edited by hand, not rewritten by a model. The edit isolates the two faults; it does not show what a model instructed to write plainly would produce.
  • Original names still visible. As in Stage 2, each translator saw the puzzle's original board names with the rename map beside them.
  • One generation per cell. Every Stage 2 limitation stands (Section 7.5); a change of a point or so in one language is within the variation between runs (Section 9.6).

9. Stage 2c: plain words, naturalness first

9.1 Why

The product owner requested Stage 2c as a further step towards natural text in every language, in his words: "I don't want a Japanese senior to open the puzzle, have a look at the story and close it right away as they found the translations too mechanical." Two findings shaped it. Stage 2b had not removed the arithmetic (Section 8.1). And the Stage 1 rubric itself rewarded number-talk: clarity carried 4 of the 10 points, and a numbered clue is maximally clear. GPT-6 Astra's Stage 1 English lead came mostly from clarity (3.89 of 4), while on naturalness Grok 4.6 was ahead (2.89 of 3, against 2.81; Table 3).

9.2 Design

Table 16. What Stage 2c changed.
Stages 1, 2 and 2bStage 2c
Writer promptplain words preferredplain words required; no arithmetic, equations or numbered positions; a clue may be told as any statement that rules in and out exactly the same arrangements
Rubricclarity 4, readability 3, naturalness 3naturalness 6, readability 2, clarity 2, read "as an older native speaker who has just opened this puzzle on a phone"
Fidelity checksame fact as the source textsame arrangements as the stored logic; reported as a gate beside the score
EnglishStage 1: nine writers, three puzzlesnine writers, three puzzles, rated blind by all nine with the live English as control
Translation sourceAstra's Stage 1 English (2b: hand-edited)the best-rated Stage 2c English of The Rope-Line Survey that passed the check, chosen by code: GPT-6 Astra's
Routesfrom Englishfrom English, and from the stored logic directly (no text)
Translatorsall nineOpus 5.5, GPT-6 Astra, GPT-6 Sol

Unchanged: effort high set and recorded on every call, blind shuffled rating by all nine models with self-ratings included and reported without, Opus 5.5 as the only fidelity judge, and the game's clue order. Stage 2c produced 27 English versions (243 ratings) and 126 translations into the same 21 languages as Stage 2 (1134 ratings). Confidence intervals use the two-way cluster bootstrap of Section 2.8 (raters and puzzles for English, raters and languages for translation; 4,000 resamples).

A first translation run was discarded. It handed the translators the stored logic, in which the "exactly one note is false" clue names its notes by internal clue number; because players see that clue first, the numbers on screen differ by one, and the GPT-6 translators attached the letters to the wrong clues. The from-logic route also saw only the board's technical labels. The second run gave every prompt the notes by on-screen position and the labels' plain meaning; all results here are from it. The first run's 30 rated versions (5 languages) are kept as a record and used only to measure run-to-run change (Section 9.6).

Scale. Because the rubric weights changed, Stage 2c scores are not on the scale of Stages 1, 2 and 2b and are never compared with them directly; comparisons across stages use ranks (Figure 14).

9.3 English

Figure 11. Stage 2c English: mean blind score out of 10 for each writer over three puzzles × nine raters (27 ratings), under the Stage 2c rubric, which weights naturalness 6, readability 2 and clarity 2 and so is not on the scale of Figures 1, 5 and 10. 95% confidence intervals: two-way cluster bootstrap resampling raters and puzzles (4,000 resamples, seed 20260925), stats.py's method; with three puzzles the intervals are wide. 'Flagged': clues the fidelity check did not judge 'same'. The last row is the live English, rated blind in the same batches (point estimate, n = 27). Source: GRIDIGMA out/s2c/en-rate/, en-check/.
3 4 5 6 7 8 9 10 Flagged GPT-6 Astra GPT-6 Astra: 8.87 (95% CI 8.35 to 9.22); naturalness 4.98 of 6 Opus 5.5 Opus 5.5: 8.80 (95% CI 8.33 to 9.19); naturalness 4.89 of 6 Fable 5.1 Fable 5.1: 8.04 (95% CI 7.26 to 8.78); naturalness 4.39 of 6 1 Grok 4.6 Grok 4.6: 7.81 (95% CI 6.74 to 8.78); naturalness 3.96 of 6 Opus 5 Opus 5: 7.78 (95% CI 6.67 to 8.93); naturalness 4.20 of 6 GPT-6 Sol GPT-6 Sol: 7.22 (95% CI 6.67 to 7.89); naturalness 3.98 of 6 1 Grok 4.7 Grok 4.7: 6.91 (95% CI 5.70 to 8.04); naturalness 3.43 of 6 Sonnet 5 Sonnet 5: 6.28 (95% CI 5.63 to 6.96); naturalness 3.70 of 6 GPT-6 Luna GPT-6 Luna: 6.13 (95% CI 5.22 to 7.11); naturalness 3.46 of 6 1 Live English Live English: 4.48; naturalness 2.30 of 6 Mean blind score out of 10, Stage 2c rubric (naturalness 6, readability 2, clarity 2)
Table 17. Stage 2c English per writer, over three puzzles and nine raters. Naturalness is out of 6. "Flagged": clues the check did not judge same. Rank: among the nine, in Stage 1 and in Stage 2c.
WriterScore [95% CI]No selfNaturalnessFlaggedRank 1 → 2c
GPT-6 Astra8.87 [8.35, 9.22]8.794.9801 → 1
Opus 5.58.80 [8.33, 9.19]8.734.8903 → 2
Fable 5.18.04 [7.26, 8.78]8.004.3915 → 3
Grok 4.67.81 [6.74, 8.78]8.003.9604 → 4
Opus 57.78 [6.67, 8.93]7.754.2006 → 5
GPT-6 Sol7.22 [6.67, 7.89]7.213.9812 → 6
Grok 4.76.91 [5.70, 8.04]6.903.4307 → 7
Sonnet 56.28 [5.63, 6.96]6.193.7008 → 8
GPT-6 Luna6.13 [5.22, 7.11]6.023.4619 → 9
Live English (control)4.482.30

Every writer dropped the arithmetic. GPT-6 Astra (8.87) and Opus 5.5 (8.80) led and cannot be separated; GPT-6 Sol, 2nd in Stage 1, fell to 6th once number-talk stopped paying. The live English scored 4.48. Agreement: Kendall's W 0.74 over the three puzzle batches (control included, as in Stage 1), ICC(2,1) 0.54 and ICC(2,k) 0.91 over the 27 rewrites.

9.4 Translation

Raters agreed with Kendall's W 0.36 over the 21 language batches (range 0.15–0.59); ICC(2,1) was 0.28 and ICC(2,k) 0.78 over the 126 translations. This is the lowest agreement in the study. Part of it is expected: the three translators are the strongest of Stage 2, so the versions differ less and leave raters less to agree on; part may come from the heavier weight on naturalness, the most subjective criterion. Either way a single rater's score says little here, the nine-rater mean is moderately reliable, and per-language winners separated by less than about a point should be read as directions only.

Figure 12. Stage 2c translation: all 126 versions, by language (rows), translator and route (columns). Each cell is the mean of nine raters' totals out of 10 under the Stage 2c rubric (not comparable with Figure 6's scale). 'From English': translated from GPT-6 Astra's Stage 2c English; 'from logic': written from the stored puzzle logic with no English text. The frame marks the language's best version among those the fidelity check passed; a corner mark flags a version with a clue judged ambiguous (no clue was judged changed). One puzzle, one generation per cell. The live Turkish, rated in the Turkish batch, scored 2.83. Source: GRIDIGMA out/s2c/rate/, check/ (run 2; run 1 excluded).
Opus 5.5 GPT-6 Astra GPT-6 Sol From English From logic From English From logic From English From logic German 7.22 German, Opus 5.5 from English: 7.22 of 10 (9 raters); naturalness 3.89 of 6 8.89 German, Opus 5.5 from the logic: 8.89 of 10 (9 raters); naturalness 5.11 of 6 8.00 German, GPT-6 Astra from English: 8.00 of 10 (9 raters); naturalness 4.00 of 6 7.56 German, GPT-6 Astra from the logic: 7.56 of 10 (9 raters); naturalness 4.00 of 6 6.44 German, GPT-6 Sol from English: 6.44 of 10 (9 raters); naturalness 3.33 of 6 8.11 German, GPT-6 Sol from the logic: 8.11 of 10 (9 raters); naturalness 4.44 of 6 Japanese 8.28 Japanese, Opus 5.5 from English: 8.28 of 10 (9 raters); naturalness 4.44 of 6 8.28 Japanese, Opus 5.5 from the logic: 8.28 of 10 (9 raters); naturalness 4.72 of 6 8.67 Japanese, GPT-6 Astra from English: 8.67 of 10 (9 raters); naturalness 4.67 of 6 7.50 Japanese, GPT-6 Astra from the logic: 7.50 of 10 (9 raters); naturalness 4.33 of 6 8.33 Japanese, GPT-6 Sol from English: 8.33 of 10 (9 raters); naturalness 4.56 of 6 6.17 Japanese, GPT-6 Sol from the logic: 6.17 of 10 (9 raters); naturalness 3.50 of 6 French 6.83 French, Opus 5.5 from English: 6.83 of 10 (9 raters); naturalness 3.67 of 6 8.11 French, Opus 5.5 from the logic: 8.11 of 10 (9 raters); naturalness 4.61 of 6 7.61 French, GPT-6 Astra from English: 7.61 of 10 (9 raters); naturalness 4.17 of 6 7.89 French, GPT-6 Astra from the logic: 7.89 of 10 (9 raters); naturalness 4.11 of 6 6.94 French, GPT-6 Sol from English: 6.94 of 10 (9 raters); naturalness 3.89 of 6 7.67 French, GPT-6 Sol from the logic: 7.67 of 10 (9 raters); naturalness 4.28 of 6 Korean 7.92 Korean, Opus 5.5 from English: 7.92 of 10 (9 raters); naturalness 4.28 of 6 8.47 Korean, Opus 5.5 from the logic: 8.47 of 10 (9 raters); naturalness 4.61 of 6 8.73 Korean, GPT-6 Astra from English: 8.73 of 10 (9 raters); naturalness 4.89 of 6 8.37 Korean, GPT-6 Astra from the logic: 8.37 of 10 (9 raters); naturalness 4.72 of 6 7.72 Korean, GPT-6 Sol from English: 7.72 of 10 (9 raters); naturalness 4.11 of 6 7.06 Korean, GPT-6 Sol from the logic: 7.06 of 10 (9 raters); naturalness 4.11 of 6; fidelity check: clue 3 ambiguous, clue 5 ambiguous Chinese (Traditional) 8.61 Chinese (Traditional), Opus 5.5 from English: 8.61 of 10 (9 raters); naturalness 4.61 of 6 7.39 Chinese (Traditional), Opus 5.5 from the logic: 7.39 of 10 (9 raters); naturalness 4.33 of 6 8.11 Chinese (Traditional), GPT-6 Astra from English: 8.11 of 10 (9 raters); naturalness 4.22 of 6 7.67 Chinese (Traditional), GPT-6 Astra from the logic: 7.67 of 10 (9 raters); naturalness 4.00 of 6 8.67 Chinese (Traditional), GPT-6 Sol from English: 8.67 of 10 (9 raters); naturalness 4.78 of 6 6.17 Chinese (Traditional), GPT-6 Sol from the logic: 6.17 of 10 (9 raters); naturalness 3.00 of 6 Chinese (Simplified) 7.39 Chinese (Simplified), Opus 5.5 from English: 7.39 of 10 (9 raters); naturalness 4.00 of 6 8.67 Chinese (Simplified), Opus 5.5 from the logic: 8.67 of 10 (9 raters); naturalness 4.89 of 6 7.89 Chinese (Simplified), GPT-6 Astra from English: 7.89 of 10 (9 raters); naturalness 4.17 of 6 7.83 Chinese (Simplified), GPT-6 Astra from the logic: 7.83 of 10 (9 raters); naturalness 4.11 of 6 8.33 Chinese (Simplified), GPT-6 Sol from English: 8.33 of 10 (9 raters); naturalness 4.61 of 6 7.78 Chinese (Simplified), GPT-6 Sol from the logic: 7.78 of 10 (9 raters); naturalness 4.06 of 6 Dutch 7.89 Dutch, Opus 5.5 from English: 7.89 of 10 (9 raters); naturalness 4.03 of 6 7.74 Dutch, Opus 5.5 from the logic: 7.74 of 10 (9 raters); naturalness 4.36 of 6; fidelity check: clue 1 ambiguous 7.74 Dutch, GPT-6 Astra from English: 7.74 of 10 (9 raters); naturalness 3.87 of 6 8.40 Dutch, GPT-6 Astra from the logic: 8.40 of 10 (9 raters); naturalness 4.57 of 6 7.68 Dutch, GPT-6 Sol from English: 7.68 of 10 (9 raters); naturalness 3.91 of 6 6.68 Dutch, GPT-6 Sol from the logic: 6.68 of 10 (9 raters); naturalness 3.39 of 6; fidelity check: clue 6 ambiguous, clue 7 ambiguous Spanish (Spain) 8.94 Spanish (Spain), Opus 5.5 from English: 8.94 of 10 (9 raters); naturalness 5.00 of 6 7.61 Spanish (Spain), Opus 5.5 from the logic: 7.61 of 10 (9 raters); naturalness 4.33 of 6 8.53 Spanish (Spain), GPT-6 Astra from English: 8.53 of 10 (9 raters); naturalness 4.56 of 6 8.42 Spanish (Spain), GPT-6 Astra from the logic: 8.42 of 10 (9 raters); naturalness 4.44 of 6 8.36 Spanish (Spain), GPT-6 Sol from English: 8.36 of 10 (9 raters); naturalness 4.50 of 6 7.28 Spanish (Spain), GPT-6 Sol from the logic: 7.28 of 10 (9 raters); naturalness 4.00 of 6 Spanish (Latin America) 7.84 Spanish (Latin America), Opus 5.5 from English: 7.84 of 10 (9 raters); naturalness 4.22 of 6 7.06 Spanish (Latin America), Opus 5.5 from the logic: 7.06 of 10 (9 raters); naturalness 3.83 of 6; fidelity check: clue 1 ambiguous 7.70 Spanish (Latin America), GPT-6 Astra from English: 7.70 of 10 (9 raters); naturalness 4.17 of 6 8.67 Spanish (Latin America), GPT-6 Astra from the logic: 8.67 of 10 (9 raters); naturalness 4.67 of 6 7.64 Spanish (Latin America), GPT-6 Sol from English: 7.64 of 10 (9 raters); naturalness 4.06 of 6 7.13 Spanish (Latin America), GPT-6 Sol from the logic: 7.13 of 10 (9 raters); naturalness 4.06 of 6 Italian 8.28 Italian, Opus 5.5 from English: 8.28 of 10 (9 raters); naturalness 4.50 of 6 8.56 Italian, Opus 5.5 from the logic: 8.56 of 10 (9 raters); naturalness 5.06 of 6 7.78 Italian, GPT-6 Astra from English: 7.78 of 10 (9 raters); naturalness 4.00 of 6 8.17 Italian, GPT-6 Astra from the logic: 8.17 of 10 (9 raters); naturalness 4.28 of 6 7.11 Italian, GPT-6 Sol from English: 7.11 of 10 (9 raters); naturalness 3.67 of 6 7.78 Italian, GPT-6 Sol from the logic: 7.78 of 10 (9 raters); naturalness 4.33 of 6 Swedish 8.11 Swedish, Opus 5.5 from English: 8.11 of 10 (9 raters); naturalness 4.11 of 6 7.78 Swedish, Opus 5.5 from the logic: 7.78 of 10 (9 raters); naturalness 4.44 of 6 8.33 Swedish, GPT-6 Astra from English: 8.33 of 10 (9 raters); naturalness 4.33 of 6 8.78 Swedish, GPT-6 Astra from the logic: 8.78 of 10 (9 raters); naturalness 4.78 of 6 6.67 Swedish, GPT-6 Sol from English: 6.67 of 10 (9 raters); naturalness 3.56 of 6 8.00 Swedish, GPT-6 Sol from the logic: 8.00 of 10 (9 raters); naturalness 4.33 of 6 Norwegian 7.67 Norwegian, Opus 5.5 from English: 7.67 of 10 (9 raters); naturalness 3.83 of 6 7.22 Norwegian, Opus 5.5 from the logic: 7.22 of 10 (9 raters); naturalness 4.06 of 6 8.28 Norwegian, GPT-6 Astra from English: 8.28 of 10 (9 raters); naturalness 4.44 of 6 8.78 Norwegian, GPT-6 Astra from the logic: 8.78 of 10 (9 raters); naturalness 4.83 of 6 7.17 Norwegian, GPT-6 Sol from English: 7.17 of 10 (9 raters); naturalness 3.94 of 6 7.67 Norwegian, GPT-6 Sol from the logic: 7.67 of 10 (9 raters); naturalness 4.11 of 6 Danish 7.39 Danish, Opus 5.5 from English: 7.39 of 10 (9 raters); naturalness 3.83 of 6 7.33 Danish, Opus 5.5 from the logic: 7.33 of 10 (9 raters); naturalness 4.33 of 6 8.00 Danish, GPT-6 Astra from English: 8.00 of 10 (9 raters); naturalness 4.33 of 6 8.67 Danish, GPT-6 Astra from the logic: 8.67 of 10 (9 raters); naturalness 4.72 of 6 7.39 Danish, GPT-6 Sol from English: 7.39 of 10 (9 raters); naturalness 4.06 of 6 7.61 Danish, GPT-6 Sol from the logic: 7.61 of 10 (9 raters); naturalness 4.11 of 6 Finnish 8.00 Finnish, Opus 5.5 from English: 8.00 of 10 (9 raters); naturalness 4.22 of 6 7.56 Finnish, Opus 5.5 from the logic: 7.56 of 10 (9 raters); naturalness 4.39 of 6 8.72 Finnish, GPT-6 Astra from English: 8.72 of 10 (9 raters); naturalness 4.72 of 6 7.89 Finnish, GPT-6 Astra from the logic: 7.89 of 10 (9 raters); naturalness 4.17 of 6 6.56 Finnish, GPT-6 Sol from English: 6.56 of 10 (9 raters); naturalness 3.61 of 6 7.39 Finnish, GPT-6 Sol from the logic: 7.39 of 10 (9 raters); naturalness 3.89 of 6 Polish 7.22 Polish, Opus 5.5 from English: 7.22 of 10 (9 raters); naturalness 3.44 of 6 8.22 Polish, Opus 5.5 from the logic: 8.22 of 10 (9 raters); naturalness 4.78 of 6 7.44 Polish, GPT-6 Astra from English: 7.44 of 10 (9 raters); naturalness 3.56 of 6 8.67 Polish, GPT-6 Astra from the logic: 8.67 of 10 (9 raters); naturalness 4.67 of 6 6.67 Polish, GPT-6 Sol from English: 6.67 of 10 (9 raters); naturalness 3.67 of 6 7.44 Polish, GPT-6 Sol from the logic: 7.44 of 10 (9 raters); naturalness 3.78 of 6 Czech 6.33 Czech, Opus 5.5 from English: 6.33 of 10 (9 raters); naturalness 3.09 of 6 8.38 Czech, Opus 5.5 from the logic: 8.38 of 10 (9 raters); naturalness 4.80 of 6 6.94 Czech, GPT-6 Astra from English: 6.94 of 10 (9 raters); naturalness 3.69 of 6 8.27 Czech, GPT-6 Astra from the logic: 8.27 of 10 (9 raters); naturalness 4.67 of 6 6.67 Czech, GPT-6 Sol from English: 6.67 of 10 (9 raters); naturalness 3.58 of 6 7.72 Czech, GPT-6 Sol from the logic: 7.72 of 10 (9 raters); naturalness 4.36 of 6 Hebrew 8.56 Hebrew, Opus 5.5 from English: 8.56 of 10 (9 raters); naturalness 4.72 of 6 7.17 Hebrew, Opus 5.5 from the logic: 7.17 of 10 (9 raters); naturalness 4.22 of 6 8.28 Hebrew, GPT-6 Astra from English: 8.28 of 10 (9 raters); naturalness 4.50 of 6 8.44 Hebrew, GPT-6 Astra from the logic: 8.44 of 10 (9 raters); naturalness 4.44 of 6 8.00 Hebrew, GPT-6 Sol from English: 8.00 of 10 (9 raters); naturalness 4.33 of 6 6.06 Hebrew, GPT-6 Sol from the logic: 6.06 of 10 (9 raters); naturalness 3.56 of 6 Arabic (Gulf) 8.32 Arabic (Gulf), Opus 5.5 from English: 8.32 of 10 (9 raters); naturalness 4.61 of 6 7.83 Arabic (Gulf), Opus 5.5 from the logic: 7.83 of 10 (9 raters); naturalness 4.67 of 6 7.67 Arabic (Gulf), GPT-6 Astra from English: 7.67 of 10 (9 raters); naturalness 3.89 of 6 8.00 Arabic (Gulf), GPT-6 Astra from the logic: 8.00 of 10 (9 raters); naturalness 4.22 of 6 7.12 Arabic (Gulf), GPT-6 Sol from English: 7.12 of 10 (9 raters); naturalness 3.83 of 6 7.11 Arabic (Gulf), GPT-6 Sol from the logic: 7.11 of 10 (9 raters); naturalness 3.89 of 6 Portuguese (Brazil) 8.17 Portuguese (Brazil), Opus 5.5 from English: 8.17 of 10 (9 raters); naturalness 4.17 of 6 8.00 Portuguese (Brazil), Opus 5.5 from the logic: 8.00 of 10 (9 raters); naturalness 4.50 of 6 8.28 Portuguese (Brazil), GPT-6 Astra from English: 8.28 of 10 (9 raters); naturalness 4.28 of 6 8.67 Portuguese (Brazil), GPT-6 Astra from the logic: 8.67 of 10 (9 raters); naturalness 4.67 of 6 8.94 Portuguese (Brazil), GPT-6 Sol from English: 8.94 of 10 (9 raters); naturalness 4.94 of 6 6.72 Portuguese (Brazil), GPT-6 Sol from the logic: 6.72 of 10 (9 raters); naturalness 4.17 of 6 Russian 8.17 Russian, Opus 5.5 from English: 8.17 of 10 (9 raters); naturalness 4.39 of 6 6.28 Russian, Opus 5.5 from the logic: 6.28 of 10 (9 raters); naturalness 3.94 of 6 7.89 Russian, GPT-6 Astra from English: 7.89 of 10 (9 raters); naturalness 4.00 of 6 8.72 Russian, GPT-6 Astra from the logic: 8.72 of 10 (9 raters); naturalness 4.83 of 6 7.67 Russian, GPT-6 Sol from English: 7.67 of 10 (9 raters); naturalness 4.06 of 6 7.61 Russian, GPT-6 Sol from the logic: 7.61 of 10 (9 raters); naturalness 4.11 of 6 Turkish 9.61 Turkish, Opus 5.5 from English: 9.61 of 10 (9 raters); naturalness 5.61 of 6 8.78 Turkish, Opus 5.5 from the logic: 8.78 of 10 (9 raters); naturalness 5.00 of 6 8.61 Turkish, GPT-6 Astra from English: 8.61 of 10 (9 raters); naturalness 4.78 of 6 8.28 Turkish, GPT-6 Astra from the logic: 8.28 of 10 (9 raters); naturalness 4.33 of 6 8.83 Turkish, GPT-6 Sol from English: 8.83 of 10 (9 raters); naturalness 4.94 of 6 7.44 Turkish, GPT-6 Sol from the logic: 7.44 of 10 (9 raters); naturalness 4.33 of 6; fidelity check: clue 6 ambiguous, clue 7 ambiguous Scale (Stage 2c rubric) 5 6 7 8 9 10 a clue judged ambiguous
Table 18. Stage 2c per translator over 21 languages. "Better route": the mean, over languages, of the translator's higher-scoring route among versions that passed the check. "Won": languages in which it made the best passing version. Route means with 95% CI; the difference is from the same paired resamples.
TranslatorBetter routeWonFrom EnglishFrom logicLogic − English
GPT-6 Astra8.45108.06 [7.72, 8.34]8.27 [7.92, 8.57]+0.21 [−0.13, +0.58]
Opus 5.58.3397.94 [7.54, 8.32]7.87 [7.32, 8.38]−0.07 [−0.68, +0.53]
GPT-6 Sol7.9427.57 [7.03, 8.05]7.27 [6.79, 7.74]−0.30 [−0.85, +0.24]

Two translators share the top. GPT-6 Astra won 10 languages with the best mean (8.45), Opus 5.5 9 (8.33), and GPT-6 Sol 2 (7.94). The single highest score of the stage was Opus 5.5's Turkish from English (9.61). Each candidate's better route scored between 7.12 and 9.61 in every language; the live Turkish, rated in the Turkish batch, scored 2.83.

Figure 13. Stage 2c translation routes. A: each candidate's mean over its 21 versions per route (189 ratings per point), and the three together (567), with 95% confidence intervals; the right column is the difference, logic minus English, with its interval from the same paired resamples (two-way cluster bootstrap over raters and languages, 4,000 resamples, seed 20260925). B: the same difference in each language (one dot per language, each the difference of two single versions); the right column counts the languages where the logic route scored higher. Stage 2c rubric. Source: GRIDIGMA out/s2c/rate/.
A. Mean by route From English From the logic 6.5 7 7.5 8 8.5 9 9.5 Logic − English [95% CI] Opus 5.5 Opus 5.5 from English: 7.94 (95% CI 7.54 to 8.32) Opus 5.5 from the logic: 7.87 (95% CI 7.32 to 8.38) −0.07 [−0.68, +0.53] GPT-6 Astra GPT-6 Astra from English: 8.06 (95% CI 7.72 to 8.34) GPT-6 Astra from the logic: 8.27 (95% CI 7.92 to 8.57) +0.21 [−0.13, +0.58] GPT-6 Sol GPT-6 Sol from English: 7.57 (95% CI 7.03 to 8.05) GPT-6 Sol from the logic: 7.27 (95% CI 6.79 to 7.74) −0.30 [−0.85, +0.24] All three All three from English: 7.86 (95% CI 7.49 to 8.17) All three from the logic: 7.80 (95% CI 7.50 to 8.04) −0.05 [−0.42, +0.32] Mean over 21 languages x 9 raters, out of 10 (Stage 2c rubric) B. Per language: from the logic minus from English, one dot per language −3 −2 −1 0 +1 +2 +3 Logic ahead Opus 5.5 Opus 5.5, German: from logic 8.89, from English 7.22 (+1.67) Opus 5.5, Japanese: from logic 8.28, from English 8.28 (+0.00) Opus 5.5, French: from logic 8.11, from English 6.83 (+1.28) Opus 5.5, Korean: from logic 8.47, from English 7.92 (+0.54) Opus 5.5, Chinese (Traditional): from logic 7.39, from English 8.61 (-1.22) Opus 5.5, Chinese (Simplified): from logic 8.67, from English 7.39 (+1.28) Opus 5.5, Dutch: from logic 7.74, from English 7.89 (-0.14) Opus 5.5, Spanish (Spain): from logic 7.61, from English 8.94 (-1.33) Opus 5.5, Spanish (Latin America): from logic 7.06, from English 7.84 (-0.79) Opus 5.5, Italian: from logic 8.56, from English 8.28 (+0.28) Opus 5.5, Swedish: from logic 7.78, from English 8.11 (-0.33) Opus 5.5, Norwegian: from logic 7.22, from English 7.67 (-0.44) Opus 5.5, Danish: from logic 7.33, from English 7.39 (-0.06) Opus 5.5, Finnish: from logic 7.56, from English 8.00 (-0.44) Opus 5.5, Polish: from logic 8.22, from English 7.22 (+1.00) Opus 5.5, Czech: from logic 8.38, from English 6.33 (+2.04) Opus 5.5, Hebrew: from logic 7.17, from English 8.56 (-1.39) Opus 5.5, Arabic (Gulf): from logic 7.83, from English 8.32 (-0.49) Opus 5.5, Portuguese (Brazil): from logic 8.00, from English 8.17 (-0.17) Opus 5.5, Russian: from logic 6.28, from English 8.17 (-1.89) Opus 5.5, Turkish: from logic 8.78, from English 9.61 (-0.83) 7 of 21 GPT-6 Astra GPT-6 Astra, German: from logic 7.56, from English 8.00 (-0.44) GPT-6 Astra, Japanese: from logic 7.50, from English 8.67 (-1.17) GPT-6 Astra, French: from logic 7.89, from English 7.61 (+0.28) GPT-6 Astra, Korean: from logic 8.37, from English 8.73 (-0.37) GPT-6 Astra, Chinese (Traditional): from logic 7.67, from English 8.11 (-0.44) GPT-6 Astra, Chinese (Simplified): from logic 7.83, from English 7.89 (-0.06) GPT-6 Astra, Dutch: from logic 8.40, from English 7.74 (+0.66) GPT-6 Astra, Spanish (Spain): from logic 8.42, from English 8.53 (-0.11) GPT-6 Astra, Spanish (Latin America): from logic 8.67, from English 7.70 (+0.97) GPT-6 Astra, Italian: from logic 8.17, from English 7.78 (+0.39) GPT-6 Astra, Swedish: from logic 8.78, from English 8.33 (+0.44) GPT-6 Astra, Norwegian: from logic 8.78, from English 8.28 (+0.50) GPT-6 Astra, Danish: from logic 8.67, from English 8.00 (+0.67) GPT-6 Astra, Finnish: from logic 7.89, from English 8.72 (-0.83) GPT-6 Astra, Polish: from logic 8.67, from English 7.44 (+1.22) GPT-6 Astra, Czech: from logic 8.27, from English 6.94 (+1.32) GPT-6 Astra, Hebrew: from logic 8.44, from English 8.28 (+0.17) GPT-6 Astra, Arabic (Gulf): from logic 8.00, from English 7.67 (+0.33) GPT-6 Astra, Portuguese (Brazil): from logic 8.67, from English 8.28 (+0.39) GPT-6 Astra, Russian: from logic 8.72, from English 7.89 (+0.83) GPT-6 Astra, Turkish: from logic 8.28, from English 8.61 (-0.33) 13 of 21 GPT-6 Sol GPT-6 Sol, German: from logic 8.11, from English 6.44 (+1.67) GPT-6 Sol, Japanese: from logic 6.17, from English 8.33 (-2.17) GPT-6 Sol, French: from logic 7.67, from English 6.94 (+0.72) GPT-6 Sol, Korean: from logic 7.06, from English 7.72 (-0.67) GPT-6 Sol, Chinese (Traditional): from logic 6.17, from English 8.67 (-2.50) GPT-6 Sol, Chinese (Simplified): from logic 7.78, from English 8.33 (-0.56) GPT-6 Sol, Dutch: from logic 6.68, from English 7.68 (-1.00) GPT-6 Sol, Spanish (Spain): from logic 7.28, from English 8.36 (-1.08) GPT-6 Sol, Spanish (Latin America): from logic 7.13, from English 7.64 (-0.51) GPT-6 Sol, Italian: from logic 7.78, from English 7.11 (+0.67) GPT-6 Sol, Swedish: from logic 8.00, from English 6.67 (+1.33) GPT-6 Sol, Norwegian: from logic 7.67, from English 7.17 (+0.50) GPT-6 Sol, Danish: from logic 7.61, from English 7.39 (+0.22) GPT-6 Sol, Finnish: from logic 7.39, from English 6.56 (+0.83) GPT-6 Sol, Polish: from logic 7.44, from English 6.67 (+0.78) GPT-6 Sol, Czech: from logic 7.72, from English 6.67 (+1.06) GPT-6 Sol, Hebrew: from logic 6.06, from English 8.00 (-1.94) GPT-6 Sol, Arabic (Gulf): from logic 7.11, from English 7.12 (-0.01) GPT-6 Sol, Portuguese (Brazil): from logic 6.72, from English 8.94 (-2.22) GPT-6 Sol, Russian: from logic 7.61, from English 7.67 (-0.06) GPT-6 Sol, Turkish: from logic 7.44, from English 8.83 (-1.39) 9 of 21 Difference in points out of 10

No general route winner. Over all 126 versions the route from English averaged 7.86 and the route from the logic 7.80 (difference −0.05, 95% CI −0.42 to +0.32). GPT-6 Astra did better from the logic (+0.21), Opus 5.5 and GPT-6 Sol from English (−0.07 and −0.30), and the better route changes from language to language (Figure 13B).

Fidelity. 121 of 126 translations had every clue judged same; 5 carried one or two ambiguous clues and none a changed one. All 5 flagged versions came from the logic route (GPT-6 Sol from the logic in Korean, GPT-6 Sol from the logic in Dutch, GPT-6 Sol from the logic in Turkish, Opus 5.5 from the logic in Spanish (Latin America), Opus 5.5 from the logic in Dutch).

9.5 Rankings across the stages

Figure 14. Ranks across stages, because the Stage 2c rubric puts its scores on a different scale. A: the nine English writers ranked by mean English score in Stage 1 and in Stage 2c (three puzzles each). B: the nine translators ranked by mean over 21 languages in Stage 2 and Stage 2b (same rubric). C: the three Stage 2c candidates ranked among themselves in Stage 2, Stage 2b and Stage 2c (Stage 2c: mean of each translator's better route per language, fidelity-passed versions). Blue: rank improved; orange: fell; grey: unchanged. Hover for the underlying means. Ranks separated by less than the intervals of Figures 1, 5, 10 and 11 are not statistically distinct. Sources: GRIDIGMA out/rate/, out/s2/, out/s2b/, out/s2c/.
A. English writers Stage 1 English Stage 2c English GPT-6 Astra, Stage 1 English: rank 1 (9.67) GPT-6 Astra, Stage 2c English: rank 1 (8.87) GPT-6 Astra GPT-6 Astra GPT-6 Sol, Stage 1 English: rank 2 (9.30) GPT-6 Sol, Stage 2c English: rank 6 (7.22) GPT-6 Sol GPT-6 Sol Opus 5.5, Stage 1 English: rank 3 (9.26) Opus 5.5, Stage 2c English: rank 2 (8.80) Opus 5.5 Opus 5.5 Grok 4.6, Stage 1 English: rank 4 (9.19) Grok 4.6, Stage 2c English: rank 4 (7.81) Grok 4.6 Grok 4.6 Fable 5.1, Stage 1 English: rank 5 (8.89) Fable 5.1, Stage 2c English: rank 3 (8.04) Fable 5.1 Fable 5.1 Opus 5, Stage 1 English: rank 6 (8.59) Opus 5, Stage 2c English: rank 5 (7.78) Opus 5 Opus 5 Grok 4.7, Stage 1 English: rank 7 (8.48) Grok 4.7, Stage 2c English: rank 7 (6.91) Grok 4.7 Grok 4.7 Sonnet 5, Stage 1 English: rank 8 (7.52) Sonnet 5, Stage 2c English: rank 8 (6.28) Sonnet 5 Sonnet 5 GPT-6 Luna, Stage 1 English: rank 9 (7.22) GPT-6 Luna, Stage 2c English: rank 9 (6.13) GPT-6 Luna GPT-6 Luna B. Translators, 21 languages Stage 2 all nine Stage 2b all nine Fable 5.1, Stage 2: rank 4 (8.35) Fable 5.1, Stage 2b: rank 5 (8.59) Fable 5.1 Fable 5.1 GPT-6 Astra, Stage 2: rank 2 (9.20) GPT-6 Astra, Stage 2b: rank 1 (9.19) GPT-6 Astra GPT-6 Astra GPT-6 Luna, Stage 2: rank 8 (7.36) GPT-6 Luna, Stage 2b: rank 8 (7.87) GPT-6 Luna GPT-6 Luna GPT-6 Sol, Stage 2: rank 3 (8.39) GPT-6 Sol, Stage 2b: rank 3 (9.05) GPT-6 Sol GPT-6 Sol Grok 4.6, Stage 2: rank 6 (7.63) Grok 4.6, Stage 2b: rank 6 (8.29) Grok 4.6 Grok 4.6 Grok 4.7, Stage 2: rank 7 (7.57) Grok 4.7, Stage 2b: rank 7 (8.18) Grok 4.7 Grok 4.7 Opus 5, Stage 2: rank 5 (8.28) Opus 5, Stage 2b: rank 4 (8.71) Opus 5 Opus 5 Opus 5.5, Stage 2: rank 1 (9.28) Opus 5.5, Stage 2b: rank 2 (9.10) Opus 5.5 Opus 5.5 Sonnet 5, Stage 2: rank 9 (7.10) Sonnet 5, Stage 2b: rank 9 (7.30) Sonnet 5 Sonnet 5 C. The three Stage 2c candidates, ranked among themselves Stage 2 Stage 2b Stage 2c better route Opus 5.5, Stage 2: 1 of 3 (9.28; best in 10 languages of all nine) Opus 5.5, Stage 2b: 2 of 3 (9.10; best in 5 languages of all nine) Opus 5.5, Stage 2c: 2 of 3 (8.33, better route; won 9 languages of 21) Opus 5.5 Opus 5.5 GPT-6 Astra, Stage 2: 2 of 3 (9.20; best in 7 languages of all nine) GPT-6 Astra, Stage 2b: 1 of 3 (9.19; best in 6 languages of all nine) GPT-6 Astra, Stage 2c: 1 of 3 (8.45, better route; won 10 languages of 21) GPT-6 Astra GPT-6 Astra GPT-6 Sol, Stage 2: 3 of 3 (8.39; best in 1 languages of all nine) GPT-6 Sol, Stage 2b: 3 of 3 (9.05; best in 7 languages of all nine) GPT-6 Sol, Stage 2c: 3 of 3 (7.94, better route; won 2 languages of 21) GPT-6 Sol GPT-6 Sol

Across the three translation stages Opus 5.5 and GPT-6 Astra stay the two strongest candidates and trade first place: Opus 5.5 led in Stage 2, GPT-6 Astra in Stages 2b and 2c. GPT-6 Sol, which rose sharply on the plain source of Stage 2b, falls back to third once naturalness is weighted first. Among English writers the change of rubric moved GPT-6 Sol from 2nd to 6th.

9.6 Limitations of Stage 2c

  • One puzzle and one generation per translation cell. The discarded first run gives a rough measure of how much one cell moves between generations. For the route from English, whose prompt changed only in how the notes were numbered, the same model moved by a median of 0.64 and up to 1.33 points over the 15 cells rated in both runs (GPT-6 Sol, German: 7.78 then 6.44); the route from the logic moved by up to 3.89 points, but its prompt changed substantially between the runs. Ranks decided by less than about a point are a direction, not a verdict.
  • A different rubric. Stage 2c totals are not comparable with those of Stages 1, 2 and 2b; only ranks and per-criterion values carry across.
  • Model raters miss the last layer of naturalness. The owner read the Turkish winner and corrected one clause ("… öbürü en sondaydı" should read "… öbürü sonuncuydu", the ordinal form beside an ordinal); none of the nine raters had noticed it. Native readers remain the real test.
  • One fidelity judge, itself a contestant, as in every stage; the English intervals rest on three puzzles.

Declarations

Data availability. Every generated version, every rating with its note, every fidelity verdict and the analysis scripts are held by Fabervant and are available to researchers on request at the address above. The puzzles' solutions and stored logic are withheld, because the puzzles are live in the game.

Competing interests. Fabervant develops GRIDIGMA and has a commercial interest in its text quality. Fabervant is not affiliated with Anthropic, OpenAI or xAI, received no funding or model access from them for this study, and none of them reviewed this report.

Use of AI tools. The generation, rating and fidelity checks in this study were performed by the models under test. The pipeline, the analysis and the drafting of this report were carried out with AI coding assistants under the author's direction; the author set the research questions and the instructions, and is responsible for the content.

Ethics. No human participants were involved and no personal data was processed.

References

  1. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems 36, Datasets and Benchmarks Track. arXiv:2306.05685.
  2. Kocmi, T., & Federmann, C. (2023). Large language models are state-of-the-art evaluators of translation quality. Proceedings of the 24th Annual Conference of the European Association for Machine Translation, 193–203. arXiv:2302.14520.
  3. Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems 37. arXiv:2404.13076.
  4. Efron, B., & Tibshirani, R. J. (1993). An Introduction to the Bootstrap. Chapman & Hall.
  5. Kendall, M. G., & Babington Smith, B. (1939). The problem of m rankings. The Annals of Mathematical Statistics, 10(3), 275–287.
  6. Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.
  7. Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163.

How to cite

Güneş, C. (2026). Nine language models rewriting and translating logic-puzzle text: A blind cross-rating study (Version 3.2). Fabervant. https://fabervant.com/research/nine-model-prose-trial/

Version history: 1.0 (24 September 2026), Stage 1. 1.1 (25 September 2026): corrected the GPT-6 cost statement (their output tokens include reasoning, so the costs are complete, not a lower bound; no figure changed) and set out the Stage 2 design. 2.0 (25 September 2026): Stage 2 results across 21 languages (Section 7), Stage 1's effort defaults established (Section 2.7), abstract and conclusion updated. 3.0 (25 September 2026): Stages 2b and 2c (Sections 8 and 9); twelve new figures; Figures 1 and 5 (formerly 2) redrawn larger; the live English point in Figure 1 corrected from 6.34 to 6.33; abstract, research questions and conclusion updated. 3.1 (27 September 2026): Stage 3, authoring the puzzle logic (Section 10, RQ7); four new figures (15–18) and Tables 19 and 20; abstract, research questions and conclusion updated. 3.2 (29 September 2026): Stage 3 moved, its findings unchanged, to a paper of its own, Six Language Models Authoring Logic Puzzles Through a Game's Production Gates; Section 10, RQ7, Figures 15–18, Tables 19 and 20 and the Stage 3 sentences of the abstract, introduction, conclusion and declarations removed, restoring the wording of version 3.0.