Fabervant Research · Technical report · Stages 1, 2, 2b and 2c
Nine Language Models Rewriting and Translating Logic-Puzzle Text: A Blind Cross-Rating Study
Stage 1: English and Turkish. Stage 2: twenty-one languages. Stages 2b and 2c: plain words and a naturalness-first rubric
These puzzles are live: play GRIDIGMA free →
Abstract
Background. Text in a logic-puzzle game must read naturally while every clue keeps one exact logical meaning. We asked which current language models do this best, in English and in many languages, and at what cost.
Methods. In Stage 1, nine models from Anthropic, OpenAI and xAI each rewrote the English of three published GRIDIGMA puzzles (27 versions) and translated every model's English into Turkish (243 versions). Every model then rated every version blind on a fixed 10-point rubric (clarity, readability, naturalness), with the text players read today inserted unmarked as a control: 2,430 ratings of the 270 versions, plus 270 of the controls. In Stage 2, the same nine models translated one English text into 21 languages at a matched reasoning effort (189 translations), and each language's nine translations were rated blind by all nine models (1,701 ratings). A separate check judged each clue as keeping, changing or blurring its meaning. Stage 2b repeated Stage 2 on a plain-words edit of the source (189 translations, 1,701 ratings). Stage 2c required plain words with no arithmetic, weighted the rubric towards naturalness (6 of 10 points), had all nine models write the English of the three puzzles and three of them translate it into the 21 languages by two routes (153 versions, 1,377 ratings). We report means with 95% two-way cluster-bootstrap confidence intervals, rater agreement, and cost at list prices.
Results. Raters agreed well in Stages 1 and 2 (Kendall's W 0.67 to 0.70; ICC for the mean of nine raters 0.90 to 0.92), less on the plain source (Stage 2b: W 0.52; ICC 0.86) and least on the Stage 2c translations (W 0.36; ICC 0.78), so per-language picks there are directions only. In Stage 1, GPT-6 Astra wrote the highest-rated English (9.67/10, 95% CI 9.33–9.93) and Opus 5.5 the highest-rated Turkish (9.18, 8.78–9.52); the pair of the two scored 9.70 with no clue flagged. Today's live text scored below the rewrites in 265 of 270 batches, and Turkish naturalness varied far more across translators than across English sources. In Stage 2, across 21 languages, Opus 5.5 (9.28, 8.90–9.57) and GPT-6 Astra (9.20, 8.80–9.54) were statistically tied at the top and clear of the other seven; only 5 of 1,701 translated clues were flagged. On the plain source (Stage 2b) the weaker translators closed much of the gap, and GPT-6 Astra, Opus 5.5 and GPT-6 Sol were tied at the top. Under the naturalness-first rubric (Stage 2c), GPT-6 Astra and Opus 5.5 tied as English writers (8.87 and 8.80) and shared the translation lead (10 and 9 of 21 languages); translating from the stored logic instead of the English made no difference on average (7.80 against 7.86).
Conclusions. For this task the translator matters more than the English writer, and two translators, Opus 5.5 and GPT-6 Astra, stay at the top through every change of language, source and rubric, with notable per-language exceptions; the mechanical text came more from the inputs (a formula clue, technical labels) than from the models. The evidence rests on model raters; native-speaker ratings follow as an addendum.
Keywords: large language models; machine translation; multilingual evaluation; Turkish; LLM-as-a-judge; controlled text generation; game localisation; logic puzzles
1. Introduction
GRIDIGMA is a logic-grid deduction game played mostly on phones, published in English and Turkish. Each puzzle has a title, a one-line description, a short story and a set of clues. The text has two jobs that pull against each other. It should read easily and sound natural, because it is what draws a player in. And each clue is a piece of logic: together the clues admit exactly one answer, so a clue that says a little more, a little less, or can be read two ways breaks the puzzle. Rewriting such text for style, or translating it, is therefore a constrained generation task in which fluency and fidelity can trade off.
Language models are now used both to produce such text and to judge it. Model judges agree with human preferences at levels comparable to agreement between humans on some tasks [1], and have been used successfully to assess translation quality [2]. They also have known biases, among them a preference for their own generations [3]. A study that uses models as judges must therefore measure how much the judges agree and whether they favour themselves.
We asked six questions:
- RQ1. Which models write the clearest and most natural English puzzle text, and which translate it into the most natural Turkish?
- RQ2. Does the quality of a Turkish translation depend more on the translator or on the English it starts from?
- RQ3. Do models keep every clue's logical meaning while rewriting and translating?
- RQ4. What does each option cost, in money and in time?
- RQ5. Does the translators' ranking hold across languages when every model runs at the same reasoning effort?
- RQ6. Do the rankings hold when the source is written in plain words and the rubric weights naturalness first?
Stage 1 (Sections 2–6) answers RQ1–RQ4 for English and Turkish. Stage 2 (Section 7) answers RQ5 across 21 languages, and Stages 2b and 2c (Sections 8 and 9) answer RQ6.
2. Methods
2.1 Design
A fully crossed design: every model wrote English for every puzzle, every model translated every model's English, and every model rated every version. Ratings were blind to authorship. The text players read today was rated alongside, unmarked, as a control.
2.2 Models
| Model | Provider | Model ID |
|---|---|---|
| Sonnet 5 | Anthropic | claude-sonnet-5 |
| Opus 5 | Anthropic | claude-opus-5 |
| Opus 5.5 | Anthropic | claude-opus-5-5 |
| Fable 5.1 | Anthropic | claude-fable-5-1 |
| Grok 4.6 | xAI | grok-4.6 |
| Grok 4.7 | xAI | grok-4.7 |
| GPT-6 Luna | OpenAI | gpt-6-luna |
| GPT-6 Sol | OpenAI | gpt-6-sol |
| GPT-6 Astra | OpenAI | gpt-6-astra |
2.3 Materials
Three published GRIDIGMA daily puzzles, all 4 × 4 grids, chosen for different kinds of clue: The Seed-Sorting Trays (8 clues, positional), The Rope-Line Survey (9 clues, exactly one of three lettered field notes is false) and The Ropewalk Ledger (9 clues). The controls were the English and Turkish texts players read in the game on 24 September 2026. The live English had been written with an earlier model-based authoring pipeline and the live Turkish translated from it under the project's Turkish style guide.
2.4 Procedure
English rewrite (9 models × 3 puzzles = 27 versions). Each model rewrote a puzzle's title, description, story and clues. The instruction, in the product owner's words, was to make the text "simple, natural and fluid. The aim is not a piece of literature: the text is there to prepare the players for the puzzle and sweep them along." Models were free to improvise the story and wording, but had to keep the number and order of clues and make each clue "state exactly the same fact as the clue it replaces – no more, no less, and with only one possible reading". To let a model check its own wording, the prompt included the puzzle's answer and each clue's stored logic.
Turkish translation (9 translators × 9 English sources × 3 puzzles = 243 versions). Each model translated each model's English, its own included. The instruction asked for "the way a Turkish writer would tell it to Turkish players, not a word-for-word rendering", under the same fidelity constraint, using the Turkish edition's own cast of names.
Every call used the same one-line system prompt ("You are a writer preparing the text of a puzzle game."), and one user message per call, with no conversation, no tools and no other context. A reply that was not valid JSON with the right number of clues was re-sent once with a one-line reminder.
2.5 Measures
Rubric. Fixed before any rating: clarity 0–4 (one reading per clue; the setting easy to grasp), readability 0–3 (flows; easy on a phone) and naturalness 0–3 (sounds the way a native speaker would say it). The total runs from 0 to 10.
Fidelity. Each English clue was compared with the puzzle's stored logic and answer, and each Turkish clue with the English clue it translated, and judged same, changed (states a different fact) or ambiguous (admits a second reading that is a different fact). Opus 5.5 performed this check; every changed verdict was re-read by hand.
2.6 Rating and blinding
For each puzzle, each rater received one English batch (the 9 English versions plus the live English) and nine Turkish batches (the 9 translations of one English source plus the live Turkish). Within every batch the 10 texts were shuffled with a deterministic per-rater seed and labelled V1–V10. Raters were told that the versions came from different writers in random order, and to judge each on its own text without rewarding length, ornament or similarity. Each returned the three criterion scores and a one-sentence note per version. In total: 270 rating calls, and 2,700 scores, 2,430 of them of rewrites. Each model also rated its own work; we report every mean with and without self-ratings.
2.7 Settings
No temperature, top-p or reasoning effort was set, so every model ran at its provider's defaults. Reasoning effort was therefore not matched across models in Stage 1. The defaults were established afterwards: high for Sonnet 5, Opus 5 and Fable 5.1, and medium for Opus 5.5 (read back from the client); at least high for both Grok models (measured: with no effort set they reason as much as at high or more); and most likely medium for the three GPT-6 models (measured from output length against each level; firm for GPT-6 Luna, not proven for Sol and Astra). Stage 2 set the effort explicitly (Section 7.1).
2.8 Statistical analysis
The unit is one rater's total for one version. A model's English score is the mean over its 3 English versions × 9 raters; its Turkish score is the mean over its 27 translations × 9 raters. The 95% confidence intervals come from a two-way cluster bootstrap [4] that resamples raters and versions independently (4,000 resamples, fixed seed), so each interval carries both rater and item variation. The stability of the best pair was estimated by resampling raters and puzzles 1,000 times and counting how often each pair ranked first. Agreement between raters was measured with Kendall's coefficient of concordance W [5], corrected for ties, within each batch of 10 versions × 9 raters. The intraclass correlation [6] was measured as ICC(2,1) for a single rater and ICC(2,k) for the mean of the nine, two-way random effects with absolute agreement, and interpreted after [7]. The controls were compared with the rewrites by counting the batches in which the control scored below the batch's mean rewrite.
2.9 Cost accounting
Cost is each provider's list price on 24 September 2026 multiplied by the tokens used, the same for every provider: reasoning tokens billed as output, no cache discounts. The prices, in USD per million input/output tokens: Sonnet 5 2/10, Opus 5 5/25, Opus 5.5 4/20, Fable 5.1 10/50, Grok 4.6 2/6, Grok 4.7 2/6, GPT-6 Luna 0.1/0.5, GPT-6 Sol 2/10, GPT-6 Astra 10/50.
3. Results
3.1 Agreement between raters
Kendall's W averaged 0.70 over the 3 English batches (range 0.59–0.79) and 0.69 over the 27 Turkish batch groups (range 0.57–0.82): strong agreement on the ordering of versions. On absolute scores, a single rater's reliability was moderate (ICC(2,1) = 0.51 for English, 0.56 for Turkish), and the mean of nine raters was excellent (ICC(2,k) = 0.90 and 0.92). Every score below is therefore a nine-rater mean.
3.2 Scores by model (RQ1)
| Model | English [95% CI] | No self | Turkish [95% CI] | No self | Gave |
|---|---|---|---|---|---|
| GPT-6 Astra | 9.67 [9.33, 9.93] | 9.63 | 9.09 [8.77, 9.37] | 9.01 | 8.14 |
| GPT-6 Sol | 9.30 [8.93, 9.63] | 9.25 | 8.93 [8.61, 9.25] | 8.91 | 7.79 |
| Opus 5.5 | 9.26 [8.56, 9.85] | 9.25 | 9.18 [8.78, 9.52] | 9.17 | 7.56 |
| Grok 4.6 | 9.19 [8.44, 9.89] | 9.17 | 7.11 [6.62, 7.62] | 7.10 | 7.64 |
| Fable 5.1 | 8.89 [8.11, 9.67] | 8.83 | 8.69 [8.40, 9.00] | 8.63 | 8.28 |
| Opus 5 | 8.59 [7.78, 9.37] | 8.50 | 8.47 [7.97, 8.92] | 8.43 | 8.16 |
| Grok 4.7 | 8.48 [7.74, 9.11] | 8.46 | 7.60 [7.07, 8.14] | 7.57 | 7.85 |
| Sonnet 5 | 7.52 [6.19, 8.70] | 7.42 | 7.48 [6.99, 7.98] | 7.35 | 8.64 |
| GPT-6 Luna | 7.22 [6.26, 8.11] | 7.21 | 7.35 [6.85, 7.84] | 7.31 | 7.86 |
The English intervals are wide because each rests on only three versions; the top four English writers are not separable at this sample size. The Turkish intervals, each resting on 27 versions, are narrower: the intervals of Opus 5.5, GPT-6 Astra and GPT-6 Sol lie entirely above those of Grok 4.7, Sonnet 5, GPT-6 Luna and Grok 4.6, while Fable 5.1 and Opus 5 overlap both groups.
Self-preference. Removing each model's rating of its own work moved its mean by at most 0.13 points (Sonnet 5, Turkish), well inside every interval. We found no meaningful self-preference in this design.
| Model | EN clarity | EN readability | EN naturalness | TR clarity | TR readability | TR naturalness |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 3.89 | 2.96 | 2.81 | 3.82 | 2.68 | 2.59 |
| GPT-6 Sol | 3.78 | 2.96 | 2.56 | 3.64 | 2.73 | 2.56 |
| Opus 5.5 | 3.56 | 2.93 | 2.78 | 3.75 | 2.81 | 2.61 |
| Grok 4.6 | 3.41 | 2.89 | 2.89 | 3.00 | 2.30 | 1.81 |
| Fable 5.1 | 3.78 | 2.48 | 2.63 | 3.65 | 2.67 | 2.38 |
| Opus 5 | 3.63 | 2.52 | 2.44 | 3.44 | 2.67 | 2.35 |
| Grok 4.7 | 3.56 | 2.67 | 2.26 | 3.27 | 2.40 | 1.93 |
| Sonnet 5 | 2.74 | 2.52 | 2.26 | 3.19 | 2.34 | 1.95 |
| GPT-6 Luna | 2.63 | 2.30 | 2.30 | 3.15 | 2.13 | 2.06 |
| Model | Seed-Sorting Trays | Rope-Line Survey | Ropewalk Ledger |
|---|---|---|---|
| GPT-6 Astra | 9.67 | 9.56 | 9.78 |
| GPT-6 Sol | 9.33 | 9.22 | 9.33 |
| Opus 5.5 | 9.78 | 8.67 | 9.33 |
| Grok 4.6 | 9.89 | 8.78 | 8.89 |
| Fable 5.1 | 8.11 | 8.89 | 9.67 |
| Opus 5 | 9.33 | 8.56 | 7.89 |
| Grok 4.7 | 8.67 | 8.78 | 8.00 |
| Sonnet 5 | 8.67 | 7.67 | 6.22 |
| GPT-6 Luna | 7.56 | 6.33 | 7.78 |
3.3 Comparison with today's text
Today's live text scored below the mean of the rewrites in its batch in 265 of 270 batches (all 27 English batches, 238 of 243 Turkish), and at or below every single rewrite in 200 of 270. Its mean scores were 5.67–6.67 in English and 4.91–5.92 in Turkish (Table 5). The raters' notes point to the same defects again and again: an ornate story, and positional clues ("left", "right", "tray 3") given without saying how the trays are laid out. The example below shows the difference on one clue.
| Puzzle | Live English | Live Turkish | Live Turkish naturalness (of 3) |
|---|---|---|---|
| The Seed-Sorting Trays | 6.67 | 5.92 | 1.83 |
| The Rope-Line Survey | 6.67 | 5.69 | 1.63 |
| The Ropewalk Ledger | 5.67 | 4.91 | 1.43 |
Live text, clue 2
The clover sorting was done at the tray immediately to the right of the basil sorting.Opus 5.5 rewrite
The clover tray number is exactly one higher than the basil tray number.Story sentence the rewrite added
The trays are numbered 1 to 4 along the bench.
3.4 Translator versus source (RQ2)
Turkish naturalness ranged from 1.81 to 2.61 out of 3 across the nine translators, but only from 2.17 to 2.33 across the nine English sources. The English writer's quality barely carried into the Turkish; the translator decided it (Figure 2). The same holds for the total: Grok 4.6 wrote the fourth-best English (9.19) but produced the weakest Turkish (7.11).
3.5 Writer and translator pairs
| # | English by | Turkish by | Score | USD / puzzle | Points / USD | Flagged |
|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | Opus 5.5 | 9.70 | 0.0829 | 117 | 0 |
| 2 | GPT-6 Sol | Opus 5.5 | 9.40 | 0.0574 | 164 | 0 |
| 3 | Opus 5.5 | Opus 5.5 | 9.37 | 0.0833 | 113 | 0 |
| 4 | Grok 4.6 | Opus 5.5 | 9.35 | 0.1170 | 80 | 1 |
| 5 | GPT-6 Astra | GPT-6 Sol | 9.35 | 0.0498 | 188 | 0 |
| 6 | GPT-6 Astra | Fable 5.1 | 9.33 | 0.1497 | 62 | 0 |
| 7 | GPT-6 Sol | GPT-6 Astra | 9.31 | 0.0590 | 158 | 0 |
| 8 | GPT-6 Astra | GPT-6 Astra | 9.31 | 0.0898 | 104 | 0 |
| 9 | GPT-6 Sol | Opus 5 | 9.31 | 0.0689 | 135 | 0 |
| 10 | GPT-6 Sol | GPT-6 Sol | 9.29 | 0.0229 | 406 | 0 |
| 11 | Grok 4.6 | GPT-6 Sol | 9.26 | 0.0869 | 107 | 1 |
| 12 | Grok 4.6 | GPT-6 Astra | 9.26 | 0.1219 | 76 | 1 |
| 13 | Fable 5.1 | Opus 5.5 | 9.20 | 0.1387 | 66 | 0 |
| 14 | Opus 5.5 | GPT-6 Astra | 9.17 | 0.0906 | 101 | 0 |
| 15 | Opus 5.5 | Fable 5.1 | 9.11 | 0.1448 | 63 | 0 |
GPT-6 Astra with Opus 5.5 ranked first in 814 of 1,000 bootstrap resamples of raters and puzzles. The next most frequent leaders were GPT-6 Sol with GPT-6 Astra (55) and GPT-6 Sol with Opus 5.5 (28).
Figure 3 shows all 81 pairs averaged over the three puzzles, with Turkish naturalness beside the total. The full 9 × 9 Turkish grids follow (Tables 7–9). Rows give whose English was translated, columns who translated it; each cell is the nine-rater mean out of 10, and the diagonal is a model translating its own English.
Table 7. Turkish grid, The Seed-Sorting Trays
| English by | Sonnet 5 | Opus 5 | Opus 5.5 | Fable 5.1 | Grok 4.6 | Grok 4.7 | GPT-6 Luna | GPT-6 Sol | GPT-6 Astra |
|---|---|---|---|---|---|---|---|---|---|
| Sonnet 5 | 7.78 | 7.89 | 7.00 | 8.33 | 6.11 | 8.89 | 8.11 | 9.78 | 9.11 |
| Opus 5 | 9.11 | 8.44 | 9.56 | 9.56 | 6.00 | 9.67 | 9.11 | 8.33 | 8.89 |
| Opus 5.5 | 7.56 | 8.78 | 9.22 | 9.67 | 7.11 | 8.11 | 8.67 | 8.33 | 8.22 |
| Fable 5.1 | 9.33 | 7.67 | 9.78 | 8.56 | 6.78 | 8.33 | 7.44 | 8.78 | 8.22 |
| Grok 4.6 | 8.33 | 9.22 | 9.67 | 8.78 | 8.67 | 5.78 | 7.33 | 9.67 | 9.33 |
| Grok 4.7 | 7.56 | 9.00 | 9.22 | 9.00 | 7.33 | 7.11 | 5.22 | 9.33 | 8.78 |
| GPT-6 Luna | 7.89 | 8.72 | 9.61 | 7.94 | 7.44 | 4.17 | 7.33 | 9.22 | 8.44 |
| GPT-6 Sol | 8.00 | 9.67 | 9.89 | 7.56 | 8.67 | 7.78 | 6.33 | 9.33 | 8.33 |
| GPT-6 Astra | 8.33 | 8.89 | 10.00 | 9.00 | 9.00 | 7.78 | 6.89 | 9.89 | 8.22 |
Table 8. Turkish grid, The Rope-Line Survey
| English by | Sonnet 5 | Opus 5 | Opus 5.5 | Fable 5.1 | Grok 4.6 | Grok 4.7 | GPT-6 Luna | GPT-6 Sol | GPT-6 Astra |
|---|---|---|---|---|---|---|---|---|---|
| Sonnet 5 | 5.78 | 9.44 | 9.11 | 8.00 | 6.00 | 7.33 | 5.56 | 7.56 | 9.44 |
| Opus 5 | 7.11 | 6.67 | 8.67 | 8.44 | 6.44 | 6.44 | 7.33 | 8.33 | 9.56 |
| Opus 5.5 | 7.44 | 7.89 | 9.67 | 8.89 | 9.11 | 8.22 | 7.67 | 8.44 | 9.56 |
| Fable 5.1 | 5.78 | 6.22 | 9.11 | 7.78 | 5.67 | 8.22 | 6.44 | 9.11 | 9.11 |
| Grok 4.6 | 6.56 | 7.56 | 9.00 | 8.11 | 6.44 | 8.00 | 5.89 | 8.78 | 9.44 |
| Grok 4.7 | 6.33 | 8.44 | 8.78 | 8.89 | 6.67 | 6.44 | 7.78 | 9.00 | 9.67 |
| GPT-6 Luna | 6.56 | 6.22 | 8.00 | 7.67 | 5.56 | 8.44 | 9.00 | 8.44 | 9.89 |
| GPT-6 Sol | 6.67 | 9.50 | 9.50 | 9.56 | 7.39 | 6.89 | 6.78 | 9.61 | 10.00 |
| GPT-6 Astra | 6.33 | 7.33 | 9.44 | 8.44 | 7.00 | 8.00 | 7.67 | 8.11 | 9.44 |
Table 9. Turkish grid, The Ropewalk Ledger
| English by | Sonnet 5 | Opus 5 | Opus 5.5 | Fable 5.1 | Grok 4.6 | Grok 4.7 | GPT-6 Luna | GPT-6 Sol | GPT-6 Astra |
|---|---|---|---|---|---|---|---|---|---|
| Sonnet 5 | 6.44 | 9.22 | 8.78 | 9.11 | 7.11 | 7.11 | 7.44 | 9.56 | 8.89 |
| Opus 5 | 8.67 | 8.67 | 9.00 | 8.78 | 6.89 | 7.78 | 7.33 | 9.00 | 9.44 |
| Opus 5.5 | 7.11 | 9.22 | 9.56 | 8.33 | 5.78 | 8.11 | 7.22 | 9.44 | 9.44 |
| Fable 5.1 | 7.67 | 9.22 | 9.67 | 8.44 | 6.67 | 7.78 | 7.78 | 9.33 | 8.11 |
| Grok 4.6 | 8.00 | 8.67 | 9.89 | 8.78 | 6.89 | 7.89 | 7.22 | 9.56 | 9.22 |
| Grok 4.7 | 7.56 | 9.44 | 7.33 | 9.22 | 8.00 | 8.56 | 7.11 | 8.78 | 8.89 |
| GPT-6 Luna | 7.56 | 8.78 | 9.44 | 8.89 | 8.00 | 7.56 | 7.00 | 7.44 | 8.78 |
| GPT-6 Sol | 8.33 | 8.78 | 9.11 | 9.44 | 6.89 | 7.78 | 6.89 | 8.89 | 9.67 |
| GPT-6 Astra | 8.11 | 9.00 | 9.78 | 9.56 | 8.44 | 7.11 | 9.78 | 9.11 | 9.22 |
3.6 Fidelity (RQ3)
| Model | EN changed | EN ambiguous | TR changed | TR ambiguous |
|---|---|---|---|---|
| Opus 5 | 0 | 0 | 0 | 0 |
| Opus 5.5 | 0 | 0 | 0 | 0 |
| Fable 5.1 | 0 | 0 | 0 | 0 |
| GPT-6 Sol | 0 | 0 | 0 | 1 |
| Sonnet 5 | 0 | 0 | 0 | 2 |
| GPT-6 Astra | 0 | 0 | 0 | 2 |
| Grok 4.7 | 0 | 0 | 2 | 0 |
| Grok 4.6 | 0 | 1 | 0 | 3 |
| GPT-6 Luna | 0 | 0 | 4 | 2 |
No English rewrite changed a clue's meaning. The six changed verdicts, all in Turkish, were confirmed by hand: Grok 4.7 wrote tepki (reaction) for tepsi (tray) in two clues of one translation, and GPT-6 Luna left an English name in four clues where the Turkish edition uses its own.
3.7 Cost and time (RQ4)
| Model | Input tokens | Output tokens | Minutes / call | USD / English version | USD / Turkish version |
|---|---|---|---|---|---|
| Sonnet 5 | 2,761 | 3,238 | 0.50 | 0.0365 | 0.0381 |
| Opus 5 | 2,693 | 1,910 | 0.40 | 0.0528 | 0.0621 |
| Opus 5.5 | 2,698 | 1,796 | 0.32 | 0.0365 | 0.0479 |
| Fable 5.1 | 2,698 | 1,650 | 0.37 | 0.0903 | 0.1116 |
| Grok 4.6 | 2,534 | 12,572 | 3.21 | 0.0701 | 0.0817 |
| Grok 4.7 | 2,704 | 16,232 | 4.38 | 0.1084 | 0.1022 |
| GPT-6 Luna | 1,408 | 711 | 0.30 | 0.0004 | 0.0005 |
| GPT-6 Sol | 1,408 | 1,195 | 0.45 | 0.0097 | 0.0153 |
| GPT-6 Astra | 1,408 | 749 | 0.42 | 0.0368 | 0.0531 |
The Grok models spent most of their output on hidden reasoning, about twenty times the visible answer, and took roughly six to fifteen times longer per call than the other models. For the GPT-6 models the reasoning tokens are counted inside the reported output tokens, which are billed in full here; only their split between reasoning and answer is not reported. At list prices the 270 versions cost USD 15.16 and the rating USD 27.63. On cost alone, GPT-6 Luna in both roles gives the most points per dollar (7.50 out of 10), but it was the weakest model and made the most meaning errors as a translator. Among the high-scoring pairs, GPT-6 Sol in both roles gives 9.29 at 406 points per dollar, and GPT-6 Astra with GPT-6 Sol 9.35 at 188 (Table 6). Figure 4 sets each model's cost against its score.
4. Discussion
Two roles, two different models. English writing and Turkish translation rewarded different models. The best English came from GPT-6 Astra and GPT-6 Sol, the best Turkish from Opus 5.5, and the best pair combined them. A pipeline that uses one model for both roles gives up quality in one of them.
The translator decides naturalness. Naturalness spread about five times wider across translators (0.80 points) than across English sources (0.16 points). A well-written source helps clarity, but it does not make a weak translator sound native. For a product in several languages, the translator is the choice to get right.
Fluency without drift. None of the 27 English rewrites changed a clue's meaning, and the best models were both the most fluent and error-free. Here, readability did not cost correctness. The Turkish errors that did occur were lexical slips (tepki for tepsi) and untranslated names, not reasoning failures, and a mechanical check against the stored logic catches both.
Reasoning volume is not quality. The Grok models reasoned the most, cost the most time, and gave the weakest Turkish. At default settings, more hidden reasoning did not buy better prose.
Model judges. Agreement among the nine raters was high, and self-ratings moved no mean by more than 0.13, so the self-preference reported in other settings [3] did not show up measurably here. We attribute this partly to the blind, shuffled batches and to a rubric fixed in advance. Agreement among models is not agreement with people, however, and the native-speaker ratings planned as an addendum have to test that directly.
5. Limitations
- Small sample. Three puzzles, all 4 × 4 grids from one game. The English intervals rest on three versions per model and do not separate the top four writers.
- No human rating. All raters were models, and none is a native Turkish speaker. So far the only native Turkish reader of the outputs is the author. The absolute numbers are machine judgements.
- The fidelity checker was also a contestant. Opus 5.5 ran the clue check and scored best on it. Every changed verdict was confirmed by hand, but same verdicts were only spot-checked (26 English clues, all agreeing).
- Unmatched settings in Stage 1. Every model ran at its provider's default reasoning effort, and these differ (Section 2.7). Stage 2 matched them; see Section 7.3 for what that changed.
- Batch composition. A Turkish batch held one English source's nine translations, so raters never compared translations of different sources side by side.
- Plain text, not production format. GRIDIGMA's authoring format carries markup that this trial left out.
- Single generation. Each version was generated once; variation across repeated generations was not measured.
6. Conclusion
For English-and-Turkish text in a logic-puzzle game, the highest-rated combination was GPT-6 Astra writing the English and Opus 5.5 translating it into Turkish (9.70/10, no clue flagged, first in 81% of resamples). Every model's rewrite outscored the text in use. Translation quality depended far more on the translator than on the source. Across 21 languages at a matched effort (Stage 2), Opus 5.5 and GPT-6 Astra were statistically tied as the best translators and led in 17 of the 21 languages between them; neither Grok model, nor GPT-6 Luna, nor Sonnet 5 led in any language. Based on Stage 1, GRIDIGMA decided not to use either Grok model for Turkish translation, and is evaluating a split pipeline with Opus 5.5 as the Turkish translator, piloted on the production format first. Stages 2b and 2c tested what the product owner found in Stage 2's source, technical words and a formula clue. On plain words, and under a rubric that weights naturalness first, Opus 5.5 and GPT-6 Astra remain the two strongest translators (GPT-6 Astra 10, Opus 5.5 9 of 21 languages in Stage 2c), and the route from the stored logic is as good as the route from English on average. Per-language choices from one puzzle and one generation remain fragile.
7. Stage 2: twenty-one languages (RQ5)
7.1 Design and protocol changes
Stage 2 asks whether Stage 1's translators keep their ranking across languages when every model runs at the same reasoning effort. One puzzle, The Rope-Line Survey, was translated from a single English source (GPT-6 Astra's Stage 1 rewrite, unchanged) by all nine models into 21 languages: German, Japanese, French, Korean, Traditional and Simplified Chinese, Dutch, Spanish for Spain and for Latin America, Italian, Swedish, Norwegian (Bokmål), Danish, Finnish, Polish, Czech, Hebrew, Arabic for the Gulf countries, Brazilian Portuguese, Russian, and Turkish again, as the row that links the two stages. That gave 189 translations. For each language, all nine models rated its nine translations blind in one batch, with self-ratings included and the rubric of Stage 1 (1,701 ratings); the Turkish batch also carried the live Turkish text as a control. Opus 5.5 checked every clue of every translation, as in Stage 1. The run took place on 24 September 2026.
The protocol changed deliberately in five respects:
- Matched effort. Every call ran at reasoning effort high, set explicitly and recorded per call with the model identifier, the time and the full token usage. Before the run, each provider was shown to act on the setting: an invalid value was rejected by the OpenAI and xAI interfaces, and output grew between low and high for every provider.
- Clue order as in the game. The clue "exactly one of notes A, B and C is false" came first, as players see it; Stage 1 had it eighth.
- Names. Outside Turkish, each translator named the four people for its own readers; Turkish kept Stage 1's Turkish cast.
- One source text instead of nine, and one puzzle instead of three.
- Corrected wording in the fidelity prompt, which in Stage 1 stated the number of translations wrongly (eight for nine); the check itself was unchanged.
Because Stage 1's defaults were already high for Sonnet 5, Opus 5, Fable 5.1 and both Grok models (Section 2.7), the effort change between the stages affected only Opus 5.5 and the three GPT-6 models.
7.2 Results across all 21 languages
Raters agreed as in Stage 1: Kendall's W averaged 0.67 over the 21 language batches (range 0.44–0.84), and ICC(2,1) was 0.54 for a single rater and ICC(2,k) 0.91 for the mean of nine, over the 189 translations.
| Translator | Score [95% CI] | No self | Clarity | Readability | Naturalness | Best in | Flagged | USD | Minutes |
|---|---|---|---|---|---|---|---|---|---|
| Opus 5.5 | 9.28 [8.90, 9.57] | 9.26 | 3.70 | 2.85 | 2.72 | 10 | 0 / 2 | 0.057 | 0.4 |
| GPT-6 Astra | 9.20 [8.80, 9.54] | 9.13 | 3.86 | 2.77 | 2.58 | 7 | 0 / 0 | 0.076 | 0.7 |
| GPT-6 Sol | 8.39 [7.90, 8.86] | 8.40 | 3.40 | 2.67 | 2.33 | 1 | 0 / 1 | 0.022 | 0.7 |
| Fable 5.1 | 8.35 [7.88, 8.77] | 8.29 | 3.43 | 2.64 | 2.27 | 2 | 0 / 0 | 0.127 | 0.4 |
| Opus 5 | 8.28 [7.77, 8.79] | 8.25 | 3.39 | 2.67 | 2.23 | 1 | 0 / 0 | 0.067 | 0.5 |
| Grok 4.6 | 7.63 [6.98, 8.21] | 7.65 | 3.20 | 2.47 | 1.96 | 0 | 0 / 0 | 0.083 | 3.5 |
| Grok 4.7 | 7.57 [6.96, 8.16] | 7.51 | 3.29 | 2.34 | 1.94 | 0 | 0 / 0 | 0.127 | 4.5 |
| GPT-6 Luna | 7.36 [6.79, 7.93] | 7.38 | 3.14 | 2.36 | 1.86 | 0 | 1 / 1 | 0.0009 | 0.6 |
| Sonnet 5 | 7.10 [6.52, 7.68] | 6.95 | 3.20 | 2.16 | 1.74 | 0 | 0 / 0 | 0.042 | 0.5 |
A tie at the top. Opus 5.5 and GPT-6 Astra cannot be separated: in 1,000 resamples of raters and languages, Opus 5.5 ranked first 612 times and GPT-6 Astra 388 times, and no other model ever ranked first. Both lie clear of the other seven. Removing self-ratings moved no translator's mean by more than 0.15; Figure 9 shows how each rater scored each translator against the other eight. Figure 6 gives every translation's score, Figure 7 each translator's criteria and Figure 8 the spread of single ratings behind each mean.
Rankings travel, with exceptions. A language's own ranking of the nine translators agreed with the overall ranking with a median Spearman correlation of 0.72, highest in Norwegian (0.97) and lowest in Traditional Chinese (0.15), where Opus 5.5 scored its lowest (7.11) because two of its clues admitted a second reading. Table 13 lists the best translator in each language.
| Language | Best | Score | Runner-up | Score |
|---|---|---|---|---|
| German | GPT-6 Sol | 9.11 | Opus 5.5 | 8.89 |
| Japanese | GPT-6 Astra | 9.89 | Fable 5.1 | 9.56 |
| French | GPT-6 Astra | 10.00 | Opus 5.5 | 9.61 |
| Korean | Opus 5.5 | 9.39 | GPT-6 Astra | 9.33 |
| Chinese (Traditional) | Opus 5 | 9.33 | Grok 4.6 | 9.22 |
| Chinese (Simplified) | Opus 5.5 | 9.50 | GPT-6 Astra | 9.38 |
| Dutch | GPT-6 Astra | 9.56 | Opus 5.5 | 9.33 |
| Spanish (Spain) | Fable 5.1 | 9.56 | GPT-6 Astra | 9.44 |
| Spanish (Latin America) | GPT-6 Astra | 10.00 | Opus 5.5 | 9.67 |
| Italian | GPT-6 Astra | 9.89 | Opus 5.5 | 9.44 |
| Swedish | GPT-6 Astra | 9.00 | three tied | 8.78 |
| Norwegian | Opus 5.5 | 10.00 | GPT-6 Astra | 9.56 |
| Danish | GPT-6 Astra | 9.56 | GPT-6 Sol | 9.33 |
| Finnish | Opus 5.5 | 9.89 | GPT-6 Sol | 9.67 |
| Polish | Opus 5.5 | 9.67 | GPT-6 Astra | 8.78 |
| Czech | Opus 5.5 | 9.56 | Fable 5.1 | 9.22 |
| Hebrew | Opus 5.5 | 9.67 | GPT-6 Astra | 8.83 |
| Arabic (Gulf) | Opus 5.5 | 9.56 | Fable 5.1 | 8.56 |
| Portuguese (Brazil) | Opus 5.5 | 9.33 | GPT-6 Astra | 9.06 |
| Russian | Opus 5.5 | 9.56 | GPT-6 Astra | 9.44 |
| Turkish | Fable 5.1 | 9.56 | Opus 5 | 9.44 |
Fidelity. Of the 189 translations, five had a flagged clue: one changed (GPT-6 Luna's Polish named the wrong pennant colour) and four ambiguous, two of them in Opus 5.5's Traditional Chinese.
Cost and time. At list prices Stage 2 cost USD 33.67: 12.65 for the translations, 19.00 for the rating and 2.02 for the fidelity checks. At effort high the Grok models again spent by far the most output on reasoning and took 3.5–4.5 minutes per translation, against 0.4–0.7 for the others.
7.3 Turkish in both stages
| Translator | Stage 1 effort | Stage 1 | Stage 2 (high) | Change |
|---|---|---|---|---|
| Sonnet 5 | high | 6.33 | 7.22 | +0.89 |
| Opus 5 | high | 7.33 | 9.44 | +2.11 |
| Opus 5.5 | medium → high | 9.44 | 8.89 | −0.56 |
| Fable 5.1 | high | 8.44 | 9.56 | +1.11 |
| Grok 4.6 | high or more | 7.00 | 8.22 | +1.22 |
| Grok 4.7 | high or more | 8.00 | 7.67 | −0.33 |
| GPT-6 Luna | medium → high | 7.67 | 8.11 | +0.44 |
| GPT-6 Sol | medium? → high | 8.11 | 8.33 | +0.22 |
| GPT-6 Astra | medium? → high | 9.44 | 9.33 | −0.11 |
| Live text (control) | 5.44 | 5.11 |
Each cell is a single translation, generated once. Without repeated generations, a change of a point or so cannot be told apart from ordinary variation between runs, so no individual change here is evidence on its own. Two patterns are still worth noting. The largest gains came from models whose effort did not change (Opus 5 +2.11, Grok 4.6 +1.22, Fable 5.1 +1.11), which points to the new clue order or to variation between generations, not to effort. And raising Opus 5.5 and the GPT-6 models to high brought no consistent gain (from −0.56 to +0.44). The live Turkish text again scored below the translations' mean in all nine rater batches.
7.4 What Stage 2 adds
Stage 1's best Turkish translator, Opus 5.5, stays at the top across 21 languages, and GPT-6 Astra, Stage 1's best English writer, turns out to be an equally strong translator. The two are statistically tied and clearly ahead of the rest, and between them they led in 17 of the 21 languages. Choosing a translator by its overall rank is therefore a sound default. Where one language matters most, it pays to check that language: in Traditional Chinese the overall leader scored 2.2 points below that language's best, because of two ambiguous clues that a fidelity check would catch before release.
7.5 Limitations of Stage 2
- One text. Every language translates the same English puzzle; a ranking can move on other texts.
- No control outside Turkish. No live text exists in the other languages, so scores are relative within a language and not comparable across languages.
- Still no human raters. A blind rating sheet for native speakers exists; their scores will be added to this report as an addendum.
- One fidelity judge, itself a contestant, as in Stage 1.
- Half points. The GPT-6 raters sometimes gave half points (196 of 5,103 sub-scores in Stage 2 (5,130, controls included), 31 of 8,100 in Stage 1); they were kept as given.
- Model identity. For the OpenAI models the recorded model identifier is the one requested; it could not be confirmed independently through our access. The Anthropic and xAI identifiers are the providers' own.
8. Stage 2b: the same run on a plain-words source
8.1 Why and what changed
Reading Stage 2, the product owner found two faults in the English it translated (GPT-6 Astra's Stage 1 rewrite) and requested a re-run. The text used technical words that few players know (alidade, theodolite, glacier axe, snow probe, bone white, ochre, plum, and rope positions called lead, upper, lower and anchor), and these reached every language. And it stated the sum clue as a formula. A formula and a dictionary term translate without error, so Stage 2 partly measured the translation of an easy text.
Stage 2b repeated Stage 2 exactly (the nine translators, the 21 languages, the prompts, the rubric, blind rating by all nine models, the fidelity check, effort high on every call) on a source edited by hand from Astra's English. Only the board's item names and one clue changed: the item names became everyday words (compass, telescope, axe, pole; first, second, third, last), and the sum clue lost its number legend. Before any translation, Opus 5.5 judged every clue of the edited source same against the stored logic.
The sum clue was still arithmetic. Stage 2 read "Number the rope positions from front to back: lead = 1, upper = 2, lower = 3, anchor = 4. Bennett's and Gareth's position numbers add up to 6." Stage 2b read "Counting from the front, Bennett's place on the rope and Gareth's add up to 6." The legend was gone; the addition was not, and every Stage 2b translation carried it. Stage 2b therefore tests plain item names, not a formula-free source. Stage 2c (Section 9) removes the arithmetic.
8.2 Results
Raters agreed less than in Stage 2: Kendall's W averaged 0.52 over the 21 language batches (range 0.25–0.76, against 0.67), and ICC(2,1) was 0.42 for a single rater and ICC(2,k) 0.86 for the mean of nine (Stage 2: 0.54 and 0.91). The translations became more alike, which leaves the raters less to order them by.
| Translator | Stage 2 | Stage 2b [95% CI] | Change | Naturalness | Clarity | Best in |
|---|---|---|---|---|---|---|
| Opus 5.5 | 9.28 | 9.10 [8.81, 9.40] | −0.17 | 2.72 → 2.59 | 3.70 → 3.81 | 10 → 5 |
| GPT-6 Astra | 9.20 | 9.19 [8.85, 9.49] | −0.02 | 2.58 → 2.64 | 3.86 → 3.81 | 7 → 6 |
| GPT-6 Sol | 8.39 | 9.05 [8.73, 9.37] | +0.66 | 2.33 → 2.62 | 3.40 → 3.63 | 1 → 7 |
| Fable 5.1 | 8.35 | 8.59 [8.25, 8.93] | +0.24 | 2.27 → 2.32 | 3.43 → 3.70 | 2 → 1 |
| Opus 5 | 8.28 | 8.71 [8.41, 8.99] | +0.42 | 2.23 → 2.45 | 3.39 → 3.59 | 1 → 2 |
| Grok 4.6 | 7.63 | 8.29 [7.84, 8.70] | +0.67 | 1.96 → 2.21 | 3.20 → 3.50 | 0 → 1 |
| Grok 4.7 | 7.57 | 8.18 [7.77, 8.52] | +0.61 | 1.94 → 2.11 | 3.29 → 3.63 | 0 → 1 |
| GPT-6 Luna | 7.36 | 7.87 [7.30, 8.43] | +0.51 | 1.86 → 2.12 | 3.14 → 3.31 | 0 → 0 |
| Sonnet 5 | 7.10 | 7.30 [6.80, 7.82] | +0.21 | 1.74 → 1.85 | 3.20 → 3.25 | 0 → 0 |
The weaker translators caught up. 7 of the nine translators scored higher on the plain source, by 0.21 to 0.67 points; the two Stage 2 leaders did not (GPT-6 Astra −0.02, Opus 5.5 −0.17). The gap between the best and the weakest translator narrowed from 2.18 to 1.88 points, and naturalness rose for 8 of the nine. The top of the ranking became a three-way tie: GPT-6 Astra 9.19 [8.85, 9.49], Opus 5.5 9.10 [8.81, 9.40] and GPT-6 Sol 9.05 [8.73, 9.37]; in 1,000 resamples of raters and languages they ranked first 569, 309 and 121 times. GPT-6 Sol gained the most among the three (+0.66).
Per-language leaders moved. The best translator changed in 16 of the 21 languages. A language's own ranking agreed with the overall one with a median Spearman correlation of 0.75.
Fidelity and control. The check flagged 1 changed and 2 ambiguous clues, all in one translation (GPT-6 Luna, Korean). The live Turkish text, rated blind in the Turkish batch, scored 4.11 (Stage 2: 5.11) and fell below the batch's translation mean in 9 of 9 rater batches. At list prices Stage 2b cost USD 33.79: 11.72 for the translations, 20.13 for the rating and 1.94 for the checks.
8.3 Limitations of Stage 2b
- Not a formula-free source. The sum clue still asked the reader to add two places (Section 8.1), so Stage 2b does not show how the translators handle a source with no arithmetic.
- Edited by hand, not rewritten by a model. The edit isolates the two faults; it does not show what a model instructed to write plainly would produce.
- Original names still visible. As in Stage 2, each translator saw the puzzle's original board names with the rename map beside them.
- One generation per cell. Every Stage 2 limitation stands (Section 7.5); a change of a point or so in one language is within the variation between runs (Section 9.6).
9. Stage 2c: plain words, naturalness first
9.1 Why
The product owner requested Stage 2c as a further step towards natural text in every language, in his words: "I don't want a Japanese senior to open the puzzle, have a look at the story and close it right away as they found the translations too mechanical." Two findings shaped it. Stage 2b had not removed the arithmetic (Section 8.1). And the Stage 1 rubric itself rewarded number-talk: clarity carried 4 of the 10 points, and a numbered clue is maximally clear. GPT-6 Astra's Stage 1 English lead came mostly from clarity (3.89 of 4), while on naturalness Grok 4.6 was ahead (2.89 of 3, against 2.81; Table 3).
9.2 Design
| Stages 1, 2 and 2b | Stage 2c | |
|---|---|---|
| Writer prompt | plain words preferred | plain words required; no arithmetic, equations or numbered positions; a clue may be told as any statement that rules in and out exactly the same arrangements |
| Rubric | clarity 4, readability 3, naturalness 3 | naturalness 6, readability 2, clarity 2, read "as an older native speaker who has just opened this puzzle on a phone" |
| Fidelity check | same fact as the source text | same arrangements as the stored logic; reported as a gate beside the score |
| English | Stage 1: nine writers, three puzzles | nine writers, three puzzles, rated blind by all nine with the live English as control |
| Translation source | Astra's Stage 1 English (2b: hand-edited) | the best-rated Stage 2c English of The Rope-Line Survey that passed the check, chosen by code: GPT-6 Astra's |
| Routes | from English | from English, and from the stored logic directly (no text) |
| Translators | all nine | Opus 5.5, GPT-6 Astra, GPT-6 Sol |
Unchanged: effort high set and recorded on every call, blind shuffled rating by all nine models with self-ratings included and reported without, Opus 5.5 as the only fidelity judge, and the game's clue order. Stage 2c produced 27 English versions (243 ratings) and 126 translations into the same 21 languages as Stage 2 (1134 ratings). Confidence intervals use the two-way cluster bootstrap of Section 2.8 (raters and puzzles for English, raters and languages for translation; 4,000 resamples).
A first translation run was discarded. It handed the translators the stored logic, in which the "exactly one note is false" clue names its notes by internal clue number; because players see that clue first, the numbers on screen differ by one, and the GPT-6 translators attached the letters to the wrong clues. The from-logic route also saw only the board's technical labels. The second run gave every prompt the notes by on-screen position and the labels' plain meaning; all results here are from it. The first run's 30 rated versions (5 languages) are kept as a record and used only to measure run-to-run change (Section 9.6).
Scale. Because the rubric weights changed, Stage 2c scores are not on the scale of Stages 1, 2 and 2b and are never compared with them directly; comparisons across stages use ranks (Figure 14).
9.3 English
| Writer | Score [95% CI] | No self | Naturalness | Flagged | Rank 1 → 2c |
|---|---|---|---|---|---|
| GPT-6 Astra | 8.87 [8.35, 9.22] | 8.79 | 4.98 | 0 | 1 → 1 |
| Opus 5.5 | 8.80 [8.33, 9.19] | 8.73 | 4.89 | 0 | 3 → 2 |
| Fable 5.1 | 8.04 [7.26, 8.78] | 8.00 | 4.39 | 1 | 5 → 3 |
| Grok 4.6 | 7.81 [6.74, 8.78] | 8.00 | 3.96 | 0 | 4 → 4 |
| Opus 5 | 7.78 [6.67, 8.93] | 7.75 | 4.20 | 0 | 6 → 5 |
| GPT-6 Sol | 7.22 [6.67, 7.89] | 7.21 | 3.98 | 1 | 2 → 6 |
| Grok 4.7 | 6.91 [5.70, 8.04] | 6.90 | 3.43 | 0 | 7 → 7 |
| Sonnet 5 | 6.28 [5.63, 6.96] | 6.19 | 3.70 | 0 | 8 → 8 |
| GPT-6 Luna | 6.13 [5.22, 7.11] | 6.02 | 3.46 | 1 | 9 → 9 |
| Live English (control) | 4.48 | 2.30 |
Every writer dropped the arithmetic. GPT-6 Astra (8.87) and Opus 5.5 (8.80) led and cannot be separated; GPT-6 Sol, 2nd in Stage 1, fell to 6th once number-talk stopped paying. The live English scored 4.48. Agreement: Kendall's W 0.74 over the three puzzle batches (control included, as in Stage 1), ICC(2,1) 0.54 and ICC(2,k) 0.91 over the 27 rewrites.
9.4 Translation
Raters agreed with Kendall's W 0.36 over the 21 language batches (range 0.15–0.59); ICC(2,1) was 0.28 and ICC(2,k) 0.78 over the 126 translations. This is the lowest agreement in the study. Part of it is expected: the three translators are the strongest of Stage 2, so the versions differ less and leave raters less to agree on; part may come from the heavier weight on naturalness, the most subjective criterion. Either way a single rater's score says little here, the nine-rater mean is moderately reliable, and per-language winners separated by less than about a point should be read as directions only.
| Translator | Better route | Won | From English | From logic | Logic − English |
|---|---|---|---|---|---|
| GPT-6 Astra | 8.45 | 10 | 8.06 [7.72, 8.34] | 8.27 [7.92, 8.57] | +0.21 [−0.13, +0.58] |
| Opus 5.5 | 8.33 | 9 | 7.94 [7.54, 8.32] | 7.87 [7.32, 8.38] | −0.07 [−0.68, +0.53] |
| GPT-6 Sol | 7.94 | 2 | 7.57 [7.03, 8.05] | 7.27 [6.79, 7.74] | −0.30 [−0.85, +0.24] |
Two translators share the top. GPT-6 Astra won 10 languages with the best mean (8.45), Opus 5.5 9 (8.33), and GPT-6 Sol 2 (7.94). The single highest score of the stage was Opus 5.5's Turkish from English (9.61). Each candidate's better route scored between 7.12 and 9.61 in every language; the live Turkish, rated in the Turkish batch, scored 2.83.
No general route winner. Over all 126 versions the route from English averaged 7.86 and the route from the logic 7.80 (difference −0.05, 95% CI −0.42 to +0.32). GPT-6 Astra did better from the logic (+0.21), Opus 5.5 and GPT-6 Sol from English (−0.07 and −0.30), and the better route changes from language to language (Figure 13B).
Fidelity. 121 of 126 translations had every clue judged same; 5 carried one or two ambiguous clues and none a changed one. All 5 flagged versions came from the logic route (GPT-6 Sol from the logic in Korean, GPT-6 Sol from the logic in Dutch, GPT-6 Sol from the logic in Turkish, Opus 5.5 from the logic in Spanish (Latin America), Opus 5.5 from the logic in Dutch).
9.5 Rankings across the stages
Across the three translation stages Opus 5.5 and GPT-6 Astra stay the two strongest candidates and trade first place: Opus 5.5 led in Stage 2, GPT-6 Astra in Stages 2b and 2c. GPT-6 Sol, which rose sharply on the plain source of Stage 2b, falls back to third once naturalness is weighted first. Among English writers the change of rubric moved GPT-6 Sol from 2nd to 6th.
9.6 Limitations of Stage 2c
- One puzzle and one generation per translation cell. The discarded first run gives a rough measure of how much one cell moves between generations. For the route from English, whose prompt changed only in how the notes were numbered, the same model moved by a median of 0.64 and up to 1.33 points over the 15 cells rated in both runs (GPT-6 Sol, German: 7.78 then 6.44); the route from the logic moved by up to 3.89 points, but its prompt changed substantially between the runs. Ranks decided by less than about a point are a direction, not a verdict.
- A different rubric. Stage 2c totals are not comparable with those of Stages 1, 2 and 2b; only ranks and per-criterion values carry across.
- Model raters miss the last layer of naturalness. The owner read the Turkish winner and corrected one clause ("… öbürü en sondaydı" should read "… öbürü sonuncuydu", the ordinal form beside an ordinal); none of the nine raters had noticed it. Native readers remain the real test.
- One fidelity judge, itself a contestant, as in every stage; the English intervals rest on three puzzles.
Declarations
Data availability. Every generated version, every rating with its note, every fidelity verdict and the analysis scripts are held by Fabervant and are available to researchers on request at the address above. The puzzles' solutions and stored logic are withheld, because the puzzles are live in the game.
Competing interests. Fabervant develops GRIDIGMA and has a commercial interest in its text quality. Fabervant is not affiliated with Anthropic, OpenAI or xAI, received no funding or model access from them for this study, and none of them reviewed this report.
Use of AI tools. The generation, rating and fidelity checks in this study were performed by the models under test. The pipeline, the analysis and the drafting of this report were carried out with AI coding assistants under the author's direction; the author set the research questions and the instructions, and is responsible for the content.
Ethics. No human participants were involved and no personal data was processed.
References
- Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems 36, Datasets and Benchmarks Track. arXiv:2306.05685.
- Kocmi, T., & Federmann, C. (2023). Large language models are state-of-the-art evaluators of translation quality. Proceedings of the 24th Annual Conference of the European Association for Machine Translation, 193–203. arXiv:2302.14520.
- Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems 37. arXiv:2404.13076.
- Efron, B., & Tibshirani, R. J. (1993). An Introduction to the Bootstrap. Chapman & Hall.
- Kendall, M. G., & Babington Smith, B. (1939). The problem of m rankings. The Annals of Mathematical Statistics, 10(3), 275–287.
- Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.
- Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163.
How to cite
Güneş, C. (2026). Nine language models rewriting and translating logic-puzzle text: A blind cross-rating study (Version 3.2). Fabervant. https://fabervant.com/research/nine-model-prose-trial/