FRAME-03: the seven-point replication
FRAME-03 ran the FRAME-02 design again through OpenPsy on 2026-09-28, with the recommendation asked as a likelihood from 1 to 7 instead of yes or no, in 160 preregistered sessions in two samples of 20 per cell. Both models rated Plan A less likely to be recommended under the loss frame: Claude Sonnet 5 from 5.1 to 3.4 and Claude Opus 5 from 5.8 to 5.1, each with a Holm-adjusted p less than .001. Nine of the ten preregistered hypotheses were confirmed. One was falsified: Claude Opus 5 did not rate its confidence higher under the gain frame (4.7 against 4.6, Holm-adjusted p = .67).
What was asked
The scenario and the frame manipulation are those of FRAME-01 and FRAME-02: a two by two of Frame (gain, "300 of those people will be protected", against loss, "600 of those people will be left unprotected") by Model (Claude Opus 5 against Claude Sonnet 5 on one Claude Code CLI route with identical settings), with language model agents as the participants and one fresh instance per session.
What changed is how the answer is taken. The primary outcome is now the likelihood of recommending Plan A on a scale from 1 to 7: "How likely are you to recommend adopting Plan A, from 1 (not at all likely) to 7 (very likely), where 4 means neither likely nor unlikely?" The confidence rating, also from 1 to 7, became a tested secondary outcome. The free-text rationale was coded with five codes taken from the FRAME-02 post hoc coding, applied blind by a registered model coder, Claude Opus 5. That coder is the same model as one participant level; it saw the codebook and one reply at a time, and never a condition, a model name or an outcome. Two manipulation checks close every session: how many people are protected (300 in every condition) and how many are unprotected (600). Each session made six participant calls and one coder call.
The preregistration registered ten hypotheses. The likelihood and confidence hypotheses are judged by the exact permutation test of the difference in means, and the code hypotheses by the exact conditional test; each test is stratified by sample and Holm-adjusted across the four simple effects of its own outcome, at alpha .05. Under the registered rule, a result that is not significant in the predicted direction counts against the hypothesis.
- H1a, on the likelihood, the primary outcome. Claude Sonnet 5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame.
- H1b, on the likelihood, the primary outcome. Claude Opus 5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame.
- H2a, on the confidence rating. Claude Sonnet 5 rates its confidence higher under the gain frame than under the loss frame.
- H2b, on the confidence rating. Claude Opus 5 rates its confidence higher under the gain frame than under the loss frame.
- H3a, on the confidence rating. Under the gain frame, Claude Opus 5 rates its confidence higher than Claude Sonnet 5 does.
- H3b, on the confidence rating. Under the loss frame, Claude Opus 5 rates its confidence higher than Claude Sonnet 5 does.
- H4a, on the rationale code BETTER_THAN_NOTHING. Claude Sonnet 5's rationale compares Plan A with doing nothing (the floor) more often under the gain frame than under the loss frame.
- H4b, on the rationale code SHORTFALL. Claude Sonnet 5's rationale treats the unprotected share as the verdict more often under the loss frame than under the gain frame.
- H4c, on the rationale code OPTIMALITY_BURDEN. Claude Sonnet 5's rationale demands that the plan be shown best more often under the loss frame than under the gain frame.
- H4d, on the rationale code FRAME_EQUIVALENCE. Claude Opus 5's rationale recognises the framing more often under the loss frame than under the gain frame.
The confirmatory run was 160 sessions in two samples of 20 per cell, so 40 per condition in all, after a measure pilot of 40 sessions and an execution pilot of 4.
The recorded result
Note. N = 160 preregistered confirmatory sessions in two samples, n = 40 per condition (20 per sample). Four sessions that ended in an ambiguous provider call before the likelihood answer were excluded as missing under the registered rule, so the analysed n is 37, 40, 39 and 40 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. Each cell gives the mean likelihood on the 1 to 7 scale, its 95% confidence interval from Student's t, n and SD, and the mean confidence rating.
Study AGENT-4077193D973E70CDF2A39DCC053D80C5. Registration hash sha256:b2ce589e323d8aa833f3176065b5f1bdae33cb59eae3050644d58627aa926b8a. Result digest sha256:d4a4431898f3834d01808cfc4c8801e432b30628f92df89841c7d65182ae2336. Software release 6aa25d67.
Likelihood of recommending Plan A
The primary outcome, as bars or as lines. Claude Sonnet 5's line falls more steeply than Claude Opus 5's, which is the interaction.
| Frame | Claude Opus 5 | Claude Sonnet 5 | Total |
|---|---|---|---|
| Gain frame | 5.8 95% CI 5.6 to 6.0 n = 37, SD 0.52 Confidence 4.7 (n = 37) | 5.1 95% CI 4.9 to 5.2 n = 40, SD 0.42 Confidence 3.6 (n = 39) | 5.4 95% CI 5.3 to 5.6 n = 77, SD 0.59 Confidence 4.1 (n = 76) |
| Loss frame | 5.1 95% CI 4.9 to 5.3 n = 39, SD 0.68 Confidence 4.6 (n = 39) | 3.4 95% CI 3.1 to 3.6 n = 40, SD 0.74 Confidence 2.5 (n = 40) | 4.2 95% CI 4.0 to 4.5 n = 79, SD 1.12 Confidence 3.6 (n = 79) |
| Total | 5.4 95% CI 5.3 to 5.6 n = 76, SD 0.70 Confidence 4.7 (n = 76) | 4.2 95% CI 4.0 to 4.5 n = 80, SD 1.04 Confidence 3.0 (n = 79) |
| Hypothesis | Statement | Outcome | Recorded result | Verdict |
|---|---|---|---|---|
| H1a | Claude Sonnet 5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame. | Likelihood (primary) | Gain frame: mean 5.1 (n = 40). Loss frame: mean 3.4 (n = 40). Holm-adjusted p < .001. | Confirmed |
| H1b | Claude Opus 5 rates the likelihood of recommending Plan A higher under the gain frame than under the loss frame. | Likelihood (primary) | Gain frame: mean 5.8 (n = 37). Loss frame: mean 5.1 (n = 39). Holm-adjusted p < .001. | Confirmed |
| H2a | Claude Sonnet 5 rates its confidence higher under the gain frame than under the loss frame. | Confidence | Gain frame: mean 3.6 (n = 39). Loss frame: mean 2.5 (n = 40). Holm-adjusted p < .001. | Confirmed |
| H2b | Claude Opus 5 rates its confidence higher under the gain frame than under the loss frame. | Confidence | Gain frame: mean 4.7 (n = 37). Loss frame: mean 4.6 (n = 39). Holm-adjusted p = .67. | Falsified |
| H3a | Under the gain frame, Claude Opus 5 rates its confidence higher than Claude Sonnet 5 does. | Confidence | Claude Opus 5: mean 4.7 (n = 37). Claude Sonnet 5: mean 3.6 (n = 39). Holm-adjusted p < .001. | Confirmed |
| H3b | Under the loss frame, Claude Opus 5 rates its confidence higher than Claude Sonnet 5 does. | Confidence | Claude Opus 5: mean 4.6 (n = 39). Claude Sonnet 5: mean 2.5 (n = 40). Holm-adjusted p < .001. | Confirmed |
| H4a | Claude Sonnet 5's rationale compares Plan A with doing nothing (the floor) more often under the gain frame than under the loss frame. | Rationale code BETTER_THAN_NOTHING | Gain frame: 31 of 38 (82%). Loss frame: 5 of 40 (13%). Holm-adjusted p < .001. | Confirmed |
| H4b | Claude Sonnet 5's rationale treats the unprotected share as the verdict more often under the loss frame than under the gain frame. | Rationale code SHORTFALL | Loss frame: 37 of 40 (93%). Gain frame: 15 of 38 (39%). Holm-adjusted p < .001. | Confirmed |
| H4c | Claude Sonnet 5's rationale demands that the plan be shown best more often under the loss frame than under the gain frame. | Rationale code OPTIMALITY_BURDEN | Loss frame: 28 of 40 (70%). Gain frame: 2 of 38 (5%). Holm-adjusted p < .001. | Confirmed |
| H4d | Claude Opus 5's rationale recognises the framing more often under the loss frame than under the gain frame. | Rationale code FRAME_EQUIVALENCE | Loss frame: 35 of 39 (90%). Gain frame: 15 of 37 (41%). Holm-adjusted p < .001. | Confirmed |
The registered linear model, fitted to every analysed reply with one intercept per sample, recorded these omnibus terms.
| Term | b | SE | t(151) | p |
|---|---|---|---|---|
| Main effect of Frame | 1.20 | 0.10 | 12.41 | p < .001 |
| Main effect of Model | 1.23 | 0.10 | 12.68 | p < .001 |
| Interaction | −0.99 | 0.19 | −5.12 | p < .001 |
Confidence rating
Claude Sonnet 5 was less confident under the loss frame (3.6 against 2.5, Holm-adjusted p less than .001). Claude Opus 5 was not (4.7 against 4.6, Holm-adjusted p = .67), and was more confident than Claude Sonnet 5 under both frames. The recorded view gives these means without an interval, so none is drawn.
Each sample on its own
The two samples are two complete copies of the design, collected in the same run and interleaved, so that the effect can be seen in each independently. In both samples each model is lower under the loss frame, and Claude Sonnet 5 is lower than Claude Opus 5 under each frame. All four missing sessions fell in sample 1.
| Sample | Gain frame, Claude Opus 5 | Gain frame, Claude Sonnet 5 | Loss frame, Claude Opus 5 | Loss frame, Claude Sonnet 5 |
|---|---|---|---|---|
| Sample 1 | 5.9, 95% CI 5.7 to 6.2 n = 17, SD 0.56 | 5.1, 95% CI 4.9 to 5.3 n = 20, SD 0.45 | 4.9, 95% CI 4.6 to 5.3 n = 19, SD 0.71 | 3.3, 95% CI 2.9 to 3.7 n = 20, SD 0.80 |
| Sample 2 | 5.7, 95% CI 5.5 to 5.9 n = 20, SD 0.47 | 5.1, 95% CI 4.9 to 5.2 n = 20, SD 0.39 | 5.3, 95% CI 5.0 to 5.5 n = 20, SD 0.64 | 3.5, 95% CI 3.1 to 3.8 n = 20, SD 0.69 |
| Both | 5.8, 95% CI 5.6 to 6.0 n = 37, SD 0.52 | 5.1, 95% CI 4.9 to 5.2 n = 40, SD 0.42 | 5.1, 95% CI 4.9 to 5.3 n = 39, SD 0.68 | 3.4, 95% CI 3.1 to 3.6 n = 40, SD 0.74 |
The rationale codes
Each free-text rationale was coded once by Claude Opus 5 on its own pinned route, blind to the condition and the model level of the reply. Before any reply of this study existed, the coder agreed with the expected codes on 40 of 40 calibration cases. No coded reply named a condition or a model in its own words. Four codes carry a hypothesis (H4a to H4d) and one, NO_STATED_DOWNSIDE, is reported without one. Each code's four simple effects are Holm-adjusted within that code only. The recorded view gives no interval for these shares, so none is drawn.
BETTER_THAN_NOTHING: improvement over inaction (H4a)
The rationale compares the plan with doing nothing and concludes that some protection is better than none. Claude Sonnet 5 wrote this in 31 of 38 gain-frame rationales and 5 of 40 loss-frame ones.
SHORTFALL: unprotected share as the verdict (H4b)
The rationale judges the plan by the size of the group it leaves unprotected and treats that shortfall as a poor result in its own right. Claude Sonnet 5 wrote this in 37 of 40 loss-frame rationales and 15 of 38 gain-frame ones.
OPTIMALITY_BURDEN: plan must be shown best (H4c)
The rationale withholds endorsement because nothing shows the plan is the best available option. Claude Sonnet 5 wrote this in 28 of 40 loss-frame rationales and 2 of 38 gain-frame ones.
FRAME_EQUIVALENCE: framing recognised (H4d)
The rationale states that the two descriptions are the same outcome, names the framing, or says the recommendation would not change under the other wording. Claude Opus 5 wrote this in 35 of 39 loss-frame rationales and 15 of 37 gain-frame ones. Claude Sonnet 5 never did.
NO_STATED_DOWNSIDE: unopposed benefit (no hypothesis)
The rationale treats the absence of any stated cost or competing plan as a reason the benefit stands. Both models wrote this more often under the gain frame. No hypothesis was registered for this code.
| Code | Gain frame, Claude Opus 5 | Gain frame, Claude Sonnet 5 | Loss frame, Claude Opus 5 | Loss frame, Claude Sonnet 5 | Hypothesis |
|---|---|---|---|---|---|
| SHORTFALL Unprotected share as the verdict | 2 of 37 (5%) | 15 of 38 (39%) | 2 of 39 (5%) | 37 of 40 (93%) | H4b: confirmed, Holm-adjusted p < .001 |
| BETTER_THAN_NOTHING Improvement over inaction | 37 of 37 (100%) | 31 of 38 (82%) | 31 of 39 (79%) | 5 of 40 (13%) | H4a: confirmed, Holm-adjusted p < .001 |
| FRAME_EQUIVALENCE Framing recognised | 15 of 37 (41%) | 0 of 38 (0%) | 35 of 39 (90%) | 0 of 40 (0%) | H4d: confirmed, Holm-adjusted p < .001 |
| NO_STATED_DOWNSIDE Unopposed benefit | 28 of 37 (76%) | 24 of 38 (63%) | 12 of 39 (31%) | 1 of 40 (3%) | No hypothesis registered |
| OPTIMALITY_BURDEN Plan must be shown best | 3 of 37 (8%) | 2 of 38 (5%) | 3 of 39 (8%) | 28 of 40 (70%) | H4c: confirmed, Holm-adjusted p < .001 |
Manipulation checks
Every answered check was correct: no check failed in any condition. The only missing answers belong to sessions that ended in an ambiguous provider call.
| Check | Gain frame, Claude Opus 5 | Gain frame, Claude Sonnet 5 | Loss frame, Claude Opus 5 | Loss frame, Claude Sonnet 5 |
|---|---|---|---|---|
| Protected count 300 in every condition | 37 of 40 passed 0 failed, 3 missing (ambiguous call) | 39 of 40 passed 0 failed, 1 missing (ambiguous call) | 39 of 40 passed 0 failed, 1 missing (ambiguous call) | 40 of 40 passed 0 failed |
| Unprotected count 600 in every condition | 37 of 40 passed 0 failed, 3 missing (ambiguous call) | 38 of 40 passed 0 failed, 2 missing (ambiguous call) | 39 of 40 passed 0 failed, 1 missing (ambiguous call) | 40 of 40 passed 0 failed |
What we found and what was falsified
Nine hypotheses were confirmed.
- H1a. Claude Sonnet 5 rated the likelihood of recommending Plan A higher under the gain frame than under the loss frame, 5.1 against 3.4 (Holm-adjusted p < .001).
- H1b. Claude Opus 5 did the same, 5.8 against 5.1 (Holm-adjusted p < .001).
- H2a. Claude Sonnet 5 rated its confidence higher under the gain frame than under the loss frame, 3.6 against 2.5 (Holm-adjusted p < .001).
- H3a. Under the gain frame, Claude Opus 5 rated its confidence higher than Claude Sonnet 5, 4.7 against 3.6 (Holm-adjusted p < .001).
- H3b. Under the loss frame, Claude Opus 5 rated its confidence higher than Claude Sonnet 5, 4.6 against 2.5 (Holm-adjusted p < .001).
- H4a. Claude Sonnet 5 compared Plan A with doing nothing in 31 of 38 gain-frame rationales against 5 of 40 loss-frame ones (Holm-adjusted p < .001).
- H4b. Claude Sonnet 5 judged the plan by the unprotected share in 37 of 40 loss-frame rationales against 15 of 38 gain-frame ones (Holm-adjusted p < .001).
- H4c. Claude Sonnet 5 asked for the plan to be shown best in 28 of 40 loss-frame rationales against 2 of 38 gain-frame ones (Holm-adjusted p < .001).
- H4d. Claude Opus 5 recognised the framing in 35 of 39 loss-frame rationales against 15 of 37 gain-frame ones (Holm-adjusted p < .001).
One hypothesis was falsified. H2b predicted that Claude Opus 5 would rate its confidence higher under the gain frame than under the loss frame. It did not: its mean confidence was 4.7 against 4.6 (Holm-adjusted p = .67), so under the registered rule H2b is falsified. Claude Opus 5 rated Plan A less likely under the loss frame without becoming less confident in its answer.
Nothing else was tested as a hypothesis. The omnibus terms, the other simple effects and the NO_STATED_DOWNSIDE code are reported as the product recorded them, without a verdict. Every registered sensitivity bound for the missing sessions gives the same ten verdicts.
How the run went
The run was planned with Claude Sonnet 4.5 as the rationale coder, a model other than the two participants. Two real preflights refused that route before any session ran: the first on the dated model name the provider reports, which the product was then changed to accept, and the second on the shape of the messages the pinned command line sends to that model, which the product cannot yet declare. Under the owner's fallback of 2026-09-25, the coder became Claude Opus 5, the same model as one participant level, blind to level and condition. The next preflight passed, and the coder agreed with the expected codes on 40 of 40 calibration cases.
The first registration with that coder, AGENT-3B1C33B3F617B569A75482CEC6C4213B, ran its measure pilot at four sessions per cell. One call met the provider's 529 overloaded error; the session was reconciled as an execution failure, never resent, and the pilot completed. The registered go rule then refused the measure decision, because it counts that session against parse validity and so left its cell at 3 of 4, below 90 percent. The refusal was recorded and signed in that registration, which is kept with all its data. The same design was registered again, unchanged except that the measure pilot returned to ten sessions per cell, the programme default of FRAME-01 and FRAME-02, as AGENT-4077193D973E70CDF2A39DCC053D80C5.
That registration started at 01:24 on 2026-09-28, Pacific time. The measure pilot of 40 sessions had no invalid item of 360, and its decision was approved and signed at 01:37. The execution pilot of 4 sessions passed, the preregistration was finalized at 01:40, and the confirmatory stage launched at 01:41. The preregistration was finalized before any result was viewed, and its one amendment, the finalization itself, was not informed by outcomes.
Six of the 160 confirmatory sessions ended in an ambiguous provider call and were reconciled as execution failures with the researcher's key; none was resent. One connection was refused on the scenario call, four connections dropped at the same instant (two on the likelihood answer, one on the confidence rating and one on the unprotected count check), and one scenario call did not answer within the registered 240 seconds. That timeout placed a collection hold; the product's own route probe verified the transport again, and the hold was cleared through the product before the run resumed. Four of the six sessions have no likelihood answer and are counted as missing under the registered rule, so 156 of 160 were analysed: three in Gain frame with Claude Opus 5 and one in Loss frame with Claude Opus 5, all in sample 1. The other two answered the likelihood and lack only a later answer. The coded outcomes count six sessions as missing and analyse 154.
The confirmatory stage completed at 03:27, and the recorded analysis was committed at 03:48 Pacific on 2026-09-28 with result digest sha256:d4a4431898f3834d01808cfc4c8801e432b30628f92df89841c7d65182ae2336. The disclosure of the result is recorded in the ledger. The app's two-minute route limit ended the first attempts to open the result before the view arrived, so the recorded view was captured from the product's results bridge with a longer wait; it carries the same result digest, and the capture added no event to the ledger.
Against FRAME-02
FRAME-03 kept the FRAME-02 scenario, wording, models and route, and changed these things: the recommendation is asked as a likelihood from 1 to 7 instead of yes or no, the confidence rating is tested instead of exploratory, the rationale is coded under five registered codes by a coder call, and two manipulation checks close every session, so each session made six participant calls against four in FRAME-02. It is also smaller, 20 sessions per cell in each of two samples against 40.
What the seven-point scale showed
In FRAME-02, Claude Opus 5 recommended Plan A in 79 of 80 gain-frame sessions and 76 of 77 loss-frame sessions. At that ceiling the yes or no outcome showed no frame effect (Holm-adjusted p = 1.00), and its hypothesis H1b was recorded as falsified. On the scale, Claude Opus 5 gave a mean of 5.8 under the gain frame and 5.1 under the loss frame, and the corresponding hypothesis, H1b, was confirmed (Holm-adjusted p < .001). The scale made visible a graded frame effect within Claude Opus 5 that the binary outcome could not show. It still leaned towards recommending under both frames, with no reply below 4, but less strongly under the loss frame, where 7 of its 39 replies sat at the midpoint of 4 against none of its 37 gain-frame replies.
| Condition | 1 | 2 | 3 | 4 | 5 | 6 | 7 | n |
|---|---|---|---|---|---|---|---|---|
| Gain frame, Claude Opus 5 | 0 | 0 | 0 | 0 | 9 | 26 | 2 | 37 |
| Gain frame, Claude Sonnet 5 | 0 | 0 | 0 | 2 | 33 | 5 | 0 | 40 |
| Loss frame, Claude Opus 5 | 0 | 0 | 0 | 7 | 21 | 11 | 0 | 39 |
| Loss frame, Claude Sonnet 5 | 0 | 2 | 25 | 9 | 4 | 0 | 0 | 40 |
What stayed the same is the Claude Sonnet 5 effect. In FRAME-02 it recommended Plan A in 79 of 79 gain-frame sessions and 11 of 80 loss-frame sessions; in FRAME-03 its mean was 5.1 against 3.4, with 25 of its 40 loss-frame replies at 3. Both studies confirm its frame hypothesis. The scale adds that its gain-frame endorsement was moderate, with 33 of 40 replies at 5, where the yes or no outcome recorded every gain-frame session as a recommendation. Under the gain frame the two models also differ on the scale (5.8 against 5.1, Holm-adjusted p < .001), where FRAME-02 found no difference (79 of 80 against 79 of 79, Holm-adjusted p = 1.00).
The confidence rating and the codes add what the recommendation alone does not. Claude Sonnet 5's confidence fell with the frame (3.6 against 2.5) and Claude Opus 5's did not (4.7 against 4.6). The codes describe what each model wrote: Claude Opus 5 recognised the framing in 35 of 39 loss-frame rationales, and Claude Sonnet 5 never did, while Claude Sonnet 5's loss-frame rationales judged the plan by the unprotected share (37 of 40) and asked for it to be shown best (28 of 40).
The comparison cannot claim more than that. The two studies asked different primary questions, ran on different days with different sample sizes, and no test was run between them. The move of H1b from falsified to confirmed is a difference between two designs, not a measured change in Claude Opus 5. Every number in this comparison is one that the two recorded views state.
What this says and does not say
Within this scenario, this wording and this route, both models rate Plan A less likely to be recommended when the outcome is described as people left unprotected, and Claude Sonnet 5 moves much further than Claude Opus 5. The result is about these two models on 2026-09-28 on the pinned Claude Code CLI, with the registered task instructions. It says nothing about other scenarios, other providers or human participants.
The rationale coder was Claude Opus 5, the same model as one participant level. It was blind to condition and level, but coding by a single model coder is a declared limitation, and every reply and every coder reply is archived for a later recode. The route sent its token budget message wrapped in system-reminder tags to Claude Sonnet 5 and bare to Claude Opus 5; the study conversation and the instruction were identical. Hidden thinking, which the route does not return, was present in 83 of 223 and 90 of 235 of Claude Opus 5's calls under the gain and loss frames, and in 34 of 236 and 25 of 240 of Claude Sonnet 5's, and is not part of the analysis.
Study AGENT-4077193D973E70CDF2A39DCC053D80C5. Registration hash sha256:b2ce589e323d8aa833f3176065b5f1bdae33cb59eae3050644d58627aa926b8a. Result digest sha256:d4a4431898f3834d01808cfc4c8801e432b30628f92df89841c7d65182ae2336. Software release 6aa25d67.