OpenPsy

FRAME-02: the replication of FRAME-01

FRAME-02 ran the FRAME-01 framing study again through OpenPsy on 2026-09-25 and 2026-09-26, with 320 preregistered sessions in two samples of 40 per condition and one hypothesis per model. Claude Sonnet 5 recommended Plan A in 79 of 79 sessions under the gain frame and in 11 of 80 under the loss frame, and OpenPsy recorded that hypothesis as confirmed. Claude Opus 5 recommended it in 79 of 80 and 76 of 77, and under the registered rule OpenPsy recorded that hypothesis as falsified. Both samples show the same pattern.

What was asked

The design is the FRAME-01 design: a two by two of Frame (gain, "300 of those people will be protected", against loss, "600 of those people will be left unprotected") by Model (Claude Opus 5 against Claude Sonnet 5 on one Claude Code CLI route with identical settings), with language model agents as the participants, one fresh instance per session, four calls per session. The preregistration registered two hypotheses. H1a: Claude Sonnet 5 recommends adopting Plan A more often under the gain frame than under the loss frame. H1b: the same for Claude Opus 5. Each is judged by the exact conditional test of that model's frame difference, stratified by sample and Holm-adjusted across the four simple effects, at alpha .05. A result that is not significant in the predicted direction counts against the hypothesis. The size, 40 per condition in each of two samples, was chosen for a simulated power of 0.94 for the registered smallest difference of thirty percentage points.

The recorded result

Recorded results of FRAME-02 as a two-by-two table. Gain frame: Claude Opus 5, 99% (79 of 80), mean confidence 5.6; Claude Sonnet 5, 100% (79 of 79), mean confidence 4.5. Loss frame: Claude Opus 5, 99% (76 of 77), mean confidence 5.2; Claude Sonnet 5, 14% (11 of 80), mean confidence 2.9. Row totals: Gain frame, 99% (158 of 159); Loss frame, 55% (87 of 157). Column totals: Claude Opus 5, 99% (155 of 157); Claude Sonnet 5, 57% (90 of 159).
The product's own recorded two-by-two summary. For Claude Sonnet 5 the registered exact test is significant, Holm-adjusted p less than .001. For Claude Opus 5 the same test finds no difference, Holm-adjusted p = 1.00. The declared Wald tests of the omnibus terms could not run because conditions sit at the ceiling, so the preregistered fallback ran: a Firth-penalized logistic regression found a significant main effect of Frame (p less than .001), no significant main effect of Model (p = .071) and a significant interaction (p less than .001). Every registered sensitivity bound for the five missing sessions gives the same verdicts.

Note. N = 320 preregistered confirmatory sessions in two samples, n = 80 per condition (40 per sample). Five sessions that ended in an ambiguous provider call were excluded as missing under the registered rule, so the analysed n is 80, 79, 77 and 80 for Gain frame with Claude Opus 5, Gain frame with Claude Sonnet 5, Loss frame with Claude Opus 5 and Loss frame with Claude Sonnet 5. Percentages are sessions recommending Plan A.

Study AGENT-B9FB8AA2D19BF8FF15BAD36154CCC6E6, registered under the title FRAME-01 as version 5 of the same experiment. Registration hash sha256:f624c956ecbb04e2188b775389c976e5f6c2c53611f9fe2c2dbda82bdec363de. Result digest sha256:2cf4606819ed0233e4169b555ea69cce5410459762c9b08c0ac48c76ac10a0a0. Software release d6834711.

Grouped bar chart of the percentage of FRAME-02 sessions recommending Plan A, by Frame and Model, with Wilson 95% confidence intervals and brackets on the significant comparisons. Gain frame: Claude Opus 5 99%, Claude Sonnet 5 100%. Loss frame: Claude Opus 5 99%, Claude Sonnet 5 14%.
The recorded percentages as grouped bars, with Wilson 95% confidence intervals and brackets on the comparisons the registered tests found significant.
FRAME-02 interaction plot in two panels. Top: the percentage of sessions recommending Plan A, with each condition as a dot and a line joining each model across the gain and loss frames. Claude Opus 5 runs flat from 99% to 99%; Claude Sonnet 5 falls from 100% to 14%, so the two lines part under the loss frame. Error bars are Wilson 95% confidence intervals centred on each dot. Bottom: the mean confidence rating from 1 to 7. Claude Opus 5 goes from 5.6 to 5.2 and Claude Sonnet 5 from 4.5 to 2.9, with 95% intervals from the recorded standard deviations.
The same recorded result as lines. Each dot is a condition's recorded value and each line joins one model across the two frames, so the interaction reads as the lines parting under the loss frame. The lower panel shows the same shape in the confidence ratings. Error bars are 95% confidence intervals centred on each dot.

Each sample on its own

The two samples are two complete copies of the design, collected in the same run and interleaved, so that the effect can be seen in each independently.

Sessions recommending Plan A, by sample
SampleGain frame, Claude Opus 5Gain frame, Claude Sonnet 5Loss frame, Claude Opus 5Loss frame, Claude Sonnet 5
Sample 140 of 40 (100%)40 of 40 (100%)38 of 39 (97%)6 of 40 (15%)
Sample 239 of 40 (98%)39 of 39 (100%)38 of 38 (100%)5 of 40 (13%)
Both79 of 80 (99%)79 of 79 (100%)76 of 77 (99%)11 of 80 (14%)

How the run went

The measure pilot, the execution pilot and the confirmatory ran through OpenPsy's command line under a supervisor from 2026-09-25 09:55 to 2026-09-26 10:30 Pacific time, in seven batches, each under a fresh sealed launch receipt. Five of the 320 confirmatory sessions ended in an ambiguous provider call: two timed out at the registered 240 seconds, one of them during a network outage on the machine, and three failed in the transport a few seconds after their write-ahead capture. None was resent. Each was reconciled as an execution failure with the researcher's key and counted as missingness under the registered rule, and each is recorded as an incident with the route probe that followed it. 315 of the 320 scheduled sessions reached a terminal state, and every condition met the registered rule of at least 90 percent analysed.

The run needed under three hours of model time. It took a day because the product re-read its whole ledger on every governance check, walked its custody for every batch receipt and ran sessions one at a time; those costs are being removed before the next study.

Against FRAME-01

The 200-session FRAME-01 run of 2026-09-24 recorded Claude Opus 5 at 50 of 50 under both frames and Claude Sonnet 5 at 49 of 50 under the gain frame against 17 of 50 under the loss frame. Its single pooled hypothesis was recorded as inconclusive because the Frame main effect could only be tested through the fallback. FRAME-02 asked one question per model and reproduced the split exactly: the frame does not move Claude Opus 5, and it moves Claude Sonnet 5 from near-unanimous recommendation to near-unanimous refusal, with a lower loss-frame rate this time, 14 percent against 34 percent, in both samples.

What this says and does not say

Within this scenario, this wording and this route, one of the two models shows a large framing effect and the other shows none. The result is about these two models on those two days on the pinned Claude Code CLI, with the registered task instructions. It says nothing about other scenarios, other providers or human participants, and the open-ended rationale was not coded. Hidden reasoning was present in about half of Claude Opus 5's calls and about a tenth of Claude Sonnet 5's, as the custody records, and is not part of the analysis.