OpenPsy

FRAME-05: the OpenAI replication

In 158 complete sessions, GPT-5.6 Sol rated Plan A more likely to be recommended under the gain frame than the loss frame, 5.75 against 3.46 on the 1–7 scale. GPT-5.6 Terra did too, 6.28 against 4.85. Two of 160 scheduled sessions remain permanently unavailable. These are post-response exploratory and sensitivity results after recovery, not a completed confirmatory replication or a model-equivalence finding.

What was asked

FRAME-05 carried forward FRAME-04’s two-by-two Frame × Model scenario, with GPT-5.6 Sol and GPT-5.6 Terra as participants. Gain says that Plan A protects 300 of 900 people; Loss says it leaves 600 unprotected. Each fresh participant had a six-turn conversation: scenario, 1–7 likelihood of recommending Plan A, free-text rationale, 1–7 confidence, and two count checks. The original two sample blocks each scheduled 20 sessions per condition, 40 per condition overall.

The original directional questions, measures, codebook, sample blocks and tests remained fixed. The participant route changed to the models’ native subscription CLIs. GPT-6 Astra was the primary rationale coder and Claude Sonnet 5 a separate independent coder, each shown the same codebook and one rationale at a time without supplied model, frame, rating or the other coder’s labels. The rationale text itself can contain clues to its frame. The original six-turn and coder protocols stayed separate from diagnostics, pilots and calibration.

  1. H1a GPT-5.6 Terra rates recommendation likelihood higher under Gain than Loss.
  2. H1b GPT-5.6 Sol rates recommendation likelihood higher under Gain than Loss.
  3. H2a GPT-5.6 Terra rates confidence higher under Gain than Loss.
  4. H2b GPT-5.6 Sol rates confidence higher under Gain than Loss.
  5. H3a GPT-5.6 Sol rates confidence higher than GPT-5.6 Terra under Gain.
  6. H3b GPT-5.6 Sol rates confidence higher than GPT-5.6 Terra under Loss.
  7. H4a GPT-5.6 Terra rationales show BETTER_THAN_NOTHING more often under Gain.
  8. H4b GPT-5.6 Terra rationales show SHORTFALL more often under Loss.
  9. H4c GPT-5.6 Terra rationales show OPTIMALITY_BURDEN more often under Loss.
  10. H4d GPT-5.6 Sol rationales show FRAME_EQUIVALENCE more often under Loss.

Recommendation and confidence comparisons use the original two-sided exact permutation tests stratified by sample. Binary rationale codes use the original exact conditional tests. Each set of four simple effects is Holm-adjusted within its own outcome and coder stream; there is no study-wide adjustment. The fixed original plan required complete-study admission. The two unavailable sessions and outcome-informed recovery mean these analyses cannot issue new prospective confirmed or falsified verdicts.

The recorded result

FRAME-05 exploratory two-by-two table. Gain Sol: mean recommendation 5.75, mean confidence 4.88, n 40. Gain Terra: 6.28 and 5.60, n 40. Loss Sol: 3.46 and 4.46, n 39. Loss Terra: 4.85 and 3.87, n 39.
The observed mean recommendation rating is higher under Gain for each participant model. The cell means describe the 158 complete sessions; they are not a prospective confirmation of the original 160-session run.

Note. N = 160 scheduled, with 40 per condition and two blocks of 20. Both Gain cells have 40 eligible responses; both Loss cells have 39. One Sol-Loss claim failed before a model request and one Terra-Loss claim has unresolved delivery; neither was resent, replaced or counted as a zero rating. The two 1–7 measures and their original nulls are retained in the private data. The recommendation charts draw newly calculated 95% Student’s t confidence intervals of the observed means, using n − 1 degrees of freedom, as in FRAME-04. These individual unadjusted intervals do not account for missingness, recovery selection or provider dependence.

FRAME-05 verified export: sha256:1a3430c1bd216c30bd13c6374a1d07887f21b691b61d1619e6f6e84ad6071c97. Exploratory analysis: sha256:0e422625c52a7c8a2cd92220b704ddca44f19efdb9906329f6b3c41466358c44. Mean confidence intervals were calculated separately for this presentation from the same verified ratings.

FRAME-05 exploratory recommendation grouped bar chart with 95% confidence intervals and significance brackets
The observed means with 95% confidence intervals. Brackets mark significant comparisons: pairwise stars use Holm-adjusted p values; the main effect of Frame uses its unadjusted omnibus p. These remain exploratory results.
FRAME-05 exploratory recommendation interaction line chart with 95% confidence intervals
The same means with 95% confidence intervals, joined across frames. The observed Gain-minus-Loss difference is 2.29 points for Sol and 1.43 for Terra. See the statistical comparisons below for the verified tests.
Statistical comparisons: Recommendation

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

Recommendation — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss2.28854.5868< .001***
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss1.42882.8658< .001***
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain-0.5250-1.0500< .001***
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss-1.3846-2.7711< .001***

Two-sided exact permutation tests stratified by sample. Differences are observed rating-point differences; the test statistic sums the within-sample differences. Error-bar overlap is not the test used for these stars.

Recommendation — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame1.8610t(153) = 17.688< .001***
main effect of Model-0.9548t(153) = -9.076< .001***
interaction0.8596t(153) = 4.085< .001***

OLS Wald t tests with one intercept per sample. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

The newly calculated recommendation mean intervals are: Gain Sol, 95% CI [5.59, 5.91]; Gain Terra, 95% CI [6.08, 6.47]; Loss Sol, 95% CI [3.23, 3.69]; Loss Terra, 95% CI [4.58, 5.11]. Each is M ± t.975, n−1 × SD/√n. Error-bar overlap is not a test of a difference. The frozen exact comparisons returned Holm-adjusted p = 2.12 × 10−19 for Sol’s recommendation difference and 9.95 × 10−13 for Terra’s. Treating the two missing ratings at the least favourable permitted endpoints leaves Gain-minus-Loss differences of at least 2.20 and 1.375 points, respectively. These missing-rating bounds do not address provider context, response dependence or selection during the recovery.

Confidence rating

FRAME-05 exploratory confidence two-by-two table: Gain Sol 4.88, Gain Terra 5.60, Loss Sol 4.46, Loss Terra 3.87
The confidence means use the same 158 complete sessions. The observed Sol Gain-minus-Loss comparison is weaker under the least-favourable missing-rating sensitivity case.
FRAME-05 exploratory confidence grouped bar chart
Confidence is 4.88 against 4.46 for Sol and 5.60 against 3.87 for Terra, Gain then Loss.
FRAME-05 exploratory confidence interaction line chart
The two observed confidence lines cross: Terra is higher under Gain and Sol is higher under Loss.
Statistical comparisons: Confidence

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

Confidence — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss0.41350.8289.016*
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss1.72823.4684< .001***
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain-0.7250-1.4500< .001***
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss0.58971.1895.002**

Two-sided exact permutation tests stratified by sample. Differences are observed rating-point differences; the test statistic sums the within-sample differences. Error-bar overlap is not the test used for these stars.

Note. The observed Sol confidence difference has adjusted p = .0160, but its least-favourable missing-rating scenario has adjusted p = .0757. Terra’s difference has adjusted p = 3.44 × 10−15 and retains its direction in the corresponding sensitivity case. These are exploratory comparisons, not a uniform Sol confidence advantage.

Rationale codes

Astra and Sonnet attempted their original assignments independently. Of 316 nominal roles on the 158 eligible rationales, 315 replies were captured and 314 produced valid codes: 158 Astra and 156 Sonnet. One earlier Sonnet action has uncertain delivery and no admissible reply; another Sonnet reply was captured but invalid and remains five null codes. The 156 valid pairs agree on 729 of 780 code decisions (93.5%), but 51 disagreements remain. Each chart below names its coder. No original labels were reconciled.

Original rationale-code streams; descriptive yes counts (Astra and Sonnet kept separate)
CodeAstra Gain SolAstra Gain TerraAstra Loss SolAstra Loss TerraSonnet Gain SolSonnet Gain TerraSonnet Loss SolSonnet Loss Terra
SHORTFALL
Unprotected share as the verdict
0 of 40 (0.0%)0 of 40 (0.0%)26 of 39 (66.7%)24 of 39 (61.5%)0 of 39 (0.0%)0 of 40 (0.0%)24 of 38 (63.2%)18 of 39 (46.2%)
BETTER_THAN_NOTHING
Improvement over inaction
1 of 40 (2.5%)11 of 40 (27.5%)2 of 39 (5.1%)24 of 39 (61.5%)1 of 39 (2.6%)1 of 40 (2.5%)1 of 38 (2.6%)6 of 39 (15.4%)
FRAME_EQUIVALENCE
Framing recognised
0 of 40 (0.0%)0 of 40 (0.0%)0 of 39 (0.0%)0 of 39 (0.0%)0 of 39 (0.0%)0 of 40 (0.0%)0 of 38 (0.0%)0 of 39 (0.0%)
NO_STATED_DOWNSIDE
Unopposed benefit
18 of 40 (45.0%)30 of 40 (75.0%)0 of 39 (0.0%)8 of 39 (20.5%)17 of 39 (43.6%)30 of 40 (75.0%)0 of 38 (0.0%)8 of 39 (20.5%)
OPTIMALITY_BURDEN
Plan must be shown best
3 of 40 (7.5%)0 of 40 (0.0%)26 of 39 (66.7%)4 of 39 (10.3%)6 of 39 (15.4%)1 of 40 (2.5%)24 of 38 (63.2%)11 of 39 (28.2%)
Original coder agreement by rationale code; 156 valid paired rationales per code
CodeAgreements / 156Observed agreementCohen’s κ
SHORTFALL14794.2%0.861
BETTER_THAN_NOTHING12781.4%0.319
FRAME_EQUIVALENCE156100.0%Undefined: both streams constant zero
NO_STATED_DOWNSIDE15599.4%0.986
OPTIMALITY_BURDEN14492.3%0.789

The five code rows contain 780 comparable decisions: 729 agreements and 51 disagreements. Ten additional code pairs are unavailable across the two rationales without valid Sonnet secondary codes. Agreement does not establish coder accuracy or participant-model equivalence.

Unprotected share as the verdict (SHORTFALL)

The rationale treats the 600 unprotected people, or the majority left unprotected, as a poor result.

Astra coded Sol 0/40 Gain and 26/39 Loss, and Terra 0/40 Gain and 24/39 Loss. Sonnet coded Sol 0/39 Gain and 24/38 Loss, and Terra 0/40 Gain and 18/39 Loss.

Astra primary Unprotected share as the verdict FRAME-05 exploratory two-by-two table
Original Astra-primary SHORTFALL decisions; each cell states yes out of valid Astra codes.
Astra primary Unprotected share as the verdict FRAME-05 exploratory bar chart
Original Astra-primary SHORTFALL shares by Frame and participant Model; exploratory descriptive display.
Astra primary Unprotected share as the verdict FRAME-05 exploratory line chart
The same original Astra-primary SHORTFALL shares joined across frames.
Statistical comparisons: SHORTFALL — Astra primary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

SHORTFALL — Astra primary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss-0.6667Exact conditional< .001***
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss-0.6154Exact conditional< .001***
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain0.0000Exact conditional1.000—
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss0.0513Exact conditional1.000—

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

SHORTFALL — Astra primary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame-4.9888χ²(1) = 88.047< .001***
main effect of Model0.1091χ²(1) = 0.011.916—
interaction-0.2182χ²(1) = 0.011.916—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

See the separate Sonnet coding for Unprotected share as the verdict

Sonnet is the independently retained secondary stream. It has one absent earlier reply and one captured-invalid later reply, so its valid cell denominators are 39 Gain Sol, 40 Gain Terra, 38 Loss Sol and 39 Loss Terra.

Sonnet secondary Unprotected share as the verdict FRAME-05 exploratory two-by-two table
Original Sonnet-secondary SHORTFALL decisions, without substitution from Astra.
Sonnet secondary Unprotected share as the verdict FRAME-05 exploratory bar chart
Separate Sonnet-secondary shares; coder differences remain visible.
Sonnet secondary Unprotected share as the verdict FRAME-05 exploratory line chart
The same separate Sonnet-secondary shares joined across frames.
Statistical comparisons: SHORTFALL — Sonnet secondary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

SHORTFALL — Sonnet secondary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss-0.6316Exact conditional< .001***
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss-0.4615Exact conditional< .001***
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain0.0000Exact conditional1.000—
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss0.1700Exact conditional.338—

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

SHORTFALL — Sonnet secondary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame-4.6063χ²(1) = 68.601< .001***
main effect of Model0.3555χ²(1) = 0.116.733—
interaction-0.6437χ²(1) = 0.096.757—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

Improvement over inaction (BETTER_THAN_NOTHING)

The rationale compares Plan A with doing nothing and treats a partial benefit as worthwhile.

Astra coded Terra 11/40 Gain and 24/39 Loss; Sonnet coded Terra 1/40 and 6/39. This is the clearest coder-dependent comparison.

Astra primary Improvement over inaction FRAME-05 exploratory two-by-two table
Original Astra-primary BETTER_THAN_NOTHING decisions; each cell states yes out of valid Astra codes.
Astra primary Improvement over inaction FRAME-05 exploratory bar chart
Original Astra-primary BETTER_THAN_NOTHING shares by Frame and participant Model; exploratory descriptive display.
Astra primary Improvement over inaction FRAME-05 exploratory line chart
The same original Astra-primary BETTER_THAN_NOTHING shares joined across frames.
Statistical comparisons: BETTER_THAN_NOTHING — Astra primary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

BETTER_THAN_NOTHING — Astra primary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss-0.0263Exact conditional1.000—
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss-0.3404Exact conditional.008**
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain-0.2500Exact conditional.008**
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss-0.5641Exact conditional< .001***

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

BETTER_THAN_NOTHING — Astra primary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame-1.0171χ²(1) = 3.091.079—
main effect of Model-2.8565χ²(1) = 36.582< .001***
interaction0.9227χ²(1) = 0.582.446—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

See the separate Sonnet coding for Improvement over inaction

Sonnet is the independently retained secondary stream. It has one absent earlier reply and one captured-invalid later reply, so its valid cell denominators are 39 Gain Sol, 40 Gain Terra, 38 Loss Sol and 39 Loss Terra.

Sonnet secondary Improvement over inaction FRAME-05 exploratory two-by-two table
Original Sonnet-secondary BETTER_THAN_NOTHING decisions, without substitution from Astra.
Sonnet secondary Improvement over inaction FRAME-05 exploratory bar chart
Separate Sonnet-secondary shares; coder differences remain visible.
Sonnet secondary Improvement over inaction FRAME-05 exploratory line chart
The same separate Sonnet-secondary shares joined across frames.
Statistical comparisons: BETTER_THAN_NOTHING — Sonnet secondary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

BETTER_THAN_NOTHING — Sonnet secondary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss-0.0007Exact conditional1.000—
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss-0.1288Exact conditional.230—
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain0.0006Exact conditional1.000—
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss-0.1275Exact conditional.322—

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

BETTER_THAN_NOTHING — Sonnet secondary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame-0.8448χ²(1) = 1.260.262—
main effect of Model-0.7928χ²(1) = 1.111.292—
interaction1.6075χ²(1) = 1.142.285—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

Framing recognised (FRAME_EQUIVALENCE)

The rationale explicitly recognises that the gain and loss descriptions state the same numerical outcome.

Neither coder applied this code to any valid rationale. Constant-zero labels do not show that the framings are equivalent in their effects.

Astra primary Framing recognised FRAME-05 exploratory two-by-two table
Original Astra-primary FRAME_EQUIVALENCE decisions; each cell states yes out of valid Astra codes.
Astra primary Framing recognised FRAME-05 exploratory bar chart
Original Astra-primary FRAME_EQUIVALENCE shares by Frame and participant Model; exploratory descriptive display.
Astra primary Framing recognised FRAME-05 exploratory line chart
The same original Astra-primary FRAME_EQUIVALENCE shares joined across frames.
Statistical comparisons: FRAME_EQUIVALENCE — Astra primary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

FRAME_EQUIVALENCE — Astra primary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss0.0000Exact conditional1.000—
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss0.0000Exact conditional1.000—
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain0.0000Exact conditional1.000—
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss0.0000Exact conditional1.000—

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

FRAME_EQUIVALENCE — Astra primary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame-0.0249χ²(1) = 0.000.986—
main effect of Model0.0000χ²(1) = 0.0001.000—
interaction0.0000χ²(1) = 0.0001.000—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

See the separate Sonnet coding for Framing recognised

Sonnet is the independently retained secondary stream. It has one absent earlier reply and one captured-invalid later reply, so its valid cell denominators are 39 Gain Sol, 40 Gain Terra, 38 Loss Sol and 39 Loss Terra.

Sonnet secondary Framing recognised FRAME-05 exploratory two-by-two table
Original Sonnet-secondary FRAME_EQUIVALENCE decisions, without substitution from Astra.
Sonnet secondary Framing recognised FRAME-05 exploratory bar chart
Separate Sonnet-secondary shares; coder differences remain visible.
Sonnet secondary Framing recognised FRAME-05 exploratory line chart
The same separate Sonnet-secondary shares joined across frames.
Statistical comparisons: FRAME_EQUIVALENCE — Sonnet secondary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

FRAME_EQUIVALENCE — Sonnet secondary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss0.0000Exact conditional1.000—
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss0.0000Exact conditional1.000—
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain0.0000Exact conditional1.000—
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss0.0000Exact conditional1.000—

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

FRAME_EQUIVALENCE — Sonnet secondary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame-0.0250χ²(1) = 0.000.986—
main effect of Model0.0250χ²(1) = 0.000.986—
interaction-0.0001χ²(1) = 0.0001.000—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

Unopposed benefit (NO_STATED_DOWNSIDE)

The rationale presents the protected group as a benefit without stating a downside.

Astra coded Sol 18/40 Gain and 0/39 Loss, and Terra 30/40 Gain and 8/39 Loss. This code carried no original directional hypothesis.

Astra primary Unopposed benefit FRAME-05 exploratory two-by-two table
Original Astra-primary NO_STATED_DOWNSIDE decisions; each cell states yes out of valid Astra codes.
Astra primary Unopposed benefit FRAME-05 exploratory bar chart
Original Astra-primary NO_STATED_DOWNSIDE shares by Frame and participant Model; exploratory descriptive display.
Astra primary Unopposed benefit FRAME-05 exploratory line chart
The same original Astra-primary NO_STATED_DOWNSIDE shares joined across frames.
Statistical comparisons: NO_STATED_DOWNSIDE — Astra primary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

NO_STATED_DOWNSIDE — Astra primary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss0.4500Exact conditional< .001***
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss0.5449Exact conditional< .001***
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain-0.3000Exact conditional.011*
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss-0.2051Exact conditional.011*

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

NO_STATED_DOWNSIDE — Astra primary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame3.3154χ²(1) = 50.776< .001***
main effect of Model-2.1773χ²(1) = 16.516< .001***
interaction1.7922χ²(1) = 1.942.163—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

See the separate Sonnet coding for Unopposed benefit

Sonnet is the independently retained secondary stream. It has one absent earlier reply and one captured-invalid later reply, so its valid cell denominators are 39 Gain Sol, 40 Gain Terra, 38 Loss Sol and 39 Loss Terra.

Sonnet secondary Unopposed benefit FRAME-05 exploratory two-by-two table
Original Sonnet-secondary NO_STATED_DOWNSIDE decisions, without substitution from Astra.
Sonnet secondary Unopposed benefit FRAME-05 exploratory bar chart
Separate Sonnet-secondary shares; coder differences remain visible.
Sonnet secondary Unopposed benefit FRAME-05 exploratory line chart
The same separate Sonnet-secondary shares joined across frames.
Statistical comparisons: NO_STATED_DOWNSIDE — Sonnet secondary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

NO_STATED_DOWNSIDE — Sonnet secondary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss0.4359Exact conditional< .001***
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss0.5449Exact conditional< .001***
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain-0.3141Exact conditional.011*
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss-0.2051Exact conditional.011*

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

NO_STATED_DOWNSIDE — Sonnet secondary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame3.2496χ²(1) = 48.259< .001***
main effect of Model-2.1839χ²(1) = 16.645< .001***
interaction1.7003χ²(1) = 1.718.190—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

Plan must be shown best (OPTIMALITY_BURDEN)

The rationale requires Plan A to be shown better than alternatives before recommending it.

For Terra, Astra coded 0/40 Gain and 4/39 Loss (adjusted p = .112); Sonnet coded 1/40 and 11/39 (adjusted p = .00405).

Astra primary Plan must be shown best FRAME-05 exploratory two-by-two table
Original Astra-primary OPTIMALITY_BURDEN decisions; each cell states yes out of valid Astra codes.
Astra primary Plan must be shown best FRAME-05 exploratory bar chart
Original Astra-primary OPTIMALITY_BURDEN shares by Frame and participant Model; exploratory descriptive display.
Astra primary Plan must be shown best FRAME-05 exploratory line chart
The same original Astra-primary OPTIMALITY_BURDEN shares joined across frames.
Statistical comparisons: OPTIMALITY_BURDEN — Astra primary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

OPTIMALITY_BURDEN — Astra primary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss-0.5917Exact conditional< .001***
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss-0.1026Exact conditional.112—
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain0.0750Exact conditional.244—
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss0.5641Exact conditional< .001***

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

OPTIMALITY_BURDEN — Astra primary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame-2.7248χ²(1) = 19.435< .001***
main effect of Model2.4068χ²(1) = 13.732< .001***
interaction-0.7547χ²(1) = 0.183.669—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

See the separate Sonnet coding for Plan must be shown best

Sonnet is the independently retained secondary stream. It has one absent earlier reply and one captured-invalid later reply, so its valid cell denominators are 39 Gain Sol, 40 Gain Terra, 38 Loss Sol and 39 Loss Terra.

Sonnet secondary Plan must be shown best FRAME-05 exploratory two-by-two table
Original Sonnet-secondary OPTIMALITY_BURDEN decisions, without substitution from Astra.
Sonnet secondary Plan must be shown best FRAME-05 exploratory bar chart
Separate Sonnet-secondary shares; coder differences remain visible.
Sonnet secondary Plan must be shown best FRAME-05 exploratory line chart
The same separate Sonnet-secondary shares joined across frames.
Statistical comparisons: OPTIMALITY_BURDEN — Sonnet secondary

Post-response exploratory tests. Pairwise p values are Holm-adjusted across the four comparisons for this outcome and coder stream. *p < .05; **p < .01; ***p < .001. A dash means no significance mark. Exact stored values are retained in the cell descriptions.

OPTIMALITY_BURDEN — Sonnet secondary — four simple effects
ComparisonObserved differenceTest statisticHolm-adjusted pStars
GPT-5.6 Sol in Gain against GPT-5.6 Sol in Loss-0.4777Exact conditional< .001***
GPT-5.6 Terra in Gain against GPT-5.6 Terra in Loss-0.2571Exact conditional.004**
GPT-5.6 Sol in Gain against GPT-5.6 Terra in Gain0.1288Exact conditional.058—
GPT-5.6 Sol in Loss against GPT-5.6 Terra in Loss0.3495Exact conditional.006**

Two-sided exact conditional tests stratified by sample. Differences are proportions, first condition minus second. Error-bar overlap is not the test used for these stars.

OPTIMALITY_BURDEN — Sonnet secondary — omnibus model terms (unadjusted)
TermCoefficientTest statisticUnadjusted pStars
main effect of Frame-2.2529χ²(1) = 26.918< .001***
main effect of Model1.5246χ²(1) = 10.598.001**
interaction0.2106χ²(1) = 0.041.839—

Firth penalized logistic likelihood-ratio tests, using the original declared saturation fallback; coefficients are on the log-odds scale. Omnibus values are unadjusted and distinct from the Holm-adjusted pairwise comparisons. There is no study-wide multiplicity adjustment.

Note. Coded percentages always use that coder’s valid original decisions, never the 160-session planned denominator. In Sol Gain and Sol Loss the Sonnet denominator is one smaller than Astra’s; in Terra Loss it is the same. FRAME_EQUIVALENCE was constantly zero in both streams, so its Cohen κ is undefined. Across the five codes κ ranges from .319 for BETTER_THAN_NOTHING to .986 for NO_STATED_DOWNSIDE among calculable codes. Agreement is neither coder accuracy nor participant-model equivalence.

Manipulation checks

Each of the 158 eligible sessions answered both retained count checks correctly: 300 protected and 600 unprotected. The two permanently unavailable sessions remain missing for both checks. Correct recall of the stated counts does not demonstrate immunity to framing.

Registered manipulation checks, correct out of 40 scheduled sessions per condition
CheckGain, SolGain, TerraLoss, SolLoss, Terra
Protected count (300)40 of 40 passed
0 failed, 0 missing
40 of 40 passed
0 failed, 0 missing
39 of 40 passed
0 failed, 1 missing
39 of 40 passed
0 failed, 1 missing
Unprotected count (600)40 of 40 passed
0 failed, 0 missing
40 of 40 passed
0 failed, 0 missing
39 of 40 passed
0 failed, 1 missing
39 of 40 passed
0 failed, 1 missing

Each sample on its own

Both original sample blocks preserve the 20-per-condition schedule. Sample 1 has 80 eligible sessions; Sample 2 has 78 because each of its Loss cells contains one permanent gap. The sample table keeps observed ratings separate instead of replacing missing answers.

Recommendation likelihood by original sample block
SampleGain, SolGain, TerraLoss, SolLoss, Terra
Sample 1M = 5.75
SD = 0.55, n = 20
M = 6.30
SD = 0.57, n = 20
M = 3.65
SD = 0.81, n = 20
M = 5.00
SD = 0.79, n = 20
Sample 2M = 5.75
SD = 0.44, n = 20
M = 6.25
SD = 0.64, n = 20
M = 3.26
SD = 0.56, n = 19
M = 4.68
SD = 0.82, n = 19

What the exploratory analysis found

The original ten directional questions are shown with observed comparisons below. None receives a new FRAME-05 confirmatory or falsified verdict: recovery and analysis followed disclosure of earlier participant responses. The frozen methods and null rows were retained, and the final available data were analysed under the reviewed exploratory amendment.

Original directional questions and observed exploratory comparisons; none receives a new confirmatory verdict
QuestionOriginal directionObserved comparisonStatus
H1aGPT-5.6 Terra rates recommendation likelihood higher under Gain than Loss.Terra recommendation Gain 6.28, Loss 4.85; adjusted p = 9.95 × 10−13.Exploratory only; no FRAME-05 confirm/falsify verdict
H1bGPT-5.6 Sol rates recommendation likelihood higher under Gain than Loss.Sol recommendation Gain 5.75, Loss 3.46; adjusted p = 2.12 × 10−19.Exploratory only; no FRAME-05 confirm/falsify verdict
H2aGPT-5.6 Terra rates confidence higher under Gain than Loss.Terra confidence Gain 5.60, Loss 3.87; adjusted p = 3.44 × 10−15.Exploratory only; no FRAME-05 confirm/falsify verdict
H2bGPT-5.6 Sol rates confidence higher under Gain than Loss.Sol confidence Gain 4.88, Loss 4.46; adjusted p = .0160; least-favourable missing-rating scenario p = .0757.Exploratory only; no FRAME-05 confirm/falsify verdict
H3aGPT-5.6 Sol rates confidence higher than GPT-5.6 Terra under Gain.Sol confidence 4.88, Terra 5.60 under Gain; direction opposite the question, adjusted p = .0000759.Exploratory only; no FRAME-05 confirm/falsify verdict
H3bGPT-5.6 Sol rates confidence higher than GPT-5.6 Terra under Loss.Sol confidence 4.46, Terra 3.87 under Loss; adjusted p = .00220; least-favourable missing-rating scenario p = .0601.Exploratory only; no FRAME-05 confirm/falsify verdict
H4aGPT-5.6 Terra rationales show BETTER_THAN_NOTHING more often under Gain.For Terra, BETTER_THAN_NOTHING: Astra 11/40 Gain and 24/39 Loss (adjusted p = .00828); Sonnet 1/40 and 6/39 (p = .230).Exploratory only; no FRAME-05 confirm/falsify verdict
H4bGPT-5.6 Terra rationales show SHORTFALL more often under Loss.For Terra, SHORTFALL: Astra 0/40 Gain and 24/39 Loss (adjusted p = 8.72 × 10−10); Sonnet 0/40 and 18/39 (p = 8.03 × 10−7).Exploratory only; no FRAME-05 confirm/falsify verdict
H4cGPT-5.6 Terra rationales show OPTIMALITY_BURDEN more often under Loss.For Terra, OPTIMALITY_BURDEN: Astra 0/40 Gain and 4/39 Loss (adjusted p = .112); Sonnet 1/40 and 11/39 (p = .00405).Exploratory only; no FRAME-05 confirm/falsify verdict
H4dGPT-5.6 Sol rationales show FRAME_EQUIVALENCE more often under Loss.FRAME_EQUIVALENCE was zero in both original coder streams in all conditions; the observed difference is zero, with no evidence for the predicted direction.Exploratory only; no FRAME-05 confirm/falsify verdict

The original linear model with sample intercepts returned the following terms for the recommendation scale. These are unadjusted exploratory model terms; the exact simple-effect comparisons are adjusted within their own outcome.

Original linear-model terms, reported here as post-response exploratory (Wald t, 153 residual degrees of freedom)
TermbSEt(153)Unadjusted p
main effect of Frame1.8610.10517.6887.94e-39
main effect of Model-0.9550.105-9.0765.27e-16
interaction0.8600.2104.0857.07e-05

The observed recommendation frame-by-model interaction is 0.860 rating points (SE 0.210, t(153) = 4.085, unadjusted p = .0000707). The slope difference is specific to these native-route model sessions. Coder-dependent rationale contrasts and a confidence comparison that changes under least-favourable missingness limit broader conclusions.

How the run went

The original design scheduled 160 main sessions, 40 in each Frame × Model condition. Both coders first passed the unchanged 40-judgment calibration with 39 correct (97.5%) and a lowest per-code accuracy of 87.5%, above the 90% overall and 75% per-code thresholds. A separate four-session feasibility, 40-session measurement pilot and four-session execution pilot passed their recorded gates. Their observations are excluded from the main-study figures.

The original main launch completed 120 six-turn sessions and then stopped after one claim failed before a model request. The first reviewed continuation claimed another session but its delivery remains unresolved. Neither claim was resent. A later, separately reviewed continuation completed the 38 sessions that had never been attempted. The final ledger holds 158 eligible sessions and two permanent gaps: one Sol Loss and one Terra Loss. The earlier stopped run and its private partial export remain preserved; no later success reclassifies either failed claim.

For those 158 rationales, Astra produced 158 valid primary codes and Sonnet produced 156 valid secondary codes. One Sonnet action remains absent after uncertain delivery; one other reply was captured but invalid, with null codes. Neither was retried or replaced. The 156 rationales with both valid sets yielded 729 agreements and 51 disagreements across 780 comparable code decisions. Five role slots are absent in the 160-session plan (four attached to the two missing participants and one uncertain Sonnet action); a sixth is the captured-invalid Sonnet reply.

The native subscription routes requested low reasoning effort, a 1,024-token output target and a 240-second logical-turn limit, with no application-level retries. The original prompts and scientific materials were frozen. Matching requested settings does not establish identical hidden provider context, internal retries, served weights or enforcement across providers. The operational deviations and all three participant roots, three coding roots, preservation audits, exploratory close and analysis are retained in the private record.

Every original rating and code family has at least 36 analysed observations in each cell. This permits the reviewed exploratory computation, not original complete-study admission. The private export contains the 160-row schedule, 158 screened designated rationales with two explicit nulls, both original 160-slot coder streams, the unchanged codebook, result and provenance. This page reports the verified observations and exploratory analysis.

Against FRAME-04

Both studies used the same 1–7 recommendation and confidence scales, scenario, five-code vocabulary, manipulation checks and two sample blocks. FRAME-05 changed the participant models from Claude Opus 5 and Claude Sonnet 5.5 to GPT-5.6 Sol and Terra and changed the route to native subscription CLIs. FRAME-04 used one Claude Sonnet 5 coder; FRAME-05 kept independent Astra-primary and Sonnet-secondary original streams. The results were collected on different days and under different provider environments.

Framing effects across all five experiments

The effect size is the gain-minus-loss difference in recommendation. Larger positive values indicate a greater observed shift. The two response formats are shown separately.

FRAME-01–02: binary recommendation; positive differences favour gain framing.
StudyParticipant modelGain: yes / nLoss: yes / nGain − loss (pp)
FRAME-01Claude Opus 550 / 5050 / 50+0.00
FRAME-01Claude Sonnet 549 / 5017 / 50+64.00
FRAME-02Claude Opus 579 / 8076 / 77+0.05
FRAME-02Claude Sonnet 579 / 7911 / 80+86.25
FRAME-03–05: recommendation on the same 1–7 response scale. FRAME-05 is exploratory.
StudyParticipant modelGain mean (n)Loss mean (n)Gain − loss (points)
FRAME-03Claude Opus 55.81 (37)5.10 (39)+0.71
FRAME-03Claude Sonnet 55.08 (40)3.38 (40)+1.70
FRAME-04Claude Opus 55.95 (39)5.28 (40)+0.67
FRAME-04Claude Sonnet 5.55.00 (40)3.90 (40)+1.10
FRAME-05GPT-5.6 Sol5.75 (40)3.46 (39)+2.29
FRAME-05GPT-5.6 Terra6.28 (40)4.85 (39)+1.43

Within the binary studies, Sonnet changed more than Opus. Within the rating studies, the largest observed shift was Sol in FRAME-05 (+2.29 points), followed by Sonnet 5 in FRAME-03 (+1.70), Terra in FRAME-05 (+1.43), Sonnet 5.5 in FRAME-04 (+1.10), and Opus 5 in FRAME-03 and FRAME-04 (+0.71 and +0.67). These describe the collected replies; they do not establish a controlled ranking across providers or studies.

Models, dates, native routes and eligibility rules differ; FRAME-05 is exploratory. The binary studies also put Opus close to the ceiling. Low within-model variation makes standardized effects large, so raw differences are the main comparison.

The FRAME-01 and FRAME-02 primary outcome was a binary Plan A recommendation, so their gain-minus-loss effects are percentage-point risk differences. FRAME-03 through FRAME-05 used a 1–7 recommendation rating, so those effects are mean-point differences. They belong in separate units; no numerical cross-scale ranking is warranted. FRAME-05 is the only post-response exploratory/sensitivity entry in this comparison and has two permanent missing sessions.

Recommendation scale distribution

Observed recommendation replies at each point of the 1–7 scale
Condition1234567n
Gain frame, GPT-5.6 Sol00001128140
Gain frame, GPT-5.6 Terra00003231440
Loss frame, GPT-5.6 Sol01231140039
Loss frame, GPT-5.6 Terra00210198039

In FRAME-05 both models have higher observed recommendation ratings under Gain: Sol 5.75 versus 3.46; Terra 6.28 versus 4.85. This does not establish that either OpenAI model is a direct substitute for a Claude model, nor that the difference between FRAME-04 and FRAME-05 is caused by provider identity. No cross-study randomisation or model-equivalence test was run.

What this says and does not say

The 158 observed sessions show a framing association under this particular wording, six-turn protocol, models and native routes. The two missing Loss sessions remain explicit, and worst-endpoint rating sensitivity retains the recommendation direction. That bound does not address the timing of outcome disclosure, unknown provider behaviour, correlated model responses, coder disagreements or generalisation to people or other scenarios. The original 160-session confirmatory study remains incomplete.