OpenPsy

Tools for Effective AI Behavioral Research.

OpenPsy is an open research platform for designing, running and analyzing behavioral experiments, including experiments in which language models take part as participants. The goal of OpenPsy is to enable researchers to understand what is happening beneath their results and to have an audit history in order to understand their research programs. OpenPsy is presently in an internal closed alpha.

Why OpenPsy?

Agentic research can be difficult to parse, monitor, and understand. Often agents will violate rules and assumptions about research design undermining results. OpenPsy was born from this difficulty and aims to assist researchers in agent to agent and agent to human research. The software is designed to improve a researcher's ability to plan, execute, and analyze agentic experiments with or without human participants.

  • Introspection

    OpenPsy enables the effective design of experiments by providing simulation prior to execution, and saved session logs for auditability.

  • Replicability

    OpenPsy saves every version of a design and changes, so a researcher can audit and export the design for replication.

  • Control

    OpenPsy enables researchers to control context, understanding exactly what content they are serving agents at each step of an experiment.

OpenPsy Builds Experiments as a Flow

In the flow editor, researchers use "blocks" to build ideal experimental process through a node graph interface. All OpenPsy aspects read from the same common database ensuring there is no drift between a design and execution. Below we illustrate an example study in the OpenPsy framework.

Example Study

1Plan

Research question

Does a gain or loss framing in a question affect how an LLM responds? Does this response vary across models?

A Classic 2x2 Research Design

Imagine a 2x2 sample experiment, two factors with two levels. The first factor is a gain versus loss frame. The second factor is model type, Claude Opus 5 or Claude Sonnet 5.

Participants are randomized across the four conditions.

Four conditions
Claude Opus 5Claude Sonnet 5
Gain frameGain + OpusGain + Sonnet
Loss frameLoss + OpusLoss + Sonnet
Factor 1 · Framing: Gain / Loss

The two scenarios differ only in the highlighted clause.

Gain frame

A new infection is expected to affect 900 people in a region. Health officials are deciding whether to adopt Plan A. If Plan A is adopted, 300 of those people will be protected.

Loss frame

A new infection is expected to affect 900 people in a region. Health officials are deciding whether to adopt Plan A. If Plan A is adopted, 600 of those people will be left unprotected.

Factor 2 · Model type: Opus / Sonnet

Only the model taking part differs. Both levels run the same procedure through the same route with the same settings.

Opus

Claude Opus 5 claude-opus-5

Sonnet

Claude Sonnet 5 claude-sonnet-5

Route
Claude Code CLI in print mode, one call at a time.
Effort
Model default.
Limits
1,024 tokens per reply, 4 minutes per call, no retries.
Session
Four calls: the scenario, then Recommendation, Rationale and Confidence.
Procedure
The same OpenPsy Flow with its simulation paused on the second of seven steps. The toolbar says Version 4 is saved and holds a Resume button, an Exit simulation button, a speed slider set to 4x slower, the step count 2 of 7 and zoom controls reading 85 percent. Under the toolbar, a strip headed Preview condition reads Frame Gain frame and Model Claude Opus 5, followed by the line: One assignment. No participant calls are made and no study data is created. In the graph below, the Random assignment block is marked as taken, the Gain frame card is outlined and reads Applying in this preview, Claude Opus 5 reads Assigned in this preview, and Loss frame and Claude Sonnet 5 are dimmed and read Not assigned in this preview. Task instructions and the three measures follow to the right, each reading 1 item and Design only. Gain frame, Loss frame, Task instructions and the three measures each carry a Prompt badge, and Claude Opus 5 and Claude Sonnet 5 each carry a Model badge. The pane across the bottom is headed Simulation step 2 of 7 with the block name Gain frame and the line Direct agent prompt, 1 item. Beside it are two buttons, Agent view, which is selected, and Library source. On the right of that header, Dry-run route is followed by: Design only. Not admitted for execution. The body of the pane opens with Gain frame, followed by its registered wording in a monospaced face: A new infection is expected to affect 900 people in a region. Health officials are deciding whether to adopt Plan A. If Plan A is adopted, 300 of those people will be protected. Below the wording, a note in smaller, lighter type reads: Specification preview. No provider is called and no participant data exists here.
Flow designer enables full simulation prior to execution. The Flow designer keeps the whole procedure in one view: random assignment to one of the four conditions, the two independent variables with their levels, the shared task instructions, and the three outcomes in order. By enabling a simulation, the researcher can evaluate their procedure prior to piloting or full execution.
The Add a Block panel in OpenPsy. Under the title, one line reads: Choose a block already saved in your Library to place it in the Flow, or create a new one below. The panel has a Close button and a search field labeled Search blocks. Under Shared steps (functions), Participant instructions is listed at version 1. Under Independent variables (IVs), Gain frame and then Loss frame are listed at version 1. Under Dependent variables (DVs), Recommendation Binary (Yes/No), Open-Ended Rationale Explanation and Confidence Rating (7 Point Scale) are listed in that order at version 1. Each group has a short description and a Show blocks from menu, which is set to This program for the first two groups and to All programs for Dependent variables (DVs). Each block has an Add to Flow button and a Fork into this program button.
Blocks Come From a Shared Library. Every step, variable and outcome is a block saved in the researcher's Library, kept in three groups: shared steps, independent variables and dependent variables. Adding a block to an experiment places an exact saved version, so the same block can be reused across experiments without retyping it.
The OpenPsy block editor headed Editing block: Confidence Rating (7 Point Scale). One line under the heading reads: Dependent variable, pinned to version 2 of its wording. The editor has an Add a block after this one button, a Move earlier button and a Remove block button, and the Move later button is unavailable. The Block title field reads Confidence Rating (7 Point Scale), and Delivered to is set to The agent. The wording box, three lines tall, reads: How confident are you in your recommendation, from 1 (not at all confident) to 7 (completely confident)? A note below says the block is pinned to Confidence Rating (7 Point Scale), version 2, that changing the wording creates the next version, and that the saved version is kept.
Versioning of Blocks Tracks Provenance. Each block in an experiment is pinned to one exact version of its wording. Editing a block creates a new version rather than overwriting the old one, so an experiment always points at the exact words it used.
The header of the experiment Gain and loss framing in a treatment choice, with the word Experiment above the title. It has Create revision, Prepare replication and Export design buttons and a menu labelled Version shown that is set to Version 4. Below it, the Source and attribution section is closed and the What changed from version 3 section is open. The open section shows a shaded panel with a green rule on its left edge. The panel says that Confidence Rating (7 Point Scale) moved from version 1 to version 2 of its wording.
Edit History Provides Auditability. Every saved version of an experiment is kept. The version panel shows which version is on screen and what changed from the one before, so any edit can be traced back to the wording it replaced.

2Execute

The Coverage & integrity panel of the Execution Monitor. Under the heading Pilot, a line reads "This phase is recorded in simulated mode." Below it is a green progress bar labelled 23/40 terminal. A section headed Cell coverage (counts only) holds a card titled Session Coverage by Frame and Model. The table has one diagonal corner header: Model is in the upper-right triangle for the columns, and Frame is in the lower-left triangle for the rows. Claude Opus 5 has a purple swatch and Claude Sonnet 5 has a blue swatch. It has four cells, each with a green progress strip: Gain frame with Claude Opus 5 shows 6 of 10 sessions complete, Gain frame with Claude Sonnet 5 shows 6 of 10, Loss frame with Claude Opus 5 shows 6 of 10, and Loss frame with Claude Sonnet 5 shows 5 of 10. There is no Total column and no totals row. A note under the table reads "Each cell counts the scheduled sessions of one condition that have finished. The table shows how many sessions are done, never how they turned out." Below the card, the terminal taxonomy shows valid-observation 23. Three tiles at the bottom show ledger freshness (evidence arriving now), envelope closure (23/23 terminal units sealed by envelope) and scoring closure (23/40 scheduled units scored).
Blinded and Secure Execution. The Execution Monitor shows how many sessions in each condition have finished and keeps results hidden until collection closes.

3Analyze

Hypothetical results in a two-by-two table and grouped bar chart. Frame labels the rows and Model labels the columns, separated by a diagonal corner. Gain frame: Claude Opus 5, 72% (36 of 50); Claude Sonnet 5, 74% (37 of 50); row total, 73% (73 of 100). Loss frame: both models, 46% (23 of 50 each); row total, 46% (46 of 100). Column totals: Opus, 59% (59 of 100); Sonnet, 60% (60 of 100). The grand-total corner is blank. The table also retains mean confidence ratings. The four bars carry Wilson 95% confidence intervals: 58.3 to 82.5%, 60.4 to 84.1%, 33.0 to 59.6% and 33.0 to 59.6%. Coloured significance brackets mark the gain-versus-loss comparison within each model with one star, and a dark bracket marks the main effect of Frame with three stars. The key defines *p less than .05, **p less than .01 and ***p less than .001.
Standardized Results Reporting. OpenPsy uses standardized formats including APA style tables and charts.

OpenPsy Roadmap

OpenPsy is in internal closed alpha. It has been tested with invented data only, and no real participant has taken part.

In Progress (Expected Live Q4 2026)

  • A researcher can create an experiment, build it as a graph from Library blocks, edit each block's wording, and save every change as a new version.
  • Any two saved versions can be compared, and the differences are written out as sentences.
  • Library wordings are stored as exact versions, and publishing a revision leaves existing experiments on the version they used.
  • A scripted test participant can take an authored study from consent to completion on a phone-sized screen.
  • Automated evals pre and post execution

Near Future

  • Export and replication of experiments by third parties
  • Automated Measure and Execution Pilot Workflows
  • Automated Power Analysis
  • Co-Pilot Assistant in application.
  • Hybrid Human/Agent Experiments

Source Code

The source code, the specification and the product documentation are being prepared for publication on GitHub.

GitHub repository opening soon

Scroll to read. Pinch to zoom.