Research dashboard
Experiment results
Compare controlled-language prompt strategies with paired effects and confidence intervals.
At a glance
Naming versus bare baseline
Each line joins a model's bare-prompt mean compliance score to its mean when ASD-STE100 is named without rules. Shaded bands show 95% confidence intervals; chevrons mark bands that extend past the 0 to 100 scale. The table below lists the exact values.
- Bare baseline
- Named without rules
- 95% confidence interval
About the experiment
The four variants are bare (prompt alone), rules (rules spelled out), named (technique named), and named_rules (name and rules together). Compliance scores run from 0 to 100; higher scores mean greater compliance. Effects are paired changes in score points.
- Named without rules
- Mean compliance score when ASD-STE100 is named without spelling out its rules.
- Bare baseline
- Mean score for the bare prompt, before naming or rules are added.
- Rule effect
- Mean paired change from spelling out the rules, averaged with and without naming.
- Naming effect
- Mean paired change from naming the technique, averaged with and without rules.
- Interaction
- The extra paired change when naming and rules appear together, beyond their separate effects.
- Paired observations
- Complete four-arm research cells used to calculate every value in the row.
A negative interaction may reflect overlap between naming the technique and spelling out its rules, but the interaction does not identify its cause.
Measured outcomes
Leaderboard
| Rank | Model | Named without rules, mean ± 95% CI | Bare baseline | Rule effect | Naming effect | Interaction | Paired sessions | Run elapsed |
|---|---|---|---|---|---|---|---|---|
| 1 | deepseek/deepseek-v4-flash-0731 | 98.9 ± 0.3 | 95.5 ± 3.7 | +2.8 ± 1.9 | +1.7 ± 1.8 | -3.4 ± 3.7 | 6 | 1h 18m |
| 2 | openai/gpt-5.6-luna | 98.0 ± 1.2 | 92.0 ± 6.0 | +5.0 ± 3.4 | +3.0 ± 2.6 | -6.0 ± 5.3 | 6 | 19m 48s |
| 3 | google/gemini-3.8-flash | 97.2 ± 1.7 | 95.7 ± 1.9 | +3.6 ± 1.4 | +0.8 ± 1.1 | -1.5 ± 2.2 | 6 | 1h 3m |
| 4 | meta-llama/llama-4-maverick | 89.3 ± 5.6 | 86.7 ± 4.5 | +9.5 ± 4.4 | +1.8 ± 3.6 | -1.8 ± 3.9 | 6 | 28m 9s |
| 5 | anthropic/claude-haiku-4.5 | 89.0 ± 12.0 | 73.3 ± 9.7 | +18.2 ± 8.0 | +7.7 ± 6.8 | -16.1 ± 14.7 | 6 | 39m 23s |
Understanding uncertainty
A 95% confidence interval (95% CI) uses the Student t critical value for the number of contributing sessions. Wider intervals indicate greater uncertainty. It is neither the range of individual scores nor a 95% probability that the fixed population effect lies inside this observed interval. Interpret estimates and uncertainty together, rather than reducing them to statistically significant or not significant.
In the leaderboard table, each value is written as a mean ± the half-width of its 95% CI, so the interval runs from the mean minus that amount to the mean plus it. ± n/a marks rows with fewer than 2 sessions, which cannot support an interval.
Only complete four-arm research cells matched within run, model, session, depth, protocol version, and scoring version are included; previews and incomplete runs are excluded. Effects are calculated within each matched cell, then repeated depths within each run and session are averaged before intervals are calculated. Sessions are treated as independent; this is not a hierarchical analysis across runs or other levels.