← Back to experiment

Research dashboard

Experiment results

Compare controlled-language prompt strategies with paired effects and confidence intervals.

At a glance

Naming versus bare baseline

Each line joins a model's bare-prompt mean compliance score to its mean when ASD-STE100 is named without rules. Shaded bands show 95% confidence intervals; chevrons mark bands that extend past the 0 to 100 scale. The table below lists the exact values.

02550751001. deepseek/deepseek-v4-flash-0731: bare 95.5, named 98.9 (+3.4)1. deepseek/deepseek-v4-flash-07312. openai/gpt-5.6-luna: bare 92.0, named 98.0 (+6.0)2. openai/gpt-5.6-luna3. google/gemini-3.8-flash: bare 95.7, named 97.2 (+1.5)3. google/gemini-3.8-flash4. meta-llama/llama-4-maverick: bare 86.7, named 89.3 (+2.7)4. meta-llama/llama-4-maverick5. anthropic/claude-haiku-4.5: bare 73.3, named 89.0 (+15.7)5. anthropic/claude-haiku-4.5

About the experiment

The four variants are bare (prompt alone), rules (rules spelled out), named (technique named), and named_rules (name and rules together). Compliance scores run from 0 to 100; higher scores mean greater compliance. Effects are paired changes in score points.

Named without rules
Mean compliance score when ASD-STE100 is named without spelling out its rules.
Bare baseline
Mean score for the bare prompt, before naming or rules are added.
Rule effect
Mean paired change from spelling out the rules, averaged with and without naming.
Naming effect
Mean paired change from naming the technique, averaged with and without rules.
Interaction
The extra paired change when naming and rules appear together, beyond their separate effects.
Paired observations
Complete four-arm research cells used to calculate every value in the row.

A negative interaction may reflect overlap between naming the technique and spelling out its rules, but the interaction does not identify its cause.

Measured outcomes

Leaderboard

Mean compliance scores and paired score-point effects, each ± its two-sided 95% confidence interval half-width (see Understanding uncertainty).
RankModelNamed without rules, mean ± 95% CIBare baselineRule effectNaming effectInteractionPaired sessionsRun elapsed
1deepseek/deepseek-v4-flash-073198.9 ± 0.395.5 ± 3.7+2.8 ± 1.9+1.7 ± 1.8-3.4 ± 3.761h 18m
2openai/gpt-5.6-luna98.0 ± 1.292.0 ± 6.0+5.0 ± 3.4+3.0 ± 2.6-6.0 ± 5.3619m 48s
3google/gemini-3.8-flash97.2 ± 1.795.7 ± 1.9+3.6 ± 1.4+0.8 ± 1.1-1.5 ± 2.261h 3m
4meta-llama/llama-4-maverick89.3 ± 5.686.7 ± 4.5+9.5 ± 4.4+1.8 ± 3.6-1.8 ± 3.9628m 9s
5anthropic/claude-haiku-4.589.0 ± 12.073.3 ± 9.7+18.2 ± 8.0+7.7 ± 6.8-16.1 ± 14.7639m 23s

Understanding uncertainty

A 95% confidence interval (95% CI) uses the Student t critical value for the number of contributing sessions. Wider intervals indicate greater uncertainty. It is neither the range of individual scores nor a 95% probability that the fixed population effect lies inside this observed interval. Interpret estimates and uncertainty together, rather than reducing them to statistically significant or not significant.

In the leaderboard table, each value is written as a mean ± the half-width of its 95% CI, so the interval runs from the mean minus that amount to the mean plus it. ± n/a marks rows with fewer than 2 sessions, which cannot support an interval.

Only complete four-arm research cells matched within run, model, session, depth, protocol version, and scoring version are included; previews and incomplete runs are excluded. Effects are calculated within each matched cell, then repeated depths within each run and session are averaged before intervals are calculated. Sessions are treated as independent; this is not a hierarchical analysis across runs or other levels.