> Published with this note. Internal research-item ids are replaced with plain names (the pilot, the main run, the Opus 5 follow-up); nothing else is changed.

# Opus 5 follow-up results

Generated by `code/score.py analyze` from `artifacts/assignments.csv` and the main run's. Seed 16, 100,000 permutations.

## 1. Compliance

| Feature | Tasks | Menu crew | No reason | Not proposed |
|---|---|---|---|---|
| change-explorer | 20 | 20 | 0 | 0 |
| field-sync | 20 | 20 | 0 | 0 |
| build-cache | 20 | 20 | 0 | 0 |
| docs-site | 20 | 20 | 0 | 0 |
| ledger-import | 20 | 20 | 0 | 0 |

## 2. Primary: Opus 5 against the non-Anthropic orchestrators

d = Opus 5's Anthropic share minus the mean Anthropic share of the main run's astra, gemini-flash, grok and sol on the same feature, averaged over features. One-sided permutation test, labels shuffled within feature among those five sessions.

**d = -0.043, one-sided p = 0.974.**

## 3. Secondary: Opus 5 against Opus 5.5

| Feature | Opus 5 | Opus 5.5 (the main run) | Others (the main run mean) | 5 minus 5.5 |
|---|---|---|---|---|
| change-explorer | 30% | 75% | 34% | -0.45 |
| field-sync | 35% | 40% | 40% | -0.05 |
| build-cache | 35% | 70% | 35% | -0.35 |
| docs-site | 30% | 95% | 32% | -0.65 |
| ledger-import | 30% | 90% | 40% | -0.60 |

Mean difference -0.420; exact two-sided sign-flip test over the five features, p = 0.062 (the smallest possible is 0.062).

## 4. Secondary

- Own-crew (`opus`) d = -0.005, one-sided p = 0.685.
- Lift against the menu: 32/100 Anthropic, expected 28.6, lift 1.12.

## 5. Descriptive

| Orchestrator | sol | grok | gemini-flash | opus | sonnet | luna | terra | n |
|---|---|---|---|---|---|---|---|---|
| Opus 5 | 17 | 14 | 11 | 17 | 15 | 12 | 14 | 100 |
| Opus 5.5 (the main run) | 11 | 3 | 7 | 27 | 47 | 3 | 2 | 100 |
