CONNECTED EXPERIMENTAL RESULTS · AUGUST 9, 2026

P1, P2, and P4 benchmark results.

These connected trials compare a direct verified answer, a taxonomy-qualified multi-model plan, and selective repair. They are publishable experimental evidence, not production certification or a universal model-ranking claim.

HEADLINE RESULTS

What the connected trials showed

Cost and latency include the complete tested chain. Quality was independently verified under the experiment contract.

P1

Direct verified control

Quality 0.92 · $0.050818 · 23.43 seconds

The baseline for the P2 comparison.

P2

Taxonomy-qualified orchestration

Quality 0.93 · $0.0121774 · 11.21 seconds

76.04% lower cost and 52.13% lower latency than P1 in this trial.

P4

Selective repair

Quality gate passed in 2 of 3 replication trials.

33.83% lower mean cost and 39.27% lower mean latency than full redo.

CONNECTED COMPARISON

Full-chain results

PolicyControlQualityCostLatencyFinding
P1 direct0.92$0.05081823,429.98 msVerified baseline
P2 orchestrationP1 direct0.93$0.012177411,214.96 msVerified; beat P1 on quality, cost, and latency
P4 selective repairP1 full redo0.9167 mean$0.02847467 mean12,727.74 ms mean2/3 quality passes; use only behind verification and redo fallback
P1 full redo0.9567 mean$0.04303167 mean20,957.02 ms meanP4 replication control

What is safe to conclude

Limits and disclosure

Method

P2 used a flagship framework with taxonomy-qualified lower-cost workers and an independent verifier. P4 repaired only failed sections, verified the repair, and retained a full-redo fallback. All displayed economics are measured per complete tested chain rather than per isolated provider call.