Benchmark results

Radium and Claude matched-tier results

Answer accuracy, throughput, latency,
and time to first text.

Model results at a glance

Accuracy
Throughput
Latency
TTFT
Radium Hal
96.25%
77/80 correct
499.5
output tok/s
2.31k ms
p50 end-to-end
1.95k ms
p50 time to first text
Claude Opus
93.75%
75/80 correct
107.69
output tok/s
1.92k ms
p50 end-to-end
1.77k ms
p50 time to first text
Radium Clarke
85%
68/80 correct
85.58
output tok/s
690.13 ms
p50 end-to-end
471.05 ms
p50 time to first text
Claude Sonnet
100%
80/80 correct
99.96
output tok/s
2.82k ms
p50 end-to-end
2.36k ms
p50 time to first text
Radium Tycho
91.25%
73/80 correct
376.1
output tok/s
1.49k ms
p50 end-to-end
1.23k ms
p50 time to first text
Claude Haiku
100%
80/80 correct
408.34
output tok/s
2.05k ms
p50 end-to-end
2.03k ms
p50 time to first text

Accuracy/tokens are from the 80-task accuracy suite; throughput/latency/TTFT are from the separate 40-task performance suite and aren't directly comparable. Scoring requires an exact JSON match; some token counts are estimated.

Six confidence intervals

Radium Hal
96.25%
Claude Opus
93.75%
Radium Clarke
85%
Claude Sonnet
100%
Radium Tycho
91.25%
Claude Haiku
100%
Pair comparisons

Matched-tier comparisons

Maximum Capability

Hal × Opus

Accuracy: Radium 96.25%; Claude 93.75%.
Difference CI90: -3.65 to 8.94 points.

Aggregate output-token throughput ratio:
4.64×. Radium p50 latency was 20.18% higher.

Balanced Performance

Clarke × Sonnet

Accuracy: Radium 85%; Claude 100%.
Difference CI90: -22.7 to -8.67 points.

Aggregate output-token throughput ratio:
0.86×. Radium p50 latency was 75.54% lower.

High-Efficiency Scale

Tycho × Haiku

Accuracy: Radium 91.25%; Claude 100%.
Difference CI90: -15.39 to -3.63 points.

Aggregate output-token throughput ratio:
0.92×. Radium p50 latency was 27.31% lower.

Confidence, parsing, tails, and tokens

Accuracy CI95
Answer parsed
Performance correct
Latency p95
Tokens/correct
Radium Hal
89.55–98.72%
97.5%
97.5% raw JSON
100%
4.02k ms
225.13 in
356.93 out · 582.05 total
Claude Opus
86.19–97.3%
93.75%
93.75% raw JSON
97.5%
2.78k ms
186.15 in
68.79 out · 254.95 total
Radium Clarke
75.59–91.21%
95%
100% raw JSON
100%
2.27k ms
162.37 in
22.97 out · 185.34 total
Claude Sonnet
95.42–100%
100%
87.5% raw JSON
100%
3.74k ms
181.5 in
80.15 out · 261.65 total
Radium Tycho
83.02–95.7%
100%
100% raw JSON
95%
3.05k ms
247.34 in
192.61 out · 439.95 total
Claude Haiku
95.42–100%
100%
0% raw JSON
100%
2.91k ms
161.55 in
260.35 out · 421.9 total

Answer results by task category

Routing & triage
Policy application
Quantitative operations
Configuration extraction
Radium Hal
100%
20/20
85%
17/20
100%
20/20
100%
20/20
Claude Opus
100%
20/20
100%
20/20
75%
15/20
100%
20/20
Radium Clarke
90%
18/20
50%
10/20
100%
20/20
100%
20/20
Claude Sonnet
100%
20/20
100%
20/20
100%
20/20
100%
20/20
Radium Tycho
100%
20/20
65%
13/20
100%
20/20
100%
20/20
Claude Haiku
100%
20/20
100%
20/20
100%
20/20
100%
20/20
Test configuration

Scoring and execution settings

01

Accuracy criteria

80 fixed-answer tasks/model. Scoring accepts a raw JSON object or a single Markdown-fenced JSON object, then requires exact fields, primitive types, and values. Equivalence requires both models to reach 95% and the 90% accuracy-difference interval to fall within ±5 points.

02

Performance criteria

Radium must have no lower correct rate, at least 20% lower p50 latency, and at least 1.25× aggregate output-token throughput.

03

Request modes

Radium Hal: disabled; Claude Opus: adaptive; Radium  Clarke: disabled; Claude Sonnet: adaptive; Radium Tycho: disabled; Claude Haiku: enabled (1024-token budget). Each result records its request mode and manual thinking budget, when applicable.

04

Execution order

Accuracy request order alternates within each pair. Performance requests use balanced blocks. 2 warmups/model are excluded. Prompts and responses are retained in the raw results.

05

Retries and interruption

Transport, 408, 429, and covered 5xx errors receive up to 2 retries. Attempts and token usage are retained. 3 consecutive exhausted transient attempts stop and checkpoint the run.