Radium and Claude matched-tier results
Answer accuracy, throughput, latency,
and time to first text.
Model results at a glance
Accuracy/tokens are from the 80-task accuracy suite; throughput/latency/TTFT are from the separate 40-task performance suite and aren't directly comparable. Scoring requires an exact JSON match; some token counts are estimated.
Six confidence intervals
Matched-tier comparisons
Hal × Opus
Accuracy: Radium 96.25%; Claude 93.75%.
Difference CI90: -3.65 to 8.94 points.
Aggregate output-token throughput ratio:
4.64×. Radium p50 latency was 20.18% higher.
Clarke × Sonnet
Accuracy: Radium 85%; Claude 100%.
Difference CI90: -22.7 to -8.67 points.
Aggregate output-token throughput ratio:
0.86×. Radium p50 latency was 75.54% lower.
Tycho × Haiku
Accuracy: Radium 91.25%; Claude 100%.
Difference CI90: -15.39 to -3.63 points.
Aggregate output-token throughput ratio:
0.92×. Radium p50 latency was 27.31% lower.
Confidence, parsing, tails, and tokens
Answer results by task category
Scoring and execution settings
Accuracy criteria
80 fixed-answer tasks/model. Scoring accepts a raw JSON object or a single Markdown-fenced JSON object, then requires exact fields, primitive types, and values. Equivalence requires both models to reach 95% and the 90% accuracy-difference interval to fall within ±5 points.
Performance criteria
Radium must have no lower correct rate, at least 20% lower p50 latency, and at least 1.25× aggregate output-token throughput.
Request modes
Radium Hal: disabled; Claude Opus: adaptive; Radium Clarke: disabled; Claude Sonnet: adaptive; Radium Tycho: disabled; Claude Haiku: enabled (1024-token budget). Each result records its request mode and manual thinking budget, when applicable.
Execution order
Accuracy request order alternates within each pair. Performance requests use balanced blocks. 2 warmups/model are excluded. Prompts and responses are retained in the raw results.
Retries and interruption
Transport, 408, 429, and covered 5xx errors receive up to 2 retries. Attempts and token usage are retained. 3 consecutive exhausted transient attempts stop and checkpoint the run.
Limitations
This vendor-authored test covers straightforward structured tasks and is not an independent evaluation. It does not measure difficult reasoning, open-ended response quality, safety, long context, tool use, multimodality, domain specialization, availability, or price. Performance requests used the recorded fixed load (4 simultaneous requests); throughput is not a maximum-capacity claim. Output-token throughput is based on provider- reported usage or flagged local estimates and is not a provider-neutral measure of visible text because tokenizers and hidden-reasoning accounting may differ. Results apply to the recorded models, endpoints, thinking modes, prompts, scoring policy, request load, retries, output limit, test window, service tiers, and client/network path. A single Markdown fence is ignored during answer scoring; other surrounding prose is not accepted. Client region was not recorded. Parameterized tasks may be correlated. A requested thinking budget is a request setting, not a verified cap on API-reported output tokens. Anthropic did not sponsor or validate this test. Replication on representative workloads is required for broader conclusions.