Tokenomics

Tokens look simple, but inside them lives a stack of cloud costs, infrastructure margins, and licensing layers.

What’s inside

Every AI platform bills you 
in tokens. None of them tell you what is in one.

AI platforms charge per token because a token roughly represents computation. What gets bundled into that number is a separate question, and it is not one 
you are invited to ask.

Understand the Radium stack.

Our technology
Token Comparison

Same class of model.
Different composition.

Proportions illustrative.
3D pie chart showing AWS/Azure, Model Licensing, and Overhead segments with AWS/Azure largest.
Anthropic / OpenAI token

Compute is one slice. The rest is cloud rent, idle capacity, licensing, training cost recovery, and platform margin, bundled into a single number on your invoice.

Pie chart showing savings from radium technology overhead reduction.
Radium Token

Compute, and the cost of running it well.

Understanding Tokens

Most infrastructure leaves
GPU capacity unused

Traditional inference systems often underutilize GPU capacity due to fragmented routing, idle cycles, and inefficient execution. This means fewer tokens are produced from the same hardware.

Rewriting the token

Radium has optimized
token througput on our GPUs

Radium optimizes routing, batching, and execution across the delivery stack. By keeping GPUs consistently utilized, Radium increases the amount of output produced from the same infrastructure.

RADIUM OS

Every layer designed for throughput.

Benchmarks

Radium OS is our proprietary architecture, built as an orchestration layer for production inference. We designed each layer of it, from the hardware we own through routing, batching and execution, to move more tokens through the same silicon.

Independent studies from Stanford CRFM and Carnegie Mellon have measured our accelerator utilisation and iteration time against major cloud platforms.

Dark background with a scattered pattern of small white dots mostly on the right side.
Inside Radium OS

Five layers, all of them ours.

Owning the stack removes the markup. Designing the stack removes the waste. Radium does both.

01

Hardware

Bare metal GPUs we own and operate. No hyperscaler in the path, and no landlord in the price.

02

Routing

Requests placed against capacity that is actually free, so work does not queue behind a busy machine while another sits idle.

03

Batching

Requests grouped so GPUs stay saturated rather than waking up for one job at a time.

04

Models

Tuned open weight models, Hal, Clarke and Tycho, matched to production workloads rather than to benchmarks.

05

Endpoint

OpenAI and Anthropic compatible, so the code you already wrote keeps working.

Benchmakrs

How Radium
benchmarks
inference

Methodology
The method behind every number we publish.
Abstract beige background with curved dotted lines forming a wave pattern.

Hal 1.0
performance
results

Resources
Twelve workloads, every percentile published.

156 requests,
zero failures

Across twelve workload configurations at one, four and eight concurrent sessions.

924 output tokens
per second

Peak sustained output throughput at eight concurrent sessions.

16,077 total tokens
per second

Peak total throughput on the largest input workload.

821 ms

Lowest median latency recorded in the run.

MEASURED

Efficiency you can measure.

Throughput claims are easy to make and hard to verify, so ours are published with the workload, the configuration and the limitations attached. The numbers below come from a single node under sustained load.

Read our benchmark report

A bill with nothing to decode.

Get your API key
Dense cluster of white dots scattered against a dark background, resembling stars or particles.