Network — live capacity across every connected site
Requests per second — by outcome
Time to first token — scheduling + prefill
Tokens per second — network total
GPU utilisation — physical sites
Busiest sites — last hour, from the usage ledger
| Site | Requests | Output tokens | GPU seconds | Avg tok/s | Charged |
|---|---|---|---|---|---|
| loading… | |||||
Fleet map — one cell per site, tinted by GPU load
By region
| Region | Sites | GPUs |
|---|
By GPU class
| GPU | Online | Share |
|---|
Enrol a new site — one line, on the machine that has the GPUs
The token is single-use and expires. Nothing on the control plane changes when a site joins — it enrols itself over the message bus.
Physical workers — real hardware, live telemetry and lifecycle control
| Site | Region | GPU | Load | VRAM free | In flight | CPU | Mem | Disk | Last seen | State | Actions |
|---|---|---|---|---|---|---|---|---|---|---|---|
| loading… | |||||||||||
Drain stops new work reaching a worker while in-flight requests finish. Revoke ejects it permanently and invalidates its credential — it must re-enrol with a fresh token to come back. Commands opens remote execution: a raw command or a fixed recipe (update the runtime, load models, set concurrency, restart, uninstall), authenticated with the worker's own secret, bounded to five minutes and 64 KB of output, and logged to the audit trail before it runs and after it returns.
Commands
Uninstall permanently removes the service, binary, config and state from the machine and cannot be undone from here — it asks you to type the site name before it sends anything.
Locations — the sites inside each, and the machines inside those
A location is a physical place - a datacenter, an office, a rack - that holds one or more sites. Every site with no assigned location still shows here, grouped separately, rather than disappearing from the hierarchy.
Registered sites — everything the control plane has ever admitted
| Site | Region | Location | Workers | GPUs | Status | Actions |
|---|---|---|---|---|---|---|
| loading… | ||||||
Placement decision — why each site is or is not eligible right now
| Site | Location | GPU | Eligible | Rejected by | Score | Load | Speed | Proximity | Cost |
|---|---|---|---|---|---|---|---|---|---|
| pick a model and evaluate | |||||||||
Selection factors
| Factor | Kind | Weight |
|---|
The six hard gates admit or reject. The four weighted factors rank whatever survives, and the result is multiplied by a queue-depth penalty so a backlog costs a worker ground quickly rather than gradually.
Scheduler health
A failover means a worker was selected, failed or refused, and the request was transparently restarted on another. It stays invisible to the customer as long as no token had been streamed yet.
Send a request — proves the whole path end to end
Usage ledger — append-only, hash-chained
Recent requests — straight from the ledger
| # | Model | GPU | In | Out | GPU ms | Batch | Charged | Margin | Status |
|---|---|---|---|---|---|---|---|---|---|
| loading… | |||||||||
Usage report — billed figures, not telemetry
| Key | Requests | Failed | In tokens | Out tokens | GPU sec | Avg tok/s | Avg TTFT | Charged | Cost | Margin |
|---|---|---|---|---|---|---|---|---|---|---|
| loading… | ||||||||||
Charged is what the customer pays, token-based. Cost is what the GPU time was worth, divided across the requests that shared the GPU. Margin can legitimately be negative for short requests on expensive hardware.
Models — catalogue and pricing
| Model | VRAM MB | In €/Mtok | Out €/Mtok | Workers |
|---|---|---|---|---|
| loading… | ||||
Add or update a model
VRAM required is a hard gate, checked per GPU: a worker is only eligible if at least one of its cards has this much VRAM free on its own - several smaller cards never combine to cover it, however fast or cheap the worker is.
Enrollment tokens — how a new GPU site joins
| Label | Site | Used | Expires | State |
|---|---|---|---|---|
| loading… | ||||
Mint a token
Shown once. Only its hash is stored, so it cannot be recovered later — mint a new one instead.
Customers — who may call the public API
| Name | Keys | Rate/min | Requests | Charged |
|---|---|---|---|---|
| loading… | ||||
Create a customer
The API key is shown once. Keys are stored hashed, exactly like enrollment tokens.
Audit log — enrollment, revocation, credential and catalogue changes
| When | Actor | Action | Subject | Detail |
|---|---|---|---|---|
| loading… | ||||