GPU Inference Network

Operations console. Paste the admin token to continue — it is exchanged for a session cookie, not kept in the page.

ADMIN_TOKEN — deploy/.env on the control-plane host
gpunet / operations
connecting
heartbeats
—
requests
—
in flight
—

Network — live capacity across every connected site

GPUs online
—
 
Sites
—
 
Regions
—
 
Generating now
—
 
Success rate
—
since start
Charged 24h
—
 

Requests per second — by outcome

Time to first token — scheduling + prefill

Tokens per second — network total

GPU utilisation — physical sites

Busiest sites — last hour, from the usage ledger

SiteRequestsOutput tokensGPU secondsAvg tok/sCharged
loading…

Fleet map — one cell per site, tinted by GPU load

idle saturated outlined = physical hardware · hover a cell for detail

By region

RegionSitesGPUs

By GPU class

GPUOnlineShare

Enrol a new site — one line, on the machine that has the GPUs

The token is single-use and expires. Nothing on the control plane changes when a site joins — it enrols itself over the message bus.

Physical workers — real hardware, live telemetry and lifecycle control

SiteRegionGPULoadVRAM freeIn flightCPUMemDiskLast seenStateActions
loading…

Drain stops new work reaching a worker while in-flight requests finish. Revoke ejects it permanently and invalidates its credential — it must re-enrol with a fresh token to come back. Commands opens remote execution: a raw command or a fixed recipe (update the runtime, load models, set concurrency, restart, uninstall), authenticated with the worker's own secret, bounded to five minutes and 64 KB of output, and logged to the audit trail before it runs and after it returns.

Commands

Uninstall permanently removes the service, binary, config and state from the machine and cannot be undone from here — it asks you to type the site name before it sends anything.

no commands run yet

Locations — the sites inside each, and the machines inside those

A location is a physical place - a datacenter, an office, a rack - that holds one or more sites. Every site with no assigned location still shows here, grouped separately, rather than disappearing from the hierarchy.

loading…

Registered sites — everything the control plane has ever admitted

SiteRegionLocationWorkersGPUsStatusActions
loading…

Placement decision — why each site is or is not eligible right now

SiteLocationGPUEligibleRejected byScoreLoadSpeedProximityCost
pick a model and evaluate

Selection factors

FactorKindWeight

The six hard gates admit or reject. The four weighted factors rank whatever survives, and the result is multiplied by a queue-depth penalty so a backlog costs a worker ground quickly rather than gradually.

Scheduler health

Selections
—
Mean latency
—
No capacity
—
Failovers
—

A failover means a worker was selected, failed or refused, and the request was transparently restarted on another. It stays invisible to the customer as long as no token had been streamed yet.

Send a request — proves the whole path end to end

Nothing sent yet.

Usage ledger — append-only, hash-chained

Records
—
Chain
—

Recent requests — straight from the ledger

#ModelGPUInOutGPU msBatchChargedMarginStatus
loading…

Usage report — billed figures, not telemetry

KeyRequestsFailedIn tokensOut tokensGPU secAvg tok/sAvg TTFTChargedCostMargin
loading…

Charged is what the customer pays, token-based. Cost is what the GPU time was worth, divided across the requests that shared the GPU. Margin can legitimately be negative for short requests on expensive hardware.

Models — catalogue and pricing

ModelVRAM MBIn €/MtokOut €/MtokWorkers
loading…

Add or update a model

VRAM required is a hard gate, checked per GPU: a worker is only eligible if at least one of its cards has this much VRAM free on its own - several smaller cards never combine to cover it, however fast or cheap the worker is.

Enrollment tokens — how a new GPU site joins

LabelSiteUsedExpiresState
loading…

Mint a token

Shown once. Only its hash is stored, so it cannot be recovered later — mint a new one instead.

Customers — who may call the public API

NameKeysRate/minRequestsCharged
loading…

Create a customer

The API key is shown once. Keys are stored hashed, exactly like enrollment tokens.

Audit log — enrollment, revocation, credential and catalogue changes

WhenActorActionSubjectDetail
loading…