The fastest inference for open models

Kimi, gpt-oss, Qwen, and Gemma behind one OpenAI-compatible endpoint. Priced under the market, streamed byte for byte.

COMING SOON

ONE JOB: INFERENCE

Nyx does one thing: serve open-weight models fast. The entire stack, from GPU kernels to the HTTP edge, is built and tuned in-house for a single workload. That focus is why our tokens arrive sooner and cost less.

Open

WEIGHTS ONLY, NO WRAPPERS

Compatible

OPENAI API, SWAP THE URL

Under

THE CHEAPEST LISTING WE FIND

US East

SERVING REGION

Rendered gold wafer with a lit die mosaic

SERVERLESS INFERENCE (NYX-1)

Swap one URL, keep your code

Every model in the catalog sits behind an OpenAI-compatible endpoint. Point your existing client at api.nyxprovider.com, pick a model, and stream. Pay per token, with no minimums and no cold starts.

COMING SOON
Graded mountain valley photograph

Output speed · kimi-k3

Ahead of the market

Nyx
Fastest cloud
Market median

Independent harness · matched prompts · streamed

DEDICATED CAPACITY

Your own pods, pinned to your traffic

Reserved GPU pods run your model alone, with nothing shared between you and your users. Flat hourly pricing, guaranteed throughput, and latency that holds at peak.

COMING SOON
Graded prairie grass photograph

Dedicated · kimi-k3

Reserve by the pod

pod-aon request
pod-b · on requestflat hourly
pod-c · on requestpinned to you

THE NYX ENGINE

Every change measured before it ships

Each engine build replays recorded traces on shadow pods before it touches traffic. Wins ship. Regressions never leave the bench. Every change has to prove itself on a real workload shape before it can take a request.

MEASURE
TUNE
VERIFY
SHIP
BUILDCHANGEMEASUREDDECISION
B-231FP8 KV-cache decodeFaster decode, quality paritySHIPPED
B-234Speculative draft, 4-token windowSlower tail under loadREJECTED
B-238Continuous batching rebalanceSteadier first token under loadSHIPPED
B-241Attention block size 32 → 64No measurable changeREJECTED
B-243Prefix cache across turnsLess prefill on agent tracesSHIPPED
Graded fog ridge photograph

Performance per dollar

Under the market

Nyx
Provider A
Provider B

Blended cost per token at matched output speed · gpt-oss-120b class

Powering inference for engineers from

UC Berkeley Stanford UT Austin UT Dallas Princeton University

Bring a workload. We will benchmark it on Nyx against your current provider and send you the numbers.

COMING SOON