The fastest inference for open models
Kimi, gpt-oss, Qwen, and Gemma behind one OpenAI-compatible endpoint. Priced under the market, streamed byte for byte.
COMING SOONONE JOB: INFERENCE
Nyx does one thing: serve open-weight models fast. The entire stack, from GPU kernels to the HTTP edge, is built and tuned in-house for a single workload. That focus is why our tokens arrive sooner and cost less.
Open
WEIGHTS ONLY, NO WRAPPERS
Compatible
OPENAI API, SWAP THE URL
Under
THE CHEAPEST LISTING WE FIND
US East
SERVING REGION
SERVERLESS INFERENCE (NYX-1)
Swap one URL, keep your code
Every model in the catalog sits behind an OpenAI-compatible endpoint. Point your existing client at api.nyxprovider.com, pick a model, and stream. Pay per token, with no minimums and no cold starts.
COMING SOON
Output speed · kimi-k3
Ahead of the market
Independent harness · matched prompts · streamed
DEDICATED CAPACITY
Your own pods, pinned to your traffic
Reserved GPU pods run your model alone, with nothing shared between you and your users. Flat hourly pricing, guaranteed throughput, and latency that holds at peak.
COMING SOON
Dedicated · kimi-k3
Reserve by the pod
THE NYX ENGINE
Every change measured before it ships
Each engine build replays recorded traces on shadow pods before it touches traffic. Wins ship. Regressions never leave the bench. Every change has to prove itself on a real workload shape before it can take a request.
Performance per dollar
Under the market
Blended cost per token at matched output speed · gpt-oss-120b class