COMPANY
One job: inference
Nyx serves open-weight models and nothing else, so every engineering hour the company has ever spent went into the serving path.
5
MODELS IN THE CATALOG
2.9×
PERFORMANCE PER DOLLAR
$0.0225
PER MILLION INPUT TOKENS
US-EAST-1
SERVING REGION
Nyx sells inference. Open-weight models behind an OpenAI-compatible endpoint at api.nyxprovider.com, serverless by the token or dedicated by the pod. That is the whole company.
There is no training cluster and no consulting arm. Revenue has to come from tokens served, and the only way to serve more tokens is to serve them faster and cheaper than the next provider. A company this size cannot afford a side quest. The serving path is the product, the roadmap, and the payroll.
The stack is built in-house from the GPU kernels to the HTTP edge. Attention kernels, KV-cache layout, the scheduler, the prefix cache, the streaming edge that puts bytes on the wire: one team owns all of it, which is why the pieces cooperate. The prefix cache persists across turns and removes 31 percent of prefill work on agent traces. First-token latency is timed from TCP connect to the first streamed byte rather than from an internal marker, so what we measure is what a caller feels. When a number needs to move, nobody files a ticket against someone else's layer.
Latency is gated rather than advertised. A figure on a marketing page is easy to write once and quietly miss forever, so ours lives in the release process instead: every engine build replays recorded traces on shadow pods, the probe percentiles are checked against the thresholds in the gate, and a build that moves one the wrong way does not ship, whatever else it improves. Keeping the site honest and keeping the fleet fast end up being the same job.
If a number cannot survive the gate, it does not go on the page.
Pricing works the same way the engineering does. Margin comes from utilization: the scheduler keeps the hardware busy, and busy hardware is what lets the gpt-oss-120b class run at $0.07 per million tokens blended, listed under the market rather than marked up to it. Nyx buys capacity and sells tokens, and the spread survives only if the fleet stays fast and full. That is where 2.9 times performance per dollar comes from. When a build earns a better cost per token, the list price follows it down.
How we work
Measured before shipped. No change to the engine reaches customer traffic on the strength of a benchmark or a code review. Shadow pods replay recorded traces against the candidate build, the replay is checked against the percentiles in the release gate, and the replay decides. A build that cannot prove a win is rejected even when it looks harmless.
Refuse past capacity. Past declared capacity the API returns an immediate 429 with a Retry-After header. Admitting one more request into a saturated fleet slows every request already in flight, and a queue that turns into someone's timeout is a worse answer than a fast, honest one your retry logic can act on.
The price is the price. The number on the pricing page is the number on the invoice. There is no long-context surcharge, no premium tier that holds the real speed behind a higher rate, and no markup on dedicated pods beyond the flat hourly price. A price that needs a footnote is a price we would rather not list.