ENTERPRISE
Dedicated capacity, priced flat
A dedicated pod runs your model alone on capacity contracted in us-east-1, at a flat hourly price with the throughput written into the order.
25%
UNDER THE CHEAPEST LISTING
2.9×
PERFORMANCE PER DOLLAR
1M
LONGEST CONTEXT WINDOW
US-EAST-1
SERVING REGION
Dedicated pods
A dedicated pod runs one model for one customer. The capacity is held for you and nothing else lands on it, so there is no other tenant's traffic beside yours and no shared queue to wait in. Billing is a flat hourly price for the pod rather than a per-token rate, which means the invoice is the same whether you saturate it at noon or leave it idle overnight.
Throughput goes in the order. Before a pod is priced we run your prompts through the model on the capacity you would be renting and record what it sustains, and that measurement at your context lengths and your concurrency, is what the order commits to. Nothing about the API surface changes: the same OpenAI-compatible chat completions route, the same byte-for-byte streaming with usage in the final chunk, and the same immediate 429 with a Retry-After header once you pass your declared ceiling. Cutover is a base URL change.
Any model in the catalog can take a pod. Four are serving today, the four live models, and kimi-k3 once its capacity lands. kimi-k3, the 1,000,000-token model, is configured and not yet serving.
Security review
A security team will ask for a list of controls. These are the ones Nyx can evidence today. Prompts and completions are never stored and never used to train anything. Logs carry metadata only, request id, model, token counts, status, and timing, and never the body of a request or a response. TLS runs end to end, from your client to the serving process. API keys are stored hashed and shown once, at creation, so a leaked key is revoked and reissued rather than recovered. Serving happens in us-east-1.
Two answers on that questionnaire are no. Nyx is not SOC 2 certified, and no audit is in progress, so any badge on this page would be decoration. Zero data retention is not contractually guaranteed either, because the compute underneath is capacity contracted from a partner rather than hardware Nyx owns, and a promise Nyx cannot enforce further down the stack is not worth signing. The model catalog reports zdr: false for that reason. The practice above holds regardless: nothing is written down and nothing is trained on.
How an order works
Bring a workload. A prompt set drawn from real traffic, the model you run now, the provider you run it on, and the shape of the load: context lengths, concurrency, tokens per day. That is enough to start.
Then we benchmark it. The same prompts run against your current provider and against Nyx on the same day, through the same harness, so the comparison is prompt for prompt rather than one vendor's marketing page against another's. You get the numbers whichever way they come out, including the runs where the incumbent wins, and you can take them with you.
Then a quote. Flat hourly on a dedicated pod, or the published per-token rates if the shared endpoint already covers the load. Those rates sit 25 percent under the cheapest listing we can find for the same model, subsidized on purpose while the capacity proves itself, which works out to a blended $0.07 per million on the gpt-oss-120b class and 2.9× the performance per dollar. Price changes come with 14 days of notice and never land mid-cycle.
A quote needs the model, the throughput, and the term. Send those three and the benchmark run comes back before the price does.