Notes from the serving path
What the team running nyxprovider.com is working on: how models get served, how latency gets measured, and how prices get set. Every claim here can be checked against /v1/models or /health.
Serving kimi-k3 at a million tokens of context
The full window is live, short prompts keep their latency, and the price sits under the cheapest listing.
The serving path, measured in microseconds
Why we publish 71µs p50 routing overhead, how loopback probes measure it, and what an immediate 429 buys you.
Why every model sits under its cheapest listing
The rule, the daily scrape behind it, and why serving efficiency pays for the discount.