Hosted inference and Composer
Point any OpenAI SDK at https://inference.nodus-compute.ai/v1 with a Nodus API key and call the catalog models
through chat completions, streamed or not, plus audio transcription, translation and speech. You pay per token from your credits at the models’ listed rates, and
every response carries a Nodus-Request-Id whose receipt stays at GET /v1/requests/{id} for 30 days. A request
whose outcome is never known is never charged.
nodus/auto (Composer) chooses one catalog model for each request and names it in the x-nodus-routed-model
response header; you pay for the model it chose plus the small routing call. Named inference endpoints give a
project its own URL with requests-per-minute, tokens-per-minute, concurrency and spend limits, and can be limited
to chosen API keys. The console playground calls the same API. See Inference.