Hosted inference and Composer

Point any OpenAI SDK at https://inference.nodus-compute.ai/v1 with a Nodus API key and call the catalog models through chat completions, streamed or not, plus audio transcription, translation and speech. You pay per token from your credits at the models’ listed rates, and every response carries a Nodus-Request-Id whose receipt stays at GET /v1/requests/{id} for 30 days. A request whose outcome is never known is never charged.

nodus/auto (Composer) chooses one catalog model for each request and names it in the x-nodus-routed-model response header; you pay for the model it chose plus the small routing call. Named inference endpoints give a project its own URL with requests-per-minute, tokens-per-minute, concurrency and spend limits, and can be limited to chosen API keys. The console playground calls the same API. See Inference.