Inference endpoints and the inference CLI
Hosted inference is an OpenAI-compatible API over the model catalog; see Inference.
- Named endpoints.
nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-cost 50gives a model its own base URL,/endpoints/<project>/support-bot/v1, and the policy that applies there: requests per minute, tokens per minute, open requests at once, the API keys allowed to call it and a total spending cap. Once the cap is spent the next request gets402 BudgetExceeded; raisespec.maxCostUSDto let requests through again.nodus stopmakes an endpoint answer503untilnodus start, andnodus get epshows whether its model can serve now and what the endpoint served in the last 24 hours. - The CLI calls hosted models.
nodus inference modelslists the models that can serve right now,nodus inference chat "..."sends one chat completion with your current context’s credential and prints the request id and the charge, andnodus inference receipt ireq_…shows what any request of the last 30 days cost. - Model names in the examples are catalog names. The guides and the Python SDK examples now use
nodus/gpt-oss-20bandnodus/gpt-oss-120b, asGET /v1/modelslists them.