Inference
View MarkdownBeta: this feature may change before general availability
Nodus inference is an OpenAI-compatible API over hosted open models. Point any OpenAI SDK at the Nodus base URL, use a Nodus API key, and pay per token from your prepaid credits.
Sign in and create a key
Section titled “Sign in and create a key”Sign in to the console, open Settings → API keys, and create a key with the inference:invoke scope.
Call a model
Section titled “Call a model”from openai import OpenAI
client = OpenAI(base_url="https://inference.nodus-compute.ai/v1", api_key="YOUR_NODUS_API_KEY")reply = client.chat.completions.create( model="nodus/gpt-oss-120b", messages=[{"role": "user", "content": "What is the capital of France?"}],)print(reply.choices[0].message.content)curl https://inference.nodus-compute.ai/v1/chat/completions \ -H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \ -d '{"model": "nodus/indra", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'Streaming works as in the OpenAI API (stream: true). Completed chat streams end with data: [DONE]; an interrupted or errored stream does not receive a generated completion marker. Every response carries a Nodus-Request-Id header.
From the CLI
Section titled “From the CLI”The CLI calls the same API with the credential of your current context:
$ nodus inference modelsMODEL NAME CONTEXT MAX OUTPUT PREVIEWnodus/indra Indra - - falseopenai/gpt-oss-120b GPT-OSS 120B 128K 64K falseopenai/gpt-oss-20b GPT-OSS 20B 128K 64K false$ nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 "What is the capital of France?"The capital of France is Paris.request ireq_01k6d2x7q9fvjt3y8m0c4r5n2e · openai/gpt-oss-20b · 84 input + 21 output tokens · $0.000015$ nodus inference receipt ireq_01k6d2x7q9fvjt3y8m0c4r5n2echat prints the answer on standard output and the request id, the model that answered, the tokens and the charge
on standard error; -o json prints the whole completion. receipt shows any request’s charge for 30 days.
Models
Section titled “Models”GET /v1/models lists the models you can call now. Each model has a catalog name such as openai/gpt-oss-120b
and a nodus/ alias such as nodus/gpt-oss-120b. Models marked preview have provisional prices.
In the console, open Inference → Models and select Indra to send a request to nodus/indra, or choose
a model from the catalog. The response summary’s Model field shows the model identifier submitted with that
request, including nodus/indra when using Indra.
The API’s routing metadata and receipt still record which underlying model served the request.
Responses, Messages and embeddings
Section titled “Responses, Messages and embeddings”Chat models also accept stateless Responses and Anthropic Messages requests, with the same model ids, credits, endpoint limits and receipt headers. Responses supports text, inline images and client function tools; stored responses, background work and hosted tools are refused. Text Completions accepts one string prompt and one choice, without echo, suffix, logprobs or best-of sampling. Messages token counting is not available.
response = client.responses.create(model="nodus/gpt-oss-20b", input="Hello", max_output_tokens=64, store=False)print(response.output_text)vectors = client.embeddings.create(model="nodus/bge-m3", input=["first document", "second document"])print(vectors.data[0].embedding)Embeddings are available when the embedding model appears in GET /v1/models. They charge only input tokens.
A model’s operations field identifies the routes it accepts, including Embeddings, and aliases lists its
alternative model ids. The console labels
embedding models as text → vectors and provides embedding examples in the model panel and endpoint API tab.
A request can carry up to 2,048 inputs, each within the model’s context window. Messages accepts the same chat
models at /v1/messages, with max_tokens, messages and optional stream: true; it does not expose private
models used by managed agents. Streaming client tool arguments remain intact when several calls interleave.
Transcribe or translate an audio file up to 25 MiB (flac, mp3, mp4, mpeg, m4a, ogg, wav or webm) with
openai/whisper-large-v3, and generate speech with canopylabs/orpheus-v1-english:
with open("meeting.m4a", "rb") as f: text = client.audio.transcriptions.create(model="nodus/whisper-large-v3", file=f)print(text.text)
speech = client.audio.speech.create(model="canopylabs/orpheus-v1-english", voice="tara", input="Your job finished.", response_format="wav")speech.write_to_file("done.wav")Upload the file itself: audio URLs and streamed transcription are not supported. Transcription and translation are billed per audio hour for the file’s length, with a 10-second minimum per request; speech is billed per million input characters, and its audio is limited to 32 MiB.
Indra: nodus/indra
Section titled “Indra: nodus/indra”Send model: "nodus/indra" and Indra picks the catalog model best suited to each request: small, fast models for
simple requests and stronger models for hard ones. The X-Nodus-Routed-Model response header names the model that
answered. If routing is unavailable, Indra answers with its default model, so requests never fail because of
routing. Indra was first called Composer, and nodus/auto still works as another name for it.
Indra comes in two tiers, and the X-Nodus-Composer-Tier response header (free or paid) names the one that
answered:
- Free Indra is for every organization without the Indra plan. It routes among the three cheapest catalog
models that are available (by blended price, three input tokens to one output token) and never draws on your
credits. It is limited per organization to 100 requests and 200,000 tokens per UTC day, and each request to 4,096
output tokens (fewer when the day has less left); past either daily limit, requests get
429with codeQuotaExceededuntil 00:00 UTC. Free Indra can also pause for the rest of the day when it is in heavy demand. - Paid Indra comes with the $20 per month Indra plan. It routes across every catalog model, and you pay the routed model’s per-token price plus the small routing charge, from the plan’s monthly Indra allowance first and then from your credits.
Calling a model by name, directly or through a named endpoint, is always billed per token from your credits.
Named endpoints and limits
Section titled “Named endpoints and limits”An InferenceEndpoint gives a model its own base URL and access policy:
$ nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-concurrent 8 --max-cost 50$ nodus wait ep/support-bot --for=jsonpath={.status.phase}=Running$ nodus get epNAME PHASE MODEL READY RPM CONCURRENT MAX-COST AGEsupport-bot Running nodus/gpt-oss-120b True 120 8 $50 12s$ nodus inference chat --endpoint support-bot "Where is my order?"Call it at https://inference.nodus-compute.ai/endpoints/<project>/support-bot/v1, or send
model: "endpoint/support-bot" on the shared base URL. Endpoints enforce rpm, tpm, maxConcurrent,
allowedKeys (API key names, --allowed-key) and an optional maxCostUSD: once the requests through an endpoint
have spent its cap, the next one gets 402 BudgetExceeded, and raising maxCostUSD lets requests through again (it
can only be raised). A new endpoint answers 503 for the few seconds until it is Running. nodus stop ep/support-bot makes it answer 503 until nodus start ep/support-bot; its Ready condition says whether its model
can serve now.
Billing per token
Section titled “Billing per token”- Before a request runs, Nodus holds its maximum cost: the input bound plus
max_tokensat the model’s rates. Lowermax_tokensto hold less. - When it finishes, you are charged the tokens used (input, cached input and output) and the rest of the hold is released. A stream you stop early is charged for the tokens it produced.
- Send an
Idempotency-Keyheader to retry safely: a repeat within 24 hours returns the same answer withIdempotent-Replayed: trueand no second charge. A request that failed without a charge (Released) frees its key, so the retry runs again. Reusing a key for a different request body returns409. GET /v1/requests/{id}returns the receipt for 30 days: itsoperation,usage,amountUSD,pricebookVersionandstate:Running,Settled,Released(no tokens, no charge) orUnknown. A request whose outcome Nodus never learned isUnknownand is never charged.
Errors
Section titled “Errors”Errors use the OpenAI shape, with the Nodus error code in both type and code:
| Status | Code | Meaning |
|---|---|---|
| 400 | Unsupported |
A feature outside per-token billing: built-in or server-side tools, a non-default service_tier, store, background, file parts, image URLs or audio URLs, a key repeated in another letter case, or an operation the model does not serve |
| 400 | Invalid |
A malformed request, or an upload that is not audio in a supported format |
| 401 | Unauthorized |
Missing or invalid API key |
| 402 | InsufficientCredits, BudgetExceeded |
Not enough credit for the hold, or a budget or endpoint cap blocks it |
| 409 | RequestInProgress |
The same Idempotency-Key is still running; retry after Retry-After |
| 409 | IdempotencyKeyReused |
The Idempotency-Key was used for a different request; send a new key |
| 413 | RequestEntityTooLarge |
A JSON body over 1 MiB or an audio file over 25 MiB |
| 429 | TooManyRequests |
An org or endpoint limit; retry after Retry-After |
| 502, 503 | Unavailable |
The model failed, is at capacity or is not available; retry shortly. You are not charged |