Skip to content

Inference

View Markdown

Beta: this feature may change before general availability

Nodus inference is an OpenAI-compatible API over hosted open models. Point any OpenAI SDK at the Nodus base URL, use a Nodus API key, and pay per token from your prepaid credits.

Sign in to the console, open Settings → API keys, and create a key with the inference:invoke scope.

from openai import OpenAI
client = OpenAI(base_url="https://inference.nodus-compute.ai/v1", api_key="YOUR_NODUS_API_KEY")
reply = client.chat.completions.create(
model="nodus/gpt-oss-120b",
messages=[{"role": "user", "content": "What is the capital of France?"}],
)
print(reply.choices[0].message.content)
Terminal window
curl https://inference.nodus-compute.ai/v1/chat/completions \
-H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "nodus/indra", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'

Streaming works as in the OpenAI API (stream: true). Completed chat streams end with data: [DONE]; an interrupted or errored stream does not receive a generated completion marker. Every response carries a Nodus-Request-Id header.

The CLI calls the same API with the credential of your current context:

Terminal window
$ nodus inference models
MODEL NAME CONTEXT MAX OUTPUT PREVIEW
nodus/indra Indra - - false
openai/gpt-oss-120b GPT-OSS 120B 128K 64K false
openai/gpt-oss-20b GPT-OSS 20B 128K 64K false
$ nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 "What is the capital of France?"
The capital of France is Paris.
request ireq_01k6d2x7q9fvjt3y8m0c4r5n2e · openai/gpt-oss-20b · 84 input + 21 output tokens · $0.000015
$ nodus inference receipt ireq_01k6d2x7q9fvjt3y8m0c4r5n2e

chat prints the answer on standard output and the request id, the model that answered, the tokens and the charge on standard error; -o json prints the whole completion. receipt shows any request’s charge for 30 days.

GET /v1/models lists the models you can call now. Each model has a catalog name such as openai/gpt-oss-120b and a nodus/ alias such as nodus/gpt-oss-120b. Models marked preview have provisional prices.

In the console, open Inference → Models and select Indra to send a request to nodus/indra, or choose a model from the catalog. The response summary’s Model field shows the model identifier submitted with that request, including nodus/indra when using Indra. The API’s routing metadata and receipt still record which underlying model served the request.

Chat models also accept stateless Responses and Anthropic Messages requests, with the same model ids, credits, endpoint limits and receipt headers. Responses supports text, inline images and client function tools; stored responses, background work and hosted tools are refused. Text Completions accepts one string prompt and one choice, without echo, suffix, logprobs or best-of sampling. Messages token counting is not available.

response = client.responses.create(model="nodus/gpt-oss-20b", input="Hello", max_output_tokens=64, store=False)
print(response.output_text)
vectors = client.embeddings.create(model="nodus/bge-m3", input=["first document", "second document"])
print(vectors.data[0].embedding)

Embeddings are available when the embedding model appears in GET /v1/models. They charge only input tokens. A model’s operations field identifies the routes it accepts, including Embeddings, and aliases lists its alternative model ids. The console labels embedding models as text → vectors and provides embedding examples in the model panel and endpoint API tab. A request can carry up to 2,048 inputs, each within the model’s context window. Messages accepts the same chat models at /v1/messages, with max_tokens, messages and optional stream: true; it does not expose private models used by managed agents. Streaming client tool arguments remain intact when several calls interleave.

Transcribe or translate an audio file up to 25 MiB (flac, mp3, mp4, mpeg, m4a, ogg, wav or webm) with openai/whisper-large-v3, and generate speech with canopylabs/orpheus-v1-english:

with open("meeting.m4a", "rb") as f:
text = client.audio.transcriptions.create(model="nodus/whisper-large-v3", file=f)
print(text.text)
speech = client.audio.speech.create(model="canopylabs/orpheus-v1-english", voice="tara",
input="Your job finished.", response_format="wav")
speech.write_to_file("done.wav")

Upload the file itself: audio URLs and streamed transcription are not supported. Transcription and translation are billed per audio hour for the file’s length, with a 10-second minimum per request; speech is billed per million input characters, and its audio is limited to 32 MiB.

Send model: "nodus/indra" and Indra picks the catalog model best suited to each request: small, fast models for simple requests and stronger models for hard ones. The X-Nodus-Routed-Model response header names the model that answered. If routing is unavailable, Indra answers with its default model, so requests never fail because of routing. Indra was first called Composer, and nodus/auto still works as another name for it.

Indra comes in two tiers, and the X-Nodus-Composer-Tier response header (free or paid) names the one that answered:

  • Free Indra is for every organization without the Indra plan. It routes among the three cheapest catalog models that are available (by blended price, three input tokens to one output token) and never draws on your credits. It is limited per organization to 100 requests and 200,000 tokens per UTC day, and each request to 4,096 output tokens (fewer when the day has less left); past either daily limit, requests get 429 with code QuotaExceeded until 00:00 UTC. Free Indra can also pause for the rest of the day when it is in heavy demand.
  • Paid Indra comes with the $20 per month Indra plan. It routes across every catalog model, and you pay the routed model’s per-token price plus the small routing charge, from the plan’s monthly Indra allowance first and then from your credits.

Calling a model by name, directly or through a named endpoint, is always billed per token from your credits.

An InferenceEndpoint gives a model its own base URL and access policy:

Terminal window
$ nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-concurrent 8 --max-cost 50
$ nodus wait ep/support-bot --for=jsonpath={.status.phase}=Running
$ nodus get ep
NAME PHASE MODEL READY RPM CONCURRENT MAX-COST AGE
support-bot Running nodus/gpt-oss-120b True 120 8 $50 12s
$ nodus inference chat --endpoint support-bot "Where is my order?"

Call it at https://inference.nodus-compute.ai/endpoints/<project>/support-bot/v1, or send model: "endpoint/support-bot" on the shared base URL. Endpoints enforce rpm, tpm, maxConcurrent, allowedKeys (API key names, --allowed-key) and an optional maxCostUSD: once the requests through an endpoint have spent its cap, the next one gets 402 BudgetExceeded, and raising maxCostUSD lets requests through again (it can only be raised). A new endpoint answers 503 for the few seconds until it is Running. nodus stop ep/support-bot makes it answer 503 until nodus start ep/support-bot; its Ready condition says whether its model can serve now.

  • Before a request runs, Nodus holds its maximum cost: the input bound plus max_tokens at the model’s rates. Lower max_tokens to hold less.
  • When it finishes, you are charged the tokens used (input, cached input and output) and the rest of the hold is released. A stream you stop early is charged for the tokens it produced.
  • Send an Idempotency-Key header to retry safely: a repeat within 24 hours returns the same answer with Idempotent-Replayed: true and no second charge. A request that failed without a charge (Released) frees its key, so the retry runs again. Reusing a key for a different request body returns 409.
  • GET /v1/requests/{id} returns the receipt for 30 days: its operation, usage, amountUSD, pricebookVersion and state: Running, Settled, Released (no tokens, no charge) or Unknown. A request whose outcome Nodus never learned is Unknown and is never charged.

Errors use the OpenAI shape, with the Nodus error code in both type and code:

Status Code Meaning
400 Unsupported A feature outside per-token billing: built-in or server-side tools, a non-default service_tier, store, background, file parts, image URLs or audio URLs, a key repeated in another letter case, or an operation the model does not serve
400 Invalid A malformed request, or an upload that is not audio in a supported format
401 Unauthorized Missing or invalid API key
402 InsufficientCredits, BudgetExceeded Not enough credit for the hold, or a budget or endpoint cap blocks it
409 RequestInProgress The same Idempotency-Key is still running; retry after Retry-After
409 IdempotencyKeyReused The Idempotency-Key was used for a different request; send a new key
413 RequestEntityTooLarge A JSON body over 1 MiB or an audio file over 25 MiB
429 TooManyRequests An org or endpoint limit; retry after Retry-After
502, 503 Unavailable The model failed, is at capacity or is not available; retry shortly. You are not charged