# Inference

> Call hosted models through an OpenAI-compatible API, route with Indra, and pay per token.

Source: https://nodus-platform-site.pages.dev/docs/guides/inference/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

**Beta:** this feature may change.

Nodus inference is an OpenAI-compatible API over hosted open models. Point any OpenAI SDK at the Nodus base URL, use a Nodus API key, and pay per token from your prepaid credits.

## Sign in and create a key

Sign in to the console, open **Settings → API keys**, and create a key with the `inference:invoke` scope.

## Call a model

```python
from openai import OpenAI

client = OpenAI(base_url="https://inference.nodus-compute.ai/v1", api_key="YOUR_NODUS_API_KEY")
reply = client.chat.completions.create(
    model="nodus/gpt-oss-120b",
    messages=[{"role": "user", "content": "What is the capital of France?"}],
)
print(reply.choices[0].message.content)
```

Terminal window

```bash
curl https://inference.nodus-compute.ai/v1/chat/completions \
  -H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \
  -d '{"model": "nodus/indra", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'
```

Streaming works as in the OpenAI API (`stream: true`). Completed chat streams end with `data: [DONE]`; an interrupted or errored stream does not receive a generated completion marker. Every response carries a `Nodus-Request-Id` header.

## From the CLI

The CLI calls the same API with the credential of your current context:

Terminal window

```console
$ nodus inference models
MODEL                 NAME           CONTEXT   MAX OUTPUT   PREVIEW
nodus/indra           Indra          -         -            false
openai/gpt-oss-120b   GPT-OSS 120B   128K      64K          false
openai/gpt-oss-20b    GPT-OSS 20B    128K      64K          false
$ nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 "What is the capital of France?"
The capital of France is Paris.
request ireq_01k6d2x7q9fvjt3y8m0c4r5n2e · openai/gpt-oss-20b · 84 input + 21 output tokens · $0.000015
$ nodus inference receipt ireq_01k6d2x7q9fvjt3y8m0c4r5n2e
```

`chat` prints the answer on standard output and the request id, the model that answered, the tokens and the charge on standard error; `-o json` prints the whole completion. `receipt` shows any request’s charge for 30 days.

## Models

`GET /v1/models` lists the models you can call now. Each model has a catalog name such as `openai/gpt-oss-120b` and a `nodus/` alias such as `nodus/gpt-oss-120b`. Models marked preview have provisional prices.

In the console, open **Inference → Models** and select Indra to send a request to `nodus/indra`, or choose a model from the catalog. The response summary’s **Model** field shows the model identifier submitted with that request, including `nodus/indra` when using Indra. The API’s routing metadata and receipt still record which underlying model served the request.

## Responses, Messages and embeddings

Chat models also accept stateless Responses and Anthropic Messages requests, with the same model ids, credits, endpoint limits and receipt headers. Responses supports text, inline images and client function tools; stored responses, background work and hosted tools are refused. Text Completions accepts one string prompt and one choice, without echo, suffix, logprobs or best-of sampling. Messages token counting is not available.

```python
response = client.responses.create(model="nodus/gpt-oss-20b", input="Hello", max_output_tokens=64, store=False)
print(response.output_text)
vectors = client.embeddings.create(model="nodus/bge-m3", input=["first document", "second document"])
print(vectors.data[0].embedding)
```

Embeddings are available when the embedding model appears in `GET /v1/models`. They charge only input tokens. A model’s `operations` field identifies the routes it accepts, including `Embeddings`, and `aliases` lists its alternative model ids. The console labels embedding models as **text → vectors** and provides embedding examples in the model panel and endpoint API tab. A request can carry up to 2,048 inputs, each within the model’s context window. Messages accepts the same chat models at `/v1/messages`, with `max_tokens`, `messages` and optional `stream: true`; it does not expose private models used by managed agents. Streaming client tool arguments remain intact when several calls interleave.

## Audio

Transcribe or translate an audio file up to 25 MiB (flac, mp3, mp4, mpeg, m4a, ogg, wav or webm) with `openai/whisper-large-v3`, and generate speech with `canopylabs/orpheus-v1-english`:

```python
with open("meeting.m4a", "rb") as f:
    text = client.audio.transcriptions.create(model="nodus/whisper-large-v3", file=f)
print(text.text)

speech = client.audio.speech.create(model="canopylabs/orpheus-v1-english", voice="tara",
                                    input="Your job finished.", response_format="wav")
speech.write_to_file("done.wav")
```

Upload the file itself: audio URLs and streamed transcription are not supported. Transcription and translation are billed per audio hour for the file’s length, with a 10-second minimum per request; speech is billed per million input characters, and its audio is limited to 32 MiB.

## Indra: `nodus/indra`

Send `model: "nodus/indra"` and Indra picks the catalog model best suited to each request: small, fast models for simple requests and stronger models for hard ones. The `X-Nodus-Routed-Model` response header names the model that answered. If routing is unavailable, Indra answers with its default model, so requests never fail because of routing. Indra was first called Composer, and `nodus/auto` still works as another name for it.

Indra comes in two tiers, and the `X-Nodus-Composer-Tier` response header (`free` or `paid`) names the one that answered:

* **Free Indra** is for every organization without the Indra plan. It routes among the three cheapest catalog models that are available (by blended price, three input tokens to one output token) and never draws on your credits. It is limited per organization to 100 requests and 200,000 tokens per UTC day, and each request to 4,096 output tokens (fewer when the day has less left); past either daily limit, requests get `429` with code `QuotaExceeded` until 00:00 UTC. Free Indra can also pause for the rest of the day when it is in heavy demand.
* **Paid Indra** comes with the $20 per month Indra plan. It routes across every catalog model, and you pay the routed model’s per-token price plus the small routing charge, from the plan’s monthly Indra allowance first and then from your credits.

Calling a model by name, directly or through a named endpoint, is always billed per token from your credits.

## Named endpoints and limits

An `InferenceEndpoint` gives a model its own base URL and access policy:

Terminal window

```console
$ nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-concurrent 8 --max-cost 50
$ nodus wait ep/support-bot --for=jsonpath={.status.phase}=Running
$ nodus get ep
NAME          PHASE     MODEL                READY   RPM   CONCURRENT   MAX-COST   AGE
support-bot   Running   nodus/gpt-oss-120b   True    120   8            $50        12s
$ nodus inference chat --endpoint support-bot "Where is my order?"
```

Call it at `https://inference.nodus-compute.ai/endpoints/<project>/support-bot/v1`, or send `model: "endpoint/support-bot"` on the shared base URL. Endpoints enforce `rpm`, `tpm`, `maxConcurrent`, `allowedKeys` (API key names, `--allowed-key`) and an optional `maxCostUSD`: once the requests through an endpoint have spent its cap, the next one gets `402 BudgetExceeded`, and raising `maxCostUSD` lets requests through again (it can only be raised). A new endpoint answers `503` for the few seconds until it is `Running`. `nodus stop ep/support-bot` makes it answer `503` until `nodus start ep/support-bot`; its `Ready` condition says whether its model can serve now.

## Billing per token

* Before a request runs, Nodus holds its maximum cost: the input bound plus `max_tokens` at the model’s rates. Lower `max_tokens` to hold less.
* When it finishes, you are charged the tokens used (input, cached input and output) and the rest of the hold is released. A stream you stop early is charged for the tokens it produced.
* Send an `Idempotency-Key` header to retry safely: a repeat within 24 hours returns the same answer with `Idempotent-Replayed: true` and no second charge. A request that failed without a charge (`Released`) frees its key, so the retry runs again. Reusing a key for a different request body returns `409`.
* `GET /v1/requests/{id}` returns the receipt for 30 days: its `operation`, `usage`, `amountUSD`, `pricebookVersion` and `state`: `Running`, `Settled`, `Released` (no tokens, no charge) or `Unknown`. A request whose outcome Nodus never learned is `Unknown` and is never charged.

## Errors

Errors use the OpenAI shape, with the Nodus error code in both `type` and `code`:

|Status|Code|Meaning|
|-|-|-|
|400|`Unsupported`|A feature outside per-token billing: built-in or server-side tools, a non-default `service_tier`, `store`, `background`, file parts, image URLs or audio URLs, a key repeated in another letter case, or an operation the model does not serve|
|400|`Invalid`|A malformed request, or an upload that is not audio in a supported format|
|401|`Unauthorized`|Missing or invalid API key|
|402|`InsufficientCredits`, `BudgetExceeded`|Not enough credit for the hold, or a budget or endpoint cap blocks it|
|409|`RequestInProgress`|The same `Idempotency-Key` is still running; retry after `Retry-After`|
|409|`IdempotencyKeyReused`|The `Idempotency-Key` was used for a different request; send a new key|
|413|`RequestEntityTooLarge`|A JSON body over 1 MiB or an audio file over 25 MiB|
|429|`TooManyRequests`|An org or endpoint limit; retry after `Retry-After`|
|502, 503|`Unavailable`|The model failed, is at capacity or is not available; retry shortly. You are not charged|
