Skip to content

Create a InferenceEndpoint

POST
/apis/nodus.dev/v1/namespaces/{namespace}/inferenceendpoints
curl --request POST \
--url https://example.com/apis/nodus.dev/v1/namespaces/example/inferenceendpoints \
--header 'Content-Type: application/json' \
--data '{ "apiVersion": "example", "kind": "example", "metadata": { "annotations": { "additionalProperty": "example" }, "creationTimestamp": "example", "deletionGracePeriodSeconds": 1, "deletionTimestamp": "example", "finalizers": [ "example" ], "generateName": "example", "generation": 1, "labels": { "additionalProperty": "example" }, "managedFields": [ { "apiVersion": "example", "fieldsType": "example", "fieldsV1": { "additionalProperty": "example" }, "manager": "example", "operation": "example", "subresource": "example", "time": "example" } ], "name": "example", "namespace": "example", "ownerReferences": [ { "apiVersion": "example", "blockOwnerDeletion": true, "controller": true, "kind": "example", "name": "example", "uid": "example" } ], "resourceVersion": "example", "selfLink": "example", "uid": "example" }, "spec": { "allowedKeys": [ "example" ], "deployment": "example", "limits": { "maxConcurrent": 1, "rpm": 1, "tpm": 1 }, "maxCostUSD": "example", "model": "example", "state": "example" }, "status": { "conditions": [ { "lastTransitionTime": "example", "message": "example", "observedGeneration": 1, "reason": "example", "status": "example", "type": "example" } ], "message": "example", "model": "example", "observedGeneration": 1, "phase": "example", "reason": "example", "usage24h": { "costUSD": "example", "errors": 1, "inputTokens": 1, "openHolds": 1, "outputTokens": 1, "requests": 1 } } }'
namespace
required
string

The project.

dryRun
string
fieldValidation
string
fieldManager
string
Idempotency-Key
string
If-Match
string
Media type application/json

InferenceEndpoint is a named access policy over one catalog model, with its own base URL /endpoints/<project>/<name>/v1 on the inference host. Requests through it, or with model: "endpoint/<name>" on the shared base URL, are attributed to its project, limited by spec.limits, accepted only from spec.allowedKeys and billed to it, so spec.maxCostUSD caps what it spends. It runs no compute: a stopped endpoint answers 503.

object
apiVersion
string
kind
string
metadata
object
annotations
object
key
additional properties
string
creationTimestamp
string
deletionGracePeriodSeconds
integer format: int64
deletionTimestamp
string
finalizers
Array<string>
nullable
generateName
string
generation
integer format: int64
labels
object
key
additional properties
string
managedFields
Array<object>
nullable
object
apiVersion
string
fieldsType
string
fieldsV1
object
key
additional properties
manager
string
operation
string
subresource
string
time
string
name
string
namespace
string
ownerReferences
Array<object>
nullable
object
apiVersion
required
string
blockOwnerDeletion
boolean
controller
boolean
kind
required
string
name
required
string
uid
required
string
resourceVersion
string
selfLink
string
uid
string
spec
required

InferenceEndpointSpec describes an InferenceEndpoint. Every field but deployment can change after create, and maxCostUSD can only be raised.

object
allowedKeys

AllowedKeys are the names of the API keys that may call the endpoint; empty allows every key of the org with the inference:invoke scope.

Array<string>
nullable
deployment

Deployment is reserved for self-hosted models and must be empty.

string
limits

InferenceEndpointLimits is spec.limits.

object
maxConcurrent

MaxConcurrent is how many requests may be open at once, 1 to 256 (default 4).

integer format: int64
rpm

RPM is the requests per minute, 1 to 100,000 (default 60).

integer format: int64
tpm

TPM is the tokens per minute, at least 1; unset is unlimited.

integer format: int64
maxCostUSD

MaxCostUSD caps what requests through the endpoint spend in total; once reached, new requests get 402 BudgetExceeded. It can only be raised.

string
model
required

Model is the catalog model the endpoint serves: its name, nodus/<name> or an alias, as GET /v1/models lists them.

string
state

State is the desired state: Running (the default) or Stopped, which answers every request with 503.

string
status

InferenceEndpointStatus is what Nodus observed about an InferenceEndpoint.

object
conditions

Conditions are Ready and ModelAvailable.

Array<object>
nullable
object
lastTransitionTime
required
string
message
required
string
observedGeneration
integer format: int64
reason
required
string
status
required
string
type
required
string
message

Message explains the phase in a sentence.

string
model

Model is the catalog name spec.model resolves to.

string
observedGeneration

ObservedGeneration is the spec generation this status reflects.

integer format: int64
phase

Phase is Running or Stopped, from the long-running vocabulary.

string
reason

Reason is the machine-readable reason for the phase.

string
usage24h

InferenceEndpointUsage is the requests through an endpoint over a window.

object
costUSD
required

CostUSD is what the requests were charged.

string
errors
required

Errors counts the requests that ended with no charge.

integer format: int64
inputTokens
required

InputTokens are the input tokens charged.

integer format: int64
openHolds
required

OpenHolds counts the requests still running, each holding its maximum cost.

integer format: int64
outputTokens
required

OutputTokens are the output tokens charged, reasoning included.

integer format: int64
requests
required

Requests counts every request, answered or not.

integer format: int64
Example generated
{
"apiVersion": "example",
"kind": "example",
"metadata": {
"annotations": {
"additionalProperty": "example"
},
"creationTimestamp": "example",
"deletionGracePeriodSeconds": 1,
"deletionTimestamp": "example",
"finalizers": [
"example"
],
"generateName": "example",
"generation": 1,
"labels": {
"additionalProperty": "example"
},
"managedFields": [
{
"apiVersion": "example",
"fieldsType": "example",
"fieldsV1": {
"additionalProperty": "example"
},
"manager": "example",
"operation": "example",
"subresource": "example",
"time": "example"
}
],
"name": "example",
"namespace": "example",
"ownerReferences": [
{
"apiVersion": "example",
"blockOwnerDeletion": true,
"controller": true,
"kind": "example",
"name": "example",
"uid": "example"
}
],
"resourceVersion": "example",
"selfLink": "example",
"uid": "example"
},
"spec": {
"allowedKeys": [
"example"
],
"deployment": "example",
"limits": {
"maxConcurrent": 1,
"rpm": 1,
"tpm": 1
},
"maxCostUSD": "example",
"model": "example",
"state": "example"
},
"status": {
"conditions": [
{
"lastTransitionTime": "example",
"message": "example",
"observedGeneration": 1,
"reason": "example",
"status": "example",
"type": "example"
}
],
"message": "example",
"model": "example",
"observedGeneration": 1,
"phase": "example",
"reason": "example",
"usage24h": {
"costUSD": "example",
"errors": 1,
"inputTokens": 1,
"openHolds": 1,
"outputTokens": 1,
"requests": 1
}
}
}

OK

Media type application/json

InferenceEndpoint is a named access policy over one catalog model, with its own base URL /endpoints/<project>/<name>/v1 on the inference host. Requests through it, or with model: "endpoint/<name>" on the shared base URL, are attributed to its project, limited by spec.limits, accepted only from spec.allowedKeys and billed to it, so spec.maxCostUSD caps what it spends. It runs no compute: a stopped endpoint answers 503.

object
apiVersion
string
kind
string
metadata
object
annotations
object
key
additional properties
string
creationTimestamp
string
deletionGracePeriodSeconds
integer format: int64
deletionTimestamp
string
finalizers
Array<string>
nullable
generateName
string
generation
integer format: int64
labels
object
key
additional properties
string
managedFields
Array<object>
nullable
object
apiVersion
string
fieldsType
string
fieldsV1
object
key
additional properties
manager
string
operation
string
subresource
string
time
string
name
string
namespace
string
ownerReferences
Array<object>
nullable
object
apiVersion
required
string
blockOwnerDeletion
boolean
controller
boolean
kind
required
string
name
required
string
uid
required
string
resourceVersion
string
selfLink
string
uid
string
spec
required

InferenceEndpointSpec describes an InferenceEndpoint. Every field but deployment can change after create, and maxCostUSD can only be raised.

object
allowedKeys

AllowedKeys are the names of the API keys that may call the endpoint; empty allows every key of the org with the inference:invoke scope.

Array<string>
nullable
deployment

Deployment is reserved for self-hosted models and must be empty.

string
limits

InferenceEndpointLimits is spec.limits.

object
maxConcurrent

MaxConcurrent is how many requests may be open at once, 1 to 256 (default 4).

integer format: int64
rpm

RPM is the requests per minute, 1 to 100,000 (default 60).

integer format: int64
tpm

TPM is the tokens per minute, at least 1; unset is unlimited.

integer format: int64
maxCostUSD

MaxCostUSD caps what requests through the endpoint spend in total; once reached, new requests get 402 BudgetExceeded. It can only be raised.

string
model
required

Model is the catalog model the endpoint serves: its name, nodus/<name> or an alias, as GET /v1/models lists them.

string
state

State is the desired state: Running (the default) or Stopped, which answers every request with 503.

string
status

InferenceEndpointStatus is what Nodus observed about an InferenceEndpoint.

object
conditions

Conditions are Ready and ModelAvailable.

Array<object>
nullable
object
lastTransitionTime
required
string
message
required
string
observedGeneration
integer format: int64
reason
required
string
status
required
string
type
required
string
message

Message explains the phase in a sentence.

string
model

Model is the catalog name spec.model resolves to.

string
observedGeneration

ObservedGeneration is the spec generation this status reflects.

integer format: int64
phase

Phase is Running or Stopped, from the long-running vocabulary.

string
reason

Reason is the machine-readable reason for the phase.

string
usage24h

InferenceEndpointUsage is the requests through an endpoint over a window.

object
costUSD
required

CostUSD is what the requests were charged.

string
errors
required

Errors counts the requests that ended with no charge.

integer format: int64
inputTokens
required

InputTokens are the input tokens charged.

integer format: int64
openHolds
required

OpenHolds counts the requests still running, each holding its maximum cost.

integer format: int64
outputTokens
required

OutputTokens are the output tokens charged, reasoning included.

integer format: int64
requests
required

Requests counts every request, answered or not.

integer format: int64
Example generated
{
"apiVersion": "example",
"kind": "example",
"metadata": {
"annotations": {
"additionalProperty": "example"
},
"creationTimestamp": "example",
"deletionGracePeriodSeconds": 1,
"deletionTimestamp": "example",
"finalizers": [
"example"
],
"generateName": "example",
"generation": 1,
"labels": {
"additionalProperty": "example"
},
"managedFields": [
{
"apiVersion": "example",
"fieldsType": "example",
"fieldsV1": {
"additionalProperty": "example"
},
"manager": "example",
"operation": "example",
"subresource": "example",
"time": "example"
}
],
"name": "example",
"namespace": "example",
"ownerReferences": [
{
"apiVersion": "example",
"blockOwnerDeletion": true,
"controller": true,
"kind": "example",
"name": "example",
"uid": "example"
}
],
"resourceVersion": "example",
"selfLink": "example",
"uid": "example"
},
"spec": {
"allowedKeys": [
"example"
],
"deployment": "example",
"limits": {
"maxConcurrent": 1,
"rpm": 1,
"tpm": 1
},
"maxCostUSD": "example",
"model": "example",
"state": "example"
},
"status": {
"conditions": [
{
"lastTransitionTime": "example",
"message": "example",
"observedGeneration": 1,
"reason": "example",
"status": "example",
"type": "example"
}
],
"message": "example",
"model": "example",
"observedGeneration": 1,
"phase": "example",
"reason": "example",
"usage24h": {
"costUSD": "example",
"errors": 1,
"inputTokens": 1,
"openHolds": 1,
"outputTokens": 1,
"requests": 1
}
}
}

An error: a metav1.Status whose reason is a registered code.

Media type application/json

Status is the error body of every API response: a Kubernetes metav1.Status (so kubectl and client-go understand it) plus three top-level extensions that those clients ignore (ADR-028).

object
apiVersion
string
code
integer format: int32
details
object
causes
Array<object>
nullable
object
field
string
message
string
reason
string
group
string
kind
string
name
string
retryAfterSeconds
integer format: int32
uid
string
docs

Docs is the URL of the code’s docs page.

string
fix

Fix says what to do next, for example a CLI command or the field to change.

string
kind
string
message
string
metadata
object
continue
string
remainingItemCount
integer format: int64
resourceVersion
string
selfLink
string
shardInfo
object
selector
required
string
reason
string
requestId

RequestID identifies the request in logs and support tickets.

string
status
string
Example generated
{
"apiVersion": "example",
"code": 1,
"details": {
"causes": [
{
"field": "example",
"message": "example",
"reason": "example"
}
],
"group": "example",
"kind": "example",
"name": "example",
"retryAfterSeconds": 1,
"uid": "example"
},
"docs": "example",
"fix": "example",
"kind": "example",
"message": "example",
"metadata": {
"continue": "example",
"remainingItemCount": 1,
"resourceVersion": "example",
"selfLink": "example",
"shardInfo": {
"selector": "example"
}
},
"reason": "example",
"requestId": "example",
"status": "example"
}