Read a InferenceEndpoint
const url = 'https://example.com/apis/nodus.dev/v1/namespaces/example/inferenceendpoints/example';const options = {method: 'GET'};
try { const response = await fetch(url, options); const data = await response.json(); console.log(data);} catch (error) { console.error(error);}curl --request GET \ --url https://example.com/apis/nodus.dev/v1/namespaces/example/inferenceendpoints/exampleParameters
Section titled “ Parameters ”Path Parameters
Section titled “ Path Parameters ”The project.
The object’s name.
Responses
Section titled “ Responses ”OK
InferenceEndpoint is a named access policy over one catalog model, with its own base URL /endpoints/<project>/<name>/v1 on the inference host. Requests through it, or with model: "endpoint/<name>" on the shared base URL, are attributed to its project, limited by spec.limits, accepted only from spec.allowedKeys and billed to it, so spec.maxCostUSD caps what it spends. It runs no compute: a stopped endpoint answers 503.
object
object
object
object
object
object
object
InferenceEndpointSpec describes an InferenceEndpoint. Every field but deployment can change after create, and maxCostUSD can only be raised.
object
AllowedKeys are the names of the API keys that may call the endpoint; empty allows every key of the org with the inference:invoke scope.
Deployment is reserved for self-hosted models and must be empty.
InferenceEndpointLimits is spec.limits.
object
MaxConcurrent is how many requests may be open at once, 1 to 256 (default 4).
RPM is the requests per minute, 1 to 100,000 (default 60).
TPM is the tokens per minute, at least 1; unset is unlimited.
MaxCostUSD caps what requests through the endpoint spend in total; once reached, new requests get 402 BudgetExceeded. It can only be raised.
Model is the catalog model the endpoint serves: its name, nodus/<name> or an alias, as GET /v1/models lists them.
State is the desired state: Running (the default) or Stopped, which answers every request with 503.
InferenceEndpointStatus is what Nodus observed about an InferenceEndpoint.
object
Conditions are Ready and ModelAvailable.
object
Message explains the phase in a sentence.
Model is the catalog name spec.model resolves to.
ObservedGeneration is the spec generation this status reflects.
Phase is Running or Stopped, from the long-running vocabulary.
Reason is the machine-readable reason for the phase.
InferenceEndpointUsage is the requests through an endpoint over a window.
object
CostUSD is what the requests were charged.
Errors counts the requests that ended with no charge.
InputTokens are the input tokens charged.
OpenHolds counts the requests still running, each holding its maximum cost.
OutputTokens are the output tokens charged, reasoning included.
Requests counts every request, answered or not.
Example generated
{ "apiVersion": "example", "kind": "example", "metadata": { "annotations": { "additionalProperty": "example" }, "creationTimestamp": "example", "deletionGracePeriodSeconds": 1, "deletionTimestamp": "example", "finalizers": [ "example" ], "generateName": "example", "generation": 1, "labels": { "additionalProperty": "example" }, "managedFields": [ { "apiVersion": "example", "fieldsType": "example", "fieldsV1": { "additionalProperty": "example" }, "manager": "example", "operation": "example", "subresource": "example", "time": "example" } ], "name": "example", "namespace": "example", "ownerReferences": [ { "apiVersion": "example", "blockOwnerDeletion": true, "controller": true, "kind": "example", "name": "example", "uid": "example" } ], "resourceVersion": "example", "selfLink": "example", "uid": "example" }, "spec": { "allowedKeys": [ "example" ], "deployment": "example", "limits": { "maxConcurrent": 1, "rpm": 1, "tpm": 1 }, "maxCostUSD": "example", "model": "example", "state": "example" }, "status": { "conditions": [ { "lastTransitionTime": "example", "message": "example", "observedGeneration": 1, "reason": "example", "status": "example", "type": "example" } ], "message": "example", "model": "example", "observedGeneration": 1, "phase": "example", "reason": "example", "usage24h": { "costUSD": "example", "errors": 1, "inputTokens": 1, "openHolds": 1, "outputTokens": 1, "requests": 1 } }}default
Section titled “ default ”An error: a metav1.Status whose reason is a registered code.
Status is the error body of every API response: a Kubernetes metav1.Status (so kubectl and client-go understand it) plus three top-level extensions that those clients ignore (ADR-028).
object
object
object
Docs is the URL of the code’s docs page.
Fix says what to do next, for example a CLI command or the field to change.
object
object
RequestID identifies the request in logs and support tickets.
Example generated
{ "apiVersion": "example", "code": 1, "details": { "causes": [ { "field": "example", "message": "example", "reason": "example" } ], "group": "example", "kind": "example", "name": "example", "retryAfterSeconds": 1, "uid": "example" }, "docs": "example", "fix": "example", "kind": "example", "message": "example", "metadata": { "continue": "example", "remainingItemCount": 1, "resourceVersion": "example", "selfLink": "example", "shardInfo": { "selector": "example" } }, "reason": "example", "requestId": "example", "status": "example"}