Skip to content

nodus create inferenceendpoint

View Markdown

Create a inferenceendpoint

nodus create inferenceendpoint NAME [flags]
nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-cost 5
--allowed-key stringArray API key name allowed to call it (repeatable)
--dry-run string[="server"] server returns the estimate without creating; client prints the manifest (default "none")
-h, --help help for inferenceendpoint
--max-concurrent int Open requests
--max-cost string Spend limit in USD
--model string Catalog model, as 'nodus inference models' lists them
-o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)
--rpm int Requests per minute
--tpm int Tokens per minute
--context string Context from the config file to use
--org string Organization (selects the context for that org)
-p, --project string Project to work in
-v, --verbose Log each API request (never credentials)
  • nodus create - Create resources from manifests or with a typed generator