Skip to content

nodus inference chat

View Markdown

Send one chat completion and print the answer; the request id and charge go to stderr

chat sends PROMPT (or standard input, with - or no argument) as one user message and prints the answer. The request is held at its maximum cost (the input plus –max-tokens of output at the model’s rates) and charged for the tokens it used; the request id, the model that answered, the tokens and the charge are printed to stderr.

nodus inference chat [PROMPT | -] [--model MODEL | --endpoint NAME] [--max-tokens N] [flags]
nodus inference chat "What is the capital of France?"
nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 "Say hello"
nodus inference chat --endpoint support-bot < question.txt
nodus inference chat -o json "Hi" | jq .usage
--endpoint string Send through this InferenceEndpoint of the project instead of --model
-h, --help help for chat
--idempotency-key string Retry safely: a repeat with the same key and prompt is answered once and charged once
--max-tokens int Most output tokens; the request's hold is sized from it (default 512)
--model string Catalog model, as 'nodus inference models' lists them; nodus/indra (Indra) routes each request (default "nodus/indra")
-o, --output string Output format: json or yaml (the whole completion)
--system string System message
--base-url string Inference origin (default: derived from the API URL)
--context string Context from the config file to use
--org string Organization (selects the context for that org)
-p, --project string Project to work in
-v, --verbose Log each API request (never credentials)
  • nodus inference - Call hosted models: list them, send a chat completion and read a request’s receipt