nodus inference chat
View Markdownnodus inference chat
Section titled “nodus inference chat”Send one chat completion and print the answer; the request id and charge go to stderr
Synopsis
Section titled “Synopsis”chat sends PROMPT (or standard input, with - or no argument) as one user message and prints the answer. The request is held at its maximum cost (the input plus –max-tokens of output at the model’s rates) and charged for the tokens it used; the request id, the model that answered, the tokens and the charge are printed to stderr.
nodus inference chat [PROMPT | -] [--model MODEL | --endpoint NAME] [--max-tokens N] [flags]Examples
Section titled “Examples” nodus inference chat "What is the capital of France?" nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 "Say hello" nodus inference chat --endpoint support-bot < question.txt nodus inference chat -o json "Hi" | jq .usageOptions
Section titled “Options” --endpoint string Send through this InferenceEndpoint of the project instead of --model -h, --help help for chat --idempotency-key string Retry safely: a repeat with the same key and prompt is answered once and charged once --max-tokens int Most output tokens; the request's hold is sized from it (default 512) --model string Catalog model, as 'nodus inference models' lists them; nodus/indra (Indra) routes each request (default "nodus/indra") -o, --output string Output format: json or yaml (the whole completion) --system string System messageOptions inherited from parent commands
Section titled “Options inherited from parent commands” --base-url string Inference origin (default: derived from the API URL) --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials)SEE ALSO
Section titled “SEE ALSO”- nodus inference - Call hosted models: list them, send a chat completion and read a request’s receipt