Inference
View MarkdownThe inference data plane speaks the OpenAI and Anthropic wire formats. nodus.llm returns the official clients,
configured with your key and the Nodus base URL, so every feature of those SDKs works unchanged.
pip install "nodus-compute[openai]" # or [anthropic]"""Chat with a catalog model through the OpenAI-compatible inference data plane.
Needs `pip install "nodus-compute[openai]"`. Run it with `python examples/python/inference/chat.py`."""
import nodus
def main() -> None: client = nodus.llm.openai() # the official OpenAI client with your Nodus key and base URL reply = client.chat.completions.create( model="nodus/gpt-oss-20b", messages=[{"role": "user", "content": "Say hello in five words."}], max_tokens=32, ) print(reply.choices[0].message.content)
if __name__ == "__main__": main()client = nodus.llm.openai(project="nlp") # usage is attributed to the projectclaude = nodus.llm.anthropic() # messages APIaclient = nodus.llm.async_openai() # AsyncOpenAIInside a Job, Function or Sandbox, nodus.llm uses the container’s built-in proxy: no key is needed, and the calls
are billed to, and capped by, the run that makes them.
Named endpoints
Section titled “Named endpoints”An InferenceEndpoint gives a model its own base URL with rate limits, allowed keys and a spending cap.
ep = nodus.InferenceEndpoint.create("support-bot", "nodus/gpt-oss-120b", rpm=120, max_concurrent=8, max_cost=50)client = ep.openai()print(ep.usage()) # requests, tokens and cost over the last 24 hours