Responses, Messages and embeddings on the inference API
Chat models accept stateless Responses, text Completions and Anthropic Messages requests, including streaming
and client function tools. Named endpoints apply their existing access rules and spending caps to these routes.
BGE-M3 embeddings charge input tokens only and accept batches of up to 2,048 inputs when the model is available.
Retrying a completed non-streamed request with its idempotency key replays the same answer without a second
charge. Stored Responses, hosted tools and Messages token counting remain unsupported.
Completed chat streams include the OpenAI terminal marker even when the upstream closes after its final usage. Interrupted streams retain their incomplete state.