Try a request, Try Indra and each model card on the Inference page open a playground. Choose a model, write
a system prompt and a message, set the temperature, max tokens and stream mode, and send it. The answer streams in
as it is generated.
Each answer ends with the model that served it (for Indra, the routed model and the free or paid tier), the
tokens used, the latency, the charge and the request id of its receipt.
Requests run with your console session and are billed to the organization, so you do not paste an API key into
the browser. Sending needs the Member role or above. The curl, Python and TypeScript tabs show the same request
for a terminal, where you use an API key with the inference:invoke scope.