Skip to Content
Hosted APIChat Completions

Chat Completions

Under maintenance. The managed Vinci API and Platform are under maintenance. These service-specific instructions are retained as reference for existing integrations when service is available. They do not promise current access or a return date. Downloadable models and direct/BYOK use of Vinci Code CLI are separate paths.

POST https://vinci.getsimpledirect.com/api/v1/chat/completions

Generate a model response for a conversation. OpenAI-compatible.

Request body

FieldTypeRequiredNotes
modelstringnoA Vinci class id — auto (the default) uses the class on the account, or pin mezzo / forte / fortissimo. Unknown ids are rejected.
messagesarrayyes{ role, content } items. role is system, user, assistant, or tool. content can be a plain string or an OpenAI-style array of parts ([{ "type": "text", "text": "…" }]) — both are accepted.
streambooleannotrue to stream chunks. Defaults to false.
temperaturenumbernoSampling temperature.
max_tokensnumbernoMax tokens to generate.
toolsarraynoOpenAI-style function/tool definitions. See Tool calling below.
tool_choicestring | objectno"auto" (default when tools is set), "none", or { "type": "function", "function": { "name": "…" } } to force a specific tool.

System messages. A system message you send is treated as additional instructions layered on top of the Vinci character — it can shape the response but is not a guarantee of behaviour or an enforcement boundary.

Response — non-streaming

{ "id": "chatcmpl-…", "object": "chat.completion", "created": 1781981326, "model": "forte", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 42, "completion_tokens": 18, "total_tokens": 60, "prompt_tokens_details": { "cached_tokens": 0 } } }

The OpenAI-compatible prompt_tokens_details.cached_tokens field reports the number of cached prompt tokens. Its value is 0 when no cached tokens are present. The existing prompt_tokens, completion_tokens, and total_tokens fields are unchanged.

Response — streaming

With "stream": true, the response is text/event-stream of chat.completion.chunk objects:

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]} data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]} data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]} data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":42,"completion_tokens":18,"total_tokens":60,"prompt_tokens_details":{"cached_tokens":0}}} data: [DONE]

Tool calling

Send a non-empty tools array and the request routes to a dedicated tool-calling path: your tools / tool_choice are forwarded straight through and real tool_calls come back, the same wire format as OpenAI’s function calling. This is the wire format coding agents use under the hood. It’s also how an MCP host drives its tools with Vinci as the model — see Use Vinci with MCP tools.

A few things specific to this path:

  • Streaming and non-streaming both work, with tool_calls deltas in the streamed chunks the same way OpenAI streams them.
  • Multi-turn tool use is supported — send the assistant’s tool_calls back in messages, followed by one { "role": "tool", "tool_call_id": "...", "content": "..." } message per result, and continue the conversation.
  • max_tokens is fit automatically to the model’s context window, accounting for your prompt and tool definitions, so large tool schemas don’t cause a hard failure.
  • This is a direct pass-through for agentic clients — the grounding features on the plain chat path (auto web search, memory, attachments) don’t apply here.

Example:

curl https://vinci.getsimpledirect.com/api/v1/chat/completions \ -H "Authorization: Bearer $VINCI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "forte", "messages": [{ "role": "user", "content": "What is the weather in Ottawa?" }], "tools": [{ "type": "function", "function": { "name": "get_weather", "description": "Get the current weather for a location.", "parameters": { "type": "object", "properties": { "location": { "type": "string" } }, "required": ["location"] } } }] }'

Response:

{ "id": "chatcmpl-…", "object": "chat.completion", "model": "forte", "choices": [{ "index": 0, "message": { "role": "assistant", "content": null, "tool_calls": [{ "id": "call_…", "type": "function", "function": { "name": "get_weather", "arguments": "{\"location\":\"Ottawa\"}" } }] }, "finish_reason": "tool_calls" }], "usage": { "prompt_tokens": 61, "completion_tokens": 14, "total_tokens": 75, "prompt_tokens_details": { "cached_tokens": 0 } } }

Run the function yourself and send the result back as a tool message to continue.

Examples

curl (Mezzo for light everyday chat):

curl https://vinci.getsimpledirect.com/api/v1/chat/completions \ -H "Authorization: Bearer $VINCI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "mezzo", "messages": [{ "role": "user", "content": "Suggest three names for a book club." }] }'

For code and file tasks, use forte.

curl (streaming):

curl -N https://vinci.getsimpledirect.com/api/v1/chat/completions \ -H "Authorization: Bearer $VINCI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "forte", "stream": true, "messages": [{ "role": "user", "content": "Write a haiku about Canada." }] }'

Python (streaming):

from openai import OpenAI client = OpenAI( base_url="https://vinci.getsimpledirect.com/api/v1", api_key="vinci_live_...", ) stream = client.chat.completions.create( model="forte", messages=[{"role": "user", "content": "Write a haiku about Canada."}], stream=True, ) for chunk in stream: if chunk.choices: # usage-only chunks may have no choices print(chunk.choices[0].delta.content or "", end="")

Errors

OpenAI-shaped: { "error": { "message": "...", "type": "..." } }.

StatustypeMeaning
400invalid_request_errorMalformed JSON, or messages missing/empty.
401authentication_errorMissing, invalid, revoked, or expired API key.
402insufficient_quotaVinci Platform credit balance is exhausted.
403permission_errorThe API key lacks the required scope.
429rate_limit_errorRate-limited, or the fleet is momentarily saturated — back off and retry.
502api_errorUpstream model error.
Last updated on