Chat Completions
Under maintenance. The managed Vinci API and Platform are under maintenance. These service-specific instructions are retained as reference for existing integrations when service is available. They do not promise current access or a return date. Downloadable models and direct/BYOK use of Vinci Code CLI are separate paths.
POST https://vinci.getsimpledirect.com/api/v1/chat/completionsGenerate a model response for a conversation. OpenAI-compatible.
Request body
| Field | Type | Required | Notes |
|---|---|---|---|
model | string | no | A Vinci class id — auto (the default) uses the class on the account, or pin mezzo / forte / fortissimo. Unknown ids are rejected. |
messages | array | yes | { role, content } items. role is system, user, assistant, or tool. content can be a plain string or an OpenAI-style array of parts ([{ "type": "text", "text": "…" }]) — both are accepted. |
stream | boolean | no | true to stream chunks. Defaults to false. |
temperature | number | no | Sampling temperature. |
max_tokens | number | no | Max tokens to generate. |
tools | array | no | OpenAI-style function/tool definitions. See Tool calling below. |
tool_choice | string | object | no | "auto" (default when tools is set), "none", or { "type": "function", "function": { "name": "…" } } to force a specific tool. |
System messages. A
systemmessage you send is treated as additional instructions layered on top of the Vinci character — it can shape the response but is not a guarantee of behaviour or an enforcement boundary.
Response — non-streaming
{
"id": "chatcmpl-…",
"object": "chat.completion",
"created": 1781981326,
"model": "forte",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "…" },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 42,
"completion_tokens": 18,
"total_tokens": 60,
"prompt_tokens_details": { "cached_tokens": 0 }
}
}The OpenAI-compatible prompt_tokens_details.cached_tokens field reports the number
of cached prompt tokens. Its value is 0 when no cached tokens are present. The
existing prompt_tokens, completion_tokens, and total_tokens fields are unchanged.
Response — streaming
With "stream": true, the response is text/event-stream of
chat.completion.chunk objects:
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":42,"completion_tokens":18,"total_tokens":60,"prompt_tokens_details":{"cached_tokens":0}}}
data: [DONE]Tool calling
Send a non-empty tools array and the request routes to a dedicated tool-calling path:
your tools / tool_choice are forwarded straight through and real tool_calls come
back, the same wire format as OpenAI’s function calling. This is the wire format coding
agents use under the hood. It’s also how an MCP host drives its tools with Vinci as
the model — see
Use Vinci with MCP tools.
A few things specific to this path:
- Streaming and non-streaming both work, with
tool_callsdeltas in the streamed chunks the same way OpenAI streams them. - Multi-turn tool use is supported — send the assistant’s
tool_callsback inmessages, followed by one{ "role": "tool", "tool_call_id": "...", "content": "..." }message per result, and continue the conversation. max_tokensis fit automatically to the model’s context window, accounting for your prompt and tool definitions, so large tool schemas don’t cause a hard failure.- This is a direct pass-through for agentic clients — the grounding features on the plain chat path (auto web search, memory, attachments) don’t apply here.
Example:
curl https://vinci.getsimpledirect.com/api/v1/chat/completions \
-H "Authorization: Bearer $VINCI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "forte",
"messages": [{ "role": "user", "content": "What is the weather in Ottawa?" }],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a location.",
"parameters": {
"type": "object",
"properties": { "location": { "type": "string" } },
"required": ["location"]
}
}
}]
}'Response:
{
"id": "chatcmpl-…",
"object": "chat.completion",
"model": "forte",
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [{
"id": "call_…",
"type": "function",
"function": { "name": "get_weather", "arguments": "{\"location\":\"Ottawa\"}" }
}]
},
"finish_reason": "tool_calls"
}],
"usage": {
"prompt_tokens": 61,
"completion_tokens": 14,
"total_tokens": 75,
"prompt_tokens_details": { "cached_tokens": 0 }
}
}Run the function yourself and send the result back as a tool message to continue.
Examples
curl (Mezzo for light everyday chat):
curl https://vinci.getsimpledirect.com/api/v1/chat/completions \
-H "Authorization: Bearer $VINCI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mezzo",
"messages": [{ "role": "user", "content": "Suggest three names for a book club." }]
}'For code and file tasks, use forte.
curl (streaming):
curl -N https://vinci.getsimpledirect.com/api/v1/chat/completions \
-H "Authorization: Bearer $VINCI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "forte",
"stream": true,
"messages": [{ "role": "user", "content": "Write a haiku about Canada." }]
}'Python (streaming):
from openai import OpenAI
client = OpenAI(
base_url="https://vinci.getsimpledirect.com/api/v1",
api_key="vinci_live_...",
)
stream = client.chat.completions.create(
model="forte",
messages=[{"role": "user", "content": "Write a haiku about Canada."}],
stream=True,
)
for chunk in stream:
if chunk.choices: # usage-only chunks may have no choices
print(chunk.choices[0].delta.content or "", end="")Errors
OpenAI-shaped: { "error": { "message": "...", "type": "..." } }.
| Status | type | Meaning |
|---|---|---|
400 | invalid_request_error | Malformed JSON, or messages missing/empty. |
401 | authentication_error | Missing, invalid, revoked, or expired API key. |
402 | insufficient_quota | Vinci Platform credit balance is exhausted. |
403 | permission_error | The API key lacks the required scope. |
429 | rate_limit_error | Rate-limited, or the fleet is momentarily saturated — back off and retry. |
502 | api_error | Upstream model error. |