API overview
Under maintenance. The managed Vinci API and Platform are under maintenance. These service-specific instructions are retained as reference for existing integrations when service is available. They do not promise current access or a return date. Downloadable models and direct/BYOK use of Vinci Code CLI are separate paths.
Vinci exposes an OpenAI-compatible HTTP API. If you’ve used the OpenAI API, you already know how to use Vinci — point your client at the Vinci base URL and use a Vinci key.
Base URL
https://vinci.getsimpledirect.com/api/v1Authentication
Every request needs a Vinci API key. Send it as a Bearer token:
Authorization: Bearer vinci_live_...An x-api-key: vinci_live_... header is also accepted if that’s easier for your client.
Issue keys on the Vinci Platform (same account as the Vinci app — same email and password). Keys are shown once and stored only as a hash. Revoke anytime; a revoked key stops working immediately. The Platform also shows your usage.
When you create a key, you may restrict it to one or more scopes: inference (call
POST /chat/completions), models (GET /models and GET /models/{model}), and
usage (read usage data). Selecting no scopes gives the key full access to all scopes;
this is the default and keeps existing keys working. Scope enforcement is server-side:
a request whose key lacks the required scope gets HTTP 403 with
{ "error": { "message": "...", "type": "permission_error" } }. For example, a key
without inference cannot call /chat/completions.
A key may expire after 30, 90, or 365 days, or never expire (the default). An expired
key is rejected exactly like a revoked key with HTTP 401 (authentication_error).
Keys are scoped to your account. Usage on a key counts against your account’s Vinci Platform credit balance, the same as the web app.
Models
The table below preserves the last documented hosted-service routing reference. Provider assignments, prices and limits are not current offers or availability claims. After maintenance, use your account and the service’s model-discovery response to confirm applicable options before sending traffic. These names are not MLE/Cyber download identifiers.
| Model id | Notes |
|---|---|
auto | Default — resolves to the class on the account. Omit model entirely for the same result. |
mezzo | Shown as Vinci Mezzo. Last documented: DeepSeek V4 Flash on DeepInfra. |
forte | Shown as Vinci Forte — the default class. Last documented: GLM 5.2 on DeepInfra. |
fortissimo | Shown as Vinci Fortissimo. Last documented: Kimi K3 on Fireworks. |
piano | Reserved in the documented routing scheme; no release date is stated. |
The last documented API pricing per 1 million tokens was $0.10 input / $0.21 output / $0.02 cached input for Mezzo, $1.07 input / $3.45 output for Forte, and $3.45 input / $17.25 output for Fortissimo.
auto is the API default, and omitting model does the same thing. An unknown id is
rejected with a validation error. auto resolves to the account preference, which is
Forte unless changed. Class names stay stable while their occupant models can rotate after
evaluation.
Discovering models
GET https://vinci.getsimpledirect.com/api/v1/models
GET https://vinci.getsimpledirect.com/api/v1/models/{model}OpenAI-compatible model listing, for clients that discover models automatically instead
of taking a hard-coded id (AnythingLLM, LibreChat, and similar). Same auth as chat
completions. GET /models returns { "object": "list", "data": [{ "id", "object": "model", "created", "owned_by" }] };
GET /models/{model} retrieves a single entry, or 404 if the id doesn’t exist.
Tool calling
Send a tools array on /chat/completions for native OpenAI-style function calling —
real tool_calls in the response, streaming or not. See
Tool calling for the request/response shape.
Streaming
Set "stream": true to receive Server-Sent Events of chat.completion.chunk objects,
terminated by data: [DONE] — the same format as OpenAI. See
Chat Completions.
Limits
These are the last documented service limits, not confirmed current quotas.
- Model limits — up to 1,000,000 tokens of context and 131,072 output tokens per request.
- Credits — usage is billed against your Vinci Platform credit balance. A
402response means the account is out of credits. - Rate limits — per-key request-rate limits smooth out bursts; exceeding them returns
429(rate_limit_error). Back off briefly and retry. - Concurrency — requests are admission-controlled; if the fleet is momentarily saturated
you also get
429. Retry in a moment.
Zero Data Retention
The API does not persist your message content. Prompts and completions are not
written to logs, analytics, or any datastore — only counts and metadata (model, token
totals, status) are recorded for metering. Model inference is served by DeepInfra or
Fireworks, depending on the class, using Canadian and US infrastructure; Zero Data
Retention is enforced with both providers. Account data and usage counters are stored in
Canada (ca-central-1). See Privacy & data retention for the full picture,
including the opt-in features that do store data you create.
The Vinci character
The hosted reference describes server-side character instructions, with client
system messages treated as additional guidance. This is intended behaviour, not a
guarantee against prompt injection or incorrect output. See The Vinci character.
Errors
Errors use OpenAI’s shape: { "error": { "message": ..., "type": ... } }. See the
error reference.