Skip to Content
Hosted APIOverview

API overview

Under maintenance. The managed Vinci API and Platform are under maintenance. These service-specific instructions are retained as reference for existing integrations when service is available. They do not promise current access or a return date. Downloadable models and direct/BYOK use of Vinci Code CLI are separate paths.

Vinci exposes an OpenAI-compatible HTTP API. If you’ve used the OpenAI API, you already know how to use Vinci — point your client at the Vinci base URL and use a Vinci key.

Base URL

https://vinci.getsimpledirect.com/api/v1

Authentication

Every request needs a Vinci API key. Send it as a Bearer token:

Authorization: Bearer vinci_live_...

An x-api-key: vinci_live_... header is also accepted if that’s easier for your client.

Issue keys on the Vinci Platform (same account as the Vinci app — same email and password). Keys are shown once and stored only as a hash. Revoke anytime; a revoked key stops working immediately. The Platform also shows your usage.

When you create a key, you may restrict it to one or more scopes: inference (call POST /chat/completions), models (GET /models and GET /models/{model}), and usage (read usage data). Selecting no scopes gives the key full access to all scopes; this is the default and keeps existing keys working. Scope enforcement is server-side: a request whose key lacks the required scope gets HTTP 403 with { "error": { "message": "...", "type": "permission_error" } }. For example, a key without inference cannot call /chat/completions.

A key may expire after 30, 90, or 365 days, or never expire (the default). An expired key is rejected exactly like a revoked key with HTTP 401 (authentication_error).

Keys are scoped to your account. Usage on a key counts against your account’s Vinci Platform credit balance, the same as the web app.

Models

The table below preserves the last documented hosted-service routing reference. Provider assignments, prices and limits are not current offers or availability claims. After maintenance, use your account and the service’s model-discovery response to confirm applicable options before sending traffic. These names are not MLE/Cyber download identifiers.

Model idNotes
autoDefault — resolves to the class on the account. Omit model entirely for the same result.
mezzoShown as Vinci Mezzo. Last documented: DeepSeek V4 Flash on DeepInfra.
forteShown as Vinci Forte — the default class. Last documented: GLM 5.2 on DeepInfra.
fortissimoShown as Vinci Fortissimo. Last documented: Kimi K3 on Fireworks.
pianoReserved in the documented routing scheme; no release date is stated.

The last documented API pricing per 1 million tokens was $0.10 input / $0.21 output / $0.02 cached input for Mezzo, $1.07 input / $3.45 output for Forte, and $3.45 input / $17.25 output for Fortissimo.

auto is the API default, and omitting model does the same thing. An unknown id is rejected with a validation error. auto resolves to the account preference, which is Forte unless changed. Class names stay stable while their occupant models can rotate after evaluation.

Discovering models

GET https://vinci.getsimpledirect.com/api/v1/models GET https://vinci.getsimpledirect.com/api/v1/models/{model}

OpenAI-compatible model listing, for clients that discover models automatically instead of taking a hard-coded id (AnythingLLM, LibreChat, and similar). Same auth as chat completions. GET /models returns { "object": "list", "data": [{ "id", "object": "model", "created", "owned_by" }] }; GET /models/{model} retrieves a single entry, or 404 if the id doesn’t exist.

Tool calling

Send a tools array on /chat/completions for native OpenAI-style function calling — real tool_calls in the response, streaming or not. See Tool calling for the request/response shape.

Streaming

Set "stream": true to receive Server-Sent Events of chat.completion.chunk objects, terminated by data: [DONE] — the same format as OpenAI. See Chat Completions.

Limits

These are the last documented service limits, not confirmed current quotas.

  • Model limits — up to 1,000,000 tokens of context and 131,072 output tokens per request.
  • Credits — usage is billed against your Vinci Platform credit balance. A 402 response means the account is out of credits.
  • Rate limits — per-key request-rate limits smooth out bursts; exceeding them returns 429 (rate_limit_error). Back off briefly and retry.
  • Concurrency — requests are admission-controlled; if the fleet is momentarily saturated you also get 429. Retry in a moment.

Zero Data Retention

The API does not persist your message content. Prompts and completions are not written to logs, analytics, or any datastore — only counts and metadata (model, token totals, status) are recorded for metering. Model inference is served by DeepInfra or Fireworks, depending on the class, using Canadian and US infrastructure; Zero Data Retention is enforced with both providers. Account data and usage counters are stored in Canada (ca-central-1). See Privacy & data retention for the full picture, including the opt-in features that do store data you create.

The Vinci character

The hosted reference describes server-side character instructions, with client system messages treated as additional guidance. This is intended behaviour, not a guarantee against prompt injection or incorrect output. See The Vinci character.

Errors

Errors use OpenAI’s shape: { "error": { "message": ..., "type": ... } }. See the error reference.

Last updated on