Connect Through the API

The optional OpenAI-compatible API connects your tools to CE’s active local model. It is disabled by default and runs on the existing CE application server.

Compatibility covers the endpoints below, not every endpoint or feature of another API provider.

Supported Endpoints

EndpointPurpose
GET /v1/modelsDiscover the active model exposed by CE.
POST /v1/chat/completionsStreaming or non-streaming chat completion.

Only the active model is served. Calls do not start, switch, or download models, administer CE, or manage Projects. Use the exact model ID returned by model discovery.

Enable Access

  1. An administrator opens Admin Settings → API Access.
  2. Generate the single CE API key and copy it immediately; plaintext is not displayed again. CE stores its SHA-256 hash.
  3. Enable API access, save, and load a compatible language model through the normal browser controls. Wait for readiness.

Requests use Authorization: Bearer <API_KEY>. This is a CE-issued key. Regeneration invalidates the previous key; revocation disables access and clears the stored hash. Keep keys out of source files, screenshots, and public logs.

Connect and Stream

Use the browser-visible CE application origin with /v1 as the client base URL, for example https://your-ce-host.example/v1. Do not connect clients to private llama-server ports. Use HTTPS when sending credentials across an untrusted network.

Supply the active model ID and your own messages. Set stream to true for streamed responses or false for a non-streaming completion. Personal System Instructions and Project Instructions are browser-chat preferences; they are not automatically injected into API requests.

Use the release-matched Bash and PowerShell examples and streaming example for exact request syntax.

Images and Limits

The API has a separate 16 MiB request-body limit, including encoded image data. Browser attachment settings do not change this boundary. Model context and available memory also constrain requests.

Inline image_url data is accepted only when the active runtime has a compatible projector. Remote image URLs are rejected. Supported image formats and conversion rules match browser image handling.

Activity and Troubleshooting

Administrator-only API Analytics reports counts, errors, durations, backend-provided usage, and recent request metadata. This tracking does not store prompts, responses, images, keys, or authorization headers. Independently configured client/proxy logs have their own behaviour.

SymptomCheck first
UnauthorizedAccess enabled, Bearer header present, and the key still current.
No active modelLoad a compatible model in CE and wait for runtime readiness.
Model rejectedUse the ID returned by /v1/models; a registry filename or previously active ID may differ.
Image rejectedActive model/projector pair, supported inline data, and serialized body size.
Interrupted or failed responseInspect runtime logs and API Analytics for backend errors or capacity problems. An interrupted response is not successful completion.

Keep raw runtime ports private; exposing them is not a workaround for API authentication or configuration errors.