Connect Through the API
The optional OpenAI-compatible API connects your tools to CE’s active local model. It is disabled by default and runs on the existing CE application server.
Compatibility covers the endpoints below, not every endpoint or feature of another API provider.
Supported Endpoints
| Endpoint | Purpose |
|---|---|
GET /v1/models | Discover the active model exposed by CE. |
POST /v1/chat/completions | Streaming or non-streaming chat completion. |
Only the active model is served. Calls do not start, switch, or download models, administer CE, or manage Projects. Use the exact model ID returned by model discovery.
Enable Access
- An administrator opens Admin Settings → API Access.
- Generate the single CE API key and copy it immediately; plaintext is not displayed again. CE stores its SHA-256 hash.
- Enable API access, save, and load a compatible language model through the normal browser controls. Wait for readiness.
Requests use Authorization: Bearer <API_KEY>. This is a CE-issued key. Regeneration invalidates the previous key; revocation disables access and clears the stored hash. Keep keys out of source files, screenshots, and public logs.
Connect and Stream
Use the browser-visible CE application origin with /v1 as the client base URL, for example https://your-ce-host.example/v1. Do not connect clients to private llama-server ports. Use HTTPS when sending credentials across an untrusted network.
Supply the active model ID and your own messages. Set stream to true for streamed responses or false for a non-streaming completion. Personal System Instructions and Project Instructions are browser-chat preferences; they are not automatically injected into API requests.
Use the release-matched Bash and PowerShell examples and streaming example for exact request syntax.
Images and Limits
The API has a separate 16 MiB request-body limit, including encoded image data. Browser attachment settings do not change this boundary. Model context and available memory also constrain requests.
Inline image_url data is accepted only when the active runtime has a compatible projector. Remote image URLs are rejected. Supported image formats and conversion rules match browser image handling.
Activity and Troubleshooting
Administrator-only API Analytics reports counts, errors, durations, backend-provided usage, and recent request metadata. This tracking does not store prompts, responses, images, keys, or authorization headers. Independently configured client/proxy logs have their own behaviour.
| Symptom | Check first |
|---|---|
| Unauthorized | Access enabled, Bearer header present, and the key still current. |
| No active model | Load a compatible model in CE and wait for runtime readiness. |
| Model rejected | Use the ID returned by /v1/models; a registry filename or previously active ID may differ. |
| Image rejected | Active model/projector pair, supported inline data, and serialized body size. |
| Interrupted or failed response | Inspect runtime logs and API Analytics for backend errors or capacity problems. An interrupted response is not successful completion. |
Keep raw runtime ports private; exposing them is not a workaround for API authentication or configuration errors.
www.llmcontroller.comLLM Controller on GitHub