Local Voice Dictation

Optional local Speech-to-Text turns a recording into text in the chat composer. Review or edit it, then click Send. It does not automatically send a message or speak the assistant’s answer.

Normal typed chat works without a speech environment, model, or running speech service.

Supported Setup

The validated v1.4 setup is Windows with NVIDIA CUDA, a separate NeMo Python environment, and NVIDIA Nemotron 3.5 ASR Streaming 0.6B as a local .nemo checkpoint. Other models must be compatible with CE’s adapter; arbitrary speech executables are not interchangeable.

Use the release-matched Voice installation guide for pinned dependencies, model download, and environment verification. Keep speech dependencies separate from CE’s application environment. AMD GPU telemetry support is not an AMD speech-support claim.

Configure and Designate

  1. Complete normal CE installation/upgrade, then prepare and verify the dedicated speech environment using the linked guide.
  2. In Admin Settings, set the speech runtime path to that environment’s Python executable, not the model file or an activation command. Choose an available speech port, separate from application/main/title ports; the default is 8082.
  3. Place the compatible checkpoint inside Models DIR and finish copying it before Rescan Models. Enable its registry row and explicitly designate it Speech-to-Text, not MMPROJ.

Speech checkpoints are excluded from language-model, title-model, and benchmark choices. The speech service binds to loopback; browsers access it through CE. Saving settings or selecting a model does not start it.

Start and Stop

An administrator opens Model → Speech-to-Text, selects the model, and clicks Start. Wait for Running, the Device name, and Ready for dictation before recording.

First startup took about 1–2 minutes on the documented reference host; CE allows up to five minutes. Timing and memory use depend on the host. The tested model occupied about 6 GB VRAM when loaded, with additional capacity needed for inference and concurrent workloads.

Speech stays loaded between recordings until an administrator clicks Stop, CE shuts down, or the process fails. There is no automatic start/stop or model unloading to make room. Restarting CE does not start speech. Stop before selecting a different speech model or changing a running runtime’s path/port.

Dictate

  1. Click Voice Dictation beside the chat composer and allow microphone access.
  2. Speak, then finish the recording. CE processes one shared transcription request at a time.
  3. Review the text appended to the composer, edit if needed, and click Send yourself.

Cancel discards the pending result. Changing conversations prevents insertion into the wrong chat. Recordings are processed in memory, not kept in an audio library. Text becomes ordinary chat content when sent.

If speech is stopped, the microphone reports that the speech-to-text model is not running. Ask an administrator to start it or continue typing.

Microphone Requirements

Microphone capture needs browser permission, Web Audio/AudioWorklet support, and a secure context: HTTPS, or localhost on the browser’s own computer. Plain HTTP to another machine’s LAN address generally does not qualify.

The microphone belongs to the browser computer, not the CE server. Check the selected input, operating-system microphone permissions, and whether the input-level meter moves. Remote HTTP can support typed chat while microphone capture remains unavailable.

Troubleshooting

SymptomCheck first
Model missing from speech pickerFile present, registry enabled, Speech-to-Text designation set, MMPROJ unset. Stop a running instance before changing selection.
Starting, failure, or timeoutOpen Logs → Speech-to-Text. Check the dedicated interpreter, dependencies, model, free port, and available VRAM. Startup timeout is five minutes; do not launch duplicate copies.
CUDA unavailable / out of memoryVerify the documented environment and inspect the actual runtime error. Other models also consume memory; CE does not automatically move or unload them.
No speech recognizedCheck mute, input selection, input-level activity, and an audible local recording before changing CUDA/dependencies.
Browser mic unavailableHTTPS/localhost, supported browser APIs, site permission, and operating-system privacy/input settings.

Logs keeps language and speech output separate. The main CE console also shows speech lifecycle and child-process messages. Follow the detailed troubleshooting guide for the specific failure. A stopped or misconfigured speech runtime does not make speech setup required for normal CE operation.