Voice (speech to text)
Voice in Imaginne is speech to text (STT): you speak, the transcription appears in the composer, and you review and send. It never sends the message on its own — it's a way to type faster, not a voice assistant. The feature is available on the Desktop, the TUI, and the web chat.
Where it works
| Surface | How to trigger | Availability |
|---|---|---|
| Desktop | Microphone button in the composer | Compiled into the app; requires microphone permission |
| Terminal (TUI) | Ctrl+G or F2 | Depends on a build with the voice tag |
| Web chat | Microphone button in the composer | Requires a browser with audio recording and microphone permission |
| VS Code | — | Not available |
The Desktop and the TUI capture audio natively and transcribe in real time — text appears as you speak. The web chat uses the browser's recorder: you record, stop, and the clip is transcribed at once. The transcription service is the same; only the capture differs. The VS Code extension has no voice dictation.
On the Desktop
- Click the microphone button in the composer (corner of the message field). If the feature is unavailable, the button is hidden.
- Speak. The partial transcription appears in the text field as you dictate.
- Click the microphone again to stop capture.
- Review the text, adjust whatever you need, and send with Enter.
The transcription lands in the same field where you type, so you can mix voice and keyboard: dictate a passage, finish by typing, and send.
On some Desktop versions there's a tooltip mentioning Ctrl+Y, but that shortcut does not trigger voice. Use the microphone button.
On the TUI
- Press Ctrl+G or F2 to start capture.
- Speak. The transcription is appended to your current prompt.
- Press Esc to cancel capture without inserting anything.
- Review and send with Enter.
Voice in the TUI depends on the binary being compiled with the voice tag. Today, that applies to the darwin-arm64 release; on other platforms voice may not be available. If the shortcuts do nothing, your build probably doesn't include voice — see Install the TUI.
Capture behavior can be tuned via IMAGINNE_VOICE_* environment variables (see Environment variables).
In the web chat
- Click the microphone button in the composer. The first time, the browser asks for microphone permission.
- Speak. The status next to the button shows "Recording… tap to stop".
- Click again to stop. You'll see "Transcribing audio…" and then the text lands in the message field.
- Review and send with Enter.
Web chat specifics:
- The transcription is inserted at the cursor and never erases what you had written — dictate a bit, fix it by typing, dictate more.
- A recording runs up to 2 minutes; past that it stops on its own and is transcribed.
- Switching conversations, closing the tab, or signing out cancels the recording and releases the microphone.
- If the browser doesn't support audio recording, the button simply doesn't appear.
Common messages: "Microphone access was blocked" (allow the microphone in your browser settings), "No microphone was found on this device", and "No speech was detected" (the recording came out empty or too short).
Cancel
On every surface, nothing is sent automatically: cancelling discards whatever hasn't been transcribed yet, and the text already in the field stays. On the Desktop and the web chat, clicking the microphone again stops capture; on the TUI, Esc cancels.
How voice is processed
The audio is transcribed by the platform's voice service, not by a chat model. The flow is: the app captures audio from the microphone, sends it to the service, and gets back the transcription, which appears in the composer. Only then, when you send, does the text become a message for the agent — the audio itself is not sent to the model.
Because it's a step separate from the chat, voice depends on your organization having the feature enabled, regardless of the model you use in the conversation.
Treat the transcription as a draft: review proper names, numbers, and technical terms before sending. You can dictate a passage, fix it by typing, and only then send — all in the same message field.
Language
The transcription follows the app's language setting (for example, PT or EN on the Desktop, in Settings → Info). Speak in the configured language for the best recognition.
Requirements
To use voice, you need:
- A microphone and permission for Imaginne to access it — from the operating system (Desktop/TUI) or from the browser (web chat). On first use, authorization is requested.
- Voice credentials enabled by your organization. The transcription is processed by the platform's voice service; if your org hasn't enabled this feature, voice stays unavailable even with a microphone.
- On the TUI, a build with the
voicetag (see above). - In the web chat, a browser that supports audio recording (recent Chrome, Edge, Firefox, and Safari).
If the button (Desktop and web chat) or the shortcuts (TUI) don't work, check in this order: microphone permission, whether the org enabled voice, and, on the TUI, whether the build has the voice tag. See Troubleshooting.
See also
Was this page helpful?
Report a problem on this pageDo not send passwords, keys, tokens, or customer data.