The voice on the other end
What you'll get. Which parts of speech run on your phone, which run in the cloud, and how to change either one.
Speech has two halves, and E:Voice lets you set them separately. Hearing you is speech-to-text. Speaking back is the voice. Either one can run on your phone or in the cloud, and they do not have to match.
Everything lives in one place: Settings → Voice/Speech, described in the menu as “STT, TTS, voice training”. Inside are six sheets — Input (“Choose how E:Voice hears you.”), Output (“Choose how E:Voice speaks back.”), Voice, Speech behavior, Voice provider and Diagnostics.
What already runs on your phone
Two speech models ship inside the app. There is nothing to download and no account needed to use them.
- Local speech recognition — the offline transcriber. Pick it in Input, under “Speech recognition”. The Input sheet shows a “Local speech model” row that reads Ready or Missing, with the note “Vosk local recognition uses the bundled offline model.”
- Kokoro local voice — the on-device voice, and the app's default for reading replies. Pick it in Output, under “Default readback provider”.
The app is plain about the boundary. Local providers are labelled “speech is transcribed on this device when the bundled model is available” and “Kokoro voice is generated on this device when the bundled model is available.”
The default way E:Voice hears you is not the local one. Out of the box, speech recognition is “E:Voice powered by ElevenLabs” — a cloud route that sends your recorded audio through the E:Voice relay and spends your plan's allowance. If you want transcription to stay on the phone, switch Input to Local speech recognition yourself. It takes one tap and it is free and unlimited.
The cloud options
Under Input you can choose Local speech recognition, OpenAI cloud speech, E:Voice powered by ElevenLabs, or ElevenLabs (your API key). Under Output: Kokoro local voice, E:Voice powered by ElevenLabs, ElevenLabs (your API key), OpenAI cloud voice, Paired runtime voice, or Muted.
Choosing a “your API key” option without a key stored sends you to the API Key Vault with a nudge: “Add an ElevenLabs key in API Key Vault to enable direct speech recognition.” A key of your own bills to your provider account instead of your E:Voice allowance. Chapter 04 covers the vault.
Premium voices, said straight
Premium voices sound better. They are also the one genuinely expensive part of the app, and the Market says so in as many words: “premium voice is the one thing that isn't cheap, so every plan is honest about how much you get.”
- Local voice is free forever, on every plan, and cannot run out.
- You can listen to every premium voice on any plan. Previewing is never blocked. Keeping one assigned to an agent is part of Premium — the app tells you so directly rather than greying the button out.
- If the allowance runs out mid-sentence, the reply finishes in the local voice and is never cut short. You get a plain-spoken heads-up in the thread, not an error.
- More premium minutes come from voice packs in the Market, or from bringing your own ElevenLabs key.
Your remaining budget is on the Market screen under “Voice budget”, shown as credits available plus anything left in your packs.
Giving an agent its own voice
Voice is set per agent, so Frank and a paired runtime can sound different. Each agent offers three routes:
- The premium voice — “Speaks with … If its account allowance runs out mid-reply, the reply finishes in the local voice and is never cut short.”
- Local voice — “Always speaks with the local voice below. Costs nothing and cannot run out.”
- Follow app default — whatever you chose in Output. The screen names it for you rather than making you remember.
Each agent also gets a local fallback voice you can pick and preview, so the drop from premium to local is a change in polish rather than a change of person. There is a Preview button next to Save voice, and preview works on every plan.
If it mishears you
Two things help. Voice training (Settings → Voice/Speech, action “Tune”) teaches local recognition how you say product names and phrases — “Corrections stay on this device and are applied before Frank sees the transcript.” And the noise filter slider in Speech behavior, at 50 by default, is worth raising in a loud truck.
The on-device models are speech models only. Vosk hears you and Kokoro speaks; neither one answers you. General-purpose AI still runs on a provider you connect or a machine you pair — there is no local language model doing the thinking, and the local model catalog in the app does not change that.
Fully offline means all three parts. Local transcription, local voice and the bundled noise library have to be present together. The Input sheet tells you when one is missing.