Documentation›Models, images, and voice
21 Models and media

Configure Voice, Dictation, and Calls

Set up speech recognition, voices, dictation, and live calls.

Local Waifu 1.7.3 macOS Windows
plain English first

Follow the bold button names and numbered steps. Technical names are included only when you need to recognize a model or file inside the app.

Local Waifu has three separate voice functions:

  1. Dictation: speech becomes editable text in the composer; it is not sent automatically.
  2. Voice output: the companion reads a response aloud.
  3. Call: a continuous speech-recognition → chat → speech loop.

Install speech and voice models

A full call needs:

  • speech recognition (STT/Whisper) to hear you;
  • a voice model (TTS/Supertonic or another configured route) to speak.

Missing STT blocks the full call. Missing high-quality TTS may fall back to an older voice, but pronunciation can be worse.

Open Settings → Models → Voice Model, install Speech recognition and Voice model, and optionally install Natural Voice. The screen reports current package sizes and disk use.

Choose a voice

  1. Open Settings → Models → Voice Model.
  2. Preview the default, female, or male voices with the play button.
  3. Select a voice for the active character.
  4. Turn on Natural voice for the warmer, more human local voice when installed.
  5. Use Test her voice, type a sentence, and select Speak.

Language support depends on the voice engine. Supertonic covers English, Polish, German, Spanish, Japanese, and Korean. Chinese requires the Natural Voice route for proper support; otherwise an English-accented fallback may be used.

Add your own voice sample

Goal

Create a character-specific voice reference.

Before you start

  • Record 10 to 30 seconds of clean speech.
  • Use one speaker in a quiet room.
  • Record in the language/accent you want to preserve.
  • File size must be under 25 MB.

Steps

  1. Open Settings → Models → Voice Model.
  2. Under Your own voice, select Add a voice from a file.
  3. Choose the audio file.
  4. Enter a short name, up to 60 characters.
  5. Wait for local processing and transcription.
  6. Preview the result and select it.

Expected result

The active character uses the new reference and Natural Voice is enabled when ready.

If something goes wrong

Use a shorter, cleaner sample. A learned English sample keeps its English accent in other languages; record again in the target language for better pronunciation.

Tune how she speaks

Available style presets include Auto, Soft & warm, Balanced, Bright & playful, Calm & even, and Confident & teasing.

Adjust:

  • speaking pace;
  • pitch;
  • expressiveness;
  • Let her mood colour her voice.

Use small changes and test after each one. Auto follows the character personality; manual presets keep your chosen delivery style.

Configure a call

Goal

Select audio devices and tune turn-taking.

Steps

  1. Open the call/voice settings.
  2. Select Microphone and Speaker, or keep System default.
  3. Select Test microphone and verify activity.
  4. Adjust How long she waits. A shorter value feels faster but can cut you off mid-thought.
  5. Enable I’m wearing headphones only when appropriate; speakers can feed her voice back into the microphone and cause self-interruption.
  6. Leave cloud recognition and cloud voice off for fully local processing, or enable them intentionally with your provider key.
  7. Start a call with the phone button or Command/Ctrl+Shift+C.

Expected result

Your speech is transcribed, answered by the active model, and spoken through the selected voice and speaker.

If something goes wrong

  • Grant microphone permission in the operating system.
  • Choose the correct input device.
  • Install STT if the call cannot hear you.
  • Install/configure TTS if text replies appear but no voice plays.
  • Increase the waiting time if you are interrupted too early.