Tune Performance, Hardware, and Network
Tune speed, memory use, hardware acceleration, and network settings.
Follow the bold button names and numbered steps. Technical names are included only when you need to recognize a model or file inside the app.
Open Settings → Hardware to inspect storage, choose compute behavior, set the local context window, reduce resource use, and control network routing.
Choose a compute device and verify acceleration
Goal
Use the best available hardware and confirm whether the active local model is running on the GPU or CPU.
Before you start
- Send at least one message so the local model is loaded before checking acceleration.
- A GPU is optional on Windows; Local Waifu can fall back to CPU.
- macOS builds require Apple Silicon and use unified memory.
Steps: macOS
- Open Settings → Hardware → Compute Performance.
- Under Compute Device, choose:
- Auto: recommended; lets the local engine select acceleration.
- CPU: forces CPU-only inference.
- Metal (GPU): allows Metal acceleration on Apple Silicon.
- Restart Local Waifu when prompted.
Steps: Windows
- Open Settings → Hardware → Compute Performance.
- Leave Compute Device on Auto-detect (Recommended) unless troubleshooting.
- Send a message to load the model.
- Under Local AI acceleration, select Check.
- Read the result:
- Running on GPU: the model is in VRAM.
- Split CPU/GPU: only part of the model fits in VRAM.
- Running on CPU: the GPU is not engaged.
- No model loaded yet: send a message, then check again.
- If NVIDIA or AMD hardware is detected but not fully used, select Retry GPU setup. The app may download a matching runtime and restart the engine.
Expected result
macOS uses Metal unless CPU-only mode was selected. Windows reports the real placement of the loaded model and, where available, the percentage held in VRAM.
If something goes wrong
- Restart the app after changing Compute Device.
- Try a smaller or more heavily quantized model if Windows reports Split CPU/GPU.
- Update GPU drivers, then use Retry GPU setup again.
- CPU fallback is functional but can be much slower.
- There is no separate Apple Neural Engine inference path in the macOS 1.7.3 build; use Auto or Metal (GPU).
Check runtime readiness
Goal
Confirm which local features can run now.
Before you start
This detailed panel is available in the Windows build. On macOS, inspect the corresponding model panels if the runtime readout is not shown.
Steps
- Open Settings → Hardware.
- Under Runtime status, select Check.
- Review Local AI engine, Memory model, Chat model, Voice output, Sees images, and Image generation.
- For an item marked Not available, open the relevant model section and install or select the missing component.
Expected result
Each feature you intend to use is marked Ready.
If something goes wrong
Restart Local Waifu after installing a runtime or model, then run Check again. If the local engine remains unavailable, verify Ollama Host and review App Logs as described in Chapter 27.
Set the context window and thinking mode
Goal
Balance conversation context, memory use, and response speed for local chat models.
Before you start
Larger context windows require more RAM or unified memory and may substantially slow replies. Cloud providers use their own limits.
Steps
- Open Settings → Hardware → Context Window.
- Start with 8K (Medium).
- Increase to 16K (High) or 32K (Ultra) only if your model and available memory can support it.
- Use Custom only when you know the model’s context limit. Windows accepts 8,192 to 262,144; the macOS field can accept 512 to 262,144, although the active model may support less.
- Turn on Deeper thinking only when you prefer more considered but slower responses.
Expected result
Local replies use the requested context size within the active model’s capabilities. Deeper thinking trades speed for more reasoning.
If something goes wrong
- Reduce Context Window if responses become slow, memory pressure rises, or the model fails to load.
- A high value is a request, not a guarantee; the active model may support less.
- Turn off Deeper thinking for faster, more responsive chat.
Reduce background resource use
Goal
Make Local Waifu release model memory sooner and perform fewer background wake-ups.
Before you start
Low Power Mode can make the next reply slower because the local model is unloaded sooner between chats.
Steps
- Open Settings → Hardware.
- Turn on Low Power Mode.
- Restart Local Waifu.
Expected result
Background work occurs less frequently and the local model leaves memory sooner after inactivity.
If something goes wrong
If first-response latency becomes inconvenient, turn off Low Power Mode and restart the app.
Configure proxy, offline mode, or a remote Ollama server
Goal
Control network behavior without disabling local AI unnecessarily.
Before you start
- Use System Proxy is enabled by default and changes require a restart.
- Offline Mode blocks cloud providers but does not block local Ollama.
- A remote Ollama Host exposes prompts to that other computer and to the intervening network; use only a host you trust.
Steps
- Open Settings → Hardware → Network.
- Leave Use System Proxy enabled if your network depends on operating-system proxy settings. Restart after changing it.
- Turn on Offline Mode to prevent cloud chat providers from running.
- To use Ollama on another computer, enter its URL in Ollama Host, for example
http://192.168.1.50:11434. - To return to the bundled local engine, select Reset beside Ollama Host.
Expected result
Local Ollama continues to work in Offline Mode. Cloud models refuse to run. A valid custom Ollama Host routes local-model requests to the specified server.
If something goes wrong
- If cloud chat reports that it is blocked, turn off Offline Mode only if sending data to the configured provider is intentional.
- If local chat goes offline after editing Ollama Host, select Reset.
- If downloads fail behind a corporate proxy, enable Use System Proxy and restart.
- Check firewall access to the remote Ollama host and port; do not expose Ollama directly to the public internet.
Storage note: Local models can occupy tens of gigabytes. The Change Path control is disabled in the macOS 1.7.3 UI. Do not rely on it for moving models. The Windows build can support moving model storage, but perform a backup first and follow the controls shown by that build.