Documentation›Models, images, and voice
26 Models and media

Use Your Own Server as a Cloud Provider

Point Local Waifu at LM Studio, vLLM, llama.cpp, or a private OpenAI-compatible server.

Local Waifu 1.7.3 macOS Windows
plain English first

Follow the bold button names and numbered steps. Technical names are included only when you need to recognize a model or file inside the app.

Cloud providers in Local Waifu are not limited to the named ones. If your server speaks the OpenAI-compatible chat API, you can add it and use it exactly like a hosted provider, with your own key or none at all.

This is the route for LM Studio, vLLM, llama.cpp server, Ollama exposed over HTTP, and anything else that implements /v1/chat/completions.

Add an endpoint

Goal

Register your own server as a usable chat option.

Before you start

  • Have the server running and reachable from this computer.
  • Know its base URL, including the /v1 part. For LM Studio on the same machine that is http://localhost:1234/v1.
  • Most local servers need no API key. Leave the key blank if so.

Steps

  1. Open Settings → Cloud AI.
  2. Scroll to Your own endpoints.
  3. Select Add an endpoint.
  4. Fill in the three fields:
    • Name: whatever you will recognise later, for example LM Studio.
    • Base URL: the address including the version path, for example http://localhost:1234/v1 or https://api.example.com/v1.
    • API key: leave it blank if the server does not need one. On an edit, leaving it blank keeps the saved key.
  5. Select Save.
  6. Select Test on the new row.
  7. Select the endpoint’s model in the model picker for chat, as you would for any provider.

Expected result

The row reports that the test succeeded, the endpoint appears in the chat model list, and replies come from your server.

If something goes wrong

  • If Test fails, open the base URL in a browser on the same machine. If that does not answer, the problem is the server or its port, not Local Waifu.
  • If the server answers in a browser but not here, check that the URL ends with the version path your server expects, usually /v1.
  • If the model list comes back empty, the server is not publishing an OpenAI-compatible /v1/models. Type the model name in by hand where the picker allows it.
  • Use Edit to correct the URL and Remove to delete the row.

Security: what the URL is allowed to be

This is the part worth reading before pasting an address.

  • Plain http:// is allowed only for the machine itself (localhost, 127.0.0.1) and for private network addresses such as 192.168.x.x, 10.x.x.x, and .local names.
  • Anything on the public internet must use https://. There is no setting that relaxes this, because a key sent over plain HTTP is a key given away.
  • A remote Ollama Host is a separate setting with its own rules. See Chapter 29.

If something goes wrong

If an address is rejected, it is almost always plain http:// pointed at a public hostname. Either move the server onto your own network or put TLS in front of it.

Offline Mode and your own endpoint

Offline Mode blocks cloud providers. It does not block an endpoint on your own machine or your own network, because that request never leaves your network. Offline Mode means off the internet, not off your own hardware.

If something goes wrong

If chat is refused while Offline Mode is on, check whether the endpoint resolves to a public address. A hostname that looks private but resolves outward is treated as a cloud provider.

What your own endpoint does not get

  • It does not become the local engine. Embeddings and memory stay on your computer, as described in Chapter 30.
  • It does not carry vision unless the server implements it and you mark the provider as vision-capable.
  • It does not survive a factory reset of the app’s settings. Endpoints are stored with the rest of your configuration, which Reset All clears, and which a full backup preserves.