Local models
Bring your own endpoint. This page is the support contract.
Local models are supportable when the boundary is clear. This is that boundary: anything on this page is supported, anything not on it is not.
What Daktyl does and does not do
- You run the model server. Daktyl never starts, stops, installs or manages an inference server. It is your process, your unit file, your GPU.
- You point the harness at it. The harness is configured with an OpenAI-compatible endpoint — a base URL, a model id, optionally a key. Daktyl reads that config; it does not write it.
- Daktyl only checks it answers. Switching a seat to a local model probes the endpoint once. If nothing answers, the switch is refused with one sentence ("llama.cpp on 127.0.0.1:8080 isn't running — start it, then switch again") and the seat stays on the model that works. If something else holds the port, the sentence names it.
- Telemetry needs a context limit. Give a custom model a context limit and the context ring works. Without one, Daktyl shows no stats rather than wrong ones.
Wiring the probe
Per harness, in ~/.daktyl/config.json, keyed by the model id's provider prefix. A
provider that is not in the map is never probed and the switch behaves as before.
"local_endpoints": {
"llamacpp": { "url": "http://127.0.0.1:8080/v1/models", "label": "llama.cpp" }
}
A recipe that works
llama-server -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen3.8-27b \
--jinja -c 32768 --host 127.0.0.1 --port 8080
Then a provider row in the harness pointing at
http://127.0.0.1:8080/v1. Reasoning traces are split off by the server's default
reasoning format, so a model's private thinking is never spoken aloud on a call.
Some local models are fine for text and poor for voice — a model that pauses for
ten seconds before its first token makes a bad phone call. Test on a call before relying on one.