model layer
…
connecting…
auto-refresh
API docs
Try it
stream
raw
Send
thinking
result appears here
Machines
free resources per host — to see what can fit a new model
Models
+ Connect backend
+ Deploy local
↑ Push all to remote
Connect a backend — paste its URL, we probe it on connect
base_url
name
model_id
API key
Advanced
type
external
llama
mode
sync
queued
concurrency
API key env var
Connect
Cancel
Deploy a local llama-server (launches + self-registers)
name
model_path
mode
queued
sync
concurrency
ctx
Deploy
Cancel
Deploy blocks while the model loads (can take 10–30s).
Edit connection —
base_url
model_id
API key
API key env var
mode
sync
queued
concurrency
tags
Load/unload control (optional) — set these to drive a native server via its node agent
agent_url
model_path
port
ctx
Save
Cancel
Push all to remote router — registers every
enabled
backend (disable one to exclude it)
remote router URL
remote admin token
reachable host
Push all
Cancel
External (paid-API) backends are pushed keyless — secrets aren't sent.
Push to remote router — register
on another router
remote router URL
remote admin token
reachable host
name on remote
Push
Cancel
Advertises this backend at
reachable host
(not loopback) so the remote can reach it. Secrets are not sent.
name
type
mode
conc
status
latency
last seen
endpoint
actions
Queue
job
model
status
enqueued
finished
Recent calls
time
model
backend
path
status
queued
latency
tokens (p/c)