docs/content/getting-started/linux.md
You can manually download the appropriate binary for your system from the releases page:
chmod +x local-ai-*
./local-ai-*
Starting the binary on its own gives you an empty server. To get a working chat right away, run LocalAI with a model name and it will download and serve it from the gallery:
./local-ai-* run qwen3-4b
Once it is ready, open the WebUI at http://localhost:8080 or send a request to the API:
curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "qwen3-4b",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Hardware requirements vary based on:
For performance benchmarks with different backends like llama.cpp, visit this link.
After installation, you can:
http://localhost:8080LocalAI accepts a single TCP listener passed through the systemd socket activation protocol. This lets systemd listen on the public port and start LocalAI only when the first client connects.
Create /etc/systemd/system/local-ai.socket:
[Unit]
Description=LocalAI API socket
[Socket]
ListenStream=8080
NoDelay=true
[Install]
WantedBy=sockets.target
Create the matching /etc/systemd/system/local-ai.service:
[Unit]
Description=LocalAI
[Service]
Type=simple
User=localai
Group=localai
ExecStart=/usr/local/bin/local-ai run
WorkingDirectory=/var/lib/local-ai
Adjust the user, binary path, working directory, and model configuration for your installation. Then enable the socket, not the service:
sudo systemctl daemon-reload
sudo systemctl enable --now local-ai.socket
The first connection to port 8080 starts local-ai.service; systemd holds that
connection until LocalAI is ready to accept it. LOCALAI_ADDRESS and
--address are ignored while an inherited listener is present. LocalAI
rejects activation with multiple stream listeners so it cannot silently choose
the wrong endpoint.
For a Podman-managed container, configure Podman to preserve and pass the systemd socket file descriptor into the container. The LocalAI process inside the container consumes the same activation protocol.
Activation needs both LISTEN_PID and LISTEN_FDS. If only one of them is set,
LocalAI ignores them and binds --address as usual. A container engine started
from a socket-activated system unit can leak a bare LISTEN_PID into every
container it spawns, and that is not an activation attempt.