Back to Localai

Measured against what each workload actually runs on

website/static/media/v4-8-0-vllm-cpp-scoreboard.html

4.8.0310 B
Original Source

vllm.cpp · throughput vs the reference engine

Measured against what each workload actually runs on

Throughput relative to the reference. 1.00 is parity, bars run from it. Higher is faster.

github.com/mudler/vllm.cppGB10 unless noted · greedy, reference in its own production config · docs/BENCHMARKS.md