Feel the power of VeloxQuant-MLX
Dial in your Mac, model and goal — see the recommended compression method, how much more context fits in your RAM, and real measured benchmarks.
Every number is computed live from the same heuristic that ships in
veloxquant recommend — a direct port of
mac_recommender.py — or read from committed benchmark data.
Nothing is faked.
Advanced — architecture shape
Recommender + compression estimates are transparent heuristics (a direct port of
veloxquant_mlx/tools/mac_recommender.py) — they state when resident RAM
savings are unlikely. Token counts and the book comparison are linear KV
extrapolations, labelled as approximations. Benchmark charts read measured values
committed under figures/. For live numbers on your own machine, run
python -m veloxquant_mlx benchmark.