Interactive · runs in your browser

Feel the power of VeloxQuant-MLX

Dial in your Mac, model and goal — see the recommended compression method, how much more context fits in your RAM, and real measured benchmarks.

Every number is computed live from the same heuristic that ships in veloxquant recommend — a direct port of mac_recommender.py — or read from committed benchmark data. Nothing is faked.

Start from a model
Advanced — architecture shape
Method

Recommender + compression estimates are transparent heuristics (a direct port of veloxquant_mlx/tools/mac_recommender.py) — they state when resident RAM savings are unlikely. Token counts and the book comparison are linear KV extrapolations, labelled as approximations. Benchmark charts read measured values committed under figures/. For live numbers on your own machine, run python -m veloxquant_mlx benchmark.