INTERACTIVE WORKBENCH

Find the right KV-cache setup for your Mac

Choose your Mac, model, and goal. We’ll recommend the best KV-cache method, estimate memory use, and show how much conversation fits.

How this works

As an LLM generates tokens, it stores key-value pairs in unified RAM so it can remember earlier context. VeloxQuant-MLX compresses this KV cache with Apple Silicon Metal kernels, helping the same Mac handle longer conversations or larger models.

Choose Quantization Algorithm
Where these numbers come from (Sizing formulas & benchmarks)

Sizing formulas and recommendations run the exact same logic as the veloxquant recommend CLI tool. Benchmark curves in Step 3 are real, measured Metal kernel results. To run native profiling on your own machine, install VeloxQuant-MLX and run veloxquant benchmark.