2026
- May 10 - VeloxQuant-MLX: Fast KV Cache Quantization for Apple Silicon
- May 12 - I Ported a Google Research Paper to Apple Silicon
- May 17 - Benchmark Results: 10 Models, 8 Compression Configs
- May 20 - TurboQuant + Metal Kernels: The Combined Writeup
- May 25 - I Wrote a Metal Kernel to Stop My Mac From OOMing
- May 28 - Hands-On: Compressing Your First LLM with VeloxQuant-MLX
- June 10 - KIVI: The Most-Cited KV Cache Baseline, Implemented in MLX
- June 20 - TensorOps Research: What We Learned Optimizing KV Caches
- August 12 - A 5.65× Metal Kernel
- August 14 - The Sign Was the Whole Paper: Debugging a KV Cache Compressor on Real Models
- August 26 - The ROM Chip That Wasn't, and the 230× Speedup That Was
- August 28 - Fusing Quantize and Pack Into One Metal Dispatch
- August 30 - Chasing the Mac-vs-CUDA Prefill Gap — and Finding a Wall Instead
- September 5 - Batching Decode Attention Across Layers Sounded Great — Until the Residual Stream Said No
- September 5 - Batching Decode Attention Across Requests Got Us 3.83x Real Throughput — Across Layers Got Us Nothing