/4 min
Running LLM Evals On-Device With MLX: Zero-Cost Batch Judging on a Mac
Yes — a quantized 3-8B judge running on MLX on Apple Silicon handles objective eval batches at effectively zero marginal cost.
Tag
2 posts.
Yes — a quantized 3-8B judge running on MLX on Apple Silicon handles objective eval batches at effectively zero marginal cost.
MLX runs on-device LLMs on the GPU for throughput; Core ML can place them on the ANE to keep the GPU free — here's how to choose.