RAG in Production
Reliability, failure modes, and the architecture that keeps retrieval-augmented generation working with real users.
29 posts
The pillars I write about — each a hub of deeper posts.
Reliability, failure modes, and the architecture that keeps retrieval-augmented generation working with real users.
29 posts
Chunking, hybrid search, reranking, and query understanding — the levers that decide whether the right context shows up.
9 posts
Qdrant deep-dives, indexing, filtering, and scaling the store behind semantic search.
6 posts
Models, dimensions, caching, and fine-tuning the representations retrieval runs on.
6 posts
Evals, hallucination detection, and monitoring so you know when the system is actually working.
5 posts
Agents, orchestration, and agentic/corrective RAG — and when the added complexity pays off.
1 post
Prompting, context management, streaming, and structured output for real applications.
8 posts
MLX, Core ML, quantization, and running inference locally on Apple Silicon.
4 posts
Caching, batching, and the infrastructure choices that keep LLM apps fast and affordable.
1 post
Document parsing, PDFs, and the pipelines that turn messy sources into clean retrievable context.
5 posts