Chapter 15 References
Chapter 15 References
Books
- Ian Goodfellow, Yoshua Bengio, and Aaron Courville, Deep Learning, MIT Press, 2016.
- Chip Huyen, Designing Machine Learning Systems, second edition, O’Reilly Media, 2024.
- Chip Huyen, AI Engineering: Building Applications with Foundation Models, O’Reilly Media, 2025.
- Daniel J. Jurafsky and James H. Martin, Speech and Language Processing, draft third edition, 2024.
Websites
- ScaNN: Scalable Nearest Neighbors — Google Research implementation and algorithm notes for partition-and-search, quantization, and rescoring.
- Accelerating Large-Scale Inference with Anisotropic Vector Quantization — the ScaNN paper.
- DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node — the DiskANN and Vamana system paper.
- DiskANN source code — Microsoft implementations of out-of-core and in-memory vector indexing.
- Efficient and Robust Approximate Nearest Neighbor Search Using HNSW — the HNSW paper.
- FAISS wiki — implementation notes for exact and approximate vector indexes, including IVF and product quantization.
- Pinecone documentation — managed vector database concepts and API guidance.
- Qdrant documentation — self-hosted vector search, collections, and payload filtering.
- Milvus documentation — distributed vector database architecture and index guidance.
- LangChain RAG tutorials — retrieval, indexing, and retrieval-chain patterns.
- LlamaIndex documentation — data connectors, indexes, query engines, and agent workflows.
- DeepSpeed documentation — ZeRO, pipeline parallelism, and large-scale training configuration.
- Megatron-LM source code — NVIDIA implementation of tensor, pipeline, sequence, and context parallelism.
- PyTorch FSDP documentation — explicit parameter, gradient, and optimizer-state sharding.
- vLLM documentation — distributed LLM inference, scheduling, paged KV caches, and supported optimizations.
- Efficient Memory Management for Large Language Model Serving with PagedAttention — the PagedAttention and vLLM systems paper.
- Orca: A Distributed Serving System for Transformer-Based Generative Models — the continuous-batching serving design.
- LangChain agents — tool-calling agent loops and runtime state.
- LlamaIndex agents — prebuilt agents and event-driven agent workflows.
- ReAct: Synergizing Reasoning and Acting in Language Models — the interleaved reasoning-and-action paper.
- Hugging Face PEFT documentation — adapter configuration and PEFT integration.
- LoRA: Low-Rank Adaptation of Large Language Models — the LoRA paper.
- QLoRA: Efficient Finetuning of Quantized LLMs — the QLoRA paper.
- Deep Reinforcement Learning from Human Preferences — the foundational RLHF paper.
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model — the DPO paper.
- Hugging Face TRL documentation — supervised, reward, DPO, and reinforcement-learning trainers.