Skip to main content

Open reference library / 523 lessons

Study the mechanics. Build the system.

Explore the full English lesson sequence from AI Engineering from Scratch, alongside our governed agent field course. The original curriculum spans mathematics, machine learning, language models, tools, agents, infrastructure, safety, and capstones.

Reference material by Rohit Ghumare and contributors, MIT licensed. Imported at revision bf7791e14076. Runnable code and outputs remain linked to the source repository.

Connect your agent to this foundation ↗ — public search and lesson retrieval, with a reusable onboarding guide.

523 lessons

00 / setup and tooling Dev Environment ~45 minutes ↗00 / setup and tooling Git & Collaboration ~30 minutes ↗00 / setup and tooling GPU Setup & Cloud ~45 minutes ↗00 / setup and tooling APIs & Keys ~30 minutes ↗00 / setup and tooling Jupyter Notebooks ~30 minutes ↗00 / setup and tooling Python Environments ~30 minutes ↗00 / setup and tooling Docker for AI ~60 minutes ↗00 / setup and tooling Editor Setup ~20 minutes ↗00 / setup and tooling Data Management ~45 minutes ↗00 / setup and tooling Terminal & Shell ~35 minutes ↗00 / setup and tooling Linux for AI ~30 minutes ↗00 / setup and tooling Debugging and Profiling ~60 minutes ↗01 / math foundations Linear Algebra Intuition ~60 minutes ↗01 / math foundations Vectors, Matrices & Operations ~60 minutes ↗01 / math foundations Matrix Transformations ~75 minutes ↗01 / math foundations Calculus for Machine Learning ~60 minutes ↗01 / math foundations Chain Rule & Automatic Differentiation ~90 minutes ↗01 / math foundations Probability and Distributions ~75 minutes ↗01 / math foundations Bayes' Theorem ~75 minutes ↗01 / math foundations Optimization ~75 minutes ↗01 / math foundations Information Theory ~60 minutes ↗01 / math foundations Dimensionality Reduction ~90 minutes ↗01 / math foundations Singular Value Decomposition ~120 minutes ↗01 / math foundations Tensor Operations ~90 minutes ↗01 / math foundations Numerical Stability ~120 minutes ↗01 / math foundations Norms and Distances ~90 minutes ↗01 / math foundations Statistics for Machine Learning ~120 minutes ↗01 / math foundations Sampling Methods ~120 minutes ↗01 / math foundations Linear Systems ~120 minutes ↗01 / math foundations Convex Optimization ~90 minutes ↗01 / math foundations Complex Numbers for AI ~60 minutes ↗01 / math foundations The Fourier Transform ~90 minutes ↗01 / math foundations Graph Theory for Machine Learning ~90 minutes ↗01 / math foundations Stochastic Processes ~75 minutes ↗02 / ml fundamentals What Is Machine Learning ~45 minutes ↗02 / ml fundamentals Linear Regression ~90 minutes ↗02 / ml fundamentals Logistic Regression ~90 minutes ↗02 / ml fundamentals Decision Trees and Random Forests ~90 minutes ↗02 / ml fundamentals Support Vector Machines ~90 minutes ↗02 / ml fundamentals K-Nearest Neighbors and Distances ~90 minutes ↗02 / ml fundamentals Unsupervised Learning ~90 minutes ↗02 / ml fundamentals Feature Engineering & Selection ~90 minutes ↗02 / ml fundamentals Model Evaluation ~90 minutes ↗02 / ml fundamentals Bias-Variance Tradeoff ~75 minutes ↗02 / ml fundamentals Ensemble Methods ~120 minutes ↗02 / ml fundamentals Hyperparameter Tuning ~90 minutes ↗02 / ml fundamentals ML Pipelines ~120 minutes ↗02 / ml fundamentals Naive Bayes ~75 minutes ↗02 / ml fundamentals Time Series Fundamentals ~90 minutes ↗02 / ml fundamentals Anomaly Detection ~75 minutes ↗02 / ml fundamentals Handling Imbalanced Data ~90 minutes ↗02 / ml fundamentals Feature Selection ~75 minutes ↗03 / deep learning core The Perceptron ~60 minutes ↗03 / deep learning core Multi-Layer Networks and Forward Pass ~90 minutes ↗03 / deep learning core Backpropagation from Scratch ~120 minutes ↗03 / deep learning core Activation Functions ~75 minutes ↗03 / deep learning core Loss Functions ~75 minutes ↗03 / deep learning core Optimizers ~75 minutes ↗03 / deep learning core Regularization ~75 minutes ↗03 / deep learning core Weight Initialization and Training Stability ~90 minutes ↗03 / deep learning core Learning Rate Schedules and Warmup ~90 minutes ↗03 / deep learning core Build Your Own Mini Framework ~120 minutes ↗03 / deep learning core Introduction to PyTorch ~75 minutes ↗03 / deep learning core Introduction to JAX ~90 minutes ↗03 / deep learning core Debugging Neural Networks ~90 minutes ↗04 / computer vision Image Fundamentals — Pixels, Channels, Color Spaces ~45 minutes ↗04 / computer vision Convolutions from Scratch ~75 minutes ↗04 / computer vision CNNs — LeNet to ResNet ~75 minutes ↗04 / computer vision Image Classification ~75 minutes ↗04 / computer vision Transfer Learning & Fine-Tuning ~75 minutes ↗04 / computer vision Object Detection — YOLO from Scratch ~75 minutes ↗04 / computer vision Semantic Segmentation — U-Net ~75 minutes ↗04 / computer vision Instance Segmentation — Mask R-CNN ~75 minutes ↗04 / computer vision Image Generation — GANs ~75 minutes ↗04 / computer vision Image Generation — Diffusion Models ~75 minutes ↗04 / computer vision Stable Diffusion — Architecture & Fine-Tuning ~75 minutes ↗04 / computer vision Video Understanding — Temporal Modeling ~45 minutes ↗04 / computer vision 3D Vision — Point Clouds & NeRFs ~45 minutes ↗04 / computer vision Vision Transformers (ViT) ~45 minutes ↗04 / computer vision Real-Time Vision — Edge Deployment ~75 minutes ↗04 / computer vision Build a Complete Vision Pipeline — Capstone ~120 minutes ↗04 / computer vision Self-Supervised Vision — SimCLR, DINO, MAE ~75 minutes ↗04 / computer vision Open-Vocabulary Vision — CLIP ~45 minutes ↗04 / computer vision OCR & Document Understanding ~45 minutes ↗04 / computer vision Image Retrieval & Metric Learning ~45 minutes ↗04 / computer vision Keypoint Detection & Pose Estimation ~45 minutes ↗04 / computer vision 3D Gaussian Splatting from Scratch ~90 minutes ↗04 / computer vision Diffusion Transformers & Rectified Flow ~75 minutes ↗04 / computer vision SAM 3 & Open-Vocabulary Segmentation ~60 minutes ↗04 / computer vision Vision-Language Models — The ViT-MLP-LLM Pattern ~75 minutes ↗04 / computer vision Monocular Depth & Geometry Estimation ~60 minutes ↗04 / computer vision Multi-Object Tracking & Video Memory ~60 minutes ↗04 / computer vision World Models & Video Diffusion ~75 minutes ↗05 / nlp foundations to advanced Text Processing — Tokenization, Stemming, Lemmatization ~45 minutes ↗05 / nlp foundations to advanced Bag of Words, TF-IDF, and Text Representation ~75 minutes ↗05 / nlp foundations to advanced Word Embeddings — Word2Vec from Scratch ~75 minutes ↗05 / nlp foundations to advanced GloVe, FastText, and Subword Embeddings ~45 minutes ↗05 / nlp foundations to advanced Sentiment Analysis ~75 minutes ↗05 / nlp foundations to advanced Named Entity Recognition ~75 minutes ↗05 / nlp foundations to advanced POS Tagging and Syntactic Parsing ~45 minutes ↗05 / nlp foundations to advanced CNNs and RNNs for Text ~75 minutes ↗05 / nlp foundations to advanced Sequence-to-Sequence Models ~75 minutes ↗05 / nlp foundations to advanced Attention Mechanism — The Breakthrough ~45 minutes ↗05 / nlp foundations to advanced Machine Translation ~75 minutes ↗05 / nlp foundations to advanced Text Summarization ~75 minutes ↗05 / nlp foundations to advanced Question Answering Systems ~75 minutes ↗05 / nlp foundations to advanced Information Retrieval and Search ~75 minutes ↗05 / nlp foundations to advanced Topic Modeling — LDA and BERTopic ~45 minutes ↗05 / nlp foundations to advanced Text Generation Before Transformers — N-gram Language Models ~45 minutes ↗05 / nlp foundations to advanced Chatbots — Rule-Based to Neural to LLM Agents ~75 minutes ↗05 / nlp foundations to advanced Multilingual NLP ~45 minutes ↗05 / nlp foundations to advanced Subword Tokenization — BPE, WordPiece, Unigram, SentencePiece ~60 minutes ↗05 / nlp foundations to advanced Structured Outputs & Constrained Decoding ~60 minutes ↗05 / nlp foundations to advanced Natural Language Inference — Textual Entailment ~60 minutes ↗05 / nlp foundations to advanced Embedding Models — The 2026 Deep Dive ~60 minutes ↗05 / nlp foundations to advanced Chunking Strategies for RAG ~60 minutes ↗05 / nlp foundations to advanced Coreference Resolution ~60 minutes ↗05 / nlp foundations to advanced Entity Linking & Disambiguation ~60 minutes ↗05 / nlp foundations to advanced Relation Extraction & Knowledge Graph Construction ~60 minutes ↗05 / nlp foundations to advanced LLM Evaluation — RAGAS, DeepEval, G-Eval ~75 minutes ↗05 / nlp foundations to advanced Long-Context Evaluation — NIAH, RULER, LongBench, MRCR ~60 minutes ↗05 / nlp foundations to advanced Dialogue State Tracking ~75 minutes ↗06 / speech and audio Audio Fundamentals — Waveforms, Sampling, Fourier Transform ~45 minutes ↗06 / speech and audio Spectrograms, Mel Scale & Audio Features ~45 minutes ↗06 / speech and audio Audio Classification — From k-NN on MFCCs to AST and BEATs ~75 minutes ↗06 / speech and audio Speech Recognition (ASR) — CTC, RNN-T, Attention ~45 minutes ↗06 / speech and audio Whisper — Architecture & Fine-Tuning ~75 minutes ↗06 / speech and audio Speaker Recognition & Verification ~45 minutes ↗06 / speech and audio Text-to-Speech (TTS) — From Tacotron to F5 and Kokoro ~75 minutes ↗06 / speech and audio Voice Cloning & Voice Conversion ~75 minutes ↗06 / speech and audio Music Generation — MusicGen, Stable Audio, Suno, and the Licensing Earthquake ~75 minutes ↗06 / speech and audio Audio-Language Models — Qwen2.5-Omni, Audio Flamingo, GPT-4o Audio ~45 minutes ↗06 / speech and audio Real-Time Audio Processing ~75 minutes ↗06 / speech and audio Build a Voice Assistant Pipeline — The Phase 6 Capstone ~120 minutes ↗06 / speech and audio Neural Audio Codecs — EnCodec, SNAC, Mimi, DAC and the Semantic-Acoustic Split ~60 minutes ↗06 / speech and audio Voice Activity Detection & Turn-Taking — Silero, Cobra, and the Flush Trick ~45 minutes ↗06 / speech and audio Streaming Speech-to-Speech — Moshi, Hibiki, and Full-Duplex Dialogue ~75 minutes ↗06 / speech and audio Voice Anti-Spoofing & Audio Watermarking — ASVspoof 5, AudioSeal, WaveVerify ~75 minutes ↗06 / speech and audio Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards ~60 minutes ↗07 / transformers deep dive Why Transformers — The Problems with RNNs ~45 minutes ↗07 / transformers deep dive Self-Attention from Scratch ~90 minutes ↗07 / transformers deep dive Multi-Head Attention ~75 minutes ↗07 / transformers deep dive Positional Encoding — Sinusoidal, RoPE, ALiBi ~45 minutes ↗07 / transformers deep dive The Full Transformer — Encoder + Decoder ~75 minutes ↗07 / transformers deep dive BERT — Masked Language Modeling ~45 minutes ↗07 / transformers deep dive GPT — Causal Language Modeling ~75 minutes ↗07 / transformers deep dive T5, BART — Encoder-Decoder Models ~45 minutes ↗07 / transformers deep dive Vision Transformers (ViT) ~45 minutes ↗07 / transformers deep dive Audio Transformers — Whisper Architecture ~45 minutes ↗07 / transformers deep dive Mixture of Experts (MoE) ~45 minutes ↗07 / transformers deep dive KV Cache, Flash Attention & Inference Optimization ~75 minutes ↗07 / transformers deep dive Scaling Laws ~45 minutes ↗07 / transformers deep dive Build a Transformer from Scratch — The Capstone ~120 minutes ↗07 / transformers deep dive Attention Variants — Sliding Window, Sparse, Differential ~60 minutes ↗07 / transformers deep dive Speculative Decoding — Draft, Verify, Repeat ~60 minutes ↗08 / generative ai Generative Models — Taxonomy & History ~45 minutes ↗08 / generative ai Autoencoders & Variational Autoencoders (VAE) ~75 minutes ↗08 / generative ai GANs — Generator vs Discriminator ~75 minutes ↗08 / generative ai Conditional GANs & Pix2Pix ~75 minutes ↗08 / generative ai StyleGAN ~45 minutes ↗08 / generative ai Diffusion Models — DDPM from Scratch ~75 minutes ↗08 / generative ai Latent Diffusion & Stable Diffusion ~75 minutes ↗08 / generative ai ControlNet, LoRA & Conditioning ~75 minutes ↗08 / generative ai Inpainting, Outpainting & Image Editing ~75 minutes ↗08 / generative ai Video Generation ~45 minutes ↗08 / generative ai Audio Generation ~45 minutes ↗08 / generative ai 3D Generation ~45 minutes ↗08 / generative ai Flow Matching & Rectified Flows ~45 minutes ↗08 / generative ai Evaluation — FID, CLIP Score, Human Preference ~45 minutes ↗08 / generative ai Visual Autoregressive Modeling (VAR): Next-Scale Prediction ~90 minutes ↗09 / reinforcement learning MDPs, States, Actions & Rewards ~45 minutes ↗09 / reinforcement learning Dynamic Programming — Policy Iteration & Value Iteration ~75 minutes ↗09 / reinforcement learning Monte Carlo Methods — Learning from Complete Episodes ~75 minutes ↗09 / reinforcement learning Temporal Difference — Q-Learning & SARSA ~75 minutes ↗09 / reinforcement learning Deep Q-Networks (DQN) ~75 minutes ↗09 / reinforcement learning Policy Gradient — REINFORCE from Scratch ~75 minutes ↗09 / reinforcement learning Actor-Critic — A2C and A3C ~75 minutes ↗09 / reinforcement learning Proximal Policy Optimization (PPO) ~75 minutes ↗09 / reinforcement learning Reward Modeling & RLHF ~45 minutes ↗09 / reinforcement learning Multi-Agent RL ~45 minutes ↗09 / reinforcement learning Sim-to-Real Transfer ~45 minutes ↗09 / reinforcement learning RL for Games — AlphaZero, MuZero, and the LLM-Reasoning Era ~120 minutes ↗10 / llms from scratch Tokenizers: BPE, WordPiece, SentencePiece ~90 minutes ↗10 / llms from scratch Building a Tokenizer from Scratch ~90 minutes ↗10 / llms from scratch Data Pipelines for Pre-Training ~90 minutes ↗10 / llms from scratch Pre-Training a Mini GPT (124M Parameters) ~120 minutes ↗10 / llms from scratch Scaling: Distributed Training, FSDP, DeepSpeed ~120 minutes ↗10 / llms from scratch Instruction Tuning (SFT) ~90 minutes ↗10 / llms from scratch RLHF: Reward Model + PPO ~90 minutes ↗10 / llms from scratch DPO: Direct Preference Optimization ~90 minutes ↗10 / llms from scratch Constitutional AI and Self-Improvement ~45 minutes ↗10 / llms from scratch Evaluation: Benchmarks, Evals, LM Harness ~90 minutes ↗10 / llms from scratch Quantization: Making Models Fit ~120 minutes ↗10 / llms from scratch Inference Optimization ~120 minutes ↗10 / llms from scratch Building a Complete LLM Pipeline ~120 minutes ↗10 / llms from scratch Open Models: Architecture Walkthroughs ~45 minutes ↗10 / llms from scratch Speculative Decoding and EAGLE-3 ~75 minutes ↗10 / llms from scratch Differential Attention (V2) ~60 minutes ↗10 / llms from scratch Native Sparse Attention (DeepSeek NSA) ~60 minutes ↗10 / llms from scratch Multi-Token Prediction (MTP) ~60 minutes ↗10 / llms from scratch DualPipe Parallelism ~60 minutes ↗10 / llms from scratch DeepSeek-V3 Architecture Walkthrough ~75 minutes ↗10 / llms from scratch Jamba — Hybrid SSM-Transformer ~60 minutes ↗10 / llms from scratch Async and Hogwild! Inference ~60 minutes ↗10 / llms from scratch Speculative Decoding and EAGLE ~75 minutes ↗10 / llms from scratch Gradient Checkpointing and Activation Recomputation ~70 minutes ↗11 / llm engineering Prompt Engineering: Techniques & Patterns ~90 minutes ↗11 / llm engineering Few-Shot, Chain-of-Thought, Tree-of-Thought ~45 minutes ↗11 / llm engineering Structured Outputs: JSON, Schema Validation, Constrained Decoding ~90 minutes ↗11 / llm engineering Embeddings & Vector Representations ~75 minutes ↗11 / llm engineering Context Engineering: Windows, Budgets, Memory, and Retrieval ~90 minutes ↗11 / llm engineering RAG (Retrieval-Augmented Generation) ~90 minutes ↗11 / llm engineering Advanced RAG (Chunking, Reranking, Hybrid Search) ~90 minutes ↗11 / llm engineering Fine-Tuning with LoRA & QLoRA ~75 minutes ↗11 / llm engineering Function Calling & Tool Use ~75 minutes ↗11 / llm engineering Evaluation & Testing LLM Applications ~45 minutes ↗11 / llm engineering Caching, Rate Limiting & Cost Optimization ~45 minutes ↗11 / llm engineering Guardrails, Safety & Content Filtering ~45 minutes ↗11 / llm engineering Building a Production LLM Application ~120 minutes ↗11 / llm engineering Model Context Protocol (MCP) ~75 minutes ↗11 / llm engineering Prompt Caching and Context Caching ~60 minutes ↗11 / llm engineering Agent State Machines — Graphs, Nodes, Checkpoints ~75 minutes ↗11 / llm engineering Agent Framework Tradeoffs — Graph, Role, and Actor Orchestration ~45 minutes ↗12 / multimodal ai Vision Transformers and the Patch-Token Primitive ~120 minutes ↗12 / multimodal ai CLIP and Contrastive Vision-Language Pretraining ~180 minutes ↗12 / multimodal ai From CLIP to BLIP-2 — Q-Former as Modality Bridge ~180 minutes ↗12 / multimodal ai Flamingo and Gated Cross-Attention for Few-Shot VLMs ~120 minutes ↗12 / multimodal ai LLaVA and Visual Instruction Tuning ~180 minutes ↗12 / multimodal ai Any-Resolution Vision: Patch-n'-Pack and NaFlex ~120 minutes ↗12 / multimodal ai Open-Weight VLM Recipes: What Actually Matters ~180 minutes ↗12 / multimodal ai LLaVA-OneVision: Single-Image, Multi-Image, Video in One Model ~180 minutes ↗12 / multimodal ai Qwen-VL Family and Dynamic-FPS Video ~120 minutes ↗12 / multimodal ai InternVL3: Native Multimodal Pretraining ~120 minutes ↗12 / multimodal ai Chameleon and Early-Fusion Token-Only Multimodal Models ~180 minutes ↗12 / multimodal ai Emu3: Next-Token Prediction for Image and Video Generation ~120 minutes ↗12 / multimodal ai Transfusion: Autoregressive Text + Diffusion Image in One Transformer ~180 minutes ↗12 / multimodal ai Show-o and Discrete-Diffusion Unified Models ~120 minutes ↗12 / multimodal ai Janus-Pro: Decoupled Encoders for Unified Multimodal Models ~120 minutes ↗12 / multimodal ai MIO and Any-to-Any Streaming Multimodal Models ~120 minutes ↗12 / multimodal ai Video-Language Models: Temporal Tokens and Grounding ~180 minutes ↗12 / multimodal ai Long-Video Understanding at Million-Token Context ~180 minutes ↗12 / multimodal ai Audio-Language Models: the Whisper to Audio Flamingo 3 Arc ~180 minutes ↗12 / multimodal ai Omni Models: Qwen2.5-Omni and the Thinker-Talker Split ~180 minutes ↗12 / multimodal ai Embodied VLAs: RT-2, OpenVLA, π0, GR00T ~180 minutes ↗12 / multimodal ai Document and Diagram Understanding ~180 minutes ↗12 / multimodal ai ColPali and Vision-Native Document RAG ~180 minutes ↗12 / multimodal ai Multimodal RAG and Cross-Modal Retrieval ~180 minutes ↗12 / multimodal ai Multimodal Agents and Computer-Use (Capstone) ~240 minutes ↗13 / tools and protocols The Tool Interface — Why Agents Need Structured I/O ~45 minutes ↗13 / tools and protocols Function Calling Deep Dive — OpenAI, Anthropic, Gemini ~75 minutes ↗13 / tools and protocols Parallel Tool Calls and Streaming with Tools ~75 minutes ↗13 / tools and protocols Structured Output — JSON Schema, Pydantic, Zod, Constrained Decoding ~75 minutes ↗13 / tools and protocols Tool Schema Design — Naming, Descriptions, Parameter Constraints ~45 minutes ↗13 / tools and protocols MCP Fundamentals: Stateless Requests and JSON-RPC ~55 minutes ↗13 / tools and protocols Building an MCP Server: Stateless Python and TypeScript ~85 minutes ↗13 / tools and protocols Building an MCP Client: Discovery, Routing, and Dual-Era Fallback ~85 minutes ↗13 / tools and protocols MCP Transports: stdio and Stateless Streamable HTTP ~65 minutes ↗13 / tools and protocols MCP Resources and Prompts: Addressable Context for Stateless Servers ~60 minutes ↗13 / tools and protocols MCP Model Input: Sampling Migration and Stateless MRTR ~75 minutes ↗13 / tools and protocols Explicit Scope and Stateless Elicitation ~60 minutes ↗13 / tools and protocols MCP Tasks Extension: Durable Work on a Stateless Core ~90 minutes ↗13 / tools and protocols MCP Apps on the Stateless Protocol ~75 minutes ↗13 / tools and protocols MCP Security: Poisoned Metadata, Routing, and MRTR State ~60 minutes ↗13 / tools and protocols MCP Authorization: CIMD, Issuer Binding, PKCE, and Step-Up ~90 minutes ↗13 / tools and protocols Stateless MCP Gateways and Registry Admission ~75 minutes ↗13 / tools and protocols MCP Auth in Production: Issuer-Bound Enrollment and Tokens ~90 minutes ↗13 / tools and protocols A2A — Agent-to-Agent Protocol ~75 minutes ↗13 / tools and protocols OpenTelemetry GenAI — Tracing Tool Calls End-to-End ~75 minutes ↗13 / tools and protocols LLM Routing Layer — LiteLLM, OpenRouter, Portkey ~45 minutes ↗13 / tools and protocols Agent Skills: Portable Contract and Runtime Boundary ~90 minutes ↗13 / tools and protocols Capstone: Stateless Tool Ecosystem ~120 minutes ↗13 / tools and protocols Skill Discovery and Progressive Disclosure ~105 minutes ↗13 / tools and protocols Skill Invocation and Routing ~105 minutes ↗13 / tools and protocols Skill Permissions, Sandboxes, and Trust ~120 minutes ↗13 / tools and protocols Skill Evals, Packaging, and Portability ~150 minutes ↗13 / tools and protocols MCP Tool Contracts and Content ~120 minutes ↗13 / tools and protocols MCP Reliability, Cancellation, and Flow Control ~120 minutes ↗13 / tools and protocols MCP Registry Supply Chain: Admission, Drift, and Rollback ~90 minutes ↗13 / tools and protocols MCP Conformance Engineering: Versioning, Evidence, and Operations ~100 minutes ↗14 / agent engineering The Agent Loop: Observe, Think, Act ~60 minutes ↗14 / agent engineering ReWOO and Plan-and-Execute: Decoupled Planning ~60 minutes ↗14 / agent engineering Reflexion: Verbal Reinforcement Learning ~60 minutes ↗14 / agent engineering Tree of Thoughts and LATS: Deliberate Search ~75 minutes ↗14 / agent engineering Self-Refine and CRITIC: Iterative Output Improvement ~60 minutes ↗14 / agent engineering Tool Use and Function Calling ~60 minutes ↗14 / agent engineering Agent Memory — Virtual Context and Memory Paging ~75 minutes ↗14 / agent engineering Memory Blocks and Sleep-Time Compute ~75 minutes ↗14 / agent engineering Hybrid Memory: Vector + Graph + KV ~75 minutes ↗14 / agent engineering Skill Libraries and Lifelong Learning (Voyager) ~75 minutes ↗14 / agent engineering Planning with HTN and Evolutionary Search ~75 minutes ↗14 / agent engineering Anthropic's Workflow Patterns: Simple Over Complex ~60 minutes ↗14 / agent engineering Stateful Graph Orchestration — Durable Execution and Checkpoints ~75 minutes ↗14 / agent engineering The Actor Model for Agents — Async Messages and Typed Runtimes ~75 minutes ↗14 / agent engineering Role-Based Agent Teams — Roles, Tasks, Processes ~75 minutes ↗14 / agent engineering OpenAI Agents SDK: Handoffs, Guardrails, Tracing ~75 minutes ↗14 / agent engineering The Harness as a Library — Subagents and Session Store ~75 minutes ↗14 / agent engineering Production Agent Runtimes — Fast Instantiation and Typed Workflows ~45 minutes ↗14 / agent engineering Benchmarks: SWE-bench, GAIA, AgentBench ~60 minutes ↗14 / agent engineering Benchmarks: WebArena and OSWorld ~60 minutes ↗14 / agent engineering Computer Use: Claude, OpenAI CUA, Gemini ~60 minutes ↗14 / agent engineering Voice Agents: Pipecat and LiveKit ~60 minutes ↗14 / agent engineering OpenTelemetry GenAI Semantic Conventions ~60 minutes ↗14 / agent engineering Agent Observability: Langfuse, Phoenix, Opik ~45 minutes ↗14 / agent engineering Multi-Agent Debate and Collaboration ~60 minutes ↗14 / agent engineering Failure Modes: Why Agents Break ~60 minutes ↗14 / agent engineering Prompt Injection and the PVE Defense ~75 minutes ↗14 / agent engineering Orchestration Patterns: Supervisor, Swarm, Hierarchical ~60 minutes ↗14 / agent engineering Production Runtimes: Queue, Event, Cron ~60 minutes ↗14 / agent engineering Eval-Driven Agent Development ~60 minutes ↗14 / agent engineering Agent Workbench Engineering: Why Capable Models Still Fail ~45 minutes ↗14 / agent engineering The Minimal Agent Workbench ~45 minutes ↗14 / agent engineering Agent Instructions as Executable Constraints ~50 minutes ↗14 / agent engineering Repo Memory and Durable State ~60 minutes ↗14 / agent engineering Initialization Scripts for Agents ~45 minutes ↗14 / agent engineering Scope Contracts and Task Boundaries ~50 minutes ↗14 / agent engineering Runtime Feedback Loops ~50 minutes ↗14 / agent engineering Verification Gates ~55 minutes ↗14 / agent engineering Reviewer Agent: Separate Builder from Marker ~55 minutes ↗14 / agent engineering Multi-Session Handoff ~50 minutes ↗14 / agent engineering The Workbench on a Real Repo ~60 minutes ↗14 / agent engineering Capstone: Ship a Reusable Agent Workbench Pack ~75 minutes ↗14 / agent engineering Frame the Task Before the Agent Writes Code ~60 minutes ↗14 / agent engineering Build an Evidence-Backed Execution Plan ~65 minutes ↗14 / agent engineering Delegate Agent Work with Isolation and Merge Contracts ~70 minutes ↗14 / agent engineering Turn Every Agent Correction into a System Improvement ~65 minutes ↗14 / agent engineering Define the Outcome Before You Choose the Output ~60 minutes ↗14 / agent engineering Discover the Workflow People Actually Perform ~70 minutes ↗14 / agent engineering Map Assumptions and Resolve the Riskiest One First ~65 minutes ↗14 / agent engineering Choose the Smallest Slice That Can Change the Decision ~65 minutes ↗14 / agent engineering Write Specifications That Preserve Judgment ~75 minutes ↗14 / agent engineering Design Success Metrics Before the Result Exists ~70 minutes ↗14 / agent engineering Choose Prototype, Pilot, or Production Deliberately ~70 minutes ↗14 / agent engineering Build a Feedback Ratchet with Ownership and Retirement ~75 minutes ↗15 / autonomous systems The Shift from Chatbots to Long-Horizon Agents ~45 minutes ↗15 / autonomous systems STaR, V-STaR, Quiet-STaR — Self-Taught Reasoning ~60 minutes ↗15 / autonomous systems AlphaEvolve — Evolutionary Coding Agents ~60 minutes ↗15 / autonomous systems Darwin Godel Machine — Open-Ended Self-Modifying Agents ~60 minutes ↗15 / autonomous systems AI Scientist v2 — Workshop-Level Autonomous Research ~60 minutes ↗15 / autonomous systems Automated Alignment Research (Anthropic AAR) ~60 minutes ↗15 / autonomous systems Recursive Self-Improvement — Capability vs Alignment ~60 minutes ↗15 / autonomous systems Bounded Self-Improvement Designs ~60 minutes ↗15 / autonomous systems The Autonomous Coding Agent Landscape (2026) ~45 minutes ↗15 / autonomous systems Permission Modes for Autonomous Agents ~45 minutes ↗15 / autonomous systems Browser Agents and Long-Horizon Web Tasks ~45 minutes ↗15 / autonomous systems Long-Running Background Agents: Durable Execution ~60 minutes ↗15 / autonomous systems Action Budgets, Iteration Caps, and Cost Governors ~60 minutes ↗15 / autonomous systems Kill Switches, Circuit Breakers, and Canary Tokens ~60 minutes ↗15 / autonomous systems Human-in-the-Loop: Propose-Then-Commit ~60 minutes ↗15 / autonomous systems Checkpoints and Rollback ~60 minutes ↗15 / autonomous systems Constitutional AI and Rule Overrides ~60 minutes ↗15 / autonomous systems Llama Guard and Input/Output Classification ~45 minutes ↗15 / autonomous systems Anthropic Responsible Scaling Policy v3.0 ~45 minutes ↗15 / autonomous systems OpenAI Preparedness Framework and DeepMind Frontier Safety Framework ~45 minutes ↗15 / autonomous systems METR Time Horizons and External Capability Evaluation ~60 minutes ↗15 / autonomous systems CAIS, CAISI, and Societal-Scale Risk ~45 minutes ↗16 / multi agent and swarms Why Multi-Agent? ~60 minutes ↗16 / multi agent and swarms Heritage of FIPA-ACL and Speech Acts ~60 minutes ↗16 / multi agent and swarms Communication Protocols ~120 minutes ↗16 / multi agent and swarms The Multi-Agent Primitive Model ~60 minutes ↗16 / multi agent and swarms Supervisor / Orchestrator-Worker Pattern ~75 minutes ↗16 / multi agent and swarms Hierarchical Architecture and Its Failure Mode ~60 minutes ↗16 / multi agent and swarms Society of Mind and Multi-Agent Debate ~60 minutes ↗16 / multi agent and swarms Role Specialization — Planner, Critic, Executor, Verifier ~60 minutes ↗16 / multi agent and swarms Parallel / Swarm / Networked Architectures ~75 minutes ↗16 / multi agent and swarms Group Chat and Speaker Selection ~60 minutes ↗16 / multi agent and swarms Handoffs and Routines — Stateless Orchestration ~60 minutes ↗16 / multi agent and swarms A2A — The Agent-to-Agent Protocol ~75 minutes ↗16 / multi agent and swarms Shared Memory and Blackboard Patterns ~75 minutes ↗16 / multi agent and swarms Consensus and Byzantine Fault Tolerance for Agents ~75 minutes ↗16 / multi agent and swarms Voting, Self-Consistency, and Debate Topology ~75 minutes ↗16 / multi agent and swarms Negotiation and Bargaining ~75 minutes ↗16 / multi agent and swarms Generative Agents and Emergent Simulation ~75 minutes ↗16 / multi agent and swarms Theory of Mind and Emergent Coordination ~75 minutes ↗16 / multi agent and swarms Swarm Optimization for LLMs (PSO, ACO) ~75 minutes ↗16 / multi agent and swarms MARL — MADDPG, QMIX, MAPPO ~90 minutes ↗16 / multi agent and swarms Agent Economies, Token Incentives, Reputation ~75 minutes ↗16 / multi agent and swarms Production Scaling — Queues, Checkpoints, Durability ~75 minutes ↗16 / multi agent and swarms Failure Modes — MAST, Groupthink, Monoculture, Cascading Errors ~75 minutes ↗16 / multi agent and swarms Evaluation and Coordination Benchmarks ~75 minutes ↗16 / multi agent and swarms Case Studies and the 2026 State of the Art ~90 minutes ↗17 / infrastructure and production Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAI ~60 minutes ↗17 / infrastructure and production Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale ~60 minutes ↗17 / infrastructure and production GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling ~75 minutes ↗17 / infrastructure and production Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill ~75 minutes ↗17 / infrastructure and production EAGLE-3 Speculative Decoding in Production ~60 minutes ↗17 / infrastructure and production Prefix-Cache Serving — RadixAttention and KV Reuse ~75 minutes ↗17 / infrastructure and production Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell ~75 minutes ↗17 / infrastructure and production Inference Metrics — TTFT, TPOT, ITL, Goodput, P99 ~60 minutes ↗17 / infrastructure and production Production Quantization — AWQ, GPTQ, GGUF K-quants, FP8, MXFP4/NVFP4 ~75 minutes ↗17 / infrastructure and production Cold Start Mitigation for Serverless LLMs ~60 minutes ↗17 / infrastructure and production Multi-Region LLM Serving and KV Cache Locality ~60 minutes ↗17 / infrastructure and production Edge Inference — Apple Neural Engine, Qualcomm Hexagon, WebGPU/WebLLM, Jetson ~60 minutes ↗17 / infrastructure and production LLM Observability Stack Selection ~60 minutes ↗17 / infrastructure and production Prompt Caching and Semantic Caching Economics ~60 minutes ↗17 / infrastructure and production Batch APIs — the 50% Discount as Industry Standard ~45 minutes ↗17 / infrastructure and production Model Routing as a Cost-Reduction Primitive ~60 minutes ↗17 / infrastructure and production Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d ~75 minutes ↗17 / infrastructure and production Production Serving Stack — KV Offloading and Cache-Aware Routing ~60 minutes ↗17 / infrastructure and production AI Gateways — LiteLLM, Portkey, Kong AI Gateway, Bifrost ~60 minutes ↗17 / infrastructure and production Shadow Traffic, Canary Rollout, and Progressive Deployment for LLMs ~60 minutes ↗17 / infrastructure and production A/B Testing LLM Features — GrowthBook, Statsig, and the Vibes Problem ~60 minutes ↗17 / infrastructure and production Load Testing LLM APIs — Why k6 and Locust Lie ~75 minutes ↗17 / infrastructure and production SRE for AI — Multi-Agent Incident Response, Runbooks, Predictive Detection ~60 minutes ↗17 / infrastructure and production Chaos Engineering for LLM Production ~60 minutes ↗17 / infrastructure and production Security — Secrets, API Key Rotation, Audit Logs, Guardrails ~60 minutes ↗17 / infrastructure and production Compliance — SOC 2, HIPAA, GDPR, PCI-DSS, EU AI Act, ISO 42001 ~60 minutes ↗17 / infrastructure and production FinOps for LLMs — Unit Economics and Multi-Tenant Attribution ~60 minutes ↗17 / infrastructure and production Self-Hosted Serving Selection — Matching Engine to Hardware and Scale ~45 minutes ↗18 / ethics safety alignment Instruction-Following as Alignment Signal ~45 minutes ↗18 / ethics safety alignment Reward Hacking and Goodhart's Law ~60 minutes ↗18 / ethics safety alignment The Direct Preference Optimization Family ~75 minutes ↗18 / ethics safety alignment Sycophancy as RLHF Amplification ~60 minutes ↗18 / ethics safety alignment Constitutional AI and RLAIF ~60 minutes ↗18 / ethics safety alignment Mesa-Optimization and Deceptive Alignment ~75 minutes ↗18 / ethics safety alignment Sleeper Agents — Persistent Deception ~60 minutes ↗18 / ethics safety alignment In-Context Scheming in Frontier Models ~60 minutes ↗18 / ethics safety alignment Alignment Faking ~60 minutes ↗18 / ethics safety alignment AI Control — Safety Despite Subversion ~75 minutes ↗18 / ethics safety alignment Scalable Oversight and Weak-to-Strong Generalization ~60 minutes ↗18 / ethics safety alignment Red-Teaming: PAIR and Automated Attacks ~75 minutes ↗18 / ethics safety alignment Many-Shot Jailbreaking ~45 minutes ↗18 / ethics safety alignment ASCII Art and Visual Jailbreaks ~60 minutes ↗18 / ethics safety alignment Indirect Prompt Injection — Production Attack Surface ~75 minutes ↗18 / ethics safety alignment Red-Team Tooling — Garak, Llama Guard, PyRIT ~75 minutes ↗18 / ethics safety alignment WMDP and Dual-Use Capability Evaluation ~60 minutes ↗18 / ethics safety alignment Frontier Safety Frameworks — RSP, PF, FSF ~75 minutes ↗18 / ethics safety alignment Anthropic's Model Welfare Program ~45 minutes ↗18 / ethics safety alignment Bias and Representational Harm in LLMs ~60 minutes ↗18 / ethics safety alignment Fairness Criteria — Group, Individual, Counterfactual ~60 minutes ↗18 / ethics safety alignment Differential Privacy for LLMs ~60 minutes ↗18 / ethics safety alignment Watermarking — SynthID, Stable Signature, C2PA ~75 minutes ↗18 / ethics safety alignment Regulatory Frameworks — EU, US, UK, Korea ~75 minutes ↗18 / ethics safety alignment EchoLeak and the Emergence of CVEs for AI ~45 minutes ↗18 / ethics safety alignment Model, System, and Dataset Cards ~60 minutes ↗18 / ethics safety alignment Data Provenance and Training-Data Governance ~60 minutes ↗18 / ethics safety alignment Alignment Research Ecosystem — MATS, Redwood, Apollo, METR ~45 minutes ↗18 / ethics safety alignment Moderation Systems — OpenAI, Perspective, Llama Guard ~60 minutes ↗18 / ethics safety alignment Dual-Use Risk — Cyber, Bio, Chem, Nuclear Uplift ~75 minutes ↗19 / capstone projects Capstone 01 — Terminal-Native Coding Agent 35 hours ↗19 / capstone projects Capstone 02 — RAG over Codebase (Cross-Repo Semantic Search) 30 hours ↗19 / capstone projects Capstone 03 — Real-Time Voice Assistant (ASR to LLM to TTS) 30 hours ↗19 / capstone projects Capstone 04 — Multimodal Document QA (Vision-First PDF, Tables, Charts) 30 hours ↗19 / capstone projects Capstone 05 — Autonomous Research Agent (AI-Scientist Class) 40 hours ↗19 / capstone projects Capstone 06 — DevOps Troubleshooting Agent for Kubernetes 30 hours ↗19 / capstone projects Capstone 07 — End-to-End Fine-Tuning Pipeline (Data to SFT to DPO to Serve) 35 hours ↗19 / capstone projects Capstone 08 — Production RAG Chatbot for a Regulated Vertical 30 hours ↗19 / capstone projects Capstone 09 — Code Migration Agent (Repo-Level Language / Runtime Upgrade) 30 hours ↗19 / capstone projects Capstone 10 — Multi-Agent Software Engineering Team 40 hours ↗19 / capstone projects Capstone 11 — LLM Observability & Eval Dashboard 25 hours ↗19 / capstone projects Capstone 12 — Video Understanding Pipeline (Scene, QA, Search) 30 hours ↗19 / capstone projects Capstone 13: Stateless MCP Server with Registry and Governance ~25 hours ↗19 / capstone projects Capstone 14 — Speculative-Decoding Inference Server 30 hours ↗19 / capstone projects Capstone 15 — Constitutional Safety Harness + Red-Team Range 25 hours ↗19 / capstone projects Capstone 16 — GitHub Issue-to-PR Autonomous Agent 30 hours ↗19 / capstone projects Capstone 17 — Personal AI Tutor (Adaptive, Multimodal, with Memory) 30 hours ↗19 / capstone projects Agent Harness Loop Contract ~90 minutes ↗19 / capstone projects Tool Registry with Schema Validation ~90 minutes ↗19 / capstone projects JSON-RPC 2.0 Over Newline-Delimited Stdio ~90 minutes ↗19 / capstone projects Function Call Dispatcher ~90 minutes ↗19 / capstone projects Plan-Execute Control Flow ~90 minutes ↗19 / capstone projects Capstone Lesson 25: Verification Gates and the Observation Budget ~90 minutes ↗19 / capstone projects Capstone Lesson 26: Sandbox Runner with Denylist and Path Jail ~90 minutes ↗19 / capstone projects Capstone Lesson 27: Eval Harness with Fixture Tasks ~90 minutes ↗19 / capstone projects Capstone Lesson 28: Observability with OTel GenAI Spans and Prometheus Metrics ~90 minutes ↗19 / capstone projects Capstone Lesson 29: End-to-End Coding Agent on the Harness ~90 minutes ↗19 / capstone projects BPE Tokenizer From Scratch ~90 minutes ↗19 / capstone projects Tokenized Dataset with Sliding Window ~90 minutes ↗19 / capstone projects Token and Positional Embeddings ~90 minutes ↗19 / capstone projects Multi-Head Self-Attention ~90 minutes ↗19 / capstone projects Transformer Block from Scratch ~90 minutes ↗19 / capstone projects GPT Model Assembly ~90 minutes ↗19 / capstone projects Training Loop and Evaluation ~90 minutes ↗19 / capstone projects Loading Pretrained Weights ~90 minutes ↗19 / capstone projects Capstone Lesson 38: Classifier Fine-Tuning by Head Swap ~90 minutes ↗19 / capstone projects Capstone Lesson 39: Instruction Tuning by Supervised Fine-Tuning ~90 minutes ↗19 / capstone projects Capstone Lesson 40: Direct Preference Optimization from Scratch ~90 minutes ↗19 / capstone projects Capstone Lesson 41: Full Evaluation Pipeline ~90 minutes ↗19 / capstone projects Large Corpus Downloader ~90 minutes ↗19 / capstone projects HDF5 Tokenized Corpus ~90 minutes ↗19 / capstone projects Cosine LR with Linear Warmup ~90 minutes ↗19 / capstone projects Gradient Clipping and Mixed Precision ~90 minutes ↗19 / capstone projects Gradient Accumulation ~90 minutes ↗19 / capstone projects Checkpoint Save and Resume ~90 minutes ↗19 / capstone projects Distributed Data Parallel and FSDP from Scratch ~90 minutes ↗19 / capstone projects Language Model Evaluation Harness ~90 minutes ↗19 / capstone projects Hypothesis Generator ~90 minutes ↗19 / capstone projects Literature Retrieval ~90 minutes ↗19 / capstone projects Experiment Runner ~90 minutes ↗19 / capstone projects Result Evaluator ~90 minutes ↗19 / capstone projects Paper Writer ~90 minutes ↗19 / capstone projects Critic Loop ~90 minutes ↗19 / capstone projects Iteration Scheduler ~90 minutes ↗19 / capstone projects End-to-End Research Demo ~90 minutes ↗19 / capstone projects Vision Encoder Patches ~90 minutes ↗19 / capstone projects Vision Transformer Encoder ~90 minutes ↗19 / capstone projects Projection Layer for Modality Alignment ~90 minutes ↗19 / capstone projects Cross-Attention Fusion ~90 minutes ↗19 / capstone projects Vision-Language Pretraining ~90 minutes ↗19 / capstone projects Multimodal Evaluation ~90 minutes ↗19 / capstone projects Chunking Strategies, Compared ~90 minutes ↗19 / capstone projects Hybrid Retrieval with BM25 and Dense Embeddings ~90 minutes ↗19 / capstone projects Cross-Encoder Reranker ~90 minutes ↗19 / capstone projects Query Rewriting: HyDE, Multi-Query, and Decomposition ~90 minutes ↗19 / capstone projects RAG Evaluation: Precision, Recall, MRR, nDCG, Faithfulness, Answer Relevance ~90 minutes ↗19 / capstone projects End-to-End RAG System ~90 minutes ↗19 / capstone projects Task Spec Format ~90 min ↗19 / capstone projects Classical Metrics ~90 min ↗19 / capstone projects Code Exec Metric ~90 min ↗19 / capstone projects Perplexity and Calibration ~90 min ↗19 / capstone projects Leaderboard Aggregation ~90 min ↗19 / capstone projects End-to-End Eval Runner ~90 min ↗19 / capstone projects Collective Ops From Scratch ~90 min ↗19 / capstone projects Data Parallel DDP From Scratch ~90 min ↗19 / capstone projects ZeRO Optimizer State Sharding ~90 min ↗19 / capstone projects Pipeline Parallel and Bubble Analysis ~90 min ↗19 / capstone projects Sharded Checkpoint and Atomic Resume ~90 min ↗19 / capstone projects End-to-End Distributed Training ~90 min ↗19 / capstone projects Capstone 82 — Jailbreak Taxonomy ~90 min ↗19 / capstone projects Capstone 83 — Prompt Injection Detector ~90 min ↗19 / capstone projects Capstone 84 — Refusal Evaluation ~90 min ↗19 / capstone projects Capstone 85 — Content Classifier Integration ~90 min ↗19 / capstone projects Capstone 86 — Constitutional Rules Engine ~90 min ↗19 / capstone projects Capstone 87 — End-to-End Safety Gate ~90 min ↗