Production AI Agents
Studies in this cluster, in series order. Each one keeps its own URL.
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
Production AI Agents
6 studies- 1.Production AI Agents — Architecture, Scaling, Caching & ReliabilityA production AI agent is a policy loop: model, tools, memory, and guardrails, with durable state, cost and latency SLOs, and cache layers. It is not a chat box glued to one completion. This hub maps architecture, memory, scale-out, caching, and reliability, and only cross-links the RAG, structured-outputs, and inference hubs.
- 2.Agent Architecture — Planner/Executor, Tool Registry & Control LoopsProduction agents need an explicit control architecture: who plans, who executes, how tools are registered and validated, and where humans gate irreversible actions. Compare ReAct, plan-then-act, and a graph, and point at structured outputs for argument shape without re-teaching constrained decoding.
- 3.Agent Memory — Working Context, Long-term Store, Compaction & SessionsAgent memory is a budgeted systems problem. Working context fits in the window. Episodes and semantic facts live outside it. Compaction must keep goals and commitments and must not leak PII across sessions. Retrieval of documents is a different hub.
- 4.Scaling Agents — Concurrency, Queues, Fan-out, Rate Limits & Cost BudgetsScaling agents is a distributed-systems problem: horizontal workers, job queues, per-tenant quotas, backpressure, and idempotent tool side effects. Every parallel tool call is a budgeted fork-join. Fan-out multiplies capability and cost together.
- 5.Caching for Agents — Prompt Cache, KV Reuse, Semantic Cache & Tool Result CacheThe wrong agent cache is a correctness bug. Separate provider prompt cache, runtime KV reuse, semantic response cache, and tool-result memoization. This page is the application policy. PagedAttention and the KV layout stay on the inference hub.
- 6.Agent Reliability — Evals, Guardrails, Tracing & Failure ModesAgents fail differently from chatbots: stuck loops, poison tool output, partial side effects, and trajectory drift. Reliability means grading whole runs, sandboxing tools, circuit breakers, and span-level traces. A fluent final string is not a pass.