<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Vibe Engines — AI Agents &amp; Tools</title>
    <link>https://vibeengines.com/topic/ai-agents</link>
    <atom:link href="https://vibeengines.com/topic/ai-agents/feed.xml" rel="self" type="application/rss+xml" />
    <description>Agentic systems — tool calling, orchestration and the interfaces that let models act — with the agent system design, the agentic interview handbook, and a tool-schema designer.</description>
    <language>en</language>
    <lastBuildDate>Thu, 24 Sep 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>Ship or Stop</title>
      <link>https://vibeengines.com/challenge/ship-or-stop</link>
      <guid isPermaLink="true">https://vibeengines.com/challenge/ship-or-stop</guid>
      <category>Challenge</category>
      <description>A self-improvement loop finished another round and every number on the dashboard went up. Write the gate that decides whether it ships: six stop conditions over two rounds of measurements, plus a canary nobody is allowed to read early. The order of the answers is part of the contract.</description>
      <pubDate>Thu, 24 Sep 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>SOC Analyst Roadmap</title>
      <link>https://vibeengines.com/roadmap/soc-analyst</link>
      <guid isPermaLink="true">https://vibeengines.com/roadmap/soc-analyst</guid>
      <category>Roadmap</category>
      <description>A step-by-step roadmap to become a security operations center (SOC) analyst in 2026 — the defensive, blue-team path. From what a SOC does, networking for defenders and the attacker playbook (MITRE ATT&amp;CK) through logs, SIEM, alert triage, investigation and incident response to SOAR automation, the AI-assisted SOC, detecting AI-era threats and the analyst career. 18 stations across 3 tracks — Security Foundations, Detection &amp; Response, The Modern &amp; AI-Era SOC.</description>
      <pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agentic Payments</title>
      <link>https://vibeengines.com/handbook/agentic-payments</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/agentic-payments</guid>
      <category>Handbook</category>
      <description>How to let an agent pay without letting it drain your account. The x402 machine-native checkout (HTTP 402: server returns payment terms, agent pays and retries with proof), signed MANDATES that scope spending like OAuth scopes for money (AP2), and spend caps enforced BEFORE execution — authorize ⟺ amount ≤ per-txn AND spent+amount ≤ total AND merchant ∈ allowed. The runaway-agent threat and why the cap belongs in the payment layer, not the prompt. Worked math plus a runnable mandate-and-cap enforcer that stops a looping agent cold at its budget.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agent Supply-Chain Security</title>
      <link>https://vibeengines.com/handbook/agent-supply-chain-security</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/agent-supply-chain-security</guid>
      <category>Handbook</category>
      <description>The other door into your agent — the one you open yourself. Every skill and MCP server you INSTALL is third-party code running with permissions, an install-time threat distinct from runtime prompt injection. Why compromise risk compounds with dependency count (1−(1−p)^N: 20 deps at 5% each ≈ 64%), how blast radius = the capabilities a compromised component holds and least-privilege (grant = needed ∩ offered) bounds it, and the two levers — vet to lower p, sandbox to bound damage. Worked math plus a runnable risk-and-blast-radius model.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>MCP Apps</title>
      <link>https://vibeengines.com/handbook/mcp-apps</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/mcp-apps</guid>
      <category>Handbook</category>
      <description>Interactive UIs inside Claude and ChatGPT — the first official UI extension to the Model Context Protocol. How a third-party UI runs as a SANDBOXED IFRAME (no ambient authority — can't touch the host DOM, storage, or network), why every capability call crosses a postMessage JSON-RPC bridge (method + params + id, response with the same id), and how the HOST MEDIATES each request against an allowlist (method ∈ allowlist ? execute : error −32601) so the app can only do what the host chose to expose. Sandbox + bridge + allowlist = the whole security model. Distinct from the MCP-server backend. Worked protocol plus a runnable host-mediation gatekeeper.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Self-Improving Agents</title>
      <link>https://vibeengines.com/handbook/self-improving-agents</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/self-improving-agents</guid>
      <category>Handbook</category>
      <description>The long one: how an agent gets better at its own job, end to end. COVERAGE vs the accuracy you can ship (a model usually already produces a right answer somewhere in its output; the missing piece is recognising it), unbiased pass@k, why the sampling curve is straight, the SELECTION bottleneck, building a VERIFIER and the two asymmetric ways it is wrong, feedback that is not a verifier ranked by independence, the five-predicate filter that decides what becomes training data, the SIX substrates a lesson can land in (context, memory, tools, scaffold, verifier, weights) and how to choose, STaR through PPO/GRPO/DPO and reward hacking, why a loop refit on its own output COLLAPSES toward its own habits, search as self-improvement, and measuring a time HORIZON rather than a score. 14 sections, 13 runnable Python blocks, 3 interactive models, 13 diagrams, 5 checkpoints and a running build log. Every figure carries the condition it was measured under.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Long-Horizon Agents</title>
      <link>https://vibeengines.com/handbook/long-horizon-agents</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/long-horizon-agents</guid>
      <category>Handbook</category>
      <description>Why a 10-step task the agent nails becomes a 100-step task it never finishes — reliability COMPOUNDS: a single run of k steps succeeds with probability p^k (95% per step, 100 steps = 0.6%). The fix is engineering, not a better model: CHECKPOINT progress so a failure doesn't restart from zero, and RESUME + retry the failed segment only — a c-step segment retried r times succeeds 1−(1−p^c)^r, so the whole task climbs to [1−(1−p^c)^r]^(k/c), turning that 0.6% into ~51% at the same per-step reliability. The METR time-horizon idea and how reliability engineering extends it. The pass@k-vs-pass^k measurement lives in the agent-evals handbook (linked, not re-derived). Worked math plus a runnable no-checkpoint-vs-checkpointed model.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agent Skills vs MCP</title>
      <link>https://vibeengines.com/handbook/agent-skills-vs-mcp</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/agent-skills-vs-mcp</guid>
      <category>Handbook</category>
      <description>Agent Skills vs the Model Context Protocol, decided by what each adds: a Skill is packaged knowledge and procedure loaded into the agent’s context to change how it behaves; MCP is a protocol that grants live tools and data over a wire at runtime. Portability, token cost, security surface, and why the two compose rather than compete.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Vibe Coding vs Spec-Driven Development</title>
      <link>https://vibeengines.com/handbook/vibe-coding-vs-spec-driven-development</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/vibe-coding-vs-spec-driven-development</guid>
      <category>Handbook</category>
      <description>Vibe coding vs spec-driven development, two ways to build with AI: vibe coding is conversational and exploratory — prompt, run, keep what works; spec-driven development writes a precise spec first and has the agent implement against it with tests as the contract. The spec→plan→tasks→implement loop, where each breaks, and how to combine them.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Claude Code vs Codex vs Gemini CLI</title>
      <link>https://vibeengines.com/handbook/claude-code-vs-codex-vs-gemini-cli</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/claude-code-vs-codex-vs-gemini-cli</guid>
      <category>Handbook</category>
      <description>Three terminal-native agentic coding tools from Anthropic, OpenAI and Google that share one loop — read the repo, plan, edit files, run commands, iterate. How they differ on model, extensibility, openness and permissions, what the reported mid-2026 market numbers say (handled with care), and how to actually choose.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an Agent Payment Gateway</title>
      <link>https://vibeengines.com/ai-system-design/agent-payment-gateway-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/agent-payment-gateway-system-design</guid>
      <category>AI System Design</category>
      <description>Build the platform that lets autonomous agents pay for things mid-task without draining an account — the agent rails, not the human card rails. See why a direct payment credential gives an agent unbounded, unscoped spending power, routing every payment through one mediated gateway, verifying a signed spending mandate (signature, scope, expiry) instead of trusting the agent's claim, enforcing per-transaction and cumulative spend caps BEFORE any settlement, machine-native x402 (HTTP 402) settlement with no human checkout, an immutable audit trail of every decision, and why the gateway must fail closed — so a runaway or hijacked agent is capped at the mandate, not the credit limit.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an MCP Security Gateway</title>
      <link>https://vibeengines.com/ai-system-design/mcp-security-gateway-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/mcp-security-gateway-system-design</guid>
      <category>AI System Design</category>
      <description>Build the security layer between an AI host and the MCP servers it connects to — the MCP-protocol-specific threat surface, not a generic LLM guardrail. See why a direct connection implicitly trusts untrusted third-party servers, routing all MCP traffic through a mediating gateway, an explicit server/tool allowlist (deny by default), scanning tool descriptions for injection (attacker-controlled text the model reads as instructions — a channel unique to MCP), issuing narrowly-scoped short-lived tokens per server so a compromised one is bounded to its slice, egress control against exfiltration through tool results, an immutable audit trail, and why the gateway must fail closed.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an Agentic Browser Security Gateway</title>
      <link>https://vibeengines.com/ai-system-design/agentic-browser-security-gateway-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/agentic-browser-security-gateway-system-design</guid>
      <category>AI System Design</category>
      <description>Build the enterprise perimeter that contains a whole fleet of agentic browsers without banning them — the fleet-level layer, not one agent's own injection defenses. See why ungoverned agentic browsers are an invisible exfiltration surface (and why a ban just creates shadow usage), routing every action through one gateway, a central org-wide policy engine, taint tracking so data from an untrusted web page can't drive a sensitive action (provenance beats detection), an action allowlist with risk classification, human-approval breakpoints for the irreversible minority, and a fleet-wide immutable audit trail.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design a Hybrid Edge-Cloud Agent</title>
      <link>https://vibeengines.com/ai-system-design/hybrid-edge-cloud-agent-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/hybrid-edge-cloud-agent-system-design</guid>
      <category>AI System Design</category>
      <description>Build an on-device agent that runs on the NPU by default and escalates hard queries to the cloud — but where privacy, not just confidence, decides what may leave the device. See why default-to-cloud is default-to-leak, answering on-device by default, a privacy classifier that gates escalation FIRST (must-stay-local data is answered locally even when the model is unsure), confidence-based escalation only for privacy-cleared queries (the cascade math lives in the LLM Router design), redaction to minimize what leaves, and an on-device escalation audit — capability and privacy reconciled by making the boundary a hard constraint confidence can never override.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Write a Verifier</title>
      <link>https://vibeengines.com/challenge/write-a-verifier</link>
      <guid isPermaLink="true">https://vibeengines.com/challenge/write-a-verifier</guid>
      <category>Challenge</category>
      <description>In RL from verifiable rewards, the verifier IS the reward — and a gameable one is worse than none, because every false positive is a lie the model learns. Write a verifier a reward-hacking agent can’t fool: recompute the score from ground truth, demand exact task coverage, and compare types-intact. Solve it in Python or TypeScript, with hidden tests.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Build Your Own Agent Loop</title>
      <link>https://vibeengines.com/challenge/build-your-own-agent-loop</link>
      <guid isPermaLink="true">https://vibeengines.com/challenge/build-your-own-agent-loop</guid>
      <category>Challenge</category>
      <description>Strip an &quot;AI agent&quot; of its mystique and what’s left is a loop: read the model’s next move — tool call or final answer — run the tool, feed the observation back, repeat until it answers or the step cap trips. Implement the whole ReAct inner loop as one pure function. Solve it in Python or TypeScript, with hidden tests.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>LLMs Get Lost in Multi-Turn Conversation</title>
      <link>https://vibeengines.com/paper/llms-lost-multi-turn</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/llms-lost-multi-turn</guid>
      <category>Paper</category>
      <description>The ICLR 2026 Outstanding Paper that measured a failure everyone had felt: give a top model a fully-specified task in ONE prompt (concat) and it shines; split the IDENTICAL requirements across several turns (sharded) and 15 leading LLMs fall apart — ~39% average drop. The decomposition is the punchline: APTITUDE (best-case ability) barely moves, but RELIABILITY craters — the spread between a model's best and worst runs roughly DOUBLES. The mechanism is premature commitment: handed partial info, the model guesses a full answer early, locks it in, and when a later turn contradicts it &quot;gets lost and does not recover.&quot; A small decay model captures the shape — concat R=p (flat), sharded R=p·(1−q)^(turns−1) (geometric decay), so longer chats get less reliable and the whole loss lives in a reliability term while the ceiling p is untouched. Fixes: front-load the spec, or consolidate/recap a long chat back into one fresh full-spec message. Worked math plus runnable code.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>GPT-5.6 System Card</title>
      <link>https://vibeengines.com/paper/gpt-5-6-system-card</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/gpt-5-6-system-card</guid>
      <category>Paper</category>
      <description>A modern system card is really a COMPUTE-ALLOCATION POLICY, not one model. TIERED ROUTING (reported Sol/Terra/Luna) sorts each query by difficulty — cheap fast tier by default, escalate the hard cases (the cascade cost math lives in the model-routing handbook — linked, not re-derived). ULTRA MODE spends test-time compute in PARALLEL: run several independent agents (reported 4) on the hardest task and a VERIFIER keeps any that solves, so the solve rate is 1−(1−p)^N. The fresh angle = why FOUR, not forty: the marginal value of the N-th agent is Δ(N)=(1−p)^(N−1)·p — geometric decay (p=0.6: +0.60/+0.24/+0.10/+0.04) while cost is linear in N, so there's a knee. Runnable proves best-of-N rises, marginal=(1−p)^(N−1)·p, diminishing returns, 4th≪1st, value/cost decreasing, and finds the knee. PREPAREDNESS: safety case must cover the STRONGEST config (top tier + full ultra), not the cheap default. Cross-links o1-system-card (scaling curve) + model-routing/ai-cost-engineering + self-consistency. Worked math plus runnable code.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agentic AI Engineer Roadmap</title>
      <link>https://vibeengines.com/roadmap/agentic-ai-engineer</link>
      <guid isPermaLink="true">https://vibeengines.com/roadmap/agentic-ai-engineer</guid>
      <category>Roadmap</category>
      <description>A step-by-step roadmap to become an agentic AI engineer in 2026 — the year's fastest-growing engineering title. From what an agent actually is, the agent loop, tool use, MCP and memory through planning, multi-agent orchestration, control planes, A2A interop, agentic coding and self-improvement to evals, long-horizon reliability, guardrails, security, supply chain and agentic payments. 18 stations across 3 tracks — Agent Foundations, Building Agents, Production &amp; Safety.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agent Time-Horizon Explorer</title>
      <link>https://vibeengines.com/tools/agent-time-horizon-explorer</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/agent-time-horizon-explorer</guid>
      <category>Tool</category>
      <description>An interactive projection of the METR “time horizon” trend — the length of task an AI agent can complete at 50% reliability, which has been doubling on a regular cadence. Set the current horizon and the doubling period, and the tool projects the calendar date when agents cross the one-hour, one-workday, one-week and beyond thresholds. Every assumption is a slider, so you can stress-test the optimistic and conservative cases yourself.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Skill Linter</title>
      <link>https://vibeengines.com/tools/skill-linter</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/skill-linter</guid>
      <category>Tool</category>
      <description>A client-side linter for agent skill files (SKILL.md). Paste your skill and it scores the structure a good skill needs — a clear name and description, an explicit “when to use” trigger, concrete examples, imperative step-by-step instructions, and a scoped length — against a 12-point rubric, with per-criterion feedback on what is missing. Everything runs in your browser; nothing is uploaded. Built to catch the quality gaps that separate a curated skill from the average public one.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agent Loop Cost Estimator</title>
      <link>https://vibeengines.com/tools/agent-loop-cost-estimator</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/agent-loop-cost-estimator</guid>
      <category>Tool</category>
      <description>An AI agent cost estimator. Enter the steps per task, the tokens added each step, and your model’s input/output prices to see the cost per task, per day and per month — and why re-sending a growing context each step makes cost scale with the square of the steps, not linearly. Shows where prompt caching helps.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The MCP Playground</title>
      <link>https://vibeengines.com/lab/mcp-playground</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/mcp-playground</guid>
      <category>Lab</category>
      <description>Don't read about MCP — watch a model use tools through it. The Model Context Protocol is a standard way to connect an AI host to external tools and data: servers advertise tools with schemas, the host discovers them, the model decides which to call and with what arguments, the server runs the tool, and the result flows back for the model to answer. Step through a real request — get the weather, then save it to notes — and see the whole discover, call, result, answer loop, with theory and a quiz.</description>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>MCP vs A2A</title>
      <link>https://vibeengines.com/handbook/mcp-vs-a2a</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/mcp-vs-a2a</guid>
      <category>Handbook</category>
      <description>Different layers, not rivals: MCP connects an agent to its own tools and data; A2A connects one autonomous agent to another. Why one protocol wasn’t enough for both jobs, and how a single agent uses both — MCP as its hands, A2A as its voice to peers.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agents vs Workflows</title>
      <link>https://vibeengines.com/handbook/agents-vs-workflows</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/agents-vs-workflows</guid>
      <category>Handbook</category>
      <description>Autonomy is a cost, not a default. Workflows hardcode the path in your own code; agents let the model decide its next step at runtime. Why it’s a spectrum, not a binary, and the decision rule for when unpredictability actually justifies an agent loop.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Structured Outputs Handbook</title>
      <link>https://vibeengines.com/handbook/structured-outputs</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/structured-outputs</guid>
      <category>Handbook</category>
      <description>How LLMs return guaranteed-valid JSON for function calling and tool use. Why prompting for JSON is unreliable, how constrained decoding masks illegal tokens to −∞ so output is valid by construction (not by hope), the spectrum from JSON mode to full schema enforcement, and the trade-offs — including why &quot;valid&quot; is not &quot;correct&quot;. With worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Agent Memory Handbook</title>
      <link>https://vibeengines.com/handbook/agent-memory</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/agent-memory</guid>
      <category>Handbook</category>
      <description>A finite context window forces a memory strategy. Short-term (context) vs long-term (external store), how eviction (FIFO) and summarization (compression) keep a growing conversation in the token budget, and how retrieval recalls a fact long after it scrolled out of context — the memory-as-OS idea. With worked budget math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Guardrails Engineering Handbook</title>
      <link>https://vibeengines.com/handbook/guardrails-engineering</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/guardrails-engineering</guid>
      <category>Handbook</category>
      <description>The input and output filters that keep an LLM system safe. Why no single filter is enough, how layering independent guardrails drives the combined miss rate down multiplicatively (defense in depth: ∏ miss_i), why false positives compound the other way (1 − ∏(1−fp_i)), the independence caveat, and where to place each layer. With worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Harness Engineering Handbook</title>
      <link>https://vibeengines.com/handbook/harness-engineering</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/harness-engineering</guid>
      <category>Handbook</category>
      <description>The loop that wraps a model into an autonomous agent — plan, act, observe — and why its most important property is bounds. A stuck agent with no hard stop loops forever, burning money, so the harness enforces max steps AND a cost budget. Geometric success math (P(done ≤ k) = 1 − (1−p)^k, E[steps] = 1/p) and a bounded loop you can run — with worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Sandboxing Handbook</title>
      <link>https://vibeengines.com/handbook/sandboxing</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/sandboxing</guid>
      <category>Handbook</category>
      <description>Safely running untrusted or agent-generated code. Why deny-by-default beats a block-list (you can't enumerate all evil, so enumerate the little good), how least-privilege allowlists shrink the escape surface, why resource caps are needed on top of policy, and the layers of isolation from seccomp to microVMs. With worked policy math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The PII in LLM Pipelines Handbook</title>
      <link>https://vibeengines.com/handbook/pii-in-llm-pipelines</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/pii-in-llm-pipelines</guid>
      <category>Handbook</category>
      <description>Handling personal data safely with the redact-before-send pattern: detect PII, replace each value with a stable placeholder before the model sees it, restore the real values in the response — so personal data never crosses the trust boundary to the provider. The leak invariant, consistent placeholders, detection limits, and defense in depth (minimization, BAA/DPA, local models). With worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Multi-Agent Orchestration Handbook</title>
      <link>https://vibeengines.com/handbook/multi-agent-orchestration</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/multi-agent-orchestration</guid>
      <category>Handbook</category>
      <description>Coordinating multiple LLM agents on one task — and when not to. Why coordination overhead grows quadratically (N(N−1)/2 pairs) while useful work divides only linearly, so past an optimal team size more agents make a system slower, pricier, and less reliable, how to find where the U-curve turns, and the patterns (orchestrator-workers, handoff, routing) that keep coordination near-linear. With worked math and a runnable optimal-team-size calculator.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Browser &amp; Computer-Use Agents Handbook</title>
      <link>https://vibeengines.com/handbook/browser-computer-use-agents</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/browser-computer-use-agents</guid>
      <category>Handbook</category>
      <description>LLM agents that operate a UI like a human — no API required. The observe-act-verify loop, why grounding (identifying which element to click on a cluttered, shifting page) is the dominant source of failure, and why grounding errors compound: task success is per-step accuracy to the power of the number of steps (0.9¹⁰ ≈ 35%), so long UI tasks are brittle. Plus the prompt-injection surface. With worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The A2A Agent Interop Handbook</title>
      <link>https://vibeengines.com/handbook/a2a-agent-interop</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/a2a-agent-interop</guid>
      <category>Handbook</category>
      <description>Letting heterogeneous agents from different teams work together. The two agreements A2A needs: capability discovery (agents advertise their skills so others can find and route to them) and a shared message schema (so one agent's output chains into another's input). Why both are required — a capable partner you can't talk to is useless — plus schema drift, capability lies, and how A2A relates to MCP. With worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an Adaptive AI Tutor</title>
      <link>https://vibeengines.com/ai-system-design/ai-tutor-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/ai-tutor-system-design</guid>
      <category>AI System Design</category>
      <description>Build an adaptive tutoring system like Khanmigo or Duolingo's AI tutor. See why a fixed question order fails every learner, how a per-skill mastery model picks the next problem, why grading routes structured answers to a deterministic checker and reserves the LLM for open-ended judgment, how the mastery feedback loop closes, why explanations are grounded in a curriculum knowledge base instead of freely generated, how a Socratic hint ladder avoids just giving away the answer, and how a long-term learner profile drives spaced repetition.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an AI Code Review Bot</title>
      <link>https://vibeengines.com/ai-system-design/code-review-bot-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/code-review-bot-system-design</guid>
      <category>AI System Design</category>
      <description>Build an automated PR review bot like CodeRabbit or Graphite. See why cheap deterministic linters run before any LLM call, how review is scoped to the diff plus just enough context, how a repo embedding index (RAG over the codebase) supplies cross-file context the diff alone can't show, why every finding is confidence-gated before posting, how a comment ledger prevents re-pushes from spamming old feedback, how secrets are redacted before they ever reach the LLM, and how developer reactions tune down false positives over time.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an Agent Control Plane</title>
      <link>https://vibeengines.com/ai-system-design/agent-control-plane-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/agent-control-plane-system-design</guid>
      <category>AI System Design</category>
      <description>Build the platform layer that runs a whole fleet of production AI agents — not how one agent coordinates a task, but how an organization deploys, versions, budgets, permissions, monitors, and can instantly kill any of potentially hundreds of independently-running agents. See why ad-hoc agent scripts are an operational blind spot, a central agent registry with versioning, server-enforced per-agent tool permissions, hard cost budgets checked per step, an out-of-band kill switch, full execution tracing, and canary rollout of new agent versions.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an Autonomous Email Agent</title>
      <link>https://vibeengines.com/ai-system-design/email-agent-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/email-agent-system-design</guid>
      <category>AI System Design</category>
      <description>Build an AI agent that triages, drafts, and sends email on a user's behalf, where sending is often irreversible. See why auto-sending everything is dangerous, a human-approval gate scoped to action reversibility, priority triage, context assembly from thread history, tone/style matching, defending against indirect prompt injection hidden in received email, and a full audit trail.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design a Sales CRM Agent</title>
      <link>https://vibeengines.com/ai-system-design/sales-crm-agent-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/sales-crm-agent-system-design</guid>
      <category>AI System Design</category>
      <description>Build an AI agent embedded in a CRM that enriches leads, scores priority, drafts outreach, and suggests next actions — writing into a shared system of record and staying a suggester, not an autonomous actor, for relationship-sensitive interactions. See multi-source enrichment with provenance tracking, lead scoring, write-back validation that fails closed, next-best-action suggestion, personalization at scale, activity logging, and stale-data safeguards.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design a Support Resolution Agent</title>
      <link>https://vibeengines.com/ai-system-design/support-resolution-agent-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/support-resolution-agent-system-design</guid>
      <category>AI System Design</category>
      <description>Build an AI agent that triages, resolves, and escalates customer support tickets end-to-end. See why answering from general knowledge alone is dangerous when the real answer depends on account-specific state, grounding responses in real account/order data, confidence-based escalation instead of guessing, tiered auto/assisted/escalate resolution, seamless handoff context, outcome-based quality tracking, tone-aware de-escalation, and identity verification against social engineering.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design a CI Test-Generation Agent</title>
      <link>https://vibeengines.com/ai-system-design/ci-test-gen-agent-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/ci-test-gen-agent-system-design</guid>
      <category>AI System Design</category>
      <description>Build an agent that automatically generates tests for new and changed code in a CI pipeline. See why manual tests leave coverage gaps, coverage-gap analysis, generating tests against the real code, mutation testing to verify a test is actually meaningful, flaky-test detection before it poisons CI signal, why a human review gate is still required, and running expensive validation in a background lane.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>SWE-bench &amp; SWE-agent</title>
      <link>https://vibeengines.com/paper/swe-agent</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/swe-agent</guid>
      <category>Paper</category>
      <description>Can a model do real software engineering? SWE-bench grades models on resolving actual GitHub issues — the patch must apply, make the failing tests pass, and break no passing tests. SWE-agent gives the model an Agent-Computer Interface (search/open/edit/run) built for the model, not humans. The strict resolution rule and why the interface matters — worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>RT-2</title>
      <link>https://vibeengines.com/paper/rt-2</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/rt-2</guid>
      <category>Paper</category>
      <description>A Vision-Language-Action model that controls a robot by emitting actions as text tokens. RT-2 discretizes each dimension of a continuous action into bins — one vocabulary token per dimension — so a web-pretrained vision-language model co-trains on internet data and robot trajectories and its semantic knowledge transfers to control, generalizing to objects and commands never seen in robot data. Action-as-token, quantization precision, and emergent generalization — worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Genie</title>
      <link>https://vibeengines.com/paper/genie</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/genie</guid>
      <category>Paper</category>
      <description>A generative interactive environment that turns an image into a frame-by-frame playable world — trained on unlabeled video with no action labels. A latent action model infers a small discrete set of actions from consecutive frames, and a dynamics model turns them into control, so a tiny codebook forces consistent, reusable latent actions to emerge unsupervised. Latent-action inference and controllable generation — worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Context Engineering Handbook</title>
      <link>https://vibeengines.com/handbook/context-engineering</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/context-engineering</guid>
      <category>Handbook</category>
      <description>The discipline that replaced prompt-tweaking: curating everything the model sees under a finite attention budget. The anatomy of a production context window, the three taxes on an overstuffed one (context rot, cost, behavioral drift), the four operations (write, select, compress, isolate), agent-specific patterns like compaction and just-in-time retrieval, six production habits, and where it sits in the prompt → context → loop → harness stack.</description>
      <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Agent Skills Handbook</title>
      <link>https://vibeengines.com/handbook/agent-skills</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/agent-skills</guid>
      <category>Handbook</category>
      <description>Packaged, on-demand expertise for AI agents: what a skill folder actually is (SKILL.md + scripts + resources), progressive disclosure and why the system prompt couldn't do it, the skills vs tools vs MCP vs fine-tuning decision table, six habits of skills that actually trigger, and why third-party skills are supply-chain dependencies to review like code.</description>
      <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Agentic Coding Handbook</title>
      <link>https://vibeengines.com/handbook/agentic-coding</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/agentic-coding</guid>
      <category>Handbook</category>
      <description>Working with coding agents — Claude Code, Cursor and their cousins — without drowning in AI slop. The loop + tools + context anatomy, the CLAUDE.md instruction file (your highest-leverage artifact), the plan-then-verify workflow, skills/subagents/hooks, the vibe-coding review dial, the security sharp edges (injection, secrets, dangerous commands), and the team norms that keep quality from eroding.</description>
      <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The AI Security Handbook</title>
      <link>https://vibeengines.com/handbook/ai-security</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/ai-security</guid>
      <category>Handbook</category>
      <description>Your app now reads the internet and believes it. Why prompt injection is a design problem filters cannot solve, the lethal trifecta (private data + untrusted content + exfiltration), defense in depth from least privilege to output validation, where PII actually leaks (logs, embeddings, weights), MCP/model supply-chain trust, and the red-team eval suite that gates CI.</description>
      <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>MCP vs Function Calling</title>
      <link>https://vibeengines.com/handbook/mcp-vs-function-calling</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/mcp-vs-function-calling</guid>
      <category>Handbook</category>
      <description>Pitted against each other, but they live at different layers — like comparing USB-C to sending data. Function calling is the model capability to emit a structured tool request; MCP is the open standard for discovering and connecting to tool servers. How MCP uses function calling, and when each matters.</description>
      <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design a Computer-Use Agent</title>
      <link>https://vibeengines.com/ai-system-design/computer-use-agent-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/computer-use-agent-system-design</guid>
      <category>AI System Design</category>
      <description>Build a computer-use / browser agent — the agent whose only API is the screen. See how the see-think-act loop works one verified action at a time, why perception mixes screenshots with accessibility trees, how set-of-marks grounding turns &quot;click Submit&quot; into coordinates, why execution lives in a disposable sandboxed VM, how every action is verified against the next frame, where approval gates catch irreversible clicks, and how task memory survives 60-step workflows — through an interactive diagram where you can change the page mid-task.</description>
      <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    </item>
  </channel>
</rss>
