<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Vibe Engines — LLM Engineering</title>
    <link>https://vibeengines.com/topic/llm-engineering</link>
    <atom:link href="https://vibeengines.com/topic/llm-engineering/feed.xml" rel="self" type="application/rss+xml" />
    <description>Building with large language models — prompting, decoding, inference serving and cost — from the handbooks down to the softmax that powers every token.</description>
    <language>en</language>
    <lastBuildDate>Mon, 27 Jul 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>ML Research Engineer Roadmap</title>
      <link>https://vibeengines.com/roadmap/research-engineer</link>
      <guid isPermaLink="true">https://vibeengines.com/roadmap/research-engineer</guid>
      <category>Roadmap</category>
      <description>A step-by-step roadmap to become a machine-learning research engineer in 2026 — the person who builds and trains models, not just features on top of them. From the math and PyTorch, deep-learning fundamentals and reading papers through data pipelines, training dynamics, distributed training and scaling laws to post-training, evaluation, efficiency, the research process and the research-engineer career. 18 stations across 3 tracks — Research Foundations, Training at Scale, Frontier &amp; the Craft.</description>
      <pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Mechanistic Interpretability</title>
      <link>https://vibeengines.com/handbook/mech-interp-for-engineers</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/mech-interp-for-engineers</guid>
      <category>Handbook</category>
      <description>The engineer's guide to reading a model — interpretability as a debugger, not a séance. Why features live in superposition (not one per neuron), how sparse autoencoders pull monosemantic features back out, how attribution and circuit tracing find why a model did something (contribution of input i to output = the summed path weight W2·W1, validated causally by ablation), feature steering as a direct intervention, and an honest map of what you can debug today versus what is still research. Worked math plus a runnable 2-layer circuit tracer.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Diffusion LLMs</title>
      <link>https://vibeengines.com/handbook/diffusion-llms</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/diffusion-llms</guid>
      <category>Handbook</category>
      <description>Text generation that isn't one token at a time. Why autoregressive decoding is latency-bound (N tokens = N sequential passes), how block diffusion denoises B tokens in K steps to emit B/K tokens per pass (⌈N/B⌉·K total passes, speedup B/K) for reported 1,000+ tokens/sec, the crucial image-vs-text split (continuous Gaussian noise over pixels vs a discrete masking corruption over tokens), and the speed-quality knob (fewer steps = faster but more parallel-commitment error). Worked math plus a runnable AR-vs-diffusion pass counter.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>RL Environments Engineering</title>
      <link>https://vibeengines.com/handbook/rl-environments-engineering</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/rl-environments-engineering</guid>
      <category>Handbook</category>
      <description>The build side of verifiable rewards — the reasoning models everyone celebrates are taught by the ENVIRONMENT, not the algorithm. Why RL learns only from reward variance p(1−p) (zero signal when every attempt passes or fails, peak at a 50% pass rate) so tasks must sit in the difficulty band, why a gameable verifier corrupts training (precision = correct ÷ all rewarded; every false positive is a lie the model learns), and how to engineer difficulty curricula and verifier soundness. The theory of reward hacking lives in the verifiable-rewards handbook; this is how to build around it. Worked math plus a runnable learning-signal and verifier-precision model.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Self-Improving Agents</title>
      <link>https://vibeengines.com/handbook/self-improving-agents</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/self-improving-agents</guid>
      <category>Handbook</category>
      <description>The long one: how an agent gets better at its own job, end to end. COVERAGE vs the accuracy you can ship (a model usually already produces a right answer somewhere in its output; the missing piece is recognising it), unbiased pass@k, why the sampling curve is straight, the SELECTION bottleneck, building a VERIFIER and the two asymmetric ways it is wrong, feedback that is not a verifier ranked by independence, the five-predicate filter that decides what becomes training data, the SIX substrates a lesson can land in (context, memory, tools, scaffold, verifier, weights) and how to choose, STaR through PPO/GRPO/DPO and reward hacking, why a loop refit on its own output COLLAPSES toward its own habits, search as self-improvement, and measuring a time HORIZON rather than a score. 14 sections, 13 runnable Python blocks, 3 interactive models, 13 diagrams, 5 checkpoints and a running build log. Every figure carries the condition it was measured under.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Autoregressive vs Diffusion LLMs</title>
      <link>https://vibeengines.com/handbook/autoregressive-vs-diffusion-llms</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/autoregressive-vs-diffusion-llms</guid>
      <category>Handbook</category>
      <description>Autoregressive vs diffusion language models, decided by how text is generated: autoregressive models emit one token at a time left-to-right, each conditioned on all previous; diffusion LLMs refine every position in parallel over a fixed number of denoising steps. Decode latency, KV cache, revisability, the quality gap, and where block diffusion fits.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>vLLM vs TGI vs SGLang</title>
      <link>https://vibeengines.com/handbook/vllm-vs-tgi-vs-sglang</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/vllm-vs-tgi-vs-sglang</guid>
      <category>Handbook</category>
      <description>Three open-source LLM serving engines that all do continuous batching but differ in their signature KV-cache trick: vLLM’s PagedAttention (removes fragmentation), SGLang’s RadixAttention (shares repeated prefixes), and TGI’s production-hardened Hugging Face integration. Where each wins, why &quot;fastest&quot; is a workload question, and how to choose.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an RL Environment Farm</title>
      <link>https://vibeengines.com/ai-system-design/rl-environment-farm-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/rl-environment-farm-system-design</guid>
      <category>AI System Design</category>
      <description>Build the infrastructure that trains reasoning models — the environment-and-verifier side of RL with verifiable rewards, not the preference-tuning of RLHF. See why running rollouts inline in the trainer is unsafe and unscalable, sandboxed rollout workers that safely execute untrusted model-generated code, a scheduler that fans rollouts across a fleet and distributes the policy, a SOUND verifier (a gameable one makes the model learn the exploit — reward hacking), a task bank that serves the productive difficulty band (a curriculum, because learning needs reward variance), and a trajectory store that closes the loop — with reward-hacking and sandbox-escape chaos.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>All At Once — Text Diffusion Decoding</title>
      <link>https://vibeengines.com/lab/text-diffusion-decoding</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/text-diffusion-decoding</guid>
      <category>Lab</category>
      <description>Don't read about diffusion LLMs — run one. An autoregressive model writes one token per pass; a text-diffusion model starts from a fully masked block and refines the whole thing in parallel over K denoising steps, emitting B/K tokens per pass. Watch a block go from ██████ to clean text, then tune the block size and step count to feel the speed-versus-coherence dial — and see why text diffusion (masking, discrete) is not image diffusion (Gaussian noise, continuous). Interactive, with theory and a quiz.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Million-Token Bill — Sparse Attention</title>
      <link>https://vibeengines.com/lab/sparse-attention-1m</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/sparse-attention-1m</guid>
      <category>Lab</category>
      <description>Don't read about long-context attention — run the numbers. Dense attention caches a key/value per token and scores every query against all of them, so at a million tokens memory grows linearly and compute grows with the square. Pull the two levers of compressed sparse attention — compress each token's K/V to a latent, and select only the top-k blocks per query — and watch the KV cache and attention FLOPs collapse, with the two savings multiplying (as in DeepSeek-V4). Drag the context to a million and see the ratios fall. Interactive, with theory and a quiz.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>DeepSeek-V4</title>
      <link>https://vibeengines.com/paper/deepseek-v4</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/deepseek-v4</guid>
      <category>Paper</category>
      <description>The efficiency sequel to V3, aimed at the real cost of long context. A ~1.6T-parameter MoE (only ~49B active per token) that serves a 1M-token window at roughly a tenth of V3.2's KV cache. The mechanism is one idea on two faces — HYBRID sparse attention: compress each token's keys/values into a small latent (memory per token: ratio = d_c/width, context-invariant), then SELECT only the top-k blocks of size b to score against, freezing the scored-key count at k·b once n&gt;k·b while dense attention keeps paying O(n²). The two discounts multiply: FLOPs ratio = (k·b/n)·(d_c/width). Reasoning is post-trained with ON-POLICY DISTILLATION — the student generates its own trajectories and a stronger teacher corrects those exact outputs — replacing R1's large-scale RL loop with a denser, more stable signal. Worked math plus runnable code.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Kimi K3</title>
      <link>https://vibeengines.com/paper/kimi-k3</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/kimi-k3</guid>
      <category>Paper</category>
      <description>The sequel to Kimi K2, and a different bet: attack decode SPEED, not KV memory. A ~2.8T open MoE with a 1M-token context — reported as the largest open model to date — built on KIMI DELTA ATTENTION (KDA), a linear-attention delta rule. Softmax attention re-reads the whole KV cache every decode step (O(n) per token, so generation slows as context grows); KDA folds the past into ONE fixed-size recurrent state S (~d×d) and reads o=S·q in constant time regardless of length. The DELTA rule (S ← S − β·(S·k − v)·kᵀ) makes each write CORRECTIVE — overwrite a key's old value instead of accumulating it — which is what plain linear attention gets wrong (write (k,a) then (k,b): delta reads b, linear reads a+b, smeared). ATTENTION RESIDUALS keep a minority of full-attention layers for the sharp exact recall a fixed state blurs — a hybrid, not a wholesale swap. Worked math plus runnable code.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Engram</title>
      <link>https://vibeengines.com/paper/engram</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/engram</guid>
      <category>Paper</category>
      <description>The memory half of the long-context problem. Softmax attention answers a query by spreading a fixed probability budget over ALL tokens, so as the haystack grows the needle dilutes and recall degrades — you pay O(n) AND lose accuracy. Engram bolts a KEY-ADDRESSED memory onto the model: write a fact under a key, retrieve it in constant time O(1) regardless of context length, no scan and no dilution (a Python dict IS the abstraction). It is CONDITIONAL — a lightweight gate decides whether a position needs a long-range recall, so local tokens pay nothing and only recall tokens incur one O(1) lookup. Because retrieval is by identity not by out-competing every other token, needle-in-a-haystack recall stops falling with length (reported ~84.2→97). Underpins DeepSeek-V4 (cheap context that stays accurate); same explicit-memory family as Titans and MemGPT, baked into the model not bolted on by a harness. Worked math plus runnable code.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Qwen4-Coder</title>
      <link>https://vibeengines.com/paper/qwen4-coder</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/qwen4-coder</guid>
      <category>Paper</category>
      <description>An open, Apache-2.0 coding model that reportedly hits ~82% on SWE-Verified while running on a MacBook — the first open, Mac-runnable model past 80% there. The enabler is the mixture-of-experts SHAPE, not a bigger model: ~32B total parameters but only ~3B active per token, which splits the two costs everyone conflates. STORAGE is set by total params (all experts must be RAM-resident since the router may pick any) — 32B at 4-bit ≈ 16GB, fits a 24GB Mac (fp16 would be 64GB and fail); quantization is what makes it a laptop model. SPEED is set by ACTIVE params (decode is bandwidth-bound and reads only the routed ~3B experts) — so a 32B model decodes at 3B throughput. The mental model: &quot;does it fit?&quot; is answered by total×bits÷8; &quot;is it fast enough?&quot; by active params vs bandwidth — dense models tie these together, MoE pulls them apart. Sequel to Qwen3; validates the local-first coding-agent stack. Worked math plus runnable code.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>GPT-5.6 System Card</title>
      <link>https://vibeengines.com/paper/gpt-5-6-system-card</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/gpt-5-6-system-card</guid>
      <category>Paper</category>
      <description>A modern system card is really a COMPUTE-ALLOCATION POLICY, not one model. TIERED ROUTING (reported Sol/Terra/Luna) sorts each query by difficulty — cheap fast tier by default, escalate the hard cases (the cascade cost math lives in the model-routing handbook — linked, not re-derived). ULTRA MODE spends test-time compute in PARALLEL: run several independent agents (reported 4) on the hardest task and a VERIFIER keeps any that solves, so the solve rate is 1−(1−p)^N. The fresh angle = why FOUR, not forty: the marginal value of the N-th agent is Δ(N)=(1−p)^(N−1)·p — geometric decay (p=0.6: +0.60/+0.24/+0.10/+0.04) while cost is linear in N, so there's a knee. Runnable proves best-of-N rises, marginal=(1−p)^(N−1)·p, diminishing returns, 4th≪1st, value/cost decreasing, and finds the knee. PREPAREDNESS: safety case must cover the STRONGEST config (top tier + full ultra), not the cheap default. Cross-links o1-system-card (scaling curve) + model-routing/ai-cost-engineering + self-consistency. Worked math plus runnable code.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Concept Circuits</title>
      <link>https://vibeengines.com/paper/concept-circuits</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/concept-circuits</guid>
      <category>Paper</category>
      <description>The foundation of mechanistic interpretability — how to read a model. Neurons are POLYSEMANTIC (one unit fires for green AND legal text AND Python) because the model stores concepts as DIRECTIONS spread across many neurons, not one-per-neuron. SUPERPOSITION: a d-dim space holds only d orthogonal directions but VASTLY more near-orthogonal ones, so a model packs in more features than dimensions and accepts small interference (interference(i,j)=feature_i·feature_j: orthogonal=0, superposed small-nonzero; 3 unit vecs at 120° in 2-D → pairwise dot −0.5). It works because features are SPARSE — few active at once, so collisions rarely bite and a lone active feature reads back cleanly. ATTRIBUTION GRAPHS trace which features caused an output (contribution = output-direction · feature-direction) and chain them across layers into a CIRCUIT. Runnable proves orthonormal self=1/cross=0, 3 features in 2-D, unit length preserved, interference bounded 0.5, clean sparse readout. First interp content sitewide; pairs the coming mech-interp-for-engineers handbook. Worked math plus runnable code.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>NPU Model-Fit Calculator</title>
      <link>https://vibeengines.com/tools/npu-model-fit-calculator</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/npu-model-fit-calculator</guid>
      <category>Tool</category>
      <description>An on-device AI planner keyed to the NPU, not the GPU. Enter your device’s TOPS (e.g. a 75-TOPS Snapdragon X2) and unified RAM, pick a model size and quantization, and see two things the VRAM-fit calculators miss: whether the weights fit in memory, and how fast the NPU can prefill a prompt — the compute-bound step that TOPS actually governs. Get an escalate-to-cloud verdict when a prompt is too long to stay responsive locally.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>EU AI Act Timeline</title>
      <link>https://vibeengines.com/tools/eu-ai-act-timeline</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/eu-ai-act-timeline</guid>
      <category>Tool</category>
      <description>An engineering-oriented timeline of the EU AI Act: the statutory milestones (entry into force, prohibitions, GPAI obligations, high-risk rules) and what each one asks a builder to produce — model cards, logging, risk documentation, human-oversight design. Filter by your system type to see which obligations land and when. This is a plain-language planning aid for engineers, not legal advice, and it flags where proposed changes (the Digital Omnibus) may move dates that are not yet settled.</description>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Batch vs Real-Time Inference</title>
      <link>https://vibeengines.com/handbook/batch-vs-realtime-inference</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/batch-vs-realtime-inference</guid>
      <category>Handbook</category>
      <description>Batch vs real-time (online) inference, decided by whether a human is waiting: batch processes many inputs together offline for maximum throughput and lowest cost per token; real-time answers one request at a time with low latency. GPU utilization, continuous batching, and why most AI products run both.</description>
      <pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Test-Time Compute &amp; Reasoning Models</title>
      <link>https://vibeengines.com/handbook/test-time-compute</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/test-time-compute</guid>
      <category>Handbook</category>
      <description>Why letting a model think longer at inference — long chain-of-thought, sampling with self-consistency, best-of-N with a verifier, and search — can beat a much larger model on hard problems. How o1/o3- and DeepSeek R1-style reasoning models are trained to use a thinking budget, the train-time vs test-time scaling trade, and when spending the extra compute actually pays.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>MoE vs Dense Models</title>
      <link>https://vibeengines.com/handbook/moe-vs-dense-models</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/moe-vs-dense-models</guid>
      <category>Handbook</category>
      <description>Mixture of Experts vs dense LLMs, decided by which parameters fire per token: a dense model runs every parameter on every token; an MoE routes each token to a few expert sub-networks, holding far more total parameters but activating only a fraction. The compute win, the memory tax (all experts must sit in VRAM), load balancing, and when each wins.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Encoder vs Decoder Models</title>
      <link>https://vibeengines.com/handbook/encoder-vs-decoder-models</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/encoder-vs-decoder-models</guid>
      <category>Handbook</category>
      <description>Encoder vs decoder Transformers, decided by which way attention flows: an encoder reads the whole input at once (bidirectional) to build representations — great for classification, NER and embeddings (BERT); a decoder generates left to right (causal) — great for chat, code and agents (GPT). Encoder-decoder models like T5 do both. When to reach for each, and why real systems chain them.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Quantization vs Distillation</title>
      <link>https://vibeengines.com/handbook/quantization-vs-distillation</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/quantization-vs-distillation</guid>
      <category>Handbook</category>
      <description>Quantization vs distillation for shrinking LLMs, decided by what you change: quantization keeps the same model but stores its weights in fewer bits (FP16 → INT8/INT4); distillation trains a smaller student model to imitate a bigger teacher. One compresses the representation, the other the architecture — the quality-vs-size frontier of each, and why teams often distill then quantize.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Deploy an LLM in a Customer’s Environment</title>
      <link>https://vibeengines.com/ai-system-design/on-prem-llm-deployment-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/on-prem-llm-deployment-system-design</guid>
      <category>AI System Design</category>
      <description>Deploy a large language model inside a customer’s own environment — the constraint a forward deployed engineer meets when a regulated customer won’t send data to a public API. See why a hosted endpoint is a non-starter, right-sizing the model and precision to the fixed hardware they actually own, serving efficiently on limited GPUs, running with no network egress (air-gapped, no phone-home), shipping model and security updates as signed bundles into a locked-down environment, and getting observability out without exfiltrating customer data.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Semantic Chunker (Token Budget)</title>
      <link>https://vibeengines.com/challenge/semantic-chunker</link>
      <guid isPermaLink="true">https://vibeengines.com/challenge/semantic-chunker</guid>
      <category>Challenge</category>
      <description>Before you embed a document for retrieval, split it into chunks that fit a token budget — without slicing a sentence. The greedy packer fills each chunk with whole sentences until the next would overflow. Solve it in Python or TypeScript, with hidden tests.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Speculative Decoding: Accept Step</title>
      <link>https://vibeengines.com/challenge/speculative-accept</link>
      <guid isPermaLink="true">https://vibeengines.com/challenge/speculative-accept</guid>
      <category>Challenge</category>
      <description>Speculative decoding speeds up LLM inference: a small draft model proposes tokens, the big target model verifies them in one pass. The accept/reject rule guarantees the output matches the target’s own distribution. Implement it. Solve it in Python or TypeScript, with hidden tests.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>KV-Cache Eviction (Attention Sinks)</title>
      <link>https://vibeengines.com/challenge/kv-eviction</link>
      <guid isPermaLink="true">https://vibeengines.com/challenge/kv-eviction</guid>
      <category>Challenge</category>
      <description>An LLM’s KV cache grows every token, so long chats must drop old entries without wrecking quality. StreamingLLM keeps the first few &quot;attention sink&quot; tokens plus a sliding window, evicting the middle. Compute the survivors. Solve it in Python or TypeScript, with hidden tests.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Constrained Decoding (Logit Masking)</title>
      <link>https://vibeengines.com/challenge/constrained-decode</link>
      <guid isPermaLink="true">https://vibeengines.com/challenge/constrained-decode</guid>
      <category>Challenge</category>
      <description>How do you force an LLM to emit only valid JSON or a token your grammar allows? Mask the logits: set every disallowed token to −∞, then take the argmax over what remains. The backbone of structured output. Solve it in Python or TypeScript, with hidden tests.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Agent Loop Cost Estimator</title>
      <link>https://vibeengines.com/tools/agent-loop-cost-estimator</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/agent-loop-cost-estimator</guid>
      <category>Tool</category>
      <description>An AI agent cost estimator. Enter the steps per task, the tokens added each step, and your model’s input/output prices to see the cost per task, per day and per month — and why re-sending a growing context each step makes cost scale with the square of the steps, not linearly. Shows where prompt caching helps.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>GPU Rental Price Reference</title>
      <link>https://vibeengines.com/tools/gpu-rental-prices</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/gpu-rental-prices</guid>
      <category>Tool</category>
      <description>A cloud GPU price reference and cost estimator. Compare approximate on-demand hourly rates for common GPUs (T4, L4, A10G, A100, H100) and estimate what a training or inference job costs from the GPU count and hours. Rates are approximate and drift over time — always confirm current pricing with your provider.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Chain-of-Thought vs Direct</title>
      <link>https://vibeengines.com/lab/cot-vs-direct</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/cot-vs-direct</guid>
      <category>Lab</category>
      <description>Don't read about chain-of-thought — watch it rescue a wrong answer. A model produces one token at a time, so asking it to blurt a final answer to a multi-step problem crams the whole computation into one step — and it often takes a tempting shortcut and gets it wrong. Asking it to think step by step lets it spend intermediate tokens working the problem, each step small enough to get right. Run the same problems both ways and see the shortcut fail and the chain succeed — made playable, with theory and a quiz.</description>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>LoRA vs Full Fine-Tuning</title>
      <link>https://vibeengines.com/handbook/lora-vs-full-fine-tuning</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/lora-vs-full-fine-tuning</guid>
      <category>Handbook</category>
      <description>LoRA freezes the base model and trains a tiny low-rank adapter; full fine-tuning updates every weight. Why a low-rank update can work at all, memory and storage trade-offs, multi-task adapter serving, and when the extra cost of full fine-tuning is worth it.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>LLM vs SLM</title>
      <link>https://vibeengines.com/handbook/llm-vs-slm</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/llm-vs-slm</guid>
      <category>Handbook</category>
      <description>The choice is about deployment constraints, not just raw capability: LLMs run in the cloud with broad reasoning; SLMs run on-device with near-zero latency, cost, and full privacy. Why the capability gap is shrinking fast, and the model-routing pattern that uses both.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>GPU vs TPU</title>
      <link>https://vibeengines.com/handbook/gpu-vs-tpu</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/gpu-vs-tpu</guid>
      <category>Handbook</category>
      <description>GPUs are flexible general-purpose processors with the CUDA ecosystem behind them; TPUs are ASICs purpose-built around the systolic array for matmul, on Google Cloud only. Why that specialization is efficient exactly where it fits, and a liability the moment it doesn’t.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>LLM API Pricing</title>
      <link>https://vibeengines.com/handbook/llm-api-pricing</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/llm-api-pricing</guid>
      <category>Handbook</category>
      <description>Every major provider’s API pricing in one comparable table, per million tokens — Anthropic Claude, OpenAI GPT, Google Gemini — pulled from official docs on 2026-07-21, not aggregators. Batch and prompt-caching discounts, a cross-provider tier map, and a fully worked chatbot cost example.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Tokenizer Efficiency Benchmark</title>
      <link>https://vibeengines.com/handbook/tokenizer-efficiency-benchmark</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/tokenizer-efficiency-benchmark</guid>
      <category>Handbook</category>
      <description>An original, reproducible benchmark of 9 real tokenizers across English prose, code, and non-English text — measuring characters-per-token and encode speed, with a fidelity check that catches the trap where a lossy tokenizer looks “efficient” only because it silently drops characters it can’t represent.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The AI Cost Engineering Handbook</title>
      <link>https://vibeengines.com/handbook/ai-cost-engineering</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/ai-cost-engineering</guid>
      <category>Handbook</category>
      <description>AI spend is a sum over tokens: input-tokens × input-price + output-tokens × output-price, per million. Why output tokens dominate the bill (usually ~5× input), how prompt caching pays off from the very first reuse, and the full lever set — shorter outputs, caching, model routing, batching, context trimming — that cuts spend without cutting quality. With worked math and a runnable cost calculator.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Model Routing Handbook</title>
      <link>https://vibeengines.com/handbook/model-routing</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/model-routing</guid>
      <category>Handbook</category>
      <description>Send each request to the cheapest model that can handle it. How a cheap-smart cascade works, why you pay the cheap model on everything so the escalation rate is the key lever, the break-even math (r* = 1 − c_cheap/c_big), predictive routers vs cascades, and the confidently-wrong pitfall. With worked math and a runnable cost model.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The LLM Observability Handbook</title>
      <link>https://vibeengines.com/handbook/llm-observability</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/llm-observability</guid>
      <category>Handbook</category>
      <description>Tracing an AI request as a tree of spans. Why total latency is the critical path — the max end time, not the sum of durations, because parallel spans overlap — while cost is the sum across every span, how to find the bottleneck span, and what to log for every LLM call (including percentiles over averages). With worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Prompt Caching Handbook</title>
      <link>https://vibeengines.com/handbook/prompt-caching</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/prompt-caching</guid>
      <category>Handbook</category>
      <description>How prompt caching actually works. Why caches key on the exact prefix and break at the first differing token, why cache-aware ordering (fixed content first, variable last) is the single biggest lever on hit rate, how TTL and cache breakpoints work, and the invisible mistakes (a timestamp up top, non-deterministic serialization) that silently kill caching. With worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Local LLM Stack Handbook</title>
      <link>https://vibeengines.com/handbook/local-llm-stack</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/local-llm-stack</guid>
      <category>Handbook</category>
      <description>Running LLMs locally with ollama and llama.cpp. The one equation that decides what fits — weight memory = params × bits per weight ÷ 8 — how quantization shrinks a 7B model from 14GB (fp16) to 3.5GB (4-bit) to fit a consumer GPU, why 4-bit is the sweet spot, and the local stack (GGUF, ollama, llama.cpp). With worked memory math and a runnable fits-in-VRAM calculator.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Voice AI Real-Time Agents Handbook</title>
      <link>https://vibeengines.com/handbook/voice-ai-realtime-agents</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/voice-ai-realtime-agents</guid>
      <category>Handbook</category>
      <description>Why a voice agent is a stopwatch, not a chat box. It's a three-stage pipeline — speech-to-text → LLM → text-to-speech — and the user waits for all three, so end-to-end latency is the sum of the stages. Human turn-taking breaks down past ~800ms, so the whole game is keeping that sum under the conversational budget: find the bottleneck (usually the LLM), shave the slowest stage first, cut network hops, and stream so the stages overlap. Plus endpointing, interruption handling, and why time-to-first-audio beats the naive sum. With worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The EU AI Act Handbook</title>
      <link>https://vibeengines.com/handbook/eu-ai-act</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/eu-ai-act</guid>
      <category>Handbook</category>
      <description>The world's first comprehensive AI law, made legible for engineers. The key idea: it regulates AI by the risk of the use case, not the technology — the same model is unregulated in a spam filter and heavily regulated in a hiring tool. Every system sorts into one of four tiers: unacceptable (banned — social scoring, manipulation, most real-time public biometric ID), high-risk (hiring, credit, medical, education, law enforcement… → risk management, data governance, human oversight, conformity assessment, registration), limited-risk (transparency — chatbots disclose they're AI, deepfakes labelled), and minimal-risk (no obligations — the vast majority). Because tiering is a concrete rule, classification is deterministic. Plus the GPAI/foundation-model layer, what engineers should actually do, and the traps (optimistic self-classification, retrofitting compliance). With a worked classification model and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The AI Product Engineering Handbook</title>
      <link>https://vibeengines.com/handbook/ai-product-engineering</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/ai-product-engineering</guid>
      <category>Handbook</category>
      <description>A demo isn't a product. Getting an impressive LLM demo takes an afternoon; turning it into something you can charge for at scale is the brutal last mile where most AI features die. Product engineering rests on two numbers a demo never shows: unit economics (every request costs tokens — cost = in·p_in + out·p_out per million, which sets your gross margin and whether you're profitable) and the eval gate (ship a change only if its measured quality clears a bar AND its margin clears a bar — blocking cheap-but-wrong and great-but-unprofitable alike). Plus the reliability stack (evals, structured outputs, retries/fallbacks, caching, guardrails, observability) that keeps a flaky, pricey, non-deterministic model dependable in production. With worked math and a runnable cost/margin/ship-decision calculator.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an Adaptive AI Tutor</title>
      <link>https://vibeengines.com/ai-system-design/ai-tutor-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/ai-tutor-system-design</guid>
      <category>AI System Design</category>
      <description>Build an adaptive tutoring system like Khanmigo or Duolingo's AI tutor. See why a fixed question order fails every learner, how a per-skill mastery model picks the next problem, why grading routes structured answers to a deterministic checker and reserves the LLM for open-ended judgment, how the mastery feedback loop closes, why explanations are grounded in a curriculum knowledge base instead of freely generated, how a Socratic hint ladder avoids just giving away the answer, and how a long-term learner profile drives spaced repetition.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design an AI Code Review Bot</title>
      <link>https://vibeengines.com/ai-system-design/code-review-bot-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/code-review-bot-system-design</guid>
      <category>AI System Design</category>
      <description>Build an automated PR review bot like CodeRabbit or Graphite. See why cheap deterministic linters run before any LLM call, how review is scoped to the diff plus just enough context, how a repo embedding index (RAG over the codebase) supplies cross-file context the diff alone can't show, why every finding is confidence-gated before posting, how a comment ledger prevents re-pushes from spamming old feedback, how secrets are redacted before they ever reach the LLM, and how developer reactions tune down false positives over time.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design a Model Registry (MLOps)</title>
      <link>https://vibeengines.com/ai-system-design/model-registry-mlops-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/model-registry-mlops-system-design</guid>
      <category>AI System Design</category>
      <description>Build the platform that versions, evaluates, promotes and can instantly roll back trained ML models — the governance layer between &quot;a model finished training&quot; and &quot;a model is safely serving production traffic.&quot; See why an artifact alone is meaningless without lineage, staged promotion through gated environments, automated evaluation gates that fail closed, shadow/canary deployment because offline metrics aren't sufficient, instant rollback via immutable versioning, and production drift monitoring.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design a GPU Cluster Scheduler</title>
      <link>https://vibeengines.com/ai-system-design/gpu-cluster-scheduler-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/gpu-cluster-scheduler-system-design</guid>
      <category>AI System Design</category>
      <description>Build a scheduler for a shared pool of expensive GPUs across many teams. See why naive FIFO scheduling starves large distributed jobs and fragments capacity, gang scheduling for all-or-nothing distributed training, topology-aware placement, priority and preemption for latency-sensitive inference versus long-running batch training, fair-share quotas across competing teams, checkpoint-and-resume, and reconciling the scheduler's bookkeeping against real cluster state.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design a Prompt Management &amp; A/B Testing Platform</title>
      <link>https://vibeengines.com/ai-system-design/prompt-mgmt-ab-platform-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/prompt-mgmt-ab-platform-system-design</guid>
      <category>AI System Design</category>
      <description>Build the platform that versions, evaluates, and safely A/B tests prompts in a production LLM application — a much faster-moving artifact than model weights, often edited by non-engineers. See why hardcoded prompts force a full deploy for every tweak, externalizing prompts as versioned config, an offline eval gate before rollout, live A/B testing, guardrail metrics that must not regress, routine one-click rollback, and a review gate suited to non-engineer editors.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Design a Multi-Tenant AI SaaS Platform</title>
      <link>https://vibeengines.com/ai-system-design/multi-tenant-ai-saas-system-design</link>
      <guid isPermaLink="true">https://vibeengines.com/ai-system-design/multi-tenant-ai-saas-system-design</guid>
      <category>AI System Design</category>
      <description>Build an AI product serving many separate customer organizations from shared infrastructure. See why unscoped shared retrieval risks cross-tenant data leaks, enforcing isolation at the data layer, per-tenant customization via configuration, noisy-neighbor prevention with per-tenant quotas, per-tenant cost attribution, data-residency routing, automated onboarding, and defense-in-depth isolation that survives a bug in any single layer.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
  </channel>
</rss>
