<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Vibe Engines — ML Foundations</title>
    <link>https://vibeengines.com/topic/ml-foundations</link>
    <atom:link href="https://vibeengines.com/topic/ml-foundations/feed.xml" rel="self" type="application/rss+xml" />
    <description>The math under the models — gradient descent, softmax, embeddings, similarity and ranking — taught as playable labs and runnable challenges, plus the research papers that build on them: reinforcement learning, generative models and graph networks.</description>
    <language>en</language>
    <lastBuildDate>Mon, 03 Aug 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>ML Fundamentals</title>
      <link>https://vibeengines.com/handbook/ml-fundamentals</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/ml-fundamentals</guid>
      <category>Handbook</category>
      <description>The seven concept pairs every practitioner is expected to have straight — how machines learn, what they predict, the two ways they miss, which mistake you can live with, how you validate, how you ensemble, and what a model is really modelling. Worked confusion matrices, real fold scores, and the failure mode behind each one.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Precision vs Recall: Drag the Threshold</title>
      <link>https://vibeengines.com/lab/precision-recall</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/precision-recall</guid>
      <category>Lab</category>
      <description>Don't read the definitions — drag the line. A spam filter scored 5,000 emails; you pick where to cut. Every bar in the chart recolours into its confusion-matrix quadrant as you move, so precision, recall and F1 stop being formulas and become regions you can see. Push it to the extremes: catch every spam and bury real mail, or never lose mail and let phishing through. Then flag nothing at all and watch the model score 85% accuracy while catching zero spam \u2014 the class-imbalance trap, in one click.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Bias vs Variance: Fit the Curve, Watch It Overfit</title>
      <link>https://vibeengines.com/lab/bias-variance</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/bias-variance</guid>
      <category>Lab</category>
      <description>Don't memorize the tradeoff \u2014 cause it. Slide the polynomial degree from 1 to 12 and watch a fit go from too stiff to bend through every point. Training error falls the whole way; test error bottoms out at degree 3 and then climbs 3.5x. Then hit resample: at degree 1 the fits barely move, at degree 12 they fan across the entire frame. That fan is variance, and you made it appear.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Cross-Validation: One Exam, or Five?</title>
      <link>https://vibeengines.com/lab/cross-validation</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/cross-validation</guid>
      <category>Lab</category>
      <description>Don't take the score on faith \u2014 see how much it could have been. Slide k to repartition 20 samples into folds, rotate the held-out fold and watch the mean assemble. Then run the payoff side by side: report a single hold-out split and your number swings from 0.790 to 0.880 depending purely on which split you happened to draw, while 5-fold sits still at 0.836 \u00b1 0.033. Plus what it costs you \u2014 k folds means k trainings.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Cosine Similarity Calculator</title>
      <link>https://vibeengines.com/tools/cosine-similarity-calculator</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/cosine-similarity-calculator</guid>
      <category>Tool</category>
      <description>An interactive cosine-similarity and vector-distance calculator. Paste two vectors and instantly see their dot product, magnitudes, cosine similarity, cosine distance, the angle between them, and Euclidean and Manhattan distance — the exact numbers behind how embedding search decides which items are “close.” See why cosine ignores magnitude and when to prefer it over Euclidean distance.</description>
      <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>RL from Verifiable Rewards (RLVR)</title>
      <link>https://vibeengines.com/handbook/rlvr-verifiable-rewards</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/rlvr-verifiable-rewards</guid>
      <category>Handbook</category>
      <description>The training technique behind modern reasoning models: reinforcement learning where the reward comes from a programmatic check (a unit test passing, a math answer matching) instead of a gameable learned reward model. How it differs from RLHF, the GRPO/PPO loop, why reasoning behaviors emerge in DeepSeek R1-Zero, and where verifiable rewards run out.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>World Models</title>
      <link>https://vibeengines.com/handbook/world-models</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/world-models</guid>
      <category>Handbook</category>
      <description>What it means for AI to learn a predictive model of an environment it can imagine inside — the basis of model-based RL, planning, and controllable simulation. The three families (latent control models like Dreamer, generative interactive video like Genie/Sora, and JEPA), how a latent world model learns and acts in imagination, and the debate over whether video generators really understand physics.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Softmax with Temperature</title>
      <link>https://vibeengines.com/challenge/softmax-temperature</link>
      <guid isPermaLink="true">https://vibeengines.com/challenge/softmax-temperature</guid>
      <category>Challenge</category>
      <description>The dial that controls how &quot;creative&quot; a language model is: temperature scales the logits before the softmax — low sharpens toward greedy, high flattens toward uniform. Implement it, numerically stable. Solve it in Python or TypeScript, with hidden tests.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Embedding Dimension Tradeoff</title>
      <link>https://vibeengines.com/tools/embedding-dim-tradeoff</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/embedding-dim-tradeoff</guid>
      <category>Tool</category>
      <description>An embedding storage calculator. Enter how many vectors you have and compare the memory footprint across common embedding dimensions and quantization levels (float32, float16, int8) — so you can weigh retrieval quality against the RAM and cost of your vector index, and see what dimension reduction or quantization buys you.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>GPU Rental Price Reference</title>
      <link>https://vibeengines.com/tools/gpu-rental-prices</link>
      <guid isPermaLink="true">https://vibeengines.com/tools/gpu-rental-prices</guid>
      <category>Tool</category>
      <description>A cloud GPU price reference and cost estimator. Compare approximate on-demand hourly rates for common GPUs (T4, L4, A10G, A100, H100) and estimate what a training or inference job costs from the GPU count and hours. Rates are approximate and drift over time — always confirm current pricing with your provider.</description>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Optimizer Race</title>
      <link>https://vibeengines.com/lab/optimizer-race</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/optimizer-race</guid>
      <category>Lab</category>
      <description>Don't read about optimizers — race them down the same hill. Training a neural network is gradient descent on a loss surface, and the optimizer decides the path. On an ill-conditioned ravine, plain SGD zigzags and crawls, momentum builds speed along the valley floor, and Adam adapts its step size per direction to head almost straight for the bottom. Step the three optimizers down the same surface and watch their paths and losses diverge — made playable, with theory and a quiz.</description>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Loss Landscape</title>
      <link>https://vibeengines.com/lab/loss-landscape</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/loss-landscape</guid>
      <category>Lab</category>
      <description>Don't read about local minima — drop a ball and watch. Training is gradient descent on a loss landscape, and the shape decides whether descent finds the best answer, gets trapped in a worse one, or crawls to a halt. Roll a ball down a convex bowl (always finds the bottom), a double-well surface (where the start decides which minimum), and a flat plateau (where the gradient vanishes and progress stalls). See why the surface's shape governs training — made playable, with theory and a quiz.</description>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Inside a Transformer: Self-Attention</title>
      <link>https://vibeengines.com/lab/transformer-forward-pass</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/transformer-forward-pass</guid>
      <category>Lab</category>
      <description>Don't read about attention — compute it, token by token. At the heart of every transformer is self-attention: each token forms a query, compares it against every token's key to get scores, softmaxes those into weights, and builds its new representation as a weighted sum of every token's value. That's how a model lets 'it' look back at 'robot.' Pick a query token and watch its scores become softmax weights become a context vector — with the real dot-product math, a quiz, and theory.</description>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>RoPE: Position by Rotation</title>
      <link>https://vibeengines.com/lab/rope-positional-lab</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/rope-positional-lab</guid>
      <category>Lab</category>
      <description>Don't read about rotary position embeddings — rotate the vectors yourself. Attention has no built-in sense of order, so RoPE injects position by rotating each token's query and key by an angle proportional to its position. Because a dot product depends only on the angle between vectors, the attention score ends up depending only on the RELATIVE distance between tokens. Slide two tokens along a sequence and watch their score depend purely on the gap — made playable, with theory and a quiz.</description>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Latent Walk</title>
      <link>https://vibeengines.com/lab/gan-latent-walk</link>
      <guid isPermaLink="true">https://vibeengines.com/lab/gan-latent-walk</guid>
      <category>Lab</category>
      <description>Don't read about latent space — walk through it. A generative model turns a vector of numbers into an image, and that space is smooth and structured: nearby codes make similar images, a straight line between two codes morphs one into the other, and specific directions correspond to meaningful attributes. Move through a toy generator's latent space, interpolate between two points, and steer a single semantic direction to see how a model organizes what it can create — made playable, with theory and a quiz.</description>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Simulated Annealing: Cool to Improve</title>
      <link>https://vibeengines.com/algorithm/simulated-annealing</link>
      <guid isPermaLink="true">https://vibeengines.com/algorithm/simulated-annealing</guid>
      <category>Algorithm</category>
      <description>Don't memorize simulated annealing — play it. Optimize a hard problem (a traveling-salesman tour) by proposing small random changes: always accept improvements, but sometimes accept a worse move too — with a probability that shrinks as a 'temperature' cools. Early heat escapes local minima; late cooling settles into a great solution. Watch the tour untangle as the temperature drops — made playable, with theory and a quiz.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Hungarian Algorithm: The Perfect Match</title>
      <link>https://vibeengines.com/algorithm/hungarian</link>
      <guid isPermaLink="true">https://vibeengines.com/algorithm/hungarian</guid>
      <category>Algorithm</category>
      <description>Don't memorize the Hungarian algorithm — play it. Given a cost matrix of workers × tasks, find the assignment of one worker per task that minimizes total cost, without trying all n! matchings. Subtract each row's minimum, then each column's, so zeros mark the cheapest options — then pick n independent zeros. Watch the matrix reduce and the optimal assignment appear — made playable, with theory and a quiz.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Synthetic Data Handbook</title>
      <link>https://vibeengines.com/handbook/synthetic-data</link>
      <guid isPermaLink="true">https://vibeengines.com/handbook/synthetic-data</guid>
      <category>Handbook</category>
      <description>Using LLMs to generate training and eval data. Why quality filtering beats raw volume (effective size = generated × pass rate), what model collapse is and why recursive training on unfiltered self-generated data shrinks diversity (Var_k = s^k · Var_0 → 0), and a safe generate-filter-mix pipeline. With worked math and runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Sora</title>
      <link>https://vibeengines.com/paper/sora</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/sora</guid>
      <category>Paper</category>
      <description>Text-to-video by treating video as a diffusion transformer over spacetime patches. Sora compresses a video into a latent volume, cuts it into patches — one token each — and denoises the sequence, so resolution, aspect ratio, and duration are all just different patch counts. Why that unifies any shape into one model, and why patch count (linear in duration, quadratic in resolution) sets the compute — worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>V-JEPA</title>
      <link>https://vibeengines.com/paper/v-jepa</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/v-jepa</guid>
      <category>Paper</category>
      <description>A self-supervised video world model that predicts in representation space, not pixels. V-JEPA masks regions of a video and predicts their features (encoder outputs) rather than reconstructing pixels — so the target drops the unpredictable detail a reconstruction loss is forced to model. Why abstraction beats reconstruction, and how a stop-gradient EMA target encoder avoids representation collapse — worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>RT-2</title>
      <link>https://vibeengines.com/paper/rt-2</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/rt-2</guid>
      <category>Paper</category>
      <description>A Vision-Language-Action model that controls a robot by emitting actions as text tokens. RT-2 discretizes each dimension of a continuous action into bins — one vocabulary token per dimension — so a web-pretrained vision-language model co-trains on internet data and robot trajectories and its semantic knowledge transfers to control, generalizing to objects and commands never seen in robot data. Action-as-token, quantization precision, and emergent generalization — worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Genie</title>
      <link>https://vibeengines.com/paper/genie</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/genie</guid>
      <category>Paper</category>
      <description>A generative interactive environment that turns an image into a frame-by-frame playable world — trained on unlabeled video with no action labels. A latent action model infers a small discrete set of actions from consecutive frames, and a dynamics model turns them into control, so a tiny codebook forces consistent, reusable latent actions to emerge unsupervised. Latent-action inference and controllable generation — worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Dense Passage Retrieval</title>
      <link>https://vibeengines.com/paper/dpr</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/dpr</guid>
      <category>Paper</category>
      <description>The dual-encoder method that made dense retrieval beat BM25 for open-domain QA. DPR encodes questions and passages into a shared vector space and matches by dot product; its efficiency trick is in-batch negatives — a batch of B pairs yields a B×B similarity matrix whose diagonal is the positive and whose off-diagonal gives B×(B−1) negatives for free. Why matching on meaning beats keywords, and how the free-negatives trick trains it — worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>VAE</title>
      <link>https://vibeengines.com/paper/vae</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/vae</guid>
      <category>Paper</category>
      <description>The Variational Autoencoder — how to turn a compressor into a generator. A plain autoencoder's latent space is full of holes you can't sample; the VAE encodes each input to a Gaussian cloud and pulls all the clouds toward a standard-normal prior, making the space smooth and samplable. Trained by maximizing the ELBO (reconstruction minus KL), with the KL in closed form and the reparameterization trick z = μ + σ·ε making sampling differentiable. The foundation of the latent space under latent diffusion — worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>PPO</title>
      <link>https://vibeengines.com/paper/ppo</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/ppo</guid>
      <category>Paper</category>
      <description>Proximal Policy Optimization — the RL algorithm behind RLHF. Naive policy gradients take steps that collapse the policy; PPO caps the step with a clipped surrogate objective, min(r·A, clip(r, 1−ε, 1+ε)·A), where r = π_new/π_old is the probability ratio and A the advantage. Once the ratio leaves the trust region 1±ε in the helpful direction the objective flattens — TRPO-level stability from a first-order clip, no hard constraint. The optimizer InstructGPT and RLHF pipelines used. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>DQN</title>
      <link>https://vibeengines.com/paper/dqn</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/dqn</guid>
      <category>Paper</category>
      <description>The Deep Q-Network that learned Atari from raw pixels and launched deep reinforcement learning. A neural net approximates the Q-function, trained toward the Bellman target y = r + γ·max Q(s′,a′) — immediate reward plus discounted best future value (just r on a terminal step). Two stabilizers made neural Q-learning trainable: experience replay (random minibatches from a buffer break frame correlation) and a frozen target network (a stationary goal so the net doesn't chase itself). The ancestor of AlphaGo and modern deep RL. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Classifier-Free Guidance</title>
      <link>https://vibeengines.com/paper/classifier-free-guidance</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/classifier-free-guidance</guid>
      <category>Paper</category>
      <description>The &quot;guidance scale&quot; slider behind every diffusion image tool. A conditional diffusion model predicts the noise twice each step — with the prompt (ε_cond) and without it (ε_uncond) — then extrapolates: ε = (1+w)·ε_cond − w·ε_uncond, equivalently ε_uncond + (1+w)·(ε_cond − ε_uncond). w=0 is plain conditioning, larger w pushes further along the conditional direction for tighter prompt-following (too far → over-saturation). No separate classifier — trained with condition-dropout — and the source of negative prompts. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Knowledge Distillation</title>
      <link>https://vibeengines.com/paper/knowledge-distillation</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/knowledge-distillation</guid>
      <category>Paper</category>
      <description>How to compress a big, accurate &quot;teacher&quot; into a small, fast &quot;student&quot; — by training on the teacher's soft probability distribution, not just the hard label. A one-hot label is a thin signal; the teacher's full distribution encodes &quot;dark knowledge&quot; (a 2 looks a bit like a 7, nothing like a cat). The trick: a temperature-scaled softmax, p_i = exp(z_i/T) / Σ exp(z_j/T), softens the distribution as T grows so those tiny wrong-class probabilities become a strong training signal — without ever changing the top class. The student minimizes KL to the softened teacher plus cross-entropy on the labels. The &quot;Distil&quot; in DistilBERT and a first-line model-compression tool. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>VQ-VAE</title>
      <link>https://vibeengines.com/paper/vq-vae</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/vq-vae</guid>
      <category>Paper</category>
      <description>The autoencoder that made the latent DISCRETE — and quietly became the tokenizer behind image generation. The encoder outputs a continuous vector, but it is quantized to the nearest entry in a learned codebook {e_1..e_K}: k = argmin_j ‖z_e − e_j‖², so each latent position becomes one of K discrete tokens. The non-differentiable argmin is trained with a straight-through estimator (copy the decoder gradient back to the encoder), plus a three-part loss: reconstruction + codebook ‖sg[z_e]−e‖² + β·commitment ‖z_e−sg[e]‖². Discrete codes let a Transformer or diffusion prior model images/audio as tokens — the lineage behind VQ-VAE-2, DALL·E, and VQGAN. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Graph Neural Networks</title>
      <link>https://vibeengines.com/paper/graph-neural-networks</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/graph-neural-networks</guid>
      <category>Paper</category>
      <description>How a neural network learns on graphs — molecules, social networks, citation webs — where a convolution has nothing to slide over. The Graph Convolutional Network reduces it to message passing: each node updates itself by mixing its own feature with a degree-normalized average of its neighbors. One layer is H' = σ(Â H W) with Â = D̃^(-1/2)(A+I)D̃^(-1/2) — self-loops (A+I) keep a node's own feature and the symmetric normalization stops high-degree hubs from dominating. Stack layers to reach further; too deep and features over-smooth. The template (gather, aggregate, update) behind GraphSAGE and GAT — and attention is message passing on a fully connected graph. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>GloVe</title>
      <link>https://vibeengines.com/paper/glove</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/glove</guid>
      <category>Paper</category>
      <description>Word vectors from GLOBAL co-occurrence statistics — the count-based cousin of word2vec. Build a co-occurrence matrix X (X_ij = how often word j appears in the context of word i), then learn vectors so a dot product recovers the LOG count: w_i·w̃_j + b_i + b̃_j ≈ log(X_ij). The insight: co-occurrence RATIOS (P(solid|ice)/P(solid|steam)) carry meaning, and taking logs turns multiplicative ratios into differences a vector space represents — which is why king−man+woman≈queen works. Fit with weighted least squares, J = Σ f(X_ij)(…−log X_ij)², where f(x)=(x/xmax)^0.75 caps frequent pairs and f(0)=0 skips the empty matrix. The standard pretrained word embedding for years. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Lottery Ticket Hypothesis</title>
      <link>https://vibeengines.com/paper/lottery-ticket-hypothesis</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/lottery-ticket-hypothesis</guid>
      <category>Paper</category>
      <description>The finding that a dense, randomly-initialized network hides a small &quot;winning ticket&quot; — a sparse subnetwork that, trained in isolation FROM THE SAME INITIAL WEIGHTS, matches the full model's accuracy. You find it by iterative magnitude pruning: train, drop the smallest-magnitude weights (a binary mask), RESET survivors to their original init, repeat. The famous twist: randomly reinitializing the same sparse structure trains worse — structure and lucky initialization are entangled. After k rounds at prune fraction p, only (1−p)^k of the weights remain. A new lens on why over-parameterization helps, and the seed of sparse-training research. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>DDIM</title>
      <link>https://vibeengines.com/paper/ddim</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/ddim</guid>
      <category>Paper</category>
      <description>Denoising Diffusion Implicit Models — the trick that made diffusion sampling deterministic and 10–50× faster, using the SAME trained DDPM network (no retraining). Each step predicts the clean image x̂0 = (x_t − √(1−ᾱ_t)·ε)/√ᾱ_t, then re-projects to an earlier step: x_{t−1} = √ᾱ_{t−1}·x̂0 + √(1−ᾱ_{t−1}−σ²)·ε + σ·z. A single knob σ = η·√((1−ᾱ_{t−1})/(1−ᾱ_t))·√(1−ᾱ_t/ᾱ_{t−1}) unifies the two samplers: η=1 recovers stochastic DDPM, η=0 gives σ=0 and the deterministic DDIM that can skip timesteps. Determinism unlocks reproducibility, latent interpolation, and DDIM inversion for editing — and reframed sampling as an ODE. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Self-Instruct</title>
      <link>https://vibeengines.com/paper/self-instruct</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/self-instruct</guid>
      <category>Paper</category>
      <description>How to bootstrap a large instruction-tuning dataset from a language model's OWN generations, starting from just 175 human-written seed tasks. The generate→filter→add loop: sample existing tasks as examples, prompt the model to write new instructions + input/output instances, filter, add survivors back, repeat — 175 seeds bloom into ~52K diverse tasks. The quantitative heart is a diversity filter: keep a new instruction only if its ROUGE-L similarity (longest-common-subsequence overlap) to every existing one is below 0.7, so the pool never collapses into near-duplicates. The seed of Alpaca and the open instruction-tuning wave, and a landmark in synthetic data. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>mixup</title>
      <link>https://vibeengines.com/paper/mixup</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/mixup</guid>
      <category>Paper</category>
      <description>The two-line data augmentation that trains on CONVEX COMBINATIONS of example pairs — blending both inputs and labels: x̃ = λ·x_i + (1−λ)·x_j, ỹ = λ·y_i + (1−λ)·y_j, with λ ~ Beta(α,α). A 70/30 blend of a cat and a dog is trained with the soft label &quot;0.7 cat, 0.3 dog.&quot; This &quot;vicinal risk minimization&quot; makes the network behave linearly between training points, smoothing its decision boundary. The payoffs, for near-zero cost and no architecture change: better generalization, honest calibration, resistance to label noise, and adversarial robustness. A default augmentation for ResNets and ViTs (often with CutMix). Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>AlphaZero</title>
      <link>https://vibeengines.com/paper/alphazero</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/alphazero</guid>
      <category>Paper</category>
      <description>The single algorithm that mastered Go, chess, and shogi from ZERO human data — pure self-play, same code for all three. A single deep net outputs a policy prior P(s,a) and a value v(s); Monte Carlo Tree Search uses them to look ahead, selecting moves by the PUCT rule: maximize Q(s,a) + c·P(s,a)·√(ΣN)/(1+N(s,a)) — exploit the mean value Q, explore high-prior under-visited moves via the bonus that fades with visits. Search yields an improved policy (visit counts), self-play yields outcomes z, and both train the net (loss = (z−v)² − π·log p). Generalized AlphaGo into a domain-agnostic recipe; the seed of MuZero and the self-play-plus-search paradigm. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>wav2vec 2.0</title>
      <link>https://vibeengines.com/paper/wav2vec2</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/wav2vec2</guid>
      <category>Paper</category>
      <description>Self-supervised speech: learn from raw, untranscribed audio, then fine-tune on as little as TEN MINUTES of labels. A CNN encodes the waveform into latent frames; a span is masked; a Transformer context net must identify each masked frame's TRUE quantized latent among K distractors via a contrastive InfoNCE loss (cosine similarity / temperature): L = −log[exp(sim(c,q_true)/κ) / Σ exp(sim(c,q̃)/κ)]. Targets are discretized through a learned codebook (product quantization + Gumbel-softmax) so the model discovers phone-like units. Brought BERT-style masked pretraining and CLIP-style contrastive learning to audio; the basis of HuBERT and multilingual XLS-R. Worked math plus runnable code.</description>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Adam</title>
      <link>https://vibeengines.com/paper/adam</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/adam</guid>
      <category>Paper</category>
      <description>The optimizer that trains almost every modern network. Momentum as the first moment, RMSProp as the second, and the bias correction most explanations skip — plus the famously robust defaults (0.9 / 0.999 / 1e-8) and AdamW. Built up from scratch with worked math and runnable code you can edit in the browser.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Dropout</title>
      <link>https://vibeengines.com/paper/dropout</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/dropout</guid>
      <category>Paper</category>
      <description>Randomly switch off half your neurons on every training step — and overfitting collapses. Why co-adaptation hurts, how inverted dropout scales survivors to keep expectations unchanged, and why the whole thing is secretly an ensemble of 2ⁿ networks. Worked math plus runnable code you can edit in the browser.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Batch Normalization</title>
      <link>https://vibeengines.com/paper/batch-norm</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/batch-norm</guid>
      <category>Paper</category>
      <description>The layer that made deep networks train fast and forgivingly. Normalize each feature across the batch to zero mean and unit variance, then a learnable scale and shift give the network its freedom back. Internal covariate shift, the train-vs-test running-statistics gotcha, and why it unlocks higher learning rates — worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>ColBERT</title>
      <link>https://vibeengines.com/paper/colbert</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/colbert</guid>
      <category>Paper</category>
      <description>Keep one vector per token, not one per document, and score by late interaction: each query token takes its best match anywhere in the document (MaxSim), summed. It recovers the term-level precision single-vector retrieval blurs away, at index-friendly speed. The math, the case where it clearly wins, and how it sits between dense and cross-encoder retrieval — with runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Sentence-BERT</title>
      <link>https://vibeengines.com/paper/sentence-bert</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/sentence-bert</guid>
      <category>Paper</category>
      <description>The model that turned BERT into fast, comparable sentence embeddings — the ancestor of every embedding model behind vector search and RAG. Why plain BERT can’t be compared without a pass per pair, the siamese fix with mean pooling, and the O(n²)→O(n) arithmetic that took a task from 65 hours to 5 seconds. Worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Masked Autoencoders</title>
      <link>https://vibeengines.com/paper/mae</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/mae</guid>
      <category>Paper</category>
      <description>BERT for images, finally working. Mask 75% of an image’s patches, encode only the visible quarter with a ViT, and reconstruct the missing pixels with a light decoder. Why images need such a high mask ratio, the asymmetric design that makes the encoder ~16× cheaper, and the strong transfer results — worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>SimCLR</title>
      <link>https://vibeengines.com/paper/simclr</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/simclr</guid>
      <category>Paper</category>
      <description>Contrastive learning that rivaled supervised vision with no labels. Two augmented views of an image are a positive pair against a batch of negatives; the NT-Xent loss pulls positives together and pushes negatives apart. The temperature, why augmentation and big batches are decisive, and how contrastive compares to masked pre-training — worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Product Quantization</title>
      <link>https://vibeengines.com/paper/product-quantization</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/product-quantization</guid>
      <category>Paper</category>
      <description>The compression that makes billion-scale vector search fit in memory. Split a vector into subvectors, quantize each with a 256-entry codebook, and store a 512-byte vector in 8 — while m small codebooks span k^m codes. The compression math, asymmetric distance lookups, and IVFPQ — worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Muon</title>
      <link>https://vibeengines.com/paper/muon</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/muon</guid>
      <category>Paper</category>
      <description>The optimizer that orthogonalizes weight-matrix updates instead of scaling them element-wise like Adam. Pushing every singular value of the momentum to 1 spreads learning across all directions, done cheaply with Newton-Schulz iteration (matmuls, no SVD). Why it pairs with Adam for 1D params, and how it scaled to trillion-parameter training — worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Flamingo</title>
      <link>https://vibeengines.com/paper/flamingo</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/flamingo</guid>
      <category>Paper</category>
      <description>The visual language model that bridges a frozen vision encoder and a frozen LLM with a small trained connector. A Perceiver Resampler compresses each image into a fixed set of tokens, and gated cross-attention (tanh gate initialized at 0) injects them without breaking the language model — enabling few-shot multimodal in-context learning. The resampler and gating math — worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>DINOv2</title>
      <link>https://vibeengines.com/paper/dinov2</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/dinov2</guid>
      <category>Paper</category>
      <description>Label-free visual features by self-distillation: a student learns to match a teacher that is an EMA of itself, with centering and sharpening to prevent collapse. DINOv2 scaled this with curated data into a general vision backbone that works frozen across detection, segmentation, and depth. The self-distillation loop and anti-collapse math — worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>3D Gaussian Splatting</title>
      <link>https://vibeengines.com/paper/gaussian-splatting</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/gaussian-splatting</guid>
      <category>Paper</category>
      <description>Real-time photorealistic radiance fields. Represent a scene as millions of 3D Gaussians and render by projecting and alpha-compositing them front-to-back — fast rasterization instead of NeRF's slow ray-marching, hitting 100+ FPS. The compositing equation (C = Σ cᵢαᵢTᵢ), the differentiable optimization that fits Gaussians to photos, and why explicit beats implicit — worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>NeRF</title>
      <link>https://vibeengines.com/paper/nerf</link>
      <guid isPermaLink="true">https://vibeengines.com/paper/nerf</guid>
      <category>Paper</category>
      <description>The paper that launched the radiance-field era. Store a 3D scene as a small MLP mapping (position, direction) → (color, density), and render photorealistic novel views by volume rendering along rays. The rendering integral (α = 1 − exp(−σδ), C = Σ Tᵢαᵢcᵢ), why positional encoding unlocks sharp detail, and how it set up Gaussian Splatting — worked math plus runnable code.</description>
      <pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate>
    </item>
  </channel>
</rss>
