<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Records of the !mmortal Data Scientist</title><description>Machine learning research notes by Taha Bouhsine — neural network interpretability, kernel methods, contrastive learning, attention mechanisms, RKHS, and transformer architectures. Long-form pieces with interactive visualisations.</description><link>https://tahabouhsine.com/</link><item><title>I Removed Every MLP from Gemma 4 12B</title><link>https://tahabouhsine.com/removing-every-mlp-from-gemma-4/</link><guid isPermaLink="true">https://tahabouhsine.com/removing-every-mlp-from-gemma-4/</guid><description>Deleting every feed-forward branch removes 8.5 billion parameters and makes local Gemma 4 inference roughly three times faster. It also turns a capable language model into a machine that emits the audio control token forever.</description><pubDate>Tue, 11 Aug 2026 19:00:00 GMT</pubDate></item><item><title>A Network Made of Parts</title><link>https://tahabouhsine.com/patch-parts/</link><guid isPermaLink="true">https://tahabouhsine.com/patch-parts/</guid><description>Cut the image into patches, run one shared kernel bank over every patch, average the results, and classify. Linearity makes the output exactly decomposable into per-patch score contributions. The architecture was also supposed to fix three failures of the whole-image network. It fixed none: softening falls with distance, and concepts remain distributed across most of the bank at every tested granularity.</description><pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate></item><item><title>How to Interrogate a Kernel Network</title><link>https://tahabouhsine.com/yat-protocol/</link><guid isPermaLink="true">https://tahabouhsine.com/yat-protocol/</guid><description>A network whose hidden units are kernel prototypes is supposed to be legible. Legible claims are cheap unless someone can check them, so this post builds the checking: five instruments that put a trained Yat network under oath, each one asking a question that only this kernel makes askable. The first instrument finds that the softening constant in the formula sits ten thousand times below the distances it is supposed to soften, so the trained network never uses it at all: a term can be load-bearing in the theory and idle in the artifact, and only an audit tells you which.</description><pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The Concept That Would Not Die</title><link>https://tahabouhsine.com/spectral-surgery/</link><guid isPermaLink="true">https://tahabouhsine.com/spectral-surgery/</guid><description>An empirical feature covariance gives a trained kernel network ranked orthogonal axes. Deleting one axis is an exact algebraic intervention; the experiment asks whether it is also a semantic one. It is not: the damage spreads broadly and a small probe recovers the targeted distinction.</description><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The Trained Network, Under Mercer&apos;s Microscope</title><link>https://tahabouhsine.com/mercer-microscope/</link><guid isPermaLink="true">https://tahabouhsine.com/mercer-microscope/</guid><description>Every hidden representation induces an empirical kernel. Decompose it into ranked modes, audit their stability and semantic evidence, then repeat the measurement on grayscale CIFAR-100 where one hundred classes leave room for a genuine concept-count test.</description><pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Lazy Training a Yat Network in JAX/Flax NNX</title><link>https://tahabouhsine.com/lazy-training-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/lazy-training-jax-flax-nnx/</guid><description>A runnable companion to the lazy-training post: the Yat layer with its frozen softplus scalars, the one-line NNX filter that trains a readout while the bank stays frozen, the per-arm learning-rate bracketing, and the movement telemetry that caught the anti-lazy power law. With the run&apos;s own prototype trajectories: a Gram matrix crystallizing as neurons accumulate, two readouts racing on frozen banks, and eight random prototypes drifting through training without ever becoming pictures.</description><pubDate>Sun, 26 Jul 2026 23:00:00 GMT</pubDate></item><item><title>How Many Random Neurons Buy a Trained One?</title><link>https://tahabouhsine.com/lazy-training/</link><guid isPermaLink="true">https://tahabouhsine.com/lazy-training/</guid><description>Freeze a bank of randomly initialized Yat units and train only the linear readout. The induced random-feature kernel is a Monte Carlo average; under finite variance its estimation error has the familiar square-root scaling. This post measures that exponent, counts how many frozen units buy each rung of an accuracy ladder, and then unfreezes the centers to measure what feature learning adds.</description><pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Measuring Attention&apos;s Geometry in JAX/Flax NNX</title><link>https://tahabouhsine.com/the-geometry-of-attention-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/the-geometry-of-attention-jax-flax-nnx/</guid><description>A runnable companion to the geometry-of-attention post: the telemetry that exports a trained head&apos;s actual query and key vectors, the offline replay that recomputes every row&apos;s winner under rescaled queries, the linear program that asks whether any key sits inside the others&apos; convex hull, and the census of who owns whom through training. Plus the toy laws animated from their real formulas: territories re-carving, a body crossing the hull, the far sky losing its sign.</description><pubDate>Sun, 19 Jul 2026 23:00:00 GMT</pubDate></item><item><title>The Geometry of Attention Is a Choice of Kernel</title><link>https://tahabouhsine.com/the-geometry-of-attention/</link><guid isPermaLink="true">https://tahabouhsine.com/the-geometry-of-attention/</guid><description>An attention head is only ever evaluated at the tokens of the sequence, but its formula accepts any query vector at all. So between the tokens there is a whole continuous space, and the head silently carves it into territories: for every possible query, some key takes the top weight. This post draws that map for the two attention laws this series has trained head to head. The dot-product law slices space into infinite wedges of sky that all meet at the origin, a gravity that reads only bearings, where a query&apos;s length is a temperature and a key inside the others&apos; hull can never win. The Yat kernel carves bounded neighborhoods around bodies that always own their ground, a gravity that reads places, where length is an address. One theorem per map, one live map per theorem; the trained transformers appear only in a coda, to confirm they were never free to draw anything else.</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Softmax-Free Attention in JAX/Flax NNX</title><link>https://tahabouhsine.com/attention-is-a-compatibility-kernel-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/attention-is-a-compatibility-kernel-jax-flax-nnx/</guid><description>A runnable companion to the compatibility-kernel post: the attention module where one branch computes softmax and the other computes the Yat kernel with no exponential anywhere, the parameter-matched training harness, the telemetry that measured our bounded-scores belief dead, and the checkpointed attention maps. Every number is from the real Kaggle runs.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The Kernel Between the Roles</title><link>https://tahabouhsine.com/attention-is-a-compatibility-kernel/</link><guid isPermaLink="true">https://tahabouhsine.com/attention-is-a-compatibility-kernel/</guid><description>The QK post in this series ended on a construction it refused to build: keep the query and key roles, but replace the bilinear-then-exponential score with a genuine Mercer kernel between them. This post builds it, with the kernel&apos;s full form: a per-head learned bias inside the square, the term the universality theorem requires, and a per-head learned softening. Because the kernel is nonnegative by construction, attention needs no softmax at all and routing becomes a literal Nadaraya-Watson smoother. Trained head to head at matched parameters and per-variant swept learning rates, the kernel transformer lands within 1.1 percent of softmax on character-level Shakespeare, and the differences that survive are the interesting part: no gauge, no max-trick, a mass channel softmax cannot represent, and two of our own assumptions measured dead.</description><pubDate>Thu, 16 Jul 2026 23:00:00 GMT</pubDate></item><item><title>An Error Controller for a Trained Net, in JAX</title><link>https://tahabouhsine.com/depth-on-demand-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/depth-on-demand-jax-flax-nnx/</guid><description>A runnable companion to the depth-on-demand post: the leapfrog classifier trained with lax.scan at fixed depth, the step-doubling controller that re-renders it to tolerance at inference, the honest work accounting (probes included), and the measured tol^(-1/3) power law. Every figure is rendered from the real Kaggle run.</description><pubDate>Thu, 16 Jul 2026 22:00:00 GMT</pubDate></item><item><title>Depth on Demand</title><link>https://tahabouhsine.com/depth-on-demand/</link><guid isPermaLink="true">https://tahabouhsine.com/depth-on-demand/</guid><description>The last post made depth a resolution: layers are time steps of a learned flow, and running more of them just renders the same trajectory finer. But every camera knows not to spend equal film on empty sky. This post gives a trained network the integrator&apos;s next tool, an error controller that chooses its own step size per input, with no retraining: the same weights, rendered to tolerance. The controller reproduces the reference verdicts at a fraction of the steps, its cost follows the integrator&apos;s textbook one-third power law, and the map of where it spends is a genuine surprise: effort tracks the stiffness of the learned flow, not the difficulty of the classification.</description><pubDate>Thu, 16 Jul 2026 21:00:00 GMT</pubDate></item><item><title>Reversible Backprop as a custom_vjp in JAX</title><link>https://tahabouhsine.com/backprop-without-the-memory-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/backprop-without-the-memory-jax-flax-nnx/</guid><description>A runnable companion to the memory post: the momentum block and its exact inverse, the custom_vjp whose backward pass reconstructs the trajectory instead of storing it, XLA&apos;s memory_analysis as the measuring instrument, and the (1/mu)^L float cliff reproduced in numpy float32. Every number is from the real Kaggle run.</description><pubDate>Thu, 16 Jul 2026 20:00:00 GMT</pubDate></item><item><title>Backprop Without the Memory</title><link>https://tahabouhsine.com/backprop-without-the-memory/</link><guid isPermaLink="true">https://tahabouhsine.com/backprop-without-the-memory/</guid><description>Training memory is a tax nobody chose: backprop must hold every activation of the forward pass hostage until the backward pass consumes it, so depth costs memory even when it costs little compute. This post spends the invertibility of the momentum residual block: it can be run backward, so the backward pass can recompute the past instead of storing it. Measured on the same network, standard backprop&apos;s activation memory grows from 13 MB to 674 MB as depth goes 8 to 512; the reversible pass holds flat at 3.2 MB, pays 24% in step time, and returns the same gradient, until a one-line arithmetic of friction and float noise says it cannot.</description><pubDate>Thu, 16 Jul 2026 19:00:00 GMT</pubDate></item><item><title>Building the Hamiltonian-Step Net in JAX/Flax NNX</title><link>https://tahabouhsine.com/a-network-that-conserves-energy-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/a-network-that-conserves-energy-jax-flax-nnx/</guid><description>A runnable companion to the energy-conservation post: the HNN pendulum field as the symplectic gradient of one learned scalar, the plain field model it beats on drift, and the leapfrog classifier whose residual block is a kick-drift-kick step of a learned potential, all as Flax NNX modules with lax.scan doing depth. Every figure is rendered from the real Kaggle run.</description><pubDate>Thu, 16 Jul 2026 17:00:00 GMT</pubDate></item><item><title>A Network Built from Hamiltonian Steps</title><link>https://tahabouhsine.com/a-network-that-conserves-energy/</link><guid isPermaLink="true">https://tahabouhsine.com/a-network-that-conserves-energy/</guid><description>A pendulum organizes its motion around energy, while an ordinary residual network has no comparable scalar. This post builds a residual block as a symplectic step of a learned Hamiltonian. The continuous field conserves that Hamiltonian exactly; the discrete network keeps it in the bounded oscillatory band predicted by symplectic integration. The learned pendulum stays within 0.6% where a plain field model drifts 36%, and the classifier survives being run at four times its training depth.</description><pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Running the Survival Trial, in JAX/Flax NNX</title><link>https://tahabouhsine.com/survival-model-on-trial-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/survival-model-on-trial-jax-flax-nnx/</guid><description>Build the Yat DeepSurv trunk and Cox loss in Flax NNX, recover exact prototype contributions and explicit row edits, then run the LR-fair five-dataset benchmark with concordance, calibration, Brier, AUC, and classical baselines.</description><pubDate>Thu, 09 Jul 2026 17:00:00 GMT</pubDate></item><item><title>A Velocity Ledger for Transformers, in JAX/Flax NNX</title><link>https://tahabouhsine.com/transformers-with-a-velocity-ledger-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/transformers-with-a-velocity-ledger-jax-flax-nnx/</guid><description>A runnable companion: the pre-norm Transformer block as a forward-Euler step, then the residual-stream velocity ledger as one line of Flax NNX state (mu = 0 recovers plain), the ngpt-lite retraction variant, best-val early-stopped training, and the depth telemetry (path length and turning angle per sub-update). Four parameter-matched char-level GPTs that tie on quality and split on dynamics: the ledger&apos;s residual-stream path is a third as long and half as sharp.</description><pubDate>Thu, 09 Jul 2026 17:00:00 GMT</pubDate></item><item><title>Solving It and Descending It, in JAX/Flax NNX</title><link>https://tahabouhsine.com/you-dont-have-to-solve-a-kernel-machine-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/you-dont-have-to-solve-a-kernel-machine-jax-flax-nnx/</guid><description>A runnable companion to the solve-vs-descend post: the Yat kernel and its Gram matrix, the exact kernel ridge solve via Cholesky, the same kernel as a Flax NNX module trained by AdamW with LR sweeps and best-epoch selection, the measured timing wall, minibatching through 511k rows, and the conv trunk the solve can never train. Every number is from the real Kaggle runs.</description><pubDate>Thu, 09 Jul 2026 17:00:00 GMT</pubDate></item><item><title>The White-Box Survival Model on Trial</title><link>https://tahabouhsine.com/survival-model-on-trial/</link><guid isPermaLink="true">https://tahabouhsine.com/survival-model-on-trial/</guid><description>Build a survival network from learned prototype patients, derive its exact risk decomposition, and benchmark it on five datasets against Cox, penalized Cox, Random Survival Forest, and ReLU DeepSurv. Calibration, editing, and shift detection are measured separately.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Transformers With a Velocity Ledger</title><link>https://tahabouhsine.com/transformers-with-a-velocity-ledger/</link><guid isPermaLink="true">https://tahabouhsine.com/transformers-with-a-velocity-ledger/</guid><description>A pre-norm Transformer&apos;s residual stream is forward Euler: x += Attn(norm x); x += MLP(norm x). So the whole integrator dictionary transfers, and the same question follows: does a velocity ledger in the residual stream do for a Transformer what it did for a ResNet? The answer splits. On quality, four variants tie. On dynamics, the ledger changes everything: the residual-stream path through depth gets dramatically shorter and straighter, reaching the same answer by a calmer journey. Same destination, gentler road.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate></item><item><title>One Kernel Family, Fitted Two Ways</title><link>https://tahabouhsine.com/you-dont-have-to-solve-a-kernel-machine/</link><guid isPermaLink="true">https://tahabouhsine.com/you-dont-have-to-solve-a-kernel-machine/</guid><description>A dense kernel-ridge solve and a learned-center Yat expansion use the same kernel family but optimize different hypothesis classes. Their predictions correlate at 0.95 on housing; the compressed model then scales through datasets the dense baseline cannot hold.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Calibrating a Bounded Net, in JAX/Flax NNX</title><link>https://tahabouhsine.com/calibration-of-a-bounded-net-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/calibration-of-a-bounded-net-jax-flax-nnx/</guid><description>A runnable companion: build the matched Yat and ReLU MLPs in Flax NNX with the same softmax head, then measure their honesty. The reliability diagram and ECE, temperature scaling fit on a held-out split, NLL and Brier, and the two out-of-distribution channels, kernel-field magnitude versus softmax confidence, all in JAX with every number from a real three-seed run.</description><pubDate>Sat, 04 Jul 2026 17:00:00 GMT</pubDate></item><item><title>Building the Second Layer by Hand, in JAX/Flax NNX</title><link>https://tahabouhsine.com/depth-by-construction-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/depth-by-construction-jax-flax-nnx/</guid><description>A runnable companion: build a whole second feature layer by hand in JAX, on top of the hand-built first. Named min-AND combinations of layer-1 edges (junctions, continuations, bends, stripes) feed the same constructed Yat head, no training anywhere. It reproduces the flat rung: 83.3% at layer 1, 82.9% with both, 78.8% from relations alone, and counts the combinatorial wall of 224 pairwise and 4,630 three-way types where construction stops.</description><pubDate>Sat, 04 Jul 2026 17:00:00 GMT</pubDate></item><item><title>Distillation as Kernel Transfer, in JAX/Flax NNX</title><link>https://tahabouhsine.com/distillation-is-kernel-transfer-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/distillation-is-kernel-transfer-jax-flax-nnx/</guid><description>A runnable companion: the five-run distillation experiment in JAX/Flax NNX. Train a teacher CNN, extract its class-similarity kernel S = E[softmax(z/T) softmax(z/T)ᵀ], train a student on nothing but pairwise relations (no labels, no soft targets), and measure it against the label ceiling and the random floor with a linear and a nearest-centroid probe. Every number is from a real run, with six GIFs that animate the kernel assembling, the temperature dial, the handoff, the spectrum inheritance, the probe race, and the inherited mistakes.</description><pubDate>Sat, 04 Jul 2026 17:00:00 GMT</pubDate></item><item><title>Editing a Deep Equilibrium Network, in JAX/Flax NNX</title><link>https://tahabouhsine.com/edit-a-fixed-point-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/edit-a-fixed-point-jax-flax-nnx/</guid><description>A runnable companion: build the weight-tied Yat equilibrium operator in Flax NNX, then teach a class by appending rows to the readout (F untouched, exact) or into the shared dynamics (one paste, present at every depth). Measure local Jacobian slopes, temper the edit gain, audit 520 old fixed points, test multiple starts, watch a finite-prefix edit evaporate, and forget by masking. Every number is from a real run.</description><pubDate>Sat, 04 Jul 2026 17:00:00 GMT</pubDate></item><item><title>Skip Connections With Inertia, in JAX/Flax NNX</title><link>https://tahabouhsine.com/momentum-resnet-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/momentum-resnet-jax-flax-nnx/</guid><description>A runnable companion: the residual block as a forward-Euler step, then the momentum residual network as a Flax NNX module with one extra state, a velocity the blocks write into. Train both on the rings task a first-order flow cannot separate exactly, watch the training crystallize, and run the trained network exactly backward until floating point, amplified by 1/mu per layer, steals the past.</description><pubDate>Sat, 04 Jul 2026 17:00:00 GMT</pubDate></item><item><title>When 80% Should Mean 80%</title><link>https://tahabouhsine.com/calibration-of-a-bounded-net/</link><guid isPermaLink="true">https://tahabouhsine.com/calibration-of-a-bounded-net/</guid><description>A network hands you a probability with every answer, and the number is the part you act on. So when this series&apos; bounded, self-explaining kernel network says 80%, is that a measurement or a mood? Five posts of evidence say it should be the honest one. This post puts that reputation through a lie-detector test, reliability diagrams, expected calibration error and temperature scaling against a matched ReLU MLP on Fashion-MNIST, and what the test found is the post.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>How Far Down Can You Build?</title><link>https://tahabouhsine.com/depth-by-construction/</link><guid isPermaLink="true">https://tahabouhsine.com/depth-by-construction/</guid><description>One hand-built feature layer matched a trained backbone at 83.3% on Fashion-MNIST, and real networks are deep. Conveniently, the recipe for a second layer has been on the shelf for half a century: vision science says edges assemble into junctions, continuations, bends and stripes. This post takes the recipe down and follows it, builds layer 2 entirely by hand with every dimension still nameable in one sentence, and measures exactly where construction stops, and why.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Distillation Is a Geometry, Not an Answer Key</title><link>https://tahabouhsine.com/distillation-is-kernel-transfer/</link><guid isPermaLink="true">https://tahabouhsine.com/distillation-is-kernel-transfer/</guid><description>What crosses the wire in knowledge distillation besides the winning class? This experiment extracts a class-similarity kernel from teacher outputs and trains a student on pairwise relations alone—no labels, class names, or target probabilities. On Fashion-MNIST, the student recovers much of the label-trained geometry and approaches the spectrum of the transferred relation matrix.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Edit One Operator, Edit Every Depth</title><link>https://tahabouhsine.com/edit-a-fixed-point/</link><guid isPermaLink="true">https://tahabouhsine.com/edit-a-fixed-point/</guid><description>One post taught and forgot classes by editing rows of a Yat network, with proofs that nothing else moved. Another melted the stack of layers into a single operator iterated to a fixed point. This is the collision. Every one of those editing proofs rested on a pasted row entering the score once, as one term in one sum, and in an equilibrium network there is no once: whatever you paste is applied at every depth and fed back into its own input, and every fixed point is free to drift. So did melting the stack melt the editability? This post pastes, deletes, and measures: every guarantee that survives is either proved inside the recursion or measured against the real run, fixed point by fixed point.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Your Skip Connection Is Half of Newton</title><link>https://tahabouhsine.com/skip-connections-are-half-of-newton/</link><guid isPermaLink="true">https://tahabouhsine.com/skip-connections-are-half-of-newton/</guid><description>A residual block x + F(x) is one forward-Euler step: depth is time, the block is a velocity, position moves directly. That is half of Newtonian mechanics. A planet does not update position from force; force updates velocity, velocity updates position, and that split is why orbits are stable. So what does the missing half cost a deep network? We let the physics make three predictions about trained networks, then check all three live in the page. One of them comes back stranger than we wrote it.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate></item><item><title>A Network That Is a Fixed Point, in JAX/Flax NNX</title><link>https://tahabouhsine.com/your-network-is-a-fixed-point-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/your-network-is-a-fixed-point-jax-flax-nnx/</guid><description>A runnable companion: build the Yat deep-equilibrium network in JAX/Flax NNX. One shared operator F(z;x)=tanh(A·φ_W(z)+Ux+z0), solved by damped iteration and trained with implicit differentiation. Measure residual convergence, local Jacobian norms, and sensitivity to initialization instead of assuming a global contraction. Plus a weight-tied maze operator that reaches 99.5% on grids larger than training by iterating longer.</description><pubDate>Wed, 01 Jul 2026 17:00:00 GMT</pubDate></item><item><title>Your Network Is a Stack of Layers. It Could Be a Fixed Point.</title><link>https://tahabouhsine.com/your-network-is-a-fixed-point/</link><guid isPermaLink="true">https://tahabouhsine.com/your-network-is-a-fixed-point/</guid><description>A deep network makes you choose its depth before you have seen the problem, and gives every layer its own weights. Share one Yat-kernel operator across depth and the stack becomes a single equation: the answer is the fixed point reached by iteration. On the measured test trajectories, the solver converges from widely separated starts and the local Jacobian norm stays below one. The same twenty-four prototypes describe every step, reaching 98.2% on two moons from 1,700 shared parameters.</description><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate></item><item><title>A White-Box Kernel FFN in JAX/Flax NNX</title><link>https://tahabouhsine.com/mlp-block-is-a-representer-theorem-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/mlp-block-is-a-representer-theorem-jax-flax-nnx/</guid><description>A runnable companion: build a transformer whose feed-forward block is a finite learned-center kernel expansion. Train it on tinyshakespeare, then read each memory slot, attribute outputs exactly, edit one slot, and test peak kernel response as an abstention score.</description><pubDate>Sat, 27 Jun 2026 17:00:00 GMT</pubDate></item><item><title>The MLP Block Can Be a Kernel Memory</title><link>https://tahabouhsine.com/mlp-block-is-a-representer-theorem/</link><guid isPermaLink="true">https://tahabouhsine.com/mlp-block-is-a-representer-theorem/</guid><description>Replace an MLP activation with a kernel and its feed-forward block becomes an explicit learned-center expansion. Its slots can be read, attributed, and edited, but this architectural parameterization is not the classical representer theorem.</description><pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate></item><item><title>What Can a Weight Be?</title><link>https://tahabouhsine.com/what-can-a-weight-be/</link><guid isPermaLink="true">https://tahabouhsine.com/what-can-a-weight-be/</guid><description>A kernel is a spectral price list: it decides which functions are affordable, and regularization sets the budget. Compare Sobolev, Gaussian, spherical, and finite examples, then connect their eigenvalues to kernel ridge shrinkage and effective dimension.</description><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Where a Weight Lives, in JAX/Flax NNX</title><link>https://tahabouhsine.com/where-does-a-weight-live-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/where-does-a-weight-live-jax-flax-nnx/</guid><description>A runnable companion: build the representer-theorem weight in JAX. A positive-definite kernel, the Gram matrix, a single linear solve for the coefficients, and the weight comes out as a combination of the data, f = sum alpha_i k(x_i, .). A linear weight cannot separate nested rings; the placed kernel weight does, read purely through the kernel as a similarity-weighted vote of the data.</description><pubDate>Thu, 25 Jun 2026 17:00:00 GMT</pubDate></item><item><title>Where Does a Weight Live?</title><link>https://tahabouhsine.com/where-does-a-weight-live/</link><guid isPermaLink="true">https://tahabouhsine.com/where-does-a-weight-live/</guid><description>A standard neuron&apos;s weight and its input never actually meet: one is a point you can see, the other an arrow off in its own space, joined only by a shadow. This is what a reproducing kernel Hilbert space fixes: it gives input and weight one shared address, where the optimal weight is built from the data itself and sits right next to it. Four interactive panels.</description><pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Constructing the Fashion-MNIST Network, in JAX/Flax NNX</title><link>https://tahabouhsine.com/train-the-features-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/train-the-features-jax-flax-nnx/</guid><description>Train a small backbone and place its prototype head, then remove training entirely: implement fixed Sobel orientation channels, pool them into 343 named features, and classify with a constructed Yat head.</description><pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate></item><item><title>How Much of a Fashion-MNIST Network Can You Build by Hand?</title><link>https://tahabouhsine.com/train-the-features/</link><guid isPermaLink="true">https://tahabouhsine.com/train-the-features/</guid><description>Construct the prototype head on random and learned features, then replace the backbone with named edge and corner measurements. On Fashion-MNIST the zero-training pipeline reaches 83.3%, versus 85.7% for the matched trained model.</description><pubDate>Mon, 22 Jun 2026 23:00:00 GMT</pubDate></item><item><title>Editing a Network by Hand, in JAX/Flax NNX</title><link>https://tahabouhsine.com/edit-a-network-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/edit-a-network-jax-flax-nnx/</guid><description>A runnable companion: build the prototype Yat-MLP in Flax NNX, then add a class by concatenating a few prototype rows and forget a class by masking them out, with no gradient steps. Class-incremental learning that matches a from-scratch build, and exact machine unlearning, both as array edits you can read. Every number is from a real run on Fashion-MNIST.</description><pubDate>Mon, 22 Jun 2026 18:00:00 GMT</pubDate></item><item><title>A Kernel&apos;s Price List, in JAX</title><link>https://tahabouhsine.com/what-can-a-weight-be-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/what-can-a-weight-be-jax-flax-nnx/</guid><description>Construct valid periodic kernel spectra, audit truncated RKHS norms for convergence, then solve kernel ridge regression and read regularization as spectral shrinkage, effective dimension, and a measured generalization curve.</description><pubDate>Mon, 22 Jun 2026 17:00:00 GMT</pubDate></item><item><title>Your Network Is a List of Pictures. You Can Edit It.</title><link>https://tahabouhsine.com/edit-a-network-by-hand/</link><guid isPermaLink="true">https://tahabouhsine.com/edit-a-network-by-hand/</guid><description>If a neuron is a labelled picture, a classifier is a list of them, and a list is something you edit. Add a class to a trained-free Yat-kernel network by placing twenty pictures, and it recognizes that class at 95% with zero gradient steps. Delete a class by removing its pictures, and it is forgotten exactly, the other classes untouched. Class-incremental learning with no penalty and machine unlearning that is instant and exact, both falling out of the architecture rather than bolted on.</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate></item><item><title>The Yat-Kernel MLP in JAX/Flax NNX</title><link>https://tahabouhsine.com/yat-mlp-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/yat-mlp-jax-flax-nnx/</guid><description>Build a finite bank of Yat-kernel sections in JAX/Flax NNX, verify the kernel, train it on two moons and Fashion-MNIST, inspect exact prototype contributions, and test the initialization control that separates visible centers from noisy ones.</description><pubDate>Thu, 18 Jun 2026 17:00:00 GMT</pubDate></item><item><title>What a Finite Kernel Buys an MLP</title><link>https://tahabouhsine.com/what-a-finite-kernel-buys-an-mlp/</link><guid isPermaLink="true">https://tahabouhsine.com/what-a-finite-kernel-buys-an-mlp/</guid><description>Replace the activation with a finite bank of learned kernel sections. The resulting MLP exposes prototypes, exact layer-local contributions, measurable geometry, and the conditions those claims require, then tests the construction on arithmetic and Fashion-MNIST.</description><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate></item><item><title>The Three States of Information, in JAX</title><link>https://tahabouhsine.com/three-states-of-information-jax/</link><guid isPermaLink="true">https://tahabouhsine.com/three-states-of-information-jax/</guid><description>A runnable companion to The Three States of Information: train tiny models in JAX and measure the three states directly: the feature-covariance spectrum collapsing from high-rank (random) to a C−1-mode frame (structured), the distributional simplicity bias that fits low-order structure first (organized), the neural-collapse simplex where class-mean cosines lock onto −1/(C−1), and the alignment/uniformity split of contrastive learning running on two separate clocks. Four live JAX visualizations, every number an eigenvalue or a loss.</description><pubDate>Sun, 07 Jun 2026 19:45:00 GMT</pubDate></item><item><title>The Three States of Information</title><link>https://tahabouhsine.com/three-states-of-information/</link><guid isPermaLink="true">https://tahabouhsine.com/three-states-of-information/</guid><description>In these training runs, representation geometry moves through three recognizable regimes: random, organized into local clusters, and globally structured around separated class means. Interactive experiments test when loss plateaus coincide with those reorganizations—and when schedules change the order.</description><pubDate>Sun, 07 Jun 2026 19:30:00 GMT</pubDate></item><item><title>Latent on the Spectrum, in JAX</title><link>https://tahabouhsine.com/latent-on-the-spectrum-jax/</link><guid isPermaLink="true">https://tahabouhsine.com/latent-on-the-spectrum-jax/</guid><description>A runnable companion to Latent on the Spectrum: build a codebook as the spectral embedding of a label kernel in JAX (classical MDS with square-root eigenvalue scaling), watch a flat spectrum become the simplex and a graded one become the horseshoe, measure kernel-target alignment, split a representation into its between-class prototype frame and within-class information spectrum, and watch neural collapse grind the information to zero.</description><pubDate>Sun, 07 Jun 2026 19:15:00 GMT</pubDate></item><item><title>Latent on the Spectrum: Why Cats Sit Closer to Dogs Than to Cars</title><link>https://tahabouhsine.com/latent-on-the-spectrum/</link><guid isPermaLink="true">https://tahabouhsine.com/latent-on-the-spectrum/</guid><description>A label-similarity kernel can be turned into a target codebook by spectral embedding: retain its leading eigenmodes, scale by their square roots, and spend a finite dimension budget. Interactive experiments move that designed geometry from a simplex toward a taxonomy, then compare it with the class-mean and within-class spectra measured in trained representations.</description><pubDate>Sun, 07 Jun 2026 19:00:00 GMT</pubDate></item><item><title>Q and K Projections in JAX/Flax NNX</title><link>https://tahabouhsine.com/qk-projections-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/qk-projections-jax-flax-nnx/</guid><description>A runnable companion to Why Attention Needs Q and K Projections: build scaled dot-product attention with separate query and key projections in Flax NNX, pull the bilinear form B = W_Q W_Kᵀ out of the module, split it into a symmetric metric and an antisymmetric directed part, wire a toy induction head, add RoPE, and measure the low-rank budget and the gauge freedom, all in plain JAX.</description><pubDate>Thu, 04 Jun 2026 19:45:00 GMT</pubDate></item><item><title>Why Attention Needs Q and K Projections</title><link>https://tahabouhsine.com/why-attention-needs-qk-projections/</link><guid isPermaLink="true">https://tahabouhsine.com/why-attention-needs-qk-projections/</guid><description>The dot product in attention is not enough by itself. Without learned query and key projections, attention can only compare tokens in the residual stream’s native geometry. With a shared projection it learns a symmetric metric. With separate Q and K projections, the score becomes a learned bilinear form x_iᵀW_QW_Kᵀx_j: directional, role-aware, low-rank, and different per head. That bilinearity is what lets attention ask one kind of question and let tokens advertise another kind of answer.</description><pubDate>Thu, 04 Jun 2026 19:30:00 GMT</pubDate></item><item><title>The Prototype Readout in JAX/Flax NNX</title><link>https://tahabouhsine.com/convex-readout-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/convex-readout-jax-flax-nnx/</guid><description>A runnable companion to The Readout is a Convex Combination of Prototypes: read the columns of W_out as output prototypes in Flax NNX, measure the convex/conic/affine/linear regimes numerically, then build a Nadaraya–Watson kernel readout that is convex by construction (nonnegative weights that sum to one, a point that never leaves the prototype hull), with the nonnegativity-vs-positive-definiteness distinction checked in code.</description><pubDate>Thu, 04 Jun 2026 19:15:00 GMT</pubDate></item><item><title>The Readout is a Convex Combination of Prototypes</title><link>https://tahabouhsine.com/readout-as-convex-combination/</link><guid isPermaLink="true">https://tahabouhsine.com/readout-as-convex-combination/</guid><description>The second linear map in a transformer MLP is a dictionary of output prototypes, one per hidden unit. If the hidden activations are nonnegative and normalized, W_out reads the active neurons as a convex combination of output prototypes. Two independent constraints, nonnegativity and summing to one, sort the readout into four regimes: convex, conic, affine, and linear. This reframes the MLP readout as the same object that makes attention legible (a weighted sum over named basis elements), connects it to feed-forward key-value memories and modern Hopfield retrieval, and shows when a kernel makes it convex by construction.</description><pubDate>Thu, 04 Jun 2026 19:00:00 GMT</pubDate></item><item><title>Auditing Latent Space Geometry in JAX</title><link>https://tahabouhsine.com/welch-bound-jax-analysis/</link><guid isPermaLink="true">https://tahabouhsine.com/welch-bound-jax-analysis/</guid><description>A runnable companion to the Welch-bound latent-space post: generate GIFs and implement the JAX metrics that tell you whether embeddings are collapsing, wasting rank, forming a simplex, or pressing against the Welch floor.</description><pubDate>Tue, 02 Jun 2026 19:00:00 GMT</pubDate></item><item><title>What Makes a Good Latent Space? The Welch Bound and the Simplex</title><link>https://tahabouhsine.com/welch-bound-good-latent-space/</link><guid isPermaLink="true">https://tahabouhsine.com/welch-bound-good-latent-space/</guid><description>The hidden codebook inside representation learning: why collapse happens, why opposition is a trap, why class means form a simplex, and why the Welch bound sets the best geometry when too many concepts share too few dimensions.</description><pubDate>Tue, 02 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Cheap Attention in JAX/Flax NNX</title><link>https://tahabouhsine.com/linear-attention-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/linear-attention-jax-flax-nnx/</guid><description>A runnable companion to Cheap Attention: implement positive-feature linear attention in JAX and Flax NNX, watch the all-pairs ledger turn into a shared feature state, and see where the N×N matrix disappears.</description><pubDate>Sun, 31 May 2026 19:30:00 GMT</pubDate></item><item><title>Cheap Attention: Linear-Time Kernel Approximation</title><link>https://tahabouhsine.com/cheap-attention-is-linear-attention/</link><guid isPermaLink="true">https://tahabouhsine.com/cheap-attention-is-linear-attention/</guid><description>A 128K-token context creates billions of pairwise questions per attention head. But the N×N matrix is not the essence of attention; it is the receipt for an infinite feature map we never wrote down. Approximate that feature map with random features, reassociate the sum, and softmax attention becomes linear-time kernel attention.</description><pubDate>Sun, 31 May 2026 00:00:00 GMT</pubDate></item><item><title>Organizing Randomness: Contrastive Learning in JAX</title><link>https://tahabouhsine.com/organizing-randomness-jax/</link><guid isPermaLink="true">https://tahabouhsine.com/organizing-randomness-jax/</guid><description>A block-by-block JAX + Optax implementation of six contrastive losses, each watched as a real animated GIF turning random 2D points into organized embeddings. The runnable companion to &quot;Untangling the Moons.&quot;</description><pubDate>Tue, 26 May 2026 19:00:00 GMT</pubDate></item><item><title>Untangling the Moons: A Visual History of Contrastive Learning</title><link>https://tahabouhsine.com/untangling-the-moons/</link><guid isPermaLink="true">https://tahabouhsine.com/untangling-the-moons/</guid><description>Eight contrastive losses, twenty years of history, and one geometric audit. Watch the losses organize the same 2D points while separating opposition, orthogonality, simplex packing, and statistical independence.</description><pubDate>Tue, 26 May 2026 00:00:00 GMT</pubDate></item><item><title>Self-Attention as Kernel Regression in JAX/Flax NNX</title><link>https://tahabouhsine.com/attention-is-kernel-jax-flax-nnx/</link><guid isPermaLink="true">https://tahabouhsine.com/attention-is-kernel-jax-flax-nnx/</guid><description>A runnable companion to Attention is Explainable Because it is a Kernel: build scaled dot-product attention from scratch in Flax NNX, prove in code that it is exactly a Nadaraya–Watson kernel smoother, watch the separate q/k projections break positive-definiteness numerically, swap the exp-dot-product kernel for Gaussian, Yat, and linear kernels to see which keep the weights a convex partition of unity, read the temperature as a kernel bandwidth, and train a single head end-to-end to route to a marked token.</description><pubDate>Thu, 14 May 2026 17:00:00 GMT</pubDate></item><item><title>What Attention Weights Can Explain</title><link>https://tahabouhsine.com/attention-is-a-kernel/</link><guid isPermaLink="true">https://tahabouhsine.com/attention-is-a-kernel/</guid><description>Self-attention has the normalized weighted-average form of a compatibility smoother. That exposes exact routing arithmetic, but it does not make attention weights causal explanations or guarantee a Mercer kernel on tokens.</description><pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate></item><item><title>What Activations Do to Geometry</title><link>https://tahabouhsine.com/activations-are-bad-for-geometry/</link><guid isPermaLink="true">https://tahabouhsine.com/activations-are-bad-for-geometry/</guid><description>ReLU, GELU, and their relatives enter a layer&apos;s Jacobian as an input-dependent row scaling. Here is when that scaling erases directions, when it merely distorts them, and what the usual repairs actually guarantee.</description><pubDate>Fri, 20 Feb 2026 00:00:00 GMT</pubDate></item></channel></rss>