01 · Metaheuristic optimization
A neural network that doesn't learn a task — it becomes the optimizer. Watch a population of guesses flow through a weight matrix toward the best answer they can find.
figure 1 · the shape
Eight parent candidates on the left. A weight matrix in the middle — drawn as edges, each one's thickness the magnitude of a weight. Eight child candidates on the right. Signals flow continuously along the edges; this is what one iteration of the algorithm looks like. There is no loss function and no labels — the network's only job is to mix the parents into the next generation.
Most uses of a neural network treat it as a function approximator. You feed it inputs, it produces outputs, and you train its weights so the outputs match labels. The Neural Network Algorithm — NNA — takes the same shape and uses it for something else entirely. Here, the weights and biases are the optimizer. Their job is not to model your data; it is to shuffle a population of guesses toward a good solution to whatever black-box problem you hand it.
Three pieces do all the work. The weight matrix mixes old candidates into new ones. The bias operator sprinkles randomness so the search keeps exploring. The transfer operator tugs candidates toward the best answer found so far. We'll meet each one as its own toy, then run them all together — but first, a primer on what NNA is borrowing from.
A neural network is a directed graph of tiny identical computations called neurons. Each neuron does the same three things: it takes a list of inputs x₁, x₂, …, xₙ, weights each one by a learned coefficient wᵢ, sums those weighted inputs together with a learnable bias b, and feeds the result through a nonlinear activation function σ. One scalar comes out.
Neurons are arranged in layers. The first layer reads the raw input — pixels of an image, tokens of a sentence, a vector of features. Each subsequent layer takes the previous layer's outputs as its inputs. The connections between two adjacent layers are encoded in a single weight matrix W: one entry per (input neuron, output neuron) pair. Computing a whole layer's output is then just a matrix-vector multiply followed by the activation: x′ = σ(W·x + b). Stack a handful of those and you have a feedforward network; information flows in one direction, input to output, layer by layer.
The activation function is what makes a network more than a stack of linear transforms. The classic choices: the sigmoid σ(z) = 1 / (1 + e⁻ᶻ), a smooth S-curve that squashes any input into (0, 1); the hyperbolic tangent tanh(z), an S-curve into (-1, 1); and ReLU max(0, z) — by far the most common today despite being embarrassingly simple. The math reason a nonlinearity matters: composing two linear maps gives you another linear map. Without σ between layers, a hundred-layer network collapses into a single matrix and can only model lines and planes. With σ in the middle, networks can in principle approximate any continuous function — the so-called universal approximation theorem.
How does a network learn? You hand it training examples and compute a loss — a number measuring how wrong the network's output was. The gradient of that loss with respect to every weight is computed by backpropagation: the chain rule, applied layer by layer from output back to input. Each weight is then nudged a small step in the direction that reduces the loss. This is gradient descent. Iterate over millions of examples and the network shapes itself to the task.
That whole machinery — the weighted sum, the activation, the layer-by-layer matrix multiply, the gradient-driven learning — is the backbone of nearly every model in modern AI: convolutional networks for vision, recurrent networks for sequences, transformers for language. They are all variations on this one shape, with different wiring patterns and different ways of computing W·x.
And it is this shape, not the learning, that NNA borrows. Look back at figure 1: input nodes, a weight matrix in the middle, output nodes on the right. The picture is identical to a single layer of a feedforward network. But the role is inverted. There is no loss, no backprop, no labels. The "inputs" are not features of a data point — they are candidate solutions to some optimization problem. The "outputs" are the next generation of candidates. The weight matrix is not learning to classify; it is learning which mixing recipe produces better candidates. NNA takes the algebra of a feedforward network and re-purposes it as a population-based optimizer.
figure 2 · the weight matrix
The weight matrix W couples the population to itself: each new candidate is a convex blend of the old ones, which is why every row has to sum to one. Click any cell — the cell you bumped grows brighter, the others in its row dim to compensate, and the row sum on the right snaps back to 1.00. This is the network's only learnable parameter, and across iterations it bends toward whatever mixing-recipe produced the best candidate.
figure 3 · the transfer operator
At every iteration each non-best candidate gets pulled part of the way toward the best solution found so far. Drag the slider to set how hard the rope pulls toward the target. Pull too hard and the swarm collapses before it has explored; pull too gently and convergence stalls. NNA starts gentle and tightens as iterations go by.
figure 4 · the bias operator
Drag the orange handle. The blue cloud around it is the bias distribution — a Gaussian over the search space. Every iteration NNA samples a tiny perturbation from this cloud and adds it to each candidate's position, which is why you see the small dots flicker into existence inside the halo. Early on the spread is wide and the swarm wanders; over time the spread decays and exploration gives way to exploitation.
Now the three pieces in a single iteration. Scroll slowly through the next figure — the animation is bound to your scroll position. Edges light up in order; child candidates emerge as their weighted blends from the parents.
figure 5 · one iteration · scroll-driven
Each child position is x′ⱼ = Σᵢ Wⱼᵢ · xᵢ — a weighted sum of every parent. The bright edges contributed the most.
Repeat that step. Then repeat again, with the weight matrix re-tilted toward what worked, the bias halo a little tighter, the transfer pull a little stronger. Scroll through the next section and watch six iterations play out in your hands.
figure 6 · convergence · scroll-driven · interactive
Drag the orange dot to move the optimum anywhere on the landscape. Slide the population up to crowd the search, or down to thin it. Lower the bias decay to make the swarm stay exploratory longer. The scroll-bar is the iteration clock — scroll back up to rewind.
Gradient descent is wonderful when you can take a derivative. Many real problems — hyperparameter search, combinatorial design, simulation tuning — give you only a black-box score. Metaheuristics exist for that world. NNA's twist is that its update rule borrows the algebra of a neural network rather than physics, biology, or thermodynamics. It's worth comparing to its closest cousin.
figure 7 · comparison
Both algorithms run continuously on the same landscape. NNA (blue) moves as a coordinated swarm — a weight matrix smoothly mixing the whole population each step. The GA (pink) crosses over pairs of parents with mutation noise, producing jumpier, less coherent motion. Both find the optimum; they take different paths there.