L1 and L2 get all the attention in regularization discussions. The L0 norm rarely does, and there’s a specific reason for that: it’s the most direct way to penalize model complexity, and also the hardest one to actually optimize.
What the L0 norm actually measures
The L0 norm counts the number of nonzero entries in a vector. Applied to a model’s coefficients, L0 regularization penalizes exactly how many features the model is using, not their magnitude the way L1 and L2 do. Conceptually, it’s the cleanest possible statement of “prefer a simpler model”: fewer active features, full stop.Why L0 regularization isn’t used directly in practice
Counting nonzero entries is a discrete, non-differentiable quantity. Standard gradient-based optimization, the backbone of how most models are trained, needs a smooth, differentiable objective to work with. L0 regularization breaks that assumption directly, which makes exact L0-penalized optimization computationally intractable for anything beyond a small number of features.
How L1 fills the gap
L1 regularization is often described as the practical stand-in for L0. It’s the tightest convex relaxation of the L0 norm, meaning it approximates the same goal, sparsity, few active features, while staying differentiable enough (almost everywhere) for standard optimization to handle. This is the real reason L1 regularization zeroes out coefficients the way it does: it’s chasing the same sparsity goal the L0 norm defines exactly, through a route that’s actually solvable.
Where L0 does get used
Exact L0 optimization shows up in specific, constrained settings, feature selection algorithms designed to handle it directly, sparse approximation problems in signal processing, and some specialized model compression techniques where the discrete feature count matters enough to justify the computational cost of approximating it more directly.
A real example: why this connects to feature selection
The feature selection approaches covered elsewhere, filter, wrapper, and embedded methods, all aim at a version of what L0 regularization defines mathematically: use fewer, better-chosen features. L1’s embedded selection behavior is the practical route to that goal that most tools actually implement, precisely because true L0 optimization is so much harder to run at scale.
A quick checklist
- Are you trying to minimize the exact count of active features, or just shrink their overall influence? That distinction is L0 versus L2’s actual goal.
- Is your feature set small enough that exact L0 optimization is computationally feasible, or does the problem call for L1’s differentiable approximation instead?
- Do you understand L1 as an approximation of L0, or as an unrelated technique? The connection changes how you’d explain why L1 produces sparse models.
- Would a specialized sparse optimization method be worth the added complexity here, or is L1 regularization’s approximation good enough?
FAQ
Is the L0 norm actually a mathematical norm?
Not in the strict sense, it fails the scaling property a true norm requires, but it’s conventionally called a norm because it fits the same conceptual role in regularization.
Why not just use L0 regularization if it’s the most direct approach?
Because optimizing it exactly is NP-hard in general, which makes it impractical for most real model training, even though it’s conceptually the cleanest formulation.
Does L1 regularization always match what L0 would have chosen?
Not exactly, it’s an approximation, not an identical result, but it’s specifically designed to be the closest differentiable stand-in available for standard optimization.

