Naive Bayesian Classifiers: How They Actually Work
Naive Bayesian classifiers are built on an assumption that’s almost always technically wrong. Here’s why it works anyway, and when it’s the right tool.
Read more →Notes on data engineering, AI and building things end to end.
Naive Bayesian classifiers are built on an assumption that’s almost always technically wrong. Here’s why it works anyway, and when it’s the right tool.
Read more →
Plenty of lines can separate two classes of points. A support vector machine specifically finds the one with the widest possible margin. Here’s how and why.
Read more →
Gradient boosting gets described often and explained rarely. Here’s how it actually works, step by step, and the hyperparameters that genuinely matter.
Read more →
A for-loop over a million rows works, just slowly. Here’s vectorization vs loops in Python, why vectorized operations win, and when a loop is still the right call.
Read more →
Summing pixel values inside a rectangle, over and over, would be unbearably slow done naively. Here’s what an integral image is, and why it makes that instant instead.
Read more →
Before running a statistical test, two competing statements need to be written down clearly. Here’s null hypothesis vs alternative hypothesis, and how to write both correctly.
Read more →
No labels, no correct answer to check, just raw data and a question about structure. Here’s what unsupervised clustering actually does, and how to judge if it’s good.
Read more →
The L2 regularization formula adds a squared-magnitude penalty to the loss function. Here’s what each term means, with a worked numeric example.
Read more →
L1 and L2 get all the attention, but the L0 norm is the most direct way to penalize model complexity. Here’s what it actually is, and why L1 approximates it instead.
Read more →
A query scanning a billion rows to find last week’s data is doing unnecessary work. Here’s how table partitioning strategies fix that, and how to pick the right key.
Read more →