The distinction usually gets taught as a definition to memorize. It’s more useful as a practical question: does your data already have the answer attached, or are you asking the model to find structure with no answer key at all.
Supervised vs unsupervised learning: what separates them
Supervised learning trains on labeled examples, inputs paired with known correct outputs, and learns to predict that output for new inputs. Unsupervised learning works with unlabeled data and looks for structure, clusters, patterns, groupings, without ever being told what the “right” grouping is. Supervised vs unsupervised learning comes down to whether ground truth exists in the training data at all.
A real example: supervised learning with labeled outcomes
The network intrusion detection project trains on the NSL-KDD benchmark, where every row of network traffic is labeled as normal or a specific attack type. That label is what makes it supervised: the model learns from examples where the correct answer was already known, then applies that learned pattern to new, unlabeled traffic.
A real example: labeled data used for triage, not just prediction
The fact-check triage NLP project is supervised in the same way, trained on the LIAR dataset’s labeled truthfulness categories. Both projects depend on someone, or some process, having already assigned the correct label before training ever starts.
Where this project set doesn’t include a pure unsupervised example
Most of the projects here, the recommenders, the classifiers, the forecasting model, are supervised or make direct use of labeled outcomes. Clustering customers without predefined segments, or reducing dimensionality to find latent structure with no target variable, are the kinds of tasks that would sit on the unsupervised side, and they ask a genuinely different question: not “predict this known outcome” but “what structure exists here that nobody labeled yet.”
Why the distinction matters for evaluation, not just training
Supervised learning can be evaluated against ground truth directly, which is exactly what enables the kind of honest, per-class evaluation used throughout these projects, comparing predictions against known correct answers. Unsupervised learning has no such ground truth to check against, which makes evaluation fundamentally harder: judging whether a clustering is “good” often relies on indirect measures or domain judgment rather than a clean right-or-wrong comparison.
A quick checklist
- Does your dataset have a known, correct output for each example, or are you looking for structure with no labels at all?
- If labels exist but are incomplete or noisy, does that change which category your problem actually falls into?
- Can you evaluate results against ground truth, or will evaluation need to rely on indirect or human judgment?
- Would combining both, using unsupervised clustering to generate candidate labels for a later supervised step, actually fit your problem better than either alone?
FAQ
Is semi-supervised learning a real third category?
Yes. It trains on a mix of labeled and unlabeled data, useful when labeling everything would be too expensive but some labeled examples exist.
Can unsupervised learning be used to help a supervised task?
Often, yes. Clustering or dimensionality reduction can generate useful features that a downstream supervised model then uses for prediction.
Is reinforcement learning supervised or unsupervised?
Neither, cleanly. It learns from reward signals through interaction with an environment, which is a distinct paradigm from both.

