Supervised Learning definition
Supervised learning is a type of machine learning in which a model is trained on labeled examples, meaning each input comes with the correct answer. The model learns the relationship between inputs and labels so it can predict the answer for new data, such as whether an email is spam or what a house will sell for.
How does supervised learning work?
You start with a dataset of input and output pairs, for instance past loan applications with a column showing whether each loan was repaid. The data is split into training, validation and test sets. A training algorithm fits a model to the training set, the validation set is used to tune settings such as tree depth or learning rate, and the test set gives a final, honest estimate of performance on applications the model has never seen.
Worked example: a lender has five years of applications labeled repaid or defaulted. Features include income, existing debt, loan amount and repayment history. A gradient-boosting model learns which combinations predict default, and the lender sets a score threshold that sends risky applications to a human underwriter instead of rejecting them automatically.
How to build a supervised learning model
- Frame the target precisely, for example "customer cancels within 30 days of the renewal date".
- Gather historical records and attach the correct label to each one.
- Split data by time where possible, so the test set reflects the future rather than the past.
- Train a simple baseline first, then stronger models such as gradient boosting.
- Pick a decision threshold that matches the business cost of false positives and false negatives.
- Deploy, log every prediction and compare it with the real outcome when it arrives.
Classification vs regression
Supervised problems come in two shapes. Classification predicts a category: fraud or legitimate, which of five product types, positive or negative sentiment. Regression predicts a number: delivery time, next week's sales, a property price. The distinction drives the metrics you use. Classification is judged with precision, recall, F1 score and ROC AUC, while regression uses mean absolute error or root mean squared error.
Common algorithms and examples
Algorithm choice matters less than most people expect. On tabular data, a well-tuned gradient-boosting model is a strong default, and a logistic regression baseline shows how much extra accuracy the added complexity is actually buying. For unstructured data such as photos or free text, fine-tuning a pretrained neural network almost always beats training one from scratch.
- Linear and logistic regression: fast, interpretable baselines.
- Decision trees, random forests and gradient boosting (XGBoost, LightGBM): strong on tabular business data.
- Support vector machines: effective on smaller datasets with many features.
- Neural networks: images, audio and text, usually fine-tuned from pretrained models.
- Real examples: spam filters, credit scoring, churn prediction, medical image triage and invoice field extraction.
Limitations of supervised learning
The biggest cost is labels. Someone has to mark thousands of examples correctly, and inconsistent labeling caps the accuracy any model can reach. Labels can also encode historical bias, for example past approval decisions that treated some groups unfairly, and the model will learn to repeat that pattern unless it is detected and corrected.
Models trained on one period may also degrade when customer behavior or market conditions change, so supervised systems need monitoring and periodic retraining. Nexzem often begins supervised projects with a written labeling guide and a small, double-checked dataset, which exposes ambiguous cases early and costs far less than relabeling a large dataset later.