Two kinds of numbers control how a model behaves, and they are set in very different ways.
Parameters
Parameters are learned automatically during training. The weights of a neural network and the coefficients of a linear regression are parameters. You never set them by hand; the optimiser finds them from data.
Hyperparameters
Hyperparameters are settings chosen before training that control how learning happens. Examples:
- the learning rate and batch size of a neural network;
- the number of trees and maximum depth in a random forest;
- the regularisation strength of a linear model;
- the number of clusters in k-means.
Tuning Hyperparameters
Good hyperparameters are found by trying options and comparing results on a validation set (never the test set). Common approaches:
- Grid search: try every combination from a list. Simple but expensive.
- Random search: sample combinations at random. Often finds good settings faster.
- Bayesian optimisation: uses past results to choose promising settings to try next.
Practical Advice
- Start with library defaults; they are usually sensible.
- Tune the few hyperparameters that matter most (for gradient boosting: learning rate, number of trees, depth).
- Use cross-validation for small datasets.
- Keep a record of every run so results can be reproduced.
Parameter Counts
When people describe a language model as having "7 billion parameters", they mean learned weights. More parameters allow more capacity, but also need more data, memory and compute.