Regularization Data Science

lecture-notes

The idea. Recall our supervised-learning-broad-framework. In regularization, we modify our loss function \(\ell(\theta)\) to penalize large parameters, so we make our new loss function \[\ell(\theta) + \alpha \text{Size}(\theta).\] This parameter \(\alpha\) is not something being learned or modified.It’s an adjustable constant which we call a hyperparameter. Here are two ways to choose \(\text{Size}\):

  • Ridge regression: \(\text{Size}(\theta) = \|\theta\|_2^2 = (\theta_1 + ... + \theta_p)^2\)
  • Lasso regression: \(\text{Size}(\theta) = \|\theta\|_1 = |\theta_1| + ... + |\theta_p|.\)

It hopefully increases the interpretability of the model and prevents \(\theta\) from exploding.