The simple linear regression model
In simple linear regression we have a single variable we would like to predict, \(y\), and a single feature \(x\) (meaning both \(x,y\in \mathbb R\)). The form of \(f\) in the supervised learning framework is \[y = f(x) + \epsilon = \beta_0 + \beta_1x + \epsilon\] where \(\beta_0,\beta_1\in\mathbb R\) are the parameters we’re trying to learn, and we assume \(\epsilon\in \mathcal N(0,\sigma^2)\) is some normally distributed error term which is independent of \(x\).
We wish to find the values of \(\beta_0,\beta_1\) which minimize the Mean Squared Error \[\text{MSE}(\beta) = \frac1n\sum_{i=1}^n (y_i - f_\beta(x_i))^2\] which in our particular case is \[\text{MSE}(\beta) = \frac1n\sum^n_{i=1}(y_i - \beta_0 - \beta_1x_i)^2.\] There is a closed form solution for the \(\beta\) which minimizes \(\text{MSE}(\beta)\) in terms of \(\{(x_i,y_i)\}_{i=1}^n\), you can find it either by calculus minimization or by projecting \(\vec y\) onto the subspace spanned by \(\vec 1\) and \(\vec x\): \[\hat\beta_0 = \overline y - \hat \beta_1\overline x\] where \(\overline x\) and \(\overline y\) are the averages of \(x\) and \(y\) respectively, and \[\hat\beta_1 = \frac{\sum_{i=1}^n(x_i - \overline x)(y_i - \overline y)}{\sum_{i=1}^n(x_i - \overline x)^2} = \frac{\text{cov}(x,y)}{\text{var}(x)}.\]
The multivariable linear regression model
Now we have \(\vec x_i = (x_1,...,x_d)\) features for every observation \(y_i\in \mathbb R\). We now wish to model \(y\) as \[y = f_\beta(x) + \epsilon = \beta_0 + \beta_1x_1 + ... + \beta_dx_d + \epsilon = \beta\cdot \vec x + \epsilon\] where \(\vec x = (1,x_1,...,x_d)\) and \(\epsilon \sim \mathcal N(0,\sigma^2)\).
If we have \(n\) observations, we package them together in a \(n\times (1 + d)\) matrix \(X\):
\begin{equation*} X = \begin{bmatrix} 1 & x_{11} & x_{12} & ... & x_{1p}\\ 1 & x_{21} & x_{22} & ... & x_{2p}\\ & & & \vdots & \\ 1 & x_{n1} & x_{n2} & ... & x_{np}\\ \end{bmatrix} \hphantom{asdfas} \vec{y} = \begin{bmatrix} y_1\\y_2\\ \vdots \\ y_n \end{bmatrix} \end{equation*}and then our MSE looks like this: