Generalized Linear Models Notes

lecture-notes

The linear model is \[Y \sim \mathcal N(x\cdot \beta, \sigma^2), ~ \mathbb E(Y|x) = x\cdot \beta.\] In lectures we also saw logistic regression, which can be framed in the following form: \[Y \sim \text{Bernoulli}(p)\] \[\log\left(\frac{1-p}{p}\right) = x\cdot \beta\] In the logistic regression case, the expected value of \(Y\) given \(x\) is simply \(p\): \(\mathbb E(Y|x) = p\). Here \(p\in [0,1]\), it represents a yes/no probability.

In both of these cases, some function of the conditional outcome is a linear function of the predictor variables, i.e. \[g(\mathbb E(Y|x)) = x\cdot \beta\] for some function \(g\). We also make an assumption regarding the distribution of \((Y|x)\) is distributed: normal in one case, Bernoulli in the other. In both cases, these distributions are part of what are known as exponential families. These include many common named distributions, and the MLE of the parameters is particularly nice for these distributions.

Natural Exponential Families

Let \(h:\mathbb R\to [0,\infty)\). The natural exponential family of distributions associated with \(h\) is a collection of distributions parameterized by the real number \(\theta\). The probability density function of the distribution \(Y\) is \[f_Y(y;\theta) = \exp(y\theta - A(\theta)) h(y)\] where \(A(\theta)\) is chosen to make the integral with respect to \(y\) equal to \(1\): \[1 = \int^\infty_{-\infty} \exp(y\theta - A(\theta))h(y)~dy\] \[\exp(A(\theta)) = \int^\infty_{-\infty} \exp(y\theta)h(y)~dy\] \[A(\theta) = \log\int^\infty_{-\infty}\exp(y\theta)h(y)~dy\] Some terminology:

  • we call \(\theta\) the natural parameter.
  • we call \(h(y)\) the base density
  • we call \(A\) the cumulant function of log-partition function.

We’ve sacrificed some generality for the sake of clarity. More generally, an exponential family of distributions has the form \[f_Y(y;\theta) = \exp(T(y)\cdot \eta(\theta) - A(\theta))\cdot h(y)\] where \(\theta\in \mathbb R^s, \eta:\mathbb R^s\to \mathbb R^d\) and \(T:\mathbb R\to \mathbb R^d\).