Logistic Regression Notes

Just as linear regression is the most well understood continuous modeling technique, logistic regression is the most well understood classification model.

Basic idea: one variable \(x\) (say, apple tree height), two categories \(A\) and \(B\) (say, whether or not tree can bear fruit). We want to model the probability that a given sample falls in \(A\) or \(B\) given observation \(x\) (probability that tree can bear fruit given its height).

Model Assumptions

  • The log odds depend linearly on the regressors.
  • The regressors are measured without error.
  • The observations are independent.
  • The data are not perfectly linearly seperable (i.e. there cannot be some vector \(\vec{v}\) with \(\vec{x} \cdot \vec{v} < 0\) for one class and \(>0\) for the other class).
    • If they are, we will have \(\vert \beta \vert \to \infty\) as we train the model. The resulting \(p_\beta\) will approach the indicator function of the set \(\{\vec{x}: \vec{x} \cdot \vec{v} > 0\}\) as we train the model. This isn’t horrible, but is probably better to use something like a support vector machine in this case.
  • We do not have multicollinearity of the regressors.
    • Just as in linear regression, multicolinearity will not impact our ability to accurately predict probabilities. It only prevents us from understanding feature importance.