A LINEAR MODEL FOR PROBABILITIES
Logistic regression
Work in progress — this demo is still being developed and checked.
The same kind of boundary, with a probability on each side. Explore how confidence, loss, and learning fit together.
Enter a point with the keyboard
From a score to a probability
p = 1 / (1 + exp(−z))
On the boundary, z = 0 and p = 0.5. Predict 1 only when p > 0.5; ties predict 0, following the lecture.
How expensive is a wrong probability?
Expected loss and the true probability η(x)
The green curve is η(−log p) + (1−η)(−log(1−p)). Its minimum is at p = η. Here η is a hypothetical true probability, not the learning rate α.
From likelihood to empirical risk
Likelihood = ∏ᵢ pᵢʸⁱ(1−pᵢ)¹⁻ʸⁱ
Log likelihood = Σᵢ [yᵢ log pᵢ + (1−yᵢ) log(1−pᵢ)]
Mean loss = −log likelihood / n
Maximizing likelihood and minimizing mean cross-entropy give the same objective. All logs are natural; loss is measured in nats.
One gradient-descent step
Each step uses all points and the same pre-update parameters. For x̃ = (x₁, x₂, 1), each contribution is (p−y)x̃. Average the contributions, then subtract α times that gradient.
Every point’s contribution to the last step
| Point | y | p | p−y | ∂ℓ/∂w₁ | ∂ℓ/∂w₂ | ∂ℓ/∂b |
|---|
A confident correct point contributes little; a confident incorrect point contributes much more. Correctly classified points can still change the weights.
A slice through the loss
Why every direction has nonnegative curvature
H = (1/n) Σᵢ pᵢ(1−pᵢ)x̃ᵢx̃ᵢᵀ
vᵀHv = (1/n) Σᵢ pᵢ(1−pᵢ)(v·x̃ᵢ)² ≥ 0
The full objective is convex. A zero gradient implies a global minimum, when attained; strict convexity and a finite minimizer are not guaranteed. For separable data, unregularized weights can grow while loss approaches zero.