The softmax function converts an N-dimensional real vector into a probability distribution: each output is positive, sums to one, and reflects the relative magnitude of inputs (a "soft" max). It is commonly written as exponentials normalized by their sum and is used to produce class probabilities from logits in multiclass classification. The piece derives the partial derivative of each softmax output with respect to each input, yielding the Jacobian matrix: ∂S_i/∂a_j = S_i(δ_{ij} - S_j), which specializes to S_i(1−S_i) on the diagonal and −S_i S_j off-diagonal. That compact form (using the Kronecker delta) makes it straightforward to reason about interactions between outputs and to propagate derivatives in further calculations.
Practical issues and composition with a linear layer are also addressed. Direct exponentiation can overflow or underflow in floating point, so subtracting the maximum input before exponentiating stabilizes computation and avoids NaNs, producing an implementation-friendly "stable softmax." For a softmax layer that follows a matrix multiply (logits = W x), the derivative needed for training is obtained by the multivariate chain rule: the Jacobian of the softmax evaluated at the logits multiplied by the Jacobian of the linear map with respect to W. Working out these Jacobians with correct indexing yields the gradients used to update the weight matrix in gradient-based optimization.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.