= Softmax = ๐งฎ
Softmax is how deep networks turn raw scores into a probability distribution โ the final layer of every classifier ๐ฏ, and the core of every attention head in a transformer ๐ค. To see what it does, picture five boba tea shops ๐ง on the same block, all competing for your dollar ๐ฐ. Five candidates: a, b, c, d, e โ different chains, different brewing styles, different pearls. A boba reviewer hands you a ๐ค๐ฉ๐ฆ๐ช๐จ๐ฉ๐ฆ๐ด๐ต ๐ค๐ฐ๐ณ๐ฆ for each โ higher means perfectly chewy "QQ" pearls with the right bite ๐ก (ask a Taiwanese friend to find out what QQ means). Negative scores are real: mushy bobas, overcooked pearls, a batch left sitting too long ๐ฅ.
How do you turn five chewiness scores into an allocation that adds to a whole dollar? You could spend everything at the chewiest shop, but that ignores how good the runners-up are ๐โโ๏ธ. Softmax is the smooth alternative ๐.
Read the diagram left to right โก๏ธ. First, raise each score to e^{x} โ this does two things: it turns negative chewiness into small positives, and it stretches the gaps between scores exponentially ๐. Then sum all five into a single total Z. Finally, divide each e^{x} by Z to get a probability. The five probabilities add up to one, so you can read them as percentages of your dollar ๐. The chewiest shop gets the biggest slice ๐ฐ โ but never the whole dollar. That's the point of softmax: it ranks confidently while still leaving room for the others ๐ค.
= Sigmoid = ๐
Sigmoid squashes any real number into a probability between 0 and 1 โ the classic activation for binary classification โ , and still the gating function inside LSTMs and GRUs. Same boba block as the previous Softmax example, narrowed to just two contenders โ a hot new shop
a with chewiness score x, and your usual go-to b whose score is pinned at zero (the neutral baseline you've come to expect) ๐.Sigmoid is just softmax with two players, one of them pinned to zero โ๏ธ.
Read the diagram left to right โก๏ธ. First, raise each score to e^{x} โ for the usual shop
b whose score is zero, this is just e^0 = 1 (the constant baseline) ๐. Then sum the two into a total Z. Finally, divide each e^{x} by Z to get a probability. The two probabilities add up to one โ the new shop wins more of your dollar when its pearls get chewier, and your usual keeps the rest ๐ธ. That's the point of sigmoid: it turns a single chewiness score into a clean 0-to-1 chance you'll try the new place over your usual ๐.https://t.me/DataScienceM ๐