TGViewer
Machine Learning Machine Learning @machinelearning9 ยท 41.7K subscribers
Post #5860 1.62K
Softmax vs Sigmoid โœ๏ธ Interact ๐Ÿ‘‰ https://byhand.ai/Khlg9b

= Softmax = ๐Ÿงฎ

Softmax is how deep networks turn raw scores into a probability distribution โ€” the final layer of every classifier ๐ŸŽฏ, and the core of every attention head in a transformer ๐Ÿค–. To see what it does, picture five boba tea shops ๐Ÿง‹ on the same block, all competing for your dollar ๐Ÿ’ฐ. Five candidates: a, b, c, d, e โ€” different chains, different brewing styles, different pearls. A boba reviewer hands you a ๐˜ค๐˜ฉ๐˜ฆ๐˜ช๐˜จ๐˜ฉ๐˜ฆ๐˜ด๐˜ต ๐˜ค๐˜ฐ๐˜ณ๐˜ฆ for each โ€” higher means perfectly chewy "QQ" pearls with the right bite ๐Ÿก (ask a Taiwanese friend to find out what QQ means). Negative scores are real: mushy bobas, overcooked pearls, a batch left sitting too long ๐Ÿฅ€.

How do you turn five chewiness scores into an allocation that adds to a whole dollar? You could spend everything at the chewiest shop, but that ignores how good the runners-up are ๐Ÿƒโ€โ™‚๏ธ. Softmax is the smooth alternative ๐ŸŒŠ.

Read the diagram left to right โžก๏ธ. First, raise each score to e^{x} โ€” this does two things: it turns negative chewiness into small positives, and it stretches the gaps between scores exponentially ๐Ÿ“ˆ. Then sum all five into a single total Z. Finally, divide each e^{x} by Z to get a probability. The five probabilities add up to one, so you can read them as percentages of your dollar ๐Ÿ“Š. The chewiest shop gets the biggest slice ๐Ÿฐ โ€” but never the whole dollar. That's the point of softmax: it ranks confidently while still leaving room for the others ๐Ÿค.

= Sigmoid = ๐Ÿ“‰

Sigmoid squashes any real number into a probability between 0 and 1 โ€” the classic activation for binary classification โœ…, and still the gating function inside LSTMs and GRUs. Same boba block as the previous Softmax example, narrowed to just two contenders โ€” a hot new shop a with chewiness score x, and your usual go-to b whose score is pinned at zero (the neutral baseline you've come to expect) ๐Ÿ“.

Sigmoid is just softmax with two players, one of them pinned to zero โš–๏ธ.

Read the diagram left to right โžก๏ธ. First, raise each score to e^{x} โ€” for the usual shop b whose score is zero, this is just e^0 = 1 (the constant baseline) ๐Ÿ›. Then sum the two into a total Z. Finally, divide each e^{x} by Z to get a probability. The two probabilities add up to one โ€” the new shop wins more of your dollar when its pearls get chewier, and your usual keeps the rest ๐Ÿ’ธ. That's the point of sigmoid: it turns a single chewiness score into a clean 0-to-1 chance you'll try the new place over your usual ๐Ÿš€.

https://t.me/DataScienceM ๐Ÿ”—
  • โค 1
More from @machinelearning9
  1. Oct 10, 2026Channel photo updated
  2. Oct 10, 2026Post #6301
  3. Oct 9, 2026Channel photo updated
  4. Oct 9, 2026Tracking experiments and versioning ML models with MLflow. ๐Ÿค– Machine learning developmentโ€ฆ
  5. Oct 8, 2026๐Ÿ”– Official PyTorch Tutorials This collection includes: ๐Ÿซก PyTorch fundamentals and torch.โ€ฆ
  6. Oct 7, 2026A distinctive bot for subscribing to courses on the most popular educational platforms acrโ€ฆ
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook โ†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 โ†’