Activation functions decide whether a neuron should "fire" and introduce non-linearity into the model — crucial for learning complex patterns.
1️⃣ Why We Need Activation Functions
Without them, neural networks are just linear regressors.
They help networks learn curves, edges, and non-linear boundaries.
2️⃣ Common Activation Functions
a) ReLU (Rectified Linear Unit)
f(x) = max(0, x) ✔️ Fast
✔️ Prevents vanishing gradients
❌ Can "die" (output 0 for all inputs if weights go bad)
b) Sigmoid
f(x) = 1 / (1 + exp(-x)) ✔️ Good for binary output
❌ Causes vanishing gradient
❌ Not zero-centered
c) Tanh (Hyperbolic Tangent)
f(x) = (exp(x) - exp(-x)) / (exp(x) + exp(-x)) ✔️ Outputs between -1 and 1
✔️ Zero-centered
❌ Still suffers vanishing gradient
d) Leaky ReLU
f(x) = x if x > 0 else 0.01 * x ✔️ Fixes dying ReLU issue
✔️ Allows small gradient for negative inputs
e) Softmax
Used in final layer for multi-class classification
✔️ Converts outputs into probability distribution
✔️ Sum of outputs = 1
3️⃣ Where to Use What?
• ReLU → Hidden layers (default choice)
• Sigmoid → Output layer for binary classification
• Tanh → Hidden layers (sometimes better than sigmoid)
• Softmax → Final layer for multi-class problems
🧪 Try This:
Build a model with:
• ReLU in hidden layers
• Softmax in output
• Use it for classifying handwritten digits (MNIST)
💬 Tap ❤️ for more!