TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 579 subscribers
Post #2074 201
Reminder:

➜ LR is primarily about the penalty L1 (lasso) or L2 (ridge)

➜ Naive Bayes is alpha

➜ decision tree is almost never used as a separate algorithm, but you still need to understand how it works

➜ random forest is primarily about max_depth, the number of estimators, max_features (you can't take all the features), min_samples_split and min_samples_leaf

➜ GBT is usually about xgboost / catboost / lightgbm, where you look at all the same things as above, plus learning_rate, alpha / lambda, the number of leaves, subsample / colsample_bytree and boosting type, if it's applicable

➜ PCA is better not to touch for time series, unless you're doing the rolling variant or using it for research purposes. But PLS is perfectly fine.

Types of PCA and when to use them:
-> linear, if linear dependencies between features are assumed
-> kernel, if the dependencies between features are non-linear
-> incremental, if you have a lot of features and samples and need to quickly run PCA
-> robust PCA, if there are outliers in the data

➜ If we're already talking about PCA, we can mention ICA, when you need statistically independent features, not just uncorrelated ones

➜ kNN is sometimes used; k-means is useful where it's obvious that the main thing is the number of clusters

➜ support vector machine is when nothing else has worked and you're just curious if this might take off. It relies on C and kernel, which are responsible for linear or non-linear dependencies

➜ NN hyperparameters are a whole separate story, because they depend on the type of network. But basically remember the combination: NN layer -> normalization layer -> dropout layer. Sometimes between normalization and dropout, or even later, an activation layer is placed. This depends on whether you need flexibility in choosing the place for activation, or you just remove it as a separate layer and set activation directly in the NN layer parameters.


••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
More from @dataxplore
  1. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  2. Sep 14, 2026Post #2188
  3. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  4. Aug 22, 2026Post #2185
  5. Aug 21, 2026Deep systemic analysis of AI constraints from context to internal weight editing. 📂 PDF #…
  6. Aug 17, 2026Adaptive Gradient Thresholding Why Fixed Gradient Clipping Kills Deep RecSys When Feedback…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →