➜ LR is primarily about the penalty L1 (lasso) or L2 (ridge)
➜ Naive Bayes is alpha
➜ decision tree is almost never used as a separate algorithm, but you still need to understand how it works
➜ random forest is primarily aboutmax_depth, the number of estimators,max_features(you can't take all the features),min_samples_splitandmin_samples_leaf
➜ GBT is usually aboutxgboost / catboost/ lightgbm, where you look at all the same things as above, pluslearning_rate,alpha / lambda, the number of leaves, subsample/ colsample_bytreeand boosting type, if it's applicable
➜ PCA is better not to touch for time series, unless you're doing the rolling variant or using it for research purposes. But PLS is perfectly fine.
Types of PCA and when to use them:
-> linear, if linear dependencies between features are assumed
-> kernel, if the dependencies between features are non-linear
-> incremental, if you have a lot of features and samples and need to quickly run PCA
-> robust PCA, if there are outliers in the data
➜ If we're already talking about PCA, we can mention ICA, when you need statistically independent features, not just uncorrelated ones
➜ kNN is sometimes used; k-means is useful where it's obvious that the main thing is the number of clusters
➜ support vector machine is when nothing else has worked and you're just curious if this might take off. It relies on C and kernel, which are responsible for linear or non-linear dependencies
➜ NN hyperparameters are a whole separate story, because they depend on the type of network. But basically remember the combination:NN layer -> normalization layer -> dropout layer. Sometimes between normalization and dropout, or even later, an activation layer is placed. This depends on whether you need flexibility in choosing the place for activation, or you just remove it as a separate layer and set activation directly in the NN layer parameters.
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
