The regression model is trained to predict the target variable. We know the input features and how to transform them, but what about the loss function? There are at least two loss functions:
* The normalised root mean square error (NRMSE), to be minimised, and
* Decomposition root sum square distance (DECOMP.RSSD), an ingenious thing described as the “Business fit”.
There are a series of tweets on how awesome this is. Assume you have spent 80% of the budget on promotion in my Telegram channel, but the model suggests that 80% of conversion came from Instagram. Obviously, you better trust my channel, not the model.
But aren’t we building the model to understand what we are doing wrong? That is correct, but there are two assumptions:
* We assume that the budget was distributed not completely at random. The current allocation is reasonable before the start and iteratively moves towards an optional solution. This is similar to the Expectation-maximisation algorithm.
* If we approach the stakeholders and propose changing the $100M budget allocation, they can be surprised and not appreciate it. Instead, a proposal to alter the budget within 3–5% usually does not raise questions.
Having at least two loss functions means obtaining an infinite number of models. What do we do with that? The Nevergrad framework powers the optimisation process with a large number (thousands to tens of thousands) of iterations and displays the results on a two-dimensional plot (one dimension per loss function).
#ArticleReview
Post #35
1.64K