Первой ответила от нас @kaleturina, затем и я добавил.
Оставляю тут текст своего коммента и Марины, чтобы потом посмотреть учли ли наши пожелания.
Мой коммент:
> I'd like to propose developing and describing a model of SLI/SLO maturity. As I see in real practice many companies can only support the very first level - define some SLOs and just watching them, because they still do not understand why they need it. The second level - do SLO based decisions based on error budget levels, if EB exhausted - dig and try to fix it. At this level I also see a lack of motivation at the teams level, so companies handle it by including a service SLO in team's KPI (I think this is a bad idea leading to "workarounds" to fit the SLIs to the KPI target). The third - a rapid reaction to SLIs drops by Multi Window Multi Burn Rate alerts, any fast burn rate alert declared as an incident (this helps with a better reaction), teams trying to fix issues before the budget exhausted, no slow burn alerts analytics, no feature freezes. The fourth is the third plus slow burn alerts analytics and decisions, analytics accepted by a company, error budget policy with feature freezers and an unfreeze tax.
I've never seen the 4th level ) Especially a real feature freezing practice.
And another one. Do we need to compensate an Error budget that was spent by service A due to failure in the underlying service B?
What if the B service is our Kubernetrs cluster or our network and all our services spent their budgets because of this falure? Should we do a feature freeze for all affected services?
Коммент Марины:
There are several topics I would personally love to see covered. It won't fit in one message, so I will split it :)
1. On-call compensation packages
Many companies expect engineers to work overtime, including handling incidents outside regular working hours, without offering any additional compensation. This approach not only leads to burnout among engineers, but also creates unrealistic expectations in upper management regarding operational costs. Both of these factors negatively impact long-term system reliability.
2. Improving the "burn rate" metric
The current concept of a "burn rate of 4" is unclear to many engineers and therefore rarely used in practice. It might be more helpful to reframe this concept—perhaps by alerting teams when they are projected to exhaust their error budget within a certain timeframe. Making the metric more intuitive could significantly improve its adoption and usefulness.
#sre #slo #allslo