TGViewer
🚀🐳 Летит Кит: SRE и не только 🚀🐳 Летит Кит: SRE и не только @letitkit · 202 subscribers
Post #67 162
3 месяца назад Neil Murphy, один из авторов SRE Book опубликовал анонс, что они пишут 2 редакцию! И спрашивал, а чего бы вы там хотели видеть?! 👁

Первой ответила от нас @kaleturina, затем и я добавил.

Оставляю тут текст своего коммента и Марины, чтобы потом посмотреть учли ли наши пожелания.

Мой коммент:

> I'd like to propose developing and describing a model of SLI/SLO maturity. As I see in real practice many companies can only support the very first level - define some SLOs and just watching them, because they still do not understand why they need it. The second level - do SLO based decisions based on error budget levels, if EB exhausted - dig and try to fix it. At this level I also see a lack of motivation at the teams level, so companies handle it by including a service SLO in team's KPI (I think this is a bad idea leading to "workarounds" to fit the SLIs to the KPI target). The third - a rapid reaction to SLIs drops by Multi Window Multi Burn Rate alerts, any fast burn rate alert declared as an incident (this helps with a better reaction), teams trying to fix issues before the budget exhausted, no slow burn alerts analytics, no feature freezes. The fourth is the third plus slow burn alerts analytics and decisions, analytics accepted by a company, error budget policy with feature freezers and an unfreeze tax.
I've never seen the 4th level ) Especially a real feature freezing practice.
And another one. Do we need to compensate an Error budget that was spent by service A due to failure in the underlying service B?
What if the B service is our Kubernetrs cluster or our network and all our services spent their budgets because of this falure? Should we do a feature freeze for all affected services?

Коммент Марины:
There are several topics I would personally love to see covered. It won't fit in one message, so I will split it :)

1. On-call compensation packages
Many companies expect engineers to work overtime, including handling incidents outside regular working hours, without offering any additional compensation. This approach not only leads to burnout among engineers, but also creates unrealistic expectations in upper management regarding operational costs. Both of these factors negatively impact long-term system reliability.

2. Improving the "burn rate" metric
The current concept of a "burn rate of 4" is unclear to many engineers and therefore rarely used in practice. It might be more helpful to reframe this concept—perhaps by alerting teams when they are projected to exhaust their error budget within a certain timeframe. Making the metric more intuitive could significantly improve its adoption and usefulness.

#sre #slo #allslo
  • 🔥 4
More from @letitkit
  1. Sep 24, 2026Раздача слонов на Claude AI Привет, киты 🐋! 1. Получить 100$ или $250 бонусных на Сlaude…
  2. Sep 22, 2026Споры про искусственный интеллект в SRE изменились. Мы уже не гадаем, заменит ли агент деж…
  3. Sep 15, 2026Toil в работе SRE. Измерять, или "мы и так знаем что работы дофига“? Привет, киты 🐳. Пора…
  4. Sep 8, 2026О важности отдыха Привет, киты! 🐳 Оказалось я забыл как отдыхать. Перед отпуском делал за…
  5. Sep 3, 2026Важное про безопасность с чужим кодом - еще раз. Привет, киты 🐳! TL; DR; на чужих AI-скил…
  6. Aug 26, 2026Реклама конференции в которой я выступаю. Если вы думали, пойти или нет, а хотели пойти с…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →