TGViewer
Channel Public Channel
DevOps & SRE notes

DevOps & SRE notes

@devops_sre_notes

Helpful articles and tools for DevOps&SRE
Subscribers
13.3K
Photos
53
Videos
0
Links
2.6K

Showing posts older than #2557 · Back to latest

Older Posts 20 shown
Post #2555 2.01K
DevOps & SRE notes How did you start your morning? Cloudflare decided that you’d had too much of the internet.
A change made to how Cloudflare's Web Application Firewall parses requests caused Cloudflare's network to be unavailable for several minutes this morning. This was not an attack; the change was deployed by our team to help mitigate the industry-wide vulnerability disclosed this week in React Server Components. We will share more information as we have it today.

https://www.cloudflarestatus.com/incidents/lfrm31y6sw9q
Cloudflare Status Cloudflare Service Issues This incident has been resolved. A change made to how Cloudflare's Web Application Firewall parses requests caused Cloudflare's network to be unavailable for…
  • 👍 4
Post #2554 1.86K
How did you start your morning? Cloudflare decided that you’d had too much of the internet.
  • 🤣 12
Post #2553 1.92K
This walkthrough from Minimal DevOps demonstrates how to implement predictive autoscaling for Kubernetes workloads. It leverages KEDA to act on forecasts generated by Prophet, allowing scaling actions to anticipate demand rather than just reacting to it.
https://minimaldevops.com/predictive-autoscaling-in-kubernetes-with-keda-and-prophet-cbccd96cf881
Medium Predictive Autoscaling in Kubernetes with Keda and Prophet Choosing proactive scaling over reactive scaling in Kubernetes
  • 👍 5
Post #2552 1.83K
This report on HackerNoon offers a detailed look at implementing graceful shutdowns for Go applications running in Kubernetes. It explains how to handle termination signals to prevent data loss and ensure service stability during updates or scaling events.
https://hackernoon.com/mastering-graceful-shutdowns-in-go-a-comprehensive-guide-for-kubernetes
Hackernoon Mastering Graceful Shutdowns in Go: A Comprehensive Guide for Kubernetes For applications deployed in orchestrated environments (e.g., Kubernetes), graceful handling of termination signals is crucial.
Post #2546 1.94K
In this piece from Slack Engineering, the authors advocate for intentionally breaking systems to improve resilience. They share a real-world incident that led them to adopt "strategic chaos" as a way to test and strengthen their recovery processes.
https://slack.engineering/break-stuff-on-purpose/
slack.engineering Break Stuff on Purpose Incidents are stressful but inevitable. Even services designed for availability will eventually encounter a failure. Engineers naturally find it daunting to defend their systems against the “infinite number of ways” things can go wrong.
  • 🔥 1
Post #2544 2K
Post #2543 2.12K
This treatise by Usama Malik explains how to create reusable infrastructure components using Terraform modules. The author highlights the benefits of modularity, such as improved efficiency, maintainability, and collaboration.
https://aws.plainenglish.io/how-to-create-reusable-infrastructure-with-terraform-modules-b4bbcf4c0ad1
Medium How to Create Reusable Infrastructure with Terraform Modules In this article, we’ll explore the benefits of Terraform modules and guide you on how to create them for building scalable and reusable…
  • 👍 4
Post #2541 2.03K
Post #2540 1.94K
Karsten Schnitter's review on the OpenSearch blog explores how to visualize metrics ingested with OpenTelemetry using OpenSearch Dashboards. The author provides examples of creating insightful visualizations for monitoring Kubernetes container metrics.
https://opensearch.org/blog/opentelemetry-metrics-visualization/
OpenSearch Using OpenSearch to visualize metrics ingested with OpenTelemetry OpenSearch is growing into a full observability platform with support for logs, metrics, and traces. Using Data Prepper, you can ingest these three signal types in OpenTelemetry (OTel) format. Logs are well supported...
  • 👍 3
Post #2539 1.9K
This piece offers a detailed look into the system architecture that powers Netflix's streaming service. It covers the company's cloud-native approach, its use of microservices, and its sophisticated content delivery network (CDN).
https://www.clickittech.com/software-development/netflix-architecture/
ClickIT Netflix Architecture | A Look Into Its 2026 System Architecture Netflix Architecture system, ensures uninterrupted streaming for a global user base and faultless content delivery across devices. Learn more
  • 👍 2
Post #2537 1.95K
Post #2536 1.9K
FluxCD UI - Coming soon

The Flux Status Page is a lightweight, mobile-friendly web interface providing real-time visibility into your GitOps pipelines. Embedded directly within the Flux Operator, it requires no additional installation steps.

Designed for DevOps engineers and platform teams, the Status Page offers direct insight into your Kubernetes clusters. It allows you to track app deployments, monitor controller readiness, and troubleshoot issues instantly, without needing to access the CLI.

Built with security in mind, the interface is strictly read-only, ensuring it never interferes with Flux controllers or compromises cluster security. Together with the Flux MCP Server, it provides a comprehensive solution for on-call monitoring and Agentic AI incident response in production environments.

https://github.com/controlplaneio-fluxcd/flux-operator/pull/488
GitHub Introduce Flux Status Page web UI by stefanprodan · Pull Request #488 · controlplaneio-fluxcd/flux-operator The Flux Status Page is a lightweight, mobile-friendly web interface providing real-time visibility into your GitOps pipelines. Embedded directly within the Flux Operator, it requires no additional...
  • 🔥 6
  • ❤ 1
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →