TGViewer
Channel Public Channel
DevOps & SRE notes

DevOps & SRE notes

@devops_sre_notes

Helpful articles and tools for DevOps&SRE
Subscribers
13.3K
Photos
54
Videos
0
Links
2.6K

Showing posts older than #2492 · Back to latest

Older Posts 20 shown
Post #2490 1.67K
In this story from the Betterstack newsletter, learn how Dropbox managed to save millions of dollars by optimizing its object storage architecture. The piece delves into the technical decisions and engineering efforts behind their impressive cost-reduction initiative.
https://newsletter.betterstack.com/p/how-dropbox-saved-millions-of-dollars
Betterstack How Dropbox Saved Millions of Dollars by Building a Load Balancer Dropbox saved resources by creating a superior version of a tool everyone uses
  • 🔥 1
Post #2489 1.76K
This technical report from Datadog offers a deep dive into managing storage for etcd, the key-value store at the heart of Kubernetes. It explains the causes of database growth and provides strategies for monitoring, defragmenting, and purging old data to maintain a healthy cluster.
https://www.datadoghq.com/blog/managing-etcd-storage/
Datadog How to support a growing Kubernetes cluster with a small etcd | Datadog Discover essential strategies for efficiently managing etcd storage in your Kubernetes clusters.
  • 👍 1
Post #2486 1.86K
This tutorial offers an interesting approach to container image distribution by using S3 as a private container registry. The author demonstrates how to set up and use an S3 bucket for storing and pulling images, providing a simple alternative to dedicated registry services.
https://ochagavia.nl/blog/using-s3-as-a-container-registry/
Adolfo Ochagavía Using S3 as a container registry For the last four months I’ve been developing a custom container image builder, collaborating with Outerbounds1. The technical details of the builder itself might be the topic of a future article, but there’s something surprising I wanted to share already:…
  • 👍 1
Post #2482 1.77K
This write-up from Prezi Engineering explains how multi-AZ deployments can lead to surprisingly high data transfer costs. It documents their journey of migrating from a costly self-hosted Prometheus setup to a more efficient monitoring solution to save on their cloud budget.
https://engineering.prezi.com/how-using-availability-zones-can-eat-up-your-budget-our-journey-from-prometheus-to-be8a816f7efe
Medium How using Availability Zones can eat up your budget — our journey from Prometheus to… Intro
  • 👍 1
Post #2481 1.72K
The Grab Engineering team shares their experience in executing a seamless database migration with zero downtime. This blogpost details the meticulous planning, tooling, and validation steps required to achieve a successful migration for a critical, high-traffic service.
https://engineering.grab.com/seamless-migration
Grab Tech How we seamlessly migrated high volume real-time streaming traffic from one service to another with zero data loss and duplication In the world of high-volume data processing, migrating services without disruption is a formidable challenge. At Grab, we recently undertook this task by spl...
  • ❤ 3
Post #2478 1.73K
This edition of the Scalable Thread newsletter breaks down effective strategies for handling sudden and unexpected bursts of traffic to your systems. It explores architectural patterns and techniques to ensure reliability and prevent service degradation during traffic spikes.
https://newsletter.scalablethread.com/p/how-to-handle-sudden-bursts-of-traffic
Scalablethread How to Handle Sudden Bursts of Traffic or "Thundering Herd Problem"? Techniques to Avoid Potential Failures Caused by Sudden Traffic Spikes
  • 👍 3
Post #2477 1.72K
This article from JP Gouin provides a deep dive into implementing GitOps at scale, with a specific focus on the cluster bootstrapping process. It covers the challenges and solutions for managing numerous Kubernetes clusters efficiently and declaratively.
https://medium.com/@jp-gouin/gitops-at-scale-clusters-bootstrapping-f36695d4340d
Medium GitOps at scale — Clusters bootstrapping Explore one approach to help infrastructure team managing their multiple environments, variants and all required applications
  • ❤ 2
Post #2474 1.83K
Elliot Graebert proposes an impact-based leveling system for engineering organizations as an alternative to traditional career ladders. This treatise discusses how focusing on impact can foster a more motivated and effective engineering culture.
https://medium.com/@elliotgraebert/an-impact-based-level-system-for-engineering-organizations-2e0f9bee20e6
Medium An impact-based level system for engineering organizations Defining L1-L6 for individual contributors and leads
  • 👍 2
  • ❤ 1
Post #2473 1.66K
Marc Christian P. Gregorio offers a practical commentary on automating centralized NAT Gateways in AWS across multiple VPCs and regions using Terraform. The solution aims to optimize costs and simplify network management for large-scale deployments.
https://medium.com/@marcchristianp.gregorio/automating-centralized-nat-gateways-in-aws-vpcs-and-region-with-terraform-69a6f90d60da
Medium Automating Centralized NAT Gateways in AWS VPCs and Region with Terraform When managing a large-scale AWS environment with multiple accounts, deploying multiple NAT gateways across various VPCs can become very…
  • 👍 3
  • ❤ 1
Post #2472 1.63K
AWS just released their postmortem (link in comment) for the October DynamoDB outage. It's thorough, technically detailed, and explains exactly what broke and how they'll "prevent" it from happening again. But this PR-approved, sanitized narrative tells us only what happened to the technology, nothing else.

https://aws.amazon.com/message/101925/
  • ❤ 2
  • 👍 2
Post #2471 1.8K
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →