TGViewer
Channel Public Channel
TechLead Bits

TechLead Bits

@techleadbits

About software development with common sense.
Thoughts, tips and useful resources on technical leadership, architecture and engineering practices.

Author: @nelia_loginova
Subscribers
517
Photos
73
Videos
0
Links
203

Showing posts older than #10 · Back to latest

Older Posts 9 shown
Post #9 248
Load Balancing

📍Round Robin. Simple, stable, well-suited for servers with identical replicas.
📍Weighted Round Robin. Adds a `weight for each replica, weight` usually correlates with the server capacity. Improves resource utilization in heterogeneous infrastructure.
📍Least connections\Load. Directs network traffic to the server with the fewest active connections\load. It can be really effective with long-lived sessions or tasks.
📍Weighted Least Connection. It’s the previous one enriched with capacity weights.
📍Hash Ring. Each host is mapped into a circle using its hashed address, each request is routed to a host by hashing some property of the request. So the balancer finds the nearest host clockwise to match requests with the server. It can work well if there is a good request attribute to hash.
📍Random Subsetting. Each client randomly shuffles the list of hosts and fills its subset by selecting available backends from the list. In real cases the load may be distributed unevenly.
📍Deterministic Subsetting. Google improvement for Random Subsetting: adds client assignment rounds, server random shuffles and allows servers to adjust weights back to clients. It has greater stability and evenness in distribution.
📍Random Choice of 2. The algorithms suggest picking up 2 servers randomly and selecting one with the least load. Simple, quite effective in many cases.

Let’s go back to that last post and figure out why Netflix wasn't satisfied and had to cook up another load balancing algorithm.
The main issue is that stateful services are asymmetric (postgres leader and followers, zookeeper leader and followers, cassandra topology with quorums, etc.), and it does matter which host to connect to. Plus, the traffic between availability zones was substantial, leading to additional latency.
So what actually was done (on Cassandra client balancing sample):
- Random Choice of 2 is extended to Random Choice of 8
- Weight for nodes is added and based on rack topology, replicas info and health state
- Selected nodes are sorted according to the weight

The improvement reduces latency by up to 40% for Netflix scenarios, which I believe is a significant achievement. Unfortunately I couldn't find the original video detailing the implementation, but there's a presentation with measurements available.

#architecture #reliability #network
  • 👍 3
  • 👀 3
Post #8 261
Reliable Stateful Systems at Netflix

There is quite an interesting long-read article explaining how Netflix implements the reliability of its stateful services.

To understand what reliability actually means the author proposes answering three fundamental questions:
- How often does the system fail?
- When it fails, how large is the blast radius?
- How long does it take to recover from an outage?
Ideally, systems don’t fail, have minimal impact on failure and recover very quickly😀. That is the simple part.

So how is that achieved? That is actually the hard part:
📍Single tenancy. There are no multi-tenant data stores. It minimizes blast impact.
📍Capacity management. Netflix programs special workload capacity models that generate specifications for a cluster according to the system requirements and SLOs.
📍Data replication. Data is replicated in 12 availability zones across 4 regions.
📍Overprovisioning. In case of one region degradation, the traffic is spread across other regions so each region must keep an extra 33% capacity reserved for failover.
📍Snapshot restoration. To replace one instance with another - load snapshot from s3 and then apply delta.
📍Performance monitoring. Continuous monitoring is essential to detect failures quickly, remediate them and recover.
📍Cache in front of services. The idea is to use caching for complex business logic rather than underlying data. Service with business logic is quite expensive to operate, so approach shifts load to the cache.
📍Reliable clients. Quite complex approach to inform clients what timeouts and level of service that they can rely on. Better to read the original.
📍Load Balancing. Netflix uses improved choice-of-2 algorithms with weight requests. It takes into account availability zones, replicas, health state.
📍Stateful APIs. All requests are idempotent with built-in work pagination. That approach requires additional idempotent tokens or global transaction ids.

#architecture #reliability #usecase
InfoQ How Netflix Ensures Highly-Reliable Online Stateful Systems Building reliable stateful services at scale isn’t a matter of building reliability into the servers, the clients, or the APIs in isolation. By combining smart and meaningful choices for each of these three components, we can build massively scalable, SLO…
  • ❤ 2
Post #7 247
Death By Meeting

I finally read "Death by Meeting: A Leadership Fable...About Solving the Most Painful Problem in Business" by P. Lencioni after it had been on my reading list for quite some time. And let me tell you my impressions.

Meetings – the bane of many of our professional lives. Who enjoys them, really? I certainly don't. So, naturally, I expect to find some insights on how to get fewer meetings. Surprisingly, the book advocates for the opposite approach - meetings are mandatory!
But not the soul-crushing, unproductive gatherings we've all come to dread. No, the book suggests a complete rebuild of our approach to meetings, transforming them into dynamic, engaging, and yes, even enjoyable experiences.

Think of it this way: movies captivate us for hours on end, holding our attention without fail. You know what makes movies so exciting? Conflict! So why not add some of that spice into our meetings? Inject some healthy debates, encourage interaction, and yeah, even a bit of arguing. Let's make decisions right then and there, turning our meetings into the place where things actually get done.

The examples and stories in the book are built around general business management, but some ideas can be adapted to the context of IT.

#booknook #softskills #management #meetings
  • ❤ 1
  • ✍ 1
Post #6 285
Scale it automagically!

Currently, there's considerable interest in automatic autoscaling, largely driven by the cost of resource usage. Let's verify how it works with dynamic load inside Kubernetes. There exists a naive belief that the HPA will deliver equivalent performance to pre-configured replicas when demand arises.

Consider a scenario where we set CPU utilization target at 80%. During peak events such as Black Friday or intense invoicing periods, service demands surge, causing CPU utilization growth to 100%.
The HPA operates on a logic where desired replicas are calculated using the formula:

desiredReplicas = ceil[currentReplicas * (currentMetricValue / desiredMetricValue)]


For instance, if we have 5 replicas and the CPU utilization grows to 100% out of the desired 80%, the calculation will result in 6 replicas as a target. Consequently, the deployment scales up to 6 replicas. However, this scaling process involves a delay as the HPA controller waits for all pods to be ready, which can be time-consuming depending on the size of the service. Additionally, there's a stabilization window with a default of 5 minutes.
After the initial scaling, it's highly likely that the adjustment won't be sufficient (let’s say you need around 50 replicas to handle the load). Therefore, subsequent iterations of scaling will occur at intervals of at least 5 minutes until the desired 80% utilization is achieved or until the maximum allowed number of replicas is reached. This delay might not meet the business requirements.

In summary, autoscaling requires time to adapt to the load. While it functions effectively in scenarios where the load changes gradually, it's less suitable for handling scheduled, heavy loads. In such cases, it's advisable to explore alternative options and proactively prepare the environment for the anticipated load.

#architecture #performance #scalability
Kubernetes Horizontal Pod Autoscaling In Kubernetes, a HorizontalPodAutoscaler automatically updates a workload resource (such as a Deployment or StatefulSet), with the aim of automatically scaling capacity to match demand. Horizontal scaling means that the response to increased load is to deploy…
Post #5 312
Code Reviews Like Humans

There's already a ton of stuff out there about code reviews, covering what they're for and how to make them better. But I want to recommend one more Better Code Reviews FTW!

One of the points that I really like here is that a code review is basically just giving feedback. And we're all pretty good at giving feedback in other areas. So why not apply that same constructive approach to code reviews? Just imagine that there is something like `Fix typo in successful`. It can be read as `Hey, you made a mistake, but I still think you’re smart! or `You make a stupid mistake, dumbass`. Impression is quite different, right? It's all about being supportive and positive with each other.

Oh, and there's a cool article that dives into the same topic if you're interested: https://mtlynch.io/human-code-reviews-1/ Check it out!

#engineering #codereview
YouTube Better Code Reviews FTW! - Tess Ferrandez-Norlander - NDC London 2024 This talk was recorded at NDC London in London, England. #ndclondon #ndcconferences #developer #softwaredeveloper Attend the next NDC conference near you: https://ndcconferences.com https://ndclondon.com/ Subscribe to our YouTube channel and learn…
  • ❤ 2
Post #4
TechLead Bits pinned «Welcome to the TechLead Bits! Navigation: #architecture - everything about software architecture, #systemdesign, #patterns, cross-cutting concerns and trade-offs #engineering - engineering practices in software development: #codereview, #ci, #refactoring…»
Post #3 377
Welcome to the TechLead Bits!

Navigation:

#architecture - everything about software architecture, #systemdesign, #patterns, cross-cutting concerns and trade-offs

#engineering - engineering practices in software development: #codereview, #ci, #refactoring, #testing, #documentation

#usecase - implementation references from real companies

#softskills - soft skills recommendations and topics: #management, #leadership, #communications, #productivity, #creativity

#booknook - books overview

#news - notable news from the industry

#offtop - some other materials and news that may not directly relate to the technical leadership

👋 Hi, I’m Nelia!

🔹 15+ years in IT
🔹 10+ years in team management
🔹 Now head of cloud platform division: we cook and deliver cloud platform services like Kubernetes, SQL/NoSQL databases, Kafka, data streaming, service mesh, backup/restore procedures and more.

My focus is on non-functional requirements — things like consistency, availability, latency, security, observability, scalability, and more.

💡 Here I'm writing about:
- Technical leadership
- Architecture and engineering practices
- How to scale systems, processes — and yourself 😉
  • ❤ 2
Post #2
Channel photo updated
Post #1
Channel created
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →