Skip to content

Load Balancing Strategies: L4 vs L7, Round-Robin, Least Connections, Consistent Hashing

What it is

Load balancing distributes requests across backend instances and removes unhealthy targets from the eligible set. An L4 load balancer forwards transport connections using addresses and ports, while an L7 load balancer understands HTTP messages and can route by hostname, path, headers, or other application attributes.

How it works

A load balancer receives traffic, selects an eligible backend, forwards the request, and observes the result. Health checks determine eligibility; a connection failure or unhealthy response removes a target until it recovers. The scheduler then applies a distribution strategy. Round-robin advances through targets in order. Least connections favors the target with the fewest active connections, which often handles uneven connection lifetimes better. Consistent hashing maps a selected key to a position on a hash ring, so adding or removing a target remaps only keys near that target; virtual nodes reduce skew caused by an uneven physical ring. None of these strategies can create capacity, so load tests must cover both the scheduler and the backend’s bottleneck.

The component flow keeps health state separate from request selection:

    flowchart LR
    Client[Clients] --> Balancer[Load balancer]
    Probe[Health probes] --> Checker[Health checker]
    Checker -->|eligible target set| Balancer
    Balancer --> Orders[Orders instance A]
    Balancer --> Payments[Payments instance B]
    Balancer --> Web[Web instance C]
    Orders -->|unhealthy signal| Checker
    Payments -->|unhealthy signal| Checker
    Web -->|unhealthy signal| Checker
  

This Envoy configuration selects an L7 route and uses active health checks with least-request scheduling:

static_resources:
  listeners:
    - name: ingress
      address:
        socket_address:
          address: 0.0.0.0
          port_value: 8080
      filter_chains:
        - filters:
            - name: envoy.filters.network.http_connection_manager
              typed_config:
                "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
                stat_prefix: ingress_http
                route_config:
                  name: local_route
                  virtual_hosts:
                    - name: services
                      domains: ["*"]
                      routes:
                        - match:
                            prefix: "/api/"
                          route:
                            cluster: api_pool
                            timeout: 5s
  clusters:
    - name: api_pool
      type: STATIC
      lb_policy: LEAST_REQUEST
      load_assignment:
        cluster_name: api_pool
        endpoints:
          - lb_endpoints:
              - endpoint:
                  address:
                    socket_address:
                      address: 10.0.1.10
                      port_value: 8080
              - endpoint:
                  address:
                    socket_address:
                      address: 10.0.1.11
                      port_value: 8080
              - endpoint:
                  address:
                    socket_address:
                      address: 10.0.1.12
                      port_value: 8080
      health_checks:
        - timeout: 1s
          interval: 10s
          unhealthy_threshold: 2
          http_health_check:
            path: /health

L4 balancing preserves the transport connection and works for TCP or UDP without parsing the application payload. L7 balancing can terminate TLS, select a service from an HTTP path, and apply retries or request policies, but it consumes more CPU and memory. Sticky sessions preserve local session state through a cookie or affinity mechanism, yet they reduce distribution freedom and can send a client to a failed backend. Prefer externalizing session state unless affinity is required by a real backend constraint.

Tradeoffs

Match the balancer layer and scheduling policy to the workload:

ChoiceGainCost or failure behaviorUse when
L4Protocol independence and lower per-request processing costCannot inspect HTTP content; usually forwards connection-level failuresTCP or UDP services need high throughput or transparent proxying
L7Host and path routing, TLS termination, and per-request policyMore CPU and memory; policy bugs affect the entire ingress pathHTTP traffic needs content-aware routing or observability
Round-robinSimple and predictable for similarly sized targetsIgnores target load and request costBackends have similar capacity and similar request cost
Least connectionsAdapts to targets with long or short connectionsDepends on timely connection accounting and may favor stale observationsConnection durations vary substantially
Weighted round-robinLets operators represent known capacity differencesWeights become stale as capacity changesA small set of targets has different provisioned limits
Consistent hashingMinimizes key remapping when membership changesSkew remains possible when key distribution or virtual-node placement is poorCache or sharded work benefits from key affinity
Session affinityKeeps stateful clients on one backend without a shared storeReduces balancing freedom and complicates failoverA backend genuinely requires local session state

When to use

  • Traffic exceeds one backend’s safe capacity or you need to remove instances without interrupting clients.
  • You need health checks and failover across availability zones or regions.
  • HTTP requests must route by host, path, method, or header.
  • You operate a cloud network load balancer for connection-level TCP or UDP traffic.
  • Cache keys or partition ownership benefit from stable placement across membership changes.

Alternatives

  • DNS round-robin — adds no proxy hop, but clients and resolvers cache answers and the DNS layer gives weak backend-health signals.
  • Client-side load balancing — removes one network hop, but requires every client to implement discovery, selection, and failover correctly.
  • Service discovery without a proxy — keeps service-to-service paths direct, but central policy and consistent traffic observation become application responsibilities.

Related