Skip to main content
FA
Faiz Akram
HomeAboutExpertiseProjectsBlogContact
FA
Faiz Akram

Senior Technical Architect specializing in enterprise-grade solutions, cloud architecture, and modern development practices.

Quick Links

Privacy PolicyTerms of ServiceBlog

Connect

© 2026 Faiz Akram. All rights reserved.

Back to Blog
Designing Production-Ready Distributed Rate Limiting Systems
System Design

Designing Production-Ready Distributed Rate Limiting Systems

F
Faiz Akram
October 7, 2026
8 min read

Modern APIs and cloud-native platforms are under constant threat from both legitimate and abusive traffic surges. Designing a distributed rate limiting system is critical to protect backend resources, prevent abuse, and deliver consistent user experience—especially as microservices and edge deployments multiply.

What Is Distributed Rate Limiting and Why Does It Matter?

Distributed rate limiting is the practice of enforcing API or service usage quotas across multiple nodes or network boundaries, rather than on a single server. In contrast to local rate limiting, which only protects resources on a per-node basis, distributed rate limiting ensures fairness and abuse prevention globally—no matter how many frontend nodes, load balancers, or edge proxies you deploy.

Example: Envoy with Redis as Distributed Rate Limiter

Here's an actual snippet from an Envoy proxy configuration (v1.26+) using Redis as the rate limit backend:

typed_extension_protocol_options:
  envoy.extensions.filters.network.redis_proxy.v3.RedisProxy:
    '@type': type.googleapis.com/envoy.extensions.filters.network.redis_proxy.v3.RedisProxy
    stat_prefix: egress_redis
    settings:
      op_timeout: 5s
    prefix_routes:
      routes:
        - prefix: "*"
          request_mirror_policy:
            cluster: rate_limit_cluster
rate_limit_service:
  grpc_service:
    envoy_grpc:
      cluster_name: rate_limit_cluster
  transport_api_version: V3

This configuration connects Envoy's rate limiting filter to a Redis-backed rate limiting service, enabling global quota management regardless of where the API is accessed.

Key insight: Distributed rate limiting is essential for any API or service that must guarantee fair access and abuse resistance at scale, especially in horizontally scaled or multi-region deployments.

Step 1: Choosing the Right Rate Limiting Algorithm for Scale

Why Algorithms Matter

The choice of rate limiting algorithm directly impacts accuracy, resource efficiency, and user experience. The most common algorithms are:

  • Fixed Window: Simple, but allows bursts at window boundaries—leads to 'thundering herd' issues.
  • Sliding Window: Smoother rate enforcement, but more complex to implement and compute.
  • Token Bucket: Allows controlled bursts, flexible refill rates—widely used for APIs.
  • Leaky Bucket: Smooths traffic by draining at a constant rate; can enforce strict pacing.

Best Practices for API Rate Limiting

In production, I recommend Token Bucket for most API workloads. It strikes a balance between flexibility and fairness. With Redis, you can use the CL.THROTTLE command (Redis 6.0+, via the Redis Bloom module) to implement efficient distributed token bucket logic:

CL.THROTTLE myapi:userid123 10 20 60 1
# 10 req/sec, burst 20, period 60s, cost 1/request

This command enforces a rate of 10 requests per second with a burst capacity of 20, over a 60-second rolling window, for the user with ID 123.

Key insight: Token Bucket is the gold standard for most real-world APIs—delivering smooth traffic, predictable limits, and burst tolerance with reasonable implementation complexity.

Step 2: Architecting a Distributed Rate Limiting Layer

Where to Enforce Limits: Edge, Service Mesh, or Backend?

You can enforce rate limits at several architectural layers:

  • API Gateway/Edge Proxy (e.g., NGINX, Envoy, Kong): Scales horizontally and shields backend from abusive clients.
  • Service Mesh (e.g., Istio, Linkerd): Enables service-to-service rate limiting, crucial for internal APIs.
  • Backend Application: Last line of defense, but expensive if abuse reaches this level.

Reference Architecture Example

A typical production architecture looks like this:

[Client] -> [Cloud Load Balancer] -> [NGINX/Envoy API Gateway] --(gRPC/HTTP)--> [Redis Cluster]
                                                               |
                                                        [Application Pods]
  • Rate limiting filter runs in the API Gateway.
  • Gateways communicate with a Redis Cluster (preferably Redis 6.2+, Cluster mode) to check and increment rate quotas.

Latency and Failure Considerations

  • Use Redis pipelining to batch checks for high-throughput environments.
  • Configure circuit breakers in Envoy/NGINX to avoid total API failure if Redis is unavailable.

Key insight: Enforce rate limiting as close to the edge as possible, but always plan for graceful degradation if your distributed rate limit store (e.g. Redis) is slow or unavailable.

Step 3: Configuring Redis for High-Performance Distributed Rate Limiting

Redis Deployment Patterns

For global rate limiting, Redis must be:

  • Highly available: Use Redis Cluster or Redis Sentinel for failover.
  • Low latency: Deploy as close to your API Gateways as possible (same AZ/region).
  • Horizontally scalable: Use Redis Cluster with hash tags to minimize cross-slot operations.

Example: Redis Cluster Helm Chart (Kubernetes)

# values.yaml for Bitnami/Redis Cluster Helm Chart
cluster:
  enabled: true
  replicas: 6
resources:
  requests:
    cpu: 500m
    memory: 1Gi
  limits:
    cpu: 2
    memory: 4Gi
networkPolicy:
  enabled: true
  • 6 Redis nodes for high throughput (10,000+ requests/sec tested with Envoy and CL.THROTTLE).
  • Resource limits prevent noisy neighbor issues in shared clusters.

Redis Security and Observability

  • Enable TLS (tls.enabled: true) for encrypted traffic.
  • Use Redis AUTH for strong authentication.
  • Expose Redis metrics via Redis Exporter to Prometheus/Grafana for real-time quota and error tracking.

Key insight: Production-grade Redis clusters for rate limiting must be both secure and observable—never expose Redis unauthenticated, and always monitor for quota exhaustion and latency spikes.

Step 4: Implementing Rate Limiting in API Gateways and Proxies

Envoy Proxy Configuration

With Envoy (v1.26+), the rate limiting filter communicates with a centralized gRPC rate limit service (commonly lyft/ratelimit).

http_filters:
- name: envoy.filters.http.ratelimit
  typed_config:
    "@type": type.googleapis.com/envoy.extensions.filters.http.ratelimit.v3.RateLimit
    domain: my_api
    failure_mode_deny: false
    rate_limit_service:
      grpc_service:
        envoy_grpc:
          cluster_name: rate_limit_cluster
      transport_api_version: V3
  • Set failure_mode_deny: false to avoid blocking all traffic if the rate limit service fails.

NGINX with Redis

With NGINX (v1.21+, plus lua-resty-redis), you can implement custom Redis-backed rate limiting using OpenResty Lua scripts.

http {
  lua_shared_dict ratelimit 10m;
  server {
    location /api/ {
      access_by_lua_file /etc/nginx/lua/ratelimit.lua;
      proxy_pass http://api_backend;
    }
  }
}

In /etc/nginx/lua/ratelimit.lua, connect to Redis and implement the token bucket logic per user/IP.

Service Mesh Patterns

  • Istio (v1.18+): Use the Envoy Rate Limit filter as an EnvoyFilter CRD, referencing a central rate limit service.
  • Linkerd: Requires third-party or sidecar-based custom integration.

Key insight: API gateways and proxies are the most effective enforcement point for distributed rate limiting—ensure your configuration supports graceful degradation on backend errors.

Step 5: Observability, Testing, and Failure Handling

Real-Time Monitoring

  • Export rate limit metrics: total requests, quota exhausted, backend errors
    • Envoy exposes these at /stats (e.g., ratelimit.ok, ratelimit.over_limit)
    • NGINX can push logs or metrics via Lua or nginx-prometheus-exporter
  • Set up Grafana dashboards for API rate limiting health

Load Testing for Rate Limits

  • Use k6 or Locust to simulate high cardinality access patterns
  • Test for race conditions and quota consistency under high concurrency (10,000+ RPS)
  • Validate failure handling: what happens if Redis is slow/unavailable?

Circuit Breaker and Fallback Logic

  • Envoy: Set failure_mode_deny: false to allow requests if backend is down (fail open), or true to block (fail closed)—choose based on API risk profile
  • NGINX: Fallback to local in-memory rate limits if Redis is unreachable

Key insight: Observability and robust fallback logic are non-negotiable for distributed rate limiting—monitor exhaustions, errors, and always rehearse failure scenarios in staging.

Comparison Table: Distributed Rate Limiting Tools and Approaches

ApproachThroughputConsistencyOperational ComplexityCloud-Native?Notes
Redis (Token Bucket)High (10k+)GoodMediumYesWell-tested, needs HA & monitoring
Envoy + Lyft/ratelimitHighGoodMediumYesgRPC-based, easy mesh integration
NGINX + Lua + RedisHighGoodHighPartialLua code required; less out-of-box
API Gateway SaaS (e.g., AWS API Gateway, Kong Cloud)Med-HighGoodLowYesManaged, but less flexible
Service Mesh (Istio)Med-HighGoodHighYesBest for internal/service-to-service
Local-only (per node)HighPoorLowYesNo cross-node fairness

Key insight: Redis-backed rate limiting (via Envoy or NGINX) offers the best mix of performance and global fairness for most cloud-native workloads; SaaS gateways trade flexibility for ease of use.

Frequently Asked Questions

Q: What is the difference between local and distributed rate limiting? A: Local rate limiting enforces quotas on a single server or pod, while distributed rate limiting uses a central store (like Redis) to enforce quotas across all nodes, ensuring global fairness and preventing circumvention by scaling horizontally.

Q: How do I prevent my rate limiting layer from becoming a single point of failure? A: Deploy your rate limit store (e.g., Redis) in a highly-available, clustered configuration, and configure your gateways/proxies for graceful degradation (fail open or fail closed) in case the backend becomes unavailable.

Q: Can I use cloud-native managed services for distributed rate limiting? A: Yes, services like AWS API Gateway, Kong Cloud, and Azure API Management provide built-in distributed rate limiting, but for maximum flexibility and control, self-managed solutions (Envoy + Redis, NGINX + Lua + Redis) are recommended at high scale.

Key Takeaways

  • Always enforce rate limits at the edge for best protection and fairness.
  • Use Token Bucket algorithms with Redis Cluster (6.0+) for reliable, high-throughput distributed rate limiting.
  • Configure your API gateways/proxies with fallback logic to handle backend failures gracefully.
  • Monitor rate limit usage, exhaustions, and latency with real-time metrics—never deploy blind.
  • Load test quota enforcement under realistic, high-concurrency scenarios to prevent race conditions and logic errors.
  • Choose between managed and self-hosted rate limiting solutions based on required flexibility, scale, and operational overhead.

Tags

clouddistributed systemsapi securityrate limitingredisenvoy

Share this article

Found it helpful? Share it with your network.

X / TwitterLinkedInFacebookWhatsApp

Related Articles

More on System Design and related topics

Designing Efficient Distributed Cache Invalidation: Patterns, Tools, and Production Tactics
System Design
September 23, 2026
8 min read

Designing Efficient Distributed Cache Invalidation: Patterns, Tools, and Production Tactics

Learn how to architect reliable distributed cache invalidation for microservices and cloud-native systems using Redis, NATS, and production patterns.

distributed systemscache invalidationcloud
Read More
Designing Production-Ready API Throttling Systems: Patterns, Tools, and Best Practices
System Design
September 15, 2026
7 min read

Designing Production-Ready API Throttling Systems: Patterns, Tools, and Best Practices

Learn how to architect production-grade API throttling systems for scale, fairness, and resilience with real-world patterns, open-source tools, and cloud services.

apisystem designrate limiting
Read More
Designing Multi-Tenant SaaS Platforms: Patterns, Isolation, and Scaling Tactics
System Design
September 7, 2026
8 min read

Designing Multi-Tenant SaaS Platforms: Patterns, Isolation, and Scaling Tactics

Learn how to architect production-grade multi-tenant SaaS platforms with strong isolation, cost efficiency, and scalable onboarding—real configs included.

cloudmulti-tenancysaas architecture
Read More