Home › Posts › AWS API Gateway throttling vs WAF rate-based rules
AWS

AWS API Gateway throttling vs WAF rate-based rules

API Gateway throttling vs WAF rate limiting: if API Gateway already throttles requests, why would we add rate limiting in AWS WAF too? It looks like the same…

Oct 7, 2026 • 8 min read • maki96milosavljevic@gmail.com

API Gateway throttling vs WAF rate limiting: if API Gateway already throttles requests, why would we add rate limiting in AWS WAF too? It looks like the same feature configured twice. You can use both since they actually protect against different problems.

In short:

  • API Gateway throttling -> protects the overall capacity of your API and backend
  • WAF rate-based rules -> stop a single source (IP, user, API key) that sends too many requests

Let’s go through two simple scenarios, see how each control works under the hood, and look at the Terraform code we use for both.

Two scenarios
Scenario 1: One source sends too many requests

Imagine a scraper, a broken retry loop in a client app, or a credential-stuffing script that hammers your /login endpoint. That single source sends thousands of requests, but overall API traffic is still well below the API Gateway throttling limits. API Gateway sees nothing unusual, so throttling never kicks in.

This is where a WAF rate-based rule helps. A rule aggregated by IP counts requests per source address and blocks that one source once it crosses the threshold. By default WAF returns 403 Forbidden, and the request never reaches your API or backend.

Scenario 2: Many users send a reasonable number of requests

Now imagine a marketing campaign or a peak hour. Thousands of users each send a normal number of requests. No single IP comes close to the WAF threshold, so WAF lets everything through. But the combined traffic is more than your Lambda functions, database or downstream services can handle.

This is where API Gateway throttling helps. It limits the total request rate at the configured scope (account, stage, method, or usage plan) and returns 429 Too Many Requests for anything above that limit. Clients should treat 429 as a signal to retry with backoff.

How API Gateway throttling works

API Gateway uses the token bucket algorithm. Think of it as a bucket of tokens where every request uses one token:

  • Rate -> how fast the bucket refills (steady-state requests per second)
  • Burst -> how big the bucket is, so a short spike can pass

For example, with rate 100 and burst 500, a full bucket lets 500 requests through at once. After that, the API accepts about 100 requests per second as the bucket refills. Everything above that gets a 429.

Throttling can be configured on several levels:

  • Account level -> the default per region is 10,000 RPS with a burst of 5,000, shared by all APIs in the account
  • Stage level -> limits for one stage (dev, qa, prod)
  • Method level -> limits for a specific resource and method, for example POST /login
  • Usage plans -> limits per client, identified by an API key

The account limit is the one people usually forget. Because it is shared, one busy API can push the others into 429 even if they have little traffic of their own. That is why we set our own limits per stage.

This is the part of our API Gateway Terraform module that does it:

resource "aws_api_gateway_method_settings" "this" {
  rest_api_id = aws_api_gateway_rest_api.this.id
  stage_name  = aws_api_gateway_stage.this.stage_name
  method_path = "*/*"

  settings {
    metrics_enabled        = true
    logging_level          = var.logging_level
    throttling_burst_limit = var.throttling_burst_limit
    throttling_rate_limit  = var.throttling_rate_limit
  }
}
variable "throttling_burst_limit" {
  type        = number
  description = "API Gateway throttling burst limit"
  default     = 500
}

variable "throttling_rate_limit" {
  type        = number
  description = "API Gateway steady-state rate limit"
  default     = 100
}

method_path = "*/*" applies the limits to every method in the stage. If one endpoint needs a different limit, add another aws_api_gateway_method_settings resource with a specific path such as login/POST. Keep in mind that stage and method limits can’t go above the account limit.

How WAF rate-based rules work

WAF doesn’t use a token bucket. A rate-based rule counts requests over a time window and groups them by an aggregation key. When one group goes over the limit, WAF applies the rule action (usually Block) to that group only, while everyone else keeps working.

The main settings are:

  • Limit -> the maximum number of requests per group within the window
  • Evaluation window -> 60, 120, 300 or 600 seconds (300 is the default)
  • Aggregation key -> how requests are grouped
  • Scope-down statement -> optional filter, so the rule only counts certain requests, for example only /login

IP is the most common aggregation key, but not the only one. You can also group by forwarded IP (X-Forwarded-For), or use custom keys such as a header, cookie, query parameter, URI path, HTTP method, and combine them up. You can also count all requests together without grouping.

Here is an example web ACL with two rules: a general limit per IP, and a stricter one for the login endpoint:

resource "aws_wafv2_web_acl" "api" {
  name  = "${var.env}-api-waf"
  scope = "REGIONAL"

  default_action {
    allow {}
  }

  rule {
    name     = "rate-limit-per-ip"
    priority = 1

    action {
      block {}
    }

    statement {
      rate_based_statement {
        limit                 = var.rate_limit_per_ip # requests per window
        evaluation_window_sec = 300
        aggregate_key_type    = "IP"
      }
    }

    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "rate-limit-per-ip"
      sampled_requests_enabled   = true
    }
  }

  rule {
    name     = "rate-limit-login"
    priority = 2

    action {
      block {}
    }

    statement {
      rate_based_statement {
        limit                 = var.rate_limit_login
        evaluation_window_sec = 60
        aggregate_key_type    = "IP"

        scope_down_statement {
          byte_match_statement {
            search_string         = "/login"
            positional_constraint = "CONTAINS"

            field_to_match {
              uri_path {}
            }

            text_transformation {
              priority = 0
              type     = "LOWERCASE"
            }
          }
        }
      }
    }

    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "rate-limit-login"
      sampled_requests_enabled   = true
    }
  }

  visibility_config {
    cloudwatch_metrics_enabled = true
    metric_name                = "${var.env}-api-waf"
    sampled_requests_enabled   = true
  }
}

To attach the web ACL to the API, we use the stage_arn output from our API Gateway module:

resource "aws_wafv2_web_acl_association" "api" {
  resource_arn = module.api_gateway.stage_arn
  web_acl_arn  = aws_wafv2_web_acl.api.arn
}
Conclusion

When comparing API Gateway throttling vs WAF, remember that WAF stops individual abusers and API Gateway protects overall capacity.

  • One IP doesn’t always mean one user. In an office or behind a corporate proxy, many people can share the same public IP. Keep that in mind when you set the limit, or group requests by API key or header instead.
  • If you have CloudFront or some other proxy in front of the API, WAF will see the proxy’s IP and not the real client IP. In that case use the forwarded IP option, but only if you trust whoever sets the X-Forwarded-For header.
  • Don’t expect exact limits. Both work on a best-effort basis, and WAF needs a few seconds to start blocking, so some extra requests will get through.
  • Throttling won’t protect your backend on its own. For that you still need things like queues, concurrency limits or Lambda reserved concurrency.
  • Don’t just pick round numbers. Check CloudWatch to see how much traffic you really get and how much your backend can handle.

Author

maki96milosavljevic@gmail.com

Practical notes about AWS, Terraform, DevOps, automation, and building systems that are easier to operate.