Skip to main content

CloudFront KeyValueStore: quotas and what they imply

When you propose terminating redirects or routing decisions at the edge — as in Bitly — a good interviewer asks what the edge store can actually hold. Knowing the real numbers turns a hand-wave into a design.

1. Hard quotas​

SpecificationLimitWhat it constrains
Total store size5 MB per KeyValueStoreThe hard ceiling on how many entries fit
Maximum key size512 bytes (UTF-8)Ample for short codes, hashes, path stems
Maximum value size1 KBLong target URLs must fit, or bypass the store
Stores per account50 (soft limit, adjustable)Multi-tenancy and vertical sharding strategy
Functions reading one storeUp to 10Staging/prod separation; sharing between hooks
Stores per function1A function cannot query two stores — namespace within one
Required runtimecloudfront-js-2.0 or laterEarlier runtimes have no KVS client

2. Capacity arithmetic​

Because the ceiling is total size, capacity depends entirely on your record shape.

Short-link case (realistic):
key: 7 bytes "aZ89xK1"
value: 60-80 bytes "https://blog.example.com/posts/how-to-scale-dynamodb"
overhead: ~15 bytes internal metadata
≈ 100 bytes/record -> 5 MB / 100 B ≈ 45,000-55,000 links

Worst case (values near the 1 KB ceiling):
≈ 1,040 bytes/record -> ≈ 4,800 records

Compact case (feature flags, A/B buckets, blocklists):
key 8 B + value 4 B + overhead ≈ 30 bytes -> ≈ 170,000 entries

So the honest headline for a shortener is roughly 50,000 hot links, which is nothing against a corpus of billions — and that is precisely why the design treats it as an L0 hot-set shield rather than a database.

3. Latency and propagation​

Read latency: sub-millisecond. The store's data is held in memory at the PoP alongside the V8 isolate running your function, so a read is an in-process memory lookup with no network hop. This is the property that makes it qualitatively different from calling DynamoDB from Lambda@Edge.

Write propagation: seconds. A PutKey through the AWS SDK or CLI replicates to all PoPs globally, typically within 2–10 seconds.

aws cloudfront-keyvaluestore put-key \
--kvs-arn "arn:aws:cloudfront::123456789012:key-value-store/my-store" \
--key "aZ89xK1" \
--value "https://example.com/destination" \
--if-match "<etag>"

The --if-match ETag gives you optimistic concurrency on the store as a whole, which matters if multiple writers update it.

Consistency: eventual, globally. Two PoPs may briefly disagree. That is fine for redirects, routing rules, A/B bucketing, and blocklists. It is not fine for anything requiring a linearizable read-modify-write — never use it as a lock, a counter, or a uniqueness claim.

4. What edge functions cannot do​

The constraint that shapes the architecture more than any quota: CloudFront Functions have no outbound network access. No fetch, no TCP, no database client, no Kafka producer. They can read the request, read the KVS, and return a request or a response.

Consequences:

  • Analytics cannot be emitted from the function. Use CloudFront Real-Time Logs instead, which stream request metadata out-of-band into Kinesis within seconds, adding zero latency. This is the correct shape anyway — counting must not sit in the redirect's critical path.
  • A miss cannot be resolved in the function. Return the unmodified request object to forward to the origin. Returning a response object terminates at the edge; returning the request continues the pipeline. That single distinction expresses hit and miss.

Lambda@Edge is the alternative that can make network calls, but it runs in a container and adds roughly 5–15ms of invocation overhead — which defeats the purpose when your entire budget is 10ms.

5. How this shapes the design​

Given ~50,000 slots and eventual consistency, the store is an L0 hot-set shield in a tiered read path:

Hot (top ~50k codes) -> edge KVS, resolved at the PoP in <1 ms
Warming -> CloudFront's standard HTTP cache (origin returns
Cache-Control: public, max-age=300 on a miss)
Cold (the long tail) -> origin: ElastiCache -> DynamoDB, ~10-30 ms

Three operating rules fall out of the quotas:

  1. Promote on observed traffic, not on creation. A background worker consuming the click stream writes a code into KVS once it exceeds a threshold (say ten clicks per minute), and evicts the coldest entry to stay under 5 MB. The store must be actively managed; it will not evict for you.
  2. Skip oversized values. A target URL above 1 KB is never promoted — the write would fail. Let those links resolve at the origin permanently.
  3. Shard across stores if you outgrow one. With a limit of one store per function, sharding means separate functions on separate cache behaviours (by path prefix), or namespacing multiple logical datasets into a single store with key prefixes.

6. What to say in an interview​

"CloudFront KVS caps at 5 MB total, with 512-byte keys and 1 KB values, which for a shortener works out to roughly 45,000–55,000 links. So it isn't the database — it's an L0 shield for the hot set, sized by the 80/20 rule, and a background worker promotes codes into it based on click velocity while evicting the coldest.

Reads are sub-millisecond because the data sits in memory at the PoP next to the V8 isolate — no network hop, unlike Lambda@Edge which adds 5–15ms of container overhead. Writes propagate globally in a few seconds and are eventually consistent, which is fine for redirects and routing but disqualifies it for anything needing a linearizable read-modify-write.

The design constraint that matters most is that edge functions have no outbound network access at all — so analytics go through CloudFront Real-Time Logs into Kinesis out-of-band, and a KVS miss is expressed by returning the unmodified request object so CloudFront forwards it to the origin."