CloudFront KeyValueStore: quotas and what they imply
When you propose terminating redirects or routing decisions at the edge — as in Bitly — a good interviewer asks what the edge store can actually hold. Knowing the real numbers turns a hand-wave into a design.
1. Hard quotas
| Specification | Limit | What it constrains |
|---|---|---|
| Total store size | 5 MB per KeyValueStore | The hard ceiling on how many entries fit |
| Maximum key size | 512 bytes (UTF-8) | Ample for short codes, hashes, path stems |
| Maximum value size | 1 KB | Long target URLs must fit, or bypass the store |
| Stores per account | 50 (soft limit, adjustable) | Multi-tenancy and vertical sharding strategy |
| Functions reading one store | Up to 10 | Staging/prod separation; sharing between hooks |
| Stores per function | 1 | A function cannot query two stores — namespace within one |
| Required runtime | cloudfront-js-2.0 or later | Earlier runtimes have no KVS client |
2. Capacity arithmetic
Because the ceiling is total size, capacity depends entirely on your record shape.
Short-link case (realistic):
key: 7 bytes "aZ89xK1"
value: 60-80 bytes "https://blog.example.com/posts/how-to-scale-dynamodb"
overhead: ~15 bytes internal metadata
≈ 100 bytes/record -> 5 MB / 100 B ≈ 45,000-55,000 links
Worst case (values near the 1 KB ceiling):
≈ 1,040 bytes/record -> ≈ 4,800 records
Compact case (feature flags, A/B buckets, blocklists):
key 8 B + value 4 B + overhead ≈ 30 bytes -> ≈ 170,000 entries
So the honest headline for a shortener is roughly 50,000 hot links, which is nothing against a corpus of billions — and that is precisely why the design treats it as an L0 hot-set shield rather than a database.
3. Latency and propagation
Read latency: sub-millisecond. The store's data is held in memory at the PoP alongside the V8 isolate running your function, so a read is an in-process memory lookup with no network hop. This is the property that makes it qualitatively different from calling DynamoDB from Lambda@Edge.
Write propagation: seconds. A PutKey through the AWS SDK or CLI replicates to all PoPs globally, typically within 2–10 seconds.
aws cloudfront-keyvaluestore put-key \
--kvs-arn "arn:aws:cloudfront::123456789012:key-value-store/my-store" \
--key "aZ89xK1" \
--value "https://example.com/destination" \
--if-match "<etag>"
The --if-match ETag gives you optimistic concurrency on the store as a whole, which matters if multiple writers update it.
Consistency: eventual, globally. Two PoPs may briefly disagree. That is fine for redirects, routing rules, A/B bucketing, and blocklists. It is not fine for anything requiring a linearizable read-modify-write — never use it as a lock, a counter, or a uniqueness claim.
4. What edge functions cannot do
The constraint that shapes the architecture more than any quota: CloudFront Functions have no outbound network access. No fetch, no TCP, no database client, no Kafka producer. They can read the request, read the KVS, and return a request or a response.
Consequences:
- Analytics cannot be emitted from the function. Use CloudFront Real-Time Logs instead, which stream request metadata out-of-band into Kinesis within seconds, adding zero latency. This is the correct shape anyway — counting must not sit in the redirect's critical path.
- A miss cannot be resolved in the function. Return the unmodified
requestobject to forward to the origin. Returning a response object terminates at the edge; returning the request continues the pipeline. That single distinction expresses hit and miss.
Lambda@Edge is the alternative that can make network calls, but it runs in a container and adds roughly 5–15ms of invocation overhead — which defeats the purpose when your entire budget is 10ms.
5. How this shapes the design
Given ~50,000 slots and eventual consistency, the store is an L0 hot-set shield in a tiered read path:
Hot (top ~50k codes) -> edge KVS, resolved at the PoP in <1 ms
Warming -> CloudFront's standard HTTP cache (origin returns
Cache-Control: public, max-age=300 on a miss)
Cold (the long tail) -> origin: ElastiCache -> DynamoDB, ~10-30 ms
Three operating rules fall out of the quotas:
- Promote on observed traffic, not on creation. A background worker consuming the click stream writes a code into KVS once it exceeds a threshold (say ten clicks per minute), and evicts the coldest entry to stay under 5 MB. The store must be actively managed; it will not evict for you.
- Skip oversized values. A target URL above 1 KB is never promoted — the write would fail. Let those links resolve at the origin permanently.
- Shard across stores if you outgrow one. With a limit of one store per function, sharding means separate functions on separate cache behaviours (by path prefix), or namespacing multiple logical datasets into a single store with key prefixes.
6. What to say in an interview
"CloudFront KVS caps at 5 MB total, with 512-byte keys and 1 KB values, which for a shortener works out to roughly 45,000–55,000 links. So it isn't the database — it's an L0 shield for the hot set, sized by the 80/20 rule, and a background worker promotes codes into it based on click velocity while evicting the coldest.
Reads are sub-millisecond because the data sits in memory at the PoP next to the V8 isolate — no network hop, unlike Lambda@Edge which adds 5–15ms of container overhead. Writes propagate globally in a few seconds and are eventually consistent, which is fine for redirects and routing but disqualifies it for anything needing a linearizable read-modify-write.
The design constraint that matters most is that edge functions have no outbound network access at all — so analytics go through CloudFront Real-Time Logs into Kinesis out-of-band, and a KVS miss is expressed by returning the unmodified request object so CloudFront forwards it to the origin."