Skip to main content

Price Tracking Service (CamelCamelCamel / Honey)

Sharpened prompt. Design a service tracking prices for 500M products across 10,000 retailers, notifying 50M users within minutes of a price drop they care about, without scraping every product every hour (which is neither affordable nor permitted), while surviving aggressive anti-bot defences and extracting a correct price from ten thousand different page layouts.

The defining constraint is economic: you cannot poll everything, so the entire design is about deciding what to look at. Everything else — extraction, alerting, storage — is downstream of that scheduling decision.

1. Problem framing​

Functional requirements​

  • Track price and availability history for a product URL.
  • Let users set alerts (absolute threshold, percentage drop, "any drop", back-in-stock).
  • Notify matching users promptly when a price changes.
  • Show price history charts and "is this actually a good deal?" context.

Non-functional requirements​

PropertyTargetConsequence
Alert latency< 10 min from price change to notificationPolling interval must be short for watched products specifically
Scrape budgetBounded and predictableAdaptive scheduling is the core algorithm
Extraction accuracy> 99% — a wrong price is worse than no priceMulti-signal extraction with validation
PolitenessNever harm a retailer's siteSame discipline as the web crawler
Alert fan-out2M users matched in seconds on a big dropInverted index over alerts, not a scan

Back-of-the-envelope​

Products: 500M tracked
Naive: poll every product hourly = 500M × 24 = 12B fetches/day
At even $0.0001/fetch (proxy + compute) = $1.2M/day. Not a business.
Reality: ~2M products have at least one active alert -> those get 15-min polling
= 2M × 96 = 192M fetches/day
The remaining 498M are polled from hourly to monthly by volatility
= ~50M fetches/day
Total ≈ 240M/day = 2,800/sec. A 50× reduction from adaptive scheduling.
Price rows: 240M observations/day, ~95% unchanged -> store only CHANGES: ~12M/day
12M × 30 B = 360 MB/day. Trivial. Storing every observation is not.
Alerts: 50M users × ~4 alerts = 200M alert rules to evaluate per price change

The 12B-versus-240M comparison is the entire pitch. Lead with it: adaptive polling is not an optimisation here, it is what makes the product exist.

2. High-level architecture​

3. Component inventory​

ComponentConcrete choiceWhy this one
SchedulerPriority queue keyed by next-check time, per-retailer politeness bucketsThe same two-tier structure as a crawler frontier
Light fetcherAsync HTTP with browser-like TLS fingerprints95% of pages need nothing more
Heavy fetcherPlaywright pool on Lambda or a container fleet~100× the cost; must be rationed
ProxiesDatacenter → ISP → residential ladder, escalating on blockCost escalates with each tier; start cheap
ExtractionJSON-LD → per-retailer template → generic model → validationLayered, cheapest and most reliable first
HistoryTimescaleDB (or Parquet on S3) storing only changes95% of observations are "no change"
Alert matchingElasticsearch percolatorInverts the problem: index the queries, run the document
NotificationThe notification systemDo not rebuild it

4. The toughest parts​

4.1 Adaptive polling: the algorithm the product is built on​

Why it's hard. Uniform polling is unaffordable and disrespectful to retailers. But a naive "poll popular items more" heuristic misses the actual objective: you want to detect price changes, and a product's change rate is only loosely correlated with its popularity. A stable staple polled hourly wastes 24 fetches a day; a volatile item polled daily misses the drop that mattered.

Solution — estimate each product's change rate and poll proportionally, weighted by how much a missed change would cost.

def next_check_interval(product) -> timedelta:
# 1. Estimate the change rate from observed history (a Poisson-ish process).
# Beta-Binomial with a weak prior handles sparse history gracefully.
changes, observations = product.changes_90d, product.observations_90d
rate = (changes + 1) / (observations + 10) # smoothed change probability

# 2. Base interval targets a fixed probability of missing a change per poll.
base_hours = clamp(TARGET_MISS_PROB / max(rate, 1e-4), 0.25, 720) # 15 min .. 30 d

# 3. Value of information: how much do we lose by being late?
urgency = 1.0
urgency *= 1 + 3.0 * math.log1p(product.active_alerts) # people are waiting
urgency *= 1 + 1.0 * math.log1p(product.page_views_7d) # people are looking
urgency *= 2.0 if product.in_deal_season else 1.0 # Black Friday etc.
urgency *= 2.0 if product.recently_changed else 1.0 # changes cluster

# 4. Retailer-level signal: some retailers reprice on a schedule.
if (phase := retailer_reprice_phase(product.retailer)) is not None:
return align_to(phase) # poll just after it

return timedelta(hours=base_hours / urgency)

Four ideas worth naming. Smoothing with a prior stops a brand-new product (zero observations) from being either ignored or hammered. Value of information is the right framing for urgency: the cost of a missed change is proportional to how many people are waiting for it, which is why an item with 5,000 alerts is polled every 15 minutes and an identical item with none is polled weekly. Change clustering — a product that just changed is far more likely to change again — is a strong empirical regularity worth exploiting. And retailer repricing phase is the highest-leverage single observation: many retailers reprice in a batch at a predictable time, so polling five minutes after that window catches nearly everything for a fraction of the cost.

Close the loop: every poll updates the change-rate estimate, so the scheduler converges on each product's real behaviour without any manual tuning. Cap both ends so nothing is polled more than every 15 minutes (politeness) or less than monthly (staleness).

4.2 Anti-bot defences, and the ladder of escalating cost​

Why it's hard. Retailers actively defend against scraping: IP-range blocks, TLS/JA4 fingerprinting, JavaScript challenges, CAPTCHAs, and behavioural detection. Each countermeasure you deploy costs more than the last — a datacenter proxy is nearly free, a residential proxy plus a headless browser plus a CAPTCHA solve is orders of magnitude more expensive. Escalating everywhere would eliminate the cost advantage from 4.1.

Solution — an escalation ladder, applied per retailer, driven by observed block rates.

TIERS = [
# (fetch method, proxy type, relative cost)
("http", "datacenter", 1),
("http", "isp", 5),
("browser", "isp", 40),
("browser", "residential", 120),
("browser+solver", "residential", 400),
]

async def fetch(url, retailer):
tier = retailer.current_tier # learned, per retailer
for t in range(tier, len(TIERS)):
method, proxy, _ = TIERS[t]
resp = await try_fetch(url, method=method, proxy=pool.get(proxy))

if resp.ok and not looks_blocked(resp):
# Success: try DE-escalating occasionally to find the cheapest tier
# that still works. Defences change in both directions.
if t == tier and random.random() < 0.02:
retailer.probe_tier(max(0, t - 1))
return resp

retailer.record_block(t)
metrics.inc("scrape.blocked", retailer=retailer.id, tier=t)

raise PermanentlyBlocked(retailer)

def looks_blocked(resp) -> bool:
# A 200 with a challenge page is the common case — never trust the status code.
return (resp.status in (403, 429, 503)
or "captcha" in resp.text.lower()[:5000]
or resp.text_length < 500
or CHALLENGE_MARKERS.search(resp.text) is not None)

Two points that matter operationally. looks_blocked must inspect content, not status — anti-bot systems commonly return 200 OK with a challenge page, and a scraper that trusts the status code silently records garbage prices, which is worse than failing. And periodic de-escalation probes stop you from paying the highest tier forever after one bad week; defences change, and without probing you never discover that the cheap path works again.

The most important paragraph in the answer, though, is the one about not scraping at all:

Before any of this, prefer the legitimate channels. Most large retailers offer affiliate product APIs (Amazon PA-API, Rakuten, CJ), merchant feeds (Google Shopping XML), or partner arrangements. These are cheaper, more accurate, more complete, and permitted. Scraping should be the fallback for retailers with no programme, done politely with a published bot identity, honouring robots.txt and Retry-After, and staying well within the traffic a normal user would generate.

Saying this unprompted is a strong signal: it shows you understand the legal and ethical dimension is a design constraint, not a footnote, and that the cheapest solution is usually the cooperative one.

4.3 Alert fan-out: matching 200 million rules in milliseconds​

Why it's hard. A popular product drops in price. Two million users have alerts on it, with heterogeneous conditions: "below $499", "20% off", "any drop", "back in stock", "below the 90-day average". Scanning 200M alert rows to find matches on every one of 12M daily price changes is 2.4 × 10^15 evaluations per day.

Solution — invert the problem with a percolator: index the queries, and run the document against them.

PUT /alerts/_doc/alert_9f3a
{
"user_id": 88213,
"query": {
"bool": {
"filter": [
{"term": {"product_id": "B08N5WRWNW"}},
{"range": {"price_cents": {"lte": 49900}}},
{"term": {"in_stock": true}}
]
}
}
}
# On a price change: one query returns every matching alert.
async def on_price_change(product_id, new_price_cents, in_stock, pct_drop):
hits = await es.search(index="alerts", body={
"query": {"percolate": {
"field": "query",
"document": {
"product_id": product_id,
"price_cents": new_price_cents,
"in_stock": in_stock,
"pct_drop": pct_drop,
"below_90d_avg": new_price_cents < await avg_90d(product_id),
}}},
"size": 10000})

# Batch into the notification pipeline; never notify inline.
await kafka.send_batch("price.alert.matched",
[{"user": h["_source"]["user_id"], "product": product_id,
"price": new_price_cents} for h in hits])

The inversion is the whole idea and it is worth stating plainly: normally you index documents and run queries; a percolator indexes queries and runs documents. Lucene's inverted index over the alert conditions makes matching sub-linear in the number of alerts, turning an impossible scan into a single indexed lookup.

Two supporting details. Precompute derived fields (pct_drop, below_90d_avg) into the percolated document, so an alert condition can reference them without the percolator needing to run computations. And route by product ID — shard the alert index by product so a price change touches exactly one shard, which keeps the query fast even at 200M rules.

Then apply the standard notification hygiene from the notification system: deduplicate (a price oscillating around a threshold must not send ten alerts), apply frequency caps, and respect quiet hours — with a deliberate exception for genuinely time-limited deals, which is a product decision worth surfacing.

4.4 Getting the right number off ten thousand different pages​

Why it's hard. There is no standard. The price might be in a <span class="a-price">, rendered by JavaScript after a fetch, split across elements ($ 49 .99), shown as a range for variants, or accompanied by three other numbers (list price, member price, price with a coupon). Extracting the wrong one produces a false alert — and a false "price dropped to $49!" that turns out to be $499 destroys user trust immediately.

Solution — layered extraction, cheapest and most authoritative first, then validate hard.

def extract_price(page, product) -> PriceObservation | None:
# 1. Structured data. Authoritative when present; a majority of large retailers
# emit it because Google Shopping requires it.
if offer := parse_schema_org_offer(page.html):
return PriceObservation(cents=to_cents(offer.price), currency=offer.currency,
in_stock=offer.availability == "InStock",
source="schema.org", confidence=0.99)

# 2. Per-retailer template, learned once and monitored for drift.
if tmpl := templates.get(product.retailer_id):
if (obs := tmpl.apply(page)) and obs.valid:
return obs._replace(source="template", confidence=0.95)

# 3. Generic model: features from DOM position, font size, currency symbol
# proximity, class-name tokens, visual prominence.
cands = price_candidates(page)
if cands and (best := model.rank(cands, context=product))[0].score > 0.8:
return best[0].observation._replace(source="ml", confidence=best[0].score)

metrics.inc("extraction.failed", retailer=product.retailer_id)
return None

def validate(obs, product) -> bool:
h = product.history
if obs.currency != product.expected_currency: return False
if obs.cents <= 0 or obs.cents > 100_000_00: return False
# A price outside 10%..300% of the 90-day median is almost always an
# extraction error (a coupon, a per-month figure, an accessory's price).
if h and not (0.10 * h.median_90d <= obs.cents <= 3.0 * h.median_90d):
quarantine(obs, reason="implausible_vs_history") # re-fetch to confirm
return False
return True

The validation-against-history check is the highest-value guard and the one that separates a working product from a trust disaster. A genuine 80% discount does happen, but it is far rarer than an extraction bug — so quarantine the observation, re-fetch (ideally through a different tier), and only publish if it reproduces. Requiring two independent confirmations before firing an alert on a dramatic drop costs one extra fetch and prevents the failure mode users never forgive.

Also monitor extraction confidence per retailer over time: a retailer redesigning their page shows up as a sudden drop in template success rate, which should page someone before a million users get wrong prices.

4.5 Storing history for 500 million products​

Why it's hard. 240M observations/day for years, and users want a two-year chart rendered instantly. Storing every observation is 87B rows/year of mostly-identical values.

Solution — store transitions, not observations; derive everything else.

-- Only rows where something CHANGED. ~95% of polls write nothing.
CREATE TABLE price_changes (
product_id BIGINT NOT NULL,
observed_at TIMESTAMPTZ NOT NULL,
price_cents INTEGER NOT NULL,
in_stock BOOLEAN NOT NULL,
seller_id INTEGER,
PRIMARY KEY (product_id, observed_at DESC)
);
SELECT create_hypertable('price_changes', 'observed_at', chunk_time_interval => '7 days');

-- Separately: proof of liveness, so a gap in changes isn't mistaken for a gap in data.
CREATE TABLE poll_log (product_id BIGINT, last_polled TIMESTAMPTZ, PRIMARY KEY (product_id));

Storing only changes cuts 87B rows/year to about 4B, and reconstructing the series is a step function — the price between two change rows is the earlier row's value. The poll_log table matters: without it you cannot distinguish "the price did not change for six weeks" from "we did not look for six weeks," and that distinction is exactly what a price-history chart needs to be honest about.

For charts, precompute daily min/max/close rollups (the same tiering idea as metrics monitoring) so a two-year chart reads 730 rows rather than reconstructing from raw changes. Age raw changes beyond a year into Parquet on object storage.

The derived signals users actually want — "is this a good price?" — come from the same data: the 90-day median, the all-time low, the percentile of the current price against its own history, and how long the current price has held. Those are what make the product useful rather than just a chart.

4.6 Products, variants, and identity across retailers​

Why it's hard. "The same product" is genuinely ambiguous. The same TV has a different SKU at every retailer, model numbers vary by region, listings bundle accessories, and marketplaces have many sellers for one item at different prices and conditions. Getting identity wrong means comparing prices for different things and alerting on nonsense.

Solution — a canonical product entity assembled from strong identifiers first, then fuzzy matching with a confidence threshold.

def resolve_product(listing) -> ProductId:
# 1. Strong identifiers. Unambiguous when present — always prefer them.
for ident in (listing.gtin, listing.upc, listing.ean, listing.isbn, listing.mpn):
if ident and (pid := index.by_identifier(ident)):
return pid

# 2. Brand + model number extracted from the title, normalised.
if (brand := listing.brand) and (model := extract_model_number(listing.title)):
if pid := index.by_brand_model(normalise(brand), normalise(model)):
return pid

# 3. Fuzzy: title embedding + image similarity + attribute overlap.
# Require a HIGH threshold — a wrong merge is worse than a missed one,
# because it produces confidently wrong price comparisons.
cands = index.ann_search(embed(listing.title), k=20)
best = max(cands, key=lambda c: match_score(listing, c), default=None)
if best and match_score(listing, best) > 0.92:
return best.product_id

return index.create_product(listing) # new canonical entity

Two points worth stating. Strong identifiers first, always — GTIN/UPC matching is exact and free, and reaching for embeddings when a barcode is present is a mistake. And the asymmetric threshold: an over-merge shows users a price for a different product (a trust-destroying bug), while an under-merge shows two entries for the same item (mildly untidy). Bias hard toward not merging — the same asymmetry as story clustering in the news aggregator.

Track variants (size, colour, condition) as children of a canonical product with their own price series, because a user alerting on "64 GB, black" does not want the 128 GB price.

5. What breaks first​

EventFirst failureMitigation
Black FridayEverything reprices at once; queue explodesPre-scale; temporarily raise polling for alerted products only; shed low-priority polls
A retailer deploys new anti-botBlock rate spikes for that retailerEscalation ladder auto-climbs; alert if it reaches the top tier
A retailer redesignsExtraction silently returns wrong pricesPer-retailer confidence monitoring; history-plausibility quarantine; two-confirmation rule
Popular product drops2M alert matches at oncePercolator returns them in one query; notification system's bulk lane absorbs the fan-out
Price oscillates around a thresholdAlert spamDedup by (user, product, price_bucket, day); require a minimum drop delta
Proxy provider outageScraping stops for affected tiersMultiple providers per tier; degrade to feeds and cheaper tiers

6. Cheat sheet​

  • The number: uniform hourly polling is 12B fetches/day; adaptive polling is ~240M. A 50× reduction is the business model.
  • Scheduling: smoothed per-product change rate × value-of-information (alerts, views, season, recent change), aligned to retailer repricing phase.
  • Fetching: escalation ladder (datacenter → ISP → browser → residential → solver) learned per retailer, with periodic de-escalation probes. Detect blocks by content, never by status code.
  • Prefer APIs and feeds over scraping wherever they exist — cheaper, more accurate, and permitted.
  • Alerts: Elasticsearch percolator inverts the problem — index the queries, percolate the document. Precompute derived fields.
  • Extraction: schema.org → template → ML → validate against price history; quarantine implausible values and require two confirmations for dramatic drops.
  • Storage: store changes, not observations, plus a separate poll log so gaps are honest. Daily rollups for charts.
  • Identity: GTIN/UPC first, brand+model second, fuzzy last with a high threshold — an over-merge is far worse than an under-merge.
  • The one-liner: "You cannot look at everything, so the product is a scheduler: estimate each item's change rate, weight it by how many people are waiting, and spend a fixed fetch budget where it buys the most information."