The problem
Competitor prices move weekly; your view of them is an intern’s quarterly spreadsheet. Repricing decisions run on stale, partial data — or on nothing.
The system
A monitoring pipeline over competitor sites and channels: change-detection so pages are fetched only when they change, tiered extraction (deterministic first, LLM fallback), entity resolution so "the same product" matches across sites, and deltas delivered as alerts, feeds, or warehouse tables.
How it's built
- Crawl fleet with per-domain politeness and change-detection gating
- Tiered extraction with confidence scores; eval set on hand-checked pages
- Entity matching across catalogs; history retained for trend analysis
- Delivery: Slack/email alerts, API, or straight into your warehouse
Delivery
Sprint proves extraction quality on your top competitors; Run operates the fleet as sites change under it.
What to expect
- Fresh competitive coverage instead of quarterly snapshots
- Extraction accuracy tracked against hand-labeled truth
- Pricing meetings start from the same live dataset