The problem
Search is only as good as the attributes behind it, and every catalog is full of missing sizes, inconsistent categories, and vendor-supplied fiction. Fixing it manually costs headcount nobody has — which is why the biggest retailers moved this exact job to LLMs.
The system
An enrichment pipeline that runs LLM extraction and normalization over product titles, descriptions, images, and vendor feeds — schema-validated, confidence-scored, spot-audited against a hand-labeled sample — feeding attribute-aware semantic search with measurable relevance, not vibes-based embeddings.
How it's built
- Attribute extraction/normalization with per-field confidence and schema validation
- Hand-labeled audit sample: enrichment accuracy reported per attribute
- Hybrid semantic + attribute search with offline relevance evals on your real queries
- Cost discipline: batch tiers and caching so the pipeline stays cheaper than headcount
Delivery
Sprint enriches one category and reports measured accuracy and search-relevance lift; Build scales across the catalog.
What to expect
- Attribute coverage and accuracy measured per field, before and after
- Search relevance evaluated on your real query set, offline first
- Enrichment unit cost that undercuts manual curation by an order of magnitude
Documented results in the wild
Independent, published deployments of this class of system — cited as market evidence that it works at scale. These are not our clients.
- Walmart LLMs created or improved 850 million pieces of catalog data — work that would otherwise have taken ~100× the headcount. CIO Dive / Walmart earnings, 2024 ↗
- Wayfair Gemini-based catalog enrichment cut listing-curation time 67% and improved some conversion rates by 2%. PYMNTS / Google Cloud, 2025 ↗