Scaling Product Data Accuracy: Website Indexers and Crons for a Leading E-Commerce Platform
How automated data pipelines kept a 92,000+ SKU marketplace accurate.
Oct 6, 2024
case Page image

Summary

Keeping product data accurate is one of the hardest operational problems a large e-commerce marketplace faces and one of the most invisible when it’s working. Our client, a leading e-commerce platform, was managing 120+ product categories, 92,000+ product catalogues, 550+ brands, and 28,000+ local sellers, with catalog and inventory data flowing in from multiple business teams on different schedules. Manually keeping that data accurate and current across the live site had become operationally unsustainable at this scale.

Bajaj Tech.AI designed and implemented an automated data pipeline combining a centralized ingestion layer, a scheduled ETL (Extract, Transform, Load) process, and a system of purpose-built indexers and crons that keeps the platform’s catalog, pricing, offers, and dealer data synchronized across the site multiple times a day, with minimal manual intervention. The result is a marketplace that stays accurate and current at a scale where manual updates were no longer viable.

Business Challenge

A marketplace of this size 120+ categories, 92,000+ catalogues, 550+ brands, and 28,000+ local sellers receives a constant stream of changes: new products, updated pricing, seller-dealer mapping changes, and time-bound offers, all originating from different business teams working on different systems and schedules.

The core challenge was extracting, transforming, and loading this volume of data into the platform’s database reliably and frequently enough to keep the live site accurate, without that process becoming a bottleneck as the catalog kept growing. Stale or inconsistent product data at this scale doesn’t just create a poor browsing experience, it erodes trust in search results, creates operational rework for the catalog and business teams, and risks customers seeing incorrect pricing or availability.

Compounding the problem, data arrived from more than one source on more than one schedule: the catalog team creating product categories and listings directly in the commerce platform, alongside separate seller-dealer mapping data. Any solution needed to reconcile both streams into a single, consistent view of the catalog not just process one feed faster.

Summary: At 92,000+ catalogues and 28,000+ sellers, manually keeping product, pricing, and dealer data accurate had become operationally unsustainable the platform needed an automated, reliable data pipeline instead.

Solution Approach

Bajaj Tech.AI designed a solution built around a centralized ingestion point and a scheduled ETL pipeline, feeding a set of purpose-built indexers and crons (a cron is a scheduled, automated task that runs without manual triggering) that keep different parts of the platform synchronized:

  • Centralized data ingestion: Dealer details, catalog data, inventory, and dealer master data are received from the business team in a structured ETL format via a centralized AWS S3 bucket, three times a day.
  • Automated processing pipeline: Files are processed daily through an automated ETL pipeline that transforms and loads the incoming data into the platform’s database.
  • Catalog and inventory indexing: A combination of bulk and incremental indexers keep product, pricing, dealer, and city-level data current bulk indexers handle full data refreshes, while incremental indexers process only what’s changed, keeping the system efficient as the catalog grows.
  • Search indexing: Data is pushed into Elasticsearch to power fast, accurate on-site search results as the catalog updates.
  • Offers and promotions indexing: Time-bound offers are validated and automatically activated or deactivated in the catalog based on offer data, removing manual toggling.
  • Automated content generation: Scheduled crons regenerate product detail pages (PDPs) and product listing pages (PLPs) in Adobe Experience Manager directly from the updated catalog data, keeping the customer-facing site aligned with backend inventory without manual publishing.
  • Downstream integrations: Additional crons handle order data synchronization, payment processing updates, and feeds to external systems like Google Shopping and the platform’s search partner, keeping reporting and third-party integrations current.

Alongside the technical build, Bajaj Tech.AI worked directly with the client’s business team to establish a structured, agreed-upon data-sharing process and cadence aligning how and when data was submitted to the pipeline, which minimized manual intervention and kept the process running smoothly on an ongoing basis.

Summary: The solution combines centralized data ingestion, a scheduled ETL pipeline, and a system of specialized indexers and crons that keep catalog, search, offers, and content in sync automatically, multiple times a day.

Business Impact & Results

The automated pipeline replaced a process that couldn’t scale manually with one built to run continuously and reliably:

  • Improved data syncing: The platform’s database now stays synchronized with the latest catalog, pricing, and dealer data automatically, reducing the manual intervention that previously created delays and inconsistencies.
  • Increased operational efficiency: Automated processing of a large, constantly-changing dataset significantly reduced the time and effort the business and catalog teams previously spent on manual updates.
  • Enhanced customer experience: With accurate, current data reliably reflected on the site, customers can search for and find products more easily, directly improving the browsing and purchase experience.
  • Ongoing performance visibility: A monthly indexer performance report and a dedicated indexer dashboard now give the business team ongoing visibility into indexer success and failure rates, enabling continuous monitoring and optimization rather than reactive troubleshooting.
Summary: The automated pipeline turned a manually unsustainable data process into a continuously monitored system, improving data accuracy, operational efficiency, and the customer search experience at scale.changes and the need for a scalable and stable solution.

Business KPIs

Monthly Indexer Performance Report: Tracks the performance of each indexer, including success and failure rates.

Image

Indexer Dashboard: Provides a detailed report of all indexers, enabling performance analysis and optimization.

Image

Key Takeaways

  • At marketplace scale 92,000+ catalogues and 28,000+ sellers in this case, manual catalog and inventory updates stop being viable, and automated, scheduled data pipelines become an operational necessity, not an optimization.
  • Separating bulk and incremental data processing (full refreshes vs. only what’s changed) keeps large-scale indexing systems efficient as catalog size grows.
  • Automating content generation (PDP/PLP pages) directly from catalog data keeps the customer-facing site aligned with backend inventory without manual publishing steps.
  • Establishing a clear, agreed-upon data-sharing process with business teams is as important as the technical pipeline itself, it’s what keeps manual intervention minimal on an ongoing basis.
  • Ongoing performance monitoring (dashboards, regular reporting) turns a one-time technical build into a system the business can continuously track and optimize.

Conclusion

For large-scale marketplaces, product data accuracy isn’t a one-time technical project, it’s an ongoing operational capability that has to scale alongside the catalog itself. By combining centralized ingestion, automated ETL processing, and purpose-built indexing and content-generation systems, Bajaj Tech.AI helped this client turn a manually unsustainable data process into a reliable, continuously monitored pipeline. Organizations managing large, fast-changing catalogs across multiple sellers or brands face a similar structural challenge and a similar automated approach.

Looking to solve a similar data accuracy or catalog-scaling challenge? Connect with our experts to explore the right solution for your organization.

Written by
Dhiraj Jha
Head - Experience & Commerce