Thank you for wanting to learn more about us. Download the PDF
here.
To stay updated with our latest content, please follow us on LinkedIn.
Oops! Something went wrong while submitting the form.
Retail

Future-proofing a next-gen platform for a leading data management player

6-month launch
for the cloud-ready MAP
Tens of millions
raised in Series A funding
Ranked #1
by Gartner and Forrester
tech stack
Polymer JS, Apache Storm, Kafka, Elasticsearch, Ngram, Kibana
AT A GLANCE

Zemoso provided end-to-end digital transformation services for an industry leader in data management and governance platforms for product information management (PIM). The company’s platform was recognized several times by reputed industry analysts (such as Gartner and Forrester) for its leadership in this space. Our client wanted to evolve their existing platform into a more intuitive, smart version of itself to retain their industry stronghold.

CORE CHALLENGEs

Our client needed to help their customers grow, protect, and forecast e-commerce sales through better data-driven decisions on pricing, selection, advertising, and supply chain - while keeping their industry-leading product information management platform agile enough to solve the data problems of tomorrow. That meant building a next-generation, on-demand, cloud-native data management and governance platform, modernizing the tech stack underneath it, and designing an API-driven, DevOps-ready architecture from the ground up.

the engineering approach

The team aligned quickly, pivoted as needed, tested rapidly, and delivered consistent, incremental wins throughout the partnership.

We helped our client speed up development and evolution cycles, consistently delivering well-designed applications. After evaluating multiple providers, we chose Amazon Web Services (AWS) for its greater flexibility, scalability, and lowest developer friction with SDKs. In partnership with their internal engineering team, we designed a multi-layered platform using a microservices architecture for a modular and agile system. The team used DevOps best practices for continuous integration and deployment to expand capabilities more quickly.

  • Upgrading the tech stack: For each functionality, we evaluated many competing tech solutions and selected the best-suited solutions with the client’s engineering leadership.
  • User interface: Polymer JS was chosen to create the next-gen user interface with added controls for tiered access protocols. It eased setting up different application elements and their relations.
  • Events stream processing: The events stream processing (key to developing deduplication functionality) was built using Apache Storm, a highly vetted solution for real-time stream processing. The team conducted a thorough evaluation between Apache Spark and Apache Storm. Apache Storm was a better fit than Spark, which is a far more complex, general purpose computation engine.
  • Messaging layer: To ensure smart deduplication whenever a record is added or updated, Kafka messaging layer transfers the data between applications. It is a fast and scalable event distributor. The ‘deduplication’ and ‘match and merge’ queues were robust and processed high volumes of data with minimal downtime or data loss. It integrated seamlessly with Apache Storm for real-time streaming data analysis, which helped their clients gain faster insights.
  • Improved searchability: For search, the team chose Elasticsearch. It is a distributed, free and open-sourced search and analytics engine that powers search solutions for global giants like Microsoft, Netflix, Slack, and Uber. This text-based, NoSQL search tool proved highly useful in indexing data points needed to fulfill search parameters. The team used advanced indexing techniques like Ngram to generate superior match results. For instance, with only the first three letters as input, the search engine could match and reflect the product name.
  • Custom deduplication algorithm: As operators, suppliers, and vendors enter product data into the platform, deduplication keeps that information accurate and reliable. We helped our MDM client ensure the same product isn't listed twice under different unique IDs by developing and deploying a deduplication algorithm on their newly upgraded tech stack. It uses similarity triggers around name, brand, and color to flag potential duplicates with a probability index across a database of over a million records. Every system change becomes an "event" that runs through a "match and merge" protocol and enters a queue, while suspect matches are routed for manual resolution.
  • Reporting dashboards: We co-created dashboards that improve access to learnings from the master data and help the client's customers visualize information and generate intelligent reports for better analysis. These surface change trends, update summaries, workflow SLAs, governance summaries, and more. Built on Kibana, a browser-based analytics and search dashboard for Elasticsearch, they enable quicker analysis, faster data compilation, and easy report sharing with stakeholders, while also using machine learning features to detect patterns in Elasticsearch data.

P.S. Since we work on early-stage products, many of them in stealth mode, we have strict Non-disclosure agreements (NDAs). The data, insights, and capabilities discussed in this blog have been anonymized to protect our client’s identity and don’t include any proprietary information.

bottom line

By rebuilding the platform's UI, event-streaming, search, and deduplication layers with a carefully evaluated modern tech stack, Zemoso helped a market-leading data management platform stay agile enough to keep its Gartner and Forrester leadership position while scaling to serve some of its largest customers across retail, CPG, healthcare, and energy.

Related resources

Got an idea?

Together, we’ll build it into a great product

DesignDesign
Zemoso Technologies
ZemosoTechnologies ZemosoTechnologies

©2025 Zemoso Technologies. All rights reserved.