Case Study · Lighting E-commerceFrance · Europe

Luxarmonie Product
Intelligence

A French lighting e-commerce was spending 40 hours a month hand-extracting specs from inconsistent supplier catalogs. We turned that into an automated extraction pipeline — and an ongoing technical review that keeps the catalog accurate as new products arrive.

−75%
Manual load
99.2%
Data integrity
30h
Reclaimed / month
Luxarmonie product-data pipeline dashboard

The Context

Luxarmonie is a French e-commerce brand selling professional and commercial lighting, sourced largely from manufacturers in China. Their growth depended on listing products fast — but every new product first had to be understood, technically, before it could be sold.

Supplier data arrived in whatever format each factory used: PDFs, spreadsheets, mixed languages and units, specs buried in product images. Someone had to read all of it, extract the real technical values, and turn them into clean listings a buyer could trust.

The Problem

40 hours a month that couldn't scale

Extracting specs by hand took roughly 40 hours every month — slow, draining, and impossible to grow without hiring. Worse, manual transcription introduced errors into technical specs, and in lighting an inaccurate lumen, CCT, or IP rating quietly erodes buyer trust.

Supplier formats inconsistent and unstructured
Manual extraction error-prone and slow
No way to scale listings without scaling headcount

The Process

01

A canonical spec model

Before automating anything, we defined the single source of truth: a normalized schema for every value that matters in a lighting listing — lumen output, CCT, CRI, IP rating, wattage, beam angle, dimensions, materials. Inconsistent supplier data now had one target to map to.

Data Modeling Spec Normalization
02

An intelligent extraction pipeline

Built in Python with the Claude API: the pipeline reads supplier PDFs and documents, extracts the technical specs wherever they hide, normalizes them into the schema, and flags anything ambiguous for human review instead of guessing. What took a day now runs in minutes.

Python Claude API PDF Extraction
03

E-commerce output + ongoing technical review

The clean data feeds straight into e-commerce-ready product descriptions. And because new products keep arriving, the engagement didn't end at delivery — it became a continuous technical review, validating specs against the suppliers in China so the catalog stays accurate as it grows.

Description Gen Continuous Review Supplier Validation

The Results

Manual load cut from 40h to 10h per month (−75%)
The bulk of spec extraction now runs automatically. The team spends its time reviewing edge cases and selling — not transcribing datasheets.
99.2% data integrity across the catalog
Normalized extraction plus human-in-the-loop validation removed the silent transcription errors that used to slip into technical specs.
Listings that scale with the catalog, not the headcount
New products are processed in minutes, and ongoing technical review keeps the data trustworthy as the catalog grows — a relationship, not a one-off delivery.

Diego transformed our product documentation process. What used to take our team 40 hours a month now runs automatically. The accuracy is remarkable — no more manual errors in technical specs.

CEO · Luxarmonie, France

Tools Used

Python Claude API PDF Extraction Data Normalization Spec Validation Technical Review

Drowning in supplier catalogs?

If your lighting catalog grows faster than your team can keep its specs accurate, I build the pipeline — and stay on as the ongoing technical review that keeps it trustworthy. Let's talk about a monthly partnership.