Evaluation of Performance
ProductHub vs. Human-Labelled Products
ProductHub was evaluated against 10,000 human-labelled products from the Australian GS1 dataset, comparing agent-generated GPC brick classifications with those assigned by humans.
- Top-1 Match: The agent and human agreed on the top brick 72% of the time.
- Top-3 Match: The correct human-assigned brick appeared within the agent's Top-3 recommendations 84% of the time.
Independent LLM Judge Review
Where the agent and human disagreed, independent LLM judges reviewed the classifications using all available product context and GPC hierarchy definitions:
- In Top-1 disagreements, LLM judges found the agent's classification correct 96–100% of the time, with human labels judged correct only 0–4%.
- In Top-3 evaluations, LLM judges confirmed the agent's classification as correct 98–100% of the time.
Key insight: Humans make material classification errors in approximately 20–30% of cases, particularly when product descriptions are incomplete, ambiguous, or outdated. ProductHub's agentic approach dramatically reduces these errors by consistently applying enriched context, hierarchy knowledge, and structured reasoning.