Client Deployment

Built a custom AI shopping assistant that outperformed the incumbent open-source model by 14%

A global eCommerce company needed to improve its AI shopping experience without increasing friction, missed sales opportunities, or operational risk. aion embedded forward-deployed engineers with the client's team and built a custom model and evaluation system that delivered measurable performance gains and a repeatable path to continuous improvement.

Engagement

The Brief

Client

A global eCommerce company

Assessment Method

Automated + human review (GEval methodology)

Delivery Model

Embedded aion forward-deployed engineers

14%Above incumbent open-source benchmark
5Capability dimensions

The Business Challenge

Context

As adoption of the company's AI shopping assistant increased, customer experience became inconsistent. The assistant frequently:

Unnecessary questions

Asked unnecessary questions that slowed purchases.

Missed checkout guidance

Missed opportunities to guide customers toward checkout.

Incorrect tool calls

Made incorrect backend tool calls.

Robotic support responses

Produced robotic responses during support interactions.

Most critically, the team had no reliable way to measure whether any change made the assistant better or worse. Without a benchmark, they couldn't confidently improve quality or scale the experience.

What aion Delivered

Approach

A Production Evaluation System

Automated and human review pipelines that continuously measure assistant performance across five capability dimensions: purchase guidance, question efficiency, tool-call accuracy, conversational quality, and support handling.

A Custom-Tuned Model

A model developed and optimized specifically for the client's shopping workflows and customer interactions.

A Fine-Tuning Strategy

A data collection, labeling, and optimization framework designed to compound performance improvements over time.

An Optimization Roadmap

A prioritized sequence of the highest-impact improvements required to reach production-grade quality.

Outcomes

Outcome

14% better than the incumbent model

Measured via structured evaluation (GEval methodology) across all five capability dimensions, the custom model outperformed the open-source model previously powering the client's production assistant by 14%.

Quality became measurable

For the first time, the client could quantify reliability and customer experience across every model release, turning model upgrades from a judgment call into a data-driven decision.

Faster improvement cycles

Evaluation pipelines eliminated guesswork, so engineering effort went to the changes that moved the metrics.

A foundation for scale

The client now owns a repeatable framework for future model upgrades, new languages, and expanded customer experiences.

The takeaway

The client didn't just receive a better model. They gained a repeatable system for measuring, improving, and scaling AI performance: every future model release can now be evaluated against clear business and customer experience outcomes.

Have a problem like this?Let’s deploy the answer.