View Project
TrackIt / AWS · 2024 · Product Manager

GenAI Model Evaluation Tool

An internal tool that lets engineers and PMs benchmark competing GenAI models on accuracy, latency, and cost in a single workflow, replacing ad-hoc spreadsheet comparisons and cutting evaluation cycles from weeks to hours.

Product StrategyUser ResearchTechnical SpecAI / LLM
GenAI Model Evaluation Tool
Overview
ClientTrackIt / AWS
Date2024
RoleProduct Manager
Scope of workProduct Strategy, User Research, Technical Spec, AI / LLM

The problem.

Teams spent 2 to 3 weeks per evaluation cycle copy-pasting prompts between playgrounds, reconciling outputs in spreadsheets, and debating cost trade-offs with incomplete data. The process was manual, inconsistent, and left no record of how decisions were made.

No shared baselineEvery evaluation started fresh, with different engineers using different prompts, different models, and different scoring criteria. Comparing results across runs was guesswork.
2 to 3 week cyclesCopy-pasting prompts between playgrounds and reconciling outputs in spreadsheets consumed weeks that should have taken hours.
No reproducibilityThere was no way to rerun an evaluation against the same conditions, making it impossible to validate improvements or catch regressions.
No audit trailWhen a model underperformed in production, teams had no record of what the evaluation showed, what was compared, or who made the call.
Cost and latency blind spotsDecisions were made on output quality alone, without visibility into cost-per-token or p95 latency across models.
Evaluation run configuration
Evaluation run configuration: prompt corpus, model selection, and scoring rubric in one view.
Approach

What we built.

A single interface for multi-model benchmarking that runs Claude, GPT-4, and Amazon Titan in parallel against the same prompt corpus, producing structured comparison outputs with quality scores, cost, and latency side by side.

01
Prompt corpus definitionUsers configure a reusable set of prompts that run consistently across every model and every evaluation cycle.
02
Multi-model parallel runsClaude, GPT-4, and Amazon Titan run simultaneously against the same corpus, with results captured in a structured comparison view.
03
Scoring rubricsHuman-in-the-loop scoring modules let teams define and apply consistent quality criteria across outputs.
04
Cost-per-token and latency metricsEvery run surfaces p95 latency and cost-per-token alongside quality scores, making trade-offs explicit.
05
One-click decision reportPMs and engineers can export a structured report capturing the evaluation, the scores, and the recommendation.
Model comparison view
Structured comparison view across models with quality scores and cost-per-token.
Cost vs accuracy scatter plot
Cost vs accuracy trade-off view, making model selection decisions explicit and shareable.
Results

The outcome.

Evaluation cycles dropped from 14 days to under 6 hours. The tool became the standard decision gate before any LLM upgrade in the platform.

It prevented two costly model regressions that would have reached production under the old process, and gave teams a shared, auditable record of every model decision made since rollout.

All Projects

Back to work