Case study

Applied AI

Vision Model Evaluation Harness

A cost-aware benchmark that compares nine vision models on extraction accuracy while fetching current per-token pricing at runtime.

Outcome

Nine vision models evaluated on accuracy and real inference cost.

Problem

Model quality claims are hard to compare and pricing changes quickly. Selecting a production OCR model required a repeatable test set, a scored baseline, and the real cost of each run rather than a demo-driven choice. The harness turns model selection into an evaluation that can be rerun when models, prompts, or prices change.

Benchmark loop

01

Freeze the task

Use the same representative receipt set and expected structured fields for every model.

02

Run every model

Normalize provider responses into one schema and retain errors instead of dropping failed cases.

03

Score against truth

Compare extracted values with the baseline field by field so one plausible answer cannot hide another miss.

04

Fetch current pricing

Calculate cost from runtime token usage and current provider pricing rather than a stale spreadsheet.

05

Choose by frontier

Compare quality and cost together; the best model is the one that meets the production threshold economically.

Decision / Accuracy is not enough

Treat cost as part of model behavior.

A model that wins a small accuracy margin can lose in production if it costs several times more at real document volume. Fetching prices at runtime keeps that tradeoff visible and makes the recommendation reproducible.

9

models, one scored baseline

9

Vision Models

Live

Pricing Lookup

1

Shared Baseline

Repeatable

Scoring Harness

Built With

TypeScriptModel EvaluationVision ModelsStructured OutputPricing APIs