Case study
Applied AI
Vision Model Evaluation Harness
A cost-aware benchmark that compares nine vision models on extraction accuracy while fetching current per-token pricing at runtime.
Outcome
Nine vision models evaluated on accuracy and real inference cost.
Problem
Model quality claims are hard to compare and pricing changes quickly. Selecting a production OCR model required a repeatable test set, a scored baseline, and the real cost of each run rather than a demo-driven choice. The harness turns model selection into an evaluation that can be rerun when models, prompts, or prices change.
Benchmark loop
01
Freeze the task
Use the same representative receipt set and expected structured fields for every model.
02
Run every model
Normalize provider responses into one schema and retain errors instead of dropping failed cases.
03
Score against truth
Compare extracted values with the baseline field by field so one plausible answer cannot hide another miss.
04
Fetch current pricing
Calculate cost from runtime token usage and current provider pricing rather than a stale spreadsheet.
05
Choose by frontier
Compare quality and cost together; the best model is the one that meets the production threshold economically.
Decision / Accuracy is not enough
Treat cost as part of model behavior.
A model that wins a small accuracy margin can lose in production if it costs several times more at real document volume. Fetching prices at runtime keeps that tradeoff visible and makes the recommendation reproducible.
9
models, one scored baseline
9
Vision Models
Live
Pricing Lookup
1
Shared Baseline
Repeatable
Scoring Harness
Built With