Meridias

Evaluation Methodology

How Meridias runs controlled benchmarks and turns them into comparable results.

The evaluation methodology defines how providers are tested, how results are normalised, and how confidence is communicated to teams using the routing layer.

Benchmark design

Each benchmark workload is versioned, repeatable, and scoped to a known category. Inputs are selected to exercise the same task shape across providers so the comparison measures provider behaviour rather than prompt drift.

Controlled execution

Runs are scheduled in controlled windows with stable infrastructure, known regions, and consistent client settings. This reduces environmental noise and makes changes in provider performance easier to interpret.

Scoring model

Raw benchmark outputs are transformed into comparable scores across latency, quality, completion success, and operational stability. Scoring weights can vary by workload category so the final ranking reflects the needs of that task.

Confidence and drift

Meridias tracks when a provider changes enough that historical results become less predictive. Confidence signals help teams understand whether a merit ranking is backed by fresh evidence or by data that may need revalidation.

Next step

Read Benchmark Workloads for the structure of the underlying workload suites.