Meridias

Benchmark Workloads

Workload suites used to evaluate providers across latency, quality, and stability.

Benchmark workloads are the repeatable units Meridias uses to measure providers. Each workload is designed to represent a real operational category rather than a generic synthetic test.

Category coverage

Workloads are grouped by task families such as structured generation, classification, extraction, and multi-step reasoning. This lets teams compare providers within the context they actually plan to route.

Dataset composition

Each workload uses curated prompts, fixtures, or request payloads that are broad enough to expose quality variance without drifting from the intended task definition. Versioned datasets make reruns comparable over time.

Success criteria

Workloads define their own acceptance rules, including output validity, response completeness, and latency thresholds. This avoids compressing very different task shapes into a single generic pass or fail score.

Refresh cadence

Benchmark suites should be rerun often enough to catch provider changes but not so frequently that minor noise overwhelms the signal. Cadence decisions should follow provider volatility and customer sensitivity.

Next step

Use Merit Rankings to see how workload outputs are summarised.