Benchmark Workloads
Workload suites used to evaluate providers across latency, quality, and stability.
Benchmark workloads are the repeatable units Meridias uses to measure providers. Each workload is designed to represent a real operational category rather than a generic synthetic test.
Category coverage
Workloads are grouped by task families such as structured generation, classification, extraction, and multi-step reasoning. This lets teams compare providers within the context they actually plan to route.
Dataset composition
Each workload uses curated prompts, fixtures, or request payloads that are broad enough to expose quality variance without drifting from the intended task definition. Versioned datasets make reruns comparable over time.
Success criteria
Workloads define their own acceptance rules, including output validity, response completeness, and latency thresholds. This avoids compressing very different task shapes into a single generic pass or fail score.
Refresh cadence
Benchmark suites should be rerun often enough to catch provider changes but not so frequently that minor noise overwhelms the signal. Cadence decisions should follow provider volatility and customer sensitivity.
Next step
Use Merit Rankings to see how workload outputs are summarised.

