AI EVALUATION · TESTING, EVALS & RED TEAMING

Prove what your AI does.
And what it can be made to do.

Any dataset becomes an evaluation. Twenty built-in evaluators score your assistants, models and agents — offline before release, online in production, and under adversarial attack.

Most AI systems reach production on the strength of a demonstration and a few dozen manual checks. That is enough to build confidence. It is not enough to answer the questions that arrive later: has quality regressed since the model changed? Can this assistant be talked into something it must never do? And if a regulator asks, what exactly can you show them?

Kosmoy answers those questions with evidence. You define what is tested, how it is scored and against which data; you run it as often as you like; and every run leaves a scorecard, a case-by-case record and a report you can hand to someone else. It is the discipline Gartner tracks as AI evaluation and observability — and in Kosmoy it lives on the same platform as the gateway, guardrails and monitoring that run your AI every day.


Any evaluation, given a dataset.

There is no fixed list of supported tests. Bring a public benchmark, your production traffic, agent traces or a generated adversarial corpus; choose the evaluators and what counts as passing; run it. The same twenty evaluators mix freely across every type, each with its own pass threshold.

Benchmark evaluation

MMLU, GSM8K, TruthfulQA or your own graded question set — accuracy against known answers, scored deterministically.

Agentic evaluation

Did the agent choose the right tools, pass the right arguments, use the results, and finish the task efficiently?

RAG evaluation

Groundedness against retrieved sources — is the system answering from its documents or from imagination?

Quality evaluation

Relevance and coherence of responses on any set of prompts, scored by an AI judge.

Safety evaluation

Toxicity, bias, PII leakage and prompt-injection resistance — on adversarial prompts or everyday traffic.

Red teaming

A labelled adversarial dataset, or seed prompts for escalating attacks. Whether the system can be made to break its rules.


Three ways to run it.

An evaluation is a reusable definition: a subject, a method and a dataset. Running it produces a numbered run with its own results, so you can compare today’s behaviour with last month’s — the second run costs one click.

Offline

Drives your system over a test dataset now and scores the responses.

Catch regressions before release. Compare models, prompts and configurations.

Online

Scores real interactions that already happened, reconstructed from production traffic.

Monitor live quality without synthetic data — evaluation joined to observability.

Red Team

Attacks your system with adversarial prompts and judges how it holds up.

Prove safety and resilience. Satisfy security review and audit.

The subject can be a Kosmoy assistant, driven end to end with its tools, retrieval and guardrails; a model accessed through one of your gateways; or a registered external agent running outside Kosmoy — so one evaluation programme covers systems across your estate, not only the ones built in-house.


Evaluation wired into the platform that runs your AI.

Test what users actually experience

Evaluations drive the complete assistant — prompt, tools, retrieval and guardrails together — not a model in isolation. Attack the bare gateway model too, and the difference quantifies exactly what your safety layer is worth.

Findings become enforcement

A red-team finding is remediation advice next to the guardrail that fixes it. Harden the policy, re-run the same evaluation, and the numbered runs show the improvement — the loop closes on one platform.

Online evaluation reads real traffic

Production interactions captured by AI Monitoring are reconstructed into cases and scored for relevance, coherence, groundedness and safety. No synthetic data, no instrumentation project.

Every run is audit evidence

Scorecards, case-level records, judge reasoning and PDF reports land in the same evidence trail as your EU AI Act risk classification and NIST AI RMF bundle — proof of testing, on demand.


The category Gartner named

Built for the AI evaluation and observability market.

In February 2026 Gartner published its first Market Guide for AI Evaluation and Observability Platforms, defining a category for tooling that automates evaluations to benchmark AI outputs against quality, fairness and accuracy, and feeds production observability data back into those evaluations to improve reliability. Gartner projects adoption by software engineering teams rising from 18% in 2025 to 60% by 2028 — the fastest-moving tooling category in enterprise AI.

A note on scope. Gartner uses “observability” narrowly here — the evaluation, tracing and quality signals that test how a model behaves. Kosmoy’s AI Monitoring & Observability layer is deliberately broader: it also covers AI FinOps, cost attribution and operational monitoring, which sit outside Gartner’s category. It is this module — evaluation and red teaming — together with the quality observability it draws on, that maps to what Gartner describes.

And it is built for exactly that: evals automated across six dimensions, model-agnostic across every provider behind your gateways (a requirement Gartner names to avoid lock-in), closing the feedback loop the guide describes as production traffic becomes online evaluation and red-team datasets. Where Kosmoy goes further than a standalone platform is enforcement — a finding becomes an enforced guardrail on the same platform, and every run becomes EU AI Act and NIST AI RMF evidence.

Source: Gartner, Market Guide for AI Evaluation and Observability Platforms (G00 doc 7387730), 2 February 2026. Kosmoy is not named in the Market Guide; this describes how the module maps to the capabilities Gartner defines for the category. GARTNER is a registered trademark of Gartner, Inc. and/or its affiliates.


Module questions, answered straight.

What is an AI evaluation platform?

An AI evaluation platform tests what AI systems actually do — before release and in production — instead of relying on demos and manual spot checks. It runs models, assistants and agents over datasets, scores the responses against defined criteria (accuracy, groundedness, tool use, safety), and keeps a run-by-run record so quality regressions and new risks are caught early. Analysts group these capabilities with observability: Gartner covers the space in its Market Guide for AI Evaluation and Observability Platforms.

What can Kosmoy evaluate?

Anything you can express as a dataset: public benchmarks such as MMLU, GSM8K or TruthfulQA, agentic tool-use tasks, RAG question sets with retrieved sources, your own production traffic, or a generated adversarial corpus. The subject can be a Kosmoy assistant driven end to end with its tools, retrieval and guardrails, a model accessed through one of your gateways, or a registered external agent running outside Kosmoy.

How is Kosmoy different from standalone LLM evaluation tools?

Standalone evaluation tools score AI systems; Kosmoy also runs them. Evaluation shares the platform with the AI Gateway, guardrails, the AI inventory and the compliance dossier — so a red-team finding becomes a guardrail change on the same platform, online evaluation reads real gateway traffic, and every scorecard lands in the evidence trail auditors ask for. The loop from test to enforcement is closed in one place, in your own Kubernetes.

Which evaluators are built in?

Twenty evaluators across six dimensions: task completion (task completion, task adherence, intent resolution, navigation efficiency), tool use (call accuracy, call success, selection, input accuracy, output utilization), quality (relevance, coherence), groundedness, reference match (exact, multiple choice, numeric) and safety (injection resistance, toxicity, bias, PII leakage, attack success). Deterministic scorers run free; AI judges handle the qualities rules cannot capture. Each evaluator has its own pass threshold.

Does Kosmoy support online evaluation of production traffic?

Yes. Online evaluation reconstructs cases from production traffic for a chosen assistant and scores them with reference-free evaluators — relevance, coherence, groundedness, tool use and the safety set. It answers a different question from offline testing: not “does it pass our suite” but “what is it doing right now, for real users”.

How does evaluation help with EU AI Act compliance?

The EU AI Act expects providers of high-risk AI systems to test for accuracy and robustness (Article 15) and to run risk management across the lifecycle (Article 9). Kosmoy evaluation runs produce exactly that evidence: named, versioned runs with scorecards, case-level records, judge reasoning and PDF reports — generated on the same platform that classifies the system's risk tier and enforces its guardrails.

Is Kosmoy in the Gartner Market Guide for AI Evaluation and Observability Platforms?

Kosmoy is not named in Gartner's Market Guide for AI Evaluation and Observability Platforms (published 2 February 2026). We reference the guide because it defines the category this module is built for — automated evals that benchmark quality, fairness and accuracy, plus observability data fed back into those evals. Kosmoy's module maps to those capabilities and adds what standalone platforms do not: the evaluation and red-teaming results feed the same gateway and guardrails that enforce the fix, and produce EU AI Act, ISO/IEC 42001 and NIST AI RMF evidence.

See an evaluation run end to end.

From dataset to scorecard to red-team report — on your assistants, your models, your rules.