TestRelic AI
Go to App
Cloud Platform

LLM Evaluations

Ingest, browse, and track LLM-evaluation runs in the TestRelic cloud platform through the same shared run UI as your tests.

TestRelic ingests LLM-evaluation runs alongside your browser and API tests. Eval runs come from the DeepEval SDK and render through the same shared run UI as regular tests — there is no separate, bespoke evaluation interface to learn.

Evaluationsshop-e2e DeepEval
nightly-eval — support-agent-v3Prompt-injection guard resists override
DeepEvalev_8a41
ci-bot·main·c91a4de·Dataset: support-qa-v2·2m 14s
67
Test cases
3
3 metrics
Cases passed
2
Cases failed
1
Total cost
$0.021
3/3

Prompt-injection guard resists override

safety
41
Pass rate
33%
1/3 metrics
Stability
38 · Volatile
Cost
$0.011
Latency
2.40s

Single-model verdict per case — for a prompts × models comparison across multiple generators, see Augur.

testrelic-deepeval SDK
TestRelic ingest
Run Detail
Repository Evaluations tab
Test Runs feed

What an eval run is

An eval run is a single execution of an LLM-evaluation suite, made up of one or more evaluation cases. Each case records whether it passed or failed against its metrics, the same way a test case records pass/fail/flaky status. Because eval runs share the run model with tests, they flow into the platform's existing views automatically.

Shared Run Detail and Session Workspace

Eval runs open in the standard Run Detail page and Session Workspace. From a run you can drill into per-case detail and review each case's Test History over time — exactly as you would for a browser or API test.

Evaluations workspace

DeepEval runs also get a dedicated Evaluations workspace focused on eval-specific analysis. At the top, an Evaluations KPI summary strip rolls up the headline numbers for the current view — pass rate, case counts, and stability at a glance — before you drill into individual runs and cases.

Repository Evaluations tab

A repository that contains eval data gets an Evaluations tab alongside its Test Cases and Test Runs tabs. The tab lists that repository's eval runs with per-repo eval stats and tags, so you can scan stability and pass rate and filter by tag without leaving the repository. Click through to per-case detail and per-case Test History. See Repositories for the full repository-detail layout.

Unified Test Runs feed

The org-wide Test Runs dashboard can include eval rows through an opt-in include evals toggle. When enabled, eval runs appear in the unified feed with proper eval naming, mixed in with your browser and API runs.

Eval Stability

Eval Stability is a metric surfaced for eval runs and repositories. It tracks how consistently eval cases pass across runs over time — analogous to flakiness and pass-rate stability for tests. Use it to spot evaluations whose outcomes drift or fluctuate between runs rather than holding steady.

Eval Stability is a qualitative indicator of consistency over a window of recent runs, not a single-run score. Like other health metrics, it updates as new eval runs are ingested.

Sending evals from the SDK

You get evals into TestRelic with the DeepEval Python SDK. Configure the SDK with your repository's API key and run your evaluations — the runs then appear in the views above.

See the DeepEval SDK for installation and configuration.

Text-to-image runs and multi-model comparison

Text-to-image visual-regression runs render through a different dedicated workspace, Augur, instead of the LLM-focused views above. When a run compares more than one generator model against the same prompts, Augur shows a prompts × models triage matrix, a per-prompt compare canvas, and prompt-phrase attribution — see Augur Multi-Model Comparison for the full workspace.

Viewing evals locally, without an upload

Studio's Evals pane reads DeepEval test-run JSON straight off your repo's disk — metric scores, category health, agent conversations, and tool calls — before or without ever uploading a run to this cloud view. Uploaded runs still land here as usual; the Studio pane is a faster local loop, not a replacement.

Was this page helpful?

On this page