List eval runs.
const url = 'https://app.everruns.com/api/v1/evals/example/runs';const options = {method: 'GET'};
try { const response = await fetch(url, options); const data = await response.json(); console.log(data);} catch (error) { console.error(error);}curl --request GET \ --url https://app.everruns.com/api/v1/evals/example/runsParameters
Section titled “ Parameters ”Path Parameters
Section titled “Path Parameters”Prefixed public identifier
Responses
Section titled “ Responses ”Success
Response wrapper for list endpoints.
All list endpoints return responses wrapped in a data field.
object
Array of items returned by the list operation.
An eval run: one execution of all/some cases.
object
Provenance for external runs: which system produced them, version, link
back, and any environment labels. None for internal runs. Open-vocab
JSON so new attribution fields need no schema change.
When the run finished.
When the run was created.
Only run cases matching these tags.
External identifier (evalrun_<32-hex>).
Model override for this run.
Case results (populated on detail view).
Result of a single case within a run.
object
Collected session file contents keyed by artifact name.
Case name (denormalized for display).
When the result was created.
Error message if errored.
The case this result is for.
External identifier (evalresult_<32-hex>).
Token usage.
Execution time in milliseconds.
External scorer metadata captured during deferred write-back.
Output tokens used.
Per-scorer results.
Session created for this case (browsable in UI).
Execution status of the case.
Session creation parameters (mirrors CreateSessionRequest).
object
Agent to work in this session.
Harness for the session. If omitted, org default harness is used.
Addressable harness name (alternative to harness_id).
Max LLM iterations per turn.
LLM model override.
System prompt override (prepended to agent prompt).
Reference to a deployed app.
object
Label-only target for externally-executed runs (e.g. imported from Mira).
Carries provider/model labels and opaque params instead of session setup:
external runs are ingested already-complete, so everruns never builds a
session from this. Mirrors a provider-agnostic (provider, model) pair.
object
Session creation parameters (mirrors CreateSessionRequest).
object
Agent to work in this session.
Harness for the session. If omitted, org default harness is used.
Addressable harness name (alternative to harness_id).
Max LLM iterations per turn.
LLM model override.
System prompt override (prepended to agent prompt).
Reference to a deployed app.
object
Label-only target for externally-executed runs (e.g. imported from Mira).
Carries provider/model labels and opaque params instead of session setup:
external runs are ingested already-complete, so everruns never builds a
session from this. Mirrors a provider-agnostic (provider, model) pair.
object
Turn count.
When the result was last updated.
Whether everruns executed this run (internal) or it was imported from
an external eval system (external).
When the run started executing.
Run lifecycle status.
Aggregate metrics (set on completion).
object
Mean case latency in milliseconds.
Mean score across cases, 0.0 to 1.0.
Mean agent turns per case.
Number of cases that errored.
Number of cases that failed.
Fraction of cases that passed, 0.0 to 1.0.
Number of cases that passed.
Total number of cases in the run.
Total input tokens across cases.
Total output tokens across cases.
Session creation parameters (mirrors CreateSessionRequest).
object
Agent to work in this session.
Harness for the session. If omitted, org default harness is used.
Addressable harness name (alternative to harness_id).
Max LLM iterations per turn.
LLM model override.
System prompt override (prepended to agent prompt).
Reference to a deployed app.
object
Label-only target for externally-executed runs (e.g. imported from Mira).
Carries provider/model labels and opaque params instead of session setup:
external runs are ingested already-complete, so everruns never builds a
session from this. Mirrors a provider-agnostic (provider, model) pair.
object
What triggered this run.
When the run was last updated.
Example
{ "data": [ { "completed_at": "2026-01-15T10:30:00Z", "created_at": "2026-01-15T10:30:00Z", "filter_tags": [ "regression", "nightly" ], "id": "evalrun_01933b5a000070008000000000000001", "model_override": "gpt-5.1", "results": [ { "artifacts": { "patch": "diff --git a/src/lib.rs b/src/lib.rs" }, "case_name": "fix-failing-test", "created_at": "2026-01-15T10:30:00Z", "error_message": "Session timed out", "eval_case_id": "evalcase_01933b5a000070008000000000000001", "id": "evalresult_01933b5a000070008000000000000001", "input_tokens": 1200, "latency_ms": 8450, "output_tokens": 850, "session_id": "session_01933b5a00007000800000000000001", "status": "pending", "target": { "type": "session" }, "target_snapshot": { "type": "session" }, "turns": 3, "updated_at": "2026-01-15T10:30:00Z" } ], "source": "internal", "started_at": "2026-01-15T10:30:00Z", "status": "pending", "summary": { "avg_latency_ms": 8450, "avg_score": 0.85, "avg_turns": 3.2, "errored": 1, "failed": 2, "pass_rate": 0.75, "passed": 9, "total": 12, "total_input_tokens": 14400, "total_output_tokens": 10200 }, "target": { "type": "session" }, "triggered_by": "manual", "updated_at": "2026-01-15T10:30:00Z" } ]}