Import externally-executed eval results.
const url = 'https://app.everruns.com/api/v1/evals/import';const options = { method: 'POST', headers: {'Content-Type': 'application/json'}, body: '{"evals":[{"cases":[{"description":"Agent fixes a failing unit test","error_message":"Session timed out","input":["Fix the failing test in src/lib.rs"],"input_tokens":1200,"latency_ms":8450,"metrics":"example","name":"fix-failing-test","output_tokens":850,"scores":[{"na":false,"pass":true,"reason":"Output contains expected text","scorer":"contains","value":1}],"status":"passed","target":{"model":"gpt-5.1","params":"example","provider":"openai"},"transcript":"example","turns":3}],"description":"Regression suite for the support agent","name":"Support agent regression","tags":["regression","nightly"]}],"source":{"metadata":"example","run_id":"run-2026-01-15-001","system":"mira","url":"https://ci.example.com/runs/42","version":"0.4.0"}}'};
try { const response = await fetch(url, options); const data = await response.json(); console.log(data);} catch (error) { console.error(error);}curl --request POST \ --url https://app.everruns.com/api/v1/evals/import \ --header 'Content-Type: application/json' \ --data '{ "evals": [ { "cases": [ { "description": "Agent fixes a failing unit test", "error_message": "Session timed out", "input": [ "Fix the failing test in src/lib.rs" ], "input_tokens": 1200, "latency_ms": 8450, "metrics": "example", "name": "fix-failing-test", "output_tokens": 850, "scores": [ { "na": false, "pass": true, "reason": "Output contains expected text", "scorer": "contains", "value": 1 } ], "status": "passed", "target": { "model": "gpt-5.1", "params": "example", "provider": "openai" }, "transcript": "example", "turns": 3 } ], "description": "Regression suite for the support agent", "name": "Support agent regression", "tags": [ "regression", "nightly" ] } ], "source": { "metadata": "example", "run_id": "run-2026-01-15-001", "system": "mira", "url": "https://ci.example.com/runs/42", "version": "0.4.0" } }'Request Bodyrequired
Section titled “Request Bodyrequired”A whole external run group: one external run, one entry per eval. Maps to
one everruns EvalRun per eval, all sharing source.run_id.
object
Evals and their case results in this run.
One eval’s worth of results within the run. The eval is upserted by name.
object
Case results for this eval.
One case result. The case is upserted by name (identity-only: everruns
never re-executes it).
object
Optional case description.
Example
Agent fixes a failing unit testError detail when the case errored.
Example
Session timed outDisplay-only input turns shown in the UI.
Example
[ "Fix the failing test in src/lib.rs"]Input tokens used.
Example
1200Execution time in milliseconds.
Example
8450Open-vocab metrics bag (cost_usd, cache/reasoning tokens, ttft, …).
Case name; the case is upserted by it.
Example
fix-failing-testOutput tokens used.
Example
850Named, attributed scores. Stored opaque; everruns does not re-grade.
A single named score from an external scorer.
object
Scorer was not applicable (excluded from aggregate).
Example
falseWhether the scorer passed.
Example
trueHuman-readable explanation of the score.
Example
Output contains expected textScorer name.
Example
containsScore value from 0.0 to 1.0.
Example
1Verdict for the case, trusted as reported.
Provider/model labels this result was produced against.
object
Model name.
Example
gpt-5.1Opaque provider parameters.
Provider name.
Example
openaiNormalized transcript (messages, tool calls, events, parts, files).
Number of agent turns taken.
Example
3Optional eval description.
Example
Regression suite for the support agentEval name; the eval is upserted by it.
Example
Support agent regressionFree-form tags for the eval.
Example
[ "regression", "nightly"]External system that produced the run.
object
Optional environment/labels (git commit, host, etc.).
Stable external run id: cross-eval group key + idempotency key.
Example
run-2026-01-15-001External system name, e.g. “mira”.
Example
miraLink back to the run in the external system.
Example
https://ci.example.com/runs/42Version of the external system.
Example
0.4.0Responses
Section titled “ Responses ”Success
Response wrapper for list endpoints.
All list endpoints return responses wrapped in a data field.
object
Array of items returned by the list operation.
An eval run: one execution of all/some cases.
object
Provenance for external runs: which system produced them, version, link
back, and any environment labels. None for internal runs. Open-vocab
JSON so new attribution fields need no schema change.
When the run finished.
When the run was created.
Only run cases matching these tags.
External identifier (evalrun_<32-hex>).
Model override for this run.
Case results (populated on detail view).
Result of a single case within a run.
object
Collected session file contents keyed by artifact name.
Case name (denormalized for display).
When the result was created.
Error message if errored.
The case this result is for.
External identifier (evalresult_<32-hex>).
Token usage.
Execution time in milliseconds.
External scorer metadata captured during deferred write-back.
Output tokens used.
Per-scorer results.
Session created for this case (browsable in UI).
Execution status of the case.
Session creation parameters (mirrors CreateSessionRequest).
object
Agent to work in this session.
Harness for the session. If omitted, org default harness is used.
Addressable harness name (alternative to harness_id).
Max LLM iterations per turn.
LLM model override.
System prompt override (prepended to agent prompt).
Reference to a deployed app.
object
Label-only target for externally-executed runs (e.g. imported from Mira).
Carries provider/model labels and opaque params instead of session setup:
external runs are ingested already-complete, so everruns never builds a
session from this. Mirrors a provider-agnostic (provider, model) pair.
object
Session creation parameters (mirrors CreateSessionRequest).
object
Agent to work in this session.
Harness for the session. If omitted, org default harness is used.
Addressable harness name (alternative to harness_id).
Max LLM iterations per turn.
LLM model override.
System prompt override (prepended to agent prompt).
Reference to a deployed app.
object
Label-only target for externally-executed runs (e.g. imported from Mira).
Carries provider/model labels and opaque params instead of session setup:
external runs are ingested already-complete, so everruns never builds a
session from this. Mirrors a provider-agnostic (provider, model) pair.
object
Turn count.
When the result was last updated.
Whether everruns executed this run (internal) or it was imported from
an external eval system (external).
When the run started executing.
Run lifecycle status.
Aggregate metrics (set on completion).
object
Mean case latency in milliseconds.
Mean score across cases, 0.0 to 1.0.
Mean agent turns per case.
Number of cases that errored.
Number of cases that failed.
Fraction of cases that passed, 0.0 to 1.0.
Number of cases that passed.
Total number of cases in the run.
Total input tokens across cases.
Total output tokens across cases.
Session creation parameters (mirrors CreateSessionRequest).
object
Agent to work in this session.
Harness for the session. If omitted, org default harness is used.
Addressable harness name (alternative to harness_id).
Max LLM iterations per turn.
LLM model override.
System prompt override (prepended to agent prompt).
Reference to a deployed app.
object
Label-only target for externally-executed runs (e.g. imported from Mira).
Carries provider/model labels and opaque params instead of session setup:
external runs are ingested already-complete, so everruns never builds a
session from this. Mirrors a provider-agnostic (provider, model) pair.
object
What triggered this run.
When the run was last updated.
Example
{ "data": [ { "completed_at": "2026-01-15T10:30:00Z", "created_at": "2026-01-15T10:30:00Z", "filter_tags": [ "regression", "nightly" ], "id": "evalrun_01933b5a000070008000000000000001", "model_override": "gpt-5.1", "results": [ { "artifacts": { "patch": "diff --git a/src/lib.rs b/src/lib.rs" }, "case_name": "fix-failing-test", "created_at": "2026-01-15T10:30:00Z", "error_message": "Session timed out", "eval_case_id": "evalcase_01933b5a000070008000000000000001", "id": "evalresult_01933b5a000070008000000000000001", "input_tokens": 1200, "latency_ms": 8450, "output_tokens": 850, "session_id": "session_01933b5a00007000800000000000001", "status": "pending", "target": { "type": "session" }, "target_snapshot": { "type": "session" }, "turns": 3, "updated_at": "2026-01-15T10:30:00Z" } ], "source": "internal", "started_at": "2026-01-15T10:30:00Z", "status": "pending", "summary": { "avg_latency_ms": 8450, "avg_score": 0.85, "avg_turns": 3.2, "errored": 1, "failed": 2, "pass_rate": 0.75, "passed": 9, "total": 12, "total_input_tokens": 14400, "total_output_tokens": 10200 }, "target": { "type": "session" }, "triggered_by": "manual", "updated_at": "2026-01-15T10:30:00Z" } ]}