Skip to content
Everruns Cloud is open in early access. Run agents without operating the platform.

Import externally-executed eval results.

POST
/v1/evals/import
curl --request POST \
--url https://app.everruns.com/api/v1/evals/import \
--header 'Content-Type: application/json' \
--data '{ "evals": [ { "cases": [ { "description": "Agent fixes a failing unit test", "error_message": "Session timed out", "input": [ "Fix the failing test in src/lib.rs" ], "input_tokens": 1200, "latency_ms": 8450, "metrics": "example", "name": "fix-failing-test", "output_tokens": 850, "scores": [ { "na": false, "pass": true, "reason": "Output contains expected text", "scorer": "contains", "value": 1 } ], "status": "passed", "target": { "model": "gpt-5.1", "params": "example", "provider": "openai" }, "transcript": "example", "turns": 3 } ], "description": "Regression suite for the support agent", "name": "Support agent regression", "tags": [ "regression", "nightly" ] } ], "source": { "metadata": "example", "run_id": "run-2026-01-15-001", "system": "mira", "url": "https://ci.example.com/runs/42", "version": "0.4.0" } }'
Media typeapplication/json

A whole external run group: one external run, one entry per eval. Maps to one everruns EvalRun per eval, all sharing source.run_id.

object
evals
required

Evals and their case results in this run.

Array<object>

One eval’s worth of results within the run. The eval is upserted by name.

object
cases
required

Case results for this eval.

Array<object>

One case result. The case is upserted by name (identity-only: everruns never re-executes it).

object
description

Optional case description.

string | null
Example
Agent fixes a failing unit test
error_message

Error detail when the case errored.

string | null
Example
Session timed out
input

Display-only input turns shown in the UI.

Array<string>
Example
[
"Fix the failing test in src/lib.rs"
]
input_tokens

Input tokens used.

integer | null format: int64
Example
1200
latency_ms

Execution time in milliseconds.

integer | null format: int64
Example
8450
metrics

Open-vocab metrics bag (cost_usd, cache/reasoning tokens, ttft, …).

name
required

Case name; the case is upserted by it.

string
Example
fix-failing-test
output_tokens

Output tokens used.

integer | null format: int64
Example
850
scores

Named, attributed scores. Stored opaque; everruns does not re-grade.

Array<object>

A single named score from an external scorer.

object
na

Scorer was not applicable (excluded from aggregate).

boolean
Example
false
pass
required

Whether the scorer passed.

boolean
Example
true
reason

Human-readable explanation of the score.

string
Example
Output contains expected text
scorer
required

Scorer name.

string
Example
contains
value
required

Score value from 0.0 to 1.0.

number format: double
Example
1
status
required

Verdict for the case, trusted as reported.

string
Allowed values: passed failed errored timeout skipped
target
required

Provider/model labels this result was produced against.

object
model
required

Model name.

string
Example
gpt-5.1
params

Opaque provider parameters.

provider
required

Provider name.

string
Example
openai
transcript

Normalized transcript (messages, tool calls, events, parts, files).

turns

Number of agent turns taken.

integer | null format: int32
Example
3
description

Optional eval description.

string | null
Example
Regression suite for the support agent
name
required

Eval name; the eval is upserted by it.

string
Example
Support agent regression
tags

Free-form tags for the eval.

Array<string>
Example
[
"regression",
"nightly"
]
source
required

External system that produced the run.

object
metadata

Optional environment/labels (git commit, host, etc.).

run_id
required

Stable external run id: cross-eval group key + idempotency key.

string
Example
run-2026-01-15-001
system
required

External system name, e.g. “mira”.

string
Example
mira
url

Link back to the run in the external system.

string | null
Example
https://ci.example.com/runs/42
version

Version of the external system.

string | null
Example
0.4.0

Success

Media typeapplication/json

Response wrapper for list endpoints. All list endpoints return responses wrapped in a data field.

object
data
required

Array of items returned by the list operation.

Array<object>

An eval run: one execution of all/some cases.

object
attribution

Provenance for external runs: which system produced them, version, link back, and any environment labels. None for internal runs. Open-vocab JSON so new attribution fields need no schema change.

completed_at

When the run finished.

string | null format: date-time
created_at
required

When the run was created.

string format: date-time
filter_tags

Only run cases matching these tags.

Array<string> | null
id
required

External identifier (evalrun_<32-hex>).

string
model_override

Model override for this run.

string | null
results

Case results (populated on detail view).

Array<object>

Result of a single case within a run.

object
artifacts

Collected session file contents keyed by artifact name.

object | null
case_name

Case name (denormalized for display).

string | null
created_at
required

When the result was created.

string format: date-time
error_message

Error message if errored.

string | null
eval_case_id
required

The case this result is for.

string
id
required

External identifier (evalresult_<32-hex>).

string
input_tokens

Token usage.

integer | null format: int64
latency_ms

Execution time in milliseconds.

integer | null format: int64
metadata

External scorer metadata captured during deferred write-back.

output_tokens

Output tokens used.

integer | null format: int64
scores

Per-scorer results.

session_id

Session created for this case (browsable in UI).

string | null
status
required

Execution status of the case.

string
Allowed values: pending running passed failed errored timeout skipped
target
One of:
One of:

Session creation parameters (mirrors CreateSessionRequest).

object
agent_id

Agent to work in this session.

string | null
harness_id

Harness for the session. If omitted, org default harness is used.

string | null
harness_name

Addressable harness name (alternative to harness_id).

string | null
max_iterations

Max LLM iterations per turn.

integer | null
model_id

LLM model override.

string | null
system_prompt

System prompt override (prepended to agent prompt).

string | null
type
required
string
Allowed values: session
target_snapshot
One of:
One of:

Session creation parameters (mirrors CreateSessionRequest).

object
agent_id

Agent to work in this session.

string | null
harness_id

Harness for the session. If omitted, org default harness is used.

string | null
harness_name

Addressable harness name (alternative to harness_id).

string | null
max_iterations

Max LLM iterations per turn.

integer | null
model_id

LLM model override.

string | null
system_prompt

System prompt override (prepended to agent prompt).

string | null
type
required
string
Allowed values: session
turns

Turn count.

integer | null format: int32
updated_at
required

When the result was last updated.

string format: date-time
source

Whether everruns executed this run (internal) or it was imported from an external eval system (external).

string
Allowed values: internal external
started_at

When the run started executing.

string | null format: date-time
status
required

Run lifecycle status.

string
Allowed values: pending running completed failed cancelled
summary
One of:

Aggregate metrics (set on completion).

object
avg_latency_ms
required

Mean case latency in milliseconds.

integer format: int64
avg_score
required

Mean score across cases, 0.0 to 1.0.

number format: double
avg_turns
required

Mean agent turns per case.

number format: double
errored
required

Number of cases that errored.

integer format: int32
failed
required

Number of cases that failed.

integer format: int32
pass_rate
required

Fraction of cases that passed, 0.0 to 1.0.

number format: double
passed
required

Number of cases that passed.

integer format: int32
total
required

Total number of cases in the run.

integer format: int32
total_input_tokens
required

Total input tokens across cases.

integer format: int64
total_output_tokens
required

Total output tokens across cases.

integer format: int64
target
One of:
One of:

Session creation parameters (mirrors CreateSessionRequest).

object
agent_id

Agent to work in this session.

string | null
harness_id

Harness for the session. If omitted, org default harness is used.

string | null
harness_name

Addressable harness name (alternative to harness_id).

string | null
max_iterations

Max LLM iterations per turn.

integer | null
model_id

LLM model override.

string | null
system_prompt

System prompt override (prepended to agent prompt).

string | null
type
required
string
Allowed values: session
triggered_by
required

What triggered this run.

string
updated_at
required

When the run was last updated.

string format: date-time
Example
{
"data": [
{
"completed_at": "2026-01-15T10:30:00Z",
"created_at": "2026-01-15T10:30:00Z",
"filter_tags": [
"regression",
"nightly"
],
"id": "evalrun_01933b5a000070008000000000000001",
"model_override": "gpt-5.1",
"results": [
{
"artifacts": {
"patch": "diff --git a/src/lib.rs b/src/lib.rs"
},
"case_name": "fix-failing-test",
"created_at": "2026-01-15T10:30:00Z",
"error_message": "Session timed out",
"eval_case_id": "evalcase_01933b5a000070008000000000000001",
"id": "evalresult_01933b5a000070008000000000000001",
"input_tokens": 1200,
"latency_ms": 8450,
"output_tokens": 850,
"session_id": "session_01933b5a00007000800000000000001",
"status": "pending",
"target": {
"type": "session"
},
"target_snapshot": {
"type": "session"
},
"turns": 3,
"updated_at": "2026-01-15T10:30:00Z"
}
],
"source": "internal",
"started_at": "2026-01-15T10:30:00Z",
"status": "pending",
"summary": {
"avg_latency_ms": 8450,
"avg_score": 0.85,
"avg_turns": 3.2,
"errored": 1,
"failed": 2,
"pass_rate": 0.75,
"passed": 9,
"total": 12,
"total_input_tokens": 14400,
"total_output_tokens": 10200
},
"target": {
"type": "session"
},
"triggered_by": "manual",
"updated_at": "2026-01-15T10:30:00Z"
}
]
}