Skip to content
Everruns Cloud is open in early access. Run agents without operating the platform.

Update scores for one eval result.

PATCH
/v1/evals/{eval_id}/runs/{run_id}/results/{result_id}/scores
curl --request PATCH \
--url https://app.everruns.com/api/v1/evals/example/runs/example/results/example/scores \
--header 'Content-Type: application/json' \
--data '{ "metadata": "example", "scores": [ { "pass": true, "reason": "Output contains expected text", "value": 1 } ], "status": "passed" }'
eval_id
required
string

Prefixed public identifier

run_id
required
string

Prefixed public identifier

result_id
required
string

Prefixed public identifier

Media typeapplication/json

Request to write external scores to one eval case result.

object
metadata

Free-form metadata attached to this resource.

scores
required

Externally computed scores to store on the result.

Array<object>

Result from a single scorer evaluation.

object
pass
required

Whether this scorer passed.

boolean
Example
true
reason
required

Human-readable explanation.

string
Example
Output contains expected text
value
required

Score value 0.0–1.0.

number format: double
Example
1
status
One of:

Current lifecycle status.

string
Allowed values: passed failed errored

Success

Media typeapplication/json

Result of a single case within a run.

object
artifacts

Collected session file contents keyed by artifact name.

object | null
case_name

Case name (denormalized for display).

string | null
created_at
required

When the result was created.

string format: date-time
error_message

Error message if errored.

string | null
eval_case_id
required

The case this result is for.

string
id
required

External identifier (evalresult_<32-hex>).

string
input_tokens

Token usage.

integer | null format: int64
latency_ms

Execution time in milliseconds.

integer | null format: int64
metadata

External scorer metadata captured during deferred write-back.

output_tokens

Output tokens used.

integer | null format: int64
scores

Per-scorer results.

session_id

Session created for this case (browsable in UI).

string | null
status
required

Execution status of the case.

string
Allowed values: pending running passed failed errored timeout skipped
target
One of:
One of:

Session creation parameters (mirrors CreateSessionRequest).

object
agent_id

Agent to work in this session.

string | null
harness_id

Harness for the session. If omitted, org default harness is used.

string | null
harness_name

Addressable harness name (alternative to harness_id).

string | null
max_iterations

Max LLM iterations per turn.

integer | null
model_id

LLM model override.

string | null
system_prompt

System prompt override (prepended to agent prompt).

string | null
type
required
string
Allowed values: session
target_snapshot
One of:
One of:

Session creation parameters (mirrors CreateSessionRequest).

object
agent_id

Agent to work in this session.

string | null
harness_id

Harness for the session. If omitted, org default harness is used.

string | null
harness_name

Addressable harness name (alternative to harness_id).

string | null
max_iterations

Max LLM iterations per turn.

integer | null
model_id

LLM model override.

string | null
system_prompt

System prompt override (prepended to agent prompt).

string | null
type
required
string
Allowed values: session
turns

Turn count.

integer | null format: int32
updated_at
required

When the result was last updated.

string format: date-time
Example
{
"artifacts": {
"patch": "diff --git a/src/lib.rs b/src/lib.rs"
},
"case_name": "fix-failing-test",
"created_at": "2026-01-15T10:30:00Z",
"error_message": "Session timed out",
"eval_case_id": "evalcase_01933b5a000070008000000000000001",
"id": "evalresult_01933b5a000070008000000000000001",
"input_tokens": 1200,
"latency_ms": 8450,
"output_tokens": 850,
"session_id": "session_01933b5a00007000800000000000001",
"status": "pending",
"target": {
"type": "session"
},
"target_snapshot": {
"type": "session"
},
"turns": 3,
"updated_at": "2026-01-15T10:30:00Z"
}

Eval result not found