Skip to content
Everruns Cloud is open in early access. Run agents without operating the platform.

Create an eval case.

POST
/v1/evals/{eval_id}/cases
curl --request POST \
--url https://app.everruns.com/api/v1/evals/example/cases \
--header 'Content-Type: application/json' \
--data '{ "artifacts": [ { "name": "patch", "path": "/workspace/fix.patch" } ], "conversation": [ { "content": "Fix the failing test in src/lib.rs" } ], "description": "Agent fixes a failing unit test", "max_turns": 10, "name": "fix-failing-test", "position": 0, "post": [ { "content": "Run the tests again and report the result" } ], "scorers": [ { "text": "tests pass", "type": "contains" } ], "tags": [ "regression", "nightly" ], "target": { "agent_id": "example", "harness_id": "example", "harness_name": "example", "max_iterations": 1, "model_id": "example", "system_prompt": "example", "type": "session" }, "timeout_seconds": 120 }'
eval_id
required
string

Prefixed public identifier

Media typeapplication/json

Request to create an eval case

object
artifacts

Session files to capture after scoring completes.

Array<object> | null

Named session file to collect after an eval case completes.

object
name
required

Export key for this artifact (for example patch or log).

string
Example
patch
path
required

Absolute path in the session filesystem.

string
Example
/workspace/fix.patch
Example
[
{
"name": "patch",
"path": "/workspace/fix.patch"
}
]
conversation
required

Input messages sent to the agent sequentially.

Array<object>

A message to send to the agent during an eval case.

object
content
required

The text content to send.

string
Example
Fix the failing test in src/lib.rs
Example
[
{
"content": "Fix the failing test in src/lib.rs"
}
]
description

Human-readable description. Safe to render in user-facing messages.

string | null
Example
Agent fixes a failing unit test
max_turns

Maximum agent turns before the case stops.

integer | null format: int32
Example
10
name
required

Human-readable name. Safe to render in user-facing messages.

string
Example
fix-failing-test
position

Display order within the eval.

integer | null format: int32
Example
0
post

Verification messages sent after conversation completes and session idles.

Array<object> | null

A message to send to the agent during an eval case.

object
content
required

The text content to send.

string
Example
Fix the failing test in src/lib.rs
Example
[
{
"content": "Run the tests again and report the result"
}
]
scorers
required

Scoring rules applied to the case output.

Array
One of:

Final assistant message contains substring.

object
text
required
string
type
required
string
Allowed values: contains
weight
number format: double
Example
[
{
"text": "tests pass",
"type": "contains"
}
]
tags

Free-form tags attached to this resource.

Array<string> | null
Example
[
"regression",
"nightly"
]
target
One of:
One of:

Session creation parameters (mirrors CreateSessionRequest).

object
agent_id

Agent to work in this session.

string | null
harness_id

Harness for the session. If omitted, org default harness is used.

string | null
harness_name

Addressable harness name (alternative to harness_id).

string | null
max_iterations

Max LLM iterations per turn.

integer | null
model_id

LLM model override.

string | null
system_prompt

System prompt override (prepended to agent prompt).

string | null
type
required
string
Allowed values: session
timeout_seconds

Per-case timeout in seconds.

integer | null format: int32
Example
120

Created

Media typeapplication/json

A single test case within an eval.

object
artifacts

Session files to collect after scoring completes.

Array<object> | null

Named session file to collect after an eval case completes.

object
name
required

Export key for this artifact (for example patch or log).

string
path
required

Absolute path in the session filesystem.

string
conversation
required

Input messages sent sequentially.

Array<object>

A message to send to the agent during an eval case.

object
content
required

The text content to send.

string
created_at
required

When the case was created.

string format: date-time
description

Optional description of what the case checks.

string | null
id
required

External identifier (evalcase_<32-hex>).

string
max_turns

Max agent turns (default: 10).

integer | null format: int32
name
required

Case name.

string
position
required

Display order.

integer format: int32
post

Verification messages sent after conversation completes and session idles. Scorers run after post messages complete (not after conversation).

Array<object> | null

A message to send to the agent during an eval case.

object
content
required

The text content to send.

string
scorers
required

Scoring rules.

Array
One of:

Final assistant message contains substring.

object
text
required
string
type
required
string
Allowed values: contains
weight
number format: double
tags

Free-form tags for filtering runs.

Array<string>
target
One of:
One of:

Session creation parameters (mirrors CreateSessionRequest).

object
agent_id

Agent to work in this session.

string | null
harness_id

Harness for the session. If omitted, org default harness is used.

string | null
harness_name

Addressable harness name (alternative to harness_id).

string | null
max_iterations

Max LLM iterations per turn.

integer | null
model_id

LLM model override.

string | null
system_prompt

System prompt override (prepended to agent prompt).

string | null
type
required
string
Allowed values: session
timeout_seconds

Per-case timeout in seconds (default: 120).

integer | null format: int32
updated_at
required

When the case was last updated.

string format: date-time
Example
{
"artifacts": [
{
"name": "patch",
"path": "/workspace/fix.patch"
}
],
"conversation": [
{
"content": "Fix the failing test in src/lib.rs"
}
],
"created_at": "2026-01-15T10:30:00Z",
"description": "Agent fixes a failing unit test",
"id": "evalcase_01933b5a000070008000000000000001",
"max_turns": 10,
"name": "fix-failing-test",
"position": 0,
"post": [
{
"content": "Run the tests again and report the result"
}
],
"scorers": [
{
"text": "tests pass",
"type": "contains"
}
],
"tags": [
"regression",
"nightly"
],
"target": {
"type": "session"
},
"timeout_seconds": 120,
"updated_at": "2026-01-15T10:30:00Z"
}