Create an eval case.
const url = 'https://app.everruns.com/api/v1/evals/example/cases';const options = { method: 'POST', headers: {'Content-Type': 'application/json'}, body: '{"artifacts":[{"name":"patch","path":"/workspace/fix.patch"}],"conversation":[{"content":"Fix the failing test in src/lib.rs"}],"description":"Agent fixes a failing unit test","max_turns":10,"name":"fix-failing-test","position":0,"post":[{"content":"Run the tests again and report the result"}],"scorers":[{"text":"tests pass","type":"contains"}],"tags":["regression","nightly"],"target":{"agent_id":"example","harness_id":"example","harness_name":"example","max_iterations":1,"model_id":"example","system_prompt":"example","type":"session"},"timeout_seconds":120}'};
try { const response = await fetch(url, options); const data = await response.json(); console.log(data);} catch (error) { console.error(error);}curl --request POST \ --url https://app.everruns.com/api/v1/evals/example/cases \ --header 'Content-Type: application/json' \ --data '{ "artifacts": [ { "name": "patch", "path": "/workspace/fix.patch" } ], "conversation": [ { "content": "Fix the failing test in src/lib.rs" } ], "description": "Agent fixes a failing unit test", "max_turns": 10, "name": "fix-failing-test", "position": 0, "post": [ { "content": "Run the tests again and report the result" } ], "scorers": [ { "text": "tests pass", "type": "contains" } ], "tags": [ "regression", "nightly" ], "target": { "agent_id": "example", "harness_id": "example", "harness_name": "example", "max_iterations": 1, "model_id": "example", "system_prompt": "example", "type": "session" }, "timeout_seconds": 120 }'Parameters
Section titled “ Parameters ”Path Parameters
Section titled “Path Parameters”Prefixed public identifier
Request Bodyrequired
Section titled “Request Bodyrequired”Request to create an eval case
object
Session files to capture after scoring completes.
Named session file to collect after an eval case completes.
object
Export key for this artifact (for example patch or log).
Example
patchAbsolute path in the session filesystem.
Example
/workspace/fix.patchExample
[ { "name": "patch", "path": "/workspace/fix.patch" }]Input messages sent to the agent sequentially.
A message to send to the agent during an eval case.
object
The text content to send.
Example
Fix the failing test in src/lib.rsExample
[ { "content": "Fix the failing test in src/lib.rs" }]Human-readable description. Safe to render in user-facing messages.
Example
Agent fixes a failing unit testMaximum agent turns before the case stops.
Example
10Human-readable name. Safe to render in user-facing messages.
Example
fix-failing-testDisplay order within the eval.
Example
0Verification messages sent after conversation completes and session idles.
A message to send to the agent during an eval case.
object
The text content to send.
Example
Fix the failing test in src/lib.rsExample
[ { "content": "Run the tests again and report the result" }]Scoring rules applied to the case output.
Final assistant message contains substring.
object
Final assistant message does NOT contain substring.
object
Final assistant message matches regex pattern.
object
Agent called named tool at least min times.
object
Agent did NOT call named tool.
object
Total tool calls within range.
object
Completed within N turns.
object
Session filesystem file contains substring.
object
Final assistant message parses as JSON matching schema.
object
Citation faithfulness: the answer’s citations
must cover the claim and be verified as supported. Scored from the
TextAnnotations on the final message — pair with the
citation_verification capability so verdicts are present.
object
Minimum number of citations the answer must carry.
Example
1Minimum fraction of citations verified entailed to pass.
Example
0.8Relative weight of this scorer in the case’s weighted average.
Example
1Citation faithfulness judged by an LLM: each cited claim/source pair is
graded by a model, so the eval works even without the
citation_verification capability.
object
Minimum judged score [0,1] to pass.
Example
0.8Rubric override; a citation-faithfulness rubric is used when absent.
Example
Score the fraction of cited claims supported by their source.Relative weight of this scorer in the case’s weighted average.
Example
1Example
[ { "text": "tests pass", "type": "contains" }]Free-form tags attached to this resource.
Example
[ "regression", "nightly"]Session creation parameters (mirrors CreateSessionRequest).
object
Agent to work in this session.
Harness for the session. If omitted, org default harness is used.
Addressable harness name (alternative to harness_id).
Max LLM iterations per turn.
LLM model override.
System prompt override (prepended to agent prompt).
Reference to a deployed app.
object
Label-only target for externally-executed runs (e.g. imported from Mira).
Carries provider/model labels and opaque params instead of session setup:
external runs are ingested already-complete, so everruns never builds a
session from this. Mirrors a provider-agnostic (provider, model) pair.
object
Per-case timeout in seconds.
Example
120Responses
Section titled “ Responses ”Created
A single test case within an eval.
object
Session files to collect after scoring completes.
Named session file to collect after an eval case completes.
object
Export key for this artifact (for example patch or log).
Absolute path in the session filesystem.
Input messages sent sequentially.
A message to send to the agent during an eval case.
object
The text content to send.
When the case was created.
Optional description of what the case checks.
External identifier (evalcase_<32-hex>).
Max agent turns (default: 10).
Case name.
Display order.
Verification messages sent after conversation completes and session idles. Scorers run after post messages complete (not after conversation).
A message to send to the agent during an eval case.
object
The text content to send.
Scoring rules.
Final assistant message contains substring.
object
Final assistant message does NOT contain substring.
object
Final assistant message matches regex pattern.
object
Agent called named tool at least min times.
object
Agent did NOT call named tool.
object
Total tool calls within range.
object
Completed within N turns.
object
Session filesystem file contains substring.
object
Final assistant message parses as JSON matching schema.
object
Citation faithfulness: the answer’s citations
must cover the claim and be verified as supported. Scored from the
TextAnnotations on the final message — pair with the
citation_verification capability so verdicts are present.
object
Minimum number of citations the answer must carry.
Minimum fraction of citations verified entailed to pass.
Relative weight of this scorer in the case’s weighted average.
Citation faithfulness judged by an LLM: each cited claim/source pair is
graded by a model, so the eval works even without the
citation_verification capability.
object
Minimum judged score [0,1] to pass.
Rubric override; a citation-faithfulness rubric is used when absent.
Relative weight of this scorer in the case’s weighted average.
Free-form tags for filtering runs.
Session creation parameters (mirrors CreateSessionRequest).
object
Agent to work in this session.
Harness for the session. If omitted, org default harness is used.
Addressable harness name (alternative to harness_id).
Max LLM iterations per turn.
LLM model override.
System prompt override (prepended to agent prompt).
Reference to a deployed app.
object
Label-only target for externally-executed runs (e.g. imported from Mira).
Carries provider/model labels and opaque params instead of session setup:
external runs are ingested already-complete, so everruns never builds a
session from this. Mirrors a provider-agnostic (provider, model) pair.
object
Per-case timeout in seconds (default: 120).
When the case was last updated.
Example
{ "artifacts": [ { "name": "patch", "path": "/workspace/fix.patch" } ], "conversation": [ { "content": "Fix the failing test in src/lib.rs" } ], "created_at": "2026-01-15T10:30:00Z", "description": "Agent fixes a failing unit test", "id": "evalcase_01933b5a000070008000000000000001", "max_turns": 10, "name": "fix-failing-test", "position": 0, "post": [ { "content": "Run the tests again and report the result" } ], "scorers": [ { "text": "tests pass", "type": "contains" } ], "tags": [ "regression", "nightly" ], "target": { "type": "session" }, "timeout_seconds": 120, "updated_at": "2026-01-15T10:30:00Z"}