Create an observer (online scoring).
const url = 'https://app.everruns.com/api/v1/observers';const options = { method: 'POST', headers: {'Content-Type': 'application/json'}, body: '{"description":"Score replies from the support agent","match":{"agent_ids":["example"],"harness_ids":["example"],"session_tags":["example"]},"name":"Support quality","sampling_rate":0.25,"scorers":[{"model_id":"example","pass_threshold":1,"rubric":"example","method":"llm_judge","key":"example","scope":"turn"}]}'};
try { const response = await fetch(url, options); const data = await response.json(); console.log(data);} catch (error) { console.error(error);}curl --request POST \ --url https://app.everruns.com/api/v1/observers \ --header 'Content-Type: application/json' \ --data '{ "description": "Score replies from the support agent", "match": { "agent_ids": [ "example" ], "harness_ids": [ "example" ], "session_tags": [ "example" ] }, "name": "Support quality", "sampling_rate": 0.25, "scorers": [ { "model_id": "example", "pass_threshold": 1, "rubric": "example", "method": "llm_judge", "key": "example", "scope": "turn" } ] }'Request Bodyrequired
Section titled “Request Bodyrequired”Request to create a new observer.
object
Human-readable description. Safe to render in user-facing messages.
Example
Score replies from the support agentWhich production sessions to score. Empty matches all org traffic.
object
Match sessions running any of these agents.
Match sessions on any of these harnesses.
Match sessions carrying any of these tags.
Human-readable name. Safe to render in user-facing messages.
Example
Support qualityFraction of matching turns to score (0.0–1.0). Defaults to 0.1.
Example
0.25Scoring rules. Must contain at least one.
One scorer inside an observer. key names the score series in listings
and future dashboards; scope selects the trace slice; method is how
it grades.
object
Deterministic rule. file_contains is rejected for observers (session
filesystems are not part of the observable trace contract).
object
Final assistant message contains substring.
object
Final assistant message does NOT contain substring.
object
Final assistant message matches regex pattern.
object
Agent called named tool at least min times.
object
Agent did NOT call named tool.
object
Total tool calls within range.
object
Completed within N turns.
object
Session filesystem file contains substring.
object
Final assistant message parses as JSON matching schema.
object
Citation faithfulness: the answer’s citations
must cover the claim and be verified as supported. Scored from the
TextAnnotations on the final message — pair with the
citation_verification capability so verdicts are present.
object
Minimum number of citations the answer must carry.
Example
1Minimum fraction of citations verified entailed to pass.
Example
0.8Relative weight of this scorer in the case’s weighted average.
Example
1Citation faithfulness judged by an LLM: each cited claim/source pair is
graded by a model, so the eval works even without the
citation_verification capability.
object
Minimum judged score [0,1] to pass.
Example
0.8Rubric override; a citation-faithfulness rubric is used when absent.
Example
Score the fraction of cited claims supported by their source.Relative weight of this scorer in the case’s weighted average.
Example
1LLM-as-judge.
object
Org model to judge with. When None, the org’s default model is used.
Judge calls go through the org’s own providers and are billed to it.
Score value at/above which pass is true.
Grading rubric shown to the judge model. Should describe what a high vs. low score means.
Stable name within the observer (score series name).
Trace slice this scorer grades.
Responses
Section titled “ Responses ”Created
An observer: online scoring config over production sessions.
object
Optional description.
External identifier (observer_<32-hex>). Shown as “id” in API.
Which sessions to score.
object
Match sessions running any of these agents.
Match sessions on any of these harnesses.
Match sessions carrying any of these tags.
Display name.
Fraction of matching turns to score (0.0–1.0), applied after match.
Scoring rules.
One scorer inside an observer. key names the score series in listings
and future dashboards; scope selects the trace slice; method is how
it grades.
object
Deterministic rule. file_contains is rejected for observers (session
filesystems are not part of the observable trace contract).
object
Final assistant message contains substring.
object
Final assistant message does NOT contain substring.
object
Final assistant message matches regex pattern.
object
Agent called named tool at least min times.
object
Agent did NOT call named tool.
object
Total tool calls within range.
object
Completed within N turns.
object
Session filesystem file contains substring.
object
Final assistant message parses as JSON matching schema.
object
Citation faithfulness: the answer’s citations
must cover the claim and be verified as supported. Scored from the
TextAnnotations on the final message — pair with the
citation_verification capability so verdicts are present.
object
Minimum number of citations the answer must carry.
Minimum fraction of citations verified entailed to pass.
Relative weight of this scorer in the case’s weighted average.
Citation faithfulness judged by an LLM: each cited claim/source pair is
graded by a model, so the eval works even without the
citation_verification capability.
object
Minimum judged score [0,1] to pass.
Rubric override; a citation-faithfulness rubric is used when absent.
Relative weight of this scorer in the case’s weighted average.
LLM-as-judge.
object
Org model to judge with. When None, the org’s default model is used.
Judge calls go through the org’s own providers and are billed to it.
Score value at/above which pass is true.
Grading rubric shown to the judge model. Should describe what a high vs. low score means.
Stable name within the observer (score series name).
Trace slice this scorer grades.
Lifecycle status.
Example
{ "description": "Grades support replies for grounded answers.", "id": "observer_01933b5a000070008000000000000001", "name": "Support answer quality", "sampling_rate": 0.1, "scorers": [ { "method": "llm_judge", "scope": "turn" } ], "status": "active"}