Subprocess scorer protocol
The subprocess scorer is EvalCore’s any-language escape hatch. It runs a
command once per case, hands it the case as JSON on stdin, and reads a
verdict as JSON from stdout. Any language that can read stdin and write
stdout can implement a scorer, so no Rust is required.
This is versioned API surface (protocol v0). Field names and semantics are stable; changes are breaking.
Configuration
Section titled “Configuration”scorers: - type: subprocess cmd: python3 scorers/faithfulness.pycmd is run via sh -c, once per case.
Input (stdin)
Section titled “Input (stdin)”The scorer receives a single JSON object on stdin:
{ "context": ["Refunds are processed within 30 days."], "expected": "30 days", "input": "How long do refunds take?", "output": "Refunds are honored within 30 days."}| Field | Type | Description |
|---|---|---|
input |
string | The case’s input, as sent to the target. |
output |
string | The target’s output text for this case. |
expected |
any JSON, or null |
The case’s expected field, verbatim. null when the case has none. |
context |
list of strings | The case’s retrieved RAG context chunks. Present only when the case carries context, and omitted entirely otherwise, so a contextless payload keeps its original shape. Since v0.6.0. |
Object keys are serialized in alphabetical order (context, expected,
input, output), so a contextless case’s payload is byte-identical to the
pre-0.6.0 shape.
The command must read stdin (even if only to discard it). A command that
exits before EvalCore finishes writing the payload can race and fail the write.
In shell, drain it with cat >/dev/null before doing anything else.
Output (stdout)
Section titled “Output (stdout)”The scorer must print a single JSON object on stdout:
{"score": 0.9, "passed": true, "reason": "grounded in the provided context"}| Field | Type | Required | Description |
|---|---|---|---|
score |
number | yes | Must be within 0.0..=1.0. A score outside that range is an error. |
passed |
bool | no | Whether the case passes. When omitted, defaults to score >= 0.5. |
reason |
string | no | Human-readable explanation, shown in reports for failing cases. Omitted when absent. |
Stdout is trimmed before parsing, so trailing newlines are fine. Only these three fields are read.
Error handling
Section titled “Error handling”Failures are data. A misbehaving scorer surfaces as a failing or errored score rather than a panic or a crashed run.
| Condition | Result |
|---|---|
| Non-zero exit | The case’s score errors; the message includes the exit status and the command’s trimmed stderr. |
| Invalid / non-JSON stdout | Errors: scorer "<cmd>" printed invalid verdict JSON: "<stdout>". |
score outside 0.0..=1.0 |
Errors: scorer "<cmd>" returned score <n> outside 0.0..=1.0. |
| Spawn failure | Errors: failed to spawn subprocess scorer: <cmd>. |
The engine converts a scorer error into a failing Score with the reason
prefixed scorer error: …, so one bad scorer never aborts the suite.
Worked example: Python
Section titled “Worked example: Python”#!/usr/bin/env python3import jsonimport sys
case = json.load(sys.stdin) # read stdin fullyoutput = case["output"]expected = case.get("expected") or ""
# toy metric: fraction of expected words present in the outputwanted = str(expected).lower().split()hits = sum(1 for w in wanted if w in output.lower())score = hits / len(wanted) if wanted else 1.0
print(json.dumps({ # write exactly one JSON object "score": score, "passed": score >= 0.8, "reason": f"{hits}/{len(wanted)} expected terms present",}))scorers: - type: subprocess cmd: python3 scorers/coverage.pyWorked example: Node
Section titled “Worked example: Node”#!/usr/bin/env nodelet raw = "";process.stdin.on("data", (chunk) => (raw += chunk)); // read stdin fullyprocess.stdin.on("end", () => { const { output, expected } = JSON.parse(raw); const want = String(expected ?? "").toLowerCase(); const passed = output.toLowerCase().includes(want); process.stdout.write( // one JSON object JSON.stringify({ score: passed ? 1.0 : 0.0, passed, reason: passed ? "expected text found" : `missing ${JSON.stringify(want)}`, }) );});scorers: - type: subprocess cmd: node scorers/contains.jsBoth examples read stdin to completion before writing, then emit exactly one
JSON object with a required score, as the protocol requires.
See also
Section titled “See also”- Custom scorers: a complete scorer built on this protocol, plus debugging tips when a command misbehaves.
- Configuration reference: the
subprocessscorer fields that point at your command. - RAG evaluation: the shipped Ragas and DeepEval shims, working implementations of this contract.
- LLM-as-judge: the built-in option when the check is a graded rubric rather than code you write.