Cache and determinism
EvalCore’s record/replay cache makes reruns free, offline, and deterministic: every cacheable target call is stored keyed by a content hash of the canonical request, so a committed cache lets CI replay a suite with zero LLM spend and no API keys. This page documents what goes into a cache key, the four modes, and the determinism guarantees the cache is built on.
Cache location
Section titled “Cache location”The cache is a single SQLite file at .evalcore/cache.db, resolved relative to
the config file’s directory. It is a project artifact, like a VCR cassette,
and committing it is the intended workflow. The directory and file are created
on demand: a purely local run (a shell target, no judge, no baseline) never
opens the store, so no .evalcore/ directory appears.
The file holds two tables: llm_cache (the record/replay entries) and runs
(saved baselines).
What is in a cache key
Section titled “What is in a cache key”A cache key is the SHA-256 of the canonical request JSON:
{"identity": <target.cache_identity()>, "input": <case input>}serde_json serializes objects with sorted keys, which is what makes the
canonical string stable. The preserve_order feature is banned workspace-wide
because enabling it would silently invalidate every cache on disk. The
identity is the target’s cache_identity(); anything that changes what would
be sent must be inside it, and secrets must never be.
Two design rules govern every identity. Unset optional fields are omitted
rather than serialized as null, so adding a new config field never
invalidates existing cassettes. Secrets and call mechanics are excluded, so
they never get persisted and changing them never re-records.
Under run.trials, trial 0 keeps this exact key, so cassettes recorded before
trials existed (or by a single-trial run) replay unchanged. Trials 1..N add the
trial index to the key, giving each trial its own cache entry, and judge and
embedding calls made while scoring trial i re-key the same way. Since v0.7.0.
See Trials and statistics.
openai-compatible identity
Section titled “openai-compatible identity”{"type": "openai-compatible", "url": "…", "model": "…", "system": "…", "params": {…}}| Field | In the key? |
|---|---|
url, model |
yes |
system |
yes, when set (omitted otherwise) |
params |
yes, when set and non-empty (omitted otherwise) |
api_key_env / the resolved key |
no (a secret) |
max_retries |
no (changes how we call, not what the model answers) |
timeout_seconds |
no (operational knob; excluded since v0.5.0 so older cassettes keep their keys) |
cost rates |
no (accounting, not request content) |
http identity
Section titled “http identity”{"http": {"url": "…", "method": "…", "headers": {…}, "body": {…}, "response_path": "…"}}| Field | In the key? |
|---|---|
url (the pre-substitution template) |
yes |
method (normalized uppercase) |
yes |
headers |
yes, when non-empty; names lowercased and sorted |
body (the pre-substitution template) |
yes, when set |
response_path |
yes, when set |
api_key_env / the resolved key |
no (a secret) |
auth_header, auth_prefix |
no |
max_retries, timeout_seconds |
no |
Header values are part of the identity, and are therefore hashed into the
committed .evalcore/cache.db, so never put secrets in headers:. Use
api_key_env, which never enters the cache. A headers: name that collides
case-insensitively with the auth header is rejected at config validation (see
the http validation rules).
Uncacheable targets
Section titled “Uncacheable targets”shell and trace targets return cache_identity() == None. They are never
cached and pass straight through in every mode. Shell targets are uncacheable by
design: local code can change behavior without the config string changing, so a
recording could go stale silently. Traces are local files, never worth caching.
The four modes
Section titled “The four modes”--cache <mode> selects behavior for cacheable targets. See the CLI
reference.
| Mode | Behavior |
|---|---|
auto (default) |
Replay on a hit; on a miss, call live and record the result. |
replay |
Cache only. A miss fails the case (cache miss for case "<id>" in replay mode — record it first with --cache auto (or live)), and it never calls live. |
live |
Always call live and overwrite the recording. |
off |
Bypass the cache entirely: no reads, no writes. |
Replayed outputs are returned verbatim, including the recorded
latency_ms and any recorded token usage, so cost accounting stays consistent
offline. Different identities never share entries: the same input under a
different model or URL misses and records separately.
Cassette-commit workflow
Section titled “Cassette-commit workflow”The cache is meant to be recorded once and committed:
- Record locally with a live-capable mode:
evalcore run evals.yaml --cache auto(records misses) or--cache live(re-records everything). - Commit
.evalcore/cache.dbalongside the config and datasets. - In CI, run
evalcore run evals.yaml --cache replay. Replay never calls the network and never reads API keys. A miss is a failure, so CI is deterministic and free.
Because --cache replay uses the optional-secret policy, unset api_key_env
variables are fine there, because the committed cassette carries every
recorded answer.
Determinism guarantees
Section titled “Determinism guarantees”Identical inputs produce identical outputs everywhere:
- Dataset order. Results stay in dataset order even though cases run concurrently, because the engine buffers completions back into input order. Reports and diffs are therefore stable.
- Pure reporters. Every reporter is a pure
&RunSummary -> Stringfunction with no I/O, no clock reads, and no environment access, so identical runs render byte-identical reports (they are snapshot-tested). - Sorted JSON keys. Canonical request JSON and every rendered JSON payload
sort object keys (
preserve_orderis banned), so cache keys and reports don’t drift with map iteration order. - No clock reads except latency measurement. Replay returns recorded
latencies, so even timing is reproducible under
--cache replay. - Deterministic retries. Transient failures back off on a fixed schedule
(500ms, 1s, 2s, … capped at 10s, honoring
Retry-After) with no jitter. - Sanitized terminal output. User-derived text in the terminal reporters (case ids, scorer reasons, target errors, gate and class labels) is neutralized against ANSI/control-sequence injection: ESC, C0/C1/DEL, CR/LF/TAB, and Unicode bidirectional controls render as visible escapes, so a hostile model output or dataset can’t move the cursor or reorder a report line. The escaping itself is deterministic; benign text is unchanged. Since v0.7.5.
A corrupt cache entry surfaces as an error telling you to delete the cache file. EvalCore never falls back to a silent live call, which would un-determinize a replay run.
Baselines
Section titled “Baselines”Baselines are stored in the same .evalcore/cache.db file, in the runs table.
--save-baseline <label> appends this run’s per-case snapshot under the label;
--baseline <label> loads the newest row with that label (labels are not
unique, so each save appends and the latest wins) and compares. Labels are
independent of each other.
A saved baseline is a pure per-case snapshot: suite-gate results are run-scoped
acceptance criteria, not case data, so they are cleared before a summary is
persisted (and old rows that predate the gates field still deserialize and
re-serialize byte-identically). Because the store also holds baselines, a
--baseline or --save-baseline flag opens .evalcore/cache.db even for a
shell target that never touches the cache.
See the CLI reference for how the baseline diff prints and how it flips the exit-code contract.
See also
Section titled “See also”- Record / replay: the day-to-day workflow these invariants exist to support.
- Running in CI: replaying committed cassettes offline, without credentials.
- CLI reference: the
--cachevalues that select the modes above. - Gates and baselines: what a saved baseline is compared against on the next run.
- Configuration reference: the target fields that feed the cache identities described here.