HTML reports
Exit codes gate the build; an HTML report is what a human opens to see why. The
--html flag writes a single, self-contained document, the shareable “here’s the
eval report” artifact a reviewer clicks straight from a PR.
evalcore run evals.yaml --cache replay --html report.htmlThis is what comes out: a real report from the shipped quickstart suite, embedded as-is:
evalcore run examples/quickstart/evals.yaml --html report.html. This exact file, unedited.It is written in addition to the primary --reporter output, never instead
of it, so you can emit terminal output to the console and an HTML file in the
same run. It composes with every reporter and every flag.
What’s in the report
Section titled “What’s in the report”The document mirrors what the terminal reporter shows, expanded for the browser:
- The header carries an overall pass/fail badge and the summary counts (passed / failed / total), plus total tokens and cost when the target reports usage.
- The gates panel has one row per suite gate with its status, floor, actual value, and any reason. It is omitted entirely when no gates are configured.
- The cases section has one expandable row per case, in dataset order. Each row
shows status, case id, latency, and cost; expanding it reveals:
- the output text,
- a scores table (each scorer’s value, pass/fail, and reason),
- for a target error, the error as the case’s reason,
- for
tracecases, the agent trajectory: each tool call with its input and output, expandable inline.
- The baseline diff appears when you pass
--baseline. It embeds the same regressed / new-failing / fixed / removed breakdown the terminal diff prints.
Self-contained and deterministic
Section titled “Self-contained and deterministic”The report is deliberately minimal and portable:
- One file, no dependencies. All CSS is inlined, there are no external requests
(no CDN, no fonts, no images), and there is no JavaScript. Expansion is driven
entirely by native HTML
<details>elements. - Air-gapped and
file://-friendly. Because nothing is fetched, it opens correctly straight from disk (file:///path/report.html) with no server and no network, which suits locked-down or offline environments. - Light and dark. It themes to the reader’s
prefers-color-schemeautomatically. - Deterministic byte-for-byte. Like every reporter, it’s a pure function of the run. The same run renders identical bytes, so it diffs and snapshot-tests cleanly.
- Safe. Every user-derived value (case ids, outputs, reasons, tool names, JSON payloads) is HTML-escaped, so even a hostile model output renders as inert text.
The GitHub Action produces it automatically
Section titled “The GitHub Action produces it automatically”The Action’s html-artifact input (default "evalcore-report") passes --html
for you and uploads the file as a CI artifact:
- uses: eval-core/evalcore@v0.7.5 with: config: evals/evals.yaml args: --cache replay --baseline main html-artifact: evalcore-report # artifact name; "" disables itTwo behaviors worth knowing:
- It uploads even when the suite fails. The report is most valuable on a red build, so the upload step runs regardless of the run’s outcome and never changes the run’s exit code. The job step summary notes that the report was uploaded.
- Set it to
""to disable. With an empty value, no--htmlis passed and nothing is uploaded. The command is byte-identical to a run without the flag.
Using it in PR review
Section titled “Using it in PR review”The workflow for a reviewer:
- A PR’s eval job runs (offline, replayed) and goes red if it regressed something.
- The reviewer opens the run’s artifacts and downloads
evalcore-report. - They open the HTML file and expand the failing case to read its output, the scorer that failed and why, and for agents the exact trajectory the run took.
- With
--baselineset, the embedded diff shows precisely which cases regressed relative to the accepted state.
The answer to “why did this fail?” is one click and one expand away, without scraping logs or re-running with more verbosity.
For the CI wiring around this, see Running in CI.
See also
Section titled “See also”- Running in CI: the GitHub Action’s
html-artifactinput in the full CI pipeline. - Gates and baselines: what the report’s gates panel and baseline diff actually mean.
- Agents and traces: the trajectory
data the report expands inline for
tracecases. - CLI reference: every flag
--htmlcomposes with, including--reporterand--baseline.