🐙Testfile

GitHub Action

The repository doubles as a GitHub Action, so running a Testfile in CI is a single step:

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: testfile-dev/testfile@main

The action installs Node, builds the runner, and executes testfile start against your repository’s Testfile. The job fails when tests fail.

Inputs

InputDefaultDescription
path.Testfile or directory containing one.
filter / filter-name / filter-tags / filter-matrix–Same as the -f/-n/-t/-m CLI filters.
changedfalseRun only tests whose inputs match files changed against the base branch (--changed).
changed-sincePR target branchBase branch/ref for changed. Defaults to github.base_ref, so on pull requests the diff is against the PR’s target branch.
fail-fastfalseAbort the whole run at the first failure.
max-parallel–Global cap on concurrently running tests.
reporter / output–Write machine-readable results (junit or json).
node-version22Node.js version for the runner.
doctortrueRun testfile doctor before the tests: every missing tool, engine or taken port that would fail the run anyway fails here instead, in one readable report.
annotationstrueEmit GitHub annotations for failed and aborted tests — each appears on the PR with the last 15 lines of its log.
summarytrueWrite a job-summary table of the run’s results (status, duration and notes per test).
statusesfalseReport one commit status per test — needs permissions: statuses: write.
status-prefixTestfile: Prefix of every status context, so the per-test statuses group together.
tokengithub.tokenToken the commit statuses are written with.
upload-runtrueUpload the recorded run folder (.testfile/runs/<id>) as a build artifact — GitHub wraps it in a zip.
artifact-nametestfile-runName of the uploaded run artifact.
variants–What tells this run apart from the other legs of a matrix, as key=value pairs separated by commas or newlines (e.g. platform=ubuntu-latest; whitespace is stripped, so values cannot contain spaces). Recorded in run.yaml; merging needs it.
labels–Extra labels to record, as key=value pairs separated by commas or newlines (e.g. tier=nightly, owner=infra). Merged with the automatic ones; a key you set yourself wins.
auto-labelstrueLabel the run with the GitHub context — see below.

What a CI run is labelled with

Every run the action records is labelled with where it came from, so a history collected from many workflows can be narrowed down afterwards (testfile serve filters by label). Each label is a key and a value; the runner takes them one --label key=value at a time.

KeyValue
triggerhow the workflow started: manual (a workflow_dispatch or repository_dispatch), schedule (a cron job), push, pull_request, or GitHub’s own name for any other event
branchthe branch the run used — on a pull request the source branch, not the ephemeral merge ref
basepull requests only: the target branch
prpull requests only: the pull request number
tagtag builds only, instead of branch
actorthe GitHub username that triggered the run
repoowner/name — worth having once runs from several repositories share a history
workflow, jobwhich workflow and job produced the run
osthe runner’s operating system
shathe short commit sha
ci-runthe Actions run id, to get from a recorded run back to its job

A label is only recorded when the context supplies it, so a run never carries an empty one. Your own labels: win over the automatic ones — setting branch=release replaces the derived value rather than clashing with it. Set auto-labels: false to record only your own.

Examples

Only the fast tests on pull requests, everything nightly:

- uses: testfile-dev/testfile@main
  if: github.event_name == 'pull_request'
  with:
    filter-tags: fast
    fail-fast: true

- uses: testfile-dev/testfile@main
  if: github.event_name == 'schedule'
  with:
    filter-tags: nightly

Only the tests a pull request could have affected, diffed against the PR’s target branch (note fetch-depth: 0 — change detection needs the base branch in the checkout, which a shallow single-commit clone doesn’t have):

- uses: actions/checkout@v7
  with:
    fetch-depth: 0
- uses: testfile-dev/testfile@main
  with:
    changed: true

Tests without inputs always run, so a suite adopts this incrementally: declare inputs on the expensive tests first. Each selected test’s log and run.yaml record why it ran — which input pattern matched how many changed files.

JUnit results for your test-report tooling:

- uses: testfile-dev/testfile@main
  with:
    reporter: junit
    output: results.xml
- uses: actions/upload-artifact@v7
  if: always()
  with:
    name: test-results
    path: results.xml

When tests fail, the action annotates the workflow run (and the PR’s files/checks views) with one error per failed or aborted test carrying the last 15 lines of its log, plus a summary notice — no extra configuration needed. The job summary shows a results table for every run, passing or failing.

A status per test

By default a commit gets one verdict from the action: the job passed or it did not. With statuses: true every test that ran reports its own commit status instead, so the pull request’s checks list names the test that broke rather than the job it was hiding in:

jobs:
  test:
    runs-on: ubuntu-latest
    # a status is a write to the repository; the default token may not have it
    permissions:
      contents: read
      statuses: write
    steps:
      - uses: actions/checkout@v7
      - uses: testfile-dev/testfile@main
        with:
          statuses: true

That produces one status per test, named after the test’s path behind the status-prefix:

StatusStateDescription
Testfile: ci/installsuccesspassed in 12.4s
Testfile: ci/checks/lintsuccesspassed in 900ms (cached)
Testfile: ci/checks/unitfailurefailed in 2.0s
Testfile: ci/checks/e2esuccessskipped

An aborted test (or a status the reporter does not recognize) maps to the error state. Each status links back to the workflow run (when the run id is available). A few details worth knowing:

Because the contexts are stable, they can be used as required checks — pick Testfile: ci/checks/unit in the branch protection rules and a pull request cannot merge while that one test is red, whatever the rest of the suite does. Keep in mind that a test which stops being selected (by a filter, or by changed) stops reporting its status too, and a required check that never arrives blocks the merge. GitHub caps a status context at 255 characters and its description at 140 — longer ones are truncated with …, so a very deep test path plus a long prefix and variants may produce a context that differs from the one you would type into a branch rule.

Bringing CI runs home

Every action run uploads the recorded run folder as a testfile-run artifact — GitHub wraps it in a zip, so the artifact is a single zip with run.yaml and the logs at its root. A manually downloaded artifact imports with testfile archive import testfile-run.zip. testfile github sync downloads the artifacts of the latest workflow runs — every artifact whose name starts with testfile-run, so the per-platform legs and the merged run all arrive — and imports them into your local run history, where testfile runs, inspect run, diff, --flaky and the TUI’s runs/tests views treat them like local runs:

export GITHUB_TOKEN=...          # a token with actions:read (GH_TOKEN works too)
export GITHUB_TOKEN=$(gh auth token)   # ... or reuse the gh CLI's login
testfile github sync you/your-repo --latest 10
testfile runs             # CI runs are now part of the history

See sharing runs for the underlying pack/import commands and the S3 variant.

Container services (postgres etc.) work out of the box on the standard ubuntu-latest runners, which ship with docker — the runner’s engine selection finds it on its own. To pin the engine explicitly (or to run services on a cluster the job can reach), set the environment variable the runner reads:

- uses: testfile-dev/testfile@main
  env:
    TESTFILE_ENGINE: kubernetes

More than one platform

The action runs on the Windows and macOS runners too, so one matrix job covers all three:

jobs:
  test:
    name: Testfile CI (${{ matrix.os }})
    strategy:
      fail-fast: false
      matrix:
        os: [ubuntu-latest, macos-latest, windows-latest]
    runs-on: ${{ matrix.os }}
    steps:
      - uses: actions/checkout@v7
      - uses: testfile-dev/testfile@main
        with:
          # what tells the legs apart when their runs are merged
          variants: platform=${{ matrix.os }}
          # artifact names are unique per workflow run
          artifact-name: testfile-run-${{ matrix.os }}

Three platforms means three recorded runs. A follow-up job combines them into a single run — one verdict, one duration, every test tagged with the platform it ran on — with testfile merge. The three-platform guided tour walks through the whole workflow, merge job included.

Two things to know before you do:

This repository’s own CI is that job — one Testfile, three platforms — plus a merge job combining the three runs and a kind cluster on the Linux leg for the kubernetes conformance case.

Letting an assistant read the failure

A red CI run usually reaches a person as a link and a wall of log. The run folder holds better material than that, and the action already uploads it, so a follow-up job can hand a failure to an assistant with the evidence attached instead of the URL.

  triage:
    needs: test
    if: failure()
    runs-on: ubuntu-latest
    permissions:
      contents: read
      pull-requests: write
    steps:
      - uses: actions/checkout@v7
      - uses: actions/download-artifact@v5
        with:
          pattern: testfile-run*
          path: runs
      # the downloaded artifacts become an ordinary local history
      - run: npx @testfile.dev/cli archive import runs/*/testfile-run*.zip
      # what failed, why, and what changed - bounded, and already prose
      - run: npx @testfile.dev/cli explain --max-failures 5 > digest.md
      - uses: anthropics/claude-code-action@v1
        with:
          anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
          prompt: |
            The CI run below failed. Work out whether each failure is real
            or a known flake, and what caused it. Do not change any code.
            $(cat digest.md)

The digest is the point: it is already the shape an assistant needs (failing tests, their reasons, log excerpts, the flaky verdict, the comparison with the previous run), it is bounded, and it costs one command rather than a prompt that has to explain the layout of .testfile/.

Give the job the MCP server instead of a digest when the assistant should be able to dig further — testfile mcp over the same imported history lets it pull a specific log or reproduce one test on its own.

Writing the conclusion back

An analysis that lives only in a job log is lost by the next run. The result format has a place for it: an optional analysis field on run.yaml, which every viewer shows next to the run — marked as somebody’s reading of it, never as a result.

node -e '
  const { readFileSync, writeFileSync } = require("node:fs");
  const { parse, stringify } = require("yaml");
  const file = process.argv[1], run = parse(readFileSync(file, "utf8"));
  run.analysis = { text: readFileSync("finding.md", "utf8"), author: "claude-code",
                   at: new Date().toISOString() };
  writeFileSync(file, stringify(run));
' .testfile/runs/<id>/run.yaml

A runner never writes that field, and neither does any viewer — the viewers are read-only. Whoever did the reading writes it, and must preserve every other field of the record.