# Testfile A declarative YAML format for running a project's tests: one file describes what to run, what it needs (services, ports, containers) and how to select subsets of it, so the same suite runs on a laptop and in any CI system. This file is the whole documentation, concatenated: the docs, the normative specification and the worked guides. An index of the same pages is at https://testfile.dev/llms.txt. # Get started Source: https://testfile.dev/start An interactive page, one question at a time: the language and its version, whether the tests run on this machine or in a container, and which PostgreSQL version they need - or every version of either, which fans the suite out over all of them. For Node.js 22 in a container with PostgreSQL 17 it produces: ```yaml version: 0 # A free port is picked per run, so two runs never collide. ports: postgres: random test: name: ci # every command below runs in this image, with the project mounted container: image: docker.io/library/node:22 sequence: - name: install command: npm ci - name: lint command: npm run lint - name: unit command: npm test # skipped when none of these changed since the last passing run inputs: ["src/**", "package-lock.json"] - name: integration services: postgres: container: image: docker.io/library/postgres:17-alpine ports: ["${{ ports.postgres }}:5432"] env: POSTGRES_PASSWORD: test POSTGRES_DB: app_test ready: # runs inside the container, so the image's own client is used exec: pg_isready -h 127.0.0.1 -p 5432 -U postgres timeout: 60s env: DATABASE_URL: postgres://postgres:test@127.0.0.1:${{ ports.postgres }}/app_test command: npm run test:integration ``` # What is a Testfile? Source: https://testfile.dev/docs/index A **Testfile** is a YAML file — named `Testfile`, `testfile.yaml` or `testfile.yml` — that describes how your project runs its tests, checked into the repository next to your code. It is built around three ideas: 1. **Tests nest.** The root of the file contains one test. Every test either runs a shell command or script, or groups nested tests that run in sequence or in parallel — together they form the test suite. A matrix can expand one test into many combinations (Node versions × databases, for example). 2. **Tests can need services.** Real test suites need the app under test, a database in a specific version, a message broker. A Testfile declares those as *services*: local processes or container images (podman/docker; running them on Kubernetes is planned). The runner starts them, waits until they are actually ready — an HTTP health check, an open TCP port, or a log line — and gracefully shuts them down afterwards, even when you hit Ctrl+C. 3. **The environment is explicit.** Environment variables, named ports (including randomly allocated free ports), and matrix values are declared in the file and injected via `${{ ... }}` templates, so runs are reproducible and parallel runs don't fight over ports. A small but complete example: ```yaml version: 0 name: my-app ports: web: random services: web: command: npm start env: PORT: ${{ ports.web }} ready: http: http://localhost:${{ ports.web }}/healthz test: name: all sequence: - name: lint command: npm run lint - name: e2e env: BASE_URL: http://localhost:${{ ports.web }} command: npm run test:e2e ``` Run it with the [`testfile` CLI](./cli): ```sh testfile start # plain output testfile tui # browse recorded runs (read-only terminal UI) ``` Continue with [Getting started](./getting-started), see [everything at once](./complete-example) in one annotated file, or read the [full specification](../spec/TESTFILE.md) and its [versioning policy](../spec/VERSIONING.md) — once version 1 ships, the format will evolve additively within it, and a [conformance suite](https://github.com/testfile-dev/testfile/blob/main/conformance/README.md) pins its semantics for alternative runner implementations. # Getting started Source: https://testfile.dev/docs/getting-started ## 1. Create a Testfile Three ways in, in increasing order of typing: [the wizard](https://testfile.dev/start) asks a handful of questions and hands you a file to copy; `testfile init` converts what your project already has; or create a file called `Testfile` (or `testfile.yaml`) in the root of your project yourself: ```yaml version: 0 name: my-project test: name: unit tests command: npm test ``` Every Testfile needs a `version` (currently always `0` while the format is under review — version 1 is targeted for Q4 2026) and exactly one root `test`. ### Importing what you have `testfile init` looks for files that already describe your setup and converts them, so the first version is rarely empty: | Source | Becomes | | ------ | ------- | | `package.json` scripts | tests for `lint`, `typecheck`, `test`, `test:*`, `build` | | `compose.yaml` / `docker-compose.yaml` | `services` — image, ports (as [random ports](./env-and-ports#named-ports)), env, volumes, command/entrypoint, `healthcheck` → [`ready`](./services#readiness-checks), `depends_on` → [`needs`](./services#service-dependencies); a published port without a healthcheck becomes a `tcp` check | | `.github/workflows/*.yaml` | one test per job, its `run:` steps as a sequence (`uses:` steps are dropped). Auto-detection picks the first workflow it finds; pass others with `--from` | | `Makefile`, `Taskfile.yaml`, `justfile` | tests for targets that look like checks (`test`, `lint`, `check`, `e2e`, …) | ```sh npx @testfile.dev/cli init # auto-detect all of the above npx @testfile.dev/cli init --from docker-compose.yml # just this file npx @testfile.dev/cli init --no-detect # package.json scripts only ``` The conversion is deliberately best-effort: anything that could not be translated — a service with neither health check nor ports, dropped action steps, a build matrix — is written into the file as a `# note:` comment and printed, so you know where to look. Read the result before trusting it. ## 2. Run it ```sh npx @testfile.dev/cli start ``` The runner finds the Testfile in the current directory, runs the suite and prints a summary. The exit code is `0` when everything passed, `1` on failure, `130` when you interrupted the run. The whole `testfile` command line ships as the [`@testfile.dev/cli`](https://www.npmjs.com/package/@testfile.dev/cli) package: `npx` runs it without installing, and `npm install -g @testfile.dev/cli` puts plain `testfile` on your `PATH` — which is how the rest of the documentation writes the commands. (A CI job that only runs the suite can use the leaner `@testfile.dev/runner` instead — see [CI systems](./ci-systems).) Other useful commands: ```sh testfile validate # check the file against the JSON schema testfile doctor # check this machine against what the file needs testfile inspect # print the expanded test suite without running it testfile tui # browse recorded runs (read-only terminal UI) ``` ## 3. Grow the suite Replace the single command with groups as your test suite grows: ```yaml version: 0 test: name: all sequence: - name: build command: npm run build - name: checks parallel: - name: lint command: npm run lint - name: unit command: npm run test:unit ``` `sequence` runs children one after another and stops at the first failure; `parallel` runs them concurrently. Groups nest arbitrarily — see [Writing tests](./writing-tests). ## 4. Editor support Most YAML language servers pick up the schema from a modeline. Add this as the first line of your Testfile to get completion and validation while typing: ```yaml # yaml-language-server: $schema=https://testfile.dev/next/testfile.schema.json ``` # CLI & TUI Source: https://testfile.dev/docs/cli One command line, two halves. `testfile start` and its siblings read the Testfile, run the processes it describes and record what happened; `testfile runs`, `explain`, `repro`, `tui`, `serve` and the rest are strictly read-only over those recorded runs — they never touch the Testfile and never start a test, so they work with any tool producing the documented [result format](https://github.com/testfile-dev/testfile/tree/main/spec), not just this runner. Under the hood that separation is real: the reading half is a set of packages ([`core`](https://github.com/testfile-dev/testfile/tree/main/testfile-ts/core), `sync`, `mcp`, `tui`, `web`) that never depend on the runner, and [`runner`](https://github.com/testfile-dev/testfile/tree/main/testfile-ts/runner) is the only one that starts anything. The whole command line ships as the [`@testfile.dev/cli`](https://www.npmjs.com/package/@testfile.dev/cli) package: `npx @testfile.dev/cli ` runs any of it without installing, and `npm install -g @testfile.dev/cli` makes it available as the plain `testfile` this page writes. This page is the guided tour; the [CLI reference](./cli-reference) lists every command with all of its arguments and options. ## Commands ```sh # the runner testfile init [path] # create a starter Testfile (from what the project has) testfile start [path] # run the test suite (default command) testfile start --verbose # also stream service output testfile start --fail-fast # abort everything at the first failure testfile start --max-parallel 4 # global cap on concurrently running tests testfile start --dry-run # print what would run, without running testfile start --watch # re-run on file changes testfile start --reporter junit --output results.xml # report for CI testfile start --variant platform=linux # tag the run (for merging later) testfile validate [path] # validate against the JSON schema testfile doctor [path] # check this machine against what the file needs testfile inspect [path] # print the expanded suite, incl. matrix instances testfile tags [path] # list the tags the suite uses testfile changes [path] # which tests a change selection would run testfile completion bash # shell completions (bash, zsh, fish) # the viewer (read-only over .testfile/) testfile runs # table of recorded runs (default command) testfile inspect run # one run in detail testfile diff # compare two runs testfile merge # combine shards / platform runs into one testfile tui # terminal UI: runs + tests, watching testfile serve # localhost REST API + web viewer testfile archive # pack/import recorded runs as archives testfile s3 # push/pull/list runs in an S3 bucket testfile github # sync/list GitHub Actions run artifacts testfile gitlab # sync/list GitLab CI job artifacts ``` Tab completion for commands and flags: ```sh source <(testfile completion bash) # bash testfile completion zsh > "${fpath[1]}/_testfile" # zsh testfile completion fish > ~/.config/fish/completions/testfile.fish ``` `--no-cache` ignores [cached results](./writing-tests#result-caching) for tests with `inputs` (fresh results still refresh the cache). `--engine` picks the container engine for the run — `podman`, `docker` or `kubernetes`; without it `TESTFILE_ENGINE` decides, and without either the first engine that responds is used ([how selection works](./services#containers)). `--fail-fast` aborts running siblings and skips everything queued as soon as one test fails. `--max-parallel` limits how many tests run at the same time across the *whole* run (group-level `maxParallel` still applies on top). `--dry-run` combines with all filter flags, so you can preview exactly what a filter expression will run; tests whose [inputs](./writing-tests#result-caching) are unchanged are marked `[cached]`, with a would-run/served-from-cache summary. ## Overriding the Testfile for one run Some runs need a file that is *almost* the committed one: the same suite against a newer database image, a fixed port instead of a random one, one command swapped for a smoke-test variant. Both `-c` and a `TESTFILE_CONFIG_` variable write a value into the loaded document for that run only — the file on disk is never touched: ```sh testfile start -c ports.db=15432 -c services.postgres.container.image=postgres:17 2 override(s) applied: ports.db, services.postgres.container.image ``` ```sh TESTFILE_CONFIG_ports__db=15432 \ TESTFILE_CONFIG_services__postgres__container__image=docker.io/library/postgres:17 \ testfile start ``` The flag is the one to reach for by hand; the variable is for the places that only have an environment — a CI job's `env:` block, a container, a `docker run -e`. **Where both name the same path, the flag wins**, and the file loses to either. That is the whole order: `-c`, then `TESTFILE_CONFIG_`, then what the Testfile says. The two forms take the same path, written the way each medium allows: `-c` separates the segments with `.` as the file reads, and the variable uses `__` (two underscores) because an environment variable name cannot hold a dot. A pasted `__` path works with `-c` too, so the two are interchangeable. Segments walk maps by key and lists by index, so anything in the document can be addressed: | `-c` | as a variable | what it changes | | ---- | ------------- | --------------- | | `-c env.LOG_LEVEL=debug` | `TESTFILE_CONFIG_env__LOG_LEVEL` | a top-level environment variable | | `-c ports.db=15432` | `TESTFILE_CONFIG_ports__db` | a named port | | `-c services.postgres.container.image=…` | `TESTFILE_CONFIG_services__postgres__container__image` | the image of one service | | `-c services.postgres.ready.timeout=5m` | `TESTFILE_CONFIG_services__postgres__ready__timeout` | that service's readiness timeout | | `-c test.sequence.1.command=…` | `TESTFILE_CONFIG_test__sequence__1__command` | the second test of the root sequence | | `-c test.maxParallel=4` | `TESTFILE_CONFIG_test__maxParallel` | a group's parallel cap | Values are read as the strings they are, so `postgres:17`, `1.20.3` and `0755` arrive unchanged. Four things are read as more than a string: - a bare `true`, `false` or `null`, - a plain number without leading zeros — `15432` becomes the integer a port wants, - something starting with `[` or `{` — `[slow, flaky]` becomes a list, `{count: 2, delay: 1s}` a map, - something starting with a quote, which is also the escape hatch when you need the literal string `"true"` or `'15432'`. Details worth knowing: - **The result is validated again.** An override that breaks the document — a list where a command belongs, a timeout that is not a duration — fails the run with the usual message instead of half-applying. - **A path that does not exist yet is created**, so `test__env__DEBUG=1` adds an `env` block to a test that had none. Lists are the exception: an index must already exist, because what a new element would mean is anyone's guess. - **Segments match loosely**, in this order: exactly, then ignoring case, then with `_` standing in for `-`. An environment variable name cannot contain a `-` at all, and Windows upper-cases the name on the way in — so `SERVICES__MY_DB__CONTAINER__IMAGE` finds `services.my-db.container.image`. With `-c` you can simply write the name as it is. - **Only the winner is reported and recorded.** A path set by both a variable and a flag appears once, as the flag — listing the loser would read as if it had an effect. - **Overrides land on the expanded document**, after [includes](./writing-tests#composing-testfiles) and `foreach` — so they can reach into what those generated. - They are applied in variable-name order, so the same environment always produces the same document. `-c` is a `start` flag; `TESTFILE_CONFIG_` applies to every command that reads the file, so `testfile validate` and `testfile inspect` see the same document a run would. `testfile start --dry-run -c …` previews a flag's effect without running anything. `testfile start` and `testfile validate` both print what was overridden, because what ran is then not quite what the file says — and the run **records** it, so the file plus the record still explain the run afterwards. `testfile inspect run ` shows it back, and [`testfile repro`](#reproducing-a-failure) puts the same variables in the command it hands you. ## Checking the machine A Testfile says what a run needs: the tools its commands call, a container engine, fixed ports, a shell, git for [change-based selection](./writing-tests#change-based-selection) — and a run that finds one of them missing says so halfway through, in the log of whichever test happened to need it first. `testfile doctor` asks the same questions up front, and only the ones this file actually raises: ```sh $ testfile doctor ✔ Testfile /home/me/app/Testfile ✔ node v22.14.0 ✔ git git version 2.43.0 ✔ shell (sh) on PATH ✔ command (npm) /usr/local/bin/npm ✘ command (pytest) not found on PATH ↳ used by app/tests/unit - install it, or use a path to it ✘ command (./scripts/e2e.sh) /home/me/app/scripts/e2e.sh is missing or not executable ↳ used by app/tests/e2e - check the path, or chmod +x it ✘ container engine (podman) not installed ✘ container engine (docker) installed, but "docker info" fails: Cannot connect to the Docker daemon ↳ is the Docker daemon running? ✘ container engine (kubernetes) kubectl not installed ✘ engine selection this Testfile starts containers, but no engine responds ↳ install/start podman or docker, or point kubectl at a cluster ✘ port web 8080 is already in use ↳ stop what listens on 8080, or declare "web: random" and template the value ✔ .testfile/ /home/me/app/.testfile is writable 7 failed, 0 warning(s), 13 checks ``` What it looks at: | Check | Fails when | | ----- | ---------- | | `node` | the Node.js running the CLI is older than 20 | | `git` | *warns* when git is missing or the folder is not inside a work tree — only `--changed` and `testfile changes` need it | | `shell (…)` | a shell a test invokes (`sh` by default, or its `shell:`) cannot be started | | `command (…)` | an executable a `command:` starts is not on `PATH`, or the path it names is missing or not executable — see below | | `container engine (…)` | one row per engine — podman, docker and kubernetes are all checked, and the first responding one is marked as what [a run would use](./services#containers). An engine that is installed but not responding (daemon down, no reachable cluster) is a warning while another covers the run, a failure when it is the pinned or only one. Files without containers only get a note | | `engine selection` | the file starts containers but no engine responds, or `TESTFILE_ENGINE` pins an engine that does not | | `port …` | a fixed `ports:` entry is already taken. `random` ports are allocated per run and never clash | | `.testfile/` | recorded runs cannot be written | ### Which commands are looked up Every `command:` in the file — tests, `setup`/`teardown` hooks, services and their `ready.exec` probe — contributes the executable it starts. A line is split at `&&`, `||`, `;` and `|`, so `cd app && pytest -q` looks for `pytest`, and a leading `FOO=bar` assignment is stepped over. A name without a separator is looked up on `PATH`; anything with one is resolved against the test's `workdir` and must be an executable file. Four things are deliberately *not* looked up, because the answer would say nothing about whether the run works: - **shell builtins and keywords** (`cd`, `echo`, `test`, `for`, …), - **commands that only exist at run time** — a first word containing a `${{ … }}` template, a `$VAR`, a substitution or quotes, - **bodies that run in a container** (a `container:` on the test or an ancestor) and the [`ready.exec` probe of a container service](./services#where-exec-runs), which runs inside that container unless it sets `host: true`: those executables live in the image, not on this machine, - **`script:` blocks and tests with a custom `shell:`** — a shell program is not a command, and another shell has its own builtins. `--json` writes the same as `{status, checks: […]}` for scripts, and the exit code is `1` when a check failed — warnings alone keep it `0`, so `testfile doctor && testfile start` is a usable pre-flight. ## Labelling runs A recorded run can carry **labels**, so it can be found again in a history that mixes branches, pull requests and nightlies: ```sh testfile start --label branch=main --label tier=nightly testfile start -l pr=42 ``` `-l/--label` takes a `key=value` pair, split at the **first** `=` so a value may contain more, and is repeatable. Both halves are trimmed, and a key may only be given once — `-l branch=main -l branch=other` is an error rather than a silent choice between the two. They land in the run's `run.yaml` as a map of strings: ```yaml labels: branch: main pr: "42" ``` Values are always strings, including numeric-looking ones. Merging a set of runs keeps the union of their labels; where two runs disagree on a key the first one wins, because what actually differs between the legs of a matrix belongs in [`--variant`](#merging-runs). Viewers show them (`testfile inspect run `, the TUI's run detail, the web viewer's run detail) and both filter by them — [`runs --filter-label`](#run-history) on the command line, a Labels chip row in the browser. The [GitHub Action](./github-action#what-a-ci-run-is-labelled-with) labels every run it records with its branch, pull request, actor and how it was triggered. ## Machine-readable reports For CI systems, `--reporter` writes the run's result after it finished: ```sh testfile start --reporter junit --output results.xml testfile start --reporter json --output results.json testfile start --reporter json # ... or to stdout ``` The JUnit XML contains one `` per **leaf** test — its name is the last path segment, the parent path becomes the classname; groups aggregate, they don't get testcases — with `` elements carrying the merged log and `` markers; the JSON report is the same record that the run's `run.yaml` stores. In watch mode the report is rewritten after every re-run. ### Streaming events while the run happens `--reporter` speaks once, at the end. `--json-stream` speaks throughout: one JSON object per line (NDJSON) on stdout, written as the run happens, so a tool — or an agent supervising a long suite — can react to the first failure instead of waiting for a summary. ```sh testfile start --json-stream | jq -c 'select(.event == "test-end" and .status == "failed")' ``` | Event | Fields | | ----- | ------ | | `run-start` | `selected` (how many tests), `at` | | `test-start` | `path`, `kind` | | `line` | `path` *or* `service`, `stream` (`stdout`/`stderr`/`system`), `text` | | `test-end` | `path`, `status`, `durationMs`, `cached`, `reason`, `error` | | `service` | `name`, `status`, `error` — once per status change | | `run-end` | `status`, `exitCode`, `runId`, `counts` (per status, leaves only), `at` | Fields that don't apply are left out rather than sent as null, and new events and fields may be added — a consumer must ignore what it doesn't know. Service output is only streamed with `-v`, as in a normal run. Because stdout belongs to the stream, everything written for a human goes to stderr; combining it with a `--reporter` that also writes to stdout is an error, so give the report a file (`--output results.json`). What the inspection commands print is available as JSON too, through the same `--json [file]` flag: a file name writes there, the bare flag writes to stdout, so a command can be piped straight into `jq`. (The commands below all take it; action commands like `merge`, `serve` or the sync commands don't.) ```sh testfile inspect --json | jq '.tests[].path' # what a filtered run would execute testfile validate --json # {path, valid} - or the errors testfile tags --json # the tag inventory testfile changes --json # what --changed selects from testfile doctor --json checks.json # machine status of this machine testfile runs --json # the recorded runs testfile inspect run --json # one run record, in full testfile diff --json # what changed between two runs testfile s3 list --json # ... and github/gitlab list ``` ## Watch mode `testfile start --watch` (or `-w`) re-runs the current selection whenever a file in the project changes — combined with filters (`-w -t fast`) it makes a tight edit-test loop: ```sh testfile start -w -f unit ``` Changes are debounced, edits made while a run is in progress trigger one re-run afterwards, and `.git/`, `node_modules/` and `.testfile/` are ignored. Every re-run is recorded in the history like a normal run. Ctrl+C while idle exits with the last run's exit code; during a run it stops the run first. `path` may be a Testfile or a directory containing one (`Testfile`, `testfile.yaml` or `testfile.yml`); it defaults to the current directory. Exit codes: `0` all tests passed (or everything skipped) · `1` failures or a service that would not start · `130` interrupted. ## Filtering `start` and `inspect` accept filters to work on a subset of the suite: ```sh testfile start -f e2e # best guess: name, tag, ... testfile start -f fast # ... tag ... testfile start -f db:postgres # ... or matrix (it has a ":") testfile start --filter-name all/checks/unit # -n: match name/path only testfile start --filter-tags "slow, nightly" # -t: tagged slow OR nightly testfile start --filter-matrix db:postgres # -m: only these matrix instances testfile start -t slow -m db:postgres -m node:22 ``` - `-f, --filter ` is the quick, best-guess filter: a value containing `:` is treated as a `key:value` matrix filter; anything else matches tests whose path contains the value **or** that carry it as a tag. Repeatable. - `-n, --filter-name ` matches case-insensitively against the test's *path* — its names joined with `/`, e.g. `all/checks/unit tests` — so a bare test name works too. A matched test runs with all its nested tests; ancestors run as scaffolding (their sequence order, services and env still apply). Repeat the flag to match more tests. - `-t, --filter-tags ` takes a comma-separated list of [tags](./writing-tests#tags) (whitespace is trimmed) and keeps tests that carry — or inherit from an ancestor — any of them. The flag can be repeated. - `-m, --filter-matrix ` keeps only matrix instances whose combination has that value. Repeating the same key ORs the values, different keys are ANDed; tests outside a matrix with that key are unaffected. `--changed` runs only tests whose [inputs](./writing-tests#result-caching) match a file changed against the base branch — the committed diff since the fork point plus everything changed locally (staged, unstaged, untracked). It needs a git checkout, not a warm cache, so it works on fresh CI clones; see [change-based selection](./writing-tests#change-based-selection). Tests without `inputs` always count as changed; it composes with the other filters, works on `inspect` for preview, and errors when nothing is selected. `--changed-since ` picks the base branch (default: the remote's default branch) and implies `--changed`. Selected tests log which pattern matched how many changed files, and record it as `reason` in `run.yaml`. `--failed` re-runs what broke last time: it keeps only tests that failed (or were aborted) in the most recent recorded run, and combines with the other filters — `testfile start --failed -t integration` re-runs only the failed integration tests. Different filter kinds are ANDed. Filters that match nothing are an error. `testfile inspect` shows each test's tags, so it's an easy way to preview what a filter will run. ## Tags `testfile tags` inventories every tag of the full expanded suite — [included Testfiles](./writing-tests#composing-testfiles) and matrix instances included — so you know what `-t` can filter on: ```sh $ testfile tags # all tags, alphabetically fast integration nightly $ testfile tags --order appearance # in document order $ testfile tags --order count # most-used first, with counts 3 fast 2 nightly 1 integration 12 tests, 4 without any tag $ testfile tags --json # machine-readable, or --json tags.json ``` Counts count runnable tests (command/script leaves, every matrix instance separately) that carry the tag directly **or inherited from an ancestor** — the same semantics `-t` filters with. The count view also reports how many tests have no tag at all; the JSON export always includes both numbers. ## Sharding across machines Large suites split across CI machines with `--shard i/n`: every shard runs the same command with a different index and executes its share of the selected tests. ```sh testfile start --shard 1/4 # on machine 1 testfile start --shard 2/4 # on machine 2, ... ``` Each shard computes the same split from the same suite, so no coordination between machines is needed. When the local [run history](#run-history) knows how long tests took, the split is **time-balanced** (longest test first into the emptiest shard) instead of by count: ``` shard 1/3: 1 of 6 tests, ~611ms of recorded work shard 3/3: 3 of 6 tests, ~48ms of recorded work ``` On a fresh machine without history — the usual CI case — the tests are dealt round-robin in suite order. To get balancing there, restore `.testfile/` from a cache or import a previous run first ([`archive import`](#sharing-runs), [`github sync`](./github-action#bringing-ci-runs-home)). Sharding composes with the filters: `-t slow --shard 2/3` shards only the slow tests. Each shard records its own run — combine them into a single result with [`merge`](#merging-runs). ## Merging runs Sharding and a matrix of CI jobs both leave you with several run folders for what is conceptually one run. `testfile merge` combines them: ```sh # shards: each ran a different part of the suite testfile merge shard-1 shard-2 shard-3 # CI artifacts, each unpacked into its own folder testfile merge downloaded/testfile-run-* ``` The result is an ordinary run — one verdict, one duration, the union of the tests — that every viewer shows like any other: ``` merged run 20260805-101500-merged passed 20260805-101500-a1c3 [platform=ubuntu-latest] 1m12s failed 20260805-101501-c3e5 [platform=windows-latest] 1m45s failed (exit code 1), 24 tests, 2m57s ``` Nothing is hidden: the merged run records which runs went into it, and every test says which one it came from. ### Variants Shards merge as they are, because no test appears twice — the group nodes around them are folded into one entry rather than reported as a clash. Jobs that run the **same** suite in different places do not — the merge would not know which `ci/unit` came from where. Tell the runs apart with `--variant`: ```sh testfile start --variant platform=linux # on the Linux job testfile start --variant platform=windows # on the Windows job ``` Variants are free-form `key=value` pairs, recorded in `run.yaml` and shown by the CLI, the TUI and the web viewer. Merging requires `path` + `variants` to be unique and refuses the merge otherwise: ``` ✘ runs 20260805-101500-a1c3 and 20260805-101501-c3e5 both recorded "ci/unit" - give the runs distinct --variant values (e.g. --variant platform=linux) ``` The [three-platform guided tour](./guides/three-platforms) walks through the whole setup in GitHub Actions. ## Changes `testfile changes` shows exactly what `--changed` selects tests from: the base branch, the current commit, and an alphabetical table of every file changed in the git diff (base → HEAD, from the fork point) or locally: ```sh $ testfile changes base: origin/main (0cc2875ab) head: 6d5dd5a91 root: /work/project file source status runner/src/cli.ts diff modified runner/src/gitchanges.ts diff added package.json local modified 3 changed files ``` `source` says where a change was found: `diff` (committed since the base branch) or `local` (dirty in the working copy — a file that is both appears once, as `local`). Options: ```sh testfile changes --changed-since origin/release-2.0 # pick the base testfile changes --files # just the paths testfile changes --json changes.json # export as JSON testfile changes --json # ... to stdout ``` To debug why a specific test would (not) run, combine the file list with `testfile inspect --changed` or `testfile start --changed --dry-run`. ## Plain output `testfile start` streams progress line by line — suitable for CI logs. Test output is prefixed with `[test name]`; service output is shown with `--verbose`, and the tail of a failing service's log is always printed. A nested summary with per-test durations is printed at the end. ## The TUI `testfile tui` opens a read-only terminal UI over the recorded runs (it never starts tests — that is `testfile start`'s job). It is a multi-page interface, mirroring the [web viewer](#the-web-viewer): pages navigate forward, `esc` walks back, and a breadcrumb on top says where you are. The [screenshots](./screenshots#the-terminal-ui) page shows every one of its pages. The **index page** has two tabs, switched with `tab` (or `1`/`2`): 1. **Runs** — a table of every recorded run (started, run id, status, duration, passed/failed counts, the remaining statuses spelled out, and variants), filtered with `/`, taking the full terminal width. Enter (or a click) on a row opens that run's page. 2. **Tests** — two tables side by side. The left lists every test path that ever ran (plus an "All tests" row on top) and acts as the filter; the right lists the matching executions across all runs. `←`/`→` (or enter and `esc`) jump between the tables; enter or a click on an execution opens that test page. The **run page** shows the suite as a tree table on the left — the run's tests with status, duration and start offset, groups indented — and, for the selected row, a tab view on the right: *Overview* (the run or test metadata, ending with the last 20 lines of the log), *Log* (the merged run log, or the selected test's log) and one tab per related service log. The **test page** (one test in one run) shows the same tab view full width; both pages breadcrumb their way back — and walking back with `esc` lands exactly where you left: the same cursor row, scroll position and tab. Every table sorts: `s` cycles the sort column, `r` reverses it, and the header shows `▲`/`▼`. `↑`/`↓`, PgUp/PgDn, `g`/`G` and the mouse wheel move the cursor; clicking a row selects it. Log panes have a visible cursor line: `↑`/`↓` move it (the view scrolls with it), `shift+↑`/`↓` grow a selection, `ctrl-c` copies the selection to the clipboard (OSC 52 — the terminal has to allow it), `←`/`→` pan long lines, `w` toggles wrapping, and `/` searches (walk the hits with `n`/`N`). In tab views, `tab` cycles the tabs. The status line at the bottom always lists the shortcuts of whatever is focused; `?` opens an overlay with every shortcut on the page. `q` (or `ctrl-c`) quits. On terminals narrower than 80 columns the side-by-side panels collapse: only the left table is shown, and enter opens the details as their own page instead. Every page watches `.testfile/runs/` — runs recorded by other processes (say, a `testfile start` in a second terminal, or a `testfile github sync`) appear live. `--view tests` opens on the Tests tab (`--view results` still works as an alias). ## Run history Every `testfile start` is recorded in a `.testfile/` folder next to the Testfile (the folder ignores itself via a generated `.gitignore`). Each run is a self-contained folder: ``` .testfile/ runs// run.yaml # the run's record junit.xml # the run as JUnit XML, for CI tooling tests/.log # merged stdout+stderr per test services/.log # log of each started service artifacts//… # files collected via `artifacts:` ``` `run.yaml` stores the run's start time, duration, status (passed/failed/aborted), exit code, whether it was cancelled, the env variables and ports provided by the Testfile, which tests were selected, the [labels and variants](#labelling-runs) attached to the run, the Testfile's whole test tree (`suite`, including tests this run did not execute), the started services, and the status/duration/log of every test that ran. Each test also records **when** it started — `startedAt` as a timestamp and `startedAfterMs` as the distance from the start of the run — so a run can be laid out on a timeline without guessing. Run ids start with the run's UTC timestamp (second granularity — order runs by `startedAt` when it matters); the last 50 runs are kept and older run folders are pruned automatically. (Histories written by older runners as one `runs.yaml` index are migrated to per-run files on first use.) Browse the history from the command line: ```sh testfile runs # table of recent runs, newest first testfile runs --json # ... as JSON (or --json runs.json) testfile tui # browse runs in the TUI testfile inspect run 20260801-1046 # one run in detail (id prefix is ok) testfile inspect run --log # merged stdout+stderr of the run testfile inspect run --log all/e2e # ... of a single test testfile inspect run --json # the whole record, for a script testfile explain # what failed in the latest run, and why testfile repro all/e2e # everything needed to reproduce one failure ``` A history that collects runs from every branch and every CI job gets long, so `runs` narrows it: ```sh testfile runs --filter-status failed testfile runs --filter-label branch=main --filter-label branch=release testfile runs --filter-label pr # any run that has a pr label testfile runs --filter-variant platform=windows ``` Several values of one filter are an **OR** (`branch=main` *or* `branch=release` above), different filters an **AND**, and an unused filter narrows nothing. `--filter-label` takes a whole [label](#labelling-runs) or just its key — the key alone asks whether the run carries that label at all, which is how you find every pull-request run. `--filter-variant` also matches the legs of a [merged run](#merging-runs), so a merged matrix is found by any platform that went into it. The filters apply to the table, to `--json` and to the `--flaky` report alike, and the footer says how much survived (`3 of 12 runs`). The detail view lists every recorded test with status, when it started (`+2.3s` into the run), duration and whether a log is available, and ends with a `timeline:` block — one fixed-width bar per test, so a sequence reads as a staircase and a parallel group as a stack: ``` timeline: ci |████████████████████████| 0ms+3.2s ci/build |█ | 60ms+40ms ci/unit | ███████████████████████| 120ms+2.9s ``` `testfile` only needs the `.testfile/` folder, so it also works when the Testfile itself has moved or changed. Compare two runs (older id first, unique prefixes are enough): ```sh testfile diff 20260801-1040 20260801-1146 ``` The diff lists newly failed, fixed and still-failing tests, tests added to or removed from the run, and significant duration changes (more than 100ms and more than 20%) of tests that passed in both runs. `--json` writes the same lists as `{base, compare, newlyFailed, fixed, stillFailing, added, removed, durations}`, which is enough to post a comment from CI. ### Digesting a run `explain` answers the three questions a red run raises — what failed, why, and what changed — in one bounded piece of markdown: ```sh testfile explain # the latest run testfile explain 20260801-1146 # a particular one testfile explain --max-failures 3 --log-lines 10 testfile explain --json # the same digest, structured ``` Each failure carries its reason, the end of its log and what the history says about it — a test that fails half the time is a different problem from one that just broke, so the digest says `known flaky — 6/12 of its recent results failed` rather than only `failed`. The verdict is the same [flaky rule](#run-history) the rest of the tooling uses. A group fails because something under it failed, so the leaves come first and the groups are the first to go when the digest has to be shorter: `--max-failures` bounds how many failures are detailed (10 by default), `--log-lines` how much log each one gets (20). What is left out is counted, never silently dropped. Log excerpts are stripped of colour — in a PR comment or a prompt, escape sequences are noise. The run before this one is compared automatically, so the digest opens with `newly failing` / `still failing` / `fixed` before it gets to the detail. ### Reproducing a failure A red test in CI is a puzzle assembled from several places: which run, which leg of a matrix, what environment, which services were up, what the log said. `repro` puts it in one place — and gives the command that reruns exactly that test, not the whole suite: ```sh testfile repro 20260801-1146 ci/e2e testfile repro ci/e2e --variant platform=windows # one leg of a merged run testfile repro ci/e2e --json # for a tool or an agent ``` ``` # reproduce ci/e2e from run 20260801-114600-9f2c status: failed (cache miss: src/**: 1 changed file) recorded: 2026-08-01T11:46:00.000Z on ci-linux labels: branch=main, pr=42 matrix: browser=firefox tags: ci, slow run it with: export DATABASE_URL=postgres://localhost:5432/test testfile start -n ci/e2e -m browser:firefox services this test needs (status in the recorded run): db — stopped the end of its log: boom: expected 4 to equal 5 ``` Everything comes from the run's own record — the viewer never reruns anything and never guesses, so what the run did not record does not appear. The environment leaves out what every run sets anyway (`CI`, `FORCE_COLOR`, `TESTFILE_OS`): what is left is what was special about this one. On a [merged run](#merging-runs) a path has one result per leg; `--variant` picks one, and without it a failing leg is chosen, since that is the one worth reproducing. Hunt down flaky tests: ```sh testfile runs --flaky # across the recorded history testfile runs --flaky --last 10 # narrowed to the 10 most recent runs ``` What a test did fifty runs ago says nothing about it now, so the verdict is decided on the **20 most recent** results per test — and below **10** there is not enough evidence to say anything at all. The same rule runs in the CLI, the TUI and the web viewer: | Failure rate of the sample | Verdict | | --- | --- | | fewer than 10 results | *no verdict* | | below 25% | healthy | | 25% to 75% | **flaky** | | above 75% | **broken** | Both verdicts are reported: a broken test is not flaky — it fails almost every time — but it is just as untrustworthy, and the report's flip count (how often the outcome changed between consecutive results) tells the two apart at a glance. The report also shows the failure rate over the sample and the latest status. Flagged tests are good candidates for a `flaky` tag and a [`retry`](./writing-tests#retries). `skipped` and `aborted` results are not evidence and never enter the sample: a skip says the test never ran, and one Ctrl+C aborts everything in flight, which would otherwise make a whole suite look flaky at once. Age is not a criterion — a long-untouched project keeps its verdicts, and `--last ` narrows the history when you only care about recent runs. ## Sharing runs Because every run is a self-contained `runs//` folder, runs can move between machines. `testfile archive` packs them as `.tgz` archives and brings them into the local history, where `runs`, `inspect run`, `diff`, `--flaky` and the TUI treat them like local runs: ```sh testfile archive pack # latest run -> testfile-run-.tgz testfile archive pack --run 20260801 -o ci.tgz testfile archive import ci.tgz # import into ./.testfile/runs/ testfile archive import testfile-run.zip # a downloaded GitHub run artifact ``` Importing skips runs that already exist locally (same id), so repeated imports are safe. Packing and importing shell out to `tar`, and importing a zip (a CI artifact) additionally needs `unzip` on the `PATH`. Both are a given on Linux and macOS; on Windows `tar` ships with the system but `unzip` does not, so zip artifacts need it installed. Everything else in the viewer — `runs`, `inspect run`, `diff`, `tui`, `serve` — is pure Node. With the [aws CLI](https://aws.amazon.com/cli/) configured, runs can be shared through S3 — for example a CI job pushes, developers pull: ```sh testfile s3 push s3://my-bucket/testfile-runs # latest run testfile s3 push s3://my-bucket/testfile-runs --run 20260801 testfile s3 list s3://my-bucket/testfile-runs # available archives testfile s3 pull s3://my-bucket/testfile-runs # newest archive testfile s3 pull s3://my-bucket/testfile-runs --run ``` And when CI is the [GitHub Action](./github-action) (which uploads every recorded run as a `testfile-run` artifact), `sync` pulls the artifacts of the latest *n* workflow runs straight into the local history. The artifact name is a **prefix**, so a [matrix over platforms](./guides/three-platforms) — `testfile-run-ubuntu-latest`, `testfile-run-macos-latest`, `testfile-run-merged` — comes along without naming each leg: ```sh export GITHUB_TOKEN=... # a token with actions:read # (GH_TOKEN works too) export GITHUB_TOKEN=$(gh auth token) # ... or reuse the gh CLI's login testfile github list owner/repo # available run artifacts testfile github sync owner/repo # latest 100 workflow runs testfile github sync owner/repo --latest 20 testfile github sync owner/repo --artifact testfile-run-merged testfile github sync owner/repo --artifact testfile-run --exact ``` Already-imported runs are skipped, so `sync` is incremental — run it again any time to top up the local history with the newest CI results. The TUI's [runs and tests views](#the-tui) pick imported runs up live. A sync narrates what it does as it works: what it is listing, how many artifacts it found (and their size, when the API says), and a `[3/12] testfile-run-macos-latest (workflow run …) …` progress line per download that tells what each artifact yielded — on a terminal the in-flight line updates in place, piped output gets plain lines. The same narration covers `gitlab sync` and `s3 pull`; the final summary of imported and skipped runs is unchanged. ## The web viewer `testfile serve` starts a small web UI over the recorded runs — the browser sibling of the TUI, view by view on the [screenshots](./screenshots#the-web-viewer) page: ```sh testfile serve # http://127.0.0.1:7357 testfile serve --port 8080 ``` - **Runs**: a table of all recorded runs; selecting one shows its details, the suite tree with this run's results on it, and the logs (merged or per test). - **Tests**: every recorded test with aggregated pass/fail counts and its executions across all runs; clicking an execution opens its own page — that test in that run, with an overview, the test's log and one tab per related service log. On a merged run the page shows every leg: one `Test log (platform=linux)`-style tab per leg (services too, when the legs recorded them), and the overview repeats the run's labels and ends with the last 20 lines of each leg's log — where a failure usually says why. - **Logs** read like logs: the colour a tool wrote (the runner asks for it — see [an isolated environment](./env-and-ports#an-isolated-environment)) is rendered rather than printed as escape sequences, and every log has a `find in log` box with `‹ ›` to walk the hits, a `wrap` toggle (on by default) and a `follow` toggle that pins the view to the end while a run is still being written. - The server watches `.testfile/runs/` and pushes changes to the browser, so runs recorded elsewhere (another terminal, `testfile github sync`) appear live. - The page follows the system theme — dark and light are both first-class (without a preference it stays dark). The logs theme too: the recorded ANSI colours are mapped onto a palette per theme, so a green test line is readable on either background. Above the tree, the run detail draws a **timeline**: one bar per test on a single axis, from the start of the run to its end, coloured by outcome — the shape of the run rather than a list of durations. A sequence reads as a staircase, a parallel group as a stack, and a merged run shows what its legs did at the same moment (`merge` recomputes each leg's offsets against the merged start, so one axis holds them all). Clicking a bar opens that test's log. Records written before the runner timed its tests simply have no timeline. The run detail draws the **suite tree** the record carries: every node of the Testfile with its kind, tags, matrix combination and declared services, indented, with the results of this run on it. Groups collapse (and `collapse all` / `expand all` do the lot), a merged run shows one line per leg under its node, and a test the run never reached — filtered out, skipped by a condition, or never started because something before it failed — keeps its place, greyed and marked `not run`. Records written before `suite` existed fall back to the tree their test paths imply. Any row with a log is clickable as a whole — the `show` link is just its label — and opens that test's log below the tree. Both tables have a filter bar above them. Nothing is selected in the multi-selects to begin with, which shows everything; the only default that narrows anything is the time window: | Filter | Applies to | Default | | ------ | ---------- | ------- | | **Started** | runs — `7 days`, `30 days`, `90 days`, `all` | last **30 days** | | **Status** | runs / tests, multi-select (several values are an OR) | everything | | **Labels** | runs, multi-select over the recorded [labels](#labelling-runs) (`branch=main`-style); a key with more than 3 values becomes a dropdown instead of a chip per value | everything | | **Variants** | runs, multi-select over `platform=linux`-style labels; a merged run matches when *any* of its legs does | everything | | **Tags** | tests, multi-select over the tags of the recorded [suite tree](https://github.com/testfile-dev/testfile/blob/main/spec/RESULTS.md) — nested tests inherit the tags of their groups | everything | | **flaky only** | tests, an on/off chip: keeps only tests the [flaky rule](#run-history) calls flaky — 25% to 75% of their last 20 results failed. Broken tests are badged but not matched by this chip | off | | **Search** | free text over run ids, test paths, statuses and variant labels | empty | Every column of every table sorts: the header is a button, clicking it flips between ascending and descending, and the arrow shows which column the table is ordered by. The runs table opens newest first, the tests table by path, and each table remembers its own order — the suite tree is the one exception, because sorting a tree would take it apart. The count on the right says how much survived (`4 of 27 runs`) and clears the filters again. A run or test opened by link stays visible even when the filters would hide it — the link should not silently open something else. The two views the CLI already had are on the same pages. In **Tests**, each test carries a **history sparkline** — one block per recorded run, newest on the right, so a test that alternates green and red looks different from one that simply broke — and a `flaky` or `broken` badge next to the tests `testfile runs --flaky` would list, decided by exactly the same rule. A test with fewer than 10 results gets no badge at all. The sparkline still draws the whole history, so the badge and the blocks can legitimately disagree about an old bad patch. The `flaky only` chip narrows the table to the flaky ones. In **Runs**, a run detail has a *compare with* picker: choose another recorded run (or press `previous run` for the one recorded before this one) and the same six sections `testfile diff a b` prints appear above the suite tree — newly failed, still failing, fixed, added, removed, and durations that moved by more than 100ms *and* more than a fifth. In a merged run the worst leg of a path decides, exactly as the run's own verdict does. Every selection is in the URL, so a view can be linked, bookmarked and reloaded — and the browser's back button walks through it: | Path | Shows | | ---- | ----- | | `/` or `/runs` | the runs table with the newest run | | `/runs/` | that run's detail | | `/tests` | the tests table | | `/tests/` | that test's executions — the test path keeps its slashes, so the URL reads like the test does | | `/runs//tests/` | one execution: that test in that run, as its own page | (The tests tab used to be called *results*; `/results/...` links keep working.) An id or test path that no longer exists falls back to the newest run (respectively the first test) instead of an error page. Whatever a run kept can be opened from the page: the artifacts of a test are links in its row, and the run detail links `run.yaml` — the record the page was built from — and `junit.xml` when the run wrote one. They all go through one endpoint, addressed exactly as `run.yaml` records the path: ``` /api/runs//artifacts/artifacts/ci-unit/report.txt /api/runs//artifacts/junit.xml /api/runs//artifacts/run.yaml ``` It reads from that one run folder and nowhere else: the id is checked as it is everywhere else, each path segment is decoded and rejected if it is `.`, `..`, or hides a separator, and the resolved path has to still be inside the folder. Nothing is served as HTML or JavaScript, so a recorded artifact can never run as a page on the viewer's own origin. The server binds to `127.0.0.1` **only** — it is never reachable from the network. It exposes a read-only REST API for other tooling: `/api/summary`, `/api/runs`, `/api/runs/`, `/api/runs//log` (`?test=` for one test), `/api/runs//artifacts/`, `/api/results` and `/api/events` (SSE). The UI itself lives in the `testfile-viewer/` workspace (React, built with Vite); `serve` picks up its build automatically. ## Talking to an AI assistant `testfile mcp` serves the recorded runs over the [Model Context Protocol](https://modelcontextprotocol.io), so an assistant that speaks MCP — Claude Code, Claude Desktop, an agent of your own — can read the history as data instead of parsing terminal output. The server also ships on its own as [`@testfile.dev/mcp`](https://www.npmjs.com/package/@testfile.dev/mcp), so a config can point at `npx` and install nothing but the server and the core reader — no `testfile` on the `PATH` needed: ```jsonc // .mcp.json, or Claude Desktop's config { "mcpServers": { "testfile": { "command": "npx", "args": ["-y", "@testfile.dev/mcp", "/path/to/your/project"] } } } ``` ```jsonc // ... or through an installed testfile CLI { "mcpServers": { "testfile": { "command": "testfile", "args": ["mcp", "/path/to/your/project"] } } } ``` | Tool | Answers | | ---- | ------- | | `list_runs` | recent runs with their status and counts; narrows by status, label or variant | | `get_run` | one run's full record | | `explain_run` | [the digest](#digesting-a-run): what failed, why, what changed — where to start when a run is red | | `repro_test` | [the repro bundle](#reproducing-a-failure) for one failure | | `get_test_log` | one test's log, whole or its tail | | `diff_runs` | what changed between two runs | | `list_tests` | every known test with its pass/fail counts and [verdict](#run-history) | | `list_flaky` | the tests the flaky rule flags | **Everything here reads.** There is deliberately no `run_tests` tool: the viewer does not run tests, and an assistant that wants to is already holding a shell — `testfile start -n `, or [`--json-stream`](#streaming-events-while-the-run-happens) to follow a run live. Keeping the server read-only means connecting it can't change anything. The transport is stdio, the history is re-read on every call (a run recorded while the assistant is connected shows up without a restart), and a tool that can't answer says so as a result the model can read rather than as a protocol error it can't. ### Skills Knowing the commands is not the same as knowing which one to reach for. This repository ships three [Claude Code skills](https://code.claude.com/docs) in `.claude/skills/`, which an assistant working in a checkout picks up automatically: | Skill | For | | ----- | --- | | `testfile-triage` | a red run: which failure is real, what caused it, whether it's worth chasing | | `testfile-run` | choosing a selection — `--changed` after an edit, `--failed` after a fix, the whole suite before declaring done | | `testfile-author` | writing or extending a Testfile, including which key replaces a shell workaround | Copy the folder into any project that uses Testfile; nothing in them is specific to this repository. They are checked by the suite: every command and flag a skill names is verified against the CLIs' own `--help`, so a renamed flag fails the build instead of turning the advice into confident nonsense. ### The documentation, for a model The website publishes itself in the [llms.txt](https://llmstxt.org) shape, so an assistant can read the documentation without scraping HTML: | File | Contents | | ---- | -------- | | [`/llms.txt`](https://testfile.dev/llms.txt) | one line per page — title, link and what the page answers — so a model can fetch the two pages it needs | | [`/llms-full.txt`](https://testfile.dev/llms-full.txt) | every page's markdown, concatenated, for reading the lot at once | Both are generated from the same content the pages are built from, and the build fails if a published page is missing from the index — an index that quietly omits a page is worse than none, because whoever reads it believes it was complete. ### Recording what a failure meant A run says what happened. It cannot say what it *meant* — that three failures share one cause, that the flake was a port collision, that the red was infrastructure and not the change. Whoever works that out can write it into the record, as an optional [`analysis`](https://github.com/testfile-dev/testfile/blob/main/spec/RESULTS.md#analysis) field on `run.yaml`: ```yaml analysis: text: | All three failures come from the port 5432 collision in the parallel group. Not caused by this change. author: claude-code at: 2026-08-01T12:10:00.000Z ``` `testfile inspect run`, `explain`, the TUI and the web viewer all show it next to the run — **marked as somebody's reading of it, never as a result**: a run whose analysis says "this is fine" is still a failed run. No runner writes the field and no viewer does either (they are read-only); whoever did the reading writes it, preserving the rest of the record. [Letting an assistant read the failure](./github-action#letting-an-assistant-read-the-failure) shows the CI shape of that loop. ## Interrupting a run The first Ctrl+C aborts running tests and shuts down all services through their configured `stop` behavior (signal, grace period, then SIGKILL; `podman stop`/`docker stop` for containers). A second Ctrl+C skips the grace period and kills everything immediately. # Screenshots Source: https://testfile.dev/docs/screenshots Both viewers read the same thing — the runs recorded under `.testfile/` (see the [CLI & TUI](./cli) page) — and neither of them ever starts a test. `testfile serve` opens the web viewer in a browser, `testfile tui` the terminal one; which one is nicer depends on where you are, not on what you can see. Every picture below is a committed screenshot from the test suites: the browser ones are the Playwright shots the viewer's suite compares against, the terminal ones are rendered from the exact frames the TUI's own suite pins. Nothing here is a mockup, and neither UI can change without these changing with it. Each suite has its own fixed little fixture, and both are recordings of a `ci` suite: three runs, older passing ones and a newer failing one, a cached test, a failing assertion, two services and a run carrying platform variants — the web viewer's fixture has that one as a run merged from a Linux and a Windows leg. ## The web viewer `testfile serve` (see [the web viewer](./cli#the-web-viewer)). ### Runs The history on the left, the selected run on the right — and that one page is most of the viewer: - the run's **labels** (who triggered it, from which branch, for which pull request) and the **files** it kept: the record the page was built from (`run.yaml`) and the reports next to it (`junit.xml`), - a **timeline** of the run's wall clock, one bar per test at its start offset, scaled to its duration and coloured by status; clicking a bar opens that test's log, - the **suite tree** with this run's results on it — kinds, tags, declared services, the cache marker, an [artifact](./writing-tests#artifacts) linked from the row of the test that produced it, and what did *not* run: `ci/e2e` is in the Testfile but was never reached, and says so, - the **services** of the run in their own table, each with its log, - and the **merged log** of everything, at the bottom. ![The web viewer's runs view: the runs table on the left, the newest run on the right with its labels, timeline, suite tree, services and merged log](../testfile-viewer/e2e/__screenshots__/runs-light.png) ### A single test's log Selecting a row in the tree — or a bar on the timeline — replaces the merged log with that test's own output. ANSI colour a tool wrote is rendered rather than printed as escape sequences. ![The run detail with a single test selected, showing only that test's log](../testfile-viewer/e2e/__screenshots__/test-log-light.png) ### Searching a log `find in log` highlights every hit and says which one is current; `‹ ›` walk them and wrap around at the end. Next to it sit a `wrap` toggle (on by default) and a `follow` toggle that pins the view to the end while a run is still being written. ![A log with the search box filled in: every match highlighted, the current one marked, and a "1 of 2" hit counter](../testfile-viewer/e2e/__screenshots__/log-search-light.png) ### Tests The other tab looks at the history the other way round: one row per test path with its last status, a sparkline of the recent results and the pass/fail counts across every run. Selecting one lists its executions. ![The tests view: every test path with last status, a history sparkline and pass/fail counts, and the executions of the selected test](../testfile-viewer/e2e/__screenshots__/tests-light.png) ### One execution An execution opens its own page — that test in that run — with a breadcrumb back out and tabs across: the overview (its metadata, including *why* it ran or was served from the cache), the test's log, and one tab per related service log. ![The page of a single test execution: breadcrumb, tabs, and the test's log open](../testfile-viewer/e2e/__screenshots__/test-page-light.png) ### A merged run Runs [merged](./guides/three-platforms) from several legs say what they were merged from and show the same test once per variant. ![A merged run: the legs it combines listed at the top, and each test once per platform variant](../testfile-viewer/e2e/__screenshots__/merged-run-light.png) ### Labels Every label of every run is a chip in the filter bar, so a branch, a trigger or a pull request narrows the table to its runs — and the text box searches them too. ![The runs view filtered by a label chip, with the selected run's labels shown as badges](../testfile-viewer/e2e/__screenshots__/labels-light.png) ### Comparing two runs A run can be compared with the one before it, or with any other from the list: what newly failed, and where a duration moved — `2.2s → 40ms` for a test that came out of the cache this time. ![A run compared with the previous one: newly failed tests and the durations that changed](../testfile-viewer/e2e/__screenshots__/run-diff-light.png) ### Sorting Every column of every table sorts, both ways, and each table keeps its own sort. ![The runs table sorted by duration instead of by start time](../testfile-viewer/e2e/__screenshots__/sorted-light.png) ### Flaky and broken Tests that pass and fail on the same input get a verdict — but only once there are enough results for it to mean anything, which this three-run fixture does not have, so its `flaky only` filter comes back empty. ![The tests view with the "flaky only" filter active and no test matching it](../testfile-viewer/e2e/__screenshots__/flaky-light.png) ## The terminal UI `testfile tui` (see [the TUI](./cli#the-tui)). The status line at the bottom always lists the shortcuts of whatever is focused. ### Runs The index page's first tab: every recorded run, full width, sorted with `s`/`r` and filtered with `/`. Enter (or a click) opens the run. ![The TUI's runs tab: a table of three recorded runs with status, duration and pass/fail counts, the cursor on the newest one](../testfile-ts/tui/screens/index-runs-light.png) ### Tests The second tab, switched with `tab` (or `1`/`2`): every test path that ever ran on the left, with its number of runs, its failures and its verdict — and on the right the matching executions across all runs. `←`/`→` jump between the two tables. ![The TUI's tests tab: test paths on the left, the executions of the selected one on the right](../testfile-ts/tui/screens/index-tests-light.png) ### A run The run page: the suite as a tree table on the left, and for the selected row a tab view on the right — overview, log, and one tab per related service log. ![The TUI's run page: the suite tree on the left, the overview of the selected row on the right](../testfile-ts/tui/screens/run-light.png) ### One test The test page (one test in one run) shows the same tab view full width, and breadcrumbs its way back — `esc` lands exactly where you left. ![The TUI's test page: breadcrumb, tabs and the overview of a single test execution](../testfile-ts/tui/screens/test-light.png) ### A log `tab` cycles to the log. It has a visible cursor line that the view scrolls with, `shift+↑`/`↓` grow a selection, `ctrl-c` copies it (OSC 52), `w` toggles wrapping and `/` searches — `n`/`N` walk the hits. ![The TUI's test page with the log tab open and the cursor line on the first line of the log](../testfile-ts/tui/screens/test-log-light.png) ### Every shortcut `?` opens an overlay with every shortcut of the current page. ![The TUI's shortcut overlay, listing the keys of the current page](../testfile-ts/tui/screens/shortcuts-light.png) ### Narrow terminals Below 80 columns the side-by-side tables collapse: only the left panel shows, and details open as their own page. ![The TUI's tests tab in a 72-column terminal: a single panel instead of two](../testfile-ts/tui/screens/index-tests-narrow-light.png) ## Dark and light Both viewers follow their surroundings — the web viewer the system theme (without a preference it stays dark), the TUI the terminal's palette. The logs follow too: the recorded ANSI colours are mapped onto a palette per theme, so a green test line is readable on either background. ![The web viewer's runs view in the dark theme](../testfile-viewer/e2e/__screenshots__/runs-dark.png) ![The TUI's run page in a dark terminal](../testfile-ts/tui/screens/run-dark.png) Every screenshot on this page exists in both themes, in [`testfile-viewer/e2e/__screenshots__/`](https://github.com/testfile-dev/testfile/tree/main/testfile-viewer/e2e/__screenshots__) and [`testfile-ts/tui/screens/`](https://github.com/testfile-dev/testfile/tree/main/testfile-ts/tui/screens) — the latter also keeps the raw frames (`.ans`) and the SVGs the PNGs are rendered from. # CLI reference Source: https://testfile.dev/docs/cli-reference The complete list of commands, arguments and options of `testfile`. The command line has two halves: [running what a Testfile describes](#running-the-suite), and [reading the runs that came out](#reading-the-runs). For guides and examples see [CLI & TUI](./cli); this page is the dry inventory. Conventions used below: - `[path]` defaults to `.` everywhere. For the suite commands it is a Testfile or a directory containing one (`Testfile`, `testfile.yaml` or `testfile.yml`); for the history commands it is a directory containing a `.testfile` folder. - *(repeatable)* flags can be passed multiple times. - Every command also accepts `-h, --help`; the binary accepts `-V, --version` and `help [command]`. ## Running the suite ``` testfile [command] [options] [path] ``` Running without a command is the same as `testfile start`. Exit codes: `0` all tests passed (or everything skipped) · `1` failures or a service that would not start · `130` interrupted. ### `testfile start [path]` Start the test suite (the default command). | Option | Description | | ------ | ----------- | | `-v, --verbose` | Also stream service output. | | `--fail-fast` | Abort the whole run at the first test failure. | | `--max-parallel ` | Global cap on concurrently running tests (group-level `maxParallel` still applies on top). | | `--dry-run` | Print what would run — with filters applied and predicted [cache](./writing-tests#result-caching) hits marked `[cached]` — without running. | | `-w, --watch` | Re-run the selection whenever files change ([watch mode](./cli#watch-mode)). | | `--no-cache` | Ignore cached results; fresh results still refresh the cache. | | `--forward-env ` | Forward matching host env vars into the [isolated test env](./env-and-ports#an-isolated-environment), e.g. `"GITHUB_*"` or `"*"`. *(repeatable)* | | `-c, --config ` | [Override a value](./cli#overriding-the-testfile-for-one-run) in the Testfile for this run, e.g. `ports.db=15432`. Beats a `TESTFILE_CONFIG_*` variable naming the same path; both beat the file. *(repeatable)* | | `--engine ` | Container engine for this run: `podman`, `docker` or `kubernetes`. Default: `$TESTFILE_ENGINE`, else the first of the three that responds. The [Testfile itself never names one](./services#containers). | | `--variant ` | Record what distinguishes this run from a sibling run — e.g. `platform=linux` for one leg of a matrix. Recorded in `run.yaml` and used by [`testfile merge`](#testfile-merge-run). *(repeatable)* | | `-l, --label ` | Record a label with the run, e.g. `branch=main`, so it can be [found again later](./cli#labelling-runs). Split at the first `=`; a key may only be given once. *(repeatable)* | | `--reporter ` | Write [machine-readable results](./cli#machine-readable-reports) after the run: `junit` or `json`. | | `--output ` | Report target file, or `-` for stdout (the default). | | `--json-stream` | Stream [NDJSON events](./cli#streaming-events-while-the-run-happens) to stdout while the run happens — `run-start`, `test-start`, `line`, `test-end`, `service`, `run-end`. Human output moves to stderr. | Plus the shared [filter options](#shared-filter-options-start-and-inspect) below. ### `testfile inspect [path]` Print the expanded test suite — matrix instances, tags and services included. Takes the shared [filter options](#shared-filter-options-start-and-inspect), so it previews exactly what a filtered `start` would execute. | Option | Description | | ------ | ----------- | | `--json [file]` | Write the suite as JSON (`{path, services, count, tests: [{path, name, kind, tags?, matrix?, services?}]}`), to a file or (without a value) stdout. | ### Shared filter options (`start` and `inspect`) | Option | Description | | ------ | ----------- | | `-f, --filter ` | Best-guess filter: `key:value` is a matrix filter, anything else matches name/path or tag. *(repeatable)* | | `-n, --filter-name ` | Only tests whose path contains this (case-insensitive). *(repeatable)* | | `-t, --filter-tags ` | Only tests tagged — directly or inherited — with any of these comma-separated [tags](./writing-tests#tags). *(repeatable)* | | `-m, --filter-matrix ` | Only matrix instances with this value; same key ORs, different keys AND. *(repeatable)* | | `--failed` | Only tests that failed (or were aborted) in the last recorded run. | | `--changed` | Only tests whose `inputs` match files [changed against the base branch](./writing-tests#change-based-selection), plus local changes. | | `--changed-since ` | Base branch/ref for `--changed`, e.g. `origin/main` (implies `--changed`). | | `--shard ` | Run only this shard of the selected tests, e.g. `2/4`. Time-balanced from the [run history](./cli#run-history) when at least half the selected tests have a recorded duration, round-robin otherwise. | ### `testfile tags [path]` List all [tags](./cli#tags) of the full suite, including [included Testfiles](./writing-tests#composing-testfiles). | Option | Description | | ------ | ----------- | | `--order ` | `alpha` (default), `appearance` (document order) or `count` (most-used first; also reports how many tests have no tag at all). | | `--json [file]` | Write the tag inventory as JSON, to a file or (without a value) stdout. | ### `testfile changes [path]` Show the files [changed against the base branch](./cli#changes) — what `--changed` selects tests from. `path` is the directory (or Testfile) whose git repository to inspect. | Option | Description | | ------ | ----------- | | `--changed-since ` | Base branch/ref to diff against (default: auto-detected from `origin/HEAD`, then `origin/main`, `origin/master`, `main`, `master`). | | `--files` | Print only the file paths, one per line. | | `--json [file]` | Write the changes as JSON, to a file or (without a value) stdout. | ### `testfile validate [path]` Validate a Testfile against the JSON schema (see [editor support](./getting-started#4-editor-support) for live validation while writing). Exits `1` when the file is rejected. | Option | Description | | ------ | ----------- | | `--json [file]` | Write the result as JSON (`{path, valid}`, plus `message` and one `errors` entry per schema violation when invalid), to a file or (without a value) stdout. | ### `testfile init [path]` Create a starter Testfile in the given directory, derived from what the project already has: `package.json` scripts plus [imported](./getting-started#1-create-a-testfile) docker-compose services, GitHub workflow steps and Make/Task/just targets. | Option | Description | | ------ | ----------- | | `--from ` | Import this file instead of the auto-detected ones: a docker-compose file, a GitHub workflow, a `Makefile`, a `Taskfile` or a `justfile`. Repeatable. | | `--no-detect` | Do not look for importable files automatically. | ### `testfile doctor [path]` Check this machine against what the Testfile needs, before a run finds out the hard way: Node.js version, git (and whether the folder is inside a work tree), every `shell:` the tests invoke, every executable a `command:` starts (on `PATH`, or as a relative/absolute path), the container engines when the file starts containers — podman, docker and kubernetes are all checked, and the report says which one [a run would pick](./services#containers) — the fixed `ports:` and a writable `.testfile/`. Exits `1` when a check fails; warnings (a missing git, for instance) do not. | Option | Description | | ------ | ----------- | | `--json [file]` | Write the checks as JSON (`{status, checks: [{name, status, detail, hint}]}`), to a file or (without a value) stdout. | ### `testfile export [path]` Write a native CI pipeline (or a bash script) that runs this suite [without the runner](./export): `` is `github` (`.github/workflows/testfile.yaml`), `gitlab` (`.gitlab-ci.yml`), `bash` (`testfile-ci.sh`) or `tekton` (`.tekton/testfile-pipeline.yaml`). The [export page](./export) documents how suites map to jobs and which features each target supports. | Option | Description | | ------ | ----------- | | `-f, --filter ` | Export only tests matching by name/path, tag, or `key:value` matrix. *(repeatable)* | | `-n, --filter-name ` | Only tests whose path contains this. *(repeatable)* | | `-t, --filter-tags ` | Only tests tagged with any of these comma-separated tags. *(repeatable)* | | `-m, --filter-matrix ` | Only matrix instances with this value. *(repeatable)* | | `-l, --label ` | Write a `TESTFILE_LABEL_*` variable into the pipeline. *(repeatable)* | | `-o, --output ` | Where to write the pipeline (default depends on the target), or `-` for stdout. | | `--image ` | Container image for GitLab jobs and Tekton steps (tekton default: `node:22`). | | `--runs-on