ci/runbook.md
Use only when asked to validate Test Optimization in this repository. Discover one small command per runner shape, complete the manifest, show the validator plan, run it after one approval, and report the diagnosis. Never modify the project to make validation pass. Applying fixes is separate work.
The validator uses private filesystem fixtures and bounded artifacts. It opens no listener, contacts no Datadog endpoint, and needs no Agent/API key. Project commands may need normal test permissions.
Use the schema/validator beside this installed runbook. At the repository root, record
git status --short as a cleanup baseline and preserve existing changes. Discovery is read-only: no
installs, setup, tests, or runner/package --version.
Inspect CI before scripts, explicitly including hidden .github/workflows/* and present GitLab,
CircleCI, Buildkite, Bitbucket, Azure, or Jenkins config. For each test job record its location, exact
command, cwd/shell/env, matrix, setup, script/runner chain, inheritance, services, and unresolved data.
Keep secret names only; executable values use dd-validation-placeholder.
Select one small representative per distinct framework/cwd/setup/wrapper/CI-env shape and record
duplicates as omissions. Include non-runnable runners with reasons; reporters are not runners.
Use CI evidence to select a focused unit test and fallback, but do not copy the CI package-manager wrapper into
existingTestCommand solely to resemble CI; keep the scaffold's direct installed runner when it preserves the
selected test's required config and setup. Avoid watch, benchmark/typecheck, snapshot-update, golden,
generated-list, export-matrix, and broad commands. Confirm filters narrow.
Seek service-free tests before builds/Docker/databases/browsers. Respect pinned runtimes/managers and
invoke pinned Yarn as node .yarn/releases/yarn-*.cjs .... When package.json requires Yarn 2 or newer
without a checked-in yarnPath, use an explicit corepack yarn ... command instead of ambient bare
yarn; the plan rejects an ambiguous ambient Yarn entrypoint. Record custom Jest runners; never use a
test-runner repository's unpublished in-repository runner implementation as evidence for the corresponding
published runner instrumentation. A project-owned wrapper around an installed supported runner is eligible
when a focused test can run; preserve the wrapper-to-runner chain and use the wrapper for CI replay. Vitest
setupFiles initialization is too late: CI must preload dd-trace/ci/init.
Before marking a command runnable, inspect its runner config and package-script expansion for local setup files, transforms, module mappings, custom environments/runners, and build outputs needed before test discovery. Confirm every statically referenced local input exists. Bypassing a package build wrapper does not make its outputs optional. If an input is missing, select another representative or record the exact setup blocker; do not defer an already-known failure to the approved live run.
Basic Reporting checks a real test with validator-applied initialization. CI wiring then checks whether the CI-shaped command carries its own initialization to the final runner. Basic Reporting never proves CI wiring. Live replay is authoritative when available; static/probe evidence only explains it. Unsafe/unavailable replay is incomplete or blocked, not a live failure.
If the Datadog run exits differently from its clean preflight, the approved validator reruns the same command once without Datadog. A changing clean result is an unstable baseline and remains inconclusive. If both clean runs agree but only the Datadog run fails, report a possible dd-trace compatibility problem; never call the failure pre-existing unless a clean run reproduces it.
Create the static network-free scaffold:
node ./node_modules/dd-trace/ci/validate-test-optimization.js --init-manifest
The scaffold is already schema-valid. Preserve its command boilerplate and edit only repository-specific command,
CI evidence, and omission fields needed for the selected representatives. Do not reconstruct the manifest from the
JSON Schema. Run --validate-manifest after each edit and follow its field-specific errors.
Use required focused existingTestCommand for the clean preflight and Basic Reporting. Prefer the resolved local
Jest, Vitest, or Mocha executable so package-manager bootstrap and home-directory cache writes cannot block the
local capability check. Preserve a package script only when a custom wrapper or required runner configuration
cannot be represented by the direct command. Use pending validator-owned preflight; exact ciWiringCommand for
the CI-shaped package-manager/wrapper command with non-secret CI env; and isolated generated scenario commands.
The local command and generated commands are Datadog-clean in the manifest and never use generated files outside
their declared scenarios. A package-manager blocker in CI replay must not replace a successful direct Basic
Reporting result. Record
CI NODE_OPTIONS/Datadog variables exactly, replacing only secret values. Validator overlays are not
CI evidence. Prefer structured command.env; if shell semantics are unsafe to represent, retain text
as evidence and mark replay unavailable.
Set preflight.maxTestCount to the smallest defensible bound for the selected representative, normally 1 for
a file-and-name-filtered test. The scaffold's 50 is only a conservative placeholder: inspect the command and
lower it before approval when the selected filter is narrower. If the clean preflight cannot determine a test count
or exceeds the approved bound, the validator stops without drawing a Test Optimization conclusion. If the package
manager cannot write its tool/cache directory, resolves an incompatible Yarn version, or Watchman cannot access its
state directory, report the concrete toolchain/execution-environment blocker. These failures happen before tests
start and are not Test Optimization evidence.
Set ciWiring.replayability explicitly. Use replayable only with a top-level ciWiringCommand that
preserves the approved CI shape. Use not_replayable only with a concrete replayBlocker explaining
the missing service, build, toolchain, or unsafe/unavailable command. A runnable framework cannot omit
this decision, and a non-replayable CI check makes full validation incomplete rather than successful.
When narrowing a broad CI command to one test, preserve the CI working directory, project/config selection, wrapper chain, and runner-specific path semantics. Inspect the selected runner config to prove the focused filter belongs to that project; an absolute repository path is not automatically a valid multi-project Jest/Vitest filter. If the approved replay finds no tests or exits before the runner produces a test result, report CI wiring as incomplete and correct the replay before recommending any Datadog CI configuration.
The selected representative and CI job must belong to the same runner project loaded by that exact CI
command. Do not pair a package test with another job merely because both eventually invoke Jest or
Vitest. If the original CI command does not execute the first representative, either select a small real
test that it does execute and use that test consistently for Basic Reporting and CI wiring, or mark CI
replay unavailable. A narrowed replay may add only a runner-supported file/name filter whose semantics
are proven by the CI-loaded config; do not invent --project, --config, --root, a different cwd, or
a wrapper bypass. In particular, do not assume a nested Vitest config's test.projects names are exposed
through a parent workspace config. Record the actual top-level project selected by the CI command.
Keep schema path fields absolute and inside the repository: repository/project roots, package/config files,
command working directories and output paths, generated test directories/files/cleanup paths, and test identity
files. Command arguments may remain relative when the runner resolves them from the command working directory;
the customer-facing plan also renders repository paths relatively for readability. Runnable entries need evidence,
setup, commands/preflight, ciWiring.initialization, replay when available, and a generated strategy. Non-runnable
entries need a status/reason. Consult the adjacent JSON Schema after field errors, then validate without execution:
node ./node_modules/dd-trace/ci/validate-test-optimization.js \
--manifest ./dd-test-optimization-validation-manifest.json --validate-manifest
For each runnable supported framework define one-test scenarios: stable basic-pass (exit 0) for
EFD, atr-fail-once (clean exit 1) for retry, and stable test-management-target (exit 0). Use
separate files or reliable filters, mirror nearby format/config, and show small printable secret-free
source in the plan. Set planned; the validator creates, verifies, runs, and cleans up. Declare exact
cleanup paths, never overwrite/delete existing files, and use suite: null unless events prove it. Every
framework entry must use its own generated files and cleanup paths in that framework's real test directory;
never share Jest and Mocha paths or reuse one framework's generated files for another runner.
For Vitest, place generated runtime tests where the selected config's literal test.include patterns accept them
and its literal test.exclude patterns do not. Do not use a typecheck-enabled project for Basic Reporting or
generated runtime tests; select an existing runtime-only config or add --typecheck.enabled=false to the approved
command. Match the generated test's ESM/CommonJS form to the nearest package.json that applies to its directory,
not only the representative project's package metadata.
node ./node_modules/dd-trace/ci/validate-test-optimization.js \
--manifest ./dd-test-optimization-validation-manifest.json \
--out ./dd-test-optimization-validation-results --print-plan
Fix placeholders, unresolved paths/files, or ambiguous scope. The command writes a bounded customer
approval checkpoint to ./dd-test-optimization-validation-results/approval-summary.md and the full
audit detail to execution-plan.md; it prints only their paths plus an agent reminder. It intentionally
does not expose the approval command in tool output. Read approval-summary.md and copy its complete
contents into the next user-facing assistant message. It contains every project command, cwd, execution
count, exact temporary test source, cleanup, outputs, and final command. Link execution-plan.md for a
user who wants the offline-fixture and integrity detail. Tool output, terminal transcripts, and collapsed
file reads do not count as showing the summary. Do not claim the file was shown, replace it with a prose
summary, or ask for approval when its contents are not visible in that message. The command also writes
approval.json plus a standard checksum list under the results directory. The approval SHA is the
SHA-256 of the exact JSON bytes and can be reproduced with the standard command printed in the detailed
plan. The validator reconstructs the JSON from current inputs before project execution; this consistency
check does not prove package provenance.
Use one approval surface. If the platform offers a command dialog without broader permissions,
submit the exact command immediately and do not ask in chat. Otherwise ask only Approve executing exactly the plan above?, then run it in the existing sandbox. New commands/resources require a plan.
If an agent platform refuses the installed validator, stop and report that its policy blocked live validation. Leave the reviewed command available for the user; do not alter the approved command or repository permissions.
After approval run only the final command; the validator owns setup, clean preflight, generated verification, offline fixtures, all checks, debug reruns, artifacts, and cleanup. Malformed, linked, incomplete, or oversized data fails closed without network fallback.
Basic pass means direct initialization reports; Basic fail/error leaves CI and advanced checks inconclusive. CI pass means replay emitted events with CI initialization; CI fail after Basic pass means CI setup did not reach the final runner; CI skip/incomplete/blocked gives no live conclusion.
Lead with verdict and compact checks table, then scope, exit code, manifest/report paths,
representative results, advanced checks, blockers, and validator How to fix. Never invent/apply fixes
or call skips failures. Link locally to dd-test-optimization-validation-results/report.md; inspect
embedded JSON/artifacts only for a specific failure and never upload them.
State whether validation coverage is complete or partial. A scenario-scoped run is partial and must
show every omitted check as NOT CHECKED; do not let an unselected CI or advanced check disappear from
the customer-facing summary. A full run is complete only when every selected check produced a result.
If no live Basic Reporting check ran, report the validation as incomplete even when discovery completed. Static CI findings are context only in that case: do not present Datadog CI changes, Git checkout changes, service naming, or other static observations as confirmed fixes. First identify the smallest runnable representative or report the concrete setup needed to obtain a live result.
Finally compare changed paths with the baseline. Remove only validation-created files; preserve prior work and leave no project changes outside declared outputs.