crates/turborepo/ARCHITECTURE.md
This document serves as a sketch of the architecture of the turbo run command
A run consists of the following steps:
futureFlags.experimentalCargoWorkspaces / futureFlags.experimentalPythonWorkspaces, Cargo workspace crates and uv workspace members)crates/turborepo/src/main.rs - Constructs TurboQueryServer (the concrete QueryServer implementation) and passes it to turborepo_lib::maincrates/turborepo-lib/src/commands/run.rs - Entry point for the run command, sets up signal handling and UIcrates/turborepo-lib/src/run/mod.rs - Core run implementationGraceful shutdown and parent-death cleanup are separate responsibilities. Graceful shutdown happens while the Turbo process is still alive, so it should be handled internally by the run and process manager. Parent-death cleanup only applies when Turbo disappears before Rust cleanup code can run.
crates/turborepo-lib/src/commands/run.rs creates one shared, process-lifetime
SignalHandler. It continuously brokers OS and in-process signals, retains
their count for force-shutdown escalation, and does not return until all
shutdown subscribers finish their cleanup work.ShutdownReason::Signal)
from close-driven shutdown (ShutdownReason::Close). Normal command
completion uses the close path to drain subscribers without printing
signal-specific shutdown UX.crates/turborepo-lib/src/run/mod.rs registers shutdown subscribers for
task processes, cache writes, and the microfrontends proxy.SIGINT/SIGTERM, Turbo enters graceful shutdown: it prints a
shutdown message, forwards SIGINT to running tasks, and waits for their
process groups to exit.Ctrl+C to
force shut down. Without a terminal on stdin, Turbo instead prints the
remaining time before the automatic force shutdown.Parent-death cleanup is not part of normal graceful shutdown. An in-process map
cannot help after SIGKILL, a crash, or OOM because the map dies with Turbo.
Turbo should not start a per-task Unix watchdog for this case. If abnormal
cleanup is required later, prefer a bounded run-level mechanism:
ProcessManager and shared by
all tasks in the run.pid, pgid, and session identity) when
spawned and unregister on normal exit or Turbo-managed shutdown.prctl(PR_SET_PDEATHSIG) as a best-effort no-helper option, but
it only signals the direct child and cannot provide delayed escalation.Regression coverage for shutdown changes should focus on observable lifecycle behavior:
turbo run signal tests should assert that descendants are not
leaked after force shutdown.crates/turborepo-daemon owns daemon lifecycle operations, server shutdown
semantics, and the daemon-backed package-discovery adapter. Both the turbo
application and the standalone turborepo-lsp binary use this shared API, so a
DaemonConnector can start the current executable with --skip-infer daemon
without requiring the LSP to depend on turborepo-lib.
The turbo application supplies its package-graph-aware change watcher and
keeps CLI-specific rendering in turborepo-lib. The standalone LSP supplies a
conservative daemon-owned watcher that emits full rediscovery for filesystem
changes. This preserves correctness when the packaged LSP hosts the daemon,
while the richer watcher avoids unnecessary rediscovery when turbo hosts it.
crates/turborepo-lib/src/run/builder.rs)Key responsibilities:
--filter)FilterMode (from turborepo-types): when no filter
or only exclude filters are active, root tasks defined in turbo.json are
auto-included. Explicit include filters or --affected suppress root task
injection. See calculate_filtered_packages and FilterMode.Run struct ready for executionWhen the affectedUsingTaskInputs future flag is enabled and --affected is
active, the run builder applies a second filtering pass after engine
construction:
turborepo-types/src/task_input_matching.rs): Each
task's inputs globs are compiled and checked against the changed files.
Shared with turbo query { affectedTasks }.turborepo-lib/src/task_change_detector.rs):
Determines directly affected tasks, handling global deps and per-task inputswith
relationships, which are not represented by task graph edges.Engine::retain_filtered_tasks): The engine retains the
selected tasks and adds their transitive execution dependencies (upstream
tasks needed as cache hits).This differs from the default --affected behavior which operates at the
package level (all tasks in changed packages run).
crates/turborepo-repository/src/package_graph/)Represents the workspace structure and package dependencies:
RepositoryKnowledge generation containing repository
root, root JavaScript execution scope, real package identities and source
boundaries, native definition paths, native workspace roots, provenance, and
required aggregate scopesrustc -vV identity directly through its external-resolution domain. Cargo
keeps missing, stale, or invalid lockfile and compiler identity failures
fatal. Core validates the combined domains and retains exact opaque
identities, definition sources, completeness, and stable fingerprints.
Repository generation owns byte-compatible per-package fingerprinting through
the cycle-free turborepo-lockfile-hash primitive; callers cannot omit it.
Open resolution-domain IDs select behavior after construction independently
of retained ToolchainId provenance, and each domain explicitly claims its
package members. Built-in JavaScript and Cargo IDs are reserved to their
canonical producers and repository roots. Core rejects duplicate IDs, package
ownership, unknown members, resolution rows outside declared membership, and
complete domains without exactly one row per member. A
single resolution owner tracks lifecycle status. Task hashing and run/task
summaries (including OpenTelemetry external-input attributes) consume the
same stored byte-compatible package resolution fingerprint and preserve
explicit unavailable states without closure fallback hashing. Query external
package listing, human names, and internal-dependent reverse indexes read
the same resolution generation (including a lazy compact reverse index)
rather than retained manifest payloads or live lockfile human-name callbacks.
N-API JavaScript lockfile package listing uses the JavaScript domain of that
generation. The JavaScript
adapter owns package-manager configuration, previous-lockfile parsing, and
resolution through the same producer used at graph construction. Core
compares the resulting normalized package identities without parser or
ecosystem knowledge; unavailable, parse, and comparison failures retain the
conservative all-packages fallback. Prune lockfile-key unions also consume
exact per-package resolution identities for retained JavaScript workspaces;
required external peer declarations come from relationship knowledge before
closure expansion consults the lockfile.
Global hashing consumes the same root resolution fingerprint as task hashing
and, when JavaScript resolution is unavailable, hashes resolution definition
sources plus the root package.json instead of reading the singleton lockfile
object. Phase 3 deletion removed closure/hash compatibility fields and
deferred resolution installation; readiness belongs to repository
construction. External declaration consumers (frameworks, boundaries, task
hashing) read the authoritative ExternalDeclarations projection built from
relationship knowledge rather than raw manifests. MFE
enablement checks exact declaration names in the same relationship generation
so internal workspace declarations and alias-key behavior remain unchanged.
MFE configuration discovery, directory ownership, and proxy execution accept
package/root scopes backed by authoritative package.json definitions;
aggregate and Cargo-manifest scopes are excluded without provenance dispatch.
Framework inference and boundaries validation consume the package-scoped
declaration projection directly;
aliases, duplicate
precedence, optional declarations, and peers remain explicit normalized
facts rather than raw manifest reads.name field
(PackageGraph::validate())RepositoryKnowledge is the crate-private authority for package identity,
paths, scope kind, and provenance during node assembly, and the resulting
PackageGraph retains that exact immutable generation. PackageGraph and
PackageTaskContext retain no manifest compatibility payloads. Consumers use
authoritative contexts, views, and immutable knowledge catalogs. A
DiscoveredPackage descriptor remains transient construction input for
JavaScript relationship classification and native-task observation; it is not
retained in the completed graph. Its PackageJson.name is non-authoritative:
no consumer may derive package identity, path, or provenance from it.
Native manifest objects do not enter repository knowledge. Repository knowledge
may retain bounded diagnostic provenance for an authoritative fact, such as an
authored JavaScript package name's source text and span. Native definition paths
must remain within the repository, including after resolving existing symlinks.
Retained graph and task-context payload deletion is complete. Discovery descriptor deletion remains pending:
Cargo.toml rediscovery names, the Cargo.lock resolution/rediscovery path, and the effective in-repository target-directory ignore prefix. PackageGraph::active_watch_spec is now a projection of only the immutable facts retained by that graph generation; it never calls live toolchains. Before the first generation is published, the watcher conservatively retains all in-repository events, closing the subscription/bootstrap race without mutable toolchain callbacks. Single-package generations retain no inactive Cargo facts.render_javascript_prune) producing typed artifacts; commands/prune.rs selects closures, performs path-safe layout, and materializes those artifacts without inline lockfile/manifest/patch format interpretation. Cargo discovery captures an immutable, generation-owned prune domain containing lockfile, root-manifest, package-directory, and post-write finalization authority. Scope contracts select JavaScript layout or an explicit native prune domain without branching on ecosystem provenance. Cargo lock pruning, extra-member selection, manifest rewriting, root/config file planning, and final lockfile canonicalization run through that graph-owned domain. Golden inventories cover standard and Docker layouts.PackageJson use is limited to construction inputs and narrowly scoped operational reads that have not yet moved into knowledge, including root package-manager configuration, prune rendering, and LSP unsaved-buffer adaptation. Boundary tag diagnostics consume optional authored-name provenance from repository knowledge only when the authored name matches the authoritative identity.Task hashing, run-cache path construction, and run-summary task directories use
a graph-created PackageTaskContext that binds identity, repository root,
directory, kind, task knowledge, and contract knowledge. Repository-wide task
namespace and external-dependency-hash enumeration is root-first, then follows
repository observation order. Pure Cargo retains the root Turbo namespace
without synthesizing a JavaScript scope, and consumers reject contexts from
another repository.
Repository-facing commands use the same optional-root construction policy as
turbo run: a missing root package.json is accepted only when Cargo support
is enabled and a root Cargo.toml exists, while malformed manifests always
fail. Devtools creates one package-graph generation per initial load or refresh
and derives both its package and task protocol graphs from that generation; the
server never performs independent package discovery. Each refresh also
re-resolves configuration, default config-file selection, and future flags.
Its watcher gives explicit custom config and repository-local config paths
physical coverage even when normal ignored-directory filtering would omit them.
Turbo configuration lookup receives knowledge-backed package and aggregate scope directories directly. Engine repository-wide task-definition enumeration and non-root missing-scope validation use the same views. The root Turbo task namespace remains available independently of whether a root JavaScript package scope was contributed; package.json scripts remain a separate temporary task-synthesis compatibility input.
Run scope resolution, affected fallback enumeration, and task-level directory
filters also consume these knowledge-backed views. Aggregate scopes remain
selectable. The always-available root Turbo task namespace remains distinct
from root JavaScript package knowledge: explicit //, repository-root
directory syntax {.}, root-task injection, and all-packages affected
fallbacks can select that namespace even in a pure native repository. This
namespace behavior does not imply that a root package owns the directory.
Repository knowledge accepts at most one physical workspace root for each open
ToolchainId, so a repository cannot combine multiple package managers for one
language. Repeated observations from one producer of the same kind and
canonical root deduplicate, while observations from different producers may
coexist. Public contributor output supplies only root kind and path; core binds
each root to the ToolchainId of the contributor whose discovery envelope
contained it. Every contributor that supplies packages must own an accepted
root, and contributed roots must remain physically within the repository.
JavaScript reports only the repository root for its authoritative package-manager
command family; if discovery reports a different family than an explicitly
resolved manager, the response is rejected. Pnpm versions and Yarn/Berry share
their respective canonical families. Cargo reports the current workspace root.
External-resolution generation validates domain roots against these authoritative
workspace roots before publishing terminal knowledge.
The package graph intentionally allows cyclic dependencies between packages —
this aligns with how npm, pnpm, and yarn handle cyclic workspace deps. Cycle
detection is deferred to the task graph layer (engine builder), since
package-level cycles only matter when they produce task-level cycles via
topological (^) dependencies.
Normalized relationship knowledge also records whether an internal relationship
orders tasks. Named core projections keep ordering, filtering, and package prune
closures acyclic while hash and affectedness projections include non-ordering
inputs. Cargo uses this distinction for cycle-closing development dependencies:
their sources still invalidate and affect consumers without creating task graph
cycles. Cargo path, development, optional, build, target-specific, and automatic
member relationships are emitted directly from cargo metadata; no synthetic
JavaScript dependency maps or behavioral affectedness callback remain.
crates/turborepo-repository/src/toolchain.rs)RepositoryContributor is a construction-time discovery abstraction. Its method
returns one envelope containing packages/scopes, workspace roots, and
ecosystem observations. The builder combines those envelopes in a local vector;
core validates them, builds immutable relationship, task,
contract, resolution, change, and prune knowledge, then drops the collection.
Runtime consumers query those retained catalogs and never dispatch through live
contributors. ToolchainId remains open provenance data rather than a closed
enum, keeping discovery extensible to future out-of-process plugin adapters.
Package-json membership is projected from real scopes with authoritative
package.json definition paths, independent of that provenance. Change
ownership and workspace path dependency splitting consume the same projection;
duplicate contributed definition owners and physical aliases of the root
definition are rejected during construction.
JavaScript discovery and typed pre-parsed manifest input both attach explicit
JavaScript task contracts; core defaults omitted contracts to empty behavior
without consulting contributor identity. Generic argv overrides for scopes
without native tasks are core execution policy; pure-root execution is one case
and does not depend on whether provenance is present.
crates/turborepo-repository/src/native_tasks.rs)The immutable native-task catalog classifies each task as Command,
Aggregate, or None. Command stores a declarative program, argument layout,
working-directory policy, and serial group. Aggregate runs no process; it adds
same-scope task dependencies. None carries classification or contract facts
without making the task runnable.
Aggregate children are validated while the catalog is built. They must be
unqualified, participating tasks in the same observed scope and cannot name the
aggregate itself. Duplicate children are removed. The engine merges these
implicit dependencies with authored dependsOn entries, deduplicates by the
effective same-scope task ID, and subjects the result to normal cycle
validation. An explicit argv command or command: null replaces the native
aggregate in either direction, while Package Configuration extends: false
can remove the inherited task. A scoped definition that adds configuration
after extends: false participates normally and restores the native aggregate.
Arguments cannot target an active aggregate because it has no process; the
error directs users to its package-qualified child tasks.
NativeCommandArguments models fixed prefixes and suffixes, pass-through
placement before or after the suffix, and optional fixed or package-manager
separators. This keeps Cargo, uv, and JavaScript argument behavior in data
rather than executor branches. Argument lookup remains task-specific, so a
dependency task does not inherit arguments supplied for another requested
task.
Behavior that varies by native task lives in NativeTaskContract: defaults,
entrypoint classification, and whether that task derives I/O. Broader behavior
such as environment projection, dependency-source participation, dynamic I/O,
pruning, and command-map capability remains on ScopeTaskContract. Engine
composition uses the task-local contract when present and the scope contract
only as a fallback. Definition memoization keys include native execution and
task-local contract facts, and it is disabled for package-scoped definitions or
per-package derived I/O. This prevents scopes with the same turbo.json chain but
different aggregates or contracts from sharing an invalid definition.
JavaScript is the first production producer. Machinery that predates the abstraction (package-manager resolution for dependency splitting and the JS lockfile closure phase) remains documented debt.
Production callers enable Cargo through PackageGraphBuilder::with_cargo
rather than constructing or retaining contributor objects. Run, watch, daemon,
prune, query, hashing, and cache paths carry only the feature decision; each
graph generation creates and drops its own Cargo contributor.
Contract-derived I/O receives the same task-scoped arguments as execution plus
a narrow, platform-aware startup-environment projection keyed by an explicit
contract domain. Dependency source participation is likewise declared by each
scope contract rather than inferred from contributor provenance.
Dependency tasks do not inherit arguments for a different requested task, each
contract can observe only the variables it declares, Windows lookup remains
case-insensitive, and every declared pattern automatically participates in task
hashing. If a user env exclusion matches a projected contract-derived I/O variable,
automatic outputs become unavailable rather than deriving cacheable paths from
an unhashed value. Derived outputs distinguish exact/resolved paths from
unavailable automatic resolution. When outputs are unavailable, the engine
disables implicit caching so a log-only hit cannot suppress execution, while
explicit outputs, cache: true, and
cache: false remain authoritative.
crates/turborepo-repository/src/cargo.rs)Behind futureFlags.experimentalCargoWorkspaces in the root turbo.json,
turbo run discovers Rust crates from a Cargo workspace at the repository
root and adds them to the package graph. Cargo workspaces can stand alone or
coexist with JavaScript workspaces; a root package.json and JavaScript package
manager are only required when JavaScript packages participate. Cargo-only
repositories may omit package.json; when one exists, it must still be valid.
CargoContributor is the second RepositoryContributor implementation.
Turborepo does not replace Cargo. Cargo is itself a build system with its
own dependency graph, scheduler, and incremental cache (target/), so the
division of labor is: Turborepo decides which crates are in scope and
whether anything changed; Cargo decides how and in what order to build.
Discovery normally uses one cargo metadata --locked --all-features
snapshot — Cargo is the only correct implementation of its own membership
semantics
(member globs, automatic path-dependency members, excludes, target-specific
dependency tables, renames). Dev-dependency edges that would form a cycle
are dropped (Cargo permits dev-dep cycles; crate edges must support
topological ^ ordering). Crate names are validated, and a crate/JS package
name collision hard-errors. Cargo contributes its already-classified native
internal relationships directly, without JavaScript dependency descriptors
or package-manager policy. The same snapshot validates resolution and every
resolved local package. --no-deps is used only to preserve error precedence
and classify memberless workspaces when locked metadata cannot be obtained.
Automatic in-repository workspace members are supported, while
excluded/non-member, outside-repository, and root-manifest local packages
hard-error because Turborepo cannot hash, watch, or prune their sources
safely. The Cargo contributor reports the current workspace root.
Package shapes: crates are classified via CargoPackageKind.
Entrypoints (crates with bin/cdylib/staticlib targets) are the
workspace's deliverables. Libraries exist in the package graph and expose
filtered build and verification tasks. Unfiltered builds prefer entrypoints
because Cargo builds their library dependency closures implicitly. A
user-named workspace scope — declared via [workspace.metadata] name in
the root Cargo.toml, a hard requirement — is an aggregate in repository
knowledge. Its normalized relationships point to every crate, and it hosts
workspace-scoped verification verbs.
Execution and entrypoint selection (NativeTaskKnowledge, native command
resolution in turborepo-task-executor, and
PackageGraph::task_entrypoint_exclusions): crate-scoped build and verification
tasks run cargo <verb> --package=<crate> --locked; entrypoints also expose
run/dev. Unfiltered builds prefer entrypoints, falling back to libraries
when the workspace has no entrypoints. Unfiltered verification uses the Cargo
workspace aggregate:
<name>#test runs cargo test --workspace --locked, <name>#lint runs
cargo clippy --workspace --locked, etc. Formatting is mutating and defaults
to uncached:
filtered runs use cargo fmt --package=<crate>, while the workspace aggregate
uses cargo fmt --all; neither form uses --locked. Other filtered runs use
their selected crates; selecting only the workspace aggregate uses its
workspace command.
RunBuilder combines filter mode with the resolved package scope to derive
task-specific exclusions. EngineBuilder applies package-level exclusions
before traversal; task-level filtering defers selection until after matching
task inputs. Exclude-only filters therefore remain exclusions rather than
being swallowed by a workspace command. Package-qualified task arguments remain
authoritative. --locked on non-formatting tasks preserves
the dependency resolution validated before task hashing. Cargo commands
(except cargo run and cargo fmt) share a mutually-exclusive serial group: concurrent
cargo processes serialize on the build-directory lock anyway, so the
executor runs one at a time without the "waiting for file lock" noise. Run
summaries read the same resolved native command catalog as execution, so
display cannot drift from execution.
Task registration (NativeTaskKnowledge): every crate implicitly
registers build; entrypoints with exactly one binary also register run
and its dev alias. Every crate and the workspace aggregate register test,
check, lint, and format. These act as empty task definitions
at the lowest precedence, so normal
tasks entries configure or override them and package configuration can
exclude them with extends: false. Registration is package-aware, so the
defaults do not make
same-named JavaScript scripts runnable without their usual turbo.json
definition. The names come from the same verb tables as command resolution
and participate in task suggestions and add-all/query graph construction.
Hashing and affectedness (HashRelationships, AffectedRelationships,
and TaskContractKnowledge): crate-scoped tasks hash their own
sources plus a conservative transitive closure of local dependencies whose
scope contracts explicitly include dependency source inputs (Cargo package
scopes opt in; the workspace aggregate opts out). Unknown participation makes
automatic inputs untracked rather than silently cacheable. The closure is
flattened, so invalidation doesn't depend on dependsOn wiring, and may
include optional or target-specific dependencies not compiled by a particular
invocation. Non-ordering relationship inputs retain
cycle-closing development edges so they still invalidate and mark their
consumers affected. Tasks also hash the
workspace files (root Cargo.toml, .cargo/config*, rust-toolchain*),
and standard Cargo/cc-rs environment inputs: rustup home/toolchain selection,
compiler and rustdoc selection and flags, Cargo build/profile/target
configuration, native compiler and
archiver settings (including target-qualified forms), and platform SDK
selection. Formatting additionally includes rustfmt.toml, .rustfmt.toml,
and RUSTFMT. Arbitrary variables consumed by project-specific build scripts
remain explicit task env configuration. The workspace aggregate hashes all
crate directories instead of default-hashing the repo root.
$TURBO_DEFAULT$ in a Cargo task's inputs means "everything turbo
derives automatically", so extra inputs (e.g. a file embedded via
include_str! from outside any crate directory) are additive.
External dependencies (turborepo-lockfiles/src/cargo.rs): locked
registry/git packages and the compiler itself flow through the same
ExternalResolutionGeneration and resolution fingerprint used by JavaScript
packages. Each crate's closure is computed
from Cargo.lock (identity = version + source + checksum, so git rev
bumps count). Source-qualified lockfile edges distinguish identical
name/version packages from different registries or git references, so each
closure follows Cargo's exact resolved package. A dependency bump therefore
only invalidates crates that actually depend on it. The complete verbose
compiler identity from rustc -vV,
including its host triple, is resolved from the repo root (so
rust-toolchain overrides apply) and added to every Cargo package's set.
This prevents compiler releases, operating systems, architectures, or host
ABIs from sharing native artifact cache entries. Explicit targets selected
through hashed task arguments or CARGO_BUILD_TARGET remain distinct;
repository build.target stays conservatively unavailable. Failure to
resolve the compiler identity is
a hard error. Every non-empty Cargo workspace must have a current
Cargo.lock: discovery runs full cargo metadata --locked --all-features
before hashing, then computes per-crate closures. Missing, stale, unparsable,
or incomplete lockfiles are hard errors. Turborepo never creates or refreshes
the source lockfile; users do that explicitly with Cargo and commit the
result.
Caching: task caches store logs plus, for entrypoint builds, exact
deliverables under the effective target directory. The rustc -vV host and
rustc --print target-list validate target triples; CLI --target wins over
CARGO_BUILD_TARGET, and the effective target adds its Cargo path segment and
platform-correct bin/cdylib/staticlib basename. CLI --target-dir wins over
CARGO_TARGET_DIR, which wins over Cargo metadata (including repository
target-dir). Target directories are accepted only when their canonical or
nearest existing path remains in the repository. No profile or platform
wildcards are cached. Automatic outputs fail closed for repository
build.target, unknown/custom targets, path escapes,
CARGO_BUILD_TARGET_DIR, compiler overrides, or when
manifests/configuration can alter
profile directories, artifact names/locations, or include unhashable external
configuration. External, included, and Cargo configuration beneath any
symlinked path component is untracked; those config paths are not emitted as
trusted inputs. Unresolved outputs
disable implicit caching unless outputs or cache behavior are configured.
Untracked inputs disable caching unless cache itself is explicitly configured;
explicit outputs alone cannot make an incomplete input hash safe. Cargo's
internal target/ state is deliberately never cached
— it is Cargo's own incremental cache, and
tarballing it fights Cargo instead of leaning on it (it is also
multi-gigabyte). For fine-grained compile caching, RUSTC_WRAPPER
(sccache) is the sound layer, and it participates in task hashes so
toggling it invalidates caches. Entrypoint run/dev tasks and library
build tasks default to cache: false: a cache hit must not suppress a
requested process, and library artifacts have no stable final path to restore.
An explicit turbo.json cache setting overrides the toolchain default.
Watch mode (ChangeKnowledge and PackageGraph::active_watch_spec,
consumed by turborepo-lib/src/package_changes_watcher.rs): discovery
observations declare workspace-definition files and build-byproduct
directories; accepted observations compose directly without reactivation by
producer identity, and the current graph generation projects one active spec. For
Cargo, any Cargo.toml or the root Cargo.lock triggers full
rediscovery (the crate set or its edges may have changed), while events
under the root target/ directory are dropped — Cargo writes there
continuously during builds, and the feedback loop must not depend on a
.gitignore entry (Cargo.toml files under target/ are build
byproducts, not workspace definition). The watcher builds its package
graph through the same construction path as a run, so watch sees the same
package set. JavaScript declares nothing extra: workspace
redefinition is caught by the change mapper's conservative
all-packages fallback. Known gap: the hash watcher's content-hash dedup
is JS-glob-based, so a no-op save inside a crate re-runs its tasks as a
fast cache hit rather than being suppressed.
Prune (PruneKnowledge and PruneDomain::{plan, finalize}, consumed by
turborepo-lib/src/commands/prune.rs): each generation-owned domain reports
what a self-contained pruned repository needs beyond copied packages. For
Cargo: the kept-member set comes from a Cargo.lock reachability walk
(not the package graph — the lockfile merges dev-dependency edges, so
members reachable only through dev-deps are retained, since kept crates'
manifests reference them), the lockfile is subset to that closure, and
the root Cargo.toml is rewritten with toml_edit (explicit members,
filtered default-members, [workspace.dependencies] path entries to
removed crates dropped — comments and formatting preserved). Ecosystem
and Cargo config files are carried over. Reachability pruning cannot see
Cargo's feature unification, so the retained Cargo domain runs cargo metadata
once in the complete output (offline first, then networked) to let Cargo
minimally sync its own lockfile; failure downgrades to a warning.
Only domains that contributed a prune plan are finalized. Finalizers
report files they may have changed, and prune copies those finalized bytes
to alternate output layers without rerunning the toolchain. Reported sources
must be regular files rather than symlinks, and paths must remain within both
output roots lexically and after resolving symlinks; invalid paths and
synchronization failures are warnings. In docker layout,
the json layer carries the root manifest, each kept crate's Cargo.toml, and
finalized lock; sources go to the full layer. A
aggregate anchored at the repo root (the Cargo workspace scope) is not
a pruneable target.
Compile cache (ScopeTaskContract::compile_cache_env, consumed by
ToolchainCommandProvider; gated by futureFlags.experimentalCargoSccache):
when enabled alongside experimentalCargoWorkspaces in a CI environment
with a linked Remote Cache, the run serves a local HTTP proxy
(turborepo-sccache-proxy) that presents an sccache-compatible webdav
storage backend and translates GET/PUT/HEAD into Remote Cache
artifact calls. Nothing needs installing: turbo embeds sccache as a
library (a Vercel fork of mozilla/sccache pinned in Cargo.toml, adding
an explicit-args entrypoint) and acts as the compiler wrapper itself —
main.rs dispatches invocations marked with TURBO_SCCACHE_WRAPPER=1
(and sccache's internal SCCACHE_START_SERVER=1 respawn) to
sccache::main_from_args, alongside the LSP and Windows ctrl-c shims.
Cargo tasks get RUSTC_WRAPPER=<turbo>, the wrapper marker,
SCCACHE_WEBDAV_ENDPOINT/SCCACHE_WEBDAV_TOKEN, and
CARGO_INCREMENTAL=0 injected at execution time; JavaScript injects
nothing. Objects are fetched lazily per rustc invocation, so nothing is
restored before a task runs; the two cache layers compose (task-cache
hit: nothing executes; miss: cargo's conservative recompiles become
downloads). The endpoint must be stable across runs because the sccache
background server captures it at startup and outlives the run: the port
is derived from the repo root and the bearer token is persisted at
.turbo/sccache-proxy-token. Injection is execution-only and does not
participate in task hashes (a compile cache is output-transparent). The
Cargo task contract decides how injection composes with the task environment: a
user-supplied RUSTC_WRAPPER or any SCCACHE_* variable signals a
competing compiler-cache configuration and suppresses the whole injected
set, while an ambient CARGO_INCREMENTAL (CI images commonly export
=0) is tolerated — injection proceeds without overriding it. Every
unmet precondition disables the proxy softly. CI-only by design: cold environments are where a compile cache
pays off, while local development is served by cargo's own incremental
compilation — which the injected CARGO_INCREMENTAL=0 would disable.
Lifecycle: started in Run::execute_visitor before the visitor,
shut down fire-and-forget after it. The proxy counts the work-unit
traffic it serves (hits/misses/stores, health-check probe excluded)
and the run summary footer reports it as a toolchain-agnostic
"Incremental cache" line — reuse below the task boundary — shown only
when the run actually exchanged work units.
A --filter that names a crate while support is disabled gets an error
hint pointing at the flag. Released turbo versions hard-error on unknown
futureFlags keys, so a repo can only adopt the flag once every consumer
(hooks, CI) runs a version that knows it.
End-to-end coverage lives in crates/turborepo/tests/cargo_workspace_test.rs
against the cargo_monorepo fixture (a mixed npm + Cargo workspace):
graph shape, execution, caching, deliverable restoration, cross-crate
invalidation, lockfile enforcement, unsupported local-package rejection,
uncached run/dev execution, and the filter hint. turbo query serves Cargo
packages through the same graph.
crates/turborepo-repository/src/uv.rs)Behind futureFlags.experimentalPythonWorkspaces, turbo run discovers
Python packages from the root uv workspace and adds them to the package graph.
uv is the only supported Python package manager. uv workspaces can stand alone
or coexist with JavaScript and Cargo workspaces; no PackageManager variant is
involved. UvContributor contributes pre-classified relationships, native
tasks, external resolution, change observations, and a prune domain through
the shared repository graph.
[tool.uv.workspace] members and exclude globs
in-process, so graph construction does not require the uv binary. When uv
is available, discovery probes its version and the Python interpreter selected
by uv python find for cache identity. Missing identities and repository-local
uv executables fail closed to uncached tasks. Names are PEP 503-normalized.
Dependencies become internal graph edges only when
their effective [tool.uv.sources] entry selects workspace = true.
Development edges that would create a cycle become non-ordering input
edges. A root [project] participates in hashing and pruning but is not a
package. A synthetic workspace package, named by [tool.turbo] name,
depends on every member and hosts workspace-wide quality tasks.[project].dependencies, recursive
[dependency-groups] (including include-group), and legacy
[tool.uv].dev-dependencies. Optional dependencies and declarations with an
environment marker are excluded. Discovery records whether a declaration
belongs to the root or a member and whether its dependency group is active by
default. For each role, a member's declarations replace the root
declarations; an empty member role inherits the root role. Ruff supplies both
lint and format roles.build for buildable members. Without a recognized
tool, all members receive fallback format and check commands. Recognized
tools add qualified lint:<tool>, format:<tool>, and check:<tool> tasks;
lint and check are same-scope aggregates of their qualified tasks.
Canonical format selects Ruff before Black and warns once per selected
scope when both are declared, while qualified tasks remain available for an
explicit choice. If every member has the same tools, owner, and non-default
activation group for a role, unfiltered runs use the workspace scope;
member-owned tools add --all-packages. Otherwise, member scopes are the
entrypoints for that role. Filtered runs use member commands. A root pytest
declaration registers a preferred-only workspace test; direct member
declarations register candidate member test tasks. The root task wins an
unfiltered run, while a package filter can select only a directly declaring
member. Member pytest tasks have no shared serial group and can run in
parallel.uv build --package=<name>, or uv run --frozen
followed by owner selection (--package <name> for a member,
--all-packages for homogeneous member ownership, and no owner flag for a
root declaration), non-default group activation (--no-default-groups --group <group>), the tool and subcommand, pass-through arguments, and the
member-directory targets. Ruff uses check/format; ty uses check.
Fallbacks remain uv format -- <dirs...> and either
uv check --frozen --package=<name> or uv check --frozen --all-packages.
Detected-tool and
fallback check commands use the uv serial group. Detected lint, check, and
test commands default to cacheable when uv and Python identities resolve;
otherwise they fail closed to uncached. The fallback uv check stays uncached
because its bundled checker is not represented in uv.lock. Builds using
uv_build cache when their sole build requirement accepts the identified uv
executable's bundled backend version. Other PEP 517 builds remain uncached,
and format commands remain uncached because they mutate source. Detected-tool commands use --frozen;
Turborepo itself never creates or updates uv.lock. Pass-through arguments are
inserted before path targets. Active aggregates reject them and name the
package-qualified child tasks that can receive them.
Pytest commands are uv run --frozen pytest for the workspace or uv run --frozen --package <name> pytest <member-dir> for members, with non-default
group activation inserted before pytest and pass-through arguments inserted
before the member target.check, check:*, and member test also include internal source closures.
Quality workspace tasks include every member's sources; a bare workspace
pytest task hashes the full repository because pytest controls collection.
Quality caches, .pytest_cache, .venv, and __pycache__ are excluded.
Path-valued uv settings, UV_NO_SYNC, UV_NO_PROJECT, active user/system uv
configuration, and any pass-through arguments make automatic inputs
untracked. Each scope also hashes its external
dependency closure from uv.lock; root-owned tools conservatively add the
workspace closure. Package identities include version, source, and artifact
hashes. Every scope also includes path-independent uv and Python identities
containing executable content hashes, Python implementation and version,
operating system, architecture, libc, variant, and host compatibility. A uv.lock change across git refs conservatively affects
all uv packages. Build output inference covers the bare command's matching
dist/ artifacts and becomes unavailable when arguments are present.pyproject.toml, the root uv.lock,
.python-version, and uv.toml. Root .venv/ and dist/, plus known quality-tool and Python cache
directories at the root and member scopes, are ignored as task byproducts.uv.lock reachability, including dependency groups and
optional extras, and preserves retained package metadata through
toml_edit. It rewrites root workspace members, removes dangling uv source
entries, and copies .python-version and uv.toml when present. Reachable
local dependencies that are not workspace members fail closed.End-to-end coverage in crates/turborepo/tests/uv_workspace_test.rs exercises
pure uv and mixed npm/uv repositories, graph shape, filtering, affectedness,
execution, and prune output. Linux Rust CI installs a pinned uv version; local
tests that execute uv skip when it is unavailable.
crates/turborepo-lib/src/engine/)The task graph is a graph of all tasks that will be part of the run and related configuration.
Due to purely historical reasons, this is referenced as "engine" throughout the codebase.
The core task graph consists of:
crates/turborepo-lib/src/engine/builder.rs)turbo.json and other configuration sources to determine task definitions^build and direct build)command override
(futureFlags.experimentalTaskCommand) in one place
(resolve_command_override, turborepo-engine's
builder/definitions.rs), across five precedence levels: Package
Configuration command → root pkg#task command → authored native task
from NativeTaskKnowledge → unscoped root default (command maps fan out by
explicit task-contract capability) → the catalog's synthesized native command. The
resolved override is authoritative in both directions — an argv executes
even where the toolchain defines nothing, an opt-out never executes even
where it does — and feeds global-deps hashing, the TUI task list, the
executor (ToolchainCommandProvider), and the task hash
(TaskHashable.commandOverride/commandOptOut). Toolchains place the
argv in their frame: cwd is the package directory, nothing is prepended,
and Cargo keeps its serial group when the override still invokes cargo.
Because an argv override is otherwise arbitrary, it does not inherit the
native command's contract-derived inputs, outputs, default-input behavior,
or hash environment; its turbo.json inputs, outputs, and env are the
authoritative task-level I/O configuration. Contract-derived task defaults and
execution-only compile-cache environment injection likewise apply only to
native-catalog-resolved commands.crates/turborepo-lib/src/run/builder.rs)futureFlags.strictTaskEntrypointSelection and is independent of
futureFlags.filterUsingTasks.turbo run build test still runs every selected build command.filterUsingTasks, package and git selectors start from the requested
task nodes rather than every dependency task already present in the package.
Missing requested nodes are dropped after selector expansion: a plain filter
runs nothing for a missing command, while trailing or leading ... may retain
executable tasks reached through that node in the Task Graph.command overrides, and native-catalog
commands all count as executable definitions. Contract-derived entrypoint
selection, including Cargo workspace/crate selection, remains authoritative
and is composed with this generic command-aware pruning.crates/turborepo-lib/src/engine/execute.rs)Task Graph Structure:
TaskId (package#task) or rootcrates/turborepo-engine/src/lib.rs)retain_affected_tasks keeps directly affected tasks, transitive dependents,
and all transitive dependencies required for normal --affected executioncreate_engine_for_subgraph and retain_watch_affected_tasks are used by
package-level and task-input watch modes, respectively. They keep changed
tasks, transitive dependents, and only cacheable upstream dependencies that
can restore outputs without forcing non-cacheable tasks to rerun. Persistent
non-interruptible tasks are excluded because watch mode cannot restart themcrates/turborepo-filewatch, crates/turborepo-lib/src/package_changes_watcher.rs)FileSystemWatcher owns the platform watcher and exposes a demand-driven
WatchSource. Scoped consumers subscribe with a WatchScope; path filtering
occurs before that consumer's bounded event channel, so irrelevant events
cannot make it lag. Package changes, package discovery, input hashing,
output-glob tracking, cookies, devtools, and daemon root monitoring all use
independent scopes; there is no repository-wide raw event broadcast..git,
paths excluded by repository and nested Git ignore rules, and toolchain
build-byproduct prefixes. Tracked files and their ancestor directories remain
relevant even when an ignore pattern matches. It always
admits .gitignore, turbo.json, turbo.jsonc, an in-repository custom
Turbo config, and toolchain workspace-definition files..gitignore
matchers before routing an event containing any .gitignore; it also applies
.git/info/exclude and core.excludesFile, but deliberately does not
interpret ripgrep .ignore files. On macOS, a global excludes file on a
different device is conservatively not applied because one FSEvents stream
cannot monitor both devices. The package scope
refreshes the active WatchSpec from the current graph generation whenever
the package graph is initialized or rediscovered. Turbo config and ecosystem definition changes
trigger full rediscovery after routing..git/info/exclude, core.excludesFile, and repository Git config
control paths are watched separately so tracked-file and exclude state stays
current without exposing .git events to normal consumers. Refreshes publish
the complete control-path and ignore snapshot as one generation and
conservatively invalidate consumers when core.worktree changes. Backend
rescan signals bypass path scopes and invoke each consumer's conservative
recovery..gitignore changes reconcile installed ordinary watches, including removing
newly ignored coverage while preserving explicit interests and cookies.
Explicit leading-wildcard inputs may require broad package coverage; these
still hard-exclude .git and node_modules.crates/turborepo-lib/src/task_graph/visitor/)The task graph visitor handles task execution:
visit (crates/turborepo-lib/src/task_graph/visitor/mod.rs)mode: "jit" or
mode: "dependencyOutputs") defer final file-input hashing until the engine
dispatches the task, after its dependencies have completed and restored any
cached outputs. Tasks that depend on deferred tasks are also deferred so their
dependency hashes are available before their own hash is calculated. Once a
deferred task has a real hash, the visitor precomputes any unblocked
non-deferred descendants instead of waiting for each descendant to be
dispatched.ExecContext for each taskcrates/turborepo-lib/src/task_graph/visitor/exec.rs)ExecContext: Holds state required to execute a taskturborepo_processstdout/stderr outputExecution Flow:
crates/turborepo-lib/src/run/cache.rs and crates/turborepo-cache/)Multi-layered caching system:
Cache restore and storage enforce filesystem boundaries. Restores are anchored to the selected restore directory, preserve safe symlinks, and reject symlink targets that escape that anchor. Cache storage rejects task outputs that resolve outside the repository root.
RunCache: High-level cache coordinationTaskCache: Individual task cache managementAsyncCache: Handles async cache operations. Supports both local filesystem and remote HTTP cachesSharedHttpClient: Process-wide lazy/activatable reqwest::Client
initialization shared by telemetry and remote-cache consumersNetwork consumers do not construct an HTTP client speculatively at process startup. Instead:
reqwest::ClientThis avoids paying client/TLS setup on invocations with no network use while still warming the client before the first network request in the common case.
turbo run builds SCM state in two stages:
.git/index and records committed blob IDs
plus modified/deleted tracked files for the whole repoThose prefixes are relative to the repo index root, which is usually the Git root. This matters when the Turbo root is nested inside a larger Git repository: the root package should scope to the nested Turbo directory, not request an untracked walk of the entire parent repository.
This keeps the cheap tracked-index work overlapped with other startup work while avoiding a repo-wide untracked walk when only a subset of packages will be hashed.
When running in a Git linked worktree (created via git worktree add), Turborepo automatically shares the local file system cache with the main worktree. This enables:
How it works:
WorktreeInfo::detect() in turborepo-scm determines if the current directory is a linked worktree using Git commands (git rev-parse --show-toplevel and git rev-parse --git-common-dir)ConfigurationOptions::resolve_cache_dir() returns the main worktree's .turbo/cache directory instead of the local oneConfiguration:
cacheDir in turbo.json disables worktree cache sharingCache writes use an atomic write pattern (write-to-temp-then-rename) for concurrent safety:
.{filename}.{pid}.{counter}.tmp)CacheWriter implements Drop to clean up temp files if finish() is not called (e.g., on error or panic)This ensures concurrent readers never see partially written cache files.
crates/turborepo-lib/src/task_hash/)Creates a "content identifier" for a specific task depending on current state of inputs:
globalDependencies, are omitted if they disappear
before hashing. Required resolution fallback files remain strict. Verbose
tracing records final-stage candidates that disappear or are rejected as
non-regular, plus path-specific hashing failures..git, that metadata participates in
the hash and transient entries use the same discovery-race behavior.inputs, glob matches still walk the
filesystem, but clean tracked matches reuse blob OIDs from the repo index
instead of re-hashing file contentsinputs entries with mode: "jit" are file
inputs hashed just before task execution. mode: "dependencyOutputs" selects
already-expanded dependency task nodes and defers the task hash because those
producers' declared outputs are not known until after dependencies complete.
In dry runs, these task hashes are reported as deferred..gitattributes marks files as text or
text=auto, git normalizes CRLF line endings to LF in blob objects. The
crlf module in turborepo-scm replicates this so turbo's file hashes
match git's regardless of the code path (git or manual/no-git after
turbo prune). .gitattributes is included in the global hash inputs
and preserved by turbo prune. Known limitations: only root-level
.gitattributes is loaded; eol= is not handled.globalConfiguration and global.inputsWhen the globalConfiguration future flag is enabled, global.inputs (formerly
globalDependencies) files are not included in the global hash. Instead,
they are prepended as implicit input globs to every task's TaskInputs during
engine construction (see prepend_global_inputs in
crates/turborepo-engine/src/task_definition.rs).
This means:
global.inputs file hashes"inputs": ["$TURBO_DEFAULT$", "!$TURBO_ROOT$/tsconfig.json"])inputs key get default: true set so package files
are still hashed alongside the global inputscapnp to serialize in memory structs for hashingcrates/turborepo-lib/src/run/summary/)The summary module is responsible for any time of summary:
--summarize--dry=jsoncrates/turborepo-lib/src/run/summary/mod.rs)Visitor::visitcrates/turborepo-lib/src/run/summary/execution.rs)--dry=json/--summarizeThe query subsystem powers turbo query (GraphQL introspection of the
package/task graph).
Crate layout:
turborepo-query-api — Trait definitions (QueryServer, QueryRun) and
shared error/result types. turborepo-lib depends on this thin interface
crate instead of the heavy implementation.turborepo-query — GraphQL implementation using async-graphql, axum, and
oxc. Implements the resolvers and HTTP server.turborepo/src/main.rs — Wires the two halves together via TurboQueryServer,
which implements QueryServer by delegating to turborepo-query.Data flow: main() constructs Arc<TurboQueryServer> → passes to
turborepo_lib::main → threaded through shim → cli::run →
commands::run → RunBuilder → Run. The Run struct stores the
query_server; the turbo query command handler uses it for direct query
execution and the local GraphQL server mode.
RunBuilder
├── Package Discovery → PackageGraph (validates package names)
├── Task Discovery → EngineBuilder
├── Task Graph Construction → Engine (built)
└── Task Graph Validation (cycles, missing deps) → Ready Engine
Process:
dependsOn configurationsEngine.execute()
├── Walker (topological order)
├── Semaphore (concurrency control)
├── Engine -[Task to Run]→ Visitor
└── Engine ←[Task Result]- Visitor
Process:
Walker traverses graph in topological orderVisitorVisitor executes task and reports back to EngineVisitor.visit()
├── Calculate Hash
├── Check Cache → Cache Hit? → Restore & Done
├── Execute Task → Create ExecContext and `exec_context.exec()`
├── Save to Cache
└── Track Results
Process:
TaskCache.restore_outputs()
├── Check caching disabled?
├── Local Cache → exists?
├── Remote Cache → exists?
├── Fetch & Extract
└── Return metadata
TaskCache.save_outputs()
├── Collect output files
├── Compress to tar
├── Save to Local Cache
└── Upload to Remote Cache (async)
RunTracker
├── Task Events → ExecutionTracker
├── State Aggregation → SummaryState
├── Summary Generation → RunSummary
└── Output (JSON/Console)
Process:
ExecutionTracker aggregates state across all tasks.turbo/runs/ and optionally printedcrates/turborepo-run-summary/src/observability/ and crates/turborepo-otel/)The observability subsystem enables exporting run metrics to external backends via OpenTelemetry.
The system uses a two-layer design:
turborepo-otel: Low-level OTLP exporter crate
turborepo-run-summary/observability: Integration layer
RunObserver trait for pluggable backendsRunSummary data into metrics payloadsotel feature flagobservability::Handle: Main entry point; wraps backend-specific implementationsRunObserver trait: Abstraction allowing future backends (Prometheus, etc.)OtelObserver: OpenTelemetry implementation of RunObserverObservability is configured via experimentalObservability.otel in turbo.json:
{
"futureFlags": {
"experimentalObservability": true
},
"experimentalObservability": {
"otel": {
"enabled": true,
"protocol": "http/protobuf",
"endpoint": "https://otel-collector.example.com:4318/v1/metrics",
"resource": {
"service.name": "turborepo"
},
"metrics": {
"runSummary": true,
"taskDetails": true,
"runAttributes": {
"id": false, // turbo.run.id — unbounded cardinality
"scmRevision": false // turbo.scm.revision — unbounded cardinality
},
"taskAttributes": {
"id": false, // turbo.task.id
"hashes": false // turbo.task.hash, turbo.task.external_inputs_hash — unbounded
}
}
}
}
}
Configuration can also be set via environment variables (TURBO_EXPERIMENTAL_OTEL_*) or CLI flags (--experimental-otel-*).
OTEL endpoints must be HTTPS URLs without userinfo. Literal private, loopback, link-local, multicast, documentation, carrier-grade NAT, and known metadata-service IP endpoints are rejected; use localhost by name for local collectors.
turbo.run.duration_ms - Run duration histogramturbo.run.tasks.attempted - Tasks attempted counterturbo.run.tasks.failed - Tasks failed counterturbo.run.tasks.cached - Cache hit counterturbo.task.duration_ms - Per-task duration histogram (when taskDetails enabled)turbo.task.cache.events - Per-task cache events (when taskDetails enabled)Duration histograms use custom millisecond buckets sized for build and task durations, rather than the OpenTelemetry SDK's default latency buckets.
Attributes with unbounded cardinality (unique run IDs, Git SHAs, content hashes) are gated behind runAttributes and taskAttributes config flags, all defaulting to false. See the Metric Attributes and Cardinality section in crates/turborepo-otel/src/lib.rs for the full attribute inventory.
RunSummary.finish()
├── observability::Handle.record(&summary)
│ ├── Convert to RunMetricsPayload
│ └── Record via OpenTelemetry instruments
└── observability::Handle.shutdown()
└── Flush pending metrics to backend
crates/turborepo-log/)Structured event system for messages intended for end users (warnings,
errors, informational output). Distinct from tracing, which remains
for developer diagnostics.
Logger — Dispatches events to registered sinks. Set globally via
init() (once, at startup) or used directly via Logger::handle()
for testing.LogHandle — Source-scoped handle for emitting events. Created via
log() (global) or Logger::handle() (specific logger). Resolves
the global logger at .emit() time, not at handle or builder
creation time — handles and builders created before init() work
once the global logger is set.LogSink — Trait for event destinations. Built-in sinks:
CollectorSink (in-memory buffer for post-run summaries) and
FileSink (newline-delimited JSON with optional size limiting).LogEvent — Structured event with level, source, message, typed
fields, and timestamp.turborepo-uiturborepo-ui handles terminal rendering (TUI, console formatting).
turborepo-log handles structured event capture and dispatch. A
terminal sink in turborepo-ui can implement LogSink to bridge
events into the rendering pipeline. turborepo-log intentionally has
no dependency on turborepo-ui — it sits at the bottom of the
dependency graph.
Subsystem / Task Executor
└── LogHandle.warn("msg").field("k", v).emit()
└── Logger.emit(&event)
├── CollectorSink → in-memory buffer → post-run summary
└── FileSink → JSONL file → external tooling