docs/plans/5039-harden-registry-changelog-index-test.md
Objective: Replace PR #5096's live-inventory changelog assertions with stable behavioral invariants; done when mutation proof, focused tests, full check, autoreview, and GitHub CI pass.
Flow mode: one-shot execution
Goal plan: docs/plans/5039-harden-registry-changelog-index-test.md
Template: docs/plans/templates/task.md
Primary template: docs/plans/templates/task.md
Applied packs:
Task source:
bun check, autoreview, and final GitHub CI.Timed checkpoint:
Completion threshold:
--check, lint, full bun check, final
autoreview, and all required PR #5096 checks pass on the final head.node .agents/skills/autogoal/scripts/check-complete.mjs docs/plans/5039-harden-registry-changelog-index-test.md passes.Verification surface:
bun test tooling/scripts/generate-ui-changelog-entries.test.mjs.node tooling/scripts/generate-ui-changelog-entries.mjs --check.pnpm lint:fix, bun check, local autoreview, and gh pr checks 5096.Constraints:
Boundaries:
tooling/scripts/generate-ui-changelog-entries.test.mjs,
the repeatedly over-threshold BlockPlaceholderPlugin test classification,
this goal ledger, and PR #5096 body/branch.Output budget strategy:
Blocked condition:
Task state:
Current verdict:
--check owns artifact currency while this test should own behaviorPre-solution issue challenge:
Completion rule:
update_goal(status: complete) while any required checklist item
remains unchecked. If an item does not apply, check it and add N/A: <reason>.update_goal(status: complete) until every completion threshold
above is satisfied, final handoff evidence is recorded, and
node .agents/skills/autogoal/scripts/check-complete.mjs docs/plans/5039-harden-registry-changelog-index-test.md passes.Start Gates:
| Gate | Applies | Evidence |
|---|---|---|
| Timed checkpoint parsed | no | N/A: no duration requested |
| Skill analysis before edits | yes | Loaded autogoal, task, testing, and tdd; use autoreview at closeout |
| Active goal checked or created | yes | Active goal created for this exact plan and threshold |
| Source of truth read before edits | yes | Read current test, index builder, CLI --check, entry README, blame, and history |
| Tracker comments and attachments read | no | N/A: follow-up comes directly from the user; PR context was read in the preceding task |
| Video transcript evidence required | no | N/A: no video evidence |
| Pre-solution issue challenge required | yes | Claim is valid: test is deterministic but coupled to legitimate inventory churn |
| Reproduction verdict before implementation | yes | Historical 23-vs-22 red is authoritative; run temporary valid-entry mutation before the refactor |
| Repro escalation ladder selected | yes | Focused Node test owns proof; browser levels are N/A |
| Suggested fix reviewed against durable boundary | yes | Keep live integration for conservation; move exact semantics to synthetic inputs |
docs/solutions checked for non-trivial existing-code work | yes | No current matching solution file; read entry README, generator owner, test history, and prior registry verification guidance |
| TDD decision before behavior change or bug fix | yes | Test-only refactor: use mutation red/green; do not fake a production-code RED cycle |
| Branch decision for code-changing task | yes | Continue existing PR #5096 branch and push the full verified checkout |
| Release artifact decision | no | N/A: tooling-test-only change; no package or user-visible registry output |
| Browser tool decision for browser surface | no | N/A: no browser surface |
| PR expectation decision | yes | Update existing PR #5096 after full bun check passes |
| Tracker sync expectation decision | yes | Sync and verify PR body; no separate issue comment |
| Output budget strategy recorded | yes | Exact reads and capped outputs only |
Work Checklist:
<video-transcripts> XML, or marked N/A with reason. N/A: no video.valid, not reproduced, invalid,
wont-fix, partially valid, or platform limitation. Feature, docs,
support, or cleanup requests with no bug claim may mark reproduction
N/A with reason.[@Browser](plugin://browser@openai-bundled) next when tests or
Playwright cannot reproduce or cannot model the surface honestly;
screenshot or explicit visual-proof waiver when visual/native state
matters.pnpm run reinstall once only for install-corruption signals..agents/**, .claude/**,
.codex/**, skills, hooks, commands, prompts, or user-action tooling. N/A: no agent tooling changes.Completion Gates:
| Gate | Applies | Required action | Evidence |
|---|---|---|---|
| Named verification threshold | yes | Run the command, proof, source audit, or artifact check named in this plan | Mutation proof, focused/full local checks, autoreview, slow-lane proof, and code-head PR CI pass |
| Pre-solution issue challenge verdict | yes | Record reporter claim, suggested fix, repro verdict, validity verdict, durable boundary, and hard-stop/pivot decision before implementation | Valid; mutation proved inventory coupling and the split test boundary is durable |
| Repro escalation ladder | yes | For bug/behavior claims, record test/source-level, Playwright, Browser, and screenshot/visual-proof outcomes or N/A/blocker reasons before not reproduced | Focused mutation red/green complete; browser/visual levels N/A |
| Bug reproduced before fix | yes | Record failing test/repro or N/A with reason | Valid temporary entry failed 15/16 at 24 !== 23 before refactor |
| Targeted behavior verification | yes | Run focused test/proof for changed behavior or record N/A | Same mutation passed 17/17 after refactor; final focused suite 17/17 |
| TypeScript or typed config changed | no | Run relevant typecheck | N/A: JavaScript test only; full repo typecheck still passed in bun check |
| Package exports or file layout changed | no | Run pnpm brl before final verification and keep generated barrel updates | N/A: no package exports or layout changed; full CI barrel guard will run |
| Package manifests, lockfile, or install graph changed | no | Run pnpm install and relevant package checks | N/A: no manifest or lockfile changes |
| Agent rules or skills changed | no | Run pnpm install and verify generated skill sync | N/A: no agent rules or skills changed |
| Workspace authority proof | yes | Run verification in the owning repo/package/app/route/tool and record cwd; do not count the wrong workspace as proof | All local proof ran in /Users/zbeyens/git/plate; remote proof will run on PR #5096 |
| Browser surface changed | no | Capture Browser Use proof or record explicit waiver/blocker | N/A: Node test tooling only |
| Browser final proof | no | Attach screenshot or exact browser verification caveat when browser proof applies | N/A: no browser-visible claim |
| CI-controlled template output changed | no | Restore generated template output or record why it is intentionally kept | N/A: no templates/** files changed |
| Package behavior or public API changed | no | Add a changeset or record why no changeset applies | N/A: test-only follow-up; existing Link changeset remains unchanged |
| User-visible registry output changed | no | Use the registry-changelog pack: add/update apps/www/src/registry/changelog/entries/*.mdx, run node tooling/scripts/generate-ui-changelog-entries.mjs --write, run node tooling/scripts/generate-ui-changelog-entries.mjs --check, or record N/A | N/A: no entry or generated output changed; generator --check passes 23/23 |
| Docs or content changed | yes | For docs-heavy work, use --template docs; for supporting public docs/content/API/example changes, load docs-creator and close the docs pack; for typo/link-only edits, record the explicit reason and proportional proof | Internal required goal ledger only; no public docs/content/API/example change |
| High-risk mini gate | no | For public API/runtime/package-boundary/browser/agent-action/command-contract changes, record realistic failure mode, proof plan, and why the chosen boundary is right; otherwise N/A | N/A: test-only refactor; generator contract unchanged |
| Agent-native review for agent/tooling changes | no | For .agents/**, .claude/**, .codex/**, skills, hooks, commands, prompts, or user-action tooling, load .agents/skills/agent-native-reviewer/SKILL.md and close accepted/actionable findings, or record N/A | N/A: no agent/tooling action surface changed |
| Local install corruption suspected | no | Run pnpm run reinstall once, rerun the exact failing command, or record N/A | N/A: no corruption signal; all checks passed |
| Autoreview for non-trivial implementation changes | yes | Load .agents/skills/autoreview/SKILL.md; use dirty local --mode local, branch/PR --mode branch --base <base>, or committed slice --mode commit --commit <ref> until no accepted/actionable findings, or record N/A for docs-only/trivial/no local patch | Local autoreview clean, no findings, patch correct at 0.86 confidence |
| PR create or update | yes | Run check before PR work and sync PR body to the task-style final handoff | Commit 2cad7474e9 pushed; PR #5096 body synced after passing bun check |
| Task-style PR body verified | yes | Verify the PR body with gh pr view --json body; it must preserve auto-release blocks when applicable, must not include a current-PR self-link, and must use the kitcn PR #270 emoji format: ๐ Fixes ..., ๐ข 95-100% confidence, Phase / ๐งช Tests / ๐ Browser table, and bold emoji Outcome/Caveat/Design/Verified sections | gh pr view 5096 --json body confirms the auto-release block, required table/sections, and no self-link |
| PR proof image hosting | no | If PR body needs browser proof, replace local image paths with hosted GitHub URLs or record N/A | N/A: no browser or visual proof applies |
| Tracker sync-back | no | Post concise issue/Linear sync after PR exists, or record N/A/blocker | N/A: PR body links #5039 and is the requested existing tracker surface; no separate comment needed |
| Final handoff contract | yes | Fill the final handoff fields below with exact PR/issue/confidence/tests/browser/outcome/caveats/design/verification content or N/A reason | Filled below with PR, confidence, tests, browser waiver, outcome, caveat, design, and verification |
| Final lint | yes | Run pnpm lint:fix or scoped equivalent | 3,285 files checked; no fixes |
| Output budget discipline | yes | Verify no unbounded high-volume command output was streamed, or record the accidental output and recovery | Exact reads and capped command output only |
| Timed checkpoint | no | If duration was requested, keep improving until elapsed, then finish the current loop cleanly; otherwise N/A | N/A: no duration requested |
| Goal plan complete | yes | Run node .agents/skills/autogoal/scripts/check-complete.mjs docs/plans/5039-harden-registry-changelog-index-test.md | Pass |
Phase / pass table:
| Phase | Status | Evidence | Next |
|---|---|---|---|
| Intake and source read | complete | generator, test, README, history, and ownership read | implementation |
| Implementation | complete | inventory literals removed; synthetic contract plus live conservation retained | verification |
| Verification | complete | mutation red/green, 17 focused tests, generator check, lint, full bun check, and autoreview pass | PR sync |
| PR / tracker sync | complete | code head 4c6d8e9b8d; body verified; required checks green; PR mergeable | closeout |
| Closeout | complete | repeated CI offender reclassified with local and remote proof | final ledger push and goal closure |
Findings:
parseRegistryChangelogEntryFiles sorts .mdx paths; the index builder sorts
outputs by date, so the test is deterministic rather than flaky.--check already owns on-disk artifact currency. The current-file test should
prove one-to-one conservation; a synthetic test should own exact ordering,
href, and component fanout semantics.Decisions and tradeoffs:
Implementation notes:
BlockPlaceholderPlugin spec as *.slow.tsx
after two independent CI samples exceeded the fast-suite file threshold.Review fixes:
Error attempts:
| Error / failed attempt | Count | Next different move | Resolution |
|---|---|---|---|
PR CI slowest-test guard sampled unrelated BlockPlaceholderPlugin.spec.tsx over the 180 ms file threshold | 2 | First rerun before changing unrelated coverage; after recurrence, follow the guard and move the repeat offender to *.slow.tsx | One attempt passed; later final-head CI reproduced at 186 ms after the earlier 192 ms failure, proving unstable fast-lane classification |
Verification evidence:
24 !== 23; no generator implementation changed.--check: 23 events from 23 sources; all projections current.pnpm lint:fix: 3,285 files checked, no fixes.bun check: exit 0; lint, 54-package build/typecheck, 3,464 fast tests,
slow tests, and slowest-test guard passed..agents/skills/autoreview/scripts/autoreview --mode local --stream-engine-output: clean, no accepted/actionable findings.31875829305, final attempt: pass in 7m; all required checks green.gh pr view 5096 --json url,headRefOid,mergeable,body,statusCheckRollup:
head 2cad7474e9, mergeable, required task body verified.pnpm test:slow -- packages/utils/src/react/plugins/BlockPlaceholderPlugin.slow.tsx --rerun-each 3:
27/27 pass; unchanged assertions run through the slow harness.pnpm test:slowest -- --top 25: 3,455 fast tests pass; no file or test
exceeds the local threshold after reclassification.bun check: exit 0; 54-package build/typecheck, 3,455 fast
tests, 352 slow tests, and slowest-test guard pass.31877053635: pass in 6m52s on code head 4c6d8e9b8d;
changeset policy passes and PR remains mergeable.Final handoff contract:
4c6d8e9b8d๐ Fixes #5039; no separate sync required24 !== 23; browser N/A--check owns projection currencybun check, autoreview, and GitHub CITask-style PR body contract:
<!-- auto-release:start --> block. If a changeset is
part of the diff and repo policy expects auto release, include that block.๐ Fixes #123 or ๐ Fixes โ N/A, then
an emoji confidence line like ๐ข 95-100% confidence.| Phase | ๐งช Tests | ๐ Browser |.Reproduced and Verified rows. Mark passing proof with ๐ข, repro or
failing proof with ๐ด, and non-applicable cells with โ N/A.**โ
Outcome**, **โ ๏ธ Caveat**,
**๐๏ธ Design**, and **๐งช Verified**.Summary / Verification PR body, an
adaptive prose body from a git helper skill, plain ## Outcome sections, or
an unrelated generated badge footer unless the caller or repo template
explicitly asks for it.gh pr view --json body output or a concise source-backed summary
of that output.Final handoff / sync:
4c6d8e9b8d, mergeable, required checks greenTimeline:
24 !== 23 failure.bun check and dirty-local autoreview passed.2cad7474e9 pushed and PR #5096 task body verified.4c6d8e9b8d passed all required PR checks, including CI in 6m52s.Reboot status:
| Question | Answer |
|---|---|
| Where am I? | Final ledger closeout after green code-head CI |
| Where am I going? | Ledger checker, docs-only push, final-head verification, and goal closure |
| What is the goal? | Stable behavioral coverage without live inventory golden assertions |
| What have I learned? | See Findings |
| What have I done? | Mutation red/green proof, test refactor, full local verification, PR sync, and remote CI |
Open risks: