Repository navigation
feat: assistant evaluation ledger and shared-format dashboard - #2
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a public home for assistant evaluation history and upkeep evidence, with a board built from Brain's shared presentation components. The initial snapshot covers four assistants, preserves 14 historical AutoLens benchmark runs, and records four separate maintenance inventory checks.
The board distinguishes missing evaluations, recorded failures, changed revisions, stale evidence and incomparable baselines. Refresh reads committed files; it never launches models or scientific compute. Existing assistant outputs remain untouched.
Closes #1 after the companion Brain and Mind registration PRs merge.
API Changes
New
pyauto-broca/python -m brocacommands:check,ingest,refresh,render. Adds version-1 immutable evaluation records and collection receipts. No modeling-library API changes. See full details below.Test Plan
Full API Changes
Added
python -m broca --root PATH check: validate records and collection receipts.python -m broca --root PATH ingest RECORD: exclusive/idempotent evidence insertion.python -m broca --root PATH refresh --workspace PATH --mind PATH [--import-history] [--pilot]: read committed evidence and retain collection receipts.python -m broca --root PATH render --brain PATH [--output dashboard.html]: render from the saved receipt without changing its freshness.docs/record-v1.md.Migration
None. Public assistants do not depend on Broca. New evaluations should run in isolated checkouts and retain operational outputs outside public assistant history.
Integration and limits
Merge Brain #511 before this PR so CI's Brain-main checkout has the Broca identity. Mind registers the repository. Broca is public at the user’s explicit request, with no Pages deployment and no scheduled or paid campaign. The first pilot is an inventory check, not a new LLM response-quality evaluation. Historical runs lacking complete environment/scorer provenance are explicitly incomparable. Root/other-organ generated instruction propagation remains a post-merge repos_sync operation; this task touched only claimed repos.
Heart
YELLOW:
manifest drift: workspace checkouts (manifest ↔ disk) — 1 mismatch(es) vs PyAutoMind/repos.yaml.STALE evidence:
release validation incomplete: no rehearsal for current source.No active freeze. Human acknowledged this exact YELLOW reason and authorized opening the PRs on 2026-10-08. No release or merge authority implied.
Generated by the PyAutoLabs agent workflow.
Public visibility requested and applied on 2026-10-08 after reviewing tracked content. Raw transcripts and private project data remain excluded.