Skip to content

feat: assistant evaluation ledger and shared-format dashboard - #2

Merged
Jammy2211 merged 2 commits into
mainfrom
feature/pyautobroca-assistant-management
Oct 8, 2026
Merged

Jammy2211 merged 2 commits into
mainfrom
feature/pyautobroca-assistant-management

Conversation

@Jammy2211

@Jammy2211 Jammy2211 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds a public home for assistant evaluation history and upkeep evidence, with a board built from Brain's shared presentation components. The initial snapshot covers four assistants, preserves 14 historical AutoLens benchmark runs, and records four separate maintenance inventory checks.

The board distinguishes missing evaluations, recorded failures, changed revisions, stale evidence and incomparable baselines. Refresh reads committed files; it never launches models or scientific compute. Existing assistant outputs remain untouched.

Closes #1 after the companion Brain and Mind registration PRs merge.

API Changes

New pyauto-broca / python -m broca commands: check, ingest, refresh, render. Adds version-1 immutable evaluation records and collection receipts. No modeling-library API changes. See full details below.

Test Plan

  • Broca unit/integration tests: 18 passed.
  • Real import: 14 historical benchmark records; four inventory records; complete collection of all four assistants.
  • Browser: five widths (390–1440px), light/dark, no page overflow, keyboard copy and exact preview/copy agreement.
  • Generated board rendered and visually inspected locally.
  • No scientific workspace smoke required: no modeling-library or workspace APIs changed. CLI collection/import/render is the relevant end-to-end smoke.
Full API Changes

Added

  • python -m broca --root PATH check: validate records and collection receipts.
  • python -m broca --root PATH ingest RECORD: exclusive/idempotent evidence insertion.
  • python -m broca --root PATH refresh --workspace PATH --mind PATH [--import-history] [--pilot]: read committed evidence and retain collection receipts.
  • python -m broca --root PATH render --brain PATH [--output dashboard.html]: render from the saved receipt without changing its freshness.
  • Record contract in docs/record-v1.md.

Migration

None. Public assistants do not depend on Broca. New evaluations should run in isolated checkouts and retain operational outputs outside public assistant history.

Integration and limits

Merge Brain #511 before this PR so CI's Brain-main checkout has the Broca identity. Mind registers the repository. Broca is public at the user’s explicit request, with no Pages deployment and no scheduled or paid campaign. The first pilot is an inventory check, not a new LLM response-quality evaluation. Historical runs lacking complete environment/scorer provenance are explicitly incomparable. Root/other-organ generated instruction propagation remains a post-merge repos_sync operation; this task touched only claimed repos.

Heart

YELLOW: manifest drift: workspace checkouts (manifest ↔ disk) — 1 mismatch(es) vs PyAutoMind/repos.yaml.
STALE evidence: release validation incomplete: no rehearsal for current source.
No active freeze. Human acknowledged this exact YELLOW reason and authorized opening the PRs on 2026-10-08. No release or merge authority implied.

Generated by the PyAutoLabs agent workflow.

Public visibility requested and applied on 2026-10-08 after reviewing tracked content. Raw transcripts and private project data remain excluded.

@Jammy2211
Jammy2211 merged commit 043f03a into main Oct 8, 2026
2 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: assistant evaluation ledger and shared-format dashboard

1 participant