diff --git a/.agents/skills/classify-note.md b/.agents/skills/classify-note/SKILL.md similarity index 86% rename from .agents/skills/classify-note.md rename to .agents/skills/classify-note/SKILL.md index 4bc40bd..19d8d0b 100644 --- a/.agents/skills/classify-note.md +++ b/.agents/skills/classify-note/SKILL.md @@ -32,11 +32,11 @@ Implemented (2026-06-13) as a focused sub-skill of ingest-source. Use it when th - Respect sensitivity overrides from `.context/routing-policy.md` (finance state etc. → Tier C, propose 20_live or park). ## Related -- [ingest-source](ingest-source.md) — calls this (or equivalent logic) as step 3. -- [extract-metadata](extract-metadata.md) — sibling for other frontmatter fields. +- [ingest-source](../ingest-source/SKILL.md) — calls this (or equivalent logic) as step 3. +- [extract-metadata](../extract-metadata/SKILL.md) — sibling for other frontmatter fields. - `.context/routing-policy.md` and `10_knowledge/index.md` — the source of truth for rules and domain inventory. ## Related -- [ingest-source](ingest-source.md) — the parent skill -- [extract-metadata](extract-metadata.md) — fills in the rest of frontmatter +- [ingest-source](../ingest-source/SKILL.md) — the parent skill +- [extract-metadata](../extract-metadata/SKILL.md) — fills in the rest of frontmatter diff --git a/.agents/skills/craft-research-loop/SKILL.md b/.agents/skills/craft-research-loop/SKILL.md index 71ce5e0..af0b76b 100644 --- a/.agents/skills/craft-research-loop/SKILL.md +++ b/.agents/skills/craft-research-loop/SKILL.md @@ -67,4 +67,4 @@ bin/craft-research-loop promote --project --trial --to knowledge|exp - `.context/workflows/craft-research-loop.md` - `.context/templates/craft-trial.md` - research-lane-loop, project-experiment-loop -- Reference implementation: `30_projects/image-generation-lab/outputs/RESEARCH_PROOF_INDEX.md` +- Reference implementation: `local-only: 30_projects//outputs/RESEARCH_PROOF_INDEX.md` diff --git a/.agents/skills/create-source-summary.md b/.agents/skills/create-source-summary.md deleted file mode 100644 index 11922c3..0000000 --- a/.agents/skills/create-source-summary.md +++ /dev/null @@ -1,32 +0,0 @@ ---- -name: create-source-summary -description: For raw items that need a separate searchable note alongside the immutable evidence, create a sibling summary file that links back to the source. -status: active ---- - -# create-source-summary - -## Purpose - -Some raw items (PDFs, long transcripts, scraped articles) are too dense for retrieval as-is. The ingest minion already creates a markdown stub for PDFs (per [ADR-007](../../DECISIONS.md)); this skill is the agent-driven version for markdown raw items that warrant a synthesized companion note. - -The original raw file is preserved as evidence. A new `note`-type file is created alongside it, containing: -- Summary of the source -- Key claims or extracted facts -- Wikilinks to existing knowledge -- Citation back to the raw source - -## Status - -**Active** — delegates directly to the `extraction-agent`. - -When invoked (e.g. from `ingest-source` during raw enrichment), immediately pass the target file to the `extraction-agent` (defined in `../../agents/extraction-agent.md`) which handles the actual deep reading and extraction workflow. - -Do not attempt to execute the extraction logic manually within this skill; the extraction agent contains the strict guardrails for formatting, routing through `01_ingest/ready/`, and maintaining evidence immutability. - -**Integration note (2026-06-13):** Now explicitly called out from the `ingest-source` skill for cases where a raw benefits from a synthesized companion `note`. This supports Fix 3 goals of increasing synthesis rate vs raw accumulation. - -## Related - -- [ingest-source](ingest-source.md) — the parent skill -- [ADR-007](../../DECISIONS.md) — defines the existing PDF stub-creation behavior diff --git a/.claude/skills/create-source-summary.md b/.agents/skills/create-source-summary/SKILL.md similarity index 81% rename from .claude/skills/create-source-summary.md rename to .agents/skills/create-source-summary/SKILL.md index 11922c3..0f22c03 100644 --- a/.claude/skills/create-source-summary.md +++ b/.agents/skills/create-source-summary/SKILL.md @@ -8,7 +8,7 @@ status: active ## Purpose -Some raw items (PDFs, long transcripts, scraped articles) are too dense for retrieval as-is. The ingest minion already creates a markdown stub for PDFs (per [ADR-007](../../DECISIONS.md)); this skill is the agent-driven version for markdown raw items that warrant a synthesized companion note. +Some raw items (PDFs, long transcripts, scraped articles) are too dense for retrieval as-is. The ingest minion already creates a markdown stub for PDFs (per [ADR-007](../../../DECISIONS.md)); this skill is the agent-driven version for markdown raw items that warrant a synthesized companion note. The original raw file is preserved as evidence. A new `note`-type file is created alongside it, containing: - Summary of the source @@ -28,5 +28,5 @@ Do not attempt to execute the extraction logic manually within this skill; the e ## Related -- [ingest-source](ingest-source.md) — the parent skill -- [ADR-007](../../DECISIONS.md) — defines the existing PDF stub-creation behavior +- [ingest-source](../ingest-source/SKILL.md) — the parent skill +- [ADR-007](../../../DECISIONS.md) — defines the existing PDF stub-creation behavior diff --git a/.agents/skills/extract-metadata.md b/.agents/skills/extract-metadata.md deleted file mode 100644 index b77a888..0000000 --- a/.agents/skills/extract-metadata.md +++ /dev/null @@ -1,26 +0,0 @@ ---- -name: extract-metadata -description: Fill in remaining frontmatter fields (source, links, etc.) and update status as the file progresses through agent enrichment. -status: stub ---- - -# extract-metadata - -## Purpose - -After [classify-note](classify-note.md) has set `domain`, `type`, and `tags`, fill in the rest of the frontmatter: -- **`source`** — preserve if already set; otherwise infer from filename, body, or capture context. -- **`links`** — preserve the minion-extracted array; add agent-proposed connections. -- **`status`** — transition `skimmed → routed` when starting work, `routed → extracted` when done. - -## Status - -**Stub** — to be expanded. Initial guidance: -- Source is provenance — preserve original URLs, file paths, or attribution. -- Don't fabricate sources. If unknown, set `source: "unknown"` and surface to user. -- Status transitions are explicit signals to `bin/prep-ingest`; don't skip the intermediate `routed` state. - -## Related - -- [ingest-source](ingest-source.md) — the parent skill -- [.context/primitives.md](../../.context/primitives.md) — schema and status lifecycle diff --git a/.claude/skills/extract-metadata.md b/.agents/skills/extract-metadata/SKILL.md similarity index 75% rename from .claude/skills/extract-metadata.md rename to .agents/skills/extract-metadata/SKILL.md index b77a888..d84599d 100644 --- a/.claude/skills/extract-metadata.md +++ b/.agents/skills/extract-metadata/SKILL.md @@ -8,7 +8,7 @@ status: stub ## Purpose -After [classify-note](classify-note.md) has set `domain`, `type`, and `tags`, fill in the rest of the frontmatter: +After [classify-note](../classify-note/SKILL.md) has set `domain`, `type`, and `tags`, fill in the rest of the frontmatter: - **`source`** — preserve if already set; otherwise infer from filename, body, or capture context. - **`links`** — preserve the minion-extracted array; add agent-proposed connections. - **`status`** — transition `skimmed → routed` when starting work, `routed → extracted` when done. @@ -22,5 +22,5 @@ After [classify-note](classify-note.md) has set `domain`, `type`, and `tags`, fi ## Related -- [ingest-source](ingest-source.md) — the parent skill -- [.context/primitives.md](../../.context/primitives.md) — schema and status lifecycle +- [ingest-source](../ingest-source/SKILL.md) — the parent skill +- [.context/primitives.md](../../../.context/primitives.md) — schema and status lifecycle diff --git a/.agents/skills/ingest-source.md b/.agents/skills/ingest-source/SKILL.md similarity index 94% rename from .agents/skills/ingest-source.md rename to .agents/skills/ingest-source/SKILL.md index 881f6aa..900939f 100644 --- a/.agents/skills/ingest-source.md +++ b/.agents/skills/ingest-source/SKILL.md @@ -10,7 +10,7 @@ status: implemented The canonical, reusable skill for the per-file judgment pass in the two-pass ingest architecture (ADR-009, ADR-011, ADR-019). It encapsulates everything between the deterministic minion pass 1 (which produces `status: skimmed` files in `01_ingest/ready/`) and the handoff to `bin/prep-ingest` (which moves to `queue/`) + minion pass 2 routing. -The skill is invoked by the [ingest-agent](../../agents/ingest-agent.md) subagent. After this skill completes its proposal + user confirmation + enrichment, the file is ready for the strict deterministic gate. +The skill is invoked by the [ingest-agent](../../../agents/ingest-agent.md) subagent. After this skill completes its proposal + user confirmation + enrichment, the file is ready for the strict deterministic gate. ## Inputs - A Markdown file in `01_ingest/ready/` with at minimum the minion-normalized frontmatter (`title`, `domain` possibly empty, `type`, `status: skimmed`, `source`, `tags`, and the `links:` array extracted from body wikilinks). @@ -94,7 +94,7 @@ The skill is invoked by the [ingest-agent](../../agents/ingest-agent.md) subagen - `bin/prep-ingest`, `bin/ingest-minion`, `.context/workflows/audit-sweep.md`, `10_knowledge/index.md`, `.context/routing-policy.md`, `01_ingest/AGENTS.md`. ## Evaluation -This skill should be exercised and measured during process evaluations (see `30_projects/mainframe-process-eval/`). Track: proposal acceptance rate, time-to-extracted, downstream routing success rate, number of needs-audit tags correctly applied, and any Tier B exceptions that later graduate to rules. +This skill should be exercised and measured during process evaluations (see `40_operations/mainframe-process-eval/`). Track: proposal acceptance rate, time-to-extracted, downstream routing success rate, number of needs-audit tags correctly applied, and any Tier B exceptions that later graduate to rules. ## Status Note This skill was promoted from stub (2026-06-13) as part of closing the gap between the detailed agent procedure and reusable skill contracts. The long-form procedure now lives here; the ingest-agent definition should remain focused on role, guardrails, batch vs. per-file mode, and orchestration. diff --git a/.agents/skills/mindgraph-retrieval/SKILL.md b/.agents/skills/mindgraph-retrieval/SKILL.md index bf145cd..698fb2b 100644 --- a/.agents/skills/mindgraph-retrieval/SKILL.md +++ b/.agents/skills/mindgraph-retrieval/SKILL.md @@ -1,6 +1,7 @@ --- name: mindgraph-retrieval -description: Use when an agent needs to retrieve durable knowledge or active project context from MainFrame's local graph-augmented search engine (MindGraph). It provides instructions on CLI/MCP commands and prevents hallucination of non-existent search verbs. Triggers on: "query mindgraph", "search mindgraph", "mindgraph query", "retrieve knowledge", "find in knowledge base", "query knowledge", "query projects", "mindgraph-refresh", "mindgraph doctor", "mindgraph status". +description: >- + Use when an agent needs to retrieve durable knowledge or active project context from MainFrame's local graph-augmented search engine (MindGraph). It provides instructions on CLI/MCP commands and prevents hallucination of non-existent search verbs. Triggers on: "query mindgraph", "search mindgraph", "mindgraph query", "retrieve knowledge", "find in knowledge base", "query knowledge", "query projects", "mindgraph-refresh", "mindgraph doctor", "mindgraph status". status: active --- diff --git a/.agents/skills/print-ready-pdf-generation/SKILL.md b/.agents/skills/print-ready-pdf-generation/SKILL.md deleted file mode 100644 index e438add..0000000 --- a/.agents/skills/print-ready-pdf-generation/SKILL.md +++ /dev/null @@ -1,63 +0,0 @@ ---- -name: print-ready-pdf-generation -description: Create, repair, and verify polished print-ready PDFs from Markdown or HTML for agent-generated deliverables. Use when Codex needs to regenerate an attachment, fix Chrome/Puppeteer/Playwright PDF layout issues, handle md-to-pdf hangs, prevent header/footer overlap, stop clipped tables or overflowing text, or leave a reusable PDF generation path for future agents. ---- - -# Print-Ready PDF Generation - -## Overview - -Use this skill to turn Markdown or HTML into a visually checked PDF that is safe to send as an attachment. Prefer a deterministic render script plus PNG inspection over one-off browser printing. - -## Workflow - -1. Find the source file. Do not edit the generated PDF directly unless the user specifically asks for binary PDF manipulation. -2. Inspect any existing generator notes, CSS, config, and prior PDF output. -3. Render the current PDF to PNG with Poppler and inspect at least the first page, every page with a table, and the last page. -4. Fix the source, CSS, or config. Common fixes: - - Remove CSS `@page margin: 0` when Chrome header/footer templates are enabled; let the PDF margin options reserve header/footer space. - - Keep body padding small when PDF margins are already set. - - Use `table-layout: fixed`, `overflow-wrap: anywhere`, and normal white-space for narrative table cells. - - Avoid global `td:nth-child(...) { white-space: nowrap; }` rules unless the table really contains short numeric values. - - Increase top/bottom PDF margins when header or footer content overlaps body text. -5. Regenerate the PDF with `scripts/render-markdown-pdf.mjs`. -6. Render the new PDF to PNG and inspect visually before delivery. Text extraction is useful for smoke checks, but it is not layout verification. - -## Script - -Run from the workspace root or the source file directory: - -```bash -node .agents/skills/print-ready-pdf-generation/scripts/render-markdown-pdf.mjs \ - path/to/source.md \ - path/to/output.pdf \ - --config path/to/pdf-config.js -``` - -The script: - -- strips YAML frontmatter from Markdown; -- converts Markdown with `marked`; -- inlines CSS from `pdf-config.js` and optional `--css` flags; -- renders with Playwright/Chromium, preferring system Chrome when bundled Playwright browsers are absent; -- supports `--keep-html` for debugging and `--fail-on-overflow` for obvious screen-layout overflow checks. - -## Verification Commands - -```bash -pdfinfo path/to/output.pdf -pdftoppm -png -r 150 path/to/output.pdf /tmp/pdf-check/page -pdftotext -layout path/to/output.pdf - -``` - -Use `view_image` on the rendered PNGs when available. Check for: - -- header/body and footer/body overlap; -- clipped tables or text running outside margins; -- awkward page starts where a heading is separated from its first paragraph; -- unreadable fonts, broken glyphs, or placeholder/source tokens; -- correct page count and intended output path. - -## Output Discipline - -Keep final artifacts where the project rules say they belong: source notes and generation assets stay with the project, while generated phone-uploadable PDFs go to the configured output vault. diff --git a/.agents/skills/print-ready-pdf-generation/agents/openai.yaml b/.agents/skills/print-ready-pdf-generation/agents/openai.yaml deleted file mode 100644 index 8b612a4..0000000 --- a/.agents/skills/print-ready-pdf-generation/agents/openai.yaml +++ /dev/null @@ -1,4 +0,0 @@ -interface: - display_name: "Print-Ready PDF Generation" - short_description: "Generate and repair polished PDFs" - default_prompt: "Use $print-ready-pdf-generation to regenerate this Markdown or HTML file as a polished, verified PDF." diff --git a/.agents/skills/print-ready-pdf-generation/scripts/render-markdown-pdf.mjs b/.agents/skills/print-ready-pdf-generation/scripts/render-markdown-pdf.mjs deleted file mode 100755 index b91a449..0000000 --- a/.agents/skills/print-ready-pdf-generation/scripts/render-markdown-pdf.mjs +++ /dev/null @@ -1,281 +0,0 @@ -#!/usr/bin/env node -import fs from "node:fs"; -import path from "node:path"; -import { createRequire } from "node:module"; -import { fileURLToPath, pathToFileURL } from "node:url"; - -const __filename = fileURLToPath(import.meta.url); -const __dirname = path.dirname(__filename); - -function usage() { - console.error(`Usage: - node render-markdown-pdf.mjs [options] - -Options: - --config CommonJS config with pdf_options and stylesheet entries - --css CSS file to inline; may be repeated - --chrome Chrome/Chromium executable path - --keep-html Write the intermediate HTML for inspection - --fail-on-overflow Exit non-zero if obvious screen overflow is detected - -The script resolves marked/playwright from local node_modules, NODE_PATH, or -the Codex bundled dependency path when present.`); -} - -function parseArgs(argv) { - const parsed = { - css: [], - config: null, - chrome: null, - keepHtml: null, - failOnOverflow: false, - positional: [], - }; - - // A value-taking option must never swallow the next flag. - // - // Found 2026-08-10: a 24 KB file named `--fail-on-overflow` had been sitting in - // the MainFrame repo root since 2026-08-04. Someone ran - // - // render-markdown-pdf.mjs in.md out.pdf --keep-html --fail-on-overflow - // - // and `--keep-html` took `--fail-on-overflow` as its output path. The stray - // file is the harmless half. The real damage is that `failOnOverflow` stayed - // false, so **the overflow check the caller explicitly asked for never ran**, - // the script exited 0, and a PDF shipped with no layout verification and no - // warning that there had been none. - // - // A guard that is requested but silently not in the path is worse than no - // guard, because it produces confidence instead of caution. - let i = 0; - const takesValue = (flag) => { - const next = argv[i + 1]; - if (next === undefined || next.startsWith("--")) { - throw new Error( - `${flag} expects a value, but got ` + - `${next === undefined ? "end of arguments" : `the flag ${next}`}. ` + - `If you meant both, write: ${flag} ${next ?? ""}`.trimEnd(), - ); - } - i += 1; - return next; - }; - - for (; i < argv.length; i += 1) { - const arg = argv[i]; - if (arg === "--css") parsed.css.push(takesValue("--css")); - else if (arg === "--config") parsed.config = takesValue("--config"); - else if (arg === "--chrome") parsed.chrome = takesValue("--chrome"); - else if (arg === "--keep-html") parsed.keepHtml = takesValue("--keep-html"); - else if (arg === "--fail-on-overflow") parsed.failOnOverflow = true; - else if (arg === "--help" || arg === "-h") { - usage(); - process.exit(0); - } else if (arg.startsWith("--")) { - throw new Error(`Unknown option: ${arg}`); - } else { - parsed.positional.push(arg); - } - } - - if (parsed.positional.length !== 2) { - usage(); - process.exit(2); - } - return parsed; -} - -function existingPaths(paths) { - return paths.filter(Boolean).filter((candidate) => fs.existsSync(candidate)); -} - -function moduleDirs() { - const dirs = []; - if (process.env.NODE_PATH) dirs.push(...process.env.NODE_PATH.split(path.delimiter)); - dirs.push( - path.resolve(process.cwd(), "node_modules"), - path.resolve(__dirname, "..", "node_modules"), - path.resolve(__dirname, "..", "..", "node_modules"), - path.join(process.env.HOME || "", ".cache/codex-runtimes/codex-primary-runtime/dependencies/node/node_modules"), - ); - return existingPaths([...new Set(dirs)]); -} - -async function importPackage(packageName) { - const errors = []; - for (const dir of moduleDirs()) { - try { - const req = createRequire(path.join(dir, "_resolver.cjs")); - const resolved = req.resolve(packageName); - return await import(pathToFileURL(resolved).href); - } catch (error) { - errors.push(`${dir}: ${error.message}`); - } - } - - try { - return await import(packageName); - } catch (error) { - errors.push(`default resolver: ${error.message}`); - } - - throw new Error(`Could not resolve ${packageName}.\n${errors.join("\n")}`); -} - -function stripFrontmatter(input) { - return input.replace(/^---\s*\r?\n[\s\S]*?\r?\n---\s*\r?\n/, ""); -} - -function readConfig(configPath) { - if (!configPath) return {}; - const absolute = path.resolve(configPath); - const req = createRequire(pathToFileURL(absolute).href); - return { - config: req(absolute), - baseDir: path.dirname(absolute), - }; -} - -function resolveRelative(filePath, baseDirs) { - if (!filePath) return null; - if (path.isAbsolute(filePath)) return filePath; - for (const baseDir of baseDirs) { - const candidate = path.resolve(baseDir, filePath); - if (fs.existsSync(candidate)) return candidate; - } - return path.resolve(baseDirs[0] || process.cwd(), filePath); -} - -function readCss(paths) { - return paths - .filter(Boolean) - .map((cssPath) => fs.readFileSync(cssPath, "utf8")) - .join("\n\n"); -} - -function findChrome(explicitPath) { - const candidates = existingPaths([ - explicitPath, - process.env.CHROME_PATH, - process.env.PUPPETEER_EXECUTABLE_PATH, - process.env.PLAYWRIGHT_CHROMIUM_EXECUTABLE_PATH, - "/Applications/Google Chrome.app/Contents/MacOS/Google Chrome", - "/Applications/Chromium.app/Contents/MacOS/Chromium", - "/Applications/Google Chrome Canary.app/Contents/MacOS/Google Chrome Canary", - "/usr/bin/google-chrome", - "/usr/bin/chromium", - "/usr/bin/chromium-browser", - ]); - return candidates[0] || null; -} - -function htmlShell({ title, css, body }) { - return ` - - - - ${title.replaceAll("&", "&").replaceAll("<", "<")} - - - -${body} - -`; -} - -async function main() { - const args = parseArgs(process.argv.slice(2)); - const inputPath = path.resolve(args.positional[0]); - const outputPath = path.resolve(args.positional[1]); - const inputDir = path.dirname(inputPath); - - const configData = readConfig(args.config); - const config = configData.config || {}; - const configBaseDir = configData.baseDir || inputDir; - const baseDirs = [configBaseDir, inputDir, process.cwd()]; - - const stylesheetEntries = [ - ...(Array.isArray(config.stylesheet) ? config.stylesheet : []), - ...(Array.isArray(config.stylesheets) ? config.stylesheets : []), - ...args.css, - ]; - const cssPaths = stylesheetEntries.map((entry) => resolveRelative(entry, baseDirs)); - const css = readCss(cssPaths); - - const raw = fs.readFileSync(inputPath, "utf8"); - const extension = path.extname(inputPath).toLowerCase(); - let body; - if (extension === ".html" || extension === ".htm") { - body = raw; - } else { - const markedModule = await importPackage("marked"); - const marked = markedModule.marked || markedModule.default || markedModule; - body = await marked.parse(stripFrontmatter(raw)); - } - - const html = htmlShell({ - title: path.basename(inputPath), - css, - body, - }); - - if (args.keepHtml) { - fs.mkdirSync(path.dirname(path.resolve(args.keepHtml)), { recursive: true }); - fs.writeFileSync(path.resolve(args.keepHtml), html); - } - - const playwrightModule = await importPackage("playwright"); - const { chromium } = playwrightModule.default || playwrightModule; - if (!chromium) { - throw new Error("Resolved playwright, but could not find chromium export."); - } - const executablePath = findChrome(args.chrome); - const launchOptions = executablePath - ? { headless: true, executablePath, args: ["--no-sandbox"] } - : { headless: true, args: ["--no-sandbox"] }; - - const browser = await chromium.launch(launchOptions); - try { - const page = await browser.newPage({ viewport: { width: 816, height: 1056 } }); - await page.emulateMedia({ media: "print" }); - await page.setContent(html, { waitUntil: "networkidle" }); - await page.evaluate(() => document.fonts && document.fonts.ready); - - const overflowing = await page.evaluate(() => - Array.from(document.querySelectorAll("body *")) - .filter((element) => element.scrollWidth > element.clientWidth + 2) - .slice(0, 10) - .map((element) => ({ - tag: element.tagName.toLowerCase(), - text: (element.textContent || "").trim().slice(0, 90), - clientWidth: element.clientWidth, - scrollWidth: element.scrollWidth, - })), - ); - - if (overflowing.length) { - console.warn("Potential overflowing elements before print:"); - console.warn(JSON.stringify(overflowing, null, 2)); - if (args.failOnOverflow) process.exitCode = 1; - } - - const pdfOptions = { - format: "Letter", - printBackground: true, - ...(config.pdf_options || {}), - path: outputPath, - }; - fs.mkdirSync(path.dirname(outputPath), { recursive: true }); - await page.pdf(pdfOptions); - } finally { - await browser.close(); - } - - if (process.exitCode) return; - console.log(`Wrote ${outputPath}`); -} - -main().catch((error) => { - console.error(error.stack || error.message || error); - process.exit(1); -}); diff --git a/.agents/skills/rename-material.md b/.agents/skills/rename-material.md deleted file mode 100644 index 3a16ade..0000000 --- a/.agents/skills/rename-material.md +++ /dev/null @@ -1,24 +0,0 @@ ---- -name: rename-material -description: Rename a file to the Mainframe convention YYYY-MM-DD__domain__type__slug.md after enrichment has resolved domain and type. -status: stub ---- - -# rename-material - -## Purpose - -Rename a file in `01_ingest/ready/` to the convention `YYYY-MM-DD__domain__type__slug.md` once the agent has resolved `domain` and `type`. This is the last step before `bin/prep-ingest` validates and moves to `queue/`. - -## Status - -**Stub** — to be expanded. Initial guidance: -- Date comes from frontmatter `captured` or `created` field; falls back to current date if missing. -- Slug is kebab-case, derived from title. -- Domain must be in the existing whitelist (subdirectories of `10_knowledge/`). -- Type must be `raw` or `note` for files routing through ingest (per ADR-007). - -## Related - -- [ingest-source](ingest-source.md) — the parent skill -- [classify-note](classify-note.md) — produces the domain/type the rename consumes diff --git a/.claude/skills/rename-material.md b/.agents/skills/rename-material/SKILL.md similarity index 83% rename from .claude/skills/rename-material.md rename to .agents/skills/rename-material/SKILL.md index 3a16ade..0a73205 100644 --- a/.claude/skills/rename-material.md +++ b/.agents/skills/rename-material/SKILL.md @@ -20,5 +20,5 @@ Rename a file in `01_ingest/ready/` to the convention `YYYY-MM-DD__domain__type_ ## Related -- [ingest-source](ingest-source.md) — the parent skill -- [classify-note](classify-note.md) — produces the domain/type the rename consumes +- [ingest-source](../ingest-source/SKILL.md) — the parent skill +- [classify-note](../classify-note/SKILL.md) — produces the domain/type the rename consumes diff --git a/.agents/skills/research-lane-loop/SKILL.md b/.agents/skills/research-lane-loop/SKILL.md index 09124b8..ce931ad 100644 --- a/.agents/skills/research-lane-loop/SKILL.md +++ b/.agents/skills/research-lane-loop/SKILL.md @@ -116,8 +116,7 @@ Do **not** add: research claims outside the synthesis note, new lanes without in - `.context/workflows/ingest-minion.md` — step 4 - `.context/workflows/epistemic-standard.md` — step 5 - `.context/workflows/research-lane-intake.md` — new lanes only -- `30_projects/research-lanes-strategy/plans/knowledge-routing.md` -- `30_projects/research-lanes-strategy/plans/first-principles-research-conventions.md` +- `local-only: 30_projects//plans/` for operator-specific lane conventions ## Evaluation diff --git a/.agents/skills/skill-creation/SKILL.md b/.agents/skills/skill-creation/SKILL.md new file mode 100644 index 0000000..051a441 --- /dev/null +++ b/.agents/skills/skill-creation/SKILL.md @@ -0,0 +1,82 @@ +--- +name: skill-creation +description: > + Use when designing, writing, auditing, or improving a SKILL.md file or any agent instruction + document in this workspace. Triggers on: "create a skill", "write a skill", "add a skill", + "audit this skill", "audit a skill", "make a new skill", "design an instruction doc". + Do NOT use for one-off task prompts or system prompts for LLMs (use prompt-creation), for + AGENTS.md or HARNESS.md system contracts, or for workflow files in .context/workflows/. +--- + +# Skill Creation + +## Purpose + +Produce a well-structured, reliable SKILL.md (or any agent instruction document) that activates at +the right time, produces consistent output, and fails gracefully at its edges. + +A skill is a training manual for an agent-employee. The failure modes are identical across SKILL.md, +AGENTS.md, CLAUDE.md, and `prompts/*.md`. Learn to diagnose by mode; the fix becomes obvious. + +## Load Order + +Always read `references/anatomy.md` first — it contains the five-component checklist, five failure +modes, the pre-ship testing protocol, and the thin-router pattern. + +Then load only the references needed: + +- `references/anatomy.md`: required for every skill-creation task +- `references/activation-language.md`: load when writing or auditing the YAML `description` block + +## Workflow + +### Writing a new skill + +1. **Name the job** — Write one sentence: "This skill owns [single task category]." If it covers more than one independently-triggerable category, split into two skills before drafting. +2. **Load references** — Load `references/anatomy.md` and `references/activation-language.md`. +3. **Draft the YAML block** — `name` (kebab-case), `description` using the template in `references/activation-language.md`: 5–7 explicit trigger phrases + at least one "Do NOT use for" clause, third-person phrasing, description text under 60 words. +4. **Write the body** — Five components in order: (1) Purpose — one paragraph written for the agent; (2) Load Order — only if external references exist, otherwise omit the section entirely; (3) Workflow — numbered imperative steps, no vague verbs; (4) Output Format — structure + explicit "Do NOT add" list; (5) Examples and edge cases — at least one happy path, at least one `If [condition], then [action]` rule. +5. **Apply the thin-router test** — If the body exceeds ~150 lines, move workflow content to `references/workflow.md` and replace with a single pointer line. +6. **Run the five-test protocol** from `references/anatomy.md`: happy path, minimal input, edge case, negative test, repeat test. +7. **Check every item** on the pre-ship checklist in `references/anatomy.md`. + +### Auditing an existing skill + +1. **Read the skill in full** before writing anything. +2. **Score each of the five components** as present / weak / missing. +3. **Run the five failure-mode diagnosis** — for each failure mode, state whether the skill has symptoms and what the fix is. +4. **Run the five-test protocol** mentally against the skill's current text. +5. **List findings** as: component, failure mode, specific line, replacement text. +6. **Apply fixes** one at a time. Re-check the affected component after each fix. + +## Output Format + +Deliver: +1. The complete SKILL.md as a ready-to-copy file block (for new skills), or the changed lines with exact replacement text (for audits). +2. A checklist showing each pre-ship item as pass / fail / note. +3. For any fail item: the specific line and the exact replacement text. + +Do NOT add: preamble explaining skill design theory, hedging disclaimers, sections not in the five-component structure, or alternative versions unless asked. + +## Examples and Edge Cases + +**Happy path:** See `examples/happy-path.md` for a complete worked example of writing a new skill from a user request through to a shipped SKILL.md with checklist. + +**If the user provides only a skill name with no description:** ask for (1) the single task category in one sentence, (2) three example user phrasings that should trigger it, (3) one request that should NOT trigger it. Do not draft until all three are answered. + +**If the user asks to audit a skill but no task context is given:** follow the Auditing workflow above. Do not ask for the intended output — diagnose from the existing text using the five-component and five-failure-mode frameworks. + +**If the requested skill overlaps with an existing skill:** name the overlapping trigger phrases, state which "Do NOT use for" clause would prevent collision, and confirm with the user before writing. + +**If the skill body grows past ~150 lines during drafting:** stop. Move workflow content to `references/workflow.md`. Replace with: "Follow the workflow in `references/workflow.md`." Note this in your checklist output. + +**If the skill needs no external references:** omit the Load Order section entirely. Do not include it with an empty list. + +## Boundaries + +This skill governs SKILL.md files and agent instruction documents only. + +Do NOT use it for: +- One-off task prompts or system prompts for LLMs → use `prompt-creation` +- Editing `AGENTS.md`, `HARNESS.md`, or `.context/workflows/` files +- Auditing or editing code files diff --git a/.agents/skills/skill-creation/examples/happy-path.md b/.agents/skills/skill-creation/examples/happy-path.md new file mode 100644 index 0000000..302484a --- /dev/null +++ b/.agents/skills/skill-creation/examples/happy-path.md @@ -0,0 +1,89 @@ +# Happy Path Example — Prompt Creation + +This file satisfies Component #5 (examples) for the `skill-creation` skill's own pre-ship checklist. + +--- + +## Scenario + +User says: "Write a prompt that gets Claude to review a pull request and flag any security issues." + +--- + +## Step 1 — Name the job + +> "This skill owns: designing a code-review prompt that targets security issues specifically." + +One task category. No split needed. + +--- + +## Step 2 — Map the six elements + +| Element | Content | +|---------|---------| +| **Role** | A senior application security engineer with experience in OWASP Top 10 and secure code review | +| **Context** | A pull request diff will be provided. The reviewer is a developer, not a security specialist. | +| **Task** | Review the diff for security vulnerabilities; categorise each by OWASP Top 10 category if applicable | +| **Format** | Bulleted list of findings; each finding: file + line range, category, description, recommended fix | +| **Constraints** | Do NOT flag style or performance issues. Do NOT praise the code. Do NOT include findings with no concrete fix. | +| **Quality standard** | Every finding must be actionable: a developer should be able to implement the fix without follow-up questions | + +--- + +## Step 3 — Draft (XML-tagged, ≥ 3 separable components) + +```xml + +You are a senior application security engineer. You specialise in secure code review and OWASP Top 10 vulnerabilities. + + + +The reader is a developer, not a security specialist. Write findings so they are understandable without prior security training. + + + +Review the pull request diff below for security vulnerabilities only. For each finding, identify the OWASP Top 10 category if applicable. + + + +- Do NOT flag style or performance issues. +- Do NOT include praise or commentary on code quality. +- Do NOT include a finding unless you can state a concrete fix. + + + +For each finding, output: +- File and line range +- OWASP category (or "Other" if none applies) +- Description: what the vulnerability is and why it matters +- Fix: the specific change to make + +If no security issues are found, output: "No security vulnerabilities found in this diff." + + +[PASTE DIFF HERE] +``` + +--- + +## Step 4 — Element annotation + +- **Role:** Named expertise area + specific knowledge domain (OWASP Top 10). Not "a helpful assistant." +- **Context:** States the reader's background so the model calibrates explanation depth. +- **Task:** "Security vulnerabilities only" scopes the task. "OWASP Top 10 category" gives a classification target. +- **Format:** Each finding's structure is explicit; no guessing required. The "no findings" fallback prevents empty output. +- **Constraints:** Three negative rules covering the most common Overachiever failures for code review prompts (style comments, praise, vague warnings). +- **Quality standard:** Embedded in output_format ("concrete fix") and constraints ("a developer should be able to implement without follow-up"). + +--- + +## Step 5 — Watch for + +Most likely failure: the model includes style or complexity comments alongside security findings. If that happens, the constraint list needs to be more specific: name the specific non-security categories to exclude. + +--- + +## Edge case + +**If the diff is too long:** the prompt needs a `` block specifying which files to prioritise. Add: `Focus on authentication, authorization, and data handling code first. If the diff exceeds 500 lines, review in sections and note any sections skipped.` diff --git a/.agents/skills/skill-creation/references/activation-language.md b/.agents/skills/skill-creation/references/activation-language.md new file mode 100644 index 0000000..9ffa2a5 --- /dev/null +++ b/.agents/skills/skill-creation/references/activation-language.md @@ -0,0 +1,63 @@ +# Activation Language Reference + +Source: `10_knowledge/software-practice/2026-06-11__software-practice__raw__instruction-design.md` + +--- + +## What activation language is + +The YAML `description` field in SKILL.md is the only surface the agent reads when deciding whether to +load a skill. It is not marketing copy. It is a routing table. + +The field must do three jobs simultaneously: +1. Fire when the user's words match the skill's domain. +2. Stay silent when the user's words almost-but-don't-quite match. +3. Communicate scope to both the agent and any human who reads the file. + +--- + +## Template + +```yaml +--- +name: your-skill-name +description: > + Use when [primary use case summary in one phrase]. Triggers on: "[phrase 1]", "[phrase 2]", + "[phrase 3]", "[phrase 4]", "[phrase 5]", "[phrase 6]", "[phrase 7]". + Do NOT use for: [adjacent thing that could false-positive 1], [adjacent thing 2], [adjacent thing 3]. +--- +``` + +--- + +## Rules + +### Volume +- List **5–7 explicit trigger phrases**. Fewer than 5 risks Silent failure. More than 7 risks Hijacker failure. +- Include the slightly-wrong phrasings users actually type ("make a skill" alongside "create a skill"). +- Include synonyms from adjacent domains ("instruction doc" alongside "SKILL.md"). + +### Negative boundaries +- Every skill needs at least **one explicit "Do NOT use for" clause**. +- Name the closest adjacent skill or task that could confuse the router. +- Example: a `prompt-creation` skill should exclude "SKILL.md files" if a separate `skill-creation` skill handles those. + +### Grammar / person +- Use **third-person, system-property phrasing**: "Generates proposals" not "I can help you with proposals." +- Use **imperative triggers** in the trigger list: "create a skill", "write a skill" (the form the user types). +- Keep the `description` as a single YAML block scalar (use `>` for multi-line). + +### Scope signal +- The description should tell a competent reader what the skill does and doesn't do in under 60 words. +- If it takes more than 60 words to define the scope, the skill probably covers too much — consider splitting. + +--- + +## Diagnosis + +| Symptom | Likely cause | Fix | +|---------|-------------|-----| +| Skill never loads | Trigger phrases too narrow or absent | Add phrases, especially how users actually phrase requests | +| Skill loads on unrelated requests | No negative boundary, or too-broad phrase | Add "Do NOT use for" clause; narrow the generic phrase | +| Two skills compete on same request | Overlapping trigger sets | Deduplicate phrases; make negative boundaries mirror each other | +| Skill loads correctly but wrong skill wins | Agent uses first-match or highest-score | Move the more specific phrase first; add a tiebreaker negative boundary | diff --git a/.agents/skills/skill-creation/references/anatomy.md b/.agents/skills/skill-creation/references/anatomy.md new file mode 100644 index 0000000..e303c7b --- /dev/null +++ b/.agents/skills/skill-creation/references/anatomy.md @@ -0,0 +1,100 @@ +# Five Components, Five Failure Modes, and the Pre-Ship Checklist + +**Authoritative synthesis:** `10_knowledge/agents/prompts/2026-06-21__agents__note__agent-instruction-design-principles.md` +Source claims below are attributed in that note with confidence labels, counterevidence, and operational examples from the Tripwire harness eval. Read the synthesis note before modifying any claim here. + +--- + +## The five components every instruction document needs + +| # | Component | In a SKILL.md | Why it matters | +|---|-----------|---------------|----------------| +| 1 | **Trigger / activation** | YAML `description` with 5–7 explicit trigger phrases + "Do NOT use for…" | Agent won't activate without the right phrases; too-broad phrases hijack unrelated tasks | +| 2 | **Overview** | One paragraph written for the agent, not the human | Sets frame and vocabulary before any rules are read | +| 3 | **Step-by-step workflow** | Numbered imperative steps | Removes ambiguity; "handle appropriately" is a Drifter waiting to happen | +| 4 | **Output format** | Exact structure, length, tone, forbidden phrases | Without it, the agent guesses — and guesses vary | +| 5 | **Examples + edge cases** | ≥ 1 happy path + ≥ 1 edge-case rule | The most commonly skipped component; the most common cause of fragile behavior | + +Skip any one → unreliable output. The gap in most shipped docs is **#5**. + +--- + +## The five failure modes + +### Failure 1 — Silent (never activates) +**Symptom:** skill should govern the interaction; agent ignores it. +**Diagnosis:** activation language too weak; doesn't contain the words the user typed. +**Fix:** add more trigger phrases, synonyms, and the slightly-wrong phrasings users actually type. + +### Failure 2 — Hijacker (fires on wrong requests) +**Symptom:** skill activates on unrelated tasks. +**Diagnosis:** activation language too broad, or missing negative boundaries. +**Fix:** add explicit "Do NOT use for [X, Y, Z]" clauses. Tighten generic phrases. + +### Failure 3 — Drifter (right doc, inconsistent output) +**Symptom:** skill activates correctly but output varies run-to-run. +**Diagnosis:** vague, non-testable instructions. "Handle appropriately." "Format nicely." +**Fix:** replace every vague rule with a specific, testable one. Leave zero room for interpretation. + +### Failure 4 — Fragile (works on clean input, breaks on edge cases) +**Symptom:** normal inputs succeed; unusual inputs collapse silently. +**Diagnosis:** edge cases not enumerated. +**Fix:** feed the doc the worst inputs imaginable. For each failure, add: `If [condition], then [specific action].` + +### Failure 5 — Overachiever (adds things not asked for) +**Symptom:** output carries unsolicited commentary, extra sections, creative additions. +**Diagnosis:** doc says what TO do, but not what NOT to do. +**Fix:** add explicit scope constraints. "Output ONLY the specified format. Do NOT add [list of forbidden extras]." + +--- + +## Three activation-language rules + +1. **Be pushy.** List 5–7 explicit trigger phrases. Include synonyms. Include the slightly-wrong phrasings users actually type. +2. **Include negative boundaries.** "Do NOT use for [similar-but-different thing]." Prevents hijacking. +3. **Write in third person.** "Generates proposals" beats "I can help you with proposals." Agent instruction parsers handle third-person system-property phrasing more reliably. + +--- + +## The five-test protocol + +Run every instruction doc through these before shipping: + +| Test | Input | Pass condition | +|------|-------|----------------| +| **Happy path** | Clean input, complete context | Expected output, no extra content | +| **Minimal input** | Absolute least information a user might provide | Asks for what it needs, doesn't invent | +| **Edge case** | Unusual, contradictory, typo-laden input | Handles explicitly, not silently mis-handles | +| **Negative test** | A request that should NOT trigger the skill | Agent correctly routes elsewhere | +| **Repeat test** | Same input, three runs | Output consistent across all three | + +Inconsistency on the repeat test = ambiguous instructions. Find the vague phrase and replace it. + +--- + +## Pre-ship checklist + +- [ ] Activation language lists 5–7 explicit trigger phrases +- [ ] Activation language includes at least one negative boundary ("Do NOT use for X") +- [ ] Activation language is third-person / system-property phrasing +- [ ] Overview paragraph speaks to the agent, not the human reader +- [ ] Workflow steps are numbered, imperative, testable +- [ ] Every vague phrase has been replaced ("handle appropriately" → specific rule) +- [ ] Output format is explicit — structure, length, tone, forbidden phrases +- [ ] At least one happy-path example exists (or is referenced from a `references/` or `examples/` file) +- [ ] At least one edge case is enumerated ("If X, then Y") +- [ ] A "what this skill does NOT do" section exists with explicit scope exclusions +- [ ] The doc passed the five-test protocol (happy, minimal, edge, negative, repeat) + +--- + +## Thin-router pattern + +A SKILL.md should contain only: +1. YAML trigger block (5–7 explicit phrases + negative boundaries). +2. One-paragraph overview. +3. A one-line pointer: "Follow the workflow in ``." + +When to use it: if the workflow body exceeds ~150 lines, move it to a `references/` file and point to it from SKILL.md. + +**Why:** Non-Claude agents don't read `.claude/skills/`. If the workflow body lives only in the skill, portability breaks. Three places to drift (skill body + prompt body + AGENTS.md routing) = eventual contradiction. diff --git a/.claude/launch.json b/.claude/launch.json deleted file mode 100644 index 5e965f0..0000000 --- a/.claude/launch.json +++ /dev/null @@ -1,33 +0,0 @@ -{ - "version": "0.0.1", - "configurations": [ - { - "name": "workstation", - "runtimeExecutable": "node", - "runtimeArgs": ["workstation/server.mjs"], - "port": 5177, - "autoPort": true - }, - { - "name": "workstation-live", - "runtimeExecutable": "zsh", - "runtimeArgs": ["-lc", "WORKSTATION_LIVE_DISPATCH=1 exec node workstation/server.mjs"], - "port": 5177, - "autoPort": true - }, - { - "name": "local-agent", - "runtimeExecutable": "node", - "runtimeArgs": ["30_projects/local-agent/workbench/server.mjs"], - "port": 5178, - "autoPort": true - }, - { - "name": "biotech-demo", - "runtimeExecutable": "python3", - "runtimeArgs": ["-m", "http.server", "8088", "--directory", "30_projects/biotech-rag-assistant/workbench/docs"], - "port": 8088, - "autoPort": true - } - ] -} diff --git a/.claude/settings.json b/.claude/settings.json deleted file mode 100644 index c48cde6..0000000 --- a/.claude/settings.json +++ /dev/null @@ -1,215 +0,0 @@ -{ - "hooks": { - "SessionStart": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "UserPromptSubmit": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "PreToolUse": [ - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - }, - { - "matcher": "Write|Edit", - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/knowledge-write-guard", - "timeout": 5 - } - ] - } - ], - "PostToolUse": [ - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "PostToolUseFailure": [ - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - }, - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/papercut auto", - "timeout": 5 - } - ] - } - ], - "PostToolBatch": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "Notification": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "Stop": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "SubagentStop": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "SessionEnd": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - }, - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/session-close --check --feed --hook-stdin", - "timeout": 5 - } - ] - } - ], - "SubagentStart": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "PermissionRequest": [ - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "PermissionDenied": [ - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "StopFailure": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ], - "PreCompact": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - }, - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/session-close --checkpoint --hook-stdin", - "timeout": 5 - } - ] - } - ], - "PostCompact": [ - { - "hooks": [ - { - "type": "command", - "command": "${CLAUDE_PROJECT_DIR}/bin/workflow-event --client claude --pixel main-claude-pixel", - "timeout": 5 - } - ] - } - ] - } -} diff --git a/.claude/skills/classify-note.md b/.claude/skills/classify-note.md deleted file mode 100644 index 4bc40bd..0000000 --- a/.claude/skills/classify-note.md +++ /dev/null @@ -1,42 +0,0 @@ ---- -name: classify-note -description: Assign domain, type, and tags to a file based on its content. Propose to user for confirmation before applying. -status: stub ---- - -# classify-note - -## Purpose - -Read a file in `01_ingest/ready/` and propose: -- **`domain`** — one of the existing `10_knowledge//` subdirectories. -- **`type`** — `raw` (unprocessed evidence) or `note` (synthesized content). -- **`tags`** — free-form, lowercase, kebab-case topic markers. - -## Status - -Implemented (2026-06-13) as a focused sub-skill of ingest-source. Use it when the parent skill (or ingest-agent in batch mode) needs a clean classification proposal. - -## Procedure (as callable sub-skill) -1. Read the full target file (and its current frontmatter + minion-extracted `links:`). -2. Read `10_knowledge/index.md` and sample 3-5 recent/representative notes from each plausible target domain for calibration. -3. Propose exactly one `domain` (existing preferred; propose new only with strong distinctness + recurrence justification + explicit user confirmation later). -4. Propose `type`: default to `raw` for captures/clippings/PDF-derived; use `note` for the operator's own synthesized thinking. -5. Propose 3-8 `tags` (kebab-case, lowercase). Always include provenance/retrieval signals (`x-capture`, `repo-clip`, `pdf`, `llm-output`, etc.) and the audit signal `needs-audit` (or `needs-verification`) when the item will be routed under Tier A/batch rules or is raw evidence. -6. Return a structured proposal object (domain, type, tags list, short rationale, any warnings). Do not mutate the file. - -## Rules -- Never invent a domain from thin air for a single file. New top-level domains are Tier C (user confirmation required; see 10_knowledge/index.md seed-domain rule and ADR-011/020). -- Tags are for retrieval and routing policy, not exhaustive keywords. -- Surface "routing-exception" signals clearly when fit is weak. -- Respect sensitivity overrides from `.context/routing-policy.md` (finance state etc. → Tier C, propose 20_live or park). - -## Related -- [ingest-source](ingest-source.md) — calls this (or equivalent logic) as step 3. -- [extract-metadata](extract-metadata.md) — sibling for other frontmatter fields. -- `.context/routing-policy.md` and `10_knowledge/index.md` — the source of truth for rules and domain inventory. - -## Related - -- [ingest-source](ingest-source.md) — the parent skill -- [extract-metadata](extract-metadata.md) — fills in the rest of frontmatter diff --git a/.claude/skills/ingest-source.md b/.claude/skills/ingest-source.md deleted file mode 100644 index 94bcfdd..0000000 --- a/.claude/skills/ingest-source.md +++ /dev/null @@ -1,100 +0,0 @@ ---- -name: ingest-source -description: Walk a single file from 01_ingest/ready/ through the full agent-driven enrichment loop — read, classify (domain/type/tags), find connections, discuss proposal with user, enrich frontmatter + Connections, rename canonically, and hand off to bin/prep-ingest + minion pass 2. This is the primary skill for the ingest-agent subagent on organic (non-batch) captures. -status: implemented ---- - -# ingest-source - -## Purpose - -The canonical, reusable skill for the per-file judgment pass in the two-pass ingest architecture (ADR-009, ADR-011, ADR-019). It encapsulates everything between the deterministic minion pass 1 (which produces `status: skimmed` files in `01_ingest/ready/`) and the handoff to `bin/prep-ingest` (which moves to `queue/`) + minion pass 2 routing. - -The skill is invoked by the [ingest-agent](../../agents/ingest-agent.md) subagent. After this skill completes its proposal + user confirmation + enrichment, the file is ready for the strict deterministic gate. - -## Inputs -- A Markdown file in `01_ingest/ready/` with at minimum the minion-normalized frontmatter (`title`, `domain` possibly empty, `type`, `status: skimmed`, `source`, `tags`, and the `links:` array extracted from body wikilinks). -- Read-only access to `10_knowledge//` (and their indexes) for calibration and connection finding. -- Optional: `bin/mindgraph query` results for graph-augmented signals (when available). - -## Outputs -- Same file (or atomically renamed version) with: - - `status: extracted` - - Fully populated `domain`, `tags`, `source`, `type` - - `links:` array extended with judgment-driven connections - - Optional `## Connections` prose section appended (never mutates original body for `type: raw`) - - Canonical filename `YYYY-MM-DD__domain__type__slug.md` -- The file remains in `01_ingest/ready/` until `bin/prep-ingest run --apply` is called by the caller. -- A clear proposal presented to the user for confirmation before any enrichment writes. - -## Procedure (step-by-step — this is the implementation of the skill) - -1. **List and select** - Enumerate files in `01_ingest/ready/`. Process **one at a time** for organic captures (batch mode uses a different table flow in the ingest-agent). Skip any file already at `status: extracted` (it is awaiting `prep-ingest`). - -2. **Read the full content** of the selected file. - -3. **Propose classification (domain, type, tags)** - - `domain`: Match against the inventory and rules in `10_knowledge/index.md`. Start with existing domains. If the topic is genuinely new, distinct, and likely to recur, **propose a new domain (or subdomain)** with short rationale and **wait for explicit user confirmation** before creating folders (ADR-011 / Tier C). Never force a weak fit. Leave `domain: ""` and stop if uncertain — the downstream gate will reject it anyway. - - `type`: `raw` for unprocessed evidence/clippings/PDF wrappers; `note` for synthesized/user-authored content. - - `tags`: Add useful retrieval tags (lowercase kebab-case). Include required signals such as `needs-audit` (or `needs-verification`) for raw or low-synthesis routed material (per routing-policy and the audit-sweep workflow). Preserve any existing good tags from the minion or prior state. - -4. **Find and propose connections** (ADR-033 dual-channel graph) - - Start with the deterministic `links:` array already populated by the ingest minion (wikilinks extracted from body, code spans stripped). - - Extend with judgment: read nearby notes in the target `10_knowledge//`, use `bin/mindgraph query "key terms"` when operational for nominations, cross-domain pointers. - - Add targets to frontmatter `links:` using **unique canonical trailing slugs** (e.g. `gxp-pharma-source-catalog`) or **full filename stems** when ambiguous. MindGraph indexes `links:` and body wikilinks equally after refresh. - - For `type: raw`, prefer frontmatter `links:` and/or an appended `## Connections` section with `[[wikilinks]]`. For notes, body wikilinks are optional when `links:` is populated. Do not wikilink `30_projects/` paths — use prose until bridge registry exists. - - Proposed connections are nominations only — they do not assert truth. - -5. **Present the full proposal to the user as a single confirm-or-correct surface** (critical judgment gate) - Show: - - Proposed full frontmatter block (`domain`, `type`, `tags` (with `needs-audit` if applicable), `source`, updated `links`). - - Proposed new canonical filename. - - Draft `## Connections` section (if any) with prose explanation of relationships. - - Open questions: new vs. known? Does this affect existing understanding? Split or keep together? Domain justification if proposing new. - - Any warnings from minion pass 1. - - **Wait for explicit user confirmation or specific corrections** before proceeding to step 6. Do not auto-apply enrichment. - -6. **Enrich after confirmation** - - If a new top-level domain was confirmed by the user: create `10_knowledge//`, `10_knowledge//raw/`, and a minimal `index.md` (see 10_knowledge/index.md rules). - - Update the file's frontmatter with the confirmed values and set `status: routed` then `status: extracted`. - - Append `## Connections` section at the bottom when useful (for raw items: frontmatter `links:` is sufficient for the graph; Connections adds human-readable context. Body is immutable evidence. For user notes/drafts: inline wikilinks are optional when `links:` is complete). - - If a raw item would benefit from a synthesized companion, prefer calling the sibling `create-source-summary` skill to produce a separate `note` rather than editing the raw. - -7. **Rename atomically** to the canonical `YYYY-MM-DD__domain__type__slug.md` (use captured date from frontmatter or today). Write the new path first, verify, then remove the old if using separate operations. - -8. **Handoff to deterministic pipeline** - - Run `bin/prep-ingest run --dry-run` (validates strict frontmatter, `status: extracted`, canonical name, known domain whitelist, no collisions). - - On clean result: `bin/prep-ingest run --apply` (moves ready/ → queue/). - - Then `bin/ingest-minion run --apply` (routes queue/ → 10_knowledge//). - - Optionally `bin/mindgraph-refresh`. - - For batch flows the caller coordinates the single review table instead of per-file discussion. - -## Guardrails (enforced by this skill and its caller) -- **Body immutability for raw**: Never rewrite the original content of `type: raw` items. All enrichment lives in frontmatter and appended sections. -- **Read-only on durable knowledge**: Only read `10_knowledge/` to find connections and calibrate proposals. All writes stay inside `01_ingest/`. -- **New domains are human**: Always Tier C. Propose with evidence from `10_knowledge/index.md`; create folders only after user OK. -- **needs-audit tagging**: For material that will be routed under Tier A or batch rules, ensure the tag is present so `bin/audit-sweep` + the epistemic auditor can verify post-placement (see `.context/workflows/audit-sweep.md`). -- **Epistemic stance**: Follow `.context/workflows/epistemic-standard.md` and `EPISTEMIC_STANCE.md`. Label claim types. Assign confidence. Preserve raw sources. MindGraph results are retrieval nominations, not verified relationships. Record provenance. -- **One file at a time for organic**: Batch mode (ADR-019) uses the table flow instead. - -## Error & Edge Handling -- Malformed or missing required frontmatter after minion: surface warning; repair in proposal if recoverable. -- Unknown domain in proposal: force user confirmation or park as exception. -- Collision on rename or prep-ingest: abort and report; do not overwrite. -- User rejects proposal: leave file as-is (or with minimal `routing_note`) and move to next or ask for guidance. - -## Related Skills & Components -- `classify-note` — focused sub-skill for domain/type/tags proposal (can be called internally). -- `extract-metadata` — optional source metadata enrichment (PDF info etc.). -- `rename-material` — the atomic rename step. -- `create-source-summary` — for turning a raw into a synthesized note sibling. -- `agents/ingest-agent.md` — the subagent that orchestrates this skill (and the batch table path). -- `bin/prep-ingest`, `bin/ingest-minion`, `.context/workflows/audit-sweep.md`, `10_knowledge/index.md`, `.context/routing-policy.md`, `01_ingest/AGENTS.md`. - -## Evaluation -This skill should be exercised and measured during process evaluations (see `30_projects/mainframe-process-eval/`). Track: proposal acceptance rate, time-to-extracted, downstream routing success rate, number of needs-audit tags correctly applied, and any Tier B exceptions that later graduate to rules. - -## Status Note -This skill was promoted from stub (2026-06-13) as part of closing the gap between the detailed agent procedure and reusable skill contracts. The long-form procedure now lives here; the ingest-agent definition should remain focused on role, guardrails, batch vs. per-file mode, and orchestration. \ No newline at end of file diff --git a/.claude/skills/source-literature.md b/.claude/skills/source-literature.md deleted file mode 100644 index f0b9631..0000000 --- a/.claude/skills/source-literature.md +++ /dev/null @@ -1,15 +0,0 @@ ---- -name: source-literature -description: Find, vet, deduplicate, and capture peer-reviewed or well-accepted literature into 00_inbox/ as raw evidence stubs for the ingest pipeline. Use when the user needs research sources, literature search, academic references, DOI/PubMed sourcing, or asks to complement ingest with upstream discovery. Triggers on "source literature", "find papers", "peer-reviewed sources", "research this topic". -status: implemented ---- - -# source-literature - -Canonical definition: [.agents/skills/source-literature/SKILL.md](../.agents/skills/source-literature/SKILL.md) - -Credibility tiers: [.agents/skills/source-literature/references/credibility-tiers.md](../.agents/skills/source-literature/references/credibility-tiers.md) - -Operator workflow: [.context/workflows/source-literature.md](../.context/workflows/source-literature.md) - -Subagent: [agents/source-literature-agent.md](../../agents/source-literature-agent.md) \ No newline at end of file diff --git a/.context/doctor/catalogue.json b/.context/doctor/catalogue.json deleted file mode 100644 index 41a6da5..0000000 --- a/.context/doctor/catalogue.json +++ /dev/null @@ -1,586 +0,0 @@ -{ - "schema_version": 1, - "catalogue_version": "0.1.0", - "doctor_version": "0.1.0", - "contract_ref": "30_projects/mainframe-process-eval/plans/scalability/mainframe-doctor-contract.md", - "checks": [ - { - "id": "AUTH-001", - "owner": "mainframe-process-eval", - "subsystem": "authority", - "layer": "operational-health", - "provider": "auth_focus", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 2, - "skip_policy": "fail", - "freshness_seconds": 604800, - "pass_condition": "structured focus authority parses and is within review window", - "remediation": "create or refresh 20_live/focus/current.yaml", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "AUTH-002", - "owner": "mainframe-process-eval", - "subsystem": "authority", - "layer": "temporal/drift", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 2, - "skip_policy": "fail", - "freshness_seconds": 604800, - "pass_condition": "focus projections carry current authority revision", - "remediation": "wire STATE/session consumers to focus revision", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "SESSION-001", - "owner": "mainframe-process-eval", - "subsystem": "session", - "layer": "operational-health", - "provider": "session_project_path", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 2, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "selected project path exists and contract chain resolves", - "remediation": "set STATE Active Project to a real project slug or structured focus", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "SESSION-002", - "owner": "mainframe-process-eval", - "subsystem": "session", - "layer": "contract/temporal", - "provider": "session_phase_alignment", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 2, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "active phase/re-entry agrees with project state", - "remediation": "align README project_state and next_action with focus/STATE selection", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "PROJECT-001", - "owner": "mainframe-process-eval", - "subsystem": "projects", - "layer": "contract", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "every project has valid state metadata and re-entry", - "remediation": "repair README frontmatter", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "PROJECT-002", - "owner": "mainframe-process-eval", - "subsystem": "projects", - "layer": "drift", - "provider": "project_index_check", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 10, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "generated index matches project authorities", - "remediation": "bin/sync-project-index --write after verifying authorities", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "PROJECT-003", - "owner": "mainframe-process-eval", - "subsystem": "projects", - "layer": "integration/operational", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "state, plan, handoff, nested repo do not contradict", - "remediation": "semantic project validation unit", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "PROJECT-004", - "owner": "mainframe-process-eval", - "subsystem": "projects", - "layer": "integration", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "fixture-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "write workflows re-read and recheck final state", - "remediation": "lifecycle transaction primitive", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "SCHED-001", - "owner": "mainframe-process-eval", - "subsystem": "scheduler", - "layer": "operational-health", - "provider": "sched_service", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 3, - "skip_policy": "fail", - "freshness_seconds": 86400, - "pass_condition": "launch service loaded, expected runs observed, last exit zero", - "remediation": "inspect launchd and eval-schedule logs", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "SCHED-002", - "owner": "mainframe-process-eval", - "subsystem": "scheduler", - "layer": "provenance/operational", - "provider": "sched_provenance", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 3, - "skip_policy": "fail", - "freshness_seconds": 86400, - "pass_condition": "latest scheduled execution from service not manual substitute", - "remediation": "separate service vs manual provenance", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "EVAL-001", - "owner": "mainframe-process-eval", - "subsystem": "evaluation", - "layer": "integration", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 604800, - "pass_condition": "final artifact and registry agree", - "remediation": "harvest order and schema alignment", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "EVAL-002", - "owner": "mainframe-process-eval", - "subsystem": "evaluation", - "layer": "contract", - "provider": "eval_registry_strict", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 15, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "typed eval-run artifacts satisfy strict schema", - "remediation": "fix metric extract headings and schema", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "EVAL-003", - "owner": "mainframe-process-eval", - "subsystem": "evaluation", - "layer": "operational-health", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 86400, - "pass_condition": "required scheduled gates ran; skips/reporter failures not green", - "remediation": "doctor SCHED/EVAL providers", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "MG-001", - "owner": "mainframe-process-eval", - "subsystem": "mindgraph", - "layer": "operational-health", - "provider": "mg_db_presence", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "canonical installed DB paths and schemas valid", - "remediation": "bin/mindgraph doctor; refresh if needed", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "MG-002", - "owner": "mainframe-process-eval", - "subsystem": "mindgraph", - "layer": "temporal", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 86400, - "pass_condition": "knowledge index freshness matches intended source scope", - "remediation": "mindgraph-refresh with provenance", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "MG-003", - "owner": "mainframe-process-eval", - "subsystem": "mindgraph", - "layer": "integration/temporal", - "provider": "mg_manifest_coverage", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 86400, - "pass_condition": "project namespaces match intended manifest and coverage policy", - "remediation": "update mindgraph-projects.json, then run the ADR-045 plan -> stage -> promote workflow", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "MG-004", - "owner": "mainframe-process-eval", - "subsystem": "mindgraph", - "layer": "recovery/integration", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "fixture-only", - "timeout_seconds": 10, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "namespace-scoped refresh cannot prune unselected namespaces", - "remediation": "engine namespace isolation tests on temp DB", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "MG-005", - "owner": "mainframe-process-eval", - "subsystem": "mindgraph", - "layer": "outcome/regression", - "provider": "unimplemented", - "required": false, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 30, - "skip_policy": "skip", - "freshness_seconds": 604800, - "pass_condition": "frozen retrieval canaries meet gates", - "remediation": "mindgraph-eval canaries", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "META-001", - "owner": "mainframe-process-eval", - "subsystem": "knowledge", - "layer": "contract/operational", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 10, - "skip_policy": "fail", - "freshness_seconds": 86400, - "pass_condition": "corpus metadata validity and quarantine within policy", - "remediation": "metadata audit", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "INGEST-001", - "owner": "mainframe-process-eval", - "subsystem": "ingest", - "layer": "integration", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "location and lifecycle status agree", - "remediation": "ingest-status / prep-ingest", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "AUDIT-001", - "owner": "mainframe-process-eval", - "subsystem": "audit", - "layer": "operational-health", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 86400, - "pass_condition": "audit worker/handoff capability is real and current", - "remediation": "epistemic audit path", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "AUDIT-002", - "owner": "mainframe-process-eval", - "subsystem": "audit", - "layer": "operational-health", - "provider": "unimplemented", - "required": false, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "skip", - "freshness_seconds": 86400, - "pass_condition": "queue age/class/throughput visible with thresholds", - "remediation": "threshold adapter on audit reporter", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "TEL-001", - "owner": "mainframe-process-eval", - "subsystem": "telemetry", - "layer": "operational-health", - "provider": "tel_hash_integrity", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 10, - "skip_policy": "fail", - "freshness_seconds": 86400, - "pass_condition": "event store parses and hash integrity holds", - "remediation": "workflow-event hash-chain repair on newest day file", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "TEL-002", - "owner": "mainframe-process-eval", - "subsystem": "telemetry", - "layer": "live smoke/temporal", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 604800, - "pass_condition": "observed client capabilities match declared adapters", - "remediation": "capability vs event coverage", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "TEL-003", - "owner": "mainframe-process-eval", - "subsystem": "telemetry", - "layer": "integration", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "client/session/task/run/artifact/verifier linkage valid", - "remediation": "identity chain", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "WS-001", - "owner": "mainframe-process-eval", - "subsystem": "workstation", - "layer": "component", - "provider": "ws_db_presence", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 3, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "workstation storage preflight succeeds", - "remediation": "create or repair 20_live/workstation/workstation.sqlite", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "WS-002", - "owner": "mainframe-process-eval", - "subsystem": "workstation", - "layer": "operational/drift", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 86400, - "pass_condition": "projection source revisions and freshness current", - "remediation": "projection refresh with authority revision", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "WS-003", - "owner": "mainframe-process-eval", - "subsystem": "workstation", - "layer": "contract/integration", - "provider": "ws_operational_rows", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "fixture/demo records not presented as sole live authority; operational rows exist or emptiness is explicit", - "remediation": "populate runs/approvals or label fixture mode", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "STRUCT-001", - "owner": "mainframe-process-eval", - "subsystem": "structure", - "layer": "static", - "provider": "structure_bounds", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 10, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "structural links and private/public boundaries valid", - "remediation": "restore root/lifecycle contracts and private boundary docs", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "STRUCT-002", - "owner": "mainframe-process-eval", - "subsystem": "structure", - "layer": "drift", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 5, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "generated docs/configs match canonical registries", - "remediation": "regenerate from registries", - "safe_fix_available": false, - "contract_version": 1 - }, - { - "id": "CLI-001", - "owner": "mainframe-process-eval", - "subsystem": "cli", - "layer": "contract", - "provider": "cli_mutation_help_safety", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 2, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "mutation-capable CLIs have safe help/dry-run and reject unknown input", - "remediation": "fix bin/mindgraph-refresh-projects argument parser", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "PATH-001", - "owner": "mainframe-process-eval", - "subsystem": "security", - "layer": "security/contract", - "provider": "unimplemented", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 2, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "slugs and resolved paths remain inside approved roots", - "remediation": "path confinement helpers", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "SEC-001", - "owner": "mainframe-process-eval", - "subsystem": "security", - "layer": "operational-health", - "provider": "sec_secret_store", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 2, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "secret references and private stores use approved locations and modes", - "remediation": "chmod 0600; move secrets to Keychain/env", - "safe_fix_available": true, - "contract_version": 1 - }, - { - "id": "TASK-001", - "owner": "mainframe-process-eval", - "subsystem": "tasks", - "layer": "contract", - "provider": "task_manifest_quarantine", - "required": true, - "mutates": false, - "isolation": "live-read-only", - "timeout_seconds": 2, - "skip_policy": "fail", - "freshness_seconds": 0, - "pass_condition": "tasks_manifest quarantine envelope present; no executable rows", - "remediation": "run bin/generate-project-tasks or Unit 2.3 wrap", - "safe_fix_available": true, - "contract_version": 1 - } - ] -} diff --git a/.context/doctor/required-invariants.json b/.context/doctor/required-invariants.json deleted file mode 100644 index 5942bfb..0000000 --- a/.context/doctor/required-invariants.json +++ /dev/null @@ -1,42 +0,0 @@ -{ - "schema_version": 1, - "manifest_version": "2026-07-15.1", - "source": "plans/scalability/mainframe-doctor-contract.md initial check registry + July 10 audit P0", - "description": "Separate required-invariant manifest. Doctor catalogue completeness is judged against this file, never against the catalogue alone.", - "required_check_ids": [ - "AUTH-001", - "AUTH-002", - "SESSION-001", - "SESSION-002", - "PROJECT-001", - "PROJECT-002", - "PROJECT-003", - "PROJECT-004", - "SCHED-001", - "SCHED-002", - "EVAL-001", - "EVAL-002", - "EVAL-003", - "MG-001", - "MG-002", - "MG-003", - "MG-004", - "MG-005", - "META-001", - "INGEST-001", - "AUDIT-001", - "AUDIT-002", - "TEL-001", - "TEL-002", - "TEL-003", - "WS-001", - "WS-002", - "WS-003", - "STRUCT-001", - "STRUCT-002", - "CLI-001", - "PATH-001", - "SEC-001", - "TASK-001" - ] -} diff --git a/.context/harness-routing.json b/.context/harness-routing.json deleted file mode 100644 index a7a46cf..0000000 --- a/.context/harness-routing.json +++ /dev/null @@ -1,69 +0,0 @@ -{ - "version": 1, - "description": "Deterministic routing hints for task packets. Graduation still requires agent-harness-eval receipts per category × profile × harness.", - "category_defaults": { - "mechanical-edit": { - "executor": "local", - "harness_recommendation": "H1-packet", - "agent_profile": "local-qwen25-coder-14b", - "allow_fusion_plan": false - }, - "multi-file-coordination": { - "executor": "local", - "harness_recommendation": "H2-repair", - "agent_profile": "local-qwen25-coder-14b", - "allow_fusion_plan": false - }, - "failure-recovery": { - "executor": "local", - "harness_recommendation": "H2-repair", - "agent_profile": "local-qwen25-coder-14b", - "allow_fusion_plan": false - }, - "scope-enforcement": { - "executor": "local", - "harness_recommendation": "H1-packet", - "agent_profile": "local-qwen25-coder-14b", - "allow_fusion_plan": false - }, - "research-synthesis": { - "executor": "cloud", - "harness_recommendation": "H1-packet", - "allow_fusion_plan": true, - "needs_deliberation": true - }, - "environment-setup": { - "executor": "auto", - "harness_recommendation": "H1-packet", - "allow_fusion_plan": false - }, - "long-horizon-build": { - "executor": "cloud", - "harness_recommendation": "H1-packet", - "allow_fusion_plan": true - } - }, - "inference_rules": [ - { - "when": { - "task_kind": "research", - "min_total_files": 0 - }, - "task_category": "research-synthesis" - }, - { - "when": { - "task_kind": "code", - "min_editable_files": 2 - }, - "task_category": "multi-file-coordination" - }, - { - "when": { - "task_kind": "code", - "max_editable_files": 1 - }, - "task_category": "mechanical-edit" - } - ] -} \ No newline at end of file diff --git a/.context/live-retention.md b/.context/live-retention.md deleted file mode 100644 index de774c2..0000000 --- a/.context/live-retention.md +++ /dev/null @@ -1,78 +0,0 @@ ---- -title: "Live surface retention and index hygiene" -domain: "knowledge-systems" -type: "note" -status: "active" -updated: "2026-07-15" -source: "ADR-045; process-eval program after Phase 2" -tags: ["20_live", "retention", "mindgraph", "hygiene"] ---- - -# Live surface retention and index hygiene - -Policy for volatile MainFrame state. Complements `20_live/AGENTS.md` (no silent -overwrite) and ADR-045. **Does not** authorize bulk deletion of history. - -## Goals - -1. Keep **MindGraph projects index** useful: coordination Markdown only, full - intended coverage, rebuildable, no telemetry firehose. -2. Keep **`20_live/`** honest: different classes age differently; derived - projections may be wiped; append-only evidence is capped/archived, not - silently rewritten. -3. Never treat “clean” as “delete until green.” - -## Retention classes - -| Class | Examples | Default rule | MindGraph? | -|-------|----------|--------------|------------| -| **A — Authority** | `20_live/focus/current.yaml`, SEC dispositions, operator cards | Keep; revise with review; cite revision | No (focus is session/doctor authority, not corpus) | -| **B — Append-only evidence** | `eval-registry/*.jsonl`, focus `decisions.jsonl` / `outcomes.jsonl`, workflow-metrics event days | Append only; **archive or compress** after soft age (see caps); never rewrite lines | No | -| **C — Derived projection** | `workstation.sqlite`, session-close feeds, staged MindGraph DBs under `~/.mindgraph/staging/` | Rebuildable; wipe/replace OK when receipt exists | No (DB is the projection) | -| **D — Reporter / schedule noise** | Many `*-scheduled-weekly.md` under process-eval, large trend dumps | Keep last **N** hot; older stay on disk or move to project archive; optional exclude from index globs later | Optional: only if under `outputs/**` and still useful for status | -| **E — High-volume ops** | `workflow-metrics/events/`, markets DBs | Size/age caps; cold days off hot path; not knowledge | **Never** into projects or knowledge index | - -## Soft caps (operator defaults — not auto-enforced yet) - -| Surface | Soft cap | Action when exceeded | -|---------|----------|----------------------| -| Eval schedule jsonl | 24 months hot | Archive older segments to `90_archive/` or compress beside file | -| Workflow event day files | 90 days hot detail | Compress or move cold days; keep chain receipts if any | -| Workstation projection DB | rebuild anytime | Prefer refuse bad seed (Unit 2.3/2.4) over partial edit | -| Process-eval scheduled outputs | last 12 weeklies hot | Older remain files but need not drive attention | -| MindGraph staging DBs | 3 newest under `~/.mindgraph/staging/` | Delete older stage DBs after promote or abandon | - -Automation of caps is a later unit. This file is the **policy authority**. - -## Projects MindGraph (separate from 20_live cleanup) - -| Rule | Detail | -|------|--------| -| Default (lean) | `30_projects/mindgraph-projects.json` — README, log, decisions, methodology, plans; **excludes outputs/** | -| Deep | `30_projects/mindgraph-projects-deep.json` or `--deep` — adds `outputs/**` for archaeology | -| Out of scope | `20_live/**` telemetry, workbenches, raw-materials, secrets | -| Mutation | Explicit staged apply: **plan → stage (temp DB) → promote** | -| Command | `bin/mindgraph-projects-apply` [ `--deep` ] | -| Success | Manifest namespaces ⊆ staged DB; no missing intended project dirs; receipt written | -| Re-stage | See HARNESS.md “When to re-stage the projects index” | - -Installed path remains `~/.mindgraph/mainframe-projects.sqlite`. Promote always -backs up the previous file first. Prefer lean for daily agent context; deep only -when hunting old eval receipts or FINDINGS. - -## What “manage staleness” means here - -| Stale kind | Response | -|------------|----------| -| Project README/plan changed | Re-stage/promote projects index (or accept lag until next apply) | -| Focus review_by past | Doctor AUTH stale/warn; rewrite focus authority | -| Eval schedule without launchd provenance | SCHED degraded (Unit 2.4) | -| Telemetry volume | Cap class E; do not index | -| Workstation shows old seeded tasks | Rebuild class C projection; do not “fix” by editing jsonl | - -## Related - -- ADR-045 -- `HARNESS.md` dual MindGraph trust profiles -- `bin/mindgraph-projects-apply` -- `20_live/AGENTS.md` diff --git a/.context/routing-policy.md b/.context/routing-policy.md index 3d6ac0a..e5c1950 100644 --- a/.context/routing-policy.md +++ b/.context/routing-policy.md @@ -42,7 +42,7 @@ Every parked file gets a `parked_reason:` frontmatter line stating why, in plain | Rule | Pattern | Disposition | | --- | --- | --- | -| P1 | Empty, untitled, or fragment notes (`Untitled*`, single-thought stubs with no source) | `parked` → `90_archive/second-brain-migration/parked/` | +| P1 | Empty, untitled, or fragment notes (`Untitled*`, single-thought stubs with no source) | `parked` → `local-only: 90_archive//parked/` | | P2 | Code, scripts, configs (`.py`, `.sh`, `.bat`, `.ps1`, `.json`, `.jsx`) with no note context | `parked` (adopt into an owning project workbench manually if wanted) | | P3 | Exact duplicate by batch-manifest or body hash | `duplicate-removed`; delete the redundant working copy — the canonical copy plus the batch `source-files/` snapshot preserve the bytes (amended 2026-06-10 per operator instruction, ADR-020) | | P4 | Binary without a convention name (images, `.docx`, `.rtf`, `.xlsx`) | `parked`; convention-named PDFs go through the minion wrapper instead | @@ -53,7 +53,7 @@ Every parked file gets a `parked_reason:` frontmatter line stating why, in plain | Rule | Pattern | Handling | | --- | --- | --- | -| B1 | Old-system operational docs: role prompts, project instructions, system audits and evals, automation guides, product references from pre-MainFrame systems | Adopt into `30_projects/second-brain-migration/raw-materials/legacy-systems//` with `status: archived` and tag `legacy-system-doc`. Everything stays in the migration project; promotion into active MainFrame projects is a later operator decision (ADR-020). | +| B1 | Old-system operational docs: role prompts, project instructions, system audits and evals, automation guides, product references from pre-MainFrame systems | Adopt into `local-only: 30_projects//raw-materials/legacy-systems//` with `status: archived` and tag `legacy-system-doc`. Everything stays in the migration project; promotion into active MainFrame projects is a later operator decision (ADR-020). | ## Routing rules (first match wins) diff --git a/.context/templates/contracts/README.md b/.context/templates/contracts/README.md new file mode 100644 index 0000000..2f21b1f --- /dev/null +++ b/.context/templates/contracts/README.md @@ -0,0 +1,134 @@ +--- +title: "Contract templates" +domain: "knowledge-systems" +type: "template" +status: "active" +source: "bin/contract-audit checks 1.5/1.6/1.7/10.5; 00_inbox/AGENTS.md and 10_knowledge/AGENTS.md as reference implementations" +tags: ["contracts", "templates", "enforcement", "tiers"] +updated: "2026-08-23" +--- + +# Contract templates + +Starter files for anything that constrains agent behaviour. They exist so a new +contract is **born enforceable** instead of being retrofitted after an incident. + +`structural-file-profile.md` (one level up) answers *what kind of file is this +and where does it belong*. These templates answer *how does a rule in it get +enforced*. Use both: the profile to place the file, a template to write it. + +## Why these exist + +Measured 2026-08-23 by `bin/contract-audit --layer document`: + +| Check | Result | +| --- | --- | +| 10.5 contracts declare a tier | **2 / 135 files** | +| 1.6 hard promises bound at clause level | 138 clauses, **124 unbound (10% bound)** | +| 1.5 normative clauses name a mechanism | 680 clauses, **39 bound (6%)** | +| 1.7 T1+ rules name an Escape | 7 rules, **0 missing** | + +1.7 is the tell. **Where the format is used it works.** The gap is propagation, +not design — so the fix is a template plus a checker, not a better argument. + +Behavioural corroboration, independent of the document layer: `AGENTS.md` says +"Avoid Subshell `cd`" and telemetry shows **5,215 `cd` command heads**, with +`cd` the **#1 repeated auto-papercut**. A rule nothing checks is a wish. + +## The clause format + +Every normative rule declares four things, as a blockquote directly under an +`##` heading: + +```markdown +## 1. The rule, stated as one sentence in the imperative. + +> **Binds:** who or what this constrains +> **Tier:** T2 (blocked) +> **Check:** `bin/some-tool --check`, a named hook, or `tests/test_x.py` +> **Escape:** the named, cheap, non-penalized way to comply when you cannot +> meet the letter of the rule + +Optional prose: the measurement or incident that made this rule necessary. +``` + +**Syntax is load-bearing.** `bin/contract-audit` parses these with real regexes +and a malformed block reads as zero adoption: + +- The block must **start** with `**Binds:**`. +- **No blank line inside the block** — a blank line ends it. Wrap continuation + lines with a leading `>`. +- `Tier:` must be literally `T0`, `T1`, `T2`, or `T3`. +- `Check:` should name something runnable: a `bin/` tool, a hook event, a + `pytest` path, `.gitignore`, or a `--check` / `--apply` / `--dry-run` flag. + Those are the tokens the audit's `MECHANISM` regex recognises. + +Verify before committing: `bin/contract-lint --file `. + +## The tier ladder + +| Tier | Means | What must be true | +| --- | --- | --- | +| **T0** advisory | Judgement. No mechanism, and none is promised. | Nothing. `Check: none` is honest and correct here. | +| **T1** detected | A violation is found **after** the fact. | A named tool or sweep reports it. | +| **T2** blocked | A violation cannot land. | A hook, gate, or guard refuses it. | +| **T3** reconciled | Drift is detected **and** repaired against a source of truth. | A reconciler runs on a schedule. | + +**T0 is a legitimate answer, and picking it honestly is the point.** "Do not +tidy this folder" carries `Check: none` in `00_inbox/AGENTS.md` and that is +correct. The defect this format fixes is not the absence of mechanisms — it is +that **nothing distinguished an advisory rule from a load-bearing one**, so a +rule that silently lost its enforcement looked exactly like one that never +needed any. + +Do not inflate tiers. A declared tier is self-asserted; claiming T2 without a +gate is the same worthless self-report as a capture writing its own +`retrieved_at`. + +## Escape is mandatory at T1 and above + +> A rule with no compliant way to fail **manufactures** violations. An agent +> that cannot comply and cannot honestly fail will produce something that +> *resembles* compliance. + +That is not theory. A source-count quota made honesty unrepresentable — no field +meant "I looked and found nothing" — and 107 fabricated citations followed. See +`10_knowledge/agents/2026-08-10__agents__note__every-rule-needs-an-honest-failure-path.md`. + +Three exits must be namable for every T1+ rule: + +1. **The agent's** — named, cheap, and non-penalized. An escape that reads as + failure is not an escape. +2. **The check's** — fail-open or fail-closed, decided and written down. +3. **The reader's** — "ran, found nothing" must be distinguishable from "did not + run". + +Diagnostic for every MUST you write: *what does an agent do when it cannot?* + +## Files here + +| Template | For | +| --- | --- | +| `agents-md.template.md` | Any `AGENTS.md` — root, lifecycle folder, project, or workbench | +| `workflow.template.md` | `.context/workflows/*.md` operator sequences | +| `skill.template.md` | `.agents/skills/**/SKILL.md` reusable agent judgement | +| `subagent.template.md` | `agents/*.md` specialized roles | +| `decision-record.template.md` | `DECISIONS.md` / project `decisions.md` entries | + +## Reference implementations + +Read these before writing a new contract — they are the format working on real +rules, not illustrations: + +- `00_inbox/AGENTS.md` — four rules spanning T0 through T2 +- `10_knowledge/AGENTS.md` — the post-incident rules, all with earned Escapes + +## Adoption is measured, not assumed + +- `bin/contract-lint --file ` — one file, before you commit it. +- `bin/contract-lint --all` — corpus adoption and the worst offenders. +- `bin/contract-audit --layer document` — checks 1.5 / 1.6 / 1.7 / 10.5 as part + of the full audit program. + +A template with no consumption path is this system's documented failure mode. If +these files stop being linted, they have already failed. diff --git a/.context/templates/contracts/agents-md.template.md b/.context/templates/contracts/agents-md.template.md new file mode 100644 index 0000000..9050a57 --- /dev/null +++ b/.context/templates/contracts/agents-md.template.md @@ -0,0 +1,71 @@ +# - Local Rules + +> [!WARNING] +> One or two lines on what this scope *is*, and the consequence of getting it +> wrong. Readers who skim one line should learn the blast radius from it. +> Delete this callout if the scope carries no standing hazard. + +Rules declare **Binds / Tier / Check / Escape**. The Escape is the named, cheap, +non-penalized way to comply when you cannot meet the letter of the rule. + +Tiers: **T0** advisory · **T1** detected · **T2** blocked · **T3** reconciled. + +--- + +## 1. + +> **Binds:** +> **Tier:** T2 (blocked) +> **Check:** `bin/` (), plus +> `tests/test_.py` as the regression +> **Escape:** it carries no penalty.> + + + +## 2. + +> **Binds:** +> **Tier:** T1 (detected) +> **Check:** `bin/` — reports violations after the fact +> **Escape:** + +## 3. + +> **Binds:** +> **Tier:** T0 (advisory) +> **Check:** none +> **Escape:** n/a + + + + diff --git a/.context/templates/contracts/decision-record.template.md b/.context/templates/contracts/decision-record.template.md new file mode 100644 index 0000000..7fe48ce --- /dev/null +++ b/.context/templates/contracts/decision-record.template.md @@ -0,0 +1,46 @@ +## ADR-: (<YYYY-MM-DD>) + +**Status**: Proposed | Accepted | Superseded by ADR-<NNN> +**Date**: <YYYY-MM-DD> + +**Context**: <What forced the decision. Measured state where it exists — counts, +dates, the failing command. A context section with no evidence produces a +decision nobody can later falsify.> + +**Decision**: <What was decided, numbered when it has parts. Each part should be +checkable against the repo six months from now.> + +**Rationale**: <Why this over the alternatives that were actually considered. +Name the rejected option; "we chose X" with no discarded Y is a description, not +a rationale.> + +**Consequences**: <What becomes true, what becomes forbidden, and what is +explicitly NOT authorized by this ADR. The last clause matters most — it is what +stops an accepted design from being read as an accepted implementation.> + +**Enforcement**: + +> **Binds:** <who or what this decision constrains in practice> +> **Tier:** T1 (detected) +> **Check:** `bin/<tool>` / `tests/test_<name>.py` / <audit check id> +> **Escape:** <how to honestly not comply — usually "raise a superseding ADR"> + +<!-- +TEMPLATE NOTES — delete before saving. + +DECISIONS.md is newest-first. Numbers are never reused and gaps stay unfilled so +external references keep matching (see the numbering note at the file's end). + +The Enforcement block is what separates a decision from an intention. On +2026-08-23 DECISIONS.md carried 60 unbound normative clauses, the largest single +block in the corpus — accepted decisions that nothing checks. + +If the decision is genuinely a direction rather than a rule, say so: + > **Tier:** T0 (advisory) + > **Check:** none +That is honest. An inflated tier is not. + +Project-only tradeoffs go in the project's decisions.md, not here. + +Verify with: bin/contract-lint --file <this file> +--> diff --git a/.context/templates/contracts/skill.template.md b/.context/templates/contracts/skill.template.md new file mode 100644 index 0000000..5bb43e1 --- /dev/null +++ b/.context/templates/contracts/skill.template.md @@ -0,0 +1,66 @@ +--- +name: <kebab-case-name> +description: "Use when <task keywords front-loaded>. Covers <X, Y, Z>. Not for <negative boundary>." +--- + +# <Skill name> + +<One paragraph: the judgement this skill encodes, and why an agent would get it +wrong without the skill. If the answer is "it would just be slower", this is a +workflow, not a skill.> + +## When this applies + +- <trigger phrasing an agent will actually encounter> +- <...> + +## When it does not + +- <the nearest neighbouring skill, and why this is not it> +- <...> + +## Procedure + +1. <Imperative step.> +2. <...> + +## Rules + +### 1. <Rule as one imperative sentence.> + +> **Binds:** any agent applying this skill +> **Tier:** T1 (detected) +> **Check:** `bin/skill-eval` <case or lint rule> +> **Escape:** <the compliant alternative — say plainly it carries no penalty> + +## Happy path + +<One worked example, end to end. Required by contract-audit 5.6.> + +## Edge case + +<One case that looks like the happy path and is not, with the correct handling. +Required by contract-audit 5.6.> + +## Verification + +<How a reviewer checks the skill was applied correctly, as a command or a +falsifiable observation.> + +<!-- +TEMPLATE NOTES — delete before saving. + +DESCRIPTION FIELD (contract-audit 5.3): front-load task keywords, then state a +negative boundary. This string is the whole trigger mechanism — an agent that +never loads the skill gets none of its judgement. + +LENGTH (5.4): long content goes in references/, not in SKILL.md. If this file +grows past roughly 200 lines, split it: + <skill>/SKILL.md + <skill>/references/<topic>.md + +SCRIPTS (5.5): skill scripts do deterministic helper work only. A script that +makes the judgement has replaced the skill. + +Verify with: bin/contract-lint --file <this file> +--> diff --git a/.context/templates/contracts/subagent.template.md b/.context/templates/contracts/subagent.template.md new file mode 100644 index 0000000..aa02948 --- /dev/null +++ b/.context/templates/contracts/subagent.template.md @@ -0,0 +1,64 @@ +--- +name: <subagent-name> +role: "<one line — the specialized job>" +status: "active" +updated: "<YYYY-MM-DD>" +--- + +# <Subagent name> + +<One paragraph: what this role is for, and what it deliberately does not do.> + +## Tools + +| Tool | Why this role needs it | +| --- | --- | +| <tool> | <reason> | + +**Denied:** <tools this role must not have, and why. An unstated denial is not +a boundary.> + +## Guardrails + +### 1. <Rule as one imperative sentence.> + +> **Binds:** this subagent +> **Tier:** T2 (blocked) +> **Check:** <tool allowlist, hook, or harness config that enforces it> +> **Escape:** <hand back to the caller with a named blocked status> + +### 2. <Scope boundary — what it may read and write.> + +> **Binds:** this subagent +> **Tier:** T1 (detected) +> **Check:** scope diff in the run receipt +> **Escape:** <alternative> + +## Procedure + +1. <Imperative step.> +2. <...> + +## Handoff + +**Returns:** <the distilled summary shape — fields, not raw output. Contract-audit +6.3: subagents return distilled summaries, not raw noise.> + +**Blocked status:** <the exact shape of an honest "I could not do this". A role +with no way to report failure will report success.> + +## Verification + +<How the caller checks the work, independent of what the subagent claims.> + +<!-- +TEMPLATE NOTES — delete before saving. + +A subagent is a ROLE (tools, guardrails, handoff shape). A repeatable procedure +belongs in a skill and is referenced from here, not restated (contract-audit 6.2). + +Write-heavy parallel subagents need conflict controls (6.4): name the worktree, +lock, or disjoint path set that keeps two instances from colliding. + +Verify with: bin/contract-lint --file <this file> +--> diff --git a/.context/templates/contracts/workflow.template.md b/.context/templates/contracts/workflow.template.md new file mode 100644 index 0000000..a3b332f --- /dev/null +++ b/.context/templates/contracts/workflow.template.md @@ -0,0 +1,75 @@ +--- +title: "<Workflow name>" +domain: "<domain>" +type: "workflow" +status: "active" +updated: "<YYYY-MM-DD>" +tags: ["workflow"] +--- + +# <Workflow name> + +**Trigger:** <the one condition that starts this. A workflow with no trigger is +a document nobody knows when to open — 18 of 42 workflows failed this on +2026-08-23.> + +**Stop state:** <what "done" looks like, and what to do when it goes wrong. +Name the blocked state explicitly: "if X, stop and record Y".> + +**Owner surface:** <path or project that owns updates to this file> + +## When not to use this + +<Negative boundary. The nearest neighbouring workflow and why this is not it.> + +## Preflight + +1. <Numbered, imperative, testable. Name the command.> +2. <...> + +## Procedure + +1. <Step. One action per step, with the exact command.> +2. <...> + +## Rules + +Rules that constrain how this workflow runs, rather than steps within it. +Omit this section if the procedure carries no normative weight. + +### 1. <Rule as one imperative sentence.> + +> **Binds:** anyone running this workflow +> **Tier:** T1 (detected) +> **Check:** `bin/<tool> --check` +> **Escape:** <the compliant alternative> + +## Verification + +<The deterministic command that proves the workflow ran and produced what it +claims. Not "review the output" — a command with an exit code.> + +## Outputs + +| Artifact | Path | Update rule | +| --- | --- | --- | +| <name> | `<path>` | append-only / replace-with-review / generated-only | + +## Escape + +<What to do when the workflow cannot complete. This is the workflow-level +honest-failure path: where to record the blocked state so a stall is visible +instead of silent.> + +<!-- +TEMPLATE NOTES — delete before saving. + +Workflows are OPERATOR SEQUENCES. If the content is reusable agent judgement it +belongs in .agents/skills/; if it is a role it belongs in agents/; if it is +deterministic it belongs in bin/. + +Do not duplicate skill bodies or project status here (contract-audit 4.5). + +Every workflow needs a trigger (4.1) and a stop state. Verify with: + bin/contract-lint --file <this file> +--> diff --git a/.context/templates/eval-output.md b/.context/templates/eval-output.md new file mode 100644 index 0000000..6c67256 --- /dev/null +++ b/.context/templates/eval-output.md @@ -0,0 +1,104 @@ +--- +title: "<Eval report title>" +domain: "<project-domain>" +type: "lab-report" +status: "active" +study_type: "exploratory" +lab_report_id: "YYYY-MM-DD-<slug>" +eval_run_id: "YYYY-MM-DD-<slug>" +protocol_ref: "plans/<file>.yaml@<git-short-sha>" +decision_sentence: "<What changes if positive / negative / inconclusive?>" +hypothesis: "<Expected before outcomes, or descriptive-only>" +disposition: "open" +project: "<project-slug>" +tags: ["evaluation", "eval-registry", "lab-report"] +updated: "YYYY-MM-DD" +source: "local" +--- + +<!-- + This file is the eval-registry specialization of the universal lab-report + convention (.context/templates/lab-report.md). Prefer bin/lab-report scaffold + for new work; keep the metric extract block for harvest. +--> + +# <Title> — YYYY-MM-DD + +## Boundary + +- Scope, DB/path, what this run does **not** claim. + +## Decision sentence + +<Repeat from frontmatter for human scan.> + +## Hypothesis + +<Expected before outcomes.> + +## Inputs + +| Field | Value | +|-------|-------| +| study_type | exploratory | +| protocol_ref | … | +| unit_of_analysis | | +| primary metric | | +| raw artifacts | `raw-materials/<run-id>/` | + +## Results + +<Tables and findings. Label claim types if interpretive.> + +## Irregularities + +<Table or "None observed." Also mirrored under metric extract.> + +## Counterevidence and limits + +<What would falsify the read; known confounds.> + +## Disposition + +`open` | `accept` | `reject` | `hold` | `iterate` — and next experiment id if iterate. + +## Metric extract (eval-registry) + +```yaml +registry: + project: <project-slug> + run_id: YYYY-MM-DD-<slug> + study_type: exploratory + protocol_ref: plans/queries.yaml@HEAD + date: YYYY-MM-DD + decision_sentence: "<one sentence>" + artifact_path: outputs/YYYY-MM-DD-<slug>.md + raw_path: raw-materials/YYYY-MM-DD-<slug>/ + environment: + git_sha: null + harness_version: null + decision_use: regression_only + prior_baseline_ref: outputs/YYYY-MM-DD-prior.md +metrics: + - name: <metric_name> + slice: <slice_id> + value: 0.0 + n: 0 + unit: ratio + ci_lower: null + ci_upper: null +irregularities: + - id: <slug> + severity: info + category: other + observation: "<what was odd>" + context: "<why it might matter or not>" + artifact_ref: "<path or null>" + resolved: false +``` + +## Claims (epistemic pass) + +| Statement | Type | Confidence | +|-----------|------|------------| +| … | observation / inference / hypothesis | high / moderate / low | \ No newline at end of file diff --git a/.context/templates/expert-field-insight.md b/.context/templates/expert-field-insight.md new file mode 100644 index 0000000..b8a80cb --- /dev/null +++ b/.context/templates/expert-field-insight.md @@ -0,0 +1,47 @@ +--- +title: "Expert field insight — <one-line claim>" +domain: "<regulated-systems | knowledge-systems | …>" +type: "note" +status: "synthesized" +source: "<path to correspondence, article, or call recap>" +expert: "<Name>" +channel: "<call | LinkedIn | article | email>" +heard: "YYYY-MM-DD" +created: "YYYY-MM-DD" +confidence: "low | moderate" # practitioner judgment is almost never high +claim_type: "source-claim" # attributed to expert; not settled fact +tags: ["expert-field-insight", "<expert-slug>", "<topic>"] +project_hooks: ["example-eval", "example-retrieval"] # optional local project slugs — tags only, not wikilinks to 30_projects +links: [] +--- + +# Expert field insight — <one-line claim> + +**Attribution:** <Name>, <role if known>, via <channel> on <date>. +**Confidence:** low/moderate — practitioner judgment / published opinion, not peer-reviewed measurement. +**Not:** a regulation, a validation result, or MainFrame policy. + +## What they said (paraphrase carefully; quote when load-bearing) + +> + +## Why it matters here + +- + +## Project hooks (optional use, not commitments) + +| Project / area | How it might apply | Do not overread as | +|---|---|---| +| example-eval | | product requirement | +| example-retrieval | | | + +## Counter / limits + +- Single expert; domain/context may not transfer. +- + +## Provenance + +- Correspondence or article: `<path>` +- Ledger: `10_knowledge/knowledge-systems/…expert-field-insights-ledger…` diff --git a/.context/templates/methodology-approach.md b/.context/templates/methodology-approach.md new file mode 100644 index 0000000..f1ef620 --- /dev/null +++ b/.context/templates/methodology-approach.md @@ -0,0 +1,52 @@ +--- +title: "Methodology approach — read first in new session" +domain: "<project-domain>" +type: "note" +status: "active" +source: "EVAL_METHODOLOGY.md eval-profile template" +tags: ["methodology", "handoff", "eval-profile", "eval-planning"] +study_type: null +eval_run_id: null +updated: "YYYY-MM-DD" +--- + +# ⚠️ Methodology approach — <Project Name> + +**New session:** Read this before any eval run, output write, or promotion claim. + +**Contracts:** [EVAL_METHODOLOGY.md](../../EVAL_METHODOLOGY.md). Optional/deferred: a longer eval-methodology workflow if installed locally. + +## Eval profile + +| Field | Value | +|-------|-------| +| **Strict profile** | yes | +| **Study types used** | exploratory / regression / confirmatory / observational / calibration | +| **Primary decision** | <what promotion or change this eval supports> | +| **Current blocker** | <calibration / hold-out / H2 / none> | + +## Do in order + +1. <project-specific sequence> +2. Write `outputs/` with [.context/templates/eval-output.md](../../.context/templates/eval-output.md) +3. `bin/eval-registry harvest` then `bin/eval-registry check --strict` + +## Do not + +- Run unlabeled study_type +- Omit metric extract or `irregularities` (use `[]` only after explicit scan) +- Promote from exploratory alone +- State confirmatory conclusions before blockers clear + +## Knowledge + +- `10_knowledge/knowledge-systems/methodology/` — design, sample size, trends, irregularities +- Project `methodology-approach.md` sections below — living posture + +## Living posture + +(Update after each eval cycle.) + +| Date | eval_run_id | study_type | Headline | Open irregularities | +|------|-------------|------------|----------|---------------------| +| — | — | — | — | — | \ No newline at end of file diff --git a/.context/templates/task-packet.md b/.context/templates/task-packet.md index 25fb667..505948e 100644 --- a/.context/templates/task-packet.md +++ b/.context/templates/task-packet.md @@ -61,6 +61,12 @@ List observable conditions that external verification can establish. Name discoveries that require escalation rather than improvisation. +## Resilience & Hazard Considerations (ADR-052) + +- **Degraded-mode behavior:** How the implementation behaves if upstream dependencies, models, or external services are unavailable (must fail closed or use explicit fallback). +- **Escape valves & error shapes:** Explicit intermediate or failure states emitted instead of guessing or swallowing exceptions. +- **Rollback / retreat pathway:** Safe revert command or worktree cleanup steps if unexpected regressions occur. + ## Expected handoff Require a concise summary of changed files, verification not claimed as run by diff --git a/.context/templates/writing-style-skill/README.md b/.context/templates/writing-style-skill/README.md deleted file mode 100644 index 4bdf954..0000000 --- a/.context/templates/writing-style-skill/README.md +++ /dev/null @@ -1,49 +0,0 @@ -# Template: personal writing-style skill - -**Status:** scaffold only (Phase 1b) — fill when you build your own voice pack -**Public intent:** ship this template with MainFrame; never ship a filled personal pack as the default skill. - -## What this is - -A copy-local recipe for a **router skill** plus genre reference pages. Operators clone the structure into a local-only skill path and fill it with their own standards. - -## Install (local only) - -```bash -# From MainFrame root — creates an ignored personal skill (do not commit) -mkdir -p .agents/skills/writing-style/references -cp .context/templates/writing-style-skill/SKILL.template.md \ - .agents/skills/writing-style/SKILL.md -cp .context/templates/writing-style-skill/references/*.template.md \ - .agents/skills/writing-style/references/ -# Rename *.template.md → short names matching SKILL.md load order -``` - -Ensure `.agents/skills/writing-style/` stays in `.gitignore` (MainFrame default). - -## Structure - -| File | Role | -| --- | --- | -| `SKILL.template.md` | Frontmatter + load order + workflow (no personal voice) | -| `references/shared-natural-voice.template.md` | Always-on craft principles | -| `references/portfolio.template.md` | Public technical / portfolio prose | -| `references/content.template.md` | Posts, newsletters, comments | -| `references/career.template.md` | Applications, outreach (name the lane you use) | -| `references/technical-professional.template.md` | Internal / professional docs | -| `references/persuasive-craft.template.md` | Asks, proposals, decisions | -| `references/humour-rhythm.template.md` | Optional rhythm / tone pass | - -Add or drop reference pages to match your surfaces. Keep the router thin; put length in references. - -## Design rules - -1. **Positive craft first** — anti-patterns only as a final lint. -2. **Truth before style** — no invented metrics, status, or credentials. -3. **Load the smallest set** of references per task. -4. **Project-specific voice packs** for one client or one private project belong under that project (e.g. `30_projects/<slug>/skills/`), not in root `.agents/skills/`. - -## Deferred - -- Example filled pages (synthetic, not a real person's voice) -- Optional eval harness hooks (`eval_writing_style`-style) as a separate template diff --git a/.context/templates/writing-style-skill/SKILL.template.md b/.context/templates/writing-style-skill/SKILL.template.md deleted file mode 100644 index 64fc42f..0000000 --- a/.context/templates/writing-style-skill/SKILL.template.md +++ /dev/null @@ -1,46 +0,0 @@ ---- -name: writing-style -description: >- - Use when writing, revising, coaching, or style-checking prose for this operator. - Loads a compact router, then task-specific reference pages. Optimize for authentic - voice, specificity, calibrated judgment, and readable prose — not detector evasion. ---- - -# Writing Style - -## Purpose - -Make prose sound like a specific person with judgment, not a language model avoiding tells. - -Core rule: write from **positive craft principles** first. Use anti-patterns only as a final lint. - -## Load order - -1. Always read `references/shared-natural-voice.md` first. -2. Load only the reference pages needed for the task (see list below). -3. Do not load every reference by default. - -Suggested mode pages (rename/add to fit your work): - -| Reference | When | -| --- | --- | -| `references/portfolio.md` | Public READMEs, ADRs, project pages, proof-of-work | -| `references/content.md` | Posts, newsletters, comments, audience content | -| `references/career.md` | Applications, recruiter mail, role-facing copy | -| `references/technical-professional.md` | Internal docs, memos, technical explanation | -| `references/persuasive-craft.md` | Proposals, asks, decision memos | -| `references/humour-rhythm.md` | Naturalness / rhythm pass when prose feels flat or too smooth | - -## Workflow - -1. Identify surface, audience, and whether the reader must **act or decide**. -2. Load shared voice + the smallest mode set. -3. If there is an ask (hire, fund, approve, adopt), also load persuasive craft. -4. Preserve truth before style — no invented credentials, metrics, status, or sources. -5. Draft around concrete evidence and the reader's actual question. -6. Final lint: banned phrases, unsupported claims, formatting, read-aloud quality. - -## Out of scope - -- Changing product claims without evidence -- Impersonating another person's private voice pack from a template diff --git a/.context/templates/writing-style-skill/references/career.template.md b/.context/templates/writing-style-skill/references/career.template.md deleted file mode 100644 index 2443b74..0000000 --- a/.context/templates/writing-style-skill/references/career.template.md +++ /dev/null @@ -1,14 +0,0 @@ -# Career / applications (fill in) - -## Surfaces - -- Cover letters, bullets, recruiter notes, fellowship answers - -## Rules - -- Truthful scope; no inflated titles or metrics -- Map evidence to the role's actual needs - -## Reject - -- _ diff --git a/.context/templates/writing-style-skill/references/content.template.md b/.context/templates/writing-style-skill/references/content.template.md deleted file mode 100644 index 3f2b5ab..0000000 --- a/.context/templates/writing-style-skill/references/content.template.md +++ /dev/null @@ -1,13 +0,0 @@ -# Content / posts (fill in) - -## Surfaces - -- e.g. LinkedIn, newsletter, blog - -## Hook → evidence → takeaway - -- _ - -## Reject - -- Engagement bait without a real point diff --git a/.context/templates/writing-style-skill/references/humour-rhythm.template.md b/.context/templates/writing-style-skill/references/humour-rhythm.template.md deleted file mode 100644 index 2e48c59..0000000 --- a/.context/templates/writing-style-skill/references/humour-rhythm.template.md +++ /dev/null @@ -1,14 +0,0 @@ -# Humour / rhythm (fill in) - -## When to load - -Prose feels too smooth, corporate, or lifeless. - -## Rules - -- Humour serves clarity or rapport; never the claim -- Prefer understatement over punchlines in professional surfaces - -## Reject - -- _ diff --git a/.context/templates/writing-style-skill/references/persuasive-craft.template.md b/.context/templates/writing-style-skill/references/persuasive-craft.template.md deleted file mode 100644 index eb90993..0000000 --- a/.context/templates/writing-style-skill/references/persuasive-craft.template.md +++ /dev/null @@ -1,14 +0,0 @@ -# Persuasive craft (fill in) - -## When to load - -Reader must decide, approve, hire, fund, or act. - -## Structure - -- Context → options → recommendation → risks → ask - -## Reject - -- Pressure without evidence -- Hidden tradeoffs diff --git a/.context/templates/writing-style-skill/references/portfolio.template.md b/.context/templates/writing-style-skill/references/portfolio.template.md deleted file mode 100644 index 50d328d..0000000 --- a/.context/templates/writing-style-skill/references/portfolio.template.md +++ /dev/null @@ -1,19 +0,0 @@ -# Portfolio / public technical prose (fill in) - -## Audience - -- _ - -## Must show - -- What was built, what was measured, what was refused -- Limits and non-claims - -## Tone - -- _ - -## Reject - -- Hype without evidence -- Fake precision diff --git a/.context/templates/writing-style-skill/references/shared-natural-voice.template.md b/.context/templates/writing-style-skill/references/shared-natural-voice.template.md deleted file mode 100644 index 30b1f4d..0000000 --- a/.context/templates/writing-style-skill/references/shared-natural-voice.template.md +++ /dev/null @@ -1,21 +0,0 @@ -# Shared natural voice (fill in) - -Replace bullets with **your** defaults. Keep this page short enough to load every time. - -## Principles - -- Prefer concrete nouns and verbs over abstract stacking. -- Specificity beats polish: one real detail > three generic intensifiers. -- Judgment is allowed: say what you think, then show why. -- Rhythm: vary sentence length; cut throat-clearing openers. -- Humour: optional; only when it clarifies stance or reduces heat. - -## Always avoid (your lint list) - -- _ -- _ -- _ - -## Signature habits (optional) - -- _ diff --git a/.context/templates/writing-style-skill/references/technical-professional.template.md b/.context/templates/writing-style-skill/references/technical-professional.template.md deleted file mode 100644 index 4f9e522..0000000 --- a/.context/templates/writing-style-skill/references/technical-professional.template.md +++ /dev/null @@ -1,14 +0,0 @@ -# Technical / professional (fill in) - -## Surfaces - -- Internal docs, design notes, memos, explanations - -## Rules - -- Define terms once; prefer operational language -- Separate observation, inference, and recommendation - -## Reject - -- _ diff --git a/.context/workflows/audit-sweep.md b/.context/workflows/audit-sweep.md deleted file mode 100644 index d102be3..0000000 --- a/.context/workflows/audit-sweep.md +++ /dev/null @@ -1,106 +0,0 @@ -# Audit Sweep Workflow - -Use this workflow to ensure that material routed under the tiered ingest policy (ADR-019) receives post-placement verification. The goal is to make the compensating control for review-after (needs-audit tags + epistemic auditor) visible and routine before relying on Tier A auto-apply at scale. - -## Philosophy -- Placement can be cheap and reversible (byte snapshots + append-only ledgers + no overwrites). -- Truth and claim quality are verified after the fact by the epistemic research system, not blocked at every file. -- The operator sees a clear pending-review surface in `20_live/epistemic-audit/` rather than having to hunt through 10_knowledge/. -- Deterministic discovery of candidates + LLM judgment for claim extraction/auditing. - -## Prerequisites -- Files routed via Tier A or batch mode should carry `needs-audit` (or `needs-verification`) in their `tags:` list (enforced in ingest-agent batch mode and per-file enrichment). -- The epistemic-research-system workbench is functional (`bin/epistemic` wrapper). -- `20_live/epistemic-audit/pending-review/` and `contradictions/` directories exist (created by mainframe_bridge). - -## Script / Entry Points -- Primary: `bin/epistemic sweep --mainframe` (orchestrator subcommand targeting MainFrame durable knowledge). -- Helper: `bin/audit-sweep` (thin deterministic wrapper; see below). -- Dry / check: `bin/audit-sweep --dry-run` or `bin/epistemic sweep --mainframe --dry-run`. -- Integration: Called from `bin/session-close --apply` (or reminded in `--check`). - -## Steps -1. **Discover candidates deterministically** - - Scan `10_knowledge/` (or a `--subset` per ADR-021) for Markdown files whose frontmatter `tags` list contains `needs-audit`. - - Also surface recent (last N days) routed raw items even without the tag for coverage sampling. - - Record the manifest: file path, date, domain, rule (if in ledger), current audit status. - -2. **Surface the worklist** - - Write or update `20_live/epistemic-audit/pending-review/sweep-YYYY-MM-DD.md` (or per-run). - - For each candidate, include: - - Link to the canonical file in 10_knowledge. - - Extracted frontmatter (title, source, tags). - - Brief context from the first section or Connections if present. - - Placeholder for auditor output (claim list + verdicts). - -3. **Run the auditor (judgment layer)** - - Delegate to the epistemic harness (Researcher/Auditor agents via `bin/epistemic`). - - Use existing `harness/auditor*.py`, `evaluator.py`, and belief_db. - - Focus on claim extraction, support/contradiction detection, confidence. - - For raw items: treat the document as source; generate or validate claims (respecting that raw body is evidence). - - Publish full audit reports via the existing `mainframe_bridge.publish_audit_report` mechanism. - -4. **Promote or park results** - - High-confidence supported claims: the auditor (or operator review) can propose promotion to `status: stable` or synthesis into a companion `note`. - - Contradictions or low-confidence: move to `contradictions/` or leave tagged `needs-audit` with notes. - - Update the source file's tags (remove `needs-audit`, add `audited-YYYY-MM-DD` or `verified` / `disputed`) **only after operator review** for Tier C items or the first calibration runs. - - Append to the relevant batch disposition-ledger where applicable (cross-reference by path or hash). - -5. **Report and close the loop** - - Emit counts: candidates found, audited, supported, contradicted, still pending. - - Include in the next `bin/workflow-report`. - - Update `20_live/epistemic-audit/` index or manifest if one exists. - - Refresh MindGraph if new synthesized notes were created. - -## Guardrails -- Never mutate raw evidence bodies. Only frontmatter tags/status and separate audit artifacts. -- The first several full sweeps after enabling Tier A are treated as calibration data (measure agreement with operator on a sample, per ADR-018/019). -- `needs-audit` is a signal for the system, not a permanent label. It should age off after verified handling. -- High-risk domains (finance sensitivity per routing-policy) remain Tier C regardless of auditor readiness. -- All automated writes to 20_live/ follow the volatility rules (dated, append or explicit snapshots). - -## Related -- [.context/routing-policy.md](../routing-policy.md) — when `needs-audit` must be applied. -- [agents/ingest-agent.md](../../agents/ingest-agent.md) — batch and per-file tagging responsibility. -- [30_projects/epistemic-research-system/](../../30_projects/epistemic-research-system/) — the auditor implementation and mainframe_bridge. -- [.context/workflows/process-evaluation.md](../process-evaluation.md) — evaluate this sweep before widening Tier A autonomy. -- [DECISIONS.md](../../DECISIONS.md) — ADR-019 and follow-ups. -- `bin/session-close` and `bin/workflow-report` for integration points. - -## Implementation Notes (Current State — 2026-06-18) - -**Done:** -- `bin/audit-sweep` — discovery, manifest to `20_live/epistemic-audit/pending-review/sweep-*.md`, `--json`, `--subset` -- `mainframe_bridge.sweep_knowledge_for_audit_tags()` — calls audit-sweep dry-run JSON - -**Not done (see `30_projects/epistemic-research-system/plans/remediation-backlog.md` Tier 1):** -- `bin/epistemic sweep --mainframe` — documented here but **not implemented** -- `audit-sweep --apply` does not invoke auditor; handoff is manual: `bin/epistemic run --source <file>` per manifest row -- Planned: `bin/epistemic audit-manifest --file <sweep.md> [--limit N]` - -**ERS paths to avoid (ADR-011):** `watch`/`daemon`, `sweep-knowledge` auto-ingest, `audit-corpus` paragraph indexing — disabled by default. - -**Still desired:** -- Hook session-close (remind in `--check`; optional `--apply`) -- workflow-report coverage metric ("X needs-audit, Y swept, Z audited") - -**Canonical workflow:** `30_projects/epistemic-research-system/plans/epistemic-workflow.md` - -## The standard this sweep enforces - -This workflow is the post-placement half of the promotion gate in -`EPISTEMIC_STANCE.md`, so the classification rules there are its input, not -background reading. Full procedure: [epistemic-standard.md](epistemic-standard.md). - -- `status: stable` and `status: audited` mean **operator-verified or sweep-cleared**, - never LLM output alone. `bin/capture-validate` R6 enforces this and additionally - rejects any file that claims a verified status while still carrying `needs-audit`. -- The sweep clears `needs-audit`. It does not confer `stable` by itself. -- **Escape:** if a claim cannot be verified, leave it `synthesized` and record why. - An unresolved item with a stated reason is a complete sweep result. - -**Stop state.** A sweep that cannot reach its sources stops and reports the gap. -It does not clear `needs-audit` on unreachable material. - -Measured 2026-08-10, before R6: 29 files claimed `stable`, none carried -verification evidence, and 14 also carried `needs-audit`. diff --git a/.context/workflows/craft-research-loop.md b/.context/workflows/craft-research-loop.md deleted file mode 100644 index 145ac9c..0000000 --- a/.context/workflows/craft-research-loop.md +++ /dev/null @@ -1,218 +0,0 @@ -# Craft Research Loop - -**Product / craft trial-and-error research.** Separate from literature and system evals. - -| Loop | Question | Ends with | -|------|----------|-----------| -| research-lane-loop | What do sources say? | Synthesis in `10_knowledge/` | -| project-experiment-loop | What does a protocol measure? | Registry + `last-eval-action.md` | -| **craft-research-loop** | What works for *this* product/stack? | Trial FINDINGS + proof index + project `next_action` | - -## Profile - -| Field | Value | -|-------|-------| -| structural_type | workflow | -| owner_surface | `30_projects/<slug>/` craft labs | -| skill | `.agents/skills/craft-research-loop/SKILL.md` | -| CLI | `bin/craft-research-loop` | - -## Purpose - -Make trial-and-error **finishable and citable** without dumping every smoke run into eval-registry or research lanes. - -```text -question → scaffold trial → run → FINDINGS (keep|kill|iterate) - → proof index / cull → log + next_action → promote? → loop -``` - -## When to use - -- Image/video generation bake-offs, runtime stack trials -- Integration prototypes (phone bridge, dashboard UX) -- “Does this setting/workflow/tool path work for us?” -- Any project with many intermediate artifacts and a need to **keep only proof** - -## When not to use - -- Peer-reviewed / institutional literature → research-lane-loop -- Harness/MindGraph/process metrics that change promotion gates → project-experiment-loop -- Pure unit-test green/red → CI - -## Loop unit - -| Field | Rule | -|-------|------| -| **Scope** | One project × **one primary question** | -| **Artifact** | `outputs/<YYYY-MM-DD>-<slug>/` (+ optional `FINDINGS.md`) | -| **Start** | Written question + success criteria *before* reviewing winners | -| **End** | Verdict recorded + proof index updated + `next_action` set | -| **Promotion** | Optional: knowledge extract, research lane, or single eval_run_id | - -## Command card - -```bash -# 0) Select project + see open trials -bin/craft-research-loop preflight --project image-generation-lab -bin/craft-research-loop triage --project image-generation-lab - -# 1) Open a trial (before running) -bin/craft-research-loop scaffold \ - --project image-generation-lab \ - --question "Does SDXL light refine beat prep-only Lanczos on green-lizard micro-detail?" \ - --decision "If no, keep prep-first defaults; if yes, change default refine path." \ - --slug "sdxl-light-vs-prep" - -# 2) Run the trial in the workbench (project-specific tools) -# ... generate, measure, save artifacts into the trial folder ... - -# 3) Close (mandatory) -bin/craft-research-loop close \ - --project image-generation-lab \ - --trial 2026-07-13-sdxl-light-vs-prep \ - --verdict keep|kill|iterate \ - --findings "One-paragraph observation + recommendation" - -# 4) Optional promotion -bin/craft-research-loop promote --project image-generation-lab --trial … --to knowledge -bin/craft-research-loop promote --project image-generation-lab --trial … --to experiment -bin/craft-research-loop promote --project image-generation-lab --trial … --to lane -``` - -## Steps - -### 0. Preflight - -- Read project README `goal` / `next_action` -- Confirm proof index path (default `outputs/RESEARCH_PROOF_INDEX.md`) -- List open trials (folders without FINDINGS or status: running) -- Dual MindGraph only if the trial depends on prior domain knowledge (optional) - -### 1. Frame the trial - -Write **one** primary question and the decision it supports. No multi-question bake-offs without a primary. - -Template: `.context/templates/craft-trial.md` (also written by `scaffold`). - -### 2. Scaffold - -Create: - -```text -30_projects/<slug>/outputs/YYYY-MM-DD-<trial-slug>/ - TRIAL.md # question, criteria, procedure (pre-results) - FINDINGS.md # filled at close (or during) - # media / logs / json as needed -``` - -Register row in proof index as **open** (or omit until close — prefer open for WIP visibility). - -### 3. Execute - -- Record exact models, params, hardware, manual interventions -- Separate **observation** from **interpretation** -- Keep failed runs (do not delete only because ugly) -- Do not change the primary question mid-trial; open a new trial instead - -### 4. Close (definition of done) - -- [ ] `FINDINGS.md` has: observations, interpretation, limitations, **verdict** -- [ ] Verdict is exactly one of: `keep` | `kill` | `iterate` -- [ ] Proof index updated (KEEP / remove / supersede) -- [ ] Ephemeral intermediates culled **or** explicitly kept with reason -- [ ] Project `log.md` one entry -- [ ] Project `next_action` updated -- [ ] Action card: `30_projects/<slug>/outputs/LAST_CRAFT_ACTION.md` - -### 5. Promote (optional, explicit) - -| Verdict path | Promotion | -|--------------|-----------| -| Durable domain lesson | `bin/extract-knowledge` or research-lane-loop capture | -| Decision-bearing metric for promotion gates | One `eval_run_id` via project-experiment-loop scaffold | -| Recurring cross-project question | research-lane-intake candidate | -| Local product default only | `decisions.md` + proof index is enough | - -### 6. Loop decision - -| Outcome | Next | -|---------|------| -| `iterate` | New trial slug; link prior FINDINGS | -| `keep` | Adopt default; optional decisions.md ADR | -| `kill` | Record anti-pattern; don’t re-run without new hypothesis | -| Knowledge gap | research-lane-loop | -| Need formal metric | project-experiment-loop | - -## Proof index contract - -Every craft lab should maintain `outputs/RESEARCH_PROOF_INDEX.md` (or path in README): - -| Section | Content | -|---------|---------| -| **Keep** | Paths that still prove a disposition | -| **Removed** | Culled ephemera + why | -| **Open** | Running trials (optional) | -| **How to re-run** | Commands / plan pointers | - -## Project-local action card - -Unlike eval portfolio triage, craft action is **per project**: - -```text -30_projects/<slug>/outputs/LAST_CRAFT_ACTION.md -``` - -Written by `close` and `triage`. Not mixed into `last-eval-action.md`. - -## Anti-patterns - -- Endless generates with no FINDINGS -- Promoting a single pretty image to a product default -- Leaving durable knowledge only in `outputs/` forever -- Stuffing bake-offs into eval-registry by default -- Multi-question “matrix” without a primary success criterion - -## Related - -- research-lane-loop, project-experiment-loop -- `30_projects/AGENTS.md` extract-knowledge rule -- Image lab reference: `outputs/RESEARCH_PROOF_INDEX.md` -- Workbench EXP template (project-local) can feed TRIAL.md - -## Dogfood reference - -| Date | Project | What ran | Smooth? | -|------|---------|----------|---------| -| 2026-07-13 | image-generation-lab | Close realism bake-off **keep**; pixel SDXL smoke **kill**; stack upgrades **keep**; scorecard **iterate** (weights); scaffold post-install | Yes — close works on legacy folders; triage prefers formal `running` | -| 2026-07-13 | integration-lab | Scaffold E01 G1; **iterate** (ttyd/tmux/tailscale missing); scaffold after-install follow-up | Yes — first trial creates proof index; blocked-on-tools is valid iterate | - -### Friction observed - -| Issue | Mitigation | -|-------|------------| -| Many legacy `outputs/` dirs look “open” | Triage prioritizes formal `TRIAL.md` running; artifact-only is low priority | -| `iterate` needs immediate scaffold | CLI tip after close; operator should scaffold next question same session | -| Close without prior TRIAL.md | Allowed; still writes FINDINGS + proof index | -| Blocked on install/operator | Use **iterate**, not hang forever on `running` | - -### Smooth path (confirmed) - -```text -preflight → close legacy keep|kill → scaffold next → (blocked?) close iterate → scaffold tighter -``` - -## Dogfood — batch20 / multi-loop (2026-07-13) - -### What worked -- Close-all legacy artifact folders with explicit keep + findings -- Multi-project preflight/triage (13 projects total across two sets) exit 0 -- Project-local LAST_CRAFT_ACTION stays out of eval portfolio - -### Friction → improve -| Issue | Fix direction | -|-------|----------------| -| Proof index append creates duplicate section headers | Structured rewrite on close | -| Iterate-for-operator-install still reads as “open work” | `blocked_on: operator` tone on action card | -| skill-eval lint red on craft skill | Schema alignment; not a close blocker | - -Cross-loop gaps: `30_projects/mainframe-process-eval/outputs/2026-07-13-three-loop-process-gaps.md`. diff --git a/.context/workflows/create-thread.md b/.context/workflows/create-thread.md deleted file mode 100644 index 94650da..0000000 --- a/.context/workflows/create-thread.md +++ /dev/null @@ -1,40 +0,0 @@ ---- -title: "Create Thread Workflow" -domain: "agents" -type: "workflow" -status: "active" -source: "ADR-038 and existing prompt/session contracts" -tags: ["threads", "handoff", "prompting"] -updated: "2026-06-28" -structural_type: "workflow" -lifecycle_scope: "root" -owner_surface: ".context/workflows/" -authority: "workflow-contract" -privacy: "public-safe" -volatility: "stable" -source_of_truth: true -update_rule: "replace-with-review" -verification: ["confirm referenced skill and workflows exist"] -related_surfaces: [".agents/skills/prompt-creation/SKILL.md", ".context/workflows/session-open.md", ".context/workflows/session-close.md"] -do_not_use_for: ["project status", "one-off implementation plans", "prompt-engineering guidance"] ---- - -# Create Thread Workflow - -Use this workflow only when the user explicitly asks for a separate thread. - -1. Load the minimum current context using `session-open` conventions and the - task's authoritative project files. For `30_projects/` planning, preserve - the required dual MindGraph query groups. -2. Use [prompt-creation](../../.agents/skills/prompt-creation/SKILL.md) to design - the seed prompt. Point to source files instead of copying large documents. -3. Include the task objective, mode when requested (for example `/plan`), - boundaries, expected first response, and completion or approval gate. -4. Create the thread in the correct project with the available thread tool, - then give it a clear title and return the created-thread reference. - -Do not copy secrets, raw transcripts, or unverified status into the prompt. A -plan-shaped handoff is context, not implementation approval. - -Related workflows: [session-open](session-open.md) and -[session-close](session-close.md). diff --git a/.context/workflows/delegate-local-task.md b/.context/workflows/delegate-local-task.md deleted file mode 100644 index 2213c5c..0000000 --- a/.context/workflows/delegate-local-task.md +++ /dev/null @@ -1,56 +0,0 @@ -# Delegate A Prepared Task - -Use this workflow when a frontier model has already planned a bounded task and -the implementation can be delegated to a local agent. - -## 1. Prepare The Packet - -1. Copy `.context/templates/task-packet.md` to - `30_projects/<project>/plans/task-packets/<task-id>.md`. -2. Resolve implementation choices. For non-trivial packets: `bin/mindgraph doctor`, dual - MindGraph queries, and fill optional `## MindGraph Query Pass` (template: - `.context/templates/mindgraph-query-pass.md`). Set `mindgraph_mode: "curated"` and - `knowledge_queries` when the packet depends on vault context. Then set editable files, acceptance criteria, - verification commands, boundaries, and stop conditions. -3. Keep `status: draft` while any material decision remains. -4. Validate the draft: - - ```bash - bin/task-packet validate 30_projects/<project>/plans/task-packets/<task-id>.md - ``` - -5. Review the packet, change it to `status: ready`, and validate with - `--require-ready`. - -## 2. Compile Project Tasks - -```bash -bin/task-packet compile -bin/generate-project-tasks -``` - -The packet manifest is generated local state. The task board may display its -summary, but the reviewed Markdown packet remains the source of truth. -After a `ready` packet is compiled, its contract hash is immutable. Retire it -or create a new `task_id`; do not rewrite or downgrade it. - -## 3. Execute In Isolation - -- Never run a first evaluation directly in an active project worktree. -- Use the Agent Harness Evaluation runner or another isolated copy/worktree. -- The executing agent may edit only `editable_files` and `create_files`. -- `read_only_files` provide context but are outside the write surface. -- Verification runs outside the agent through argv execution with no shell. - -## 4. Accept Or Reject - -Accept a run only when: - -- the external verifier passes; -- changed files remain inside the packet scope; -- no commits were created unless the packet explicitly delegates commits; -- the agent's completion claim matches the verifier result; -- human diff review finds no repair is required. - -Record run evidence in a receipt. Do not rewrite the ready packet with run -results. diff --git a/.context/workflows/deterministic-tool-standard.md b/.context/workflows/deterministic-tool-standard.md new file mode 100644 index 0000000..10b89af --- /dev/null +++ b/.context/workflows/deterministic-tool-standard.md @@ -0,0 +1,136 @@ +# Deterministic Tool & Execution Integrity Standard + +**Operating standard for executable tools, integration bridges, and evaluation tests.** This workflow enforces that every script in `bin/` or `scripts/` physically executes the computation it claims, tests real operational behavior rather than tautological mocks, and never allows agents to simulate deterministic outputs in prose. + +## Profile + +| Field | Value | +|---|---| +| structural_type | workflow | +| owner_surface | `AGENTS.md` (item 10) + `HARNESS.md` (ADR-051) | +| related | `epistemic-standard.md`, `eval-methodology.md`, `lab-report.md`, `process-evaluation.md` | + +--- + +## 1. Purpose & The Anti-Simulation Rule + +LLMs naturally suffer from a **"Scaffold-as-Execution" bias**: they generate schemas, parsers, and data packets, treat the pipeline as complete, and then synthesize plausible-looking evaluation tables or verdicts in Markdown without ever executing the underlying engine. + +This standard establishes the **Anti-Simulation Rule**: +> **Never simulate compute.** A tool must physically invoke the model, compiler, database, or evaluation engine it represents, or fail closed. An agent must never author an audit table, verification scorecard, or verdict claiming deterministic tool authority unless backed by a persistent execution receipt on disk. + +--- + +## 2. The Three-Layer Tool Taxonomy + +Every script, CLI command, and function in MainFrame must strictly belong to and be named after one of three layers: + +``` +┌────────────────────────────────────────────────────────┐ +│ Layer 1: Data Transformer / Packet Builder │ +│ Prefix: parse-*, build-*, format-*, stage-* │ +│ Role: Pure data transformation & serialization. │ +│ Rule: MUST NOT be named 'audit', 'verify', or │ +│ 'evaluate'. MUST NOT emit verdicts. │ +└────────────────────────────────────────────────────────┘ + │ + ▼ +┌────────────────────────────────────────────────────────┐ +│ Layer 2: Operational Engine Runner │ +│ Prefix: run-*, audit-*, verify-*, execute-* │ +│ Role: Physically executes the engine/model/DB. │ +│ Rule: MUST fail closed if engine is unavailable. │ +│ MUST write a persistent trace receipt. │ +└────────────────────────────────────────────────────────┘ + │ + ▼ +┌────────────────────────────────────────────────────────┐ +│ Layer 3: Agent Qualitative Judgment │ +│ Prefix: synthesize-*, review-*, appraise-* │ +│ Role: LLM reasoning, contextual trade-offs. │ +│ Rule: MUST be labeled 'evaluator: agent_llm'. │ +│ MUST NOT claim deterministic tool authority. │ +└────────────────────────────────────────────────────────┘ +``` + +### Naming & Responsibility Rules: +1. **Transformers (`build-*`, `format-*`, `parse-*`)**: + - Only transforms inputs (e.g. Markdown → JSON request packet). + - Must document clearly: *"Prepares packet for `<engine>`; does not execute `<engine>`."* +2. **Runners (`run-*`, `audit-*`, `verify-*`)**: + - Ingests the packet, executes the deterministic engine (e.g. `claim_audit_lab`, PyTorch model, SQLite query, external subprocess), and writes output traces. + - Must import or execute the real engine binary. +3. **Agent Heuristics (`synthesize-*`, `appraise-*`)**: + - Contextual synthesis that cannot be evaluated deterministically. + - Must be transparently attributed to LLM judgment. + +--- + +## 3. The Fail-Closed Boundary + +When a runner depends on external engines, local models, or environment packages (e.g. CAL, spaCy, PyTorch, Ollama, MLX, SQLite): + +1. **Preflight Dependency Check**: Check that the engine binary, model weights, or Python package are importable and healthy before processing. +2. **No Mock Fallbacks in Production**: If the engine is missing or unreachable, the runner must exit with a non-zero status and print an explicit remediation message (e.g., `EngineUnavailableError: CAL v1 engine dependencies missing in environment`). +3. **Never Fallback to Agent Simulation**: Do not catch engine failure and replace it with a simulated result or blank "pass" object. + +### 3.1 Degraded Runtimes, Exit Codes, and Escape-Valve Contracts (ADR-052) + +Complex systems run as broken systems (Cook 1998). Software and CLI tools must maintain resilience and clear envelope visibility when running in degraded environments: + +1. **Standard Exit Code Contract**: + - `0`: Clean execution, zero policy violations / verification passed. + - `1`: Operational success, but verification findings or lint/contract errors detected (actionable findings). + - `2`: Tool crash, preflight failure, missing dependency, or unhandled exception (tool did not run cleanly; do not read as a verification result). +2. **Escape-Valve Contract for Quality Gates**: + - Every gate validator (e.g., `bin/capture-validate`, `01_ingest/minion.py`, `bin/prep-ingest`) must provide structured, valid escape routes for partial or uncertain data (e.g., `type: hypothesis`, `status: queued`, `needs-audit`). + - Rigid gates that lack escape valves create perverse incentives at the sharp end, forcing agents to fabricate metadata to satisfy the gate. A documented hypothesis or pending item is always superior to an invented source. +3. **Edge-of-the-Envelope Telemetry**: + - When tools operate in degraded mode (e.g., daemon offline fallback, cached vector lookup, context truncation), the runner must write explicit diagnostic telemetry into the persistent receipt (`degraded: true`, `fallback_used: "local_sqlite"`, `boundary_warning: "token compaction near limit"`). + +--- + +## 4. Falsification & Adversarial Test Discipline + +A test suite that only tests positive happy paths or asserts against internal dictionary keys is **tautological**—it proves the code can run its own lines, not that it correctly validates reality. + +Every test file under `tests/` for tools or evaluation bridges must implement at least **two falsification controls**: + +### Control A: The Disconnection / Engine Failure Test +* Mock or simulate the backend engine being absent, raising an error, or timing out. +* **Assertion**: The tool must raise an exception or exit non-zero. It must **not** return a success code or partial fake verdict. + +### Control B: The Adversarial / Contradictory Input Test +* Feed the tool deliberately false data (e.g. invalid DOI, contradicted claim, ungrounded passage). +* **Assertion**: The tool must output a negative verdict (`contradicted`, `insufficient`, `rejected`, or exit non-zero). If it returns `supported` or `ok`, the test fails. + +--- + +## 5. The "Receipt or It Didn't Happen" Standard + +To prevent agents from hallucinating evaluation results in summaries: + +1. **Persistent Trace Files**: Every runner must write its output to disk under `outputs/` or `20_live/` (e.g. `<timestamp>-<target>-trace.json`). +2. **Mandatory Receipt Fields**: + - `run_id` / `timestamp` + - `engine_version` / `rules_version` + - `input_hash` (SHA256 of source file/packet) + - `verdicts` / `scores` (raw numbers, not summarized prose) +3. **Citation in Reports**: Any agent outputting an "Audit Table" or "Verification Scorecard" in a note or PR must include: + ```markdown + **Verification Receipt:** `outputs/cal-trials/2026-08-21-report-trace.json` (SHA256: `a1b2c3...`) + ``` + +--- + +## 6. Authoring Checklist for New Scripts & Tests + +Before committing any new script in `bin/` or `scripts/`, verify: + +- [ ] **1. Taxonomy Alignment:** Is the tool correctly named as a Transformer (`build-*`) or Runner (`run-*`/`audit-*`)? +- [ ] **2. Engine Invocation:** Does the runner actually import/execute the backend engine, or does it stop at JSON serialization? +- [ ] **3. Fail-Closed Check:** Does it exit non-zero if the engine is missing or misconfigured? +- [ ] **4. Exit Code Contract:** Does it adhere to the `0` (pass) / `1` (findings) / `2` (crash) exit code contract? +- [ ] **5. Escape Valve Support:** If this tool gates workflow or metadata, does it support valid intermediate states (`type: hypothesis`, `status: queued`) without forcing fabricated values? +- [ ] **6. Receipt Generation:** Does execution leave a deterministic, hashed trace file on disk? +- [ ] **7. Falsification Test in `tests/`:** Is there a test asserting that invalid/contradictory input produces a failure verdict? diff --git a/.context/workflows/eval-schedule.md b/.context/workflows/eval-schedule.md deleted file mode 100644 index 0962bf7..0000000 --- a/.context/workflows/eval-schedule.md +++ /dev/null @@ -1,69 +0,0 @@ -# Scheduled Evaluation Workflow - -Use this workflow to run MainFrame eval suites on a fixed cadence and feed -`20_live/eval-registry/` with trend data. - -Binding: [process-evaluation.md](process-evaluation.md), [eval-methodology.md](eval-methodology.md), lane EV01. - -## Cadence - -| Cadence | When (launchd default) | Suite | -| --- | --- | --- | -| `daily` | 06:15 every day | ingest dry-run, mindgraph dry-run, project index check, eval-registry status | -| `weekly` | 07:30 Sundays | daily suite + unittest + **4-query fused MindGraph regression probe** + single eval-registry harvest | -| `monthly` | manual | weekly suite + `workflow-report --days 7 --json` | - -Weekly/monthly runs also write an observational report to -`30_projects/mainframe-process-eval/outputs/YYYY-MM-DD-scheduled-<cadence>.md` -and re-harvest the registry. - -## Commands - -```bash -# Health (exit 1 if stale/missing launchd/failed weekly) -bin/eval-schedule check - -# Run once (operator or CI) -bin/eval-schedule run --cadence weekly -bin/eval-schedule run --cadence weekly --full-probe # 12 queries, fused+expanded (~2 min) -bin/eval-schedule run --cadence daily --dry-run - -# Install macOS launchd agents (daily + weekly) -bin/eval-schedule install --cadence both -bin/eval-schedule status -bin/eval-schedule uninstall --cadence both -``` - -Operator card (read when check fails): `20_live/eval-registry/OPERATOR.md` - -Logs: `20_live/eval-registry/logs/{daily,weekly}.{log,err}` -Manifest: `20_live/eval-registry/schedule-runs.jsonl` - -## Visibility hooks (do not skip) - -| When | Surface | -| --- | --- | -| Session start | `bin/session-open` prints eval-schedule health | -| Session end | `bin/session-close --check` warns if `bin/eval-schedule check` fails; the SessionEnd hook lands the same result in the tracker feed (ADR-040) | -| Compaction | `bin/session-close --checkpoint` (PreCompact hook) snapshots weekly-eval staleness into the draft + tracker feed | -| Handoff draft | `20_live/last-handoff-draft.md` includes eval-schedule status | -| Architecture | ADR-036 in `DECISIONS.md` | -| Research lane | EV01 `lanes/scheduled-process-evaluation/` | - -## Weekly review (human, ~15 min) - -1. `bin/eval-schedule status` — last run green? -2. `bin/eval-registry status` — new metrics/irregularities? -3. Read the latest `scheduled-weekly` output in `mainframe-process-eval/outputs/`. -4. Pick **one** improvement slice; rerun the same cadence after the fix. -5. Run `bin/mindgraph-audit-links --dry-run --scope 10_knowledge --output-dir 01_ingest/audit-receipts` when graph hygiene is in scope. Review its JSON/Markdown action queue as an advisory observation; do not turn findings into a weekly green score or blocking gate until a false-positive baseline has been measured. - -## Guardrails - -- Scheduled runs are **regression signal**, not promotion by themselves. -- MindGraph probe is skipped when `~/.mindgraph/mainframe.sqlite` is missing. -- Default weekly probe runs four fused regression queries (`q04_memory…`, `q07_ai_detection`, `q11/q12` scope negatives, ~25s). Use `--full-probe` for the complete matrix. -- Raw evidence leaves (informational) are retained as counts and classification queue items; they never fail the weekly health check. Curated-note zero-outbound, dangling/ambiguous links, freshness, and retrieval regressions remain separate signals. -- `evaluation-feedback.md` files are excluded from harvest (not eval runs). -- Do not schedule behavioral `skill-eval` cases until receipt automation exists (EV02). -- Failed steps exit non-zero so launchd logs surface breakage. diff --git a/.context/workflows/extract-knowledge.md b/.context/workflows/extract-knowledge.md deleted file mode 100644 index 4b86240..0000000 --- a/.context/workflows/extract-knowledge.md +++ /dev/null @@ -1,23 +0,0 @@ -# Extract Knowledge Workflow - -Use this workflow at project closeout or after a repeated pattern becomes clear. - -## Script -- Command: `bin/extract-knowledge` -- Check mode: `bin/extract-knowledge --project <slug> --domain <domain> --title <title> --check` -- Write mode: `bin/extract-knowledge --project <slug> --domain <domain> --title <title> --write` -- Scaffolds a note with correct metadata; knowledge content stays manual -- Run `bin/mindgraph-refresh` after filling in the note - -## Steps -1. Review the project `README.md`, `log.md`, `decisions.md`, and final outputs. -2. Identify reusable practices, failure modes, command patterns, and standards. -3. Create or update a note in `10_knowledge/` with the standard metadata schema from `.context/primitives.md`. -4. Link the note's `source` to the project path or preserved raw evidence. -5. Keep claims calibrated: separate observed facts from inferences. -6. Run `bin/mindgraph-refresh` after durable knowledge notes change. - -## Guardrails -- Extract only reusable knowledge. Do not turn every project detail into a durable note. -- Preserve source material; extracted text is a working copy, not the source of truth. -- If the note touches finance, legal, health, or other high-risk live data, verify against source documents before promotion. diff --git a/.context/workflows/lab-report.md b/.context/workflows/lab-report.md deleted file mode 100644 index 14de45a..0000000 --- a/.context/workflows/lab-report.md +++ /dev/null @@ -1,140 +0,0 @@ -# Lab report convention - -**Universal experiment tracking.** Every decision-bearing test, matrix cell family, or live probe that should be remembered gets one lab report — the same skeleton a careful experimentalist would keep in a lab notebook. - -## Profile - -| Field | Value | -|-------|-------| -| structural_type | workflow | -| owner_surface | `EVAL_METHODOLOGY.md` + project `outputs/lab-reports/` | -| related | eval-methodology, project-experiment-loop, craft-research-loop, epistemic-standard | - -## Purpose - -Stop losing experiments as unnamed logs, half-filled scorecards, or chat history. A lab report forces: - -1. **Question and decision first** (before results) -2. **One primary factor** (or an explicit multi-factor design) -3. **Raw paths** separate from interpretation -4. **Irregularities** (even "probably nothing") -5. **Disposition + next experiment** - -## Relationship to other templates - -| Artifact | Use when | -|----------|----------| -| **lab-report** (this) | Default for any measured experiment / test campaign | -| `eval-output.md` | Same content shape specialized for **eval-registry** harvest (metric YAML required) | -| `craft-trial.md` | Product/craft bake-offs with keep\|kill\|iterate (still encouraged to use lab-report sections 1–10) | -| JSON receipts | Machine-checkable cell-level evidence; **not** a substitute for the report | - -**Rule:** If you would write a methods section in a paper, write a lab report. If it is only `unittest` green in CI with no design question, skip. - -## Where reports live - -```text -30_projects/<slug>/ - outputs/ - lab-reports/ - YYYY-MM-DD-<slug>.md # one report per experiment identity - raw-materials/ - YYYY-MM-DD-<slug>/ # logs, receipts copies, scorecards -``` - -Shorthand allowed: `outputs/YYYY-MM-DD-<slug>.md` for eval-profile projects already using that path (project-experiment-loop scaffold). Prefer `outputs/lab-reports/` for new work. - -## Lifecycle - -```text -scaffold → (optional) design freeze → execute → fill results → disposition → harvest? -``` - -```bash -# Scaffold -bin/lab-report scaffold \ - --project agent-harness-eval \ - --title "coder-14b-verify-ablation" \ - --question "Does hiding run_verify increase false completions on task 07?" \ - --decision "If yes, require verify tool for multi-file coding profiles" \ - --study-type exploratory - -# After run -# edit the report; move/link raw receipts into raw-materials/<id>/ - -# Check completeness -bin/lab-report check --project agent-harness-eval --id 2026-08-08-coder-14b-verify-ablation - -# List open reports -bin/lab-report list --project agent-harness-eval --open - -# Eval-profile: also harvest when metrics are final -bin/project-experiment-loop close --project agent-harness-eval -``` - -## Required fields (minimum bar) - -A report is incomplete if any are missing: - -- [ ] `lab_report_id` / `eval_run_id` -- [ ] `study_type` -- [ ] `decision_sentence` -- [ ] `hypothesis` (or explicit "descriptive only — no hypothesis") -- [ ] primary metric + unit of analysis + n -- [ ] one primary independent factor (or stated multi-factor design) -- [ ] results table or explicit "no cells completed" -- [ ] `irregularities` section (may be empty list with statement "none observed") -- [ ] limitations / does-not-prove -- [ ] `disposition` + `next_experiment` - -## Study types - -Same vocabulary as `EVAL_METHODOLOGY.md`: - -| Type | Use | -|------|-----| -| exploratory | First look, small n, generate next protocol | -| confirmatory | Pre-registered contrast | -| regression | Frozen baseline after change | -| observational | Telemetry / before-after windows | -| calibration | Human gold / scorer agreement | - -## Dispositions - -| Value | Meaning | -|-------|---------| -| `open` | Running or unfilled | -| `accept` | Evidence supports the decision sentence's positive branch (still respect study_type limits) | -| `reject` | Evidence supports the negative branch or fails pre-registered gate | -| `hold` | Interesting but blocked (n, confound, tooling) | -| `iterate` | Close this id; open a new lab report with a sharper question | - -Do **not** use `accept` on a single exploratory run to claim production graduation. - -## One factor per phase - -Default for harness and model work (agent-harness-eval rule): change **either** model/path **or** harness factor, not both, in one `lab_report_id`. Multi-factor matrices need an explicit design section and still one **primary** metric. - -## Privacy - -| Report content | Allowed | -|----------------|---------| -| Methods, n, metrics, disposition | yes in private project | -| Public-safe restatement | only after sanitization (no absolute paths, no fixture gold) | -| Raw agent transcripts | stay in raw-materials / gitignored live trees | - -## Retrofit - -Existing scorecards and sprint rollups can be **linked** from a new lab report rather than rewritten wholesale. Minimum retrofit: - -1. Scaffold lab report with the real question/decision. -2. Point results section at existing paths. -3. Fill disposition + irregularities + does-not-prove. - -## Related - -- Template: `.context/templates/lab-report.md` -- CLI: `bin/lab-report` -- Eval contract: `EVAL_METHODOLOGY.md` -- Experiment loop: `.context/workflows/project-experiment-loop.md` -- Epistemic labels: `.context/workflows/epistemic-standard.md` diff --git a/.context/workflows/local-coder-run.md b/.context/workflows/local-coder-run.md deleted file mode 100644 index 7d2cd98..0000000 --- a/.context/workflows/local-coder-run.md +++ /dev/null @@ -1,71 +0,0 @@ -# Local Coder Run Workflow - -Use this workflow when assigning a bounded code change to the local -Aider/Ollama coder. The goal is to use the local model where judgment is -useful while keeping deterministic operations, verification, and acceptance -under explicit operator control. - -## 1. Route The Task - -- Run deterministic operations directly. Patch application, file moves, - formatting, generated indexes, and exact scripted rewrites do not need a - model unless they produce a conflict that requires judgment. -- Use the local coder for a bounded implementation or repair where the target - files and acceptance condition are already known. -- Keep one run to one coherent change. Split broad cleanup across independent - modules into separate runs. - -## 2. Preflight - -1. Inspect `git status` and preserve existing work. -2. Give the model one to three tightly coupled files when possible. -3. State the exact behavioral change, the no-commit boundary, and the - verification commands that will be run after the model exits. -4. Stop and split the task if Aider reports that estimated context exceeds the - model limit. Do not proceed with a known-overflow prompt. - -## 3. Run - -Use the configured local model with auto-commit disabled. A typical command is: - -```bash -aider --model local --no-auto-commits --yes-always \ - path/to/file.py path/to/test_file.py \ - --message "Make the bounded change. Do not commit. Do not claim tests passed." -``` - -Treat Aider output conservatively: - -- A suggested shell command is not an executed command. -- An applied edit is evidence of a change, not evidence that the change is - correct. -- An edit-format failure may leave partial edits behind. Inspect the diff - before retrying. -- A run with no model output is incomplete even if the prompt was accepted. - -## 4. Verify Outside The Model - -After Aider exits: - -1. Inspect `git diff` and `git status`. -2. Run the narrowest relevant test, type-check, or lint command directly. -3. Run the broader project verification chain when the change touches shared - behavior. -4. Revert or repair only the local coder's incorrect edits. Preserve unrelated - work already in the tree. - -The task is complete only when the external checks and human diff review pass. - -## 5. Record The Outcome - -Record: - -- task scope and files supplied; -- whether Aider reported a context warning; -- observed edits and edit-format failures; -- verification commands and results; -- remaining manual repair; -- next safe action. - -Use `bin/workflow-report --days 1` for telemetry coverage, but sample the diff -and verification output before drawing conclusions about task quality. diff --git a/.context/workflows/mindgraph-refresh.md b/.context/workflows/mindgraph-refresh.md deleted file mode 100644 index 56e412b..0000000 --- a/.context/workflows/mindgraph-refresh.md +++ /dev/null @@ -1,95 +0,0 @@ -# MindGraph Refresh Workflow - -MindGraph is a complementary retrieval layer for Mainframe, not the source of truth. - -## Defaults -- Knowledge database: `~/.mindgraph/mainframe.sqlite` -- Knowledge ingest scope: `10_knowledge/` -- Projects database: `~/.mindgraph/mainframe-projects.sqlite` -- Projects ingest scope: optional ignored local manifest at `30_projects/mindgraph-projects.json`; if absent, `bin/mindgraph-refresh-projects` uses its built-in curated list. `--full` discovers all project coordination surfaces and `MAINFRAME_MINDGRAPH_PROJECTS="slug-a slug-b"` overrides the manifest for a focused pass -- Wrapper: `bin/mindgraph` -- Refresh commands: `bin/mindgraph-refresh` (knowledge), `bin/mindgraph-refresh-projects` (projects) -- Advisory graph audit: `bin/mindgraph-audit-links --dry-run` reads source Markdown with the canonical parser/resolver and emits JSON/Markdown action-queue receipts; it does not read eval snapshots or mutate SQLite. `bin/mindgraph-refresh --audit-links` runs it after a successful knowledge refresh without making findings a gate. -- Query station: workstation interpreter over both DBs for `knowledge`, `projects`, and grouped `federated`; use explicit CLI queries for any station mode that has not shipped yet -- Operating contract: `HARNESS.md` (dual-index routing) -- MCP (ADR-049 default): shared daemon at `http://127.0.0.1:8000/mcp`. Clients use - `bin/mindgraph mcp-proxy` (stdio) or a streamable-HTTP URL. Ensure - `bin/mindgraph daemon-health` is ok before MCP work. Every MCP call selects - `scope` = `knowledge` or `projects`. Root `.mcp.json` / `.mcp.json.example` - point at the proxy; do not configure per-session `serve-mcp` for daily use. -- Binary resolution: install `mindgraph` on `PATH`, set `MINDGRAPH_BIN=/path/to/mindgraph`, or set local repo config `git config mainframe.mindgraphBin /path/to/mindgraph` - -## Steps -1. Add or update durable Markdown notes in `10_knowledge/`. -2. Run `bin/mindgraph-refresh`. - When link hygiene is in scope, use `bin/mindgraph-refresh --audit-links --audit-output-dir <receipt-dir>`; review its advisory receipt separately from index freshness. -3. Add or update project-layer notes (README, `outputs/`, `plans/`, `decisions.md`) in `30_projects/`. -4. Run `bin/mindgraph-refresh-projects`; it generates one temporary manifest and calls `mindgraph ingest-many` so repeated filenames across projects do not collide. -5. Query with the workstation Query Station when available, or with `bin/mindgraph query "<question>"` for the knowledge DB and `MINDGRAPH_DB_PATH="$HOME/.mindgraph/mainframe-projects.sqlite" bin/mindgraph query "<question>"` for the projects DB. Preserve both result groups with trust labels. MCP clients use the shared daemon via `.mcp.json` (`mcp-proxy`) with explicit `scope` per call. - -## Query Pass Template - -**Canonical copy-paste:** [`.context/templates/mindgraph-query-pass.md`](../templates/mindgraph-query-pass.md) -Also embedded in [`.context/templates/task-packet.md`](../templates/task-packet.md) as an optional section. - -Minimal YAML-style block for plans and handoffs: - -```markdown -## MindGraph Query Pass -intent: -doctor: # overall line from `bin/mindgraph doctor` -knowledge_query: -projects_query: -knowledge_nominations: -- title: ... - path: ... - reason: ... -project_nominations: -- title: ... - path: ... - reason: ... -weak_or_excluded: -- ... -source_inspection_required: -- ... -``` - -## Guardrails -- Do not ingest the full vault by default; operating contracts and empty indexes add retrieval noise. -- Do not merge `mainframe.sqlite` and `mainframe-projects.sqlite` into one blended ranking. The station/interpreter may group, annotate, and bridge nominations, but the DBs stay physically and epistemically separate. -- Project DB refresh uses multi-root namespaced ingest; still inspect source files before treating station output as complete project context. -- MindGraph nominations are not verification. Treat returned chunks as candidates to inspect. -- Graph audit findings are also advisory nominations. Inspect source notes before any relationship edit; raw evidence leaves remain an informational queue, not a health failure. -- Keep the SQLite database outside the repo so Git history stays clean. -- Refresh does not hot-reload an already running shared daemon. Stop it before - refresh and restart it afterward when fresh file handles are required. PID - and log state are separate from both lifecycle indexes. - -## First-Time Agent / Troubleshooting (added 2026-06-23) - -When an agent or operator is starting fresh or after a long gap, the mandatory dual MindGraph planning hook can become an environment archaeology exercise. The following checklist and known issues reduce that friction. - -### Quick first-run checklist -1. `bin/mindgraph doctor` — confirm dual `~/.mindgraph` indexes (not workspace stubs) -2. Official CLI: `bin/mindgraph query "..." --json --db ~/.mindgraph/mainframe.sqlite` -3. Projects: `bin/mindgraph query "..." --json --db ~/.mindgraph/mainframe-projects.sqlite` -4. Always run **both** and keep groups labeled (`durable_knowledge` vs `project_status`) -5. Emit a MindGraph Query Pass (template above / `.context/templates/mindgraph-query-pass.md`) -6. Use `--json` for agent processing - -### Known friction points (from 2026-06-23 session) -- Workspace root `mainframe*.sqlite` files are usually 4 KB stubs with no tables. The real indexes live in `~/.mindgraph/`. Doctor warns; query fail-fasts if you hit a stub. -- `bin/mindgraph query` reloads the embedding model on every cold process. Plan for latency on discovery passes (or use MCP warm path when available). -- `python` may not be in PATH (use `python3`). -- Project-internal docs often live one level deeper (e.g. `workbench/docs/ENGINE_NOTES.md`). -- Early broad queries return weak or tangential hits. Use exact project slugs, phase names, and file stems once you have them from the lane tracker or prior notes. - -### Recommended mitigations -- Start research or project planning with doctor + dual query + Query Pass record. -- Record new friction in `10_knowledge/agents/` + MH01 lane (`mindgraph-agent-harness-discovery`). -- **Shipped (2026-07-13):** `bin/mindgraph doctor|status`; query fail-fast on missing tables. -- Shipped as opt-in: loopback shared MCP with explicit scope/trust. Operational - activation and latency-budget measurement remain open. -- Keep this section and the companion friction note up to date. - -See also the dedicated research lane `30_projects/research-lanes-strategy/lanes/mindgraph-agent-harness-discovery/README.md`. MindGraph is the retrieval engine of the system; making first contact reliable is high-leverage. diff --git a/.context/workflows/process-evaluation.md b/.context/workflows/process-evaluation.md new file mode 100644 index 0000000..ecab08f --- /dev/null +++ b/.context/workflows/process-evaluation.md @@ -0,0 +1,121 @@ +# MainFrame Process Evaluation Workflow + +Use this workflow to evaluate a MainFrame process before changing it. The goal +is to measure current behavior, identify the smallest useful improvement, and +rerun the same evaluation after the change. + +Eval-profile binding: [EVAL_METHODOLOGY.md](../../EVAL_METHODOLOGY.md). The +longer eval-methodology workflow is an optional/deferred component and is +not required for this public-core slice. Process evals use `study_type: +observational`. Every `outputs/YYYY-MM-DD-evaluation.md` requires metric extract, +`irregularities`, and `bin/eval-registry harvest`. + +## CLI + +```bash +bin/process-eval preflight # standard non-mutating check pack +bin/process-eval status # project + latest output pointers +bin/process-eval close [--write] # loop-decision scaffold +``` + +Owner operation: `40_operations/mainframe-process-eval/`. +Loop catalogue: `40_operations/mainframe-process-eval/workbench/loop-eval/`. +Root command: `bin/process-eval` (shim; implementation is operation-owned). + +## Loop unit + +| Field | Rule | +|-------|------| +| Unit | One observational pass: question → baseline → checks → samples → ≤2 slices → rerun → output | +| Entry | Cadence (≈1 week / 5 sessions), post-change verification, or operator request | +| Exit | Dated output harvested; loop decision recorded (`bin/process-eval close`) | + +## Evaluation Loop + +1. Write one evaluation question. + - Good: "Does the ingest path move supported captures forward without + losing provenance or creating avoidable manual work?" + - Avoid: "Is MainFrame good?" +2. Capture a baseline before editing files. +3. Run the relevant deterministic checks: + + ```bash + python3 -m unittest discover -s tests + bin/ingest-minion run --dry-run + bin/mindgraph-refresh --dry-run + bin/sync-project-index --check + bin/session-open --json + bin/session-close --check + bin/eval-schedule check + bin/workflow-report --days 7 --json + ``` + +Scheduled regression (ADR-036): `bin/eval-schedule run --cadence weekly` and +`20_live/eval-registry/OPERATOR.md` for the standing review ritual. + +4. Sample real outcomes. For each evaluated workflow, inspect at least three + representative cases when available: + - a normal case; + - a boundary or ambiguous case; + - a known failure or high-friction case. +5. Score the process on: + - correctness and safety; + - provenance and reversibility; + - throughput and backlog; + - operator friction; + - observability; + - handoff and reentry quality. +6. Classify each finding: + - `code-defect`: implementation does not match the contract; + - `process-gap`: the contract lacks a needed step or boundary; + - `adoption-gap`: a sound workflow exists but is not being used; + - `telemetry-gap`: current signals cannot support the conclusion; + - `intentional-backlog`: queued work is expected and should not be treated + as a defect. +7. Select one or two improvement slices. Prefer changes that make future + evaluation easier or remove repeated friction without weakening safety. +8. Apply the change, then rerun the same checks and samples + (`bin/process-eval preflight`). +9. Save the baseline, change, rerun result, and next action in a dated project + output. Record accepted architecture or workflow changes in `DECISIONS.md`. +10. **Close the pass** with an explicit loop decision: + ```bash + bin/process-eval close --run-id YYYY-MM-DD-… --question "…" --write + ``` + Choices: `same_slice` | `new_question` | `promote_pattern` | `park` | `handoff`. +11. Harvest: `bin/eval-registry harvest` (project mainframe-process-eval). + +## Action surface + +- Dated files under `local-only: 40_operations/mainframe-process-eval/outputs/` +- `improvement-backlog/items.md` for nominations +- Optional loop-eval catalogue re-score when process taxonomy changes +- `bin/mainframe-doctor` and `bin/eval-schedule check` remain adjacent health tools + +## Promotion Test + +When a repeated pattern appears during evaluation, place it at the right layer: + +| Pattern | Destination | +| --- | --- | +| Fixed, deterministic operation | `bin/` script with tests | +| Operator-driven sequence using existing tools | `.context/workflows/` | +| Repeated agent judgment, domain rules, or tool strategy | `.agents/skills/` | +| Role with its own tools, guardrails, and procedure | `agents/` subagent | +| One-project experiment or uncertain practice | `30_projects/<slug>/` | + +Prefer improving an existing workflow or skill over creating a near-duplicate. +Promote a new skill only when there are concrete trigger examples, repeated +judgment that general models would otherwise rediscover, and a way to evaluate +the skill against representative outputs. + +## Guardrails + +- Tool-call success is not task success. +- Speed is not the only quality measure. +- Do not treat inbox or queue size alone as a failure; separate expected + migration backlog from stuck work. +- Keep telemetry metadata-only. Do not add prompts, file contents, command + output, or model responses to workflow logs. +- Preserve raw evidence and baseline artifacts before changing the process. +- Do not tune against one convenient example. Keep boundary and failure cases. diff --git a/.context/workflows/project-experiment-loop.md b/.context/workflows/project-experiment-loop.md deleted file mode 100644 index 20d5036..0000000 --- a/.context/workflows/project-experiment-loop.md +++ /dev/null @@ -1,195 +0,0 @@ -# Project Experiment Loop - -**One measured experiment pass.** Complements the research-lane-loop (literature → knowledge). -This loop owns **testing and results**: design → run → output → registry → **action**. - -## Profile - -| Field | Value | -|-------|-------| -| structural_type | workflow | -| owner_surface | `EVAL_METHODOLOGY.md` + eval-profile projects | -| related | eval-methodology, eval-schedule, eval-registry, research-lane-loop | - -## Purpose - -Make every decision-bearing test: - -1. **Identifiable** (`eval_run_id`) -2. **Protocol-pinned** (`protocol_ref`) -3. **Harvested** into `20_live/eval-registry/` -4. **Actionable** (triage card with one next step — not a silent green log) - -```text -preflight → design gate → execute → write output → harvest → triage/action → loop? -``` - -## When to use - -- Eval-profile projects (`*-eval`, `scaffold-claims-study`, `eval-profile` tag) -- MindGraph canaries, harness matrices, process baselines -- Any measurement that should change a promotion, architecture, or process decision - -## When not to use - -- Pure literature capture → `research-lane-loop` -- Single unit test green with no report (keep as CI) -- One-off task-packet verification (local gate only) - -## Separation - -| Need | Loop | -|------|------| -| External knowledge | research-lane-loop → `10_knowledge/` | -| Measured result | **this loop** → `outputs/` + registry | -| Project gap routing | PRP01 project-research-pipeline | - -## Command card - -```bash -# 0) Health -bin/mindgraph doctor -bin/eval-schedule check -bin/project-experiment-loop preflight --project mindgraph-eval - -# 1) Fresh MindGraph regression canary (probe + envelope + doctor + harvest + triage) -bin/project-experiment-loop canary --fresh - -# 2) Or scaffold a custom experiment output, run manually, then close -bin/project-experiment-loop scaffold \ - --project mindgraph-eval \ - --study-type exploratory \ - --title "my-slice" \ - --decision "If X, do Y" -# ... run experiment, fill outputs/ ... -bin/project-experiment-loop close --project mindgraph-eval - -# 3) Action ALL MainFrame evals (default triage = portfolio) -bin/project-experiment-loop portfolio -# same as: bin/project-experiment-loop triage --scope portfolio -cat 20_live/eval-registry/last-eval-action.md - -# MindGraph-only action card (optional) -bin/project-experiment-loop triage --scope canary -``` - -## Steps - -### 0. Preflight - -- Read `30_projects/<slug>/methodology-approach.md` -- `bin/mindgraph doctor` when retrieval is in scope -- `bin/eval-registry status` / `bin/eval-schedule check` - -### 1. Design gate - -- [ ] `decision_sentence` -- [ ] `study_type` -- [ ] `eval_run_id` (`YYYY-MM-DD-<slug>` or timestamped) -- [ ] `protocol_ref` pinned -- [ ] Primary metric + unit of analysis -- [ ] Irregularity watch list - -### 2. Execute - -- Raw under `raw-materials/<eval_run_id>/` when applicable -- Do not change protocol mid-run without a new `eval_run_id` - -### 3. Write output - -Use the **lab report** convention (`.context/templates/lab-report.md`, workflow `lab-report.md`) so every experiment has researcher-grade tracking. - -- Prefer: `bin/lab-report scaffold …` → `outputs/lab-reports/<id>.md` -- Eval-registry specialization: `.context/templates/eval-output.md` (also emitted by `project-experiment-loop scaffold`) -- Minimum bar: `bin/lab-report check --project <slug> --id <id>` - -### 4. Harvest - -```bash -bin/eval-registry harvest -bin/eval-registry check --strict # eval-profile -``` - -### 5. Triage / action (non-optional for canaries) - -Running is not enough. After harvest: - -1. Compare to prior run (pass/fail, metric deltas) -2. Write `20_live/eval-registry/last-canary-action.md` -3. Set **one** next action (investigate / no-op keep cadence / open improvement slice) -4. Optional: update project `next_action` when severity ≥ medium - -`bin/project-experiment-loop triage` and `canary --fresh` do this automatically. - -### 6. Loop decision - -| Outcome | Next | -|---------|------| -| Regression green, no high irregularities | Keep cadence; no product change | -| Regression red or unprotected scope rise | Investigate before ingest/harness changes | -| Exploratory interesting | Design confirmatory follow-up with new `eval_run_id` | -| Knowledge gap | Emit research-lane candidate → research-lane-loop | - -## Automation model (all MainFrame evals) - -| Layer | What runs | What was missing | -|-------|-----------|------------------| -| **Schedule** | weekly: process suite + MindGraph canaries + harvest | Already automated | -| **Registry** | `runs.jsonl` / `metrics.jsonl` / `irregularities.jsonl` | Already automated | -| **Action** | Human reads OPERATOR.md | **Often skipped** | - -**Right idea:** keep execution scheduled; automate **portfolio triage**, not more silent metrics. - -| Card | Path | Scope | -|------|------|--------| -| **Portfolio (primary)** | `20_live/eval-registry/last-eval-action.md` | MindGraph + process suite + every eval-profile project | -| Canary pointer | `last-canary-action.md` | Points at portfolio card | - -- After weekly suite: `triage --scope portfolio` -- On session-open: surface `last-eval-action.md` -- `canary --fresh` for MindGraph package; `portfolio` anytime for action refresh - -Eval-profile projects covered: `mindgraph-eval`, `mainframe-process-eval`, `agent-harness-eval`, `agent-tracker-eval`, `scaffold-claims-study`, `skill-eval-workshop`, plus any `tags: [eval-profile]`. - -## Related - -- `EVAL_METHODOLOGY.md`, `.context/workflows/eval-methodology.md` -- `.context/workflows/eval-schedule.md`, `process-evaluation.md` -- `.context/workflows/research-lane-loop.md` -- `bin/project-experiment-loop`, `bin/eval-registry`, `bin/eval-schedule` -- Operator card: `20_live/eval-registry/OPERATOR.md` - -## Dogfood — batch20 process gaps (2026-07-13) - -### What worked -- `canary --fresh` end-to-end green (probe failures 0) -- Portfolio triage gives a single operator card for all eval-profile projects -- Harvest + check --strict stay green when outputs are registry-shaped - -### Friction → improve -| Issue | Fix direction | -|-------|----------------| -| Scaffold harvested with placeholder metrics | Close must re-harvest or gate placeholders | -| High irregularities stay open after product disposition | Irregularity lifecycle: waived / accepted_risk / superseded | -| Operator-gated studies inflate severity | `operator_gated` flag in registry / triage sections | -| Canary log buffering when backgrounded | Unbuffered progress / JSONL | - -Full note: `30_projects/mainframe-process-eval/outputs/2026-07-13-three-loop-process-gaps.md`. - -## Irregularity lifecycle (G1) - -```bash -# Accept historical risk so portfolio severity tracks actionable work only -bin/eval-registry dispose \ - --project agent-harness-eval \ - --run-id <run> \ - --id <irregularity_id> \ - --status accepted_risk|waived|superseded \ - --reason "..." - -bin/eval-registry list-open-high -bin/project-experiment-loop portfolio -``` - -Statuses: `open` | `accepted_risk` | `waived` | `superseded` | `resolved`. -Store: `20_live/eval-registry/irregularity-dispositions.jsonl`. diff --git a/.context/workflows/research-lane-intake.md b/.context/workflows/research-lane-intake.md deleted file mode 100644 index ff93e52..0000000 --- a/.context/workflows/research-lane-intake.md +++ /dev/null @@ -1,142 +0,0 @@ -# Research Lane Intake Workflow - -Use this workflow when a new research question surfaces that merits its own bounded lane in the portfolio tracker. - -## Purpose -Turn recurring, decision-impacting, or cross-project open questions into **lane trackers** (`lanes/<slug>/README.md`) with minimal friction, while preserving all rules: -- Tracker only (no knowledge content here). -- All durable captures still route `00_inbox/` → ingest-minion → `10_knowledge/`. -- First-principles convention. -- Dual MindGraph query pass before committing a lane. -- WIP discipline and operator confirmation. - -This is the **explicit "add new questions to research lanes" step**. - -## When to use (the trigger step) -- During or after a source-literature run, new side questions appear that deserve tracking. -- PRP01 / project sweep identifies a gap that is not covered by an existing lane. -- Agent runs, eval outputs, or MindGraph queries (esp. repeated weak_fit or scope warnings) reveal a missing foundation or specialization area. -- Operator session surfaces a high-stakes open question. -- Synthesis notes or plans name a reusable research strand. - -**Do not** create a lane for one-off curiosity or transient lookup. - -## Steps -1. **Emit the candidate** (anywhere) - Add a block like this in the source artifact (run note, project plan, synthesis, log, agent output): - - ```markdown - ## Research Lane Candidate - - - **lane_id**: C32 # optional; script suggests next - - **slug**: intent-graphs-rag-routing - - **central_question**: "How can intent and goal graphs drive routing and planning across specialized retrievers?" - - **decision**: "Decide whether and how to make intent graphs first-class in MainFrame retrieval and agent harnesses." - - **knowledge_domain**: "graph-memory" - - **stakes**: high - - **priority**: next - - **trigger**: "surfaced during multi-RAG architecture planning in mindgraph-eval + C28/C29 work" - - **handoff**: "30_projects/mindgraph-eval, 30_projects/mindgraph" - ``` - - (Scan-friendly: any file containing one or more such blocks can be fed to `bin/lane-intake scan`.) - -2. **Dual MindGraph query pass** (mandatory) - ```bash - bin/mindgraph query "<question keywords>" --db ~/.mindgraph/mainframe.sqlite --json --top-k 6 - bin/mindgraph query "<question keywords>" --db ~/.mindgraph/mainframe-projects.sqlite --json --top-k 6 - ``` - Record nominations + trust labels. Treat as nominations only. - -3. **Run intake** - ```bash - # Propose / dry-run (recommended first) - bin/lane-intake scaffold \ - --question "..." \ - --decision "..." \ - --domain "graph-memory" \ - --trigger "path or description" \ - --slug "my-topic" \ - --priority next - - # Or parse emitters from a file - bin/lane-intake scan 30_projects/research-lanes-strategy/plans/some-plan.md - - # Apply (creates folder + README + receipt + log append) - bin/lane-intake scaffold ... --apply - ``` - -4. **Review outputs** - - New `lanes/<slug>/README.md` (populated from template). - - `raw-materials/YYYY-MM-DD__proposed-lane-<slug>.md` (receipt with context, MindGraph results, suggested master-plan row). - - Append in `log.md`. - -5. **Commit to master plan** - Copy the suggested table row from the receipt into the correct tier/section of `plans/research-lanes-master-plan.md`. - Update any "active trackers" lists if appropriate. - -6. **Handoff** - - Freeze the plan-freezing checklist on the lane card (`first-principles-research-conventions.md`). - - Run research with `.context/workflows/research-lane-loop.md` (source-literature → synthesis → close), not capture-only. - - Use the lane README as the brief; update the capture index only as part of loop step 6. - -## Guardrails -- Script refuses to overwrite existing slugs. -- Never auto-runs source-literature or ingests. -- New domains still require confirmation per ingest rules. -- Respect WIP cap (document in receipt). -- Always preserve provenance (trigger + dual-query output) in the receipt. -- Lane creation is a **portfolio decision**, not an automatic side effect of every question. - -## Integration points (where the step is called) -- Source-literature workflow (post run note): scan for candidates and propose. -- Project research pipeline (PRP01): after gap sweep, promote to intake. -- Session-close / process-evaluation: harvest open questions. -- Agent instructions: when you surface a recurring decision question, emit a candidate block and/or run `bin/lane-intake scan`. - -## Related -- `30_projects/research-lanes-strategy/plans/research-lanes-master-plan.md` -- `30_projects/research-lanes-strategy/plans/project-research-pipeline.md` -- `30_projects/research-lanes-strategy/plans/knowledge-routing.md` -- `30_projects/research-lanes-strategy/lanes/_TEMPLATE.md` -- `.context/workflows/source-literature.md` -- `bin/mindgraph` (strict CLI only) -- `.agents/skills/mindgraph-retrieval/SKILL.md` - -## Quick usage (operator or agent) - -After any research activity surfaces a question: - -1. Append a candidate block to the relevant note or plan. -2. `bin/lane-intake scan that-file.md` # review output + receipt -3. `bin/lane-intake scan that-file.md --apply` -4. Paste the suggested row into master-plan.md under the right tier. -5. Dual-query evidence and receipt already captured for provenance. - -Example command that created C32: -``` -bin/lane-intake scaffold \ - --question "..." --decision "..." --domain knowledge-systems \ - --trigger "..." --slug research-lane-intake --lane-id C32 --apply -``` - -## Claim discipline - -Lane intake produces claim-bearing output, so -[epistemic-standard.md](epistemic-standard.md) and `EPISTEMIC_STANCE.md` bind -every capture and synthesis this workflow creates. - -- **No source quotas, ever.** Do not set or infer a target count of sources for a - lane. Ask a coverage question instead: is the decision answerable with what we - found? A documented gap is a valid, and often better, lane outcome. -- Captures route through `01_ingest/`, where `bin/capture-validate` checks - provenance. Never write a `type: raw` file into `10_knowledge/` directly. -- **Escape:** if a lane cannot be answered from real sources, close it as - `blocked` with the gap named. That is a complete result, not a failed one. - -**Stop state.** A lane that cannot find real evidence stops and says so. It does -not fill to a batch size. - -Background: a source count in `bin/research-lane-loop` produced 107 captures -citing papers that do not exist -(`20_live/security/2026-08-09__fabricated-source-captures-in-10-knowledge.md`). diff --git a/.context/workflows/research-lane-loop.md b/.context/workflows/research-lane-loop.md deleted file mode 100644 index 2c76390..0000000 --- a/.context/workflows/research-lane-loop.md +++ /dev/null @@ -1,359 +0,0 @@ -# Research Lane Loop - -**One complete research pass.** Starts at source-literature and ends at knowledge synthesis. Designed to be easy to run and safe to loop. - -## Profile - -| Field | Value | -|-------|-------| -| structural_type | workflow | -| owner_surface | `30_projects/research-lanes-strategy/` + `.context/workflows/` | -| authority | workflow-contract | -| volatility | stable | -| related_surfaces | source-literature, ingest-minion, epistemic-standard, research-lane-intake | - -## Purpose - -Replace ad-hoc “capture then maybe synthesize later” runs with a single unit of work: - -```text -select → brief → preflight → source-literature → ingest → synthesize → close → loop? -``` - -Tracker stays a link index. Durable claims live only in `10_knowledge/`. - -## When to use - -- Completing research for one lane phase (foundation, taxonomy, specialization, or application/eval). -- Weekly portfolio cadence: pick a P0/P1 lane and finish one pass end-to-end. -- Operator says “run the research loop”, “full loop”, or “source to synthesis” for a lane. - -## When not to use - -- **New lane only** → `.context/workflows/research-lane-intake.md` (then come back here). -- **File already in hand** → drop in `00_inbox/` and start at **Ingest** (step 4). -- **Project closeout extraction** → `.context/workflows/extract-knowledge.md`. -- **Portfolio governance only** (priority, archive, kill) → `bin/lane-intake` + master plan; no sources needed. - -## Loop unit (what one iteration is) - -| Field | Rule | -|-------|------| -| **Scope** | Exactly **one** lane × **one** phase | -| **Batch size** | **3–4** accepted sources (plus 1 failure/counterevidence when stakes are high) | -| **Start gate** | Lane README has decision, `knowledge_domain`, `capture_tags`, and a stop condition | -| **End gate** | Synthesis note written (or explicit skip logged) **and** tracker capture index updated | -| **Default phases** | `foundation` → `taxonomy` → `specialization` → `application` (see first-principles conventions) | - -One iteration is **not** “finish the whole lane.” Multi-phase lanes re-enter this loop once per phase. - -## Command card (copy/paste) - -```bash -# 0) Select + classify resume (preferred entry) -bin/research-lane-loop doctor # top P0 active lane + class A–D -# or: bin/research-lane-loop preflight --slug <slug> -# or: bin/lane-intake list --priority P0 --status active - -# 0b) Portfolio hygiene — stale capture indexes (class D) -bin/research-lane-loop audit-indexes --status active -# bin/research-lane-loop audit-indexes --status active --repair-index - -# 2) Preflight MindGraph (replace QUESTION / DOMAIN) -bin/mindgraph query "QUESTION" --db ~/.mindgraph/mainframe.sqlite --json --top-k 6 -bin/mindgraph query "QUESTION" --db ~/.mindgraph/mainframe-projects.sqlite --json --top-k 6 - -# 4) Ingest after inbox stubs exist -bin/ingest-minion run --dry-run -bin/ingest-minion run --apply -# If files landed in 01_ingest/ready/: ingest-agent → bin/prep-ingest run --apply → ingest-minion again -export UNPAYWALL_EMAIL="${UNPAYWALL_EMAIL:-you@example.com}" -bin/post-route-enrich --subset DOMAIN - -# 5) Optional synthesis candidates -bin/suggest-synthesis --domain DOMAIN --dry-run - -# After synthesis note written -bin/mindgraph-refresh -``` - -Agent judgment steps (1 brief, 3 source-literature, 5 synthesize, 6 close) use the skills listed under **Related**. -**Deterministic preflight:** `bin/research-lane-loop` (tests: `tests/test_research_lane_loop.py`, `tests/test_lane_intake.py`). - ---- - -## Steps - -### 0. Select lane + phase - -```bash -bin/research-lane-loop doctor -# or explicit: -bin/research-lane-loop preflight --slug <slug> -bin/lane-intake list --priority P0 --status active -``` - -Record from preflight output (do not guess): - -- `lane_id`, slug, path to `lanes/<slug>/README.md` -- **Phase** for this pass: `foundation` | `taxonomy` | `specialization` | `application` -- Whether the lane’s **stop condition** can be met this pass -- **Resume class** A–D and **recommended step** from `bin/research-lane-loop` - -| Resume class | Signal | Enter loop at | -|--------------|--------|---------------| -| **A Fresh** | No inbox stubs; few/no knowledge raws | Step 1 → 2 → 3 | -| **B Partial batch** | Matching stubs in `00_inbox/` | Step 4 (ingest) then 5 | -| **C Ingested** | Lane-tagged raws/notes in `10_knowledge/` | Step 5 (or next phase if stop met) | -| **D Stale tracker** | Capture index lists `00_inbox/` for files already in knowledge | `audit-indexes --repair-index`, then re-preflight | - -**Dogfood note (2026-07-13):** P0 list was empty until priority aliases (`immediate`≡P0). Class D false WIP showed on C28 until `audit-indexes --repair-index`. Prefer `bin/research-lane-loop doctor` over hand-classifying. - -If the plan-freezing checklist on the card is empty, fill it first (first-principles conventions). Do not search externally until step 2 is complete (unless class B/C/D). - -### 1. Brief (read-only) - -From the lane README, copy as the run brief (do not re-invent): - -| Field | Source | -|-------|--------| -| Central question / decision | README body | -| `knowledge_domain` | frontmatter | -| `capture_tags` | frontmatter | -| Stakes | frontmatter | -| Stop condition | first-principles section | -| Phase tag | chosen in step 0 | - -Every capture in this pass must include: - -```yaml -tags: ["research-lane", "lane-<id>", "<phase>", "needs-audit", "<topic>"] -``` - -Example: `["research-lane", "lane-c28", "application", "needs-audit", "agent-memory"]`. - -### 2. Preflight dedup - -Before any external search: - -1. Dual MindGraph query (knowledge + projects) with the lane question keywords. -2. Grep / list `10_knowledge/<knowledge_domain>/` for existing raws and notes. -3. Check domain source catalogs if present. -4. Note gaps and **reuse hits** (do not re-capture). - -Write gap notes into the eventual source-literature run note. Treat MindGraph rows as **nominations only**. - -### 3. Source-literature (**loop start** — skip if resume class B/C) - -Follow `.context/workflows/source-literature.md` and `.agents/skills/source-literature/SKILL.md`: - -1. Frame PICO from the lane brief (step 1). -2. Search, tier, and filter (credibility tiers). -3. Confirm candidates when >3 sources, any Tier C/E, or new domain. -4. Write **3–4** `type: raw` stubs to `00_inbox/`. -5. Write run note: `00_inbox/YYYY-MM-DD__source-literature-run__<lane-or-topic-slug>.md`. -6. **Atomic batch rule:** every row marked ACCEPT in the run note must have a real file on disk before handoff. If a stub fails to materialize, either write it now or demote to REJECT/`deferred-captures-backlog` — never leave “accepted” orphans (MH01 Li survey was orphaned 19 days). - -**Batch shape (default):** - -| Slot | Count | Purpose | -|------|-------|---------| -| Foundation / phase-primary | 2–3 | Definitions, mechanisms, or phase-specific core | -| Taxonomy or standard | 0–1 | Map the domain (taxonomy phase; optional otherwise) | -| Counterevidence / failure | 0–1 | Required when stakes = high | -| Official/current docs | as needed | Tools, regulators, APIs (date-stamped) | - -Do **not** synthesize in stubs. Stop at inbox + run note. - -### 4. Ingest pipeline - -```bash -bin/ingest-minion run --dry-run -bin/ingest-minion run --apply -``` - -If anything is in `01_ingest/ready/`: - -1. Run ingest-agent enrichment (`agents/ingest-agent.md`). -2. `bin/prep-ingest run --apply` -3. `bin/ingest-minion run --apply` again - -Then enrich full text and refresh search: - -```bash -bin/post-route-enrich --subset <knowledge_domain> -``` - -Optional for high-stakes: `bin/audit-sweep --apply --subset <knowledge_domain>`. - -### 5. Knowledge synthesis (**loop end**) - -**Write a synthesis note** in the knowledge domain when **any** of these is true: - -| Trigger | Action | -|---------|--------| -| ≥3 related raws share this `lane-*` tag (this pass or cumulative for the phase) | Write / update phase synthesis note | -| Phase stop condition is met with current evidence | Write synthesis that answers the decision | -| `bin/suggest-synthesis --domain <domain>` flags this cluster | Prefer writing now over parking | - -**Skip synthesis only if** sources were pure gap-fill under an existing current note **and** the stop condition is already satisfied. Log the skip in `log.md` with a link to the existing note. - -**Note path:** - -```text -10_knowledge/<domain>/YYYY-MM-DD__<domain>__note__<lane-or-phase-slug>.md -``` - -**Required note shape:** - -1. Decision this note supports (one sentence from lane card). -2. Claims labeled per `.context/workflows/epistemic-standard.md` (observation / source-claim / inference / hypothesis + confidence). -3. Links to the raw stubs consumed (`links:` + body citations). -4. **Adverse Findings & Limitations** (boundary conditions, literature gaps, compensating controls) — required by first-principles conventions. -5. Same `lane-*` + phase tags as captures. -6. Explicit **stop-condition check**: met / not met / deferred (with reason). - -Then: - -```bash -bin/mindgraph-refresh -``` - -### 6. Tracker close (mandatory) - -Do not leave the loop mid-air: - -1. **Lane README** — append rows to **Captured knowledge** (paths + type + status only; no claim copy). -2. **Phase status** — mark phase done / in progress / blocked. -3. **Project `log.md`** — one entry: lane, phase, N sources, synthesis path, stop-condition result. -4. **Project `README.md` `next_action`** — what the next loop should be. -5. **Side questions** — emit `## Research Lane Candidate` blocks; `bin/lane-intake scan <run-note>` if any deserve new lanes. -6. **Deferred sources** — append to `plans/deferred-captures-backlog.md` (do not stall the loop). - -### 7. Loop decision (close the *pass*, not necessarily the *lane*) - -| Outcome | Next | -|---------|------| -| Phase stop condition **met**; more phases remain | Same lane, **next phase** → restart at step 0 | -| Phase incomplete (need more sources) | Same lane + phase → step 2/3 | -| Blocked | Log + deferred backlog; pick another P0 lane | -| New high-priority gap | `research-lane-intake`, then loop that lane | -| **Lane** stop condition **met** | Enter **lane close** (below) — do **not** auto-archive mid-pass | - -#### Lane close (separate closing micro-loop — optional after step 7) - -Archiving is **not** part of every research pass. It is a portfolio decision once the **lane** stop condition is met: - -```text -stop check → handoff block on lane README → bin/lane-intake archive <slug> # dry-run - → bin/lane-intake archive <slug> --apply → portfolio README/log -``` - -**Dogfood (2026-07-13):** C28 closed this way → `lanes/completed/cognitive-agent-memory-architectures/`. - -| Keep archive separate because | If you fold archive into every loop | -|-------------------------------|-------------------------------------| -| Most passes only finish a *phase* | Agents will archive too early | -| Archive moves cards to `lanes/completed/` | Loses WIP signal mid-research | -| Needs operator confirmation for portfolio | Silent archive is hard to reverse | - -**Rule of thumb:** research-lane-loop ends at **synthesis + tracker close + decision**. Archive is the **closing loop** for a finished lane (step 7 terminal branch), not a default step 0–6 action. - -**Looping is intentional.** Prefer many short end-to-end passes over one open-ended search that never synthesizes. - ---- - -## Resume map (if interrupted) - -| Last completed step | Resume at | -|---------------------|-----------| -| Stubs in `00_inbox/` only | Step 4 (ingest) | -| Ingested, no synthesis | Step 5 | -| Synthesis written, tracker stale | Step 6 | -| Tracker closed, stop not met | Step 7 → new iteration step 2 or 3 | -| Run note ACCEPT list longer than files on disk | Write missing stubs or demote to deferred; then step 4 | -| Capture index shows `00_inbox/` but files live in `10_knowledge/` | Step 6 repair only (class D) | - -Do not re-run source-literature for the same DOIs/titles already in the vault. - -## Process metrics (log after each run) - -Record in project `log.md` one line each: - -- **select_friction:** did P0 list work, or need fallback? -- **resume_class:** A/B/C/D -- **sources_accepted / sources_ingested** -- **synthesis:** path or skip -- **minutes_approx:** rough wall time (optional) -- **process_fix:** one sentence if workflow/tooling should change - -## Definition of done (one iteration) - -- [ ] 3–4 raws (or justified smaller batch) routed to `10_knowledge/<domain>/` -- [ ] Run note preserved (inbox or routed with stubs) -- [ ] Synthesis note written **or** explicit skip logged with existing note path -- [ ] `bin/mindgraph-refresh` after durable note change -- [ ] Lane capture index + phase status updated -- [ ] `log.md` line + `next_action` set -- [ ] Loop decision recorded (next phase / next lane / archive / blocked) - -## Anti-patterns - -- Stopping after inbox capture (“we’ll synthesize later”) without a dated next_action -- Synthesizing inside `lanes/` or `plans/` -- Re-capturing catalog or vault hits -- Skipping dual MindGraph preflight -- Mixing two phases in one batch without tagging -- Treating MindGraph hits as verified claims -- Leaving paywalled full-text as a hard stop (stub + abstract is enough to continue; flag `full-text-pending`) - -## Related - -| Piece | Role | -|-------|------| -| `.context/workflows/source-literature.md` | Step 3 detail | -| `.agents/skills/source-literature/SKILL.md` | Discovery judgment | -| `.context/workflows/ingest-minion.md` | Step 4 detail | -| `.context/workflows/epistemic-standard.md` | Step 5 claim discipline | -| `.context/workflows/research-lane-intake.md` | New lanes from side questions | -| `30_projects/research-lanes-strategy/plans/knowledge-routing.md` | Tracker vs knowledge boundary | -| `30_projects/research-lanes-strategy/plans/first-principles-research-conventions.md` | Phase model + adverse findings | -| `30_projects/research-lanes-strategy/plans/lane-ingest-process-notes.md` | Retrospectives / friction | -| `.agents/skills/research-lane-loop/SKILL.md` | Agent orchestration of this loop | -| `bin/lane-intake` | Select, priority, archive | -| `bin/suggest-synthesis` | Synthesis candidate surfacing | - -## Worked pattern - -C28 application/eval pass (2026-07-12) is the reference full loop: preflight → 6 sources → ingest + enrich → synthesis design note → stop condition met → handoff to `mindgraph-eval`. See project `log.md` that day. - -## Portfolio doctor + handoff (G2/G4/G8) - -```bash -bin/research-lane-loop doctor --all-active -bin/research-lane-loop preflight --slug <slug> # shows next_phase, blocker, suggested_command -``` - -**Typed handoffs** (gate vs opportunity vs application, etc.) are defined in: - -→ **`.context/workflows/research-project-handoff.md`** -→ template: `.context/templates/research-handoff.md` - -```bash -# Gate: clears/blocks a project decision -bin/research-lane-loop handoff-project --slug <lane> --to <project> \ - --kind gate --note "..." --evidence "10_knowledge/..." - -# Application: planned build phase -bin/research-lane-loop handoff-project --slug <lane> --to <project> \ - --kind application --note "..." - -# Opportunity: consider later — does NOT steal busy next_action -bin/research-lane-loop handoff-project --slug <lane> --to <project> \ - --kind opportunity --note "..." - -# Constraint / experiment / craft / knowledge / split / close — see workflow -``` - -**Rule:** research-lane-loop ends at synthesis + tracker close + decision. -Handoff is how that decision becomes project work **without** treating every signal as “do this next.” diff --git a/.context/workflows/research-project-handoff.md b/.context/workflows/research-project-handoff.md deleted file mode 100644 index 910376a..0000000 --- a/.context/workflows/research-project-handoff.md +++ /dev/null @@ -1,189 +0,0 @@ ---- -title: "Research ↔ project handoff" -domain: "knowledge-systems" -type: "workflow" -status: "active" -structural_type: workflow -lifecycle_scope: root -owner_surface: ".context/workflows/research-project-handoff.md" -authority: workflow-contract -volatility: stable -updated: "2026-07-13" -tags: ["research-lane", "handoff", "projects", "workflow"] ---- - -# Research ↔ project handoff - -## Purpose - -Move value **out of a research lane** into a **project surface** without collapsing every signal into the same `next_action` rewrite. - -Handoff is **not** “more literature.” It is a typed packet that answers: - -> Who should care, how hard, and what loop runs next? - -## Scope and authority - -- **From:** `30_projects/research-lanes-strategy/lanes/` (+ durable notes under `10_knowledge/`) -- **To:** one `30_projects/<slug>/` (or explicit multi-consumer list in the receipt) -- **Does not replace:** research-lane-loop (literature), craft-research-loop, project-experiment-loop, or `lane-intake archive` -- **Does not mean:** auto-code, auto-archive, or silent promotion of claims - -Related: `.context/workflows/research-lane-loop.md` (step 7), `craft-research-loop.md`, `project-experiment-loop.md`. - -## When to hand off - -Use a handoff when literature (or synthesis) changes what a **project** should do, decide, try, measure, or avoid — and continuing the same literature pass would only add footnotes. - -Do **not** hand off when: - -- You still need 3–4 sources for the **current phase** stop condition -- The “target” is only “me someday” with no project slug -- You are dumping unprocessed inbox stubs (finish ingest/synthesis first) - -## Handoff kinds (taxonomy) - -Kinds are the core of this workflow. Pick **one primary kind** per handoff. Secondary tags optional in the receipt. - -| Kind | Intent | Urgency | Typical destination behavior | -|------|--------|---------|------------------------------| -| **`gate`** | Literature (or stop-condition) **clears or blocks a project decision gate** — go / no-go / promotion / protocol freeze | High | Set or replace project `next_action` with the gate decision + evidence links | -| **`application`** | Planned **build/use** work that was always the lane’s application phase (checklist, wiring, product default) | High–medium | Primary `next_action` for implementation or craft scaffold | -| **`opportunity`** | Something **interesting** related to a project — try, consider, or queue when WIP allows | Low | Log + optional ideas card; **do not steal** primary `next_action` unless project is idle | -| **`constraint`** | Bound future work: don’t do X without Y; failure mode; compliance/risk | Medium | Append to project `decisions.md` or README risks; may demote a planned path | -| **`experiment`** | Ready for a **measured** decision (protocol, metric, eval_run_id) | Medium–high | Point at `project-experiment-loop scaffold` (or canary/matrix) | -| **`craft`** | Ready for product **trial-and-error** (bake-off, stack trial, prototype) | Medium | Point at `craft-research-loop scaffold` with one question | -| **`knowledge`** | Durable lesson for the vault; project is only a **consumer**, not the worker | Low | Ensure note is in `10_knowledge/`; project log cites path; no forced `next_action` | -| **`split`** | Signal is really a **new research question** | Medium | `## Research Lane Candidate` + `bin/lane-intake scan` — not a project implement task | -| **`close`** | **Lane** stop condition met; consumers notified before/alongside archive | Portfolio | Handoff block on lane README + archive micro-loop | - -### How to choose (quick rules) - -1. **Does a decision wait on this?** → `gate` -2. **Was this always “phase = application” on the lane card?** → `application` -3. **Must we measure before changing process/product?** → `experiment` -4. **Must we try settings/stack before claiming “what works”?** → `craft` -5. **Is it a hard bound or anti-pattern?** → `constraint` -6. **Nice to explore if capacity exists?** → `opportunity` -7. **Wrong project; needs its own literature thread?** → `split` -8. **Lane finished for real?** → `close` (+ archive when operator agrees) - -If two apply, pick the **highest urgency that changes a decision**. Put the other as a secondary note in the receipt. - -## Other situations (explicit) - -| Situation | Kind | Notes | -|-----------|------|--------| -| Eval promotion blocked on missing literature | `gate` or `experiment` | Cite protocol_ref / irregularity id | -| Canary green; optional product bake-off | `craft` or `opportunity` | Don’t inflate severity | -| Found competitor/stack pattern mid-build | `opportunity` or `constraint` | Don’t restart whole lane unless decision depends on it | -| Multi-project consumer (e.g. G43 methodology) | `knowledge` + optional multi-`to` list | One primary project still owns the receipt | -| Project discovers a knowledge gap while building | **Reverse handoff** → research-lane-loop / intake | See “Inbound” below | -| Side question during source-literature | `split` | Don’t bury in project next_action | - -## Inbound (project → research) - -Handoff is bidirectional in spirit: - -```text -project friction / unknown - → research-lane-loop (existing lane) or research-lane-intake (new) - → later: outbound handoff back to project -``` - -Do not use `handoff-project` for inbound gaps. Use lane intake or a lane preflight resume class A/B. - -## Procedure (outbound) - -### 0. Preflight - -```bash -bin/research-lane-loop preflight --slug <lane> -# Prefer blocker=project, or explicit stop-condition met for application/close -``` - -Confirm: synthesis path(s), decision sentence, target project exists, kind chosen. - -### 1. Write the packet (mental or receipt) - -Every handoff answers: - -| Field | Required | -|-------|----------| -| `kind` | yes — one of the taxonomy | -| `from_lane` | yes — slug / lane_id | -| `to_project` | yes — one primary slug | -| `decision_or_signal` | yes — one sentence | -| `evidence` | yes — knowledge paths (not “vibes”) | -| `urgency` | yes — high / medium / low | -| `next_loop` | yes — none / craft / experiment / implement / operator / archive | -| `does_not_mean` | yes — one negative boundary | - -Template: `.context/templates/research-handoff.md`. - -### 2. Apply via CLI - -```bash -bin/research-lane-loop handoff-project \ - --slug <lane> \ - --to <project> \ - --kind gate|application|opportunity|constraint|experiment|craft|knowledge|split|close \ - --note "one sentence decision or signal" \ - --evidence "10_knowledge/.../note.md" \ - [--urgency high|medium|low] \ - [--dry-run] -``` - -### 3. Kind-specific routing (CLI + operator) - -| Kind | Lane side | Project side | -|------|-----------|--------------| -| `gate` | Log + phase note “gate cleared/blocked” | **Replace** `next_action` with gate sentence + evidence | -| `application` | Log; literature stop for this phase | **Set** `next_action` to concrete build/craft step | -| `opportunity` | Log only | Append **Consider** entry to `log.md` (and `ideas/` if project uses it); **do not** overwrite a busy `next_action` | -| `constraint` | Log | Append `decisions.md` or risk note; optional demote conflicting next_action | -| `experiment` | Log | Set `next_action` to experiment-loop scaffold command | -| `craft` | Log | Set `next_action` to craft scaffold command | -| `knowledge` | Log | Cite path in project log; no forced next_action | -| `split` | Emit candidate; optional intake | No project next_action change | -| `close` | Handoff block + archive checklist | Notify consumers in log only | - -### 4. Close the research pass - -Still complete research-lane-loop step 6–7: capture index, phase status, research-lanes-strategy `log.md`, loop decision. - -### 5. Verification - -- [ ] Receipt exists (CLI prints path under project `outputs/` or lane `log.md`) -- [ ] Kind matches urgency (opportunity did not hijack active WIP next_action) -- [ ] Evidence paths resolve under `10_knowledge/` or project outputs -- [ ] Next loop is named (or explicit `none`) -- [ ] Archive only if kind is `close` **and** operator accepts lane stop - -## Anti-patterns - -- Every interesting paper becomes a **gate** (severity inflation) -- Every handoff overwrites project `next_action` (kills real WIP) -- Handoff without evidence paths (“trust me”) -- Using handoff instead of synthesis (skipping step 5) -- Archiving the lane because one opportunity was filed -- Filing `split` as `application` on a random active project - -## Relation to the three research modes - -```text -literature ──handoff──► project surface - │ - ├─ gate / application / constraint → implement or decide - ├─ craft → craft-research-loop - ├─ experiment → project-experiment-loop - ├─ opportunity → consider queue - ├─ knowledge → consume note only - └─ split / close → intake or archive -``` - -## Update discipline - -- Workflow is stable; extend kinds only with a short DECISIONS note if taxonomy changes. -- Receipts are project/lane logs — append-only. -- CLI defaults must stay conservative for `opportunity` (no next_action steal). diff --git a/.context/workflows/session-close.md b/.context/workflows/session-close.md index 1648c62..cab51ee 100644 --- a/.context/workflows/session-close.md +++ b/.context/workflows/session-close.md @@ -30,6 +30,6 @@ A session-close workflow should always update state, write a concise handoff not 3. **Regenerate indexes:** Run `bin/sync-project-index --write` when project metadata changed. The generated `30_projects/index.md` is local and ignored by Git. 4. **Refresh retrieval:** Run `bin/mindgraph-refresh` when durable knowledge changed. 5. **Review telemetry:** Run `bin/workflow-report --days 1` when diagnosing process friction. -6. **Audit surface (post-ingest verification):** When durable knowledge was changed or a batch/tiered route occurred, run `bin/audit-sweep --dry-run` (or `--apply` after review). See [.context/workflows/audit-sweep.md](.context/workflows/audit-sweep.md). This is the compensating control for review-after (ADR-019). +6. **Audit surface (post-ingest verification):** When durable knowledge was changed or a batch/tiered route occurred, run the optional/deferred audit-sweep tool if installed (`bin/audit-sweep --dry-run`, or `--apply` after review). This is the compensating control for review-after (ADR-019). It is not part of the first public-core slice. 7. **Handoff digest (operator load reduction):** `bin/session-close --apply` now writes `20_live/last-handoff-draft.md` with recent signals from ingest-status, audit-sweep, etc. + a template for STATE.md narrative. Review/edit it into STATE.md to lower manual transcription. 8. **Commit:** Ensure any living documents have their timestamps or logs updated. diff --git a/.context/workflows/session-guide.md b/.context/workflows/session-guide.md deleted file mode 100644 index 261b0f9..0000000 --- a/.context/workflows/session-guide.md +++ /dev/null @@ -1,113 +0,0 @@ -# Session Guide - -What a normal MainFrame working session looks like from the operator's seat. -Scripts handle the deterministic steps; judgment stays manual. Telemetry -records itself. The per-step contracts live in the other files in this -folder — this page is the end-to-end view. - -## 1. Open - -```bash -bin/session-open -``` - -Loads context in a fixed order (root contract, state, active project README). -It picks the active project from `STATE.md` — if that is stale, override with -`--project <slug>` now and correct `STATE.md` before you close. Details: -`session-open.md`. - -## 2. Work - -- Capture without ceremony: drop files into `00_inbox/`. No frontmatter is - required at capture time; pass 1 is suggestion-first (ADR-011). -- Project work happens inside `30_projects/<slug>/` workbenches, which the - outer repo ignores. -- Live-state updates in `20_live/` get dates and append-only timelines or - snapshots, never silent overwrites. -- Telemetry needs nothing from you: Claude Code and Codex hooks append - redacted events (hashes, zones, durations — never prompt or file text) to - `20_live/workflow-metrics/events/`. - -## 3. Ingest pass (only when you choose to run one) - -Not every session includes ingest. When one does: - -```bash -bin/ingest-minion run --dry-run # warnings are suggestions, not rejects -bin/ingest-minion run --apply -``` - -The ingest-agent then enriches `01_ingest/ready/` files (domain, type, tags, -connections — the judgment middle), `bin/prep-ingest` validates them into -`queue/`, and minion pass 2 applies the strict `queue/ -> 10_knowledge/` -gate. Details: `ingest-minion.md`; migration drops follow -`rolling-second-brain-migration.md`. - -Migration batches and aged backlogs skip the per-file loop: batch mode -(ADR-019) classifies the lane under `.context/routing-policy.md`, you approve -one review table, and exceptions stay in `ready/` tagged `routing-exception`. -Routed clips carry `needs-audit`, which the epistemic research system sweeps -continuously (see [.context/workflows/audit-sweep.md](.context/workflows/audit-sweep.md) and `bin/audit-sweep`) — review-after instead of approve-before. Your first full-table -review is also the calibration baseline for how much autonomy each rule earns. Run the sweep regularly (integrated into session-close) and review the pending-review surface before widening Tier A usage. - -To see where the backlog actually stands first: - -```bash -bin/ingest-status -``` - -It separates batch-registered migration files (intentional backlog tracked by -an append-only disposition ledger) from organic captures going stale. - -## 4. Close - -```bash -bin/session-close --check -bin/session-close --apply # index sync, MindGraph refresh, telemetry report -``` - -The judgment steps stay yours and are only reminded, never automated: - -1. Update `STATE.md`: active project, what changed, what remains, the next - reentry point. -2. Record architecture or workflow changes in `DECISIONS.md`. -3. Review the staged diff before committing — keep private content out of - the tracked surface. - -Details: `session-close.md`. - -## Weekly, or when something feels off - -```bash -bin/workflow-report --days 7 -bin/ingest-status -``` - -Read the per-client quality block before trusting aggregates, then run the -evaluation loop in `process-evaluation.md` before changing any process. - -## Known telemetry limits (as of 2026-06-09) - -- Codex has no `SessionEnd`, `PostToolUseFailure`, or `PostToolBatch` hook - events, and it silently ignores unrecognized names. Its session-close - coverage is structurally 0%; judge close habits from the claude row, not - the aggregate. `.codex/hooks.json` now wires its real vocabulary - (`Stop`, `SubagentStart`/`SubagentStop`, `PermissionRequest`, - `UserPromptSubmit`, compaction); treat those rows as unverified until the - first events appear, and expect Codex to re-prompt once to trust the - changed hooks. -- Codex tool failures are derived from `tool_response` exit codes inside - `PostToolUse` (best effort); its payloads carry no durations. -- Permission and approval events (`PermissionRequest`, `Notification`) - cannot fire while sessions run with permissions bypassed, so any state - built on them stays unverified until a default-permission session runs. -- Tool failures are execution events, not task quality. - -## What never happens automatically - -- No script writes `STATE.md` narrative, `DECISIONS.md`, or knowledge - content. -- Pass 1 never rejects a capture; strict rejects exist only at the - `queue/ -> 10_knowledge/` gate. -- Nothing moves or deletes raw evidence; the reports (`workflow-report`, - `ingest-status`) are read-only. diff --git a/.context/workflows/session-lifecycle.md b/.context/workflows/session-lifecycle.md index dc37c4c..d0efc7b 100644 --- a/.context/workflows/session-lifecycle.md +++ b/.context/workflows/session-lifecycle.md @@ -20,7 +20,7 @@ related_surfaces: do_not_use_for: - "project outcome work itself" - "replacing eval-schedule or mainframe-doctor" -updated: "2026-07-23" +updated: "2026-09-04" --- # Session lifecycle loop @@ -61,10 +61,14 @@ open (focus/STATE → contract chain) → work → close (check|checkpoint|apply ### Open -1. Prefer structured focus: `20_live/focus/current.yaml` (ADR-044 / MPE-024). -2. Run `bin/session-open` (or `--json` / `--project <slug>`). -3. Load only the progressive chain the script prints (AGENTS → STATE → project → plan → …). -4. Fail closed if project path does not resolve (Unit 2.2). +1. Run `bin/session-open --json` for arrival, or add `--project <slug> --task + "request"` for project resume. Use `--intent resume` for recorded focus. +2. Read required content batches separately; recover any truncated response. + Arrival stops at the lifecycle map and recorded focus. Project-specific + diagnosis or recommendations require the resume route and reconstruction. +3. Add `--path <repo-relative-path>` with an explicit project for deeper rules. +4. Report missing relevant context before dependent work. Keep focus freshness + and adjacent scheduler warnings separate from permission and task readiness. Detail: `.context/workflows/session-open.md`. diff --git a/.context/workflows/session-open.md b/.context/workflows/session-open.md index 2b42fe9..3e0e138 100644 --- a/.context/workflows/session-open.md +++ b/.context/workflows/session-open.md @@ -1,21 +1,113 @@ -# Session Open Workflow - -Part of the **session lifecycle loop**: `.context/workflows/session-lifecycle.md` -(open → work → close → loop decision). - -A session-open workflow should load only the minimum useful context in a fixed order. This implements progressive context disclosure and prevents the model from consuming broad ambient context that is unrelated to the immediate work. - -## Script -- Command: `bin/session-open` -- Auto-detect project: reads `## Active Project` from `STATE.md` -- Override: `--project <slug>` -- Content dump: `--print-contents` -- Structured output: `--json` - -## Suggested Order: -1. Load `AGENTS.md` (root control file) -2. Load `STATE.md` (workspace state file) -3. Load the selected project `README.md` (if working on a specific project; see `30_projects/AGENTS.md`) -4. Load the phase plan (if applicable) -5. Load task-local docs/code -6. Load evidence ONLY when needed. +--- +title: "Session open" +domain: "knowledge-systems" +type: "workflow" +status: "active" +source: "AGENTS.md, HARNESS.md, and bin/session-open" +tags: ["session", "onboarding", "context-routing"] +updated: "2026-09-05" +structural_type: "workflow" +lifecycle_scope: "root" +owner_surface: "HARNESS.md" +authority: "workflow-contract" +privacy: "public-safe" +volatility: "stable" +source_of_truth: true +update_rule: "replace-with-review" +verification: ["uvx --with pytest pytest tests/test_session_open.py tests/test_focus_authority.py -q"] +related_surfaces: ["bin/session-open", ".context/workflows/session-lifecycle.md", ".context/workflows/project-resume-and-candidate-lifecycle.md"] +do_not_use_for: ["system health", "project acceptance", "authorization", "evidence verification"] +--- + +# Session open + +**Trigger:** starting or resuming MainFrame work. +**Owner surface:** `HARNESS.md`. +**Stop state:** the request, governing sources, applicable constraints, relevant +unknowns and next useful step are established for the selected route. Missing +context is reported before dependent claims or action. + +## Choose the route + +```bash +bin/session-open --json +bin/session-open --project my-project --task "onboarding" --json +bin/session-open --intent resume --json +``` + +Without a project, the default is **arrival**: root AGENTS, the harness's Session +orientation section, narrative state, this workflow and available structured +focus. Describe the lifecycle map and recorded focus with freshness caveats, +then stop. Project-specific diagnosis or recommendations require **resume**. + +A named `--project` implies resume. Explicit `--intent resume` selects structured +focus, then the STATE fallback. Resume requires the full harness, project lifecycle +and applicable local contracts, reconstruction workflow, README and existing +PROJECT metadata, methodology note, current log and latest decision. Older log +and decision entries remain available when the current task needs them. + +Add `--path` with `--project` to include deeper ancestor contracts. The path must +exist inside that project; use an existing parent for a new file. Sibling trees, +raw evidence and evaluator bodies are not recursively loaded. + +`--task` nominates an active plan by unique overlap with its filename/title. +This is a lexical navigation aid, not a relevance verdict. Verify the candidate. +Ties, no match or no task leave selection unresolved and list the candidates; +there is no alphabetical first-plan fallback. Use `--plan <repo-relative-path>` +with `--project` for a known plan inside its `plans/` directory. + +Operation-aware selection is deferred. For an operation, read `40_operations/AGENTS.md`, +its owning README and applicable local contracts directly without reclassifying it. + +## Read complete, bounded context + +The JSON lists numbered `read_batches`, each at most 8,000 UTF-8 content bytes. +Read **one batch per tool response**, retaining the same routing arguments: + +```bash +bin/session-open --project my-project --task "onboarding" --read-batch 1 --json +``` + +Continue through every required batch. Allow enough tool output for the batch +(for example, 10,000 output tokens); do not combine many batches into one response. +Use `--expect-hash <source_sha256>` from the listing to reject source changes +between calls. If a response is truncated, recover the missing content before +relying on it. A byte limit cannot guarantee another tool's output setting. +`--print-contents` prints the first batch only and identifies the continuation. + +The route emits `context_status: unread` when paths are valid, or `incomplete` +when required context is missing, unreadable or invalid. It never reports that an +agent read or understood the content. The agent checks actual receipt of required +context and completes the route's `stop_when` conditions before giving its answer. + +For resume, reconstruct current coordination, Git/worktrees, candidate/experiment +identity and direct evidence under the project-resume workflow. Use valid existing +receipts for status questions; rerun checks for changed source, missing evidence +or an explicit test request. After reconstruction, answer the request. Load an +action workflow when that action is authorized. + +## Output meanings and fallback + +- `ok`: required context files can be read and project/task/plan paths resolve. + Exit 1 indicates an unresolved prerequisite. File presence and readability do + not establish comprehension, project readiness, system health or authorization. +- `focused_project`: recorded attention during arrival; `project` is the project + actually entered by resume. Selecting either route never changes focus. +- `missing_required`, `read_errors`, `intent_error`, `project_error`, `path_error` + and `plan_error`: unresolved prerequisites. `reading_verified` is always false. +- `degraded`, `focus_errors` and `focus_warnings`: context/freshness concerns. + Keep stale attention distinct from missing task authority. +- `eval_schedule_ok`: adjacent scheduled-evaluation health. `null` means the check + was unavailable or timed out. This warning does not initiate unrelated repairs. + +### Required context remains visible + +> **Binds:** callers consuming `bin/session-open` output +> **Tier:** T2 (CLI blocks missing/unreadable required context and invalid paths; reading itself is T0) +> **Check:** `tests/test_session_open.py`, `tests/test_focus_authority.py`; live CLI exits and content batches +> **Escape:** read the named source files directly; report unresolved context before dependent claims or action + +When routing is unavailable, use the same source order directly. If the arrival +section is missing, read the full harness. A missing project prerequisite keeps +resume incomplete; it never licenses skipping the relevant contract. The larger +open/work/close loop lives in `session-lifecycle.md`. diff --git a/.context/workflows/source-literature.md b/.context/workflows/source-literature.md deleted file mode 100644 index abbe22d..0000000 --- a/.context/workflows/source-literature.md +++ /dev/null @@ -1,105 +0,0 @@ -# Source Literature Workflow - -Use this workflow when you need peer-reviewed or well-accepted sources **before** ingest. It complements `ingest-minion` by handling discovery, credibility gating, and inbox capture. - -## Defaults - -- Skill: `.agents/skills/source-literature/SKILL.md` -- Subagent: `agents/source-literature-agent.md` -- Credibility reference: `.agents/skills/source-literature/references/credibility-tiers.md` -- Output: `00_inbox/` raw stubs + run note -- Downstream: `.context/workflows/ingest-minion.md` → `ingest-source` skill - -## When to use - -- Starting research on a new topic (e.g. data integrity, salary negotiation). -- Gap-filling an existing domain collection (check vault first). -- User asks for "peer-reviewed sources" or "well-accepted literature." - -## When not to use - -- File already in hand → drop in `00_inbox/` and run ingest directly. -- Fast social/article clip → `x-bookmark-web-clipper` workflow. -- Claim extraction from existing notes → epistemic audit / extraction-agent path. - -## Steps - -1. **Frame** — Write research question, stakes, target domain, exclusions in the run note. -2. **Dedup** — Search `10_knowledge/<domain>/`, source catalogs, and MindGraph before external search. -3. **Search** — Invoke `source-literature` skill (or source-literature-agent). Agent returns candidate table with tier labels. -4. **Confirm** — User approves candidates (required when >3 sources, any Tier C/E, or new domain proposed). -5. **Capture** — Agent writes stubs to `00_inbox/` as `YYYY-MM-DD__<domain>__raw__<slug>.md`. -6. **Run note** — Agent writes `YYYY-MM-DD__source-literature-run__<topic-slug>.md` with queries, accept/reject log, file list. -7. **Ingest handoff**: - ```sh - bin/ingest-minion run --dry-run - bin/ingest-minion run --apply - ``` -8. **Agent enrich** — Invoke ingest-agent on `01_ingest/ready/` per `agents/ingest-agent.md`. -9. **Post-route enrich** — After stubs land in `10_knowledge/<domain>/`: - ```sh - export UNPAYWALL_EMAIL=you@example.com # free API; optional but recommended - bin/post-route-enrich --subset <domain> - ``` - Fetches OA full text, appends `## Full text extract`, then runs `bin/mindgraph-refresh` so deeper body terms enter the search index. -10. **Optional audit** — `bin/audit-sweep --apply --subset <domain>` for `needs-audit` items. -11. **Capture surfaced questions as lanes** — Any side questions, gaps, or recurring uncertainties that emerged during the run should be emitted as `## Research Lane Candidate` blocks (see `.context/workflows/research-lane-intake.md`). Run `bin/lane-intake scan <run-note>` (or scaffold directly). This is the explicit step for "adding new questions to research lanes". Dual MindGraph pass and receipt are produced automatically. - -## Capture filename convention - -``` -YYYY-MM-DD__<domain>__raw__<author-or-body>-<short-slug>-<year>.md -``` - -Examples: - -- `2026-06-17__negotiation__raw__small-salary-negotiation-2007.md` -- `2026-06-17__regulated-systems__raw__pda-data-integrity-history-2018.md` - -## Run note convention - -``` -YYYY-MM-DD__source-literature-run__<topic-slug>.md -``` - -Include: question, stakes, queries, candidate table, rejects with reason codes, captures written, ingest status. - -## Guardrails - -- Captures are `type: raw` — bibliographic stubs, not synthesis. -- All captures carry `needs-audit` until epistemic review. -- Institutional guidance (FDA, MHRA, WHO) is Tier D — authoritative for expectations, not compliance proof. -- Do not route directly to `10_knowledge/` — inbox → ingest pipeline only. -- New domains require user confirmation before folder creation (same rule as ingest-agent). - -## Pipeline position - -``` -Research question - → source-literature (this workflow) - → 00_inbox/ - → ingest-minion - → ingest-source - → 10_knowledge/<domain>/ - → knowledge synthesis (required end of a research-lane pass) - → tracker close + optional research-lane-intake for side questions -``` - -For a **full research-lane pass** (source-literature through synthesis, loopable by phase), use `.context/workflows/research-lane-loop.md` and `.agents/skills/research-lane-loop/SKILL.md`. This workflow remains the discovery step only. -## Claim discipline - -Source discovery feeds the capture pipeline, so credibility tiers govern **what to -retrieve** and GRADE certainty governs **what confidence to assign afterwards**. -Both are defined in `EPISTEMIC_STANCE.md`; procedure in -[epistemic-standard.md](epistemic-standard.md). - -- Never record a source you did not fetch. A `retrieved_at` an agent writes about - itself is worth nothing; captures claiming it were citing hard 404s. -- Verify identifiers, not bylines. The identifier-shape check found all 107 - fabricated captures in the 2026-08-09 finding; the non-human-author heuristic - caught 6 and missed 76. -- **Escape:** finding one real source and recording the gap beats finding three - that resolve to nothing. There is no target count and never should be. - -**Stop state.** If a search returns nothing usable, close the pass with the gap -named. Do not lower the tier bar to fill a lane. diff --git a/.context/workflows/system-integrity-audit.md b/.context/workflows/system-integrity-audit.md deleted file mode 100644 index c750125..0000000 --- a/.context/workflows/system-integrity-audit.md +++ /dev/null @@ -1,121 +0,0 @@ -# System Integrity Audit Workflow - -Use this workflow to audit and review system architectures, software workflows, or AI data pipelines from a data-integrity perspective. This protocol evaluates how reliably a system captures evidence, traces decisions, maintains rule versioning, and handles data preservation. - -## Related docs - -| Context / Template | Path | -|:---|:---| -| Meeting lifecycle | [boardy-meeting-lifecycle.md](boardy-meeting-lifecycle.md) | -| Website audit | [website-audit.md](website-audit.md) | -| Epistemic standard | [epistemic-standard.md](epistemic-standard.md) | - ---- - -## 1. Scope & Calibration Boundary - -Every review must start with a clear definition of bounds to manage liability and establish peer-to-peer collaboration: -- **Data-Integrity Focus:** Review is strictly limited to how data flows, where it is recorded, how it is modified, and how decisions are reconstructed. -- **Not a Legal/Compliance Opinion:** State clearly that this is an architectural peer review, not a formal legal counsel, privacy audit (e.g. GDPR/CCPA certification), or regulatory inspection. -- **Execution vs. Documentation:** Clarify whether the review is based on written documentation/diagrams alone or verified via a live trace. - ---- - -## 2. Step-by-Step Audit Procedure - -Follow this sequence when auditing a client's system architecture: - -### Step 2.1: Map the Anatomy of the System -Identify the core components of the target system: -1. **Raw Inputs:** The original user data, uploads, API requests, or events. -2. **Transient/Processing Layer:** Intermediate APIs, transcription services, worker queues, LLM parsers, or temporary databases. -3. **Derived Outputs:** Summaries, dashboards, notifications, or reports generated for the user. -4. **Decision Gates:** Where human review or approval blocks automated downstream actions. - -### Step 2.2: Identify the Survival Boundary -Determine the relationship between raw inputs and derived outputs: -- **The Survival Test:** If the raw input is deleted (e.g., for data minimization or privacy), does the system preserve enough structured evidence to independently reconstruct and defend the output? -- If the original evidence is deleted and the client only retains scores/verdicts, highlight that the retained records cannot be audited or replayed. - -### Step 2.3: Execute the R1–R8 Checklist -Evaluate the system against the **R1–R8 System Integrity Checklist** (Section 3). For each point: -- Mark as **PASS** if the system meets the condition. -- Mark as **GAP** if the system relies on unverified assertions or manual workflows. -- Detail the exact risk and remediation for each GAP. - -### Step 2.4: Define the Synthetic Walkthrough -Design a concrete "verification trace" to test the system's actual behavior on a single synthetic record. Outlining this walkthrough forces the client to move from *assertions of design* to *operational proof*. - -### Step 2.5: Draft the Review -Package the findings into the standard audit report layout (Section 4). Maintain a peer-to-peer, collaborative, and constructive tone. - ---- - -## 3. The R1–R8 System Integrity Checklist - -Use this standard checklist to evaluate the robustness of any data pipeline or decision-making system. - -| Code | Principle | Audit Question | Pass Condition | -| :--- | :--- | :--- | :--- | -| **R1** | **Raw Source Capture** | Can we identify the authoritative, raw source records that triggered the system run? | The raw inputs (e.g. raw transcripts, uploaded files, API payloads) are preserved with unique IDs, hashes, and timestamps. | -| **R2** | **Logic/Rule Versioning** | Can we recover the exact rule schema, prompt template, model parameters, or code commit in effect at the time of execution? | The run log records the exact version identifier or snapshot of the governing logic. | -| **R3** | **Run Initialization Log** | Is the start of the workflow recorded immediately in a non-repudiable log? | The database writes a run entry at the initialization timestamp before downstream processing. | -| **R4** | **Sequenced Execution Path** | Does the logic trace a strict, ordered progression of validation nodes? | The execution path or trace shows that prerequisites were satisfied before proceeding. | -| **R5** | **Outcome-to-Source Traceability** | Is every rating, flag, or verdict dimension supported by specific, attributable evidence? | The system output links specific claims back to segments of the source data or database states. | -| **R6** | **Controlled System Overrides** | Are exceptions, manual corrections, or rule overrides recorded transparently? | Adjustments to system state are logged as separate, signed events (`PATCH` transactions) rather than silently overwriting the original output. | -| **R7** | **Attributable Approval Gates** | Is human oversight a hard gate (rather than a retrospective override), and is it signed by a verified account? | The workflow requires a named human supervisor to confirm the verdict before triggering downstream automation. | -| **R8** | **Exportable Evidence Package** | Does the final output package contain verifiable links or hashes of all run elements? | The delivered report or API packet contains cryptographic hashes of the raw input, the active policy version, the run logs, and the approval events. | - ---- - -## 4. Audit Report Template - -When drafting the final deliverable for a client, copy and use the format below. - -```markdown -# System Integrity & Data-Flow Review: [System Name] - -*Prepared by Cameron Sanderson* -*Date: [Date]* - ---- - -## Scope & Boundaries -I reviewed the [System Name] architecture and data-flow specifications from a data-integrity and decision-reconstruction perspective. - -This review covers system architecture, record linkage, and data-preservation design. It is not a legal opinion, a privacy assessment, or a formal compliance certification. The findings below are based on the provided technical package and should be verified on a live system trace. - ---- - -## Start Here: The Core Survival Test -[Identify the main point of vulnerability in the system. Typically, this is whether the raw source data survives deletion/minimization, and what the retained summaries actually prove if audited.] - ---- - -## Architectural Findings - -### 1. [Finding Heading - e.g., Logic Versioning Gaps] -[Analyze what is currently happening vs. what is required. Reference R2 or other R-principles.] - -### 2. [Finding Heading - e.g., Audit Trail Attributability] -[Analyze the storage of audit logs. Highlight if logs are stored in soft formats like spreadsheets or databases without write-once protection.] - -### 3. [Finding Heading - e.g., Human-in-the-Loop Boundaries] -[Detail whether human review is a hard gate blocking automation, or merely a retrospective override.] - ---- - -## Reconstruction Trace Checklist -During the upcoming walkthrough, we should trace a single synthetic execution to verify the following parameters: - -- [ ] **Authoritative Source:** Verify what remains in client and system storage post-deletion. -- [ ] **Logic Versioning:** Trace one automated output back to the exact rule schema, prompt file, and model version. -- [ ] **Logging & Error Capture:** Force a processing error and verify that it triggers an alert and logs a failure rather than reporting false success. -- [ ] **Access & Credentials:** Confirm the actual permissions of the system's service accounts and verify that they match the stated access control boundaries. -- [ ] **Human Sign-Off:** Confirm that downstream automation is blocked until a user approval signature is written to the audit log. - ---- - -## Summary Assessment -[Summarize the recommendations. Emphasize separating data minimization (deleting transient files) from decision preservation (keeping audit trails, hashes, and human sign-off records intact).] -``` diff --git a/.context/workflows/workflow-telemetry.md b/.context/workflows/workflow-telemetry.md deleted file mode 100644 index 2b8ab77..0000000 --- a/.context/workflows/workflow-telemetry.md +++ /dev/null @@ -1,28 +0,0 @@ -# Workflow Telemetry Workflow - -This workflow records local process metrics so future sessions can improve the way work is done. It does not record prompts, file contents, command output, or tool responses. - -## Data Path -- Claude Code hooks live in `.claude/settings.json`. -- Codex hooks live in `.codex/hooks.json`. -- Antigravity hooks live in `.antigravity/settings.json` and `.antigravity/hooks.json`. -- Hook events call `bin/workflow-event`. -- Redacted JSONL is appended under ignored `20_live/workflow-metrics/events/`. -- Summaries are produced with `bin/workflow-report`. - -## Captured Fields -- Hook event name -- Tool name -- Duration when the client provides it -- Success or failure for post-tool events -- Permission mode and effort level when present -- Redacted path zone, file extension, command head, and hashed identifiers -- Optional allowlisted `process_id` / `process_ids` values for catalogue attribution - -## Guardrails -- Telemetry is append-only local state. -- Never log raw prompts, raw command text, file contents, tool output, or model responses. -- Process IDs must be stable catalogue-style identifiers with an approved prefix (`cli-`, `script-`, `ingest-`, `workflow-`, `skill-`, or `agent-`), not raw commands, paths, prompts, project names, people names, or free-text task descriptions. -- `bin/workflow-report` keeps redacted defaults; use `--by-process` when a process rollup is needed. -- Use `bin/workflow-report --input-signals` only for aggregate redacted input rollups such as safe command heads, lifecycle path zones, and file extensions; it must not expose command hashes, prompt hashes, full paths, or raw command text. -- Use reports to spot workflow friction, not to treat speed as the only measure of quality. diff --git a/.context/workflows/x-bookmark-web-clipper.md b/.context/workflows/x-bookmark-web-clipper.md deleted file mode 100644 index a17e844..0000000 --- a/.context/workflows/x-bookmark-web-clipper.md +++ /dev/null @@ -1,126 +0,0 @@ -# X Bookmark Web Clipper Workflow - -Use this workflow when capturing X bookmarks into MainFrame as raw inbox evidence with Obsidian Web Clipper. - -## Purpose - -Capture each bookmarked X post, plus any directly linked article or readable source page, into `00_inbox/` without changing the raw captured files. This is a fast-capture workflow; normalization and routing happen later through the ingest workflow. - -## Defaults - -- Browser: the user's real Chrome session. -- Source list: `https://x.com/i/bookmarks`. -- Capture action: Obsidian Web Clipper `Save file...`. -- Temporary save location: `~/Downloads`. -- Final destination: `00_inbox/`. -- Run note location: `00_inbox/YYYY-MM-DD__x-bookmark-web-clipper-run.md`. - -## Setup Checks - -1. Confirm Chrome is already authenticated to X. -2. Confirm Obsidian Web Clipper is installed and available from Chrome Extensions. -3. Record a Downloads baseline before clipping: - - ```sh - find ~/Downloads -maxdepth 1 -type f -print | sort - ``` - -4. Create a run note with start time, baseline files, and a capture table with columns for bookmark URL, linked article URL, post clip, article clip, cleanup, and notes. -5. Run one pilot bookmark before the full batch. - -## Clipping Loop - -For each bookmark: - -1. Open the bookmarked status URL directly when possible. -2. Click Chrome toolbar Extensions. -3. Click Obsidian Web Clipper. -4. Click `Save file...`. -5. Verify a new `.md` file appears in Downloads. -6. Record the post URL, generated filename, and any capture quirks in the run note. - -Direct status URLs are more reliable than clipping from the scrolling bookmark timeline because they make Web Clipper target the intended post instead of a nearby item. - -## Linked Articles - -If the bookmarked post links to a readable article: - -1. Open the article or article card. -2. Prefer the canonical readable page when available. -3. For X Articles, click the article card from the post. If that fails, try the visible focus-mode URL or direct `/article/<id>` URL. -4. Save the article page with Obsidian Web Clipper `Save file...`. -5. Record the article URL and generated filename in the same run-note row as the source post. - -If the link is a GitHub repository, product page, course page, or project page rather than an article, capture it only when it is the key linked source or provenance for the bookmark. For external lead forms, capture the visible page only; do not enter personal data or submit forms. - -## X Article Quirks - -- Some X Article cards are wrapped in a video or image surface. Click the lower title/card area if the first click only starts video playback. -- If a card opens `/photo/1`, search the exact article title to recover a direct `/article/<id>` URL. -- If exact-title search does not expose a readable article and the card remains media-bound, log the article as partial and keep the bookmark row as a failure/partial. -- Reply bookmarks may save as the parent post with the bookmarked reply included in Comments. Log that quirk rather than rewriting the raw capture. - -## Batch Move - -Move only new Web Clipper Markdown files from Downloads into `00_inbox/`. Preserve filenames and avoid overwrites with suffixes: - -```sh -python3 - <<'PY' -from pathlib import Path - -src_dir = Path.home() / "Downloads" -dst_dir = Path.home() / "Desktop" / "MainFrame" / "00_inbox" - -for src in sorted(src_dir.glob("*.md")): - dst = dst_dir / src.name - if dst.exists(): - stem, suffix = dst.stem, dst.suffix - i = 2 - while True: - candidate = dst_dir / f"{stem}-{i}{suffix}" - if not candidate.exists(): - dst = candidate - break - i += 1 - src.rename(dst) - print(f"MOVED\t{dst.name}") -PY -``` - -After moving, update the run note with final `00_inbox/` filenames. - -## Completion Policy - -- Do not unbookmark or remove X bookmarks unless the user explicitly confirms that cleanup action during the run. -- If the user confirms cleanup, unbookmark only after the post and required article/source sidecars are confirmed in `00_inbox/`. -- Failed or partially clipped bookmarks stay bookmarked and are listed in the run note. -- If no cleanup confirmation is given, leave all X bookmarks intact and rely on the run note for future duplicate avoidance. - -## Failure Handling - -- If Web Clipper stays open after `Save file...`, check Downloads before clicking again. Duplicate files can be discarded only after verifying they are duplicate captures. -- If Chrome or the accessibility tree stops exposing windows, restart or reactivate Chrome and re-check the current URL before continuing. -- If an article is behind a login, form, Telegram gate, or bio-only promise, do not chase it unless the user explicitly asks. Capture the visible post and log the limitation. -- If Downloads contains non-baseline files that are not Web Clipper Markdown files, leave them alone. - -## Verification - -At the end: - -1. Confirm Downloads has no new Web Clipper Markdown left behind: - - ```sh - find ~/Downloads -maxdepth 1 -name '*.md' -print - ``` - -2. Confirm all moved filenames listed in the run note exist in `00_inbox/`. -3. Confirm the run note lists every processed bookmark and every failure/partial. -4. Record whether bookmark cleanup was attempted. -5. Leave raw captures unedited as evidence. - -## Privacy And Provenance - -- Treat all clipped files in `00_inbox/` as private raw evidence. -- Do not edit generated clip bodies after capture. -- Do not enter personal data, submit forms, follow gates, or alter accounts while clipping. -- Preserve source URLs in the run note so later ingest can distinguish raw evidence from extracted working copies. diff --git a/.githooks/post-commit b/.githooks/post-commit deleted file mode 100755 index 69c4594..0000000 --- a/.githooks/post-commit +++ /dev/null @@ -1,24 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -ROOT="$(git rev-parse --show-toplevel)" - -# MindGraph ingests 10_knowledge/ only. Bail unless the just-committed change -# actually touched that scope, so commits to live state, project files, or -# tooling don't trigger a no-op refresh (and the concurrent-refresh window that -# rapid unrelated commits would otherwise open). Capture-then-grep avoids the -# pipefail + `grep -q` SIGPIPE gotcha; --root makes the initial commit correct. -CHANGED="$(git diff-tree --no-commit-id --name-only -r --root HEAD 2>/dev/null || true)" - -LOG_DIR="$ROOT/20_live/workflow-metrics" -mkdir -p "$LOG_DIR" - -# 1. MindGraph refresh for 10_knowledge/ -if grep -q '^10_knowledge/' <<<"$CHANGED"; then - nohup "$ROOT/bin/mindgraph-refresh" >>"$LOG_DIR/mindgraph-refresh.log" 2>&1 & -fi - -# 2. Skill evaluation static checks (lint + stub) on commit -if grep -q '^\.agents/skills/' <<<"$CHANGED"; then - nohup "$ROOT/bin/skill-eval" stub --changed "$CHANGED" >>"$LOG_DIR/skill-eval.log" 2>&1 & -fi diff --git a/.githooks/pre-commit b/.githooks/pre-commit deleted file mode 100755 index c18156c..0000000 --- a/.githooks/pre-commit +++ /dev/null @@ -1,6 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -ROOT="$(git rev-parse --show-toplevel)" - -git -C "$ROOT" diff --check --cached diff --git a/.github/workflows/public-core.yml b/.github/workflows/public-core.yml new file mode 100644 index 0000000..52fcb9a --- /dev/null +++ b/.github/workflows/public-core.yml @@ -0,0 +1,39 @@ +name: Public core + +on: + pull_request: + branches: [main] + push: + branches: [main] + +permissions: + contents: read + +jobs: + public-core: + runs-on: ubuntu-latest + steps: + - name: Check out repository + uses: actions/checkout@v4 + + - name: Set up Python + uses: actions/setup-python@v5 + with: + python-version: "3.12" + + - name: Install pytest + run: python -m pip install --disable-pip-version-check pytest + + - name: Run public tests + run: python -m pytest tests 40_operations/mainframe-process-eval/tests -q + + - name: Exercise public CLIs + run: | + ./bin/process-eval --help + ./bin/process-eval status --json + ./bin/contract-lint --help + ./bin/contract-lint --selftest + ./bin/session-open --help + ./bin/session-close --help + ./bin/ingest-minion --help + ./bin/work-inventory --help diff --git a/.gitignore b/.gitignore index 651fe40..b0d4fb5 100644 --- a/.gitignore +++ b/.gitignore @@ -1,185 +1,79 @@ -# OS generated files -.DS_Store -.DS_Store? - -# Local scratch (client PDFs, renders, throwaway verification) — never commit -tmp/ -._* -.Spotlight-V100 -.Trashes -ehthumbs.db -Thumbs.db +# Public MainFrame product gitignore. +# This file is generated into the public candidate. It is not a copy of the +# private workspace ignore list. -# Python generated files +# Generated / runtime +.DS_Store __pycache__/ *.pyc - -# Personal Working State & Logs -STATE.md -01_ingest/ingest-log.md -01_ingest/prep-ingest-log.md - -# Local tool configuration and generated telemetry -.mcp.json -.claude/settings.local.json -20_live/workflow-metrics/ -# Obsidian workspace state — records recent files and layout, which would leak -# private project names into the public repo -.obsidian/ - -# Local databases *.sqlite *.sqlite-shm *.sqlite-wal -# Inbox (ignore all personal captures). core.ignorecase=true here, so a lowercase -# index.md whitelist would also un-ignore an uppercase INDEX.md (a personal catalogue -# lives at 00_inbox/INDEX.md). Keep only the .gitkeep / AGENTS.md control files. +# Credentials +.env +.env.* +!.env.example +*.key +*.pem +*.p12 +*.pfx +credentials.json +*.token +.netrc +.mcp.json + +# Operator local state +STATE.md + +# Lifecycle contents stay local; contracts/templates remain tracked 00_inbox/* !00_inbox/.gitkeep !00_inbox/AGENTS.md - -# Ingest Pipelines (ignore files being processed) 01_ingest/ready/* 01_ingest/queue/* 01_ingest/processing/* 01_ingest/renamed/* 01_ingest/rejected/* -01_ingest/audit-receipts/* - -# Knowledge Domains (ignore personal topics and the generated local index, keep public control files) 10_knowledge/* !10_knowledge/index.template.md !10_knowledge/AGENTS.md - -# Live State (ignore personal dashboards and logs, keep base control files) 20_live/* !20_live/index.md !20_live/AGENTS.md - -# Projects (ignore personal projects and the generated local index, keep public control files) 30_projects/* !30_projects/index.template.md !30_projects/AGENTS.md -# No project-content exceptions. Claim Audit Lab retired prototypes stay on disk -# under the project tree (private) — do not re-add a 30_projects whitelist. - -# Quant-markets-lab + private ops workstreams — never published from MainFrame. -# Project trees under 30_projects/* are already ignored; these patterns block -# operational scripts/orchestrators that live at MainFrame root. -bin/market-watcher -bin/markets-track -bin/price-snapshotter -bin/quant-portfolio-producer -bin/valuation-fetch -bin/income-application -bin/nd-discovery -bin/scheduling/ -bin/repo-radar-weekly -scripts/route_quant_markets_batch.py -scripts/*quant_market* -scripts/markets_registry.py -bin/*market-watcher* -tests/test_markets_registry.py -tests/test_quant_strategy_validation.py - -# Archive (ignore archived files, keep base control files) +40_operations/* +!40_operations/AGENTS.md +!40_operations/README.md +!40_operations/mainframe-process-eval/ +40_operations/mainframe-process-eval/* +!40_operations/mainframe-process-eval/README.md +!40_operations/mainframe-process-eval/AGENTS.md +!40_operations/mainframe-process-eval/src/ +!40_operations/mainframe-process-eval/src/** +!40_operations/mainframe-process-eval/tests/ +!40_operations/mainframe-process-eval/tests/** +!40_operations/mainframe-process-eval/methodology/ +!40_operations/mainframe-process-eval/methodology/** +!40_operations/mainframe-process-eval/workbench/ +40_operations/mainframe-process-eval/workbench/* +!40_operations/mainframe-process-eval/workbench/loop-eval/ +40_operations/mainframe-process-eval/workbench/loop-eval/* +!40_operations/mainframe-process-eval/workbench/loop-eval/README.md +!40_operations/mainframe-process-eval/workbench/loop-eval/THRESHOLDS.md +!40_operations/mainframe-process-eval/workbench/loop-eval/schema/ +!40_operations/mainframe-process-eval/workbench/loop-eval/schema/** +!40_operations/mainframe-process-eval/workbench/loop-eval/scripts/ +!40_operations/mainframe-process-eval/workbench/loop-eval/scripts/** +!40_operations/mainframe-process-eval/workbench/loop-eval/compositions/ +!40_operations/mainframe-process-eval/workbench/loop-eval/compositions/** +40_operations/**/__pycache__/ 90_archive/* !90_archive/index.md !90_archive/AGENTS.md -# Personal skills — methodology, workflow knowledge, and prompt craft stay local, never public -.agents/skills/writing-style/ -.claude/skills/writing-style/ -.agents/skills/prompt-creation/ -.claude/skills/prompt-creation/ -.agents/skills/skill-creation/ -.claude/skills/skill-creation/ -# Project-specific skill parked at root (canonical home: 30_projects/income-engine/) -.agents/skills/income-engine-process-control/ -.claude/skills/income-engine-process-control/ -scripts/eval_writing_style.py -scripts/ai_detect_check.py -scripts/eval_samples/ - -# Private project tooling stays local -bin/epistemic - -# Project-specific or archived root workflows (prefer 30_projects/<slug>/ or 90_archive/) -.context/workflows/rolling-second-brain-migration.md -.context/workflows/image-lab-operating.md -.context/workflows/pixel-sprite-generation.md - -# Additional dev caches and local agent client configs (prevent leakage of settings/telemetry paths) -.ruff_cache/ -.antigravity/ -.codex/ - -# Volatile generated artifacts from session/audit/synthesis tools (defense-in-depth under 20_live/*) -# These must never enter the public repo even if the broad 20_live/* rule is bypassed -20_live/last-handoff-draft.md -20_live/epistemic-audit/ -20_live/epistemic-audit/pending-review/sweep-*.md -20_live/epistemic-audit/pending-review/bridge-sweep-*.md - -# Workstation local state & config -workstation/settings.json -workstation/workstation/tasks.json -workstation/workstation/agents.json - -# Private outreach apparatus — kept local, never public -.context/workflows/boardy-meeting-lifecycle.md -.context/workflows/contact-website-gauge.md -.context/workflows/website-audit.md -bin/contact-gauge -bin/website-audit -scripts/contact_gauge.py -scripts/website_audit.py - -# Private methodology tooling — kept local, never public -EVAL_METHODOLOGY.md -.context/workflows/eval-methodology.md -.context/workflows/process-evaluation.md -.context/templates/eval-output.md -.context/templates/methodology-approach.md -bin/eval-registry -scripts/eval_registry.py -tests/test_eval_registry.py - -# Local voice extracts for writing-style idiolect (private) -.context/voice-corpus/ - -# determinism-cost sealed instrument (DISCLOSURE) -30_projects/determinism-cost/studies/S3-discriminative-instrument/sealed/corpus-*/ -30_projects/determinism-cost/studies/S3-discriminative-instrument/sealed/**/claims.json -30_projects/determinism-cost/studies/S3-discriminative-instrument/sealed/**/documents.json -30_projects/determinism-cost/studies/S3-discriminative-instrument/sealed/**/worlds.json - - -# determinism-cost sealed-derived materials -30_projects/determinism-cost/studies/S3-discriminative-instrument/pilot/blind-sheet-12.md -30_projects/determinism-cost/studies/S3-discriminative-instrument/agent-panel/items-blind.jsonl -30_projects/determinism-cost/outputs/ - -# Secrets — pattern-based, everywhere in the tree. -# Until 2026-08-10 no secret pattern existed here at all. Real API keys were safe -# only incidentally, because they happened to live under 30_projects/*, which is -# ignored wholesale. A .env in bin/, workstation/, 01_ingest/ or the repo root -# would have been committed. Protection that holds by coincidence does not -# generalize, and the coincidence is invisible until it stops being true. -.env -.env.* -!.env.example -*.key -*.pem -*.p12 -*.pfx -credentials.json -client_secret*.json -service-account*.json -*secret*.yaml -*secret*.yml -*.token -.netrc -30_projects/determinism-cost/studies/S3-discriminative-instrument/pilot/batch_01.html -.aider* +# Synthetic demo evidence stays local +examples/demo-mainframe/**/outputs/ +examples/demo-mainframe/**/raw-materials/ diff --git a/.grok/hooks/mainframe-telemetry.json b/.grok/hooks/mainframe-telemetry.json deleted file mode 100644 index d1cc3b7..0000000 --- a/.grok/hooks/mainframe-telemetry.json +++ /dev/null @@ -1,162 +0,0 @@ -{ - "hooks": { - "SessionStart": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "UserPromptSubmit": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "PreToolUse": [ - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "PostToolUse": [ - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "PostToolUseFailure": [ - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "PermissionDenied": [ - { - "matcher": "*", - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "Stop": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "StopFailure": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "Notification": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "SubagentStart": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "SubagentStop": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "PreCompact": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "PostCompact": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ], - "SessionEnd": [ - { - "hooks": [ - { - "type": "command", - "command": "${GROK_WORKSPACE_ROOT}/bin/workflow-event --client grok --pixel main-grok-build", - "timeout": 5 - } - ] - } - ] - } -} diff --git a/00_inbox/AGENTS.md b/00_inbox/AGENTS.md index 7dedb0d..d64bfab 100644 --- a/00_inbox/AGENTS.md +++ b/00_inbox/AGENTS.md @@ -39,6 +39,37 @@ it governed, silently.** Reserved: `AGENTS.md`, `README.md`, `index.md`, Deliberately checked at the gate rather than here. This folder is meant to be frictionless, and a capture zone that argues with you stops being used. +**The Escape did not work until 2026-08-27.** `hypothesis` was absent from +`ALLOWED_TYPES` in `01_ingest/minion.py`, so taking the documented escape route +produced `unsupported type: hypothesis` and stranded the file. Six documents +promised it — AGENTS.md principle 11, two workflow standards, `ingest-minion.md`, +this rule, and `bin/capture-validate`'s own fix message. None of them was the set +the gate reads. + +That is the exact failure principle 11 names: an escape valve that does not open +leaves fabricating the citation as the only way past the check. It is now in +`ALLOWED_TYPES` and deliberately not in `KNOWLEDGE_TYPES` — a hypothesis is a +valid capture and is not durable knowledge, so it moves through the pipeline and +stops before `10_knowledge/` for a human to decide. + +## 2b. A source lead is not a source. + +> **Binds:** captures written by `bin/deep-research-ingest leads` +> **Tier:** T2 (blocked) at the routing gate, same as rule 2 +> **Check:** `bin/capture-validate` R1 — a lead carries `asserted_doi` / +> `asserted_url` / `asserted_arxiv`, never `url` / `doi` / `authors` / `year` / +> `journal` / `venue` +> **Escape:** n/a. Retrieving the thing is the only way a lead becomes a capture. + +A lead records that a language model *asserted* a source exists. Promoting one +means retrieving it and recording `fetch_method` plus a `retrieval_receipt`. + +If the asserted identifier resolves to something unrelated to the claim, that is +a finding — record it and stop. Do not go looking for a better source to put in +its place. On 2026-08-19 the resolver did exactly that and attached a +cultural-studies interview to a report on agent runtimes. See +`local-only: 20_live/provenance/` for the dated ingest-audit record. + ## 3. Nothing here may be cited. > **Binds:** any agent answering a question diff --git a/01_ingest/minion.py b/01_ingest/minion.py index c37d775..dab53e1 100755 --- a/01_ingest/minion.py +++ b/01_ingest/minion.py @@ -17,7 +17,17 @@ ROOT = Path(__file__).resolve().parents[1] REQUIRED_KEYS = ("title", "domain", "type", "status", "source", "tags") -ALLOWED_TYPES = {"raw", "note", "live", "project", "decision"} +# `hypothesis` is the escape valve six documents already promise: AGENTS.md +# principle 11, `.context/workflows/deterministic-tool-standard.md`, +# `deterministic-automation-standard.md`, `ingest-minion.md`, `00_inbox/AGENTS.md` +# rule 2, and `bin/capture-validate`'s own fix message ("clear these fields and +# set type: hypothesis"). Until 2026-08-27 it was in none of them that mattered — +# this set — so following the documented escape produced `unsupported type: +# hypothesis` and the file was stuck. +# +# That is the failure mode principle 11 exists to prevent: an escape route that +# does not work leaves fabricating the citation as the only way past the gate. +ALLOWED_TYPES = {"raw", "note", "live", "project", "decision", "plan", "handoff", "hypothesis"} ALLOWED_STATUSES = { "queued", "skimmed", @@ -29,6 +39,11 @@ "archived", "parked", } +# What may route into 10_knowledge/. Deliberately narrower than ALLOWED_TYPES: +# `hypothesis` is a valid capture and is NOT durable knowledge, so it stays +# schema-valid, moves through the pipeline, and stops at this gate with the +# "not routable in v1" warning until a human decides otherwise. A thought with +# no source is welcome; it just does not get filed as something known. KNOWLEDGE_TYPES = {"note", "raw"} RAW_PDF_RE = re.compile( r"^(?P<date>\d{4}-\d{2}-\d{2})__(?P<domain>[^_]+)__raw__(?P<slug>.+)\.pdf$", @@ -358,6 +373,19 @@ def _infer_title(body_lines: list[str], path: Path) -> str: return slug_to_title(path.stem) +DEEP_RESEARCH_INDICATOR_RE = re.compile( + r"[\ue200\uE200]cite|[\ue200\uE200]filecite|\*\*Source appendix|## Source appendix", + re.IGNORECASE, +) + + +def is_deep_research_body(body: str, tags: list[str]) -> bool: + """Check if content exhibits ChatGPT deep research markers or tags.""" + if any(t.lower() in {"deep-research", "chatgpt-deep-research", "synthetic-report"} for t in tags): + return True + return bool(DEEP_RESEARCH_INDICATOR_RE.search(body)) + + def normalize_metadata( metadata: dict[str, Any], body: str, @@ -382,6 +410,11 @@ def normalize_metadata( if not isinstance(md.get("tags"), list): md["tags"] = [] + if is_deep_research_body(body, md["tags"]): + for req_tag in ("deep-research", "needs-audit"): + if req_tag not in md["tags"]: + md["tags"].append(req_tag) + body_links = extract_wikilinks(body) existing_links = md.get("links") if isinstance(md.get("links"), list) else [] seen: set[str] = set() @@ -784,14 +817,21 @@ def _stage_inbox_markdown( domain = parsed.metadata.get("domain") item_type = parsed.metadata.get("type") + project = parsed.metadata.get("project") is_known_knowledge_target = ( isinstance(domain, str) and domain in domains and isinstance(item_type, str) and item_type in KNOWLEDGE_TYPES ) + is_known_project_plan = ( + isinstance(item_type, str) + and item_type in {"plan", "handoff"} + and isinstance(project, str) + and (self.root / "30_projects" / project).is_dir() + ) - if was_strict_valid and is_known_knowledge_target: + if was_strict_valid and (is_known_knowledge_target or is_known_project_plan): md = normalize_metadata( parsed.metadata, body, @@ -840,6 +880,12 @@ def _stage_inbox_markdown( source.name, force_skimmed=True, ) + if is_deep_research_body(body, md["tags"]) and severity == "info": + message = ( + "normalize frontmatter and stage for agent enrichment (deep research " + "report detected; run `bin/deep-research-ingest leads` to classify it as " + "inference and file its asserted sources — CAL is not in this path)" + ) target = self.ready / source.name if target.exists(): result.add( @@ -944,8 +990,24 @@ def _route_markdown( self._reject(source, result, apply, f"invalid frontmatter: {exc}") return - domain = metadata["domain"] - item_type = metadata["type"] + domain = metadata.get("domain", "") + item_type = metadata.get("type", "") + project = metadata.get("project") + + # Project plan routing + if item_type == "plan" and isinstance(project, str) and project: + project_dir = self.root / "30_projects" / project + if project_dir.is_dir(): + target = project_dir / "plans" / source.name + if target.exists(): + result.add("blocked", source, target, "project plan destination already exists", "error") + return + result.add("route", source, target, f"route plan markdown to 30_projects/{project}/plans") + if apply: + target.parent.mkdir(parents=True, exist_ok=True) + shutil.move(str(source), target) + return + if domain not in domains: self._reject(source, result, apply, f"unknown knowledge domain: {domain}") return diff --git a/10_knowledge/AGENTS.md b/10_knowledge/AGENTS.md index 37156c9..86ff03b 100644 --- a/10_knowledge/AGENTS.md +++ b/10_knowledge/AGENTS.md @@ -8,8 +8,8 @@ Every rule below declares four things. **Escape** is not a loophole: it is the named, cheap, non-penalized way to comply when you cannot meet the letter of the rule. A rule without one manufactures violations, because an agent that cannot comply and cannot honestly fail will produce something that *looks* like -compliance. See -[every-rule-needs-an-honest-failure-path](agents/2026-08-10__agents__note__every-rule-needs-an-honest-failure-path.md). +compliance. See root `AGENTS.md` principle 11 (degraded-mode and escape-valve discipline) +and `.context/workflows/deterministic-tool-standard.md`. Tiers: **T0** advisory · **T1** detected · **T2** blocked · **T3** reconciled. @@ -59,15 +59,16 @@ evidence, and 14 simultaneously carried `needs-audit`. > **Binds:** anything reading this folder, including MindGraph > **Tier:** T1 (detected) > **Check:** `mindgraph` attaches `provenance_warning` to every chunk of a -> quarantined document, not only the first +> quarantined document, not only the first, held by +> `mindgraph/tests/test_citation_trust.py` > **Escape:** re-source the claim against real literature and write a new note. > The quarantined body may well be correct; it simply has no source behind it. ## 5. Raw bodies are immutable. Frontmatter may change. > **Binds:** any edit to a `type: raw` file -> **Tier:** **T0 (advisory). Nothing checks this.** -> **Check:** none +> **Tier:** T0 (advisory) +> **Check:** none — nothing enforces this today, said plainly > **Escape:** n/a Stated honestly rather than dressed up. Status changes, tags, links and appended diff --git a/20_live/AGENTS.md b/20_live/AGENTS.md index 5f86db5..c9921d1 100644 --- a/20_live/AGENTS.md +++ b/20_live/AGENTS.md @@ -9,3 +9,9 @@ 3. **Verify Before Promotion:** High-risk domain data (finance) must be verified against source documents before updating the "Compiled Truth". 4. **Retention classes (ADR-045):** Follow `.context/live-retention.md`. Do not bulk-index this folder into MindGraph. Treat derived projections as rebuildable; append-only evidence as archive-capable; high-volume ops as capped, never knowledge. 5. **Projects index apply:** Use `bin/mindgraph-projects-apply` (plan → stage → promote). Do not “clean” live state by wiping evidence to make an index look healthy. +6. **Handoffs have a lifecycle, not a paragraph:** `20_live/handoffs/*.md` is the + **open** set; `20_live/handoffs/archive/` holds consumed and superseded ones. + Create with `bin/handoff emit`, close with `bin/handoff consume <id> --apply`. + Never hand-edit `status` or move files between the two by hand — transitions + stamp `consumed_on` / `superseded_by` and preserve the body byte-for-byte. + STATE.md's Current Handoff section is generated by `bin/handoff state-block`. diff --git a/20_live/index.md b/20_live/index.md deleted file mode 100644 index 5af2e41..0000000 --- a/20_live/index.md +++ /dev/null @@ -1,17 +0,0 @@ -# Live State Index - -Information that changes, expires, or needs repeated updates. - -Example domains (local only — not published): - -- Operational dashboards and status boards -- Active research watchlists -- Time-bounded experiment notes - -## Rule - -Do not let dashboards become archives. Promote durable insights to `10_knowledge/`. - -## Live databases (local) - -Live SQLite files and telemetry stay gitignored. Query tools, when installed, use separate live indexes from durable MindGraph knowledge/project databases. diff --git a/30_projects/AGENTS.md b/30_projects/AGENTS.md index 60a1303..bd4c181 100644 --- a/30_projects/AGENTS.md +++ b/30_projects/AGENTS.md @@ -1,217 +1,15 @@ -# 30_projects - Project Lifecycle Rules +# Projects — Local Rules -This directory contains active work with concrete outcomes. Project folders may include the full local workbench for that outcome, but the outer MainFrame repo treats project contents as private ignored state. +`30_projects/` holds bounded outcome work. Project slugs share one namespace +with `40_operations/`. Duplicate slugs fail closed. -## Project States (ADR-041 / ADR-046: evidence-based, dual-pool WIP) +Each project README is direct authority for identity and lifecycle state. +Generated indexes and retrieval results do not replace it. -`bin/sync-project-index --check` enforces these semantics and fails loudly -when a state contradicts observed activity (file mtimes + nested-repo -commits, bounded scan — the same derivation `bin/session-close --checkpoint` -uses). Self-reported `updated:` is display metadata only; activity truth -comes from evidence. +> **Binds:** agents creating or resuming a project +> **Tier:** T0 (advisory in the public reference implementation) +> **Check:** `tests/test_lifecycle_identity.py` +> **Escape:** if authority is missing, label the field unknown and stop rather than guessing -- `active`: Work is moving now. Requires activity evidence within the last - 14 days **and** a `next_action`. Dual-pool WIP (ADR-046): - - **Product seats:** at most **5** `active` projects with `wip_class: - product` (default). - - **Eval seats:** measurement / eval / instrument projects - (`wip_class: eval`, or slug ending in `-eval`, or known eval slugs) - may stay `active` **without consuming product seats**. - - **Anchor:** strategic hub (`wip_class: anchor`, default for - `income-engine`) stays `active` and consumes **neither** product nor - total seats — other projects are meant to feed it. - - **Total ceiling:** at most **10** product+eval projects may be `active`. - Activating past a cap means pausing one first — visibly. -- `paused`: Deliberately shelved; the healthy default for real-but-not-now - work. Requires a `next_action` reentry pointer. No evidence requirement. -- `planned`: Registered but not started; name the activation gate in - `next_action`. -- `blocked`: Waiting on an external dependency or unresolved decision. -- `suspended`: Indefinite hold, heavier than `paused` (whole system parked, - e.g. epistemic-research-system). -- `shipped`: The outcome is complete enough to preserve, but may still - receive maintenance. -- `trashed`: The project should leave active navigation and move to - `90_archive/` (use the archive-project workflow). - -Any other state string is a checker error. - -High-churn private projects may carry a **nested local git repo** at the -project root (ADR-042). The outer MainFrame repo ignores `30_projects/*`, so -nothing conflicts, and nested commits become the project's activity evidence. -Local-only — no remotes by default; any future remote must be private and pass -a leak-detection review first. - -## Required README Metadata -Each project folder must have a `README.md` with YAML frontmatter: - -```yaml ---- -title: "Project name" -domain: "Broad area" -type: "project" -status: "active" -project_state: "active" -wip_class: "product" # optional: product (default) | eval | anchor -goal: "Outcome this project is meant to produce" -next_action: "Single next step" -updated: "YYYY-MM-DD" -source: "local" -tags: [] ---- -``` - -`wip_class` is optional. When omitted: `income-engine` → `anchor`; slugs -ending in `-eval` and the known eval/instrument set (`claim-audit-lab`, -`skill-eval-workshop`, `scaffold-claims-study`, `reliability-eval-framework`, -`verified-done`) → `eval`; everything else → `product`. - -## Workbench Layout -Project folders contain both coordination records and the working project itself. The four coordination entries are required for tracked outcomes; lighter experiments may start with fewer (see create-project workflow for light mode + graduation). The rest exist as needed: - -```text -30_projects/<slug>/ - README.md # required — frontmatter above - log.md # required for full projects — append-only work log... - decisions.md # required for full projects - plans/ # required for full projects - ... - raw-materials/ # optional - outputs/ # optional - workbench/ # optional — nested repo... -``` - -**Synthesis & extraction (Fix 3):** After meaningful work, use `bin/extract-knowledge --project <slug> --domain <...> --write` (or the new audit-sweep synthesis signals) to push reusable lessons into `10_knowledge/`. Do not leave durable knowledge trapped in project workbenches. - -**Craft / trial-and-error research:** Use `.context/workflows/craft-research-loop.md` and `bin/craft-research-loop` for product bake-offs and stack trials (image lab, integration prototypes). Close every trial with keep|kill|iterate + proof index. Do **not** put craft smokes on `last-eval-action.md` unless promoted via experiment-loop. - -**Lab reports (all measured experiments):** Decision-bearing tests use the universal lab-report notebook (question, design, results, irregularities, limits, disposition). Scaffold with `bin/lab-report scaffold --project <slug> …`; convention in `.context/workflows/lab-report.md`. Eval-registry harvest still uses the metric YAML block (eval-output specialization). - -The `workbench/` directory may be a nested Git repository, a local source tree, or a collection of drafts and artifacts. A workbench may keep its own internal records (STATUS, DECISIONS, ADRs); the project-level files above stay the coordination surface and point into the workbench rather than duplicating it. Keep reusable lessons in `10_knowledge/` only after an explicit extraction step. - -`30_projects/index.md` is a generated local index and is ignored by Git because it can list private projects. The public repo keeps `30_projects/index.template.md` to document the shape without exposing the live project inventory. - -## Planning Standard - -All planning documents live in `plans/`. Flat plan files are fine for small or single-track work. Phased work uses `plans/phases/phase-<n>-<slug>.md`, one file per phase, following this template: - -```text -# Phase <N> — <Title> - -Status: planned | active | complete -Started: YYYY-MM-DD or — -Completed: YYYY-MM-DD or — - -## Goal -What the phase produces and why, naming the source of the work -(analysis, finding, decision, or MindGraph query). - -## Non-Negotiable Boundaries -Constraints the phase must not cross: contracts, dependencies, -scope exclusions. - -## Unit Stance -Unit ordering and rationale. Build in testable units; stop at each -green boundary before the next unit; do not stack untested units. - -## Unit Plan - -### Unit 1 — <Title> -Scope: one or two lines. - -- [ ] Task checkboxes - -Green boundary: - -- [ ] Unit-specific checks -- [ ] Standard verification chain (see Verification) green - -## Verification -The commands run at every unit boundary. - -## Tie-Off Review - -- [ ] All deliverables present and tested -- [ ] Planning changelog / project log updated -- [ ] Master plan (or README next_action) updated -- [ ] Handoff notes written below -- [ ] Any blocked item is explicit and does not hide behind a green - phase status - -## Handoff Notes -(written at tie-off) -``` - -Rules: -- The frame is fixed: `Goal` first; `Tie-Off Review` and `Handoff Notes` last. Phase-specific sections (fixture expectations, file maps, impact tables) may be inserted between `Unit Plan` and `Tie-Off Review`. -- Log planning changes in `plans/CHANGELOG.md` when the project keeps one (`YYYY-MM-DD | [scope] | description`, newest first); otherwise in the project `log.md`. -- Completed or historical plans are never rewritten to a newer template. The standard binds new and still-active plans only. -- **MindGraph Querying**: Prior to initializing any new project plan or phase, query both the Knowledge and Projects MindGraph indexes through the MindGraph Query Station when available, or with `bin/mindgraph query` plus `MINDGRAPH_DB_PATH="$HOME/.mindgraph/mainframe-projects.sqlite"` as the CLI equivalent. Document the query strings, durable-knowledge nominations, project-context nominations, weak/excluded hits, and files that still need source inspection in the plan's `Goal` or the project `log.md`. - -## Delegation Packet Standard - -When a prepared implementation task is delegated to a local agent, store its -reviewed contract at: - -```text -30_projects/<slug>/plans/task-packets/<task-id>.md -``` - -Start from `.context/templates/task-packet.md`. A packet is written or refined -by a frontier model and reviewed by the operator before its status becomes -`ready`. It must resolve scope, implementation choices, acceptance criteria, -external verification, and stop conditions. Run state and outcomes belong in -evaluation receipts, never in the ready packet. - -Validate and compile packets with: - -```bash -bin/task-packet validate <packet-path> -bin/task-packet validate --require-ready <packet-path> -bin/task-packet compile -``` - -Only `ready` packets may execute. The compiled -`30_projects/task_packets_manifest.json` is generated local state for runners -and the workstation; the compiler rejects changes to a previously compiled -ready contract. Retire it or create a new task id instead. The Markdown packet -remains the source of truth. See -`.context/workflows/delegate-local-task.md`. - -## Local Coder Capabilities & Delegation Limits - -Based on empirical runs in `agent-harness-eval` (H1-packet multi-path hard screens 2026-07-18 + earlier matrices). Full per-model Can/Cannot: -`30_projects/agent-harness-eval/outputs/CAPABILITIES_CARD.md`. -Usage wishlist: `30_projects/agent-harness-eval/outputs/MAINFRAME_USAGE_ROADMAP.md`. -**Real-tree dispatch** only for tuples in `workstation/server/graduated-tuples.json` (currently `aider` × `local-qwen25-coder-14b` × `mechanical-edit`). - -### 1. Model Routing Matrix -- **Multi-File Coordination** (up to 4 files, real-tree): Route to **graduated** `local-qwen25-coder-14b` unless operator adds another tuple. Hard-screen evidence also supports **`local-qwen35-9b`** for *isolated/pilot* multi-file (8/8, two-file 2/2, lighter RAM) — **not** auto-graduated for live trees. -- **Do not multi-file:** `local-qwen3-14b` (hard two-file 0/2; H2 does not rescue; historical false-completion risk). Same restriction for Gemma-class / devstral / mistral-small on hard suite. -- **Efficiency multi-file (pilot):** Prefer `local-qwen35-9b` when RAM is tight; prefer `local-qwen25-coder-14b` for default reliability / graduated path. -- **Large context / search-heavy reference:** Prefer `local-qwen35-9b` or `local-qwen3-14b` for *read* load; keep editable surface single-file unless profile is 14b-coder or pilot 9b multi-file. Qwen2.5-coder remains fragile under huge read-only context — keep refs minimal. -- **Constrained dry prose / filler:** Prefer `local-qwen35-9b`. **Do not** use `local-qwen25-coder-14b` for P1-style filler (0 auto-clean on fixture-safety). -- **Claim extraction:** Prefer MLX Llama-3.1-8B + light scaffold (`format_only`); see capabilities card — not the coding profiles. - -### 2. Scope Boundaries -- **Editable Limit:** Maximum of **4 files** for `local-qwen25-coder-14b` (and for pilot `local-qwen35-9b` multi-file). **1 file** for `local-qwen3-14b` and other non-multi-file profiles. -- **Commit Restriction:** Local agents are strictly prohibited from creating Git commits. The operator maintains exclusive authority over git tree state and merges. - -### 3. Context Limits -- **Read-Only Volume:** For `local-qwen25-coder-14b`, reference context files (`read_only_files`) must be minimal and clean (under 100 lines total). -- **Aider Warning:** Stop and split the task immediately if Aider reports that the estimated context exceeds the model's limit. - -### 4. Verification Requirements -- Every packet **must** declare at least one deterministic verification command. -- Runs are accepted only when external verification commands pass with exit code `0`. -- If verification fails, a single fresh-context repair run (`H2-repair` logic) may be performed with the error output and current diff appended. H2 helps **near-miss coders** (e.g. 30b residual); it does **not** rescue wrong model class on multi-file. - -## Agent Protocol -1. Create projects with the `create-project` workflow. -2. Prior to writing plans or task packets, query both MindGraph databases (Knowledge & Projects) for relevant prior context, patterns, or similar work. Preserve the output as a grouped `MindGraph Query Pass`; do not collapse durable knowledge and project status into one unlabelled summary. -3. Update a project's `README.md`, `log.md`, and `decisions.md` rather than copying status into multiple places. -4. Regenerate the local `30_projects/index.md` with `bin/sync-project-index --write`; do not hand-edit it. -5. When archiving, use the `archive-project` workflow so knowledge extraction and status cleanup happen first. -6. Preserve project history. Move or append; do not silently overwrite logs or decisions. -7. New phase plans follow the Planning Standard above; do not retrofit completed plans. -8. Delegate local-agent implementation only from a reviewed `ready` task packet; keep verification outside the executing agent. +Do not copy another operator's project contents into this directory. Use +`30_projects/index.template.md` as the navigation template. diff --git a/30_projects/index.template.md b/30_projects/index.template.md index bcf2a2f..e4d8309 100644 --- a/30_projects/index.template.md +++ b/30_projects/index.template.md @@ -12,4 +12,4 @@ bin/sync-project-index --write | Project | State | Goal | Next action | Updated | | --- | --- | --- | --- | --- | -| [Example Project](example-project/README.md) | active | Outcome this project is meant to produce. | Single next step. | YYYY-MM-DD | +| Example Project (`local-only: 30_projects/<slug>/README.md`) | active | Outcome this project is meant to produce. | Single next step. | YYYY-MM-DD | diff --git a/40_operations/AGENTS.md b/40_operations/AGENTS.md new file mode 100644 index 0000000..c2c5f0c --- /dev/null +++ b/40_operations/AGENTS.md @@ -0,0 +1,66 @@ +# 40_operations - Local Rules + +> [!WARNING] +> This folder holds standing coordination systems. A path here is not +> evidence that a system is active, healthy, focused, approved, or safe to +> automate. + +Rules declare **Binds / Tier / Check / Escape**. The Escape is the named, cheap, +non-penalized way to comply when the required authority or evidence is absent. + +Tiers: **T0** advisory · **T1** detected · **T2** blocked · **T3** reconciled. + +--- + +## 1. Keep physical admission bounded. + +> **Binds:** any agent proposing to create or move another record under `40_operations/` +> **Tier:** T1 (detected) +> **Check:** `bin/work-inventory --check` +> **Escape:** leave the prospective entity in its current canonical location and record the unmet review gate instead of expanding the admitted set. + +Admission of one operation does not migrate other projects automatically. +MainFrame Process Evaluation is the portable public example. No further move +is implicit in an operation classification, a completed cycle, or a successful +check. + +## 2. Admit operations only with explicit identity and direct authority. + +> **Binds:** any agent or operator proposing an operation under this directory +> **Tier:** T2 (blocked on duplicate or invalid typed identity) +> **Check:** `bin/work-inventory --check` and `tests/test_lifecycle_identity.py` +> **Escape:** retain the existing record unchanged, or label the proposed identity `UNKNOWN`; do not create a duplicate or infer classification from recurrence or path. + +An admitted operation needs a stable slug, its own authority `README.md`, an +explicit `record_type: operation`, and the classification fields in this +directory's README. A slug must be unique across `30_projects/` and +`40_operations/`; a copy, alias, or symlink cannot act as a migration. + +## 3. Preserve authority and keep projections non-authoritative. + +> **Binds:** any reader or writer of an operation coordination record +> **Tier:** T0 (advisory) +> **Check:** none +> **Escape:** report the field as `UNKNOWN` and consult the direct authority; do not fill it from retrieval, a generated index, a schedule, or a past handoff. + +Each operation README owns its coordination fields. Local focus authority, +when present, remains `local-only: 20_live/focus/`. Ledgers, receipts, and +nested repositories retain their own authority. Retrieval results are +nominations, never proof. + +## 4. Keep WIP and lifecycle semantics orthogonal to location. + +> **Binds:** any agent or operator assessing an operation's lifecycle or WIP effect +> **Tier:** T1 (detected) +> **Check:** `bin/work-inventory --check` +> **Escape:** leave the assessment unresolved and preserve the current WIP/focus decision rather than granting an exemption because the record is an operation. + +`wip_class`, not `record_type` or folder location, determines seat treatment. +Finishing one recurring cycle is not closure; pause, decommission, and archive +actions require their own explicit authority. + +## Verification + +Run `bin/contract-lint --file 40_operations/AGENTS.md` after changing these +local rules. The linter verifies contract structure; it does not establish +operation health. diff --git a/40_operations/README.md b/40_operations/README.md new file mode 100644 index 0000000..34617a6 --- /dev/null +++ b/40_operations/README.md @@ -0,0 +1,90 @@ +--- +title: "MainFrame Operations" +domain: "knowledge-systems" +type: "lifecycle" +status: "active" +source: "public-export-transform" +tags: ["mainframe", "operations", "lifecycle"] +structural_type: "project-entry" +lifecycle_scope: "lifecycle" +owner_surface: "40_operations/" +authority: "operating-policy" +privacy: "public-safe" +volatility: "stable" +source_of_truth: true +update_rule: "replace-with-review" +verification: + - "bin/contract-lint --file 40_operations/AGENTS.md" +related_surfaces: + - "AGENTS.md" + - "30_projects/AGENTS.md" + - "DECISIONS.md" +do_not_use_for: + - "automatic project migration or lifecycle classification" + - "focus, health, approval, or completion authority" +--- + +# MainFrame Operations + +`40_operations/` is the lifecycle home for standing control planes, recurring +programs, and management systems. + +This public tree includes the operations-layer contracts and the portable +MainFrame Process Evaluation engine. Other admitted operation trees, raw +evidence, and local receipts remain local-only unless explicitly allowlisted. + +The folder separates a continuing loop from a bounded project outcome; it +does not make any operation healthier, focused, approved, or exempt from +WIP limits. + +## Scope and authority + +This folder owns the coordination entrypoint for an admitted operation. Each +operation's own `README.md` is its direct authority for identity, lifecycle +state, goal, next action, classification, and WIP class. Existing ledgers, +receipts, schedules, external systems, and nested repositories remain their +own authorities. Local focus authority, when present, lives under +`local-only: 20_live/focus/`. + +MindGraph is retrieval only: a result may nominate a coordination surface but +cannot establish lifecycle, WIP, focus, health, approval, or completion. + +## Admission and identity + +Admission of one operation does not migrate other projects automatically. +A recurring loop is not an operation merely because it repeats. Further +admission requires an explicitly reviewed decision. + +An admitted operation uses a stable slug and direct `README.md`, with at +least these classification fields in addition to its lifecycle fields: + +```yaml +record_type: operation +work_kind: control_plane | relationship_pipeline | research_program | evaluation_program | product_asset | experiment | external_coordinator +portfolio_role: primary | support | maintenance | waiting | parked +authority_mode: local_coordination | nested_repo | external_workspace +wip_class: product | eval | anchor +``` + +The slug is a cross-lifecycle identity. It must be unique across +`30_projects/` and `40_operations/`; two live copies, aliases, or symlinked +copies do not constitute a migration. Folder location never determines WIP, +focus, health, approval, or lifecycle state. + +## Lifecycle semantics + +Operations use explicit lifecycle states rather than the fact that a recurring +cycle exists. `active` requires an explicit next action and current evidence; +`paused`, `blocked`, and `suspended` state the different ways a loop is not +running; `planned` names its activation gate. An operation is not `shipped` +because one cycle finished. + +`wip_class` remains orthogonal to folder location. + +## Update discipline + +Keep durable local rules in [AGENTS.md](AGENTS.md). Put status and cycle +history in admitted operation records or their append-only ledgers, and record +accepted architecture changes in root [DECISIONS.md](../DECISIONS.md). Do not +place secrets, unrelated raw evidence, mailbox content, or generated +projections here. diff --git a/40_operations/mainframe-process-eval/AGENTS.md b/40_operations/mainframe-process-eval/AGENTS.md new file mode 100644 index 0000000..7e4faa0 --- /dev/null +++ b/40_operations/mainframe-process-eval/AGENTS.md @@ -0,0 +1,52 @@ +# MainFrame Process Evaluation — Local Rules + +> [!WARNING] +> This operation evaluates MainFrame processes. A path here is not proof that +> a process is healthy, that an evaluation passed, or that Git contains the +> evaluation history. + +Rules declare **Binds / Tier / Check / Escape**. Tiers: **T0** advisory · +**T1** detected · **T2** blocked · **T3** reconciled. + +--- + +## 1. Keep portable engine and local evidence distinct. + +> **Binds:** any agent adding files under `40_operations/mainframe-process-eval/` +> **Tier:** T2 (blocked from tracking evidence) +> **Check:** `tests/test_operations_git_boundary.py`; `.gitignore` default-deny allowlist +> **Escape:** leave evidence, receipts, outputs, raw-materials, logs, and +> transcripts untracked; do not force-add them to private Git + +Private `mainframe-live` tracks only the allowlisted portable surfaces +(README, AGENTS, `src/`, `tests/`, `methodology/`, loop-eval workbench). +Git is not the evaluation archive. + +## 2. Depend downward on MainFrame lifecycle substrate. + +> **Binds:** MPE portable source under `src/` +> **Tier:** T1 (detected) +> **Check:** `tests/test_lifecycle_substrate_isolation.py` +> **Escape:** call `lifecycle_identity` / `lifecycle_runtime` / `migration_lease`; +> do not make those modules import this operation + +## 3. Fail closed when direct operation authority is missing. + +> **Binds:** `bin/process-eval` and MPE writers +> **Tier:** T2 (blocked) +> **Check:** `40_operations/mainframe-process-eval/tests/test_process_eval.py` +> **Escape:** report `status: blocked` / missing identity; do not guess a path + +## 4. Do not treat this operation as a standalone GitHub owner. + +> **Binds:** agents proposing a `mainframe-process-eval` repository +> **Tier:** T0 (advisory) +> **Check:** none +> **Escape:** keep the operation inside private `mainframe-live`; evidence stays local + +## Verification + +```bash +uvx --with pytest pytest tests/test_operations_git_boundary.py tests/test_lifecycle_substrate_isolation.py 40_operations/mainframe-process-eval/tests -q +bin/contract-lint --file 40_operations/mainframe-process-eval/AGENTS.md +``` diff --git a/40_operations/mainframe-process-eval/README.md b/40_operations/mainframe-process-eval/README.md new file mode 100644 index 0000000..a2d84d7 --- /dev/null +++ b/40_operations/mainframe-process-eval/README.md @@ -0,0 +1,43 @@ +--- +title: "MainFrame Process Evaluation" +domain: "knowledge-systems" +type: operation +status: "active" +record_type: operation +work_kind: "evaluation_program" +portfolio_role: "maintenance" +authority_mode: "local_coordination" +wip_class: "eval" +goal: "Evaluate MainFrame operating processes with repeatable baselines without publishing private evidence." +updated: "2026-09-10" +source: "public-export-transform" +tags: ["mainframe", "evaluation", "synthetic"] +--- + +# MainFrame Process Evaluation + +This admitted operation is the portable self-evaluation engine for MainFrame. + +There is no standalone GitHub owner. Generic lifecycle substrate remains +MainFrame-owned. Raw evaluation evidence is local-only. Git does not contain +complete evaluation history. + +## Portable implementation + +- [`src/`](src/) — evaluator CLI implementation +- [`tests/`](tests/) — portable engine tests +- [`methodology/methodology-approach.md`](methodology/methodology-approach.md) +- [`workbench/loop-eval/`](workbench/loop-eval/) — schema and scorer + +`bin/process-eval` is a repository-level shim that delegates here. + +## Local-only evidence + +`local-only: outputs/`, `local-only: raw-materials/`, receipts, logs, and +campaign artefacts are not part of the public tree. Synthetic demos must be +labelled synthetic. + +## What this public surface is not + +It is not a report of any private MainFrame's quality. It does not include +private migration archaeology. diff --git a/40_operations/mainframe-process-eval/methodology/methodology-approach.md b/40_operations/mainframe-process-eval/methodology/methodology-approach.md new file mode 100644 index 0000000..676d8ad --- /dev/null +++ b/40_operations/mainframe-process-eval/methodology/methodology-approach.md @@ -0,0 +1,97 @@ +--- +title: "Methodology approach — read first in new session" +domain: "knowledge-systems" +type: "note" +status: "active" +source: "G43–G45 synthesis session 2026-06-20" +tags: ["methodology", "handoff", "eval-planning"] +updated: "2026-06-20" +--- + +# ⚠️ Methodology approach — MainFrame Process Eval + +**New session:** Read this before treating eval outputs as causal proof of process changes. Playbook: `10_knowledge/knowledge-systems/methodology/2026-06-20__knowledge-systems__note__scientific-method-experiment-design-synthesis.md`. Contract: `.context/workflows/process-evaluation.md`. + +## Current posture (2026-06-20 eval) + +- 168 tests green; duration coverage ~75% (up from ~20%). +- `workflow-report` now separates tool failures (79) vs StopFailures (410) vs policy blocks (0). +- **Next slice:** Claude StopFailure investigation (410 events, 2 sessions) — need safe reason/kind field if hook exposes it. + +## Study type: observational friction radar + +This is **not** an RCT. Use language: *associated with*, *trend*, *signal* — not *caused*. + +| Dimension | Use for | +|-----------|---------| +| Correctness / safety | Tests, dry runs, blockers | +| Provenance | Raw preservation, append-only | +| Throughput | Inbox/ingest counts | +| Friction | Repeated manual steps, failures | +| Observability | Tags, duration, pairing | +| Improvement readiness | Bugs, process gaps, workflow/skill candidates, parked questions | + +## Cadence (keep) + +1. Baseline before changes +2. ≤2 improvement slices at a time +3. Re-run after change +4. Repeat after 1 week or 5 sessions + +## Promotion gate (G45) + +Promote pattern only when: 3+ real tasks (or high-risk mandatory path), recognizable I/O, clear layer, success/boundary/failure cases evaluable, no duplicate workflow. + +## Improvement tracking method + +Use `improvement-backlog/items.md` for MainFrame process findings that might +turn into a bug fix, automation, workflow, skill, documentation update, +project-local experiment, or explicit no-action decision. + +Treat every item as a nomination until triaged. Do not promote or implement +from a single attractive observation unless it is a high-risk mandatory path. + +Process-eval finding classes: + +| Class | Meaning | +|-------|---------| +| `code-defect` | Implementation does not match the contract. | +| `process-gap` | The contract lacks a needed step or boundary. | +| `adoption-gap` | A sound workflow exists but is not being used. | +| `telemetry-gap` | Current signals cannot support the conclusion. | +| `intentional-backlog` | Queued work is expected, not a defect. | + +Practical backlog categories: + +| Category | Use for | +|----------|---------| +| `bug` | A broken deterministic check, parser, script, or stale state condition. | +| `automation-candidate` | Repeated deterministic work that might belong in `bin/`. | +| `workflow-candidate` | Operator-driven sequence using existing tools. | +| `skill-candidate` | Repeated agent judgment, domain rules, or tool strategy. | +| `documentation-gap` | Missing or stale guidance, handoff, or boundary text. | +| `research-question` | Useful but not ready for implementation or promotion. | +| `declined` | Explicitly not worth promoting now; keep the rationale visible. | + +Every backlog item needs source evidence: eval output, MindGraph query, +telemetry summary, command output, project log, or decision reference. If the +evidence is only a retrieval nomination, mark it as such and inspect the source +before selecting the item. + +## StopFailure slice frame + +1. **Observation** — concentration in `main-claude-pixel`, 2 sessions +2. **Hypothesis** — specific hook/session pattern (document, don't assume) +3. **Instrumentation** — `stop_failure_kind` enum if payload allows +4. **Re-baseline** — one change only, narrow time window + +## Do not + +- Claim a workflow change "improved throughput X%" without a single-variable window and before/after note. +- Copy prompts or private telemetry into tracked reports. + +## Related + +- Latest: `local-only: outputs/2026-06-20-evaluation.md` +- `10_knowledge/knowledge-systems/methodology/2026-06-21__knowledge-systems__note__eval-sample-size-and-significance.md` (§3.4) +- `10_knowledge/knowledge-systems/methodology/2026-06-21__knowledge-systems__note__eval-cross-project-trends-synthesis.md` (§3.1, §5 portfolio rollup) diff --git a/40_operations/mainframe-process-eval/src/__init__.py b/40_operations/mainframe-process-eval/src/__init__.py new file mode 100644 index 0000000..19562fc --- /dev/null +++ b/40_operations/mainframe-process-eval/src/__init__.py @@ -0,0 +1,6 @@ +"""Portable MainFrame Process Evaluation engine. + +Owned by the `mainframe-process-eval` operation. Depends on MainFrame +lifecycle substrate (`lifecycle_identity`, `lifecycle_runtime`, +`migration_lease`). Must not be imported by that substrate. +""" diff --git a/40_operations/mainframe-process-eval/src/process_eval.py b/40_operations/mainframe-process-eval/src/process_eval.py new file mode 100755 index 0000000..bb702b4 --- /dev/null +++ b/40_operations/mainframe-process-eval/src/process_eval.py @@ -0,0 +1,293 @@ +#!/usr/bin/env python3 +"""Process-evaluation loop CLI — deterministic preflight + loop decision surface. + +Pairs with `.context/workflows/process-evaluation.md`. Implementation is +owned by the `mainframe-process-eval` operation. Does not replace +`bin/mainframe-doctor` (health vector) or `bin/eval-schedule` (scheduled suite). + +Commands: + preflight Run the standard observational check pack (non-mutating). + status Point at project surfaces and last eval artifacts. + close Print loop-decision template after an eval pass. +""" + +from __future__ import annotations + +import argparse +import json +import subprocess +import sys +from datetime import date +from pathlib import Path +from typing import Any + + +ROOT = Path(__file__).resolve().parents[3] +sys.path.insert(0, str(Path(__file__).resolve().parent)) +sys.path.insert(0, str(ROOT / "scripts")) +from lifecycle_identity import IdentityError # noqa: E402 +from runtime import mpe_writer, resolve_mpe # noqa: E402 + +WORKFLOW = ROOT / ".context" / "workflows" / "process-evaluation.md" +# Compatibility labels for callers that import this module. Every command +# below resolves the live path from the shared lifecycle identity instead of +# using these historical constants as authority. +PROJECT = ROOT / "30_projects" / "mainframe-process-eval" +OUTPUTS = PROJECT / "outputs" +BACKLOG = PROJECT / "improvement-backlog" / "items.md" + + +def live_project() -> Path: + return resolve_mpe(ROOT).path + + +def live_outputs() -> Path: + return live_project() / "outputs" + + +def live_backlog() -> Path: + return live_project() / "improvement-backlog" / "items.md" + + +def run_cmd(cmd: list[str], *, timeout: int = 180) -> subprocess.CompletedProcess[str]: + return subprocess.run( + cmd, + cwd=str(ROOT), + capture_output=True, + text=True, + timeout=timeout, + ) + + +def preflight_steps() -> list[tuple[str, list[str]]]: + py = sys.executable + project = live_project() + return [ + ("unittest", [py, "-m", "unittest", "discover", "-s", "tests", "-q"]), + ("ingest_minion_dry_run", [str(ROOT / "bin" / "ingest-minion"), "run", "--dry-run"]), + ("mindgraph_refresh_dry_run", [str(ROOT / "bin" / "mindgraph-refresh"), "--dry-run"]), + ("sync_project_index_check", [str(ROOT / "bin" / "sync-project-index"), "--check"]), + ("session_open_json", [str(ROOT / "bin" / "session-open"), "--json"]), + ("session_close_check", [str(ROOT / "bin" / "session-close"), "--check"]), + ("eval_schedule_check", [str(ROOT / "bin" / "eval-schedule"), "check"]), + ("workflow_report_7d", [str(ROOT / "bin" / "workflow-report"), "--days", "7", "--json"]), + ("mainframe_doctor_quick", [str(ROOT / "bin" / "mainframe-doctor")]), + ( + "loop_catalogue", + [ + py, + str( + project + / "workbench" + / "loop-eval" + / "scripts" + / "score_catalogue.py" + ), + ], + ), + ] + + +def cmd_preflight(args: argparse.Namespace) -> int: + try: + outputs = live_outputs() + except IdentityError as exc: + print(f"process-eval: BLOCKED — {exc}", file=sys.stderr) + return 1 + results: list[dict[str, Any]] = [] + worst = 0 + for name, cmd in preflight_steps(): + if args.skip_doctor and name == "mainframe_doctor_quick": + results.append({"name": name, "exit": None, "skipped": True}) + continue + if args.skip_unittest and name == "unittest": + results.append({"name": name, "exit": None, "skipped": True}) + continue + try: + proc = run_cmd(cmd, timeout=args.timeout) + code = proc.returncode + except subprocess.TimeoutExpired: + code = 124 + proc = None # type: ignore + except FileNotFoundError: + code = 127 + proc = None # type: ignore + # Tools that exit 1 when work remains / health is degraded — still "ran". + advisory = name in { + "session_close_check", + "mainframe_doctor_quick", + "sync_project_index_check", + "eval_schedule_check", + } + hard_fail = code not in (0, None) and not (advisory and code == 1) + results.append( + { + "name": name, + "exit": code, + "skipped": False, + "advisory": bool(advisory and code == 1), + "cmd": cmd, + "stdout_tail": (proc.stdout or "")[-400:] if proc else "", + "stderr_tail": (proc.stderr or "")[-400:] if proc else "", + } + ) + if hard_fail: + worst = max(worst, 1) + + if args.json: + print(json.dumps({"results": results, "ok": worst == 0}, indent=2)) + else: + print("process-eval preflight") + print(f"workflow: {WORKFLOW.relative_to(ROOT)}") + for r in results: + if r.get("skipped"): + print(f" - {r['name']}: SKIP") + continue + if r["exit"] == 0: + flag = "OK" + elif r.get("advisory"): + flag = f"ADVISORY exit={r['exit']}" + else: + flag = f"FAIL exit={r['exit']}" + print(f" - {r['name']}: {flag}") + print() + print("next: sample ≥3 real cases → score → ≤2 improvement slices → rerun") + print(" then: bin/process-eval close") + print(f"outputs: {outputs.relative_to(ROOT)}/") + if worst == 0: + print("preflight: OK (advisory non-zeros do not fail the pack)") + return worst + + +def cmd_status(args: argparse.Namespace) -> int: + try: + project = live_project() + outputs = project / "outputs" + backlog = project / "improvement-backlog" / "items.md" + except IdentityError as exc: + payload = {"ok": False, "error": str(exc)} + if args.json: + print(json.dumps(payload, indent=2)) + else: + print(f"process-eval status: BLOCKED — {exc}") + return 1 + latest = None + if outputs.is_dir(): + mds = sorted(outputs.glob("*.md"), key=lambda p: p.stat().st_mtime, reverse=True) + latest = mds[0] if mds else None + loop_index = project / "workbench" / "loop-eval" / "catalogue" / "INDEX.md" + payload = { + "project": str(project.relative_to(ROOT)), + "workflow": str(WORKFLOW.relative_to(ROOT)) if WORKFLOW.exists() else None, + "backlog": str(backlog.relative_to(ROOT)) if backlog.exists() else None, + "latest_output": str(latest.relative_to(ROOT)) if latest else None, + "loop_eval": ( + str(loop_index.relative_to(ROOT)) if loop_index.is_file() else None + ), + } + if args.json: + print(json.dumps(payload, indent=2)) + else: + print("process-eval status") + for k, v in payload.items(): + print(f" {k}: {v}") + return 0 + + +def cmd_close(args: argparse.Namespace) -> int: + """Emit loop-decision scaffold (operator fills after eval).""" + today = date.today().isoformat() + if args.write: + writer_context = mpe_writer(ROOT, "process-eval.close") + else: + try: + record = resolve_mpe(ROOT) + except IdentityError as exc: + print(f"process-eval close: BLOCKED — {exc}", file=sys.stderr) + return 1 + writer_context = None + + if writer_context is not None: + context = writer_context.__enter__() + record, lease = context + project = record.path + else: + project = record.path + lease = None + outputs = project / "outputs" + output_rel = f"{outputs.relative_to(ROOT)}/{args.run_id or f'{today}-…'}.md" + text = f"""# process-eval close — loop decision ({today}) + +## Pass identity +- eval_run_id: {args.run_id or f"{today}-process-eval"} +- question: {args.question or "(one evaluation question)"} +- output: {output_rel} + +## Checks +- [ ] Baseline captured before changes +- [ ] Deterministic pack rerun (`bin/process-eval preflight`) +- [ ] ≥3 sampled cases (normal / boundary / failure) when available +- [ ] Findings classified (code-defect | process-gap | adoption-gap | telemetry-gap | intentional-backlog) +- [ ] ≤2 improvement slices selected (or none) +- [ ] Metric extract + irregularities; `bin/eval-registry harvest` + +## Loop decision (pick one) +- [ ] **same_slice** — continue same process area next pass +- [ ] **new_question** — new observational question +- [ ] **promote_pattern** — pattern ready for bin/workflow/skill (record destination) +- [ ] **park** — blocked or low value; reason: … +- [ ] **handoff** — ownership moves to project: … + +## Destinations (if promote) +- bin / workflow / skill / agent / project-local / no-action + +## Next action +- … + +Workflow: .context/workflows/process-evaluation.md +""" + try: + if args.write: + out = outputs / f"{args.run_id or today}-process-eval-close.md" + outputs.mkdir(parents=True, exist_ok=True) + out.write_text(text, encoding="utf-8") + assert lease is not None + lease.record_write(out, action="process-eval-close") + print(f"wrote {out.relative_to(ROOT)}") + else: + print(text) + return 0 + finally: + if writer_context is not None: + writer_context.__exit__(None, None, None) + + +def main() -> int: + parser = argparse.ArgumentParser( + description="MainFrame process-evaluation loop CLI", + ) + sub = parser.add_subparsers(dest="cmd", required=True) + + p = sub.add_parser("preflight", help="Run non-mutating check pack") + p.add_argument("--json", action="store_true") + p.add_argument("--skip-unittest", action="store_true") + p.add_argument("--skip-doctor", action="store_true") + p.add_argument("--timeout", type=int, default=300) + p.set_defaults(func=cmd_preflight) + + s = sub.add_parser("status", help="Project + artifact pointers") + s.add_argument("--json", action="store_true") + s.set_defaults(func=cmd_status) + + c = sub.add_parser("close", help="Print or write loop-decision scaffold") + c.add_argument("--run-id", default="") + c.add_argument("--question", default="") + c.add_argument("--write", action="store_true", help="Write under the resolved MPE lifecycle record") + c.set_defaults(func=cmd_close) + + args = parser.parse_args() + return int(args.func(args)) + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/40_operations/mainframe-process-eval/src/runtime.py b/40_operations/mainframe-process-eval/src/runtime.py new file mode 100644 index 0000000..8d465a6 --- /dev/null +++ b/40_operations/mainframe-process-eval/src/runtime.py @@ -0,0 +1,39 @@ +"""MPE-owned convenience wrappers over MainFrame lifecycle substrate.""" + +from __future__ import annotations + +from contextlib import contextmanager +from pathlib import Path +from typing import Iterator + +from lifecycle_identity import LifecycleRecord +from lifecycle_runtime import lifecycle_writer, resolve_lifecycle +from migration_lease import MigrationLease + +SLUG = "mainframe-process-eval" + + +def resolve_mpe(root: Path | str) -> LifecycleRecord: + return resolve_lifecycle(root, SLUG, expected_record_type=None) + + +def mpe_path(root: Path | str, *parts: str) -> Path: + return resolve_mpe(root).path.joinpath(*parts) + + +@contextmanager +def mpe_writer( + root: Path | str, + holder: str, + *, + pause_token: str | None = None, + state_root: Path | str | None = None, +) -> Iterator[tuple[LifecycleRecord, MigrationLease]]: + with lifecycle_writer( + root, + SLUG, + holder, + pause_token=pause_token, + state_root=state_root, + ) as result: + yield result diff --git a/40_operations/mainframe-process-eval/tests/test_process_eval.py b/40_operations/mainframe-process-eval/tests/test_process_eval.py new file mode 100644 index 0000000..0b151a0 --- /dev/null +++ b/40_operations/mainframe-process-eval/tests/test_process_eval.py @@ -0,0 +1,131 @@ +"""Unit tests for bin/process-eval.""" + +from __future__ import annotations + +import argparse +import importlib.util +import json +import subprocess +import unittest +from importlib.machinery import SourceFileLoader +from pathlib import Path +from tempfile import TemporaryDirectory +import sys +from unittest.mock import MagicMock, patch + +ROOT = Path(__file__).resolve().parents[3] +LOADER = SourceFileLoader( + "process_eval", + str(ROOT / "40_operations" / "mainframe-process-eval" / "src" / "process_eval.py"), +) +SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) +assert SPEC and SPEC.loader +pe = importlib.util.module_from_spec(SPEC) +sys.modules[SPEC.name] = pe +SPEC.loader.exec_module(pe) + + +class TestProcessEval(unittest.TestCase): + def test_preflight_steps_structure(self): + steps = pe.preflight_steps() + self.assertTrue(len(steps) > 5) + for name, cmd in steps: + self.assertIsInstance(name, str) + self.assertIsInstance(cmd, list) + self.assertTrue(len(cmd) > 0) + + def test_cmd_status_runs(self): + args = argparse.Namespace(json=False) + ret = pe.cmd_status(args) + self.assertEqual(ret, 0) + + def test_cmd_status_loop_eval_absent(self) -> None: + with TemporaryDirectory() as td: + root = Path(td) + project = root / "40_operations" / "mainframe-process-eval" + project.mkdir(parents=True) + args = argparse.Namespace(json=True) + with ( + patch.object(pe, "ROOT", root), + patch.object(pe, "WORKFLOW", root / "missing.md"), + patch.object(pe, "live_project", lambda: project), + patch("builtins.print") as printer, + ): + ret = pe.cmd_status(args) + self.assertEqual(ret, 0) + payload = json.loads(printer.call_args[0][0]) + self.assertIsNone(payload["loop_eval"]) + + def test_cmd_status_loop_eval_present(self) -> None: + with TemporaryDirectory() as td: + root = Path(td) + project = root / "40_operations" / "mainframe-process-eval" + index = project / "workbench" / "loop-eval" / "catalogue" / "INDEX.md" + index.parent.mkdir(parents=True) + index.write_text("# catalogue\n", encoding="utf-8") + args = argparse.Namespace(json=True) + with ( + patch.object(pe, "ROOT", root), + patch.object(pe, "WORKFLOW", root / "missing.md"), + patch.object(pe, "live_project", lambda: project), + patch("builtins.print") as printer, + ): + ret = pe.cmd_status(args) + self.assertEqual(ret, 0) + payload = json.loads(printer.call_args[0][0]) + self.assertEqual( + payload["loop_eval"], + "40_operations/mainframe-process-eval/workbench/loop-eval/catalogue/INDEX.md", + ) + + def test_cmd_close_runs(self): + args = argparse.Namespace( + run_id="2026-08-18-test-run", + question="Is the process verified?", + decision="Accept current baseline", + write=False, + ) + ret = pe.cmd_close(args) + self.assertEqual(ret, 0) + + @patch.object(pe, "run_cmd") + def test_cmd_preflight_all_pass(self, mock_run): + mock_proc = MagicMock(spec=subprocess.CompletedProcess) + mock_proc.returncode = 0 + mock_proc.stdout = "OK" + mock_proc.stderr = "" + mock_run.return_value = mock_proc + + args = argparse.Namespace( + json=False, + fast=False, + stop_on_failure=False, + skip_doctor=False, + skip_unittest=False, + timeout=180, + ) + ret = pe.cmd_preflight(args) + self.assertEqual(ret, 0) + + @patch.object(pe, "run_cmd") + def test_cmd_preflight_failure_with_stop(self, mock_run): + mock_proc_fail = MagicMock(spec=subprocess.CompletedProcess) + mock_proc_fail.returncode = 1 + mock_proc_fail.stdout = "" + mock_proc_fail.stderr = "Error in step" + mock_run.return_value = mock_proc_fail + + args = argparse.Namespace( + json=False, + fast=False, + stop_on_failure=True, + skip_doctor=False, + skip_unittest=False, + timeout=180, + ) + ret = pe.cmd_preflight(args) + self.assertEqual(ret, 1) + + +if __name__ == "__main__": + unittest.main() diff --git a/40_operations/mainframe-process-eval/workbench/loop-eval/README.md b/40_operations/mainframe-process-eval/workbench/loop-eval/README.md new file mode 100644 index 0000000..29e78a0 --- /dev/null +++ b/40_operations/mainframe-process-eval/workbench/loop-eval/README.md @@ -0,0 +1,79 @@ +--- +title: "Loop Eval Workstation" +domain: "knowledge-systems" +type: "workbench" +status: "active" +project: "mainframe-process-eval" +updated: "2026-07-23" +tags: ["loop-eval", "catalogue", "process-eval"] +--- + +# Loop Eval Workstation + +Mini evaluation workbench inside **mainframe-process-eval** for cataloguing +MainFrame loops, underspecified “loop-shaped” processes, and promotion +candidates — plus thresholds and automated checks. + +## Layout + +```text +workbench/loop-eval/ + README.md ← you are here + THRESHOLDS.md ← promotion bar + testing system + schema/loop-entry.schema.json + catalogue/ + entries.json ← machine authority (tests read this) + entries.yaml ← twin for editors + INDEX.md ← human summary + 01-well-defined.md + 02-underspecified.md + 03-promotion-candidates.md + compositions/ + README.md ← bigger programs that chain loops + scripts/ + score_catalogue.py ← score + tier consistency (also used by tests) +``` + +## Three catalogue tiers + +| Tier | File | Meaning | +|------|------|---------| +| **A — well-defined** | `01-well-defined.md` | First-class loops meeting the threshold | +| **B — underspecified** | `02-underspecified.md` | Named or used as loops but missing contract pieces | +| **C — promotion candidates** | `03-promotion-candidates.md` | Repeated workflows/processes that *could* become loops | + +Compositions (e.g. source → archive) are **not** a fourth loop type; see +`compositions/`. + +## Operator commands + +```bash +# Score + tier consistency (exit 1 on violations) +python3 40_operations/mainframe-process-eval/workbench/loop-eval/scripts/score_catalogue.py + +# Same checks via unittest (repo root) +python3 -m unittest tests.test_loop_catalogue -v +``` + +## How to add or reclassify an entry + +1. Edit `catalogue/entries.json` (source of truth for tests). +2. Optionally sync `entries.yaml` via PyYAML for readability. +3. Fill checklist booleans honestly; do not invent surfaces. +4. Run `score_catalogue.py` — checks score/tier/surface/edge rules. +5. Update tier markdown + INDEX when classifications change. +6. Log a one-line note in project `log.md` when tier changes. + +## Relationship to durable docs + +| Surface | Role | +|---------|------| +| This workbench | Lab: inventory, scores, promotion experiments | +| `.context/loops/` (later) | Promoted durable catalogue for all agents | +| `improvement-backlog/items.md` | MPE/LOOP backlog items when action is selected | + +## Non-goals + +- Do not invent a fourth first-class “meta-loop” (see three-loop gaps 2026-07-13). +- Do not promote from a single attractive observation (methodology G45). +- Do not copy prompts or private telemetry into catalogue files. diff --git a/40_operations/mainframe-process-eval/workbench/loop-eval/THRESHOLDS.md b/40_operations/mainframe-process-eval/workbench/loop-eval/THRESHOLDS.md new file mode 100644 index 0000000..4eb8b13 --- /dev/null +++ b/40_operations/mainframe-process-eval/workbench/loop-eval/THRESHOLDS.md @@ -0,0 +1,89 @@ +--- +title: "Loop promotion thresholds and testing system" +domain: "knowledge-systems" +type: "workbench" +status: "active" +updated: "2026-07-23" +--- + +# Loop thresholds and testing system + +## Checklist (10 binary criteria) + +Every catalogue entry scores **0–10**. One point each: + +| # | id | Criterion | +|---|-----|-----------| +| 1 | `unit_of_work` | One pass is named and bounded | +| 2 | `entry_exit` | Clear entry and exit conditions | +| 3 | `close_or_loop_decision` | Explicit close and/or next-loop decision | +| 4 | `workflow_contract` | Workflow file under `.context/workflows/` | +| 5 | `skill_or_agent_procedure` | Skill and/or subagent procedure exists | +| 6 | `bin_cli` | Deterministic CLI under `bin/` | +| 7 | `dogfood_evidence` | ≥1 dated dogfood / eval receipt (or scheduled automation proof) | +| 8 | `promotion_destinations` | Named destinations (knowledge, project, registry, archive, …) | +| 9 | `anti_scope` | Explicit does-not-own / when-not-to-use | +| 10 | `eval_or_action_surface` | Action card, registry, proof index, or equivalent | + +## Tier thresholds + +| Tier | Score | Extra gates | +|------|-------|-------------| +| **A — well_defined** | **≥ 8 / 10** | Must have `unit_of_work`, `entry_exit`, `workflow_contract`, and (`bin_cli` **or** `skill_or_agent_procedure`). Must have `dogfood_evidence`. | +| **B — underspecified** | **4–7** **or** named as a loop while failing A gates | Claims loop-shaped behavior but missing contract pieces. | +| **C — promotion candidate** | **any**, usually **3–7** | Not yet a loop; repeated operator/agent process with promotion potential. Track `repeat_signal` (low/medium/high). | + +Reclassify when evidence changes — score is not permanent. + +## Promotion bar (C → B or A) + +A candidate may be **selected for implementation** only if **all** hold: + +1. **Repeat threshold:** `repeat_signal: high` **or** ≥ **3** independent real uses documented **or** high-risk mandatory path (safety/provenance). +2. **Non-duplication:** does not re-implement an existing tier-A loop’s question (route there instead). +3. **Layer fit:** destination is clear (workflow vs skill vs bin vs composition-only). +4. **Evaluable cases:** success, boundary, and failure cases can be stated in one page. +5. **Test plan:** how we will know promotion worked (CLI preflight, unittest, dogfood receipt). +6. **Owner:** a project or lifecycle surface that will maintain it. + +After implementation, re-score. **Promotion to tier A** requires meeting the tier-A table above, not just “we wrote a workflow.” + +## Testing system + +### Automated (every change to `entries.yaml`) + +| Check | Tool | +|-------|------| +| Schema: required fields, enums, checklist keys | `tests/test_loop_catalogue.py` + JSON Schema | +| Score == count of true checklist items | `score_catalogue.py` | +| Tier consistent with score + A gates | `score_catalogue.py` | +| Surface paths that are set must exist on disk | `test_loop_catalogue.py` | +| Unique `id` | tests | + +```bash +python3 -m unittest tests.test_loop_catalogue -v +python3 40_operations/mainframe-process-eval/workbench/loop-eval/scripts/score_catalogue.py +``` + +### Semi-automated / observational + +| Check | Cadence | +|-------|---------| +| Dogfood receipt still valid (path exists, not contradicted) | when reclassifying | +| Composition edges resolve to catalogue ids | when editing compositions | +| No fourth first-class loop without ADR | process-eval decision | + +### Manual gate (operator) + +- Approve promotion implementation (≤2 active process-eval slices). +- Accept ADR only if durable `.context/loops/` catalogue is published. + +## Composition rule (bigger “loops”) + +Compositions chain tier-A/B loops and workflows. They: + +- **do not** need a single mega-CLI on day one +- **must** list ordered stages and terminal conditions +- **must not** blur the three-loop questions (literature / measure / craft) + +See `compositions/README.md`. diff --git a/40_operations/mainframe-process-eval/workbench/loop-eval/compositions/README.md b/40_operations/mainframe-process-eval/workbench/loop-eval/compositions/README.md new file mode 100644 index 0000000..b11ce08 --- /dev/null +++ b/40_operations/mainframe-process-eval/workbench/loop-eval/compositions/README.md @@ -0,0 +1,71 @@ +--- +title: "Loop compositions (bigger programs)" +domain: "knowledge-systems" +type: "workbench" +status: "active" +updated: "2026-07-23" +--- + +# Compositions — bigger “loops” without a fourth loop type + +A **composition** is an ordered chain of catalogue entries (tier A/B loops + workflows). +It answers “how does X get from start to archive?” without merging literature, craft, and measurement into one god-procedure. + +## Rules + +1. Stages reference catalogue `id`s from `entries.json`. +2. Terminal stages are explicit (archive, stop, blocked). +3. Optional branches are labeled (craft vs experiment). +4. No new first-class loop unless tier-A gates + ADR. + +## Draft compositions (baseline) + +### `comp-research-lane-lifecycle` + +**Intent:** Source discovery through portfolio archive for one research lane. + +```text +cand-lane-intake-archive # intake (once) + ↓ +loop-research-lane × N # foundation → … (includes source-literature step) + ↓ +cand-research-handoff # when knowledge should hit a project + ↓ + ┌─ loop-craft ────────────────┐ + ├─ loop-experiment ───────────┤ optional application branches + └─ implement / operator ──────┘ + ↓ +cand-lane-intake-archive # archive when lane stop condition met +``` + +**Note:** `cand-source-literature` and `loop-ingest-pipeline` run *inside* research-lane passes, not as peer outer stages. + +### `comp-weekly-control-plane` + +**Intent:** Honest scheduled health → action. + +```text +cand-eval-schedule-control # launchd daily/weekly + ↓ +loop-experiment (canary/triage) # portfolio action card + ↓ +loop-process-evaluation # optional deeper process slice +``` + +### `comp-standing-operation` (example) + +**Intent:** Recurring signals feed a standing operation. Admission of one +operation does not migrate other projects automatically. + +```text +cand-intake + ↓ +loop-operating-cycle + ↓ +(operation next_action) +``` + +## Next + +- Optional `compositions.yaml` with machine-checked stage ids. +- Link from durable `.context/loops/` when published. diff --git a/40_operations/mainframe-process-eval/workbench/loop-eval/schema/loop-entry.schema.json b/40_operations/mainframe-process-eval/workbench/loop-eval/schema/loop-entry.schema.json new file mode 100644 index 0000000..ecdc3fe --- /dev/null +++ b/40_operations/mainframe-process-eval/workbench/loop-eval/schema/loop-entry.schema.json @@ -0,0 +1,100 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "mainframe-loop-entry", + "title": "MainFrame loop catalogue entry", + "type": "object", + "required": [ + "id", + "title", + "tier", + "kind", + "question", + "unit_of_work", + "checklist", + "status", + "last_reviewed" + ], + "additionalProperties": true, + "properties": { + "id": { + "type": "string", + "pattern": "^(loop|wf|comp|cand)-[a-z0-9]+([a-z0-9-]*[a-z0-9])?$" + }, + "title": { "type": "string", "minLength": 3 }, + "tier": { + "type": "string", + "enum": ["well_defined", "underspecified", "candidate"] + }, + "kind": { + "type": "string", + "enum": ["loop", "workflow", "composition", "skill_process"] + }, + "question": { "type": "string" }, + "unit_of_work": { "type": "string" }, + "entry": { "type": "string" }, + "exit": { "type": "string" }, + "does_not_own": { + "type": "array", + "items": { "type": "string" } + }, + "surfaces": { + "type": "object", + "properties": { + "workflow": { "type": ["string", "null"] }, + "skill": { "type": ["string", "null"] }, + "bin": { "type": ["string", "null"] }, + "owner": { "type": ["string", "null"] } + } + }, + "edges": { + "type": "object", + "properties": { + "feeds": { "type": "array", "items": { "type": "string" } }, + "fed_by": { "type": "array", "items": { "type": "string" } } + } + }, + "checklist": { + "type": "object", + "required": [ + "unit_of_work", + "entry_exit", + "close_or_loop_decision", + "workflow_contract", + "skill_or_agent_procedure", + "bin_cli", + "dogfood_evidence", + "promotion_destinations", + "anti_scope", + "eval_or_action_surface" + ], + "additionalProperties": false, + "properties": { + "unit_of_work": { "type": "boolean" }, + "entry_exit": { "type": "boolean" }, + "close_or_loop_decision": { "type": "boolean" }, + "workflow_contract": { "type": "boolean" }, + "skill_or_agent_procedure": { "type": "boolean" }, + "bin_cli": { "type": "boolean" }, + "dogfood_evidence": { "type": "boolean" }, + "promotion_destinations": { "type": "boolean" }, + "anti_scope": { "type": "boolean" }, + "eval_or_action_surface": { "type": "boolean" } + } + }, + "score": { "type": "integer", "minimum": 0, "maximum": 10 }, + "repeat_signal": { + "type": "string", + "enum": ["low", "medium", "high", "n/a"] + }, + "status": { + "type": "string", + "enum": ["active", "draft", "parked", "declined", "superseded"] + }, + "evidence": { + "type": "array", + "items": { "type": "string" } + }, + "notes": { "type": "string" }, + "last_reviewed": { "type": "string", "pattern": "^\\d{4}-\\d{2}-\\d{2}$" } + } +} diff --git a/40_operations/mainframe-process-eval/workbench/loop-eval/scripts/score_catalogue.py b/40_operations/mainframe-process-eval/workbench/loop-eval/scripts/score_catalogue.py new file mode 100644 index 0000000..9c566c5 --- /dev/null +++ b/40_operations/mainframe-process-eval/workbench/loop-eval/scripts/score_catalogue.py @@ -0,0 +1,247 @@ +#!/usr/bin/env python3 +"""Score loop catalogue entries and enforce tier thresholds. + +Usage (from MainFrame root): + python3 40_operations/mainframe-process-eval/workbench/loop-eval/scripts/score_catalogue.py + python3 .../score_catalogue.py --json +""" + +from __future__ import annotations + +import argparse +import json +import sys +from pathlib import Path +from typing import Any + +ROOT = Path(__file__).resolve().parents[5] +CATALOGUE_JSON = ( + Path(__file__).resolve().parents[1] / "catalogue" / "entries.json" +) +CATALOGUE_YAML = ( + Path(__file__).resolve().parents[1] / "catalogue" / "entries.yaml" +) +CATALOGUE = CATALOGUE_JSON if CATALOGUE_JSON.exists() else CATALOGUE_YAML +CHECKLIST_KEYS = [ + "unit_of_work", + "entry_exit", + "close_or_loop_decision", + "workflow_contract", + "skill_or_agent_procedure", + "bin_cli", + "dogfood_evidence", + "promotion_destinations", + "anti_scope", + "eval_or_action_surface", +] + +# Tier A required checklist keys (in addition to score ≥ 8) +TIER_A_REQUIRED = ( + "unit_of_work", + "entry_exit", + "workflow_contract", + "dogfood_evidence", +) + + +def load_catalogue(path: Path = CATALOGUE) -> dict[str, Any]: + text = path.read_text(encoding="utf-8") + if path.suffix.lower() == ".json": + data = json.loads(text) + else: + try: + import yaml # type: ignore + except ImportError as exc: # pragma: no cover + raise SystemExit( + "entries.json missing and PyYAML not installed; " + "prefer catalogue/entries.json" + ) from exc + data = yaml.safe_load(text) + if not isinstance(data, dict) or "entries" not in data: + raise SystemExit(f"invalid catalogue: {path}") + return data + + +def score_entry(entry: dict[str, Any]) -> int: + cl = entry.get("checklist") or {} + return sum(1 for k in CHECKLIST_KEYS if cl.get(k) is True) + + +def tier_a_gates_ok(entry: dict[str, Any]) -> bool: + cl = entry.get("checklist") or {} + if not all(cl.get(k) is True for k in TIER_A_REQUIRED): + return False + # bin OR skill + if not (cl.get("bin_cli") or cl.get("skill_or_agent_procedure")): + return False + return score_entry(entry) >= 8 + + +def expected_tier(entry: dict[str, Any]) -> str: + """Suggest tier from score + gates (does not auto-rename candidates).""" + declared = entry.get("tier") + s = score_entry(entry) + if declared == "candidate": + return "candidate" + if tier_a_gates_ok(entry): + return "well_defined" + if s >= 4: + return "underspecified" + return "underspecified" + + +def validate_entry(entry: dict[str, Any], root: Path = ROOT) -> list[str]: + problems: list[str] = [] + eid = entry.get("id", "?") + cl = entry.get("checklist") or {} + for k in CHECKLIST_KEYS: + if k not in cl or not isinstance(cl[k], bool): + problems.append(f"{eid}: checklist missing bool {k}") + + computed = score_entry(entry) + if entry.get("score") is not None and int(entry["score"]) != computed: + problems.append( + f"{eid}: score field {entry.get('score')} != computed {computed}" + ) + + tier = entry.get("tier") + if tier == "well_defined": + if not tier_a_gates_ok(entry): + problems.append( + f"{eid}: tier well_defined but fails A gates/score " + f"(score={computed})" + ) + elif tier == "underspecified": + if tier_a_gates_ok(entry): + problems.append( + f"{eid}: tier underspecified but meets A gates — reclassify?" + ) + if computed < 4 and entry.get("status") != "draft": + problems.append( + f"{eid}: underspecified with score {computed} < 4 " + "(raise evidence or mark candidate/draft)" + ) + elif tier == "candidate": + pass + else: + problems.append(f"{eid}: unknown tier {tier!r}") + + surfaces = entry.get("surfaces") or {} + for key in ("workflow", "skill", "bin", "owner"): + rel = surfaces.get(key) + if not rel: + continue + # owner may be multi-token description + if key == "owner" and (" " in rel or "+" in rel): + continue + path = root / rel + if not path.exists(): + problems.append(f"{eid}: surface {key} path missing: {rel}") + + return problems + + +def validate_catalogue(data: dict[str, Any], root: Path = ROOT) -> list[str]: + problems: list[str] = [] + entries = data.get("entries") or [] + ids: list[str] = [] + for entry in entries: + if not isinstance(entry, dict): + problems.append("non-mapping entry") + continue + eid = str(entry.get("id", "")) + if not eid: + problems.append("entry missing id") + continue + if eid in ids: + problems.append(f"duplicate id: {eid}") + ids.append(eid) + problems.extend(validate_entry(entry, root=root)) + + # edge references should resolve when present + idset = set(ids) + for entry in entries: + if not isinstance(entry, dict): + continue + edges = entry.get("edges") or {} + for direction in ("feeds", "fed_by"): + for ref in edges.get(direction) or []: + if ref not in idset: + problems.append( + f"{entry.get('id')}: edge {direction} unknown id {ref}" + ) + return problems + + +def summary(data: dict[str, Any]) -> dict[str, Any]: + entries = [e for e in data.get("entries") or [] if isinstance(e, dict)] + by_tier: dict[str, list[str]] = { + "well_defined": [], + "underspecified": [], + "candidate": [], + } + scored = [] + for e in entries: + s = score_entry(e) + scored.append({"id": e.get("id"), "tier": e.get("tier"), "score": s}) + t = e.get("tier") or "candidate" + if t in by_tier: + by_tier[t].append(f"{e.get('id')} ({s}/10)") + return { + "n": len(entries), + "by_tier_counts": {k: len(v) for k, v in by_tier.items()}, + "by_tier": by_tier, + "scores": scored, + } + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--json", action="store_true") + parser.add_argument("--catalogue", type=Path, default=CATALOGUE) + parser.add_argument( + "--out", + type=Path, + help="optional path to write the JSON report (local-only; not a Git surface)", + ) + args = parser.parse_args() + + data = load_catalogue(args.catalogue) + # stamp computed scores into report only + for e in data.get("entries") or []: + if isinstance(e, dict): + e["_computed_score"] = score_entry(e) + + problems = validate_catalogue(data, root=ROOT) + summ = summary(data) + report = {"summary": summ, "problems": problems, "ok": not problems} + + if args.json or args.out: + payload = json.dumps(report, indent=2) + "\n" + if args.json or not args.out: + print(payload, end="") + if args.out: + args.out.parent.mkdir(parents=True, exist_ok=True) + args.out.write_text(payload, encoding="utf-8") + else: + try: + shown = args.catalogue.resolve().relative_to(ROOT) + except ValueError: + shown = args.catalogue + print(f"catalogue: {shown}") + print(f"entries: {summ['n']} tiers: {summ['by_tier_counts']}") + for tier, rows in summ["by_tier"].items(): + print(f"\n[{tier}]") + for row in rows: + print(f" - {row}") + if problems: + print(f"\nPROBLEMS ({len(problems)}):") + for p in problems: + print(f" - {p}") + else: + print("\nOK — tier/score/surface checks passed") + return 1 if problems else 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/AGENTS.md b/AGENTS.md index a89b069..ae9ea83 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -9,22 +9,55 @@ This system organizes knowledge by its **information lifecycle** first, and topi - `10_knowledge/`: Durable, slower-moving knowledge. - `20_live/`: Volatile personal and project state, active research. - `30_projects/`: Active work with outcomes. +- `40_operations/`: Standing coordination systems under the local pilot rules. - `90_archive/`: Preserved material without cluttering active navigation. +## Session Entry + +Run `bin/session-open --json` for workspace arrival. For project work, use +`--project <slug> --task "brief task description"`; use `--intent resume` to +resume the recorded focus. Add `--path <repo-relative-path>` for deeper rules. +Read every required content batch separately with `--read-batch N`, repeating +the routing arguments. Recover truncated reads before relying on them. +The command lists and serves context; it cannot verify reading or understanding. +Follow `.context/workflows/session-open.md` for the route's stopping point. + +> **Binds:** agents starting or resuming MainFrame work +> **Tier:** T0 (advisory) +> **Check:** none — invocation and reading are not enforced by the command +> **Escape:** if the command is unavailable, read `HARNESS.md` and the applicable directory contracts directly; report unresolved prerequisites before dependent work + ## Agent Behavior -1. **Harness contract:** Read `HARNESS.md` for task-category harness rules, dual MindGraph indexes, and local/cloud delegation boundaries. +1. **Harness contract:** For workspace arrival, read `HARNESS.md`'s **Session orientation** section. For project resume, implementation, evaluation, delegation, or other action, read the full harness and applicable directory contracts. Arrival may report the lifecycle map and recorded focus, then stops; project-specific diagnosis or recommendations require the resume route. This explicitly narrows the former unconditional full-harness reading requirement for arrival only (ADR-059). 2. **MindGraph Planning Hook:** Always query the dual MindGraph databases (`mainframe.sqlite` and `mainframe-projects.sqlite`) when planning or starting work on any project in `30_projects/` to leverage existing workspace context. Use the MindGraph Query Station when available, or the CLI equivalent, and preserve the two result groups with trust labels instead of blending them into one answer. 3. **MindGraph MCP (shared daemon):** Operational clients connect through the loopback shared daemon (`http://127.0.0.1:8000/mcp`) via `bin/mindgraph mcp-proxy` or a streamable-HTTP URL — never by spawning `serve-mcp` per session. Every MCP `query` / `graph_neighbors` call **must** pass `scope` = `knowledge` or `projects`. Do not invent `both`. Prefer CLI `bin/mindgraph query --db …` when MCP is unavailable. See ADR-049 and `.agents/skills/mindgraph-retrieval/SKILL.md`. -4. **Centralized Skills:** Use `.context/workflows/` for operator-driven sequences and `.agents/skills/` for reusable agent skills rather than duplicating instructions. -5. **Subagents:** Named subagent definitions live in `agents/` (e.g. `agents/ingest-agent.md`). Subagents are first-class collaborators; their roles, tools, guardrails, and procedures are defined there. -6. **Local Constraints:** Respect local `AGENTS.md` files in subdirectories—they contain overriding rules for sensitive or volatile data. -7. **Immutability:** Do not silently overwrite history. If a file is in `20_live`, use snapshots or append-only timelines. -8. **Provenance:** Preserve raw sources as evidence. Extracted text is a searchable working copy, not the source of truth. +4. **MindGraph engine edits:** The engine has one source tree — `30_projects/mindgraph/workbench/`, which is also the tree that publishes to GitHub. Root `mindgraph/` is a **promoted artifact**; editing it directly is drift by definition, and doing so through 2026-08 left the published MindGraph without the guardrails built after the fabricated-citations incident. Change the workbench, then run `bin/mindgraph-promote`. `bin/mindgraph-promote --check` and `tests/test_mindgraph_promotion.py` fail on drift. See `mindgraph/AGENTS.md`. +5. **Centralized Skills:** Use `.context/workflows/` for operator-driven sequences and `.agents/skills/` for reusable agent skills rather than duplicating instructions. +6. **Subagents:** Named subagent definitions live in `agents/` (e.g. `agents/ingest-agent.md`). Subagents are first-class collaborators; their roles, tools, guardrails, and procedures are defined there. +7. **Local Constraints:** Respect local `AGENTS.md` files in subdirectories—they contain overriding rules for sensitive or volatile data. +8. **Immutability:** Do not silently overwrite history. If a file is in `20_live`, use snapshots or append-only timelines. +9. **Provenance:** Preserve raw sources as evidence. Extracted text is a searchable working copy, not the source of truth. +10. **Execution Honesty:** Never simulate compute or mock operational completeness. Scripts, tools, and tests must physically invoke the underlying engines, models, or databases they represent or fail closed (`status: blocked`). Do not format evaluation tables or claim-audit verdicts in Markdown without a verifiable machine trace on disk. See `.context/workflows/deterministic-tool-standard.md`. +11. **Degraded-Mode & Escape Valve Discipline:** Complex systems run as broken systems (Cook 1998). Tools, scripts, and validation gates must never assume 100% operational availability or faultless inputs. When dependencies fail, tools must fail closed or report explicit degraded status. All validation gates must provide explicit, valid intermediate escape hatches (e.g. `type: hypothesis`, `status: queued`, `needs-audit`) so agents and operators are never pressured into deceptive workarounds or fabricated metadata to pass a check. Enforced via ADR-052 and `.context/workflows/deterministic-tool-standard.md`. +12. **Project Resume Reconstruction:** Treat casual project-resume language such as “take a look,” “where are we,” “read the handoff,” or “decide what to do next” as authorization to inspect, diagnose, and recommend only. Before proposing work, reconstruct the current coordination files, every nested Git repository and registered worktree, dirty/concurrent state, candidate and experiment identity, and direct test/receipt evidence. Handoffs, READMEs, plans, generated indexes, and MindGraph results are leads; they do not override current source, Git, manifests, or execution evidence. Report the reconstructed workstate before recommendations, and follow `.context/workflows/project-resume-and-candidate-lifecycle.md`. + +> **Binds:** any agent starting or resuming work in `30_projects/` +> **Tier:** T0 (advisory MainFrame-wide; a project-local contract may raise the tier) +> **Check:** none at the root; managed projects may provide candidate/experiment validators +> **Escape:** if direct authority is missing or conflicting, label the field `UNKNOWN` or the worktree `UNMANAGED`, remain read-only, and stop comparison or promotion rather than guessing + +13. **Repository reconciliation:** GitHub is authoritative for registered repository state. Query `bin/repo-reconcile` for derived drift; verify the actual repository before consequential operations. `DIRTY` and `DIVERGED` are observations, not license to rebase, force-push, or pick a side. See `.context/workflows/repo-reconciliation.md` and ADR-060. + +> **Binds:** agents inspecting or mutating GitHub-facing checkouts registered in `20_live/github-portfolio/registry.json` +> **Tier:** T0 (advisory) for reading the derived receipt; T2 (blocked) for unattended repair +> **Check:** `bin/repo-reconcile --check`; `tests/test_repo_reconcile.py` +> **Escape:** inspect `git -C <path> status` / `git fetch` directly; leave a checkout unregistered if it must not be classified ## Structural File Discipline - Structural files include operating contracts, local `AGENTS.md` files, `HARNESS.md`, decision records, workflows, skills, subagent definitions, templates, configs, indexes, manifests, hooks, scripts, and project/workbench files that define process or verification behavior. - Use `.context/templates/structural-file-profile.md` when creating or auditing structural files. Project-local files follow the same framework even when they are ignored or private. +- **Contract templates:** start new contracts from `.context/templates/contracts/`. Every normative rule declares **Binds / Tier / Check / Escape**, with tiers **T0** advisory · **T1** detected · **T2** blocked · **T3** reconciled. `T0` with `Check: none` is an honest answer; an inflated tier is not. Check the file before committing with `bin/contract-lint --file <path>`; `.githooks/pre-commit` runs `--changed` so an adopted contract cannot silently lose its enforcement to a blank line. Reference implementations: `00_inbox/AGENTS.md`, `10_knowledge/AGENTS.md`. - Put stable always-on rules in `AGENTS.md`; put long operator sequences in `.context/workflows/`; put repeated agent judgment in `.agents/skills/`; put specialized roles in `agents/`; put deterministic enforcement in `bin/`, scripts, hooks, config, or tests. - Keep root and lifecycle contracts durable. Put volatile status in `STATE.md`, `20_live/`, project `README.md`, or project `log.md`. - Record accepted architecture or workflow changes in `DECISIONS.md`; record project-only tradeoffs in the project's `decisions.md`. @@ -33,12 +66,56 @@ This system organizes knowledge by its **information lifecycle** first, and topi `grep` here is a **ugrep wrapper that honours `.gitignore`**, and `10_knowledge/`, `20_live/` and `30_projects/` are all ignored. A repo-root search returns zero -hits for content that demonstrably exists. Use `command grep`, an explicit path, -or MindGraph. +hits for content that demonstrably exists. This is the trap that makes an agent +confidently answer "that doesn't exist" about something that does. + +### Never run a bare recursive `grep` from the repo root + +> **Binds:** any client issuing a Bash command in this repo +> **Tier:** T2 (blocked) +> **Check:** `bin/bash-discipline-guard` (PreToolUse hook on Bash), with +> `tests/test_bash_discipline_guard.py` as the regression +> **Escape:** `bin/vault-grep`, `command grep`, `git grep`, or `grep -r` with an +> explicit path. All four are allowed and none carries a penalty. + +**Use `bin/vault-grep` first.** It is the purpose-built fix: `rg --no-ignore` +scoped to the repo with the cache and dependency trees already excluded. + +```bash +bin/vault-grep -l "Idle Capacity Engine" # whole-vault file search +bin/vault-grep -n "def parse_frontmatter" scripts/ # only this directory +``` + +It passes arguments to `rg` from the MainFrame root. Explicit paths limit the +search; without paths, ripgrep searches MainFrame or piped stdin using its normal +rules. It requires `rg` on PATH and fails loudly when unavailable. + +Alternatives when `vault-grep` does not fit: `command grep`, an explicit path, or +MindGraph. Never a bare repo-root `grep`. Recursive searches from the root also time out: `.venv` and `node_modules` trees dominate the file count. Scope to the directory you mean. +## Command Execution & Path Discipline + +**Workspace Root Execution:** Always execute commands from the MainFrame repository root using repo-relative paths (e.g. `bin/sync-project-index`). + +**Running the tests:** `uvx --with pytest pytest tests/`. There is no project virtualenv and no pytest on any system python, so a bare `pytest tests/` fails with `No module named pytest`. This line read `pytest tests/` until 2026-08-26, which is a rule that could not be followed as written — recorded as a papercut against this file. + +### Do not prefix commands with `cd <dir> && ...` + +> **Binds:** any client issuing a Bash command in this repo +> **Tier:** T2 (blocked) for a relative target; T1 (detected) otherwise +> **Check:** `bin/bash-discipline-guard` (PreToolUse hook on Bash), with +> `tests/test_bash_discipline_guard.py` as the regression. Allowed-but-noted +> usages are counted in `20_live/workflow-metrics/bash-discipline.jsonl` +> **Escape:** run from the root with a repo-relative path, use the tool's own +> `--root` flag, use `git -C <dir>`, or make the target absolute +> (`cd "$(git rev-parse --show-toplevel)/dir" && ...`). Setting +> `MAINFRAME_BASH_DISCIPLINE=0` disables the guard for a deliberate exception. + +Compound subshell directory changes create brittle relative path assumptions and fail across tool harnesses. The guard denies only a relative target, because telemetry records command *heads* and cannot distinguish a legitimate `cd` from a broken one — the counter exists to make that promotable on evidence rather than opinion. + ## Metadata & Updating - Every finalized note must contain the standard metadata schema defined in `.context/primitives.md`. - Ensure changes to architecture or workflow are recorded in `DECISIONS.md`. diff --git a/CHANGELOG.md b/CHANGELOG.md deleted file mode 100644 index cc7a4c6..0000000 --- a/CHANGELOG.md +++ /dev/null @@ -1,82 +0,0 @@ -# Changelog - -## 2026-07-13 — Process gap P0–P2 implementation - -- **G1** Irregularity lifecycle: `bin/eval-registry dispose|list-open-high` + `20_live/eval-registry/irregularity-dispositions.jsonl`; portfolio triage ignores accepted_risk/waived/superseded. -- **G2/G8** Lane preflight emits `next_phase`, `blocker`, `suggested_command`; `bin/research-lane-loop doctor --all-active`. -- **G3** Harvest re-upgrades placeholder scaffolds; `--project` / `--force` flags. -- **G4** `research-lane-loop handoff-project --slug --to <project>`. -- **G5** Craft close `--blocked-on operator` keeps severity info. -- **G6** Operator-gated highs no longer alone force portfolio severity high. -- **G7** skill-eval linter aliases Procedure/Guardrails/Use for; craft + research-lane skills lint PASS. -- **G9–G11, G13** Canary unbuffered; PATH-safe wrappers; structured proof index; freshness watch-only. -- Gap analysis: `30_projects/mainframe-process-eval/outputs/2026-07-13-three-loop-process-gaps.md`. - - -## Unreleased - -### Added - -- ADR-043 (Accepted) in [DECISIONS.md](DECISIONS.md): research-lane-loop as the completion unit for portfolio research — one lane × one phase from source-literature through knowledge synthesis. Workflow `.context/workflows/research-lane-loop.md`, skill `.agents/skills/research-lane-loop/SKILL.md`; pointers from HARNESS, knowledge-routing, research-lanes-strategy README/template, source-literature handoff. -- Research-loop hardening (2026-07-13): `bin/lane-intake` P0↔`immediate` aliases + status prefix match; `bin/research-lane-loop` preflight/doctor/audit-indexes (resume class A–D, stale capture-index repair); tests `tests/test_lane_intake.py`, `tests/test_research_lane_loop.py`. -- MindGraph doctor (MH01, 2026-07-13): `bin/mindgraph doctor|status` dual-index diagnostics; query fail-fast on missing `documents_fts`/`vec_chunks`/etc.; workspace stub detection; tests `mindgraph/tests/test_doctor.py`. Archive clarified as lane-close micro-loop (not default research pass step). -- Query Pass templates (2026-07-13): `.context/templates/mindgraph-query-pass.md`; task-packet optional section + validator; create-project / mindgraph-refresh / delegate-local-task wired. C28 research lane archived to `lanes/completed/` after lane-close dogfood. -- Project experiment loop (2026-07-13): `.context/workflows/project-experiment-loop.md` + `bin/project-experiment-loop` (preflight/scaffold/canary/close/triage). MindGraph canaries write `20_live/eval-registry/last-canary-action.md`; weekly eval-schedule runs triage after harvest; session-open surfaces the card. Fresh dogfood run `2026-07-13T145406Z-experiment-loop-canary-*`. -- Portfolio eval action (2026-07-13): `bin/project-experiment-loop portfolio` triages **all** eval-profile projects into `20_live/eval-registry/last-eval-action.md`; registry harvest/check skip non-eval outputs so hygiene is actionable. -- Craft research loop (2026-07-13): `.context/workflows/craft-research-loop.md`, skill, template, `bin/craft-research-loop` (preflight/scaffold/close/triage/promote). Project-local `outputs/LAST_CRAFT_ACTION.md`. Dogfood on image-generation-lab: closed realism bake-off as keep; scaffolded `2026-07-13-sdxl-install-scorecard-verify`. -- ADR-029 (Accepted) in [DECISIONS.md](DECISIONS.md): MainFrame Epistemic Standard — 12-source epistemics canon (local `10_knowledge/knowledge-systems/epistemics/`), expanded [EPISTEMIC_STANCE.md](EPISTEMIC_STANCE.md) with GRADE certainty and promotion gate, workflow at `.context/workflows/epistemic-standard.md`, GRADE mapping in credibility-tiers, agent wiring. -- Local epistemics sub-wiki (gitignored `10_knowledge/`, 2026-06-17): 12 raw stubs (SEP, GRADE, Ioannidis, Nickerson, OSC, Cochrane, CEBM, CASP, Guba/Lincoln, Merton, CASP); synthesis notes `mainframe-epistemic-standard` and `epistemics-source-map`; source-literature run note. -- ADR-028 (Accepted) in [DECISIONS.md](DECISIONS.md): `source-literature` workflow complements ingest — credibility tiers, dedup, bibliographic capture to `00_inbox/` before `ingest-minion`. Skill at `.agents/skills/source-literature/`, workflow at `.context/workflows/source-literature.md`, subagent at `agents/source-literature-agent.md`, Claude mirror at `.claude/skills/source-literature.md`. - -- Project-level local-agent delegation packets and private harness evaluation (2026-06-14): added `bin/task-packet`, `.context/templates/task-packet.md`, `.context/workflows/delegate-local-task.md`, packet-aware task manifest generation, deterministic validation tests, and ADR-024. The ignored `agent-harness-eval` project supplies disposable Git fixtures, Stub/Aider adapters, external verification, receipts, harness variants, reports, and a sanitized capability summary for the Local Agent workstation. -- `bin/aider-watcher` telemetry hardening and `tests/test_aider_watcher.py`: Local Coder runs now keep the explicit `client: local` tag, emit one start per session, capture newly created histories, record allowlisted context/edit/approval diagnostics plus observed edits and verification commands, and support isolated history replay. Added `.context/workflows/local-coder-run.md` to keep deterministic operations and external verification outside the model. -- `10_knowledge/regulated-systems/` knowledge domain (2026-06-13): GxP/pharma regulatory rules — cGMP records & document control, data integrity / ALCOA+, electronic records & audit trails (21 CFR Part 11 / EU Annex 11), quality risk management & CSV, and cosmetics/OTC GMP — with a source catalog and a rule→control crosswalk grounding the `biotech-rag-assistant` project. Registered in the tracked [10_knowledge/index.md](10_knowledge/index.md) and indexed into MindGraph; domain contents are gitignored per the processed-knowledge convention. -- ADR-022 (Accepted) in [DECISIONS.md](DECISIONS.md): shared-component reuse via pinned git submodules with automated Dependabot bump-PRs, canonical = each component's projects-root/GitHub repo. Codifies how `evidence-bundler` / `claim-audit-lab` / `apparatus-contracts` / `research-scaffold-harness` stay current across consumers (`scaffold-claims-study` today; the Biotech RAG Assistant per its own ADR-015). -- Planning Standard in [30_projects/AGENTS.md](30_projects/AGENTS.md) (2026-06-12): the four coordination entries (`README.md`, `log.md`, `decisions.md`, `plans/`) are now explicitly required per project, and phased work follows a documented phase-plan template (`Goal` / `Non-Negotiable Boundaries` / `Unit Stance` / `Unit Plan` with per-unit green-boundary checklists / `Verification` / `Tie-Off Review` / `Handoff Notes`), generalized from the format proven in two project workbenches. Completed or historical plans are never retrofitted (Agent Protocol rule 6). Missing coordination files were scaffolded in the two non-compliant private projects. -- Phase 5 dogfood completed (2026-05-28). The full v2 pipeline (`00_inbox/` → minion normalize → `01_ingest/ready/` → ingest-agent enrich → `bin/prep-ingest` → `01_ingest/queue/` → minion route → `10_knowledge/<domain>/`) was exercised end-to-end against a sandbox covering all four entry shapes (no-frontmatter clipping, partial-frontmatter draft, already-`status: extracted` note, convention-named raw PDF) and against a real captured note (`mindgraph-integration-notes.md` routed into `10_knowledge/ai-systems/`). All six acceptance criteria from [planning/mainframe-agent-ingest-plan.md](../../planning/mainframe-agent-ingest-plan.md) Phase 5 passed. -- `tests/test_ingest_minion.py::test_wikilinks_inside_code_spans_are_ignored`: covers the Phase 5 dogfood edge case where wikilinks that appear inside fenced code blocks or inline code spans (i.e. discussion of the syntax, not real connections) are excluded from the `links:` array. -- `bin/prep-ingest` (backed by [01_ingest/prep_ingest.py](01_ingest/prep_ingest.py)): deterministic `01_ingest/ready/` → `01_ingest/queue/` gate for the ADR-009 two-pass design. Validates strict frontmatter, `status: "extracted"`, canonical filename (`YYYY-MM-DD__domain__type__slug.md`), domain in the `10_knowledge/` whitelist, and no queue collision before promoting a file. Same dry-run-first CLI shape as `bin/ingest-minion`. -- `tests/test_prep_ingest.py` coverage for promotion, dry-run, partial frontmatter, non-extracted status, malformed filenames, filename/frontmatter domain mismatch, unknown domain, destination collision, and empty-directory cases. -- Agent-ingest v2 pass-1 normalization in [01_ingest/minion.py](01_ingest/minion.py): missing/partial frontmatter is filled with deterministic defaults and routed to `01_ingest/ready/` with `status: "skimmed"` instead of being rejected. Strict-valid files continue to stage to `01_ingest/queue/` for pass-2 routing. Body `[[wikilinks]]` are extracted into a `links:` array during normalization. -- New `normalize` event kind for files routed to `01_ingest/ready/`. -- `tests/test_ingest_minion.py` coverage for inbox normalization (no frontmatter, partial frontmatter, wikilink extraction) and `status: extracted` direct-to-queue routing. -- ADR-009 (Accepted) in [DECISIONS.md](DECISIONS.md): two-pass ingest with agent-driven middle. Extends ADR-007 by adding a sub-agent enrichment step between deterministic minion passes; preserves the strict `queue/ → 10_knowledge/` quality gate. -- ADR-010 (Accepted) in [DECISIONS.md](DECISIONS.md): cross-tool agent layout (`.agents/skills/` for portable skills, `agents/` for subagent definitions). Keeps Claude-Code-specific config under `.claude/`. -- [agents/ingest-agent.md](agents/ingest-agent.md): subagent definition for the ingest enrichment middle pass — role, tools, procedure, guardrails. -- [01_ingest/AGENTS.md](01_ingest/AGENTS.md): defensive constraints for the ingest layer (no auto-routing, no body modification, raw items immutable, no domain guessing, read-only access to durable knowledge). -- [.agents/skills/](.agents/skills/) stubs for the planned ingest skill set: `ingest-source`, `rename-material`, `classify-note`, `extract-metadata`, `create-source-summary`. -- Status lifecycle extension in [.context/primitives.md](.context/primitives.md): added `skimmed`, `routed`, `extracted`, `synthesized`, `parked`; added `links:` field populated by minion link extraction. -- `bin/session-open` for deterministic session context loading in a fixed order with auto-detection of active project from `STATE.md`. -- `bin/session-close` for end-of-session checks and downstream script triggers (`sync-project-index`, `mindgraph-refresh`, `workflow-report`) with `--check`/`--apply` modes. -- `bin/extract-knowledge` for validating prerequisites and scaffolding knowledge notes extracted from projects with correct metadata. -- `unittest` coverage for session-open (14 tests), session-close (13 tests), and extract-knowledge (14 tests). -- ADR-008 in [DECISIONS.md](DECISIONS.md) for the session lifecycle scripts boundary. -- Script sections in workflow docs for `session-open`, `session-close`, and `extract-knowledge`. -- `bin/ingest-minion` for dry-run-first routing from `00_inbox/` and `01_ingest/queue/` into existing `10_knowledge/<domain>/` folders. -- Markdown frontmatter validation against the standard Mainframe metadata keys defined in [.context/primitives.md](.context/primitives.md). -- Raw PDF handling that preserves the PDF under `10_knowledge/<domain>/raw/` and writes a MindGraph-compatible Markdown stub beside it. -- Ingest workflow documentation in [.context/workflows/ingest-minion.md](.context/workflows/ingest-minion.md). -- ADR-007 in [DECISIONS.md](DECISIONS.md) for the deterministic ingest Minion v1 boundary. -- `unittest` coverage for ingest routing guardrails, including dry runs, missing metadata, unknown domains, raw PDF stubs, and destination collisions. -- Public `README.md` covering the lifecycle model, metadata schema, deterministic scripts, safe operating rules, and the MindGraph boundary. -- MIT `LICENSE`. - -### Changed - -- [01_ingest/minion.py](01_ingest/minion.py): `extract_wikilinks()` now strips fenced code blocks and inline code spans before scanning for `[[wikilink]]` targets, so notes that *discuss* the wikilink syntax don't pollute the `links:` array with example targets. Discovered while dogfooding Phase 5 against a real captured note. -- [agents/ingest-agent.md](agents/ingest-agent.md): step 8 (hand-off) now references the real `bin/prep-ingest run --dry-run` / `--apply` commands instead of placeholder wording. -- [01_ingest/minion.py](01_ingest/minion.py): split frontmatter parsing into permissive `read_frontmatter()` + strict `validate_strict()`; added `extract_wikilinks()`, `render_frontmatter()`, and `normalize_metadata()`. Status enum extended to include the v2 lifecycle values (`skimmed`, `routed`, `extracted`, `synthesized`, `parked`) per ADR-009. -- [.context/workflows/ingest-minion.md](.context/workflows/ingest-minion.md): documents the now-current two-pass behavior; the "Pending changes (ADR-009)" preamble is removed. -- Root [AGENTS.md](AGENTS.md): updated centralized-skills reference from `/skills` to `.agents/skills/`; added `agents/` line for subagent definitions per ADR-010. -- Removed empty `skills/` folder; replaced by `.agents/skills/` per ADR-010. -- Replaced stale legacy naming in `20_live/AGENTS.md` and `.context/workflows/session-open.md` so guidance refers to the current Mainframe primitives. -- Reframed the README MindGraph section to make the paired-but-separate-repo relationship with MindGraph explicit. - -### Security - -- Hardened `bin/workflow-event` command-head redaction. Leading shell env-var assignments (e.g. `FOO=/tmp/bar cmd`) are skipped and path-shaped heads collapse to `<path>`, so filesystem basenames no longer leak into telemetry. Covered by new `tests/test_workflow_event.py`. - -### Notes - -- The ingest Minion is manual and deterministic in v1. It does not process `20_live/` state or `30_projects/` records. -- MindGraph remains a complementary retrieval layer. Raw evidence and markdown files in the lifecycle tree remain the source of truth. diff --git a/DECISIONS.md b/DECISIONS.md index 95bb50d..33103c1 100644 --- a/DECISIONS.md +++ b/DECISIONS.md @@ -1,585 +1,47 @@ -# Architecture Decision Records (ADRs) +# Architecture Decision Records (public subset) -> **Public copy.** Private project names are replaced with stable placeholders -> (`<private-hub>`, `<private-eval-a>`, …). The same placeholder always means the same -> project, so the reasoning still follows. Public projects (MainFrame, MindGraph, -> Claim Audit Lab, Evidence Bundler, verified-done) are named normally. +This file is a **curated public extract**. It is produced by a named export +transform. It is not the complete private decision log and does not include +private project names, personal operations, or unpublished experiments. +## ADR-064: Public MainFrame Is A Fresh-Tree Product, Not A Git Mirror (2026-09-10) -## ADR-049: Operational default is shared daemon + mcp-proxy (2026-08-03) +**Status**: Accepted for the private→public publication boundary. This does +not authorize an unattended public push. -**Status**: Accepted +**Context**: Private `mainframe-live` history contains machine-path backlog. +`.gitignore` is a tracking policy, not a publication policy. The public +repository must not inherit private Git history. -**Context**: ADR-048 shipped the shared loopback daemon as opt-in while leaving -`.mcp.json` on per-client `serve-mcp --db`. In practice Grok, Claude Code, and -other stdio clients each spawned a full MiniLM process (~1.4 GB), so multiple -agent sessions multiplied RAM while the shared daemon either sat idle or ran -*in addition* to the duplicates. - -**Decision**: Make the shared daemon the operational MCP backend for MainFrame: - -1. Root `.mcp.json` (and examples) launch `bin/mindgraph mcp-proxy --url http://127.0.0.1:8000/mcp`. -2. One `serve-daemon` process owns both indexes; clients are thin stdio proxies. -3. MCP tools require explicit `scope` of `knowledge` or `projects` (no blend). -4. `serve-mcp --db` remains available for single-DB debugging only, not daily clients. -5. Optional login LaunchAgent may keep the daemon up; proxy still does not auto-start it. - -**Consequences**: Clients must restart MCP after config change. Agents must pass -`scope` on every shared-MCP call. Daemon must be healthy (`daemon-start` / -`daemon-health`) before proxy connections succeed. Embedder cost is paid once -per machine, not once per agent session. - -**Client rollout (same day)**: Wired proxy/URL for MainFrame `.mcp.json`, Grok -`~/.grok/config.toml`, Claude Code (project `.mcp.json`), Claude Desktop, -Cursor `~/.cursor/mcp.json`, Antigravity `~/.gemini/**/mcp_config.json`, and -Codex `~/.codex/config.toml` (`url = http://127.0.0.1:8000/mcp`). Operating -contracts: `AGENTS.md` item 3, `HARNESS.md` shared-MCP bullet, skill -`mindgraph-retrieval`. - -## ADR-048: Opt-in loopback shared MindGraph MCP daemon (2026-08-03) - -**Status**: Superseded in part by ADR-049 (operational default); engine contract still applies - -**Decision**: Promote the tested workbench Streamable HTTP slice into root -`mindgraph/` while retaining `serve-mcp --db` as the default-compatible stdio -path. The shared daemon binds only to loopback, opens the durable and project -indexes read-only, requires one explicit `knowledge` or `projects` scope per -call, exposes the corresponding trust profile, and is supervised through -explicit PID/status/health/stop commands. The stdio proxy uses the official MCP -SDK and does not auto-start the daemon. - -**Consequences**: Engine remains local-first and scope-explicit. ADR-049 moves -MainFrame client config onto proxy + daemon; `serve-mcp` stays for compatibility -and isolated tests. Auto-start-from-proxy, authentication, remote exposure, -concurrency guarantees, and RAM/latency claims remain deferred or measured -elsewhere. Workbench remains the design/test source for future upgrades; root -`mindgraph/` is the promoted operational engine. - -## ADR-047: Separable C-0 Filtering and Context Allocation (2026-07-27) - -**Status**: Accepted -**Date**: 2026-07-27 - -**Context**: The dual-gate confirmatory protocol measures two apparatus independently — C-0 source eligibility and Speaker context allocation — across four conditions: baseline, C-0 only, Speaker only, both. `apply_dual_gate_governance` in `src/mindgraph/query.py` performs both jobs in one call: it requires a non-empty eligibility manifest, filters candidates against it, and only then applies seat and character budgets. Consequences: - -- **Speaker-only is unavailable.** The helper refuses to run without a manifest and filters before seating, so there is no path that allocates over unfiltered candidates. -- **C-0-only is a fudge.** It can be approximated only by setting seat/char limits high enough to "effectively disable" them, which is a parameter choice rather than an absent step. - -Without both arms, any measured difference cannot be attributed to a specific gate, and the protocol's per-gate decision rules (high-risk exposure, required-proposition coverage) cannot be evaluated. The evaluation project identified this dependency (W5) but explicitly declined to design it. - -**Decision**: - -1. **Extract two primitives** from the existing helper, preserving current behaviour exactly: - - `_filter_by_c0_eligibility(results, manifest)` — manifest validation, identity/path/hash matching, and attachment of the consumed `eligibility_run_id`. - - `_allocate_context_budget(results, max_seats, max_chars, quiet_keywords)` — seat shortlisting, the existing `quiet_keywords` swap heuristic, and the character budget with its truncation rule. -2. **Express all four arms as exact compositions** of those primitives, so a condition differs by the *presence or absence of a step*, never by parameter values: - - | Arm | Composition | - | --- | --- | - | baseline | `run_query` — neither primitive | - | C-0 only | filter | - | Speaker only | allocate | - | both | filter → allocate (`apply_dual_gate_governance`, unchanged) | - -3. **Expose two new public entry points**: `filter_by_c0_eligibility` and `allocate_ungoverned_context`. `apply_dual_gate_governance` keeps its name, signature, and behaviour. -4. **Constrain the ungoverned allocator** — it is an evaluation instrument, not a product capability: - - named for the absence it carries, so `grep ungoverned` finds every call site; - - requires keyword-only `evaluation_use_only: Literal[True]` with **no default**, so it cannot be called accidentally or positionally, and the acknowledgement is visible at the call site; - - never sets `eligibility_run_id`, and asserts the returned rows carry `None`, so no future refactor can stamp governance identity onto ungoverned rows; - - raises when handed an eligibility manifest — passing one signals the caller wanted the governed path; - - stays out of `__all__`, the CLI, and the MCP surface. The evaluation harness imports it directly from the module. - -**Rationale**: Allocation is **subtractive** — `allocate(results, …) ⊆ results`. It cannot admit a source that ungated `run_query` did not already return, so exposing it adds no retrieval reach beyond what the baseline path already provides. The security delta against today is zero; the genuine risk is *misinterpretation* — a caller concluding that "allocated" implies "governed" — which the naming, the explicit acknowledgement argument, and the null-provenance assertion address directly. - -Extraction rather than reimplementation keeps the governed path byte-identical. The existing `tests/test_governance.py` cases must pass **unmodified**; that is the regression proof, and adapting them would void it. - -**Consequences**: - -- The confirmatory 2×2 becomes measurable, and the "effectively disabled limits" workaround for C-0-only is removed rather than documented. -- `apply_dual_gate_governance` gains no new behaviour; callers are unaffected. -- **This authorizes measurement only.** Promoting Speaker-only allocation to any production or default path is a separate decision requiring its own evidence; a passing 2×2 arm is not that evidence. -- Work proceeds workbench-first under a task packet, then a **separate root-promotion packet**. Root promotion requires: the five existing governance tests green and unmodified in both trees, the new allocation tests green, and no change to CLI/MCP response shape. The promotion packet is written after workbench is green rather than now, so it describes a real diff instead of a predicted one. -- ADR-035's retrieval model and ADR-033's link resolution are untouched; this is a consumer-side boundary change only. - -## ADR-046: Dual-Pool WIP — Product Seats vs Eval Seats (2026-07-23) - -**Status**: Accepted -**Date**: 2026-07-23 -**Context**: ADR-041’s hard cap of 5 active projects forced constant pause/swap churn once MainFrame accumulated parallel evaluation suites (harness, tracker, process-eval, mindgraph-eval, claim-audit, <private-claims>, etc.) alongside product work (<private-product>, portfolio, <private-hub>, labs). Eval projects are often correctly “on” for scheduled probes and multi-week measurement programs; treating them as interchangeable product WIP seats made the cap unmaintainable without lying about state. Operator intent: <private-hub> is the strategic hub other projects serve — it should stay active without WIP-swap thrash. - -**Decision**: -1. **Dual pool** for `project_state: active`: - - **Product WIP cap = 5** — outcome/product projects (`wip_class: product`, default). - - **Total active ceiling = 10** — product + eval combined. - - **Eval actives do not consume product seats.** - - **Anchor** (`wip_class: anchor`, default for `<private-hub>`): always-on strategic hub; consumes **neither** product nor total seats. -2. **`wip_class`** optional frontmatter: `product` | `eval` | `anchor`. When omitted: - - `<private-hub>` → `anchor` - - slug ends with `-eval` → `eval` - - known instrument slugs → `eval` (`claim-audit-lab`, `<private-eval-e>`, `<private-claims>`, `<private-eval-c>`, `verified-done`) - - else → `product` -3. **Enforcement** remains in `bin/sync-project-index --check` (and Focus Board total cap for product+eval). Evidence rules and vocabulary from ADR-041 are unchanged. -4. **Focus vs activation** (ADR-044) still separate — eval/anchor being active does not mean they are focus primary. - -**Rationale**: Keeps a tight product focus budget while letting measurement infrastructure and the income hub stay honestly active. Total ceiling of 10 still prevents “everything active” sprawl among product+eval work. - -**Consequences**: Operators pause product projects to free product seats; eval suites may remain active up to the residual total budget; <private-hub> should not be paused to free WIP. Workstation `ACTIVE_CAP` tracks the **product+eval total** ceiling (10). Explicit `wip_class` overrides heuristics when a slug is misclassified. - -## ADR-045: Live Retention Classes and Staged Projects MindGraph Apply (2026-07-15) - -**Status**: Accepted -**Date**: 2026-07-15 -**Context**: After Phase 2 containment, the projects MindGraph manifest covers 28/28 real projects, but the installed projects DB still has incomplete namespaces. Operators need live-index coverage without dumping volatile `20_live/` telemetry into retrieval, and without one-shot mutation of `~/.mindgraph/mainframe-projects.sqlite` without a prove-then-promote path. Different live surfaces go stale at different rates. - -**Decision**: -1. **Retention classes** for `20_live/` and related volatile state are policy authority in `.context/live-retention.md`: Authority (A), Append-only evidence (B), Derived projection (C), Reporter noise (D), High-volume ops (E). Soft caps are documented there; automation is later. -2. **Projects MindGraph stays coordination-only** — manifest-scoped Markdown under `30_projects/`. Never bulk-ingest `20_live` telemetry/events into the projects (or knowledge) index as a hygiene shortcut. -3. **Staged apply is required** for projects index mutation: `bin/mindgraph-projects-apply` supports `--plan` → `--stage` (temp/staging DB + receipt) → `--promote` (backup installed, replace only from a green receipt). Bare `mindgraph-refresh-projects --apply` remains available for experts but the recommended path is the staged packet. -4. **Staleness responses are class-specific** (rebuild projections, archive append-only, rewrite focus authority, re-stage index) — not a single “clean 20_live” cron. - -**Rationale**: Separates “make retrieval cover projects” from “manage live operational volume,” keeps indexes rebuildable and trust-labeled, and makes promote an explicit operator boundary with evidence. - -**Consequences**: Operators follow live-retention policy before/around index work. Stage receipts live under `20_live/system-health/mindgraph-projects-apply/`. Promoting requires a green stage. Soft retention caps are not auto-enforced yet. MG-004 namespace isolation tests remain a separate later unit. - -## ADR-044: Focus Authority, Doctor Contract, and Separate Activation (2026-07-14) - -**Status**: Accepted -**Date**: 2026-07-14 -**Context**: July 10 system audit and scalability program left three design drafts (`plans/scalability/canonical-authority-map.md`, `mainframe-doctor-contract.md`, `p0-repair-packet.md`) unapproved while `<private-eval-b>` stayed paused. Live false-green still reproduces: `bin/session-open --json` reports `ok: true` for compound `STATE.md` focus whose derived project path does not exist. Implementation of `mainframe-doctor` and focus writers was blocked on operator decisions. - -**Decision**: -1. **Focus authority location:** Canonical current focus lives under `20_live/focus/` — at minimum `current.yaml`, with append-only `decisions.jsonl` and `outcomes.jsonl` for history. Project READMEs continue to own project lifecycle state; focus only allocates discretionary attention. -2. **STATE.md role:** Human handoff **narrative** (what changed / remains / blocked / reentry). It must cite the focus revision when structured focus exists. It is **not** the parseable primary project identifier for session tooling after migration. -3. **Session open:** After implementation, reads structured focus first; validates that the primary project path and contract chain resolve. Missing/invalid focus or project is non-green. During a single migration window, `STATE.md` narrative parsing may remain a read-only fallback with an explicit degraded/unknown signal. -4. **Focus vs activation:** Selecting or changing focus does **not** change `project_state`. WIP activation/pause under ADR-041 remains a separate explicit operator action. -5. **Doctor program:** Accept the doctor contract as the health model (vector of claims; required `unknown` cannot aggregate to healthy; reporters ≠ checks). Accept the P0 repair packet as sequenced design input subordinate to the post-audit execution overlay. No repair writers or `bin/mainframe-doctor` ship under this ADR alone. -6. **Weekly human context (default, minimal):** When a focus decision is written, capture `capacity`, `fixed_commitments`, and `deliberate_deprioritizations` unless the operator supplies a fuller `operator_context`. Horizon goals may be empty without inventing them. - -**Rationale**: Separates attention allocation from project lifecycle (ADR-041 WIP), kills compound-STATE false greens without making narrative the authority, and unblocks Phase 0 without authorizing mutations. - -**Consequences**: Phase 0 design contracts are **accepted** with these bindings. Next executable work requires an explicit WIP swap and selection of ≤2 units (recommended first pair after activation: Unit 1.1 frozen baseline receipt, then Unit 1.3 fixture-only doctor shell). Schema and writers remain unbuilt until those units. ADR-041 WIP cap and project README authority are unchanged. - -## ADR-Handoff — Typed research↔project handoffs (2026-07-13) - -**Context**: After research-lane dogfood, “handoff to project” meant both gate-clearing decisions and low-urgency opportunities. One-size application handoffs either stole WIP `next_action` or buried real gates. - -**Decision**: Formalize kinds (`gate`, `application`, `opportunity`, `constraint`, `experiment`, `craft`, `knowledge`, `split`, `close`) in `.context/workflows/research-project-handoff.md` with CLI routing that does **not** overwrite busy project `next_action` for `opportunity`/`knowledge`. - -**Consequences**: Operators pick kind before handoff; receipts land in project `outputs/`; reverse gaps still use lane intake / research-lane-loop, not handoff-project. - - -This file captures meaningful project choices, especially trade-offs affecting reproducibility, scope, evidence quality, or system behavior. - -Numbering note (2026-07-02): entries are newest-first; ADR-023 exists only as "ADR-023 (follow-up)" and ADR-026 was never assigned — both gaps stay unfilled so numbers keep matching external references. - -## ADR-043: Research Lane Loop (Source-Literature → Synthesis) -**Status**: Accepted -**Date**: 2026-07-13 -**Context**: Completing research for portfolio lanes required stitching source-literature (stops at inbox), ingest-minion, optional synthesis, and tracker updates from separate docs. Operators and agents often stopped after capture; the successful C28 full loop lived only in project log prose. -**Decision**: Define one **loopable research pass** as the completion unit for research-lanes work: -1. Workflow: `.context/workflows/research-lane-loop.md` (command card, resume map, definition of done). -2. Skill: `.agents/skills/research-lane-loop/SKILL.md` (agent orchestration). -3. Unit: one lane × one phase (foundation → taxonomy → specialization → application); default 3–4 sources; start at source-literature; end at knowledge synthesis + tracker close; then decide next phase/lane/archive. -4. Source-literature remains the discovery skill and hands off into loop steps 4–7 rather than treating inbox as done. -**Rationale**: Easy to run, easy to resume, hard to skip synthesis. Priority (which lane) stays in `bin/lane-intake`; how a run finishes is now explicit and repeatable. -**Consequences**: Weekly cadence prefers full loops over open-ended search. Project ADR-009 in `30_projects/<private-research>/decisions.md` mirrors this for the tracker. Synthesis skip requires an existing note that already meets the stop condition plus a log line. - -## ADR-042: Nested Local Repos For High-Churn Private Projects -**Status**: Accepted -**Date**: 2026-07-03 -**Context**: `30_projects/*` is intentionally gitignored so project contents stay private, but that makes high-churn projects invisible to every git-based activity signal — the 2026-07-01 audit misread <private-hub> (the busiest control plane) as stale for exactly this reason. quant-markets-lab already carried a nested repo as precedent without conflicts. -**Decision**: High-churn private projects may initialize a nested local git repository at the project root. <private-hub> is now one (initial commit f519bb6, 234 files). Nested repos are **local-only — no remotes**. If a remote is ever wanted it must be private and the history must pass the leak-detection grep (prior-username/local-path rule) first. Nested `git log -1` timestamps count as activity evidence for ADR-041 states and the ADR-040 checkpoint derivation. -**Rationale**: History and diffs for the places where the most decisions happen, plus an honest machine-readable activity signal, without weakening the outer repo's privacy boundary. -**Consequences**: The evidence scanners (sync-project-index, session-close) prefer nested-commit recency over raw mtimes when fresher. Operators must remember nested repos have their own working-tree hygiene; MainFrame's session-close does not yet check nested-tree dirtiness (candidate for a later phase). - -## ADR-041: Evidence-Based Project Activity States With WIP Cap -**Status**: Accepted -**Date**: 2026-07-03 -**Context**: The 2026-07-01 audit found 21+ of 25 projects self-reporting `project_state: active`, several with no file activity for weeks (content-engine 27d idle) and two with missing or malformed frontmatter. The `updated:` field is hand-typed and drifts. Meanwhile the operator's real problem is deciding where to put focus — a state layer where everything is "active" answers nothing. -**Decision**: -1. **State vocabulary**: `active`, `paused`, `planned`, `blocked`, `suspended`, `shipped`, `trashed` — semantics recorded in `30_projects/AGENTS.md`. Any other string is a checker error. -2. **Evidence rule**: `active` requires activity evidence within 14 days — newest of bounded file-mtime scan and nested-repo `git log -1` (same derivation as the ADR-040 checkpoint) — plus a `next_action`. `paused` requires a `next_action` reentry pointer. Self-reported `updated:` carries no authority. -3. **WIP cap (original)**: at most **5** projects `active`; the checker fails loudly on breach. Activating a sixth means pausing one first. **Superseded for pool structure by ADR-046** (product seats remain 5; total ceiling 10; eval seats separate). -4. **Enforcement**: `bin/sync-project-index --check` validates all rules and the generated index gains an `Evidence` column (last observed activity date). `--write` still writes but repeats the problems on stderr. -5. **Sweep applied 2026-07-03**: active set reduced 21 → 5 (<private-eval-a>, claim-audit-lab, <private-hub>, <private-eval-b>, <private-research>); 16 projects paused with reentry pointers; <private-eval-d> and <private-eval-e> got compliant frontmatter. -**Rationale**: Truth from evidence, not self-report (truth-layer design principle 2). The cap makes the focus decision explicit and visible instead of deferred, and gives the workstation's game layer a real mechanic (desks = WIP cap, Phase 4 Unit 7). -**Consequences**: `session-close`'s sync-project-index auto action now stays "needed" while any state lies, so drift is loud at every close and in the tracker feed. Editing a README to pause a project bumps its mtime, so freshly-swept projects show today's evidence date until they decay naturally. The audit's <private-hub> false positive is fixed by construction: nested-repo commits are first-class evidence. **See ADR-046** for dual-pool product vs eval seats. - -## ADR-040: Session Checkpoints Attach To Compaction Events -**Status**: Accepted -**Date**: 2026-07-02 -**Context**: The 2026-07-01 bird's-eye audit found the state layer drifting out of truth because manual rituals (STATE.md narrative, session-close) run slower than the work rate. `bin/session-close` is operator-run; nothing fires it automatically. Compaction cannot be triggered *from* a shell script (it is a context operation inside the agent client), so the integration is inverted: the ritual hooks onto compaction and session-end events that already carry telemetry hooks. -**Decision**: -1. **`bin/session-close --checkpoint`** — a fast (<1s), prompt-free evidence snapshot. It derives active projects from `30_projects/*` file mtimes plus nested-repo `git log -1` (bounded scan, no self-reported `updated:` fields), summarizes today's telemetry zones, and reads weekly-eval staleness directly from `schedule-runs.jsonl` (no subprocess chain). It appends a dated snapshot block to `20_live/last-handoff-draft.md` and a machine record to `20_live/workstation/session-close-feed.jsonl`. -2. **PreCompact hook** runs `session-close --checkpoint --hook-stdin` after the existing `workflow-event` telemetry hook: every mid-session compaction becomes a state checkpoint taken before context is summarized. `--hook-stdin` keeps only derived fields (sha256[:16] session hash — same scheme as `bin/workflow-event` — event name, trigger); prompt or transcript content is never copied. -3. **SessionEnd hook** runs `session-close --check --feed --hook-stdin`: the full check result (pending autos, warnings, eval staleness) is appended to the same feed. Feed mode exits 0 once the record is written — the outcome lives in the record, so a session end never reports a hook failure for pending rituals. -4. **Draft file ownership**: the digest (`--apply`) owns the scaffold above the `## Session Checkpoints (auto)` heading; the checkpoint owns everything below it (derived-active line rebuilt each run, newest five snapshots kept). Each writer preserves the other's region. STATE.md narrative remains human-approved — checkpoints draft, the operator promotes at true session close. -**Rationale**: Design principle 1 of the truth-layer plan — make rituals cheaper than skipping them by automating the draft and keeping the human on the approve step. Attaching to compaction converts the operator's existing habit (compacting long sessions) into automatic state capture, and the JSONL feed gives the workstation tracker (Phase 4 Unit 6 Focus Board) one evidence surface for close-list and attention signals. -**Consequences**: `20_live/last-handoff-draft.md` gains a machine-owned tail section; both files stay gitignored under `20_live/*`. The feed grows one line per compaction/session-end. Derived-vs-declared drift is now measured on every checkpoint — the first live run immediately flagged `<private-hub>` as hotter than the declared active project, confirming the audit's activity-blindness finding. Phase 2 (evidence-based activity states, WIP cap) consumes the same derivation. Hook timeouts stay at 5s; measured checkpoint cost is ~0.6s. - -## ADR-039: System Integrity Audit Workflow and Generalized R1–R8 Reconstruction Rules -**Status**: Accepted -**Date**: 2026-06-29 -**Context**: Modern software workflow and AI pipeline audits (such as reviews for Eldorado Node and Agent Trust Gate) were previously written using domain-specific terms (like pharma QC's LIMS, OOS, and SOP versioning). We need a generalized, system-agnostic procedure and template to standardise data-integrity audits and allow agents and operators to audit technical architectures consistently. (Originally misnumbered ADR-025 and appended at the file bottom; renumbered and moved 2026-07-02 — ADR-025 is "MindGraph Source Lives In A MainFrame Project Workbench".) -**Decision**: -1. Created `.context/workflows/system-integrity-audit.md` to serve as the canonical procedure and template for system integrity audits. -2. Generalised the R1–R8 pharma-inspired checklists into eight system-agnostic data-integrity principles (Raw Source Capture, Logic/Rule Versioning, Run Initialization Log, Sequenced Execution Path, Outcome-to-Source Traceability, Controlled System Overrides, Attributable Approval Gates, and Exportable Evidence Package). -3. Integrated the new workflow into existing discovery and engagement files (`.context/workflows/<private-workflow>.md` and `30_projects/<private-hub>/<private-subarea>/meeting-audit-protocol.md`) for high visibility during client and peer engagements. -**Rationale**: Standardising the data-integrity rules into a reusable system-agnostic format makes it easy to run audits on diverse client architectures while keeping the successful "authoritative vs. derived" and "verification trace" structures from previous reviews. -**Consequences**: Future workflow audits will reference this standard protocol and template. Subagents can consume `.context/workflows/system-integrity-audit.md` directly to perform initial data-integrity evaluations. - -## ADR-038: Thread Creation Uses Existing Prompt And Session Contracts -**Status**: Accepted -**Date**: 2026-06-28 -**Context**: New-thread prompts were being assembled from live project context, but MainFrame had no short workflow connecting thread creation to the existing prompt-design and session lifecycle contracts. -**Decision**: Add `.context/workflows/create-thread.md` as a lightweight pointer. Prompt-design judgment remains in `.agents/skills/prompt-creation/`; context loading and handoff state remain in `session-open` and `session-close`; a separate thread is created only when the user explicitly requests one. -**Rationale**: A pointer makes thread handoffs consistent without duplicating prompt-engineering guidance or creating another large procedure. -**Consequences**: New thread prompts should name the objective, relevant authority, boundaries, expected first response, and completion or approval gate. Plan-shaped handoffs remain context rather than implementation approval. - -## ADR-037: Structural Files Use Shared Profile Schema -**Status**: Accepted -**Date**: 2026-06-24 -**Context**: MainFrame has a growing set of structural files: root/lifecycle contracts, project-local `AGENTS.md` files, `HARNESS.md`, workflows, skills, subagents, templates, manifests, configs, workbench contracts, and eval-methodology files. The 2026-06-23 structural catalog showed that many are intentionally ignored/private, but still architecturally meaningful. -**Decision**: -1. Add `.context/templates/structural-file-profile.md` as the shared schema for creating, auditing, and tightening structural files. -2. Treat project-local structural files as part of the same framework even when they are private or ignored by the outer repo. -3. Keep `AGENTS.md` as the compact always-on contract and `HARNESS.md` as the MainFrame harness-policy contract; route long procedures to workflows, repeated agent judgment to skills, specialized roles to `agents/`, and deterministic enforcement to scripts/config/tests/hooks. -4. Catalogue project structural files with tracked/ignored status and authority/trust labels before promoting or tightening them. -**Rationale**: A shared profile reduces drift without turning every contract into a long manual. It also prevents ignored project files from disappearing from architecture reviews while preserving the public/private boundary. -**Consequences**: Future structural-file audits should include the profile fields or explain why a file type intentionally omits them. Root contracts now point to the schema. Tightening local `AGENTS.md`, workbench contracts, methodology files, and templates should start from the profile rather than ad hoc rewriting. - -## ADR-034: Query-Time Semantic Association -**Status**: Accepted (shipped 2026-06-19) -**Date**: 2026-06-19 -**Context**: MindGraph uses semantic search for query→chunk ranking and explicit edges for document→document traversal (`--expand`). Cross-domain material often co-ranks on well-formed queries but has no wikilink path — so `--expand` cannot surface it and `graph_neighbors` from a single note cannot either. The operator wants deeper connection discovery without adding a parallel trust taxonomy (MainFrame lifecycle trust already lives in index scope, result metadata, and Query Station grouping). -**Decision**: -1. Add a fourth retrieval signal **`associated`**: from fused seed documents, embed a per-doc association text (title + primary chunk), run vec kNN, promote to doc level, append results with `signal="associated"`, `semantic_distance`, and `weak_fit` — same `QueryResult` shape as fused/expanded rows. -2. **Append-only semantics:** Association does not enter RRF fusion math (mirrors Phase 3 `--expand`). CLI flag `--associate`; MCP `query` parameter parity. -3. **Per-index execution:** Association runs inside one SQLite scope at a time; federated grouping remains Query Station responsibility (ADR-032). -4. **Phase record:** `30_projects/mindgraph/plans/phases/phase-9-semantic-association.md`. -**Rationale**: Reuses existing chunk embeddings and semantic ranking primitives — no offline semantic edge table, no LLM entity extraction. Closes the doc-to-doc discovery gap while keeping signals inspectable. -**Consequences**: `Signal` type gains `"associated"`. mindgraph-eval gains cross-domain probes with no wikilink path. Latency budget must be measured on full corpus before defaulting `--associate` in agent workflows. Offline precomputed semantic edges remain deferred. - -## ADR-036: Scheduled MainFrame Eval Suites With launchd And Session Visibility -**Status**: Accepted -**Date**: 2026-06-22 -**Context**: ADR-018 established the process-evaluation loop, and G46 added `bin/eval-registry`, but eval runs still depended on operator memory. A one-off weekly script would silently rot without launchd, staleness checks, or session hooks. -**Decision**: -1. **`bin/eval-schedule`** runs daily/weekly suites, writes `20_live/eval-registry/schedule-runs.jsonl`, and installs macOS launchd agents (`com.mainframe.eval-schedule.{daily,weekly}`). -2. **`bin/eval-schedule check`** exits non-zero when launchd is missing, the weekly run is stale (>8 days), or the last weekly run failed. -3. **Session hooks** surface health: `bin/session-open` prints eval status; `bin/session-close --check` warns when unhealthy; handoff digest includes `eval-schedule status`. -4. **Operator card** at `20_live/eval-registry/OPERATOR.md` documents the weekly review ritual and commands. -5. Weekly MindGraph probe defaults to a **four-query fused regression** (~25s); `--full-probe` retains the full matrix. -**Rationale**: Measurement only improves the system when it runs on cadence and surfaces failures before they become folklore. Wiring eval health into session open/close makes neglect visible without blocking unrelated work. -**Consequences**: EV01 lane owns the ritual; `<private-eval-b>` receives dated `scheduled-weekly` outputs; harvest hygiene excludes `evaluation-feedback.md`. Promotions from scheduled runs still require human review per ADR-018. - -## ADR-035: MindGraph Retrieval Model — Hybrid Explicit Graph + Chunk RAG -**Status**: Accepted -**Date**: 2026-06-19 -**Context**: After ADR-033 link fixes and federation/graph-RAG literature review, the operator asked whether MainFrame should adopt a different retrieval architecture (GraphRAG, LightRAG, vector-only, HippoRAG PPR, newer embedders) instead of the current MindGraph model. -**Decision**: **Keep the hybrid model** as the MainFrame default: -1. **Per-scope SQLite** with FTS5 + chunk embeddings + explicit operator-authored edges (plus ADR-033 frontmatter links). -2. **RRF fusion** for query→chunk ranking; graph BFS for explicit expansion; **semantic association** (ADR-034) for implicit doc neighborhoods — four inspectable signals, not one merged score. -3. **Do not** replace dual lifecycle indexes with a single LLM entity graph or GraphRAG community vault. -4. **Empirical upgrades only** for embedding model and optional cross-encoder rerank — gated by `mindgraph-eval` A/B on a frozen query set, not ad hoc swaps. -5. **Planning reference:** `30_projects/mindgraph/plans/retrieval-model-review.md` for landscape comparison and evolution order. -**Rationale**: MainFrame's job is lifecycle-aware **nomination**, not answer generation. The current shape matches hybrid-memory "narrow then expand" and ADR-032 federated Query Station. GraphRAG-style systems optimize global private-corpus QA with LLM-extracted graphs — different trust and cost profile. Gaps (sparse cross-domain links, no doc-to-doc semantic hop) are addressable incrementally without architectural replacement. -**Consequences**: Phase 9 ships association before embedder migration or GraphRAG community experiments. Optional overview/community layers stay opt-in and low-promotion. mindgraph-eval owns comparative measurements; README/portfolio claims remain "design intent" until probes pass. - -## ADR-033: Dual-Channel Graph Links and Canonical Slug Resolution -**Status**: Accepted -**Date**: 2026-06-19 -**Context**: MainFrame notes commonly declare relationships in frontmatter `links:` (especially syntheses written directly into `10_knowledge/`) while MindGraph only indexed body `[[wikilinks]]`. Authors also link using Obsidian-style trailing slugs (`gxp-pharma-source-catalog`) while files use the canonical `YYYY-MM-DD__domain__type__slug` stem. The mismatch produced rich metadata graphs with sparse traversable edges and widespread phantom links. -**Decision**: -1. **Dual-channel graph ingest:** `extract_document_graph_edges` indexes both frontmatter `links:` and body wikilinks, deduplicating on `target_id` (one edge per target; body `relationship_type` wins when frontmatter had none). -2. **Canonical slug resolution:** `LinkResolver` resolves unique trailing slugs from canonical filename stems (`…__slug` after the type segment) in addition to full stem, sibling path, and unique title matches. Ambiguous slugs remain dangling. -3. **Authoring contract** (documented in `.context/primitives.md` § Link Convention): either channel is sufficient; prefer unique trailing slugs or full stems; do not wikilink `30_projects/` from knowledge notes until bridge registry exists. -**Rationale**: Fixes the convention mismatch without requiring a corpus-wide rewrite first. Frontmatter-only syntheses immediately contribute to `--expand`; trailing-slug links match how operators already write `links:` arrays. -**Consequences**: `bin/mindgraph-refresh` may change edge counts on unchanged file bodies when frontmatter `links:` resolve newly. mindgraph-eval baselines that measure graph degree should be re-run after refresh. Project cross-references stay prose or future bridges — not auto-indexed wikilinks. - -## ADR-031: MindGraph First-Class Workspace Integration -**Status**: Accepted -**Date**: 2026-06-19 -**Context**: MindGraph active source code lived inside the private project layer at `30_projects/mindgraph/workbench/`. This created stale paths, made wrapper and workstation configurations complex, and mixed workspace retrieval infrastructure with a development/upgrade sandbox. **Decision**: -1. Copy the active engine source directories (`src/`, `tests/`, `pyproject.toml`, `README.md`, `LICENSE`, `.gitignore`) into a new, first-class root directory `mindgraph/`. -2. Initialize a dedicated virtual environment in `mindgraph/.venv/` and install the package locally. -3. Update `bin/mindgraph` wrapper to check `mindgraph/.venv/bin/mindgraph` first, falling back to the project workbench version if needed. -4. Modify all references in global/project documentation, workflows, and conventions to point to the root-level `mindgraph/` engine. -5. Extend the local workstation dashboard API (`workstation/server.mjs`) and UI panel (`workstation/components/mindgraph-panel.mjs`) to support direct index refreshing and semantic graph neighbor navigation. -**Rationale**: Moving the active engine to the root establishes MindGraph as standard MainFrame retrieval infrastructure, separating it from the sandbox upgrade workbench. -**Consequences**: Future engine features or upgrades are developed/tested in the `30_projects/mindgraph/` sandbox, then promoted to the root `mindgraph/` directory. All daily execution and workstation interfaces target the root engine. - -## ADR-032: MindGraph Query Station Interpreter -**Status**: Accepted for planning -**Date**: 2026-06-19 -**Context**: ADR-030 made dual MindGraph querying mandatory for project planning, but the current user-facing workstation and MCP examples still behave mostly like single-DB clients. A review of local reality also found that `mainframe-projects.sqlite` can be incomplete because the current projects refresh ingests one project root at a time and prunes against each root separately. The operator wants a structural query station for agents and humans while preserving strong separation between durable knowledge and active project state. -**Decision**: Create a planned MindGraph Query Station as a MainFrame interpreter over separate stores, not as a merged database. The station should expose modes such as `knowledge`, `projects`, `federated`, `apply`, `extract`, and `trace`; group results by lifecycle/trust zone; emit copyable `MindGraph Query Pass` blocks for plans and task packets; and preserve per-index rank, source root, namespace/project, trust profile, path, query string, and fit/scope warnings. Engine prerequisites and response contracts are planned in `30_projects/mindgraph/`; the human/agent UI belongs in `workstation/` and is coordinated by `30_projects/<private-eval-a>/`. -**Rationale**: MainFrame's useful boundary is that `10_knowledge/` answers durable learning questions while `30_projects/` answers active work/status questions. A station should make that boundary easier to use, not dissolve it. Fixing multi-root project ingest and explicit provenance first prevents the UI from confidently presenting incomplete or mislabelled project context. -**Consequences**: Agents should use the station when available, or a CLI-equivalent dual query pass for station modes that have not shipped yet. The first implementation slice ships namespaced `ingest-many`, provenance-rich query rows, manifest-backed project refresh, and workstation `knowledge`/`projects`/`federated` modes with copyable Query Pass output. Workstation implementations must not compare raw semantic distances across indexes as if they were calibrated together, infer trust from path strings, or treat retrieval nominations as verification. Cross-index hard bridges require explicit approval; soft bridges remain provisional suggestions. - -## ADR-030: MindGraph Querying Integrated into Project Planning -**Status**: Accepted -**Date**: 2026-06-18 -**Context**: While MainFrame has dual MindGraph indexes (`mainframe.sqlite` for durable knowledge and `mainframe-projects.sqlite` for active projects), querying them was not structured as an obligatory step during project creation, master planning, phase planning, or task packet preparation. This created a risk of agents or operators duplicating work, missing historical lessons, or violating existing patterns. -**Decision**: Make querying both MindGraph databases an explicit, required step in the planning and project setup workflows: -1. **Global Contract**: Add a MindGraph sourcing rule to `AGENTS.md` (global). -2. **Project Contract**: Update `30_projects/AGENTS.md` to require querying before initializing new plans, task packets, or project phases, and updating the planning rules. -3. **Setup Workflow**: Update `.context/workflows/create-project.md` to incorporate a mandatory MindGraph scan step. -**Rationale**: By structuring MindGraph retrieval as a pre-requisite for planning, we guarantee that all active work leverages the synthesized lessons in `10_knowledge/` and active project context in `30_projects/` from the outset. -**Consequences**: Planning documents must document the query strings used and their key findings. Agents must run these query sweeps before suggesting architectures or writing task packets. - -## ADR-029: MainFrame Epistemic Standard -**Status**: Accepted -**Date**: 2026-06-17 -**Context**: `EPISTEMIC_STANCE.md` had four core rules but no canonical methodology sources, no operational confidence language, and no structured appraisal workflow. The operator required truth-seeking to rank above vibes across all MainFrame work. Existing machinery (`needs-audit`, audit-sweep, credibility tiers) lacked an evidence base in established epistemology and research methodology. -**Decision**: (1) Capture a 12-source epistemics canon via source-literature into `10_knowledge/knowledge-systems/epistemics/` (local). (2) Write synthesis notes `note__mainframe-epistemic-standard` and `note__epistemics-source-map`. (3) Expand `EPISTEMIC_STANCE.md` with claim types, GRADE certainty language, evidence minimums, disconfirmation duty, and promotion gate. (4) Add `.context/workflows/epistemic-standard.md`. (5) Extend credibility-tiers with GRADE operational mapping. (6) Wire agents (`extraction-agent`, `ingest-agent`, `source-literature-agent`, `ingest-source`, `AGENTS.md`). -**Rationale**: Grounds existing post-placement audit model (ADR-019) in peer-reviewed and institutional sources. GRADE certainty gives shared vocabulary for synthesized claims. CASP-adapted checklist makes "truth over vibes" procedurally enforceable for operators and agents. -**Consequences**: All claim-bearing work uses confidence language. `stable` promotion requires operator verification or audit clearance — never LLM alone. Epistemics sub-wiki is the local evidence base; tracked repo carries the contract. Retrofitting existing domain notes with confidence labels is operator backlog. Epistemic research system (suspended) may later consume GRADE scores in auditor scoring. - -## ADR-028: Source-Literature Workflow Complements Ingest -**Status**: Accepted -**Date**: 2026-06-17 -**Context**: Ingest (`ingest-minion` + `ingest-source`) assumes files already exist. Peer-reviewed and institutional source discovery — credibility tiers, deduplication against `10_knowledge/`, bibliographic capture — had no operator workflow or agent skill. Negotiation gap-fill and data-integrity research exposed the gap directly. -**Decision**: Add upstream source acquisition as a paired workflow + skill: `.context/workflows/source-literature.md`, `.agents/skills/source-literature/` (with `references/credibility-tiers.md`), and `agents/source-literature-agent.md`. Captures land in `00_inbox/` as `type: raw` stubs; handoff to existing ingest pipeline unchanged. -**Rationale**: Discovery and enrichment are different lifecycles (ADR-009 minion vs subagent split). Credibility gating requires judgment; routing stays deterministic. Tiers label evidence type, not verified truth — consistent with `EPISTEMIC_STANCE.md`. -**Consequences**: Agents invoke `source-literature` before ingest when the user needs literature search. Run notes document queries, accept/reject decisions, and dedup hits. `.claude/skills/source-literature.md` mirrors the canonical skill. First exercised on negotiation compensation/initiation gap-fill and data-integrity strands (PDA 2018, Schneier & Kelsey 1999). - -## ADR-027: Knowledge Index Is Local, Template Is Public -**Status**: Accepted -**Date**: 2026-06-17 -**Context**: `10_knowledge/index.md` lists the live durable-knowledge domain inventory. Like the project index, it can expose private or still-forming areas before they are ready for the public repo. -**Decision**: Ignore the live `10_knowledge/index.md` and track `10_knowledge/index.template.md` instead. The template preserves the domain-inventory shape, promotion rules, seed-domain rule, and navigation aids without publishing the current inventory. -**Rationale**: The public repo should document the operating pattern without leaking private knowledge architecture or forcing every local domain change into Git history. -**Consequences**: Operators maintain `10_knowledge/index.md` locally. Public docs and fresh checkouts use `10_knowledge/index.template.md` as the scaffold. Any tool or workflow that mentions the live index should treat it as local state. - -## ADR-025: MindGraph Source Lives In A MainFrame Project Workbench -**Status**: Accepted -**Date**: 2026-06-17 -**Context**: MindGraph started as a portfolio live asset, but MainFrame now uses it as operational retrieval infrastructure and evaluates it through `30_projects/mindgraph-eval/`. Keeping the engine source under `projects/portfolio/live-asset/mindgraph/` created stale path references, split planning surfaces, and made MainFrame wrapper configuration depend on a different workspace. -**Decision**: Move MindGraph's active source into `30_projects/mindgraph/workbench/` as a nested project workbench with fresh post-migration Git history. Keep `30_projects/mindgraph/` as the coordination surface (`README.md`, `AGENTS.md`, `decisions.md`, `log.md`, `plans/`) and keep `30_projects/mindgraph-eval/` as the separate evaluation harness for retrieval-quality probes and raw run artifacts. The root `bin/mindgraph` wrapper now prefers the local workbench venv before falling back to `mainframe.mindgraphBin` or `PATH`. -**Rationale**: The project-workbench pattern keeps active source close to the system that exercises it while preserving MainFrame's public/private boundary. Separating source from eval avoids mixing implementation decisions with measured retrieval evidence. A fresh workbench history matches the existing MainFrame migration convention and keeps pre-migration history in the portfolio repo. -**Consequences**: Portfolio registry files point to MainFrame instead of carrying an active source copy. Active MindGraph plans and path references now resolve under `30_projects/mindgraph/`. Historical eval reports may still mention the old portfolio path because they record the environment that produced those runs. +1. Public `camerontjs-dot/MainFrame` is a curated reference implementation. +2. Export uses a positive allowlist over a fresh tree (no `.git` copy). +3. Files that need redaction use named deterministic transforms. Transform + failure blocks export; unsanitized private bytes are not a fallback. +4. Operation admission is not whole-tree Git admission and is not whole-tree + public admission. MPE portable engine/methodology/tests may be public; + MPE evidence remains local. +5. This does not generalize every future operation into the public tree. -## ADR-020: Seed Domains, Legacy-System Consolidation, And Duplicate Deletion -**Status**: Accepted -**Date**: 2026-06-10 -**Context**: Batch pass 2 (ADR-019 batch mode) surfaced three gaps. The domain-creation criteria (5+ meaningful files, recurring use) blocks legitimately new research areas at the moment exploration starts — the first two captures of a distinct topic had no honest destination. Old-system operational docs were being adopted into multiple active projects, scattering pre-MainFrame material the operator has not yet decided to promote. And exact-duplicate working copies were being parked even though batch `source-files/` snapshots already preserve the bytes and the canonical copy is routed. -**Decision**: (1) **Seed domains**: a new top-level `10_knowledge/` domain may be created from one or two strong captures when the topic is clearly distinct from existing domains and active research is expected to bring more — operator confirmation still required (ADR-011 unchanged). Seed domains that stay small without recurring use merge back into an index entry or park. First instance: `software-practice`. (2) **Legacy consolidation**: all old-system documents adopt into `30_projects/<private-migration>/raw-materials/legacy-systems/<system>/` (rule B1); nothing legacy lands in active projects until the operator explicitly promotes it. The pass-1 adoptions into content-engine relocate accordingly. (3) **Duplicate deletion**: exact-duplicate working copies (batch-manifest or body-hash matches) are deleted rather than parked; the ledger records `duplicate-removed` with the canonical path. -**Rationale**: Seed domains keep the knowledge architecture honest about new research instead of forcing weak fits into old domains or stranding captures in exception queues. A single legacy section preserves everything ("don't lose my old stuff") while deferring promotion judgment to the operator. Deleting exact duplicates is safe because provenance lives in the snapshots and the canonical copy; parking them would re-create the clutter the migration exists to remove. -**Consequences**: `10_knowledge/index.md` documents the seed-domain rule and the new domain. `routing-policy.md` gains R11 and B1, P6 (stale generated artifacts), and an amended P3. Relocations of previously adopted legacy docs are recorded as append-only ledger rows, not edits to prior rows. +**Verification**: the private exporter on `camerontjs-dot/mainframe-live` +(`bin/public-export`) plus candidate leak scan and clean-candidate tests. +Those exporter files stay private; they are not part of this public tree. -## ADR-019: Tiered Ingest Autonomy With Review-After And Epistemic Audit Integration -**Status**: Accepted -**Date**: 2026-06-10 -**Context**: The two-pass ingest design (ADR-009, ADR-011) gates every file on per-file operator confirmation. At organic capture volume that is fine; against the registered migration backlog (ADR-016/017) it stalls — lanes age while the operator owes synchronous attention on every file. The safety case for confirm-before-apply is weaker than it looks: raw bytes are batch-snapshotted and hash-registered, routing never overwrites, and the disposition ledger is append-only, so a misplaced clip is cheap to detect and reverse. Separately, the epistemic research system project now runs a continuous local-LLM audit loop — claim extraction, confidence scoring, `needs-audit` tag sweeps of `10_knowledge/`, and a pending-review surface in `20_live/epistemic-audit/` — that can verify routed material after placement. -**Decision**: Ingest classification moves from approve-before to review-after, governed by the rules in `.context/routing-policy.md` and three tiers: -- **Tier A (rule-matched, auto-apply)**: files matching a named policy rule get frontmatter, canonical rename, `bin/prep-ingest`, and pass-2 routing without per-file confirmation. Routed raw clips are tagged `needs-audit` so the epistemic audit sweep verifies them after placement. Ledger rows cite the rule that fired. -- **Tier B (exception)**: no rule match, weak fit, or a proposed new domain. The file stays in `01_ingest/ready/` with a `routing-exception` tag and a one-line `routing_note:`; the operator clears these from a single review table. -- **Tier C (always human)**: new top-level domains (ADR-011 unchanged), anything destined for `20_live/`, promotion to `stable`, and `rejected` dispositions. -For registered migration batches, the ingest-agent classifies the whole lane in one pass and emits one review table. The operator's first full-table review doubles as the ADR-018 baseline: agreement is measured from their corrections, and auto-apply at scale is enabled per rule only where agreement is high. Rules below threshold stay Tier B until refined. -**Rationale**: Confirmation-before-action duplicates protections the system already provides after the fact, and it prices operator attention into every file instead of every rule. Named rules amortize judgment; the epistemic audit loop restores verification where confirmation was removed, so placement becomes cheap and reversible while truth-claims still get audited continuously. Machine-routed and machine-audited material stays labeled (`needs-audit`, audit statuses), preserving the epistemic stance rather than bypassing it. -**Consequences**: `routing-policy.md` becomes a maintained contract — repeated Tier B corrections should graduate into named rules by commit. Ledger rows and frontmatter must keep machine placement distinguishable from human-verified knowledge; `status: stable` remains exclusively human. The epistemic research system's write surface into `10_knowledge/` stays limited to labeled claim and audit artifacts governed by its own project ADRs; classification itself is not delegated to its local models until an ADR-018 evaluation shows rule-level agreement with operator choices. The operator's role shifts from pipeline operation to exception clearing and policy maintenance. +## ADR-063: Selectively Track Portable MPE Engine On Private mainframe-live (public summary) -## ADR-018: Process Evaluation Uses Baseline, Outcome Samples, And Reruns -**Status**: Accepted -**Date**: 2026-06-07 -**Context**: MainFrame has deterministic tests, dry-run checks, workflow telemetry, and focused evaluations for MindGraph and writing style, but no shared system-level loop for deciding whether an operating process is effective. Tool telemetry measures activity and some failures, but it does not establish task quality, workflow adoption, or outcome correctness. -**Decision**: MainFrame process evaluation will use a baseline -> focused change -> rerun loop documented in `.context/workflows/process-evaluation.md`. Each evaluation starts with a specific question, combines deterministic checks with representative outcome samples, classifies findings by cause, and limits implementation to one or two improvement slices before rerunning the same checks. Repeated patterns are promoted according to their nature: deterministic operations to `bin/`, operator sequences to `.context/workflows/`, repeated agent judgment to `.agents/skills/`, and uncertain experiments to private project workbenches. -**Rationale**: This preserves the value of telemetry without mistaking activity for quality. A shared evaluation loop also makes improvements comparable over time and reduces the chance that one-off fixes become untested process changes. -**Consequences**: Evaluation reports should state telemetry coverage limits, preserve baseline evidence, distinguish intentional backlog from defects, and record accepted workflow changes in `DECISIONS.md`. +Private `mainframe-live` allowlists portable MPE source under +`40_operations/mainframe-process-eval/`. Evidence stays local. There is no +standalone public MPE repository. -## ADR-017: Migration Batches Preserve Source Bytes And Use Append-Only File Ledgers -**Status**: Accepted -**Date**: 2026-06-06 -**Context**: ADR-016 established immutable, hash-backed batch registration for repeated second-brain imports. The June 6 migration sweep showed that a manifest alone can detect later file replacement but cannot restore the replaced bytes. It also showed that lifecycle handling drifted because the immutable manifest could not record later per-file decisions without becoming stale. -**Decision**: Each second-brain migration batch stores a batch-local `source-files/` snapshot under `30_projects/<private-migration>/raw-materials/batches/<batch-id>/` and initializes an append-only `disposition-ledger.csv` with one `unresolved` row per registered file. The manifest remains immutable classification evidence. Later decisions such as `retained`, `promoted`, `archived`, `duplicate-removed`, `parked`, `unverified`, `superseded`, or `rejected` append new ledger rows rather than mutating the manifest. -**Rationale**: Recoverable source bytes preserve provenance and make later audits or rollback possible even if the flat inbox changes. A separate ledger keeps lifecycle handling inspectable without rewriting historical registration evidence. -**Consequences**: Registration remains non-destructive to `00_inbox/` and still refuses overwrites. Existing pre-correction batches may need ledger backfills from already-recorded project evidence, and ambiguous rows should stay unresolved instead of being inferred. Source snapshot backfills are only safe when the current source still hash-matches the registered manifest. +## ADR-062 / ADR-061 (public summary) -## ADR-016: Rolling Second-Brain Migration Uses Immutable Batch Registration -**Status**: Accepted -**Date**: 2026-06-05 -**Context**: Repeated imports from the previous second brain mix current state, historical snapshots, durable knowledge, project records, templates, and generated artifacts. -**Decision**: Register each drop as an immutable, hash-backed batch before routing or cleanup. Each batch receives a unique ID, complete file manifest, exact-duplicate groups, and reconciliation record. Reconcile live state claim by claim and require dated evidence before promotion. -**Rationale**: Bulk ingest or whole-file source-of-truth selection would create competing state records and could silently discard useful history. -**Consequences**: `00_inbox/` remains the safe arrival boundary. Registration does not move or edit source files. Content Engine state remains anchored to its external production workspace. Durable knowledge follows confirmation-gated ingest, while volatile state uses dated snapshots or append-only timelines. +`40_operations/` is the lifecycle home for standing operations. MainFrame +Process Evaluation is an admitted operation. Lifecycle location does not +itself determine WIP, focus, health, or approval. Direct README authority +remains authoritative. -## ADR-001: Centralized vs Local Agent Instructions -**Status**: Accepted -**Date**: 2026-05-20 -**Context**: Should each lifecycle folder contain its own `AGENTS.md` file, or should all instructions be centralized? -**Decision**: We will enforce a global `AGENTS.md` at the root, and keep repeatable workflows in `skills/`. We will ONLY place local `AGENTS.md` files in subdirectories that have special defensive rules or processes (e.g. `20_live`). -**Rationale**: Prevents agent context bloat and repetitive instructions while still allowing for localized safety rules. +## ADR-056 (public summary) -## ADR-002: MindGraph Integration Strategy -**Status**: Accepted -**Date**: 2026-05-20 -**Context**: Should the MindGraph GraphRAG system be deeply coupled into the Mainframe ingest pipeline? -**Decision**: MindGraph will remain a separate project and will be wired in as a complementary RAG feature. The Mainframe will have its own ingest pipeline, but will use the same tracking/metadata schema so MindGraph can index it effectively. -**Rationale**: Keeps the Mainframe ingestion simple and markdown-first, separating the graph database concerns from the file organization concerns. - -## ADR-003: Automation Tooling for Ingest -**Status**: Superseded by ADR-007 -**Date**: 2026-05-20 -**Context**: The `01_ingest` pipeline requires metadata validation and graph extraction. Relying entirely on LLMs for this is token-heavy. -**Decision**: We will plan to create lightweight CLI/bash scripts ("Minions") in a future session to handle deterministic routing. -**Rationale**: Saves tokens and increases reliability for repetitive, rule-based operations. - -## ADR-004: Project Lifecycle Automation -**Status**: Accepted; tracked-index portion superseded by ADR-015 -**Date**: 2026-05-22 -**Context**: The `30_projects` area needs low-friction status recall without requiring agents to copy the same status into multiple files by hand. -**Decision**: Project folders will use `README.md` frontmatter as the single machine-readable status source. The project README extends the standard metadata schema with `project_state`, `goal`, `next_action`, and `updated`. The `bin/sync-project-index` script generates `30_projects/index.md` from those fields. ADR-015 changes the generated index from tracked public state to ignored local state. -**Rationale**: Keeps navigation accurate while preserving one editable project status surface. The generated index prevents manual drift. - -## ADR-005: Mainframe MindGraph Operating Boundary -**Status**: Accepted -**Date**: 2026-05-22 -**Context**: The MindGraph MCP wrapper is now complete in the portfolio asset, but Mainframe still needs a local integration policy. -**Decision**: Mainframe will use MindGraph as an external complementary retrieval layer through `bin/mindgraph`, `bin/mindgraph-refresh`, and the optional `.mcp.json.example`. The default database is `~/.mindgraph/mainframe.sqlite`. The default ingest scope is `10_knowledge/`, not the vault root. -**Rationale**: `10_knowledge/` keeps search focused on durable notes and avoids adding operating contracts, workflow files, and empty index stubs to retrieval results. The database stays outside the repo so Git history stays clean. - -## ADR-006: Workflow Telemetry Split -**Status**: Accepted; project-index hook portion superseded by ADR-015 -**Date**: 2026-05-22 -**Context**: We want to improve workflow efficiency over time, including tool-call patterns, without confusing Git hooks with AI-client hooks or leaking sensitive content into logs. -**Decision**: Git hooks enforce deterministic repository hygiene through staged diff checks before commit and nonblocking MindGraph refresh after commit. Tool-call telemetry lives in agent-client hook configuration (`.claude/settings.json` and `.codex/hooks.json`), which calls `bin/workflow-event` for session and tool lifecycle events. Telemetry is metadata-only, append-only, local, and ignored under `20_live/workflow-metrics/`. -**Rationale**: Git hooks can see repository transitions but not AI tool-call intent or duration. Agent-client hooks can see tool lifecycle metadata, including post-tool timing, so they are the right layer for workflow measurement. Keeping logs redacted and ignored respects the volatility constraints of `20_live`. -**Amendment 2026-06-04**: ADR-015 removes project-index freshness from the Git pre-commit hook because the generated project index is private local state. -**Amendment 2026-06-11**: Added Antigravity configuration files (`.antigravity/settings.json` and `.antigravity/hooks.json`) with hook telemetry targeting `--client antigravity` to match the Claude and Codex configuration layouts. -**Amendment 2026-06-13**: Added the Aider history watcher as a metadata-only `client: local` source. An explicit `--client` tag is authoritative over parent-process inference. The watcher records only allowlisted run diagnostics, observed edit counts, and command hashes/heads; it does not store model output, command output, file contents, or raw commands. Existing histories are skipped at watcher startup, newly created histories are read from their beginning, and reviewed histories may be replayed into an isolated telemetry directory for evaluation. - - -## ADR-007: Deterministic Ingest Minion V1 -**Status**: Accepted -**Date**: 2026-05-23 -**Context**: Files captured in `00_inbox/` need deterministic staging, metadata validation, routing, and raw-evidence stub generation without spending LLM tokens on repeatable work. -**Decision**: Mainframe will use `bin/ingest-minion` as a manual, dry-run-first CLI for the v1 ingest path. The script stages files through `01_ingest/queue/`, validates Markdown against the approved metadata schema, routes `note` and `raw` Markdown into existing `10_knowledge/<domain>/` directories, and converts convention-named PDFs into immutable raw files plus MindGraph-compatible Markdown stubs. -**Rationale**: A manual CLI keeps ingest behavior inspectable and low-risk while preserving provenance. Existing knowledge-domain directories act as the whitelist, and MindGraph refresh remains a separate workflow. - -## ADR-008: Session Lifecycle Scripts -**Status**: Accepted -**Date**: 2026-05-27 -**Context**: The session-open, session-close, and extract-knowledge workflows are manual checklists in `.context/workflows/`. Their deterministic steps (file existence checks, downstream script invocation, context ordering, scaffold generation) can be scripted without removing judgment from the agent. -**Decision**: Add `bin/session-open`, `bin/session-close`, and `bin/extract-knowledge` as self-contained Python scripts following the existing conventions (check/apply modes, structured result objects, no shared library). The scripts automate only deterministic operations. Narrative judgment (STATE.md writing, DECISIONS.md review, knowledge content) remains explicitly manual. -**Rationale**: Consistent with ADR-007's approach of scripting deterministic work while keeping judgment manual. Session boundaries are the highest-frequency workflows and the most prone to step omission. - -## ADR-009: Two-Pass Ingest with Agent-Driven Middle -**Status**: Accepted -**Date**: 2026-05-27 -**Context**: The v1 ingest minion (ADR-007) is a strict file sorter — files in `00_inbox/` without complete YAML frontmatter are rejected to `01_ingest/rejected/`. This defeats `00_inbox/` as a fast capture zone. The original second-brain planning (`second-brain-redesign/raw-processing.md`, local notes) defined a `new → skimmed → routed → extracted → synthesized` lifecycle with judgment-driven enrichment between deterministic passes; only the deterministic pass exists today. -**Decision**: Extend the ingest pipeline to a two-pass architecture: -1. **Minion pass 1 (deterministic):** Normalize frontmatter (instead of rejecting), extract `[[wikilinks]]` from body into a `links:` array, stage to `01_ingest/ready/` with `status: skimmed`. -2. **Ingest-agent (subagent, judgment):** Reads files in `ready/`, classifies (domain/type/tags), proposes connections (using `links:` + MindGraph when operational), discusses with user, enriches, renames to convention, sets `status: extracted`. -3. **`bin/prep-ingest` (deterministic):** Validates extracted files and moves to `01_ingest/queue/`. -4. **Minion pass 2 (deterministic):** Existing strict routing from `queue/ → 10_knowledge/<domain>/`. Unchanged from ADR-007. - -Extends the status enum in `.context/primitives.md` with: `skimmed`, `routed`, `extracted`, `synthesized`, `parked`. -**Rationale**: Applies the "Minions vs Sub-agents" routing rule from `second-brain-redesign/gbrain-adaptations.md` (local notes) — deterministic work in scripts, judgment work in sub-agents. Preserves the v1 quality gate at `queue/ → 10_knowledge/` (the strict ADR-007 routing is unchanged) while loosening the entry point so `00_inbox/` can be a real capture zone. Deterministic link extraction (also from gbrain-adaptations §2) means connection-finding doesn't depend on MindGraph being operational. - -## ADR-010: Cross-Tool Agent Layout (`.agents/` and `agents/`) -**Status**: Accepted -**Date**: 2026-05-27 -**Context**: The empty `live-asset/Mainframe/skills/` folder doesn't follow a clear convention. Mainframe is used across multiple agent tools (Claude, Codex, occasionally Google), so a Claude-Code-native layout (`.claude/agents/`, `.claude/skills/`) would not be portable. We need a layout that's recognized across agent tooling. -**Decision**: Adopt the cross-tool convention: -- **`.agents/skills/`** — reusable skill definitions (dotfile because skills are agent configuration, not user-facing content). -- **`agents/`** — top-level named subagent definitions (visible at root because subagents are first-class collaborators). - -Claude-Code-specific settings continue to live in `.claude/`. The empty `skills/` folder is deleted and replaced by `.agents/skills/`. Root `AGENTS.md` is updated to reference the new layout. -**Rationale**: `.agents/` and `agents/` are recognized across Claude, Codex, and other agent runtimes. Keeping multi-tool config under `.agents/` and Claude-specific config under `.claude/` cleanly separates portable agent contracts from tool-specific settings. The `agents/` folder being visible signals that subagents are part of the system's public contract, not internal config. - -## ADR-011: Suggestion-First Ingest Intake and Agent-Led Domain Promotion -**Status**: Accepted -**Date**: 2026-05-29 -**Context**: The first real inbox run showed that pass-1 rejects are too harsh for a capture zone. Recoverable Markdown frontmatter issues and raw PDFs without convention names should guide the agent instead of moving evidence to `01_ingest/rejected/`. The same run also exposed a fresh-vault problem: with no `10_knowledge/<domain>/` folders, strict routing cannot complete, but automatic domain creation would guess at the knowledge architecture. -**Decision**: Pass 1 of `bin/ingest-minion` is suggestion-first. Recoverable Markdown problems stage to `01_ingest/ready/` with warnings, while unsupported files and PDFs needing domain or filename decisions stay in `00_inbox/` with concrete suggestions. Strict rejects remain reserved for pass 2 from `01_ingest/queue/`. Domain and subdomain creation belongs to the ingest-agent: it may propose a new domain when a topic is distinct and likely to recur, but it must justify the proposal against `10_knowledge/index.md` and wait for user confirmation before creating folders or setting `status: extracted`. -**Rationale**: The minion should stay deterministic and nonjudgmental, while the agent handles taxonomy judgment. This keeps fast capture low-friction, preserves raw evidence, avoids weak domain guesses, and still protects durable knowledge with the strict queue gate. - -## ADR-012: Optional Source Metadata on Raw Evidence Wrappers -**Status**: Accepted -**Date**: 2026-05-29 -**Context**: Web clippings and binary documents often carry useful provenance fields beyond the required Mainframe routing schema. For PDFs, common document metadata can provide titles, authors, subjects, keywords, and creation/modification dates, but those fields may be absent, stale, or tool-generated. -**Decision**: Keep the required metadata schema small, but allow optional source metadata fields such as `author`, `published`, `created`, `modified`, `retrieved_at`, `source_type`, `description`, and `keywords`. The ingest minion may extract common PDF Info dictionary fields with stdlib-only best-effort parsing and add them to suggestions or generated raw stubs. These fields are advisory and must not replace the raw evidence file or be treated as verified publication facts without review. -**Rationale**: Optional metadata improves search and triage without making the ingest gate brittle. Keeping extraction best-effort and stdlib-only preserves the deterministic, dependency-light minion boundary while maintaining provenance discipline. - -## ADR-013: AI Business and Knowledge Systems Domains -**Status**: Accepted -**Date**: 2026-05-31 -**Context**: The ready ingest queue contains recurring captures that do not fit cleanly into the existing `agents`, `ai-detection`, `finance`, or `humour` domains. Several notes concern AI product/business strategy rather than agent implementation, while another cluster concerns vault design, retrieval, and personal knowledge systems. -**Decision**: Add two top-level knowledge domains: `ai-business` for AI product strategy, app-layer defensibility, monetization, consulting, and AI-native business models; and `knowledge-systems` for personal knowledge management, Obsidian/vault design, retrieval-first organization, and AI-assisted knowledge workflows. -**Rationale**: Keeping these as separate domains avoids overloading `agents` with business and PKM material, improves retrieval, and follows ADR-011 by making domain promotion agent-led and user-confirmed before deterministic routing. - -## ADR-014: Bundled Writing Style Skill -**Status**: Accepted -**Date**: 2026-05-31 -**Context**: Writing guidance was split across portfolio, content-engine, career, and natural-voice contexts. The recurring need is to apply shared prose principles while still respecting genre-specific constraints without loading every reference into the skill body. -**Decision**: Add `.agents/skills/writing-style/` as a bundled skill with a compact router in `SKILL.md` and separate reference pages for shared natural voice, portfolio writing, content-engine writing, career writing, technical/professional writing, and humour/rhythm. The skill explicitly optimizes for authentic voice, specificity, calibrated judgment, and readable prose rather than AI-detector evasion. -**Rationale**: Keeping `SKILL.md` small follows the centralized skill pattern from ADR-010 while allowing task-specific references to be loaded only when needed. A positive craft-first rule prevents the guidance from collapsing into brittle anti-pattern chasing. - -## ADR-015: Private Project Workbenches and Public Index Template -**Status**: Accepted -**Date**: 2026-06-04 -**Context**: MainFrame is easier to use when the active project workbench can live inside the project layer itself. At the same time, the generated project index can expose private project names and next actions if tracked in the public repository. -**Decision**: Project folders under `30_projects/<slug>/` may contain the full local workbench for that outcome, including nested Git repositories, source trees, drafts, artifacts, and project-local plans. The outer MainFrame repo continues to ignore `30_projects/*`. The generated `30_projects/index.md` is local/private and ignored by Git. The public repo tracks `30_projects/index.template.md` to document the index shape without exposing the live project inventory. If project-layer MindGraph indexing is added, it should use a separate project-scoped database or scope from the default durable-knowledge index and surface project-context trust labels. -**Rationale**: This keeps MainFrame useful as the organizing root for real work while preserving the public/private boundary. Durable, reusable knowledge still moves into `10_knowledge/` only through an explicit extraction step, so active project context does not collapse into verified knowledge. - -## ADR-021: Focus Query Sourcing, Subset Corpus Auditing, launchd Daemonisation, sqlite3 Online Backups, and Evaluation Benchmarking -**Status**: Accepted -**Date**: 2026-06-11 -**Context**: In Phase 6 of the Epistemic Research System development, we need to allow directed web research via focus queries, target specific directories for corpus auditing, schedule a background watcher daemon robustly on macOS, secure SQLite database data with nightly backups in WAL mode, and measure system accuracy on claim extraction and audit verdicts. -**Decision**: -1. **Focus Sourcing**: Added `--focus` option to the CLI `run` and `research-topic` subcommands to interpolate focus instructions into the prompt template. -2. **Corpus Subsets**: Added `--subset` option to `audit-corpus` to limit indexing to a specific subdirectory of `10_knowledge/`. -3. **launchd Daemonisation**: Added `daemon` subcommand to orchestrator CLI with `install`, `uninstall`, `start`, and `stop` actions, managing a plist dynamically referenced to the active virtual environment's Python executable and script paths. -4. **Online Backups**: Implemented a daily backup hook using `sqlite3.Connection.backup()` inside the watcher loop, writing WAL-checkpointed database copies under `90_archive/epistemic/backups/`. -5. **Evaluation Framework**: Created `harness/evaluator.py` to calculate claim extraction Precision/Recall/F1 and support verdict accuracy against manual ground truth lists, with automated fallback logic for offline/stub mode. -**Rationale**: These features ensure high-fidelity directed research, resource-conscious local corpus indexing, zero-configuration background operation, robust transactional database backups, and measurable system iteration quality. -**Consequences**: The database backups will run automatically on the first watch loop run of a calendar day. When testing offline, the evaluator will automatically mock the LLM output with corresponding ground truth claims for end-to-end integration validation. - -## ADR-022: Shared-component reuse via pinned git submodules with automated bump PRs -**Status**: Accepted -**Date**: 2026-06-13 -**Context**: Several projects reuse the same building blocks — `evidence-bundler`, `claim-audit-lab`, `apparatus-contracts`, `research-scaffold-harness`. Each is an independent project under `30_projects/<name>/workbench/` whose working copy is the clone of a per-component GitHub repo (`github.com/camerontjs-dot/<name>`). `<private-claims>` already consumes them as git submodules under `workbench/components/`. As more projects reuse them (the Biotech RAG Assistant intends to vendor `evidence-bundler` and `claim-audit-lab`), we need one rule for keeping every copy current. -**Decision**: -1. **Canonical = the projects-root working copy of each component, which is the clone of its own GitHub repo.** Fixes happen there and are pushed to that component's GitHub repo; one source of truth per component. Consumers never edit a vendored copy. -2. **Consumers vendor via git submodules** under `workbench/components/<name>`, mirroring `<private-claims>`. No copy-paste, no path symlinks, no package publishing. -3. **Pin submodules to release tags** (e.g. `claim-audit-lab v0.2.0`), not a moving `main`, so a consumer takes shared-code changes as deliberate, reviewable bumps. (Today `<private-claims>` pins CAL/apparatus to tags but EB to `main`; tags are the target state.) -4. **Automate "keep current" with Dependabot** in each consumer repo (`package-ecosystem: "gitsubmodule"`): it opens a PR when a tracked submodule's upstream advances; a human reviews and merges. A scheduled GitHub Action running `git submodule update --remote` + opening a PR is an acceptable equivalent. Manual fallback: `git submodule update --remote <path>` then commit the gitlink bump. -5. **Prerequisite**: a consumer must be a GitHub repo for the automation to run; components moving from `main`-tracking to tags need release tags cut on their GitHub repos. -**Rationale**: Submodules keep one source of truth with explicit, pinned, reviewable version references — the same evidence discipline the components themselves embody. Tag pinning plus Dependabot gives controlled propagation (fix once in canonical → push → auto-PR into every consumer) without a consumer silently drifting onto an unreviewed shared-code change. -**Rejected alternatives**: copy-paste vendoring (drifts, no provenance); path symlinks (not portable, break on clone/CI); publishing each component to a package index (heavier release process than warranted while the components are pre-1.0 and co-evolving); a single monorepo (loses the independent per-component repos and their histories/tags); tracking `main` everywhere (consumers take unreviewed changes the moment upstream moves). -**Consequences**: `<private-claims>` should gain a Dependabot config and move EB from `main` to a tag. The Biotech RAG Assistant adds `evidence-bundler` and `claim-audit-lab` submodules under `workbench/components/` once it is on GitHub and those components are stable/tagged (its own ADR-015 records the application and its by-reference predecessor, ADR-001/006). Component repos should cut release tags so consumers can pin. This ADR governs how shared code propagates; it does not itself wire any new submodule. - -## ADR-023 (follow-up): Audit Sweep Tooling and Skill Promotion for Ingest (Stabilization of ADR-019) -**Status**: Accepted -**Date**: 2026-06-13 -**Context**: After landing the tiered ingest + review-after model (ADR-019) and discovering real `needs-audit` tags already present on hundreds of routed raws (but no routine surfacing or processing), plus the epistemic research system still completing "final audit integration", the compensating control for batch/Tier A routing was not yet operational. At the same time the detailed per-file judgment procedure remained embedded in a long subagent definition instead of being promoted to reusable skills per the process-evaluation promotion test and ADR-010. -**Decision**: -1. Created `.context/workflows/audit-sweep.md` and `bin/audit-sweep` (dry-run/apply/JSON/--subset support) for deterministic discovery of `needs-audit` + recent routed items, manifest generation into `20_live/epistemic-audit/pending-review/`, and handoff to the auditor. -2. Extended `30_projects/<private-harness>/workbench/mainframe_bridge/` with `sweep_knowledge_for_audit_tags()` to allow the harness to drive or be driven by the sweep. -3. Integrated the sweep into `session-close`, `session-guide`, README, and reinforced tagging obligations in the ingest-agent. -4. Promoted the core per-file enrichment loop into a first-class implemented skill at `.agents/skills/ingest-source.md` (with sub-skill `classify-note` also expanded). The `agents/ingest-agent.md` per-file procedure was thinned to delegation + guardrails/mode selection only. Batch mode remains in the agent definition for now. -5. Verified the scanner immediately surfaces real backlog (226+ explicit needs-audit items). -**Rationale**: Makes the post-placement audit safety net (the key assumption of ADR-019) actually usable and routine today. Moves repeated judgment out of the subagent contract into the skill layer exactly as the system's own promotion rules prescribe. Keeps the new tool lightweight and deterministic while delegating LLM claim work to the existing epistemic harness. -**Consequences**: -- Before any further widening of Tier A or heavy batch usage, run `bin/audit-sweep` (and drive auditor passes on the manifests) on the existing backlog, especially high-volume domains. Treat early sweeps as calibration data. -- Future ingest-agent definitions and workflows should reference the skills rather than re-describing the steps. -- Add coverage metrics and sweep calls to workflow-report / process evaluations. -- The epistemic orchestrator can grow a dedicated `sweep --mainframe` / audit-focused mode that consumes the manifests or directly queries tags. -- Record ongoing audit adoption and any rule refinements in `<private-eval-b>/` outputs and future DECISIONS entries. No change to the core routing-policy or primitives schema. - -## ADR-024: Reviewed Task Packets and Isolated Local-Agent Evaluation -**Status**: Accepted -**Date**: 2026-06-14 -**Context**: Local Coder telemetry can reveal execution and capture failures, but it does not establish which task categories are safe to delegate or whether a harness change improves outcomes. Project plans also need a stable implementation boundary that can be handed to a local agent without delegating unresolved design choices. -**Decision**: -1. Prepared local-agent work is expressed as a reviewed Markdown task packet under `30_projects/<slug>/plans/task-packets/`. Only packets marked `ready` may execute; run state belongs in receipts rather than mutating the packet. -2. `bin/task-packet` validates paths, scope separation, required sections, workdirs, and argv-safe verification commands, then compiles packet summaries into ignored local state for the task board. -3. Local-agent evaluation runs in disposable Git copies with hidden external verification, exact scope scoring, no agent commits, private raw artifacts, and versioned receipts. -4. Harness changes are evaluated one at a time against a fixed baseline. Capability graduation is specific to task category, model profile, and harness version. -5. The Local Agent workstation may read only a sanitized aggregate capability file. It must not expose prompts, transcripts, receipt paths, diffs, or private verifier output. -6. MindGraph retrieval is curated outside the executing agent and remains a context nomination, never verification. -**Rationale**: This separates planning authority, execution, verification, and promotion. It makes delegation reusable across projects while preventing synthetic success, model self-report, or a polished workstation display from becoming unsupported evidence of capability. -**Consequences**: Projects gain a stricter preparation step before local delegation. The private `<private-harness-eval>` project owns experimental adapters, fixtures, receipts, and harness variants; only proven reusable validators and workflows are promoted into MainFrame. +Operations are distinct from projects. They share one slug namespace across +`30_projects/` and `40_operations/`. Duplicate identity fails closed. diff --git a/EVAL_METHODOLOGY.md b/EVAL_METHODOLOGY.md new file mode 100644 index 0000000..81b36c8 --- /dev/null +++ b/EVAL_METHODOLOGY.md @@ -0,0 +1,107 @@ +# Eval Methodology Contract + +This document defines how MainFrame designs, runs, records, and interprets **evaluations** — distinct from claim discipline in [EPISTEMIC_STANCE.md](EPISTEMIC_STANCE.md). + +## Separation of duties + +| Question | Contract | Workflow | +|----------|----------|----------| +| Was the run designed and analyzed honestly? | **This file** | optional/deferred eval-methodology workflow | +| Are the written conclusions supported? | EPISTEMIC_STANCE | [.context/workflows/epistemic-standard.md](.context/workflows/epistemic-standard.md) | +| Is MainFrame process healthy? | process-eval plan | [.context/workflows/process-evaluation.md](.context/workflows/process-evaluation.md) | + +Eval results are **observations or hypotheses** until replication, calibration, or hold-out confirms them. Never promote from a single exploratory run. + +## Lab report convention (universal) + +Every decision-bearing test or experiment is tracked as a **lab report** — the same skeleton a careful experimentalist would keep (question, design, results, irregularities, limits, disposition, next experiment). + +| Piece | Path | +|-------|------| +| Template | [`.context/templates/lab-report.md`](.context/templates/lab-report.md) | +| Workflow | optional/deferred lab-report workflow (`bin/lab-report` when installed) | +| CLI | `bin/lab-report scaffold \| check \| list` | +| Default location | `30_projects/<slug>/outputs/lab-reports/<lab_report_id>.md` | +| Raw | `30_projects/<slug>/raw-materials/<lab_report_id>/` | + +Eval-registry outputs (`.context/templates/eval-output.md`) **implement** this convention and add the mandatory metric YAML harvest block. Craft trials may use `craft-trial.md` but should still answer the lab-report minimum bar when the trial is measurement-like. + +```bash +bin/lab-report scaffold --project example-project \ + --title "my-experiment" \ + --question "..." --decision "..." --study-type exploratory +bin/lab-report check --project example-project --id YYYY-MM-DD-my-experiment +``` + +## Core rules + +1. **Decision sentence first** — Every eval slice states what will change if the result is positive, negative, or inconclusive. +2. **Study type required** — `confirmatory`, `exploratory`, `regression`, `observational`, or `calibration`. No unlabeled runs. +3. **Protocol pin** — `protocol_ref` (frozen query set, harness version, rubric version, script hash) before peeking at results. +4. **Unit of analysis fixed** — Per query, per claim, per case, per session window — declared before analysis. +5. **Raw before summary** — Artifacts in `raw-materials/` or redacted telemetry paths; summaries in `outputs/` never replace raw. +6. **Lab report required** — Decision-bearing runs write a lab report (or eval-output specialization); unnamed log dumps are not enough. +7. **Metric extract mandatory** — Every eval-profile / promotion-relevant output ends with a `## Metric extract (eval-registry)` YAML block. +8. **Irregularities explicit** — Every extract includes `irregularities:` — use `[]` only when none were observed; omitting the field is a protocol violation. +9. **Track everything odd** — Timing glitches, parse warnings, hash mismatches, off-by-one counts, flaky reruns, scope warnings, tool exit 127, YAML scalar quirks — log as irregularities even if "irrelevant" to the headline metric. +10. **No significance theater** — Effect sizes, CIs, and raw counts over binary p-value promotion. See methodology synthesis on sample size. +11. **Registry harvest** — Run `bin/eval-registry harvest` after writing or updating eval outputs; `bin/eval-registry check --strict` before session close when eval work occurred. + +## Eval-profile projects (strict) + +Applies to projects matching `*-eval`, `scaffold-claims-study`, and any project with `tags: [eval-profile]` in README: + +- `methodology-approach.md` — read-first in new session +- `outputs/*.md` — frontmatter `study_type`, `protocol_ref`, `eval_run_id` +- Metric extract + irregularities on every report +- `decisions.md` conclusions that depend on evals cite `eval_run_id` + study_type + +Project rules: [30_projects/AGENTS.md](30_projects/AGENTS.md) § Eval profile. + +## Knowledge base + +Canonical playbooks (local, `10_knowledge/knowledge-systems/methodology/`): + +- Scientific method & experiment design synthesis +- Eval sample size and significance +- Cross-eval trends and registry +- Irregularity and anomaly tracking + +Lane trackers, when present, live under the local research-program operation. + +## Registry + +Append-only eval state: `20_live/eval-registry/` + +- `runs.jsonl` — one row per harvested run +- `metrics.jsonl` — long-format metric rows +- `irregularities.jsonl` — every logged irregularity + +Harvest: `bin/eval-registry harvest`. Status: `bin/eval-registry status`. + +## Skill & Agent-as-Evaluator Standards + +For validating portable agent skills (e.g., prompt-creation, skill-creation) without external API costs, the following standards apply: + +1. **Agent-as-Evaluator Separation**: + - The workbench/runner is restricted to deterministic, zero-cost operations (static structure linting, trigger phrase density validation, receipt log checks, gate math). + - The active agent session serves as the reasoning evaluator, executing the mock runs, judging correctness, and writing JSON receipts. +2. **Pesticide Paradox Mitigation**: + - **Trigger Density Control**: Centralized skill triggers are strictly capped at 5 to 7 trigger phrases/keywords to prevent trigger bloat and false-activation noise. + - **Dynamic Case Expansion**: Whenever an agent fails a skill constraint in production, a representative test case must be immediately transcribed and added to the workbench cases. +3. **The 5-Case Protocol**: All evaluated skills must be measured against 5 distinct case classes to prevent overfitting: + - *Silent Negative* (activates only on target domain, stays silent on adjacent tasks). + - *Happy Path* (performs correctly when complete information is present). + - *Minimal Input* (seeks clarification instead of making assumptions). + - *Edge Cases* (resilient to contradictory or vague goals). + - *Overachiever / Constraints* (adheres to strict length limits and forbidden phrases). +4. **Graduation Gates**: + - Skills cannot be promoted to global production without passing their designated gate thresholds (typically $\ge 80\%$ overall pass rate, with $100\%$ success on critical paths like Happy Path and Silent Negative). + - **Cross-Model Calibration**: Must compare performance (using the `compare` command) between models (e.g., Claude Code vs. Codex) to record formatting consistency and token efficiency. + +## Related + +- [HARNESS.md](HARNESS.md) — harness eval program layer +- [DECISIONS.md](DECISIONS.md) ADR-040 +- `.context/templates/eval-output.md` +- `.context/templates/methodology-approach.md` diff --git a/HARNESS.md b/HARNESS.md index e48a372..deda063 100644 --- a/HARNESS.md +++ b/HARNESS.md @@ -1,6 +1,26 @@ # MainFrame Harness Operating Contract -This file defines how MainFrame treats **agent harnesses** across lifecycle folders. It complements `AGENTS.md` (global agent behavior) and `STATE.md` (current focus). Agents and operators should read this when preparing delegation, evaluating capability, or choosing local vs cloud execution. +This file defines how MainFrame treats **agent harnesses** across lifecycle folders. It complements `AGENTS.md` (global agent behavior), `20_live/focus/current.yaml` (structured focus), and `STATE.md` (handoff narrative). Start with `bin/session-open`; its reading order and output meanings are in `.context/workflows/session-open.md`. + +## Session orientation + +Workspace arrival describes the lifecycle map, recorded focus and relevant +uncertainty, then stops. Read root `AGENTS.md`, this section, `STATE.md`, available +structured focus, and `.context/workflows/session-open.md`. The remaining harness +sections are deferred until project resume or an action makes them applicable. +Recorded focus is attention metadata; it does not authorize work or establish +project readiness. A workspace arrival does not diagnose the focused project. + +For a named project, use `bin/session-open --project <slug> --task "request"`. +That route requires the full harness, lifecycle/local contracts and reconstruction +workflow. Use `--intent resume` to select recorded focus explicitly. Read complete +required batches, keep unknowns visible, and follow current authority into evidence. +The route does not grant implementation, evaluation, delegation or release authority. + +> **Binds:** agents choosing arrival versus project/action context +> **Tier:** T0 (advisory reading and scope policy) +> **Check:** none for actual reading; `tests/test_session_open.py` checks emitted routes and bounded content +> **Escape:** read the full harness when the section or router is unavailable; report missing relevant context and stop before dependent claims or action ## Definition @@ -15,7 +35,7 @@ session lifecycle, client differences, task categories, and promotion gates. This file owns durable harness policy. It should not absorb project status, repo-specific build commands, raw telemetry, or current scoreboards. Current -evaluation results live in `30_projects/agent-harness-eval/`; durable patterns +evaluation results live in `local-only: harness-evaluation project/`; durable patterns live in `10_knowledge/agents/`; structural-file updates should be checked against `.context/templates/structural-file-profile.md`. @@ -26,7 +46,7 @@ Harness knowledge is split on purpose — different lifecycle, different MindGra | Layer | Location | MindGraph DB | Trust | Contents | |-------|----------|--------------|-------|----------| | **Patterns** | `10_knowledge/agents/` — e.g. `harness-engineering-by-task-category` | `mainframe.sqlite` (default) | Durable, synthesized | Task categories, five subsystems, local/cloud integration patterns, skill routing heuristics | -| **Program** | `30_projects/agent-harness-eval/outputs/` — e.g. `harness-evaluation-program` | `mainframe-projects.sqlite` | Project status, dated | H0/H1/H2/H3 variants, sealed cases, graduation matrix, next gate | +| **Program** | `local-only: harness-evaluation outputs/` — e.g. `harness-evaluation-program` | `mainframe-projects.sqlite` | Project status, dated | H0/H1/H2/H3 variants, sealed cases, graduation matrix, next gate | | **Contract** | `HARNESS.md` (this file) | Not indexed | Operating policy | Rules that do not change every eval run | Volatile telemetry (`20_live/workflow-metrics/`) is intentionally **outside** default MindGraph scope. It shows activity, not graduated capability. @@ -35,7 +55,7 @@ Volatile telemetry (`20_live/workflow-metrics/`) is intentionally **outside** de | Client | Typical use | Harness surface | Capability evidence | |--------|-------------|-----------------|---------------------| -| **Local Coder** (Aider + Ollama) | Bounded code edits | Task packets, file allowlists, post-hoc verification | `agent-harness-eval` receipts + graduation gate | +| **Local Coder** (Aider + Ollama) | Bounded code edits | Task packets, file allowlists, post-hoc verification | `local-only: harness-evaluation project` receipts + graduation gate | | **Claude Code** | Planning, multi-file judgment | AGENTS.md, skills, hooks | Redacted telemetry; no auto-graduation | | **Codex** | Implementation slices | Hooks, permissions | Redacted telemetry; vocabulary TBD | | **Grok Build** | Cloud agent sessions | Native project hooks, skills | Redacted telemetry (`client: grok`) | @@ -44,19 +64,30 @@ Do not infer delegation authority from live telemetry or workstation display alo ## Planning, execution, verification -1. **Planning authority** — Frontier/cloud agents and the operator produce reviewed task packets. Unresolved design stays out of execution. -2. **Execution** — Local or cloud agent runs inside scope. Run state goes to receipts, not packet mutation. -3. **Verification** — External deterministic checks (tests, linters, scope diff, claim-accuracy). MindGraph supplies **context nominations only**, never verification. +1. **State reconstruction on resume** — Before planning, resolve the current project, repository/worktree, candidate, experiment, dirty/concurrent, and evidence state. Casual “take a look / decide next” language is read-only; it does not authorize implementation or lifecycle mutation. Follow `.context/workflows/project-resume-and-candidate-lifecycle.md`. +2. **Planning authority** — Frontier/cloud agents and the operator produce reviewed task packets. Unresolved design stays out of execution. +3. **Execution** — Local or cloud agent runs inside scope. Run state goes to receipts, not packet mutation. +4. **Verification** — External deterministic checks (tests, linters, scope diff, claim-accuracy). MindGraph supplies **context nominations only**, never verification. + +### Execution honesty and tool discipline (ADR-051) + +1. **Tool taxonomy:** Explicitly distinguish **Transformers** (data formatting/packet builders — never named `audit` or `evaluate`), **Runners** (execute real models, CLI subprocesses, or databases; fail closed if unavailable), and **Agent Judgment** (explicitly qualitative LLM synthesis). +2. **Falsification testing:** Tests for evaluators, verifiers, and bridges must include adversarial and disconnection cases (asserting failure when the backend is missing or evidence is contradictory), not merely tautological assertions on internal dictionary keys. +3. **Live trace receipts:** Evaluation claims and scorecards must reference a persistent execution receipt on disk with an exact path and hash. See `.context/workflows/deterministic-tool-standard.md`. +4. **Degraded-mode resilience and envelope visibility (ADR-052):** Tools and harness hooks must operate cleanly in degraded states (e.g. offline daemon, unindexed cache). Fallbacks must be deterministic, emitting clear diagnostic receipts. Telemetry and tool errors must expose the "edge of the envelope" (boundary conditions, rate limits, token pressure, timeout thresholds) so agents and operators can calibrate risk rather than encountering silent failure walls. ### Focus and system health (ADR-044) -- **Focus authority (accepted, not yet implemented):** operator-approved primary attention will live under `20_live/focus/` (`current.yaml` + decision/outcome history). Project READMEs keep lifecycle state; focus does not activate projects (WIP remains ADR-041). -- **STATE.md:** human handoff narrative; after migration it cites focus revision rather than acting as the parseable project id. -- **Doctor:** health is a **vector** of claims (`bin/mainframe-doctor` contract accepted; binary ships with Unit 1.3 after WIP activation). Required `unknown` must not aggregate to healthy. Until the doctor exists, treat `session-open ok: true` with a missing project path as a known false-green (reproduced 2026-07-14). -- Design surfaces: `30_projects/mainframe-process-eval/plans/scalability/`. +- **Focus authority:** operator-approved primary attention lives under `20_live/focus/` (`current.yaml` + decision/outcome history). Arrival displays focus without entering its project. `bin/session-open --intent resume` prefers structured focus unless `--project` names the task's project. Project coordination files keep lifecycle state; focus does not activate projects (WIP remains ADR-041). +- **STATE.md:** human handoff narrative that cites the focus revision. Session-open retains it as a visible fallback when structured focus cannot supply a project. +- **Doctor:** `bin/mainframe-doctor --quick` reports a vector of health checks. Required `unknown` must not aggregate to healthy. Session-open checks required context paths separately; `ok: true` does not establish system health, fresh focus, project readiness, or authorization. +- Design surfaces: `40_operations/mainframe-process-eval/` (operation-owned plans stay local). Workflows: +- `.context/workflows/session-open.md` — ordered context references, applicable contracts, and explicit missing-prerequisite reporting +- `.context/workflows/project-resume-and-candidate-lifecycle.md` — read-only project reconstruction, immutable candidate identity, active experiment visibility, and comparison gates (ADR-054) +- `.context/workflows/deterministic-tool-standard.md` — deterministic tool taxonomy, fail-closed boundaries, and falsification test standards (ADR-051) - `.context/workflows/delegate-local-task.md` — packet prep and isolated execution - `.context/workflows/local-coder-run.md` — live local coder discipline - `.context/workflows/source-literature.md` — peer-reviewed / institutional source discovery before ingest (ADR-028) @@ -65,6 +96,7 @@ Workflows: - `.context/workflows/epistemic-standard.md` — claim classification, evidence appraisal, confidence language, promotion gate (ADR-029) - `.context/workflows/ingest-minion.md` — deterministic inbox → `10_knowledge/` routing - `.context/workflows/eval-schedule.md` — scheduled eval suites, launchd, weekly review ritual (ADR-036) +- `.context/workflows/repo-reconciliation.md` — GitHub-authoritative fetch/classify of registered checkouts; gated fast-forward only (ADR-060) - `.context/workflows/project-experiment-loop.md` — measured experiment pass (design → run → harvest → **triage/action**); `bin/project-experiment-loop` - `.context/workflows/lab-report.md` — universal researcher lab-report notebook for every decision-bearing test; `bin/lab-report scaffold|check|list` - `20_live/eval-registry/last-eval-action.md` — portfolio triage for **all** MainFrame evals (action layer); `last-canary-action.md` is a pointer @@ -76,7 +108,7 @@ Optimize and graduate harnesses **per task category × model profile × harness ## Graduation (local delegation) -Real delegation requires a tuple that passes the strict gate in `agent-harness-eval`: +Real delegation requires a tuple that passes the strict gate in `local-only: harness-evaluation project`: - ≥ 8/10 held-out verified passes - Zero scope violations @@ -89,9 +121,9 @@ Local Coder evaluation is **three layers**. Do not collapse them into one score. | Layer | Question | Owner / artifacts | |-------|----------|-------------------| -| **Public standards** | Is the model coding-agent-shaped? | Calibration only: Aider Polyglot; optional SWE-bench Verified lite. Plan: `30_projects/agent-harness-eval/plans/coding-standard-benchmark-calibration.md` | -| **Sealed H1 suite** | Does *our* stack pass *our* edit jobs under packet + external verify? | `agent-harness-eval` coding-backend hard screen / multi-path scorecards | -| **Live receipts** | May it touch real trees under this contract? | This section + verified-done live path | +| **Public standards** | Is the model coding-agent-shaped? | Calibration only: Aider Polyglot; optional SWE-bench Verified lite. Plan: `local-only: harness-evaluation plans/` | +| **Sealed H1 suite** | Does *our* stack pass *our* edit jobs under packet + external verify? | `local-only: harness-evaluation project` coding-backend hard screen / multi-path scorecards | +| **Live receipts** | May it touch real trees under this contract? | This section + local verification-demo path | HumanEval/MBPP-style completion benches are **not** sufficient for promotion. Public leaderboards do **not** override sealed false-completion or scope gates. Sealed matrices do **not** replace live receipts for graduation. @@ -107,11 +139,11 @@ Minimum live receipt fields: - unified ledger label when available (`verified-supported` | `honestly-abstained` | `unsupported-assertion` | `out-of-scope`) - trajectory summary when the adapter supports tools: commit-class tags (`lookup` / `verify` / `commit` / `finish`) and `commit_fired` -Public demo instrumentation: `30_projects/verified-done` (`runner/run.py live`, `LEDGER.md`). Private lab matrices remain in `agent-harness-eval`. +Public demo instrumentation: `local-only: verification-demo project` (`runner/run.py live`, `LEDGER.md`). Private lab matrices remain in `local-only: harness-evaluation project`. The Local Agent workstation may display only **sanitized aggregates** from the evaluator. Raw prompts, transcripts, diffs, and verifier output stay in the private eval project. -Current matrix status and next gates live in `30_projects/agent-harness-eval/README.md`, `methodology-approach.md`, and dated outputs. This file records the graduation rule, not the live scoreboard. +Current matrix status and next gates live in `local-only: harness-evaluation README`, `methodology-approach.md`, and dated outputs. This file records the graduation rule, not the live scoreboard. ## MindGraph: two indexes @@ -142,7 +174,10 @@ Do **not** re-stage for pure `20_live` telemetry, workstation UI, or knowledge-o - **Tool contract:** MCP `query(question, scope, …)` and `graph_neighbors(doc_id, scope)` require `scope` ∈ {`knowledge`, `projects`}. Response includes `trust_profile`. No blended scope. - **Debug only:** `serve-mcp --db <path>` for single-DB isolation; never the daily default. -MindGraph's active engine source lives in `mindgraph/`. That source project is not automatically folded into the durable knowledge index. Keep sandbox/upgrade work in `30_projects/mindgraph/` and retrieval-quality measurements in `30_projects/mindgraph-eval/`. +MindGraph's active engine source lives in `30_projects/mindgraph/workbench/`. +Root `mindgraph/` is the promoted operational artifact; update it through +`bin/mindgraph-promote`. Engine source is outside the durable knowledge index. +Retrieval-quality measurements live in `30_projects/mindgraph-eval/`. Query intent routing: diff --git a/README.md b/README.md index 7ebdfbe..f9e196d 100644 --- a/README.md +++ b/README.md @@ -1,221 +1,117 @@ -# Mainframe +# MainFrame -Mainframe is the markdown-first workspace I use to organize knowledge by information lifecycle before topic. It solves a practical recall problem: quick captures, durable notes, live state, and project work need different update rules, but they still need to stay easy to find. +MainFrame is a local-first, Markdown-centered operating environment for agent-assisted knowledge work. It separates capture, durable knowledge, volatile state, bounded projects, standing operations, and archive so each can have different update and authority rules. -The source of truth is the file tree. Scripts and MindGraph can index, check, or summarize parts of the tree, but they do not replace the notes, raw sources, decisions, or project records stored here. +This repository is a **reference implementation**. It contains portable contracts, deterministic control surfaces, synthetic examples, and a bounded self-evaluation engine. It does not contain private knowledge, live state, project evidence, or evaluation history. -## Lifecycle model - -| Path | Purpose | Update rule | -| --- | --- | --- | -| `00_inbox/` | Fast capture zone for unsorted material. | Treat as temporary intake. Move through `01_ingest/` before promoting. | -| `01_ingest/` | Normalization, validation, routing, and rejected items. | Use deterministic workflows where possible. Preserve routing logs locally. | -| `10_knowledge/` | Durable, slower-moving notes and raw evidence references. | Notes need standard metadata. The live `index.md` is local/private; the repo tracks `index.template.md`. Extracted text must point back to source evidence. | -| `20_live/` | Volatile personal and project state, and active research. | Use append-only timelines or explicit snapshots. Do not silently overwrite current state. | -| `30_projects/` | Active work with outcomes and local project workbenches. | Project `README.md` metadata drives the ignored local `30_projects/index.md`; the public repo tracks `30_projects/index.template.md`. | -| `90_archive/` | Preserved material that should not clutter active navigation. | Archive without deleting raw evidence or rewriting history. | - -Local `AGENTS.md` files may add stricter rules inside a lifecycle folder. The main examples today are `20_live/AGENTS.md` and `30_projects/AGENTS.md`. - -## Metadata - -Finalized markdown notes use the schema defined in [.context/primitives.md](.context/primitives.md): - -```yaml ---- -title: "Name of the file" -domain: "Broad area" -type: raw | note | live | project | decision -status: queued | active | stable | archived -source: "URL or local path to raw evidence" -tags: ["sensitivity", "etc"] ---- -``` - -Project records extend that schema with `project_state`, `goal`, `next_action`, and `updated`. See `30_projects/AGENTS.md` for the exact project README shape. - -Architecture and workflow changes belong in [DECISIONS.md](DECISIONS.md). Claim discipline is defined in [EPISTEMIC_STANCE.md](EPISTEMIC_STANCE.md). Operational procedure: [.context/workflows/epistemic-standard.md](.context/workflows/epistemic-standard.md) (ADR-029). Canonical sources live locally in `10_knowledge/knowledge-systems/epistemics/`. - -## Deterministic scripts - -| Script | What it does | -| --- | --- | -| `bin/ingest-minion` | Routes files from `00_inbox/` and `01_ingest/queue/` through the v1 ingest path. Defaults to dry-run behavior unless `--apply` is passed. | -| `bin/fetch-source-text` | Pulls OA full text (or best excerpt) into `type: raw` stubs under `10_knowledge/`. Uses Europe PMC, direct PDF/HTML, and Unpaywall when `UNPAYWALL_EMAIL` is set. | -| `bin/post-route-enrich` | Post-ingest step: `fetch-source-text --apply` for a domain or file, then `mindgraph-refresh`. Companion to `.context/workflows/ingest-minion.md` step 8. | -| `bin/second-brain-batch` | Registers a rolling migration drop as an immutable batch with a manifest, byte-preserving `source-files/` snapshot, duplicate report, and append-only `disposition-ledger.csv`. | -| `bin/sync-project-index` | Generates or checks the ignored local `30_projects/index.md` from project README metadata. `--check` also enforces evidence-based project states: `active` needs activity within 14 days (file mtimes + nested-repo commits), valid state vocabulary, reentry pointers, dual-pool WIP (5 product seats + eval seats, total ceiling 10; ADR-041 / ADR-046). | -| `bin/mindgraph-refresh` | Refreshes the external durable-knowledge MindGraph database from `10_knowledge/`. Supports `--dry-run`. | -| `bin/mindgraph-refresh-projects` | Refreshes the projects MindGraph DB. Lean default (no bulk outputs); `--deep` adds outputs. Supports `--dry-run`, `--full`, `--apply`. | -| `bin/mindgraph-projects-apply` | **Recommended** staged apply (ADR-045): `--plan` → `--stage` → `--promote --receipt …`. Optional `--deep`. See HARNESS for when to re-stage. | -| `bin/workflow-report` | Summarizes redacted local workflow telemetry and reports coverage limits, overall and per client. Supports `--json`. | -| `bin/ingest-status` | Reports ingest lane ages and splits inbox backlog into batch-registered migration files vs organic captures. Supports `--json`. | -| `bin/audit-sweep` | Deterministic discovery of `needs-audit` tagged items (and recent routed material) in `10_knowledge/`. Writes a manifest into `20_live/epistemic-audit/pending-review/` and hands off to the epistemic auditor (`bin/epistemic`). Supports `--dry-run`, `--apply`, `--json`, `--subset`. Companion workflow: `.context/workflows/audit-sweep.md`. | -| `bin/session-open` | Loads session context files in a fixed order. Auto-detects active project from `STATE.md`; supports `--project`, `--print-contents`, and `--json`. | -| `bin/session-close` | Runs end-of-session checks and triggers downstream scripts. `--check` reports what needs doing; `--apply` runs auto actions; `--checkpoint` takes a fast evidence snapshot (wired to the PreCompact hook); `--feed` appends results to `20_live/workstation/session-close-feed.jsonl` for the workstation tracker (ADR-040). | -| `bin/extract-knowledge` | Validates prerequisites and scaffolds a knowledge note from a project. `--check` validates; `--write` creates the scaffold. | -| `bin/eval-schedule` | Runs scheduled MainFrame eval suites (`daily` / `weekly` / `monthly`), writes `20_live/eval-registry/schedule-runs.jsonl`, installs macOS launchd agents, and exposes `check` for staleness. Surfaced in `session-open` / `session-close`. See `20_live/eval-registry/OPERATOR.md` and `.context/workflows/eval-schedule.md`. | -| `bin/eval-registry` | Harvests metric extracts from eval-profile project outputs into `20_live/eval-registry/`. | - -The ingest Minion workflow is documented in [.context/workflows/ingest-minion.md](.context/workflows/ingest-minion.md). ADR-007 in [DECISIONS.md](DECISIONS.md) records why v1 is manual, dry-run-first, and limited to deterministic routing. ADR-008 records the session lifecycle scripts boundary. -ADR-011 records the suggestion-first inbox rule: first-pass intake should guide the ingest-agent instead of dead-lettering recoverable captures, while strict rejects remain part of the `queue/ → 10_knowledge/` gate. -ADR-019 adds tiered batch ingest for registered migration backlogs: files matching a named rule in [.context/routing-policy.md](.context/routing-policy.md) route in bulk behind one review table, exceptions queue for the operator, and routed clips are tagged `needs-audit` for continuous post-route auditing. - -## Safe operating rules - -Preserve provenance. Raw sources are evidence. Extracted text and generated stubs are searchable working copies, not replacements for the original material. - -Do not silently overwrite history. If a destination already exists, the deterministic ingest path blocks instead of replacing it. Project logs, decisions, and live records should append or snapshot. - -Treat `20_live/` as volatile. Current-state claims need dates, and high-risk domains need source-backed verification before promotion. - -Keep generated and personal state out of the tracked surface. The repo ignores inbox captures, ingest queues, processed knowledge domains, live telemetry, project contents, archive contents, local MCP config, local databases, and `STATE.md`. - -## Common workflows - -Start an ingest pass with a dry run: - -```bash -bin/ingest-minion run --dry-run -``` - -Warnings in the dry run are suggestions for the ingest-agent. They do not move files to `01_ingest/rejected/` during pass 1. - -Apply the planned ingest moves after reviewing the dry run: - -```bash -bin/ingest-minion run --apply -``` - -Route a raw PDF without a convention-named domain only after the destination domain exists: - -```bash -bin/ingest-minion run --dry-run --domain ai-systems -bin/ingest-minion run --apply --domain ai-systems -``` - -Refresh MindGraph after durable knowledge changes: - -```bash -bin/mindgraph-refresh -``` - -Preview the MindGraph refresh command path: +## Run the public core ```bash -bin/mindgraph-refresh --dry-run -``` - -Query the Mainframe MindGraph database: +git clone https://github.com/camerontjs-dot/MainFrame.git +cd MainFrame -```bash -bin/mindgraph query "agentic design patterns" +./bin/process-eval status --json +uvx --with pytest pytest tests 40_operations/mainframe-process-eval/tests -q ``` -Register a second-brain migration drop without mutating the source inbox: +The test surface exercises lifecycle identity, fail-closed duplicate and missing authority behavior, session routing, the process-eval shim, and the bounded public MPE profile. `process-eval status` should resolve `40_operations/mainframe-process-eval` without requiring private evaluation output. -```bash -bin/second-brain-batch \ - --batch-id YYYY-MM-DD-NNN \ - --source 00_inbox \ - --source-label "old second-brain export" -``` +## Why it exists -Regenerate the local project index after project README metadata changes: +Captures, durable notes, volatile status, bounded projects, and standing operations need different update rules. Mixing them in one pile makes recall and safe updates harder. MainFrame organizes work by **information lifecycle first, topic second**, with explicit authority and fail-closed tools around identity and publication. -```bash -bin/sync-project-index --write -``` +## Architecture -Check that the local project index is current: - -```bash -bin/sync-project-index --check +```text +PRIVATE / USER MAINFRAME PUBLIC MAINFRAME +living knowledge + evidence portable architecture +projects, operations, runtime contracts, engines, tests + | ^ + | positive fresh-tree publication | + +-----------------------------------------+ ``` -Review local workflow telemetry: +The public tree is cut from a private working tree through a positive allowlist into a fresh directory before publication. Private Git history does not cross that boundary. -```bash -bin/workflow-report --days 7 -``` +## Lifecycle model -Check ingest lane ages and migration-backlog composition: +| Path | Purpose | +| --- | --- | +| `00_inbox/` | Fast capture | +| `01_ingest/` | Normalization and routing | +| `10_knowledge/` | Durable knowledge | +| `20_live/` | Volatile-state interfaces and templates, not someone else's live state | +| `30_projects/` | Bounded outcome work | +| `40_operations/` | Standing systems and recurring programs | +| `90_archive/` | Retired or preserved material | -```bash -bin/ingest-status -``` +Lifecycle zones may carry local `AGENTS.md` contracts that refine root policy. Location does not itself decide WIP, focus, health, or approval. Direct file authority outranks indexes and retrieval. -Run a process evaluation with a baseline, representative outcome samples, and -a post-change rerun: +## Projects vs operations -See [.context/workflows/process-evaluation.md](.context/workflows/process-evaluation.md). +Slugs are unique across `30_projects/` and `40_operations/`. A project is a bounded outcome. An operation is a standing loop. Duplicate slugs and missing README authority fail closed. -The end-to-end operator walkthrough for a working session is -[.context/workflows/session-guide.md](.context/workflows/session-guide.md). +The implementation and tests are in `scripts/lifecycle_identity.py` and `tests/test_lifecycle_identity.py`. -**Client configuration note (Fix 5):** Hook telemetry for Claude, Codex, Antigravity, and Aider is currently in per-client dotfiles (.claude/, .codex/, .antigravity/). A future unification generator could centralize definitions (see .context/ or scripts/) to reduce maintenance while keeping `bin/workflow-event` as the single source for redacted events. +## MainFrame Process Evaluation -Load session context at the start of a work session: +The portable evaluation engine lives at `40_operations/mainframe-process-eval/`. Generic lifecycle substrate stays at the repository root; evaluation evidence stays local. -```bash -bin/session-open -``` +The public profile includes: -Load context with file contents for a specific project: +- operation identity and status; +- process-evaluation methodology; +- the real loop-evaluation scorer; +- labelled synthetic evaluation cases; +- a bounded `close --write` path whose output remains local-only. -```bash -bin/session-open --project my-project --print-contents -``` - -Check what needs doing before closing a session: - -```bash -bin/session-close --check -``` +Malformed synthetic catalogue input fails closed. The installed private MainFrame uses a larger `process-eval preflight` pack with additional operational tools. That full pack is not claimed as publicly reproducible here. -Run end-of-session auto actions (index sync, MindGraph refresh, telemetry): +## Agent and runtime integration -```bash -bin/session-close --apply -``` +Agents read contracts from `AGENTS.md`, `.context/workflows/`, and `.agents/skills/`. `bin/session-open` lists context; it does not establish that an agent understood it. Client-specific runtimes are optional. -Validate prerequisites for extracting knowledge from a project: +## Authority boundaries -```bash -bin/extract-knowledge --project my-project --domain ai-systems --title "Lessons from My Project" --check -``` +- Direct README and contract files are authority. +- Retrieval nominates evidence; it does not establish truth or lifecycle state. +- Deterministic tools prefer `--check` or dry-run behavior and fail closed on missing identity, duplicate slugs, or broken inputs. +- Publication uses a positive allowlist. A newly tracked private file does not become public by default. -Scaffold the knowledge note after validation passes: +## MindGraph retrieval -```bash -bin/extract-knowledge --project my-project --domain ai-systems --title "Lessons from My Project" --write -``` +MindGraph is an optional public component and is not included in this first vNext core. When the retrieval engine is present, returned chunks are nominations. Inspect the underlying note before treating a result as true. -Run the full test suite: +## Evidence in this tree -```bash -python3 -m unittest discover -s tests -``` +The public tests and synthetic fixtures are meant to be inspected, not just counted: -Review navigation/synthesis signals and generate handoff draft: +- `tests/test_lifecycle_identity.py` exercises project/operation identity and failure cases. +- `tests/test_public_mpe_synthetic.py` exercises the bounded public evaluation profile. +- `40_operations/mainframe-process-eval/tests/test_process_eval.py` exercises operation-owned evaluator behavior. +- `tests/test_session_open.py` and `tests/test_session_open_routes.py` exercise session context and routing behavior. +- `examples/demo-mainframe/` provides synthetic project, operation, and evaluation fixtures. -```bash -bin/knowledge-report --domain agents -bin/session-close --apply # writes 20_live/last-handoff-draft.md (ignored) -bin/audit-sweep --dry-run --subset regulated-systems -``` +The public branch is also exercised in GitHub Actions using the same public-core test and CLI surfaces. -## MindGraph boundary +## What is deliberately not included -Mainframe is the markdown-first workspace. MindGraph is the complementary retrieval engine. The engine source is embedded directly at the root under `mindgraph/`, with `30_projects/mindgraph/` maintained as an upgrade sandbox and `30_projects/mindgraph-eval/` kept separate for retrieval-quality measurement. The operating boundary is recorded in ADR-031 and other decisions in [DECISIONS.md](DECISIONS.md). +- Private knowledge bases, inbox contents, and archives +- Live focus, handoff, or telemetry state +- Private project contents +- Workstation UI +- Real evaluation outputs, transcripts, and baselines +- Private Git history +- The private-to-public exporter itself -By default, `bin/mindgraph-refresh` ingests `10_knowledge/` into `~/.mindgraph/mainframe.sqlite` with durable-knowledge provenance. `bin/mindgraph-refresh-projects` ingests an optional ignored local manifest (`30_projects/mindgraph-projects.json`) or its built-in curated project list into `~/.mindgraph/mainframe-projects.sqlite` with project-status provenance. The wrapper `bin/mindgraph` resolves the real MindGraph binary from `MINDGRAPH_BIN`, the root venv at `mindgraph/.venv/bin/mindgraph`, `git config mainframe.mindgraphBin`, or `PATH`. +## Limitations -The MindGraph Query Station is a MainFrame interpreter over the separate knowledge and project databases, not a merged graph. It lives in the `workstation/` component, which is not part of this public skeleton. Its v1 modes are `knowledge`, `projects`, and grouped `federated`; for unfinished station modes, run explicit CLI queries against both DBs and keep results grouped by lifecycle/trust zone. +- This is not an autonomous agent framework or enterprise orchestrator. +- Full process-eval `preflight` talks to adjacent MainFrame tools; some steps are machine-bound and are outside the bounded public profile. +- MindGraph, Workstation, and many operator CLIs are not in this first vNext core. +- `session-open` intentionally fails closed when local `STATE.md` authority is absent. +- No claim is made that a private MainFrame is healthy or that Git contains evaluation history. -Root MindGraph also provides an opt-in loopback Streamable HTTP daemon and an -official-SDK stdio proxy. It opens both indexes read-only but requires one -explicit scope per call and returns the scope trust profile; it never blends -the stores. The default MCP example remains the single-database stdio path. +## Publication boundary -Returned chunks are retrieval nominations, not verification. Inspect the underlying note and source evidence before treating a result as true. +The public repository keeps its own Git history. A private exporter builds a fresh candidate from an exact private source identity, a positive manifest, and curated public variants. Publication remains a deliberate PR into this repository rather than a mirror or history copy. diff --git a/agents/ingest-agent.md b/agents/ingest-agent.md index e65457a..882fee8 100644 --- a/agents/ingest-agent.md +++ b/agents/ingest-agent.md @@ -36,7 +36,7 @@ Do not invoke automatically on minion runs. The user controls when judgment work For organic (non-batch) captures the ingest-agent now delegates the detailed enrichment loop to the reusable skill: -**Invoke the `ingest-source` skill** (see [.agents/skills/ingest-source.md](../.agents/skills/ingest-source.md)) on one file at a time. +**Invoke the `ingest-source` skill** (see [.agents/skills/ingest-source/SKILL.md](../.agents/skills/ingest-source/SKILL.md)) on one file at a time. The skill performs: - Read + classification proposal (domain/type/tags, using `classify-note` sub-skill where helpful) @@ -63,9 +63,9 @@ Use when the target files belong to a registered batch (check `bin/ingest-status 1. **Load policy** — read [.context/routing-policy.md](../.context/routing-policy.md). Evaluation order: sensitivity overrides (S), then park rules (P), then routing rules (R); first match wins; no match → Tier B. 2. **Classify the whole lane in one pass** — for each file record: matched rule (or none), proposed domain, type, tags, canonical filename, tier. -3. **Emit one review table** — write `review-table.md` into the batch folder (`30_projects/second-brain-migration/raw-materials/batches/<batch-id>/`), or `01_ingest/` for non-batch backlogs. One row per file (file → rule → destination → tier), tier counts at the top, Tier B rows grouped with their open questions, Tier C rows listed separately and never auto-applied. +3. **Emit one review table** — write `review-table.md` into the local-only batch folder (`local-only: 30_projects/<batch-project>/raw-materials/batches/<batch-id>/`), or `01_ingest/` for non-batch backlogs. One row per file (file → rule → destination → tier), tier counts at the top, Tier B rows grouped with their open questions, Tier C rows listed separately and never auto-applied. 4. **One approval** — the user approves the table as a whole with line-item corrections. The first full-table review doubles as the ADR-018 baseline: record corrections per rule in the table file so rule-level agreement is measurable. -5. **Apply Tier A** — set frontmatter, rename to convention, run `bin/prep-ingest run --apply`, then `bin/ingest-minion run --apply`. **Explicitly ensure `needs-audit` (or `needs-verification`) is present in the `tags:` list** for any raw or low-synthesis capture being auto-routed. This is the signal for the post-placement epistemic audit sweep (see [.context/workflows/audit-sweep.md](../.context/workflows/audit-sweep.md)). Append disposition-ledger rows citing the rule that fired (for example `rule:R3`). Apply P-rule parks per policy with `parked` ledger rows. +5. **Apply Tier A** — set frontmatter, rename to convention, run `bin/prep-ingest run --apply`, then `bin/ingest-minion run --apply`. **Explicitly ensure `needs-audit` (or `needs-verification`) is present in the `tags:` list** for any raw or low-synthesis capture being auto-routed. This is the signal for the post-placement epistemic audit sweep (optional/deferred: `bin/audit-sweep` and its workflow are not in this public-core slice). Append disposition-ledger rows citing the rule that fired (for example `rule:R3`). Apply P-rule parks per policy with `parked` ledger rows. 6. **Leave Tier B in `ready/`** — tag `routing-exception`, add a one-line `routing_note:`. They surface in `bin/ingest-status` and the next review table; corrections that repeat should graduate into named policy rules by commit. For any Tier A items you do apply in batch, double-check that `needs-audit` (or equivalent) made it into the tags list so the new sweep workflow can pick them up. 7. **Report** — counts routed/parked/excepted per rule, corrections per rule, then run `bin/post-route-enrich --subset <domain>` for each domain that received raw stubs (or `bin/mindgraph-refresh` if notes only). @@ -86,7 +86,7 @@ See [01_ingest/AGENTS.md](../01_ingest/AGENTS.md) for the defensive constraints ## Related -- [planning/mainframe-agent-ingest-plan.md](../../../planning/mainframe-agent-ingest-plan.md) — design plan (v2) +- `local-only: planning notes` — historical ingest-agent design plan; not part of this tree - [DECISIONS.md](../DECISIONS.md) — ADR-009 (two-pass design), ADR-010 (layout convention) - [.context/workflows/ingest-minion.md](../.context/workflows/ingest-minion.md) — the deterministic counterpart - [.context/primitives.md](../.context/primitives.md) — schema and status lifecycle diff --git a/bin/aider-watcher b/bin/aider-watcher deleted file mode 100755 index d713571..0000000 --- a/bin/aider-watcher +++ /dev/null @@ -1,496 +0,0 @@ -#!/usr/bin/env python3 -"""Tail Aider histories and emit conservative, redacted workflow telemetry.""" - -from __future__ import annotations - -import argparse -import hashlib -import json -import re -import subprocess -import sys -import time -from collections.abc import Callable, Iterable -from pathlib import Path - - -POLL_INTERVAL = 1.0 -DISCOVERY_INTERVAL = 10.0 -# Default to the MainFrame root that contains this bin/ wrapper. -_MAINFRAME_ROOT = Path(__file__).resolve().parent.parent -DEFAULT_WORKSPACES = (_MAINFRAME_ROOT,) - -SESSION_START_RE = re.compile(r"^# aider chat started at (.*)") -MODEL_RE = re.compile(r"^>\s*Model:\s*(\S+)") -PROMPT_RE = re.compile(r"^####\s+(.+)") -RUN_RE = re.compile(r"^/run\s+(.+)") -BANG_RE = re.compile(r"^!\s*(.+)") -SHELL_RE = re.compile(r"^\$\s+(.+)") -PROPOSED_COMMAND_RE = re.compile(r"^>\s+(.+)") -SHELL_APPROVAL_RE = re.compile( - r"^>\s*Run shell commands\?.*\[[^]]+\]:\s*(\S+)", re.IGNORECASE -) -DIFF_RE = re.compile(r"^diff --git") -APPLIED_EDIT_RE = re.compile( - r"(?:Applied edit to\s+|SEARCH/REPLACE blocks? were applied successfully)", - re.IGNORECASE, -) -APPLIED_COUNT_RE = re.compile(r"(\d+)\s+SEARCH/REPLACE blocks?", re.IGNORECASE) -CONTEXT_LIMIT_RE = re.compile( - r"(?:context.*exceeds.*limit|token limit.*exceeded)", re.IGNORECASE -) -EDIT_FAILURE_RE = re.compile( - r"(?:did not conform to the edit format|SEARCH/REPLACE block failed|" - r"SearchReplaceNoExactMatch)", - re.IGNORECASE, -) -FAILURE_RE = re.compile( - r"(?:traceback|command failed|exit code [1-9])", re.IGNORECASE -) - -VERIFICATION_TOKENS = { - "check", - "clippy", - "ctest", - "eslint", - "jest", - "lint", - "mocha", - "mypy", - "phpunit", - "pytest", - "rspec", - "ruff", - "test", - "tox", - "tsc", - "typecheck", - "unittest", - "vitest", -} -KNOWN_COMMAND_HEADS = VERIFICATION_TOKENS | { - "bash", - "cargo", - "git", - "make", - "node", - "npm", - "npx", - "pnpm", - "python", - "python3", - "sh", - "yarn", -} - -Emitter = Callable[[dict[str, object]], None] - - -def new_state(size: int = 0) -> dict[str, object]: - return { - "size": size, - "partial": b"", - "session_id": None, - "session_started": False, - "model": "unknown", - "diagnostics_seen": set(), - "pending_shell_command": None, - } - - -def extract_direct_command(line: str) -> str | None: - for pattern in (RUN_RE, BANG_RE, SHELL_RE): - match = pattern.match(line.strip()) - if match: - return match.group(1).strip() - return None - - -def extract_proposed_command(line: str) -> str | None: - match = PROPOSED_COMMAND_RE.match(line.strip()) - if not match: - return None - candidate = match.group(1).strip() - tokens = candidate.split() - if not tokens: - return None - head = Path(tokens[0]).name.lower() - if head not in KNOWN_COMMAND_HEADS: - return None - return candidate - - -def is_test_command(command: str) -> bool: - tokens = command.split() - if not tokens: - return False - normalized = { - Path(token.strip(".,:;()[]{}")).name.lower() - for token in tokens - } - return bool(normalized & VERIFICATION_TOKENS) - - -def emit_event(event_dict: dict[str, object]) -> None: - try: - bin_dir = Path(__file__).parent - proc = subprocess.run( - [ - sys.executable, - str(bin_dir / "workflow-event"), - "--client", - "local", - ], - input=json.dumps(event_dict), - capture_output=True, - text=True, - check=False, - ) - if proc.returncode != 0: - print( - f"[aider-watcher] workflow-event failed: {proc.stderr.strip()}", - file=sys.stderr, - ) - except Exception as exc: - print( - f"[aider-watcher] error calling workflow-event: {exc}", - file=sys.stderr, - ) - - -def session_id_for(filepath: Path, timestamp: str) -> str: - material = f"{filepath.resolve()}\0{timestamp}" - return f"aider-{hashlib.sha256(material.encode()).hexdigest()[:16]}" - - -def event_base(filepath: Path, state: dict[str, object]) -> dict[str, object]: - return { - "session_id": state.get("session_id") or "aider-default", - "source": "aider", - "cwd": str(filepath.parent), - "transcript_path": str(filepath), - } - - -def ensure_session_started( - filepath: Path, state: dict[str, object], emitter: Emitter -) -> None: - if state.get("session_started"): - return - payload = event_base(filepath, state) - payload.update( - { - "hook_event_name": "SessionStart", - "model": state.get("model") or "unknown", - } - ) - emitter(payload) - state["session_started"] = True - - -def end_session( - filepath: Path, - state: dict[str, object], - emitter: Emitter, - reason: str, -) -> None: - if not state.get("session_started"): - return - payload = event_base(filepath, state) - payload.update({"hook_event_name": "SessionEnd", "reason": reason}) - emitter(payload) - state["session_started"] = False - - -def emit_diagnostic( - filepath: Path, - state: dict[str, object], - emitter: Emitter, - diagnostic: str, -) -> None: - seen = state["diagnostics_seen"] - assert isinstance(seen, set) - if diagnostic in seen: - return - ensure_session_started(filepath, state, emitter) - payload = event_base(filepath, state) - payload.update( - {"hook_event_name": "Diagnostic", "diagnostic": diagnostic} - ) - emitter(payload) - seen.add(diagnostic) - - -def emit_command( - filepath: Path, - state: dict[str, object], - emitter: Emitter, - command: str, -) -> None: - ensure_session_started(filepath, state, emitter) - payload = event_base(filepath, state) - payload.update( - { - "hook_event_name": "PostToolUse", - "tool_name": "Bash", - "tool_input": {"command": command}, - "tool_purpose": ( - "verification" if is_test_command(command) else "operation" - ), - } - ) - emitter(payload) - - -def process_line( - line: str, - filepath: Path, - state: dict[str, object], - emitter: Emitter = emit_event, -) -> None: - trimmed = line.strip() - if not trimmed: - return - - start_match = SESSION_START_RE.match(trimmed) - if start_match: - end_session(filepath, state, emitter, "next_session_started") - timestamp = start_match.group(1).strip() - state.update( - { - "session_id": session_id_for(filepath, timestamp), - "session_started": False, - "model": "unknown", - "diagnostics_seen": set(), - "pending_shell_command": None, - } - ) - return - - model_match = MODEL_RE.match(trimmed) - if model_match: - state["model"] = model_match.group(1).strip() - ensure_session_started(filepath, state, emitter) - return - - prompt_match = PROMPT_RE.match(trimmed) - if prompt_match: - ensure_session_started(filepath, state, emitter) - payload = event_base(filepath, state) - payload.update( - { - "hook_event_name": "UserPromptSubmit", - "prompt": prompt_match.group(1).strip(), - } - ) - emitter(payload) - return - - approval_match = SHELL_APPROVAL_RE.match(trimmed) - if approval_match: - response = approval_match.group(1).lower() - pending = state.get("pending_shell_command") - if response in {"y", "yes", "d"} and isinstance(pending, str): - emit_command(filepath, state, emitter, pending) - elif response in {"n", "no"}: - emit_diagnostic( - filepath, - state, - emitter, - "shell_command_denied", - ) - state["pending_shell_command"] = None - return - - direct_command = extract_direct_command(trimmed) - if direct_command: - emit_command(filepath, state, emitter, direct_command) - return - - proposed_command = extract_proposed_command(trimmed) - if proposed_command: - state["pending_shell_command"] = proposed_command - return - - if CONTEXT_LIMIT_RE.search(trimmed): - emit_diagnostic( - filepath, state, emitter, "context_limit_exceeded" - ) - return - - if EDIT_FAILURE_RE.search(trimmed): - emit_diagnostic(filepath, state, emitter, "edit_format_failure") - return - - if DIFF_RE.match(trimmed) or APPLIED_EDIT_RE.search(trimmed): - ensure_session_started(filepath, state, emitter) - count_match = APPLIED_COUNT_RE.search(trimmed) - observed_count = int(count_match.group(1)) if count_match else 1 - payload = event_base(filepath, state) - payload.update( - { - "hook_event_name": "PostToolUse", - "tool_name": "Edit", - "tool_input": {}, - "observed_count": observed_count, - } - ) - emitter(payload) - return - - if FAILURE_RE.search(trimmed): - emit_diagnostic(filepath, state, emitter, "command_failure") - - -def process_bytes( - data: bytes, - filepath: Path, - state: dict[str, object], - emitter: Emitter = emit_event, - *, - final: bool = False, -) -> None: - partial = state.get("partial") - assert isinstance(partial, bytes) - parts = (partial + data).split(b"\n") - complete = parts if final else parts[:-1] - state["partial"] = b"" if final else parts[-1] - for raw_line in complete: - process_line( - raw_line.rstrip(b"\r").decode("utf-8", errors="ignore"), - filepath, - state, - emitter, - ) - - -def read_growth( - filepath: Path, - state: dict[str, object], - emitter: Emitter = emit_event, - *, - final: bool = False, -) -> None: - current_size = filepath.stat().st_size - size = state.get("size") - assert isinstance(size, int) - if current_size < size: - state.update(new_state()) - size = 0 - if current_size > size: - with filepath.open("rb") as handle: - handle.seek(size) - data = handle.read() - state["size"] = current_size - process_bytes(data, filepath, state, emitter, final=final) - elif final and state.get("partial"): - process_bytes(b"", filepath, state, emitter, final=True) - - -def scan_files(workspaces: Iterable[Path]) -> list[Path]: - found: set[Path] = set() - for workspace in workspaces: - root_history = workspace / ".aider.chat.history.md" - if root_history.exists(): - found.add(root_history) - projects_dir = workspace / "30_projects" - if projects_dir.exists(): - found.update(projects_dir.rglob(".aider.chat.history.md")) - return sorted(found) - - -def replay(paths: Iterable[Path], emitter: Emitter = emit_event) -> None: - for path in paths: - state = new_state() - read_growth(path, state, emitter, final=True) - end_session(path, state, emitter, "replay_eof") - - -def watch( - workspaces: Iterable[Path], - poll_interval: float, - discovery_interval: float, -) -> None: - print("[aider-watcher] Starting Aider chat history watcher...", flush=True) - files_state: dict[Path, dict[str, object]] = {} - for path in scan_files(workspaces): - size = path.stat().st_size - files_state[path] = new_state(size=size) - print( - f"[aider-watcher] Watching existing file: {path} " - f"(size: {size} bytes)", - flush=True, - ) - - next_discovery = time.monotonic() - while True: - try: - now = time.monotonic() - if now >= next_discovery: - current_paths = set(scan_files(workspaces)) - for path in current_paths - files_state.keys(): - files_state[path] = new_state() - print( - f"[aider-watcher] Discovered new file: {path}", - flush=True, - ) - for path in set(files_state) - current_paths: - del files_state[path] - next_discovery = now + discovery_interval - - for path, state in list(files_state.items()): - try: - read_growth(path, state) - except FileNotFoundError: - del files_state[path] - except Exception as exc: - print( - f"[aider-watcher] Error reading file {path}: {exc}", - file=sys.stderr, - flush=True, - ) - except Exception as exc: - print( - f"[aider-watcher] Loop error: {exc}", - file=sys.stderr, - flush=True, - ) - time.sleep(poll_interval) - - -def main() -> int: - parser = argparse.ArgumentParser() - parser.add_argument( - "--replay", - action="append", - type=Path, - help="parse an existing history from the beginning, then exit", - ) - parser.add_argument( - "--workspace", - action="append", - type=Path, - help="workspace root to scan; may be repeated", - ) - parser.add_argument("--poll-interval", type=float, default=POLL_INTERVAL) - parser.add_argument( - "--discovery-interval", - type=float, - default=DISCOVERY_INTERVAL, - ) - args = parser.parse_args() - - if args.replay: - replay(args.replay) - return 0 - - watch( - args.workspace or DEFAULT_WORKSPACES, - max(0.1, args.poll_interval), - max(args.poll_interval, args.discovery_interval), - ) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/bin/audit-sweep b/bin/audit-sweep deleted file mode 100755 index 12e4bb2..0000000 --- a/bin/audit-sweep +++ /dev/null @@ -1,256 +0,0 @@ -#!/usr/bin/env python3 -"""Discover needs-audit (and recent routed) items in 10_knowledge/ and surface them -for the epistemic auditor. Deterministic discovery + manifest; delegates LLM -judgment to the epistemic-research-system orchestrator. - -PROVENANCE AUDIT (added after the 2026-08-09 fabricated-citation finding, which -this sweep could not have caught: it discovers items by tag and freshness, and -never asks whether a cited source exists). Three companion tools do that, and -should be run on the same cadence as this one: - - bin/citation-resolvability-sweep --run # do the identifiers resolve? - bin/title-existence-check --from-list ... # do the papers exist? - bin/capture-validate --knowledge # is any citation unearned? - -All three run positive AND negative controls in-batch and report BROKEN rather -than clean when a control fails. See -20_live/security/2026-08-09__fabricated-source-captures-in-10-knowledge.md. - -Follows the style and guardrails of other MainFrame minions (dry-run first, -no shared library, structured reporting, respect for volatility in 20_live/). - -Usage: - bin/audit-sweep --dry-run - bin/audit-sweep --apply # write manifest + hand off to auditor - bin/audit-sweep --json - bin/audit-sweep --subset agents # limit to one domain (per ADR-021) -""" - -from __future__ import annotations - -import argparse -import json -import sys -from datetime import date, datetime -from pathlib import Path -from typing import Any - -ROOT = Path(__file__).resolve().parents[1] -KNOWLEDGE = ROOT / "10_knowledge" -LIVE_AUDIT = ROOT / "20_live" / "epistemic-audit" -EPISTEMIC_BIN = ROOT / "bin" / "epistemic" - - -def parse_simple_frontmatter_tags(text: str) -> list[str]: - """Very lightweight frontmatter tag extractor. Avoids full YAML dependency. - Looks for the first tags: [...] or tags: list under --- ... --- . - Returns lowercased tag list or []. - """ - if not text.startswith("---"): - return [] - end = text.find("\n---", 3) - if end == -1: - return [] - header = text[3:end] - tags: list[str] = [] - in_tags = False - for raw_line in header.splitlines(): - line = raw_line.strip() - if line.lower().startswith("tags:"): - in_tags = True - # tags: [a, b] form - rest = line.split(":", 1)[1].strip() - if rest.startswith("["): - content = rest.strip("[]").strip() - if content: - tags = [t.strip().strip('"').strip("'").lower() for t in content.split(",") if t.strip()] - continue - if in_tags: - if line.startswith("-"): - t = line.lstrip("- ").strip().strip('"').strip("'").lower() - if t: - tags.append(t) - elif ":" in line and not line.startswith(" "): - # next top-level key - break - return tags - - -def find_knowledge_files(subset: str | None = None) -> list[Path]: - """Return candidate .md files under 10_knowledge/ (or a domain subset).""" - base = KNOWLEDGE - if subset: - base = base / subset - if not base.is_dir(): - return [] - if not base.is_dir(): - return [] - files: list[Path] = [] - for p in base.rglob("*.md"): - if p.is_file() and not p.name.startswith(".") and "index.md" not in p.name.lower(): - # Skip the raw/ subdirs? No — we want the wrapper notes too, but typically the canonical md lives beside or above raw/. - files.append(p) - return sorted(files) - - -def collect_candidates(subset: str | None, max_days: int = 365) -> list[dict[str, Any]]: - """Scan for files carrying needs-audit (or needs-verification) and recent routed items.""" - today = date.today() - candidates: list[dict[str, Any]] = [] - for path in find_knowledge_files(subset): - try: - text = path.read_text(encoding="utf-8", errors="replace") - except Exception: - continue - - tags = parse_simple_frontmatter_tags(text) - has_needs_audit = any("needs-audit" in t or "needs-verification" in t for t in tags) - - mtime = datetime.fromtimestamp(path.stat().st_mtime).date() - age_days = max((today - mtime).days, 0) - - # Include if it has the explicit signal, or if it is recent and looks routed (convention name or raw tag) - is_recent = age_days <= max_days - looks_routed = "__" in path.name and ("raw" in path.name or any("raw" in t or "routed" in t for t in tags)) - - if has_needs_audit or (is_recent and looks_routed): - domain_guess = path.relative_to(KNOWLEDGE).parts[0] if path.is_relative_to(KNOWLEDGE) else "unknown" - candidates.append({ - "path": str(path.relative_to(ROOT)), - "domain": domain_guess, - "name": path.name, - "age_days": age_days, - "tags": tags, - "has_needs_audit": has_needs_audit, - "mtime": mtime.isoformat(), - }) - - # Sort: explicit needs-audit first, then by age asc (newest first among recent) - candidates.sort(key=lambda c: (not c["has_needs_audit"], c["age_days"])) - return candidates - - -def nonnegative_int(value: str) -> int: - """Parse a zero-or-greater integer for age-window arguments.""" - parsed = int(value) - if parsed < 0: - raise argparse.ArgumentTypeError("must be zero or greater") - return parsed - - -def write_manifest(candidates: list[dict[str, Any]], run_date: str) -> Path: - LIVE_AUDIT.mkdir(parents=True, exist_ok=True) - (LIVE_AUDIT / "pending-review").mkdir(exist_ok=True) - manifest_path = LIVE_AUDIT / "pending-review" / f"sweep-{run_date}.md" - lines = [ - "# Audit Sweep Manifest", - "", - f"Generated: {run_date} by bin/audit-sweep", - f"Total candidates: {len(candidates)}", - "", - "This file lists items that carried `needs-audit` (or were recently routed).", - "The epistemic auditor should process these and publish reports under this directory or via mainframe_bridge.", - "", - "## Candidates", - "", - ] - if not candidates: - lines.append("_No candidates found in this sweep._") - else: - lines.append("| Path | Domain | Age (days) | Has needs-audit | Tags |") - lines.append("| --- | --- | --- | --- | --- |") - for c in candidates: - tags_str = ", ".join(c["tags"][:6]) + ("..." if len(c["tags"]) > 6 else "") - has = "yes" if c["has_needs_audit"] else "no (recent sample)" - lines.append(f"| {c['path']} | {c['domain']} | {c['age_days']} | {has} | {tags_str} |") - - lines.extend([ - "", - "## Next Actions", - "- Run the auditor on high-priority items (explicit needs-audit).", - "- Review published audit reports in 20_live/epistemic-audit/.", - "- After verification, remove `needs-audit` from source frontmatter (or mark `audited-YYYY-MM-DD`).", - "- Record any policy or process findings in DECISIONS.md or the mainframe-process-eval project.", - "", - "## Synthesis Opportunities (raw accumulation signal)", - "Many `raw` items without companion synthesized `note` entries increase the burden on recall and the epistemic auditor.", - "Consider using `bin/extract-knowledge` patterns or the `create-source-summary` skill (via ingest-source) to produce higher-value notes from clusters of related raws in high-volume domains.", - "See also: local 10_knowledge/index.md promotion rules (template: 10_knowledge/index.template.md) and process-evaluation for measuring synthesis rate over time.", - ]) - content = "\n".join(lines) - manifest_path.write_text(content, encoding="utf-8") - return manifest_path - - -def handoff_to_auditor(manifest: Path, dry_run: bool) -> str: - """Call the epistemic system for judgment. In real use this would invoke - specific auditor entrypoints with focus on the manifest or a subset. - """ - if not EPISTEMIC_BIN.exists(): - return "epistemic bin not found; skipping LLM pass (install/activate the venv in the workbench)." - - if dry_run: - return f"DRY-RUN: would invoke {EPISTEMIC_BIN} with sweep focused on {manifest}" - - # Best-effort handoff. The real orchestrator may grow a `sweep` or `audit-mainframe` subcommand. - # For now we simply note that the operator (or a future orchestrator call) should process the manifest. - # We could do: subprocess with --focus or a dedicated mode. - return f"Manifest ready at {manifest}. Run: bin/epistemic sweep --mainframe --manifest {manifest} (or equivalent orchestrator command) to drive the auditor." - - -def main() -> int: - parser = argparse.ArgumentParser(description="Surface needs-audit items for epistemic review.") - parser.add_argument("--dry-run", action="store_true", help="Report what would be done; do not write manifests.") - parser.add_argument("--apply", action="store_true", help="Write the sweep manifest and trigger auditor handoff.") - parser.add_argument("--json", action="store_true", help="Output machine-readable summary.") - parser.add_argument("--subset", help="Limit scan to one 10_knowledge/<domain> (e.g. agents, finance).") - parser.add_argument( - "--max-days", - type=nonnegative_int, - default=365, - help="Include routed sample items up to this age; explicit audit signals are always included.", - ) - args = parser.parse_args() - - if not args.dry_run and not args.apply: - # Default to a safe reporting mode - args.dry_run = True - - candidates = collect_candidates(args.subset, args.max_days) - run_date = date.today().isoformat() - - if args.json: - print(json.dumps({ - "run_date": run_date, - "subset": args.subset, - "candidate_count": len(candidates), - "explicit_needs_audit": sum(1 for c in candidates if c["has_needs_audit"]), - "candidates": candidates[:50], # cap for readability - }, indent=2)) - return 0 - - print(f"Audit sweep — {run_date}") - print(f"Subset: {args.subset or 'all 10_knowledge'}") - print(f"Candidates found: {len(candidates)} (explicit needs-audit: {sum(1 for c in candidates if c['has_needs_audit'])})") - for c in candidates[:20]: - flag = " [needs-audit]" if c["has_needs_audit"] else "" - print(f" {c['path']} ({c['age_days']}d){flag}") - - if len(candidates) > 20: - print(f" ... and {len(candidates)-20} more") - - if args.dry_run: - print("\n--dry-run: no files written. Use --apply to surface a manifest in 20_live/epistemic-audit/pending-review/.") - return 0 - - # --apply path - manifest = write_manifest(candidates, run_date) - note = handoff_to_auditor(manifest, dry_run=False) - print(f"\nManifest written: {manifest}") - print(note) - print("Next: review the manifest, run the auditor on the listed items, then update source tags after verification.") - return 0 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/capture-validate b/bin/capture-validate deleted file mode 100755 index 5d54ea2..0000000 --- a/bin/capture-validate +++ /dev/null @@ -1,391 +0,0 @@ -#!/usr/bin/env python3 -"""G2 — a capture may not wear a citation it did not earn. - -Gate for the ingest pipeline, built after the 2026-08-09 integrity finding -(`20_live/security/2026-08-09__fabricated-source-captures-in-10-knowledge.md`): -50 captures in `10_knowledge/` carried invented titles, invented institutional -"authors", and identifiers that do not resolve. - -The root cause was not a bad model. The capture schema **required** `url`, -`authors` and `year`, nothing verified them, and the research loop asked for a -count of sources. Filling those fields with plausible text was the only way to -comply. This validator removes that path: - - no successful fetch -> no citation fields -> it is a hypothesis, not a source - -A capture with no retrieval is still welcome. It just cannot claim provenance. - -Rules - R1 CITATION-WITHOUT-FETCH fetch_method none/absent but url|doi|authors|year set - R2 FABRICATED-SHAPE identifier is a slug or unfilled placeholder - R3 NO-RECEIPT literature capture with no retrieval_receipt - R4 NON-HUMAN-AUTHOR "author" is an institution, committee, or venue - R5 PLACEHOLDER-TEXT template text left in any frontmatter value - R7 BARE-DOMAIN citation URL resolves only to a domain root - -Usage: - bin/capture-validate 00_inbox/foo.md ... - bin/capture-validate --inbox # gate 00_inbox/ before ingest - bin/capture-validate --ingest-ready # gate 01_ingest/ready/ - bin/capture-validate --knowledge # audit 10_knowledge/ (reports only) - bin/capture-validate --inbox --strict # promote R3/R4 warnings to errors - bin/capture-validate --inbox --json -""" - -from __future__ import annotations - -import argparse -import importlib.util -import json -import re -import sys -from dataclasses import dataclass -from importlib.machinery import SourceFileLoader -from pathlib import Path -from typing import Any -from urllib.parse import urlsplit - -ROOT = Path(__file__).resolve().parents[1] - -# Reuse the sweep's detectors rather than re-deriving them: one definition of -# "fabricated shape" for the whole repo. -_LOADER = SourceFileLoader("_sweep", str(ROOT / "bin" / "citation-resolvability-sweep")) -_SPEC = importlib.util.spec_from_loader(_LOADER.name, _LOADER) -assert _SPEC and _SPEC.loader -sweep = importlib.util.module_from_spec(_SPEC) -sys.modules[_SPEC.name] = sweep -_SPEC.loader.exec_module(sweep) - -CITATION_FIELDS = ("url", "doi", "authors", "author", "year", "journal", "venue") - -# A fetch actually happened. Anything else is "we did not retrieve this." -REAL_FETCH = { - "direct-html", "pdf", "api", "manual-paste", "browser", "webfetch", - "crossref", "pubmed", "local-file", "direct-pdf", "arxiv-api", -} -NO_FETCH = {"", "none", "n/a", "na", "pending", "unknown", "null", "-"} - -# The tell that found the original batch: 21 captures, 21 distinct "authors", -# zero human names. A conference is a venue, not an author. -NON_HUMAN_AUTHOR_RE = re.compile( - r"(consortium|task\s*force|working\s*group|standards?\s*group|committee|" - r"conference|symposium|joint\s*\w+|initiative|alliance|board|council)\b", - re.I, -) -HUMAN_NAME_HINT_RE = re.compile(r"[A-Z][a-z]+,\s*[A-Z]|\bet\s+al\.|[A-Z]\.\s*[A-Z][a-z]+") - -PLACEHOLDER_TEXT_RE = re.compile( - r"\b(TODO|TBD|FIXME|lorem ipsum|placeholder|your[-_ ]?name|xxxx+|" - r"example\.com|<[a-z-]+>)\b", - re.I, -) - -# Statuses that assert a human or a deterministic sweep has checked the content. -# `synthesized` is deliberately absent: EPISTEMIC_STANCE.md's promotion table -# defines it as "claim types labeled; confidence assigned", which an LLM may do -# on its own. `stable` and `audited` are the ones that claim someone looked. -VERIFIED_STATUSES = {"stable", "audited"} - -SEVERITY_ORDER = {"error": 0, "warn": 1} - - -@dataclass -class Finding: - rule: str - severity: str - message: str - detail: str = "" - - -def parse_frontmatter(text: str) -> dict[str, str]: - """Lightweight scalar frontmatter reader. Lists collapse to their raw text.""" - if not text.startswith("---"): - return {} - end = text.find("\n---", 3) - if end == -1: - return {} - out: dict[str, str] = {} - for line in text[3:end].splitlines(): - m = re.match(r"^([A-Za-z_][A-Za-z0-9_]*):\s*(.*)$", line) - if m: - out[m.group(1).strip().lower()] = m.group(2).strip().strip('"').strip("'") - return out - - -def body_fetch_method(text: str) -> str | None: - """Captures also record fetch method in the body: '**Fetch method:** none'.""" - m = re.search(r"\*\*Fetch method:\*\*\s*(.+)", text, re.I) - return m.group(1).strip().lower() if m else None - - -def had_real_fetch(fm: dict[str, str], text: str) -> bool: - for value in (fm.get("fetch_method"), body_fetch_method(text)): - if value and value.strip().lower() not in NO_FETCH: - return True - if fm.get("retrieval_receipt"): - return True - # An explicit "full text retrieved / verified" marker also counts. - if re.search(r"\*\*Full[- ]text:\*\*\s*verified|full-text-retrieved", text, re.I): - return True - return False - - -def nonempty_citation_fields(fm: dict[str, str]) -> dict[str, str]: - return { - k: v for k in CITATION_FIELDS - if (v := fm.get(k, "").strip()) and v.lower() not in {"none", "n/a", "unknown", "[]"} - } - - -def is_literature(fm: dict[str, str]) -> bool: - blob = f"{fm.get('source_type', '')} {fm.get('type', '')} {fm.get('tags', '')}".lower() - return any(w in blob for w in ("literature", "paper", "preprint", "journal", - "peer-review", "regulation", "standard")) - - -def is_bare_domain_url(value: str) -> bool: - """Return whether an HTTP(S) URL names only its domain root. - - A root URL may resolve successfully while failing to identify the document - that supports a claim. That is a provenance warning, not evidence that the - source itself is fabricated, so callers should keep this check warning-only. - """ - try: - parsed = urlsplit(value.strip()) - except ValueError: - return False - return ( - parsed.scheme in {"http", "https"} - and bool(parsed.netloc) - and parsed.path in {"", "/"} - and not parsed.query - and not parsed.fragment - ) - - -def validate(path: Path, strict: bool = False) -> list[Finding]: - try: - text = path.read_text(encoding="utf-8", errors="replace") - except OSError as exc: - return [Finding("R0", "error", f"unreadable: {exc}")] - - fm = parse_frontmatter(text) - if not fm: - return [] - - findings: list[Finding] = [] - fetched = had_real_fetch(fm, text) - cited = nonempty_citation_fields(fm) - - # R1 — the core rule. Scoped to captures that actually claim an external - # source. A MainFrame-authored spec or template legitimately carries an - # `authors:` line with nothing to fetch; flagging those would drown the - # signal in noise and teach people to ignore the validator. - claims_external = bool( - (fm.get("url", "") or fm.get("doi", "") or fm.get("source", "")).strip().startswith("http") - ) - if not fetched and cited and claims_external: - findings.append(Finding( - "R1", "error", - "citation fields present but nothing was fetched", - "fields: " + ", ".join(f"{k}={v!r}" for k, v in cited.items()) - + "\n Fix: either record a real retrieval (fetch_method + " - "retrieval_receipt), or clear these fields and set type: hypothesis.", - )) - - # R2 — fabricated identifier shape, no network needed. - for field in ("url", "doi"): - val = fm.get(field, "").strip() - if val and (reason := sweep.structural_verdict(val)): - findings.append(Finding("R2", "error", f"{field} is not a real identifier", - f"{val}\n {reason}")) - for url in sweep.URL_RE.findall(text): - url = sweep.strip_trailing_punct(url) - if (reason := sweep.structural_verdict(url)): - findings.append(Finding("R2", "error", "fabricated identifier in body", - f"{url}\n {reason}")) - - # R3 — a literature claim needs a receipt. - if is_literature(fm) and not fm.get("retrieval_receipt") and cited: - findings.append(Finding( - "R3", "error" if strict else "warn", - "literature capture has no retrieval_receipt", - "Add retrieval_receipt: \"<ISO8601> HTTP <status> sha256:<hash of what " - "was fetched>\" — or drop the citation fields.", - )) - - # R4 — "author" that is not a person. - authors = fm.get("authors") or fm.get("author") or "" - if authors and NON_HUMAN_AUTHOR_RE.search(authors) and not HUMAN_NAME_HINT_RE.search(authors): - findings.append(Finding( - "R4", "error" if strict else "warn", - "author looks like an institution or venue, not a researcher", - f"authors: {authors!r}\n Every fabricated capture in the 2026-08-09 " - "finding had a byline of exactly this shape and zero named researchers.", - )) - - # R5 — unfilled template text. - for key, val in fm.items(): - if val and PLACEHOLDER_TEXT_RE.search(val): - findings.append(Finding("R5", "warn", f"placeholder text in '{key}'", repr(val))) - - # R7 — a resolving homepage is not a document citation. Limit this to - # frontmatter source fields so ordinary prose/code examples do not become - # citation findings. This deliberately remains a warning: a domain root - # can be a useful navigation source, but it cannot by itself substantiate a - # document-level claim. - seen_bare_urls: set[str] = set() - for field in ("source", "source_url", "url"): - for value in sweep.URL_RE.findall(fm.get(field, "")): - value = sweep.strip_trailing_punct(value) - if value in seen_bare_urls or not is_bare_domain_url(value): - continue - seen_bare_urls.add(value) - findings.append(Finding( - "R7", "warn", "citation URL is a bare domain", - f"{field}: {value}\n Fix: record the specific document URL, or mark the source as a lead/manual-review item.", - )) - - # R6 — a verified status must be earned. - # - # EPISTEMIC_STANCE.md's promotion gate says `stable` is "Operator-verified or - # audit-sweep cleared — **never** from LLM output alone". Measured on - # 2026-08-10, nothing in the repo enforced that: 29 files claimed `stable`, - # all 29 carried no verification evidence of any kind, and **16 of them - # simultaneously carried the `needs-audit` tag**. The highest trust tier in - # the system was 55% self-contradictory. - # - # R6a needs no judgement at all. A file cannot be both operator-verified and - # awaiting audit; one of the two fields is simply false, and which one does - # not matter for the purpose of refusing to trust it. - status = fm.get("status", "").strip().strip('"').lower() - tags = fm.get("tags", "") - if status in VERIFIED_STATUSES and "needs-audit" in tags: - findings.append(Finding( - "R6", "error", - f"status '{status}' contradicts the needs-audit tag", - "A note cannot be operator-verified and awaiting audit at the same " - "time. Drop needs-audit, or lower the status to 'synthesized'.", - )) - - # R6b — the receipt must name something that can be checked. A bare - # `verified_by: operator` is still self-asserted, and this investigation has - # already established what a self-asserted field is worth: captures claiming - # `retrieved_at` were citing hard 404s. So a receipt pointing at an audit - # artifact must point at one that exists. That does not make the claim - # unforgeable — it makes it falsifiable, which is the whole difference - # between a receipt and a decoration. - if status in VERIFIED_STATUSES: - receipt = (fm.get("audit_receipt") or fm.get("verified_by") or "").strip() - if not receipt: - findings.append(Finding( - "R6", "error" if strict else "warn", - f"status '{status}' with no verification evidence", - "Add audit_receipt: <path to the audit artifact> or verified_by: " - "<who> plus verified_on: <date>. Per EPISTEMIC_STANCE.md this " - "status may never come from LLM output alone.", - )) - elif "/" in receipt and not (ROOT / receipt.strip('"')).exists(): - findings.append(Finding( - "R6", "error", - f"audit_receipt points at a file that does not exist", - f"{receipt}\n A receipt naming a missing artifact is worse " - "than no receipt: it reads as verification to everything " - "downstream and cannot be checked by anyone in a hurry.", - )) - - return findings - - -def gather(args: argparse.Namespace) -> list[Path]: - paths: list[Path] = [Path(p) for p in args.paths] - if args.inbox: - paths += sorted((ROOT / "00_inbox").glob("*.md")) - if args.ingest_ready: - paths += sorted((ROOT / "01_ingest" / "ready").rglob("*.md")) - if args.knowledge: - paths += sorted((ROOT / "10_knowledge").rglob("*.md")) - return paths - - -def main() -> int: - ap = argparse.ArgumentParser(description=__doc__, - formatter_class=argparse.RawDescriptionHelpFormatter) - ap.add_argument("paths", nargs="*") - ap.add_argument("--inbox", action="store_true") - ap.add_argument("--ingest-ready", action="store_true") - ap.add_argument("--knowledge", action="store_true") - ap.add_argument("--strict", action="store_true", - help="promote R3/R4 warnings to errors") - ap.add_argument("--json", action="store_true") - args = ap.parse_args() - - paths = gather(args) - if not paths: - ap.error("nothing to validate — pass paths or --inbox / --ingest-ready / --knowledge") - - def label(p: Path) -> str: - """Repo-relative when possible, absolute otherwise. - - `Path.relative_to` raises for anything outside ROOT, which used to abort - the whole run with a traceback the moment someone validated a scratch - file or a staging directory outside the repo. A validator that dies on an - out-of-tree path cannot be used as a control, which is how - `bin/contract-audit`'s positive control found this. - """ - try: - return str(p.relative_to(ROOT)) - except ValueError: - return str(p) - - results: dict[str, list[Finding]] = {} - for p in paths: - if fs := validate(p, strict=args.strict): - results[label(p)] = fs - - errors = sum(1 for fs in results.values() for f in fs if f.severity == "error") - warns = sum(1 for fs in results.values() for f in fs if f.severity == "warn") - - if args.json: - print(json.dumps({ - "checked": len(paths), - "files_with_findings": len(results), - "errors": errors, "warnings": warns, - "findings": {k: [vars(f) for f in v] for k, v in results.items()}, - }, indent=2)) - return 1 if errors else 0 - - for path, fs in sorted(results.items()): - print(f"\n{path}") - for f in sorted(fs, key=lambda x: SEVERITY_ORDER[x.severity]): - print(f" [{f.severity.upper():5}] {f.rule} {f.message}") - if f.detail: - for line in f.detail.splitlines(): - print(f" {line}") - - print(f"\nchecked {len(paths)} file(s): {errors} error(s), {warns} warning(s) " - f"across {len(results)} file(s)") - if errors: - print("\nA capture may not wear a citation it did not earn. " - "Fetch it, or drop the citation fields.") - return 1 if errors else 0 - - -if __name__ == "__main__": - # Exit codes are part of the contract, because a caller must be able to tell - # "ran and found problems" from "did not run". Those were both 1 until - # 2026-08-10, which means a crashed validator was indistinguishable from a - # validator reporting findings — the captive-portal failure mode, in the tool - # built to catch captive-portal failure modes. - # 0 ran, nothing wrong - # 1 ran, found errors - # 2 did not run (bad usage, unreadable path, internal failure) - try: - sys.exit(main()) - except SystemExit: - raise - except Exception as exc: # noqa: BLE001 — report, never masquerade as a result - print(f"capture-validate FAILED TO RUN: {type(exc).__name__}: {exc}", - file=sys.stderr) - print("This is not a clean result. Do not read it as one.", file=sys.stderr) - sys.exit(2) diff --git a/bin/citation-resolvability-sweep b/bin/citation-resolvability-sweep deleted file mode 100755 index 4283aea..0000000 --- a/bin/citation-resolvability-sweep +++ /dev/null @@ -1,334 +0,0 @@ -#!/usr/bin/env python3 -"""G5 — resolution-test every citation-shaped URL in 10_knowledge/. - -Built after the 2026-08-09 integrity finding -(`20_live/security/2026-08-09__fabricated-source-captures-in-10-knowledge.md`), -which found 21 fabricated DOI citations with a hand-rolled one-off probe. This -turns that probe into a standing control. - -Two independent detectors, deliberately kept separate: - - 1. STRUCTURAL (offline, high precision, zero network): - DOI suffixes that are human-readable slugs rather than registrant-issued - identifiers. `10.1145/retrieval-granularity-2025` is not a DOI shape any - registration agency emits. Catches fabrication without touching the wire. - - 2. NETWORK (resolution test): - Does the identifier resolve at all? doi.org answers 3xx for a registered - DOI and 404 for one that was never registered. - -**Positive controls run in the same batch.** If a control fails, the run is -reported as BROKEN-PROBE, never as a clean sweep. A network outage that silently -reported "all fine" would be worse than not running at all — that is the failure -mode this guard exists to prevent. - -Usage: - bin/citation-resolvability-sweep --dry-run # what would be tested - bin/citation-resolvability-sweep --structural-only # offline detector only - bin/citation-resolvability-sweep --run # full sweep (network) - bin/citation-resolvability-sweep --run --limit 50 # bounded smoke test - bin/citation-resolvability-sweep --run --json OUT.json -""" - -from __future__ import annotations - -import argparse -import json -import re -import sys -import urllib.error -import urllib.request -from collections import defaultdict -from concurrent.futures import ThreadPoolExecutor, as_completed -from datetime import datetime, timezone -from pathlib import Path -from typing import Any - -ROOT = Path(__file__).resolve().parents[1] -KNOWLEDGE = ROOT / "10_knowledge" - -URL_RE = re.compile(r"""https?://[^\s>\]"'`,]+""") - -# DOI suffixes legitimately contain parentheses — 10.1016/0895-4356(90)90159-M is -# a real Elsevier DOI. A regex that stops at the first ')' truncates it and the -# truncated form 404s, which would be reported as a fabricated citation. That is -# a false accusation, so parens are balanced rather than excluded. -# Markdown/HTML also leak into scraped URLs: '"' (a stray "), a trailing -# backslash, a ']( ' from a markdown link. All stripped before probing. -HTML_JUNK_RE = re.compile(r"(&#\d+;?|"?|&?|\\+)$") - - -def balance_parens(url: str) -> str: - """Drop a trailing ')' only when it does not close a '(' inside the URL.""" - while url.endswith(")") and url.count("(") < url.count(")"): - url = url[:-1] - return url - - -# Literal placeholders. Distinct from a slug DOI: these are not even trying to -# look real — they are template text that was never filled in. -PLACEHOLDER_RE = re.compile( - r"(?:/10\.x{3,}/)" # 10.xxxx/ — unfilled registrant prefix - r"|(?:\bx{4,}\b)" # 2403.xxxx, 2510.xxxxx - r"|(?:\b0{5,}\b)" # 2309.00000, s42256-024-00000-0 - r"|(?:\b(\d)\1{5,}\b)" # 3468.888888 - r"|placeholder|dummy-|/example[-./]", - re.I, -) - -# Hosts that carry a bibliographic claim. Social/code/vendor links are excluded: -# a dead github link is link rot, a dead DOI is a citation that never existed. -CITATION_HOSTS = ( - "doi.org", - "arxiv.org", - "aclanthology.org", - "openreview.net", - "ncbi.nlm.nih.gov", - "dl.acm.org", - "ieeexplore.ieee.org", - "link.springer.com", - "sciencedirect.com", - "nature.com", - "jmlr.org", - "proceedings.mlr.press", - "pubmed.ncbi.nlm.nih.gov", -) - -# Known-good, stable, and cheap. If any of these fails the probe is broken. -POSITIVE_CONTROLS = [ - "https://doi.org/10.1000/182", # the DOI Handbook's own DOI - "https://doi.org/10.48550/arXiv.2112.07618", - "https://arxiv.org/abs/1706.03762", # Attention Is All You Need - "https://aclanthology.org/2020.acl-main.677/", -] - -# Must NOT resolve. Guards the opposite failure: a probe that answers 200 to -# everything (captive portal, proxy interception) would report a clean sweep. -NEGATIVE_CONTROLS = [ - "https://doi.org/10.1145/this-doi-was-invented-for-a-control-2026", - "https://arxiv.org/abs/9999.99999", -] - -UA = "MainFrame-citation-sweep/1.0 (integrity audit; local)" - -# A registered DOI suffix is registrant-assigned and effectively opaque. -# Fabrications read like article titles: all-lowercase alphabetic words joined by -# hyphens, often with a trailing year. This is the detector that found the -# original 21. -# -# Deliberately narrow. Real suffixes carry structure that breaks the pattern: -# 10.1146/annurev-psych-033020-014116 digit runs that are not years -# 10.1108/ijcma-02-2018-0027 interleaved numeric segments -# 10.18653/v1/2023.emnlp-main.397 dots and path segments -# 10.1038/s41586-020-2649-2 alphanumeric registrant tokens -# A handful of publishers do mint readable suffixes, so a structural hit is a -# "check this", not a verdict — it is paired with the network probe below. -DOI_SLUG_RE = re.compile( - r"^10\.\d{4,9}/[a-z]+(?:-[a-z]+){1,}(?:-(?:19|20)\d{2})?$" -) - - -def strip_trailing_punct(url: str) -> str: - url = url.split("](")[0] # markdown link that swallowed its target - url = HTML_JUNK_RE.sub("", url) - url = url.rstrip(".,;:'\"`]}>!*_") - return balance_parens(url) - - -def is_citation_url(url: str) -> bool: - return any(h in url for h in CITATION_HOSTS) - - -def collect_urls() -> dict[str, set[str]]: - """url -> set of repo-relative file paths citing it.""" - index: dict[str, set[str]] = defaultdict(set) - for path in KNOWLEDGE.rglob("*.md"): - try: - text = path.read_text(encoding="utf-8", errors="replace") - except OSError: - continue - rel = str(path.relative_to(ROOT)) - for raw in URL_RE.findall(text): - url = strip_trailing_punct(raw) - if is_citation_url(url): - index[url].add(rel) - return index - - -def doi_suffix(url: str) -> str | None: - m = re.search(r"doi\.org/(10\.\d{4,9}/\S+)", url) - return m.group(1) if m else None - - -def structural_verdict(url: str) -> str | None: - """Offline detector. Returns a reason string when the shape is wrong.""" - if PLACEHOLDER_RE.search(url): - return "unfilled placeholder identifier (template text, never a real ID)" - suffix = doi_suffix(url) - if suffix and DOI_SLUG_RE.match(suffix.lower()): - return "slug-style DOI suffix (not a registrant-issued identifier)" - # arXiv IDs have exactly two legal forms. Rather than enumerate fabrication - # shapes (the first version matched only hyphenated slugs and so missed - # `abs/physalign` and `abs/track4animate3d`), assert the legal form and - # reject everything else. Allow-list, not deny-list. - m = re.search(r"arxiv\.org/(?:abs|pdf)/([^/?#]+)", url) - if m: - ident = m.group(1).removesuffix(".pdf") - modern = re.fullmatch(r"\d{4}\.\d{4,5}(v\d+)?", ident) - legacy = re.fullmatch(r"[a-z-]+(\.[A-Z]{2})?/\d{7}(v\d+)?", ident) - if not (modern or legacy): - return ("not a valid arXiv ID (must be YYMM.NNNNN or " - "archive/YYMMNNN)") - return None - - -def probe(url: str, timeout: float = 20.0) -> dict[str, Any]: - """Resolution test. 3xx counts as resolved — doi.org redirects on success.""" - - class NoRedirect(urllib.request.HTTPRedirectHandler): - def redirect_request(self, *_args, **_kwargs): # noqa: ANN002 - return None - - opener = urllib.request.build_opener(NoRedirect) - req = urllib.request.Request(url, method="HEAD", headers={"User-Agent": UA}) - try: - with opener.open(req, timeout=timeout) as resp: - return {"url": url, "status": resp.status, "resolved": True} - except urllib.error.HTTPError as exc: - if 300 <= exc.code < 400: - return {"url": url, "status": exc.code, "resolved": True} - if exc.code in (401, 403, 405, 429): - # Bot-blocked, auth-required, or HEAD-unsupported. Not evidence - # either way — arXiv's /auth/show-endorsers/ URLs are real pages - # behind a login, and calling them fabricated would be wrong. - return { - "url": url, - "status": exc.code, - "resolved": None, - "note": "inconclusive (blocked, auth-required, or HEAD unsupported)", - } - return {"url": url, "status": exc.code, "resolved": False} - except Exception as exc: # noqa: BLE001 - network errors are data here - return { - "url": url, - "status": None, - "resolved": None, - "note": f"inconclusive ({type(exc).__name__}: {exc})", - } - - -def run_probes(urls: list[str], workers: int = 12) -> dict[str, dict[str, Any]]: - out: dict[str, dict[str, Any]] = {} - with ThreadPoolExecutor(max_workers=workers) as pool: - futures = {pool.submit(probe, u): u for u in urls} - done = 0 - for fut in as_completed(futures): - res = fut.result() - out[res["url"]] = res - done += 1 - if done % 50 == 0: - print(f" ... {done}/{len(urls)}", file=sys.stderr, flush=True) - return out - - -def main() -> int: - ap = argparse.ArgumentParser(description=__doc__) - ap.add_argument("--run", action="store_true", help="execute the network sweep") - ap.add_argument("--dry-run", action="store_true", help="show scope only") - ap.add_argument("--structural-only", action="store_true") - ap.add_argument("--limit", type=int, default=0) - ap.add_argument("--workers", type=int, default=12) - ap.add_argument("--json", type=str, default="") - args = ap.parse_args() - - if not (args.run or args.dry_run or args.structural_only): - ap.error("choose one of --run / --dry-run / --structural-only") - - index = collect_urls() - urls = sorted(index) - print(f"citation-shaped URLs in 10_knowledge: {len(urls)}") - print(f"files citing them: {len({f for s in index.values() for f in s})}") - - structural = {u: r for u in urls if (r := structural_verdict(u))} - print(f"\nSTRUCTURAL detector (offline): {len(structural)} bad-shape DOIs") - for u, reason in sorted(structural.items()): - print(f" FABRICATED-SHAPE {u}\n {reason}") - for f in sorted(index[u]): - print(f" cited by: {f}") - - if args.dry_run or args.structural_only: - by_host: dict[str, int] = defaultdict(int) - for u in urls: - by_host[re.sub(r"^https?://([^/]+).*", r"\1", u)] += 1 - print("\nwould probe, by host:") - for h, n in sorted(by_host.items(), key=lambda kv: -kv[1]): - print(f" {n:5d} {h}") - return 0 - - target = urls[: args.limit] if args.limit else urls - controls = POSITIVE_CONTROLS + NEGATIVE_CONTROLS - print(f"\nprobing {len(target)} URLs + {len(controls)} controls " - f"({args.workers} workers)...", file=sys.stderr) - - results = run_probes(target + controls, workers=args.workers) - - # --- control gate: evaluate BEFORE reporting any finding --------------- - pos_fail = [u for u in POSITIVE_CONTROLS if results.get(u, {}).get("resolved") is not True] - neg_fail = [u for u in NEGATIVE_CONTROLS if results.get(u, {}).get("resolved") is not False] - probe_ok = not pos_fail and not neg_fail - - print("\n=== CONTROL GATE ===") - for u in POSITIVE_CONTROLS: - r = results.get(u, {}) - print(f" [+] {'PASS' if r.get('resolved') is True else 'FAIL'} {r.get('status')} {u}") - for u in NEGATIVE_CONTROLS: - r = results.get(u, {}) - print(f" [-] {'PASS' if r.get('resolved') is False else 'FAIL'} {r.get('status')} {u}") - - if not probe_ok: - print("\n*** BROKEN PROBE — results below are NOT a clean sweep. ***") - print("*** Do not record this run as evidence of anything. ***") - - unresolved = {u: results[u] for u in target if results.get(u, {}).get("resolved") is False} - inconclusive = {u: results[u] for u in target if results.get(u, {}).get("resolved") is None} - resolved_n = len(target) - len(unresolved) - len(inconclusive) - - print("\n=== RESULT ===") - print(f" probed: {len(target)}") - print(f" resolved: {resolved_n}") - print(f" UNRESOLVED: {len(unresolved)}") - print(f" inconclusive: {len(inconclusive)} (blocked / timeout — not evidence)") - - if unresolved: - print("\n--- UNRESOLVED (identifier does not exist) ---") - for u, r in sorted(unresolved.items()): - flag = " [ALSO BAD SHAPE]" if u in structural else "" - print(f"\n {r['status']} {u}{flag}") - for f in sorted(index[u]): - print(f" cited by: {f}") - - payload = { - "run_at": datetime.now(timezone.utc).isoformat(), - "probe_ok": probe_ok, - "control_failures": {"positive": pos_fail, "negative": neg_fail}, - "counts": { - "citation_urls_total": len(urls), - "probed": len(target), - "resolved": resolved_n, - "unresolved": len(unresolved), - "inconclusive": len(inconclusive), - "structural_bad_shape": len(structural), - }, - "structural": {u: {"reason": r, "files": sorted(index[u])} for u, r in structural.items()}, - "unresolved": {u: {**r, "files": sorted(index[u])} for u, r in unresolved.items()}, - "inconclusive": {u: {**r, "files": sorted(index[u])} for u, r in inconclusive.items()}, - } - if args.json: - Path(args.json).write_text(json.dumps(payload, indent=2), encoding="utf-8") - print(f"\nwrote {args.json}") - - return 0 if probe_ok else 2 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/contract-audit b/bin/contract-audit deleted file mode 100755 index a8f0955..0000000 --- a/bin/contract-audit +++ /dev/null @@ -1,1274 +0,0 @@ -#!/usr/bin/env python3 -"""Audit MainFrame's own conventions — by measuring state, not by reading prose. - -## Why this exists, and why it is built the way it is - -`10_knowledge/agents/prompts/2026-06-22__agents__note__mainframe-agent-architecture-audit-checklist.md` -is a sixty-item convention audit written on 2026-06-22. Every box in it is still -unchecked. Its item 1.5 reads: - - [ ] Safety rules that need enforcement have hooks, scripts, config, permissions, or tests. - -and its item 10.5: - - [ ] Enforcement need -> config/hook/script/test, not prose only. - -Seven weeks later, 107 captures citing papers that do not exist were found in -`10_knowledge/`, because the rule that should have stopped them existed only as -prose. **The checklist named the defect and was never run.** - -So the first design constraint is that this file is a program, not a document. -A checklist that requires a human to sit down and read sixty items is a checklist -that does not get read. - -## The second design constraint: an audit that reads documents would have passed - -This is the part that matters. Run the 2026-06-22 checklist by hand against the -repo as it stood on 2026-08-08 — the day before the fabrications were found — and -it very plausibly comes back clean: - - * `01_ingest/AGENTS.md` says raw captures route through ingest. tick - * `EPISTEMIC_STANCE.md` says LLM inferences are not truth. tick - * `agents/ingest-agent.md` names guardrails, tools, and handoff shape. tick - * `.context/workflows/audit-sweep.md` exists and is referenced. tick - -Meanwhile 449 raw captures had bypassed ingest entirely and 107 of the files -those contracts governed carried invented DOIs. **Every document said the right -thing. The state was wrong.** Document-reading audits measure whether you wrote -the rule down, which is not the question. - -Hence three layers, and only the last two are worth anything: - - Layer A DOCUMENT does the contract say the right thing? - Cheap, and the weakest evidence there is. Reported, but - never sufficient — every Layer A pass is labelled as - document-level so it cannot be mistaken for a control. - - Layer B STATE run the enforcing tool and report what it found. - This is what catches the thing nobody anticipated, - because it asks what is true rather than what is written. - - Layer C REPLAY a positive control. Known past failures are replayed - against the current checks. If a check does not fire on - a failure it is supposed to catch, this audit reports - BROKEN rather than clean. - -Layer C is the direct answer to "an audit would have missed the made-up sources." -It would have. So the audit is required to prove, every run, that it still -detects the failures we already know about. An audit that cannot fail its own -controls is not evidence — the same rule the citation sweep runs under, where a -probe with no negative control reports a captive portal as a clean bill of health. - -## What it deliberately does not do - -It does not grade judgement. Roughly half the checklist is unautomatable — -"AGENTS.md is concise and current", "subagents return distilled summaries" — and -silently dropping those items is how a checklist decays into a linter and stops -asking the hard questions. They are carried in the register as OPEN with a review -date, visible and explicitly not machine-checked. - -Usage: - bin/contract-audit # full audit, human-readable - bin/contract-audit --selftest # Layer C only: are the checks alive? - bin/contract-audit --layer state # only the checks that measure state - bin/contract-audit --register out.md # write the contract register - bin/contract-audit --json out.json - bin/contract-audit --fail-on-broken # CI: nonzero if a control fails -""" - -from __future__ import annotations - -import argparse -import json -import os -import re -import subprocess -import sys -from dataclasses import dataclass, field -from datetime import date, timedelta -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] - -# How long a judgement item stays fresh before it wants a human again. -REVIEW_INTERVAL_DAYS = 90 - -CHECKLIST = ( - ROOT / "10_knowledge/agents/prompts" - / "2026-06-22__agents__note__mainframe-agent-architecture-audit-checklist.md" -) - - -# ---------------------------------------------------------------- result types - - -@dataclass -class Finding: - item: str # checklist coordinate, e.g. "1.5" - layer: str # document | state | judgement - title: str - status: str # PASS | FAIL | OPEN | BROKEN | SKIP - detail: str = "" - evidence: list[str] = field(default_factory=list) - # Would this check have fired during the 2026-08-09 fabrication incident? - # "no" on a document-layer item is expected and is itself the finding. - caught_incident: str = "n/a" # yes | no | n/a - review_by: str = "" - - -RESULTS: list[Finding] = [] - - -def add(f: Finding) -> Finding: - RESULTS.append(f) - return f - - -def review_date() -> str: - return (date.today() + timedelta(days=REVIEW_INTERVAL_DAYS)).isoformat() - - -def read(path: Path) -> str: - try: - return path.read_text(encoding="utf-8", errors="replace") - except OSError: - return "" - - -def run(cmd: list[str], timeout: int = 120) -> tuple[int, str]: - """Run an enforcing tool. A missing tool is a FAIL, not a crash. - - Deliberately does NOT set `MAINFRAME_KNOWLEDGE_WRITE=1`. An earlier version - did, so that any tool the audit invoked could write freely. Nothing it - invokes needs that: every call here is `--help`, `--check`, or a read-only - validation pass. What the override actually bought was a latent bypass of - the write guard, living inside the tool whose job is to check that guards - are in place. An audit must not run with privileges the thing it audits - would not have. - """ - try: - p = subprocess.run( - cmd, cwd=ROOT, capture_output=True, text=True, timeout=timeout, - ) - return p.returncode, (p.stdout or "") + (p.stderr or "") - except FileNotFoundError: - return 127, f"tool not found: {cmd[0]}" - except subprocess.TimeoutExpired: - return 124, f"timed out after {timeout}s: {' '.join(cmd)}" - - -# ------------------------------------------------------- contract file discovery - - -# Files that carry normative language but are deliberately NOT contracts. -# Naming them is the point: an exclusion that is argued for is auditable, an -# exclusion that is an oversight is a blind spot wearing the same clothes. -NON_CONTRACT = { - "log.md": "append-only work log — narrates what happened, binds nothing", - "log_new.md": "same", - "STATE.md": "current focus; declared volatile by AGENTS.md", - "CHANGELOG.md": "history", - "README.md": "orientation for humans, not agent-binding", - "EVAL_METHODOLOGY.md": "method description; its rules live in workflows", -} - - -def contract_files() -> list[Path]: - """Every file that carries normative weight over agent behaviour. - - The first version of this function listed AGENTS.md, HARNESS.md, - EPISTEMIC_STANCE.md, DECISIONS.md, workflows and agents/ — the file types - that came to mind. Measured against the repo, that missed **291 normative - clauses in 51 files**, the largest block of them inside `.agents/skills/`, - which is where the ingest procedure actually lives. - - That is the deny-list mistake for the fourth time in this investigation, this - time committed by the tool built to catch it. Every previous instance had the - same fix and so does this one: **enumerate what a contract is, then take - everything that qualifies**, rather than listing the ones you remember. - - A contract here is: any file under a directory whose purpose is to govern - agent behaviour (`.agents/`, `.context/`, `agents/`), plus the named root - contracts, minus an explicit, argued exclusion list. - """ - out: list[Path] = [] - for p in ROOT.rglob("AGENTS.md"): - if "node_modules" in p.parts or "90_archive" in p.parts: - continue - out.append(p) - for name in ("HARNESS.md", "EPISTEMIC_STANCE.md", "DECISIONS.md"): - if (ROOT / name).exists(): - out.append(ROOT / name) - for base in (".context", ".agents", "agents"): - d = ROOT / base - if not d.exists(): - continue - for p in d.rglob("*.md"): - if "node_modules" in p.parts or p.name in NON_CONTRACT: - continue - out.append(p) - return list(dict.fromkeys(out)) - - -def coverage_gap() -> tuple[int, int, list[str]]: - """How much normative language does this audit NOT look at? - - Reported every run. An audit that does not state its own coverage invites - the reader to assume it is total, which is the failure mode that let a - checklist full of correct rules sit next to 449 files that ignored them. - """ - covered = {p.resolve() for p in contract_files()} - missed, total = [], 0 - seen: set[Path] = set() - for p in list(ROOT.glob("*.md")) + list((ROOT / ".agents").rglob("*.md")) \ - + list((ROOT / ".context").rglob("*.md")): - if not p.is_file() or p.resolve() in covered or p.resolve() in seen: - continue - if "90_archive" in p.parts or "node_modules" in p.parts: - continue - seen.add(p.resolve()) - n = sum(1 for ln in read(p).splitlines() if NORMATIVE.search(ln)) - if n: - total += n - why = NON_CONTRACT.get(p.name, "NOT EXCLUDED — genuine blind spot") - missed.append(f"{n:>3} clauses {p.relative_to(ROOT)} ({why})") - return len(missed), total, missed - - -NORMATIVE = re.compile( - r"\b(MUST(?! not)|MUST NOT|NEVER|ALWAYS|SHALL|REQUIRED|PROHIBITED" - r"|[Dd]o not|must not|may not|never|only|strictly)\b" -) - -# A clause is "bound to a mechanism" if it names something that can run. -MECHANISM = re.compile( - r"(bin/[a-z0-9._-]+" # a tool - r"|\.claude/settings" # config - r"|PreToolUse|PostToolUse|hook" - r"|pytest|test_[a-z_]+|selftest" - r"|--check\b|--apply\b|--dry-run\b|--fail-on" - r"|\.gitignore" - r")" -) - -# Tier declarations, once contracts start carrying them. -# -# Tolerant of the forms contracts actually use: inside a blockquote, and with -# the colon inside or outside the bold markers. The first version required -# `Tier**:` exactly and reported 0/121 against files that all declared a tier -# correctly — a checker that only recognises its author's preferred syntax -# reports the corpus as non-compliant and is worse than no checker, because the -# failure looks like a finding. -TIER_RE = re.compile(r"^[>\s]*(?:\*\*)?Tier:?(?:\*\*)?:?\s*(T[0-3])\b", re.M) - -# The strongest promises a contract can make. Deliberately wider than NORMATIVE: -# "immutable" and "read-only" are absolute claims even though they use no modal -# verb, and both appear on rules that were violated in the 2026-08-09 incident. -HARD_CLAUSE = re.compile( - r"\b(MUST NOT|MUST|NEVER|Never|never|PROHIBITED|prohibited|SHALL" - r"|may not|must not|strictly|immutable|read-only|only to|exclusive)\b" -) - - -# --------------------------------------------------------------- Layer A: docs - - -def layer_a_documents() -> None: - root_agents = ROOT / "AGENTS.md" - text = read(root_agents) - lines = text.splitlines() - - add(Finding( - "1.1", "document", "AGENTS.md is concise (<200 lines)", - "PASS" if len(lines) < 200 else "FAIL", - f"{len(lines)} lines", [str(root_agents.relative_to(ROOT))], - caught_incident="no", - )) - - routes = { - "HARNESS.md": "HARNESS.md" in text, - "workflows": ".context/workflows" in text, - "skills": ".agents/skills" in text, - "agents/": "agents/" in text, - "local AGENTS.md": "local `AGENTS.md`" in text or "local AGENTS.md" in text, - "epistemic stance": "EPISTEMIC_STANCE" in text, - "MindGraph": "MindGraph" in text, - } - missing = [k for k, v in routes.items() if not v] - add(Finding( - "1.2", "document", "AGENTS.md routes to every layer", - "PASS" if not missing else "FAIL", - "all seven routes present" if not missing else f"missing: {', '.join(missing)}", - caught_incident="no", - )) - - # 1.3 — durable rules only. Dated lines and status words are the smell. - volatile = [ - ln.strip() for ln in lines - if re.search(r"\b20\d\d-\d\d-\d\d\b", ln) - or re.search(r"\b(currently|as of|this week|next gate|in progress)\b", ln, re.I) - ] - add(Finding( - "1.3", "document", "AGENTS.md holds durable rules, not live status", - "PASS" if not volatile else "FAIL", - "no dated or status-bearing lines" if not volatile - else f"{len(volatile)} volatile line(s)", - volatile[:3], caught_incident="no", - )) - - # 3.1 — the four claim types. - ep = read(ROOT / "EPISTEMIC_STANCE.md") - types = [t for t in ("Observation", "Source-claim", "Inference", "Hypothesis") - if t.lower() not in ep.lower()] - add(Finding( - "3.1", "document", "EPISTEMIC_STANCE covers all four claim types", - "PASS" if not types else "FAIL", - "observation / source-claim / inference / hypothesis all present" - if not types else f"missing: {', '.join(types)}", - caught_incident="no", - )) - - # 3.2 — claim-bearing workflows point at the standard. - # - # The threshold is stated rather than tuned in silence. A first version - # matched any mention of "source" and reported 30 gaps, most of which were - # workflows that say "source" once while archiving a project. That is the - # deny-list mistake in a different costume: a detector that fires on - # everything is as uninformative as one that fires on nothing, and it looks - # more diligent. Density >= 4 is the line between "produces claims" and - # "mentions the word". - CLAIM_DENSITY = 4 - wf = sorted((ROOT / ".context/workflows").glob("*.md")) - claim_bearing, pointing = [], [] - for p in wf: - t = read(p) - density = len(re.findall(r"\b(claim|evidence|synthes)\w*", t, re.I)) - if density >= CLAIM_DENSITY: - claim_bearing.append(p) - if "epistemic-standard" in t or "EPISTEMIC_STANCE" in t: - pointing.append(p) - gap = [p.name for p in claim_bearing if p not in pointing] - add(Finding( - "3.2", "document", "Claim-bearing workflows cite the epistemic standard", - "PASS" if not gap else "FAIL", - f"{len(pointing)}/{len(claim_bearing)} workflows with claim density " - f">= {CLAIM_DENSITY} cite the standard", - gap, caught_incident="no", - )) - - # 4.x — workflow shape. - no_trigger, no_command, no_stop = [], [], [] - for p in wf: - t = read(p) - if not re.search(r"^#{1,3}\s*(when to use|trigger|use this|scope)", t, re.I | re.M) \ - and not re.search(r"^Use (this|the) workflow", t, re.M): - no_trigger.append(p.name) - if "```bash" not in t and "bin/" not in t: - no_command.append(p.name) - # Deliberately strict. A loose version matching a bare "fail" anywhere - # scored 19/37, because almost every document mentions failure in - # passing. What the reader needs is an answer to "this went wrong, now - # what", and that requires a named state or a conditional instruction, - # not the word. The strict count is 12/37, and it is the true one. - if not re.search( - r"(stop state|blocked state|## stop|if .{0,30}fails?" - r"|do not proceed|abort|halt|escalate)", t, re.I - ): - no_stop.append(p.name) - - add(Finding( - "4.1", "document", "Every workflow names a trigger", - "PASS" if not no_trigger else "FAIL", - f"{len(wf) - len(no_trigger)}/{len(wf)} have one", no_trigger[:8], - caught_incident="no", - )) - add(Finding( - "4.3", "document", "Every workflow names explicit commands", - "PASS" if not no_command else "FAIL", - f"{len(wf) - len(no_command)}/{len(wf)} name a command", no_command[:8], - caught_incident="no", - )) - add(Finding( - "4.4", "document", "Every workflow names a stop/blocked state", - "PASS" if not no_stop else "FAIL", - f"{len(wf) - len(no_stop)}/{len(wf)} name one", no_stop[:8], - caught_incident="no", - )) - - # 5.x — skills. - sk = ROOT / ".agents/skills" - dirs = [p for p in sk.iterdir() if p.is_dir()] if sk.exists() else [] - flat = [p for p in sk.glob("*.md")] if sk.exists() else [] - no_skillmd = [p.name for p in dirs if not (p / "SKILL.md").exists()] - add(Finding( - "5.1", "document", "Every skill directory has SKILL.md", - "PASS" if not no_skillmd else "FAIL", - f"{len(dirs) - len(no_skillmd)}/{len(dirs)} directories conform", no_skillmd, - caught_incident="no", - )) - - bad_fm = [] - for p in dirs: - s = p / "SKILL.md" - if not s.exists(): - continue - head = read(s)[:1200] - if not re.search(r"^name:\s*\S", head, re.M) or not re.search(r"^description:\s*\S", head, re.M): - bad_fm.append(p.name) - add(Finding( - "5.2", "document", "Every SKILL.md declares name and description", - "PASS" if not bad_fm else "FAIL", - f"{len(dirs) - len(bad_fm)}/{len(dirs)} valid", bad_fm, - caught_incident="no", - )) - - long_skills = [ - f"{p.name} ({len(read(p / 'SKILL.md').splitlines())} lines)" - for p in dirs - if (p / "SKILL.md").exists() - and len(read(p / "SKILL.md").splitlines()) > 300 - and not (p / "references").exists() - ] - add(Finding( - "5.4", "document", "Long skill content lives in references/", - "PASS" if not long_skills else "FAIL", - "no oversized SKILL.md without references/" if not long_skills - else f"{len(long_skills)} oversized", long_skills, - caught_incident="no", - )) - - legacy = [p.name for p in flat if p.suffix == ".md"] - add(Finding( - "5.7", "document", "Legacy flat skills are migrated or documented", - "OPEN" if legacy else "PASS", - f"{len(legacy)} flat skill file(s) remain alongside {len(dirs)} directories; " - "no note in .agents/skills/ marks them as intentionally legacy", - legacy, caught_incident="no", review_by=review_date(), - )) - - # 6.1 — subagent definitions. - incomplete = [] - for p in sorted((ROOT / "agents").glob("*.md")): - t = read(p) - want = { - "tools": bool(re.search(r"^tools:", t, re.M)), - "guardrails": "uardrail" in t or "afety" in t, - "handoff": bool(re.search(r"(handoff|outputs?|report)", t, re.I)), - } - if not all(want.values()): - incomplete.append(f"{p.name}: missing {', '.join(k for k, v in want.items() if not v)}") - add(Finding( - "6.1", "document", "Subagents define tools, guardrails, and handoff", - "PASS" if not incomplete else "FAIL", - f"{len(list((ROOT / 'agents').glob('*.md'))) - len(incomplete)} complete", - incomplete, caught_incident="no", - )) - - # 8.5 — hooks bounded. - settings = ROOT / ".claude/settings.json" - unbounded = [] - try: - cfg = json.loads(read(settings) or "{}") - for ev, arr in (cfg.get("hooks") or {}).items(): - for entry in arr: - for hk in entry.get("hooks", []): - if "timeout" not in hk: - unbounded.append(f"{ev}:{entry.get('matcher', '*')}") - except json.JSONDecodeError as exc: - unbounded.append(f"unparseable settings.json: {exc}") - add(Finding( - "8.5", "document", "Every hook is timeout-bounded", - "PASS" if not unbounded else "FAIL", - "all hooks declare a timeout" if not unbounded - else f"{len(unbounded)} hook(s) with no timeout — a hang blocks the session", - unbounded[:10], caught_incident="no", - )) - - # 8.3 — secrets excluded. - gi = read(ROOT / ".gitignore") - patterns = [".env", "*.key", "credentials", "secret"] - absent = [p for p in patterns if p not in gi] - add(Finding( - "8.3", "document", "Secrets are excluded from version control", - "PASS" if len(absent) <= 1 else "FAIL", - f".gitignore covers {len(patterns) - len(absent)}/{len(patterns)} secret patterns", - absent, caught_incident="no", - )) - - -# ------------------------------------------------------- Layer A': clause tiers - - -def layer_a_clauses() -> None: - """Item 1.5 / 10.5 — the two the incident turned on. - - Counted honestly. A normative clause with no mechanism is not automatically - a defect: "do not bury dissent to preserve narrative coherence" cannot have a - script and should not pretend to. The defect is that **nothing distinguishes - an advisory rule from a load-bearing one**, so a rule that silently lost its - enforcement looks exactly like a rule that never needed any. - """ - files = contract_files() - total = bound = 0 - tiered = 0 - worst: list[tuple[int, str]] = [] - - for p in files: - text = read(p) - if TIER_RE.search(text): - tiered += 1 - n = b = 0 - for ln in text.splitlines(): - if NORMATIVE.search(ln): - n += 1 - if MECHANISM.search(ln): - b += 1 - total += n - bound += b - if n - b: - worst.append((n - b, str(p.relative_to(ROOT)))) - - pct = 100 * bound / max(total, 1) - add(Finding( - "1.5", "document", "Normative clauses name an enforcing mechanism", - "FAIL", - f"{total} normative clauses across {len(files)} contract files; " - f"{bound} ({pct:.0f}%) name a runnable mechanism, {total - bound} are prose only. " - "Not all of these need one — but nothing in the corpus says which do.", - [f"{n} unbound {path}" for n, path in sorted(worst, reverse=True)[:8]], - caught_incident="no", - )) - - # 1.6 — hard mandatory language with nothing behind it. - # - # This is the closest machine-checkable proxy for the thing that actually - # went wrong. Once tiers exist, the audit can only verify that a tier was - # *declared*, not that it was declared *honestly* — a load-bearing rule - # marked T0 goes green, which is the same worthless self-assertion as a - # `retrieved_at` field the capture wrote about itself. - # - # Language, though, is not self-serving in the same way. A file that says - # NEVER and MUST NOT is making a strong promise to whoever reads it. If it - # makes several such promises and names no mechanism at all, the promise is - # louder than the enforcement, and a reader is entitled to be misled. - # Granularity is the whole check. A first version asked whether the FILE - # named any mechanism, and passed clean. Tested against the rule that - # actually broke — `01_ingest/AGENTS.md` rule 5, "writes only to files inside - # 01_ingest/", the rule the 21 direct writes walked straight through — it - # did not fire, because that file mentions `bin/prep-ingest` several lines - # away on an unrelated clause. One tool named anywhere laundered every - # unenforced promise in the file. - # - # So the unit is the clause, not the file. A promise is bound only if the - # sentence making it says what enforces it. - hard_unbound: list[str] = [] - hard_total = 0 - for p in contract_files(): - for ln in read(p).splitlines(): - s = ln.strip() - if not HARD_CLAUSE.search(s): - continue - # numbered/bulleted rules are the normative unit; prose is context - if not re.match(r"^(\d+\.|[-*]|\|)", s): - continue - hard_total += 1 - if not MECHANISM.search(s): - hard_unbound.append(f"{p.relative_to(ROOT)} :: {s[:88]}") - pct = 100 * (hard_total - len(hard_unbound)) / max(hard_total, 1) - add(Finding( - "1.6", "document", "Each hard promise names its own enforcement", - "FAIL" if hard_unbound else "PASS", - f"{hard_total} clauses use MUST / NEVER / PROHIBITED / immutable / read-only " - f"language; {len(hard_unbound)} of them name no mechanism in the clause itself " - f"({pct:.0f}% bound). These are the strongest promises the corpus makes and the " - "least is behind them.", - hard_unbound[:10], caught_incident="yes", - )) - - # 1.7 — every enforced rule names an honest-failure path. - # - # The lesson the incident actually taught, and the one most likely to be - # dropped from a contract retrofit because it looks like documentation of an - # exception. It is not. A rule that demands an outcome an agent cannot - # legitimately produce does not get obeyed and does not get refused: it gets - # *simulated*. The source quota did not encourage fabrication, it made - # honesty unrepresentable, because no field meant "I looked and found - # nothing". 107 captures followed. - # - # So a rule at T1 or above with a Check and no Escape is not half-finished. - # It is the dangerous configuration: enforced, unfollowable, and therefore - # productive of compliant-looking output. - # - # See 10_knowledge/agents/2026-08-10__agents__note__every-rule-needs-an-honest-failure-path.md - RULE_BLOCK = re.compile( - r"(?:^[>\s]*\*\*Binds:\*\*.*?)(?=^\s*$|^##)", re.M | re.S - ) - tiered_rules = 0 - missing_escape: list[str] = [] - for p in contract_files(): - text = read(p) - if not TIER_RE.search(text): - continue - for block in RULE_BLOCK.findall(text): - m = TIER_RE.search(block) - if not m or m.group(1) == "T0": - continue # T0 is advisory; no escape required, none promised - tiered_rules += 1 - if not re.search(r"^[>\s]*\*\*Escape:\*\*\s*\S", block, re.M): - first = block.strip().splitlines()[0][:70] - missing_escape.append(f"{p.relative_to(ROOT)} :: {first}") - add(Finding( - "1.7", "document", "Every enforced rule names an honest-failure path", - "FAIL" if missing_escape else "PASS", - f"{tiered_rules} rule(s) at T1 or above; {len(missing_escape)} name no Escape. " - "A rule that is enforced but cannot be honestly failed produces " - "compliant-looking output, which is how a source quota produced 107 " - "invented papers." - if tiered_rules else - "no rules declare a tier yet, so this cannot be checked. That is the " - "finding, not a pass.", - missing_escape[:8], caught_incident="yes", - )) - - n_files, n_clauses, detail = coverage_gap() - unexplained = [d for d in detail if "genuine blind spot" in d] - add(Finding( - "0.0", "document", "This audit states its own coverage", - "PASS" if not unexplained else "FAIL", - f"{n_clauses} normative clauses in {n_files} file(s) sit outside the contract " - f"scan. {len(unexplained)} of those file(s) have no argued exclusion — " - "an unexplained gap is indistinguishable from an oversight." - if unexplained else - f"{n_clauses} clauses in {n_files} file(s) are outside scope, every one of them " - "under a named and argued exclusion.", - (unexplained or detail)[:8], caught_incident="n/a", - )) - - add(Finding( - "10.5", "document", "Contracts declare an enforcement tier", - "FAIL", - f"{tiered}/{len(files)} contract files declare a tier (T0 advisory / T1 detected / " - "T2 blocked / T3 reconciled). Without a declared tier, an unenforced mandatory " - "rule is indistinguishable from an advisory one — which is exactly how " - "'raw captures route through ingest' stayed true on paper while 449 bypassed it.", - caught_incident="no", - )) - - -# -------------------------------------------------------------- Layer B: state - - -def layer_b_state() -> None: - """Run the enforcing tools. This is the layer with evidentiary value.""" - - # 3.4 / ingest boundary — does every raw capture have an ingest record? - code, out = run([str(ROOT / "bin/knowledge-reconcile")]) - m = re.search(r"VIOLATIONS: (\d+)", out) - n = int(m.group(1)) if m else -1 - add(Finding( - "3.4", "state", "Every raw capture in 10_knowledge has an ingest record", - "PASS" if n == 0 else ("FAIL" if n > 0 else "BROKEN"), - f"{n} `type: raw` capture(s) with no ingest record" - if n >= 0 else f"bin/knowledge-reconcile did not report a count (exit {code})", - [ln for ln in out.splitlines() if "by month" in ln or "by domain" in ln], - caught_incident="yes", - )) - - # 3.5 — is `stable` actually gated, AND does the existing population comply? - # - # Two questions, and the second is the one that bites. A gate that only - # guards new writes leaves whatever was already inside untouched — the same - # split the write guard has, where layer 2 stops new raws and layer 3 - # reconciles the 449 that were already there. Reporting only "a gate exists" - # would be a document-layer answer wearing a state-layer badge. - validator = read(ROOT / "bin/capture-validate") - gated = "VERIFIED_STATUSES" in validator and "R6" in validator - - violations = 0 - files = 0 - for p in (ROOT / "10_knowledge").rglob("*.md"): - head = read(p)[:1200] - m = re.search(r'^status:\s*"?([a-z-]+)', head, re.M) - if not m or m.group(1) not in ("stable", "audited"): - continue - files += 1 - tagline = re.search(r"^tags:.*$", head, re.M) - contradicts = bool(tagline and "needs-audit" in tagline.group(0)) - unevidenced = not re.search(r"^(audit_receipt|verified_by):\s*\S", head, re.M) - if contradicts or unevidenced: - violations += 1 - - add(Finding( - "3.5", "state", "`status: stable` is gated, and the existing population complies", - "PASS" if (gated and violations == 0) else "FAIL", - (f"gate: {'R6 in bin/capture-validate' if gated else 'NONE'}. " - f"{violations}/{files} file(s) claiming stable or audited fail it — no " - "verification evidence, or a needs-audit tag contradicting the status. " - "New writes are blocked; the standing population is the backlog."), - caught_incident="no", - )) - - # 1.5-state — do the provenance controls actually still run? - for tool, args, label, item in ( - ("bin/capture-validate", ["--help"], "provenance validator", "1.5a"), - ("bin/citation-resolvability-sweep", ["--help"], "citation sweep", "1.5b"), - ("bin/knowledge-write-guard", [], "write guard (hook)", "1.5c"), - ): - path = ROOT / tool - if not path.exists(): - add(Finding(item, "state", f"{label} is present and runnable", "FAIL", - f"{tool} does not exist", caught_incident="yes")) - continue - if not os.access(path, os.X_OK): - add(Finding(item, "state", f"{label} is present and runnable", "FAIL", - f"{tool} is not executable", caught_incident="yes")) - continue - add(Finding(item, "state", f"{label} is present and runnable", "PASS", - tool, caught_incident="yes")) - - # 8.5-state — is the write guard actually wired into settings, not just present? - cfg = read(ROOT / ".claude/settings.json") - wired = "knowledge-write-guard" in cfg - add(Finding( - "1.5d", "state", "The write guard is wired into settings, not merely present", - "PASS" if wired else "FAIL", - "PreToolUse hook on Write|Edit" if wired else - "bin/knowledge-write-guard exists but no hook invokes it — a control that is " - "installed but not wired is prose with a shebang", - caught_incident="yes", - )) - - # 7.2 / 7.3 — project state truth. - code, out = run([str(ROOT / "bin/sync-project-index"), "--check"], timeout=180) - add(Finding( - "7.2", "state", "Project states match observed activity", - "PASS" if code == 0 else "FAIL", - "bin/sync-project-index --check green" if code == 0 - else f"exit {code}: " + " / ".join( - ln.strip() for ln in out.splitlines() if ln.strip())[:400], - caught_incident="no", - )) - - # 9.1 — telemetry stays out of the knowledge index. - gi = read(ROOT / ".gitignore") - ignored = "20_live/workflow-metrics" in gi or "20_live/" in gi - add(Finding( - "9.1", "state", "Volatile telemetry is excluded from version control", - "PASS" if ignored else "FAIL", - "20_live telemetry is gitignored" if ignored else "20_live is tracked", - caught_incident="no", - )) - - # 10.1 — "repeated agent mistake -> fix the closest contract or tool." - # - # This item sat in the judgement pile because nothing recorded agent - # mistakes, which made it unfalsifiable in both directions: no evidence of - # repeated mistakes, and no evidence of none. `bin/papercut` supplies the - # observational axis the audit otherwise lacks entirely (limit L6), so the - # item becomes a state check — with the crucial caveat that an empty log is - # NOT a pass. Silence here means nobody is recording, not that nothing broke. - store = ROOT / "20_live/papercuts" - entries = 0 - if store.exists(): - for p in store.glob("*.jsonl"): - entries += sum(1 for ln in read(p).splitlines() if ln.strip()) - receipt = store / "last-harvest.json" - if entries == 0: - add(Finding( - "10.1", "state", "Repeated friction is recorded and harvested into fixes", - "FAIL", - "no papercuts recorded. This is an absence of evidence, not evidence of " - "absence — friction that nobody logs looks exactly like friction that did " - "not happen, which is the condition that let a flag-named HTML file sit in " - "the repo root for six days.", - caught_incident="no", - )) - else: - try: - r = json.loads(read(receipt) or "{}") - except json.JSONDecodeError: - r = {} - cands = r.get("graduation_candidates", 0) - harvested = r.get("entries_seen", 0) - stale_by = entries - harvested - add(Finding( - "10.1", "state", "Repeated friction is recorded and harvested into fixes", - "PASS" if (receipt.exists() and stale_by <= 0 and cands == 0) else "FAIL", - f"{entries} papercut(s) recorded; last harvest saw {harvested}; " - f"{cands} graduation candidate(s) awaiting a fix" - + ("" if stale_by <= 0 else f"; {stale_by} unharvested since"), - caught_incident="no", - )) - - # 8.7 — automations declare what they do when they cannot run. - # - # The third exit from the honest-failure-path lesson, applied to `bin/` - # rather than to prose. A tool that catches every exception and exits 0 has - # made a fail-open choice; a tool that exits 1 for both "found problems" and - # "crashed" has made its result unreadable. Neither is wrong in itself. Both - # are wrong undeclared, because the caller cannot tell which they got. - # - # `bin/capture-validate` shipped with exactly this defect and it took a - # positive control to notice. - tools, undeclared, single_code = [], [], [] - bindir = ROOT / "bin" - for p in sorted(bindir.iterdir()): - if not p.is_file() or not os.access(p, os.X_OK) or p.suffix == ".pyc": - continue - text = read(p) - if not text.startswith("#!/usr/bin/env python"): - continue - tools.append(p.name) - # Catch `return 0 if ok else 2` as well as plain literals. The first - # version missed the conditional form and reported a tool that does - # signal failure as one that never can. - codes = set(re.findall(r"(?:sys\.)?exit\(\s*(\d)\s*\)", text)) - codes |= set(re.findall(r"return\s+(\d)\b", text)) - codes |= set(re.findall(r"else\s+(\d)\b", text)) - swallows = bool(re.search(r"except\s+(Exception|BaseException)?\s*:", text)) - declared = bool(re.search( - r"fail[- ]open|fail[- ]closed|fails open|fails closed" - r"|did not run|exit contract|Exit codes", text, re.I)) - if swallows and not declared: - undeclared.append(p.name) - elif len(codes - {"0"}) < 1 and swallows: - single_code.append(p.name) - - add(Finding( - "8.7", "state", "Automations declare what they do when they cannot run", - "PASS" if not undeclared else "FAIL", - f"{len(tools)} python tools in bin/; {len(undeclared)} catch every exception " - "without stating fail-open or fail-closed anywhere in the file. A caller " - "cannot tell 'ran and found nothing' from 'did not run', and the quieter " - "reading always looks like the better news.", - undeclared[:12], caught_incident="yes", - )) - - # 2.6 — does HARNESS.md embed numbers that should live in project outputs? - harness = read(ROOT / "HARNESS.md") - stale = [ - ln.strip() for ln in harness.splitlines() - if re.search(r"\b\d+/\d+\b", ln) and re.search(r"\b20\d\d-\d\d-\d\d\b", ln) - ] - add(Finding( - "2.6", "document", "HARNESS.md cites project outputs rather than embedding scores", - "PASS" if len(stale) <= 2 else "FAIL", - f"{len(stale)} line(s) embed a dated score directly in the contract", - stale[:4], caught_incident="no", - )) - - -# ------------------------------------------------------------- Layer C: replay - - -@dataclass -class Incident: - key: str - name: str - when: str - # The state signature the incident left behind, and the check that must see it. - probe: str - expect: str - - -# Known failures. Each must still be detectable. This is the positive control: -# an audit that cannot detect the failures we already suffered is not evidence -# that there are none. -INCIDENTS = [ - Incident( - "fabricated-citation", "107 captures citing papers that do not exist", - "2026-08-09", - "capture-validate sees a fabricated identifier in a synthetic capture", - "R2", - ), - Incident( - "ingest-bypass", "449 raw captures written past the ingest boundary", - "2026-08-10", - "knowledge-reconcile reports unlogged raw captures when they exist", - "VIOLATIONS", - ), - Incident( - "direct-write", "21 raw captures hand-written into 10_knowledge", - "2026-08-09", - "knowledge-write-guard denies a new type: raw under 10_knowledge/", - "deny", - ), - Incident( - "source-quota", "a source-count quota that made fabrication the compliant output", - "2026-08-09", - "research-lane-loop contains no live numeric source quota", - "absent", - ), -] - - -def layer_c_replay() -> bool: - """Replay known failures against the current checks. Returns True if all alive.""" - all_alive = True - - # 1. capture-validate must still flag a fabricated identifier. - # - # The control lives inside the repo on purpose. The first version wrote to - # $TMPDIR and the control came back BROKEN — not because the detector had - # failed, but because capture-validate crashed on `relative_to` for any path - # outside ROOT and exited 1, which was also its "found errors" code. So a - # crash and a finding were the same signal. That bug was invisible until - # something tried to use the validator as a control, which is the argument - # for having controls at all. - scratch = ROOT / ".contract-audit-control.md" - scratch.write_text( - "---\ntitle: \"Control — this citation is invented\"\n" - "type: \"raw\"\ndomain: \"knowledge-systems\"\n" - "authors: \"Synthetic Control Task Force\"\nyear: 2025\n" - "url: \"https://doi.org/10.1145/this-identifier-is-a-control-2025\"\n---\n\n" - "# Control\n\nPositive control for bin/contract-audit --selftest.\n", - encoding="utf-8", - ) - code, out = run([str(ROOT / "bin/capture-validate"), str(scratch)]) - # exit 1 = ran and found errors. exit 2 = did not run. Only 1 proves it works. - alive = code == 1 and ("R2" in out or "fabricat" in out.lower()) - all_alive &= alive - add(Finding( - "C1", "replay", "Control: fabricated identifier is still detected", - "PASS" if alive else "BROKEN", - "capture-validate flagged a known-bad synthetic DOI (exit 1)" if alive else - f"capture-validate did NOT flag a synthetic fabricated DOI (exit {code}). The " - "detector that found the 107 is dead or changed shape — treat every clean sweep " - "since as void.", - [ln.strip() for ln in out.splitlines() if "R2" in ln][:2], - caught_incident="yes", - )) - scratch.unlink(missing_ok=True) - - # 1b. NEGATIVE control. A detector that flags everything is as useless as one - # that flags nothing, and it looks far more diligent. This is the check the - # 2026-08-09 sweep would have failed without in-batch negative controls. - clean = ROOT / ".contract-audit-control-clean.md" - clean.write_text( - "---\ntitle: \"Attention Is All You Need\"\n" - "type: \"raw\"\ndomain: \"knowledge-systems\"\n" - "authors: \"Vaswani, A.; Shazeer, N.; Parmar, N.\"\nyear: 2017\n" - "url: \"https://arxiv.org/abs/1706.03762\"\n" - "fetch_method: \"control — not a real retrieval\"\n" - "retrieval_receipt: \"control\"\n---\n\n# Negative control\n", - encoding="utf-8", - ) - code_n, out_n = run([str(ROOT / "bin/capture-validate"), str(clean)]) - quiet = "R2" not in out_n - all_alive &= quiet - add(Finding( - "C1b", "replay", "Negative control: a real identifier is NOT flagged", - "PASS" if quiet else "BROKEN", - "capture-validate left a genuine arXiv ID alone" if quiet else - "capture-validate flagged a REAL identifier (arXiv 1706.03762) as fabricated. " - "A detector with a false-positive rate this visible would quarantine real " - "literature — the exact mistake the 2026-08-09 triage had to undo by hand.", - [ln.strip() for ln in out_n.splitlines() if "R2" in ln][:2], - caught_incident="n/a", - )) - clean.unlink(missing_ok=True) - - # 2. reconcile must be able to say a number at all. - code, out = run([str(ROOT / "bin/knowledge-reconcile")]) - alive = bool(re.search(r"VIOLATIONS: \d+", out)) - all_alive &= alive - add(Finding( - "C2", "replay", "Control: reconciliation still reports a count", - "PASS" if alive else "BROKEN", - "knowledge-reconcile produced a violation count" if alive else - "knowledge-reconcile produced no count. A reconciler that cannot report is " - "indistinguishable from a clean repo — the captive-portal failure mode.", - caught_incident="yes", - )) - - # 3. the write guard must still deny. - payload = json.dumps({ - "tool_name": "Write", - "tool_input": { - "file_path": str(ROOT / "10_knowledge/knowledge-systems/__control__.md"), - "content": "---\ntype: \"raw\"\ntitle: \"control\"\n---\n\nbody\n", - }, - }) - try: - env = {k: v for k, v in os.environ.items() if k != "MAINFRAME_KNOWLEDGE_WRITE"} - p = subprocess.run( - [str(ROOT / "bin/knowledge-write-guard")], input=payload, - capture_output=True, text=True, timeout=10, cwd=ROOT, env=env, - ) - alive = '"deny"' in p.stdout - detail_out = p.stdout[:200] - except (OSError, subprocess.SubprocessError) as exc: - alive, detail_out = False, str(exc) - all_alive &= alive - add(Finding( - "C3", "replay", "Control: write guard still denies a direct raw write", - "PASS" if alive else "BROKEN", - "guard returned permissionDecision: deny" if alive else - "guard did NOT deny a new type: raw under 10_knowledge/. It fails open by design, " - "so a silent breakage looks exactly like a working guard. That is why this control " - "exists. " + detail_out, - caught_incident="yes", - )) - - # 4. the quota must stay dead. - loop = read(ROOT / "bin/research-lane-loop") - live_quota = [ - ln.strip() for ln in loop.splitlines() - if re.search(r"(need|fill to|at least|minimum).{0,20}\b\d+\b.{0,20}(raw|source)", ln, re.I) - and not ln.strip().startswith("#") - ] - alive = not live_quota - all_alive &= alive - add(Finding( - "C4", "replay", "Control: no live source quota has returned", - "PASS" if alive else "BROKEN", - "no numeric source quota outside comments" if alive else - f"{len(live_quota)} line(s) reintroduce a source count — the root cause of the " - "2026-08-09 incident", - live_quota[:3], caught_incident="yes", - )) - - # 5. The clause-level enforcement check must still flag the specific rule - # that failed. `01_ingest/AGENTS.md` rule 5 — "writes only to files inside - # 01_ingest/" — is the rule the 21 direct writes of 2026-08-09 walked - # through. A file-level version of check 1.6 passed this file clean, - # because it mentions `bin/prep-ingest` a few lines away on an unrelated - # clause. That near-miss is why the control is pinned to a named clause - # and not to a percentage: a summary statistic can improve while the one - # rule you care about quietly stops being covered. - ingest_rules = read(ROOT / "01_ingest/AGENTS.md") - target = [ - ln.strip() for ln in ingest_rules.splitlines() - if "Read-only access" in ln or "immutable" in ln - ] - flagged = [s for s in target if HARD_CLAUSE.search(s) and not MECHANISM.search(s)] - alive = bool(flagged) - all_alive &= alive - add(Finding( - "C5", "replay", "Control: the clause that failed on 2026-08-09 is still flagged", - "PASS" if alive else "BROKEN", - f"{len(flagged)} of {len(target)} target clause(s) still detected as unenforced" - if alive else - "the enforcement check no longer flags 01_ingest/AGENTS.md rules 3 and 5. Either " - "they gained real enforcement (verify it, then update this control) or check 1.6 " - "has been loosened and its pass rate no longer means anything.", - flagged[:2], caught_incident="yes", - )) - - return all_alive - - -# ---------------------------------------------------------- Layer D: judgement - - -# Items that cannot be machine-checked and must not be silently dropped. -# Dropping them is how a checklist decays into a linter. -JUDGEMENT = [ - ("1.1b", "AGENTS.md is current — the rules still describe how work is actually done"), - ("1.4", "Local AGENTS files add scoped constraints rather than duplicate root rules"), - ("2.1", "HARNESS.md separates patterns, program status, and operating policy"), - ("2.2", "Task categories are explicit and still match how work is routed"), - ("2.3", "Local/cloud delegation boundaries are current"), - ("2.4", "Verification is external to the executing agent in practice, not just on paper"), - ("2.5", "MindGraph is described and used as nomination only, never verification"), - ("3.3", "High-stakes claims have evidence minimums that are actually applied"), - ("4.2", "Workflow steps are numbered, imperative, and testable"), - ("4.5", "Workflows do not duplicate skill bodies or project status"), - ("5.3", "Skill descriptions front-load task keywords and negative boundaries"), - ("5.5", "Skill scripts are used only for deterministic helper work"), - ("5.6", "Each skill has at least one happy path and one edge case"), - ("6.2", "Repeatable procedures are delegated to skills, not restated in agents/"), - ("6.3", "Subagents return distilled summaries, not raw noise"), - ("6.4", "Write-heavy parallel subagents have conflict controls"), - ("7.4", "Plans include dual MindGraph query passes"), - ("7.5", "Workbench truth is inspected before readiness claims"), - ("7.6", "Durable lessons are extracted to 10_knowledge/"), - ("8.1", "Active Codex config is known and reviewed"), - ("8.2", "Sandbox/approval settings match trust level"), - ("8.4", "MCP servers are scoped, documented, and consent-aware"), - ("8.6", "Tool allowlists/denylists exist for powerful external actions"), - ("9.2", "Receipts capture verifier result, scope diff, and completion truth"), - ("9.3", "Dashboards do not claim correctness from activity"), - ("9.4", "Eval outputs are dated and separate from durable pattern notes"), - # 10.1 moved to the state layer on 2026-08-10 — bin/papercut records the - # friction that makes it checkable. It is the only item to have graduated - # out of this list, and the route it took is the route the others need: - # not a better checklist, an instrument that produces evidence. - ("10.2", "Architecture tradeoff -> DECISIONS.md"), - ("10.3", "Durable lesson -> 10_knowledge/"), - ("10.4", "Active status -> project README/log or 20_live"), -] - - -def layer_d_judgement() -> None: - rb = review_date() - for item, title in JUDGEMENT: - add(Finding(item, "judgement", title, "OPEN", - "not machine-checkable — requires a human read", - caught_incident="no", review_by=rb)) - - -# ------------------------------------------------------------------- reporting - - -SYMBOL = {"PASS": "PASS", "FAIL": "FAIL", "OPEN": "OPEN", "BROKEN": "BROKEN", "SKIP": "skip"} - - -def report(controls_alive: bool) -> None: - by_layer: dict[str, list[Finding]] = {} - for f in RESULTS: - by_layer.setdefault(f.layer, []).append(f) - - order = ["replay", "state", "document", "judgement"] - names = { - "replay": "LAYER C — CONTROLS (do the checks still fire on known failures?)", - "state": "LAYER B — STATE (what is actually true in the repo)", - "document": "LAYER A — DOCUMENT (what the contracts say; weakest evidence)", - "judgement": "LAYER D — JUDGEMENT (carried open, not machine-checkable)", - } - - for layer in order: - items = by_layer.get(layer, []) - if not items: - continue - print(f"\n{names[layer]}") - print("-" * 78) - if layer == "judgement": - print(f" {len(items)} items OPEN, review by {items[0].review_by}") - for f in items: - print(f" {f.item:<6} {f.title}") - continue - for f in sorted(items, key=lambda x: (x.status != "BROKEN", x.status != "FAIL", x.item)): - print(f" {SYMBOL[f.status]:<6} {f.item:<6} {f.title}") - if f.detail: - for line in _wrap(f.detail, 68): - print(f" {line}") - for ev in f.evidence[:6]: - print(f" - {ev[:100]}") - - counts = {s: sum(1 for f in RESULTS if f.status == s) for s in ("PASS", "FAIL", "OPEN", "BROKEN")} - print("\n" + "=" * 78) - print(f" PASS {counts['PASS']} FAIL {counts['FAIL']} " - f"OPEN {counts['OPEN']} BROKEN {counts['BROKEN']}") - - if not controls_alive: - print("\n *** CONTROLS FAILED — THIS AUDIT IS NOT EVIDENCE ***") - print(" A check that no longer fires on a failure it caught before cannot") - print(" support a claim that the failure is absent. Fix the control first.") - else: - n = sum(1 for f in RESULTS if f.layer == "replay") - print(f"\n Controls alive: all {n} still fire, including the negative control.") - print(" This bounds rot in the checks. It says nothing about failures we") - print(" have not had yet — Layer C is backward-looking by construction.") - - doc_pass = sum(1 for f in RESULTS if f.layer == "document" and f.status == "PASS") - print(f"\n Note: {doc_pass} document-layer PASSes mean the contracts say the right") - print(" thing. On 2026-08-08 they also said the right thing. Weight Layer B.") - - -def _wrap(text: str, width: int) -> list[str]: - words, lines, cur = text.split(), [], "" - for w in words: - if len(cur) + len(w) + 1 > width: - lines.append(cur) - cur = w - else: - cur = f"{cur} {w}".strip() - if cur: - lines.append(cur) - return lines - - -def write_register(path: Path) -> None: - lines = [ - "---", - 'title: "MainFrame contract register"', - 'domain: "agents"', - 'type: "index"', - f'created: "{date.today().isoformat()}"', - f'as_of: "{date.today().isoformat()}"', - 'source: "Generated by bin/contract-audit — do not hand-edit"', - "---", - "", - "# MainFrame contract register", - "", - "Generated by `bin/contract-audit`. Every row is one item from the", - "2026-06-22 architecture audit checklist, with the layer that checks it and", - "what that check found.", - "", - "**Read the layer column before the status column.** A document-layer PASS", - "means a file says the right thing; it is not evidence the rule holds. Every", - "document-layer item passed on 2026-08-08, the day before 107 fabricated", - "citations were found in the corpus those documents govern.", - "", - "| Item | Layer | Rule | Status | Would it have caught 2026-08-09? | Review by |", - "|------|-------|------|--------|----------------------------------|-----------|", - ] - for f in sorted(RESULTS, key=lambda x: [int(n) if n.isdigit() else 99 - for n in re.findall(r"\d+", x.item)] or [99]): - caught = {"yes": "yes", "no": "**no**", "n/a": "—"}[f.caught_incident] - lines.append( - f"| {f.item} | {f.layer} | {f.title} | {f.status} | {caught} | {f.review_by or '—'} |" - ) - lines += [ - "", - "## Judgement items", - "", - f"{sum(1 for f in RESULTS if f.layer == 'judgement')} items cannot be machine-checked.", - "They are carried OPEN rather than dropped, because dropping them is how a", - "checklist quietly becomes a linter and stops asking the questions that need a", - f"person. Review interval: {REVIEW_INTERVAL_DAYS} days.", - "", - ] - path.write_text("\n".join(lines) + "\n", encoding="utf-8") - - -def main() -> int: - ap = argparse.ArgumentParser( - description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter - ) - ap.add_argument("--selftest", action="store_true", help="Layer C controls only") - ap.add_argument("--layer", choices=["document", "state", "judgement", "replay"]) - ap.add_argument("--register", type=str, default="") - ap.add_argument("--json", type=str, default="") - ap.add_argument("--fail-on-broken", action="store_true") - args = ap.parse_args() - - if not CHECKLIST.exists(): - print(f"WARNING: source checklist not found at {CHECKLIST.relative_to(ROOT)}", - file=sys.stderr) - - controls_alive = True - if args.selftest: - controls_alive = layer_c_replay() - else: - controls_alive = layer_c_replay() - if args.layer in (None, "state"): - layer_b_state() - if args.layer in (None, "document"): - layer_a_documents() - layer_a_clauses() - if args.layer in (None, "judgement"): - layer_d_judgement() - - report(controls_alive) - - if args.register: - write_register(Path(args.register)) - print(f"\nwrote register: {args.register}") - - if args.json: - Path(args.json).write_text(json.dumps({ - "as_of": date.today().isoformat(), - "controls_alive": controls_alive, - "findings": [vars(f) for f in RESULTS], - }, indent=2), encoding="utf-8") - print(f"wrote json: {args.json}") - - if args.fail_on_broken and not controls_alive: - return 2 - return 0 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/contract-lint b/bin/contract-lint new file mode 100755 index 0000000..c31c294 --- /dev/null +++ b/bin/contract-lint @@ -0,0 +1,867 @@ +#!/usr/bin/env python3 +"""Check one contract file against the Binds/Tier/Check/Escape format, before it lands. + +## The gap this fills + +`bin/contract-audit` measures the whole corpus and reports adoption. That is the +right instrument for *how bad is it*, and the wrong one for *is the file I am +about to commit correct*. Measured 2026-08-23: + + 10.5 contracts declare a tier ................ 2 / 135 files + 1.6 hard promises bound at clause level ..... 138 clauses, 124 unbound (10%) + 1.5 normative clauses name a mechanism ...... 680 clauses, 39 bound (6%) + 1.7 T1+ rules name an Escape ............... 7 rules, 0 missing + +1.7 is the tell. **Where the format is used it works.** The gap is propagation, +so the fix is a per-file check cheap enough to run on every edit, plus templates +at `.context/templates/contracts/`. + +## The failure this is really built to catch + +The clause format is parsed with real regexes, and a `**Binds:**` block is +terminated by a blank line. So this: + + > **Binds:** any agent writing here + > **Tier:** T2 (blocked) + > + > **Check:** bin/some-tool + > **Escape:** route through 00_inbox/ + +parses as a block with no Check and no Escape. It *looks* adopted to a human +reader and reads as unenforced to every tool. A contract can lose its +enforcement to one blank line, silently — which is check 1.5's stated defect +("a rule that silently lost its enforcement looks exactly like a rule that never +needed any") reproduced at the level of typography. + +That is L1 below, and it is the whole reason this runs pre-commit. + +## The ratchet + +Corpus-wide adoption is 2/135, so a lint that demanded adoption everywhere would +fail 133 files on the first run and be disabled within a day. Instead: + + a file that has NOT adopted the format ....... reported, never blocking + a file that HAS adopted it ................... may not regress + +Regression is the thing worth blocking, because it is the thing that happens +silently. Growth is left to the templates and to 10.5's corpus number. + +## Fail-closed, on purpose + +Unlike `bin/knowledge-write-guard` (fail-open — a broken guard must not block +work), this is a **checker**, and a checker that cannot read its input must not +report clean. Unreadable file, bad regex, unexpected shape: exit 2, say so. + + exit 0 checked, no blocking findings + exit 1 checked, blocking findings (regression, or --strict with findings) + exit 2 could not check — never confuse this with a pass + +Verify: bin/contract-lint --selftest +""" + +from __future__ import annotations + +import argparse +import os +import re +import subprocess +import sys +from dataclasses import dataclass +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] + +sys.path.insert(0, str(ROOT)) +from scripts.mainframe_paths import rel_to_root # noqa: E402 + +# ---------------------------------------------------------------- the grammar +# +# These MUST stay byte-identical to the ones in bin/contract-audit, or lint and +# audit disagree about what adoption means and the number nobody trusts is the +# one that gets ignored. tests/test_contract_lint.py extracts both and asserts +# equality, so drift fails a test rather than being discovered in an incident. + +NORMATIVE = re.compile( + r"\b(MUST(?! not)|MUST NOT|NEVER|ALWAYS|SHALL|REQUIRED|PROHIBITED" + r"|[Dd]o not|must not|may not|never|only|strictly)\b" +) + +MECHANISM = re.compile( + r"(bin/[a-z0-9._-]+" # a tool + r"|\.claude/settings" # config + r"|PreToolUse|PostToolUse|hook" + r"|pytest|test_[a-z_]+|selftest" + r"|--check\b|--apply\b|--dry-run\b|--fail-on" + r"|\.gitignore" + r")" +) + +TIER_RE = re.compile(r"^[>\s]*(?:\*\*)?Tier:?(?:\*\*)?:?\s*(T[0-3])\b", re.M) + +HARD_CLAUSE = re.compile( + r"\b(MUST NOT|MUST|NEVER|Never|never|PROHIBITED|prohibited|SHALL" + r"|may not|must not|strictly|immutable|read-only|only to|exclusive)\b" +) + +RULE_BLOCK = re.compile( + r"(?:^[>\s]*\*\*Binds:\*\*.*?)(?=^\s*$|^##)", re.M | re.S +) + +# Local to the lint: locating block starts so a truncated block can be reported +# with a line number the author can jump to. +BINDS_LINE = re.compile(r"^[>\s]*\*\*Binds:\*\*", re.M) +FIELD = { + "Tier": re.compile(r"^[>\s]*\*\*Tier:\*\*\s*\S", re.M), + "Check": re.compile(r"^[>\s]*\*\*Check:\*\*\s*\S", re.M), + "Escape": re.compile(r"^[>\s]*\*\*Escape:\*\*\s*\S", re.M), +} +CHECK_NONE = re.compile(r"^[>\s]*\*\*Check:\*\*\s*(none|n/?a)\b", re.M | re.I) + +# A field `--fix` scaffolded and nobody finished. This exists so the scaffold is +# safe to insert: `**Escape:** TODO` satisfies FIELD["Escape"] and would silence +# L5, turning a repair tool into a way to launder an unenforced rule. L9 keeps +# the file blocking until a human writes the words. +PLACEHOLDER = re.compile( + r"^[>\s]*\*\*(Binds|Tier|Check|Escape):\*\*\s*(?:TODO|FIXME|<[^>\n]*>)", re.M +) + +# Canonical field order, used when scaffolding a missing field back into place. +FIELD_ORDER = ("Binds", "Tier", "Check", "Escape") + +FIELD_HINT = { + "Binds": "TODO — who or what this constrains", + "Tier": "TODO — T0 advisory / T1 detected / T2 blocked / T3 reconciled", + "Check": "TODO — name a runnable tool, hook, or test; if none exists, the honest tier is T0", + "Escape": "TODO — the named, cheap, non-penalized way to comply when you cannot meet the letter", +} + +# Any line that *announces* a tier, however it is punctuated. L6 is then "this +# line declares a tier and TIER_RE will not match it", which is the only +# definition that matters — the audit is the reader being satisfied here. +# +# An earlier version tried to express this as one negative-lookahead regex and +# fired on every correct block: `Tier:?(?:\*\*)?:?\s*` can backtrack to consume +# only `Tier:`, leaving `**` to satisfy `(?!T[0-3])(\S+)`. Caught by the +# selftest's negative control, which is what that control is for. +TIER_LINE = re.compile(r"^[>\s]*(?:\*\*)?\s*Tier\s*(?:\*\*)?\s*:", re.I) + +# A contract file is anything carrying normative weight over agent behaviour. +# Enumerated, not remembered — the deny-list mistake has been made four times in +# this investigation, including once inside the tool built to catch it. +CONTRACT_GLOBS = ( + "AGENTS.md", "HARNESS.md", "EPISTEMIC_STANCE.md", "DECISIONS.md", + "**/AGENTS.md", "**/decisions.md", + ".context/workflows/*.md", ".agents/skills/**/SKILL.md", "agents/*.md", +) +SKIP_DIRS = { + ".git", ".venv", "node_modules", "__pycache__", "90_archive", + ".pytest_cache", ".ruff_cache", "tmp", +} +# Templates describe the format; they are not themselves in force. +SKIP_PARTS = (".context/templates/",) + + +@dataclass(frozen=True) +class Finding: + code: str + line: int + blocking: bool + message: str + detail: str = "" + + +class Unreadable(RuntimeError): + """Raised when a file cannot be checked. Never reported as clean.""" + + +# ------------------------------------------------------------------- checking + + +def read(path: Path) -> str: + try: + return path.read_text(encoding="utf-8") + except Exception as exc: # noqa: BLE001 — surfaced as exit 2, never as clean + raise Unreadable(f"{path}: {exc}") from exc + + +def line_of(text: str, index: int) -> int: + return text.count("\n", 0, index) + 1 + + +def adopted(text: str) -> bool: + """Has this file taken on the format at all?""" + return bool(TIER_RE.search(text)) + + +def lint_text(text: str) -> list[Finding]: + """Every finding for one contract file's contents. + + `blocking` marks the ones that fail an adopted file. Non-blocking findings + are reported for every file and block nothing, because 133 of 135 files have + not adopted the format yet and a lint that fails all of them gets turned off. + """ + out: list[Finding] = [] + has_format = adopted(text) + + matches = list(RULE_BLOCK.finditer(text)) + lines = text.splitlines() + + for m in matches: + block = m.group(0) + ln = line_of(text, m.start()) + first = block.strip().splitlines()[0][:60] + + missing = [name for name, rx in FIELD.items() if not rx.search(block)] + + # L1 — the fields exist, just past a blank line that ended the block. + # + # This is the silent-zero-adoption trap and the reason this runs + # pre-commit: the file looks adopted to a human and reads as unenforced + # to every tool. Distinguished from L2 by looking at what follows the + # block — "you forgot the Escape" and "your Escape is unparsed" need + # different fixes, and only one of them is invisible. + if missing: + after = "\n".join(lines[line_of(text, m.end()) - 1:][:6]) + stranded = [n for n in missing if FIELD[n].search(after)] + if stranded: + out.append(Finding( + "L1", ln, True, + f"rule block ends early — {', '.join(stranded)} " + "follow(s) a blank line and will not be parsed", + "remove the blank line, or prefix it with '>' to keep the " + "block open", + )) + remaining = [n for n in missing if n not in stranded] + if remaining: + out.append(Finding( + "L2", ln, True, + f"rule block missing {', '.join(remaining)}", + first, + )) + + tm = TIER_RE.search(block) + tier = tm.group(1) if tm else None + + if tier and tier != "T0": + if CHECK_NONE.search(block): + out.append(Finding( + "L3", ln, True, + f"{tier} declares 'Check: none' — a tier above T0 claims a " + "mechanism exists; if there is none the honest tier is T0", + first, + )) + elif not MECHANISM.search(block): + out.append(Finding( + "L4", ln, False, + f"{tier} names no runnable mechanism in the block", + "Check should name a bin/ tool, hook, pytest path, or " + "--check/--apply/--dry-run flag", + )) + if not FIELD["Escape"].search(block): + out.append(Finding( + "L5", ln, True, + f"{tier} names no Escape — an enforced rule with no honest " + "way to fail manufactures violations", + first, + )) + + # L9 — a scaffolded field nobody filled in. Always blocking: an unfinished + # rule is the one state that must not be committable, or `--fix` becomes a + # way to make an unenforced rule look enforced. + for m in PLACEHOLDER.finditer(text): + out.append(Finding( + "L9", line_of(text, m.start()), True, + f"{m.group(1)} is still a scaffold placeholder", + "written by `--fix`; replace it with the real words before committing", + )) + + # L6 — a Tier line that is not T0..T3 reads as no tier at all. + for i, raw in enumerate(text.splitlines(), start=1): + if not TIER_LINE.match(raw): + continue + if TIER_RE.match(raw): + continue + value = raw.split(":", 1)[-1].strip().strip("*").strip() or "(empty)" + out.append(Finding( + "L6", i, True, + f"Tier value {value!r} is not T0/T1/T2/T3", + "the audit's TIER_RE will not match it, so the rule reads as untiered", + )) + + # L7 — hard promises standing outside any rule block. Reported everywhere, + # blocking only where the file has shown it knows how to bind them. + covered = set() + for m in matches: + lo = line_of(text, m.start()) + covered.update(range(lo - 6, lo + m.group(0).count("\n") + 2)) + + loose: list[str] = [] + for i, raw in enumerate(text.splitlines(), start=1): + s = raw.strip() + if not HARD_CLAUSE.search(s): + continue + if not re.match(r"^(\d+\.|[-*]|\|)", s): + continue # numbered/bulleted rules are the normative unit + if i in covered or MECHANISM.search(s): + continue + loose.append(f"L{i}: {s[:80]}") + if loose: + out.append(Finding( + "L7", 0, False, + f"{len(loose)} hard clause(s) (MUST/NEVER/immutable/read-only) name " + "no mechanism and sit outside any rule block", + "\n".join(f" {x}" for x in loose[:6]), + )) + + # L8 — the file makes normative claims and has not adopted the format. + # Never blocking. This is the corpus-wide backlog, not a defect in this edit. + if not has_format: + n = sum(1 for ln in text.splitlines() if NORMATIVE.search(ln)) + if n: + out.append(Finding( + "L8", 0, False, + f"{n} normative clause(s), no Binds/Tier/Check/Escape block", + "start from .context/templates/contracts/agents-md.template.md", + )) + + return out + + +# ------------------------------------------------------------------ repairing +# +# Adoption is 14/153 and the corpus is growing faster than adoption (114 -> 153 +# since 2026-08-23, while adopted went 4 -> 14). A lint that only names what is +# wrong leaves the compliant route strictly longer than the non-compliant one: +# write the sentence, then recall four field names, then author a mechanism. +# So the lint repairs what it can. +# +# What it may repair is bounded by honesty, not by difficulty: +# +# typography the author already wrote the right words and lost them to a +# blank line (L1) or a bold marker (L6). Intent is on the page; +# repairing it invents nothing. Fixed outright. +# content a missing Escape cannot be invented. Scaffolded as a visibly +# unfinished TODO that L9 blocks on, so starting is one command +# and finishing dishonestly is not available. +# +# That boundary is the whole design. Cross it and this becomes a tool for +# manufacturing green contracts, which is the failure it was built against. + +MAX_PASSES = 40 + + +def _prefix_of(line: str) -> str: + """The '> ' quoting a rule block is written with, so repairs match it.""" + return re.match(r"^([>\s]*)", line).group(1) or "> " + + +def fix_pass(text: str) -> tuple[str, str | None]: + """Apply at most one repair. Returns (text, description or None if settled). + + One edit per pass, because every insertion shifts the line numbers the next + repair would be computed from. The caller loops. + """ + lines = text.splitlines() + trailing = "\n" if text.endswith("\n") else "" + + def joined() -> str: + return "\n".join(lines) + trailing + + # R1 — a tier value wrapped in bold, which TIER_RE cannot see. This is the + # defect found in 10_knowledge/AGENTS.md rule 5: the file others are told to + # copy had silently lost a tier to markup. + for i, raw in enumerate(lines): + if not TIER_LINE.match(raw) or TIER_RE.match(raw): + continue + fixed = re.sub(r"(\*\*Tier:\*\*\s*)\*\*(.*?)\*\*\s*$", r"\1\2", raw) + if fixed != raw: + lines[i] = fixed + return joined(), f"L{i + 1}: unbolded the tier value so it parses as a tier" + + # R2 — a blank line severing a block from its own Check and Escape. The + # silent-zero-adoption trap; the fields are right there, six lines down. + for m in RULE_BLOCK.finditer(text): + block = m.group(0) + missing = [n for n, rx in FIELD.items() if not rx.search(block)] + if not missing: + continue + bl = line_of(text, m.end()) + after = "\n".join(lines[bl - 1:][:6]) + stranded = [n for n in missing if FIELD[n].search(after)] + if stranded and bl - 1 < len(lines) and not lines[bl - 1].strip(): + lines[bl - 1] = _prefix_of(lines[line_of(text, m.start()) - 1]).rstrip() or ">" + return joined(), (f"L{bl}: reopened the rule block — a blank line was " + f"hiding {', '.join(stranded)}") + + # R3 — a field that is genuinely absent. Scaffolded, never invented. + for m in RULE_BLOCK.finditer(text): + block = m.group(0) + start = line_of(text, m.start()) + span = block.rstrip("\n").count("\n") + 1 + bl = start + span + after = "\n".join(lines[bl - 1:][:6]) if bl - 1 < len(lines) else "" + + present: dict[str, int] = {} + for name, rx in FIELD.items(): + hit = rx.search(block) + if hit: + present[name] = start + block.count("\n", 0, hit.start()) + present["Binds"] = start + + for name in FIELD_ORDER: + if name in present or FIELD.get(name, PLACEHOLDER).search(after): + continue # present, or stranded past a blank line — R2's job + before = [n for n in FIELD_ORDER[:FIELD_ORDER.index(name)] if n in present] + anchor_line = present[before[-1]] if before else start + prefix = _prefix_of(lines[start - 1]) + lines.insert(anchor_line, f"{prefix}**{name}:** {FIELD_HINT[name]}") + return joined(), (f"L{anchor_line + 1}: scaffolded a missing {name} " + f"— fill it in, L9 blocks until you do") + + return text, None + + +def fix_text(text: str) -> tuple[str, list[str]]: + """Repair until settled. Returns (text, what changed).""" + repairs: list[str] = [] + for _ in range(MAX_PASSES): + text, what = fix_pass(text) + if what is None: + return text, repairs + repairs.append(what) + repairs.append(f"stopped after {MAX_PASSES} passes — repair did not settle") + return text, repairs + + +SCAFFOLD = """## N. State the rule as one sentence in the imperative. + +{prefix}**Binds:** {binds} +{prefix}**Tier:** {tier} +{prefix}**Check:** {check} +{prefix}**Escape:** {escape} +""" + + +TIER_LABEL = {"T0": "advisory", "T1": "detected", "T2": "blocked", "T3": "reconciled"} + + +def scaffold(tier: str) -> str: + """A correct empty block. T0 comes pre-filled, because 'no mechanism' is a + legitimate answer and making the honest tier the cheapest one to write is + the point.""" + hint = dict(FIELD_HINT) + if tier == "T0": + hint["Check"] = "none" + hint["Escape"] = "n/a" + return SCAFFOLD.format( + prefix="> ", + binds=hint["Binds"], + tier=f"{tier} ({TIER_LABEL[tier]})", + check=hint["Check"], + escape=hint["Escape"], + ) + + +def contract_files() -> list[Path]: + seen: dict[Path, None] = {} + for pattern in CONTRACT_GLOBS: + for p in ROOT.glob(pattern): + if not p.is_file(): + continue + rel = p.relative_to(ROOT).as_posix() + if any(part in SKIP_DIRS for part in p.relative_to(ROOT).parts): + continue + if any(rel.startswith(s) for s in SKIP_PARTS): + continue + seen[p] = None + return sorted(seen) + + +def is_contract(path: Path) -> bool: + try: + rel = path.resolve().relative_to(ROOT).as_posix() + except ValueError: + return False # outside the repo: not our contract corpus + if any(rel.startswith(s) for s in SKIP_PARTS): + return False + name = path.name + return ( + name in {"AGENTS.md", "HARNESS.md", "EPISTEMIC_STANCE.md", + "DECISIONS.md", "decisions.md", "SKILL.md"} + or rel.startswith(".context/workflows/") + or (rel.startswith("agents/") and name.endswith(".md")) + ) + + +# -------------------------------------------------------------------- reports + + +def report(path: Path, findings: list[Finding], text: str, quiet: bool) -> bool: + """Print one file's result. Returns True if it has blocking findings.""" + rel = rel_to_root(path, ROOT) + blocking = [f for f in findings if f.blocking] if adopted(text) else [] + + if not findings: + if not quiet: + state = "adopted" if adopted(text) else "no normative clauses" + print(f" OK {rel} ({state})") + return False + + if quiet and not blocking: + return False + + print(f"\n {'FAIL' if blocking else 'note'} {rel}" + f"{'' if adopted(text) else ' [not adopted — reported, not blocking]'}") + for f in findings: + mark = "!" if (f.blocking and adopted(text)) else "-" + where = f":{f.line}" if f.line else "" + print(f" {mark} {f.code}{where} {f.message}") + if f.detail: + for dl in f.detail.splitlines(): + print(f" {dl.strip() if not dl.startswith(' ') else dl}") + return bool(blocking) + + +def changed_files() -> list[Path]: + """Staged contract files, for pre-commit use.""" + try: + out = subprocess.run( + ["git", "-C", str(ROOT), "diff", "--cached", "--name-only", + "--diff-filter=ACM"], + capture_output=True, text=True, timeout=15, check=True, + ).stdout + except Exception as exc: # noqa: BLE001 — a checker must not report clean + raise Unreadable(f"git diff --cached failed: {exc}") from exc + return [ROOT / line for line in out.splitlines() if line.strip()] + + +# ------------------------------------------------------- grammar drift control + + +# The patterns that must stay identical to bin/contract-audit's. RULE_BLOCK and +# the rest are shared vocabulary; a divergence means lint and audit disagree +# about what a rule *is*. +GRAMMAR = { + "NORMATIVE": NORMATIVE, + "MECHANISM": MECHANISM, + "TIER_RE": TIER_RE, + "HARD_CLAUSE": HARD_CLAUSE, + "RULE_BLOCK": RULE_BLOCK, +} + + +def audit_patterns() -> dict[str, str]: + """Extract `NAME = re.compile("...")` pattern literals from bin/contract-audit. + + Uses the AST rather than a regex over source, because the audit writes its + patterns as implicitly-concatenated string literals across many lines and + any textual comparison would be measuring formatting, not grammar. + """ + import ast + + audit = ROOT / "bin" / "contract-audit" + try: + tree = ast.parse(audit.read_text(encoding="utf-8")) + except Exception as exc: # noqa: BLE001 — surfaced as exit 2, never as clean + raise Unreadable(f"{audit}: {exc}") from exc + + found: dict[str, str] = {} + for node in ast.walk(tree): + if not isinstance(node, ast.Assign) or len(node.targets) != 1: + continue + target = node.targets[0] + if not isinstance(target, ast.Name) or target.id not in GRAMMAR: + continue + call = node.value + if not (isinstance(call, ast.Call) and call.args): + continue + try: + found[target.id] = ast.literal_eval(call.args[0]) + except Exception: # noqa: BLE001 — a non-literal pattern is a mismatch + continue + return found + + +def grammar_drift() -> list[tuple[str, str]]: + """Names where this file and the audit disagree, with why. Empty == in sync.""" + theirs = audit_patterns() + out: list[tuple[str, str]] = [] + for name, rx in GRAMMAR.items(): + if name not in theirs: + out.append((name, "not found in bin/contract-audit")) + elif theirs[name] != rx.pattern: + out.append((name, "pattern differs")) + return out + + +# ------------------------------------------------------------------- selftest + + +SELFTEST_GOOD = """## 1. Raw captures enter through ingest. + +> **Binds:** any client writing a raw file +> **Tier:** T2 (blocked) +> **Check:** `bin/knowledge-write-guard` (PreToolUse hook on Write) +> **Escape:** write to `00_inbox/` and run `bin/ingest-minion run --apply` +""" + +# A *real* blank line, not a '>'-prefixed one. `RULE_BLOCK`'s terminator is +# `^\s*$`, so a line containing '>' keeps the block open and is harmless; only +# a genuinely empty line severs it. Getting this fixture wrong the first time is +# what proved the check needed to distinguish L1 from L2 at all. +SELFTEST_TRUNCATED = """## 1. Raw captures enter through ingest. + +> **Binds:** any client writing a raw file +> **Tier:** T2 (blocked) + +> **Check:** `bin/knowledge-write-guard` +> **Escape:** write to `00_inbox/` +""" + +SELFTEST_CONTINUED = """## 1. Raw captures enter through ingest. + +> **Binds:** any client writing a raw file +> **Tier:** T2 (blocked) +> +> **Check:** `bin/knowledge-write-guard` +> **Escape:** write to `00_inbox/` +""" + +SELFTEST_NO_ESCAPE = """## 1. Raw captures enter through ingest. + +> **Binds:** any client writing a raw file +> **Tier:** T2 (blocked) +> **Check:** `bin/knowledge-write-guard` +""" + +SELFTEST_FAKE_T2 = """## 1. Raw captures enter through ingest. + +> **Binds:** any client +> **Tier:** T2 (blocked) +> **Check:** none +> **Escape:** ask an operator +""" + +SELFTEST_HONEST_T0 = """## 1. Do not tidy this folder. + +> **Binds:** any agent +> **Tier:** T0 (advisory) +> **Check:** none +> **Escape:** n/a +""" + + +SELFTEST_BOLD_TIER = """## 1. Do not tidy this folder. + +> **Binds:** any agent +> **Tier:** **T0 (advisory). Nothing checks this.** +> **Check:** none +> **Escape:** n/a +""" + +SELFTEST_SCAFFOLDED = """## 1. Raw captures enter through ingest. + +> **Binds:** any client writing a raw file +> **Tier:** T2 (blocked) +> **Check:** `bin/knowledge-write-guard` +> **Escape:** TODO — the named, cheap, non-penalized way to comply +""" + + +def selftest() -> int: + """Prove the lint still fires on the shapes it exists to catch. + + A checker that passes everything is indistinguishable from a checker that is + broken, so each case asserts a specific code fires — and the honest-T0 and + good cases assert that nothing fires, which is the negative control. + """ + cases = [ + ("clean block fires nothing", SELFTEST_GOOD, None), + ("blank line truncates a block", SELFTEST_TRUNCATED, "L1"), + ("'>' continuation line is not a truncation", SELFTEST_CONTINUED, None), + ("T2 without an Escape", SELFTEST_NO_ESCAPE, "L5"), + ("T2 claiming 'Check: none'", SELFTEST_FAKE_T2, "L3"), + ("honest T0 fires nothing", SELFTEST_HONEST_T0, None), + ("a bolded tier value reads as untiered", SELFTEST_BOLD_TIER, "L6"), + ("an unfilled scaffold blocks", SELFTEST_SCAFFOLDED, "L9"), + ] + failed = 0 + print("contract-lint selftest") + for name, text, expect in cases: + codes = {f.code for f in lint_text(text)} + ok = (expect in codes) if expect else not codes + print(f" {'PASS' if ok else 'FAIL'} {name}" + f"{'' if ok else f' (expected {expect or "no findings"}, got {sorted(codes) or "none"})'}") + failed += 0 if ok else 1 + + # Repair controls. + # + # `--fix` writes to contract files, so each repair asserts two things: that + # it closes the finding it claims to, and that it does not close one it must + # not. The last case is the one that matters — a repair that silenced L5 by + # writing a plausible Escape would be this tool defeating its own purpose. + repairs = [ + ("unbolding a tier clears L6", SELFTEST_BOLD_TIER, "L6", None), + ("reopening a block clears L1", SELFTEST_TRUNCATED, "L1", None), + ("scaffolding a missing Escape clears L5", SELFTEST_NO_ESCAPE, "L5", "L9"), + ] + for name, text, closes, leaves in repairs: + before = {f.code for f in lint_text(text)} + fixed, what = fix_text(text) + after = {f.code for f in lint_text(fixed)} + settled, _ = fix_pass(fixed) + ok = (closes in before and closes not in after and bool(what) + and (leaves is None or leaves in after) + and settled == fixed) + print(f" {'PASS' if ok else 'FAIL'} {name}" + f"{'' if ok else f' (before {sorted(before)}, after {sorted(after)})'}") + failed += 0 if ok else 1 + + # The repair boundary, stated as a control: a scaffolded rule must never + # reach zero findings on its own. Cheap to start, impossible to finish + # dishonestly — that property is the reason --fix is allowed to write at all. + fixed, _ = fix_text(SELFTEST_NO_ESCAPE) + honest = bool([f for f in lint_text(fixed) if f.blocking]) + print(f" {'PASS' if honest else 'FAIL'} a scaffold cannot pass as a finished rule") + failed += 0 if honest else 1 + + # Negative control on the grammar itself. + # + # Lint and audit must agree on what adoption means, or the number nobody + # trusts is the one that gets ignored. Compared by extracting the audit's + # own `X = re.compile(...)` literals from its AST, not by substring search: + # a fuzzy match here would pass while the two drifted, which is the exact + # class of false-green this whole program exists to kill. + audit = ROOT / "bin" / "contract-audit" + if not audit.is_file(): + print( + " SKIP grammar matches bin/contract-audit " + "(optional/deferred; not in this tree)" + ) + else: + try: + mismatch = grammar_drift() + except Unreadable as exc: + print(f" FAIL grammar matches bin/contract-audit (unreadable: {exc})") + failed += 1 + else: + ok = not mismatch + print(f" {'PASS' if ok else 'FAIL'} grammar matches bin/contract-audit " + f"({'in sync, ' + str(len(GRAMMAR)) + ' patterns' if ok else 'DRIFTED'})") + for name, why in mismatch: + print(f" {name}: {why}") + failed += 0 if ok else 1 + + total = len(cases) + len(repairs) + 2 # + scaffold-honesty + grammar drift + print(f"\n {total - failed}/{total} passed") + return 1 if failed else 0 + + +# ----------------------------------------------------------------------- main + + +def main() -> int: + ap = argparse.ArgumentParser( + description="Check contract files against the Binds/Tier/Check/Escape format.", + ) + g = ap.add_mutually_exclusive_group(required=True) + g.add_argument("--file", action="append", metavar="PATH", + help="lint one contract file (repeatable)") + g.add_argument("--all", action="store_true", + help="lint the whole contract corpus and report adoption") + g.add_argument("--changed", action="store_true", + help="lint staged contract files (pre-commit mode)") + g.add_argument("--selftest", action="store_true", + help="prove the lint still fires on known-bad shapes") + g.add_argument("--scaffold", metavar="TIER", nargs="?", const="T2", + choices=("T0", "T1", "T2", "T3"), + help="print a correct empty rule block to paste (default T2)") + ap.add_argument("--fix", action="store_true", + help="repair what can be repaired without inventing content, " + "and scaffold the rest as TODO (blocks on L9 until filled). " + "Use with --file or --changed.") + ap.add_argument("--strict", action="store_true", + help="exit 1 on any finding, including in unadopted files") + ap.add_argument("--quiet", action="store_true", + help="print only files with blocking findings") + args = ap.parse_args() + + if args.selftest: + return selftest() + + if args.scaffold: + print(scaffold(args.scaffold), end="") + return 0 + + if args.fix and args.all: + # 153 files, 228 findings. A corpus-wide rewrite is a decision, not a + # lint invocation, and it would bury the one edit worth reviewing. + print("contract-lint: --fix needs --file or --changed, not --all", + file=sys.stderr) + return 2 + + if args.all: + paths = contract_files() + elif args.changed: + paths = [p for p in changed_files() if p.exists() and is_contract(p)] + if not paths: + return 0 + else: + paths = [] + for raw in args.file: + p = Path(raw) + p = p if p.is_absolute() else (Path.cwd() / p) + if not p.exists(): + print(f"contract-lint: no such file: {raw}", file=sys.stderr) + return 2 + paths.append(p.resolve()) + + blocking = 0 + findings_total = 0 + adopted_n = 0 + + for p in paths: + text = read(p) + if args.fix: + fixed, repairs = fix_text(text) + if repairs: + p.write_text(fixed, encoding="utf-8") + print(f"\n FIXED {rel_to_root(p, ROOT)}") + for r in repairs: + print(f" + {r}") + text = fixed + if adopted(text): + adopted_n += 1 + fs = lint_text(text) + findings_total += len(fs) + if report(p, fs, text, args.quiet): + blocking += 1 + if args.strict and fs: + blocking += 0 if adopted(text) else 1 + + if args.all: + pct = 100 * adopted_n / max(len(paths), 1) + print(f"\n {adopted_n}/{len(paths)} contract file(s) declare a tier " + f"({pct:.0f}% adopted) · {findings_total} finding(s) · " + f"{blocking} file(s) blocking") + print(" Unadopted files are reported, never blocking. Adopted files " + "may not regress.") + elif not args.quiet and not findings_total: + print(f" {len(paths)} file(s) checked, no findings.") + + return 1 if blocking else 0 + + +if __name__ == "__main__": + try: + sys.exit(main()) + except Unreadable as exc: + print(f"contract-lint: could not check — {exc}", file=sys.stderr) + print("This is exit 2, not a pass.", file=sys.stderr) + sys.exit(2) + except KeyboardInterrupt: + sys.exit(2) + except Exception as exc: # noqa: BLE001 — fail closed, never silently clean + print(f"contract-lint: unexpected failure — {exc!r}", file=sys.stderr) + print("This is exit 2, not a pass.", file=sys.stderr) + sys.exit(2) diff --git a/bin/craft-research-loop b/bin/craft-research-loop deleted file mode 100755 index 1243b70..0000000 --- a/bin/craft-research-loop +++ /dev/null @@ -1,789 +0,0 @@ -#!/usr/bin/env python3 -"""Craft research loop — product/trial-and-error research with finishable FINDINGS. - -Separate from research-lane-loop (literature) and project-experiment-loop (evals). - -Commands: - preflight --project <slug> - scaffold --project <slug> --question ... --decision ... [--slug ...] - close --project <slug> --trial <id> --verdict keep|kill|iterate --findings ... - triage --project <slug> - promote --project <slug> --trial <id> --to knowledge|experiment|lane - -See .context/workflows/craft-research-loop.md -""" - -from __future__ import annotations - -import argparse -import re -import sys -from datetime import date, datetime, timezone -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -PROJECTS = ROOT / "30_projects" -TEMPLATE = ROOT / ".context" / "templates" / "craft-trial.md" -WORKFLOW = ROOT / ".context" / "workflows" / "craft-research-loop.md" -TODAY = date.today().isoformat() - - -def project_dir(slug: str) -> Path: - p = PROJECTS / slug - if not p.is_dir(): - raise SystemExit(f"Project not found: {p}") - return p - - -def parse_frontmatter(path: Path) -> dict[str, str]: - if not path.exists(): - return {} - text = path.read_text(encoding="utf-8", errors="replace") - if not text.startswith("---"): - return {} - end = text.find("\n---", 3) - if end < 0: - return {} - out: dict[str, str] = {} - for line in text[3:end].splitlines(): - if ":" not in line: - continue - k, v = line.split(":", 1) - out[k.strip()] = v.strip().strip('"').strip("'") - return out - - -def set_frontmatter_field(path: Path, key: str, value: str) -> None: - text = path.read_text(encoding="utf-8") - if not text.startswith("---"): - return - end = text.find("\n---", 3) - if end < 0: - return - block = text[3:end] - body = text[end:] - pattern = rf"(?m)^{re.escape(key)}:\s*.*$" - line = f'{key}: "{value}"' - if re.search(pattern, block): - block = re.sub(pattern, line, block) - else: - block = block.rstrip() + "\n" + line + "\n" - path.write_text("---" + block + body, encoding="utf-8") - - -def proof_index_path(proj: Path) -> Path: - # Prefer outputs/RESEARCH_PROOF_INDEX.md; allow README pointer later - return proj / "outputs" / "RESEARCH_PROOF_INDEX.md" - - -def action_card_path(proj: Path) -> Path: - return proj / "outputs" / "LAST_CRAFT_ACTION.md" - - -def list_trial_dirs(proj: Path) -> list[Path]: - outputs = proj / "outputs" - if not outputs.is_dir(): - return [] - trials: list[Path] = [] - for p in sorted(outputs.iterdir()): - if not p.is_dir(): - continue - if p.name.startswith("."): - continue - # Heuristic: trial folders are dated or contain TRIAL/FINDINGS - if re.match(r"^\d{4}-\d{2}-\d{2}", p.name) or (p / "TRIAL.md").exists() or ( - p / "FINDINGS.md" - ).exists(): - trials.append(p) - return trials - - -def trial_status(trial: Path) -> str: - findings = trial / "FINDINGS.md" - trial_md = trial / "TRIAL.md" - if findings.exists(): - text = findings.read_text(encoding="utf-8", errors="replace").lower() - # closed only when verdict is explicit keep|kill|iterate - for v in ("keep", "kill", "iterate"): - if re.search(rf"`{v}`|\*\*{v}\*\*|verdict:\s*{v}|\bverdict\b[^\n]*\b{v}\b", text): - # open findings template has `open` - if v == "open": - continue - return v - if re.search(r"`open`|verdict:\s*open|status:\s*\*\*open\*\*", text): - return "running" - if "verdict" in text: - return "findings-present" - return "findings-present" - if trial_md.exists(): - fm = parse_frontmatter(trial_md) - v = (fm.get("verdict") or "").lower() - if v in {"keep", "kill", "iterate"}: - return v - st = (fm.get("status") or "").lower() - if st in {"complete", "closed"}: - return "closed" - return "running" - return "artifact-only" - - -def cmd_preflight(project: str) -> int: - proj = project_dir(project) - fm = parse_frontmatter(proj / "README.md") - proof = proof_index_path(proj) - trials = list_trial_dirs(proj) - open_t = [t for t in trials if trial_status(t) in {"running", "open", "artifact-only", "findings-present"}] - closed = [t for t in trials if trial_status(t) in {"keep", "kill", "iterate", "closed"}] - - print(f"# Craft research preflight — {project}") - print() - print(f"- **path:** `{proj.relative_to(ROOT)}`") - print(f"- **project_state:** {fm.get('project_state') or fm.get('status') or '?'}") - print(f"- **next_action:** {(fm.get('next_action') or '')[:160]}") - print(f"- **proof index:** {'yes' if proof.exists() else 'MISSING (will create on first close)'}") - print(f"- **trial folders:** {len(trials)} (open-ish={len(open_t)}, closed-ish={len(closed)})") - if open_t: - print() - print("## Open / incomplete trials") - for t in open_t[:12]: - print(f"- `{t.relative_to(ROOT)}` status={trial_status(t)}") - print() - print("## Next") - print("1. scaffold a trial (or close an open one)") - print("2. run workbench procedure") - print("3. close with --verdict keep|kill|iterate") - print(f"Workflow: `{WORKFLOW.relative_to(ROOT)}`") - return 0 - - -def cmd_scaffold( - project: str, - question: str, - decision: str, - slug: str | None, - hypothesis: str | None, -) -> int: - proj = project_dir(project) - out = proj / "outputs" - out.mkdir(parents=True, exist_ok=True) - slug_clean = re.sub(r"[^a-z0-9-]+", "-", (slug or question[:40]).lower()).strip("-")[:48] - trial_id = f"{TODAY}-{slug_clean}" - trial_dir = out / trial_id - if trial_dir.exists(): - print(f"Refusing overwrite: {trial_dir}", file=sys.stderr) - return 1 - trial_dir.mkdir(parents=True) - - fm_domain = parse_frontmatter(proj / "README.md").get("domain", "general") - trial_body = f"""--- -title: "Craft trial — {slug_clean}" -domain: "{fm_domain}" -type: "project" -status: "running" -craft_trial_id: "{trial_id}" -project: "{project}" -verdict: "open" -decision_sentence: "{decision.replace('"', "'")}" -tags: ["craft-research", "trial"] -updated: "{TODAY}" -source: "bin/craft-research-loop scaffold" ---- - -# Craft trial — {slug_clean} - -## Question - -{question} - -## Decision this trial supports - -{decision} - -## Hypothesis - -{hypothesis or "(state expected result before ranking outputs)"} - -## Success criteria - -(define before reviewing winners) - -## Fixed inputs - -- assets: -- models / stack: -- hardware: -- software versions: - -## Variables - -- independent: -- controlled: -- observed: - -## Procedure - -1. -2. -3. - -## Run table - -| Run | Config | Result | Notes | -|-----|--------|--------|-------| - -## Observations - -(facts only — fill during/after runs) - -## Interpretation - -(inferences labeled) - -## Limitations - -## Verdict - -`open` - -### If keep / kill / iterate - -(fill at close) - -## Promotion - -- [ ] none (local product only) -- [ ] decisions.md -- [ ] extract-knowledge / research-lane-loop -- [ ] project-experiment-loop -""" - (trial_dir / "TRIAL.md").write_text(trial_body, encoding="utf-8") - (trial_dir / "FINDINGS.md").write_text( - f"""# FINDINGS — {trial_id} - -Status: **open** (fill at close via `bin/craft-research-loop close`) - -## Observations - -## Interpretation - -## Limitations - -## Verdict - -`open` - -## Next trial (if iterate) - -## Promotion - -none yet -""", - encoding="utf-8", - ) - - # ensure proof index exists - proof = proof_index_path(proj) - if not proof.exists(): - proof.write_text( - f"""# Research / craft proof index — {project} - -Curated evidence only. Ephemeral intermediates should be culled at trial close. - -## Open trials - -| Trial | Question | Status | -|-------|----------|--------| -| `{trial_id}` | {question[:80]} | running | - -## Keep (canonical proof) - -| Path | Why it stays | -|------|----------------| -| | | - -## Removed - -| Path | Why removed | -|------|-------------| -| | | - -## How to re-run - -See project `plans/` and latest `LAST_CRAFT_ACTION.md`. - -Updated: {TODAY} -""", - encoding="utf-8", - ) - else: - text = proof.read_text(encoding="utf-8") - if "## Open trials" in text and trial_id not in text: - text = text.replace( - "## Open trials", - f"## Open trials\n\n| `{trial_id}` | {question[:60]} | running |", - 1, - ) - # if table header missing, still ok - proof.write_text(text, encoding="utf-8") - elif "## Open trials" not in text: - proof.write_text( - text.rstrip() - + f"\n\n## Open trials\n\n| Trial | Question | Status |\n|-------|----------|--------|\n| `{trial_id}` | {question[:60]} | running |\n", - encoding="utf-8", - ) - - next_a = f"Execute craft trial `{trial_id}` then close with keep|kill|iterate." - set_frontmatter_field(proj / "README.md", "next_action", next_a) - set_frontmatter_field(proj / "README.md", "updated", TODAY) - write_action_card( - proj, - severity="info", - primary=next_a, - extras=[ - f"Scaffolded `{trial_dir.relative_to(ROOT)}`", - f"Question: {question}", - ], - ) - print(f"Wrote {trial_dir.relative_to(ROOT)}/TRIAL.md") - print(f"craft_trial_id={trial_id}") - print("Next: execute procedure, then:") - print( - f" bin/craft-research-loop close --project {project} " - f"--trial {trial_id} --verdict keep|kill|iterate --findings \"...\"" - ) - return 0 - - -def ensure_findings( - trial_dir: Path, - trial_id: str, - verdict: str, - findings: str, - question: str, -) -> None: - path = trial_dir / "FINDINGS.md" - body = f"""# FINDINGS — {trial_id} - -Closed: {datetime.now(timezone.utc).strftime('%Y-%m-%d %H:%M UTC')} -Verdict: **{verdict}** - -## Question - -{question} - -## Observations - -{findings} - -## Interpretation - -(see observations; expand if needed) - -## Limitations - -Single-trial craft result — not a formal eval-registry study unless promoted. - -## Verdict - -`{verdict}` - -## Promotion - -- none unless explicitly run `craft-research-loop promote` -""" - path.write_text(body, encoding="utf-8") - - -def update_proof_index_on_close( - proj: Path, trial_id: str, verdict: str, findings: str -) -> None: - """G11: maintain Keep/Removed/Open tables without duplicate section noise.""" - proof = proof_index_path(proj) - rel = f"outputs/{trial_id}/" - why = findings.replace("\n", " ").replace("|", "/")[:120] - header = ( - f"# Research / craft proof index — {proj.name}\n\n" - "Curated evidence only. Ephemeral intermediates should be culled at trial close.\n\n" - f"Last close: {TODAY}\n\n" - ) - - def _parse_table_rows(section_body: str) -> list[tuple[str, str]]: - rows: list[tuple[str, str]] = [] - for line in section_body.splitlines(): - if not line.strip().startswith("|"): - continue - cells = [c.strip() for c in line.strip().strip("|").split("|")] - if len(cells) < 2: - continue - if cells[0].lower() in {"path", "trial"} or set(cells[0]) <= {"-"}: - continue - rows.append((cells[0], cells[1])) - return rows - - def _extract_sections(src: str) -> dict[str, str]: - parts = re.split(r"(?m)^(## .+)$", src) - sections: dict[str, str] = {"_preamble": parts[0] if parts else ""} - i = 1 - while i < len(parts) - 1: - sections[parts[i].strip()] = parts[i + 1] - i += 2 - return sections - - raw = proof.read_text(encoding="utf-8") if proof.exists() else header - sections = _extract_sections(raw) - keep_key = next((k for k in sections if k.lower().startswith("## keep")), "## Keep") - rem_key = next((k for k in sections if k.lower().startswith("## removed")), "## Removed") - open_key = next((k for k in sections if "open" in k.lower()), "## Open trials") - - keep_rows = _parse_table_rows(sections.get(keep_key, "")) - rem_rows = _parse_table_rows(sections.get(rem_key, "")) - open_rows = _parse_table_rows(sections.get(open_key, "")) - - def _drop(rows: list[tuple[str, str]], needle: str) -> list[tuple[str, str]]: - return [r for r in rows if needle not in r[0] and needle not in r[1]] - - keep_rows = _drop(keep_rows, trial_id) - rem_rows = _drop(rem_rows, trial_id) - open_rows = _drop(open_rows, trial_id) - - if verdict == "keep": - keep_rows.append((f"`{rel}`", why)) - elif verdict == "kill": - rem_rows.append((f"`{rel}`", f"killed: {why}")) - else: - open_rows.append((f"`{trial_id}`", "iterate — see FINDINGS; closed→next")) - - def _fmt(title: str, h0: str, h1: str, rows: list[tuple[str, str]]) -> str: - out_lines = [title, "", f"| {h0} | {h1} |", "|------|-----|"] - for a, b in rows: - out_lines.append(f"| {a} | {b} |") - out_lines.append("") - return "\n".join(out_lines) - - body = ( - header - + _fmt("## Keep", "Path", "Why", keep_rows) - + "\n" - + _fmt("## Removed", "Path", "Why", rem_rows) - + "\n" - + _fmt("## Open trials", "Trial", "Status", open_rows) - + f"\n<!-- craft-research-loop close {TODAY} {trial_id} {verdict} -->\n" - ) - extra = [] - for k, v in sections.items(): - if k in {"_preamble", keep_key, rem_key, open_key}: - continue - if k.startswith("## How") or "re-run" in k.lower(): - extra.append(f"{k}\n{v.rstrip()}\n") - if extra: - body = body.rstrip() + "\n\n" + "\n".join(extra) + "\n" - proof.write_text(body if body.endswith("\n") else body + "\n", encoding="utf-8") - -def append_log(proj: Path, entry: str) -> None: - log = proj / "log.md" - block = f"\n## {TODAY} | craft-research-loop\n\n{entry}\n" - if log.exists(): - log.write_text(log.read_text(encoding="utf-8") + block, encoding="utf-8") - else: - log.write_text("# Log\n" + block, encoding="utf-8") - - -def write_action_card( - proj: Path, *, severity: str, primary: str, extras: list[str] | None = None -) -> None: - lines = [ - f"# Last craft action — {proj.name}", - f"", - f"Generated: {datetime.now(timezone.utc).strftime('%Y-%m-%d %H:%M UTC')}", - f"", - f"## Severity: **{severity}**", - f"", - f"## Primary next action", - f"", - primary, - f"", - f"## Notes", - f"", - ] - for e in extras or []: - lines.append(f"- {e}") - lines.extend( - [ - "", - "Project-local (not on last-eval-action.md).", - f"Workflow: `{WORKFLOW.relative_to(ROOT)}`", - "", - ] - ) - out = action_card_path(proj) - out.parent.mkdir(parents=True, exist_ok=True) - out.write_text("\n".join(lines), encoding="utf-8") - - -def cmd_close( - project: str, - trial: str, - verdict: str, - findings: str, - blocked_on: str | None = None, - blocked_reason: str | None = None, -) -> int: - verdict = verdict.lower().strip() - if verdict not in {"keep", "kill", "iterate"}: - print("verdict must be keep|kill|iterate", file=sys.stderr) - return 2 - proj = project_dir(project) - trial_dir = proj / "outputs" / trial - if not trial_dir.is_dir(): - # allow bare slug match - matches = [p for p in list_trial_dirs(proj) if p.name == trial or p.name.endswith(trial)] - if not matches: - print(f"Trial folder not found under outputs/: {trial}", file=sys.stderr) - return 1 - trial_dir = matches[0] - trial = trial_dir.name - - question = "" - trial_md = trial_dir / "TRIAL.md" - if not trial_md.exists(): - # Legacy artifact folder close — write minimal TRIAL.md for loop shape - trial_md.write_text( - f"""--- -title: "Craft trial — {trial}" -domain: "general" -type: "project" -status: "complete" -craft_trial_id: "{trial}" -project: "{project}" -verdict: "{verdict}" -decision_sentence: "Retroactive craft close" -tags: ["craft-research", "trial", "retroactive"] -updated: "{TODAY}" -source: "bin/craft-research-loop close" ---- - -# Craft trial — {trial} - -## Question - -(retroactive close — see FINDINGS) - -## Verdict - -`{verdict}` -""", - encoding="utf-8", - ) - if trial_md.exists(): - t = trial_md.read_text(encoding="utf-8", errors="replace") - m = re.search(r"## Question\s*\n+(.+?)(\n## |\Z)", t, re.S) - if m: - question = m.group(1).strip() - set_frontmatter_field(trial_md, "verdict", verdict) - set_frontmatter_field(trial_md, "status", "complete") - set_frontmatter_field(trial_md, "updated", TODAY) - - ensure_findings(trial_dir, trial, verdict, findings, question or "(see FINDINGS)") - update_proof_index_on_close(proj, trial, verdict, findings) - - blocked_on = (blocked_on or "").strip().lower() or None - if verdict == "iterate" and blocked_on == "operator": - reason = (blocked_reason or "").strip() or "operator step required" - next_a = ( - f"Waiting on operator ({reason}) before continuing from `{trial}`." - ) - sev = "info" - elif verdict == "iterate": - next_a = f"Open next craft trial continuing from `{trial}` (iterate)." - sev = "medium" - elif verdict == "keep": - next_a = f"Adopt keep from `{trial}`; optional decisions.md + next product step." - sev = "info" - else: - next_a = f"Record kill from `{trial}`; do not re-run without new hypothesis." - sev = "medium" - - set_frontmatter_field(proj / "README.md", "next_action", next_a) - set_frontmatter_field(proj / "README.md", "updated", TODAY) - - append_log( - proj, - f"- **Trial:** `{trial}`\n- **Verdict:** {verdict}\n- **Findings:** {findings[:300]}\n" - f"- **Proof index:** updated\n- **Action:** {next_a}" - + (f"\n- **blocked_on:** {blocked_on} ({blocked_reason})" if blocked_on else ""), - ) - extras = [f"Closed `{trial}` with verdict={verdict}", findings[:200]] - if blocked_on: - extras.insert(0, f"blocked_on={blocked_on}: {blocked_reason or ''}") - write_action_card( - proj, - severity=sev, - primary=next_a, - extras=extras, - ) - print(f"Closed {trial_dir.relative_to(ROOT)} verdict={verdict}") - print(f"Wrote {action_card_path(proj).relative_to(ROOT)}") - print(f"next_action: {next_a}") - if verdict == "iterate": - print("Tip: scaffold the follow-up trial now with a tighter question.") - return 0 - - -def cmd_triage(project: str) -> int: - proj = project_dir(project) - fm = parse_frontmatter(proj / "README.md") - trials = list_trial_dirs(proj) - open_t = [] - closed = [] - for t in trials: - st = trial_status(t) - if st in {"keep", "kill", "iterate", "closed"}: - closed.append((t, st)) - else: - open_t.append((t, st)) - - # Prefer formal running trials (TRIAL.md) over legacy artifact-only folders - running = [(t, s) for t, s in open_t if s == "running"] - running.sort(key=lambda x: x[0].name, reverse=True) # newest date-slug first - legacy = [(t, s) for t, s in open_t if s != "running"] - - proof = proof_index_path(proj) - extras = [ - f"proof_index={'yes' if proof.exists() else 'MISSING'}", - f"open_trials={len(open_t)} (formal_running={len(running)}) closed_ish={len(closed)}", - ] - for t, st in (running + legacy)[:10]: - extras.append(f"open: {t.name} ({st})") - - na = fm.get("next_action") or "" - if running: - primary = ( - f"Execute then close formal trial `{running[0][0].name}` " - f"(keep|kill|iterate)." - ) - sev = "info" - elif na and "craft trial" in na.lower(): - primary = na - sev = "info" - elif legacy: - primary = ( - f"Backfill craft close on legacy proof `{legacy[0][0].name}` " - f"or ignore if already in RESEARCH_PROOF_INDEX Keep table." - ) - sev = "low" - elif na: - primary = na - sev = "info" - else: - primary = "Scaffold a craft trial from README goal (one question)." - sev = "medium" - - write_action_card(proj, severity=sev, primary=primary, extras=extras) - text = action_card_path(proj).read_text(encoding="utf-8") - print(text) - return 0 - - -def cmd_promote(project: str, trial: str, to: str) -> int: - proj = project_dir(project) - trial_dir = proj / "outputs" / trial - if not trial_dir.is_dir(): - print(f"Trial not found: {trial_dir}", file=sys.stderr) - return 1 - rel = trial_dir.relative_to(ROOT) - if to == "knowledge": - print("Promotion: durable knowledge") - print(f" 1. Review FINDINGS in {rel}") - print( - f" 2. bin/extract-knowledge --project {project} --domain <domain> " - f"--title \"craft lesson from {trial}\" --write" - ) - print(" 3. Or capture literature via research-lane-loop if sources external") - elif to == "experiment": - print("Promotion: formal measured experiment") - print( - f" bin/project-experiment-loop scaffold --project {project} " - f"--study-type exploratory --title \"{trial}\" " - f"--decision \"Graduate craft trial {trial} to measured decision\"" - ) - print(" Then run protocol, fill metrics, close/harvest.") - elif to == "lane": - print("Promotion: research lane") - print(" Emit ## Research Lane Candidate in FINDINGS or log, then:") - print(f" bin/lane-intake scan 30_projects/{project}/outputs/{trial}/FINDINGS.md") - else: - print("to must be knowledge|experiment|lane", file=sys.stderr) - return 2 - append_log(proj, f"- promote `{trial}` → **{to}** (instructions emitted; not auto-executed)") - return 0 - - -def main() -> None: - p = argparse.ArgumentParser(description="Craft research loop CLI") - sub = p.add_subparsers(dest="cmd", required=True) - - pf = sub.add_parser("preflight") - pf.add_argument("--project", required=True) - - sc = sub.add_parser("scaffold") - sc.add_argument("--project", required=True) - sc.add_argument("--question", required=True) - sc.add_argument("--decision", required=True) - sc.add_argument("--slug", default=None) - sc.add_argument("--hypothesis", default=None) - - cl = sub.add_parser("close") - cl.add_argument("--project", required=True) - cl.add_argument("--trial", required=True, help="outputs/<trial-id> folder name") - cl.add_argument("--verdict", required=True, choices=["keep", "kill", "iterate"]) - cl.add_argument("--findings", required=True) - cl.add_argument( - "--blocked-on", - default=None, - choices=["operator", "install", "external"], - help="G5: iterate waiting on non-agent work (keeps severity info)", - ) - cl.add_argument("--blocked-reason", default=None) - - tr = sub.add_parser("triage") - tr.add_argument("--project", required=True) - - pr = sub.add_parser("promote") - pr.add_argument("--project", required=True) - pr.add_argument("--trial", required=True) - pr.add_argument("--to", required=True, choices=["knowledge", "experiment", "lane"]) - - args = p.parse_args() - if args.cmd == "preflight": - raise SystemExit(cmd_preflight(args.project)) - if args.cmd == "scaffold": - raise SystemExit( - cmd_scaffold( - args.project, - args.question, - args.decision, - args.slug, - args.hypothesis, - ) - ) - if args.cmd == "close": - raise SystemExit( - cmd_close( - args.project, - args.trial, - args.verdict, - args.findings, - blocked_on=getattr(args, "blocked_on", None), - blocked_reason=getattr(args, "blocked_reason", None), - ) - ) - if args.cmd == "triage": - raise SystemExit(cmd_triage(args.project)) - if args.cmd == "promote": - raise SystemExit(cmd_promote(args.project, args.trial, args.to)) - p.error(f"unknown {args.cmd}") - - -if __name__ == "__main__": - main() diff --git a/bin/delegate-coder b/bin/delegate-coder deleted file mode 100755 index 95e20a7..0000000 --- a/bin/delegate-coder +++ /dev/null @@ -1,576 +0,0 @@ -#!/usr/bin/env python3 -"""Automate task packet delegation, execution, and repair for local coder agents.""" - -from __future__ import annotations - -import argparse -import json -import os -import re -import shlex -import subprocess -import sys -from datetime import UTC, datetime -from pathlib import Path - -# Profile details mapped to Ollama models and settings -PROFILES = { - "local-qwen25-coder-14b": { - "model": "ollama_chat/qwen2.5-coder:14b", - "ollama_model": "qwen2.5-coder:14b", - "edit_format": "diff", - "map_tokens": 4096, - }, - "local-qwen3-14b": { - "model": "ollama_chat/qwen3:14b", - "ollama_model": "qwen3:14b", - "edit_format": "diff", - "map_tokens": 4096, - }, - "local-qwen35-9b": { - "model": "ollama_chat/qwen3.5:9b", - "ollama_model": "qwen3.5:9b", - "edit_format": "whole", - "map_tokens": 4096, - }, - "local-qwen3-coder-30b": { - "model": "ollama_chat/qwen3-coder:30b", - "ollama_model": "qwen3-coder:30b", - "edit_format": "diff", - "map_tokens": 4096, - }, -} - -def detect_project(cwd: Path) -> tuple[str, Path]: - """Find the MainFrame root and current project slug if running within one.""" - curr = cwd.resolve() - root = None - for _ in range(10): - if (curr / "30_projects").is_dir(): - root = curr - break - if curr.parent == curr: - break - curr = curr.parent - - if root: - try: - rel = cwd.relative_to(root / "30_projects") - parts = rel.parts - if parts: - return parts[0], root - except ValueError: - pass - return "local-agent", root - - # Fallback: directory that contains this bin/ wrapper (MainFrame root). - default_root = Path(__file__).resolve().parent.parent - if default_root.is_dir(): - return "local-agent", default_root - return "local-agent", cwd - -def check_git_status(cwd: Path) -> bool: - """Check if git working tree has uncommitted changes.""" - proc = subprocess.run( - ["git", "status", "--porcelain"], - cwd=cwd, - capture_output=True, - text=True, - check=False - ) - if proc.returncode != 0: - return False - return bool(proc.stdout.strip()) - -def health_check(profile_name: str, profile: dict[str, any]) -> tuple[bool, str]: - """Verify that Aider and Ollama with the requested model are available.""" - if not shutil_which("aider"): - return False, "aider CLI not found in PATH" - if not shutil_which("ollama"): - return False, "ollama CLI not found in PATH" - - # Check Ollama status - proc = subprocess.run( - ["ollama", "list"], - capture_output=True, - text=True, - check=False - ) - if proc.returncode != 0: - return False, "Failed to communicate with Ollama daemon (ollama list failed)" - - model = profile["ollama_model"] - if model and model not in proc.stdout: - return False, f"Ollama model '{model}' is not pulled/installed. Run 'ollama pull {model}' first." - - return True, "health check passed" - -def shutil_which(cmd: str) -> bool: - """Simple check for executable presence.""" - for path in os.environ.get("PATH", "").split(os.pathsep): - p = Path(path) / cmd - if p.is_file() and os.access(p, os.X_OK): - return True - return False - -def get_git_diff(cwd: Path, editable_files: list[str]) -> str: - """Retrieve current git diff including any untracked files in editable_files.""" - proc = subprocess.run(["git", "diff", "HEAD"], cwd=cwd, capture_output=True, text=True, check=False) - diff = proc.stdout or "" - - for f in editable_files: - status_proc = subprocess.run(["git", "status", "--porcelain", f], cwd=cwd, capture_output=True, text=True, check=False) - if status_proc.stdout.startswith("??"): - diff_proc = subprocess.run(["git", "diff", "--no-index", "--", "/dev/null", f], cwd=cwd, capture_output=True, text=True, check=False) - if diff_proc.stdout: - diff += "\n" + diff_proc.stdout - return diff - -def render_prompt(packet: dict[str, any], repair_context: str = "") -> str: - """Generate the instructions file content for Aider.""" - metadata = packet["metadata"] - sections = packet["sections"] - if repair_context: - return ( - "Perform one fresh-context repair attempt. Use only the failed " - "verifier summary, current diff, and scope below. Do not broaden " - "scope, commit, or run/claim external verification.\n\n" - f"Editable files: {json.dumps(metadata['editable_files'])}\n" - f"Files allowed to be created: {json.dumps(metadata['create_files'])}\n" - f"Read-only context: {json.dumps(metadata['read_only_files'])}\n\n" - "## Failed verification and current diff\n" - f"{repair_context}\n\n" - "End with exactly these fields:\n" - "TASK_STATUS: complete|blocked\n" - "VERIFICATION_CLAIM: not_run\n" - "TASK_SUMMARY: one concise line\n" - ) - - rendered_sections = "\n\n".join( - f"## {name}\n{sections[name]}" for name in sections - ) - return ( - "Implement this reviewed task packet exactly. Do not broaden scope, " - "commit, or run/claim external verification.\n\n" - f"Editable files: {json.dumps(metadata['editable_files'])}\n" - f"Files allowed to be created: {json.dumps(metadata['create_files'])}\n" - f"Read-only context: {json.dumps(metadata['read_only_files'])}\n\n" - f"{rendered_sections}\n\n" - "End with exactly these fields:\n" - "TASK_STATUS: complete|blocked\n" - "VERIFICATION_CLAIM: not_run\n" - "TASK_SUMMARY: one concise line\n" - ) - -def parse_claims(stdout: str) -> dict[str, str | None]: - """Extract completion, verification, and summary fields from output.""" - status = re.search(r"^TASK_STATUS:\s*(complete|blocked)\s*$", stdout, re.MULTILINE) - verification = re.search(r"^VERIFICATION_CLAIM:\s*(not_run|passed|failed)\s*$", stdout, re.MULTILINE) - summary = re.search(r"^TASK_SUMMARY:\s*(.+)$", stdout, re.MULTILINE) - return { - "completion": status.group(1) if status else None, - "verification": verification.group(1) if verification else None, - "summary": summary.group(1).strip() if summary else None, - } - -def run_verification_commands(workdir: Path, commands: list[str]) -> list[tuple[str, int, str]]: - """Execute each verification command and gather output.""" - results = [] - for cmd in commands: - print(f"\n\033[1;34m[Verification] Running command:\033[0m {cmd}") - argv = shlex.split(cmd) - proc = subprocess.run( - argv, - cwd=workdir, - capture_output=True, - text=True, - check=False - ) - combined_output = f"--- STDOUT ---\n{proc.stdout or ''}\n--- STDERR ---\n{proc.stderr or ''}" - results.append((cmd, proc.returncode, combined_output)) - if proc.returncode == 0: - print("\033[1;32m[Verification] Command passed.\033[0m") - else: - print(f"\033[1;31m[Verification] Command failed with exit code {proc.returncode}.\033[0m") - return results - -def run_aider( - workdir: Path, - profile_name: str, - profile: dict[str, any], - prompt_path: Path, - raw_dir: Path, - attempt: int, - editable_files: list[str], - read_only_files: list[str] -) -> tuple[int, str]: - """Execute Aider with standard parameters, streaming output to console.""" - cmd = [ - "aider", - "--model", profile["model"], - "--edit-format", profile["edit_format"], - "--map-tokens", str(profile["map_tokens"]), - "--no-auto-commits", - "--no-dirty-commits", - "--no-auto-lint", - "--no-auto-test", - "--no-restore-chat-history", - "--no-suggest-shell-commands", - "--no-analytics", - "--no-pretty", - "--no-stream", - "--yes-always", - "--message-file", str(prompt_path), - "--chat-history-file", str(raw_dir / f"chat-attempt-{attempt}.md"), - "--input-history-file", str(raw_dir / f"input-attempt-{attempt}"), - "--llm-history-file", str(raw_dir / f"llm-attempt-{attempt}.log"), - ] - - # Use absolute paths for target files to make Aider editing unambiguous - for f in editable_files: - cmd.extend(["--file", str(workdir / f)]) - for f in read_only_files: - cmd.extend(["--read", str(workdir / f)]) - - print(f"\n\033[1;34m[Aider] Starting Aider execution (Attempt {attempt})...\033[0m") - print(f"Command: {' '.join(shlex.quote(arg) for arg in cmd)}") - - proc = subprocess.Popen( - cmd, - cwd=workdir, - stdout=subprocess.PIPE, - stderr=subprocess.STDOUT, - text=True, - bufsize=1 - ) - - stdout_lines = [] - assert proc.stdout is not None - for line in proc.stdout: - sys.stdout.write(line) - sys.stdout.flush() - stdout_lines.append(line) - proc.wait() - - return proc.returncode, "".join(stdout_lines) - -def main() -> int: - parser = argparse.ArgumentParser(description="Delegate and execute small coding tasks using local coder agents.") - parser.add_argument("--title", required=True, help="Title of the task") - parser.add_argument("--goal", required=True, help="Goal/outcome description") - parser.add_argument("--implementation", required=True, help="Required implementation details") - parser.add_argument("--editable", required=True, nargs="+", help="Editable file path(s) relative to project root") - parser.add_argument("--read-only", nargs="+", default=[], help="Read-only context file path(s) relative to project root") - parser.add_argument("--verify", required=True, nargs="+", help="Deterministic verification command(s)") - parser.add_argument("--project", help="Project slug (e.g. content-engine). Automatically detected if omitted") - parser.add_argument("--task-id", help="Task ID. Autogenerated if omitted") - parser.add_argument("--profile", default="local-qwen25-coder-14b", choices=list(PROFILES.keys()), help="Model profile to use") - parser.add_argument("--task-category", help="Harness task category override (see .context/harness-routing.json)") - parser.add_argument("--run", action="store_true", help="Immediately run and verify the task packet on active worktree") - - args = parser.parse_args() - profile_overridden = any( - arg == "--profile" or arg.startswith("--profile=") - for arg in sys.argv[1:] - ) - - # 1. Project Detection - project_slug, root_dir = detect_project(Path.cwd()) - if args.project: - project_slug = args.project - - project_dir = root_dir / "30_projects" / project_slug - if not project_dir.is_dir(): - print(f"\033[1;31mError: Project folder '{project_dir}' does not exist.\033[0m") - return 1 - - routing_policy_path = root_dir / ".context" / "harness-routing.json" - routing_metadata = { - "task_kind": "code", - "editable_files": args.editable, - "create_files": [], - } - routing_hints = {} - if routing_policy_path.is_file(): - try: - task_packet_proc = subprocess.run( - [ - sys.executable, - str(root_dir / "bin" / "task-packet"), - "routing-hints", - "--policy", - str(routing_policy_path), - "--metadata-json", - json.dumps(routing_metadata), - ], - capture_output=True, - text=True, - check=False, - ) - if task_packet_proc.returncode == 0 and task_packet_proc.stdout.strip(): - routing_hints = json.loads(task_packet_proc.stdout) - elif task_packet_proc.returncode != 0: - stderr = task_packet_proc.stderr.strip() - print( - "\033[1;33mWarning: routing-hints failed " - f"(exit {task_packet_proc.returncode})" - + (f": {stderr}" if stderr else "") - + ".\033[0m" - ) - except (json.JSONDecodeError, OSError) as exc: - print(f"\033[1;33mWarning: routing-hints failed: {exc}.\033[0m") - routing_hints = {} - - if args.task_category: - inferred_category = routing_hints.get("task_category") - if inferred_category and inferred_category != args.task_category: - print( - "\033[1;33mWarning: --task-category overrides inferred " - f"'{inferred_category}' for this packet shape.\033[0m" - ) - routing_metadata["task_category"] = args.task_category - try: - override_proc = subprocess.run( - [ - sys.executable, - str(root_dir / "bin" / "task-packet"), - "routing-hints", - "--policy", - str(routing_policy_path), - "--metadata-json", - json.dumps(routing_metadata), - ], - capture_output=True, - text=True, - check=False, - ) - if override_proc.returncode == 0 and override_proc.stdout.strip(): - routing_hints = json.loads(override_proc.stdout) - else: - routing_hints["task_category"] = args.task_category - except (json.JSONDecodeError, OSError): - routing_hints["task_category"] = args.task_category - - recommended_profile = routing_hints.get("agent_profile") - if ( - recommended_profile - and recommended_profile in PROFILES - and not profile_overridden - and args.profile != recommended_profile - ): - print( - f"\033[1;33mRouting: {routing_hints.get('task_category', 'unknown')} " - f"→ profile '{recommended_profile}' " - f"(harness {routing_hints.get('harness_recommendation', 'n/a')}).\033[0m" - ) - args.profile = recommended_profile - elif len(args.editable) > 1 and args.profile != "local-qwen25-coder-14b": - print( - "\033[1;33mWarning: Coordinated multi-file changes are restricted to " - "'local-qwen25-coder-14b'. Forcing profile to 'local-qwen25-coder-14b'.\033[0m" - ) - args.profile = "local-qwen25-coder-14b" - elif args.profile == "local-qwen25-coder-14b": - # Check context limits warning - read_only_lines = 0 - for f in args.read_only: - f_path = project_dir / f - if f_path.is_file(): - with open(f_path, "rb") as fh: - read_only_lines += sum(1 for _ in fh) - if read_only_lines > 100: - print(f"\033[1;33mWarning: Read-only context size is {read_only_lines} lines. Qwen2.5-coder is sensitive to contexts > 100 lines. Consider trimming context.\033[0m") - - # 2. Task ID Setup - task_id = args.task_id - if not task_id: - timestamp = datetime.now(UTC).strftime("%Y%m%dt%H%M%Sz") - task_id = f"coder-task-{timestamp}" - - print(f"\033[1;32mConfiguring task packet: {task_id} under project: {project_slug}\033[0m") - - # 3. Create Task Packet File - packet_dir = project_dir / "plans" / "task-packets" - packet_dir.mkdir(parents=True, exist_ok=True) - packet_path = packet_dir / f"{task_id}.md" - - packet_metadata = { - "packet_version": 1, - "task_id": task_id, - "project_slug": project_slug, - "title": args.title, - "status": "ready", - "task_kind": "code", - "agent_profile": args.profile, - "workdir": ".", - "timeout_seconds": 900, - "editable_files": args.editable, - "create_files": [], - "read_only_files": args.read_only, - "verification_commands": args.verify, - "mindgraph_mode": "off", - "knowledge_queries": [], - } - for key in ( - "task_category", - "executor", - "harness_recommendation", - "allow_fusion_plan", - "needs_deliberation", - ): - if key in routing_hints: - packet_metadata[key] = routing_hints[key] - - packet_sections = { - "Goal": args.goal, - "Context and plan references": "Delegated via bin/delegate-coder.", - "Required implementation": args.implementation, - "Non-negotiable boundaries": "Do not commit. Do not claim tests passed.", - "Acceptance criteria": f"Verification commands must exit with 0:\n" + "\n".join(f"- {c}" for c in args.verify), - "Stop conditions": "Estimated context exceeds limits, edit failures, or unexpected exceptions.", - "Expected handoff": "Concise summary of changes and verification outcome.", - } - - # Format markdown - yaml_lines = ["---"] - for k, v in packet_metadata.items(): - yaml_lines.append(f"{k}: {json.dumps(v)}") - yaml_lines.append("---") - - md_content = "\n".join(yaml_lines) + "\n\n# Task Packet\n\n" - for name, content in packet_sections.items(): - md_content += f"## {name}\n\n{content}\n\n" - - packet_path.write_text(md_content, encoding="utf-8") - print(f"Written packet to {packet_path}") - - # 4. Validate and Compile Packet - print("\n\033[1;34m[Validate & Compile] Running validation...\033[0m") - val_proc = subprocess.run( - [sys.executable, str(root_dir / "bin" / "task-packet"), "validate", str(packet_path), "--require-ready"], - check=False - ) - if val_proc.returncode != 0: - print("\033[1;31mError: Task packet validation failed.\033[0m") - return 1 - - comp_proc = subprocess.run( - [sys.executable, str(root_dir / "bin" / "task-packet"), "compile"], - check=False - ) - if comp_proc.returncode != 0: - print("\033[1;31mError: Manifest compilation failed.\033[0m") - return 1 - - print("\033[1;32mPacket successfully validated and compiled into manifest.\033[0m") - - if not args.run: - print("\033[1;32mDelegation setup complete. To run it, run again with --run flag, or use the harness eval workbench.\033[0m") - return 0 - - # 5. Live Execution Loop (if --run set) - print("\n\033[1;34m[Execution] Starting execution flow...\033[0m") - - # Git Clean check - if check_git_status(root_dir): - print("\033[1;33mWarning: The working tree has uncommitted changes.\033[0m") - if sys.stdin.isatty(): - ans = input("Do you want to proceed with the run anyway? [y/N]: ").strip().lower() - if ans not in ("y", "yes"): - print("Aborting.") - return 0 - else: - print("Non-interactive terminal: Proceeding with warnings.") - - # Health check - profile = PROFILES[args.profile] - healthy, msg = health_check(args.profile, profile) - if not healthy: - print(f"\033[1;31mHealth Check Failed: {msg}\033[0m") - return 1 - - # Setup runs folder - run_id = datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ") + f"-{task_id}" - raw_dir = project_dir / "raw-materials" / "runs" / run_id - raw_dir.mkdir(parents=True, exist_ok=True) - - # Write first attempt prompt - prompt_file = raw_dir / "prompt-attempt-1.md" - prompt_file.write_text(render_prompt({"metadata": packet_metadata, "sections": packet_sections}), encoding="utf-8") - - # RUN ATTEMPT 1 - aider_code, stdout = run_aider( - workdir=project_dir, - profile_name=args.profile, - profile=profile, - prompt_path=prompt_file, - raw_dir=raw_dir, - attempt=1, - editable_files=args.editable, - read_only_files=args.read_only - ) - - # Save stdout/stderr - (raw_dir / "agent-attempt-1.stdout").write_text(stdout, encoding="utf-8") - - # Parse claims - claims = parse_claims(stdout) - print(f"\n\033[1;34m[Analysis] Aider attempt 1 finished. Parsing claims...\033[0m") - print(f"TASK_STATUS claim: {claims['completion']}") - print(f"TASK_SUMMARY claim: {claims['summary']}") - - # Run Verifiers - verifications = run_verification_commands(project_dir, args.verify) - passed = all(res[1] == 0 for res in verifications) - - # Save verifier logs - for idx, (cmd, rc, output) in enumerate(verifications, start=1): - (raw_dir / f"verify-1-{idx}.stdout").write_text(output, encoding="utf-8") - - if passed: - print("\n\033[1;32m[Success] Task completed and all verification commands passed successfully!\033[0m") - return 0 - - # RUN ATTEMPT 2 (REPAIR LOOP) - print("\n\033[1;33m[Repair] Verification failed. Initiating H2-repair loop...\033[0m") - failures = [(cmd, rc, output) for cmd, rc, output in verifications if rc != 0] - current_diff = get_git_diff(project_dir, args.editable) - - repair_summary = [] - for cmd, rc, output in failures: - repair_summary.append(f"Command: {json.dumps(cmd)}\nExit: {rc}\nOutput:\n{output[:4000]}") - repair_ctx_str = "\n\n".join(repair_summary) + "\n\nCurrent diff:\n" + current_diff[:8000] - - repair_prompt = render_prompt({"metadata": packet_metadata, "sections": packet_sections}, repair_context=repair_ctx_str) - prompt_file_2 = raw_dir / "prompt-attempt-2.md" - prompt_file_2.write_text(repair_prompt, encoding="utf-8") - - aider_code_2, stdout_2 = run_aider( - workdir=project_dir, - profile_name=args.profile, - profile=profile, - prompt_path=prompt_file_2, - raw_dir=raw_dir, - attempt=2, - editable_files=args.editable, - read_only_files=args.read_only - ) - - # Save repair logs - (raw_dir / "agent-attempt-2.stdout").write_text(stdout_2, encoding="utf-8") - - # Re-run verifications - verifications_2 = run_verification_commands(project_dir, args.verify) - for idx, (cmd, rc, output) in enumerate(verifications_2, start=1): - (raw_dir / f"verify-2-{idx}.stdout").write_text(output, encoding="utf-8") - - passed_2 = all(res[1] == 0 for res in verifications_2) - if passed_2: - print("\n\033[1;32m[Success] Task completed and verified successfully after repair attempt!\033[0m") - return 0 - else: - print("\n\033[1;31m[Failure] Task failed verification even after repair attempt. Inspect files and diff.\033[0m") - return 1 - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/eval-schedule b/bin/eval-schedule deleted file mode 100755 index 4e4b7fb..0000000 --- a/bin/eval-schedule +++ /dev/null @@ -1,4 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail -ROOT="$(cd "$(dirname "$0")/.." && pwd)" -exec python3 "$ROOT/scripts/eval_schedule.py" "$@" \ No newline at end of file diff --git a/bin/extract-knowledge b/bin/extract-knowledge deleted file mode 100755 index dff3104..0000000 --- a/bin/extract-knowledge +++ /dev/null @@ -1,243 +0,0 @@ -#!/usr/bin/env python3 -"""Validate prerequisites and scaffold a knowledge note from a project.""" - -from __future__ import annotations - -import argparse -import json -import re -import sys -from dataclasses import dataclass, field -from datetime import date -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -PROJECTS_DIR = ROOT / "30_projects" -KNOWLEDGE_DIR = ROOT / "10_knowledge" -SAFE_SEGMENT = re.compile(r"^[a-z0-9][a-z0-9._-]*$") - - -@dataclass(frozen=True) -class Prerequisite: - name: str - path: str - exists: bool - required: bool - message: str - - -@dataclass -class ExtractResult: - prerequisites: list[Prerequisite] = field(default_factory=list) - target_path: str | None = None - input_error: str | None = None - collision: bool = False - written: bool = False - - @property - def ok(self) -> bool: - if self.input_error or self.collision: - return False - return all(p.exists for p in self.prerequisites if p.required) - - -def slugify(title: str) -> str: - slug = title.lower().strip() - slug = re.sub(r"[^a-z0-9]+", "-", slug) - return slug.strip("-") - - -def scaffold_content(title: str, domain: str, project_slug: str, - tags: list[str]) -> str: - tag_str = json_tags(tags) - return ( - f"---\n" - f"title: {json.dumps(title)}\n" - f"domain: {json.dumps(domain)}\n" - f'type: "note"\n' - f'status: "queued"\n' - f"source: {json.dumps(f'30_projects/{project_slug}/README.md')}\n" - f"tags: {tag_str}\n" - f"---\n" - f"\n" - f"# {title}\n" - f"\n" - ) - - -def json_tags(tags: list[str]) -> str: - return json.dumps(tags) - - -def _contained_directory(base: Path, candidate: Path) -> bool: - if not candidate.is_dir(): - return False - try: - candidate.resolve(strict=True).relative_to(base.resolve(strict=True)) - except (OSError, ValueError): - return False - return True - - -def _contained_file(base: Path, candidate: Path) -> bool: - if not candidate.is_file(): - return False - try: - candidate.resolve(strict=True).relative_to(base.resolve(strict=True)) - except (OSError, ValueError): - return False - return True - - -def check_prerequisites(root: Path, project_slug: str, - domain: str) -> list[Prerequisite]: - prereqs: list[Prerequisite] = [] - - projects_dir = root / "30_projects" - proj_dir = projects_dir / project_slug - project_exists = _contained_directory(projects_dir, proj_dir) - prereqs.append(Prerequisite( - name="project directory", - path=f"30_projects/{project_slug}", - exists=project_exists, - required=True, - message="project directory must exist inside 30_projects/", - )) - - readme = proj_dir / "README.md" - prereqs.append(Prerequisite( - name="project README", - path=f"30_projects/{project_slug}/README.md", - exists=project_exists and _contained_file(proj_dir, readme), - required=True, - message="project README is the extraction source", - )) - - log = proj_dir / "log.md" - prereqs.append(Prerequisite( - name="project log", - path=f"30_projects/{project_slug}/log.md", - exists=project_exists and _contained_file(proj_dir, log), - required=False, - message="log.md is optional but useful for extraction", - )) - - decisions = proj_dir / "decisions.md" - prereqs.append(Prerequisite( - name="project decisions", - path=f"30_projects/{project_slug}/decisions.md", - exists=project_exists and _contained_file(proj_dir, decisions), - required=False, - message="decisions.md is optional but useful for extraction", - )) - - knowledge_dir = root / "10_knowledge" - domain_dir = knowledge_dir / domain - prereqs.append(Prerequisite( - name="knowledge domain", - path=f"10_knowledge/{domain}", - exists=_contained_directory(knowledge_dir, domain_dir), - required=True, - message="target domain must exist inside 10_knowledge/", - )) - - return prereqs - - -def run_extract(root: Path, project_slug: str, domain: str, title: str, - tags: list[str], write: bool = False) -> ExtractResult: - result = ExtractResult() - if not SAFE_SEGMENT.fullmatch(project_slug): - result.input_error = ( - "project slug must be one path segment using lowercase letters, " - "digits, dots, dashes, or underscores" - ) - return result - if not SAFE_SEGMENT.fullmatch(domain): - result.input_error = ( - "domain must be one path segment using lowercase letters, digits, " - "dots, dashes, or underscores" - ) - return result - if not title.strip() or not slugify(title) or any( - character in title for character in ("\x00", "\n", "\r") - ): - result.input_error = "title must be a non-empty single-line title" - return result - result.prerequisites = check_prerequisites(root, project_slug, domain) - - filename = f"{date.today().isoformat()}__{domain}__note__{slugify(title)}.md" - target = root / "10_knowledge" / domain / filename - result.target_path = f"10_knowledge/{domain}/{filename}" - - if target.exists() or target.is_symlink(): - result.collision = True - - if not result.ok: - return result - - if write: - content = scaffold_content(title, domain, project_slug, tags) - try: - with target.open("x", encoding="utf-8") as handle: - handle.write(content) - except FileExistsError: - result.collision = True - else: - result.written = True - - return result - - -def print_result(result: ExtractResult) -> None: - for p in result.prerequisites: - status = "OK" if p.exists else ("MISSING" if p.required else "WARN") - print(f" [{status:>7}] {p.path} — {p.message}") - - print() - if result.input_error: - print(f" ERROR: {result.input_error}") - if result.target_path: - print(f" target: {result.target_path}") - if result.collision: - print(" ERROR: target file already exists; will not overwrite") - if result.written: - print(" WROTE scaffold — fill in knowledge content, then run bin/mindgraph-refresh") - elif result.ok and not result.written: - print(" CHECK passed — run with --write to create the scaffold") - elif not result.ok: - print(" CHECK failed — resolve required prerequisites first") - - -def main() -> int: - parser = argparse.ArgumentParser( - description="Scaffold a knowledge note extracted from a project." - ) - parser.add_argument("--project", required=True, - help="project slug under 30_projects/") - parser.add_argument("--domain", required=True, - help="target domain under 10_knowledge/") - parser.add_argument("--title", required=True, - help="title for the new knowledge note") - mode = parser.add_mutually_exclusive_group() - mode.add_argument("--check", action="store_true", - help="validate prerequisites only (default)") - mode.add_argument("--write", action="store_true", - help="create the scaffolded note file") - parser.add_argument("--tags", default="extracted", - help="comma-separated tags (default: extracted)") - args = parser.parse_args() - - tags = [t.strip() for t in args.tags.split(",") if t.strip()] - write = args.write - - result = run_extract(ROOT, args.project, args.domain, args.title, - tags, write=write) - print_result(result) - - return 0 if result.ok else 1 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/bin/fetch-source-text b/bin/fetch-source-text deleted file mode 100755 index cb59558..0000000 --- a/bin/fetch-source-text +++ /dev/null @@ -1,10 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -if ROOT="$(git rev-parse --show-toplevel 2>/dev/null)"; then - : -else - ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -fi - -exec python3 "$ROOT/scripts/fetch_source_text.py" "$@" \ No newline at end of file diff --git a/bin/generate-project-tasks b/bin/generate-project-tasks deleted file mode 100755 index 549da62..0000000 --- a/bin/generate-project-tasks +++ /dev/null @@ -1,522 +0,0 @@ -#!/usr/bin/env python3 -"""Crawl project directories for task lists and compile them into a unified manifest. - -Unit 2.3: the compiled manifest is a **quarantined projection**. It is not -reviewed executable task authority. Ready packets remain authoritative only via -their Markdown contracts / task_packets_manifest, not via this file's rows. -""" - -from __future__ import annotations - -import argparse -import json -import re -import sys -from datetime import datetime, timezone -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] -PROJECTS_DIR = ROOT / "30_projects" -MANIFEST_PATH = PROJECTS_DIR / "tasks_manifest.json" -PACKET_MANIFEST_PATH = PROJECTS_DIR / "task_packets_manifest.json" -QUARANTINE_BASELINE_PATH = PROJECTS_DIR / "tasks_manifest.quarantine-baseline.json" - -GENERATOR_VERSION = "2.3.0-quarantine" - -# Only explicit checkbox lines become tasks. Decision-log bullets (promoted -# ADRs and the like) are records, not work items. -MAX_TASK_TITLE_LENGTH = 200 -SKIP_TITLE_PREFIXES = ("**Promoted", "ADR-") - - -def normalize_task_title(raw_title: str) -> str | None: - title = raw_title.strip() - if not title or title.startswith(SKIP_TITLE_PREFIXES): - return None - if len(title) > MAX_TASK_TITLE_LENGTH: - title = title[: MAX_TASK_TITLE_LENGTH - 1].rstrip() + "…" - return title - - -def parse_frontmatter(text: str) -> dict[str, str]: - if not text.startswith("---\n"): - return {} - end = text.find("\n---", 4) - if end == -1: - return {} - metadata = {} - for line in text[4:end].splitlines(): - if ":" not in line: - continue - key, value = line.split(":", 1) - value = value.strip().strip('"').strip("'") - metadata[key.strip()] = value - return metadata - - -def extract_readme_tasks(readme_path: Path, project_slug: str) -> list[dict]: - if not readme_path.exists(): - return [] - - text = readme_path.read_text(encoding="utf-8") - lines = text.splitlines() - - tasks = [] - in_tasks_section = False - task_counter = 1 - - for line in lines: - stripped = line.strip() - # Toggle section - if stripped.startswith("## "): - header = stripped[3:].strip().lower() - if header in ("tasks", "todo", "todo list", "checklist", "roadmap"): - in_tasks_section = True - else: - in_tasks_section = False - elif stripped.startswith("# "): - in_tasks_section = False - - if in_tasks_section: - match = re.match(r"^[-*]\s+\[([ x/])\]\s+(.+)$", stripped) - if match: - state_char = match.group(1) - title = normalize_task_title(match.group(2)) - if title is None: - continue - - # Deduce status - if state_char == "x": - status = "done" - elif state_char == "/": - status = "active" - else: - status = "active" - - tasks.append( - { - "id": f"task-{project_slug}-readme-{task_counter}", - "title": title, - "status": status, - "priority": "normal", - "owner_agent_id": "agent-local-coder", - "phase": "README Tasks", - "next_action": title, - "blocker": "", - "unresolved_decisions": [], - "mainframe_target": f"30_projects/{project_slug}/README.md", - "evidence": [], - "updated": "today", - } - ) - task_counter += 1 - - return tasks - - -def packet_tasks(packet_manifest_path: Path = PACKET_MANIFEST_PATH) -> list[dict]: - if not packet_manifest_path.exists(): - return [] - payload = json.loads(packet_manifest_path.read_text(encoding="utf-8")) - tasks = [] - for packet in payload.get("packets", []): - status = packet.get("status", "draft") - if status == "retired": - task_status = "done" - elif status == "ready": - task_status = "active" - else: - task_status = "review" - tasks.append( - { - "id": packet["id"], - "title": packet["title"], - "status": task_status, - "priority": "normal", - "owner_agent_id": "agent-local-coder", - "phase": f"Delegation packet ({status})", - "goal": packet.get("goal", ""), - "next_action": packet.get("next_action", ""), - "blocker": "", - "unresolved_decisions": [], - "mainframe_target": packet["packet_path"], - "evidence": [ - { - "label": "Delegation packet contract", - "session_id": "task-packet", - "event_id": packet["task_id"], - } - ], - "updated": "generated", - "task_kind": packet.get("task_kind"), - "agent_profile": packet.get("agent_profile"), - "packet_path": packet["packet_path"], - "packet_status": status, - "authority_class": "packet_candidate", - "executable": False, - } - ) - return tasks - - -def extract_phase_tasks(plan_path: Path, project_slug: str) -> list[dict]: - try: - content = plan_path.read_text(encoding="utf-8") - except Exception as e: - print(f"Error reading {plan_path}: {e}", file=sys.stderr) - return [] - - # Try to parse frontmatter or top metadata lines - metadata = {} - if content.startswith("---\n") or content.startswith("---\r\n"): - metadata = parse_frontmatter(content) - else: - # Fallback to simple colon matching in the first 15 lines - for line in content.splitlines()[:15]: - match = re.match(r"^([a-zA-Z0-9_-]+):\s*(.+)$", line.strip()) - if match: - key = match.group(1).lower() - val = match.group(2).strip().strip('"').strip("'") - metadata[key] = val - match_bold = re.match(r"^\*\*([a-zA-Z0-9_-]+):\*\*\s*(.+)$", line.strip()) - if match_bold: - key = match_bold.group(1).lower() - val = match_bold.group(2).strip().strip('"').strip("'") - metadata[key] = val - - # Check plan status - status_raw = metadata.get("status", "active").lower() - phase_status = "active" - if any(k in status_raw for k in ("complete", "shipped", "done", "retired")): - phase_status = "complete" - elif any(k in status_raw for k in ("planned", "draft", "backlog", "not started")): - phase_status = "planned" - - # H1 title - phase_title = plan_path.stem - for line in content.splitlines(): - h1_match = re.match(r"^#\s+(.+)$", line) - if h1_match: - phase_title = h1_match.group(1).strip() - break - - # Extract goal - goal = "" - goal_match = re.search(r"## Goal\s+(.+?)(?=\n##|\Z)", content, re.DOTALL) - if goal_match: - goal = goal_match.group(1).strip().replace("\n", " ") - if len(goal) > 200: - goal = goal[:197] + "..." - - # Parse lines - lines = content.splitlines() - tasks = [] - current_section = "" - current_unit = "" - unit_tasks_buffer = [] - - # Heuristic list of headers that contain tasks - task_header_keywords = ("unit", "step", "milestone", "implementation", "task", "checklist", "order", "run") - exclude_header_keywords = ("reference", "boundary", "starting point", "measured", "purpose", "non-negotiable") - - def is_task_section(sec_name): - sec_lower = sec_name.lower() - if any(kw in sec_lower for kw in exclude_header_keywords): - return False - if any(kw in sec_lower for kw in task_header_keywords): - return True - return False - - for idx, line in enumerate(lines): - line_num = idx + 1 - stripped = line.strip() - - # Detect headers - if stripped.startswith("## "): - current_section = stripped[3:].strip() - current_unit = "" - elif stripped.startswith("### "): - current_unit = stripped[4:].strip() - elif stripped.startswith("- **M") or stripped.startswith("* **M"): - m_match = re.match(r"^[-*]\s+\*\*([^*]+)\*\*(.*)$", stripped) - if m_match: - current_unit = m_match.group(1).strip() + m_match.group(2).strip() - - # Parse tasks in valid sections - sec_context = current_unit or current_section - if sec_context and is_task_section(sec_context): - # Check for checkbox - cb_match = re.match(r"^[-*]\s+\[([ x/])\]\s+(.+)$", stripped) - if cb_match: - state_char = cb_match.group(1) - task_title = normalize_task_title(cb_match.group(2)) - if task_title is None: - continue - - prefix = f"[{current_unit}] " if current_unit else "" - full_title = f"{prefix}{task_title}" - # Unit prefixes can push a capped title past the serving - # gate (240); truncate rather than lose the task downstream. - if len(full_title) > 240: - full_title = full_title[:239].rstrip() + "…" - - task_state = "active" - if state_char == "x": - task_state = "done" - elif state_char == "/": - task_state = "active" - else: - if phase_status == "complete": - task_state = "done" - elif phase_status == "planned": - task_state = "backlog" - else: - task_state = "active" - - unit_tasks_buffer.append({ - "line_num": line_num, - "title": full_title, - "status": task_state, - "unit": current_unit or current_section, - "is_checkbox": True - }) - # Plain bullets are intentionally not captured: plan sections mix - # work items with decision-log and status notes, and only the - # operator's explicit checkboxes are commitments. - - # Refine status of active items using the unit order heuristic - if phase_status == "active" and unit_tasks_buffer: - units_ordered = [] - unit_to_tasks = {} - for t in unit_tasks_buffer: - u = t["unit"] - if u not in unit_to_tasks: - units_ordered.append(u) - unit_to_tasks[u] = [] - unit_to_tasks[u].append(t) - - active_unit_idx = -1 - for i, u in enumerate(units_ordered): - has_active = any(t["status"] == "active" for t in unit_to_tasks[u]) - if has_active: - active_unit_idx = i - break - - if active_unit_idx != -1: - for i in range(active_unit_idx + 1, len(units_ordered)): - for t in unit_to_tasks[units_ordered[i]]: - if t["status"] == "active": - t["status"] = "backlog" - - # Construct final task records - target_rel_prefix = f"30_projects/{project_slug}/plans" - if project_slug == "local-agent": - target_rel_prefix = "30_projects/local-agent/workbench/planning" - - for t in unit_tasks_buffer: - plan_rel_path = f"{target_rel_prefix}/{plan_path.name}" - tasks.append({ - "id": f"task-{project_slug}-{plan_path.stem}-l{t['line_num']}", - "title": t["title"], - "status": t["status"], - "priority": metadata.get("priority", "normal"), - "owner_agent_id": "agent-local-coder", - "phase": phase_title, - "goal": goal, - "next_action": t["title"], - "blocker": "", - "unresolved_decisions": [], - "mainframe_target": f"{plan_rel_path}#L{t['line_num']}", - "evidence": [], - "updated": metadata.get("last_updated") or metadata.get("updated") or "today" - }) - - return tasks - - -def _tag_plan_item(task: dict) -> dict: - task = dict(task) - task.setdefault("authority_class", "plan_item") - task["executable"] = False - return task - - -def compile_tasks( - projects_dir: Path = PROJECTS_DIR, - packet_manifest_path: Path = PACKET_MANIFEST_PATH, -) -> list[dict]: - compiled_tasks = [] - - # Scan projects - for p_dir in sorted(projects_dir.iterdir()): - if not p_dir.is_dir(): - continue - if p_dir.name.startswith("."): - continue - - readme_path = p_dir / "README.md" - project_tasks = [] - - # 1. Parse README metadata for defaults - metadata = {} - project_title = p_dir.name - if readme_path.exists(): - try: - readme_text = readme_path.read_text(encoding="utf-8") - metadata = parse_frontmatter(readme_text) - project_title = metadata.get("title") or project_title - except Exception as e: - print(f"Error reading README for {p_dir.name}: {e}", file=sys.stderr) - - # 2. Extract tasks from plans/ and plans/phases/ (recursively) - plan_files = [] - plans_dir = p_dir / "plans" - if plans_dir.exists(): - plan_files.extend(plans_dir.glob("*.md")) - plan_files.extend(plans_dir.glob("phases/*.md")) - - # Special case: local-agent planning files - if p_dir.name == "local-agent": - la_planning = p_dir / "workbench" / "planning" - if la_planning.exists(): - plan_files.extend(la_planning.glob("*.md")) - plan_files.extend(la_planning.glob("phases/*.md")) - - for plan_file in plan_files: - if plan_file.name == "README.md": - continue - plan_tasks = extract_phase_tasks(plan_file, p_dir.name) - project_tasks.extend(plan_tasks) - - # 3. Fallback to README metadata if no tasks extracted from plan files - if not project_tasks: - status = metadata.get("status") or metadata.get("project_state") or "active" - status = status.lower() - if status == "shipped": - status = "done" - if status not in ("active", "review", "blocked", "done", "backlog"): - status = "active" - project_tasks.append( - { - "id": f"project-{p_dir.name}", - "title": project_title, - "status": status, - "priority": metadata.get("priority", "normal"), - "owner_agent_id": "agent-local-coder", - "phase": metadata.get("project_state", "Active Project"), - "next_action": metadata.get("next_action", "Explore next steps."), - "blocker": metadata.get("blocker", ""), - "unresolved_decisions": metadata.get("unresolved_decisions", []), - "mainframe_target": f"30_projects/{p_dir.name}/README.md", - "evidence": [], - "updated": metadata.get("updated") or "unknown", - } - ) - - compiled_tasks.extend(_tag_plan_item(t) for t in project_tasks) - - compiled_tasks.extend(packet_tasks(packet_manifest_path)) - return compiled_tasks - - -def build_quarantine_envelope(tasks: list[dict]) -> dict: - now = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") - baseline_sha = None - if QUARANTINE_BASELINE_PATH.exists(): - try: - baseline_sha = json.loads(QUARANTINE_BASELINE_PATH.read_text(encoding="utf-8")).get( - "baseline_sha256_pre_quarantine" - ) - except Exception: - baseline_sha = None - for t in tasks: - t.setdefault("authority_class", "plan_item") - t["executable"] = False - return { - "schema_version": 1, - "generator": GENERATOR_VERSION, - "generated_at": now, - "quarantine": { - "status": "quarantined", - "reason": ( - "Unit 2.3: checkbox/plan projection is not reviewed executable " - "task authority" - ), - "baseline_sha256": baseline_sha, - "task_count": len(tasks), - "as_of": now, - "executable_authority": ( - "reviewed ready task packets only " - "(plans/task-packets + task_packets_manifest)" - ), - "plan_item_vs_executable": ( - "plan_item=checkbox/plan extract; packet_candidate=packet row; " - "none executable while quarantine.status=quarantined" - ), - "unit": "2.3", - }, - "tasks": tasks, - } - - -def is_quarantined_manifest(payload: dict) -> bool: - q = payload.get("quarantine") if isinstance(payload, dict) else None - return isinstance(q, dict) and q.get("status") == "quarantined" - - -def main(argv: list[str] | None = None) -> int: - parser = argparse.ArgumentParser( - description=( - "Compile project checkboxes into a quarantined tasks_manifest " - "(not executable authority)." - ) - ) - parser.add_argument( - "--check", - action="store_true", - help="Verify existing manifest is quarantined; do not write", - ) - parser.add_argument( - "--dry-run", - action="store_true", - help="Compile and print counts without writing", - ) - args = parser.parse_args(argv) - - if args.check: - if not MANIFEST_PATH.exists(): - print("error: tasks_manifest.json missing", file=sys.stderr) - return 1 - payload = json.loads(MANIFEST_PATH.read_text(encoding="utf-8")) - if not is_quarantined_manifest(payload): - print( - "error: tasks_manifest.json is not quarantined " - "(refuse treating as executable)", - file=sys.stderr, - ) - return 1 - n = len(payload.get("tasks") or []) - print(f"ok: quarantined projection with {n} tasks") - return 0 - - compiled_tasks = compile_tasks() - envelope = build_quarantine_envelope(compiled_tasks) - - if args.dry_run: - print( - f"dry-run: would write {len(compiled_tasks)} quarantined tasks " - f"to {MANIFEST_PATH.relative_to(ROOT)}" - ) - return 0 - - MANIFEST_PATH.write_text(json.dumps(envelope, indent=2) + "\n", encoding="utf-8") - print( - f"Compiled {len(compiled_tasks)} quarantined plan_items/packet_candidates " - f"to {MANIFEST_PATH.relative_to(ROOT)} " - f"(not executable authority; generator {GENERATOR_VERSION})" - ) - return 0 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/ingest-status b/bin/ingest-status deleted file mode 100755 index 3b94789..0000000 --- a/bin/ingest-status +++ /dev/null @@ -1,297 +0,0 @@ -#!/usr/bin/env python3 -"""Report ingest lane ages and migration-batch dispositions without moving files.""" - -from __future__ import annotations - -import argparse -import csv -import hashlib -import json -from collections import Counter -from datetime import date, datetime -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -BATCHES_REL = Path("30_projects/second-brain-migration/raw-materials/batches") -LANES = ( - "00_inbox", - "01_ingest/ready", - "01_ingest/queue", - "01_ingest/rejected", -) -UNRESOLVED = "unresolved" - - -def sha256_of(path: Path) -> str: - digest = hashlib.sha256() - with path.open("rb") as handle: - for chunk in iter(lambda: handle.read(65536), b""): - digest.update(chunk) - return digest.hexdigest() - - -def visible_files(directory: Path) -> list[Path]: - """Flat scan of one lane, matching the ingest minion's visibility rule.""" - if not directory.is_dir(): - return [] - return [ - path - for path in sorted(directory.iterdir(), key=lambda item: item.name) - if path.is_file() and not path.name.startswith(".") - ] - - -def visible_dirs(directory: Path) -> list[Path]: - if not directory.is_dir(): - return [] - return [ - path - for path in sorted(directory.iterdir(), key=lambda item: item.name) - if path.is_dir() and not path.name.startswith(".") - ] - - -def file_age_days(path: Path, today: date) -> int: - modified = datetime.fromtimestamp(path.stat().st_mtime).date() - return max((today - modified).days, 0) - - -def scan_lane(root: Path, lane: str, today: date) -> dict[str, object]: - directory = root / lane - files = visible_files(directory) - ages = [file_age_days(path, today) for path in files] - oldest_age = max(ages) if ages else None - return { - "name": lane, - "exists": directory.is_dir(), - "files": len(files), - "directories": len(visible_dirs(directory)), - "age_days_0_7": sum(age <= 7 for age in ages), - "age_days_8_30": sum(7 < age <= 30 for age in ages), - "age_days_31_plus": sum(age > 30 for age in ages), - "oldest_file_age_days": oldest_age, - } - - -def load_batches(root: Path) -> dict[str, object]: - """Read every registered batch's manifest and append-only disposition ledger. - - The manifest is immutable registration evidence; the ledger appends later - per-file decisions, so the last ledger row for a path wins. - """ - batches_dir = root / BATCHES_REL - batches: list[dict[str, object]] = [] - hash_index: dict[str, str] = {} - name_index: dict[tuple[str, str], str] = {} - disposition_by_name: dict[tuple[str, str], str] = {} - disposition_by_hash: dict[str, str] = {} - invalid_dirs: list[str] = [] - - if not batches_dir.is_dir(): - return { - "batches": batches, - "hash_index": hash_index, - "name_index": name_index, - "disposition_by_name": disposition_by_name, - "disposition_by_hash": disposition_by_hash, - "invalid_dirs": invalid_dirs, - } - - for batch_dir in sorted(path for path in batches_dir.iterdir() if path.is_dir()): - manifest_path = batch_dir / "manifest.csv" - if not manifest_path.is_file(): - invalid_dirs.append(batch_dir.name) - continue - - batch_id = batch_dir.name - registered = 0 - with manifest_path.open(encoding="utf-8", newline="") as handle: - for row in csv.DictReader(handle): - relative_path = (row.get("relative_path") or "").strip() - file_hash = (row.get("sha256") or "").strip() - if not relative_path or not file_hash: - continue - registered += 1 - hash_index[file_hash] = batch_id - name_index[(batch_id, relative_path)] = file_hash - - ledger_rows = 0 - latest_in_batch: dict[str, str] = {} - ledger_path = batch_dir / "disposition-ledger.csv" - if ledger_path.is_file(): - with ledger_path.open(encoding="utf-8", newline="") as handle: - for row in csv.DictReader(handle): - relative_path = (row.get("relative_path") or "").strip() - disposition = (row.get("disposition") or "").strip() - file_hash = (row.get("sha256") or "").strip() - if not relative_path or not disposition: - continue - ledger_rows += 1 - latest_in_batch[relative_path] = disposition - disposition_by_name[(batch_id, relative_path)] = disposition - if file_hash: - disposition_by_hash[file_hash] = disposition - - dispositions = Counter(latest_in_batch.values()) - batches.append({ - "id": batch_id, - "registered_files": registered, - "ledger_rows": ledger_rows, - "has_ledger": ledger_path.is_file(), - "dispositions": [ - {"name": name, "count": count} - for name, count in dispositions.most_common() - ], - }) - - return { - "batches": batches, - "hash_index": hash_index, - "name_index": name_index, - "disposition_by_name": disposition_by_name, - "disposition_by_hash": disposition_by_hash, - "invalid_dirs": invalid_dirs, - } - - -def classify_inbox(root: Path, batch_data: dict[str, object]) -> dict[str, object]: - """Split 00_inbox into batch-registered migration files and organic captures. - - Matching is scoped to the inbox because registration is non-destructive - there; once the minion normalizes a file into 01_ingest the bytes (and so - the hash) legitimately drift from the registered snapshot. - """ - hash_index = batch_data["hash_index"] - name_index = batch_data["name_index"] - disposition_by_name = batch_data["disposition_by_name"] - disposition_by_hash = batch_data["disposition_by_hash"] - - registered = 0 - organic = 0 - matched_batches: set[str] = set() - dispositions: Counter[str] = Counter() - - for path in visible_files(root / "00_inbox"): - file_hash = sha256_of(path) - batch_id = hash_index.get(file_hash) - if batch_id is None: - organic += 1 - continue - registered += 1 - matched_batches.add(batch_id) - if (batch_id, path.name) in name_index: - disposition = disposition_by_name.get( - (batch_id, path.name), - disposition_by_hash.get(file_hash, UNRESOLVED), - ) - else: - disposition = disposition_by_hash.get(file_hash, UNRESOLVED) - dispositions[disposition] += 1 - - return { - "batch_registered": registered, - "organic": organic, - "batches_matched": sorted(matched_batches), - "registered_by_disposition": [ - {"name": name, "count": count} - for name, count in dispositions.most_common() - ], - } - - -def build_status(root: Path, today: date | None = None) -> dict[str, object]: - today = today or date.today() - batch_data = load_batches(root) - return { - "as_of": today.isoformat(), - "lanes": [scan_lane(root, lane, today) for lane in LANES], - "inbox_composition": classify_inbox(root, batch_data), - "batches": batch_data["batches"], - "invalid_batch_dirs": batch_data["invalid_dirs"], - } - - -def print_status(status: dict[str, object]) -> None: - print(f"Ingest status, as of {status['as_of']}") - print("") - print("Lanes (visible files, flat scan):") - for lane in status["lanes"]: - if not lane["exists"]: - print(f"- {lane['name']}: missing") - continue - line = ( - f"- {lane['name']}: {lane['files']} file(s); " - f"age <=7d: {lane['age_days_0_7']}, " - f"8-30d: {lane['age_days_8_30']}, " - f">30d: {lane['age_days_31_plus']}" - ) - if lane["oldest_file_age_days"] is not None: - line += f"; oldest {lane['oldest_file_age_days']}d" - if lane["directories"]: - line += f" (+{lane['directories']} unscanned subdirectories)" - print(line) - - composition = status["inbox_composition"] - print("") - print("Inbox composition:") - print( - f"- batch-registered migration files: {composition['batch_registered']}" - + ( - f" (batches: {', '.join(composition['batches_matched'])})" - if composition["batches_matched"] - else "" - ) - ) - print(f"- organic captures (no batch match): {composition['organic']}") - for item in composition["registered_by_disposition"]: - print(f" - {item['name']}: {item['count']}") - - if status["batches"]: - print("") - print("Migration batches (latest ledger row per file wins):") - for batch in status["batches"]: - dispositions = ", ".join( - f"{item['name']} {item['count']}" - for item in batch["dispositions"] - ) or "no ledger rows" - ledger_note = "" if batch["has_ledger"] else " [no ledger file]" - print( - f"- {batch['id']}: {batch['registered_files']} registered; " - f"{dispositions}{ledger_note}" - ) - - if status["invalid_batch_dirs"]: - print("") - print( - "Batch directories without a manifest (skipped): " - + ", ".join(status["invalid_batch_dirs"]) - ) - - print("") - print( - "Interpretation: batch-registered files are intentional migration " - "backlog, not stuck intake. Organic captures aging past 30 days and " - "ready-lane age measure the agent-enrichment middle. This report " - "moves nothing." - ) - - -def main() -> int: - parser = argparse.ArgumentParser() - parser.add_argument("--root", type=Path, default=ROOT) - parser.add_argument( - "--json", action="store_true", help="emit a structured summary" - ) - args = parser.parse_args() - - status = build_status(args.root) - if args.json: - print(json.dumps(status, indent=2)) - else: - print_status(status) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/bin/install-hooks b/bin/install-hooks deleted file mode 100755 index 032cf93..0000000 --- a/bin/install-hooks +++ /dev/null @@ -1,6 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -ROOT="$(git rev-parse --show-toplevel)" -git -C "$ROOT" config core.hooksPath .githooks -echo "Installed Mainframe hooks: core.hooksPath=.githooks" diff --git a/bin/knowledge-reconcile b/bin/knowledge-reconcile deleted file mode 100755 index e5bec37..0000000 --- a/bin/knowledge-reconcile +++ /dev/null @@ -1,162 +0,0 @@ -#!/usr/bin/env python3 -"""Does every raw capture in 10_knowledge/ have an ingest record? - -The third and most important layer of the ingest boundary, because it is the only -one that cannot be bypassed by a tool nobody remembered to guard. The PreToolUse -hook guards agent writes; the provenance gate guards `minion.py`; **this checks -the resulting state** and therefore catches whatever got in by any route at all, -including routes that do not exist yet. - -## What it found on 2026-08-10, the first time it ran - - 10_knowledge/ files: 2,879 - with an ingest-log record: 1,747 (61%) - with none: 1,132 (39%) - - of those, `type: raw`: 449 <- protocol violations - written in the preceding 10 days: 395 - -Notes and run-notes dominate the unlogged population and that is correct: they are -authored in place by design. **Raw captures are the violation**, because a raw is -by definition something retrieved from outside, and the ingest gate is where -provenance gets checked. - -The 2026-08-07 fabricated-citation batch sits inside that 449. - -## Reading the output - -A finding here is not an accusation of fabrication. It says the file skipped the -check, so nobody knows either way. Run `bin/capture-validate` on the listed paths -to find out which of them actually carry unearned citations. - -Usage: - bin/knowledge-reconcile # summary - bin/knowledge-reconcile --list # every unlogged raw - bin/knowledge-reconcile --since 2026-08 # only recent - bin/knowledge-reconcile --json out.json - bin/knowledge-reconcile --fail-on-new # CI: nonzero if any unlogged raw exists -""" - -from __future__ import annotations - -import argparse -import collections -import json -import re -import sys -from datetime import date -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] -KNOWLEDGE = ROOT / "10_knowledge" -LOGS = [ROOT / "01_ingest" / "ingest-log.md", ROOT / "01_ingest" / "prep-ingest-log.md"] - -# Types authored directly in 10_knowledge by design. Not violations. -AUTHORED_TYPES = { - "note", "index", "source-literature-run", "synthesis", "project", - "lab-report", "adr", "decision", "incident-finding", "draft", -} - -FILENAME_RE = re.compile(r"([0-9A-Za-z_.\-]+\.md)") -TYPE_RE = re.compile(r'^type:\s*"?([A-Za-z-]+)"?', re.M) -DATE_RE = re.compile(r"(\d{4}-\d{2}-\d{2})") - - -def logged_basenames() -> set[str]: - names: set[str] = set() - for log in LOGS: - if log.exists(): - names |= set(FILENAME_RE.findall(log.read_text(encoding="utf-8", errors="replace"))) - return names - - -def doc_type(path: Path) -> str | None: - try: - head = path.read_text(encoding="utf-8", errors="replace")[:1500] - except OSError: - return None - m = TYPE_RE.search(head) - return m.group(1).lower() if m else None - - -def main() -> int: - ap = argparse.ArgumentParser( - description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter - ) - ap.add_argument("--list", action="store_true") - ap.add_argument("--since", type=str, default="", help="YYYY-MM or YYYY-MM-DD prefix") - ap.add_argument("--json", type=str, default="") - ap.add_argument("--fail-on-new", action="store_true") - args = ap.parse_args() - - logged = logged_basenames() - if not logged: - print("ERROR: no ingest log found or it is empty. Cannot reconcile — " - "reporting nothing rather than reporting everything as a violation.", - file=sys.stderr) - return 2 - - total = 0 - by_type: dict[str, list[int]] = collections.defaultdict(lambda: [0, 0]) - violations: list[Path] = [] - - for path in KNOWLEDGE.rglob("*.md"): - total += 1 - dtype = doc_type(path) or "(none)" - is_logged = path.name in logged - by_type[dtype][0 if is_logged else 1] += 1 - if not is_logged and dtype == "raw": - if args.since and not (m := DATE_RE.search(path.name)) or ( - args.since and not m.group(1).startswith(args.since) - ): - continue - violations.append(path) - - unlogged_total = sum(v[1] for v in by_type.values()) - print(f"10_knowledge files: {total}") - print(f" with an ingest record: {total - unlogged_total}" - f" ({100 * (total - unlogged_total) / max(total, 1):.0f}%)") - print(f" with none: {unlogged_total}" - f" ({100 * unlogged_total / max(total, 1):.0f}%)") - - print(f"\n{'type':<24}{'logged':>8}{'unlogged':>10}{'%':>6}") - for t, (a, b) in sorted(by_type.items(), key=lambda kv: -(kv[1][0] + kv[1][1])): - if a + b < 5: - continue - note = " <- authored in place, expected" if t in AUTHORED_TYPES else "" - print(f"{t:<24}{a:>8}{b:>10}{100 * b / (a + b):>5.0f}%{note}") - - print(f"\n*** VIOLATIONS: {len(violations)} `type: raw` capture(s) with no ingest record ***") - if violations: - print(" A raw is by definition retrieved from outside. Skipping ingest skips") - print(" the provenance gate, so nobody knows whether its citation is real.") - months = collections.Counter( - m.group(1)[:7] for p in violations if (m := DATE_RE.search(p.name)) - ) - print("\n by month: " + ", ".join(f"{k} {v}" for k, v in sorted(months.items()))) - domains = collections.Counter( - str(p.relative_to(KNOWLEDGE)).split("/")[0] for p in violations - ) - print(" by domain: " + ", ".join(f"{k} {v}" for k, v in domains.most_common(6))) - print(f"\n Next: bin/capture-validate {' '.join(str(p.relative_to(ROOT)) for p in violations[:2])} ...") - - if args.list: - print() - for p in sorted(violations): - print(f" {p.relative_to(ROOT)}") - - if args.json: - Path(args.json).write_text(json.dumps({ - "as_of": date.today().isoformat(), - "total": total, - "unlogged": unlogged_total, - "violations": [str(p.relative_to(ROOT)) for p in sorted(violations)], - "by_type": {t: {"logged": a, "unlogged": b} for t, (a, b) in by_type.items()}, - }, indent=2), encoding="utf-8") - print(f"\nwrote {args.json}") - - return 1 if (args.fail_on_new and violations) else 0 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/knowledge-report b/bin/knowledge-report deleted file mode 100755 index 86de0b2..0000000 --- a/bin/knowledge-report +++ /dev/null @@ -1,89 +0,0 @@ -#!/usr/bin/env python3 -"""Lightweight navigation and synthesis report for 10_knowledge/. - -Helps with recall surface (Fix 5) and raw-vs-synthesis visibility (Fix 3). -Uses the audit-sweep machinery for candidate discovery + simple domain stats. - -Usage: - bin/knowledge-report --json - bin/knowledge-report --domain agents --max-days 30 -""" - -from __future__ import annotations - -import argparse -import json -import subprocess -import sys -from datetime import date -from pathlib import Path -from typing import Any - -ROOT = Path(__file__).resolve().parents[1] -KNOWLEDGE = ROOT / "10_knowledge" - - -def domain_stats(subset: str | None = None) -> dict[str, Any]: - base = KNOWLEDGE - if subset: - base = base / subset - if not base.is_dir(): - return {"error": "no such domain"} - - files = list(base.rglob("*.md")) - raws = [f for f in files if "__raw__" in f.name or "/raw/" in str(f)] - notes = [f for f in files if "__note__" in f.name] - total = len(files) - return { - "domain": subset or "all", - "total_md": total, - "raw_like": len(raws), - "note_like": len(notes), - "raw_note_ratio": round(len(raws) / max(len(notes), 1), 2) if total else 0, - } - - -def main() -> int: - parser = argparse.ArgumentParser() - parser.add_argument("--json", action="store_true") - parser.add_argument("--domain", help="limit to one domain") - parser.add_argument("--max-days", type=int, default=30) - args = parser.parse_args() - - stats = domain_stats(args.domain) - # Leverage audit-sweep for synthesis candidates (needs-audit raws are prime targets) - try: - cmd = [str(ROOT / "bin" / "audit-sweep"), "--dry-run", "--json", "--max-days", str(args.max_days)] - if args.domain: - cmd.extend(["--subset", args.domain]) - proc = subprocess.run(cmd, capture_output=True, text=True, timeout=30) - audit_data = json.loads(proc.stdout) if proc.stdout.strip() else {} - except Exception: - audit_data = {"candidate_count": 0, "explicit_needs_audit": 0} - - report = { - "date": date.today().isoformat(), - "domain_stats": stats, - "audit_candidates": { - "total": audit_data.get("candidate_count", 0), - "explicit_needs_audit": audit_data.get("explicit_needs_audit", 0), - }, - "synthesis_note": "High raw:note ratio or many needs-audit raws = good targets for ingest-source + create-source-summary to produce durable notes.", - "recommended": "bin/audit-sweep --apply --subset <domain> ; bin/extract-knowledge for projects ; review local 10_knowledge/index.md promotion rules (template: 10_knowledge/index.template.md)", - } - - if args.json: - print(json.dumps(report, indent=2)) - return 0 - - print(f"Knowledge Report — {report['date']}") - print(f"Domain: {stats.get('domain')}") - print(f" Total md: {stats.get('total_md')} raw-like: {stats.get('raw_like')} note-like: {stats.get('note_like')} ratio: {stats.get('raw_note_ratio')}") - print(f"Audit candidates (last {args.max_days}d): {audit_data.get('candidate_count', 0)} (explicit needs-audit: {audit_data.get('explicit_needs_audit', 0)})") - print(report["synthesis_note"]) - print("Next:", report["recommended"]) - return 0 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/knowledge-write-guard b/bin/knowledge-write-guard deleted file mode 100755 index ca95c0a..0000000 --- a/bin/knowledge-write-guard +++ /dev/null @@ -1,154 +0,0 @@ -#!/usr/bin/env python3 -"""PreToolUse hook — a new raw capture may not be hand-written into 10_knowledge/. - -## Why this exists - -`01_ingest` was documented as the mandatory path into durable knowledge. Measured -on 2026-08-10, it was not being treated as mandatory: - - 10_knowledge/ holds 2,879 files. - 1,132 (39%) have no ingest-log record at all. - 449 of those are `type: raw` — captures that MUST route through 01_ingest. - 395 of the 449 were written in the preceding ten days. - -The 2026-08-07 fabricated-citation batch is inside that population (see -`20_live/security/2026-08-09__fabricated-source-captures-in-10-knowledge.md`). -Its 21 files were written straight into `10_knowledge/`, which is exactly why the -"never auto-route without confirmation" guarantee never engaged: nothing was -routed. And the provenance gate added to `01_ingest/minion.py` cannot help, because -a direct write never reaches the minion. - -A convention that 39% of writes ignore is not a control. This makes it one. - -## What it blocks, and what it deliberately does not - -BLOCKS creating a **new** `type: raw` file under `10_knowledge/`. - That is the capture-bypass path and the only thing measured as a problem. - -ALLOWS edits to existing files. Remediation, quarantine banners, status changes - and re-sourcing all need this — the 2026-08-09 cleanup touched 144 files - and a blanket ban would have blocked all of it. - -ALLOWS notes, syntheses, indexes and run-notes. Those are authored in place by - design; the research loop's own step 5 tells you to write one. 86% of - notes have no ingest record and that is correct, not a violation. - -ALLOWS anything when MAINFRAME_KNOWLEDGE_WRITE=1, which is how the minion and a - deliberate operator get through. - -## Fail-open, on purpose - -If this guard throws, cannot parse its input, or is unsure, it **allows the write** -and says nothing. A guard that breaks the session when it has a bug is worse than -the problem it prevents. Its job is to make the wrong path inconvenient and -visible, not to be a load-bearing security boundary. - -Install as a PreToolUse hook matching Write|Edit. -""" - -from __future__ import annotations - -import json -import os -import re -import sys - -KNOWLEDGE_DIR = "10_knowledge/" - -# Authored-in-place types. Anything here is a synthesis product, not a capture. -AUTHORED_TYPES = { - "note", "index", "source-literature-run", "synthesis", "project", - "lab-report", "adr", "decision", "incident-finding", "draft", -} - -DENY_MESSAGE = """\ -Blocked: a new raw capture may not be written directly into 10_knowledge/. - - {path} - -Raw captures must enter through the ingest pipeline, which is where provenance is -checked: - - 1. write the capture to 00_inbox/ - 2. bin/ingest-minion run --dry-run # see what would route - 3. bin/ingest-minion run --apply # routes, with the provenance gate - -The gate rejects any capture carrying url/doi/authors/year without a -retrieval_receipt. That check is the entire point, and a direct write skips it. -On 2026-08-09 this bypass was found to account for 449 unlogged raw captures, -including a batch of 107 with citations to papers that do not exist. - -If this file is a synthesis rather than a capture, set `type: note` (or another -authored type) and it will be allowed. - -To override deliberately: MAINFRAME_KNOWLEDGE_WRITE=1 -""" - - -def allow() -> None: - sys.exit(0) - - -def deny(path: str) -> None: - print(json.dumps({ - "hookSpecificOutput": { - "hookEventName": "PreToolUse", - "permissionDecision": "deny", - "permissionDecisionReason": DENY_MESSAGE.format(path=path), - } - })) - sys.exit(0) - - -def declared_type(content: str) -> str | None: - """Read `type:` from frontmatter, if there is any.""" - if not content.startswith("---"): - return None - end = content.find("\n---", 3) - head = content[:end] if end != -1 else content[:1500] - m = re.search(r'^type:\s*"?([A-Za-z-]+)"?', head, re.M) - return m.group(1).lower() if m else None - - -def main() -> None: - if os.environ.get("MAINFRAME_KNOWLEDGE_WRITE") == "1": - allow() - - try: - payload = json.load(sys.stdin) - except Exception: # noqa: BLE001 — unparseable input is not the user's problem - allow() - - tool = payload.get("tool_name") or payload.get("toolName") or "" - args = payload.get("tool_input") or payload.get("toolInput") or {} - path = str(args.get("file_path") or args.get("path") or "") - - if KNOWLEDGE_DIR not in path.replace(os.sep, "/"): - allow() - - # Edits to existing material are how remediation happens. Never block them. - if tool != "Write": - allow() - if os.path.exists(path): - allow() - - content = str(args.get("content") or "") - dtype = declared_type(content) - - # No frontmatter at all, or an authored type: not a capture. Let it through. - if dtype is None or dtype in AUTHORED_TYPES: - allow() - - if dtype == "raw": - deny(path) - - allow() - - -if __name__ == "__main__": - try: - main() - except SystemExit: - raise - except Exception: # noqa: BLE001 — a broken guard must never block work - sys.exit(0) diff --git a/bin/lab-report b/bin/lab-report deleted file mode 100755 index 6b09ab7..0000000 --- a/bin/lab-report +++ /dev/null @@ -1,416 +0,0 @@ -#!/usr/bin/env python3 -"""Scaffold and check researcher-style lab reports for project experiments. - -Universal notebook layer for measured tests. Complements: - - project-experiment-loop (eval-registry harvest) - - craft-research-loop (keep|kill|iterate product trials) - -Commands: - scaffold --project <slug> --title ... --question ... --decision ... - check --project <slug> --id <lab_report_id> - list --project <slug> [--open] -""" - -from __future__ import annotations - -import argparse -import re -import sys -from datetime import date -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] -PROJECTS = ROOT / "30_projects" -TEMPLATE = ROOT / ".context" / "templates" / "lab-report.md" -WORKFLOW = ROOT / ".context" / "workflows" / "lab-report.md" - -REQUIRED_HEADINGS = [ - "## 1. Question", - "## 2. Decision this report supports", - "## 3. Hypothesis", - "## 5. Materials and methods", - "## 6. Results", - "## 7. Irregularities and anomalies", - "## 9. Limitations", - "## 10. Disposition and next experiment", -] - -FRONTMATTER_KEYS = [ - "lab_report_id", - "study_type", - "decision_sentence", - "disposition", -] - - -def slugify(title: str) -> str: - s = re.sub(r"[^a-z0-9-]+", "-", title.lower()).strip("-") - return s[:48] or "experiment" - - -def parse_frontmatter(text: str) -> dict[str, str]: - if not text.startswith("---"): - return {} - parts = text.split("---", 2) - if len(parts) < 3: - return {} - meta: dict[str, str] = {} - for line in parts[1].splitlines(): - if ":" not in line: - continue - k, v = line.split(":", 1) - meta[k.strip()] = v.strip().strip('"').strip("'") - return meta - - -def report_paths(project: str) -> tuple[Path, Path, Path]: - proj = PROJECTS / project - lab_dir = proj / "outputs" / "lab-reports" - raw_root = proj / "raw-materials" - return proj, lab_dir, raw_root - - -def cmd_scaffold(args: argparse.Namespace) -> int: - proj, lab_dir, raw_root = report_paths(args.project) - if not proj.is_dir(): - print(f"Project not found: {proj}", file=sys.stderr) - return 1 - - today = date.today().isoformat() - slug = slugify(args.title) - report_id = f"{today}-{slug}" - lab_dir.mkdir(parents=True, exist_ok=True) - raw_dir = raw_root / report_id - raw_dir.mkdir(parents=True, exist_ok=True) - - protocol = args.protocol_ref or f"30_projects/{args.project}/methodology-approach.md" - study = args.study_type - question = args.question.strip() - decision = args.decision.strip().replace('"', "'") - hypothesis = (args.hypothesis or "Descriptive measurement — no directional hypothesis.").replace( - '"', "'" - ) - metric = args.primary_metric or "TBD" - unit = args.unit_of_analysis or "TBD" - - body = f"""--- -title: "Lab report — {args.title}" -domain: "agent-operations" -type: "lab-report" -status: "running" -lab_report_id: "{report_id}" -eval_run_id: "{report_id}" -project: "{args.project}" -study_type: "{study}" -protocol_ref: "{protocol}" -decision_sentence: "{decision}" -hypothesis: "{hypothesis}" -primary_metric: "{metric}" -unit_of_analysis: "{unit}" -disposition: "open" -privacy: "private-local" -tags: ["lab-report", "experiment"] -updated: "{today}" -source: "bin/lab-report scaffold" ---- - -# Lab report — {args.title} - -**ID:** `{report_id}` -**Project:** `{args.project}` -**Study type:** {study} -**Disposition:** `open` - ---- - -## 1. Question - -{question} - -## 2. Decision this report supports - -{decision} - -## 3. Hypothesis - -{hypothesis} - -## 4. Background / prior art (optional) - -- - -## 5. Materials and methods - -### 5.1 System under test - -| Field | Value | -|-------|-------| -| Stack / models | | -| Harness / tools | | -| Code / config pin | `{protocol}` | -| Hardware / host constraints | | - -### 5.2 Design - -| Field | Value | -|-------|-------| -| study_type | {study} | -| unit_of_analysis | {unit} | -| primary metric | {metric} | -| secondary metrics | | -| n / replicates | | -| factors (independent) | | -| controlled / held fixed | | -| randomization / seeds | | -| blinding / sealing | | - -### 5.3 Procedure - -1. Preflight -2. Execute -3. Score -4. Fill this report; store raw under `raw-materials/{report_id}/` - -```bash -# Commands: - -``` - -## 6. Results - -| Cell / condition | n | Primary metric | Secondary | Notes | -|------------------|---:|---------------:|----------:|-------| - -### Raw artifacts - -| Kind | Path | -|------|------| -| receipts / logs | `raw-materials/{report_id}/` | - -## 7. Irregularities and anomalies - -| id | severity | category | observation | artifact_ref | resolved | -|----|----------|----------|-------------|--------------|----------| - -None observed: **no** (fill rows) / **yes** (delete table and write "None observed.") - -## 8. Interpretation - -| Statement | Type | Confidence | -|-----------|------|------------| -| | observation / inference / hypothesis | | - -## 9. Limitations / what this does **not** prove - -- - -## 10. Disposition and next experiment - -| Field | Value | -|-------|-------| -| disposition | open | -| why | | -| next_experiment | | -| registry harvest? | {"yes if decision-bearing" if study != "observational" else "optional"} | - -## 11. Metric extract (eval-registry) - -```yaml -registry: - project: {args.project} - run_id: {report_id} - study_type: {study} - protocol_ref: {protocol} - date: {today} - decision_sentence: "{decision}" - artifact_path: outputs/lab-reports/{report_id}.md - raw_path: raw-materials/{report_id}/ - decision_use: {"exploratory_only" if study in {"exploratory", "observational"} else "regression_only"} -metrics: - - name: {metric if metric != "TBD" else "placeholder_metric"} - slice: all - value: 0 - n: 0 - unit: count -irregularities: - - id: scaffold-placeholder - severity: info - category: protocol - observation: "Scaffold only — replace after real run" - context: "{report_id}" - resolved: false -``` - -## 12. Provenance - -| Field | Value | -|-------|-------| -| operator / agent | | -| started | {today} | -| finished | | -| git_sha | | -| privacy | private-local | -| related reports | | -""" - - out_path = lab_dir / f"{report_id}.md" - if out_path.exists(): - print(f"Refusing overwrite: {out_path}", file=sys.stderr) - return 1 - out_path.write_text(body, encoding="utf-8") - print(f"Wrote {out_path.relative_to(ROOT)}") - print(f"Raw dir {raw_dir.relative_to(ROOT)}") - print(f"lab_report_id={report_id}") - print("Convention: .context/workflows/lab-report.md") - print("Next: run experiment, fill results/disposition, then:") - print(f" bin/lab-report check --project {args.project} --id {report_id}") - if args.project.endswith("-eval") or args.project in { - "scaffold-claims-study", - "skill-eval-workshop", - "verified-done", - "agent-harness-eval", - }: - print(f" bin/project-experiment-loop close --project {args.project}") - return 0 - - -def find_report(project: str, report_id: str) -> Path | None: - _, lab_dir, _ = report_paths(project) - candidates = [ - lab_dir / f"{report_id}.md", - PROJECTS / project / "outputs" / f"{report_id}.md", - ] - for p in candidates: - if p.is_file(): - return p - # fuzzy: suffix match - if lab_dir.is_dir(): - for p in lab_dir.glob("*.md"): - if report_id in p.stem: - return p - out = PROJECTS / project / "outputs" - if out.is_dir(): - for p in out.glob("*.md"): - if report_id in p.stem: - return p - return None - - -def cmd_check(args: argparse.Namespace) -> int: - path = find_report(args.project, args.id) - if not path: - print(f"Report not found for id={args.id} under {args.project}", file=sys.stderr) - return 1 - text = path.read_text(encoding="utf-8") - meta = parse_frontmatter(text) - errors: list[str] = [] - warnings: list[str] = [] - - for k in FRONTMATTER_KEYS: - if k == "lab_report_id" and not meta.get(k) and not meta.get("eval_run_id"): - errors.append("missing lab_report_id or eval_run_id") - elif k != "lab_report_id" and not meta.get(k): - errors.append(f"missing frontmatter: {k}") - - for h in REQUIRED_HEADINGS: - if h not in text: - # allow short "Limitations" without full string - if h.startswith("## 9.") and "## 9." in text: - continue - errors.append(f"missing section: {h}") - - if "None observed" not in text and "irregularities:" not in text.lower(): - if "| id | severity |" not in text and "Irregularities" in text: - warnings.append("irregularities table empty — state 'None observed' if truly none") - - disp = meta.get("disposition", "open") - if disp == "open" and "<!--" not in text: - if re.search(r"\|[^|]+\|[^|]+\|", text) and "TBD" not in meta.get("primary_metric", ""): - warnings.append("disposition still open — set accept/reject/hold/iterate when done") - - if "does **not** prove" not in text and "does not prove" not in text.lower(): - warnings.append("add explicit does-not-prove / limitations language") - - print(f"check: {path.relative_to(ROOT)}") - print(f" disposition={disp} study_type={meta.get('study_type')}") - for e in errors: - print(f" ERROR: {e}") - for w in warnings: - print(f" WARN: {w}") - if errors: - return 1 - print(" OK (minimum bar)") - return 0 - - -def cmd_list(args: argparse.Namespace) -> int: - proj, lab_dir, _ = report_paths(args.project) - if not proj.is_dir(): - print(f"Project not found: {proj}", file=sys.stderr) - return 1 - paths: list[Path] = [] - if lab_dir.is_dir(): - paths.extend(sorted(x for x in lab_dir.glob("*.md") if x.name != "README.md")) - # Also list root outputs that explicitly declare lab_report_id (not all study_type docs) - out = proj / "outputs" - if out.is_dir(): - for p in sorted(out.glob("*.md")): - if p.parent.name == "lab-reports": - continue - t = p.read_text(encoding="utf-8", errors="replace")[:800] - if "lab_report_id:" in t or ( - "type: \"lab-report\"" in t or "type: lab-report" in t - ): - paths.append(p) - if not paths: - print("(no lab reports found)") - return 0 - for p in paths: - meta = parse_frontmatter(p.read_text(encoding="utf-8", errors="replace")) - disp = meta.get("disposition") or meta.get("status") or "?" - if args.open and disp not in {"open", "running", "active"}: - continue - rid = meta.get("lab_report_id") or meta.get("eval_run_id") or p.stem - print(f"{rid}\t{disp}\t{p.relative_to(ROOT)}") - return 0 - - -def main() -> int: - ap = argparse.ArgumentParser(description=__doc__) - sub = ap.add_subparsers(dest="cmd", required=True) - - sc = sub.add_parser("scaffold", help="Create outputs/lab-reports/<id>.md + raw dir") - sc.add_argument("--project", required=True) - sc.add_argument("--title", required=True) - sc.add_argument("--question", required=True) - sc.add_argument("--decision", required=True) - sc.add_argument( - "--study-type", - default="exploratory", - choices=["exploratory", "confirmatory", "regression", "observational", "calibration"], - ) - sc.add_argument("--hypothesis", default=None) - sc.add_argument("--protocol-ref", default=None) - sc.add_argument("--primary-metric", default=None) - sc.add_argument("--unit-of-analysis", default=None) - - ch = sub.add_parser("check", help="Minimum-bar completeness check") - ch.add_argument("--project", required=True) - ch.add_argument("--id", required=True, help="lab_report_id") - - ls = sub.add_parser("list", help="List lab reports for a project") - ls.add_argument("--project", required=True) - ls.add_argument("--open", action="store_true", help="Only open/running dispositions") - - args = ap.parse_args() - if args.cmd == "scaffold": - return cmd_scaffold(args) - if args.cmd == "check": - return cmd_check(args) - if args.cmd == "list": - return cmd_list(args) - return 2 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/lane-intake b/bin/lane-intake deleted file mode 100755 index 24d3167..0000000 --- a/bin/lane-intake +++ /dev/null @@ -1,716 +0,0 @@ -#!/usr/bin/env python3 -"""Research Lane Intake — scaffold trackers from surfaced questions. - -This is the explicit step that turns new research questions into -lanes/<slug>/README.md entries (tracker only). - -Usage examples: - bin/lane-intake scaffold \ - --question "How should intent graphs drive RAG routing?" \ - --decision "Decide on intent-graph integration in retrieval + planning." \ - --domain "graph-memory" \ - --trigger "mindgraph-eval multi-RAG planning 2026-06-24" \ - --slug "intent-graph-rag" \ - --stakes high --priority P1 --apply - - bin/lane-intake scan some-plan-or-run-note.md --apply - - bin/lane-intake list --status active --priority P0 --format table - bin/lane-intake archive dcf-valuation --apply - bin/lane-intake set-priority some-slug P0 --apply - -Priority model (P0–P4) is project-needs driven. Archive moves completed -phases to lanes/completed/ (all research retained; see lanes/completed/README.md). -Always performs dual MindGraph query pass when possible. -Respects MainFrame boundaries: only writes tracker skeleton + receipt + log append. -Durable knowledge still requires 00_inbox → ingest pipeline. -""" - -from __future__ import annotations - -import argparse -import json -import re -import subprocess -import sys -from datetime import date -from pathlib import Path -from textwrap import dedent -from typing import Any - -ROOT = Path(__file__).resolve().parents[1] -LANES_DIR = ROOT / "30_projects" / "research-lanes-strategy" / "lanes" -COMPLETED_DIR = LANES_DIR / "completed" -MASTER_PLAN = ROOT / "30_projects" / "research-lanes-strategy" / "plans" / "research-lanes-master-plan.md" -LOG = ROOT / "30_projects" / "research-lanes-strategy" / "log.md" -RAW_MATERIALS = ROOT / "30_projects" / "research-lanes-strategy" / "raw-materials" -TEMPLATE = LANES_DIR / "_TEMPLATE.md" -MINDGRAPH = ROOT / "bin" / "mindgraph" - -TODAY = date.today().isoformat() - - -def run_dual_mindgraph(query: str, top_k: int = 6) -> dict[str, Any]: - """Run dual queries. Returns dict with knowledge + projects results or errors.""" - results: dict[str, Any] = {"query": query, "knowledge": [], "projects": [], "errors": []} - if not MINDGRAPH.exists(): - results["errors"].append("bin/mindgraph not found") - return results - - for label, db in [ - ("knowledge", str(Path.home() / ".mindgraph" / "mainframe.sqlite")), - ("projects", str(Path.home() / ".mindgraph" / "mainframe-projects.sqlite")), - ]: - cmd = [ - str(MINDGRAPH), - "query", - query, - "--db", - db, - "--json", - "--top-k", - str(top_k), - ] - try: - out = subprocess.run(cmd, capture_output=True, text=True, timeout=120) - if out.returncode == 0: - try: - data = json.loads(out.stdout) - results[label] = data if isinstance(data, list) else [data] - except Exception: - results[label] = [{"raw": out.stdout[:2000]}] - else: - results["errors"].append(f"{label}: {out.stderr[:300]}") - except Exception as e: - results["errors"].append(f"{label}: {e}") - return results - - -def find_lane_dir(slug: str) -> Path | None: - """Find lane dir in main lanes/ or completed/.""" - for base in (LANES_DIR, COMPLETED_DIR): - p = base / slug - if p.exists() and (p / "README.md").exists(): - return p - return None - - -def parse_frontmatter(path: Path) -> dict[str, str]: - """Lightweight frontmatter parser (no PyYAML dep).""" - if not path.exists(): - return {} - text = path.read_text() - if not text.startswith("---"): - return {} - m = re.search(r"^---\s*\n(.*?)\n---", text, re.DOTALL) - if not m: - return {} - fm = {} - for line in m.group(1).splitlines(): - if ":" not in line: - continue - k, v = line.split(":", 1) - k = k.strip() - v = v.strip().strip('"').strip("'") - fm[k] = v - return fm - - -def next_lane_id(prefix: str = "C") -> str: - """Heuristic: scan master plan for highest ID with given prefix.""" - if not MASTER_PLAN.exists(): - return f"{prefix}01" - text = MASTER_PLAN.read_text() - ids = re.findall(rf"\b{prefix}(\d+)\b", text) - if not ids: - return f"{prefix}01" - max_n = max(int(x) for x in ids) - return f"{prefix}{max_n + 1:02d}" - - -def slugify(text: str) -> str: - s = text.lower().strip() - s = re.sub(r"[^a-z0-9]+", "-", s) - return s.strip("-")[:60] - - -def load_template() -> str: - if TEMPLATE.exists(): - return TEMPLATE.read_text() - # Fallback minimal - return dedent("""\ - --- - title: "Lane: <name>" - domain: "general-research" - type: "project" - status: "active" - lane_id: "CXX" - knowledge_domain: "<primary-10_knowledge-domain>" - capture_tags: ["research-lane", "lane-cxx", "<topic>"] - stakes: "high" - cadence: "weekly" - priority: "next" - handoff: "" - updated: "YYYY-MM-DD" - tags: ["research-lane"] - --- - - ## Central question - - (one sentence) - - ## Decision this lane supports - - (one sentence) - - ## Knowledge routing - - - **Captures:** `00_inbox/` → `bin/ingest-minion` → `10_knowledge/<knowledge_domain>/` - - **Synthesis:** `10_knowledge/<knowledge_domain>/` when 3+ related stubs - - **Routing doc:** `plans/knowledge-routing.md` - - **Research convention:** `plans/first-principles-research-conventions.md` - - ## First-principles pass - - - **Foundation question:** ... - - **Primitive map:** ... - - **Specialization path:** ... - - **Stop condition:** ... - - ## Captured knowledge (index only) - - | Date | Path | Phase | Type | Status | - |------|------|-------|------|--------| - | — | — | — | — | no captures yet | - - ## Spin-out project - - `30_projects/<name>` — implementation and experiments live there. - """) - - -def render_lane_readme(meta: dict[str, str]) -> str: - tpl = load_template() - # Very lightweight substitution for the critical fields - replacements = { - "<name>": meta.get("title", meta.get("slug", "New Lane")), - "CXX": meta.get("lane_id", "C??"), - "<primary-10_knowledge-domain>": meta.get("knowledge_domain", "knowledge-systems"), - "lane-cxx": f"lane-{meta.get('lane_id', 'c??').lower()}", - "<topic>": meta.get("slug", "research"), - "YYYY-MM-DD": TODAY, - } - content = tpl - for k, v in replacements.items(): - content = content.replace(k, v) - - # Force key fields in frontmatter - content = re.sub( - r'lane_id: ".*?"', - f'lane_id: "{meta.get("lane_id", "C??")}"', - content, - ) - content = re.sub( - r'knowledge_domain: ".*?"', - f'knowledge_domain: "{meta.get("knowledge_domain", "knowledge-systems")}"', - content, - ) - content = re.sub( - r'priority: ".*?"', - f'priority: "{meta.get("priority", "P1")}"', - content, - ) - content = re.sub( - r'stakes: ".*?"', - f'stakes: "{meta.get("stakes", "high")}"', - content, - ) - if meta.get("project_ties"): - # inject or update optional project_ties line - if "project_ties:" in content: - content = re.sub(r'project_ties: ".*?"', f'project_ties: "{meta["project_ties"]}"', content) - else: - content = re.sub( - r'(priority: "[^"]*")', - r'\1\nproject_ties: "' + meta["project_ties"] + '"', - content, - ) - - # Inject actual question/decision after headers - q = meta.get("central_question", "(fill in)") - d = meta.get("decision", "(fill in)") - handoff = meta.get("handoff", "") - - # Replace the placeholder sections if present - content = re.sub( - r"## Central question\n\n\(one sentence\)", - f"## Central question\n\n{q}", - content, - ) - content = re.sub( - r"## Decision this lane supports\n\n\(one sentence\)", - f"## Decision this lane supports\n\n{d}", - content, - ) - if handoff: - content = re.sub( - r"## Spin-out project\n\n`30_projects/<name>`.*", - f"## Spin-out project\n\n{handoff}", - content, - flags=re.DOTALL, - ) - - # Ensure capture_tags include both raw lane id and lane-<id> tag - lid = meta.get("lane_id", "").lower() - if lid: - lane_tag = f"lane-{lid}" if not lid.startswith("lane-") else lid - content = re.sub( - r'capture_tags: \["research-lane", "[^"]*", "[^"]*"\]', - f'capture_tags: ["research-lane", "{lid}", "{lane_tag}", "{meta.get("slug", "")}"]', - content, - ) - - # Add trigger note at end if present - trigger = meta.get("trigger") - if trigger: - content += f"\n\n## Intake provenance\n\n- Trigger: {trigger}\n- Created: {TODAY} via bin/lane-intake\n" - - return content - - -def write_receipt(meta: dict[str, str], mg_results: dict[str, Any]) -> Path: - RAW_MATERIALS.mkdir(parents=True, exist_ok=True) - slug = meta.get("slug", "new-lane") - path = RAW_MATERIALS / f"{TODAY}__proposed-lane-{slug}.md" - row = f"| {meta.get('lane_id','C??')} | `{slug}` | {meta.get('central_question','')} | {meta.get('priority','next')} | {meta.get('stakes','high')} | {meta.get('handoff','')} | proposed |" - - receipt = dedent(f"""\ - --- - title: "Proposed lane: {meta.get('slug')}" - type: "proposed-lane" - created: "{TODAY}" - trigger: "{meta.get('trigger', '')}" - --- - - # Proposed Research Lane — {meta.get('lane_id')} {slug} - - **Central question:** {meta.get('central_question')} - - **Decision supported:** {meta.get('decision')} - - **Knowledge domain:** {meta.get('knowledge_domain')} - **Stakes / Priority:** {meta.get('stakes')} / {meta.get('priority')} - - ## Suggested master-plan row (copy into correct tier) - {row} - - ## Dual MindGraph pass (nominations only) - ```json - {json.dumps(mg_results, indent=2)[:4000]} - ``` - - ## Next - 1. Review this receipt + dual results. - 2. `bin/lane-intake scaffold ... --apply` (if not already). - 3. Insert row into research-lanes-master-plan.md. - 4. Apply first-principles pass before any source batch. - 5. Use lane README as brief for source-literature. - """) - path.write_text(receipt) - return path - - -def append_log(meta: dict[str, str]): - if not LOG.exists(): - return - entry = dedent(f""" - - ## {TODAY} | New lane via intake — {meta.get('lane_id')} {meta.get('slug')} - - - **Added** {meta.get('lane_id')} `{meta.get('slug')}` — {meta.get('central_question')} - - **Trigger:** {meta.get('trigger', 'operator / process')} - - **Trackers:** lane README created; master-plan row suggested in receipt; log updated. - - **MindGraph:** dual pass recorded in receipt. - """) - with LOG.open("a") as f: - f.write(entry) - - -# Priority aliases: list --priority P0 also matches legacy "immediate", etc. -PRIORITY_ALIASES: dict[str, set[str]] = { - "p0": {"p0", "immediate"}, - "p1": {"p1", "next", "high"}, - "p2": {"p2", "medium"}, - "p3": {"p3", "later", "low"}, - "p4": {"p4", "monitor", "alongside", "deferred", "curiosity"}, -} - - -def _priority_matches(prio: str, priority_filter: str) -> bool: - """Match P0–P4 filters against both Pn labels and legacy vocabulary.""" - p = (prio or "").strip().lower() - f = priority_filter.strip().lower() - if not f: - return True - if p == f or p.startswith(f): - return True - # Normalize "P0 critical" → p0 - token = p.split()[0] if p else "" - aliases = PRIORITY_ALIASES.get(f, {f}) - return token in aliases or p in aliases - - -def _status_matches(status: str, status_filter: str) -> bool: - """Prefix match so 'active — specialization pass done' matches --status active.""" - s = (status or "").strip().lower() - f = status_filter.strip().lower() - if not f: - return True - return s == f or s.startswith(f) or s.startswith(f + " ") or s.startswith(f + "—") or s.startswith(f + "-") - - -def list_lanes(status_filter: str | None = None, priority_filter: str | None = None, fmt: str = "table") -> list[dict[str, str]]: - """Scan lanes/ and lanes/completed/ and return filtered list.""" - results = [] - for base, loc in [(LANES_DIR, "lanes"), (COMPLETED_DIR, "completed")]: - if not base.exists(): - continue - for d in sorted(base.iterdir()): - if not d.is_dir() or d.name.startswith(".") or d.name == "completed": - continue - readme = d / "README.md" - if not readme.exists(): - continue - fm = parse_frontmatter(readme) - if not fm: - continue - st = fm.get("status", "unknown") - prio = fm.get("priority", "P?") - lid = fm.get("lane_id", "") - if status_filter and not _status_matches(st, status_filter): - continue - if priority_filter and not _priority_matches(prio, priority_filter): - continue - results.append({ - "id": lid, - "slug": d.name, - "status": st, - "priority": prio, - "domain": fm.get("knowledge_domain", ""), - "location": loc, - "path": str(readme), - }) - # sort by priority rough (P0 first), then id - prio_order = {"P0": 0, "P1": 1, "immediate": 0, "P2": 2, "P3": 3, "P4": 4, "next": 1, "later": 3, "high": 1, "monitor": 4} - results.sort(key=lambda r: (prio_order.get(r["priority"].split()[0] if r["priority"] else "P9", 9), r.get("id", ""))) - return results - - -def format_list(results: list[dict[str, str]], fmt: str = "table") -> str: - if fmt == "json": - return json.dumps(results, indent=2) - if not results: - return "No matching lanes." - lines = ["ID\tPriority\tStatus\tLocation\tSlug"] - for r in results: - lines.append(f"{r['id']}\t{r['priority']}\t{r['status']}\t{r['location']}\t{r['slug']}") - return "\n".join(lines) - - -def archive_lane(slug: str, apply: bool = False) -> dict[str, Any]: - """Move a lane from lanes/ to lanes/completed/. Updates frontmatter status.""" - src = LANES_DIR / slug - dst_dir = COMPLETED_DIR / slug - readme = src / "README.md" - result = {"slug": slug, "moved": False, "updated": False, "error": None} - - if not src.exists() or not readme.exists(): - # maybe already archived? - if (COMPLETED_DIR / slug).exists(): - result["error"] = "already in completed/" - return result - result["error"] = "lane not found in lanes/" - return result - - fm = parse_frontmatter(readme) - content = readme.read_text() - - if apply: - COMPLETED_DIR.mkdir(parents=True, exist_ok=True) - # update frontmatter status to complete if not already - if fm.get("status") != "complete": - content = re.sub(r'(?m)^status:\s*".*?"', 'status: "complete"', content) - content = re.sub(r'(?m)^status:\s*\S+', 'status: "complete"', content) - readme.write_text(content) # write before move - result["updated"] = True - - # move - dst_dir.parent.mkdir(parents=True, exist_ok=True) - src.rename(dst_dir) - result["moved"] = True - - # log - append_log({"lane_id": fm.get("lane_id", ""), "slug": slug, "central_question": "(archived)", "trigger": "bin/lane-intake archive"}) - # best-effort master plan + lanes/README update (status column) - _update_master_plan_status(slug, "complete") - _update_lanes_readme_status(slug, "complete") - - result["dst"] = str(dst_dir) - return result - - -def _update_master_plan_status(slug: str, new_status: str): - if not MASTER_PLAN.exists(): - return - text = MASTER_PLAN.read_text() - # crude but effective: target the row containing the slug - pattern = rf"(\| .*?`?{re.escape(slug)}`? .*?\| )(parked|active|complete|\*\*complete\*\*|\*\*active\*\*)(\s*\|)" - def repl(m): - return m.group(1) + new_status + m.group(3) - new_text = re.sub(pattern, repl, text, flags=re.IGNORECASE) - if new_text != text: - MASTER_PLAN.write_text(new_text) - - -def _update_lanes_readme_status(slug: str, new_status: str): - lr = LANES_DIR / "README.md" - if not lr.exists(): - return - text = lr.read_text() - # look for the table row with the slug folder - pattern = rf"(\| .*?`?{re.escape(slug)}/?`? .*?\| )(parked|active|complete|\*\*.*?\*\*)(\s*\|)" - def repl(m): - return m.group(1) + new_status + m.group(3) - new_text = re.sub(pattern, repl, text, flags=re.IGNORECASE) - if new_text != text: - lr.write_text(new_text) - - -def set_priority(slug: str, new_priority: str, apply: bool = False) -> dict[str, Any]: - """Update priority in a lane's frontmatter (works for main or completed).""" - lane_dir = find_lane_dir(slug) - result = {"slug": slug, "priority": new_priority, "updated": False, "error": None} - if not lane_dir: - result["error"] = "lane not found" - return result - readme = lane_dir / "README.md" - content = readme.read_text() - fm = parse_frontmatter(readme) - - if apply: - # replace priority line - content = re.sub(r'(?m)^priority:\s*".*?"', f'priority: "{new_priority}"', content) - content = re.sub(r'(?m)^priority:\s*\S+', f'priority: "{new_priority}"', content) - readme.write_text(content) - result["updated"] = True - # optional: touch updated date - content2 = readme.read_text() - content2 = re.sub(r'(?m)^updated:\s*".*?"', f'updated: "{TODAY}"', content2) - readme.write_text(content2) - - # best effort update tables - _update_master_plan_priority(slug, new_priority) - _update_lanes_readme_priority(slug, new_priority) - - result["path"] = str(readme) - return result - - -def _update_master_plan_priority(slug: str, new_prio: str): - if not MASTER_PLAN.exists(): - return - text = MASTER_PLAN.read_text() - pattern = rf"(\| .*?`?{re.escape(slug)}`? .*?\| )(P[0-9]|immediate|next|later|monitor|high|medium|low|.*?)( \|)" - def repl(m): - return m.group(1) + new_prio + m.group(3) - new_text = re.sub(pattern, repl, text, flags=re.IGNORECASE) - if new_text != text: - MASTER_PLAN.write_text(new_text) - - -def _update_lanes_readme_priority(slug: str, new_prio: str): - lr = LANES_DIR / "README.md" - if not lr.exists(): - return - # lanes/README currently doesn't have priority column in main table; skip or extend later - pass - - -def scaffold(meta: dict[str, str], apply: bool = False) -> dict[str, Any]: - slug = meta.get("slug") or slugify(meta.get("central_question", "new-research-lane")) - meta["slug"] = slug - - if not meta.get("lane_id"): - # choose prefix from domain or default C - prefix = "C" - if "fn" in meta.get("knowledge_domain", "").lower() or "finance" in meta.get("knowledge_domain", "").lower(): - prefix = "FN" - elif "data" in meta.get("knowledge_domain", "").lower(): - prefix = "DATA" - meta["lane_id"] = next_lane_id(prefix) - - lane_dir = LANES_DIR / slug - readme_path = lane_dir / "README.md" - - result = { - "slug": slug, - "lane_id": meta["lane_id"], - "lane_dir": str(lane_dir), - "readme": str(readme_path), - "created": False, - "receipt": None, - "log_appended": False, - } - - mg = run_dual_mindgraph(meta.get("central_question", slug)) - receipt_path = write_receipt(meta, mg) - result["receipt"] = str(receipt_path) - - if lane_dir.exists() and readme_path.exists(): - result["error"] = "lane already exists" - return result - - if apply: - lane_dir.mkdir(parents=True, exist_ok=True) - content = render_lane_readme(meta) - readme_path.write_text(content) - result["created"] = True - append_log(meta) - result["log_appended"] = True - - return result - - -def parse_candidate_block(text: str) -> dict[str, str] | None: - """Very simple parser for the ## Research Lane Candidate block.""" - m = re.search(r"## Research Lane Candidate\s*(.*?)(?=\n## |\Z)", text, re.DOTALL | re.IGNORECASE) - if not m: - return None - block = m.group(1) - meta = {} - aliases = { - "question": "central_question", - "central question": "central_question", - "lane": "lane_id", - "id": "lane_id", - "dom": "knowledge_domain", - "domain": "knowledge_domain", - } - for line in block.splitlines(): - if ":" not in line: - continue - k, v = line.split(":", 1) - k = k.strip().strip("*").strip("-").lstrip("-").strip().lower().replace(" ", "_") - v = v.strip().strip('"').strip("'") - k = aliases.get(k, k) - if k in {"lane_id", "slug", "central_question", "decision", "knowledge_domain", "stakes", "priority", "trigger", "handoff", "project_ties"}: - meta[k] = v - # convenience: if only "central_question" provided as "question" shorthand - if "central_question" not in meta and "question" in meta: - meta["central_question"] = meta.pop("question") - return meta or None - - -def scan(path: Path, apply: bool = False) -> list[dict[str, Any]]: - text = path.read_text() - candidates = [] - # support one or more blocks - for m in re.finditer(r"## Research Lane Candidate\s*(.*?)(?=\n## |\Z)", text, re.DOTALL | re.IGNORECASE): - block = m.group(0) - meta = parse_candidate_block(block) or {} - if "central_question" not in meta: - continue - res = scaffold(meta, apply=apply) - candidates.append(res) - return candidates - - -def main(): - p = argparse.ArgumentParser(description="Research Lane Intake tool") - sub = p.add_subparsers(dest="cmd", required=True) - - sc = sub.add_parser("scaffold", help="Create a lane from explicit fields") - sc.add_argument("--question", required=True) - sc.add_argument("--decision", required=True) - sc.add_argument("--domain", required=True, dest="knowledge_domain") - sc.add_argument("--trigger", default="manual") - sc.add_argument("--slug") - sc.add_argument("--lane-id", dest="lane_id") - sc.add_argument("--stakes", default="high") - sc.add_argument("--priority", default="P1") - sc.add_argument("--handoff", default="") - sc.add_argument("--project-ties", dest="project_ties", default="") - sc.add_argument("--apply", action="store_true") - - sn = sub.add_parser("scan", help="Parse Research Lane Candidate blocks from a file and scaffold") - sn.add_argument("path") - sn.add_argument("--apply", action="store_true") - - lst = sub.add_parser("list", help="List lanes (main + completed) filtered by status/priority. Supports researching everything with priority sequencing.") - lst.add_argument("--status", help="Filter e.g. active, complete, parked") - lst.add_argument( - "--priority", - help="Filter e.g. P0, P1 (aliases: P0=immediate, P1=next|high, P4=monitor|alongside|deferred)", - ) - lst.add_argument("--format", choices=["table", "json"], default="table") - lst.add_argument("--all", action="store_true", help="Include completed by default; use --status complete to focus") - - arc = sub.add_parser("archive", help="Move lane to lanes/completed/ (for complete phases). Updates status in key docs.") - arc.add_argument("slug") - arc.add_argument("--apply", action="store_true") - - sp = sub.add_parser("set-priority", help="Update priority on a lane (main or completed) based on current project needs. Example: P0 for active project handoffs.") - sp.add_argument("slug") - sp.add_argument("priority") - sp.add_argument("--apply", action="store_true") - - args = p.parse_args() - - if args.cmd == "scaffold": - meta = { - "central_question": args.question, - "decision": args.decision, - "knowledge_domain": args.knowledge_domain, - "trigger": args.trigger, - "slug": args.slug, - "lane_id": args.lane_id, - "stakes": args.stakes, - "priority": args.priority, - "handoff": args.handoff, - "project_ties": args.project_ties, - } - res = scaffold(meta, apply=args.apply) - print(json.dumps(res, indent=2)) - if not args.apply: - print("\nDry run. Re-run with --apply to write files.") - - elif args.cmd == "scan": - pth = Path(args.path) - if not pth.exists(): - print(f"File not found: {pth}") - sys.exit(1) - results = scan(pth, apply=args.apply) - print(json.dumps(results, indent=2)) - if not args.apply: - print("\nDry run. Re-run with --apply to write files.") - - elif args.cmd == "list": - res = list_lanes(status_filter=args.status, priority_filter=args.priority, fmt=args.format) - print(format_list(res, args.format)) - print(f"\nTotal matching: {len(res)} (use --status / --priority to focus; all research retained under priority model)") - - elif args.cmd == "archive": - res = archive_lane(args.slug, apply=args.apply) - print(json.dumps(res, indent=2)) - if not args.apply: - print("\nDry run. Re-run with --apply to move to lanes/completed/ and update trackers.") - - elif args.cmd == "set-priority": - res = set_priority(args.slug, args.priority, apply=args.apply) - print(json.dumps(res, indent=2)) - if not args.apply: - print("\nDry run. Re-run with --apply to update priority (reflects project needs).") - - -if __name__ == "__main__": - main() diff --git a/bin/mainframe-doctor b/bin/mainframe-doctor deleted file mode 100755 index 3fe9fb2..0000000 --- a/bin/mainframe-doctor +++ /dev/null @@ -1,12 +0,0 @@ -#!/usr/bin/env bash -# PATH-safe wrapper for scripts/mainframe_doctor -set -euo pipefail -ROOT="$(cd "$(dirname "$0")/.." && pwd)" -export PYTHONPATH="${ROOT}/scripts${PYTHONPATH:+:$PYTHONPATH}" -if [[ -x /opt/homebrew/bin/python3 ]]; then - exec /opt/homebrew/bin/python3 -m mainframe_doctor "$@" -elif [[ -x /usr/bin/python3 ]]; then - exec /usr/bin/python3 -m mainframe_doctor "$@" -else - exec python3 -m mainframe_doctor "$@" -fi diff --git a/bin/mindgraph b/bin/mindgraph deleted file mode 100755 index 093fab3..0000000 --- a/bin/mindgraph +++ /dev/null @@ -1,59 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" -ROOT="$(cd "$SCRIPT_DIR/.." && pwd)" -MINDGRAPH_BIN="${MINDGRAPH_BIN:-}" -DB_PATH="${MINDGRAPH_DB_PATH:-$HOME/.mindgraph/mainframe.sqlite}" -LOCAL_MINDGRAPH_BIN="$ROOT/mindgraph/.venv/bin/mindgraph" -FALLBACK_MINDGRAPH_BIN="$ROOT/30_projects/mindgraph/workbench/.venv/bin/mindgraph" - -if [[ -z "$MINDGRAPH_BIN" ]]; then - if [[ -x "$LOCAL_MINDGRAPH_BIN" ]]; then - MINDGRAPH_BIN="$LOCAL_MINDGRAPH_BIN" - elif [[ -x "$FALLBACK_MINDGRAPH_BIN" ]]; then - MINDGRAPH_BIN="$FALLBACK_MINDGRAPH_BIN" - fi -fi - -if [[ -z "$MINDGRAPH_BIN" ]]; then - MINDGRAPH_BIN="$(git -C "$ROOT" config --get mainframe.mindgraphBin 2>/dev/null || true)" -fi - -if [[ -z "$MINDGRAPH_BIN" ]]; then - if command -v mindgraph >/dev/null 2>&1; then - MINDGRAPH_BIN="$(command -v mindgraph)" - else - echo "MindGraph binary not found. Create mindgraph/.venv, set MINDGRAPH_BIN=/path/to/mindgraph, git config mainframe.mindgraphBin /path/to/mindgraph, or install mindgraph on PATH." >&2 - exit 1 - fi -fi - -if [[ ! -x "$MINDGRAPH_BIN" ]]; then - echo "MindGraph binary is not executable: $MINDGRAPH_BIN" >&2 - exit 1 -fi - -if [[ $# -eq 0 || "${1:-}" == "--help" || "${1:-}" == "-h" ]]; then - exec "$MINDGRAPH_BIN" "$@" -fi - -for arg in "$@"; do - if [[ "$arg" == "--help" || "$arg" == "-h" || "$arg" == "--db" ]]; then - exec "$MINDGRAPH_BIN" "$@" - fi -done - -case "$1" in - # doctor/status inspect dual ~/.mindgraph indexes by default — do not force --db. - doctor|status) - exec "$MINDGRAPH_BIN" "$@" - ;; - init|ingest|ingest-many|query|neighbors|serve-mcp) - mkdir -p "$(dirname "$DB_PATH")" - exec "$MINDGRAPH_BIN" "$@" --db "$DB_PATH" - ;; - *) - exec "$MINDGRAPH_BIN" "$@" - ;; -esac diff --git a/bin/mindgraph-audit-links b/bin/mindgraph-audit-links deleted file mode 100755 index 8b214be..0000000 --- a/bin/mindgraph-audit-links +++ /dev/null @@ -1,19 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -if ROOT="$(git rev-parse --show-toplevel 2>/dev/null)"; then - : -else - ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -fi - -PYTHON="${MAINFRAME_PYTHON:-}" -if [[ -z "$PYTHON" && -x "$ROOT/mindgraph/.venv/bin/python" ]]; then - PYTHON="$ROOT/mindgraph/.venv/bin/python" -fi -if [[ -z "$PYTHON" && -x "$ROOT/30_projects/mindgraph/workbench/.venv/bin/python" ]]; then - PYTHON="$ROOT/30_projects/mindgraph/workbench/.venv/bin/python" -fi -PYTHON="${PYTHON:-python3}" - -exec "$PYTHON" "$ROOT/scripts/mindgraph_link_audit.py" --root "$ROOT" "$@" diff --git a/bin/mindgraph-idle-lifecycle b/bin/mindgraph-idle-lifecycle deleted file mode 100755 index aa82112..0000000 --- a/bin/mindgraph-idle-lifecycle +++ /dev/null @@ -1,67 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -PLIST="$HOME/Library/LaunchAgents/com.user.mindgraph-daemon.plist" -STATE_DIR="$HOME/.mindgraph/run" -LABEL="com.user.mindgraph-daemon" -DOMAIN="gui/$(id -u)" - -usage() { - echo "usage: bin/mindgraph-idle-lifecycle activate --confirm-no-active-clients [IDLE_SECONDS]" >&2 - echo " bin/mindgraph-idle-lifecycle rollback PLIST_BACKUP" >&2 -} - -case "${1:-}" in - activate) - if [[ "${2:-}" != "--confirm-no-active-clients" ]]; then - echo "refusing: activation unloads the live daemon; confirm a verified idle window" >&2 - usage - exit 2 - fi - idle_seconds="${3:-900}" - if [[ ! "$idle_seconds" =~ ^[0-9]+$ ]] || (( idle_seconds < 60 )); then - echo "refusing: IDLE_SECONDS must be an integer >= 60" >&2 - exit 2 - fi - if [[ ! -f "$PLIST" ]]; then - echo "refusing: expected plist not found: $PLIST" >&2 - exit 2 - fi - if [[ -e "${PLIST}.disabled" ]]; then - echo "refusing: disabled plist target already exists: ${PLIST}.disabled" >&2 - exit 2 - fi - if ! /usr/libexec/PlistBuddy -c 'Print :KeepAlive' "$PLIST" 2>/dev/null | grep -qx true; then - echo "refusing: expected KeepAlive=true plist was not found" >&2 - exit 2 - fi - stamp="$(date -u +%Y%m%dT%H%M%SZ)" - backup="${PLIST}.keepalive-backup-${stamp}" - cp -p "$PLIST" "$backup" - launchctl bootout "$DOMAIN/$LABEL" - mv "$PLIST" "${PLIST}.disabled" - mkdir -p "$STATE_DIR" - printf '%s\n' "$idle_seconds" > "$STATE_DIR/idle-lifecycle.enabled" - chmod 600 "$STATE_DIR/idle-lifecycle.enabled" - echo "idle lifecycle activated; rollback backup: $backup" - ;; - rollback) - backup="${2:-}" - if [[ -z "$backup" || ! -f "$backup" ]]; then - echo "refusing: provide the exact backup path emitted by activate" >&2 - exit 2 - fi - rm -f "$STATE_DIR/idle-lifecycle.enabled" - if [[ -f "${PLIST}.disabled" ]]; then - mv "${PLIST}.disabled" "$PLIST" - else - cp -p "$backup" "$PLIST" - fi - launchctl bootstrap "$DOMAIN" "$PLIST" - echo "KeepAlive LaunchAgent restored from: $backup" - ;; - *) - usage - exit 2 - ;; -esac diff --git a/bin/mindgraph-live b/bin/mindgraph-live deleted file mode 100755 index 93dcdd5..0000000 --- a/bin/mindgraph-live +++ /dev/null @@ -1,203 +0,0 @@ -#!/usr/bin/env bash -# mindgraph-live: Procedure to query live knowledge data (workstation.sqlite) or finance-specific (fund.sqlite) -# using MindGraph-style fused lexical + semantic search where possible. -# -# This allows MindGraph-like access to volatile/live data without polluting the durable index. -# Usage examples: -# bin/mindgraph-live "AAPL price history or recent metrics" --scope finance -# bin/mindgraph-live "current active tasks or agent status" --scope live -# bin/mindgraph-live "GME strategy performance" --scope finance --json -# bin/mindgraph-live "recent runs or errors" --scope live -# -# Scopes: -# finance (default for finance queries): 20_live/markets/fund.sqlite + registry text -# live / workstation: 20_live/workstation/workstation.sqlite -# both: federated -# -# For full semantic, uses the same embedder as main MindGraph if available. -# Results include source, snippet, score, and trust note (volatile). - -set -euo pipefail - -ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -PYTHON_BIN="$ROOT/mindgraph/.venv/bin/python" -LIVE_DB_FINANCE="$ROOT/20_live/markets/fund.sqlite" -LIVE_DB_WORKSTATION="$ROOT/20_live/workstation/workstation.sqlite" -REGISTRY_DIR="$ROOT/20_live/markets/registry" - -SCOPE="finance" -QUERY="" -JSON_OUT=0 -TOP_K=5 - -while [[ $# -gt 0 ]]; do - case "$1" in - --scope) SCOPE="$2"; shift 2 ;; - --json) JSON_OUT=1; shift ;; - --top-k) TOP_K="$2"; shift 2 ;; - -h|--help) - echo "mindgraph-live \"your question\" [--scope finance|live|both] [--json] [--top-k N]" - exit 0 - ;; - *) QUERY="$1"; shift ;; - esac -done - -if [[ -z "$QUERY" ]]; then - echo "Usage: bin/mindgraph-live \"query text\" [options]" - echo "See --help" - exit 1 -fi - -if [[ ! -x "$PYTHON_BIN" ]]; then - echo "MindGraph python not found at $PYTHON_BIN" - exit 1 -fi - -$PYTHON_BIN - "$QUERY" "$SCOPE" "$JSON_OUT" "$TOP_K" "$LIVE_DB_FINANCE" "$LIVE_DB_WORKSTATION" "$REGISTRY_DIR" << 'PYEOF' -import sys -import json -import sqlite3 -import os -from pathlib import Path -from datetime import datetime - -query = sys.argv[1] -scope = sys.argv[2] -json_out = bool(int(sys.argv[3])) -top_k = int(sys.argv[4]) -db_finance = sys.argv[5] -db_work = sys.argv[6] -registry_dir = sys.argv[7] - -def load_text_rows(db_path, queries): - """Return list of (source, text, meta) from SQL queries that return text content.""" - rows = [] - if not Path(db_path).exists(): - return rows - try: - conn = sqlite3.connect(db_path) - for sql, label in queries: - try: - cur = conn.execute(sql) - for r in cur.fetchall(): - if r and r[0]: - text = str(r[0]) - meta = r[1] if len(r) > 1 else "" - rows.append((label, text[:2000], meta or "")) - except Exception: - pass - conn.close() - except Exception: - pass - return rows - -def collect_finance_texts(): - texts = [] - # tickers - texts += load_text_rows(db_finance, [ - ("SELECT symbol || ' ' || COALESCE(company_name,'') || ' ' || COALESCE(sector,'') || ' ' || COALESCE(notes,'') FROM tickers", "ticker") - ]) - # recent prices / history summary - texts += load_text_rows(db_finance, [ - ("SELECT 'Price ' || symbol || ' ' || price || ' ' || COALESCE(market_date,'') FROM prices LIMIT 50", "price") - ]) - # strategy metrics - texts += load_text_rows(db_finance, [ - ("SELECT 'Strategy ' || COALESCE(strategy,'') || ' ' || COALESCE(ticker,'') || ' oos:' || COALESCE(oos_return,'') || ' dsr:' || COALESCE(dsr_p,'') || ' ' || COALESCE(promotion_status,'') FROM strategy_metrics LIMIT 100", "metric") - ]) - # registry text files (hypotheses, runs, etc.) - for f in Path(registry_dir).glob("*.jsonl"): - try: - with open(f) as fh: - for line in fh: - if line.strip(): - try: - obj = json.loads(line) - text = json.dumps(obj)[:1500] - texts.append((f"registry/{f.name}", text, str(f))) - except: - pass - except: - pass - # reports - for md in Path(db_finance).parent.glob("reports/*.md"): - try: - with open(md) as fh: - texts.append(("report", fh.read()[:2000], str(md))) - except: - pass - return texts - -def collect_live_texts(): - texts = [] - # tasks, projects, agents from workstation - texts += load_text_rows(db_work, [ - ("SELECT title || ' ' || COALESCE(goal,'') || ' ' || COALESCE(next_action,'') FROM projects", "project"), - ("SELECT title || ' ' || COALESCE(description,'') FROM tasks", "task"), - ("SELECT name || ' ' || role || ' ' || status || ' ' || location FROM agent_profiles", "agent"), - ]) - # runs and events if text - texts += load_text_rows(db_work, [ - ("SELECT COALESCE(output,'') || ' ' || COALESCE(error,'') FROM runs LIMIT 30", "run"), - ]) - return texts - -def simple_rank(query, texts, top_k): - """Very lightweight lexical + length bias rank (no heavy deps if no embedder).""" - qlower = query.lower() - scored = [] - for src, txt, meta in texts: - score = 0.0 - tlower = txt.lower() - # lexical overlap - for w in qlower.split(): - if w and w in tlower: - score += 1.0 / (1 + tlower.count(w)) - # boost recent-ish or key fields (heuristic) - if "strategy" in src or "metric" in src: - score *= 1.2 - if score > 0: - scored.append((score, src, txt[:300].replace('\n',' '), meta)) - scored.sort(reverse=True) - return scored[:top_k] - -# Collect -all_texts = [] -if scope in ("finance", "both"): - all_texts += collect_finance_texts() -if scope in ("live", "workstation", "both"): - all_texts += collect_live_texts() - -if not all_texts: - print("No live data found for scope", scope) - sys.exit(0) - -results = simple_rank(query, all_texts, top_k) - -if json_out: - out = [] - for score, src, snippet, meta in results: - out.append({ - "score": round(score, 4), - "source": src, - "snippet": snippet, - "meta": meta, - "trust": "volatile-live", - "note": "Live data - promote summaries to 10_knowledge after review" - }) - print(json.dumps(out, indent=2)) -else: - print(f"Live MindGraph-style query (scope={scope}): {query}") - print(f"Found {len(results)} candidates (lexical rank; use --json for machine)\n") - for i, (score, src, snippet, meta) in enumerate(results, 1): - print(f"{i}. [{src}] score={score:.3f}") - print(f" {snippet[:200]}...") - if meta: - print(f" meta: {meta}") - print(" trust: volatile-live (not durable)\n") - -PYEOF -chmod +x bin/mindgraph-live -echo "Created bin/mindgraph-live" -ls -l bin/mindgraph-live \ No newline at end of file diff --git a/bin/mindgraph-projects-apply b/bin/mindgraph-projects-apply deleted file mode 100755 index 33ce6cd..0000000 --- a/bin/mindgraph-projects-apply +++ /dev/null @@ -1,12 +0,0 @@ -#!/usr/bin/env bash -# Staged apply for projects MindGraph index (ADR-045 / live-retention). -set -euo pipefail -ROOT="$(cd "$(dirname "$0")/.." && pwd)" -export PYTHONPATH="${ROOT}/scripts${PYTHONPATH:+:$PYTHONPATH}" -if [[ -x /opt/homebrew/bin/python3 ]]; then - exec /opt/homebrew/bin/python3 "$ROOT/scripts/mindgraph_projects_apply.py" "$@" -elif [[ -x /usr/bin/python3 ]]; then - exec /usr/bin/python3 "$ROOT/scripts/mindgraph_projects_apply.py" "$@" -else - exec python3 "$ROOT/scripts/mindgraph_projects_apply.py" "$@" -fi diff --git a/bin/mindgraph-refresh b/bin/mindgraph-refresh deleted file mode 100755 index e2d1840..0000000 --- a/bin/mindgraph-refresh +++ /dev/null @@ -1,75 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)" -SCOPE="${MAINFRAME_MINDGRAPH_SCOPE:-$ROOT/10_knowledge}" -DB_PATH="${MINDGRAPH_DB_PATH:-$HOME/.mindgraph/mainframe.sqlite}" -DRY_RUN=0 -AUDIT_LINKS=0 -AUDIT_OUTPUT_DIR="${MAINFRAME_MINDGRAPH_AUDIT_DIR:-$ROOT/01_ingest/audit-receipts}" - -usage() { - cat <<'EOF' -Usage: bin/mindgraph-refresh [--dry-run] [--audit-links] [--audit-output-dir DIR] - -Refresh the durable knowledge index. --audit-links runs the separate, -read-only advisory source audit after a successful refresh; findings never -block ingest and the audit never mutates SQLite or source notes. -EOF -} - -while [[ $# -gt 0 ]]; do - case "$1" in - --dry-run) DRY_RUN=1 ;; - --audit-links) AUDIT_LINKS=1 ;; - --audit-output-dir) - if [[ $# -lt 2 ]]; then - echo "--audit-output-dir requires a directory" >&2 - exit 2 - fi - AUDIT_OUTPUT_DIR="$2" - shift - ;; - --help|-h) usage; exit 0 ;; - *) echo "unknown option: $1" >&2; usage >&2; exit 2 ;; - esac - shift -done - -if [[ ! -d "$SCOPE" ]]; then - echo "MindGraph ingest scope does not exist: $SCOPE" >&2 - exit 1 -fi - -if [[ "$DRY_RUN" -eq 1 ]]; then - echo "DB: $DB_PATH" - echo "Scope: $SCOPE" - echo "Index ID: mainframe-knowledge" - echo "Trust profile: durable_knowledge" - echo "Commands:" - echo " $ROOT/bin/mindgraph init" - echo " $ROOT/bin/mindgraph ingest \"$SCOPE\" --index-id mainframe-knowledge --trust-profile durable_knowledge --namespace knowledge --source-root \"$SCOPE\" --display-prefix 10_knowledge" - if [[ "$AUDIT_LINKS" -eq 1 ]]; then - echo " $ROOT/bin/mindgraph-audit-links --dry-run --scope \"$SCOPE\" --output-dir \"$AUDIT_OUTPUT_DIR\"" - fi - exit 0 -fi - -mkdir -p "$(dirname "$DB_PATH")" - -if [[ ! -f "$DB_PATH" ]]; then - "$ROOT/bin/mindgraph" init -fi - -"$ROOT/bin/mindgraph" ingest "$SCOPE" \ - --index-id mainframe-knowledge \ - --trust-profile durable_knowledge \ - --namespace knowledge \ - --source-root "$SCOPE" \ - --display-prefix 10_knowledge - -if [[ "$AUDIT_LINKS" -eq 1 ]]; then - if ! "$ROOT/bin/mindgraph-audit-links" --dry-run --scope "$SCOPE" --output-dir "$AUDIT_OUTPUT_DIR"; then - echo "WARNING: advisory graph audit could not complete; refresh remains successful" >&2 - fi -fi diff --git a/bin/mindgraph-refresh-projects b/bin/mindgraph-refresh-projects deleted file mode 100755 index 9d2b919..0000000 --- a/bin/mindgraph-refresh-projects +++ /dev/null @@ -1,246 +0,0 @@ -#!/usr/bin/env bash -# Refresh the projects MindGraph index from coordination surfaces. -# Fail-closed CLI (Unit 2.1): --help is non-mutating; unknown flags rejected; -# flags are order-independent; live ingest requires explicit --apply. -set -euo pipefail - -ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)" -DB_PATH="${MINDGRAPH_DB_PATH:-$HOME/.mindgraph/mainframe-projects.sqlite}" -# Default lean manifest; --deep switches to archaeology profile unless env overrides. -PROJECT_MANIFEST="${MAINFRAME_MINDGRAPH_PROJECTS_MANIFEST:-$ROOT/30_projects/mindgraph-projects.json}" -DEEP_MANIFEST="$ROOT/30_projects/mindgraph-projects-deep.json" -DRY_RUN=0 -FULL=0 -DEEP=0 -APPLY=0 - -usage() { - cat <<'EOF' -Usage: bin/mindgraph-refresh-projects [options] - -Refresh (or preview) the projects MindGraph index from 30_projects/ -coordination surfaces. Trust profile: project_status (see HARNESS.md). - -Options (order-independent): - -h, --help Show this help and exit 0 (never mutates) - --dry-run Preview scopes and intended commands; no DB writes - --full Include every project directory (not just the curated list) - --deep Use mindgraph-projects-deep.json (includes outputs/**) - --apply Perform the live ingest (required for mutation) - -Default manifest (lean): README, AGENTS, log, decisions, methodology, plans — -no bulk outputs. Deep profile adds outputs/** for archaeology. - -Environment: - MINDGRAPH_DB_PATH Target SQLite DB - (default: ~/.mindgraph/mainframe-projects.sqlite) - MAINFRAME_MINDGRAPH_PROJECTS_MANIFEST Path to curated manifest JSON - MAINFRAME_MINDGRAPH_PROJECTS Space-separated project slug override - -Safety: - - Unknown flags fail closed (exit 2). - - Without --apply, this command will not mutate any database. - - Prefer --dry-run first. Point MINDGRAPH_DB_PATH at a temporary DB for experiments. - - Recommended: bin/mindgraph-projects-apply (plan → stage → promote). -EOF -} - -while [[ $# -gt 0 ]]; do - case "$1" in - -h|--help) - usage - exit 0 - ;; - --dry-run) - DRY_RUN=1 - shift - ;; - --full) - FULL=1 - shift - ;; - --deep) - DEEP=1 - shift - ;; - --apply) - APPLY=1 - shift - ;; - -*) - echo "error: unknown option: $1" >&2 - echo "hint: use --help for usage; mutation requires --apply" >&2 - exit 2 - ;; - *) - echo "error: unexpected argument: $1" >&2 - echo "hint: this command takes flags only; use --help" >&2 - exit 2 - ;; - esac -done - -if [[ "$DRY_RUN" -eq 1 && "$APPLY" -eq 1 ]]; then - echo "error: refuse --dry-run together with --apply; choose one" >&2 - exit 2 -fi - -if [[ "$DRY_RUN" -eq 0 && "$APPLY" -eq 0 ]]; then - echo "error: refusing to mutate without --apply" >&2 - echo "hint: preview with --dry-run, or pass --apply for live ingest" >&2 - echo "hint: set MINDGRAPH_DB_PATH to a temporary DB when testing" >&2 - exit 2 -fi - -# --deep selects the deep manifest only when the env did not already pin one -if [[ "$DEEP" -eq 1 ]]; then - if [[ -z "${MAINFRAME_MINDGRAPH_PROJECTS_MANIFEST:-}" ]]; then - PROJECT_MANIFEST="$DEEP_MANIFEST" - fi - if [[ ! -f "$PROJECT_MANIFEST" ]]; then - echo "error: deep manifest not found: $PROJECT_MANIFEST" >&2 - exit 2 - fi -fi - -TMP_MANIFEST="$(mktemp "${TMPDIR:-/tmp}/mindgraph-projects.XXXXXX.json")" -trap 'rm -f "$TMP_MANIFEST"' EXIT - -python3 - "$ROOT" "$TMP_MANIFEST" "$PROJECT_MANIFEST" "$FULL" "${MAINFRAME_MINDGRAPH_PROJECTS:-}" <<'PY' -import json -import sys -from pathlib import Path - -root = Path(sys.argv[1]).resolve() -out_path = Path(sys.argv[2]) -manifest_path = Path(sys.argv[3]).expanduser() -full = sys.argv[4] == "1" -override = sys.argv[5].strip() -projects_dir = root / "30_projects" - -default_projects = [ - "agent-harness-eval", - "agent-tracker-eval", - "local-agent", - "mindgraph", - "mindgraph-eval", - "mainframe-process-eval", -] -include = [ - "README.md", - "AGENTS.md", - "log.md", - "decisions.md", - "methodology-approach.md", - "plans/*.md", - "plans/**/*.md", -] -exclude = [ - ".git/*", - ".git/**/*", - ".venv/*", - ".venv/**/*", - "__pycache__/*", - "__pycache__/**/*", - ".pytest_cache/*", - ".pytest_cache/**/*", - "workbench/*", - "workbench/**/*", - "raw-materials/*", - "raw-materials/**/*", - "outputs/*", - "outputs/**/*", -] - -if full: - projects = sorted(p.name for p in projects_dir.iterdir() if p.is_dir()) -elif override: - projects = override.split() -elif manifest_path.exists(): - payload = json.loads(manifest_path.read_text()) - if not isinstance(payload, dict): - raise SystemExit(f"project manifest must be a JSON object: {manifest_path}") - projects = payload.get("projects", default_projects) - include = payload.get("include", include) - exclude = payload.get("exclude", exclude) -else: - projects = default_projects - -if not isinstance(projects, list) or not all(isinstance(p, str) for p in projects): - raise SystemExit("project manifest field 'projects' must be a string list") -if not isinstance(include, list) or not all(isinstance(p, str) for p in include): - raise SystemExit("project manifest field 'include' must be a string list") -if not isinstance(exclude, list) or not all(isinstance(p, str) for p in exclude): - raise SystemExit("project manifest field 'exclude' must be a string list") - -scopes = [] -for slug in projects: - project_path = projects_dir / slug - if not project_path.is_dir(): - print(f"warn: project scope missing, skipping: {project_path}", file=sys.stderr) - continue - scopes.append( - { - "namespace": slug, - "root": str(project_path), - "source_root": str(project_path), - "display_prefix": f"30_projects/{slug}", - } - ) - -if not scopes: - raise SystemExit("MindGraph projects ingest scope is empty.") - -out = { - "index_id": "mainframe-projects", - "trust_profile": "project_status", - "include": include, - "exclude": exclude, - "scopes": scopes, -} -out_path.write_text(json.dumps(out, indent=2) + "\n") -PY - -if [[ "$DRY_RUN" -eq 1 ]]; then - echo "DB: $DB_PATH" - echo "Trust profile: project_status (see HARNESS.md)" - echo "Mode: dry-run (no DB mutation)" - if [[ "$FULL" -eq 1 ]]; then - echo "Scope mode: full project directory list (include patterns from manifest/defaults)" - elif [[ -n "${MAINFRAME_MINDGRAPH_PROJECTS:-}" ]]; then - echo "Scope mode: MAINFRAME_MINDGRAPH_PROJECTS override" - else - echo "Scope mode: curated manifest$([ "$DEEP" -eq 1 ] && echo ' (deep)' || echo ' (lean/default)')" - echo "Manifest: $PROJECT_MANIFEST" - fi - echo "Scopes:" - python3 - "$TMP_MANIFEST" <<'PY' -import json -import sys -payload = json.load(open(sys.argv[1])) -for scope in payload["scopes"]: - print(f" {scope['namespace']}: {scope['root']}") -PY - echo "Includes:" - python3 - "$TMP_MANIFEST" <<'PY' -import json -import sys -payload = json.load(open(sys.argv[1])) -for pattern in payload["include"]: - print(f" {pattern}") -PY - echo "Commands (not run):" - echo " MINDGRAPH_DB_PATH=\"$DB_PATH\" $ROOT/bin/mindgraph init" - echo " MINDGRAPH_DB_PATH=\"$DB_PATH\" $ROOT/bin/mindgraph ingest-many <generated-project-manifest> --allow-failures" - echo "To apply against this DB: MINDGRAPH_DB_PATH=\"$DB_PATH\" $0 --apply${FULL:+ --full}" - exit 0 -fi - -# --apply path only below -mkdir -p "$(dirname "$DB_PATH")" - -if [[ ! -f "$DB_PATH" ]]; then - MINDGRAPH_DB_PATH="$DB_PATH" "$ROOT/bin/mindgraph" init -fi - -MINDGRAPH_DB_PATH="$DB_PATH" "$ROOT/bin/mindgraph" ingest-many "$TMP_MANIFEST" --allow-failures diff --git a/bin/mindgraph-repair-connectivity b/bin/mindgraph-repair-connectivity deleted file mode 100755 index 571b43a..0000000 --- a/bin/mindgraph-repair-connectivity +++ /dev/null @@ -1,191 +0,0 @@ -#!/usr/bin/env python3 -import sys -import re -import json -from pathlib import Path -from datetime import datetime - -# Whitelist of known domains to check (subdirectories of 10_knowledge) -_MAINFRAME_ROOT = Path(__file__).resolve().parent.parent -KNOWLEDGE_ROOT = _MAINFRAME_ROOT / "10_knowledge" - -def parse_frontmatter(content: str) -> tuple[dict[str, str], str]: - """Parse YAML frontmatter from markdown content.""" - meta = {} - body = content - if content.startswith("---"): - parts = content.split("---", 2) - if len(parts) >= 3: - yaml_content = parts[1] - body = parts[2] - for line in yaml_content.splitlines(): - if ":" in line: - key, val = line.split(":", 1) - key = key.strip() - val = val.strip().strip('"').strip("'") - meta[key] = val - return meta, body - -def find_atlas_file(domain_dir: Path) -> Path | None: - """Find the existing atlas/index file for a domain.""" - for path in domain_dir.glob("*.md"): - if path.name.startswith("."): - continue - # Check filename rules - name = path.name.lower() - if name == "index.md" or name.endswith("-index.md") or name.endswith("_index.md") or "atlas" in name or "source-map" in name: - if "llamaindex" not in name and "index-id" not in name: - return path - # Check frontmatter tags - try: - content = path.read_text(encoding="utf-8") - meta, _ = parse_frontmatter(content) - tags = meta.get("tags", "") - if "source-index" in tags or "atlas" in tags or "topic-index" in tags: - return path - except Exception: - pass - return None - -def repair_domain(domain_dir: Path, *, dry_run: bool = False) -> tuple[int, int, bool]: - """ - Check for missing links in a domain's files and repair its atlas note. - Returns (total_files, missing_files_count, was_repaired). - """ - domain = domain_dir.name - md_files = [p for p in domain_dir.glob("*.md") if not p.name.startswith(".") and p.name != ".gitkeep"] - - if not md_files: - return 0, 0, False - - atlas_path = find_atlas_file(domain_dir) - - # If no atlas exists, we will create a new one - is_new = atlas_path is None - if is_new: - today = datetime.now().strftime("%Y-%m-%d") - atlas_name = f"{today}__{domain}__note__{domain}-atlas.md" - atlas_path = domain_dir / atlas_name - - # Get all file stems (slugs) in the domain - stems = {p.stem: p for p in md_files if p != atlas_path} - - # Check which stems are already linked in the atlas file - linked_stems = set() - atlas_content = "" - - if not is_new and atlas_path.exists(): - atlas_content = atlas_path.read_text(encoding="utf-8") - # Search for file stems in the content - for stem in stems: - # Match wikilinks [[stem]] or simple filename strings - if stem in atlas_content: - linked_stems.add(stem) - - missing_stems = sorted(list(set(stems.keys()) - linked_stems)) - - if not missing_stems: - return len(md_files), 0, False - - print(f"Domain '{domain}': Found {len(missing_stems)} unlinked file(s).") - - # Build links block for the missing ones - links_lines = [] - for stem in missing_stems: - file_path = stems[stem] - try: - content = file_path.read_text(encoding="utf-8") - meta, _ = parse_frontmatter(content) - title = meta.get("title", stem) - links_lines.append(f"- [[{stem}]] — {title}") - print(f" * Adding: {stem} -> '{title}'") - except Exception: - links_lines.append(f"- [[{stem}]]") - print(f" * Adding: {stem}") - - links_block = "\n".join(links_lines) - - if is_new: - # Generate complete new atlas file - today = datetime.now().strftime("%Y-%m-%d") - new_content = f"""--- -title: "{domain.replace('-', ' ').title()} Topic and Source Index" -domain: "{domain}" -type: "note" -status: "synthesized" -source: "connectivity-repair-minion" -tags: ["{domain}", "source-index", "atlas", "needs-audit"] -links: [] -updated: "{today}" ---- - -# {domain.replace('-', ' ').title()} Topic and Source Index - -Central index of resources and notes for the `{domain}` domain. - -## Unsorted / Ingested Sources - -{links_block} -""" - if dry_run: - print(f" -> [dry-run] would create {atlas_path.name} with {len(missing_stems)} links") - else: - atlas_path.write_text(new_content, encoding="utf-8") - print(f" -> Created new atlas file at {atlas_path.name}") - else: - # Append to existing atlas file - today = datetime.now().strftime("%Y-%m-%d") - - # Check if the file ends with a newline - if not atlas_content.endswith("\n"): - atlas_content += "\n" - - append_content = f"\n\n## Unsorted / Ingested Sources (Appended {today})\n\n{links_block}\n" - if dry_run: - print(f" -> [dry-run] would append {len(missing_stems)} links to {atlas_path.name}") - else: - atlas_path.write_text(atlas_content + append_content, encoding="utf-8") - print(f" -> Appended {len(missing_stems)} links to {atlas_path.name}") - - return len(md_files), len(missing_stems), True - -def main(): - dry_run = "--dry-run" in sys.argv or "-n" in sys.argv - if "--help" in sys.argv or "-h" in sys.argv: - print( - "Usage: mindgraph-repair-connectivity [--dry-run|-n]\n" - "Append missing domain file stems into each 10_knowledge/*/index.md atlas.\n" - "Default applies writes. Prefer --dry-run first." - ) - return - - if not KNOWLEDGE_ROOT.is_dir(): - print(f"Error: {KNOWLEDGE_ROOT} is not a directory.", file=sys.stderr) - sys.exit(1) - - print("=== MindGraph Connectivity Repair Minion ===") - if dry_run: - print("(dry-run: no files will be written)\n") - else: - print() - - repaired_count = 0 - total_checked = 0 - - for subdir in sorted(KNOWLEDGE_ROOT.iterdir()): - if subdir.is_dir() and not subdir.name.startswith("."): - total_checked += 1 - try: - files_count, missing_count, repaired = repair_domain( - subdir, dry_run=dry_run - ) - if repaired: - repaired_count += 1 - except Exception as e: - print(f"Error checking domain {subdir.name}: {e}", file=sys.stderr) - - action = "Would repair" if dry_run else "Repaired/Created" - print(f"\nDone. Checked {total_checked} domains. {action} {repaired_count} atlas indexes.") - -if __name__ == "__main__": - main() diff --git a/bin/model-disk-status b/bin/model-disk-status deleted file mode 100755 index b669d90..0000000 --- a/bin/model-disk-status +++ /dev/null @@ -1,84 +0,0 @@ -#!/usr/bin/env bash -# model-disk-status — free-space gates + large model cache summary -# Policy: 30_projects/agent-harness-eval/plans/model-disk-budget-policy.md -set -euo pipefail - -DATA_VOL="/System/Volumes/Data" -INVENTORY="${MAINFRAME_ROOT:-$HOME/Desktop/MainFrame}/30_projects/agent-harness-eval/outputs/MODEL_DISK_INVENTORY.md" -POLICY="${MAINFRAME_ROOT:-$HOME/Desktop/MainFrame}/30_projects/agent-harness-eval/plans/model-disk-budget-policy.md" - -echo "=== Model disk status ===" -echo "Policy: $POLICY" -echo "Inventory: $INVENTORY" -echo - -# Free space on Data volume (GB available) -if df -g "$DATA_VOL" >/dev/null 2>&1; then - # macOS df -g: size used avail in GiB - read -r _ size used avail _ < <(df -g "$DATA_VOL" | tail -1 | awk '{print $1,$2,$3,$4}') - # Prefer 1024-block parse for free bytes - avail_kb=$(df -k "$DATA_VOL" | tail -1 | awk '{print $4}') - avail_gb=$(python3 -c "print(round($avail_kb/1024/1024, 1))") - used_pct=$(df -h "$DATA_VOL" | tail -1 | awk '{print $5}') -else - avail_gb="?" - used_pct="?" -fi - -echo "Data volume free: ${avail_gb} GB (capacity used ${used_pct})" -echo - -# Gate -gate="OK" -if python3 -c "import sys; sys.exit(0 if float('${avail_gb}')>=40 else 1)" 2>/dev/null; then - gate="OK (≥40 GB free)" -elif python3 -c "import sys; sys.exit(0 if float('${avail_gb}')>=25 else 1)" 2>/dev/null; then - gate="WATCH (25–40 GB) — no speculative multi-GB pulls" -elif python3 -c "import sys; sys.exit(0 if float('${avail_gb}')>=15 else 1)" 2>/dev/null; then - gate="SOFT BLOCK (15–25 GB) — no multi-GB pulls without free-up" -elif python3 -c "import sys; sys.exit(0 if float('${avail_gb}')>=5 else 1)" 2>/dev/null; then - gate="HARD BLOCK (5–15 GB) — stop downloads; incompletes only" -else - gate="EMERGENCY (<5 GB) — free space before evals/downloads" -fi -echo "Gate: $gate" -echo - -echo "=== Large caches ===" -for p in \ - "$HOME/.ollama" \ - "$HOME/.cache/huggingface" \ - "$HOME/Desktop/image-generation-lab/comfyui/models" -do - if [ -e "$p" ]; then - du -sh "$p" 2>/dev/null || true - else - echo "missing $p" - fi -done -echo - -echo "=== Ollama models ===" -if command -v ollama >/dev/null 2>&1; then - ollama list 2>/dev/null || true -else - echo "ollama not installed" -fi -echo - -echo "=== HF hub top entries ===" -if [ -d "$HOME/.cache/huggingface/hub" ]; then - du -sh "$HOME/.cache/huggingface/hub"/models--* 2>/dev/null | sort -hr | head -15 -else - echo "no HF hub cache" -fi -echo - -# Incomplete blobs (safe prune anytime) -inc=$(find "$HOME/.cache/huggingface/hub" -name '*.incomplete' 2>/dev/null | wc -l | tr -d ' ') -echo "HF incomplete blobs: $inc (safe to delete anytime; not evidence)" -echo - -echo "=== Reminder ===" -echo "Benchmark first → update MODEL_DISK_INVENTORY.md → drop_candidate only after phase evidence." -echo "Do not dual-load Ollama + heavy MLX server when free < 25 GB." diff --git a/bin/model-memory-status b/bin/model-memory-status deleted file mode 100755 index d31e063..0000000 --- a/bin/model-memory-status +++ /dev/null @@ -1,218 +0,0 @@ -#!/usr/bin/env python3 -"""model-memory-status — unified-memory pressure + dual-stack watch. - -Companion to bin/model-disk-status. -Exit codes: - 0 OK (no dual stack; memory gate OK or idle) - 1 WATCH (memory tight — do not start second stack) - 2 SOFT BLOCK (free more memory before another load) - 3 HARD BLOCK (critically low free memory) - 4 DUAL STACK (Ollama model loaded AND mlx_lm.server up) - -Policy: 30_projects/agent-harness-eval/plans/model-disk-budget-policy.md -""" -from __future__ import annotations - -import json -import re -import subprocess -import sys -import urllib.error -import urllib.request -from pathlib import Path - - -def _run(cmd: list[str], timeout: float = 5.0) -> str: - try: - return subprocess.check_output(cmd, text=True, timeout=timeout, stderr=subprocess.DEVNULL) - except (subprocess.CalledProcessError, FileNotFoundError, subprocess.TimeoutExpired): - return "" - - -def mem_total_gb() -> float: - out = _run(["sysctl", "-n", "hw.memsize"]).strip() - try: - return int(out) / 1e9 - except ValueError: - return 0.0 - - -def vm_stats() -> dict[str, float]: - """Return selected vm_stat quantities in GB.""" - out = _run(["vm_stat"]) - page = 4096 - m = re.search(r"page size of (\d+)", out) - if m: - page = int(m.group(1)) - pages: dict[str, int] = {} - for line in out.splitlines(): - if ":" not in line: - continue - k, v = line.split(":", 1) - v = v.strip().rstrip(".").replace(",", "") - if v.isdigit(): - pages[k.strip()] = int(v) - - def gb(key: str) -> float: - return pages.get(key, 0) * page / 1e9 - - compressor = 0.0 - for k, n in pages.items(): - if "compressor" in k.lower() and "occupied" in k.lower(): - compressor = n * page / 1e9 - break - - free = gb("Pages free") + gb("Pages speculative") - inactive = gb("Pages inactive") - purgeable = gb("Pages purgeable") - return { - "free_spec": free, - "inactive": inactive, - "active": gb("Pages active"), - "wired": gb("Pages wired down"), - "compressor": compressor, - "purgeable": purgeable, - "avail_ish": free + inactive + purgeable, - } - - -def http_json(url: str, timeout: float = 2.0) -> dict | list | None: - try: - with urllib.request.urlopen(url, timeout=timeout) as resp: - return json.loads(resp.read().decode("utf-8")) - except (urllib.error.URLError, TimeoutError, json.JSONDecodeError, OSError): - return None - - -def ollama_loaded() -> list[dict]: - data = http_json("http://127.0.0.1:11434/api/ps") - if not isinstance(data, dict): - return [] - return list(data.get("models") or []) - - -def mlx_server_up() -> tuple[bool, list[str], float | None]: - """Return (up, model_ids, rss_gb).""" - data = http_json("http://127.0.0.1:8080/v1/models") - ids: list[str] = [] - if isinstance(data, dict): - ids = [str(x.get("id", "")) for x in (data.get("data") or [])] - # pgrep process RSS (KB on macOS ps -o rss) - rss_gb = None - ps = _run(["ps", "-axo", "rss,command"]) - for line in ps.splitlines(): - if "mlx_lm.server" in line and "rg" not in line: - parts = line.strip().split(None, 1) - if parts and parts[0].isdigit(): - rss_gb = int(parts[0]) / 1024 / 1024 - break - up = data is not None or rss_gb is not None - return up, ids, rss_gb - - -def pgrep_running(pattern: str) -> bool: - # Use pgrep if available - out = _run(["pgrep", "-f", pattern]) - return bool(out.strip()) - - -def main() -> int: - print("=== Model memory status ===\n") - - total = mem_total_gb() - st = vm_stats() - print(f"hw.memsize: {total:.1f} GB") - print(f"free+spec: {st['free_spec']:.1f} GB") - print(f"inactive: {st['inactive']:.1f} GB") - print(f"active: {st['active']:.1f} GB") - print(f"wired: {st['wired']:.1f} GB") - print(f"compressor: {st['compressor']:.1f} GB") - print(f"purgeable: {st['purgeable']:.1f} GB") - print( - f"avail-ish*: {st['avail_ish']:.1f} GB " - "(*free+inactive+purgeable; watch metric, not exact API)" - ) - - avail = st["avail_ish"] - if avail >= 8: - mem_gate, mem_code = "OK", 0 - elif avail >= 4: - mem_gate, mem_code = "WATCH — avoid starting a second heavy model stack", 1 - elif avail >= 2: - mem_gate, mem_code = "SOFT BLOCK — unload something before another load", 2 - else: - mem_gate, mem_code = "HARD BLOCK — free memory before evals", 3 - print(f"\nMemory gate: {mem_gate}") - print(f"(memory exit would be {mem_code}: 0=OK 1=WATCH 2=SOFT 3=HARD)\n") - - print("=== Loaded inference stacks ===") - ollama = ollama_loaded() - if ollama: - print("Ollama LOADED:") - for m in ollama: - size = m.get("size_vram") or m.get("size") or 0 - print(f" - {m.get('name')} ~{float(size)/1e9:.1f} GB") - else: - # distinguish unreachable vs idle - probe = http_json("http://127.0.0.1:11434/api/tags") - if probe is None: - print("Ollama: not reachable on :11434") - else: - print("Ollama: idle (no models loaded)") - - mlx_up, mlx_ids, mlx_rss = mlx_server_up() - if mlx_up: - print("MLX server: UP on :8080") - prefer = [ - i - for i in mlx_ids - if any(k in i.lower() for k in ("gemma", "mistral", "llama", "qwen", "phi")) - ] - show = (prefer or mlx_ids)[:8] - if show: - print(" models (sample):", ", ".join(show)) - if mlx_rss is not None: - print(f" mlx_lm.server RSS ≈ {mlx_rss:.1f} GB") - else: - print("MLX server: down (no response on :8080)") - - if pgrep_running("agent_harness_eval.coding_backend"): - print("coding_backend: RUNNING") - if pgrep_running("aider --model") or pgrep_running("/aider"): - print("aider: RUNNING") - - print("\n=== Dual-stack rule ===") - dual = bool(ollama) and mlx_up - if dual: - print("FAIL: Ollama model loaded AND mlx_lm.server up — dual stack.") - print(" Unload Ollama example:") - name = (ollama[0].get("name") if ollama else "MODEL") or "MODEL" - print( - " curl -s http://127.0.0.1:11434/api/generate " - f"-d '{{\"model\":\"{name}\",\"prompt\":\"\",\"keep_alive\":0}}'" - ) - print(" Stop MLX: pkill -f mlx_lm.server") - print("\nRule: never leave Ollama + mlx_lm.server both holding multi-GB models.") - print("Fair H1: Ollama half → keep_alive=0 → MLX half → stop server.") - return 4 - - if ollama: - print("OK: Ollama-only load") - elif mlx_up: - print("OK: MLX-server-only load") - else: - print("OK: no heavy stacks detected") - - print("\nRule: never leave Ollama + mlx_lm.server both holding multi-GB models.") - print("Fair H1: Ollama half → keep_alive=0 → MLX half → stop server.") - print("Disk: bin/model-disk-status") - - # Prefer dual-stack code if both; else memory gate - return mem_code if not dual else 4 - - -if __name__ == "__main__": - try: - raise SystemExit(main()) - except BrokenPipeError: - raise SystemExit(0) diff --git a/bin/papercut b/bin/papercut deleted file mode 100755 index 75f94d0..0000000 --- a/bin/papercut +++ /dev/null @@ -1,753 +0,0 @@ -#!/usr/bin/env python3 -"""Record friction encountered during work, and harvest it into fixes. - -## The gap this fills - -MainFrame's convention audit (`bin/contract-audit`) reads the repo. Every one of -its checks looks at the same files with the same assumptions, which is a single -detector — and this investigation has now watched a single detector under-count -four separate times. The audit's own write-up names the missing axis: - - L6 - No independent second opinion. Cross-checking a contract against how - agents actually behaved would be the independent axis, and it does not - exist yet. - -This is that axis. Not another read of the contracts: a record of what actually -happened while working under them. - -## Why agents do not report friction on their own - -They push through. A dead-end tool call, a flag that silently did the wrong -thing, a documented path that does not exist — all of it gets worked around -in-context and never surfaces, because the task still completed. The work looks -clean from the outside and the friction is invisible to everyone who was not -watching the reasoning. - -Physical evidence, found 2026-08-10: a 24 KB file named `--fail-on-overflow` -sitting in the repo root since 2026-08-04. It is an HTML render of -`cameron-originated-prospect-nominations-priority-a-2026-08-04.md`. Something -passed `--fail-on-overflow` where an output path belonged, the renderer wrote -its output to a file named after the flag, and whoever hit it re-ran with the -right arguments and moved on. Six days and an unknown number of `git status` -calls later, nobody had recorded it. - -That is not a rare event. That is the normal event, and it is invisible. - -## Capture is trivial. The harvest is the product. - -The failure mode this tool must avoid is the one MainFrame demonstrably has: -records that are written and never read. The 2026-06-22 architecture checklist -sat unrun for seven weeks and named the exact defect that then caused a -107-capture fabrication incident. `needs-audit` tags were applied faithfully and -were invisible at query time because MindGraph returns body chunks. A papercut -log with no consumption path is that failure a third time. - -So `harvest` is the real command. It clusters repeats, because **one papercut is -noise and the same papercut three times is a defect with an address** — which is -precisely checklist item 10.1, "repeated agent mistake -> add or update closest -AGENTS.md, workflow, skill, hook, or script". That item is unautomatable today -only because nothing records agent mistakes. - -## Scope: what, who, when - -Three lanes. They differ in who writes them, when, and what the harvest is -allowed to conclude from a repeat. - - LANE WHO WHEN REPEAT MEANS - auto the hook at the failure an address worth explaining - observed any agent when noticed, or a defect with an address; - prompted at close graduates to a fix - suggestion anyone any time one opinion, held more - than once. Never graduates. - -**auto** is the only lane that does not depend on anyone remembering anything. -`PostToolUseFailure` fires, a stub is written, nobody is asked. It captures the -address and cannot capture the reason. - -**observed** is the lane with the value in it, and the one that needs a -judgement call: - - RECORD a tool that failed in a way the docs did not predict - a documented path, flag, or command that does not work as written - a contract rule that could not be followed as literally stated - a dead end that cost real time and would cost the next agent the same - - DO NOT your own mistakes that the system handled correctly - anything you can fix in under a minute. Fix it instead - - The bar is: **would the next agent hit this too?** If not, it is not one. - -**suggestion** is deliberately unbounded. "There is a better way to do this" is -welcome even with nothing broken behind it, because the alternative is that the -observation nobody could quite justify gets written nowhere. It lives in the -same store and behind the same command so there is one habit and one place to -look, and it is quarantined from the graduation count so an opinion can never -masquerade as evidence: - - bin/papercut --kind suggestion "workflows should all open with a trigger" - -That split is MainFrame's own epistemic line. An observation is a thing that -happened; a suggestion is a hypothesis about what would be better. Three -recordings of the first are three findings. Three recordings of the second are -still one opinion. - -## Honest limits - -Self-reported friction is unfalsifiable, the same way a self-asserted -`retrieved_at` is. The difference is stakes: a `retrieved_at` becomes a citation -that downstream work trusts, whereas a papercut is a triage signal that gets -verified before anything is built on it. Low cost if wrong. It is evidence about -where to look, never evidence that something is true. - -Records land in `20_live/papercuts/` — volatile, gitignored, append-only. This is -telemetry, not knowledge. It must never be written into `10_knowledge/`. - -Usage: - bin/papercut "yarn web:test needs a workspace-relative path; root-relative finds no files" - bin/papercut --where bin/capture-validate --kind broken "crashes on paths outside the repo" - bin/papercut harvest # clusters, repeats, graduation candidates - bin/papercut harvest --since 2026-08 - bin/papercut list --limit 20 - bin/papercut harvest --json out.json -""" - -from __future__ import annotations - -import argparse -import json -import os -import re -import sys -from collections import Counter, defaultdict -from datetime import datetime -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] -STORE = ROOT / "20_live" / "papercuts" - -# A cluster this size stops being noise and becomes a defect with an address. -# Deliberately low: the cost of investigating a false repeat is minutes, and the -# cost of a real one going unnoticed is however long it has already been going -# unnoticed. -GRADUATION_THRESHOLD = 2 - -# Max auto-stubs per (tool, command-head) per day, so a flapping command cannot -# bury the hand-written entries. Announced in the harvest, never applied silently. -AUTO_CAP = 3 - -# Two populations, split on MainFrame's own epistemic line. -# -# OBSERVATION kinds record that something happened. They are evidence, so a -# repeat is a defect with an address and can graduate to a fix on its own. -# -# The SUGGESTION kind records that something could be better. That is a -# hypothesis, and a hypothesis repeated three times is still one opinion, not -# three findings. It never auto-graduates and never counts toward a graduation -# threshold, because letting opinion into that count would destroy the only -# signal the harvest has. -# -# Same store and same command on purpose: one habit, one place to look. If -# suggestions lived somewhere else they would be written nowhere. -OBSERVATION_KINDS = { - "friction": "worked, but cost more than it should have", - "dead-end": "a path the docs suggested that goes nowhere", - "broken": "a tool or command that failed outright", - "stale": "documentation or a contract that no longer matches reality", - "surprise": "behaved correctly but not as any document described", - "contract-gap": "a rule that could not be followed as literally written", - "auto": "recorded by the PostToolUseFailure hook, no agent involved", -} - -JUDGEMENT_KINDS = { - "suggestion": "there is a better way to do this. An opinion, and welcome", -} - -KINDS = {**OBSERVATION_KINDS, **JUDGEMENT_KINDS} - -STOPWORDS = { - "the", "a", "an", "and", "or", "but", "is", "are", "was", "were", "be", - "to", "of", "in", "on", "at", "for", "with", "it", "this", "that", "i", - "so", "not", "no", "if", "when", "then", "than", "as", "by", "from", - "would", "could", "should", "does", "did", "has", "have", "had", "its", -} - - -def store_path(when: datetime) -> Path: - return STORE / f"{when:%Y-%m}.jsonl" - - -def record(args: argparse.Namespace) -> int: - # Local time with offset, matching `logged_at` in 20_live/workflow-metrics/. - # - # The first version stamped UTC. Correlating a papercut with the tool failure - # that caused it is the entire point of `papercut prompt`, and after 20:00 - # Eastern the two systems disagree about what day it is — so ten papercuts - # written this evening sat in the same file as failures they were about and - # matched none of them. Two clocks in two formats is not a correlation key. - now = datetime.now().astimezone() - STORE.mkdir(parents=True, exist_ok=True) - entry = { - "ts": now.isoformat(timespec="seconds"), - "client": args.client or os.environ.get("MAINFRAME_CLIENT", "unknown"), - "session": os.environ.get("CLAUDE_SESSION_ID", "")[:8], - "kind": args.kind, - "where": args.where or "", - "message": args.message.strip(), - } - if not entry["message"]: - print("papercut: nothing to record", file=sys.stderr) - return 2 - - path = store_path(now) - with path.open("a", encoding="utf-8") as fh: - fh.write(json.dumps(entry, ensure_ascii=False) + "\n") - - print(f"recorded ({entry['kind']}) -> {path.relative_to(ROOT)}") - return 0 - - -# Set by load(). The reader's exit: a count that silently excluded records is -# indistinguishable from a count that had nothing to exclude, and the smaller -# number always looks like better news. -LOAD_SKIPPED = 0 -STORE_MISSING = False - - -def load(since: str = "") -> list[dict]: - global LOAD_SKIPPED, STORE_MISSING - LOAD_SKIPPED = 0 - STORE_MISSING = not STORE.exists() - if STORE_MISSING: - return [] - out: list[dict] = [] - for p in sorted(STORE.glob("*.jsonl")): - if since and p.stem < since[:7]: - continue - try: - lines = p.read_text(encoding="utf-8", errors="replace").splitlines() - except OSError as exc: - LOAD_SKIPPED += 1 - print(f"papercut: could not read {p.name}: {exc}", file=sys.stderr) - continue - for line in lines: - line = line.strip() - if not line: - continue - try: - out.append(json.loads(line)) - except json.JSONDecodeError: - # One malformed line must not hide the rest, and must not vanish - # from the count either. Silently dropping records is how a log - # becomes reassuring. - LOAD_SKIPPED += 1 - print(f"papercut: skipped a malformed line in {p.name}", file=sys.stderr) - return out - - -def signature(message: str) -> frozenset[str]: - """Content words, for clustering near-duplicates written by different agents. - - Deliberately crude. The job is to notice that three agents hit the same wall - in three different phrasings, not to do topic modelling. A cluster is a - prompt to go look, and a human looks at the raw messages anyway. - """ - words = re.findall(r"[a-z0-9_./-]{3,}", message.lower()) - return frozenset(w for w in words if w not in STOPWORDS) - - -def cluster(entries: list[dict]) -> list[list[dict]]: - clusters: list[list[dict]] = [] - sigs: list[frozenset[str]] = [] - for e in entries: - s = signature(e["message"]) - placed = False - for i, other in enumerate(sigs): - overlap = len(s & other) - smaller = min(len(s), len(other)) or 1 - if overlap / smaller >= 0.5: - clusters[i].append(e) - sigs[i] = s & other or s - placed = True - break - if not placed: - clusters.append([e]) - sigs.append(s) - return clusters - - -def harvest(args: argparse.Namespace) -> int: - all_entries = load(args.since) - - # Split before anything else. Suggestions are hypotheses and must never - # reach a graduation count: three people wanting the same change is one - # opinion held three times, not three findings, and mixing them would make - # the only signal this tool produces unreadable. - entries = [e for e in all_entries if e.get("kind") not in JUDGEMENT_KINDS] - ideas = [e for e in all_entries if e.get("kind") in JUDGEMENT_KINDS] - - # Auto stubs know an address but no reason. Surfaced separately so the - # short list of "repeated and still unexplained" is obvious at a glance. - autos = [e for e in entries if e.get("kind") == "auto"] - undiagnosed: Counter = Counter( - (e.get("where", ""), e.get("head", "")) for e in autos - ) - if not all_entries: - # Three different worlds, and they must not print the same sentence. - if STORE_MISSING: - print(f"No store at {STORE.relative_to(ROOT)}. Nothing has ever been recorded,") - print("and nothing is wired to record it. This is a setup problem, not a result.") - return 2 - if LOAD_SKIPPED: - print(f"The store exists but {LOAD_SKIPPED} record(s) could not be read,") - print("and nothing readable remains. Do not read this as a clean log.") - return 2 - print("Store present and empty. No papercuts recorded.") - print() - print("That is not the same as no friction. It is more likely that nothing") - print("is calling `bin/papercut`, which is the state this tool exists to end.") - return 0 - - if LOAD_SKIPPED: - print(f"WARNING: {LOAD_SKIPPED} record(s) unreadable and excluded from every") - print("count below. The numbers are a floor, not a total.\n") - - groups = sorted(cluster(entries), key=len, reverse=True) - repeats = [g for g in groups if len(g) >= GRADUATION_THRESHOLD] - singles = [g for g in groups if len(g) < GRADUATION_THRESHOLD] - - print(f"{len(entries)} papercut(s) across {len(groups)} distinct issue(s)") - print(f"span: {entries[0]['ts'][:10]} .. {entries[-1]['ts'][:10]}") - - by_kind = Counter(e["kind"] for e in entries) - print("by kind: " + ", ".join(f"{k} {v}" for k, v in by_kind.most_common())) - - where = Counter(e["where"] for e in entries if e["where"]) - if where: - print("most-implicated: " + ", ".join(f"{k} ({v})" for k, v in where.most_common(5))) - - print() - print("=" * 78) - print(f"GRADUATION CANDIDATES — seen {GRADUATION_THRESHOLD}+ times") - print("=" * 78) - if not repeats: - print(" none yet. Every papercut so far is a one-off.") - for g in repeats: - addr = Counter(e["where"] for e in g if e["where"]).most_common(1) - target = addr[0][0] if addr else "(no path recorded)" - clients = sorted({e["client"] for e in g}) - print(f"\n [{len(g)}x] {target}") - print(f" kinds: {', '.join(sorted({e['kind'] for e in g}))}" - f" clients: {', '.join(clients)}") - for e in g[:4]: - print(f" {e['ts'][:10]} {e['message'][:88]}") - if len(g) > 4: - print(f" ... and {len(g) - 4} more") - print(f" -> checklist 10.1: this wants a fix in the closest AGENTS.md,") - print(f" workflow, skill, hook, or script — not another note.") - - # Second graduation axis: the same address, hit by *different* problems. - # Message clustering only catches "we keep tripping on the same thing". Two - # unrelated bugs in one tool is the other signal — the tool is the defect, - # not either bug. Both `bin/papercut` entries below are distinct bugs, and - # counting only message-similarity would have called them one-offs forever. - by_addr: dict[str, list[dict]] = defaultdict(list) - for e in entries: - if e["where"]: - by_addr[e["where"]].append(e) - # A hot address means several *diagnosed* problems in one place. Auto stubs - # carry no diagnosis, and their messages are terse enough that "timed out" - # and "timed out after 120s" score as two different problems when they are - # one problem twice. Counting them here inflated `Bash` to three distinct - # issues that were all the same timeout. Auto entries have their own - # section; they do not get to vote in this one. - hot = {} - for addr, es in by_addr.items(): - diagnosed = [e for e in es if e.get("kind") != "auto"] - if len(diagnosed) >= GRADUATION_THRESHOLD and len( - {frozenset(signature(x["message"])) for x in diagnosed} - ) > 1: - hot[addr] = diagnosed - if hot: - print() - print("=" * 78) - print(f"HOT ADDRESSES — {GRADUATION_THRESHOLD}+ *distinct* problems in one place") - print("=" * 78) - for addr, es in sorted(hot.items(), key=lambda kv: -len(kv[1])): - print(f"\n [{len(es)} distinct] {addr}") - for e in es[:4]: - print(f" {e['kind']:<13} {e['message'][:76]}") - print(" -> the address is the defect, not any single entry.") - - if undiagnosed: - repeated = [(k, n) for k, n in undiagnosed.most_common() - if n >= GRADUATION_THRESHOLD] - if repeated: - print() - print("=" * 78) - print("AUTO-CAUGHT, REPEATED, STILL UNEXPLAINED") - print("=" * 78) - print(" The hook knows these failed more than once. It cannot know why.") - print(" A sentence each turns an address into a fixable defect.\n") - for (tool, head), n in repeated: - label = f"{tool}" + (f" ({head}…)" if head else "") - capped = " [at daily cap, true count is higher]" if n >= AUTO_CAP else "" - print(f" {n:>3}x {label}{capped}") - print(f'\n bin/papercut --kind broken --where <tool> "why this matters"') - - if ideas: - print() - print("-" * 78) - print(f"SUGGESTIONS ({len(ideas)}) — opinions, not findings") - print("-" * 78) - print(" Never graduate on their own. A repeated suggestion is one opinion") - print(" held more than once, and needs a person to agree with it.\n") - for e in ideas[-10:]: - w = f" [{e['where']}]" if e.get("where") else "" - print(f" {e['ts'][:10]}{w}") - print(f" {e['message'][:88]}") - if len(ideas) > 10: - print(f" ... and {len(ideas) - 10} more") - - if singles and not args.repeats_only: - print() - print("-" * 78) - print(f"ONE-OFFS ({len(singles)}) — watch, do not act yet") - print("-" * 78) - for g in singles[:15]: - e = g[0] - w = f" [{e['where']}]" if e["where"] else "" - print(f" {e['ts'][:10]} {e['kind']:<13}{w}") - print(f" {e['message'][:92]}") - if len(singles) > 15: - print(f" ... and {len(singles) - 15} more (use --json for all)") - - if args.json: - Path(args.json).write_text(json.dumps({ - "as_of": datetime.now().astimezone().isoformat(timespec="seconds"), - "total": len(entries), - "distinct": len(groups), - "graduation_candidates": [ - {"count": len(g), "where": (Counter(e["where"] for e in g if e["where"]) - .most_common(1) or [("", 0)])[0][0], - "messages": [e["message"] for e in g]} - for g in repeats - ], - }, indent=2), encoding="utf-8") - print(f"\nwrote {args.json}") - - # A harvest receipt is what makes checklist 10.1 machine-checkable: the audit - # can ask whether harvesting is happening, rather than trusting that it is. - receipt = STORE / "last-harvest.json" - receipt.write_text(json.dumps({ - "harvested_at": datetime.now().astimezone().isoformat(timespec="seconds"), - # ALL entries, including suggestions. `entries` here is the observation - # subset, and using it left every suggestion looking permanently - # unharvested to contract-audit 10.1, which counts raw lines. Two - # denominators for the same log is how a staleness check cries wolf - # until nobody reads it. - "entries_seen": len(all_entries), - "observations_seen": len(entries), - "suggestions_seen": len(ideas), - "unreadable_skipped": LOAD_SKIPPED, - "distinct_issues": len(groups), - "graduation_candidates": len(repeats), - }, indent=2), encoding="utf-8") - - return 0 - - -def auto(args: argparse.Namespace) -> int: - """PostToolUseFailure hook. Record that a tool failed, with no agent involved. - - ## Why this is not optional - - Everything else in this file depends on an agent choosing to write something - down, and the whole reason the tool exists is that they reliably do not. The - moment of friction is mid-task with a workaround already in mind. Asking for - a note there is a T0 advisory, and this repo measured what those are worth: - 613 normative clauses, 6% enforcement. - - So the mechanical half is taken away from the agent entirely. A failed tool - call writes its own stub. Who: nobody. When: at the moment of failure. - - ## What a stub is and is not - - An auto entry records the **address**: which tool, what the command started - with, when. It cannot record the diagnosis, because the hook payload does not - contain one and `20_live/workflow-metrics/` is metadata-only by contract - (item 9.1). `run_terminal_command failed, head "set"` will never tell you - that `grep` here is a ugrep wrapper honouring .gitignore. - - That is the division of labour, and it is deliberate: - - auto THAT it failed. Free, objective, complete. - agent WHY it matters. Costly, subjective, and the part with value. - - `harvest` shows auto stubs that repeat and still carry no diagnosis. Those - are the ones worth a human sentence. The agent is handed a short list at - session close instead of being asked to remember anything. - - ## Flood control - - Capped at AUTO_CAP entries per (tool, command-head) per day. A flapping - command would otherwise bury every hand-written papercut, and a log nobody - can read is the failure mode this tool was built to avoid. The cap is - announced in the harvest rather than applied silently, because a silent cap - reads as "that only happened three times". - - Fails open and silent. A hook that breaks the session to record a complaint - about a broken session is not a trade worth making. - """ - try: - payload = json.load(sys.stdin) - except Exception: # noqa: BLE001 - return 0 - - tool = payload.get("tool_name") or payload.get("toolName") or "" - if not tool: - return 0 - - tin = payload.get("tool_input") or payload.get("toolInput") or {} - head = "" - for key in ("command", "file_path", "path", "pattern", "url"): - val = tin.get(key) - if isinstance(val, str) and val.strip(): - head = val.strip().split()[0][:48] - break - - err = payload.get("tool_response") or payload.get("error") or "" - if isinstance(err, dict): - err = err.get("error") or err.get("stderr") or "" - err = re.sub(r"\s+", " ", str(err)).strip()[:180] - - now = datetime.now().astimezone() - today = now.strftime("%Y-%m-%d") - same = [ - e for e in load(today[:7]) - if e.get("kind") == "auto" and e.get("where") == tool - and e.get("head", "") == head and e.get("ts", "").startswith(today) - ] - if len(same) >= AUTO_CAP: - return 0 - - STORE.mkdir(parents=True, exist_ok=True) - entry = { - "ts": now.isoformat(timespec="seconds"), - "client": os.environ.get("MAINFRAME_CLIENT", "unknown"), - "session": os.environ.get("CLAUDE_SESSION_ID", "")[:8], - "kind": "auto", - "where": tool, - "head": head, - "message": err or f"{tool} failed with no error text captured", - "diagnosed": False, - } - with store_path(now).open("a", encoding="utf-8") as fh: - fh.write(json.dumps(entry, ensure_ascii=False) + "\n") - return 0 - - -def prompt(args: argparse.Namespace) -> int: - """Show today's tool failures that carry no explanation. - - ## Why this exists rather than "remember to run bin/papercut" - - The moment of friction is not the moment of reflection. An agent that hits a - broken tool is mid-task, has a workaround in mind, and will take it. That is - the whole observation behind this tool: models push through and never mention - the problem. A convention that asks them to stop and log at exactly the worst - moment is a T0 advisory, and this repo has 613 normative clauses and a - measured 6% enforcement rate to show what those are worth. - - So do not rely on remembering. `20_live/workflow-metrics/` already records - every `PostToolUseFailure` automatically, and `bin/session-close` already - runs at SessionEnd and PreCompact, when the agent is reflecting anyway. - This command joins the two: it hands over the list of things that broke, so - the agent supplies diagnosis rather than recall. - - ## What each layer can and cannot do - - workflow-metrics THAT something failed. automatic, objective, high - volume, metadata-only by contract (item 9.1), so it - records `run_terminal_command failed, command_head - "set"` and can never tell you why. - - bin/papercut WHAT was wrong and why it will bite the next agent. - Semantic, low volume, self-reported. - - Neither derives from the other. Telemetry cannot know that `grep` here is a - ugrep wrapper honouring .gitignore; a papercut cannot know that the same - command failed eleven times this week. - - ## This is a prompt, not a gate - - Most tool failures deserve no papercut. A typo, a path that did not exist - yet, a deliberate probe: all expected, none worth recording. Demanding an - entry per failure would flood the log and make the harvest useless, which is - the failure mode that matters more than under-recording. The list is offered, - not enforced. - - Correlation is coarse and stated as such: papercuts carry no telemetry id, so - matching is by tool name and day. It surfaces candidates. It proves nothing. - """ - events_dir = ROOT / "20_live" / "workflow-metrics" / "events" - if not events_dir.exists(): - print("No telemetry to read.") - return 0 - - day = args.since or datetime.now().strftime("%Y-%m-%d") - files = [p for p in sorted(events_dir.glob("*.jsonl")) if p.stem >= day[:10]] - if not files: - print(f"No telemetry on or after {day}.") - return 0 - - failures: Counter = Counter() - total = 0 - for p in files: - for ln in p.read_text(encoding="utf-8", errors="replace").splitlines(): - if '"PostToolUseFailure"' not in ln and '"success": false' not in ln: - continue - try: - d = json.loads(ln) - except json.JSONDecodeError: - continue - if d.get("event") != "PostToolUseFailure" and d.get("success") is not False: - continue - total += 1 - head = (d.get("input_summary") or {}).get("command_head") or "" - failures[(d.get("tool_name", "?"), head)] += 1 - - recorded = load(day[:7]) - explained = {e["where"].split("/")[-1].lower() for e in recorded if e["where"]} - explained |= {w for e in recorded for w in re.findall(r"[a-z_-]{4,}", e["message"].lower())} - - print(f"{total} tool failure(s) recorded since {day}, " - f"{len(failures)} distinct shape(s).") - print(f"{len(recorded)} papercut(s) written in the same period.") - - if not failures: - print("\nNothing broke. No prompt needed.") - return 0 - - unexplained = [ - (tool, head, n) for (tool, head), n in failures.most_common() - if not ({tool.lower(), head.lower()} & explained) - ] - - print("\n" + "-" * 74) - print("FAILED, WITH NO PAPERCUT EXPLAINING IT") - print("-" * 74) - if not unexplained: - print(" none. Every failure shape has a matching papercut.") - for tool, head, n in unexplained[:12]: - label = f"{tool}" + (f" ({head}…)" if head else "") - print(f" {n:>3}x {label}") - - print("\n Most of these are expected: typos, missing paths, deliberate probes.") - print(" Record only the ones the next agent would hit too:") - print(' bin/papercut --kind broken --where <tool> "what went wrong and why"') - return 0 - - -def show_list(args: argparse.Namespace) -> int: - entries = load(args.since) - for e in entries[-args.limit:]: - w = f" [{e['where']}]" if e["where"] else "" - print(f"{e['ts'][:16]} {e['kind']:<13}{w}\n {e['message']}") - print(f"\n{len(entries)} total") - return 0 - - -SUBCOMMANDS = {"harvest", "list", "prompt", "auto"} - - -def main() -> int: - """Dispatch by hand rather than with argparse subparsers. - - Subparsers plus a positional `message` cannot coexist: argparse consumes the - first positional as the subcommand name and rejects any message that is not - "harvest" or "list". The first six real papercuts recorded with this tool all - bounced off that, and `--where --fail-on-overflow` bounced off argparse - reading a leading `--` as a flag. - - Recorded as papercuts against this file, which is the correct outcome, but - the lesson is sharper than the bug: **a tool for capturing friction cannot - itself have any.** Anything that makes recording harder than pushing through - guarantees the tool goes unused and the log reads as an absence of problems. - """ - argv = sys.argv[1:] - - if argv and argv[0] in SUBCOMMANDS: - ap = argparse.ArgumentParser(prog=f"papercut {argv[0]}") - ap.add_argument("--since", default="", help="YYYY-MM") - ap.add_argument("--json", default="") - ap.add_argument("--limit", type=int, default=20) - ap.add_argument("--repeats-only", action="store_true") - args = ap.parse_args(argv[1:]) - return {"harvest": harvest, "list": show_list, - "prompt": prompt, "auto": auto}[argv[0]](args) - - if not argv or argv[0] in ("-h", "--help"): - print(__doc__) - print("kinds:") - for k, v in sorted(KINDS.items()): - print(f" {k:<14} {v}") - return 0 - - # Everything else is a record. Options may appear anywhere; whatever is not - # an option is the message, joined. `--where` accepts a leading `--` value. - kind, where, client, parts = "friction", "", "", [] - i = 0 - while i < len(argv): - tok = argv[i] - if tok.startswith("--") and "=" in tok: - key, _, val = tok.partition("=") - tok, nxt = key, val - elif tok in ("--kind", "--where", "--client"): - nxt = argv[i + 1] if i + 1 < len(argv) else "" - i += 1 - else: - parts.append(tok) - i += 1 - continue - if tok == "--kind": - if nxt not in KINDS: - print(f"papercut: unknown kind '{nxt}'. Choose from: " - f"{', '.join(sorted(KINDS))}", file=sys.stderr) - return 2 - kind = nxt - elif tok == "--where": - where = nxt - elif tok == "--client": - client = nxt - i += 1 - - return record(argparse.Namespace( - message=" ".join(parts), kind=kind, where=where, client=client)) - - -if __name__ == "__main__": - # Exit contract, stated because this tool runs in two very different places. - # - # 0 did what was asked - # 2 could not run (bad usage, unwritable store, internal failure) - # - # EXCEPT in `auto`, which is a PostToolUseFailure hook and must be - # **silent and always 0**. The docstring already promised that a hook must - # never break the session to record a complaint about a broken session, and - # the first version did not implement it: any exception printed to stderr - # and exited 2, injecting noise into the session at the exact moment - # something had already gone wrong. Found by sweeping bin/ for exactly this. - # - # Fail-open in the hook, fail-loud everywhere else. A dropped papercut costs - # one record. A hook that shouts costs the session. - is_hook = len(sys.argv) > 1 and sys.argv[1] == "auto" - try: - sys.exit(main()) - except KeyboardInterrupt: - sys.exit(0 if is_hook else 130) - except Exception as exc: # noqa: BLE001 - if is_hook: - sys.exit(0) - print(f"papercut: failed to record ({type(exc).__name__}: {exc})", file=sys.stderr) - print("This is not a clean result. Do not read it as one.", file=sys.stderr) - sys.exit(2) diff --git a/bin/post-route-enrich b/bin/post-route-enrich deleted file mode 100755 index d5c2cd3..0000000 --- a/bin/post-route-enrich +++ /dev/null @@ -1,74 +0,0 @@ -#!/usr/bin/env bash -# Post-route enrichment: pull OA full text into raw stubs, then refresh MindGraph. -# -# Run after bin/ingest-minion run --apply routes files into 10_knowledge/. -# Uses UNPAYWALL_EMAIL or MAINFRAME_UNPAYWALL_EMAIL when set (free API). -# -# Usage: -# bin/post-route-enrich --subset healthcare-practice -# bin/post-route-enrich --subset healthcare-practice --dry-run -# bin/post-route-enrich --file 10_knowledge/.../stub.md -set -euo pipefail - -if ROOT="$(git rev-parse --show-toplevel 2>/dev/null)"; then - : -else - ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -fi - -DRY_RUN=0 -SUBSET="" -FILE="" -FORCE=0 - -while [[ $# -gt 0 ]]; do - case "$1" in - --dry-run) DRY_RUN=1; shift ;; - --subset) SUBSET="${2:-}"; shift 2 ;; - --file) FILE="${2:-}"; shift 2 ;; - --force) FORCE=1; shift ;; - -h|--help) - sed -n '2,12p' "$0" - exit 0 - ;; - *) - echo "Unknown argument: $1" >&2 - exit 2 - ;; - esac -done - -if [[ -z "$SUBSET" && -z "$FILE" ]]; then - echo "post-route-enrich: pass --subset <domain> or --file <path>" >&2 - exit 2 -fi - -FETCH_ARGS=() -if [[ -n "$SUBSET" ]]; then - FETCH_ARGS+=(--subset "$SUBSET") -fi -if [[ -n "$FILE" ]]; then - FETCH_ARGS+=(--file "$FILE") -fi -if [[ "$FORCE" -eq 1 ]]; then - FETCH_ARGS+=(--force) -fi - -if [[ "$DRY_RUN" -eq 1 ]]; then - FETCH_ARGS+=(--dry-run) - echo "post-route-enrich dry-run" - "$ROOT/bin/fetch-source-text" "${FETCH_ARGS[@]}" - echo "" - echo "Would run: bin/mindgraph-refresh" - exit 0 -fi - -echo "post-route-enrich: fetch-source-text" -"$ROOT/bin/fetch-source-text" --apply "${FETCH_ARGS[@]}" - -echo "" -echo "post-route-enrich: mindgraph-refresh" -"$ROOT/bin/mindgraph-refresh" - -echo "" -echo "post-route-enrich: done" \ No newline at end of file diff --git a/bin/prep-ingest b/bin/prep-ingest deleted file mode 100755 index 8fb73ec..0000000 --- a/bin/prep-ingest +++ /dev/null @@ -1,10 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -if ROOT="$(git rev-parse --show-toplevel 2>/dev/null)"; then - : -else - ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -fi - -exec python3 "$ROOT/01_ingest/prep_ingest.py" "$@" diff --git a/bin/process-eval b/bin/process-eval index 478cd29..4320921 100755 --- a/bin/process-eval +++ b/bin/process-eval @@ -1,230 +1,27 @@ #!/usr/bin/env python3 -"""Process-evaluation loop CLI — deterministic preflight + loop decision surface. +"""MainFrame entrypoint for the process-evaluation loop. -Pairs with `.context/workflows/process-evaluation.md` and the -mainframe-process-eval project. Does not replace `bin/mainframe-doctor` -(health vector) or `bin/eval-schedule` (scheduled suite). - -Commands: - preflight Run the standard observational check pack (non-mutating). - status Point at project surfaces and last eval artifacts. - close Print loop-decision template after an eval pass. +Implementation is owned by `40_operations/mainframe-process-eval/src/`. +This shim exists so operators keep a stable global command. """ from __future__ import annotations -import argparse -import json -import subprocess import sys -from datetime import date from pathlib import Path -from typing import Any - ROOT = Path(__file__).resolve().parents[1] -WORKFLOW = ROOT / ".context" / "workflows" / "process-evaluation.md" -PROJECT = ROOT / "30_projects" / "mainframe-process-eval" -OUTPUTS = PROJECT / "outputs" -BACKLOG = PROJECT / "improvement-backlog" / "items.md" - - -def run_cmd(cmd: list[str], *, timeout: int = 180) -> subprocess.CompletedProcess[str]: - return subprocess.run( - cmd, - cwd=str(ROOT), - capture_output=True, - text=True, - timeout=timeout, +SRC = ROOT / "40_operations" / "mainframe-process-eval" / "src" +if not SRC.is_dir(): + sys.stderr.write( + "process-eval: BLOCKED — MPE portable source is missing under " + "40_operations/mainframe-process-eval/src\n" ) + raise SystemExit(1) - -def preflight_steps() -> list[tuple[str, list[str]]]: - py = sys.executable - return [ - ("unittest", [py, "-m", "unittest", "discover", "-s", "tests", "-q"]), - ("ingest_minion_dry_run", [str(ROOT / "bin" / "ingest-minion"), "run", "--dry-run"]), - ("mindgraph_refresh_dry_run", [str(ROOT / "bin" / "mindgraph-refresh"), "--dry-run"]), - ("sync_project_index_check", [str(ROOT / "bin" / "sync-project-index"), "--check"]), - ("session_open_json", [str(ROOT / "bin" / "session-open"), "--json"]), - ("session_close_check", [str(ROOT / "bin" / "session-close"), "--check"]), - ("eval_schedule_check", [str(ROOT / "bin" / "eval-schedule"), "check"]), - ("workflow_report_7d", [str(ROOT / "bin" / "workflow-report"), "--days", "7", "--json"]), - ("mainframe_doctor_quick", [str(ROOT / "bin" / "mainframe-doctor")]), - ( - "loop_catalogue", - [ - py, - str( - PROJECT - / "workbench" - / "loop-eval" - / "scripts" - / "score_catalogue.py" - ), - ], - ), - ] - - -def cmd_preflight(args: argparse.Namespace) -> int: - results: list[dict[str, Any]] = [] - worst = 0 - for name, cmd in preflight_steps(): - if args.skip_doctor and name == "mainframe_doctor_quick": - results.append({"name": name, "exit": None, "skipped": True}) - continue - if args.skip_unittest and name == "unittest": - results.append({"name": name, "exit": None, "skipped": True}) - continue - try: - proc = run_cmd(cmd, timeout=args.timeout) - code = proc.returncode - except subprocess.TimeoutExpired: - code = 124 - proc = None # type: ignore - except FileNotFoundError: - code = 127 - proc = None # type: ignore - # Tools that exit 1 when work remains / health is degraded — still "ran". - advisory = name in { - "session_close_check", - "mainframe_doctor_quick", - "sync_project_index_check", - "eval_schedule_check", - } - hard_fail = code not in (0, None) and not (advisory and code == 1) - results.append( - { - "name": name, - "exit": code, - "skipped": False, - "advisory": bool(advisory and code == 1), - "cmd": cmd, - "stdout_tail": (proc.stdout or "")[-400:] if proc else "", - "stderr_tail": (proc.stderr or "")[-400:] if proc else "", - } - ) - if hard_fail: - worst = max(worst, 1) - - if args.json: - print(json.dumps({"results": results, "ok": worst == 0}, indent=2)) - else: - print("process-eval preflight") - print(f"workflow: {WORKFLOW.relative_to(ROOT)}") - for r in results: - if r.get("skipped"): - print(f" - {r['name']}: SKIP") - continue - if r["exit"] == 0: - flag = "OK" - elif r.get("advisory"): - flag = f"ADVISORY exit={r['exit']}" - else: - flag = f"FAIL exit={r['exit']}" - print(f" - {r['name']}: {flag}") - print() - print("next: sample ≥3 real cases → score → ≤2 improvement slices → rerun") - print(" then: bin/process-eval close") - print(f"outputs: {OUTPUTS.relative_to(ROOT)}/") - if worst == 0: - print("preflight: OK (advisory non-zeros do not fail the pack)") - return worst - - -def cmd_status(args: argparse.Namespace) -> int: - latest = None - if OUTPUTS.is_dir(): - mds = sorted(OUTPUTS.glob("*.md"), key=lambda p: p.stat().st_mtime, reverse=True) - latest = mds[0] if mds else None - payload = { - "project": str(PROJECT.relative_to(ROOT)), - "workflow": str(WORKFLOW.relative_to(ROOT)) if WORKFLOW.exists() else None, - "backlog": str(BACKLOG.relative_to(ROOT)) if BACKLOG.exists() else None, - "latest_output": str(latest.relative_to(ROOT)) if latest else None, - "loop_eval": str( - (PROJECT / "workbench" / "loop-eval" / "catalogue" / "INDEX.md").relative_to(ROOT) - ), - } - if args.json: - print(json.dumps(payload, indent=2)) - else: - print("process-eval status") - for k, v in payload.items(): - print(f" {k}: {v}") - return 0 - - -def cmd_close(args: argparse.Namespace) -> int: - """Emit loop-decision scaffold (operator fills after eval).""" - today = date.today().isoformat() - text = f"""# process-eval close — loop decision ({today}) - -## Pass identity -- eval_run_id: {args.run_id or f"{today}-process-eval"} -- question: {args.question or "(one evaluation question)"} -- output: 30_projects/mainframe-process-eval/outputs/{args.run_id or f"{today}-…"} .md - -## Checks -- [ ] Baseline captured before changes -- [ ] Deterministic pack rerun (`bin/process-eval preflight`) -- [ ] ≥3 sampled cases (normal / boundary / failure) when available -- [ ] Findings classified (code-defect | process-gap | adoption-gap | telemetry-gap | intentional-backlog) -- [ ] ≤2 improvement slices selected (or none) -- [ ] Metric extract + irregularities; `bin/eval-registry harvest` - -## Loop decision (pick one) -- [ ] **same_slice** — continue same process area next pass -- [ ] **new_question** — new observational question -- [ ] **promote_pattern** — pattern ready for bin/workflow/skill (record destination) -- [ ] **park** — blocked or low value; reason: … -- [ ] **handoff** — ownership moves to project: … - -## Destinations (if promote) -- bin / workflow / skill / agent / project-local / no-action - -## Next action -- … - -Workflow: .context/workflows/process-evaluation.md -""" - if args.write: - out = OUTPUTS / f"{args.run_id or today}-process-eval-close.md" - OUTPUTS.mkdir(parents=True, exist_ok=True) - out.write_text(text, encoding="utf-8") - print(f"wrote {out.relative_to(ROOT)}") - else: - print(text) - return 0 - - -def main() -> int: - parser = argparse.ArgumentParser( - description="MainFrame process-evaluation loop CLI", - ) - sub = parser.add_subparsers(dest="cmd", required=True) - - p = sub.add_parser("preflight", help="Run non-mutating check pack") - p.add_argument("--json", action="store_true") - p.add_argument("--skip-unittest", action="store_true") - p.add_argument("--skip-doctor", action="store_true") - p.add_argument("--timeout", type=int, default=300) - p.set_defaults(func=cmd_preflight) - - s = sub.add_parser("status", help="Project + artifact pointers") - s.add_argument("--json", action="store_true") - s.set_defaults(func=cmd_status) - - c = sub.add_parser("close", help="Print or write loop-decision scaffold") - c.add_argument("--run-id", default="") - c.add_argument("--question", default="") - c.add_argument("--write", action="store_true", help=f"Write under {OUTPUTS}") - c.set_defaults(func=cmd_close) - - args = parser.parse_args() - return int(args.func(args)) - +sys.path.insert(0, str(ROOT / "scripts")) +sys.path.insert(0, str(SRC)) +from process_eval import main # noqa: E402 if __name__ == "__main__": raise SystemExit(main()) diff --git a/bin/project-experiment-loop b/bin/project-experiment-loop deleted file mode 100755 index 272bd80..0000000 --- a/bin/project-experiment-loop +++ /dev/null @@ -1,1139 +0,0 @@ -#!/usr/bin/env python3 -"""Project experiment loop — measured runs with registry harvest + triage. - -Complements research-lane-loop (literature). This tool forces: - design identity (eval_run_id) → execute → harvest → action card - -Commands: - preflight --project <slug> - scaffold --project <slug> --study-type ... --title ... --decision ... - canary --fresh [--full-probe] # MindGraph regression package - close --project <slug> # harvest + check - triage [--scope canary|portfolio] # action card (default: portfolio) - portfolio # alias for triage --scope portfolio - -See .context/workflows/project-experiment-loop.md -""" - -from __future__ import annotations - -import argparse -import json -import os -import re -import subprocess -import sys -from datetime import date, datetime, timezone -from pathlib import Path -from typing import Any - - -ROOT = Path(__file__).resolve().parents[1] -PROJECTS = ROOT / "30_projects" -REGISTRY = ROOT / "20_live" / "eval-registry" -ACTION_CARD = REGISTRY / "last-canary-action.md" # MindGraph section (compat) -EVAL_ACTION_CARD = REGISTRY / "last-eval-action.md" # all MainFrame evals -TEMPLATE = ROOT / ".context" / "templates" / "eval-output.md" -WORKFLOW = ROOT / ".context" / "workflows" / "project-experiment-loop.md" -MG_EVAL = PROJECTS / "mindgraph-eval" -PROCESS_EVAL = PROJECTS / "mainframe-process-eval" -EVAL_PROFILE_EXTRA = {"scaffold-claims-study", "skill-eval-workshop"} -PROBE = MG_EVAL / "scripts" / "retrieval_quality_probe.py" -LIVE_ENV = MG_EVAL / "scripts" / "live_envelope_probe.py" -GRAPH_HEALTH = MG_EVAL / "scripts" / "graph_health_trend.py" -MINDGRAPH_PROJECT = ROOT / "mindgraph" - - -def run_cmd( - cmd: list[str], - *, - cwd: Path | None = None, - timeout: int | None = 1200, -) -> subprocess.CompletedProcess[str]: - return subprocess.run( - cmd, - cwd=str(cwd or ROOT), - capture_output=True, - text=True, - timeout=timeout, - ) - - -def mindgraph_uv_python(script: Path, *args: str) -> list[str] | None: - if not MINDGRAPH_PROJECT.is_dir(): - return None - uv = subprocess.run(["which", "uv"], capture_output=True, text=True) - if uv.returncode != 0: - return None - return [ - "uv", - "run", - "--project", - str(MINDGRAPH_PROJECT), - "python", - str(script), - *args, - ] - - -def parse_frontmatter(path: Path) -> dict[str, str]: - if not path.exists(): - return {} - text = path.read_text(encoding="utf-8", errors="replace") - if not text.startswith("---"): - return {} - end = text.find("\n---", 3) - if end < 0: - return {} - block = text[3:end] - out: dict[str, str] = {} - for line in block.splitlines(): - if ":" not in line: - continue - k, v = line.split(":", 1) - out[k.strip()] = v.strip().strip('"').strip("'") - return out - - -def cmd_preflight(project: str) -> int: - proj = PROJECTS / project - if not proj.is_dir(): - print(f"Project not found: {proj}", file=sys.stderr) - return 1 - meth = proj / "methodology-approach.md" - readme = proj / "README.md" - outputs = proj / "outputs" - fm = parse_frontmatter(readme) if readme.exists() else {} - - print(f"# Experiment preflight — {project}") - print() - print(f"- **path:** `{proj.relative_to(ROOT)}`") - print(f"- **project_state:** {fm.get('project_state') or fm.get('status') or '?'}") - print(f"- **next_action:** {fm.get('next_action', '(none)')[:160]}") - print(f"- **methodology-approach.md:** {'yes' if meth.exists() else 'MISSING'}") - print(f"- **outputs/:** {len(list(outputs.glob('*.md'))) if outputs.is_dir() else 0} md files") - - # recent registry rows for project - runs_file = REGISTRY / "runs.jsonl" - recent = [] - if runs_file.exists(): - for line in runs_file.read_text(encoding="utf-8", errors="replace").splitlines(): - if not line.strip(): - continue - try: - row = json.loads(line) - except json.JSONDecodeError: - continue - if row.get("project") == project or row.get("registry", {}).get("project") == project: - recent.append(row) - # harvest may flatten keys - if isinstance(row, dict) and project in json.dumps(row): - if row not in recent: - recent.append(row) - print(f"- **registry lines mentioning project:** ~{len(recent)} (approx scan)") - - print() - print("## Required next (experiment loop)") - print("1. Design: decision_sentence, study_type, eval_run_id, protocol_ref") - print("2. Execute + write outputs/ with metric extract") - print("3. `bin/project-experiment-loop close --project", project + "`") - print("4. `bin/project-experiment-loop triage` # action card") - print() - print(f"Workflow: `{WORKFLOW.relative_to(ROOT)}`") - - # doctor if mindgraph-eval - if project == "mindgraph-eval": - print() - print("## MindGraph doctor (retrieval experiments)") - doc = run_cmd([str(ROOT / "bin" / "mindgraph"), "doctor"], timeout=60) - print(doc.stdout[-800:] if doc.stdout else doc.stderr[-400:]) - if doc.returncode != 0: - print("(doctor non-zero — fix indexes before canary)", file=sys.stderr) - return doc.returncode - return 0 - - -def cmd_scaffold( - project: str, - study_type: str, - title: str, - decision: str, - protocol_ref: str | None, -) -> int: - proj = PROJECTS / project - if not proj.is_dir(): - print(f"Project not found: {proj}", file=sys.stderr) - return 1 - today = date.today().isoformat() - slug = re.sub(r"[^a-z0-9-]+", "-", title.lower()).strip("-")[:48] - run_id = f"{today}-{slug}" - out_dir = proj / "outputs" - raw_dir = proj / "raw-materials" / run_id - out_dir.mkdir(parents=True, exist_ok=True) - raw_dir.mkdir(parents=True, exist_ok=True) - - protocol = protocol_ref or f"30_projects/{project}/methodology-approach.md" - body = f"""--- -title: "{title}" -domain: "knowledge-systems" -type: "lab-report" -status: "running" -study_type: "{study_type}" -lab_report_id: "{run_id}" -eval_run_id: "{run_id}" -protocol_ref: "{protocol}" -decision_sentence: "{decision.replace('"', "'")}" -hypothesis: "Descriptive measurement — replace with directional hypothesis before run when applicable" -disposition: "open" -project: "{project}" -tags: ["evaluation", "eval-registry", "project-experiment-loop", "lab-report"] -updated: "{today}" -source: "bin/project-experiment-loop scaffold" ---- - -<!-- Universal notebook form: .context/templates/lab-report.md · bin/lab-report --> - -# {title} — {today} - -## Boundary - -- Scope and non-claims go here. -- MindGraph / harness results are nominations or measurements — not automatic truth. - -## Decision sentence - -{decision} - -## Hypothesis - -<!-- Expected result before outcomes --> - -## Inputs - -| Field | Value | -|-------|-------| -| study_type | {study_type} | -| eval_run_id | {run_id} | -| protocol_ref | {protocol} | -| unit_of_analysis | | -| primary metric | | -| raw artifacts | `raw-materials/{run_id}/` | - -## Results - -<!-- Fill after execution --> - -## Irregularities - -<!-- Table or "None observed." --> - -## Counterevidence and limits - -<!-- Fill after execution --> - -## Disposition - -`open` — set accept | reject | hold | iterate when closing. - -## Metric extract (eval-registry) - -```yaml -registry: - project: {project} - run_id: {run_id} - study_type: {study_type} - protocol_ref: {protocol} - date: {today} - decision_sentence: "{decision.replace('"', "'")}" - artifact_path: outputs/{run_id}.md - raw_path: raw-materials/{run_id}/ - decision_use: {"regression_only" if study_type in {"regression", "observational"} else "exploratory_only"} -metrics: - - name: placeholder_metric - slice: all - value: 0 - n: 0 - unit: count -irregularities: - - id: scaffold-placeholder - severity: info - category: protocol - observation: "Scaffold only — replace metrics after real run" - context: "{run_id}" - resolved: false -``` - -## Claims (epistemic pass) - -| Statement | Type | Confidence | -|-----------|------|------------| -| | | | -""" - out_path = out_dir / f"{run_id}.md" - if out_path.exists(): - print(f"Refusing overwrite: {out_path}", file=sys.stderr) - return 1 - out_path.write_text(body, encoding="utf-8") - print(f"Wrote {out_path.relative_to(ROOT)}") - print(f"Raw dir {raw_dir.relative_to(ROOT)}") - print(f"eval_run_id={run_id}") - print("Next: execute experiment, edit metrics, then:") - print(f" bin/lab-report check --project {project} --id {run_id}") - print(f" bin/project-experiment-loop close --project {project}") - return 0 - - -def cmd_close(project: str) -> int: - # Project-scoped harvest upgrades placeholder scaffolds (G3) without - # force-reharvesting the entire portfolio. - harvest = run_cmd( - [ - str(ROOT / "bin" / "eval-registry"), - "harvest", - "--project", - project, - ], - timeout=300, - ) - print(harvest.stdout or harvest.stderr) - if harvest.returncode != 0: - return harvest.returncode - check = run_cmd( - [str(ROOT / "bin" / "eval-registry"), "check", "--strict"], - timeout=120, - ) - print(check.stdout or check.stderr) - # strict may fail other projects; still run triage - triage_code = cmd_triage(write=True) - if check.returncode != 0: - print( - "Note: eval-registry check --strict reported issues " - "(may be other projects).", - file=sys.stderr, - ) - return triage_code if harvest.returncode == 0 else harvest.returncode - - -def _latest_output(glob_pat: str) -> Path | None: - paths = sorted(MG_EVAL.glob(glob_pat), key=lambda p: p.stat().st_mtime, reverse=True) - return paths[0] if paths else None - - -def _extract_metric_block(path: Path) -> dict[str, Any] | None: - text = path.read_text(encoding="utf-8", errors="replace") - m = re.search( - r"## Metric extract \(eval-registry\)\s*```yaml\s*(.*?)```", - text, - re.S, - ) - if not m: - return None - # light parse for key metrics - block = m.group(1) - metrics: list[dict[str, Any]] = [] - current: dict[str, Any] | None = None - for line in block.splitlines(): - if line.strip().startswith("- name:"): - if current: - metrics.append(current) - current = {"name": line.split(":", 1)[1].strip()} - elif current and line.strip().startswith("value:"): - val = line.split(":", 1)[1].strip() - try: - current["value"] = float(val) if "." in val else int(val) - except ValueError: - current["value"] = val - elif current and line.strip().startswith("slice:"): - current["slice"] = line.split(":", 1)[1].strip() - if current: - metrics.append(current) - fm = parse_frontmatter(path) - return { - "path": str(path.relative_to(ROOT)), - "eval_run_id": fm.get("eval_run_id", path.stem), - "study_type": fm.get("study_type", ""), - "decision_sentence": fm.get("decision_sentence", ""), - "metrics": metrics, - "updated": fm.get("updated", ""), - } - - - -def _sev_rank(s: str) -> int: - return {"info": 0, "low": 1, "medium": 2, "high": 3}.get(s, 0) - - -def _bump(sev: str, new: str) -> str: - return new if _sev_rank(new) > _sev_rank(sev) else sev - - -def discover_eval_projects() -> list[Path]: - out: list[Path] = [] - if not PROJECTS.is_dir(): - return out - for child in sorted(PROJECTS.iterdir()): - if not child.is_dir(): - continue - if child.name.endswith("-eval") or child.name in EVAL_PROFILE_EXTRA: - out.append(child) - continue - readme = child / "README.md" - if readme.exists() and "eval-profile" in readme.read_text(encoding="utf-8", errors="replace")[:2500]: - out.append(child) - return out - - -def _latest_evalish_output(project: Path) -> Path | None: - """Prefer harvestable metric extracts over newer bare handoff notes.""" - outputs = project / "outputs" - if not outputs.is_dir(): - return None - scored: list[tuple[int, float, Path]] = [] - for p in outputs.glob("*.md"): - name = p.name.lower() - if any( - x in name - for x in ( - "walkthrough", - "blueprint", - "architecture-plan", - "start_here", - "__note__", - ) - ): - continue - try: - head = p.read_text(encoding="utf-8", errors="replace")[:4000] - except OSError: - continue - has_extract = "Metric extract (eval-registry)" in head - has_run_id = bool(re.search(r"(?m)^eval_run_id:\s*\S", head)) - name_hit = any( - x in name for x in ("scheduled", "probe", "compare-", "core-h", "scorecard", "route") - ) - if not (has_extract or has_run_id or name_hit or "handoff" in name): - continue - # Rank: suite/scorecard extracts first; handoffs with extracts still - # lose to non-handoff extracts so triage does not latch on lane notes. - if has_extract and "handoff" not in name: - rank = 0 - elif has_extract and "handoff" in name: - rank = 2 - elif has_run_id and "handoff" not in name: - rank = 1 - elif "handoff" in name: - rank = 4 - else: - rank = 3 - try: - mtime = p.stat().st_mtime - except OSError: - mtime = 0.0 - scored.append((rank, -mtime, p)) - if not scored: - return None - scored.sort() - return scored[0][2] - - -def _load_dispositions() -> dict[str, dict[str, Any]]: - path = REGISTRY / "irregularity-dispositions.jsonl" - out: dict[str, dict[str, Any]] = {} - if not path.exists(): - return out - for line in path.read_text(encoding="utf-8", errors="replace").splitlines(): - if not line.strip(): - continue - try: - row = json.loads(line) - except json.JSONDecodeError: - continue - key = row.get("key") or "|".join( - [ - str(row.get("project", "")), - str(row.get("run_id", "")), - str(row.get("irregularity_id", "")), - ] - ) - out[str(key)] = row - return out - - -_CLOSED_LIFECYCLE = frozenset({"accepted_risk", "waived", "superseded", "resolved"}) - - -def _irregularity_actionable(row: dict[str, Any], dispositions: dict[str, dict[str, Any]]) -> bool: - if row.get("resolved") is True: - return False - status = str(row.get("lifecycle_status") or "").lower() - if status in _CLOSED_LIFECYCLE: - return False - key = "|".join( - [ - str(row.get("project", "")), - str(row.get("run_id", "")), - str(row.get("irregularity_id") or row.get("id") or ""), - ] - ) - d = dispositions.get(key) - if d and str(d.get("status", "")).lower() in _CLOSED_LIFECYCLE: - return False - return True - - -def _open_high_irregs_from_registry(project: str, limit: int = 5) -> list[dict[str, Any]]: - """Latest high irregularities still actionable (G1 lifecycle-aware).""" - path = REGISTRY / "irregularities.jsonl" - if not path.exists(): - return [] - dispositions = _load_dispositions() - latest: dict[str, dict[str, Any]] = {} - for line in path.read_text(encoding="utf-8", errors="replace").splitlines(): - if not line.strip(): - continue - try: - row = json.loads(line) - except json.JSONDecodeError: - continue - if row.get("project") != project: - continue - if str(row.get("severity", "")).lower() != "high": - continue - key = "|".join( - [ - str(row.get("project", "")), - str(row.get("run_id", "")), - str(row.get("irregularity_id") or ""), - ] - ) - latest[key] = row - rows = [r for r in latest.values() if _irregularity_actionable(r, dispositions)] - return rows[-limit:] - - -def _is_operator_gated_context(next_action: str, row: dict[str, Any] | None = None) -> bool: - """G6: operator-gated work should not drive portfolio severity alone.""" - if row and ( - row.get("operator_gated") is True - or str(row.get("category", "")).lower() in {"operator", "calibration", "human_gold"} - ): - return True - na = (next_action or "").lower() - needles = ( - "operator", - "calibration", - "human-gold", - "human gold", - "blind", - "cameron", - "iphone", - "tailscale", - "login", - ) - return any(n in na for n in needles) - - -def _schedule_health() -> dict[str, Any]: - proc = run_cmd([str(ROOT / "bin" / "eval-schedule"), "check"], timeout=60) - summary = (proc.stdout or "").strip().splitlines() - return { - "ok": proc.returncode == 0, - "summary": summary[0] if summary else "unknown", - "detail": "\n".join(summary[:8]), - } - - -def _mindgraph_canary_section() -> tuple[list[str], list[str], str]: - """Returns (lines, actions, severity).""" - probe = _latest_output("outputs/*probe*.md") - envelope = _latest_output("outputs/*live-envelope*.md") - health = _latest_output("outputs/*graph-health*.md") - lines: list[str] = ["### MindGraph canaries (`mindgraph-eval`)", ""] - actions: list[str] = [] - severity = "info" - for label, p in ( - ("retrieval probe", probe), - ("live envelope", envelope), - ("graph health", health), - ): - lines.append(f"- **{label}:** `{p.relative_to(ROOT) if p else '(none)'}`") - - probe_data = _extract_metric_block(probe) if probe else None - env_data = _extract_metric_block(envelope) if envelope else None - if probe_data: - lines.append(f"- probe run_id: `{probe_data['eval_run_id']}`") - failures = 0 - for m in probe_data.get("metrics") or []: - name = m.get("name", "") - val = m.get("value") - lines.append(f" - {name}: {val}") - if name == "probe_query_failures" and isinstance(val, (int, float)): - failures = val - if val > 0: - severity = "high" - actions.append( - "[mindgraph-eval] Investigate probe_query_failures before changing ingest/embedder." - ) - if name in { - "probe_unprotected_scope_candidates", - "unprotected_scope_candidates", - } and isinstance(val, (int, float)) and val > 0: - severity = _bump(severity, "medium") - actions.append( - "[mindgraph-eval] Review unprotected scope candidates on negative live/inbox queries." - ) - # G13: freshness noise alone is not a red canary — low watch only. - # Split snapshot lag (frozen eval DB behind live tree) vs live-index lag - # (installed/frozen corpus missing recent hash matches) so triage does not - # read like a ranking failure. - if name == "probe_freshness_changed_or_missing" and isinstance(val, (int, float)): - if val > 0 and failures == 0: - severity = _bump(severity, "low") - kind = "index lag" - # Prefer explicit submetrics when present in the same extract. - missing = None - hash_mm = None - docs = None - for m2 in probe_data.get("metrics") or []: - n2 = m2.get("name", "") - if n2 == "probe_freshness_missing_from_db": - missing = m2.get("value") - elif n2 == "probe_freshness_hash_mismatches": - hash_mm = m2.get("value") - elif n2 == "probe_db_documents": - docs = m2.get("value") - if isinstance(missing, (int, float)) and missing >= 100: - kind = "snapshot lag (frozen eval DB behind live 10_knowledge/)" - elif isinstance(hash_mm, (int, float)) and hash_mm > 0 and ( - not isinstance(missing, (int, float)) or missing < 50 - ): - kind = "live/frozen hash drift (re-ingest or re-snapshot)" - detail = f"n={int(val)}" - if missing is not None or hash_mm is not None: - detail += ( - f" missing={missing if missing is not None else '?'} " - f"hash_mm={hash_mm if hash_mm is not None else '?'}" - ) - if docs is not None: - detail += f" db_docs={docs}" - actions.append( - f"[mindgraph-eval] Freshness {kind}: {detail} " - "(watch only while query/rank stay green; " - "refresh snapshots or `bin/mindgraph-refresh` — not a ranking defect)." - ) - if failures == 0: - actions.append( - "[mindgraph-eval] Probe green — keep weekly canary cadence." - ) - else: - severity = "medium" - actions.append( - "[mindgraph-eval] No probe output — run `bin/project-experiment-loop canary --fresh`." - ) - - if env_data: - lines.append(f"- live envelope run_id: `{env_data['eval_run_id']}`") - for m in env_data.get("metrics") or []: - lines.append(f" - {m.get('name')}: {m.get('value')}") - if m.get("name") in {"envelope_failures", "intent_failures"} and m.get("value", 0): - severity = "high" - actions.insert( - 0, - "[mindgraph-eval] Live envelope/intent failed — re-check intent install.", - ) - elif envelope: - # many envelope reports lack metric extract; use body pass line - body = envelope.read_text(encoding="utf-8", errors="replace") - if "failed=0" in body or "passed=" in body: - lines.append("- live envelope: report present (metric extract may be thin)") - if "failed=" in body and "failed=0" not in body: - # try parse failed=N - m = re.search(r"failed=(\d+)", body) - if m and int(m.group(1)) > 0: - severity = "high" - actions.insert(0, "[mindgraph-eval] Live envelope reported failures.") - - doc = run_cmd([str(ROOT / "bin" / "mindgraph"), "doctor", "--json"], timeout=60) - doctor_ok = doc.returncode == 0 - lines.append(f"- mindgraph doctor: {'OK' if doctor_ok else 'FAIL'}") - if not doctor_ok: - severity = "high" - actions.insert( - 0, - "[mindgraph-eval] Fix `bin/mindgraph doctor` failures before retrieval work.", - ) - lines.append("") - return lines, actions, severity - - -def _process_eval_section() -> tuple[list[str], list[str], str]: - lines = ["### Scheduled process suite (`mainframe-process-eval`)", ""] - actions: list[str] = [] - severity = "info" - sched = REGISTRY / "schedule-runs.jsonl" - last = None - if sched.exists(): - for line in sched.read_text(encoding="utf-8", errors="replace").splitlines(): - if not line.strip(): - continue - try: - last = json.loads(line) - except json.JSONDecodeError: - continue - health = _schedule_health() - lines.append(f"- eval-schedule check: {'OK' if health['ok'] else 'FAIL'} — {health['summary']}") - if not health["ok"]: - severity = "high" - actions.append( - "[mainframe-process-eval] `bin/eval-schedule check` failed — read OPERATOR.md and fix launchd/stale weekly." - ) - if last: - lines.append( - f"- last schedule run: `{last.get('run_id')}` cadence={last.get('cadence')} " - f"all_passed={last.get('all_passed')}" - ) - if last.get("all_passed") is False: - severity = _bump(severity, "high") - failed = [ - s.get("name") - for s in last.get("steps") or [] - if not s.get("skipped") and s.get("exit_code", 0) != 0 - ] - actions.append( - f"[mainframe-process-eval] Last suite failed steps: {', '.join(failed) or 'unknown'} — fix before process changes." - ) - else: - actions.append( - "[mainframe-process-eval] Last scheduled suite green — keep cadence." - ) - else: - severity = _bump(severity, "medium") - actions.append("[mainframe-process-eval] No schedule-runs.jsonl rows — install/run eval-schedule.") - - latest = _latest_evalish_output(PROCESS_EVAL) - if latest: - lines.append(f"- latest process output: `{latest.relative_to(ROOT)}`") - lines.append("") - return lines, actions, severity - - -def _per_project_section(project: Path) -> tuple[list[str], list[str], str]: - slug = project.name - if slug in {"mindgraph-eval", "mainframe-process-eval"}: - return [], [], "info" - lines = [f"### `{slug}`", ""] - actions: list[str] = [] - severity = "info" - fm = parse_frontmatter(project / "README.md") - state = fm.get("project_state") or fm.get("status") or "?" - next_action = fm.get("next_action", "") - lines.append(f"- state: **{state}**") - if next_action: - lines.append(f"- readme next_action: {next_action[:140]}") - meth = (project / "methodology-approach.md").exists() - lines.append(f"- methodology-approach.md: {'yes' if meth else 'MISSING'}") - if not meth and state == "active": - severity = "medium" - actions.append(f"[{slug}] Add methodology-approach.md (eval-profile requirement).") - - latest = _latest_evalish_output(project) - if latest: - data = _extract_metric_block(latest) - lines.append(f"- latest evalish output: `{latest.relative_to(ROOT)}`") - if data: - lines.append( - f" - run_id=`{data.get('eval_run_id')}` study_type=`{data.get('study_type')}`" - ) - if data.get("decision_sentence"): - lines.append(f" - decision: {data['decision_sentence'][:120]}") - else: - severity = _bump(severity, "medium") - actions.append( - f"[{slug}] Latest output lacks metric extract — retrofit or re-run via experiment loop scaffold." - ) - else: - outs = list((project / "outputs").glob("*.md")) if (project / "outputs").is_dir() else [] - if state == "active" and not outs: - severity = _bump(severity, "medium") - actions.append( - f"[{slug}] Active eval project with no outputs — plan first experiment or pause project." - ) - elif state == "paused" and not outs: - lines.append("- no eval outputs yet (paused — OK)") - elif not latest and outs: - lines.append(f"- outputs present ({len(outs)}) but none look registry-shaped") - severity = _bump(severity, "low") - actions.append( - f"[{slug}] Convert decision-bearing outputs to eval-output template + harvest." - ) - - high = _open_high_irregs_from_registry(slug) - operator_high = [h for h in high if _is_operator_gated_context(next_action, h)] - agent_high = [h for h in high if h not in operator_high] - if agent_high: - severity = "high" - lines.append(f"- open high irregularities (actionable): {len(agent_high)}") - actions.insert( - 0, - f"[{slug}] Resolve open high irregularity: " - f"{agent_high[-1].get('observation') or agent_high[-1].get('irregularity_id')}", - ) - elif operator_high: - # G6: operator-gated highs inform but do not max severity alone - severity = _bump(severity, "medium") - lines.append( - f"- open high irregularities (operator-gated, not agent-runnable): {len(operator_high)}" - ) - actions.append( - f"[{slug}] Operator-gated high irregularity remains " - f"({operator_high[-1].get('observation') or operator_high[-1].get('irregularity_id')}) " - "— schedule human step; do not block automated loops." - ) - elif high: - lines.append(f"- open high irregularities (registry, filtered): 0 actionable of {len(high)}") - - # project-specific nudges - if slug == "claim-audit-lab" or slug == "scaffold-claims-study": - if "calibration" in next_action.lower() or "gold" in next_action.lower() or "blind" in next_action.lower(): - severity = _bump(severity, "medium") - actions.append( - f"[{slug}] Operator-blocked calibration/gold step is the experiment — track completion as calibration study_type." - ) - if slug == "agent-harness-eval" and state == "paused": - actions.append( - "[agent-harness-eval] Paused with prior H1 rejection — next experiment is H4-asymmetric when reactivated." - ) - if slug == "skill-eval-workshop" and state == "paused": - if latest and _extract_metric_block(latest): - # Observational status present — only remind reentry, not missing outputs. - actions.append( - "[skill-eval-workshop] Paused with registry status only — reentry is master-plan " - "workbench modules/cases/githooks before any skill graduation claim." - ) - else: - actions.append( - "[skill-eval-workshop] Paused with 0 registry-shaped outputs — resume workbench " - "before claiming skill graduation." - ) - - lines.append("") - return lines, actions, severity - - -def cmd_triage(*, write: bool = True, scope: str = "portfolio") -> int: - """Build action card for MindGraph canaries and/or all MainFrame evals.""" - now = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M UTC") - all_actions: list[str] = [] - severity = "info" - sections: list[str] = [] - - if scope in {"canary", "portfolio"}: - lines, actions, sev = _mindgraph_canary_section() - sections.extend(lines) - all_actions.extend(actions) - severity = _bump(severity, sev) - - if scope == "portfolio": - lines, actions, sev = _process_eval_section() - sections.extend(lines) - all_actions.extend(actions) - severity = _bump(severity, sev) - - sections.append("### Other eval-profile projects") - sections.append("") - for proj in discover_eval_projects(): - lines, actions, sev = _per_project_section(proj) - if not lines: - continue - sections.extend(lines) - all_actions.extend(actions) - severity = _bump(severity, sev) - - # registry check summary (non-strict noise reduced by harvest skip) - check = run_cmd( - [str(ROOT / "bin" / "eval-registry"), "check", "--strict"], - timeout=120, - ) - check_out = (check.stdout or "").strip() - sections.append("### Registry hygiene") - sections.append("") - if check.returncode == 0: - sections.append("- eval-registry check --strict: **OK**") - all_actions.append("[registry] Hygiene OK — keep writing registry-shaped outputs only.") - else: - # count problems - probs = [ln for ln in check_out.splitlines() if ln.strip().startswith("- ")] - n = len(probs) - sections.append(f"- eval-registry check --strict: **{n} problem(s)** (legacy/non-eval outputs or missing fields)") - severity = _bump(severity, "low" if n < 10 else "medium") - # show top 5 - for ln in probs[:5]: - sections.append(f" {ln}") - if n > 5: - sections.append(f" - … +{n - 5} more") - all_actions.append( - f"[registry] Clean or retrofit {n} non-conforming outputs " - "(or keep them skipped via harvest filters)." - ) - sections.append("") - - # de-dupe actions - seen: set[str] = set() - uniq: list[str] = [] - for a in all_actions: - if a not in seen: - seen.add(a) - uniq.append(a) - - # Prefer actionable failures as primary; demote pure watch / keep-cadence lines. - def _action_rank(a: str) -> int: - al = a.lower() - if any(k in al for k in ("investigate", "failed", "fix `", "fix ", "resolve")): - return 0 - if "watch only" in al or "keep weekly" in al or "keep cadence" in al: - return 3 - if "lacks metric extract" in al or "convert decision" in al or "paused with 0" in al: - return 1 - return 2 - - primary = min(uniq, key=_action_rank) if uniq else "No action extracted." - - title = ( - "Last canary action" - if scope == "canary" - else "Last eval action (MainFrame portfolio)" - ) - header = [ - f"# {title} — {now}", - "", - f"Generated by `bin/project-experiment-loop triage --scope {scope}`.", - "Runs may be automated; **this card is the action layer** so results are useful.", - "", - f"## Severity: **{severity}**", - "", - "## Primary next action", - "", - primary, - "", - "## Sections", - "", - ] - footer = [ - "## Full action list", - "", - ] - for a in uniq: - footer.append(f"- {a}") - footer.extend( - [ - "", - "## Operator ritual", - "", - "1. Read this card on session-open or after weekly / `canary --fresh`.", - "2. If severity ≥ medium, put the primary action on the named project's `next_action`.", - "3. Literature gaps → research-lane-loop; measurements → this experiment loop.", - "4. Do not open new retrieval features while MindGraph probe_query_failures > 0.", - "", - f"Workflow: `{WORKFLOW.relative_to(ROOT)}`", - "", - ] - ) - text = "\n".join(header + sections + footer) - print(text) - if write: - REGISTRY.mkdir(parents=True, exist_ok=True) - if scope == "canary": - ACTION_CARD.write_text(text, encoding="utf-8") - print(f"Wrote {ACTION_CARD.relative_to(ROOT)}", file=sys.stderr) - else: - EVAL_ACTION_CARD.write_text(text, encoding="utf-8") - print(f"Wrote {EVAL_ACTION_CARD.relative_to(ROOT)}", file=sys.stderr) - # keep canary card in sync as MindGraph subsection snapshot for old links - # write thin pointer canary card for back-compat paths - ACTION_CARD.write_text( - f"# Last canary action — see portfolio card\n\n" - f"Superseded by portfolio triage. Read:\n\n" - f"`{EVAL_ACTION_CARD.relative_to(ROOT)}`\n\n" - f"Severity: **{severity}**\n\nPrimary: {primary}\n", - encoding="utf-8", - ) - print(f"Updated {ACTION_CARD.relative_to(ROOT)} (pointer)", file=sys.stderr) - return 0 if severity != "high" else 2 - - -def cmd_canary(*, fresh: bool, full_probe: bool) -> int: - """Run MindGraph regression package: doctor + probe + envelope + harvest + triage.""" - if not fresh: - print("Use --fresh to run a new eval_run_id package.", file=sys.stderr) - return 2 - - # G9: unbuffered progress for background monitors - os.environ.setdefault("PYTHONUNBUFFERED", "1") - try: - sys.stdout.reconfigure(line_buffering=True) # type: ignore[attr-defined] - except Exception: - pass - - stamp = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H%M%SZ") - base_id = f"{stamp}-experiment-loop-canary" - print(f"# Project experiment canary — {base_id}", flush=True) - print(flush=True) - - # 0 doctor - print("## Step 0 — mindgraph doctor") - doc = run_cmd([str(ROOT / "bin" / "mindgraph"), "doctor"], timeout=60) - print(doc.stdout or doc.stderr) - if doc.returncode != 0: - print("Doctor failed — aborting canary.", file=sys.stderr) - return doc.returncode - - # 1 retrieval probe (scheduled query set by default) - probe_id = f"{base_id}-probe" - print(f"## Step 1 — retrieval probe ({probe_id})") - probe_cmd = [ - sys.executable, - str(PROBE), - "--run-id", - probe_id, - "--registry", - "--fused-only", - "--query-ids", - "q04_memory_cite_forget,q07_ai_detection,q11_negative_live_state,q12_negative_inbox", - ] - if full_probe: - probe_cmd = [ - sys.executable, - str(PROBE), - "--run-id", - probe_id, - "--registry", - ] - pr = run_cmd(probe_cmd, timeout=1200) - print(pr.stdout[-2000:] if pr.stdout else "") - if pr.stderr: - print(pr.stderr[-800:], file=sys.stderr) - if pr.returncode != 0: - print(f"Probe exit {pr.returncode}", file=sys.stderr) - # continue to harvest whatever was written - - # 2 live envelope - env_id = f"{base_id}-live-envelope" - print(f"## Step 2 — live envelope ({env_id})") - live_cmd = mindgraph_uv_python(LIVE_ENV, "--run-id", env_id) - if live_cmd: - lr = run_cmd(live_cmd, timeout=1200) - print(lr.stdout[-1500:] if lr.stdout else "") - if lr.returncode != 0: - print(f"Live envelope exit {lr.returncode}", file=sys.stderr) - else: - print("SKIP live envelope (uv/mindgraph project missing)", file=sys.stderr) - - # 3 graph health (optional, fast) - if GRAPH_HEALTH.exists(): - print("## Step 3 — graph health trend") - gh = run_cmd([sys.executable, str(GRAPH_HEALTH)], timeout=120) - print(gh.stdout[-1200:] if gh.stdout else "") - - # 4 harvest + triage - print("## Step 4 — harvest + triage") - code = cmd_close("mindgraph-eval") - - # record experiment loop receipt - receipt = REGISTRY / "experiment-loop-runs.jsonl" - REGISTRY.mkdir(parents=True, exist_ok=True) - with receipt.open("a", encoding="utf-8") as fh: - fh.write( - json.dumps( - { - "base_id": base_id, - "probe_id": probe_id, - "envelope_id": env_id, - "finished_at": datetime.now(timezone.utc).isoformat(), - "doctor_ok": doc.returncode == 0, - "probe_exit": pr.returncode, - "kind": "mindgraph-canary", - } - ) - + "\n" - ) - print(f"Appended {receipt.relative_to(ROOT)}") - return 0 if pr.returncode == 0 and code in (0, 2) else pr.returncode or code - - -def main() -> None: - p = argparse.ArgumentParser(description="Project experiment loop CLI") - sub = p.add_subparsers(dest="cmd", required=True) - - pf = sub.add_parser("preflight", help="Project experiment readiness") - pf.add_argument("--project", required=True) - - sc = sub.add_parser("scaffold", help="Scaffold outputs/<eval_run_id>.md") - sc.add_argument("--project", required=True) - sc.add_argument("--study-type", required=True) - sc.add_argument("--title", required=True) - sc.add_argument("--decision", required=True) - sc.add_argument("--protocol-ref", default=None) - - cl = sub.add_parser("close", help="Harvest registry + triage") - cl.add_argument("--project", required=True) - - tr = sub.add_parser( - "triage", - help="Write action card (default: all MainFrame evals)", - ) - tr.add_argument("--no-write", action="store_true") - tr.add_argument( - "--scope", - choices=["portfolio", "canary"], - default="portfolio", - help="portfolio = all evals (default); canary = MindGraph only", - ) - - sub.add_parser( - "portfolio", - help="Alias for: triage --scope portfolio (all MainFrame evals)", - ) - - cy = sub.add_parser( - "canary", - help="Fresh MindGraph regression: doctor+probe+envelope+harvest+triage", - ) - cy.add_argument( - "--fresh", - action="store_true", - help="Required: allocate new eval_run_id package", - ) - cy.add_argument( - "--full-probe", - action="store_true", - help="Full query set instead of weekly 4-query canary", - ) - - args = p.parse_args() - if args.cmd == "preflight": - raise SystemExit(cmd_preflight(args.project)) - if args.cmd == "scaffold": - raise SystemExit( - cmd_scaffold( - args.project, - args.study_type, - args.title, - args.decision, - args.protocol_ref, - ) - ) - if args.cmd == "close": - raise SystemExit(cmd_close(args.project)) - if args.cmd == "triage": - raise SystemExit( - cmd_triage(write=not args.no_write, scope=args.scope) - ) - if args.cmd == "portfolio": - raise SystemExit(cmd_triage(write=True, scope="portfolio")) - if args.cmd == "canary": - raise SystemExit(cmd_canary(fresh=args.fresh, full_probe=args.full_probe)) - p.error(f"unknown {args.cmd}") - - -if __name__ == "__main__": - main() diff --git a/bin/research-lane-loop b/bin/research-lane-loop deleted file mode 100755 index 85beae1..0000000 --- a/bin/research-lane-loop +++ /dev/null @@ -1,1316 +0,0 @@ -#!/usr/bin/env python3 -"""Research-lane-loop preflight and health checks. - -Hardens the operator loop from `.context/workflows/research-lane-loop.md` -by classifying resume state before agents re-run source-literature. - -Commands: - bin/research-lane-loop preflight [--slug SLUG | --priority P0] [--json] - bin/research-lane-loop audit-indexes [--status active] [--json] - bin/research-lane-loop doctor # same as: preflight --priority P0 (top pick) - -Does not write knowledge or run source-literature. Dry diagnostics only -unless --repair-index is passed on audit-indexes (rewrites stale 00_inbox -paths in lane capture tables when the file already exists under 10_knowledge/). -""" - -from __future__ import annotations - -import argparse -import importlib.util -import json -import re -import sys -from dataclasses import asdict, dataclass, field -from importlib.machinery import SourceFileLoader -from pathlib import Path -from typing import Any - - -ROOT = Path(__file__).resolve().parents[1] -LANES_DIR = ROOT / "30_projects" / "research-lanes-strategy" / "lanes" -COMPLETED_DIR = LANES_DIR / "completed" -INBOX = ROOT / "00_inbox" -KNOWLEDGE = ROOT / "10_knowledge" -WORKFLOW = ROOT / ".context" / "workflows" / "research-lane-loop.md" - -# Load lane-intake list helpers without shelling out. -_LOADER = SourceFileLoader("lane_intake_mod", str(ROOT / "bin" / "lane-intake")) -_SPEC = importlib.util.spec_from_loader(_LOADER.name, _LOADER) -assert _SPEC and _SPEC.loader -lane_intake = importlib.util.module_from_spec(_SPEC) -sys.modules[_SPEC.name] = lane_intake -_SPEC.loader.exec_module(lane_intake) - - -# G1 — never quota evidence. -# -# This loop used to emit "need ~3–4 raws" and "fill to batch size" whenever a -# lane held fewer than three captures. On 2026-08-07 an agent obeyed that -# instruction across lanes C68–C74 and produced exactly three captures per lane, -# with invented titles, invented institutional authors, and DOIs that do not -# resolve. See 20_live/security/2026-08-09__fabricated-source-captures-in-10-knowledge.md. -# -# The agent was not malfunctioning. A quota for evidence, plus mandatory citation -# fields, plus no fetch verification, leaves fabrication as the only compliant -# response — there was no way to answer "I looked, and only one source exists." -# -# The threshold below is now a trigger for a COVERAGE REVIEW, not a target to -# reach. Nothing downstream may require a count of sources, and every prompt this -# tool emits must leave "record the gap" available as a valid outcome. -COVERAGE_REVIEW_THRESHOLD = 3 - -NO_QUOTA_NOTICE = ( - "# NO QUOTA: capture only what you actually retrieved. Finding one real " - "source and recording the gap is a BETTER result than three invented ones. " - "If you cannot fetch it, do not give it a citation." -) - -INBOX_PATH_RE = re.compile(r"`?(00_inbox/[A-Za-z0-9_./-]+\.md)`?") -KNOWLEDGE_PATH_RE = re.compile(r"`?(10_knowledge/[A-Za-z0-9_./-]+\.md)`?") -LANE_TAG_RE = re.compile(r"lane-[a-z0-9]+", re.I) - - -@dataclass -class StaleRef: - tracker_path: str - basename: str - knowledge_path: str | None - still_in_inbox: bool - - -@dataclass -class PreflightReport: - slug: str - lane_id: str - priority: str - status: str - knowledge_domain: str - capture_tags: list[str] - resume_class: str - resume_reason: str - inbox_stubs: list[str] = field(default_factory=list) - knowledge_raws: list[str] = field(default_factory=list) - knowledge_notes: list[str] = field(default_factory=list) - stale_index_refs: list[dict[str, Any]] = field(default_factory=list) - recommended_step: str = "" - command_card: list[str] = field(default_factory=list) - warnings: list[str] = field(default_factory=list) - # G2: machine-readable phase / routing (avoids class-C parking lot) - phase_status: dict[str, str] = field(default_factory=dict) - next_phase: str = "" - blocker: str = "none" # literature|project|operator|none - suggested_command: str = "" - handoff: str = "" - - -def parse_frontmatter(path: Path) -> dict[str, str]: - return lane_intake.parse_frontmatter(path) - - -def parse_capture_tags(fm: dict[str, str], readme_text: str) -> list[str]: - raw = fm.get("capture_tags", "") - tags: list[str] = [] - if raw: - # capture_tags: ["a", "b"] or JSON-ish - for m in re.finditer(r'["\']([^"\']+)["\']', raw): - tags.append(m.group(1)) - if not tags: - # fallback: lane-xx from body/frontmatter - for m in LANE_TAG_RE.finditer(readme_text + " " + json.dumps(fm)): - tags.append(m.group(0).lower()) - # de-dupe preserve order - seen: set[str] = set() - out: list[str] = [] - for t in tags: - tl = t.lower() - if tl not in seen: - seen.add(tl) - out.append(tl) - return out - - -def find_lane_readme(slug: str) -> Path | None: - for base in (LANES_DIR, COMPLETED_DIR): - p = base / slug / "README.md" - if p.exists(): - return p - return None - - -def basename_of(path_str: str) -> str: - return Path(path_str).name - - -def _topic_slug(basename: str) -> str: - """Best-effort topic slug from a MainFrame filename stem.""" - stem = Path(basename).stem - # YYYY-MM-DD__domain__type__slug OR YYYY-MM-DD__source-literature-run__slug - parts = stem.split("__") - if len(parts) >= 4: - return parts[-1].lower().replace("_", "-") - if len(parts) == 3: - return parts[-1].lower().replace("_", "-") - return stem.lower().replace("_", "-") - - -def resolve_knowledge_candidate(basename: str, domain: str | None) -> Path | None: - """Find basename under 10_knowledge, preferring domain. - - Also fuzzy-matches when ingest rewrote type (raw run-note → note) or - inserted a domain segment (source-literature-run__X → domain__note__source-literature-run--X). - """ - if domain: - preferred = KNOWLEDGE / domain / basename - if preferred.exists(): - return preferred - preferred_raw = KNOWLEDGE / domain / "raw" / basename - if preferred_raw.exists(): - return preferred_raw - matches = list(KNOWLEDGE.rglob(basename)) - if matches: - return matches[0] - - # Fuzzy: topic slug contained in any knowledge filename under domain (or all) - topic = _topic_slug(basename) - if len(topic) < 8: - return None - search_roots = [KNOWLEDGE / domain] if domain and (KNOWLEDGE / domain).exists() else [KNOWLEDGE] - fuzzy: list[Path] = [] - for root in search_roots: - for p in root.rglob("*.md"): - name = p.name.lower().replace("_", "-") - if topic in name or topic.replace("-", "") in name.replace("-", ""): - fuzzy.append(p) - if not fuzzy: - return None - # Prefer same domain; prefer shorter path - fuzzy.sort(key=lambda p: (0 if domain and f"/{domain}/" in str(p) else 1, len(str(p)))) - return fuzzy[0] - - -def scan_stale_inbox_refs(readme_text: str, domain: str | None) -> list[StaleRef]: - refs: list[StaleRef] = [] - seen: set[str] = set() - for m in INBOX_PATH_RE.finditer(readme_text): - tracker_path = m.group(1) - if tracker_path in seen: - continue - seen.add(tracker_path) - base = basename_of(tracker_path) - in_inbox = (INBOX / base).exists() or (ROOT / tracker_path).exists() - know = resolve_knowledge_candidate(base, domain) - # Stale = listed as inbox but already in knowledge (or gone from inbox) - if know is not None or not in_inbox: - refs.append( - StaleRef( - tracker_path=tracker_path, - basename=base, - knowledge_path=str(know.relative_to(ROOT)) if know else None, - still_in_inbox=in_inbox, - ) - ) - return refs - - -def find_inbox_stubs_for_lane(tags: list[str], slug: str, lane_id: str) -> list[str]: - if not INBOX.exists(): - return [] - needles = {t.lower() for t in tags} - for t in list(needles): - if t.startswith("lane-"): - needles.add(t[5:]) - elif t and t not in {"research-lane", "needs-audit", "high", "medium", "low"}: - needles.add(f"lane-{t}") - needles.add(slug.lower()) - if lane_id: - needles.add(lane_id.lower()) - needles.add(f"lane-{lane_id.lower()}") - # lane-mh01 style - compact = re.sub(r"[^a-z0-9]", "", lane_id.lower()) - if compact: - needles.add(f"lane-{compact}") - hits: list[str] = [] - for p in sorted(INBOX.glob("*.md")): - try: - text = p.read_text(encoding="utf-8", errors="ignore")[:8000].lower() - except OSError: - continue - name = p.name.lower() - if any(n in name or n in text for n in needles if len(n) >= 3): - # require at least a lane tag or slug signal for short ids - hits.append(str(p.relative_to(ROOT))) - return hits - - -def _classify_knowledge_path(path: Path, head: str) -> str | None: - """Return 'raw' or 'note' or None.""" - name = path.name - if "source-literature-run" in name: - return "raw" - if "__raw__" in name: - return "raw" - if "__note__" in name or "note__" in name: - return "note" - if "type: raw" in head or 'type: "raw"' in head: - return "raw" - if "type: note" in head or 'type: "note"' in head: - return "note" - return None - - -def paths_from_capture_index(readme_text: str) -> list[str]: - """Extract 10_knowledge paths listed in a lane capture-index table.""" - found: list[str] = [] - seen: set[str] = set() - for m in KNOWLEDGE_PATH_RE.finditer(readme_text): - p = m.group(1) - if p not in seen: - seen.add(p) - found.append(p) - # file:///…/10_knowledge/… links in capture tables - for m in re.finditer(r"10_knowledge/[A-Za-z0-9_./-]+\.md", readme_text): - p = m.group(0) - if p not in seen: - seen.add(p) - found.append(p) - return found - - -def find_knowledge_files_for_lane( - domain: str, - tags: list[str], - slug: str, - *, - readme_text: str | None = None, -) -> tuple[list[str], list[str]]: - """Inventory knowledge files for a lane. - - Match strategy (any hit counts): - 1. Paths already listed in the lane capture index - 2. Domain files whose head/name match lane-* tags, slug, or other capture_tags - """ - raws: list[str] = [] - notes: list[str] = [] - seen: set[str] = set() - - def add_path(rel: str, head: str = "", path: Path | None = None) -> None: - if rel in seen: - return - p = path or (ROOT / rel) - if not p.exists(): - return - if not head: - try: - head = p.read_text(encoding="utf-8", errors="ignore")[:4000].lower() - except OSError: - return - kind = _classify_knowledge_path(p, head) - if kind == "raw": - raws.append(rel) - seen.add(rel) - elif kind == "note": - notes.append(rel) - seen.add(rel) - - # 1) Capture index is ground truth when present - if readme_text: - for rel in paths_from_capture_index(readme_text): - add_path(rel) - - if not domain: - return raws, notes - domain_dir = KNOWLEDGE / domain - if not domain_dir.exists(): - return raws, notes - - # 2) Broader tag/slug matching (not only lane-* — older captures often omit it) - needles: list[str] = [] - for t in tags: - tl = t.lower().strip() - if tl and tl not in {"research-lane", "needs-audit", "high", "medium", "low"}: - needles.append(tl) - if tl.startswith("lane-"): - needles.append(tl[5:]) - else: - needles.append(f"lane-{tl}") - if slug: - needles.append(slug.lower().replace("_", "-")) - needles.append(slug.lower()) - # de-dupe - needles = list(dict.fromkeys(n for n in needles if len(n) >= 3)) - - for p in sorted(domain_dir.rglob("*.md")): - if p.name == "index.md": - continue - rel = str(p.relative_to(ROOT)) - if rel in seen: - continue - try: - head = p.read_text(encoding="utf-8", errors="ignore")[:4000].lower() - except OSError: - continue - name = p.name.lower() - if not any(n in head or n in name for n in needles): - continue - add_path(rel, head=head, path=p) - return raws, notes - - -def classify_resume( - inbox_stubs: list[str], - knowledge_raws: list[str], - knowledge_notes: list[str], - stale_refs: list[StaleRef], -) -> tuple[str, str]: - """Return (class, reason). Classes A/B/C/D per research-lane-loop workflow.""" - if stale_refs and any(s.knowledge_path for s in stale_refs): - # D can combine with others; surface D first if index lies - if not inbox_stubs and knowledge_raws: - return ( - "D", - "Capture index still lists 00_inbox/ paths that already exist under 10_knowledge/", - ) - if not inbox_stubs and not knowledge_raws: - return ( - "D", - "Capture index lists 00_inbox/ paths that are missing from inbox (stale or moved)", - ) - - if inbox_stubs: - return ("B", f"{len(inbox_stubs)} inbox stub(s) match this lane — ingest before new search") - - if len(knowledge_raws) >= 3 and not knowledge_notes: - return ( - "C", - f"{len(knowledge_raws)} knowledge raws and no lane synthesis note — write synthesis", - ) - - if len(knowledge_raws) >= 3 and knowledge_notes: - # may still need another phase - return ( - "C", - f"{len(knowledge_raws)} raws + {len(knowledge_notes)} note(s) present — check stop condition / next phase before new search", - ) - - if knowledge_raws and not knowledge_notes: - return ( - "C", - f"{len(knowledge_raws)} knowledge raw(s) without synthesis — synthesise, or " - f"review coverage for unanswered sub-questions", - ) - - if knowledge_raws or knowledge_notes: - return ( - "C", - f"{len(knowledge_raws)} raw(s) + {len(knowledge_notes)} note(s) — review coverage " - f"or check stop condition", - ) - - if stale_refs: - return ("D", "Stale capture-index inbox paths present") - - return ("A", "No inbox stubs and no knowledge raws/notes for this lane — run source-literature") - - -def recommended_step( - resume_class: str, - domain: str, - *, - has_notes: bool = False, - raw_count: int = 0, - note_count: int = 0, -) -> tuple[str, list[str]]: - if resume_class == "A": - return ( - "Step 1–3: brief + preflight + source-literature " - "(capture what exists — no target count)", - [ - 'bin/mindgraph query "LANE QUESTION" --db ~/.mindgraph/mainframe.sqlite --json --top-k 6', - 'bin/mindgraph query "LANE QUESTION" --db ~/.mindgraph/mainframe-projects.sqlite --json --top-k 6', - "# then source-literature skill → 00_inbox/", - NO_QUOTA_NOTICE, - ], - ) - if resume_class == "B": - return ( - "Step 4: ingest pipeline (skip new search)", - [ - "bin/ingest-minion run --dry-run", - "bin/ingest-minion run --apply", - f"bin/post-route-enrich --subset {domain or 'DOMAIN'}", - ], - ) - if resume_class == "C": - if has_notes and note_count >= 1 and raw_count >= 3: - return ( - "Step 7: stop-condition check — synthesis already present; next phase, archive, or handoff", - [ - "# read lane stop condition + latest note", - "bin/lane-intake archive SLUG --apply # only if lane stop condition met", - "# else: set next phase and re-run doctor", - ], - ) - if raw_count < COVERAGE_REVIEW_THRESHOLD: - return ( - "Step 1–3: coverage check — which of this lane's sub-questions are " - "still unanswered, and does real literature exist for them?", - [ - 'bin/mindgraph query "LANE QUESTION" --db ~/.mindgraph/mainframe.sqlite --json --top-k 6', - "# source-literature skill — capture only sources you actually retrieved", - "# if a sub-question has no findable literature, record the gap in the " - "lane README and proceed to synthesis. A documented gap is a result.", - NO_QUOTA_NOTICE, - ], - ) - return ( - "Step 5: knowledge synthesis + mindgraph-refresh", - [ - f"# write 10_knowledge/{domain or 'DOMAIN'}/YYYY-MM-DD__…__note__….md", - "bin/mindgraph-refresh", - ], - ) - # D - return ( - "Step 6: repair capture index paths (00_inbox → 10_knowledge), then re-preflight", - [ - "bin/research-lane-loop audit-indexes --status active # dry-run", - "bin/research-lane-loop audit-indexes --status active --repair-index", - "bin/research-lane-loop preflight --slug SLUG", - ], - ) - - - -def parse_phase_status(readme_text: str) -> dict[str, str]: - """Parse phase tables from lane README (G2). - - Only rows under a Phase status heading (or a `| Phase | Status |` header) - are considered — capture-index tables (`| Date | Path |`) are ignored. - """ - out: dict[str, str] = {} - in_phase_table = False - for line in readme_text.splitlines(): - stripped = line.strip() - # Enter phase table via heading or header row - if re.search(r"^##+\s*phase\s+status", stripped, re.I): - in_phase_table = True - continue - if re.search(r"^##+\s+", stripped) and not re.search(r"phase\s+status", stripped, re.I): - # left the phase section - if in_phase_table: - in_phase_table = False - continue - if not stripped.startswith("|"): - continue - cells = [c.strip() for c in stripped.strip("|").split("|")] - if len(cells) < 2: - continue - phase, status = cells[0], cells[1] - header_l = phase.lower() - if header_l in {"phase", "--------", "---"} or set(phase) <= {"-"}: - # detect explicit Phase|Status header even without heading - if header_l == "phase" and "status" in status.lower(): - in_phase_table = True - continue - if "status" in header_l: - continue - # Skip capture-index style tables - if header_l in {"date", "path", "type"} or phase.startswith("20"): - in_phase_table = False - continue - if not in_phase_table: - # only accept known phase names outside an explicit table - if phase.lower() not in { - "foundation", - "taxonomy", - "specialization", - "application", - }: - continue - # normalize status tokens - st = status.lower() - if "✅" in status or "done" in st or "complete" in st or "met" in st: - norm = "done" - elif "next" in st or "⏳" in status or "todo" in st or "open" in st: - norm = "next" - elif "active" in st or "wip" in st or "progress" in st: - norm = "active" - elif "skip" in st or "n/a" in st or "defer" in st: - norm = "skipped" - else: - norm = status.strip() or "unknown" - out[phase] = norm - return out - - -def infer_next_phase(phase_status: dict[str, str]) -> str: - order = [ - "Foundation", - "Taxonomy", - "Specialization", - "Application", - "foundation", - "taxonomy", - "specialization", - "application", - ] - # prefer explicit next - for name, st in phase_status.items(): - if st == "next": - return name - # first not done in preferred order - lower_map = {k.lower(): (k, v) for k, v in phase_status.items()} - for key in ("foundation", "taxonomy", "specialization", "application"): - if key in lower_map: - name, st = lower_map[key] - if st not in {"done", "skipped"}: - return name - for name, st in phase_status.items(): - if st not in {"done", "skipped"}: - return name - return "archive_or_handoff" if phase_status else "" - - -def infer_blocker( - resume_class: str, - next_phase: str, - handoff: str, - *, - raw_count: int, - note_count: int, -) -> str: - """literature | project | operator | none""" - np = next_phase.lower() - if resume_class == "B": - return "literature" # inbox awaiting ingest (still research pipeline) - if resume_class == "A": - return "literature" - if resume_class == "D": - return "literature" - if resume_class == "C": - if note_count == 0 and raw_count >= 1: - return "literature" - if "application" in np or "specialization" in np: - if handoff and ("30_projects" in handoff or "project" in handoff.lower()): - return "project" - return "project" - if "archive" in np or "handoff" in np: - return "operator" if not handoff else "project" - if raw_count < 3: - return "literature" - return "none" - return "none" - - -def build_suggested_command( - slug: str, - resume_class: str, - blocker: str, - next_phase: str, - handoff: str, -) -> str: - if resume_class == "B": - return "bin/ingest-minion run --dry-run && bin/ingest-minion run --apply" - if resume_class == "D": - return f"bin/research-lane-loop audit-indexes --status active --repair-index" - if resume_class == "A" or (blocker == "literature" and resume_class != "C"): - return f"# source-literature for {slug} (capture what exists; no target count)" - if blocker == "literature" and resume_class == "C": - return f"# coverage check or synthesis for {slug}; then update tracker" - if blocker == "project": - target = "" - # first 30_projects path token in handoff - m = re.search(r"30_projects/([a-z0-9_-]+)", handoff or "") - if m: - target = m.group(1) - note = f"advance to {next_phase or 'application'}" - return ( - f"bin/research-lane-loop handoff-project --slug {slug} " - f"--to {target} --note {note!r}" - ) - return f"# application work for {slug} (set project next_action)" - if blocker == "operator": - return f"# operator decision: archive or continue {slug}" - if next_phase and "archive" in next_phase.lower(): - return f"bin/lane-intake archive {slug} --apply # only if stop condition met" - return f"bin/research-lane-loop preflight --slug {slug}" - - -def preflight_slug(slug: str) -> PreflightReport: - readme = find_lane_readme(slug) - if not readme: - raise SystemExit(f"Lane not found: {slug}") - text = readme.read_text(encoding="utf-8", errors="ignore") - fm = parse_frontmatter(readme) - domain = fm.get("knowledge_domain", "") - tags = parse_capture_tags(fm, text) - lane_id = fm.get("lane_id", "") - inbox = find_inbox_stubs_for_lane(tags, slug, lane_id) - raws, notes = find_knowledge_files_for_lane( - domain, tags, slug, readme_text=text - ) - stale = scan_stale_inbox_refs(text, domain) - rclass, reason = classify_resume(inbox, raws, notes, stale) - # If D and B both true, prefer B (ingest) but warn about D - warnings: list[str] = [] - if rclass == "B" and stale: - warnings.append( - "Also has stale capture-index inbox paths — repair after ingest (class D)" - ) - if rclass == "C" and stale: - warnings.append("Capture index still lists old 00_inbox/ paths — repair at step 6") - # upgrade emphasis: still C for synthesis but flag D repair needed - step, card = recommended_step( - rclass, - domain, - has_notes=bool(notes), - raw_count=len(raws), - note_count=len(notes), - ) - card = [c.replace("SLUG", slug) for c in card] - - phase_status = parse_phase_status(text) - next_phase = infer_next_phase(phase_status) - handoff = fm.get("handoff", "") - blocker = infer_blocker( - rclass, - next_phase, - handoff, - raw_count=len(raws), - note_count=len(notes), - ) - # When class C with notes and application next, prefer project blocker - if rclass == "C" and notes and not next_phase: - # read "Phase status" prose fallback - if re.search(r"application\s*[|:]\s*(next|todo|open)", text, re.I): - next_phase = "Application" - blocker = infer_blocker(rclass, next_phase, handoff, raw_count=len(raws), note_count=len(notes)) - suggested = build_suggested_command(slug, rclass, blocker, next_phase, handoff) - # enrich command card with suggested - if suggested and suggested not in card: - card = [suggested] + card - - return PreflightReport( - slug=slug, - lane_id=lane_id, - priority=fm.get("priority", ""), - status=fm.get("status", ""), - knowledge_domain=domain, - capture_tags=tags, - resume_class=rclass, - resume_reason=reason, - inbox_stubs=inbox, - knowledge_raws=raws, - knowledge_notes=notes, - stale_index_refs=[asdict(s) for s in stale], - recommended_step=step, - command_card=card, - warnings=warnings, - phase_status=phase_status, - next_phase=next_phase, - blocker=blocker, - suggested_command=suggested, - handoff=handoff, - ) - - -def pick_top_lane(priority: str, status: str) -> str | None: - rows = lane_intake.list_lanes(status_filter=status, priority_filter=priority) - # prefer main lanes/ over completed - main = [r for r in rows if r.get("location") == "lanes"] - pool = main or rows - if not pool: - return None - return pool[0]["slug"] - - -def format_preflight(report: PreflightReport) -> str: - lines = [ - f"# Research-lane-loop preflight — {report.lane_id or '?'} / {report.slug}", - "", - f"- **priority:** {report.priority}", - f"- **status:** {report.status}", - f"- **domain:** {report.knowledge_domain}", - f"- **capture_tags:** {', '.join(report.capture_tags) or '(none)'}", - f"- **resume_class:** **{report.resume_class}** — {report.resume_reason}", - f"- **recommended:** {report.recommended_step}", - f"- **next_phase:** {report.next_phase or '(unknown)'}", - f"- **blocker:** {report.blocker}", - f"- **suggested_command:** `{report.suggested_command}`" if report.suggested_command else "- **suggested_command:** (none)", - f"- **handoff:** {report.handoff or '(none)'}", - "", - f"## Inventory", - f"- inbox stubs: {len(report.inbox_stubs)}", - f"- knowledge raws (tag match): {len(report.knowledge_raws)}", - f"- knowledge notes (tag match): {len(report.knowledge_notes)}", - f"- stale index refs: {len(report.stale_index_refs)}", - ] - if report.phase_status: - lines.append("") - lines.append("## Phase status") - for ph, st in report.phase_status.items(): - lines.append(f"- {ph}: {st}") - if report.warnings: - lines.append("") - lines.append("## Warnings") - for w in report.warnings: - lines.append(f"- {w}") - if report.inbox_stubs: - lines.append("") - lines.append("## Inbox stubs") - for p in report.inbox_stubs: - lines.append(f"- `{p}`") - if report.stale_index_refs: - lines.append("") - lines.append("## Stale capture-index paths") - for s in report.stale_index_refs: - kp = s.get("knowledge_path") or "(not found in 10_knowledge/)" - lines.append(f"- `{s['tracker_path']}` → `{kp}`") - if report.knowledge_raws: - lines.append("") - lines.append("## Knowledge raws (sample)") - for p in report.knowledge_raws[:12]: - lines.append(f"- `{p}`") - if len(report.knowledge_raws) > 12: - lines.append(f"- … +{len(report.knowledge_raws) - 12} more") - if report.knowledge_notes: - lines.append("") - lines.append("## Knowledge notes") - for p in report.knowledge_notes[:8]: - lines.append(f"- `{p}`") - lines.append("") - lines.append("## Command card") - lines.append("```bash") - lines.extend(report.command_card) - lines.append("```") - lines.append("") - lines.append(f"Workflow: `{WORKFLOW.relative_to(ROOT)}`") - return "\n".join(lines) - - -def cmd_handoff_project( - slug: str, - to_project: str, - note: str, - dry_run: bool, - *, - kind: str = "application", - evidence: str | None = None, - urgency: str | None = None, -) -> int: - """Outbound research → project handoff (typed). - - Kinds: see `.context/workflows/research-project-handoff.md` - """ - from datetime import date as _date - - kinds = { - "gate", - "application", - "opportunity", - "constraint", - "experiment", - "craft", - "knowledge", - "split", - "close", - } - kind = (kind or "application").lower().strip() - if kind not in kinds: - print(f"kind must be one of {sorted(kinds)}", file=sys.stderr) - return 2 - - readme = find_lane_readme(slug) - if not readme: - print(f"Lane not found: {slug}", file=sys.stderr) - return 1 - - proj = ROOT / "30_projects" / to_project - if kind != "split" and not proj.is_dir(): - print(f"Project not found: {proj}", file=sys.stderr) - return 1 - - today = _date.today().isoformat() - default_urgency = { - "gate": "high", - "application": "high", - "opportunity": "low", - "constraint": "medium", - "experiment": "medium", - "craft": "medium", - "knowledge": "low", - "split": "medium", - "close": "high", - } - urg = (urgency or default_urgency[kind]).lower() - evidence = (evidence or "").strip() - note = (note or "").strip() or "see lane README / synthesis" - - next_loop = { - "gate": "implement", - "application": "implement", - "opportunity": "none", - "constraint": "none", - "experiment": "experiment", - "craft": "craft", - "knowledge": "none", - "split": "intake", - "close": "archive", - }[kind] - - does_not = { - "gate": "not optional flavor text — a decision depends on this", - "application": "not a new literature search; implement or craft the application phase", - "opportunity": "not a forced WIP steal; consider when capacity allows", - "constraint": "not a full product redesign — a bound on existing plans", - "experiment": "not a craft smoke; open a measured eval_run_id", - "craft": "not registry metrics by default; one craft question", - "knowledge": "not project labor — consume the durable note", - "split": "not this project's next_action — candidate for research-lane-intake", - "close": "not a phase pause — lane stop / archive path", - }[kind] - - handoff_msg = ( - f"[{kind}/{urg}] lane `{slug}` → `{to_project}`: {note}" - + (f" | evidence: {evidence}" if evidence else "") - ) - - def _set_na(path: Path, value: str) -> None: - raw = path.read_text(encoding="utf-8") - if raw.startswith("---"): - parts = raw.split("---", 2) - if len(parts) >= 3: - head, body = parts[1], parts[2] - if re.search(r"^next_action:\s*", head, re.M): - head = re.sub( - r"^next_action:\s*.*$", - f'next_action: "{value}"', - head, - count=1, - flags=re.M, - ) - else: - head = head.rstrip() + f'\nnext_action: "{value}"\n' - if re.search(r"^updated:\s*", head, re.M): - head = re.sub( - r"^updated:\s*.*$", - f'updated: "{today}"', - head, - count=1, - flags=re.M, - ) - path.write_text(f"---{head}---{body}", encoding="utf-8") - return - path.write_text(raw + f"\n\nnext_action: {value}\n", encoding="utf-8") - - def _append_log(path: Path, title: str, body: str) -> None: - entry = f"\n## {today} | {title}\n\n{body}\n" - if path.exists(): - path.write_text(path.read_text(encoding="utf-8") + entry, encoding="utf-8") - else: - path.write_text("# Log\n" + entry, encoding="utf-8") - - if kind == "close": - lane_next = ( - f"Lane stop / close path — handoff to consumers done; " - f"run bin/lane-intake archive {slug} when operator accepts." - ) - elif kind == "split": - lane_next = ( - f"Split candidate emitted toward research intake " - f"(related project signal: {to_project})." - ) - elif kind in {"application", "gate"}: - lane_next = ( - f"{kind}: with `{to_project}` — literature stop for this signal; " - f"re-enter research-lane-loop only for a new phase/question." - ) - else: - lane_next = ( - f"Handoff ({kind}) filed to `{to_project}`; " - f"lane may continue other phases or wait on consumer." - ) - - proj_na: str | None = None - steal_next = kind in {"gate", "application", "experiment", "craft"} - if kind == "experiment": - proj_na = ( - f"From lane `{slug}` [{kind}]: {note}. " - f"Next: bin/project-experiment-loop scaffold --project {to_project} " - f'--study-type exploratory --title "from-{slug}" ' - f'--decision "{note[:80]}"' - ) - elif kind == "craft": - proj_na = ( - f"From lane `{slug}` [{kind}]: {note}. " - f"Next: bin/craft-research-loop scaffold --project {to_project} " - f'--question "{note[:100]}" ' - f'--decision "If yes, adopt; if no, iterate."' - ) - elif kind in {"gate", "application"}: - proj_na = f"From lane `{slug}` [{kind}/{urg}]: {note}" - if evidence: - proj_na += f" Evidence: {evidence}" - - if dry_run: - print(f"dry-run kind={kind} urgency={urg} next_loop={next_loop}") - print(f" lane next_action → {lane_next}") - if steal_next and proj_na: - print(f" project next_action REPLACE → {proj_na[:200]}") - elif kind == "opportunity": - print(" project: append CONSIDER to log only (no next_action steal)") - elif kind == "constraint": - print(" project: append decisions/risk log (no automatic next_action steal)") - elif kind == "knowledge": - print(" project: cite evidence in log only") - elif kind == "split": - print(" project: no next_action change; emit lane candidate") - print(f" does_not_mean: {does_not}") - return 0 - - lane_log = readme.parent / "log.md" - _append_log( - lane_log, - f"handoff-project [{kind}] → `{to_project}`", - f"- {handoff_msg}\n" - f"- next_loop: {next_loop}\n" - f"- does_not_mean: {does_not}\n" - f"- lane next_action: {lane_next}\n", - ) - _set_na(readme, lane_next) - - if kind == "split": - if proj.is_dir() and (proj / "README.md").exists(): - _append_log( - proj / "log.md", - f"research split signal from `{slug}`", - f"- {handoff_msg}\n- Action: consider research-lane-intake, not implement.\n", - ) - print(f"Handed off [{kind}] lane `{slug}` (split — intake, not implement)") - print(f" does_not_mean: {does_not}") - return 0 - - proj_readme = proj / "README.md" - if not proj_readme.exists(): - print(f"Missing {proj_readme}", file=sys.stderr) - return 1 - - pfm = parse_frontmatter(proj_readme) - prev = pfm.get("next_action", "") - - if steal_next and proj_na: - _set_na(proj_readme, proj_na[:500]) - action_note = f"next_action replaced (prior kept in log): {prev[:120]}" - elif kind == "opportunity": - action_note = "CONSIDER only — primary next_action unchanged" - elif kind == "constraint": - action_note = "constraint logged — review decisions.md / risks" - if not prev.strip(): - _set_na( - proj_readme, - f"From lane `{slug}` [constraint]: review bound — {note}", - ) - action_note += " (set next_action because project was idle)" - elif kind == "knowledge": - action_note = "knowledge cite only" - elif kind == "close": - action_note = "lane close notify — no forced implement" - else: - action_note = "logged" - - receipt_dir = proj / "outputs" - receipt_dir.mkdir(parents=True, exist_ok=True) - receipt_path = receipt_dir / f"{today}-handoff-{kind}-{slug}.md" - receipt = f"""--- -title: "Research handoff — {kind} — {slug} → {to_project}" -domain: "knowledge-systems" -type: "project" -status: "active" -handoff_kind: "{kind}" -from_lane: "{slug}" -to_project: "{to_project}" -urgency: "{urg}" -next_loop: "{next_loop}" -updated: "{today}" -source: "bin/research-lane-loop handoff-project" -tags: ["research-handoff", "handoff", "{kind}"] ---- - -# Research handoff — {kind} - -## Decision or signal - -{note} - -## Kind - -| Field | Value | -|-------|-------| -| kind | {kind} | -| urgency | {urg} | -| next_loop | {next_loop} | -| from_lane | {slug} | -| to_project | {to_project} | - -## Evidence - -- {evidence or "(none provided — add paths)"} - -## Does not mean - -- {does_not} - -## Routing applied - -- {action_note} -- prior project next_action: {prev[:200] if prev else "(empty)"} -""" - receipt_path.write_text(receipt, encoding="utf-8") - - log_body = ( - f"- {handoff_msg}\n" - f"- next_loop: **{next_loop}**\n" - f"- does_not_mean: {does_not}\n" - f"- routing: {action_note}\n" - f"- receipt: `{receipt_path.relative_to(ROOT)}`\n" - ) - if kind == "opportunity": - log_body = f"### CONSIDER\n\n{log_body}" - if kind == "constraint": - log_body = f"### CONSTRAINT\n\n{log_body}" - _append_log(proj / "log.md", f"research handoff [{kind}] from `{slug}`", log_body) - - if kind == "constraint": - dec = proj / "decisions.md" - block = ( - f"\n## {today} — constraint from lane `{slug}`\n\n" - f"**Signal:** {note}\n\n" - f"**Evidence:** {evidence or 'see handoff receipt'}\n\n" - f"**Receipt:** `{receipt_path.relative_to(ROOT)}`\n" - ) - if dec.exists(): - dec.write_text(dec.read_text(encoding="utf-8") + block, encoding="utf-8") - else: - dec.write_text("# Decisions\n" + block, encoding="utf-8") - - print(f"Handed off [{kind}/{urg}] lane `{slug}` → project `{to_project}`") - print(f" next_loop: {next_loop}") - print(f" receipt: {receipt_path.relative_to(ROOT)}") - print(f" routing: {action_note}") - print(f" does_not_mean: {does_not}") - return 0 - - -def audit_indexes(status: str | None, repair: bool) -> dict[str, Any]: - rows = lane_intake.list_lanes(status_filter=status, priority_filter=None) - # only main lanes for repair by default - findings: list[dict[str, Any]] = [] - repaired: list[str] = [] - for r in rows: - if r.get("location") != "lanes": - continue - readme = Path(r["path"]) - text = readme.read_text(encoding="utf-8", errors="ignore") - fm = parse_frontmatter(readme) - domain = fm.get("knowledge_domain", "") - stale = scan_stale_inbox_refs(text, domain) - fixable = [s for s in stale if s.knowledge_path] - if not stale: - continue - entry = { - "slug": r["slug"], - "lane_id": r.get("id"), - "stale_count": len(stale), - "fixable_count": len(fixable), - "stale": [asdict(s) for s in stale], - } - findings.append(entry) - if repair and fixable: - new_text = text - for s in fixable: - assert s.knowledge_path - # replace bare and backticked forms - new_text = new_text.replace(s.tracker_path, s.knowledge_path) - if new_text != text: - readme.write_text(new_text, encoding="utf-8") - repaired.append(r["slug"]) - return { - "lanes_with_stale_refs": len(findings), - "repaired": repaired, - "repair_applied": repair, - "findings": findings, - } - - -def main() -> None: - p = argparse.ArgumentParser(description="Research-lane-loop preflight and index audit") - sub = p.add_subparsers(dest="cmd", required=True) - - pf = sub.add_parser("preflight", help="Classify resume class A–D for a lane") - pf.add_argument("--slug", help="Lane slug (directory name)") - pf.add_argument("--priority", default="P0", help="When --slug omitted, pick top match") - pf.add_argument("--status", default="active", help="Status filter for auto-pick") - pf.add_argument("--json", action="store_true") - - doc = sub.add_parser( - "doctor", - help="Preflight top P0 active lane, or --all-active portfolio table (G8)", - ) - doc.add_argument("--json", action="store_true") - doc.add_argument( - "--all-active", - action="store_true", - help="Table of all active lanes: class, phase, blocker, suggested command", - ) - doc.add_argument( - "--status", - default="active", - help="With --all-active, lane status filter (default active)", - ) - - au = sub.add_parser("audit-indexes", help="Find stale 00_inbox paths in lane trackers") - au.add_argument("--status", default="active") - au.add_argument("--json", action="store_true") - au.add_argument( - "--repair-index", - action="store_true", - help="Rewrite fixable 00_inbox → 10_knowledge paths in trackers", - ) - - ho = sub.add_parser( - "handoff-project", - help="Typed research→project handoff (see research-project-handoff.md)", - ) - ho.add_argument("--slug", required=True, help="Lane slug") - ho.add_argument("--to", required=True, dest="to_project", help="30_projects/<slug>") - ho.add_argument( - "--kind", - default="application", - choices=[ - "gate", - "application", - "opportunity", - "constraint", - "experiment", - "craft", - "knowledge", - "split", - "close", - ], - help="Handoff kind (default application)", - ) - ho.add_argument("--note", default="", help="One-sentence decision or signal") - ho.add_argument("--evidence", default=None, help="Knowledge path(s) supporting the handoff") - ho.add_argument( - "--urgency", - default=None, - choices=["high", "medium", "low"], - help="Override default urgency for kind", - ) - ho.add_argument("--dry-run", action="store_true") - - args = p.parse_args() - - if args.cmd in {"preflight", "doctor"}: - if args.cmd == "doctor" and getattr(args, "all_active", False): - rows = lane_intake.list_lanes( - status_filter=getattr(args, "status", "active"), - priority_filter=None, - ) - main = [r for r in rows if r.get("location") == "lanes"] - reports = [] - for r in main: - try: - reports.append(preflight_slug(r["slug"])) - except SystemExit: - continue - if args.json: - print(json.dumps([asdict(x) for x in reports], indent=2)) - else: - print( - f"# Research-lane doctor — all {getattr(args, 'status', 'active')} " - f"({len(reports)} lanes)" - ) - print() - print( - f"{'ID':6} {'class':5} {'blocker':10} {'next_phase':18} " - f"{'slug':36} suggested" - ) - print("-" * 110) - for rep in reports: - print( - f"{(rep.lane_id or '?'):6} {rep.resume_class:5} " - f"{rep.blocker:10} {(rep.next_phase or '-'):18} " - f"{rep.slug:36} {rep.suggested_command[:40]}" - ) - print() - print( - "blocker=literature|project|operator|none · " - "Workflow: .context/workflows/research-lane-loop.md" - ) - return - - slug = getattr(args, "slug", None) - priority = getattr(args, "priority", "P0") - status = getattr(args, "status", "active") - if args.cmd == "doctor": - slug = None - priority = "P0" - status = "active" - if not slug: - slug = pick_top_lane(priority, status) - if not slug: - print( - f"No lanes match --priority {priority} --status {status}. " - "Try: bin/lane-intake list --status active", - file=sys.stderr, - ) - sys.exit(1) - report = preflight_slug(slug) - if args.json: - print(json.dumps(asdict(report), indent=2)) - else: - print(format_preflight(report)) - return - - if args.cmd == "audit-indexes": - result = audit_indexes(args.status, repair=args.repair_index) - if args.json: - print(json.dumps(result, indent=2)) - else: - print( - f"Lanes with stale inbox refs: {result['lanes_with_stale_refs']} " - f"(status={args.status})" - ) - for f in result["findings"]: - print( - f"- {f.get('lane_id') or '?'} {f['slug']}: " - f"{f['stale_count']} stale, {f['fixable_count']} fixable" - ) - for s in f["stale"][:5]: - print(f" {s['tracker_path']} → {s.get('knowledge_path') or '?'}") - if args.repair_index: - print(f"\nRepaired: {', '.join(result['repaired']) or '(none)'}") - else: - print("\nDry-run only. Re-run with --repair-index to rewrite fixable paths.") - return - - if args.cmd == "handoff-project": - raise SystemExit( - cmd_handoff_project( - args.slug, - args.to_project, - args.note, - args.dry_run, - kind=getattr(args, "kind", "application"), - evidence=getattr(args, "evidence", None), - urgency=getattr(args, "urgency", None), - ) - ) - - p.error(f"unknown command {args.cmd}") - - -if __name__ == "__main__": - main() diff --git a/bin/second-brain-batch b/bin/second-brain-batch deleted file mode 100755 index 45b8109..0000000 --- a/bin/second-brain-batch +++ /dev/null @@ -1,357 +0,0 @@ -#!/usr/bin/env python3 -"""Register an immutable snapshot of a rolling second-brain inbox drop.""" - -from __future__ import annotations - -import argparse -import csv -import hashlib -import re -import shutil -import sys -from collections import Counter -from datetime import datetime -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -PROJECT_DIR = ROOT / "30_projects" / "second-brain-migration" -BATCHES_DIR = PROJECT_DIR / "raw-materials" / "batches" -REGISTER_PATH = PROJECT_DIR / "raw-materials" / "batch-register.md" -SOURCE_SNAPSHOT_DIRNAME = "source-files" -LEDGER_FILENAME = "disposition-ledger.csv" -TEXT_EXTENSIONS = {".md", ".txt", ".csv", ".tsv", ".json", ".jsx", ".html"} -ISO_DATE_RE = re.compile(r"\b20\d{2}-[01]\d-[0-3]\d\b") -MONTH_DATE_RE = re.compile( - r"\b(?:January|February|March|April|May|June|July|August|September|" - r"October|November|December) [0-3]?\d, 20\d{2}\b" -) -BATCH_ID_RE = re.compile(r"^20\d{2}-[01]\d-[0-3]\d-\d{3}$") -MANIFEST_FIELDS = [ - "batch_id", - "source", - "relative_path", - "size_bytes", - "sha256", - "created_at", - "modified_at", - "extension", - "embedded_dates", - "category", - "proposed_destination", - "review_status", - "duplicate_group", -] -LEDGER_FIELDS = [ - "recorded_at", - "batch_id", - "relative_path", - "sha256", - "disposition", - "destination", - "evidence", - "notes", -] - - -def sha256(path: Path) -> str: - digest = hashlib.sha256() - with path.open("rb") as handle: - for chunk in iter(lambda: handle.read(1024 * 1024), b""): - digest.update(chunk) - return digest.hexdigest() - - -def embedded_dates(path: Path) -> str: - values = ISO_DATE_RE.findall(path.name) - if path.suffix.lower() in TEXT_EXTENSIONS: - try: - sample = path.read_text(encoding="utf-8", errors="replace")[:32768] - except OSError: - sample = "" - values.extend(ISO_DATE_RE.findall(sample)) - values.extend(MONTH_DATE_RE.findall(sample)) - return "; ".join(dict.fromkeys(values[:8])) - - -def classify(path: Path) -> tuple[str, str]: - name = path.name.lower() - rel = path.as_posix().lower() - - if name == ".ds_store" or "vault-lint-workspace/" in rel: - return "generated-artifact", "park pending artifact review" - if ( - name.startswith("cb-") - or "ce-state" in name - or "content-pipeline" in name - or "newsletter" in name - or "_draft_" in name - or "_deepdive_" in name - or "_quicktake_" in name - ): - return "project-specific", "reconcile with 30_projects/content-engine" - if any( - token in name - for token in ( - "state", - "status", - "positions", - "allocations", - "rebalance", - "watchlist", - "options-log", - "crypto-log", - "identity", - "deployment", - "housing-indicators", - "earnings-tracker", - ) - ): - return "live-state-candidate", "reconcile before 20_live promotion" - if any( - token in name - for token in ("template", "prompt", "migration_plan", "raw_catalogue", "taxonomy") - ) or path.suffix.lower() in {".jsx", ".docx"}: - return "template-or-tool", "park pending operational review" - if any( - token in name - for token in ("daily brief", "daily-brief", "weekly-market", "research leads") - ): - return "dated-snapshot", "archive after reconciliation" - if any( - token in name - for token in ( - "framework", - "strategies", - "strategy", - "technical indicators", - "greeks reference", - "asset allocation", - ) - ): - return "durable-knowledge-candidate", "review through 01_ingest" - return "raw-research-or-evidence", "review through 01_ingest" - - -def fmt_timestamp(seconds: float) -> str: - return datetime.fromtimestamp(seconds).astimezone().isoformat(timespec="seconds") - - -def ensure_batch_not_registered(batch_id: str) -> None: - text = REGISTER_PATH.read_text(encoding="utf-8") - if f"| `{batch_id}` |" in text: - raise ValueError(f"batch already registered: {batch_id}") - - -def append_register(batch_id: str, source: str, file_count: int, manifest: Path) -> None: - text = REGISTER_PATH.read_text(encoding="utf-8") - row = ( - f"| `{batch_id}` | {datetime.now().astimezone().date().isoformat()} | " - f"`{source}` | {file_count} | registered | " - f"`{manifest.relative_to(PROJECT_DIR)}` |\n" - ) - REGISTER_PATH.write_text(text.rstrip() + "\n" + row, encoding="utf-8") - - -def snapshot_path(batch_dir: Path, relative_path: str) -> Path: - return batch_dir / SOURCE_SNAPSHOT_DIRNAME / Path(relative_path) - - -def copy_source_snapshot(source_path: Path, destination_path: Path) -> None: - destination_path.parent.mkdir(parents=True, exist_ok=True) - shutil.copy2(source_path, destination_path) - - -def append_disposition_entries( - ledger_path: Path, entries: list[dict[str, str]] -) -> None: - file_exists = ledger_path.exists() - with ledger_path.open("a", encoding="utf-8", newline="") as handle: - writer = csv.DictWriter(handle, fieldnames=LEDGER_FIELDS) - if not file_exists: - writer.writeheader() - if entries: - writer.writerows(entries) - - -def initialize_disposition_ledger( - batch_id: str, batch_dir: Path, records: list[dict[str, str | int]] -) -> Path: - ledger_path = batch_dir / LEDGER_FILENAME - recorded_at = datetime.now().astimezone().isoformat(timespec="seconds") - entries = [ - { - "recorded_at": recorded_at, - "batch_id": batch_id, - "relative_path": str(record["relative_path"]), - "sha256": str(record["sha256"]), - "disposition": "unresolved", - "destination": "", - "evidence": manifest_relpath(batch_id), - "notes": "Initialized at registration; no lifecycle decision recorded.", - } - for record in records - ] - append_disposition_entries(ledger_path, entries) - return ledger_path - - -def manifest_relpath(batch_id: str) -> str: - return f"raw-materials/batches/{batch_id}/manifest.csv" - - -def register_batch(batch_id: str, source_path: Path, source_label: str) -> Path: - if not BATCH_ID_RE.fullmatch(batch_id): - raise ValueError("batch ID must use YYYY-MM-DD-NNN form") - if not REGISTER_PATH.exists(): - raise FileNotFoundError( - "second-brain migration project is missing its batch register" - ) - ensure_batch_not_registered(batch_id) - if not source_path.is_dir(): - raise ValueError(f"source is not a directory: {source_path}") - - batch_dir = BATCHES_DIR / batch_id - if batch_dir.exists(): - raise ValueError(f"batch output already exists: {batch_dir}") - - files = sorted(path for path in source_path.rglob("*") if path.is_file()) - records: list[dict[str, str | int]] = [] - batch_dir.mkdir(parents=True) - for path in files: - stat = path.stat() - relative_path = path.relative_to(source_path).as_posix() - category, destination = classify(Path(relative_path)) - snapshot_file = snapshot_path(batch_dir, relative_path) - copy_source_snapshot(path, snapshot_file) - records.append( - { - "batch_id": batch_id, - "source": source_label, - "relative_path": relative_path, - "size_bytes": stat.st_size, - "sha256": sha256(snapshot_file), - "created_at": fmt_timestamp( - getattr(stat, "st_birthtime", stat.st_ctime) - ), - "modified_at": fmt_timestamp(stat.st_mtime), - "extension": path.suffix.lower() or "[none]", - "embedded_dates": embedded_dates(path), - "category": category, - "proposed_destination": destination, - "review_status": "unreviewed", - "duplicate_group": "", - } - ) - - hash_counts = Counter(str(record["sha256"]) for record in records) - duplicate_hashes = sorted(value for value, count in hash_counts.items() if count > 1) - duplicate_ids = { - value: f"dup-{index:03d}" for index, value in enumerate(duplicate_hashes, 1) - } - for record in records: - record["duplicate_group"] = duplicate_ids.get(str(record["sha256"]), "") - - manifest = batch_dir / "manifest.csv" - with manifest.open("w", encoding="utf-8", newline="") as handle: - writer = csv.DictWriter(handle, fieldnames=MANIFEST_FIELDS) - writer.writeheader() - if records: - writer.writerows(records) - ledger = initialize_disposition_ledger(batch_id, batch_dir, records) - - categories = Counter(str(record["category"]) for record in records) - duplicates = [record for record in records if str(record["duplicate_group"])] - report_lines = [ - f"# Batch Report - {batch_id}", - "", - f"- Registered: {datetime.now().astimezone().date().isoformat()}", - f"- Source: `{source_label}`", - f"- Files: {len(records)}", - f"- Source snapshot: `raw-materials/batches/{batch_id}/{SOURCE_SNAPSHOT_DIRNAME}/`", - f"- Exact duplicate files: {len(duplicates)}", - f"- Exact duplicate groups: {len(duplicate_ids)}", - f"- Disposition ledger: `{ledger.relative_to(PROJECT_DIR)}`", - "- Status: registered; classification and reconciliation pending", - "", - "## Initial Classification", - "", - "| Candidate category | Files |", - "|---|---:|", - ] - report_lines.extend( - f"| {category} | {count} |" for category, count in sorted(categories.items()) - ) - report_lines.extend( - [ - "", - "## Guardrails", - "", - "- Categories are triage suggestions, not routing decisions.", - "- Registration copies recoverable source bytes into the batch-local snapshot.", - "- No source files were moved, renamed, edited, or deleted.", - "- The disposition ledger is append-only; add later lifecycle decisions as new rows.", - "- Live-state candidates require claim-level reconciliation and source verification.", - "- Content Engine candidates must be compared with its external production workspace.", - "", - ] - ) - (batch_dir / "batch-report.md").write_text( - "\n".join(report_lines), encoding="utf-8" - ) - - duplicate_lines = [ - f"# Exact Duplicates - {batch_id}", - "", - "Hash-backed duplicate groups. No files have been deleted.", - "", - ] - if not duplicate_ids: - duplicate_lines.append("_No exact duplicates found._") - else: - for digest in duplicate_hashes: - duplicate_lines.extend( - [ - f"## {duplicate_ids[digest]}", - "", - f"`sha256:{digest}`", - "", - ] - ) - duplicate_lines.extend( - f"- `{record['relative_path']}`" - for record in records - if record["sha256"] == digest - ) - duplicate_lines.append("") - (batch_dir / "duplicates.md").write_text( - "\n".join(duplicate_lines), encoding="utf-8" - ) - - append_register(batch_id, source_label, len(records), manifest) - return batch_dir - - -def main() -> int: - parser = argparse.ArgumentParser( - description="Register a read-only snapshot of a second-brain inbox batch." - ) - parser.add_argument("--batch-id", required=True) - parser.add_argument("--source", default=str(ROOT / "00_inbox")) - parser.add_argument("--source-label") - args = parser.parse_args() - - source = Path(args.source).expanduser().resolve() - source_label = args.source_label or str(source) - try: - output = register_batch(args.batch_id, source, source_label) - except (FileNotFoundError, OSError, ValueError) as exc: - print(f"second-brain-batch: {exc}", file=sys.stderr) - return 1 - - print(f"registered {args.batch_id}: {output}") - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/bin/session-close b/bin/session-close index 5543edb..2deed96 100755 --- a/bin/session-close +++ b/bin/session-close @@ -20,6 +20,11 @@ ROOT = Path(__file__).resolve().parents[1] DRAFT_REL = Path("20_live") / "last-handoff-draft.md" FEED_REL = Path("20_live") / "workstation" / "session-close-feed.jsonl" CHECKPOINT_HEADING = "## Session Checkpoints (auto)" +# Mirrors STALE_DAYS in bin/handoff. +STALE_HANDOFF_DAYS = 21 +# The producer registry. Every entry must resolve to an action below; the +# coverage test is what stops it from becoming another unread checklist. +CONSUMPTION_REL = Path(".context") / "consumption.json" SNAPSHOT_CAP = 5 ACTIVITY_WINDOW_DAYS = 14 # Bounded per-project scan keeps --checkpoint inside the 5s hook budget even @@ -30,6 +35,11 @@ SKIP_DIR_NAMES = { ".pytest_cache", "dist", "build", ".cache", } WEEKLY_EVAL_STALE_DAYS = 8 # matches bin/eval-schedule check (ADR-036) +# Matches GRADUATION_THRESHOLD in bin/papercut: two entries is the point at +# which a cluster can exist, so it is the point at which the last harvest's +# conclusions may already be out of date. Any higher and the prompt arrives +# after the repeat it exists to catch. +PAPERCUT_HARVEST_TRIGGER = 2 @dataclass(frozen=True) @@ -177,31 +187,33 @@ def derive_active_projects(root: Path, """ if now_ts is None: now_ts = datetime.now().timestamp() - projects_dir = root / "30_projects" - if not projects_dir.is_dir(): - return [] + lifecycle_dirs = [root / "30_projects", root / "40_operations"] candidates: list[dict[str, object]] = [] - for entry in sorted(projects_dir.iterdir()): - if not entry.is_dir() or entry.name.startswith("."): - continue - signals: list[tuple[str, float]] = [] - mtime = latest_mtime(entry) - if mtime is not None: - signals.append(("files", mtime)) - commit = nested_repo_last_commit(entry, run_command) - if commit is not None: - signals.append(("commit", commit)) - if not signals: + for lifecycle_dir in lifecycle_dirs: + if not lifecycle_dir.is_dir() or lifecycle_dir.is_symlink(): continue - signal, ts = max(signals, key=lambda item: item[1]) - age_hours = max(0.0, (now_ts - ts) / 3600) - if age_hours > window_days * 24: - continue - candidates.append({ - "project": entry.name, - "signal": signal, - "age_hours": round(age_hours, 1), - }) + for entry in sorted(lifecycle_dir.iterdir()): + if not entry.is_dir() or entry.is_symlink() or entry.name.startswith("."): + continue + signals: list[tuple[str, float]] = [] + mtime = latest_mtime(entry) + if mtime is not None: + signals.append(("files", mtime)) + commit = nested_repo_last_commit(entry, run_command) + if commit is not None: + signals.append(("commit", commit)) + if not signals: + continue + signal, ts = max(signals, key=lambda item: item[1]) + age_hours = max(0.0, (now_ts - ts) / 3600) + if age_hours > window_days * 24: + continue + candidates.append({ + "project": entry.name, + "path": str(entry.relative_to(root)), + "signal": signal, + "age_hours": round(age_hours, 1), + }) candidates.sort(key=lambda item: item["age_hours"]) return candidates[:limit] @@ -477,6 +489,353 @@ def papercut_action(root: Path) -> "CloseAction": ) +def papercut_harvest_action(root: Path) -> "CloseAction": + """Prompt a harvest when the log has grown past the last reading of it. + + `papercut_action` above closes the recording half of the loop. Nothing closed + the reading half: `bin/papercut harvest` is the step that turns a pile of + entries into "this address is the defect", and the only thing that ever + called for it was `bin/contract-audit` item 10.1 — a check that reports the + backlog after the fact rather than at the moment someone could act on it. + + Deliberately `manual`, never `auto`. Harvesting writes + `20_live/papercuts/last-harvest.json`, and 10.1 reads that receipt. Running + the harvest from `--apply` would satisfy the audit with a receipt attesting + to a report that no one read — a green check standing in for the work it was + built to measure. The receipt has to mean a person looked. + + Triggers on *unread growth*, not on open issues. A papercut log is + append-only with no resolved state, so "graduation candidates exist" stays + true forever once a cluster forms, and a demand that can never be satisfied + becomes wallpaper within a week. Standing candidates are reported in the + reason line either way; only new material makes the action needed. + """ + store = root / "20_live" / "papercuts" + if not store.exists(): + return CloseAction( + name="papercut-harvest", kind="manual", needed=False, + reason="no papercut store yet", + ) + + total = 0 + diagnosed = 0 + for path in sorted(store.glob("*.jsonl")): + try: + lines = path.read_text(encoding="utf-8", errors="replace").splitlines() + except OSError: + continue + for ln in lines: + if not ln.strip(): + continue + total += 1 + try: + kind = json.loads(ln).get("kind", "") + except (json.JSONDecodeError, AttributeError): + # Unreadable lines count toward the total and not toward the + # diagnosed set. A parse failure is not evidence of an auto stub. + continue + if kind not in ("auto", "suggestion"): + diagnosed += 1 + + if total == 0: + return CloseAction( + name="papercut-harvest", kind="manual", needed=False, + reason="papercut store is empty — nothing to harvest", + ) + + receipt = store / "last-harvest.json" + try: + r = json.loads(receipt.read_text(encoding="utf-8")) if receipt.exists() else None + except (OSError, json.JSONDecodeError): + r = None + + if not isinstance(r, dict): + return CloseAction( + name="papercut-harvest", kind="manual", needed=True, + reason=(f"{total} papercut(s) recorded and never harvested — " + "run `bin/papercut harvest`"), + ) + + cands = r.get("graduation_candidates", 0) + standing = f"; {cands} standing graduation candidate(s)" if cands else "" + + # Receipts written before the auto hold-out landed carry no `auto_stubs` + # key, so diagnosed-at-harvest is unknowable from them. Fall back to total + # growth rather than guessing, which errs toward prompting. + if "auto_stubs" in r: + seen = r.get("observations_seen", 0) - r.get("auto_stubs", 0) + new = diagnosed - seen + unit = "diagnosed papercut(s)" + else: + new = total - r.get("entries_seen", 0) + unit = "papercut(s)" + + if new >= PAPERCUT_HARVEST_TRIGGER: + return CloseAction( + name="papercut-harvest", kind="manual", needed=True, + reason=(f"{new} new {unit} since the last harvest " + f"({r.get('harvested_at', 'unknown')[:10]}) — " + f"enough to form a cluster; run `bin/papercut harvest`{standing}"), + ) + return CloseAction( + name="papercut-harvest", kind="manual", needed=False, + reason=(f"harvest current as of {r.get('harvested_at', 'unknown')[:10]} " + f"({total} recorded, {max(new, 0)} unread){standing}"), + ) + + +def _handoff_json(root: Path, run_command: RunCommand) -> list[dict] | None: + """Open handoffs via bin/handoff, or None when the tool cannot be read.""" + try: + proc = run_command([str(root / "bin" / "handoff"), "list", "--json"]) + except Exception: + return None + if proc.returncode != 0: + return None + try: + payload = json.loads(proc.stdout) + except (json.JSONDecodeError, TypeError): + return None + return payload if isinstance(payload, list) else None + + +def handoff_lifecycle_action(root: Path, run_command: RunCommand) -> CloseAction: + """Report the open handoff set and what closing the session owes it. + + A handoff is an object with a lifecycle (open -> consumed/superseded), so + session close has two obligations: consume what this session finished, and + emit a handoff for work that outlives it. + """ + handoffs = _handoff_json(root, run_command) + if handoffs is None: + return CloseAction( + name="handoff-lifecycle", kind="warn", needed=True, + reason="could not read open handoffs — run `bin/handoff check`", + ) + if not handoffs: + return CloseAction( + name="handoff-lifecycle", kind="manual", needed=True, + reason="no open handoffs — emit one with `bin/handoff emit` if work " + "outlives this session", + ) + stale = [h for h in handoffs if (h.get("age_days") or 0) > STALE_HANDOFF_DAYS] + stale_note = f", {len(stale)} stale" if stale else "" + return CloseAction( + name="handoff-lifecycle", kind="manual", needed=True, + reason=f"{len(handoffs)} open handoff(s){stale_note} — " + "`bin/handoff consume <id> --apply` for anything this session " + "finished; `bin/handoff emit` for what it starts", + ) + + +def handoff_digest_lines(root: Path, run_command: RunCommand) -> list[str]: + """Open handoffs as digest bullets, newest-blocking-first.""" + handoffs = _handoff_json(root, run_command) + if handoffs is None: + return ["(could not read handoffs — run `bin/handoff check`)"] + if not handoffs: + return ["None open."] + rank = {"high": 0, "medium": 1, "low": 2} + handoffs.sort(key=lambda h: (rank.get(str(h.get("urgency")), 3), h.get("updated", ""))) + lines = [] + for h in handoffs: + age = h.get("age_days") + age_s = f", {age}d open" if age is not None else "" + lines.append( + f"- {h.get('to_project','?')} — {h.get('urgency','?')} · " + f"`{h.get('next_loop','?')}`{age_s} — {h.get('id','?')}" + ) + return lines + + +def _tool_json(root: Path, argv: list[str], run_command: RunCommand): + """A tool's --json payload, or None when it cannot be read. + + None is deliberately distinct from an empty result: "the check did not run" + and "the check found nothing" must never render the same, which is the + distinction bin/capture-validate had to be taught the hard way. + """ + try: + proc = run_command([str(root / "bin" / argv[0]), *argv[1:]]) + except Exception: # noqa: BLE001 — a close check must not crash the hook + return None + if not (proc.stdout or "").strip(): + return None + try: + return json.loads(proc.stdout) + except (json.JSONDecodeError, TypeError): + return None + + +def idle_runs_action(root: Path, run_command: RunCommand) -> CloseAction: + """Idle runs that produced an artifact and never got a terminal state. + + The rule this enforces is the one this repo breaks most: nothing may be + generated that has no scheduled consumption pass. bin/idle-dispatch was + built to close the actuation gap and reproduced it within three days — + 33 starts against 18 finishes. The dispatcher was never the missing piece. + """ + payload = _tool_json(root, ["idle-queue", "backlog", "--json"], run_command) + if payload is None: + return CloseAction( + name="idle-runs", kind="warn", needed=True, + reason="could not read the idle run backlog — run `bin/idle-queue backlog`", + ) + open_n = int(payload.get("open") or 0) + if not open_n: + return CloseAction( + name="idle-runs", kind="manual", needed=False, + reason="no idle runs waiting on a reader", + ) + stale_n = int(payload.get("stale") or 0) + oldest = payload.get("oldest_days") + detail = f", {stale_n} past {payload.get('stale_days')}d" if stale_n else "" + if oldest: + detail += f" (oldest {oldest}d)" + return CloseAction( + name="idle-runs", kind="manual", needed=True, + reason=f"{open_n} idle run(s) unreviewed{detail} — " + "`bin/idle-queue finish <run> --outcome <label>`, or " + "`bin/idle-queue abandon <run>` if nobody is going to read it", + ) + + +def scheduled_jobs_action(root: Path, run_command: RunCommand) -> CloseAction: + """launchd jobs that failed, half-succeeded, or went quiet. + + `degraded` is the state worth surfacing separately: measured 2026-08-26, + repo-radar wrote its digest and then died on DNS, and the only thing recorded + anywhere was the exit 1. Silence read as health for three days. + """ + payload = _tool_json(root, ["scheduled-run", "--json"], run_command) + if payload is None: + return CloseAction( + name="scheduled-jobs", kind="warn", needed=True, + reason="could not read launchd job health — run `bin/scheduled-run --check`", + ) + bad = [j for j in payload if j.get("status") in ("degraded", "failed", "stale")] + if not bad: + return CloseAction( + name="scheduled-jobs", kind="warn", needed=False, + reason=f"{len(payload)} scheduled job(s), all healthy", + ) + summary = ", ".join(f"{j.get('label','?').removeprefix('com.mainframe.')} " + f"{j.get('status')}" for j in bad[:4]) + return CloseAction( + name="scheduled-jobs", kind="warn", needed=True, + reason=f"{len(bad)}/{len(payload)} scheduled job(s) need attention — " + f"{summary} — `bin/scheduled-run --check`", + ) + + +def ingest_lane_action(root: Path, changed: list[str]) -> CloseAction: + """What this session added to the ingest lane and did not route. + + Deliberately scoped to this session's own contribution. The standing lane is + 282 files deep; reporting that number every close would be a permanent red, + and a permanent red is a warning nobody reads. What a session owes is what it + added. + """ + added = [f for f in changed if f.startswith("01_ingest/") or f.startswith("00_inbox/")] + if not added: + return CloseAction( + name="ingest-lane", kind="manual", needed=False, + reason="this session added nothing to the ingest lane", + ) + return CloseAction( + name="ingest-lane", kind="manual", needed=True, + reason=f"{len(added)} file(s) added to the ingest lane this session — " + "route or reject them rather than leaving them to age", + ) + + +def leak_scan_action(root: Path, run_command: RunCommand) -> CloseAction: + """Machine paths or credentials this session could still remove. + + Scoped to the working tree with `--no-history`, for two reasons. It runs in + 0.27s against 13.9s for the full scan, and this hook has a budget. And + history is a 127-string backlog needing a filter-repo rewrite — reporting it + at every close would be a permanent red, which is a warning nobody reads. + Growth of that backlog is guarded by `.leak-scan-baseline` on the weekly + run instead. + + Untracked files count here. Until 2026-08-27 the scanner never read them, + so it reported clean on files it had never opened. + """ + try: + proc = run_command([str(root / "bin" / "repo-leak-scan"), + str(root), "--no-history"]) + except Exception: # noqa: BLE001 — a close check must not crash the hook + return CloseAction( + name="leak-scan", kind="warn", needed=True, + reason="could not run the leak scan — run `bin/repo-leak-scan --no-history`", + ) + if proc.returncode == 0: + return CloseAction( + name="leak-scan", kind="warn", needed=False, + reason="no machine paths or credentials in the working tree", + ) + if proc.returncode != 1: + # ADR-052: 2 means the tool could not run. That is never a pass. + return CloseAction( + name="leak-scan", kind="warn", needed=True, + reason=f"leak scan could not complete (exit {proc.returncode}) — " + "run `bin/repo-leak-scan --no-history`", + ) + offenders = sorted({ + line.split(":", 1)[0] + for line in (proc.stdout or "").splitlines() + if ":" in line and not line.startswith((" ", "\x1b", "=", "T", "#")) + }) + named = ", ".join(offenders[:3]) if offenders else "see the scan output" + return CloseAction( + name="leak-scan", kind="warn", needed=True, + reason=f"leak(s) in the working tree — {named} — " + "remove them or declare them in `.leak-scan-allow`", + ) + + +def brainstorm_sync_action(root: Path, run_command: RunCommand) -> CloseAction: + """Drift between the brainstorm dump, its mirror, and SOURCE.md. + + Deliberately runs without `--fetch`: pin, orphans and index are local reads + that cost nothing, and a SessionEnd hook must not depend on the network. + The `behind` check reports not-assessable here and is covered by the weekly + job instead. + + Measured 2026-08-27, all four had drifted at once — the mirror 30 commits + behind, the pin three days stale, a capture orphaned in the mirror for nine + days, and 8 of 14 captures missing from the index. + """ + payload = _tool_json(root, ["brainstorm-sync", "--json"], run_command) + if payload is None: + return CloseAction( + name="brainstorm-sync", kind="warn", needed=True, + reason="could not read brainstorm sync state — run `bin/brainstorm-sync`", + ) + drifted = [c for c in payload if c.get("state") == "drift"] + if not drifted: + return CloseAction( + name="brainstorm-sync", kind="warn", needed=False, + reason="brainstorm mirror, pin and index agree", + ) + return CloseAction( + name="brainstorm-sync", kind="warn", needed=True, + reason=f"{len(drifted)} brainstorm drift(s) — " + + "; ".join(str(c.get("detail", "")) for c in drifted[:2]), + ) + + +def consumption_registry(root: Path) -> list[dict]: + """Registered producers. Raises nothing — an unreadable registry is reported + by the coverage test, not by crashing a SessionEnd hook.""" + path = root / CONSUMPTION_REL + try: + return json.loads(path.read_text(encoding="utf-8")).get("producers", []) + except Exception: # noqa: BLE001 + return [] + + def check_session(root: Path, project: str | None = None, run_command: RunCommand = default_run_command, apply: bool = False) -> SessionCloseResult: @@ -492,9 +851,14 @@ def check_session(root: Path, project: str | None = None, else: result.actions.append(CloseAction( name="state-md-narrative", kind="manual", needed=True, - reason="update STATE.md narrative (what changed, what remains, blockers, handoff)", + reason="update STATE.md narrative (what changed, what remains, blockers); " + "the Current Handoff section comes from `bin/handoff state-block`", )) + # Handoff lifecycle: work that outlives the session becomes a linked file + # with a consumption state, not another paragraph in STATE.md. + result.actions.append(handoff_lifecycle_action(root, run_command)) + if project is None: project = detect_project_from_state(state_path) if state_exists else None @@ -529,6 +893,15 @@ def check_session(root: Path, project: str | None = None, )) result.actions.append(papercut_action(root)) + result.actions.append(papercut_harvest_action(root)) + + # Producers registered in .context/consumption.json. Each must land an + # action here or tests/test_session_close.py fails. + result.actions.append(idle_runs_action(root, run_command)) + result.actions.append(scheduled_jobs_action(root, run_command)) + result.actions.append(leak_scan_action(root, run_command)) + result.actions.append(brainstorm_sync_action(root, run_command)) + result.actions.append(ingest_lane_action(root, changed)) knowledge_changed = any(f.startswith("10_knowledge/") for f in changed) if knowledge_changed: @@ -613,6 +986,9 @@ def check_session(root: Path, project: str | None = None, digest_lines.append(f"Date: {date.today().isoformat()}") digest_lines.append("Active project: " + (project or "(from STATE.md)")) digest_lines.append("") + digest_lines.append("## Open handoffs") + digest_lines.extend(handoff_digest_lines(root, run_command)) + digest_lines.append("") digest_lines.append("## Quick Signals (paste/edit into STATE.md)") # Ingest status try: @@ -697,7 +1073,7 @@ def main() -> int: help="fast prompt-free snapshot: append a dated block to " "20_live/last-handoff-draft.md and a record to the " "tracker feed (wired to the PreCompact hook)") - parser.add_argument("--project", help="project slug under 30_projects/") + parser.add_argument("--project", help="lifecycle slug under 30_projects/ or 40_operations/") parser.add_argument("--feed", action="store_true", help="with --check/--apply: append the machine-readable " "result to 20_live/workstation/session-close-feed.jsonl " diff --git a/bin/session-open b/bin/session-open index 4f7f8a0..6025800 100755 --- a/bin/session-open +++ b/bin/session-open @@ -1,5 +1,5 @@ #!/usr/bin/env python3 -"""Load session context files in a fixed order for progressive disclosure. +"""List session context and applicable contracts for progressive disclosure. Unit 2.2: fail closed when a selected project does not resolve. MPE-024: prefer structured focus (20_live/focus/current.yaml); STATE.md is @@ -9,6 +9,7 @@ narrative fallback and must not invent compound slugs. from __future__ import annotations import argparse +import hashlib import json import re import subprocess @@ -23,8 +24,10 @@ if str(SCRIPTS) not in sys.path: sys.path.insert(0, str(SCRIPTS)) from focus_authority import load_focus # noqa: E402 +from lifecycle_identity import IdentityError, resolve_record # noqa: E402 _SLUG_RE = re.compile(r"^[A-Za-z0-9][A-Za-z0-9._-]*$") +BATCH_BYTES = 8000 @dataclass(frozen=True) @@ -34,6 +37,8 @@ class ContextEntry: exists: bool size: int | None note: str + required: bool = False + section: str | None = None @dataclass @@ -49,19 +54,55 @@ class SessionOpenResult: focus_path: str | None = None focus_errors: list[str] = field(default_factory=list) focus_warnings: list[str] = field(default_factory=list) + task_path: str | None = None + path_error: str | None = None + intent: str = "arrival" + intent_error: str | None = None + task: str | None = None + focused_project: str | None = None + plan: str | None = None + plan_status: str = "not_requested" + plan_error: str | None = None + plan_candidates: list[dict] = field(default_factory=list) + read_batches: list[dict] = field(default_factory=list) + read_errors: list[str] = field(default_factory=list) + + @property + def missing_required(self) -> list[str]: + return [e.path for e in self.entries if e.required and not e.exists] @property def ok(self) -> bool: - required = {"AGENTS.md", "STATE.md"} - found = {e.path for e in self.entries if e.exists} - if not required.issubset(found): + if self.missing_required: return False - if self.project_error: + if self.project_error or self.path_error or self.intent_error or self.plan_error or self.read_errors: return False if self.project is not None and self.project_exists is False: return False return True + @property + def context_status(self) -> str: + return "unread" if self.ok else "incomplete" + + @property + def stop_when(self) -> list[str]: + common = [ + "Read every required batch completely; recover any truncated output before relying on it.", + "State the requested task, governing sources, applicable constraints, relevant unknowns, and next useful step.", + "Report incomplete context before dependent claims or action; file presence does not establish reading or readiness.", + ] + if self.intent == "arrival": + return common + [ + "Describe the lifecycle map and recorded focus with its freshness caveat, then stop at workspace orientation.", + "Before project-specific diagnosis or recommendations, switch to resume with that project.", + ] + return common + [ + "Verify the relevant plan, coordination, Git/worktrees, candidate/experiment identity and direct receipts under the reconstruction workflow.", + "Use existing valid receipts for status questions; rerun checks only for changed source, missing evidence, or an explicit test request.", + "After reconstruction, answer the request; load an action workflow only when that action is authorized.", + ] + def is_valid_project_slug(label: str) -> bool: label = label.strip() @@ -91,33 +132,171 @@ def detect_project_from_state(state_path: Path) -> str | None: return None -def find_active_phase_plan(plans_dir: Path) -> Path | None: +def entry(order: int, rel: str, root: Path, note: str, + *, required: bool = False, section: str | None = None) -> ContextEntry: + full = root / rel + exists = full.is_file() + size = full.stat().st_size if exists else None + return ContextEntry(order=order, path=rel, exists=exists, size=size, + note=note, required=required, section=section) + + +def _entry_bytes(root: Path, item: ContextEntry) -> tuple[bytes, str]: + path = root / item.path + if not path.resolve().is_relative_to(root.resolve()): + raise ValueError("context file resolves outside MainFrame") + data = path.read_bytes() + source_hash = hashlib.sha256(data).hexdigest() + data.decode("utf-8") # A non-text prerequisite must not become a partial read. + if item.section == "@latest": + lines = data.splitlines(keepends=True) + headings = [i for i, line in enumerate(lines) if line.startswith(b"## ")] + if len(headings) > 1: + return b"".join(lines[:headings[1]]), source_hash + elif item.section: + lines = data.splitlines(keepends=True) + marker = ("## " + item.section).encode() + start = next((i for i, line in enumerate(lines) if line.strip() == marker), None) + if start is not None: + end = next((i for i in range(start + 1, len(lines)) if lines[i].startswith(b"## ")), len(lines)) + return b"".join(lines[start:end]), source_hash + # Older contracts are read in full rather than silently omitted. + return data, source_hash + + +def build_read_batches(result: SessionOpenResult, root: Path) -> None: + for item in result.entries: + if not item.exists: + continue + try: + data, source_hash = _entry_bytes(root, item) + start = 0 + while start < len(data): + end = min(start + BATCH_BYTES, len(data)) + if end < len(data): + newline = data.rfind(b"\n", start, end) + if newline >= start: + end = newline + 1 + else: + while end > start and data[end] & 0xC0 == 0x80: + end -= 1 + result.read_batches.append({ + "id": len(result.read_batches) + 1, "path": item.path, + "required": item.required, "section": item.section, + "start_byte": start, "end_byte": end, "bytes": end - start, + "source_sha256": source_hash, + }) + start = end + except (OSError, ValueError, UnicodeError) as exc: + result.read_errors.append(f"{item.path}: {exc}") + + +def read_batch(result: SessionOpenResult, root: Path, batch_id: int) -> dict: + if batch_id < 1 or batch_id > len(result.read_batches): + raise ValueError(f"batch must be between 1 and {len(result.read_batches)}") + batch = result.read_batches[batch_id - 1] + item = next(e for e in result.entries if e.path == batch["path"]) + data, source_hash = _entry_bytes(root, item) + if source_hash != batch["source_sha256"]: + raise ValueError("source changed after routing; reopen the context list") + return {**batch, "content": data[batch["start_byte"]:batch["end_byte"]].decode("utf-8")} + + +def select_plan(result: SessionOpenResult, root: Path, project_dir: Path, + explicit_plan: str | None) -> Path | None: + plans_dir = project_dir / "plans" + if plans_dir.exists() and (not plans_dir.is_dir() or not plans_dir.resolve().is_relative_to(project_dir.resolve())): + result.plan_error = "plans/ must be a directory inside the selected project" + result.plan_status = "invalid" + return None + if explicit_plan: + candidate = (root / explicit_plan).resolve() + if not candidate.is_relative_to(plans_dir.resolve()) or not candidate.is_file(): + result.plan_error = "--plan must be an existing file inside the selected project's plans/" + result.plan_status = "invalid" + return None + result.plan_status = "explicit" + return candidate if not plans_dir.is_dir(): + result.plan_status = "none" return None - for plan in sorted(plans_dir.glob("*.md")): - text = plan.read_text(encoding="utf-8") - if not text.startswith("---\n"): + tokens = set(re.findall(r"[a-z0-9]+", (result.task or "").lower())) - { + "the", "a", "an", "and", "or", "for", "to", "of", "in", "work", "project", "mainframe", "process", "eval", + } + candidates = [] + for plan in sorted(plans_dir.rglob("*.md")): + if not plan.is_file() or not plan.resolve().is_relative_to(plans_dir.resolve()): continue - end = text.find("\n---", 4) - if end == -1: + try: + with plan.open(encoding="utf-8", errors="replace") as handle: + text = handle.read(4096) + except OSError as exc: + result.plan_error = f"cannot inspect plan metadata: {plan.relative_to(root)}: {exc}" + result.plan_status = "incomplete" + return None + fm = text.split("---", 2)[1] if text.startswith("---\n") and text.count("---") >= 2 else "" + if not re.search(r"^status:\s*[\"']?active[\"']?\s*$", fm, re.MULTILINE): continue - for fm_line in text[4:end].splitlines(): - if ":" not in fm_line: - continue - key, value = fm_line.split(":", 1) - if key.strip() == "status" and value.strip().strip('"').strip("'") == "active": - return plan + title_match = re.search(r"^title:\s*(.+)$", fm, re.MULTILINE) + title = title_match.group(1).strip().strip("\"'") if title_match else plan.stem + terms = set(re.findall(r"[a-z0-9]+", (plan.stem + " " + title).lower())) + candidates.append({"path": str(plan.relative_to(root)), "title": title, "matched_terms": sorted(tokens & terms)}) + candidates.sort(key=lambda p: (-len(p["matched_terms"]), p["path"])) + result.plan_candidates = candidates + if not candidates: + result.plan_status = "none" + elif tokens and candidates[0]["matched_terms"] and (len(candidates) == 1 or len(candidates[0]["matched_terms"]) > len(candidates[1]["matched_terms"])): + result.plan_status = "task_match_candidate" + return root / candidates[0]["path"] + else: + result.plan_status = "ambiguous" if tokens else "task_needed" return None -def entry(order: int, rel: str, root: Path, note: str) -> ContextEntry: - full = root / rel - exists = full.exists() - size = full.stat().st_size if exists else None - return ContextEntry(order=order, path=rel, exists=exists, size=size, note=note) +def _project_contracts(result: SessionOpenResult, root: Path, + project_dir: Path, order: int) -> int: + """Inspect only the selected project's ancestors, never sibling trees.""" + order += 1 + lifecycle_contract = ( + "40_operations/AGENTS.md" + if project_dir.parent.name == "40_operations" + else "30_projects/AGENTS.md" + ) + result.entries.append(entry(order, lifecycle_contract, root, + "project lifecycle contract", required=True)) + directories = [project_dir] + if result.task_path is not None and result.path_error is None: + target = (root / result.task_path).resolve() + project_root = project_dir.resolve() + if not target.is_relative_to(project_root): + result.path_error = "--path must stay inside the selected project" + elif not target.exists(): + result.path_error = "--path does not exist; use an existing parent directory" + else: + directory = target if target.is_dir() else target.parent + result.task_path = str(target.relative_to(root)) + current = project_dir + for part in directory.relative_to(project_root).parts: + current = current / part + directories.append(current) + + for directory in directories: + contract = directory / "AGENTS.md" + # A broken symlink or a directory named AGENTS.md must remain visible. + if contract.exists() or contract.is_symlink(): + order += 1 + result.entries.append(entry(order, str(contract.relative_to(root)), + root, "local contract", required=True)) + order += 1 + result.entries.append(entry( + order, ".context/workflows/project-resume-and-candidate-lifecycle.md", + root, "project reconstruction before planning or acting", required=True, + )) + return order -def _apply_project(result: SessionOpenResult, root: Path, order: int) -> int: +def _apply_project(result: SessionOpenResult, root: Path, order: int, + explicit_plan: str | None = None) -> int: if not result.project: return order raw = result.project.strip() @@ -125,7 +304,7 @@ def _apply_project(result: SessionOpenResult, root: Path, order: int) -> int: result.project_error = ( "compound or invalid project label cannot resolve to a project slug; " "set structured focus primary.project or STATE Active Project to a " - "single 30_projects/<slug> name" + "single lifecycle slug" ) result.degraded = True result.project_exists = False @@ -142,10 +321,18 @@ def _apply_project(result: SessionOpenResult, root: Path, order: int) -> int: return order slug = raw - project_dir = root / "30_projects" / slug - readme_rel = f"30_projects/{slug}/README.md" + try: + record = resolve_record(root, slug) + except IdentityError as exc: + result.project_error = f"selected lifecycle identity is unresolved: {exc}" + result.project_exists = False + result.degraded = True + return order + project_dir = record.path + order = _project_contracts(result, root, project_dir, order) + readme_rel = f"{record.relative_path(root)}/README.md" result.project_path = readme_rel - readme_entry = entry(order + 1, readme_rel, root, "active project") + readme_entry = entry(order + 1, readme_rel, root, "selected project", required=True) order += 1 result.entries.append(readme_entry) result.project_exists = readme_entry.exists @@ -153,27 +340,65 @@ def _apply_project(result: SessionOpenResult, root: Path, order: int) -> int: result.project_error = f"selected project path missing: {readme_rel}" result.degraded = True else: - plans_dir = project_dir / "plans" - phase_plan = find_active_phase_plan(plans_dir) + project_metadata = project_dir / "PROJECT.md" + if project_metadata.exists(): + order += 1 + result.entries.append(entry( + order, str(project_metadata.relative_to(root)), root, + "project coordination metadata (public-mirror convention)", required=True, + )) + for filename in ("methodology-approach.md", "log.md", "decisions.md"): + path = project_dir / filename + if path.exists() or path.is_symlink(): + order += 1 + result.entries.append(entry(order, str(path.relative_to(root)), root, + "project coordination; latest entry only for logs/decisions; follow relevant older evidence", + required=True, section="@latest" if filename in {"log.md", "decisions.md"} else None)) + phase_plan = select_plan(result, root, project_dir, explicit_plan) if phase_plan: order += 1 plan_rel = str(phase_plan.relative_to(root)) - result.entries.append(entry(order, plan_rel, root, "active phase plan")) + result.plan = plan_rel + result.entries.append(entry(order, plan_rel, root, + "explicit plan" if explicit_plan else "task-matched plan candidate; verify against this task", + required=bool(explicit_plan))) return order -def build_context(root: Path, project: str | None = None) -> SessionOpenResult: +def build_context(root: Path, project: str | None = None, + task_path: str | None = None, *, intent: str | None = None, + task: str | None = None, plan: str | None = None) -> SessionOpenResult: + root = root.resolve() result = SessionOpenResult() + result.intent = intent or ("resume" if project else "arrival") + result.task = task + if result.intent not in {"arrival", "resume"}: + result.intent_error = "intent must be arrival or resume" + if result.intent == "arrival" and (project or task_path or plan): + result.intent_error = "arrival does not enter a project; use --intent resume with --project" + if plan and not project: + result.plan_error = "--plan requires an explicit --project" + result.task_path = task_path + if task_path is not None and project is None: + result.path_error = "--path requires an explicit --project" order = 0 order += 1 - result.entries.append(entry(order, "AGENTS.md", root, "root control file")) + result.entries.append(entry(order, "AGENTS.md", root, "root control file", required=True)) + + order += 1 + result.entries.append(entry(order, "HARNESS.md", root, "harness contract", required=True, + section="Session orientation" if result.intent == "arrival" else None)) + + order += 1 + result.entries.append(entry(order, "STATE.md", root, "workspace narrative", required=True)) order += 1 - result.entries.append(entry(order, "STATE.md", root, "workspace state")) + result.entries.append(entry(order, ".context/workflows/session-open.md", root, + "routing, complete reads and stopping point", required=True)) # 1) Explicit flag wins - if project: + if project and result.intent == "resume": result.project = project result.project_source = "flag" else: @@ -184,17 +409,18 @@ def build_context(root: Path, project: str | None = None) -> SessionOpenResult: result.focus_revision = focus.revision result.focus_errors = list(focus.errors) result.focus_warnings = list(focus.warnings) + result.focused_project = focus.primary_project order += 1 result.entries.append( entry(order, result.focus_path, root, "structured focus authority") ) - if focus.ok and focus.primary_project: + if result.intent == "resume" and focus.ok and focus.primary_project: result.project = focus.primary_project result.project_source = "focus" - elif focus.errors: + if focus.errors: result.degraded = True # 3) STATE narrative fallback (migration window) - if result.project is None: + if result.intent == "resume" and result.project is None: detected = detect_project_from_state(root / "STATE.md") if detected: result.project = detected @@ -205,12 +431,15 @@ def build_context(root: Path, project: str | None = None) -> SessionOpenResult: "using STATE.md fallback; structured focus did not supply primary" ] - order = _apply_project(result, root, order) + if result.intent == "resume": + order = _apply_project(result, root, order, explicit_plan=plan) + if not result.project: + result.project_error = "resume requires a project; use --project when focus cannot resolve one" order += 1 result.entries.append( ContextEntry(order=order, path="(task-local)", exists=False, size=None, - note="load task-local docs manually when ready") + note="load relevant docs/code; use --project and --path for deeper contracts") ) order += 1 @@ -219,11 +448,22 @@ def build_context(root: Path, project: str | None = None) -> SessionOpenResult: note="evidence loading deferred") ) + build_read_batches(result, root) + result.degraded = bool(result.degraded or result.focus_errors or result.focus_warnings or + result.missing_required or result.path_error or + result.project_error or result.intent_error or result.plan_error or result.read_errors) return result def print_result(result: SessionOpenResult, root: Path, print_contents: bool = False) -> None: + print(f"intent: {result.intent}; context_status: {result.context_status} (reading is not verified)") + if result.intent_error: + print(f"intent_error: {result.intent_error}") + if result.plan_error: + print(f"plan_error: {result.plan_error}") + for error in result.read_errors: + print(f"read_error: {error}") if result.project: print(f"project: {result.project} (from {result.project_source})") else: @@ -232,12 +472,18 @@ def print_result(result: SessionOpenResult, root: Path, print(f"focus_revision: {result.focus_revision}") if result.project_error: print(f"project_error: {result.project_error}") + if result.task_path: + print(f"task_path: {result.task_path}") + if result.path_error: + print(f"path_error: {result.path_error}") + if result.missing_required: + print(f"missing_required: {', '.join(result.missing_required)}") if result.focus_errors: print(f"focus_errors: {'; '.join(result.focus_errors)}") if result.focus_warnings: print(f"focus_warnings: {'; '.join(result.focus_warnings)}") if result.degraded: - print("status: DEGRADED — session context is not fully valid") + print("status: DEGRADED — inspect context errors and freshness warnings") print() for e in result.entries: @@ -249,48 +495,145 @@ def print_result(result: SessionOpenResult, root: Path, else: print(f" {e.order}. {'[MISSING]':>12} {e.path} — {e.note}") - if print_contents: - print() - for e in result.entries: - if not e.exists or e.path.startswith("("): - continue - full = root / e.path - print(f"--- BEGIN {e.path} ---") - print(full.read_text(encoding="utf-8"), end="") - print(f"--- END {e.path} ---") - print() + print(f"\nplan_status: {result.plan_status}; selected: {result.plan or 'none'}") + for candidate in result.plan_candidates: + print(f" plan candidate: {candidate['path']} — {candidate['title']}") + print(f"\nread_batches: {len(result.read_batches)}; at most {BATCH_BYTES} content bytes each") + for batch in result.read_batches: + print(f" {batch['id']}. {batch['path']} [{batch['start_byte']}:{batch['end_byte']}] required={batch['required']}") + for condition in result.stop_when: + print(f" stop when: {condition}") + if print_contents and result.read_batches: + batch = read_batch(result, root, 1) + print(f"\n--- BEGIN {batch['path']} ---") + print(batch['content'], end="") + print(f"\n--- END {batch['path']} ---") + print("Only batch 1 printed. Read the other required batches separately with --read-batch N.") + + +def scheduler_status(root: Path) -> tuple[bool | None, str]: + try: + proc = subprocess.run([str(root / "bin/eval-schedule"), "check"], cwd=root, + capture_output=True, text=True, timeout=10) + except (OSError, subprocess.TimeoutExpired) as exc: + return None, f"unavailable: {type(exc).__name__}" + lines = (proc.stdout or proc.stderr).strip().splitlines() + return proc.returncode == 0, lines[0] if lines else "unknown" + + +def brain_status(root: Path) -> dict[str, Any]: + """Retrieve fast (<2ms, 0 GPU) Brain executive briefing from local SQLite state.""" + try: + sys_path_brain = str(root / "30_projects/mainframe-brain") + if sys_path_brain not in sys.path: + sys.path.insert(0, sys_path_brain) + from engine import storage + storage.seed_initial_state() + attention = storage.get_attention() + cards = storage.get_action_cards(status="pending") + lessons = storage.get_lessons(limit=1) + p0 = next((a for a in attention if a["priority"] == "P0"), None) + return { + "available": True, + "p0_project": p0["project"] if p0 else None, + "p0_reason": p0["reason"] if p0 else None, + "pending_action": cards[0]["action"] if cards else None, + "recent_lesson": f"[{lessons[0]['domain']}] {lessons[0]['rule']}" if lessons else None, + } + except Exception: + return {"available": False} + + +def record_session_arrival(root: Path, result: SessionOpenResult) -> None: + """Ingest session arrival event into Brain sensory afference (<2ms, fails open).""" + try: + sys_path_brain = str(root / "30_projects/mainframe-brain") + if sys_path_brain not in sys.path: + sys.path.insert(0, sys_path_brain) + from engine import storage + storage.add_event( + source="session", + summary=f"Session arrival: {result.intent} (project={result.project or 'none'})", + details=f"task={result.task or 'none'}, ok={result.ok}", + ) + except Exception: + pass def main() -> int: + parser = argparse.ArgumentParser( - description="Load session context files in fixed order." + description="List required context paths; does not establish project readiness." ) - parser.add_argument("--project", help="project slug under 30_projects/") + parser.add_argument("--project", help="project slug under 30_projects/ or 40_operations/") + parser.add_argument("--intent", choices=("arrival", "resume"), + help="default: arrival; --project implies resume") + parser.add_argument("--task", help="brief task phrase used to nominate a relevant active plan") + parser.add_argument("--plan", help="explicit existing plan path inside the selected project's plans/") + parser.add_argument("--path", help="existing path inside --project; relative to MainFrame or absolute") parser.add_argument("--print-contents", action="store_true", - help="print file contents after the checklist") + help="print only the first bounded content batch after the checklist") + parser.add_argument("--read-batch", type=int, help="print one numbered content batch; repeat the same routing arguments") + parser.add_argument("--expect-hash", help="with --read-batch, reject a changed source SHA-256 from the earlier listing") parser.add_argument("--json", action="store_true", help="emit structured JSON instead of human text") args = parser.parse_args() + if args.path is not None and args.project is None: + parser.error("--path requires --project") + if args.expect_hash and args.read_batch is None: + parser.error("--expect-hash requires --read-batch") + if args.read_batch is not None and args.print_contents: + parser.error("choose --read-batch or --print-contents") + + result = build_context(ROOT, project=args.project, task_path=args.path, + intent=args.intent, task=args.task, plan=args.plan) + record_session_arrival(ROOT, result) + + if args.read_batch is not None: + try: + batch = read_batch(result, ROOT, args.read_batch) + if args.expect_hash and args.expect_hash != batch["source_sha256"]: + raise ValueError("source hash differs from the earlier listing; reopen context") + except (OSError, ValueError, UnicodeError) as exc: + print(json.dumps({"ok": False, "context_status": "incomplete", "read_error": str(exc)})) + return 1 + if args.json: + print(json.dumps({"ok": result.ok, "context_status": result.context_status, + "reading_verified": False, "batch": batch}, ensure_ascii=False)) + else: + print(f"batch {batch['id']}: {batch['path']}; source_sha256={batch['source_sha256']}") + print(batch["content"], end="") + print("\nReading is not verified; complete every required batch and the task's stopping conditions.") + return 0 if result.ok else 1 - result = build_context(ROOT, project=args.project) - - eval_proc = subprocess.run( - [str(ROOT / "bin" / "eval-schedule"), "check"], - cwd=str(ROOT), - capture_output=True, - text=True, - ) - eval_ok = eval_proc.returncode == 0 - eval_summary = eval_proc.stdout.strip().splitlines()[0] if eval_proc.stdout.strip() else "unknown" + eval_ok, eval_summary = scheduler_status(ROOT) + b_stat = brain_status(ROOT) if args.json: out = { "ok": result.ok, + "intent": result.intent, + "intent_error": result.intent_error, + "task": result.task, + "context_status": result.context_status, + "reading_verified": False, + "stop_when": result.stop_when, + "focused_project": result.focused_project, + "plan": result.plan, + "plan_status": result.plan_status, + "plan_error": result.plan_error, + "plan_candidates": result.plan_candidates, + "read_batches": result.read_batches, + "read_errors": result.read_errors, + "batch_content_bytes_max": BATCH_BYTES, "project": result.project, "project_source": result.project_source, "project_path": result.project_path, "project_exists": result.project_exists, "project_error": result.project_error, + "task_path": result.task_path, + "path_error": result.path_error, + "missing_required": result.missing_required, "degraded": result.degraded, "focus_revision": result.focus_revision, "focus_path": result.focus_path, @@ -299,12 +642,19 @@ def main() -> int: "entries": [asdict(e) for e in result.entries], "eval_schedule_ok": eval_ok, "eval_schedule_summary": eval_summary, + "brain": b_stat, } json.dump(out, sys.stdout, indent=2) print() else: print_result(result, ROOT, print_contents=args.print_contents) print() + if b_stat.get("available") and b_stat.get("p0_project"): + print(f"brain: [P0] {b_stat['p0_project']} — {b_stat['p0_reason']}") + if b_stat.get("pending_action"): + print(f"brain-action: pending — {b_stat['pending_action'][:140]}") + if b_stat.get("recent_lesson"): + print(f"brain-lesson: {b_stat['recent_lesson'][:140]}") if eval_ok: print(f"eval-schedule: {eval_summary}") action_card = ROOT / "20_live" / "eval-registry" / "last-eval-action.md" @@ -332,7 +682,8 @@ def main() -> int: if primary: print(f"eval-action: {primary[:160]}") else: - print("eval-schedule: ATTENTION — run `bin/eval-schedule check` and read 20_live/eval-registry/OPERATOR.md") + print(f"eval-schedule: ATTENTION — {eval_summary}; affects scheduled-evaluation health, not permission for an unrelated task") + if not result.ok: print("session-open: FAIL — context is not valid (see project_error / missing entries)") diff --git a/bin/skill-eval b/bin/skill-eval deleted file mode 100755 index 796f5fd..0000000 --- a/bin/skill-eval +++ /dev/null @@ -1,14 +0,0 @@ -#!/usr/bin/env bash -# G10: PATH-safe skill-eval wrapper. -set -euo pipefail -ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -VENV_PY="$ROOT/30_projects/skill-eval-workshop/workbench/.venv/bin/python" -if [[ -x "$VENV_PY" ]]; then - exec "$VENV_PY" -m skill_eval "$@" -elif [[ -x /opt/homebrew/bin/python3 ]]; then - exec /opt/homebrew/bin/python3 -m skill_eval "$@" -elif [[ -x /usr/bin/python3 ]]; then - exec /usr/bin/python3 -m skill_eval "$@" -else - exec python3 -m skill_eval "$@" -fi diff --git a/bin/suggest-synthesis b/bin/suggest-synthesis deleted file mode 100755 index f5eea2e..0000000 --- a/bin/suggest-synthesis +++ /dev/null @@ -1,191 +0,0 @@ -#!/usr/bin/env python3 -"""Lightweight recurring synthesis trigger. - -Combines knowledge-report (raw/note ratios + candidates), audit-sweep (needs-audit -surface), and a targeted MindGraph query for "synthesis opportunities / gaps". - -Writes a dated draft artifact into 20_live/ (following last-handoff-draft and -knowledge-inquiries conventions). Intended to be called from session-close or -a launchd equivalent. - -Dry-run first by design. Respects MainFrame boundaries (20_live/ for output, -10_knowledge/ scope for MindGraph by default). - -Usage: - bin/suggest-synthesis --dry-run - bin/suggest-synthesis --apply --domain knowledge-systems - bin/suggest-synthesis --json --subset agents -""" - -from __future__ import annotations - -import argparse -import json -import subprocess -import sys -from datetime import date, datetime -from pathlib import Path -from typing import Any - -ROOT = Path(__file__).resolve().parents[1] -LIVE = ROOT / "20_live" -KNOWLEDGE = ROOT / "10_knowledge" -MINDGRAPH_BIN = ROOT / "bin" / "mindgraph" - -SYNTHESIS_DIR = LIVE # or LIVE / "synthesis-drafts" if we want a subdir later -TODAY = date.today().isoformat() - - -def run_cmd(cmd: list[str], timeout: int = 60) -> dict[str, Any]: - """Run a command, return parsed json if possible or raw output.""" - try: - proc = subprocess.run( - cmd, capture_output=True, text=True, timeout=timeout, cwd=ROOT - ) - out = proc.stdout.strip() - if out and (out.startswith("{") or out.startswith("[")): - try: - return {"ok": True, "data": json.loads(out), "raw": out} - except Exception: - pass - return {"ok": proc.returncode == 0, "raw": out, "stderr": proc.stderr.strip()} - except Exception as e: - return {"ok": False, "error": str(e)} - - -def get_stats_and_candidates(domain: str | None) -> dict[str, Any]: - """Leverage existing tools for ratios + audit surface.""" - results: dict[str, Any] = {} - # knowledge-report style stats (reuse its domain_stats logic via call or duplicate lightly) - try: - kr_cmd = [str(ROOT / "bin" / "knowledge-report"), "--json"] - if domain: - kr_cmd.extend(["--domain", domain]) - kr = run_cmd(kr_cmd) - results["knowledge_report"] = kr - except Exception as e: - results["knowledge_report"] = {"error": str(e)} - - # audit-sweep for candidates (the same mechanism the manifest used) - try: - as_cmd = [str(ROOT / "bin" / "audit-sweep"), "--dry-run", "--json"] - if domain: - as_cmd.extend(["--subset", domain]) - asw = run_cmd(as_cmd) - results["audit_sweep"] = asw - except Exception as e: - results["audit_sweep"] = {"error": str(e)} - - return results - - -def mindgraph_synthesis_query(domain: str | None) -> dict[str, Any]: - """Targeted MindGraph query for synthesis opportunities / gaps in the corpus.""" - q = "synthesis opportunities gaps hybrid memory architecture dashboard" - if domain: - q = f"{domain} {q}" - cmd = [str(MINDGRAPH_BIN), "query", q, "--json", "--top-k", "5", "--expand"] - # The wrapper will resolve the real binary - return run_cmd(cmd, timeout=90) - - -def write_draft(domain: str | None, data: dict[str, Any]) -> Path | None: - """Write a dated synthesis draft with proper MainFrame frontmatter.""" - slug = (domain or "knowledge-systems").replace("/", "-") - fname = f"{TODAY}__suggest-synthesis__{slug}.md" - target = SYNTHESIS_DIR / fname - - frontmatter = f"""--- -title: "Synthesis Opportunities — {domain or 'MainFrame'} ({TODAY})" -domain: "{domain or 'knowledge-systems'}" -type: "live" -status: "active" -source: "bin/suggest-synthesis (knowledge-report + audit-sweep + mindgraph query)" -tags: ["synthesis", "audit-sweep", "mindgraph", "gaps", "self-application"] -links: ["hybrid-memory-in-practice", "mainframe-dashboard-mindgraph-interface-ideas", "static-living-memory-architecture"] -as_of: "{TODAY}" ---- - -# Synthesis Opportunities — {domain or 'MainFrame'} ({TODAY}) - -Generated by `bin/suggest-synthesis --apply`. - -This is a volatile scratchpad capturing raw/note ratios, explicit needs-audit items, and MindGraph-nominated synthesis/gap signals. Promote durable insights to 10_knowledge/ only after review (see promotion rules in local 10_knowledge/index.md; public shape in 10_knowledge/index.template.md). - -""" - - body = "## Tool Signals\n\n" - body += "```json\n" + json.dumps(data, indent=2, default=str)[:8000] + "\n```\n\n" - - body += "## Recommended Next Actions\n" - body += "- Review the audit-sweep manifest (if any) under 20_live/epistemic-audit/pending-review/\n" - body += "- Run targeted `bin/audit-sweep --apply` on high-signal domains\n" - body += "- Feed promising clusters to `bin/extract-knowledge` or the create-source-summary skill\n" - body += "- Update the hybrid-memory-in-practice note or add cross-domain links after verification\n" - body += "- Consider a mindgraph-eval baseline re-run with any newly promoted notes\n" - - body += "\n## Open Signals from MindGraph (top synthesis/gap candidates)\n" - # The data already contains the query results; the caller can enrich if wanted. - - try: - target.parent.mkdir(parents=True, exist_ok=True) - target.write_text(frontmatter + body) - return target - except Exception as e: - print(f"Failed to write draft: {e}", file=sys.stderr) - return None - - -def main() -> int: - parser = argparse.ArgumentParser(description="Suggest synthesis opportunities using existing tools + MindGraph.") - parser.add_argument("--json", action="store_true", help="Emit structured JSON only") - parser.add_argument("--domain", "--subset", dest="domain", help="Limit to one 10_knowledge domain") - parser.add_argument("--dry-run", action="store_true", default=True, help="Default: show what would be done") - parser.add_argument("--apply", action="store_true", help="Write the dated draft to 20_live/") - args = parser.parse_args() - - if args.apply: - args.dry_run = False - - print(f"suggest-synthesis — domain={args.domain or 'all'} dry_run={args.dry_run}") - - signals = get_stats_and_candidates(args.domain) - mg = mindgraph_synthesis_query(args.domain) - signals["mindgraph"] = mg - - if args.json: - print(json.dumps({"domain": args.domain, "signals": signals}, indent=2, default=str)) - return 0 - - print("\n=== Knowledge Report / Audit Signals (truncated) ===") - print(json.dumps(signals.get("knowledge_report", {}), indent=2, default=str)[:1500]) - print("\n=== Audit Sweep Candidates (top) ===") - audit = signals.get("audit_sweep", {}) - if isinstance(audit.get("data"), dict): - cands = audit["data"].get("candidates", [])[:5] - for c in cands: - print(f" - {c.get('path')} tags={c.get('tags')}") - else: - print(audit.get("raw", str(audit))[:800]) - - print("\n=== MindGraph Synthesis Query (top fused/expanded) ===") - mg_raw = mg.get("raw") or str(mg) - print(mg_raw[:1200]) - - if args.apply: - draft_path = write_draft(args.domain, signals) - if draft_path: - print(f"\nDraft written: {draft_path}") - # Optional: trigger a mindgraph refresh so the new draft note is findable (even though 20_live is normally out of scope) - # We skip auto-refresh here to stay conservative; operator can run bin/mindgraph-refresh if desired. - print("Reminder: run bin/mindgraph-refresh if you want this draft indexed in a test DB.") - return 0 - else: - return 1 - - print("\n(dry-run complete — use --apply to write the draft)") - return 0 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/sync-project-index b/bin/sync-project-index deleted file mode 100755 index ab40547..0000000 --- a/bin/sync-project-index +++ /dev/null @@ -1,380 +0,0 @@ -#!/usr/bin/env python3 -"""Generate or check the local/private 30_projects/index.md. - -`--check` also validates project-state truth (ADR-041 / ADR-046): `active` -requires activity evidence (file mtimes or nested-repo commits) within the -window, active counts respect dual-pool WIP caps, and every project needs -valid frontmatter. Evidence rules match bin/session-close's checkpoint -derivation. -""" - -from __future__ import annotations - -import argparse -import difflib -import os -import re -import subprocess -import sys -from datetime import date, datetime -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -PROJECTS_DIR = ROOT / "30_projects" -INDEX_PATH = PROJECTS_DIR / "index.md" - -VALID_STATES = { - "active", "paused", "planned", "blocked", "suspended", "shipped", "trashed", -} -VALID_WIP_CLASSES = {"product", "eval", "anchor"} - -# ADR-046 dual-pool WIP (+ anchor): -# - product seats stay tight (operator focus) -# - eval/instrument projects may stay active without consuming product seats -# - anchor (income-engine) is always-on strategic hub — not product, not total -# - total active hard ceiling still bounds sprawl for product+eval -PRODUCT_WIP_CAP = 5 -TOTAL_ACTIVE_CAP = 10 -# Backward-compatible alias used by older callers/tests. -WIP_CAP = TOTAL_ACTIVE_CAP - -# Slugs that are measurement/eval infrastructure even without a *-eval suffix. -DEFAULT_EVAL_SLUGS = frozenset({ - "claim-audit-lab", - "reliability-eval-framework", - "scaffold-claims-study", - "skill-eval-workshop", - "verified-done", -}) - -# Strategic hubs that should stay active; they do not consume product or total seats. -DEFAULT_ANCHOR_SLUGS = frozenset({ - "income-engine", -}) - -ACTIVITY_WINDOW_DAYS = 14 -# Bounded scan, parity with bin/session-close latest_mtime. -SCAN_ENTRY_CAP = 600 -SKIP_DIR_NAMES = { - ".git", "node_modules", ".venv", "venv", "__pycache__", - ".pytest_cache", "dist", "build", ".cache", -} - - -def parse_frontmatter(path: Path) -> dict[str, str]: - text = path.read_text(encoding="utf-8") - if not text.startswith("---\n"): - return {} - end = text.find("\n---", 4) - if end == -1: - return {} - metadata: dict[str, str] = {} - for line in text[4:end].splitlines(): - if ":" not in line: - continue - key, value = line.split(":", 1) - value = value.strip().strip('"').strip("'") - metadata[key.strip()] = value - return metadata - - -def first_heading(path: Path) -> str: - for line in path.read_text(encoding="utf-8").splitlines(): - match = re.match(r"^#\s+(.+)$", line) - if match: - return match.group(1).strip() - return path.parent.name - - -def latest_mtime(project_dir: Path) -> float | None: - newest: float | None = None - scanned = 0 - stack = [project_dir] - while stack and scanned < SCAN_ENTRY_CAP: - current = stack.pop() - try: - with os.scandir(current) as entries: - for entry in entries: - scanned += 1 - if scanned >= SCAN_ENTRY_CAP: - break - if entry.name in SKIP_DIR_NAMES or entry.name.startswith("."): - continue - try: - if entry.is_dir(follow_symlinks=False): - stack.append(Path(entry.path)) - elif entry.is_file(follow_symlinks=False): - mtime = entry.stat(follow_symlinks=False).st_mtime - if newest is None or mtime > newest: - newest = mtime - except OSError: - continue - except OSError: - continue - return newest - - -def nested_repo_last_commit(project_dir: Path) -> float | None: - if not (project_dir / ".git").exists(): - return None - try: - proc = subprocess.run( - ["git", "-C", str(project_dir), "log", "-1", "--format=%ct"], - capture_output=True, text=True, timeout=10, - ) - except (OSError, subprocess.TimeoutExpired): - return None - if proc.returncode != 0 or not proc.stdout.strip(): - return None - try: - return float(proc.stdout.strip().splitlines()[-1]) - except ValueError: - return None - - -def activity_evidence(project_dir: Path) -> float | None: - """Newest evidence timestamp: file mtimes or nested-repo last commit.""" - signals = [ts for ts in (latest_mtime(project_dir), - nested_repo_last_commit(project_dir)) - if ts is not None] - return max(signals) if signals else None - - -def project_rows(projects_dir: Path = PROJECTS_DIR) -> list[dict[str, str]]: - rows: list[dict[str, str]] = [] - for readme in sorted(projects_dir.glob("*/README.md")): - metadata = parse_frontmatter(readme) - evidence = activity_evidence(readme.parent) - rows.append( - { - "name": metadata.get("title") or first_heading(readme), - "path": readme.parent.name, - "state": metadata.get("project_state") or metadata.get("status") or "unknown", - "goal": metadata.get("goal") or "", - "next_action": metadata.get("next_action") or "", - "updated": metadata.get("updated") or metadata.get("last_updated") or "", - "evidence": (date.fromtimestamp(evidence).isoformat() - if evidence is not None else ""), - } - ) - return rows - - -def cell(value: str) -> str: - return value.replace("|", "\\|").replace("\n", " ").strip() or "-" - - -def render(projects_dir: Path = PROJECTS_DIR) -> str: - lines = [ - "# Projects Index", - "", - "Generated from project README metadata. This local file is ignored by Git; do not edit by hand.", - "Run `bin/sync-project-index --write` to regenerate it.", - "`Evidence` is the last observed activity (file mtimes / nested-repo commit), not self-report.", - "", - ] - rows = project_rows(projects_dir) - if not rows: - lines.extend(["_No project README files found._", ""]) - return "\n".join(lines) - - lines.extend( - [ - "| Project | State | Goal | Next action | Updated | Evidence |", - "| --- | --- | --- | --- | --- | --- |", - ] - ) - for row in rows: - project = f"[{cell(row['name'])}]({cell(row['path'])}/README.md)" - lines.append( - f"| {project} | {cell(row['state'])} | {cell(row['goal'])} | " - f"{cell(row['next_action'])} | {cell(row['updated'])} | " - f"{cell(row['evidence'])} |" - ) - lines.append("") - return "\n".join(lines) - - -def resolve_wip_class(slug: str, metadata: dict[str, str]) -> tuple[str, list[str]]: - """Return (wip_class, problems). Explicit frontmatter wins; else heuristics.""" - problems: list[str] = [] - raw = (metadata.get("wip_class") or "").strip().lower() - if raw: - if raw not in VALID_WIP_CLASSES: - problems.append( - f"{slug}: unknown wip_class '{raw}' " - f"(valid: {', '.join(sorted(VALID_WIP_CLASSES))})") - return "product", problems - return raw, problems - if slug in DEFAULT_ANCHOR_SLUGS: - return "anchor", problems - if slug.endswith("-eval") or slug in DEFAULT_EVAL_SLUGS: - return "eval", problems - return "product", problems - - -def validate_projects( - projects_dir: Path = PROJECTS_DIR, - now_ts: float | None = None, - window_days: int = ACTIVITY_WINDOW_DAYS, - wip_cap: int | None = None, - product_wip_cap: int = PRODUCT_WIP_CAP, - total_active_cap: int | None = None, - only: str | None = None, -) -> list[str]: - """Project-state truth rules (ADR-041 / ADR-046). Returns human-readable problems. - - With `only`, validate that single project directory and skip the - cross-project WIP-cap rules; a missing directory is itself a problem. - - Dual-pool WIP (ADR-046): - - product actives ≤ product_wip_cap (default 5) - - total actives (product+eval) ≤ total_active_cap (default 10; legacy alias: wip_cap) - - eval actives do not consume product seats - - anchor actives (income-engine) consume neither product nor total seats - """ - if now_ts is None: - now_ts = datetime.now().timestamp() - if total_active_cap is None: - total_active_cap = TOTAL_ACTIVE_CAP if wip_cap is None else wip_cap - - problems: list[str] = [] - active_product: list[str] = [] - active_eval: list[str] = [] - active_anchor: list[str] = [] - if only is not None: - target = projects_dir / only - if not target.is_dir(): - return [f"{only}: project directory does not exist"] - entries = [target] - else: - entries = sorted(projects_dir.iterdir()) - for entry in entries: - if not entry.is_dir() or entry.name.startswith("."): - continue - readme = entry / "README.md" - if not readme.exists(): - problems.append(f"{entry.name}: missing README.md") - continue - metadata = parse_frontmatter(readme) - state = (metadata.get("project_state") or metadata.get("status") or "").strip() - if not state: - problems.append( - f"{entry.name}: missing project_state frontmatter " - "(README must start with a `---` block)") - continue - if state not in VALID_STATES: - problems.append( - f"{entry.name}: unknown project_state '{state}' " - f"(valid: {', '.join(sorted(VALID_STATES))})") - continue - wip_class, class_problems = resolve_wip_class(entry.name, metadata) - problems.extend(class_problems) - next_action = (metadata.get("next_action") or "").strip() - if state == "active": - if wip_class == "eval": - active_eval.append(entry.name) - elif wip_class == "anchor": - active_anchor.append(entry.name) - else: - active_product.append(entry.name) - evidence = activity_evidence(entry) - if evidence is None: - problems.append(f"{entry.name}: active but no activity evidence found") - else: - age_days = (now_ts - evidence) / 86400 - if age_days > window_days: - problems.append( - f"{entry.name}: active but idle {age_days:.1f}d " - f"(> {window_days}d window) — pause it with a " - "next_action reentry pointer or resume work") - if not next_action: - problems.append(f"{entry.name}: active without next_action") - elif state == "paused" and not next_action: - problems.append( - f"{entry.name}: paused without next_action reentry pointer") - - if only is not None: - return problems - - # Anchors are permanent strategic hubs; product+eval share the capped pools. - active_capped = active_product + active_eval - if len(active_product) > product_wip_cap: - problems.append( - f"product WIP cap breach: {len(active_product)} active product " - f"projects (cap {product_wip_cap}): {', '.join(active_product)} " - f"[eval active={len(active_eval)}; anchor active={len(active_anchor)} " - f"not counted toward product seats]") - if len(active_capped) > total_active_cap: - problems.append( - f"total active cap breach: {len(active_capped)} product+eval actives " - f"(cap {total_active_cap}): product=[{', '.join(active_product)}] " - f"eval=[{', '.join(active_eval)}] " - f"[anchor active={len(active_anchor)} not counted toward total]") - return problems - - -def main() -> int: - parser = argparse.ArgumentParser() - group = parser.add_mutually_exclusive_group(required=True) - group.add_argument("--check", action="store_true", help="fail if index is stale") - group.add_argument("--write", action="store_true", help="write generated index") - parser.add_argument( - "--project", - metavar="SLUG", - help="with --check: validate one project's state only " - "(skips index staleness and the WIP cap)", - ) - args = parser.parse_args() - - if args.project: - if args.write: - parser.error("--project only applies to --check") - problems = validate_projects(only=args.project) - if problems: - print("project-state problems (ADR-041):", file=sys.stderr) - for problem in problems: - print(f" - {problem}", file=sys.stderr) - return 1 - print(f"{args.project}: project state verified") - return 0 - - generated = render() - problems = validate_projects() - if args.write: - INDEX_PATH.write_text(generated, encoding="utf-8") - print(f"wrote {INDEX_PATH.relative_to(ROOT)}") - if problems: - print("project-state problems (--check fails until fixed):", - file=sys.stderr) - for problem in problems: - print(f" - {problem}", file=sys.stderr) - return 0 - - failed = False - current = INDEX_PATH.read_text(encoding="utf-8") if INDEX_PATH.exists() else "" - if current != generated: - failed = True - print("local 30_projects/index.md is stale; run `bin/sync-project-index --write`.", file=sys.stderr) - for line in difflib.unified_diff( - current.splitlines(), - generated.splitlines(), - fromfile="current", - tofile="generated", - lineterm="", - ): - print(line, file=sys.stderr) - if problems: - failed = True - print("project-state problems (ADR-041):", file=sys.stderr) - for problem in problems: - print(f" - {problem}", file=sys.stderr) - if failed: - return 1 - print("local 30_projects/index.md is current; project states verified") - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/bin/task-packet b/bin/task-packet deleted file mode 100755 index 1783870..0000000 --- a/bin/task-packet +++ /dev/null @@ -1,811 +0,0 @@ -#!/usr/bin/env python3 -"""Validate and compile MainFrame project delegation packets.""" - -from __future__ import annotations - -import argparse -import hashlib -import json -import os -import re -import shlex -import sys -import tempfile -from pathlib import Path -from typing import Any - - -ROOT = Path(__file__).resolve().parents[1] -PROJECTS_DIR = ROOT / "30_projects" -DEFAULT_OUTPUT = PROJECTS_DIR / "task_packets_manifest.json" - -REQUIRED_KEYS = ( - "packet_version", - "task_id", - "project_slug", - "title", - "status", - "task_kind", - "agent_profile", - "workdir", - "timeout_seconds", - "editable_files", - "create_files", - "read_only_files", - "verification_commands", - "mindgraph_mode", - "knowledge_queries", -) -OPTIONAL_ROUTING_KEYS = ( - "task_category", - "executor", - "harness_recommendation", - "allow_fusion_plan", - "needs_deliberation", -) -ALLOWED_KEYS = REQUIRED_KEYS + OPTIONAL_ROUTING_KEYS -LIST_KEYS = { - "editable_files", - "create_files", - "read_only_files", - "verification_commands", - "knowledge_queries", -} -REQUIRED_SECTIONS = ( - "Goal", - "Context and plan references", - "Required implementation", - "Non-negotiable boundaries", - "Acceptance criteria", - "Stop conditions", - "Expected handoff", -) -# Allowed but not required — planning observability (MH01 / ADR-030). -OPTIONAL_SECTIONS = ( - "MindGraph Query Pass", -) -STATUSES = {"draft", "ready", "retired"} -TASK_KINDS = {"code", "research", "analysis", "operations"} -MINDGRAPH_MODES = {"off", "curated"} -TASK_CATEGORIES = { - "mechanical-edit", - "multi-file-coordination", - "failure-recovery", - "scope-enforcement", - "research-synthesis", - "environment-setup", - "long-horizon-build", -} -EXECUTORS = {"local", "cloud", "auto"} -HARNESS_RECOMMENDATIONS = { - "H0-current", - "H1-packet", - "H2-repair", - "H3-mindgraph", -} -BOOL_KEYS = {"allow_fusion_plan", "needs_deliberation"} -ROUTING_POLICY_PATH = ROOT / ".context" / "harness-routing.json" -SHELL_OPERATORS = ("&&", "||", ";", "|", "`", "$(", ">", "<", "\n", "\r") -SAFE_ID = re.compile(r"^[a-z0-9][a-z0-9._-]*$") - - -class PacketError(ValueError): - """Raised when a packet violates the project convention.""" - - -def _parse_scalar(value: str, key: str) -> Any: - value = value.strip() - if key in LIST_KEYS: - try: - parsed = json.loads(value) - except json.JSONDecodeError as exc: - raise PacketError(f"{key} must be a JSON array: {exc.msg}") from exc - if not isinstance(parsed, list) or not all( - isinstance(item, str) for item in parsed - ): - raise PacketError(f"{key} must be a JSON array of strings") - return parsed - if key == "timeout_seconds": - try: - return int(value) - except ValueError as exc: - raise PacketError("timeout_seconds must be an integer") from exc - if key in BOOL_KEYS: - normalized = value.strip().lower() - if normalized in {"true", "yes"}: - return True - if normalized in {"false", "no"}: - return False - raise PacketError(f"{key} must be true or false") - if ( - len(value) >= 2 - and value[0] == value[-1] - and value[0] in {'"', "'"} - ): - return value[1:-1] - return value - - -def parse_packet_text(text: str, source: str = "<packet>") -> dict[str, Any]: - if not text.startswith("---\n"): - raise PacketError(f"{source}: packet must start with YAML frontmatter") - end = text.find("\n---\n", 4) - if end == -1: - raise PacketError(f"{source}: frontmatter closing delimiter is missing") - - metadata: dict[str, Any] = {} - for number, raw_line in enumerate(text[4:end].splitlines(), start=2): - line = raw_line.strip() - if not line or line.startswith("#"): - continue - if ":" not in line: - raise PacketError(f"{source}:{number}: expected key: value") - key, value = line.split(":", 1) - key = key.strip() - if key in metadata: - raise PacketError(f"{source}:{number}: duplicate key {key}") - metadata[key] = _parse_scalar(value, key) - - body = text[end + 5 :] - sections: dict[str, str] = {} - current: str | None = None - lines: list[str] = [] - for raw_line in body.splitlines(): - if raw_line.startswith("## "): - if current is not None: - if current in sections: - raise PacketError(f"{source}: duplicate section: {current}") - sections[current] = "\n".join(lines).strip() - current = raw_line[3:].strip() - lines = [] - elif current is not None: - lines.append(raw_line) - if current is not None: - if current in sections: - raise PacketError(f"{source}: duplicate section: {current}") - sections[current] = "\n".join(lines).strip() - - return { - "metadata": metadata, - "sections": sections, - "source": source, - } - - -def load_packet(path: Path) -> dict[str, Any]: - try: - raw = path.read_bytes() - except OSError as exc: - raise PacketError(f"{path}: cannot read packet: {exc}") from exc - try: - text = raw.decode("utf-8") - except UnicodeDecodeError as exc: - raise PacketError(f"{path}: packet is not valid UTF-8: {exc}") from exc - packet = parse_packet_text(text, str(path)) - packet["path"] = path - packet["source_sha256"] = hashlib.sha256(raw).hexdigest() - return packet - - -def _validate_relative_path(value: str, field: str) -> None: - path = Path(value) - if not value or path.is_absolute(): - raise PacketError(f"{field} entries must be non-empty relative paths") - if "\\" in value or any( - part in {"", ".", ".."} for part in value.split("/") - ): - raise PacketError(f"{field} contains unsafe path: {value}") - - -def _first_path_overlap( - left: list[str], - right: list[str], - *, - same_collection: bool = False, -) -> tuple[str, str] | None: - for left_index, left_value in enumerate(left): - left_parts = tuple(left_value.split("/")) - start = left_index + 1 if same_collection else 0 - for right_value in right[start:]: - right_parts = tuple(right_value.split("/")) - common = min(len(left_parts), len(right_parts)) - if left_parts[:common] == right_parts[:common]: - return left_value, right_value - return None - - -def command_argv(command: str) -> list[str]: - if not command.strip(): - raise PacketError("verification_commands may not contain empty commands") - if any(operator in command for operator in SHELL_OPERATORS): - raise PacketError( - f"verification command contains a forbidden shell operator: {command}" - ) - try: - argv = shlex.split(command) - except ValueError as exc: - raise PacketError(f"invalid verification command: {exc}") from exc - if not argv: - raise PacketError("verification command produced no arguments") - return argv - - -def validate_packet( - packet: dict[str, Any], - *, - projects_dir: Path = PROJECTS_DIR, - workspace_root: Path | None = None, - require_ready: bool = False, -) -> list[str]: - errors: list[str] = [] - metadata = packet["metadata"] - sections = packet["sections"] - - missing_keys = [] - for key in REQUIRED_KEYS: - if key not in metadata: - errors.append(f"missing frontmatter key: {key}") - missing_keys.append(key) - unknown_keys = sorted(set(metadata) - set(ALLOWED_KEYS)) - for key in unknown_keys: - errors.append(f"unknown frontmatter key: {key}") - for section in REQUIRED_SECTIONS: - if not sections.get(section, "").strip(): - errors.append(f"missing or empty section: {section}") - allowed_sections = set(REQUIRED_SECTIONS) | set(OPTIONAL_SECTIONS) - for section in sorted(set(sections) - allowed_sections): - errors.append(f"unknown packet section: {section}") - # Empty optional sections are ignored; if present with mindgraph curated - # mode, encourage a filled pass (warning as error only when empty placeholder). - qp = sections.get("MindGraph Query Pass", "").strip() - if ( - metadata.get("mindgraph_mode") == "curated" - and qp - and "knowledge_query:" in qp - and "REPLACE" in qp.upper() - ): - errors.append( - "MindGraph Query Pass still contains REPLACE placeholders " - "while mindgraph_mode is curated" - ) - if missing_keys: - return errors - - if str(metadata["packet_version"]) != "1": - errors.append("packet_version must be 1") - if metadata["status"] not in STATUSES: - errors.append(f"status must be one of: {', '.join(sorted(STATUSES))}") - if require_ready and metadata["status"] != "ready": - errors.append("only packets with status ready may execute") - if metadata["task_kind"] not in TASK_KINDS: - errors.append( - f"task_kind must be one of: {', '.join(sorted(TASK_KINDS))}" - ) - if metadata["mindgraph_mode"] not in MINDGRAPH_MODES: - errors.append( - "mindgraph_mode must be one of: " - + ", ".join(sorted(MINDGRAPH_MODES)) - ) - if not SAFE_ID.fullmatch(str(metadata["task_id"])): - errors.append("task_id must use lowercase letters, digits, dots, dashes, or underscores") - if not SAFE_ID.fullmatch(str(metadata["project_slug"])): - errors.append("project_slug must use lowercase letters, digits, dots, dashes, or underscores") - for key in ("title", "agent_profile", "workdir"): - if not isinstance(metadata[key], str) or not metadata[key].strip(): - errors.append(f"{key} must be a non-empty string") - if not isinstance(metadata["timeout_seconds"], int) or not ( - 1 <= metadata["timeout_seconds"] <= 7200 - ): - errors.append("timeout_seconds must be between 1 and 7200") - - path_fields = ("editable_files", "create_files", "read_only_files") - valid_paths: dict[str, list[str]] = {} - for field in path_fields: - if len(metadata[field]) != len(set(metadata[field])): - errors.append(f"{field} may not contain duplicate paths") - valid_paths[field] = [] - for value in metadata[field]: - try: - _validate_relative_path(value, field) - except PacketError as exc: - errors.append(str(exc)) - else: - valid_paths[field].append(value) - - for field in path_fields: - overlap = _first_path_overlap( - valid_paths[field], valid_paths[field], same_collection=True - ) - if overlap and overlap[0] != overlap[1]: - errors.append( - f"{field} contains overlapping paths: {overlap[0]} and {overlap[1]}" - ) - for left_field, right_field in ( - ("editable_files", "create_files"), - ("editable_files", "read_only_files"), - ("create_files", "read_only_files"), - ): - overlap = _first_path_overlap( - valid_paths[left_field], valid_paths[right_field] - ) - if overlap: - errors.append( - f"{left_field} and {right_field} may not overlap: " - f"{overlap[0]} and {overlap[1]}" - ) - if not metadata["editable_files"] and not metadata["create_files"]: - errors.append("at least one editable_files or create_files entry is required") - if not metadata["verification_commands"]: - errors.append("at least one verification command is required") - for command in metadata["verification_commands"]: - try: - command_argv(command) - except PacketError as exc: - errors.append(str(exc)) - if len(metadata["verification_commands"]) != len( - set(metadata["verification_commands"]) - ): - errors.append("verification_commands may not contain duplicates") - - if metadata["mindgraph_mode"] == "curated" and not metadata["knowledge_queries"]: - errors.append("curated MindGraph mode requires at least one knowledge query") - if metadata["mindgraph_mode"] == "off" and metadata["knowledge_queries"]: - errors.append("knowledge_queries must be empty when mindgraph_mode is off") - - task_category = metadata.get("task_category") - if task_category is not None: - if not isinstance(task_category, str) or task_category not in TASK_CATEGORIES: - errors.append( - "task_category must be one of: " - + ", ".join(sorted(TASK_CATEGORIES)) - ) - executor = metadata.get("executor") - if executor is not None: - if not isinstance(executor, str) or executor not in EXECUTORS: - errors.append( - f"executor must be one of: {', '.join(sorted(EXECUTORS))}" - ) - harness_recommendation = metadata.get("harness_recommendation") - if harness_recommendation is not None: - if ( - not isinstance(harness_recommendation, str) - or harness_recommendation not in HARNESS_RECOMMENDATIONS - ): - errors.append( - "harness_recommendation must be one of: " - + ", ".join(sorted(HARNESS_RECOMMENDATIONS)) - ) - for key in BOOL_KEYS: - value = metadata.get(key) - if value is not None and not isinstance(value, bool): - errors.append(f"{key} must be true or false") - - if workspace_root is None: - project_root = projects_dir / metadata["project_slug"] - workdir = project_root / metadata["workdir"] - else: - project_root = workspace_root - workdir = workspace_root / metadata["workdir"] - workdir_value = str(metadata["workdir"]) - try: - if workdir_value != ".": - _validate_relative_path(workdir_value, "workdir") - except PacketError as exc: - errors.append(str(exc)) - else: - if not project_root.is_dir(): - errors.append(f"project root does not exist: {project_root}") - elif not workdir.is_dir(): - errors.append(f"workdir does not exist: {workdir}") - - if metadata["project_slug"] and packet.get("path") and workspace_root is None: - expected = projects_dir / metadata["project_slug"] / "plans" / "task-packets" - try: - packet["path"].resolve().relative_to(expected.resolve()) - except ValueError: - errors.append(f"packet must live under {expected}") - - return errors - - -def packet_summary(packet: dict[str, Any], root: Path = ROOT) -> dict[str, Any]: - metadata = packet["metadata"] - path = packet["path"] - try: - relative_path = str(path.resolve().relative_to(root.resolve())) - except ValueError: - relative_path = str(path) - contract_metadata = { - key: value - for key, value in metadata.items() - if key != "status" - } - contract_payload = { - "metadata": contract_metadata, - "sections": packet["sections"], - } - contract_sha256 = hashlib.sha256( - json.dumps( - contract_payload, - sort_keys=True, - separators=(",", ":"), - ).encode("utf-8") - ).hexdigest() - summary = { - "id": f"packet-{metadata['project_slug']}-{metadata['task_id']}", - "task_id": metadata["task_id"], - "project_slug": metadata["project_slug"], - "title": metadata["title"], - "status": metadata["status"], - "task_kind": metadata["task_kind"], - "agent_profile": metadata["agent_profile"], - "workdir": metadata["workdir"], - "timeout_seconds": metadata["timeout_seconds"], - "mindgraph_mode": metadata["mindgraph_mode"], - "goal": packet["sections"]["Goal"], - "next_action": ( - "Delegate the reviewed task packet." - if metadata["status"] == "ready" - else "Review and mark the task packet ready." - ), - "packet_path": relative_path, - "editable_files": metadata["editable_files"], - "create_files": metadata["create_files"], - "read_only_files": metadata["read_only_files"], - "verification_commands": metadata["verification_commands"], - "contract_sha256": contract_sha256, - "source_sha256": packet["source_sha256"], - } - for key in OPTIONAL_ROUTING_KEYS: - if key in metadata: - summary[key] = metadata[key] - return summary - - -def _coerce_policy_int(value: Any, field: str) -> int: - try: - return int(value) - except (TypeError, ValueError) as exc: - raise PacketError(f"routing policy field {field} must be an integer") from exc - - -def load_routing_policy(path: Path = ROUTING_POLICY_PATH) -> dict[str, Any]: - try: - payload = json.loads(path.read_text(encoding="utf-8")) - except OSError as exc: - raise PacketError(f"cannot read routing policy: {path}: {exc}") from exc - except json.JSONDecodeError as exc: - raise PacketError(f"invalid routing policy JSON: {path}: {exc}") from exc - if not isinstance(payload, dict): - raise PacketError("routing policy root must be an object") - return payload - - -def infer_routing_hints( - metadata: dict[str, Any], - *, - policy_path: Path = ROUTING_POLICY_PATH, -) -> dict[str, Any]: - """Suggest routing metadata from packet shape and harness-routing policy.""" - if not policy_path.is_file(): - return {} - policy = load_routing_policy(policy_path) - hints: dict[str, Any] = {} - if metadata.get("task_category"): - hints["task_category"] = metadata["task_category"] - else: - editable_count = len(metadata.get("editable_files", [])) - create_count = len(metadata.get("create_files", [])) - file_count = editable_count + create_count - task_kind = metadata.get("task_kind") - for rule in policy.get("inference_rules", []): - when = rule.get("when", {}) - if task_kind and when.get("task_kind") and when["task_kind"] != task_kind: - continue - min_files = when.get("min_editable_files") - if min_files is not None and editable_count < _coerce_policy_int( - min_files, "when.min_editable_files" - ): - continue - max_files = when.get("max_editable_files") - if max_files is not None and editable_count > _coerce_policy_int( - max_files, "when.max_editable_files" - ): - continue - min_total = when.get("min_total_files") - if min_total is not None and file_count < _coerce_policy_int( - min_total, "when.min_total_files" - ): - continue - if "task_category" in rule: - hints["task_category"] = rule["task_category"] - break - - category = hints.get("task_category") or metadata.get("task_category") - defaults = policy.get("category_defaults", {}).get(category or "", {}) - if isinstance(defaults, dict): - for key in OPTIONAL_ROUTING_KEYS: - if key not in metadata and key in defaults: - hints[key] = defaults[key] - if "agent_profile" in defaults and "agent_profile" not in metadata: - hints["agent_profile"] = defaults["agent_profile"] - return hints - - -def discover_packets(projects_dir: Path = PROJECTS_DIR) -> list[Path]: - return sorted(projects_dir.glob("*/plans/task-packets/*.md")) - - -def _manifest_packets(payload: Any, output: Path) -> dict[str, dict[str, Any]]: - if not isinstance(payload, dict): - raise PacketError(f"existing manifest root must be an object: {output}") - if payload.get("packet_version") != 1: - raise PacketError(f"existing manifest packet_version must be 1: {output}") - packets = payload.get("packets") - if not isinstance(packets, list): - raise PacketError(f"existing manifest packets must be an array: {output}") - - prior_packets: dict[str, dict[str, Any]] = {} - for index, item in enumerate(packets): - if not isinstance(item, dict): - raise PacketError( - f"existing manifest packet {index} must be an object: {output}" - ) - packet_id = item.get("id") - status = item.get("status") - contract_sha256 = item.get("contract_sha256") - if not isinstance(packet_id, str) or not packet_id: - raise PacketError( - f"existing manifest packet {index} has no valid id: {output}" - ) - if packet_id in prior_packets: - raise PacketError( - f"existing manifest contains duplicate packet id {packet_id}: {output}" - ) - if status not in STATUSES: - raise PacketError( - f"existing manifest packet {packet_id} has invalid status: {output}" - ) - if not isinstance(contract_sha256, str) or not re.fullmatch( - r"[a-f0-9]{64}", contract_sha256 - ): - raise PacketError( - f"existing manifest packet {packet_id} has no valid contract hash: " - f"{output}" - ) - prior_packets[packet_id] = item - return prior_packets - - -def _load_prior_packets(output: Path) -> dict[str, dict[str, Any]]: - if output.is_symlink(): - raise PacketError(f"existing manifest may not be a symlink: {output}") - try: - payload = json.loads(output.read_text(encoding="utf-8")) - except OSError as exc: - raise PacketError(f"cannot read existing manifest: {output}: {exc}") from exc - except json.JSONDecodeError as exc: - raise PacketError(f"invalid existing manifest JSON: {output}: {exc}") from exc - return _manifest_packets(payload, output) - - -def _write_manifest_atomic(output: Path, payload: dict[str, Any]) -> None: - output.parent.mkdir(parents=True, exist_ok=True) - temporary: Path | None = None - try: - with tempfile.NamedTemporaryFile( - mode="w", - encoding="utf-8", - dir=output.parent, - prefix=f".{output.name}.", - suffix=".tmp", - delete=False, - ) as handle: - temporary = Path(handle.name) - json.dump(payload, handle, indent=2, sort_keys=True) - handle.write("\n") - handle.flush() - os.fsync(handle.fileno()) - temporary.replace(output) - finally: - if temporary is not None: - temporary.unlink(missing_ok=True) - - -def compile_packets( - *, - projects_dir: Path = PROJECTS_DIR, - output: Path = DEFAULT_OUTPUT, - include_drafts: bool = True, -) -> dict[str, Any]: - summaries: list[dict[str, Any]] = [] - problems: list[dict[str, Any]] = [] - seen_packet_ids: set[str] = set() - prior_packets: dict[str, dict[str, Any]] = {} - if output.exists() or output.is_symlink(): - try: - prior_packets = _load_prior_packets(output) - except PacketError as exc: - return { - "packet_version": 1, - "packets": [], - "invalid": [{"path": str(output), "errors": [str(exc)]}], - } - for path in discover_packets(projects_dir): - try: - packet = load_packet(path) - errors = validate_packet(packet, projects_dir=projects_dir) - except PacketError as exc: - errors = [str(exc)] - packet = None - if errors: - problems.append({"path": str(path), "errors": errors}) - continue - assert packet is not None - summary = packet_summary(packet, root=projects_dir.parent) - seen_packet_ids.add(summary["id"]) - prior = prior_packets.get(summary["id"]) - if prior and prior.get("status") in {"ready", "retired"}: - if ( - prior.get("contract_sha256") - and prior["contract_sha256"] != summary["contract_sha256"] - ): - problems.append( - { - "path": str(path), - "errors": [ - "compiled ready packet contract is immutable; " - "create a new task_id for revised work" - ], - } - ) - continue - if summary["status"] == "draft": - problems.append( - { - "path": str(path), - "errors": ["compiled ready packet may not return to draft"], - } - ) - continue - if prior.get("status") == "retired" and summary["status"] != "retired": - problems.append( - { - "path": str(path), - "errors": ["retired packet may not be reactivated"], - } - ) - continue - if include_drafts or packet["metadata"]["status"] == "ready": - summaries.append(summary) - - for packet_id, prior in sorted(prior_packets.items()): - if ( - prior.get("status") in {"ready", "retired"} - and packet_id not in seen_packet_ids - ): - problems.append( - { - "path": prior.get("packet_path", packet_id), - "errors": [ - "compiled ready packet may not be deleted; " - "retire it and preserve the contract" - ], - } - ) - - payload = { - "packet_version": 1, - "packets": summaries, - "invalid": problems, - } - if problems: - return payload - _write_manifest_atomic(output, payload) - return payload - - -def _print_errors(errors: list[str]) -> None: - for error in errors: - print(f"ERROR: {error}", file=sys.stderr) - - -def main(argv: list[str] | None = None) -> int: - parser = argparse.ArgumentParser(description=__doc__) - subparsers = parser.add_subparsers(dest="command", required=True) - - validate_parser = subparsers.add_parser("validate") - validate_parser.add_argument("path", type=Path) - validate_parser.add_argument("--require-ready", action="store_true") - - inspect_parser = subparsers.add_parser( - "inspect", - help="validate one packet and emit its canonical summary as JSON", - ) - inspect_parser.add_argument("path", type=Path) - inspect_parser.add_argument( - "--projects-dir", - type=Path, - default=PROJECTS_DIR, - ) - inspect_parser.add_argument("--require-ready", action="store_true") - - compile_parser = subparsers.add_parser("compile") - compile_parser.add_argument("--projects-dir", type=Path, default=PROJECTS_DIR) - compile_parser.add_argument("--output", type=Path, default=DEFAULT_OUTPUT) - compile_parser.add_argument("--ready-only", action="store_true") - - hints_parser = subparsers.add_parser("routing-hints") - hints_parser.add_argument( - "--policy", - type=Path, - default=ROUTING_POLICY_PATH, - ) - hints_parser.add_argument("--metadata-json", required=True) - - args = parser.parse_args(argv) - if args.command == "routing-hints": - try: - metadata = json.loads(args.metadata_json) - except json.JSONDecodeError as exc: - print(f"ERROR: invalid metadata JSON: {exc}", file=sys.stderr) - return 1 - if not isinstance(metadata, dict): - print("ERROR: metadata JSON must be an object", file=sys.stderr) - return 1 - try: - hints = infer_routing_hints(metadata, policy_path=args.policy) - except PacketError as exc: - print(f"ERROR: {exc}", file=sys.stderr) - return 1 - print(json.dumps(hints, indent=2, sort_keys=True)) - return 0 - - if args.command in {"validate", "inspect"}: - try: - packet = load_packet(args.path) - errors = validate_packet( - packet, - projects_dir=( - args.projects_dir - if args.command == "inspect" - else PROJECTS_DIR - ), - require_ready=args.require_ready, - ) - except PacketError as exc: - errors = [str(exc)] - if errors: - _print_errors(errors) - return 1 - if args.command == "inspect": - print( - json.dumps( - packet_summary(packet, root=args.projects_dir.parent), - sort_keys=True, - ) - ) - return 0 - print(f"Valid task packet: {args.path}") - return 0 - - payload = compile_packets( - projects_dir=args.projects_dir, - output=args.output, - include_drafts=not args.ready_only, - ) - print( - f"Compiled {len(payload['packets'])} packets " - f"with {len(payload['invalid'])} invalid packet(s) to {args.output}" - ) - if payload["invalid"]: - for problem in payload["invalid"]: - _print_errors( - [f"{problem['path']}: {message}" for message in problem["errors"]] - ) - return 1 - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/bin/title-existence-check b/bin/title-existence-check deleted file mode 100755 index 5ce384a..0000000 --- a/bin/title-existence-check +++ /dev/null @@ -1,232 +0,0 @@ -#!/usr/bin/env python3 -"""G6 — does a capture's claimed paper actually exist? - -Automates the manual check run on 2026-08-09, which found that **21 of 21** -claimed titles did not exist: every exact-title search returned real work on the -subject by named authors, and never the cited paper. See -`20_live/security/2026-08-09__g6-title-existence-check.md`. - -## This tool nominates. It does not convict. - -Crossref does not index arXiv preprints or much of the ACL Anthology. "Not found -in Crossref" is therefore **not** evidence of fabrication, and a tool that said -otherwise would falsely accuse most of the machine-learning literature. Measured -during development: - - Crossref's own relevance score is worthless as a discriminator. - A fabricated title scored 40.8; "Attention Is All You Need" scored 33.1. - -So this checks **title similarity** against two sources (Crossref and arXiv), and -reports a band rather than a verdict: - - likely-real a near-exact title match exists somewhere - inconclusive partial match only, or plausible coverage gap -> human check - no-match nothing resembling this title in either source -> human check - -Only `no-match` is worth a person's attention, and even then the output is a -question, not an accusation. - -**Controls run in the same batch.** Known-real titles must land in `likely-real` -and a known-fabricated title must land in `no-match`. If they do not, the run is -reported BROKEN and its results are not evidence. Same discipline as -`bin/citation-resolvability-sweep`. - -Usage: - bin/title-existence-check --file 10_knowledge/domain/foo.md - bin/title-existence-check --from-list 20_live/security/sweeps/<file>.txt - bin/title-existence-check --title "Some Paper Title" - bin/title-existence-check --from-list ... --json out.json -""" - -from __future__ import annotations - -import argparse -import difflib -import json -import os -import re -import sys -import urllib.parse -import urllib.request -import xml.etree.ElementTree as ET -from pathlib import Path -from typing import Any - -ROOT = Path(__file__).resolve().parents[1] -# Crossref's "polite pool" gives better rate limits when the User-Agent carries a -# contact address. Read it from the environment so a real address is never committed. -# Matches the UNPAYWALL_EMAIL convention in scripts/fetch_source_text.py. -_CONTACT = os.environ.get("MAINFRAME_CONTACT_EMAIL", "mainframe@local") -UA = f"MainFrame-audit/1.0 (integrity audit; mailto:{_CONTACT})" - -# Calibrated on the 2026-08-09 corpus: fabricated titles scored 0.44-0.50, -# real titles 0.65-1.00. The bands are deliberately wide, with everything -# between them routed to a human rather than guessed at. -LIKELY_REAL = 0.85 -NO_MATCH = 0.55 - -REAL_CONTROLS = [ - "Attention Is All You Need", - "Computing Inter-Rater Reliability and Its Variance in the Presence of High Agreement", -] -FAKE_CONTROL = "Chunk Granularity and Retrieval Floors in Multi-Stage RAG/NLI Pipelines" - - -def normalize(s: str) -> str: - return " ".join(re.sub(r"[^a-z0-9 ]", " ", s.lower()).split()) - - -def similarity(a: str, b: str) -> float: - return difflib.SequenceMatcher(None, normalize(a), normalize(b)).ratio() - - -def _get(url: str, timeout: float = 25.0) -> bytes | None: - try: - req = urllib.request.Request(url, headers={"User-Agent": UA}) - with urllib.request.urlopen(req, timeout=timeout) as resp: - return resp.read() - except Exception: # noqa: BLE001 — a dead source is inconclusive, not proof - return None - - -def crossref_best(title: str) -> tuple[float, str]: - q = urllib.parse.urlencode( - {"query.bibliographic": title, "rows": 5, "select": "title,DOI"} - ) - raw = _get(f"https://api.crossref.org/works?{q}") - if raw is None: - return (-1.0, "(crossref unreachable)") - try: - items = json.loads(raw)["message"]["items"] - except Exception: # noqa: BLE001 - return (-1.0, "(crossref unparseable)") - best, best_t = 0.0, "" - for it in items: - for t in it.get("title") or []: - if (s := similarity(title, t)) > best: - best, best_t = s, t - return best, best_t - - -def arxiv_best(title: str) -> tuple[float, str]: - q = urllib.parse.urlencode( - {"search_query": f'ti:"{title}"', "max_results": 5} - ) - raw = _get(f"http://export.arxiv.org/api/query?{q}") - if raw is None: - return (-1.0, "(arxiv unreachable)") - try: - root = ET.fromstring(raw) - except ET.ParseError: - return (-1.0, "(arxiv unparseable)") - ns = {"a": "http://www.w3.org/2005/Atom"} - best, best_t = 0.0, "" - for entry in root.findall("a:entry", ns): - node = entry.find("a:title", ns) - if node is not None and node.text: - if (s := similarity(title, node.text)) > best: - best, best_t = s, " ".join(node.text.split()) - return best, best_t - - -def check(title: str) -> dict[str, Any]: - cs, ct = crossref_best(title) - as_, at = arxiv_best(title) - best = max(cs, as_) - source, match = ("crossref", ct) if cs >= as_ else ("arxiv", at) - - if best < 0: - band = "inconclusive" - elif best >= LIKELY_REAL: - band = "likely-real" - elif best < NO_MATCH: - band = "no-match" - else: - band = "inconclusive" - - return { - "title": title, "band": band, "best_similarity": round(best, 3), - "best_source": source, "best_match": match, - "crossref": round(cs, 3), "arxiv": round(as_, 3), - } - - -def title_of(path: Path) -> str | None: - try: - text = path.read_text(encoding="utf-8", errors="replace") - except OSError: - return None - if m := re.search(r'^title:\s*"?(.+?)"?\s*$', text[:2000], re.M): - return m.group(1) - if m := re.search(r"^#\s+(.+)$", text, re.M): - return m.group(1) - return None - - -def main() -> int: - ap = argparse.ArgumentParser( - description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter - ) - ap.add_argument("--title", action="append", default=[]) - ap.add_argument("--file", action="append", default=[]) - ap.add_argument("--from-list", type=str, default="") - ap.add_argument("--json", type=str, default="") - args = ap.parse_args() - - targets: list[tuple[str, str]] = [(t, "(--title)") for t in args.title] - paths = [Path(f) for f in args.file] - if args.from_list: - paths += [ROOT / line for line in Path(args.from_list).read_text().split()] - for p in paths: - if t := title_of(p): - targets.append((t, str(p))) - - if not targets: - ap.error("nothing to check — pass --title, --file, or --from-list") - - print("=== CONTROL GATE ===", file=sys.stderr) - ctrl_ok = True - for c in REAL_CONTROLS: - r = check(c) - ok = r["band"] == "likely-real" - ctrl_ok &= ok - print(f" [+] {'PASS' if ok else 'FAIL'} {r['band']:<13} {r['best_similarity']:.2f} {c[:50]}", - file=sys.stderr) - rf = check(FAKE_CONTROL) - ok = rf["band"] == "no-match" - ctrl_ok &= ok - print(f" [-] {'PASS' if ok else 'FAIL'} {rf['band']:<13} {rf['best_similarity']:.2f} " - f"{FAKE_CONTROL[:50]}", file=sys.stderr) - - if not ctrl_ok: - print("\n*** BROKEN CHECK — results below are NOT evidence. ***", file=sys.stderr) - - results = [] - for title, origin in targets: - r = check(title) - r["file"] = origin - results.append(r) - flag = {"no-match": "NO MATCH ", "inconclusive": "inconclusive", - "likely-real": "likely-real "}[r["band"]] - print(f"{flag} {r['best_similarity']:.2f} {title[:64]}") - if r["band"] != "likely-real": - print(f" closest ({r['best_source']}): {r['best_match'][:64]}") - print(f" {origin}") - - counts = {b: sum(1 for r in results if r["band"] == b) - for b in ("no-match", "inconclusive", "likely-real")} - print(f"\nchecked {len(results)}: " + ", ".join(f"{v} {k}" for k, v in counts.items())) - print("\nA `no-match` is a question for a human, not a verdict. Crossref does not " - "index arXiv or much of ACL.") - - if args.json: - Path(args.json).write_text(json.dumps( - {"controls_ok": ctrl_ok, "counts": counts, "results": results}, indent=2 - ), encoding="utf-8") - print(f"wrote {args.json}") - - return 0 if ctrl_ok else 2 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/bin/vault-grep b/bin/vault-grep index 7cbffcc..582b660 100755 --- a/bin/vault-grep +++ b/bin/vault-grep @@ -10,6 +10,9 @@ if ! command -v rg >/dev/null 2>&1; then exit 1 fi +# Let ripgrep parse patterns, option values, paths, and stdin itself. Its default +# filesystem search is MainFrame; an explicit path must not add a root search. +cd "$ROOT" exec rg --no-ignore \ -g '!.git/' \ -g '!.venv/' \ @@ -19,4 +22,4 @@ exec rg --no-ignore \ -g '!.grok/' \ -g '!node_modules/' \ -g '!.ruff_cache/' \ - "$@" "$ROOT" + "$@" diff --git a/bin/work-inventory b/bin/work-inventory new file mode 100755 index 0000000..5535405 --- /dev/null +++ b/bin/work-inventory @@ -0,0 +1,82 @@ +#!/usr/bin/env python3 +"""Validate the typed cross-lifecycle work inventory. + +This command is a derived check. Direct README authorities remain the path +resolver; the inventory only reports whether the records and manifest agree. +""" + +from __future__ import annotations + +import argparse +import json +import sys +from pathlib import Path + + +ROOT = Path(__file__).resolve().parents[1] +SCRIPTS = ROOT / "scripts" +if str(SCRIPTS) not in sys.path: + sys.path.insert(0, str(SCRIPTS)) + +from work_inventory import inventory_payload, validate_manifest # noqa: E402 + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--check", action="store_true", help="validate direct identities and optional typed manifest") + parser.add_argument("--json", action="store_true", help="emit JSON") + parser.add_argument("--root", type=Path, default=ROOT, help="MainFrame root (for shadow fixtures)") + parser.add_argument("--manifest", type=Path, default=None, help="typed projects manifest to validate") + parser.add_argument("--output", type=Path, default=None, help="write JSON inventory receipt") + parser.add_argument("--human-output", type=Path, default=None, help="write a human-readable inventory") + args = parser.parse_args(argv) + root = args.root.expanduser().resolve() + payload = inventory_payload(root) + if args.manifest: + payload["manifest"] = validate_manifest(root, args.manifest) + payload["ok"] = bool(payload["ok"] and payload["manifest"]["ok"]) + if args.output: + args.output.parent.mkdir(parents=True, exist_ok=True) + args.output.write_text(json.dumps(payload, indent=2, sort_keys=True) + "\n", encoding="utf-8") + if args.human_output: + args.human_output.parent.mkdir(parents=True, exist_ok=True) + lines = [ + "# MainFrame Work Inventory", + "", + f"- root: `{root}`", + f"- records: {payload['count']}", + f"- status: `{'OK' if payload['ok'] else 'BLOCKED'}`", + "", + "| id | record_type | path | lifecycle_state | README SHA-256 |", + "| --- | --- | --- | --- | --- |", + ] + for record in payload["records"]: + lines.append( + "| {id} | {record_type} | `{path}` | {lifecycle_state} | `{readme_sha256}` |".format( + id=record.get("id") or "", + record_type=record.get("record_type") or "", + path=record.get("path") or "", + lifecycle_state=record.get("lifecycle_state") or "", + readme_sha256=record.get("readme_sha256") or "", + ) + ) + if payload.get("issues"): + lines.extend(["", "## Issues", ""]) + lines.extend(f"- {issue}" for issue in payload["issues"]) + args.human_output.write_text("\n".join(lines) + "\n", encoding="utf-8") + if args.json: + print(json.dumps(payload, indent=2, sort_keys=True)) + else: + print("mainframe work inventory") + print(f" root: {root}") + print(f" records: {payload['count']}") + print(f" status: {'OK' if payload['ok'] else 'BLOCKED'}") + for issue in payload.get("issues", []): + print(f" issue: {issue}") + if payload.get("manifest"): + print(f" manifest: {'OK' if payload['manifest']['ok'] else 'BLOCKED'}") + return 0 if payload["ok"] else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/bin/workflow-event b/bin/workflow-event deleted file mode 100755 index b8fe446..0000000 --- a/bin/workflow-event +++ /dev/null @@ -1,633 +0,0 @@ -#!/usr/bin/env python3 -"""Append redacted agent-client hook event telemetry to 20_live.""" - -from __future__ import annotations - -import hashlib -import json -import os -import re -import sys -import fcntl -from datetime import datetime -from pathlib import Path -from zoneinfo import ZoneInfo - - -ROOT = Path(__file__).resolve().parents[1] -DEFAULT_DIR = ROOT / "20_live" / "workflow-metrics" / "events" -TZ = ZoneInfo("America/Toronto") -MAX_EVENT_LINE_BYTES = 256 * 1024 -MAX_HOOK_PAYLOAD_CHARS = 1024 * 1024 - -# A leading shell token shaped like NAME=value is an inline env assignment -# (e.g. ``FOO=/tmp/bar cmd``). Skip it when picking the command head so the -# assignment value does not leak into telemetry. -ENV_ASSIGNMENT_RE = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*=") -# A safe command head is a plain program name: alphanumerics with the usual -# punctuation, no slashes, no path separators. Anything outside this shape is -# collapsed before it is written to disk. -SAFE_HEAD_RE = re.compile(r"^[A-Za-z0-9][A-Za-z0-9._+-]*$") -SAFE_PROCESS_ID_RE = re.compile( - r"^(?:cli|script|ingest|workflow|skill|agent)-[a-z0-9][a-z0-9._-]{0,72}$" -) -SHELL_TOOL_NAMES = { - "Bash", - "exec_command", - "functions.exec_command", - "run_terminal_command", -} -ALLOWED_DIAGNOSTICS = { - "command_failure", - "context_limit_exceeded", - "edit_format_failure", - "shell_command_denied", -} -ALLOWED_TOOL_PURPOSES = {"operation", "verification"} -EVENT_NAME_ALIASES = { - "session_start": "SessionStart", - "user_prompt_submit": "UserPromptSubmit", - "pre_tool_use": "PreToolUse", - "post_tool_use": "PostToolUse", - "post_tool_use_failure": "PostToolUseFailure", - "permission_request": "PermissionRequest", - "permission_denied": "PermissionDenied", - "stop": "Stop", - "stop_failure": "StopFailure", - "notification": "Notification", - "subagent_start": "SubagentStart", - "subagent_stop": "SubagentStop", - "subagent_end": "SubagentStop", - "pre_compact": "PreCompact", - "post_compact": "PostCompact", - "session_end": "SessionEnd", -} - - -def digest(value: object) -> str | None: - if value in (None, ""): - return None - return hashlib.sha256(str(value).encode("utf-8")).hexdigest()[:16] - - -def path_class(path_value: object) -> dict[str, str | None]: - if not path_value: - return {"path_zone": None, "extension": None} - try: - path = Path(str(path_value)).expanduser() - if not path.is_absolute(): - path = (ROOT / path).resolve() - relative = path.relative_to(ROOT) - zone = relative.parts[0] if relative.parts else "." - except Exception: - zone = "external" - path = Path(str(path_value)) - return {"path_zone": zone, "extension": path.suffix or None} - - -def command_head(command: str) -> str | None: - """Pick the bare program name from a shell command line. - - Skips leading env-var assignments (``FOO=/tmp/bar cmd`` -> ``cmd``) so - inline assignment values do not leak. Collapses path-shaped heads to - ``"<path>"`` and anything else unexpected to ``"<other>"`` so only the - shape of the call is recorded, never a filesystem basename. - """ - tokens = command.split() - while tokens and ENV_ASSIGNMENT_RE.match(tokens[0]): - tokens.pop(0) - if not tokens: - return None - head = tokens[0] - if "/" in head: - return "<path>" - if not SAFE_HEAD_RE.match(head): - return "<other>" - return head - - -def bash_summary(tool_input: dict) -> dict[str, object]: - command = str(tool_input.get("command") or tool_input.get("cmd") or "").strip() - return { - "command_head": command_head(command) if command else None, - "command_hash": digest(command), - "timeout": tool_input.get("timeout"), - "background": tool_input.get("run_in_background"), - } - - -def _process_id_candidates(value: object) -> list[str]: - if value in (None, ""): - return [] - if isinstance(value, str): - return [part.strip() for part in value.split(",")] - if isinstance(value, list): - return [ - str(item).strip() - for item in value - if item not in (None, "") - ] - return [str(value).strip()] - - -def process_ids(raw: dict) -> list[str]: - """Return validated catalogue/process ids from allowlisted fields/env vars.""" - candidates: list[str] = [] - for key in ("process_id", "process_ids", "catalogue_id", "catalogue_ids"): - candidates.extend(_process_id_candidates(raw.get(key))) - for key in ("MAINFRAME_PROCESS_ID", "MAINFRAME_CATALOGUE_ID"): - candidates.extend(_process_id_candidates(os.environ.get(key))) - - ids: list[str] = [] - seen: set[str] = set() - for candidate in candidates: - if not SAFE_PROCESS_ID_RE.match(candidate): - continue - if candidate in seen: - continue - ids.append(candidate) - seen.add(candidate) - return ids - - -def input_summary(tool_name: str | None, tool_input: dict) -> dict[str, object]: - if tool_name in SHELL_TOOL_NAMES: - return bash_summary(tool_input) - if tool_name in { - "Read", - "Edit", - "Write", - "MultiEdit", - "read_file", - "search_replace", - "apply_patch", - }: - return path_class( - tool_input.get("file_path") - or tool_input.get("path") - ) - if tool_name in {"Glob", "Grep", "grep", "list_dir"}: - summary = path_class( - tool_input.get("path") - or tool_input.get("directory") - ) - summary["pattern_hash"] = digest(tool_input.get("pattern")) - summary["glob_hash"] = digest(tool_input.get("glob")) - return summary - if tool_name: - return {"input_hash": digest(json.dumps(tool_input, sort_keys=True, default=str))} - return {} - - -PERMISSION_HINTS = ("permission", "approve", "approval", "allow", "needs your") -IDLE_HINTS = ("waiting for", "is waiting", "idle") - - -def notification_kind(message: object) -> str | None: - """Coarsely classify a notification without storing its text. - - ``permission`` means the agent is blocked on a human decision (review - table); ``idle`` means it is waiting on input; ``other`` is anything else. - """ - if not message: - return None - lowered = str(message).lower() - if any(hint in lowered for hint in PERMISSION_HINTS): - return "permission" - if any(hint in lowered for hint in IDLE_HINTS): - return "idle" - return "other" - - -def response_success(tool_response: object) -> bool | None: - """Best-effort tool outcome from a post-tool payload. - - Claude Code signals failures with a separate PostToolUseFailure event, - but Codex reports the outcome inside PostToolUse's tool_response. Look - for an exit code or an explicit success flag, at the top level or one - level down in ``metadata``; return None when the shape is unrecognized - so the caller keeps the event-name default. - """ - if not isinstance(tool_response, dict): - return None - containers = [tool_response] - if isinstance(tool_response.get("metadata"), dict): - containers.append(tool_response["metadata"]) - for container in containers: - for key in ("exit_code", "exitCode", "returncode"): - value = container.get(key) - if isinstance(value, int) and not isinstance(value, bool): - return value == 0 - if isinstance(container.get("success"), bool): - return container["success"] - return None - - -def normalize_hook_payload(raw: dict) -> dict: - """Normalize supported client hook envelopes to MainFrame's schema. - - Grok Build uses camelCase keys and lower snake-case event names, while the - existing Claude/Codex hooks use snake_case keys and canonical event names. - Only fields consumed by ``safe_event`` are copied; raw content remains - subject to the same redaction boundary. - """ - normalized = dict(raw) - aliases = { - "hookEventName": "hook_event_name", - "sessionId": "session_id", - "transcriptPath": "transcript_path", - "workspaceRoot": "workspace_root", - "permissionMode": "permission_mode", - "toolName": "tool_name", - "toolInput": "tool_input", - "toolUseId": "tool_use_id", - "toolResponse": "tool_response", - "durationMs": "duration_ms", - "stopHookActive": "stop_hook_active", - "modelId": "model", - } - for source, target in aliases.items(): - if target not in normalized and source in raw: - normalized[target] = raw[source] - if not normalized.get("cwd"): - normalized["cwd"] = raw.get("workspaceRoot") - event_name = normalized.get("hook_event_name") - if isinstance(event_name, str): - normalized["hook_event_name"] = EVENT_NAME_ALIASES.get( - event_name, - event_name, - ) - return normalized - - -def safe_event(raw: dict, client: str | None = None, pixel: str | None = None) -> dict[str, object]: - now = datetime.now(TZ) - event_name = raw.get("hook_event_name") - tool_name = raw.get("tool_name") - tool_input = raw.get("tool_input") if isinstance(raw.get("tool_input"), dict) else {} - event = { - "logged_at": now.isoformat(timespec="seconds"), - "as_of": now.date().isoformat(), - "event": event_name, - "client": client, - "pixel": pixel, # unique/trackable identifier for the local agents pixel agent tracker - "session_hash": digest(raw.get("session_id")), - "transcript_hash": digest(raw.get("transcript_path")), - "cwd_zone": path_class(raw.get("cwd")).get("path_zone"), - "permission_mode": raw.get("permission_mode"), - "effort": (raw.get("effort") or {}).get("level") - if isinstance(raw.get("effort"), dict) - else None, - "tool_name": tool_name, - "tool_use_hash": digest(raw.get("tool_use_id")), - "duration_ms": raw.get("duration_ms"), - "input_summary": input_summary(tool_name, tool_input), - } - safe_process_ids = process_ids(raw) - if safe_process_ids: - event["process_id"] = safe_process_ids[0] - event["process_ids"] = safe_process_ids - if raw.get("source") == "aider": - event["source"] = "aider" - # Event-specific signals. Prompt and notification text are never stored - # verbatim - only a hash, a length, or a coarse kind - so the live tracker - # can light up planning / approval / handoff states without leaking content. - if event_name == "SessionStart": - event["source"] = raw.get("source") - event["model"] = raw.get("model") # best-effort; absent on some clients - elif event_name == "SessionEnd": - event["reason"] = raw.get("reason") - elif event_name == "UserPromptSubmit": - prompt = raw.get("prompt") - event["prompt_hash"] = digest(prompt) - event["prompt_chars"] = len(prompt) if isinstance(prompt, str) else None - elif event_name == "Notification": - event["notification_hash"] = digest(raw.get("message")) - event["notification_kind"] = notification_kind(raw.get("message")) - elif event_name == "PermissionRequest": - # A dedicated approval event (Claude Code and Codex). Reuse the - # notification_kind vocabulary so downstream consumers have one - # "awaiting a human decision" signal. - event["notification_kind"] = "permission" - elif event_name in {"Stop", "SubagentStop"}: - event["stop_hook_active"] = raw.get("stop_hook_active") - elif event_name in {"PreCompact", "PostCompact"}: - event["trigger"] = raw.get("trigger") - elif event_name == "Diagnostic": - diagnostic = raw.get("diagnostic") - if diagnostic in ALLOWED_DIAGNOSTICS: - event["diagnostic"] = diagnostic - tool_purpose = raw.get("tool_purpose") - if tool_purpose in ALLOWED_TOOL_PURPOSES: - event["tool_purpose"] = tool_purpose - observed_count = raw.get("observed_count") - if isinstance(observed_count, int) and not isinstance(observed_count, bool): - event["observed_count"] = max(1, observed_count) - if event_name in {"PostToolUseFailure", "PermissionDenied", "StopFailure"}: - event["success"] = False - elif event_name == "PostToolUse": - derived = response_success(raw.get("tool_response")) - event["success"] = True if derived is None else derived - return event - - -def parse_client(argv: list[str]) -> str | None: - """Read ``--client NAME`` (or ``--client=NAME``) from the hook command. - - The client cannot be inferred reliably from the payload, so each agent's - hook config passes it explicitly (``--client claude`` / ``--client codex``). - """ - for index, arg in enumerate(argv): - if arg == "--client" and index + 1 < len(argv): - return argv[index + 1] - if arg.startswith("--client="): - return arg.split("=", 1)[1] - return None - - -def parse_pixel(argv: list[str]) -> str | None: - """Read optional ``--pixel ID`` for unique per-agent/pixel tracking. - - Allows the local agents pixel agent tracker to distinguish specific - "pixels" (named agent instances or setups) while still grouping by client. - Example hook: .../workflow-event --client claude --pixel desktop-main-claude - """ - for index, arg in enumerate(argv): - if arg == "--pixel" and index + 1 < len(argv): - return argv[index + 1] - if arg.startswith("--pixel="): - return arg.split("=", 1)[1] - return None - - -def get_client_override() -> str | None: - import subprocess - import os - curr = os.getpid() - for _ in range(10): - if curr <= 1: - break - try: - res = subprocess.run( - ["ps", "-p", str(curr), "-o", "ppid,command"], - capture_output=True, text=True, check=False - ) - if res.returncode != 0: - break - lines = res.stdout.strip().splitlines() - if len(lines) < 2: - break - parts = lines[1].split(None, 1) - ppid = int(parts[0]) - cmd = parts[1].lower() - if "antigravity" in cmd: - return "antigravity" - if "codex" in cmd: - return "codex" - curr = ppid - except Exception: - break - return None - - -def detect_loop(log_path: Path, session_hash: str) -> dict[str, object]: - """Read the last few events in log_path for the same session to detect loops. - - If we see multiple consecutive failures or retries in the same session, we flag it. - """ - if not log_path.exists(): - return {"is_loop": False, "loop_count": 0} - try: - lines = [] - with log_path.open("r", encoding="utf-8") as f: - f.seek(0, 2) - file_size = f.tell() - buffer_size = 4096 - pos = file_size - while pos > 0 and len(lines) < 15: - pos = max(0, pos - buffer_size) - f.seek(pos) - chunk = f.read(buffer_size) - lines = chunk.splitlines() + lines - lines = [l for l in lines if l.strip()] - - session_events = [] - for line in reversed(lines): - try: - ev = json.loads(line) - if ev.get("session_hash") == session_hash: - session_events.append(ev) - except Exception: - continue - if len(session_events) >= 6: - break - - failures = 0 - for ev in session_events: - if ev.get("success") is False or ev.get("event") == "PostToolUseFailure": - failures += 1 - else: - break - if failures >= 2: - return {"is_loop": True, "loop_count": failures} - except Exception: - pass - return {"is_loop": False, "loop_count": 0} - - -def calculate_duration_from_pre_tool( - log_path: Path, - session_hash: str | None, - tool_use_hash: str | None, - now: datetime, -) -> int | None: - """Read back logs in log_path to calculate tool duration from PreToolUse. - - Used as fallback when the agent-client hook does not supply duration_ms. - """ - if not log_path.exists() or not session_hash or not tool_use_hash: - return None - try: - lines = [] - with log_path.open("r", encoding="utf-8") as f: - f.seek(0, 2) - file_size = f.tell() - buffer_size = 4096 - pos = file_size - while pos > 0 and len(lines) < 50: - pos = max(0, pos - buffer_size) - f.seek(pos) - chunk = f.read(buffer_size) - lines = chunk.splitlines() + lines - lines = [l for l in lines if l.strip()] - - for line in reversed(lines): - try: - ev = json.loads(line) - if ( - ev.get("session_hash") == session_hash - and ev.get("tool_use_hash") == tool_use_hash - and ev.get("event") == "PreToolUse" - and ev.get("logged_at") - ): - pre_time = datetime.fromisoformat(ev["logged_at"]) - diff = now - pre_time - return int(diff.total_seconds() * 1000) - except Exception: - continue - except Exception: - pass - return None - - -def compute_hash(event_data: dict, prev_hash: str) -> str: - event_str = json.dumps(event_data, sort_keys=True, separators=(",", ":")) - h = hashlib.sha256((event_str + prev_hash).encode("utf-8")) - return h.hexdigest()[:16] - - -def _last_event_from_locked_file(handle) -> dict | None: - """Return the final non-empty JSONL record without reading the whole log.""" - handle.seek(0, os.SEEK_END) - end = handle.tell() - if end == 0: - return None - - start = max(0, end - MAX_EVENT_LINE_BYTES - 1) - handle.seek(start) - tail = handle.read(end - start).rstrip() - if not tail: - if end > MAX_EVENT_LINE_BYTES: - raise ValueError("telemetry tail exceeds size limit") - return None - - newline = tail.rfind(b"\n") - if newline < 0: - if start > 0: - raise ValueError("final telemetry record exceeds size limit") - line = tail - else: - line = tail[newline + 1 :] - if len(line) > MAX_EVENT_LINE_BYTES: - raise ValueError("final telemetry record exceeds size limit") - value = json.loads(line.decode("utf-8")) - if not isinstance(value, dict): - raise ValueError("final telemetry record is not an object") - return value - - -def post_or_append_event(event: dict, log_path: Path, allow_post: bool = True) -> bool: - import urllib.request - import urllib.error - - url = "http://127.0.0.1:5177/api/log-event" - event_without_hash = {k: v for k, v in event.items() if k != "hash_chain"} - - # Try POSTing to server. An explicit WORKFLOW_METRICS_DIR override means - # "this directory is the destination" (tests rely on this), so the caller - # disables the POST and the event appends locally regardless of whether a - # workstation server happens to be listening. - if allow_post: - try: - payload = json.dumps({"event": event_without_hash}, sort_keys=True, separators=(",", ":")).encode("utf-8") - req = urllib.request.Request(url, data=payload, headers={"Content-Type": "application/json"}) - with urllib.request.urlopen(req, timeout=1.0) as response: - if response.status == 200: - res_data = json.loads(response.read().decode("utf-8")) - if res_data.get("ok"): - event["hash_chain"] = res_data.get("hash_chain") - return True - except Exception as e: - print(f"Server log-event endpoint unavailable ({e}); falling back to local file append", file=sys.stderr) - - # Fallback: hold one cross-process lock across tail read, hash calculation, - # and append. Without this critical section concurrent hooks can fork the - # linear chain by deriving two records from the same predecessor. - try: - with log_path.open("a+b") as handle: - fcntl.flock(handle.fileno(), fcntl.LOCK_EX) - last_event = _last_event_from_locked_file(handle) - prev_hash = "genesis-seed" - if last_event is not None: - previous = last_event.get("hash_chain") - if not isinstance(previous, str) or not previous: - raise ValueError("final telemetry record has no hash_chain") - prev_hash = previous - - event["hash_chain"] = compute_hash(event_without_hash, prev_hash) - encoded = ( - json.dumps(event, sort_keys=True, separators=(",", ":")) + "\n" - ).encode("utf-8") - if len(encoded) > MAX_EVENT_LINE_BYTES: - raise ValueError("telemetry record exceeds size limit") - handle.write(encoded) - handle.flush() - return True - except Exception as exc: - print(f"workflow-event local append failed: {exc}", file=sys.stderr) - return False - - -def main() -> int: - if any(arg in {"--help", "-h"} for arg in sys.argv[1:]): - print("Read agent-client hook JSON from stdin and append redacted JSONL telemetry.") - return 0 - - client = parse_client(sys.argv[1:]) - pixel = parse_pixel(sys.argv[1:]) - if client is None: - override = get_client_override() - if os.environ.get("ANTIGRAVITY_AGENT") == "1": - client = "antigravity" - elif override: - client = override - try: - payload = sys.stdin.read(MAX_HOOK_PAYLOAD_CHARS + 1) - if not payload.strip(): - return 0 - if len(payload) > MAX_HOOK_PAYLOAD_CHARS: - raise ValueError("hook payload exceeds size limit") - raw = normalize_hook_payload(json.loads(payload)) - if client == "grok": - raw.setdefault("source", "grok") - out_dir = Path(os.environ.get("WORKFLOW_METRICS_DIR", DEFAULT_DIR)) - out_dir.mkdir(parents=True, exist_ok=True) - event = safe_event(raw, client, pixel=pixel) - log_path = out_dir / f"{event['as_of']}.jsonl" - - if event.get("session_hash"): - loop_info = detect_loop(log_path, event["session_hash"]) - event.update(loop_info) - - if event.get("event") in {"PostToolUse", "PostToolUseFailure"} and event.get("duration_ms") is None: - logged_at_str = event.get("logged_at") - if isinstance(logged_at_str, str): - try: - now_dt = datetime.fromisoformat(logged_at_str) - calc_duration = calculate_duration_from_pre_tool( - log_path, - event.get("session_hash"), - event.get("tool_use_hash"), - now_dt, - ) - if calc_duration is not None: - event["duration_ms"] = calc_duration - except Exception: - pass - - allow_post = "WORKFLOW_METRICS_DIR" not in os.environ - success = post_or_append_event(event, log_path, allow_post) - if not success: - return 1 - - - # PermissionRequest is telemetry only. Permission decisions remain with - # the native client; this recorder must never poll approval files, - # delay the client, or return an approve/deny decision. - except Exception as exc: - print(f"workflow-event telemetry skipped: {exc}", file=sys.stderr) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/bin/workflow-report b/bin/workflow-report deleted file mode 100755 index db9ab03..0000000 --- a/bin/workflow-report +++ /dev/null @@ -1,510 +0,0 @@ -#!/usr/bin/env python3 -"""Summarize local workflow telemetry without exposing raw prompts or outputs.""" - -from __future__ import annotations - -import argparse -import json -import re -from collections import Counter, defaultdict -from datetime import date, timedelta -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -DEFAULT_DIR = ROOT / "20_live" / "workflow-metrics" / "events" -SAFE_PROCESS_ID_RE = re.compile( - r"^(?:cli|script|ingest|workflow|skill|agent)-[a-z0-9][a-z0-9._-]{0,72}$" -) - - -def load_events(log_dir: Path, days: int) -> list[dict]: - start = date.today() - timedelta(days=days - 1) - events: list[dict] = [] - for path in sorted(log_dir.glob("*.jsonl")): - try: - if date.fromisoformat(path.stem) < start: - continue - except ValueError: - continue - for line in path.read_text(encoding="utf-8").splitlines(): - if not line.strip(): - continue - try: - event = json.loads(line) - except json.JSONDecodeError: - continue - if isinstance(event, dict): - events.append(event) - return events - - -def percentage(numerator: int, denominator: int) -> float | None: - if denominator == 0: - return None - return round((numerator / denominator) * 100, 1) - - -def named_counts(counter: Counter[str]) -> list[dict[str, object]]: - return [ - {"name": name, "count": count} - for name, count in counter.most_common() - ] - - -def event_process_ids(event: dict) -> list[str]: - values = event.get("process_ids") - if isinstance(values, list): - ids = [str(value) for value in values if value] - else: - ids = [] - single = event.get("process_id") - if single: - ids.insert(0, str(single)) - - unique: list[str] = [] - seen: set[str] = set() - for process_id in ids: - if not SAFE_PROCESS_ID_RE.match(process_id): - continue - if process_id in seen: - continue - unique.append(process_id) - seen.add(process_id) - return unique - - -def client_quality(events: list[dict]) -> list[dict[str, object]]: - """Per-client quality rollup; a session counts under every client tag it carries.""" - grouped: dict[str, list[dict]] = defaultdict(list) - for event in events: - grouped[event.get("client") or "unknown"].append(event) - - rollup: list[dict[str, object]] = [] - for client, client_events in grouped.items(): - sessions = { - event.get("session_hash") - for event in client_events - if event.get("session_hash") - } - starts = { - event.get("session_hash") - for event in client_events - if event.get("event") == "SessionStart" and event.get("session_hash") - } - ends = { - event.get("session_hash") - for event in client_events - if event.get("event") == "SessionEnd" and event.get("session_hash") - } - post_events = [ - event for event in client_events - if event.get("event") in {"PostToolUse", "PostToolUseFailure"} - ] - timed_post_events = sum( - 1 - for event in post_events - if isinstance(event.get("duration_ms"), (int, float)) - ) - closed = len(starts & ends) - rollup.append({ - "name": client, - "events": len(client_events), - "event_types": sorted( - {event.get("event") or "unknown" for event in client_events} - ), - "sessions": len(sessions), - "sessions_with_start": len(starts), - "sessions_with_end": len(ends), - "closed_started_sessions": closed, - "session_close_coverage_pct": percentage(closed, len(starts)), - "post_tool_events": len(post_events), - "timed_post_tool_events": timed_post_events, - "duration_coverage_pct": percentage( - timed_post_events, len(post_events) - ), - "tool_failures": sum( - 1 for event in client_events if event.get("success") is False - ), - }) - rollup.sort(key=lambda item: item["events"], reverse=True) - return rollup - - -def process_rollup(events: list[dict]) -> list[dict[str, object]]: - grouped: dict[str, list[dict]] = defaultdict(list) - for event in events: - for process_id in event_process_ids(event): - grouped[process_id].append(event) - - rollup: list[dict[str, object]] = [] - for process_id, process_events in grouped.items(): - sessions = { - event.get("session_hash") - for event in process_events - if event.get("session_hash") - } - clients = Counter( - event.get("client") or "unknown" - for event in process_events - ) - rollup.append({ - "process_id": process_id, - "events": len(process_events), - "sessions": len(sessions), - "tool_failures": sum( - 1 for event in process_events if event.get("success") is False - ), - "event_types": sorted( - {event.get("event") or "unknown" for event in process_events} - ), - "clients": named_counts(clients), - }) - rollup.sort(key=lambda item: item["events"], reverse=True) - return rollup - - -def input_signal_rollup(events: list[dict]) -> dict[str, object]: - command_heads: Counter[str] = Counter() - path_zones: Counter[str] = Counter() - extensions: Counter[str] = Counter() - redacted_input_events = 0 - - for event in events: - summary = event.get("input_summary") - if not isinstance(summary, dict) or not summary: - continue - redacted_input_events += 1 - command_head = summary.get("command_head") - if isinstance(command_head, str) and command_head: - command_heads[command_head] += 1 - path_zone = summary.get("path_zone") - if isinstance(path_zone, str) and path_zone: - path_zones[path_zone] += 1 - extension = summary.get("extension") - if isinstance(extension, str) and extension: - extensions[extension] += 1 - - return { - "redacted_input_events": redacted_input_events, - "command_heads": named_counts(command_heads), - "path_zones": named_counts(path_zones), - "file_extensions": named_counts(extensions), - } - - -def build_summary( - events: list[dict], - days: int, - *, - by_process: bool = False, - input_signals: bool = False, -) -> dict[str, object]: - by_event = Counter(event.get("event") or "unknown" for event in events) - by_tool = Counter(event.get("tool_name") or "none" for event in events) - by_client = Counter(event.get("client") or "unknown" for event in events) - failures = [event for event in events if event.get("success") is False] - failures_by_tool = Counter( - event.get("tool_name") or "unknown" for event in failures - ) - diagnostics = Counter( - event.get("diagnostic") - for event in events - if event.get("diagnostic") - ) - verification_events = sum( - 1 for event in events if event.get("tool_purpose") == "verification" - ) - post_events = [ - event for event in events - if event.get("event") in {"PostToolUse", "PostToolUseFailure"} - ] - durations = [ - { - "duration_ms": event.get("duration_ms"), - "tool_name": event.get("tool_name") or "unknown", - } - for event in events - if isinstance(event.get("duration_ms"), (int, float)) - ] - durations.sort(key=lambda item: item["duration_ms"], reverse=True) - - sessions = { - event.get("session_hash") - for event in events - if event.get("session_hash") - } - start_sessions = { - event.get("session_hash") - for event in events - if event.get("event") == "SessionStart" and event.get("session_hash") - } - end_sessions = { - event.get("session_hash") - for event in events - if event.get("event") == "SessionEnd" and event.get("session_hash") - } - - pairing: dict[str, Counter[str]] = defaultdict(Counter) - outcome_only_source_events = 0 - for event in events: - session = event.get("session_hash") - if not session: - continue - if event.get("event") == "PreToolUse": - pairing[session]["pre"] += 1 - elif event.get("event") in {"PostToolUse", "PostToolUseFailure"}: - if event.get("source") == "aider": - outcome_only_source_events += 1 - continue - pairing[session]["outcome"] += 1 - - deltas = [ - counts["pre"] - counts["outcome"] - for counts in pairing.values() - if counts["pre"] or counts["outcome"] - ] - client_tagged = sum( - count for name, count in by_client.items() if name != "unknown" - ) - process_tagged = sum(1 for event in events if event_process_ids(event)) - timed_post_events = sum( - 1 - for event in post_events - if isinstance(event.get("duration_ms"), (int, float)) - ) - closed_started_sessions = len(start_sessions & end_sessions) - - summary: dict[str, object] = { - "days": days, - "events": len(events), - "sessions": len(sessions), - "tool_failures": len(failures), - "events_by_type": named_counts(by_event), - "tool_calls": named_counts(by_tool), - "clients": named_counts(by_client), - "failures_by_tool": named_counts(failures_by_tool), - "diagnostics": named_counts(diagnostics), - "verification_events": verification_events, - "slowest_tool_calls": durations[:10], - "telemetry_quality": { - "client_tagged_events": client_tagged, - "client_tag_coverage_pct": percentage(client_tagged, len(events)), - "process_tagged_events": process_tagged, - "process_tag_coverage_pct": percentage(process_tagged, len(events)), - "post_tool_events": len(post_events), - "timed_post_tool_events": timed_post_events, - "duration_coverage_pct": percentage( - timed_post_events, len(post_events) - ), - "sessions_with_start": len(start_sessions), - "sessions_with_end": len(end_sessions), - "closed_started_sessions": closed_started_sessions, - "session_close_coverage_pct": percentage( - closed_started_sessions, len(start_sessions) - ), - "tool_sessions": len(deltas), - "unbalanced_tool_sessions": sum(delta != 0 for delta in deltas), - "unmatched_pre_tool_events": sum(max(delta, 0) for delta in deltas), - "unmatched_outcome_events": sum(max(-delta, 0) for delta in deltas), - "outcome_only_source_events": outcome_only_source_events, - "by_client": client_quality(events), - }, - } - if by_process: - summary["processes"] = process_rollup(events) - if input_signals: - summary["input_signals"] = input_signal_rollup(events) - return summary - - -def format_coverage(value: float | None) -> str: - return "n/a" if value is None else f"{value:.1f}%" - - -def print_summary(summary: dict[str, object]) -> None: - quality = summary["telemetry_quality"] - assert isinstance(quality, dict) - - print(f"Workflow telemetry summary, last {summary['days']} days") - print(f"Events: {summary['events']}") - print(f"Sessions: {summary['sessions']}") - print(f"Tool failures: {summary['tool_failures']}") - print("") - print("Events by type:") - for item in summary["events_by_type"]: - print(f"- {item['name']}: {item['count']}") - print("") - print("Tool calls:") - for item in summary["tool_calls"][:10]: - print(f"- {item['name']}: {item['count']}") - - failures_by_tool = summary["failures_by_tool"] - if failures_by_tool: - print("") - print("Tool failures by type:") - for item in failures_by_tool: - print(f"- {item['name']}: {item['count']}") - - diagnostics = summary["diagnostics"] - if diagnostics: - print("") - print("Run diagnostics:") - for item in diagnostics: - print(f"- {item['name']}: {item['count']}") - - if summary["verification_events"]: - print("") - print(f"Observed verification events: {summary['verification_events']}") - - slowest = summary["slowest_tool_calls"] - if slowest: - print("") - print("Slowest timed tool calls:") - for item in slowest: - print(f"- {item['tool_name']}: {item['duration_ms']} ms") - - print("") - print("Telemetry quality:") - print( - "- Client tags: " - f"{quality['client_tagged_events']}/{summary['events']} " - f"({format_coverage(quality['client_tag_coverage_pct'])})" - ) - print( - "- Duration coverage: " - f"{quality['timed_post_tool_events']}/{quality['post_tool_events']} " - f"post-tool events " - f"({format_coverage(quality['duration_coverage_pct'])})" - ) - print( - "- Session close coverage: " - f"{quality['closed_started_sessions']}/{quality['sessions_with_start']} " - f"started sessions " - f"({format_coverage(quality['session_close_coverage_pct'])})" - ) - print( - "- Tool event pairing: " - f"{quality['unbalanced_tool_sessions']} unbalanced session(s), " - f"{quality['unmatched_pre_tool_events']} unmatched pre-tool event(s), " - f"{quality['unmatched_outcome_events']} unmatched outcome event(s)" - ) - if quality["outcome_only_source_events"]: - print( - "- Outcome-only source events: " - f"{quality['outcome_only_source_events']} " - "(excluded from pre/post pairing)" - ) - - by_client = quality.get("by_client") or [] - if by_client: - print("") - print("Telemetry quality by client:") - for item in by_client: - print( - f"- {item['name']}: {item['events']} events, " - f"{item['sessions']} session(s); " - f"close {item['closed_started_sessions']}/{item['sessions_with_start']} " - f"({format_coverage(item['session_close_coverage_pct'])}); " - f"durations {item['timed_post_tool_events']}/{item['post_tool_events']} " - f"({format_coverage(item['duration_coverage_pct'])}); " - f"failures {item['tool_failures']}" - ) - - processes = summary.get("processes") or [] - if processes: - print("") - print("Processes:") - for item in processes[:20]: - clients = ", ".join( - f"{client['name']}:{client['count']}" - for client in item["clients"] - ) - print( - f"- {item['process_id']}: {item['events']} events, " - f"{item['sessions']} session(s), failures {item['tool_failures']}; " - f"clients {clients}" - ) - - input_signals = summary.get("input_signals") - if isinstance(input_signals, dict): - print("") - print("Redacted input signals:") - print( - "- Events with redacted input summaries: " - f"{input_signals['redacted_input_events']}" - ) - for label, key in ( - ("Command heads", "command_heads"), - ("Path zones", "path_zones"), - ("File extensions", "file_extensions"), - ): - items = input_signals.get(key) or [] - if items: - values = ", ".join( - f"{item['name']}:{item['count']}" for item in items[:10] - ) - print(f"- {label}: {values}") - - print("") - print( - "Interpretation: tool failures measure execution events, not task quality. " - "A client with started sessions but no SessionEnd events may never emit " - "that event; read per-client coverage before trusting the aggregate. " - "Use telemetry coverage and outcome sampling before drawing process conclusions." - ) - - -def main() -> int: - parser = argparse.ArgumentParser() - parser.add_argument("--days", type=int, default=7) - parser.add_argument("--dir", type=Path, default=DEFAULT_DIR) - parser.add_argument( - "--json", action="store_true", help="emit a structured summary" - ) - parser.add_argument( - "--by-process", - action="store_true", - help="include privacy-safe process_id rollups when events provide them", - ) - parser.add_argument( - "--input-signals", - action="store_true", - help="include aggregate redacted command-head/path-zone input signals", - ) - args = parser.parse_args() - - events = load_events(args.dir, args.days) - if not events: - if args.json: - print(json.dumps( - build_summary( - [], - args.days, - by_process=args.by_process, - input_signals=args.input_signals, - ), - indent=2, - )) - else: - print( - f"No workflow telemetry found in {args.dir} " - f"for the last {args.days} days." - ) - return 0 - - summary = build_summary( - events, - args.days, - by_process=args.by_process, - input_signals=args.input_signals, - ) - if args.json: - print(json.dumps(summary, indent=2)) - else: - print_summary(summary) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/examples/demo-mainframe/30_projects/example-project/README.md b/examples/demo-mainframe/30_projects/example-project/README.md new file mode 100644 index 0000000..431a297 --- /dev/null +++ b/examples/demo-mainframe/30_projects/example-project/README.md @@ -0,0 +1,17 @@ +--- +title: "Example Project" +domain: "knowledge-systems" +type: project +status: "active" +project_state: "active" +record_type: project +goal: "Synthetic bounded-outcome example for public MainFrame identity tests." +next_action: "Do not treat this as real work." +updated: "2026-09-10" +source: "synthetic" +tags: ["synthetic", "demo"] +--- + +# Example Project + +Synthetic project. Not a private MainFrame record. diff --git a/examples/demo-mainframe/40_operations/demo-eval/README.md b/examples/demo-mainframe/40_operations/demo-eval/README.md new file mode 100644 index 0000000..486e064 --- /dev/null +++ b/examples/demo-mainframe/40_operations/demo-eval/README.md @@ -0,0 +1,22 @@ +--- +title: "Demo Eval" +domain: "knowledge-systems" +type: operation +status: "active" +project_state: "active" +record_type: operation +work_kind: "evaluation_program" +portfolio_role: "maintenance" +authority_mode: "local_coordination" +wip_class: "eval" +goal: "Synthetic self-evaluation target used by public MPE feasibility tests." +next_action: "Keep outputs local-only." +updated: "2026-09-10" +source: "synthetic" +tags: ["synthetic", "demo"] +--- + +# Demo Eval + +Synthetic evaluation operation. Any generated output must remain local-only +and labelled synthetic. diff --git a/examples/demo-mainframe/40_operations/demo-eval/catalogue/entries.json b/examples/demo-mainframe/40_operations/demo-eval/catalogue/entries.json new file mode 100644 index 0000000..c458a28 --- /dev/null +++ b/examples/demo-mainframe/40_operations/demo-eval/catalogue/entries.json @@ -0,0 +1,57 @@ +{ + "schema_version": 1, + "synthetic": true, + "label": "SYNTHETIC — not a private MainFrame evaluation", + "entries": [ + { + "id": "loop-synthetic-capture", + "title": "Synthetic capture loop", + "tier": "well_defined", + "kind": "loop", + "question": "Does the synthetic capture loop declare a unit of work?", + "unit_of_work": "one labelled synthetic capture", + "entry": "synthetic inbox item exists", + "exit": "synthetic item is classified or rejected", + "status": "active", + "last_reviewed": "2026-09-10", + "score": 8, + "surfaces": {}, + "checklist": { + "unit_of_work": true, + "entry_exit": true, + "close_or_loop_decision": true, + "workflow_contract": true, + "skill_or_agent_procedure": false, + "bin_cli": true, + "dogfood_evidence": true, + "promotion_destinations": true, + "anti_scope": true, + "eval_or_action_surface": false + } + }, + { + "id": "loop-synthetic-thin", + "title": "Synthetic thin loop", + "tier": "underspecified", + "kind": "loop", + "question": "Is this synthetic loop still underspecified?", + "unit_of_work": "unclear", + "status": "draft", + "last_reviewed": "2026-09-10", + "score": 4, + "surfaces": {}, + "checklist": { + "unit_of_work": true, + "entry_exit": true, + "close_or_loop_decision": false, + "workflow_contract": true, + "skill_or_agent_procedure": false, + "bin_cli": false, + "dogfood_evidence": false, + "promotion_destinations": false, + "anti_scope": true, + "eval_or_action_surface": false + } + } + ] +} diff --git a/examples/demo-mainframe/40_operations/demo-eval/methodology/README.md b/examples/demo-mainframe/40_operations/demo-eval/methodology/README.md new file mode 100644 index 0000000..5866f37 --- /dev/null +++ b/examples/demo-mainframe/40_operations/demo-eval/methodology/README.md @@ -0,0 +1,3 @@ +# Synthetic methodology + +Observational only. Do not claim causal improvement from this demo. diff --git a/examples/demo-mainframe/40_operations/example-operation/README.md b/examples/demo-mainframe/40_operations/example-operation/README.md new file mode 100644 index 0000000..9f0ae29 --- /dev/null +++ b/examples/demo-mainframe/40_operations/example-operation/README.md @@ -0,0 +1,21 @@ +--- +title: "Example Operation" +domain: "knowledge-systems" +type: operation +status: "active" +project_state: "active" +record_type: operation +work_kind: "evaluation_program" +portfolio_role: "maintenance" +authority_mode: "local_coordination" +wip_class: "eval" +goal: "Synthetic standing-loop example for public MainFrame identity tests." +next_action: "Do not treat this as real work." +updated: "2026-09-10" +source: "synthetic" +tags: ["synthetic", "demo"] +--- + +# Example Operation + +Synthetic operation. Not a private MainFrame record. diff --git a/examples/demo-mainframe/README.md b/examples/demo-mainframe/README.md new file mode 100644 index 0000000..8e8e015 --- /dev/null +++ b/examples/demo-mainframe/README.md @@ -0,0 +1,18 @@ +# Synthetic demo MainFrame + +**This directory is synthetic.** It exists so identity and process-evaluation +tools can be exercised without anyone's private knowledge, projects, or +evaluation evidence. + +Do not treat files here as a claim that a real MainFrame passed an evaluation. + +```text +examples/demo-mainframe/ +├── 30_projects/example-project/README.md +└── 40_operations/ + ├── example-operation/README.md + └── demo-eval/README.md +``` + +Outputs written while trying these examples belong under a local-only +`outputs/` path and must stay untracked. diff --git a/mindgraph/.gitignore b/mindgraph/.gitignore deleted file mode 100644 index 9cbca18..0000000 --- a/mindgraph/.gitignore +++ /dev/null @@ -1,25 +0,0 @@ -# Python -__pycache__/ -*.py[cod] -*.egg-info/ -.pytest_cache/ -.mypy_cache/ -.ruff_cache/ - -# Virtual environments -.venv/ -venv/ - -# Local databases and build artifacts -*.sqlite -*.sqlite-journal -*.db -build/ - -# OS -.DS_Store - -# Internal-only project files (kept locally for ongoing work, not published) -AGENTS.md -CHANGELOG.md -planning/ diff --git a/mindgraph/LICENSE b/mindgraph/LICENSE deleted file mode 100644 index 911ba58..0000000 --- a/mindgraph/LICENSE +++ /dev/null @@ -1,21 +0,0 @@ -MIT License - -Copyright (c) 2026 camerontjs-dot - -Permission is hereby granted, free of charge, to any person obtaining a copy -of this software and associated documentation files (the "Software"), to deal -in the Software without restriction, including without limitation the rights -to use, copy, modify, merge, publish, distribute, sublicense, and/or sell -copies of the Software, and to permit persons to whom the Software is -furnished to do so, subject to the following conditions: - -The above copyright notice and this permission notice shall be included in all -copies or substantial portions of the Software. - -THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR -IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, -FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE -AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER -LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, -OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE -SOFTWARE. diff --git a/mindgraph/README.md b/mindgraph/README.md deleted file mode 100644 index 6b4f0da..0000000 --- a/mindgraph/README.md +++ /dev/null @@ -1,372 +0,0 @@ -# MindGraph - -MindGraph is a local, graph-augmented retrieval engine for personal Markdown knowledge bases. It ingests a directory of notes, extracts a typed `[[link]]` document graph, chunks the body text, embeds the chunks with a small CPU model, and stores everything in one SQLite file. The retrieval surface combines vector similarity, lexical search, and graph traversal over the same store. - -This is the engine I run against Mainframe, my own Markdown knowledge base. Any vault of Markdown notes with `[[wikilink]]` syntax (Obsidian, Foam, Logseq with the right setting) works the same way. - -## What this does - -- Parses Markdown files. Reads optional YAML frontmatter for a `title` and a `domain`. -- Splits each note into a Truth body and an optional Timeline section on a `---` rule followed by a `## Timeline` heading. -- Extracts `[[target]]` and `[[target]] (relationship)` links as typed graph edges, with link targets normalized to add `.md` when missing. -- Computes a stable document ID from `sha256(relative_path)` for ordinary single-root ingest, or from `index_id + namespace + source_path` for scoped multi-root ingest. -- Skips re-embedding when the content hash matches an existing row. -- Chunks the Truth body into paragraphs packed up to `max_chars`, keeping paragraphs whole. -- Embeds chunks with a selectable model (`--embedder`: `minilm`, `bge-small`, `e5-small`; default MiniLM) at 384 dimensions per DB. -- Optional `--embed-template mainframe` prefixes domain/type/title at ingest and `[intent=query]` at query time. -- Writes documents, chunks, embeddings, FTS5 rows, and edges to one SQLite file with `sqlite-vec` and FTS5 attached. - -## What it ranks - -`mindgraph query <text>` returns ranked chunks. Two signals run over the same SQLite store and fuse with Reciprocal Rank Fusion at the canonical `k = 60`: - -- Lexical: FTS5 BM25 over the chunked Truth text -- Semantic: `sqlite-vec` cosine over `vec_chunks` - -Each result carries a `signal` label (`lexical`, `semantic`, `fused`, `expanded`, or `associated`), a `rrf_score`, and the per-signal `lexical_rank` and `semantic_rank` integers, so the attribution is mechanically verifiable. Ties break by `(doc_id, chunk_index)` lexicographic for deterministic output. Free-text queries pass through an FTS5 sanitizer that strips operator characters and the uppercase keywords `AND`, `OR`, `NOT`, `NEAR`, then OR-joins the surviving tokens. - -Result rows also carry trust and provenance metadata for consumers that need to decide what to inspect next: `doc_type`, `domain`, `status`, `index_id`, `trust_profile`, `namespace`, `source_root`, `source_path`, `display_path`, `semantic_distance`, `weak_fit`, and `query_scope_warning`. `weak_fit` marks semantic-only rows beyond the current distance threshold. `query_scope_warning` appears when the query itself seems to ask for inbox, live/current, or project-status state that may belong in a different lifecycle database. - -`mindgraph query --expand` appends graph-walk results. After the fused list returns, the query walks outbound `[[link]]` edges from each fused result to a bounded depth (`--depth N`, default 1, cap 3) and appends walked documents with `signal = "expanded"` and an `expansion_depth` integer. The walk is outbound only, deduplicates against the fused set, terminates at dangling edges, and does not interact with the RRF math. `--expand-top-k N` (default 20) caps appended expanded rows. - -`mindgraph query --associate` appends semantic doc-neighbor results (ADR-034). From fused seeds (not expand results), embeds title + chunk excerpt per seed, runs vec kNN, and appends rows with `signal = "associated"`, `association_depth = 1`, and `semantic_distance`. `--associate-top-k` and `--associate-seed-k` cap output and seed count. - -`mindgraph neighbors <doc_id>` lists outbound edges for a document, including dangling edges (links to files that do not exist as documents). - -## What this does not do yet - -- There is no LLM generation step. MindGraph retrieves and ranks. It does not write summaries, answers, or explanations. -- Renaming a file produces a new document ID. The old document remains in the database until a future cleanup pass prunes orphans. -- Only Markdown is a first-class input. PDFs and other formats are out of scope for this asset. -- Use a **separate SQLite file per embedder** (`--embedder` sets `vec_chunks` dimensions at init). Re-ingest the full scope into each eval DB; do not swap models in-place on one DB. -- Lexical-only results surface chunk index 0 because there is no semantic ranking to pick a better chunk from. Fused and semantic results surface the best-ranked chunk per document. -- Retrieval is nomination, not verification. A retrieved chunk is a candidate for a reader to read, not a verified source for any claim. - -## Useful commands - -```bash -mindgraph init --db mindgraph.sqlite -mindgraph ingest path/to/your/vault --db mindgraph.sqlite -mindgraph ingest path/to/your/vault --db mindgraph.sqlite --embedder bge-small --embed-template mainframe -mindgraph ingest path/to/your/vault --db mindgraph.sqlite --verbose -mindgraph ingest path/to/your/vault --db mindgraph.sqlite --index-id mainframe-knowledge --trust-profile durable_knowledge --namespace knowledge --display-prefix 10_knowledge -mindgraph ingest-many path/to/manifest.json --db mindgraph.sqlite -mindgraph query "what does this vault say about X" --db mindgraph.sqlite -mindgraph query "..." --db mindgraph.sqlite --top-k 5 --json -mindgraph query "..." --db mindgraph.sqlite --top-k 5 --json --envelope -mindgraph query "..." --db mindgraph.sqlite --expand -mindgraph query "..." --db mindgraph.sqlite --expand --depth 2 --expand-top-k 10 -mindgraph query "..." --db mindgraph.sqlite --associate --associate-top-k 10 -mindgraph query "..." --db mindgraph.sqlite --embedder e5-small --embed-template mainframe -mindgraph neighbors <doc_id> --db mindgraph.sqlite -mindgraph neighbors <doc_id> --db mindgraph.sqlite --json -mindgraph serve-mcp --db mindgraph.sqlite -mindgraph serve-mcp --db mindgraph.sqlite --verbose -``` - -The commands above are the compatibility path: one database, stdio transport, -and legacy list-shaped query/neighbor JSON by default. - -### Optional shared daemon - -The opt-in shared server loads one embedder and opens the durable and project -indexes read-only. It binds only to loopback and requires every tool call to -select `knowledge` (`durable_knowledge`) or `projects` (`project_status`). Its -responses expose the selected scope and trust profile and never blend stores. - -```bash -mindgraph daemon-start -mindgraph daemon-status -mindgraph daemon-health -mindgraph mcp-proxy --url http://127.0.0.1:8000/mcp -mindgraph daemon-stop -``` - -The default remains persistent and does not auto-start or idle-exit. Explicit -idle mode (`serve-daemon --idle-seconds N`) exits only after the grace elapses -with no in-flight tool call and no unexpired client lease. A proxy started with -`mcp-proxy --auto-start` serializes startup under a state-directory lock, waits -for health, and holds a renewable lease for the stdio client's lifetime. - -The installed MainFrame LaunchAgent currently uses `KeepAlive=true` and -`RunAtLoad=true`; that supervision would immediately undo an idle exit. Do not -enable idle mode without disabling that job in a verified idle window. From -MainFrame root, the single explicit activation step is: - -```bash -bin/mindgraph-idle-lifecycle activate --confirm-no-active-clients 900 -``` - -The command backs up and unloads the job, renames its plist to `.disabled`, and -writes the marker that makes existing `mcp-proxy` clients use locked auto-start -plus leases. It is never run automatically. Roll back with the exact backup -path printed by activation: - -```bash -bin/mindgraph-idle-lifecycle rollback /exact/plist.backup-path -``` - -Direct Streamable HTTP clients cannot start a process at a dead URL. They must -run `bin/mindgraph daemon-start --idle-seconds 900` before connecting and may -lose an idle session after the grace unless they explicitly maintain a lease at -`/lifecycle/lease`. Transparent direct-HTTP wakeup needs a separate always-on -proxy or launchd socket-activation design and is not claimed here. Refresh does -not hot-reload a running daemon. Authentication, remote exposure, socket -activation, concurrency guarantees, and RAM/latency claims remain unestablished. - -First ingest on a fresh machine downloads and caches the embedding model. First query also loads the model to embed the query text, and logs the same line. Subsequent runs reuse the cached model. - -## Try it - -A small seven-file Markdown vault under `examples/example-vault/` exercises every retrieval path the engine exposes: lexical-only matches, semantic-only matches, fused matches, dangling graph edges, and graph expansion. Run the sequence below from the asset root after `pip install -e .` to reproduce the captures. The captures regenerate from `scripts/run_example_smoke.py` against a fresh `/tmp/mindgraph-example/db.sqlite`. - -### Initialize the database - -``` -$ mindgraph init --db /tmp/mindgraph-example/db.sqlite -INFO mindgraph | Initialized database at /tmp/mindgraph-example/db.sqlite -``` - -### Ingest the example vault - -The first ingest downloads the MiniLM model and logs one canonical line so the first-run latency is visible. - -``` -$ mindgraph ingest examples/example-vault --db /tmp/mindgraph-example/db.sqlite -INFO mindgraph | Loading embedding model (all-MiniLM-L6-v2)... -INFO mindgraph | ingested: balancing-loops.md (1 chunks, 1 edges) -INFO mindgraph | ingested: bounded-rationality.md (1 chunks, 1 edges) -INFO mindgraph | ingested: feedback-loops.md (1 chunks, 2 edges) -INFO mindgraph | ingested: mental-models-overview.md (1 chunks, 1 edges) -INFO mindgraph | ingested: reinforcing-loops.md (1 chunks, 1 edges) -INFO mindgraph | ingested: systems-archetypes.md (1 chunks, 1 edges) -INFO mindgraph | ingested: unrelated-noise.md (1 chunks, 0 edges) -INFO mindgraph | Done. total=7 ingested=7 skipped=0 failed=0 -``` - -### Query a unique keyword - -`antinet` appears in only one note (`mental-models-overview.md`). It lands at rank 1 as a fused hit because the keyword also pulls the doc on the semantic side. The second and third results carry `signal=semantic` because they have no lexical match. - -``` -$ mindgraph query antinet --db /tmp/mindgraph-example/db.sqlite --top-k 3 -#1 signal=fused rrf_score=0.032787 lex_rank=1 sem_rank=1 - path: mental-models-overview.md - title: Mental models overview - chunk_index: 0 - excerpt: I keep a running file of mental models I have found useful, separate from the structural systems-thinking notes... -#2 signal=semantic rrf_score=0.016129 lex_rank=- sem_rank=2 - path: feedback-loops.md - title: Feedback loops - chunk_index: 0 - excerpt: A feedback loop is a circular causal structure where the output of a process becomes part of its own input on a later pass... -#3 signal=semantic rrf_score=0.015873 lex_rank=- sem_rank=3 - path: bounded-rationality.md - title: Bounded rationality - chunk_index: 0 - excerpt: Herbert Simon coined this term to describe how people decide when full information and unlimited compute are not available... -``` - -### Query a concept term - -`satisficing` is unique to `bounded-rationality.md`. The signal attribution shows the same pattern: a fused top hit plus two semantic-only neighbors that the embedder pulled in by topical similarity. - -``` -$ mindgraph query satisficing --db /tmp/mindgraph-example/db.sqlite --top-k 3 -#1 signal=fused rrf_score=0.032787 lex_rank=1 sem_rank=1 - path: bounded-rationality.md - title: Bounded rationality - chunk_index: 0 - excerpt: Herbert Simon coined this term to describe how people decide when full information and unlimited compute are not available... -#2 signal=semantic rrf_score=0.016129 lex_rank=- sem_rank=2 - path: mental-models-overview.md - title: Mental models overview - chunk_index: 0 - excerpt: I keep a running file of mental models I have found useful, separate from the structural systems-thinking notes... -#3 signal=semantic rrf_score=0.015873 lex_rank=- sem_rank=3 - path: balancing-loops.md - title: Balancing loops - chunk_index: 0 - excerpt: A balancing loop is a feedback structure that pushes a system back toward a setpoint... -``` - -### Query that fuses lexical and semantic signals - -`balancing feedback` matches multiple notes both ways. The top three all carry `signal=fused` with different `lex_rank` and `sem_rank` values, so the attribution is mechanically checkable. - -``` -$ mindgraph query "balancing feedback" --db /tmp/mindgraph-example/db.sqlite --top-k 3 -#1 signal=fused rrf_score=0.032522 lex_rank=1 sem_rank=2 - path: feedback-loops.md - title: Feedback loops - chunk_index: 0 - excerpt: A feedback loop is a circular causal structure where the output of a process becomes part of its own input on a later pass... -#2 signal=fused rrf_score=0.032522 lex_rank=2 sem_rank=1 - path: balancing-loops.md - title: Balancing loops - chunk_index: 0 - excerpt: A balancing loop is a feedback structure that pushes a system back toward a setpoint... -#3 signal=fused rrf_score=0.031258 lex_rank=5 sem_rank=3 - path: systems-archetypes.md - title: Systems archetypes - chunk_index: 0 - excerpt: A systems archetype is a recurring pattern of feedback structure that shows up across very different domains... -``` - -### Walk the graph from a seed - -With `--top-k 1` scoping the fused step to just the seed, `--expand --depth 2` walks the outbound `[[link]]` edges from `feedback-loops.md`. Depth-1 hits are the two `(illustrates)` targets; the depth-2 hit is `systems-archetypes.md`, reached through both one-hop docs and deduplicated. Walked rows carry `signal=expanded` and a `depth=N` suffix, with `lex_rank` and `sem_rank` blank because expansion is a labeled append, not a rerank. - -``` -$ mindgraph query "feedback loops" --db /tmp/mindgraph-example/db.sqlite --top-k 1 --expand --depth 2 -#1 signal=fused rrf_score=0.032787 lex_rank=1 sem_rank=1 - path: feedback-loops.md - title: Feedback loops - chunk_index: 0 - excerpt: A feedback loop is a circular causal structure where the output of a process becomes part of its own input on a later pass... -#2 signal=expanded rrf_score=0.000000 lex_rank=- sem_rank=- depth=1 - path: reinforcing-loops.md - title: Reinforcing loops - chunk_index: 0 - excerpt: A reinforcing loop is a feedback structure where a change in one direction produces more change in the same direction on the next cycle... -#3 signal=expanded rrf_score=0.000000 lex_rank=- sem_rank=- depth=1 - path: balancing-loops.md - title: Balancing loops - chunk_index: 0 - excerpt: A balancing loop is a feedback structure that pushes a system back toward a setpoint... -#4 signal=expanded rrf_score=0.000000 lex_rank=- sem_rank=- depth=2 - path: systems-archetypes.md - title: Systems archetypes - chunk_index: 0 - excerpt: A systems archetype is a recurring pattern of feedback structure that shows up across very different domains... -``` - -### Inspect outbound edges and surface a dangling target - -`systems-archetypes.md` links to `[[unicycle-mental-model]] (cites)`, but the target file does not exist in the vault. `mindgraph neighbors` returns the row with `target_path: (dangling)` so broken links surface during traversal rather than getting silently dropped. The `doc_id` argument is the hex ID from any earlier `--json` query output for the source document. - -``` -$ mindgraph neighbors c8a1be119b7ad0c3 --db /tmp/mindgraph-example/db.sqlite -#1 -> 885be168decf4005 rel=cites - target_path: (dangling) -``` - -## MCP - -MindGraph ships MCP transports around the same retrieval code used by the CLI. -It does not add ranking behavior, change the database schema, or turn retrieved -chunks into verified claims. - -### Operational default: shared daemon + proxy (MainFrame) - -One loopback daemon owns both indexes so agent sessions do not each load MiniLM: - -```bash -bin/mindgraph daemon-start # or LaunchAgent com.user.mindgraph-daemon -bin/mindgraph daemon-health # expect knowledge + projects scopes -# stdio clients: -bin/mindgraph mcp-proxy --url http://127.0.0.1:8000/mcp -# streamable-HTTP clients (e.g. Codex): -# url = "http://127.0.0.1:8000/mcp" -``` - -Shared tools require explicit `scope` of `knowledge` or `projects` (no blend). -Root `.mcp.json` points at `mcp-proxy`. Do not configure daily clients on -`serve-mcp` — that path loads a full embedder per process. - -### Legacy single-DB stdio server - -For isolated debugging of one database: - -```bash -.venv/bin/mindgraph init --db /tmp/mindgraph-mcp/db.sqlite -.venv/bin/mindgraph ingest examples/example-vault --db /tmp/mindgraph-mcp/db.sqlite -.venv/bin/mindgraph serve-mcp --db /tmp/mindgraph-mcp/db.sqlite --verbose -``` - -Stdout is reserved for MCP protocol frames; logs go to stderr. - -### Client config (stdio proxy) - -```json -{ - "mcpServers": { - "mindgraph": { - "command": "/absolute/path/to/MainFrame/bin/mindgraph", - "args": ["mcp-proxy", "--url", "http://127.0.0.1:8000/mcp"] - } - } -} -``` - -Claude Code, Claude Desktop, Cursor, Grok, Antigravity, Cline, and other -stdio-MCP clients can share this shape. Codex may use the HTTP URL directly. - -#### Idle lifecycle safety contract - -- Idle time starts after the last completed tool request or lease activity. -- In-flight `query` and `graph_neighbors` calls block shutdown. -- Renewable leases protect opt-in stdio proxy sessions; expired leases stop - protecting crashed clients. -- A standard direct HTTP session is not a lease. Direct clients either tolerate - reconnect/manual start or explicitly use the lease endpoint. Health probes do - not reset idle time. -- `daemon-status` combines PID-file and health evidence. `daemon-stop` refuses - a mismatched or identity-less healthy listener rather than signaling an - ambiguous process. - -### Tools - -`query` runs the same retrieval path as `mindgraph query --json`. Parameters: `question`, `lexical_top_k`, `semantic_top_k`, `final_top_k`, `expand`, `expand_depth`, `expand_top_k`, `associate`, `associate_top_k`, `associate_seed_k`, and `envelope`. By default, the MCP response content is a JSON array of `QueryResult` records. With `envelope=true` (or CLI `--json --envelope`), it returns an object containing `schema_version`, `intent_resolution`, `routing`, and `results`; legacy list output remains unchanged when the flag is omitted. `routing` is single-database metadata for the bound index (not multi-index federation). In a smoke run against the example vault, the default list path matched the CLI JSON output exactly: - -```json -{ - "question": "feedback loops", - "final_top_k": 1, - "expand": true, - "expand_depth": 2 -} -``` - -Observed result paths from that smoke were `feedback-loops.md`, `reinforcing-loops.md`, `balancing-loops.md`, and `systems-archetypes.md`, with `expansion_depth` values `[0, 1, 1, 2]`. - -`graph_neighbors` runs the same lookup as `mindgraph neighbors --json`. Parameter: `doc_id`. The MCP response content is a JSON array of `NeighborResult` records, including dangling edges with `target_path = null`. In the same smoke, calling `graph_neighbors` with `doc_id = "c8a1be119b7ad0c3"` returned the single dangling edge the CLI lookup also returns. - -### claude.ai web - -The stdio transport does not connect directly to claude.ai web. The web product requires a remote transport such as Streamable HTTP or SSE over HTTPS. Exposing a local SQLite-backed knowledge base over the public internet would contradict MindGraph's local-first framing, so a remote transport is out of scope for now. - -## Architecture - -MindGraph runs entirely locally. There is no service to start and no remote dependency at retrieval time. - -1. **Ingestion.** `parser.parse_document` reads YAML frontmatter, splits Truth from Timeline, and returns a `ParsedDocument`. `parser.extract_document_graph_edges` collects `GraphEdge` records from both frontmatter `links:` and body `[[wikilinks]]` (ADR-033). `LinkResolver` resolves unique canonical trailing slugs (`…__slug`) as well as full stems and titles. `parser.chunk_truth` packs paragraphs into bounded chunks. -2. **Storage.** `db.init_db` creates the schema: `documents`, `documents_fts` (FTS5 over title and Truth content), `chunks`, `vec_chunks` (sqlite-vec, 384-dim), and `edges`. Foreign keys are on. -3. **Re-ingest.** `db.get_document_hash` compares the stored hash to the freshly computed one. Unchanged files exit before parsing or embedding. -4. **Retrieval.** `query.fetch_lexical_ranking` runs FTS5 BM25 over `documents_fts`. `query.fetch_semantic_ranking` runs a `sqlite-vec` KNN over `vec_chunks` and promotes to document granularity by keeping the best chunk per document. `query.rrf_fuse` combines the two ranked lists at the canonical `k = 60` and returns deterministic, signal-attributed results. -5. **Graph lookup.** `query.list_neighbors` resolves edges from a source document against the `documents` table, preserving dangling edges with `null` resolved paths. -6. **Graph expansion.** `query.expand_results` walks outbound edges from the fused result set in a deterministic BFS, deduplicating against the seed doc_ids and skipping dangling targets. Walked documents are appended with `signal = "expanded"`, `expansion_depth` set to the walk distance, and sorted by `(expansion_depth, doc_id, chunk_index)` before the `--expand-top-k` cut. - -## Test discipline - -The parser test suite (`tests/test_parser.py`) covers frontmatter parsing, the page-model split rule, internal `---` rules that must not split, the link extraction regex including nested-bracket rejection, chunk packing, and end-to-end `parse_document`. The ingest test suite (`tests/test_ingest.py`) covers the ingest happy path against a fixture vault. - -The query test suite (`tests/test_query.py`) covers FTS5 input sanitization, RRF fusion math, lexical and semantic ranking against a small ingested vault, end-to-end signal attribution, and the CLI surface. It uses a deterministic `KeywordEmbedder` stub so semantic similarity is reproducible without depending on the real MiniLM model. - -The neighbors test suite (`tests/test_neighbors.py`) covers outbound edge resolution, dangling edges, multiple relationship types per edge, sources with no edges, and the CLI surface. - -The expand test suite (`tests/test_expand.py`) covers one-hop walk, two-hop walk, no-expand and depth-0 equivalence to the un-expanded path, dedup of walked targets already in the fused set, the `--expand-top-k` cap, dangling-edge termination, and unreachable-doc absence. The CLI tests also assert the `expansion_depth` field in JSON output and the hard `--depth` cap of 3. - -The examples test suite (`tests/test_examples.py`) ingests the committed `examples/example-vault/` into a temp database with the same deterministic `KeywordEmbedder` and asserts each retrieval path: the unique lexical keyword lands on the expected doc, the semantic-only synonym path lands on `bounded-rationality.md` with `signal=semantic`, the fused query lands on `balancing-loops.md`, the dangling edge from `systems-archetypes.md` is preserved by `list_neighbors`, and the depth-2 expansion topology reaches the expected docs and skips the dangling target. - -The MCP test suite (`tests/test_mcp.py`) uses the official Python MCP SDK's in-memory client plus the deterministic `KeywordEmbedder` stub. It asserts server startup, missing-DB startup failure, `query` and `graph_neighbors` output shape parity with CLI JSON, opt-in envelope metadata, expansion parameter routing, dangling-edge preservation, and a clean tool error for unknown `doc_id` values. - -Run the full suite from the asset root: - -```bash -.venv/bin/python -m pytest -``` - -## Design notes - -`DECISIONS.md` is the architectural decision log. Each entry records what was decided, what was rejected, and why, so the reasoning is recoverable without reading the code or chasing planning files. diff --git a/mindgraph/examples/README.md b/mindgraph/examples/README.md deleted file mode 100644 index 07126db..0000000 --- a/mindgraph/examples/README.md +++ /dev/null @@ -1,5 +0,0 @@ -# MindGraph example vault - -A seven-file Markdown vault under `example-vault/` that exercises every retrieval path MindGraph exposes: lexical-only matches, semantic-only matches, fused matches, dangling graph edges, and graph expansion from a seed document. The walkthrough that ingests this vault and shows the captured output for each path lives in the asset README's "Try it" section. - -The CI-safe smoke test under `../tests/test_examples.py` asserts the expected retrieval surface against this same vault using the deterministic `KeywordEmbedder` stub. The real-model captures in the README come from `../scripts/run_example_smoke.py`, which uses `sentence-transformers/all-MiniLM-L6-v2`. diff --git a/mindgraph/examples/example-vault/balancing-loops.md b/mindgraph/examples/example-vault/balancing-loops.md deleted file mode 100644 index 4544840..0000000 --- a/mindgraph/examples/example-vault/balancing-loops.md +++ /dev/null @@ -1,10 +0,0 @@ ---- -title: Balancing loops -domain: systems-thinking ---- - -A balancing loop is a feedback structure that pushes a system back toward a setpoint. A thermostat is the classic example: the room cools, the heater turns on, the room warms, the heater turns off. Predator-prey populations behave the same way at a longer time scale. - -The visible signature of a balancing loop is oscillation around the setpoint with delay. The delay matters. A balancing loop with no delay produces smooth correction. A balancing loop with a long delay can swing wildly, overshoot, and look like instability before it settles. - -Balancing loops paired with reinforcing loops produce most of the [[systems-archetypes]] (related) catalogued in the literature. diff --git a/mindgraph/examples/example-vault/bounded-rationality.md b/mindgraph/examples/example-vault/bounded-rationality.md deleted file mode 100644 index a890f5d..0000000 --- a/mindgraph/examples/example-vault/bounded-rationality.md +++ /dev/null @@ -1,10 +0,0 @@ ---- -title: Bounded rationality -domain: decision-science ---- - -Herbert Simon coined this term to describe how people decide when full information and unlimited compute are not available. Rather than searching for an optimal choice, decision-makers tend to pick the first option that clears an acceptance threshold. Simon called the strategy satisficing, a portmanteau of satisfying and sufficing. - -The behavior shows up across many domains. Where the textbook expects an optimizer, bounded rationality predicts a satisficer who stops as soon as good enough is found. Later work by Kahneman and Tversky on cognitive shortcuts built on the same observation from a different angle. - -A satisficer is not irrational. They are economizing on a resource the optimizer pretends is free. The connection to [[feedback-loops]] (related) is that real decision-makers operate inside a balancing loop where search cost grows with each candidate evaluated. diff --git a/mindgraph/examples/example-vault/feedback-loops.md b/mindgraph/examples/example-vault/feedback-loops.md deleted file mode 100644 index 58dcccf..0000000 --- a/mindgraph/examples/example-vault/feedback-loops.md +++ /dev/null @@ -1,10 +0,0 @@ ---- -title: Feedback loops -domain: systems-thinking ---- - -A feedback loop is a circular causal structure where the output of a process becomes part of its own input on a later pass. The two basic shapes are reinforcing loops, which amplify a change in the same direction, and balancing loops, which push the system back toward a setpoint. - -Most real systems contain both kinds in interaction. A thermostat is a balancing loop on temperature. A viral product launch is a reinforcing loop on adoption. The interesting behavior usually comes from how these loops layer. - -See [[reinforcing-loops]] (illustrates) and [[balancing-loops]] (illustrates) for the two basic patterns. diff --git a/mindgraph/examples/example-vault/mental-models-overview.md b/mindgraph/examples/example-vault/mental-models-overview.md deleted file mode 100644 index 4af918a..0000000 --- a/mindgraph/examples/example-vault/mental-models-overview.md +++ /dev/null @@ -1,10 +0,0 @@ ---- -title: Mental models overview -domain: epistemics ---- - -I keep a running file of mental models I have found useful, separate from the structural systems-thinking notes. The two categories overlap but the framings are different. A mental model is a portable lens. A systems pattern is a structural claim about how parts connect. - -The current shortlist includes inversion (Charlie Munger's "always invert"), the map-territory distinction, Chesterton's fence, and the Antinet method for note-taking that Scott Scheper revived from Niklas Luhmann's working style. Each one has paid for itself more than once. - -The Antinet entry is the longest because I am still figuring out how it interacts with [[feedback-loops]] (related) in my own thinking workflow. diff --git a/mindgraph/examples/example-vault/reinforcing-loops.md b/mindgraph/examples/example-vault/reinforcing-loops.md deleted file mode 100644 index 2480a19..0000000 --- a/mindgraph/examples/example-vault/reinforcing-loops.md +++ /dev/null @@ -1,10 +0,0 @@ ---- -title: Reinforcing loops -domain: systems-thinking ---- - -A reinforcing loop is a feedback structure where a change in one direction produces more change in the same direction on the next cycle. Compound interest is the textbook example. So is word-of-mouth growth, the rich-get-richer dynamic, and a fire that spreads faster as the heat dries out the fuel ahead of it. - -Reinforcing loops are not inherently good or bad. They are the engine behind virtuous cycles and vicious spirals alike. What matters is direction and what eventually stops them. Most reinforcing loops are bounded by a balancing loop somewhere downstream, even if the bound is far enough away to feel unlimited inside the run. - -A reinforcing loop in isolation is one of the [[systems-archetypes]] (related) that recur across domains. diff --git a/mindgraph/examples/example-vault/systems-archetypes.md b/mindgraph/examples/example-vault/systems-archetypes.md deleted file mode 100644 index 2b55307..0000000 --- a/mindgraph/examples/example-vault/systems-archetypes.md +++ /dev/null @@ -1,10 +0,0 @@ ---- -title: Systems archetypes -domain: systems-thinking ---- - -A systems archetype is a recurring pattern of feedback structure that shows up across very different domains. The catalogue includes shifting the burden, limits to growth, tragedy of the commons, escalation, success to the successful, and fixes that fail. Each archetype has a characteristic causal-loop diagram and a characteristic failure mode. - -The point of learning archetypes is not to label every system as one of them. It is to recognize the structural shape early enough to intervene before the failure mode plays out. Most archetypes have a documented high-leverage intervention point that beats fighting the symptoms. - -For the underlying mental model behind why these patterns generalize, see [[unicycle-mental-model]] (cites). diff --git a/mindgraph/examples/example-vault/unrelated-noise.md b/mindgraph/examples/example-vault/unrelated-noise.md deleted file mode 100644 index 7a2f4e1..0000000 --- a/mindgraph/examples/example-vault/unrelated-noise.md +++ /dev/null @@ -1,10 +0,0 @@ ---- -title: Saturday morning pancakes -domain: cooking ---- - -My current pancake recipe is one cup of flour, one cup of buttermilk, one egg, a tablespoon of melted butter, a teaspoon of baking powder, a half-teaspoon of baking soda, and a pinch of salt. Mix the wet and dry separately, then fold together until just combined. Lumps are fine. Lumps are good. - -Let the batter rest for at least ten minutes before cooking. This is the non-negotiable step. Cooked straight away the pancakes are dense; rested for ten minutes they puff up cleanly. I have no idea why the rest matters that much but it does. - -A medium pan on medium-low heat, a film of butter, and patience. Flip when the bubbles in the top stop being replaced. diff --git a/mindgraph/notes/2026-06-19-retrieval-embedder-session.md b/mindgraph/notes/2026-06-19-retrieval-embedder-session.md deleted file mode 100644 index ba454e0..0000000 --- a/mindgraph/notes/2026-06-19-retrieval-embedder-session.md +++ /dev/null @@ -1,91 +0,0 @@ -# Session notes — 2026-06-19 retrieval & embedder research - -Operator session spanning MindGraph graph fixes, retrieval architecture ADRs, embedder literature, and vault ingest. - -## Shipped (tracked in git) - -### ADR-033 — Dual-channel graph + slug resolution - -- `extract_document_graph_edges()` merges frontmatter `links:` and body wikilinks; dedupes on `target_id`. -- `LinkResolver` resolves unique trailing slugs from canonical filename stems. -- Tests in `mindgraph/tests/test_parser.py`; ingest wired in `cli.py`. -- Authoring contract in `.context/primitives.md`; ingest-source skill updated (`.agents/` + `.claude/` mirror). - -**Post-ship:** `bin/mindgraph-refresh` run — 594 docs, edge count increased (903 edges vs 854 pre-ADR-033 on prior refresh). - -### ADR-034 / ADR-035 — Retrieval model direction - -- **ADR-034:** Query-time semantic association (`--associate`) — fourth signal, append-only. -- **ADR-035:** Keep hybrid FTS + chunk semantic + explicit graph; embedder changes gated on eval. - -Planning lives in `30_projects/mindgraph/plans/` (local, gitignored). - -## Vault work (local, gitignored) - -### Embedder source literature - -Eight peer-reviewed stubs ingested to `10_knowledge/graph-memory/`: - -| Slug | arXiv | -|------|-------| -| mteb-massive-text-embedding-benchmark | 2210.07316 | -| sentence-bert-sentence-embeddings | 1908.10084 | -| beir-zero-shot-retrieval-benchmark | 2104.08663 | -| e5-text-embeddings | 2212.03533 | -| bge-c-pack-embeddings | 2309.07597 | -| refine-scarce-data-embedding-finetune | 2410.12890 | -| clp-finetuning-embeddings | 2412.17364 | -| sparse-meets-dense-hybrid-retrieval | 2401.04055 | - -### Synthesis - -`10_knowledge/graph-memory/2026-06-19__graph-memory__note__text-embedding-models-for-mindgraph.md` - -**Working conclusions:** - -1. Hybrid FTS + dense + explicit graph is literature-supported — tune the dense channel, don't swap architecture. -2. MTEB/BEIR inform model choice; MainFrame probes (q01–q18) are the promotion gate. -3. MainFrame-native embedder = fine-tune/template-tune (REFINE/CLP), not train-from-scratch (~584 docs). -4. Recommended path: template prefixes → Phase A bake-off → ensemble or swap → Phase C fine-tune. - -Run note: `embedder-retrieval-literature-run`. Eval runbooks updated in `30_projects/mindgraph-eval/plans/`. - -## Shipped (2026-06-19 continuation) - -### Track 1 — `--embedder` + template prefixes -- `mindgraph/embedders.py` registry (minilm, bge-small, e5-small). -- `--embedder` on init/ingest/query/serve-mcp; `--embed-template mainframe`. -- Per-DB `embedding_dims` in `index_meta`; parallel DBs for bake-off. - -### Track 2 — Phase 9 `--associate` (ADR-034) -- `associate_results()` in `query.py`; CLI/MCP flags shipped. -- `tests/test_associate.py`; 154 tests pass. - -### Track 3 — Embedder bake-off automation -- `30_projects/mindgraph-eval/scripts/embedder_bakeoff.py`. - -### Track 4 — Research resolutions -- `plans/research-tracks/2026-06-19-vocabulary-vs-structure.md` -- `plans/research-tracks/2026-06-19-finetune-timing.md` - -### Track 5 — Eval probes + graph baseline -- `association-probes.yaml` for Phase 9 eval. -- Graph baseline `2026-06-19-post-adr033-graph-baseline` (952 edges, 594 docs). -- Link cultivation on federation + embedder synthesis notes. - -## Open / next - -| Item | Blocker | -|------|---------| -| Parallel BGE/E5 DB ingest | Operator GPU/time for model download | -| Phase A bake-off run | Parallel DBs | -| Phase B ensemble | Multi-vec storage ADR | -| Association eval pass | Run `association-probes.yaml` matrix | - -## Discussion threads offered - -1. Unblock bake-off (`--embedder` flag) -2. Template prefix experiment (no weight change) -3. Dimension/schema for BGE non-384d -4. Fine-tune timing vs Phase 9 association -5. Vocabulary vs structure bottleneck \ No newline at end of file diff --git a/mindgraph/pyproject.toml b/mindgraph/pyproject.toml deleted file mode 100644 index ce478fd..0000000 --- a/mindgraph/pyproject.toml +++ /dev/null @@ -1,32 +0,0 @@ -[project] -name = "mindgraph" -version = "0.2.0" -description = "A Graph-Augmented Personal Knowledge Engine" -readme = "README.md" -license = "MIT" -license-files = ["LICENSE"] -requires-python = ">=3.10" -dependencies = [ - "typer>=0.12.0", - "sqlite-vec>=0.1.6", - "sentence-transformers>=3.0.0", - "pyyaml>=6.0.1", - "pydantic>=2.0.0", - "mcp>=1.0.0", -] - -[project.optional-dependencies] -dev = [ - "pytest>=8.0.0", -] - -[project.scripts] -mindgraph = "mindgraph.cli:app" - -[build-system] -requires = ["hatchling"] -build-backend = "hatchling.build" - -[tool.pytest.ini_options] -pythonpath = ["src"] -testpaths = ["tests"] diff --git a/mindgraph/scripts/run_example_smoke.py b/mindgraph/scripts/run_example_smoke.py deleted file mode 100644 index 55d50b9..0000000 --- a/mindgraph/scripts/run_example_smoke.py +++ /dev/null @@ -1,267 +0,0 @@ -"""Regenerates the asset README's "Try it" captures. - -Runs the handoff prompt's smoke sequence against the committed -examples/example-vault/ vault, using the real MiniLM model. Prints fenced -Markdown blocks formatted for direct paste into the asset README's "Try it" -section. - -First-run latency note: the first ingest downloads the MiniLM model on a fresh -machine; the script keeps the `Loading embedding model (all-MiniLM-L6-v2)...` -log line on the first block so the cost is visible. Subsequent blocks filter -the log line to reduce noise. - -Usage (from the asset root): - - .venv/bin/python scripts/run_example_smoke.py > /tmp/captures.md - -Inspect `/tmp/captures.md` and paste blocks into the README's "Try it" section. -""" - -import shutil -import subprocess -import sys -from pathlib import Path - -from mindgraph import parser - -ASSET_ROOT = Path(__file__).resolve().parent.parent -VAULT_PATH = ASSET_ROOT / "examples" / "example-vault" -DB_DIR = Path("/tmp/mindgraph-example") -DB_PATH = DB_DIR / "db.sqlite" -DB_PATH_STR = str(DB_PATH) - - -def reset_db_dir() -> None: - if DB_DIR.exists(): - shutil.rmtree(DB_DIR) - DB_DIR.mkdir(parents=True) - - -_NOISE_SUBSTRINGS = ( - "HTTP Request:", - "huggingface_hub.utils._http", - "sentence_transformers.base.model", - "Loading weights:", - "Batches:", - "Warning: You are sending unauthenticated", - "transformers_modules", -) - - -def _filter_noise(stderr: str, *, keep_loading_log: bool) -> str: - """Drop HuggingFace and sentence-transformers chatter. - - Keeps the canonical `Loading embedding model (all-MiniLM-L6-v2)...` line - when `keep_loading_log` is True so first-run latency stays visible in the - initial capture block. - """ - kept: list[str] = [] - for line in stderr.splitlines(): - stripped = line.strip() - if not stripped: - continue - if "Loading embedding model" in stripped: - if keep_loading_log: - kept.append(line) - continue - if any(noise in stripped for noise in _NOISE_SUBSTRINGS): - continue - kept.append(line) - return "\n".join(kept) - - -def run(cmd: list[str], *, keep_loading_log: bool = False) -> tuple[str, str]: - """Run a subprocess and capture stdout + stderr. - - Filters HuggingFace and sentence-transformers chatter from stderr so the - README captures stay readable. Keeps the canonical model-loading log line - when `keep_loading_log` is True. - """ - proc = subprocess.run(cmd, capture_output=True, text=True) - stdout = proc.stdout - stderr = _filter_noise(proc.stderr, keep_loading_log=keep_loading_log) - if proc.returncode != 0: - print( - f"WARNING: command exited with code {proc.returncode}: " - f"{' '.join(cmd)}", - file=sys.stderr, - ) - return stdout, stderr - - -def emit_block( - heading: str, display_cmd: list[str], stdout: str, stderr: str -) -> None: - """Print a README-ready subsection: heading + fenced code block.""" - print(f"### {heading}") - print() - print("```") - print("$ " + " ".join(display_cmd)) - if stderr.strip(): - print(stderr.rstrip()) - if stdout.strip(): - print(stdout.rstrip()) - print("```") - print() - - -def main() -> None: - reset_db_dir() - mindgraph_bin = str(ASSET_ROOT / ".venv" / "bin" / "mindgraph") - - # 1. init - out, err = run([mindgraph_bin, "init", "--db", DB_PATH_STR]) - emit_block( - "Initialize the database", - ["mindgraph", "init", "--db", DB_PATH_STR], - out, - err, - ) - - # 2. ingest — keep the loading log on the first model use for honesty - out, err = run( - [mindgraph_bin, "ingest", str(VAULT_PATH), "--db", DB_PATH_STR], - keep_loading_log=True, - ) - emit_block( - "Ingest the example vault", - [ - "mindgraph", - "ingest", - "examples/example-vault", - "--db", - DB_PATH_STR, - ], - out, - err, - ) - - # 3. lexical-heavy query (unique keyword 'antinet' in mental-models-overview) - out, err = run( - [ - mindgraph_bin, - "query", - "antinet", - "--db", - DB_PATH_STR, - "--top-k", - "3", - ] - ) - emit_block( - "Query a unique keyword", - [ - "mindgraph", - "query", - "antinet", - "--db", - DB_PATH_STR, - "--top-k", - "3", - ], - out, - err, - ) - - # 4. concept query (paraphrase-friendly term) - out, err = run( - [ - mindgraph_bin, - "query", - "satisficing", - "--db", - DB_PATH_STR, - "--top-k", - "3", - ] - ) - emit_block( - "Query a concept term", - [ - "mindgraph", - "query", - "satisficing", - "--db", - DB_PATH_STR, - "--top-k", - "3", - ], - out, - err, - ) - - # 5. fused query - out, err = run( - [ - mindgraph_bin, - "query", - "balancing feedback", - "--db", - DB_PATH_STR, - "--top-k", - "3", - ] - ) - emit_block( - "Query that fuses lexical and semantic signals", - [ - "mindgraph", - "query", - "balancing feedback", - "--db", - DB_PATH_STR, - "--top-k", - "3", - ], - out, - err, - ) - - # 6. expansion query — walk two hops from the seed - out, err = run( - [ - mindgraph_bin, - "query", - "feedback loops", - "--db", - DB_PATH_STR, - "--top-k", - "1", - "--expand", - "--depth", - "2", - ] - ) - emit_block( - "Walk the graph from a seed", - [ - "mindgraph", - "query", - "feedback loops", - "--db", - DB_PATH_STR, - "--top-k", - "1", - "--expand", - "--depth", - "2", - ], - out, - err, - ) - - # 7. neighbors lookup — surfaces the dangling unicycle-mental-model target - archetypes_id = parser.compute_doc_id("systems-archetypes.md") - out, err = run( - [mindgraph_bin, "neighbors", archetypes_id, "--db", DB_PATH_STR] - ) - emit_block( - "Inspect outbound edges and surface a dangling target", - ["mindgraph", "neighbors", archetypes_id, "--db", DB_PATH_STR], - out, - err, - ) - - -if __name__ == "__main__": - main() diff --git a/mindgraph/src/mindgraph/__init__.py b/mindgraph/src/mindgraph/__init__.py deleted file mode 100644 index c82131a..0000000 --- a/mindgraph/src/mindgraph/__init__.py +++ /dev/null @@ -1,6 +0,0 @@ -"""MindGraph: A Graph-Augmented Personal Knowledge Engine.""" - -__version__ = "0.2.0" - -from . import intent -from . import routing diff --git a/mindgraph/src/mindgraph/cli.py b/mindgraph/src/mindgraph/cli.py deleted file mode 100644 index 0021b8a..0000000 --- a/mindgraph/src/mindgraph/cli.py +++ /dev/null @@ -1,1197 +0,0 @@ -import json -import logging -import os -import sys -from dataclasses import dataclass, field -from fnmatch import fnmatch -from pathlib import Path - -import typer - -from mindgraph import daemon, db, embedders, idle_lifecycle, mcp_proxy, mcp_server, parser -from mindgraph import query as query_mod -from mindgraph.exceptions import EmbeddingError, IngestionError, MindgraphError -from mindgraph import intent as intent_mod -from mindgraph import routing as routing_mod - -app = typer.Typer( - name="mindgraph", - help="A Graph-Augmented Personal Knowledge Engine", - add_completion=False, -) - -logger = logging.getLogger("mindgraph") - - -_NOISY_LOGGERS = ( - "httpx", - "httpcore", - "huggingface_hub", - "huggingface_hub.utils._http", - "sentence_transformers", - "sentence_transformers.base.model", - "transformers", -) - - -def _configure_logging(verbose: bool) -> None: - os.environ.setdefault("HF_HUB_DISABLE_PROGRESS_BARS", "1") - os.environ.setdefault("TOKENIZERS_PARALLELISM", "false") - logging.basicConfig( - level=logging.WARNING, - format="%(asctime)s %(levelname)-7s %(name)s | %(message)s", - datefmt="%H:%M:%S", - ) - logger.setLevel(logging.DEBUG if verbose else logging.INFO) - for noisy_logger in _NOISY_LOGGERS: - logging.getLogger(noisy_logger).setLevel( - logging.WARNING if verbose else logging.ERROR - ) - - -def _load_embedder(embedder: str | None = None): - spec = embedders.resolve_embedder(embedder) - logger.info("Loading embedding model (%s)...", spec.model_id) - try: - return embedders.load_sentence_embedder(spec) - except EmbeddingError: - raise - except Exception as exc: - raise EmbeddingError( - f"Failed to load embedding model {spec.model_id!r}: " - f"{type(exc).__name__}: {exc}" - ) from exc - - -def _encode_without_progress(embedder, texts): - """Encode text while suppressing sentence-transformers progress output.""" - try: - try: - return embedder.encode( - texts, convert_to_numpy=True, show_progress_bar=False - ) - except TypeError: - return embedder.encode(texts, convert_to_numpy=True) - except Exception as exc: - raise EmbeddingError( - f"Embedding model failed to encode input: {type(exc).__name__}: {exc}" - ) from exc - - -@dataclass(frozen=True) -class IngestScope: - root: Path - index_id: str | None = None - trust_profile: str | None = None - namespace: str | None = None - source_root: Path | None = None - display_prefix: str | None = None - include_globs: tuple[str, ...] = field(default_factory=tuple) - exclude_globs: tuple[str, ...] = field(default_factory=tuple) - - -def _matches_any(path: str, patterns: tuple[str, ...]) -> bool: - return any(fnmatch(path, pattern) for pattern in patterns) - - -def _markdown_files_for_scope(scope: IngestScope) -> list[Path]: - all_md = sorted(scope.root.rglob("*.md")) - selected: list[Path] = [] - for md_file in all_md: - rel_path = md_file.relative_to(scope.root).as_posix() - if scope.include_globs and not _matches_any(rel_path, scope.include_globs): - continue - if scope.exclude_globs and _matches_any(rel_path, scope.exclude_globs): - continue - selected.append(md_file) - return selected - - -def _display_path(prefix: str | None, source_path: str) -> str: - if not prefix: - return source_path - return f"{prefix.rstrip('/')}/{source_path}" - - -def _apply_scope_provenance( - parsed: parser.ParsedDocument, scope: IngestScope -) -> parser.ParsedDocument: - has_provenance = any( - ( - scope.index_id, - scope.trust_profile, - scope.namespace, - scope.source_root, - scope.display_prefix, - ) - ) - if not has_provenance: - return parsed - - source_path = parsed.path - namespace = scope.namespace or "" - doc_id = parsed.id - if scope.index_id and namespace: - doc_id = parser.compute_scoped_doc_id(scope.index_id, namespace, source_path) - display_path = _display_path(scope.display_prefix, source_path) - source_root = scope.source_root or scope.root - return parsed.model_copy( - update={ - "id": doc_id, - "path": display_path, - "index_id": scope.index_id, - "trust_profile": scope.trust_profile, - "namespace": namespace or None, - "source_root": str(source_root), - "source_path": source_path, - "display_path": display_path, - } - ) - - -def _doc_id_for_scope_file(scope: IngestScope, source_path: str) -> str: - namespace = scope.namespace or "" - if scope.index_id and namespace: - return parser.compute_scoped_doc_id(scope.index_id, namespace, source_path) - return parser.compute_doc_id(source_path) - - -def _ingest_scopes( - scopes: list[IngestScope], - db_path: str, - *, - embedder: str | None = None, - embed_template: str | None = None, -) -> dict[str, int]: - stats = {"total": 0, "ingested": 0, "skipped": 0, "pruned": 0, "failed": 0} - scope_files: list[tuple[IngestScope, Path]] = [] - for scope in scopes: - md_files = _markdown_files_for_scope(scope) - scope_files.extend((scope, md_file) for md_file in md_files) - md_files = [md_file for _, md_file in scope_files] - stats["total"] = len(md_files) - - if not md_files: - logger.warning( - "No markdown files found in authoritative ingest scope(s); " - "pruning stale indexed documents" - ) - - spec = embedders.resolve_embedder(embedder) - template = embedders.resolve_embed_template(embed_template) - conn = db.init_db(db_path, embedding_dims=spec.dimensions) - model = None - parsed_docs: list[tuple[Path, parser.ParsedDocument]] = [] - - try: - for scope, md_file in scope_files: - relative_path = md_file.relative_to(scope.root).as_posix() - try: - body_bytes = md_file.read_bytes() - parsed = parser.parse_document(relative_path, body_bytes) - parsed_docs.append( - (md_file, _apply_scope_provenance(parsed, scope)) - ) - except MindgraphError as e: - logger.error("failed: %s — %s", relative_path, e) - stats["failed"] += 1 - except Exception: - logger.exception("unexpected failure: %s", relative_path) - stats["failed"] += 1 - - link_resolver = parser.LinkResolver.from_documents( - parsed for _, parsed in parsed_docs - ) - - for md_file, parsed in parsed_docs: - relative_path = parsed.path - try: - edges = parser.extract_document_graph_edges( - parsed, - link_resolver=link_resolver, - ) - - existing_hash = db.get_document_hash(conn, parsed.id) - if existing_hash == parsed.content_hash: - logger.debug("skipped (unchanged): %s", relative_path) - # Re-resolve edges (a newly-added note can resolve a target - # that was previously dangling), but only write when the - # resolved edge set actually differs from what is stored. - # This keeps the common no-op refresh read-only instead of - # running a DELETE+INSERT transaction per unchanged document. - new_edge_keys = { - (e.target_id, e.relationship_type) for e in edges - } - if new_edge_keys != db.get_outbound_edge_keys(conn, parsed.id): - with conn: - db.replace_edges(conn, parsed.id, edges) - stats["skipped"] += 1 - continue - - chunks = parser.chunk_truth(parsed.truth_text) - embeddings: list[list[float]] = [] - if chunks: - if model is None: - model = _load_embedder(embedder) - encode_chunks = [ - embedders.format_passage_text( - spec, - chunk, - template=template, - title=parsed.title, - domain=parsed.metadata.get("domain"), - doc_type=parsed.metadata.get("type"), - ) - for chunk in chunks - ] - raw = _encode_without_progress(model, encode_chunks) - embeddings = [row.tolist() for row in raw] - - with conn: - db.upsert_document(conn, parsed) - db.insert_chunks_and_embeddings( - conn, parsed.id, chunks, embeddings - ) - db.insert_edges(conn, edges) - - logger.info( - "ingested: %s (%d chunks, %d edges)", - relative_path, - len(chunks), - len(edges), - ) - stats["ingested"] += 1 - except MindgraphError as e: - logger.error("failed: %s — %s", relative_path, e) - stats["failed"] += 1 - except Exception: - logger.exception("unexpected failure: %s", relative_path) - stats["failed"] += 1 - - # Prune documents whose source file no longer exists. The selected file - # walk is authoritative for this ingest root, including files that failed - # to parse in this run; a partial refresh should not delete still-present - # rows just because one source is temporarily malformed. - # Materialize the id list before deleting so we don't mutate a live - # cursor. Inbound edges are left dangling by design (see delete_document). - on_disk_ids = { - _doc_id_for_scope_file(scope, md_file.relative_to(scope.root).as_posix()) - for scope, md_file in scope_files - } - db_ids = [row["id"] for row in conn.execute("SELECT id FROM documents")] - orphans = [doc_id for doc_id in db_ids if doc_id not in on_disk_ids] - if orphans: - with conn: - for doc_id in orphans: - db.delete_document(conn, doc_id) - stats["pruned"] += 1 - logger.info("pruned %d orphaned document(s)", stats["pruned"]) - finally: - conn.close() - - return stats - - -def _ingest_directory( - directory: Path, - db_path: str, - *, - index_id: str | None = None, - trust_profile: str | None = None, - namespace: str | None = None, - source_root: Path | None = None, - display_prefix: str | None = None, - include_globs: tuple[str, ...] | None = None, - exclude_globs: tuple[str, ...] | None = None, - embedder: str | None = None, - embed_template: str | None = None, -) -> dict[str, int]: - return _ingest_scopes( - [ - IngestScope( - root=directory, - index_id=index_id, - trust_profile=trust_profile, - namespace=namespace, - source_root=source_root, - display_prefix=display_prefix, - include_globs=include_globs or (), - exclude_globs=exclude_globs or (), - ) - ], - db_path, - embedder=embedder, - embed_template=embed_template, - ) - - -def _coerce_globs(value, field_name: str) -> tuple[str, ...]: - if value is None: - return () - if not isinstance(value, list): - raise IngestionError(f"manifest field {field_name!r} must be a list") - globs: list[str] = [] - for item in value: - if not isinstance(item, str) or not item.strip(): - raise IngestionError( - f"manifest field {field_name!r} must contain non-empty strings" - ) - globs.append(item.strip()) - return tuple(globs) - - -def _resolve_manifest_path(value: str, manifest_dir: Path) -> Path: - path = Path(value).expanduser() - if not path.is_absolute(): - path = manifest_dir / path - return path.resolve() - - -def _load_ingest_manifest(manifest_path: Path) -> list[IngestScope]: - try: - payload = json.loads(manifest_path.read_text()) - except json.JSONDecodeError as e: - raise IngestionError(f"manifest is not valid JSON: {e}") from e - if not isinstance(payload, dict): - raise IngestionError("manifest must be a JSON object") - - scopes_payload = payload.get("scopes") - if not isinstance(scopes_payload, list) or not scopes_payload: - raise IngestionError("manifest must include a non-empty 'scopes' list") - - manifest_dir = manifest_path.resolve().parent - scopes: list[IngestScope] = [] - for idx, entry in enumerate(scopes_payload, start=1): - if not isinstance(entry, dict): - raise IngestionError(f"manifest scope #{idx} must be an object") - root_value = entry.get("root") or entry.get("path") - if not isinstance(root_value, str) or not root_value.strip(): - raise IngestionError(f"manifest scope #{idx} requires a root path") - root = _resolve_manifest_path(root_value, manifest_dir) - if not root.is_dir(): - raise IngestionError(f"manifest scope #{idx} root is not a directory: {root}") - source_root_value = entry.get("source_root") - source_root = ( - _resolve_manifest_path(source_root_value, manifest_dir) - if isinstance(source_root_value, str) and source_root_value.strip() - else root - ) - scopes.append( - IngestScope( - root=root, - index_id=entry.get("index_id") or payload.get("index_id"), - trust_profile=entry.get("trust_profile") - or payload.get("trust_profile"), - namespace=entry.get("namespace"), - source_root=source_root, - display_prefix=entry.get("display_prefix"), - include_globs=_coerce_globs( - entry.get("include") or payload.get("include"), "include" - ), - exclude_globs=_coerce_globs( - entry.get("exclude") or payload.get("exclude"), "exclude" - ), - ) - ) - return scopes - - -@app.command() -def init( - db_path: str = typer.Option("mindgraph.sqlite", "--db", help="Path to SQLite DB."), - embedder: str | None = typer.Option( - None, - "--embedder", - help="Embedder key (minilm, bge-small, e5-small). Sets vec_chunks dimensions.", - ), - verbose: bool = typer.Option(False, "--verbose", "-v"), -): - """Initialize the MindGraph database.""" - _configure_logging(verbose) - try: - spec = embedders.resolve_embedder(embedder) - db.init_db(db_path, embedding_dims=spec.dimensions).close() - logger.info( - "Initialized database at %s (embedding_dims=%d, embedder=%s)", - db_path, - spec.dimensions, - spec.key, - ) - except MindgraphError as e: - logger.error(str(e)) - raise typer.Exit(code=1) - - -def _format_db_doctor_block(report: dict) -> str: - status = "OK" if report.get("ok") else "FAIL" - role = report.get("role") or "db" - trust = report.get("trust_profile") or "-" - lines = [ - f"## {role} [{status}]", - f" path: {report.get('path')}", - f" trust_profile: {trust}", - f" exists: {report.get('exists')}", - ] - if report.get("size_bytes") is not None: - size = report["size_bytes"] - if size >= 1_048_576: - size_h = f"{size / 1_048_576:.1f} MiB" - elif size >= 1024: - size_h = f"{size / 1024:.1f} KiB" - else: - size_h = f"{size} B" - lines.append(f" size: {size_h} ({size} bytes)") - if report.get("mtime_iso"): - lines.append(f" mtime: {report['mtime_iso']}") - if report.get("embedding_dims") is not None: - lines.append(f" embedding_dims: {report['embedding_dims']}") - missing = report.get("tables_missing") or [] - if missing: - lines.append(f" tables_missing: {', '.join(missing)}") - else: - lines.append(" tables_missing: (none)") - counts = report.get("counts") or {} - if counts: - parts = [ - f"{k}={v}" for k, v in counts.items() if v is not None - ] - if parts: - lines.append(f" counts: {', '.join(parts)}") - for issue in report.get("issues") or []: - lines.append(f" issue: {issue}") - for warn in report.get("warnings") or []: - lines.append(f" warning: {warn}") - return "\n".join(lines) - - -@app.command("doctor") -@app.command("status") -def doctor( - db_path: str | None = typer.Option( - None, - "--db", - help="Inspect a single DB. Default: dual MainFrame indexes under ~/.mindgraph/.", - ), - workspace: Path | None = typer.Option( - None, - "--workspace", - help="Also scan this directory for tiny stub mainframe*.sqlite files (default: cwd).", - ), - as_json: bool = typer.Option(False, "--json", help="Emit JSON diagnostics."), - verbose: bool = typer.Option(False, "--verbose", "-v"), -): - """First-contact diagnostics for MindGraph indexes (MH01). - - Reports authoritative DB paths, sizes, required tables (documents_fts, - vec_chunks, …), row counts, and workspace stub traps — without loading the - embedding model. - """ - _configure_logging(verbose) - reports: list[dict] = [] - if db_path: - reports.append(db.inspect_database(db_path)) - else: - for spec in db.default_dual_db_specs(): - report = db.inspect_database( - spec["path"], - role=spec["role"], - trust_profile=spec["trust_profile"], - ) - report["refresh_hint"] = spec.get("refresh_hint") - reports.append(report) - - scan_root = workspace if workspace is not None else Path.cwd() - stubs = db.find_workspace_stub_sqlite([scan_root]) - - # Exit non-zero only when a checked index is unusable. Workspace stubs are - # loud warnings (common on MainFrame checkouts) but not hard failures when - # ~/.mindgraph indexes are healthy. - dbs_ok = all(r.get("ok") for r in reports) - payload = { - "ok": dbs_ok, - "databases": reports, - "workspace_stubs": stubs, - "hints": [ - "Authoritative DBs: ~/.mindgraph/mainframe.sqlite (durable_knowledge)", - "Authoritative DBs: ~/.mindgraph/mainframe-projects.sqlite (project_status)", - "Never query workspace-root mainframe*.sqlite stubs", - "Refresh: bin/mindgraph-refresh && bin/mindgraph-refresh-projects", - ], - } - - if as_json: - typer.echo(json.dumps(payload, indent=2)) - else: - typer.echo("# MindGraph doctor") - typer.echo("") - for report in reports: - typer.echo(_format_db_doctor_block(report)) - if report.get("refresh_hint") and not report.get("ok"): - typer.echo(f" refresh_hint: {report['refresh_hint']}") - typer.echo("") - if stubs: - typer.echo("## workspace stubs (do not query these)") - for stub in stubs: - typer.echo(f" path: {stub['path']} size={stub['size_bytes']} B") - typer.echo(f" warning: {stub['warning']}") - typer.echo("") - if dbs_ok: - msg = "Overall: OK — dual indexes look query-ready." - if stubs: - msg += " (workspace stubs present — ignore them; use ~/.mindgraph paths)" - typer.echo(msg) - else: - typer.echo("Overall: FAIL — fix database issues above before planning queries.") - typer.echo("Hints:") - for h in payload["hints"]: - typer.echo(f" - {h}") - - if not dbs_ok: - raise typer.Exit(code=1) - - -@app.command() -def ingest( - directory: Path = typer.Argument( - ..., exists=True, file_okay=False, dir_okay=True, readable=True - ), - db_path: str = typer.Option("mindgraph.sqlite", "--db", help="Path to SQLite DB."), - index_id: str | None = typer.Option( - None, "--index-id", help="Optional lifecycle/index identifier." - ), - trust_profile: str | None = typer.Option( - None, "--trust-profile", help="Optional trust profile for every document." - ), - namespace: str | None = typer.Option( - None, "--namespace", help="Optional namespace for scoped document IDs." - ), - source_root: Path | None = typer.Option( - None, "--source-root", help="Absolute source root stored as provenance." - ), - display_prefix: str | None = typer.Option( - None, "--display-prefix", help="Path prefix shown to query clients." - ), - embedder: str | None = typer.Option( - None, - "--embedder", - help="Embedder key (minilm, bge-small, e5-small) or MINDGRAPH_EMBEDDER env.", - ), - embed_template: str | None = typer.Option( - None, - "--embed-template", - help="Optional passage/query template (none, mainframe).", - ), - verbose: bool = typer.Option(False, "--verbose", "-v"), -): - """Ingest a directory of markdown files.""" - _configure_logging(verbose) - try: - stats = _ingest_directory( - directory, - db_path, - index_id=index_id, - trust_profile=trust_profile, - namespace=namespace, - source_root=source_root, - display_prefix=display_prefix, - embedder=embedder, - embed_template=embed_template, - ) - logger.info( - "Done. total=%d ingested=%d skipped=%d pruned=%d failed=%d", - stats["total"], - stats["ingested"], - stats["skipped"], - stats["pruned"], - stats["failed"], - ) - if stats["failed"]: - raise typer.Exit(code=1) - except MindgraphError as e: - logger.error(str(e)) - raise typer.Exit(code=1) - - -@app.command("ingest-many") -def ingest_many( - manifest: Path = typer.Argument( - ..., exists=True, file_okay=True, dir_okay=False, readable=True - ), - db_path: str = typer.Option("mindgraph.sqlite", "--db", help="Path to SQLite DB."), - allow_failures: bool = typer.Option( - False, - "--allow-failures", - help="Keep a partial index when some files fail to parse or ingest.", - ), - embedder: str | None = typer.Option( - None, - "--embedder", - help="Embedder key (minilm, bge-small, e5-small) or MINDGRAPH_EMBEDDER env.", - ), - embed_template: str | None = typer.Option( - None, - "--embed-template", - help="Optional passage/query template (none, mainframe).", - ), - verbose: bool = typer.Option(False, "--verbose", "-v"), -): - """Ingest multiple markdown roots from a JSON manifest as one index.""" - _configure_logging(verbose) - try: - scopes = _load_ingest_manifest(manifest) - stats = _ingest_scopes( - scopes, - db_path, - embedder=embedder, - embed_template=embed_template, - ) - logger.info( - "Done. scopes=%d total=%d ingested=%d skipped=%d pruned=%d failed=%d", - len(scopes), - stats["total"], - stats["ingested"], - stats["skipped"], - stats["pruned"], - stats["failed"], - ) - if stats["failed"] and not allow_failures: - raise typer.Exit(code=1) - except MindgraphError as e: - logger.error(str(e)) - raise typer.Exit(code=1) - - -def _format_query_result_block(idx: int, result) -> str: - lex = result.lexical_rank if result.lexical_rank is not None else "-" - sem = result.semantic_rank if result.semantic_rank is not None else "-" - header = ( - f"#{idx} signal={result.signal} rrf_score={result.rrf_score:.6f} " - f"lex_rank={lex} sem_rank={sem}" - ) - if result.expansion_depth > 0: - header = f"{header} depth={result.expansion_depth}" - if result.semantic_distance is not None: - header = f"{header} dist={result.semantic_distance:.4f}" - if result.weak_fit: - header = f"{header} [weak-fit]" - excerpt = result.chunk_text.strip().replace("\n", " ") - if len(excerpt) > 280: - excerpt = excerpt[:277] + "..." - meta_parts = [ - f"{label}={value}" - for label, value in ( - ("type", result.doc_type), - ("domain", result.domain), - ("status", result.status), - ) - if value - ] - meta_line = f" meta: {' '.join(meta_parts)}\n" if meta_parts else "" - provenance_parts = [ - f"{label}={value}" - for label, value in ( - ("index", result.index_id), - ("trust", result.trust_profile), - ("namespace", result.namespace), - ) - if value - ] - provenance_line = ( - f" provenance: {' '.join(provenance_parts)}\n" - if provenance_parts - else "" - ) - return ( - f"{header}\n" - f" path: {result.path}\n" - f" title: {result.title}\n" - f"{meta_line}" - f"{provenance_line}" - f" chunk_index: {result.chunk_index}\n" - f" excerpt: {excerpt}" - ) - - -def _format_scope_warning(warning) -> str: - return ( - "scope warning: " - f"{warning.message} " - f"recommended_trust_profile={warning.recommended_trust_profile}" - ) - - -def _format_neighbor_block(idx: int, neighbor) -> str: - rel = neighbor.relationship_type or "(no relationship)" - target_path = neighbor.target_path or "(dangling)" - return ( - f"#{idx} -> {neighbor.target_id} rel={rel}\n" - f" target_path: {target_path}" - ) - - -@app.command() -def query( - question: str = typer.Argument(..., help="The free-text query."), - db_path: str = typer.Option("mindgraph.sqlite", "--db", help="Path to SQLite DB."), - lexical_top_k: int = typer.Option( - query_mod.DEFAULT_LEXICAL_TOP_K, - "--lexical-top-k", - help="Top-k for the FTS5 ranking before fusion.", - ), - semantic_top_k: int = typer.Option( - query_mod.DEFAULT_SEMANTIC_TOP_K, - "--semantic-top-k", - help="Top-k for the vec_chunks ranking before fusion.", - ), - final_top_k: int = typer.Option( - query_mod.DEFAULT_FINAL_TOP_K, - "--top-k", - help="Top-k for the fused output.", - ), - expand: bool = typer.Option( - False, - "--expand", - help="Walk outbound graph edges from Phase 2 results and append expanded matches.", - ), - expand_depth: int = typer.Option( - query_mod.DEFAULT_EXPAND_DEPTH, - "--depth", - min=1, - max=3, - help="Walk depth when --expand is set. Default 1, hard cap 3.", - ), - expand_top_k: int = typer.Option( - query_mod.DEFAULT_EXPAND_TOP_K, - "--expand-top-k", - help="Cap on the number of appended expanded results.", - ), - associate: bool = typer.Option( - False, - "--associate", - help="Append semantic doc-neighbor matches from fused seeds (ADR-034).", - ), - associate_top_k: int = typer.Option( - query_mod.DEFAULT_ASSOCIATE_TOP_K, - "--associate-top-k", - help="Cap on appended associated results.", - ), - associate_seed_k: int = typer.Option( - query_mod.DEFAULT_ASSOCIATE_SEED_K, - "--associate-seed-k", - help="How many fused rows seed association (default min(5, top-k)).", - ), - embedder: str | None = typer.Option( - None, - "--embedder", - help="Embedder key (minilm, bge-small, e5-small) or MINDGRAPH_EMBEDDER env.", - ), - embed_template: str | None = typer.Option( - None, - "--embed-template", - help="Optional query template (none, mainframe).", - ), - as_json: bool = typer.Option( - False, "--json", help="Emit machine-readable JSON instead of text." - ), - envelope: bool = typer.Option( - False, - "--envelope", - help="With --json, emit intent metadata plus results instead of the legacy result list.", - ), - no_intent: bool = typer.Option( - False, "--no-intent", help="Skip intent graph resolution." - ), - intent_db: str = typer.Option( - "~/.mindgraph/mainframe-intent.sqlite", "--intent-db", help="Path to intent graph DB for resolution." - ), - verbose: bool = typer.Option(False, "--verbose", "-v"), -): - """Run a fused lexical + semantic query against an ingested database. - - Pass --expand for graph BFS matches; --associate for semantic doc neighbors. - Pass --envelope with --json to include intent graph resolution metadata. - Plain --json preserves the legacy result-list contract for existing callers. - Text output still shows intent resolution by default. Use --no-intent to skip. - """ - _configure_logging(verbose) - try: - conn = db.get_db(db_path, read_only=True) - # MH01: fail fast on stub/incomplete DBs instead of soft FTS degradation. - db.validate_query_schema(conn, db_path) - except MindgraphError as e: - logger.error(str(e)) - raise typer.Exit(code=1) - - try: - try: - spec = embedders.resolve_embedder(embedder) - template = embedders.resolve_embed_template(embed_template) - model = _load_embedder(spec.key) - formatted_question = embedders.format_query_text( - spec, question, template=template - ) - results = query_mod.run_query( - conn, - formatted_question, - model, - lexical_top_k=lexical_top_k, - semantic_top_k=semantic_top_k, - final_top_k=final_top_k, - expand=expand, - expand_depth=expand_depth, - expand_top_k=expand_top_k, - associate=associate, - associate_top_k=associate_top_k, - associate_seed_k=associate_seed_k, - embedder_spec=spec, - embed_template=template, - ) - except MindgraphError as e: - logger.error(str(e)) - raise typer.Exit(code=1) - - resolution = None - if not no_intent and (envelope or not as_json): - intent_path = os.path.expanduser(intent_db) - intent_conn = None - if os.path.exists(intent_path): - try: - intent_conn = intent_mod.open_intent_store(intent_path) - except Exception as exc: - logger.warning("Failed to open intent DB: %s", exc) - if intent_conn: - try: - resolution = intent_mod.resolve_intent( - intent_conn, - formatted_question, - limits=intent_mod.TraversalLimits(max_depth=2, max_nodes=32), - ) - finally: - try: - intent_conn.close() - except Exception: - pass - - if as_json: - if envelope: - # Match MCP envelope shape so CLI and tool callers share one parser. - if resolution is not None: - resolution_payload = mcp_server._intent_resolution_payload( - resolution - ) - reason = "intent_resolved" - warnings = list(resolution.warnings) - else: - intent_path = Path(os.path.expanduser(intent_db)) - if intent_path.exists(): - reason = "intent_resolution_skipped_or_failed" - warnings = ["intent_resolution_unavailable"] - else: - reason = "intent_store_missing" - warnings = [] - resolution_payload = None - out = { - "schema_version": routing_mod.SCHEMA_VERSION, - "intent_resolution": resolution_payload, - "routing": { - "mode": "single_database", - "selected_retrievers": ["cli-bound-db"], - "reason_codes": [reason], - "warnings": warnings, - }, - "results": [r.model_dump() for r in results], - } - else: - out = [r.model_dump() for r in results] - typer.echo(json.dumps(out, indent=2, default=str)) - return - - if resolution: - typer.echo("=== Intent Resolution from graph ===") - typer.echo( - f"graph: {resolution.graph_id}@{resolution.graph_version}" - ) - typer.echo( - f"outcome: {resolution.outcome} " - f"(method: {resolution.resolution_method})" - ) - if resolution.matched_goal_ids: - typer.echo(f"matched goals: {list(resolution.matched_goal_ids)}") - if resolution.capability_hints: - typer.echo(f"hints: {list(resolution.capability_hints)}") - if resolution.warnings: - typer.echo(f"warnings: {list(resolution.warnings)}") - typer.echo("--- results below ---") - elif not no_intent: - typer.echo("(no intent DB or resolution; legacy results)") - - warning = query_mod.classify_query_scope(question) - if warning is not None: - typer.echo(_format_scope_warning(warning)) - - if not results: - typer.echo("(no candidate found)") - return - - for idx, result in enumerate(results, start=1): - typer.echo(_format_query_result_block(idx, result)) - finally: - try: - conn.close() - except Exception: # noqa: BLE001 - pass - - -@app.command() -def neighbors( - doc_id: str = typer.Argument(..., help="The source document ID."), - db_path: str = typer.Option("mindgraph.sqlite", "--db", help="Path to SQLite DB."), - as_json: bool = typer.Option( - False, "--json", help="Emit machine-readable JSON instead of text." - ), - verbose: bool = typer.Option(False, "--verbose", "-v"), -): - """List outbound edges from a document. Preserves dangling edges.""" - _configure_logging(verbose) - try: - conn = db.get_db(db_path, read_only=True) - except MindgraphError as e: - logger.error(str(e)) - raise typer.Exit(code=1) - try: - results = query_mod.list_neighbors(conn, doc_id) - except MindgraphError as e: - logger.error(str(e)) - conn.close() - raise typer.Exit(code=1) - finally: - try: - conn.close() - except Exception: # noqa: BLE001 - pass - - if as_json: - typer.echo(json.dumps([n.model_dump() for n in results], indent=2)) - return - - if not results: - typer.echo("(no outbound edges)") - return - - for idx, neighbor in enumerate(results, start=1): - typer.echo(_format_neighbor_block(idx, neighbor)) - - -@app.command("serve-mcp") -def serve_mcp( - db_path: str = typer.Option("mindgraph.sqlite", "--db", help="Path to SQLite DB."), - embedder: str | None = typer.Option( - None, - "--embedder", - help="Embedder key (minilm, bge-small, e5-small) or MINDGRAPH_EMBEDDER env.", - ), - embed_template: str | None = typer.Option( - None, - "--embed-template", - help="Optional query template (none, mainframe).", - ), - intent_db: str = typer.Option( - "~/.mindgraph/mainframe-intent.sqlite", - "--intent-db", - help="Path to intent graph DB used for MCP query envelope metadata.", - ), - verbose: bool = typer.Option(False, "--verbose", "-v"), -): - """Start a stdio MCP server for one ingested MindGraph database.""" - _configure_logging(verbose) - conn = None - try: - conn = mcp_server.open_database(db_path) - spec = embedders.resolve_embedder(embedder) - template = embedders.resolve_embed_template(embed_template) - model = _load_embedder(spec.key) - server = mcp_server.create_server( - conn, - model, - embedder_spec=spec, - embed_template=template, - intent_db_path=intent_db, - log_level="DEBUG" if verbose else "INFO", - ) - mcp_server.run_stdio(server) - except MindgraphError as e: - logger.error(str(e)) - typer.echo(str(e), err=True) - raise typer.Exit(code=1) - finally: - if conn is not None: - conn.close() - - -SCOPE_SPEC_HELP = ( - "Serve an arbitrary named scope: NAME=PATH or NAME:TRUST_PROFILE=PATH. " - "Repeatable. Trust profile defaults to NAME. When any --scope is given it " - "replaces the default knowledge/projects pair." -) - - -def parse_scope_specs(values: list[str] | None) -> dict[str, tuple[str, str]]: - """Parse `NAME=PATH` / `NAME:TRUST=PATH` specs into {name: (path, trust)}. - - Split on the first `=` only, so database paths may contain any character. - """ - scopes: dict[str, tuple[str, str]] = {} - for raw in values or []: - name_part, sep, db_path = raw.partition("=") - if not sep or not name_part.strip() or not db_path.strip(): - raise MindgraphError( - f"invalid --scope {raw!r}; expected NAME=PATH or NAME:TRUST_PROFILE=PATH" - ) - name, _, trust = name_part.partition(":") - name, trust = name.strip(), trust.strip() - if not name: - raise MindgraphError(f"invalid --scope {raw!r}; scope name is empty") - if name in scopes: - raise MindgraphError(f"duplicate --scope name: {name}") - scopes[name] = (db_path.strip(), trust or name) - return scopes - - -@app.command("serve-daemon") -def serve_daemon( - knowledge_db: str = typer.Option("~/.mindgraph/mainframe.sqlite", "--knowledge-db"), - projects_db: str = typer.Option("~/.mindgraph/mainframe-projects.sqlite", "--projects-db"), - scope: list[str] = typer.Option(None, "--scope", help=SCOPE_SPEC_HELP), - host: str = typer.Option("127.0.0.1", "--host"), - port: int = typer.Option(8000, "--port"), - path: str = typer.Option("/mcp", "--path"), - embedder: str | None = typer.Option(None, "--embedder"), - idle_seconds: float | None = typer.Option(None, "--idle-seconds", min=60.0), - lease_ttl: float = typer.Option(90.0, "--lease-ttl", min=30.0), -): - """Run the explicit-scope shared MCP server in the foreground.""" - conns = [] - try: - requested = parse_scope_specs(scope) or { - "knowledge": (knowledge_db, "durable_knowledge"), - "projects": (projects_db, "project_status"), - } - scopes = {} - for name, (db_path, trust_profile) in requested.items(): - conn = mcp_server.open_database_readonly(db_path) - conns.append(conn) - scopes[name] = (conn, trust_profile) - spec = embedders.resolve_embedder(embedder) - lifecycle = ( - idle_lifecycle.IdleLifecycle(idle_seconds, lease_ttl=lease_ttl) - if idle_seconds is not None else None - ) - server = mcp_server.create_shared_server( - scopes, - _load_embedder(spec.key), host=host, port=port, path=path, - embedder_spec=spec, lifecycle=lifecycle, - ) - if lifecycle: - lifecycle.start() - try: - mcp_server.run_streamable_http(server) - finally: - if lifecycle: - lifecycle.stop() - except MindgraphError as exc: - typer.echo(str(exc), err=True) - raise typer.Exit(1) - finally: - for conn in conns: - conn.close() - - -@app.command("mcp-proxy") -def mcp_proxy_command( - url: str = typer.Option("http://127.0.0.1:8000/mcp", "--url"), - lease: bool = typer.Option(False, "--lease"), - auto_start: bool | None = typer.Option(None, "--auto-start/--no-auto-start"), - state_dir: Path = typer.Option(Path("~/.mindgraph/run").expanduser(), "--state-dir"), - idle_seconds: float = typer.Option(900.0, "--idle-seconds", min=60.0), -): - """Proxy stdio MCP to an already-running shared daemon.""" - try: - activated_grace = daemon.idle_opt_in(state_dir) - if auto_start is None: - auto_start = activated_grace is not None - if activated_grace is not None and idle_seconds == 900.0: - idle_seconds = activated_grace - if auto_start: - health_url = daemon.health_url_for_mcp(url) - daemon_host, daemon_port, daemon_path = daemon.daemon_endpoint_args(url) - command = [ - sys.executable, "-m", "mindgraph.cli", "serve-daemon", - "--idle-seconds", str(idle_seconds), - "--host", daemon_host, "--port", str(daemon_port), - "--path", daemon_path, - ] - result = daemon.start(state_dir, command, health_url=health_url) - if result["status"] == "start_failed": - raise MindgraphError(f"shared daemon failed to start: {result}") - mcp_proxy.run_proxy_sync(url, lease=lease or auto_start) - except Exception as exc: - typer.echo(f"MCP proxy failed: {exc}", err=True) - raise typer.Exit(1) - - -@app.command("daemon-start") -def daemon_start( - state_dir: Path = typer.Option(Path("~/.mindgraph/run").expanduser(), "--state-dir"), - knowledge_db: str = typer.Option("~/.mindgraph/mainframe.sqlite", "--knowledge-db"), - projects_db: str = typer.Option("~/.mindgraph/mainframe-projects.sqlite", "--projects-db"), - scope: list[str] = typer.Option(None, "--scope", help=SCOPE_SPEC_HELP), - host: str = typer.Option("127.0.0.1", "--host"), - port: int = typer.Option(8000, "--port"), - path: str = typer.Option("/mcp", "--path"), - idle_seconds: float | None = typer.Option(None, "--idle-seconds", min=60.0), -): - try: - parse_scope_specs(scope) - except MindgraphError as exc: - typer.echo(str(exc), err=True) - raise typer.Exit(1) - command = [sys.executable, "-m", "mindgraph.cli", "serve-daemon", - "--host", host, "--port", str(port), "--path", path] - if scope: - for spec in scope: - command.extend(["--scope", spec]) - else: - command.extend(["--knowledge-db", knowledge_db, - "--projects-db", projects_db]) - if idle_seconds is not None: - command.extend(["--idle-seconds", str(idle_seconds)]) - health_endpoint = daemon.health_url(host, port) - typer.echo(json.dumps(daemon.start(state_dir, command, health_url=health_endpoint))) - - -@app.command("daemon-status") -def daemon_status( - state_dir: Path = typer.Option(Path("~/.mindgraph/run").expanduser(), "--state-dir"), - url: str = typer.Option("http://127.0.0.1:8000/health", "--url"), -): - typer.echo(json.dumps(daemon.status(state_dir, url))) - - -@app.command("daemon-health") -def daemon_health(url: str = typer.Option("http://127.0.0.1:8000/health", "--url")): - result = daemon.health(url) - typer.echo(json.dumps(result)) - if result.get("status") != "ok": - raise typer.Exit(1) - - -@app.command("daemon-stop") -def daemon_stop( - state_dir: Path = typer.Option(Path("~/.mindgraph/run").expanduser(), "--state-dir"), - url: str = typer.Option("http://127.0.0.1:8000/health", "--url"), -): - typer.echo(json.dumps(daemon.stop(state_dir, health_url=url))) - - -if __name__ == "__main__": - app() diff --git a/mindgraph/src/mindgraph/daemon.py b/mindgraph/src/mindgraph/daemon.py deleted file mode 100644 index bb1a7f8..0000000 --- a/mindgraph/src/mindgraph/daemon.py +++ /dev/null @@ -1,195 +0,0 @@ -"""Explicit process lifecycle for the opt-in shared MCP daemon.""" - -import json -import fcntl -import os -import signal -import subprocess -import time -import urllib.request -import urllib.parse -from pathlib import Path - - -def paths(state_dir: Path) -> tuple[Path, Path]: - return state_dir / "mindgraph-daemon.pid", state_dir / "mindgraph-daemon.log" - - -def lock_path(state_dir: Path) -> Path: - return state_dir / "mindgraph-daemon.start.lock" - - -def identity_path(state_dir: Path) -> Path: - return state_dir / "mindgraph-daemon.identity.json" - - -def read_identity(state_dir: Path) -> dict | None: - try: - payload = json.loads(identity_path(state_dir).read_text()) - except (FileNotFoundError, json.JSONDecodeError): - return None - return payload if isinstance(payload, dict) else None - - -def read_pid(state_dir: Path) -> int | None: - try: - return int(paths(state_dir)[0].read_text().strip()) - except (FileNotFoundError, ValueError): - return None - - -def alive(pid: int | None) -> bool: - if pid is None: - return False - try: - waited, _status = os.waitpid(pid, os.WNOHANG) - if waited == pid: - return False - except ChildProcessError: - pass - try: - os.kill(pid, 0) - return True - except ProcessLookupError: - return False - - -def is_mindgraph_health(payload: dict | None) -> bool: - return bool( - payload - and payload.get("status") == "ok" - and isinstance(payload.get("scopes"), list) - ) - - -def status(state_dir: Path, health_url: str | None = None) -> dict: - pid = read_pid(state_dir) - tracked = alive(pid) - observed = health(health_url) if health_url else None - if is_mindgraph_health(observed): - observed_pid = observed.get("pid") - return { - "status": "running", - "pid": observed_pid, - "supervision": "pid_file" if tracked and observed_pid == pid else "external", - "pid_file_pid": pid, - "health": observed, - } - result = {"status": "running" if tracked else "stopped", "pid": pid} - if health_url: - result["health"] = observed - return result - - -def start( - state_dir: Path, - command: list[str], - *, - health_url: str | None = None, - startup_timeout: float = 20.0, -) -> dict: - state_dir.mkdir(parents=True, exist_ok=True) - with lock_path(state_dir).open("a+") as lock: - fcntl.flock(lock, fcntl.LOCK_EX) - current = status(state_dir, health_url) - if current["status"] == "running": - return current - pid_file, log_file = paths(state_dir) - with log_file.open("ab") as log: - proc = subprocess.Popen( - command, stdin=subprocess.DEVNULL, stdout=log, stderr=log, - start_new_session=True, - ) - pid_file.write_text(f"{proc.pid}\n") - identity_path(state_dir).write_text(json.dumps({ - "pid": proc.pid, - "command": command, - }) + "\n") - if not health_url: - return {"status": "started", "pid": proc.pid} - deadline = time.monotonic() + startup_timeout - while time.monotonic() < deadline: - observed = health(health_url, timeout=0.25) - if is_mindgraph_health(observed): - return {"status": "started", "pid": proc.pid, "health": observed} - if not alive(proc.pid): - break - time.sleep(0.1) - return { - "status": "start_failed", - "pid": proc.pid, - "health": health(health_url, timeout=0.25), - } - - -def stop( - state_dir: Path, - timeout: float = 5.0, - *, - health_url: str | None = None, -) -> dict: - pid = read_pid(state_dir) - if health_url: - observed = health(health_url) - if is_mindgraph_health(observed) and observed.get("pid") is None: - return {"status": "refused_unverified_listener", "pid": pid} - observed_pid = observed.get("pid") if is_mindgraph_health(observed) else None - if observed_pid is not None and observed_pid != pid: - return { - "status": "refused_pid_mismatch", - "pid": pid, - "observed_pid": observed_pid, - } - if not alive(pid): - paths(state_dir)[0].unlink(missing_ok=True) - identity_path(state_dir).unlink(missing_ok=True) - return {"status": "stopped", "pid": pid} - identity = read_identity(state_dir) - if identity is None or identity.get("pid") != pid: - return {"status": "refused_unverified_pid", "pid": pid} - os.kill(pid, signal.SIGTERM) - deadline = time.monotonic() + timeout - while alive(pid) and time.monotonic() < deadline: - time.sleep(0.05) - if alive(pid): - return {"status": "stop_timeout", "pid": pid} - paths(state_dir)[0].unlink(missing_ok=True) - identity_path(state_dir).unlink(missing_ok=True) - return {"status": "stopped", "pid": pid} - - -def health(url: str, timeout: float = 1.0) -> dict: - try: - with urllib.request.urlopen(url, timeout=timeout) as response: - return json.loads(response.read()) - except Exception as exc: - return {"status": "unhealthy", "error": str(exc)} - - -def health_url_for_mcp(url: str) -> str: - parsed = urllib.parse.urlsplit(url) - if parsed.scheme not in {"http", "https"} or not parsed.netloc: - raise ValueError(f"invalid MCP URL: {url}") - return urllib.parse.urlunsplit((parsed.scheme, parsed.netloc, "/health", "", "")) - - -def daemon_endpoint_args(url: str) -> tuple[str, int, str]: - parsed = urllib.parse.urlsplit(url) - if parsed.scheme != "http" or parsed.hostname not in {"127.0.0.1", "localhost", "::1"}: - raise ValueError("auto-start requires a loopback http MCP URL") - return parsed.hostname, parsed.port or 80, parsed.path or "/mcp" - - -def health_url(host: str, port: int) -> str: - rendered_host = f"[{host}]" if ":" in host else host - return f"http://{rendered_host}:{port}/health" - - -def idle_opt_in(state_dir: Path) -> float | None: - """Return the explicitly activated idle grace, or None when disabled.""" - marker = state_dir / "idle-lifecycle.enabled" - try: - value = float(marker.read_text().strip()) - except (FileNotFoundError, ValueError): - return None - return value if value >= 60 else None diff --git a/mindgraph/src/mindgraph/db.py b/mindgraph/src/mindgraph/db.py deleted file mode 100644 index fe92f90..0000000 --- a/mindgraph/src/mindgraph/db.py +++ /dev/null @@ -1,577 +0,0 @@ -import os -import sqlite3 -import struct -import time -from collections.abc import Iterable -from pathlib import Path -from typing import Any -from urllib.parse import quote - -import sqlite_vec - -from mindgraph.exceptions import DatabaseError -from mindgraph.models import GraphEdge, ParsedDocument - -# Tables required for fused lexical + semantic query (MH01 first-contact contract). -REQUIRED_QUERY_TABLES = frozenset( - { - "documents", - "documents_fts", - "chunks", - "vec_chunks", - "edges", - } -) - -# Heuristic: workspace-root stub SQLite files from early scaffolding are tiny. -STUB_SIZE_BYTES = 16_384 # 16 KiB - - -def _configure_connection( - conn: sqlite3.Connection, *, read_only: bool -) -> sqlite3.Connection: - conn.enable_load_extension(True) - sqlite_vec.load(conn) - conn.enable_load_extension(False) - conn.row_factory = sqlite3.Row - conn.execute("PRAGMA foreign_keys = ON") - if read_only: - conn.execute("PRAGMA query_only = ON") - # Force SQLite to read the schema now. A WAL database can connect in - # mode=ro and then fail on its first query when the containing directory - # cannot host a shared-memory file; callers need that failure here so - # the immutable snapshot fallback can be selected deterministically. - conn.execute("SELECT 1 FROM sqlite_master LIMIT 1").fetchone() - else: - # WAL lets the long-lived MCP reader and the ingest/refresh writer - # coexist without `database is locked` errors. Journal mode is a - # persistent property of the file, so issuing it on every writable - # connection is idempotent and migrates pre-existing rollback journals. - conn.execute("PRAGMA journal_mode = WAL") - return conn - - -def _open_read_only(db_path: str) -> sqlite3.Connection: - """Open a query-only database, falling back to an immutable snapshot. - - Normal ``mode=ro`` remains preferred because it observes a live WAL. The - immutable fallback is only used when SQLite cannot create/read WAL shared - memory in a non-writable runtime (for example, a sandbox-mounted index). - """ - if db_path == ":memory:": - raise DatabaseError(":memory: cannot be opened read-only") - - resolved = Path(db_path).expanduser().resolve() - base_uri = f"file:{quote(str(resolved), safe='/')}?mode=ro" - last_error: Exception | None = None - for immutable in (False, True): - conn: sqlite3.Connection | None = None - uri = f"{base_uri}&immutable=1" if immutable else base_uri - try: - conn = sqlite3.connect(uri, uri=True, timeout=30.0) - return _configure_connection(conn, read_only=True) - except (sqlite3.Error, RuntimeError) as exc: - last_error = exc - if conn is not None: - try: - conn.close() - except sqlite3.Error: - pass - assert last_error is not None - raise DatabaseError(f"Failed to open database at {db_path}: {last_error}") - - -def get_db( - db_path: str = "mindgraph.sqlite", *, read_only: bool = False -) -> sqlite3.Connection: - """Connect to SQLite, load sqlite-vec, and apply the requested access mode.""" - if read_only: - return _open_read_only(db_path) - - try: - conn = sqlite3.connect(db_path, timeout=30.0) - _configure_connection(conn, read_only=False) - if conn.execute( - "SELECT 1 FROM sqlite_master WHERE type = 'table' AND name = 'documents'" - ).fetchone(): - with conn: - _ensure_document_provenance_columns(conn) - return conn - except (sqlite3.Error, RuntimeError) as e: - raise DatabaseError(f"Failed to open database at {db_path}: {e}") from e - - -DEFAULT_EMBEDDING_DIMS = 384 - - -def init_db( - db_path: str = "mindgraph.sqlite", *, embedding_dims: int = DEFAULT_EMBEDDING_DIMS -) -> sqlite3.Connection: - """Initialize the database schema for MindGraph.""" - conn = get_db(db_path) - - with conn: - conn.execute(""" - CREATE TABLE IF NOT EXISTS index_meta ( - key TEXT PRIMARY KEY, - value TEXT NOT NULL - ) - """) - conn.execute(""" - CREATE TABLE IF NOT EXISTS documents ( - id TEXT PRIMARY KEY, - title TEXT, - path TEXT, - domain TEXT, - content_hash TEXT NOT NULL, - index_id TEXT, - trust_profile TEXT, - namespace TEXT, - source_root TEXT, - source_path TEXT, - display_path TEXT, - timeline_text TEXT, - metadata_json TEXT, - created_at TEXT, - updated_at TEXT - ) - """) - _ensure_document_provenance_columns(conn) - - conn.execute(""" - CREATE VIRTUAL TABLE IF NOT EXISTS documents_fts USING fts5( - id UNINDEXED, - title, - content - ) - """) - - conn.execute(""" - CREATE TABLE IF NOT EXISTS chunks ( - rowid INTEGER PRIMARY KEY AUTOINCREMENT, - doc_id TEXT, - chunk_index INTEGER, - text TEXT, - FOREIGN KEY (doc_id) REFERENCES documents(id) ON DELETE CASCADE - ) - """) - - vec_exists = conn.execute( - "SELECT 1 FROM sqlite_master WHERE type = 'table' AND name = 'vec_chunks'" - ).fetchone() - stored_dims = get_embedding_dims(conn) - if stored_dims is None: - if vec_exists: - stored_dims = DEFAULT_EMBEDDING_DIMS - else: - stored_dims = embedding_dims - set_embedding_dims(conn, stored_dims) - elif stored_dims != embedding_dims: - raise DatabaseError( - f"Database {db_path} was initialized with embedding_dims=" - f"{stored_dims}; requested {embedding_dims}. Use a separate DB " - "per embedder dimension." - ) - - if not vec_exists: - conn.execute( - f""" - CREATE VIRTUAL TABLE vec_chunks USING vec0( - embedding float[{stored_dims}] - ) - """ - ) - - conn.execute(""" - CREATE TABLE IF NOT EXISTS edges ( - source_id TEXT, - target_id TEXT, - relationship_type TEXT, - PRIMARY KEY (source_id, target_id, relationship_type) - ) - """) - - return conn - - -def list_table_names(conn: sqlite3.Connection) -> set[str]: - rows = conn.execute( - "SELECT name FROM sqlite_master WHERE type IN ('table', 'virtual table')" - ).fetchall() - return {row["name"] if isinstance(row, sqlite3.Row) else row[0] for row in rows} - - -def missing_required_tables( - conn: sqlite3.Connection, required: frozenset[str] = REQUIRED_QUERY_TABLES -) -> list[str]: - existing = list_table_names(conn) - return sorted(required - existing) - - -def validate_query_schema( - conn: sqlite3.Connection, - db_path: str, - *, - required: frozenset[str] = REQUIRED_QUERY_TABLES, -) -> None: - """Fail fast when a DB cannot support fused query (MH01 doctor contract). - - Raises DatabaseError with an actionable message instead of soft-degrading - to semantic-only / empty FTS results on stub databases. - """ - missing = missing_required_tables(conn, required) - if missing: - raise DatabaseError( - f"Database at {db_path} is not a usable MindGraph index; " - f"missing tables: {', '.join(missing)}. " - "Authoritative DBs live under ~/.mindgraph/ " - "(not workspace-root stub *.sqlite files). " - "Run: bin/mindgraph doctor && bin/mindgraph-refresh" - ) - - -def _count(conn: sqlite3.Connection, sql: str) -> int | None: - try: - row = conn.execute(sql).fetchone() - if row is None: - return None - return int(row[0]) - except sqlite3.Error: - return None - - -def inspect_database( - db_path: str, - *, - role: str | None = None, - trust_profile: str | None = None, -) -> dict[str, Any]: - """Return a doctor diagnostic payload for one database path. - - Does not load the embedding model. Safe for first-contact preflight. - """ - expanded = os.path.expanduser(db_path) - path = Path(expanded) - report: dict[str, Any] = { - "path": str(path), - "role": role, - "trust_profile": trust_profile, - "exists": path.exists(), - "size_bytes": None, - "mtime_iso": None, - "ok": False, - "issues": [], - "warnings": [], - "tables_present": [], - "tables_missing": sorted(REQUIRED_QUERY_TABLES), - "counts": {}, - "embedding_dims": None, - "likely_stub": False, - } - - if not path.exists(): - report["issues"].append("file_missing") - return report - - try: - stat = path.stat() - except OSError as e: - report["issues"].append(f"stat_failed:{e}") - return report - - report["size_bytes"] = stat.st_size - report["mtime_iso"] = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime(stat.st_mtime)) - if stat.st_size <= STUB_SIZE_BYTES: - report["likely_stub"] = True - report["warnings"].append( - f"file is very small ({stat.st_size} bytes); may be a workspace stub, " - "not an authoritative index" - ) - - conn: sqlite3.Connection | None = None - try: - conn = get_db(str(path), read_only=True) - tables = list_table_names(conn) - report["tables_present"] = sorted(tables) - missing = missing_required_tables(conn) - report["tables_missing"] = missing - if missing: - report["issues"].append("missing_required_tables") - report["embedding_dims"] = get_embedding_dims(conn) - report["counts"] = { - "documents": _count(conn, "SELECT COUNT(*) FROM documents") - if "documents" in tables - else None, - "chunks": _count(conn, "SELECT COUNT(*) FROM chunks") - if "chunks" in tables - else None, - "edges": _count(conn, "SELECT COUNT(*) FROM edges") - if "edges" in tables - else None, - "vec_chunks": _count(conn, "SELECT COUNT(*) FROM vec_chunks") - if "vec_chunks" in tables - else None, - } - docs = report["counts"].get("documents") - if docs == 0 and not missing: - report["warnings"].append("index has zero documents — run mindgraph-refresh") - report["ok"] = not report["issues"] - except DatabaseError as e: - report["issues"].append("open_failed") - report["warnings"].append(str(e)) - except sqlite3.Error as e: - report["issues"].append("sqlite_error") - report["warnings"].append(str(e)) - finally: - if conn is not None: - try: - conn.close() - except Exception: # noqa: BLE001 - pass - - return report - - -def default_dual_db_specs() -> list[dict[str, str]]: - """Authoritative MainFrame dual-index locations (MH01 / HARNESS).""" - home = Path.home() / ".mindgraph" - return [ - { - "role": "knowledge", - "trust_profile": "durable_knowledge", - "path": str(home / "mainframe.sqlite"), - "refresh_hint": "bin/mindgraph-refresh", - }, - { - "role": "projects", - "trust_profile": "project_status", - "path": str(home / "mainframe-projects.sqlite"), - "refresh_hint": "bin/mindgraph-refresh-projects", - }, - ] - - -def find_workspace_stub_sqlite(search_roots: list[Path] | None = None) -> list[dict[str, Any]]: - """Detect tiny workspace-root *.sqlite files that look authoritative but are not.""" - roots = search_roots or [Path.cwd()] - hits: list[dict[str, Any]] = [] - for root in roots: - if not root.is_dir(): - continue - for name in ("mainframe.sqlite", "mainframe-projects.sqlite"): - p = root / name - if not p.is_file(): - continue - try: - size = p.stat().st_size - except OSError: - continue - if size <= STUB_SIZE_BYTES: - hits.append( - { - "path": str(p.resolve()), - "size_bytes": size, - "warning": ( - "Workspace-root sqlite looks like a stub. " - "Query ~/.mindgraph/*.sqlite instead." - ), - } - ) - return hits - - -def get_embedding_dims(conn: sqlite3.Connection) -> int | None: - row = conn.execute( - "SELECT value FROM index_meta WHERE key = 'embedding_dims'" - ).fetchone() - if row is None: - return None - try: - return int(row["value"]) - except (TypeError, ValueError): - return None - - -def set_embedding_dims(conn: sqlite3.Connection, dims: int) -> None: - conn.execute( - """ - INSERT INTO index_meta (key, value) VALUES ('embedding_dims', ?) - ON CONFLICT(key) DO UPDATE SET value = excluded.value - """, - (str(dims),), - ) - - -def _ensure_document_provenance_columns(conn: sqlite3.Connection) -> None: - """Add provenance columns to older databases without requiring a rebuild.""" - existing = { - row["name"] - for row in conn.execute("PRAGMA table_info(documents)").fetchall() - } - for column in ( - "index_id", - "trust_profile", - "namespace", - "source_root", - "source_path", - "display_path", - ): - if column not in existing: - conn.execute(f"ALTER TABLE documents ADD COLUMN {column} TEXT") - - -def _serialize_embedding(vec: list[float]) -> bytes: - """Pack a float vector into the bytes format sqlite-vec expects.""" - return struct.pack(f"{len(vec)}f", *vec) - - -def get_document_hash(conn: sqlite3.Connection, doc_id: str) -> str | None: - row = conn.execute( - "SELECT content_hash FROM documents WHERE id = ?", (doc_id,) - ).fetchone() - return row["content_hash"] if row else None - - -def _delete_document_artifacts(conn: sqlite3.Connection, doc_id: str) -> None: - """Remove chunks, vec_chunks, FTS rows, and outgoing edges for a doc.""" - chunk_rowids = [ - row["rowid"] - for row in conn.execute( - "SELECT rowid FROM chunks WHERE doc_id = ?", (doc_id,) - ) - ] - if chunk_rowids: - placeholders = ",".join("?" for _ in chunk_rowids) - conn.execute( - f"DELETE FROM vec_chunks WHERE rowid IN ({placeholders})", - chunk_rowids, - ) - conn.execute("DELETE FROM chunks WHERE doc_id = ?", (doc_id,)) - conn.execute("DELETE FROM documents_fts WHERE id = ?", (doc_id,)) - conn.execute("DELETE FROM edges WHERE source_id = ?", (doc_id,)) - - -def delete_document(conn: sqlite3.Connection, doc_id: str) -> None: - """Fully remove a document from the index, including the documents row. - - Drops the chunks, vec rows, FTS row, and outbound edges (via - `_delete_document_artifacts`) and then the `documents` row itself. Inbound - edges (where this doc is the *target*) are intentionally left in place so - they become dangling, matching the engine's link-resolution model where an - unresolved target stays dangling rather than being guessed. Used by ingest - to prune deleted/renamed source files. - """ - _delete_document_artifacts(conn, doc_id) - conn.execute("DELETE FROM documents WHERE id = ?", (doc_id,)) - - -def upsert_document(conn: sqlite3.Connection, doc: ParsedDocument) -> None: - """Insert or replace a document and clear any prior chunks/edges/FTS rows.""" - import json - from datetime import datetime, timezone - - _delete_document_artifacts(conn, doc.id) - - now = datetime.now(timezone.utc).isoformat() - existing = conn.execute( - "SELECT created_at FROM documents WHERE id = ?", (doc.id,) - ).fetchone() - created_at = existing["created_at"] if existing else now - - conn.execute( - """ - INSERT OR REPLACE INTO documents - ( - id, title, path, domain, content_hash, index_id, trust_profile, - namespace, source_root, source_path, display_path, - timeline_text, metadata_json, created_at, updated_at - ) - VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) - """, - ( - doc.id, - doc.title, - doc.path, - doc.metadata.get("domain"), - doc.content_hash, - doc.index_id, - doc.trust_profile, - doc.namespace, - doc.source_root, - doc.source_path, - doc.display_path, - doc.timeline_text, - json.dumps(doc.metadata, default=str), - created_at, - now, - ), - ) - - conn.execute( - "INSERT INTO documents_fts (id, title, content) VALUES (?, ?, ?)", - (doc.id, doc.title, doc.truth_text), - ) - - -def insert_chunks_and_embeddings( - conn: sqlite3.Connection, - doc_id: str, - chunks: list[str], - embeddings: list[list[float]], -) -> None: - if len(chunks) != len(embeddings): - raise DatabaseError( - f"chunk/embedding count mismatch for {doc_id}: " - f"{len(chunks)} chunks vs {len(embeddings)} embeddings" - ) - for idx, (text, embedding) in enumerate(zip(chunks, embeddings)): - cursor = conn.execute( - "INSERT INTO chunks (doc_id, chunk_index, text) VALUES (?, ?, ?)", - (doc_id, idx, text), - ) - rowid = cursor.lastrowid - conn.execute( - "INSERT INTO vec_chunks (rowid, embedding) VALUES (?, ?)", - (rowid, _serialize_embedding(embedding)), - ) - - -def insert_edges(conn: sqlite3.Connection, edges: Iterable[GraphEdge]) -> None: - for edge in edges: - conn.execute( - """ - INSERT OR IGNORE INTO edges (source_id, target_id, relationship_type) - VALUES (?, ?, ?) - """, - (edge.source_id, edge.target_id, edge.relationship_type), - ) - - -def get_outbound_edge_keys( - conn: sqlite3.Connection, source_id: str -) -> set[tuple[str, str | None]]: - """Return the set of (target_id, relationship_type) for a source's edges. - - Lets the ingest skip-path detect when a re-resolved edge set is identical to - what is already stored, so it can avoid a redundant DELETE+INSERT write. - """ - return { - (row["target_id"], row["relationship_type"]) - for row in conn.execute( - "SELECT target_id, relationship_type FROM edges WHERE source_id = ?", - (source_id,), - ) - } - - -def replace_edges( - conn: sqlite3.Connection, source_id: str, edges: Iterable[GraphEdge] -) -> None: - """Replace all outbound edges for one source document.""" - conn.execute("DELETE FROM edges WHERE source_id = ?", (source_id,)) - insert_edges(conn, edges) - - -if __name__ == "__main__": - init_db() - print("Database schema initialized successfully.") diff --git a/mindgraph/src/mindgraph/embedders.py b/mindgraph/src/mindgraph/embedders.py deleted file mode 100644 index 78bd782..0000000 --- a/mindgraph/src/mindgraph/embedders.py +++ /dev/null @@ -1,137 +0,0 @@ -"""Embedding model registry and text formatting for ingest/query.""" - -from __future__ import annotations - -import os -from dataclasses import dataclass -from typing import Literal - -from mindgraph.exceptions import EmbeddingError - -EmbedTemplate = Literal["none", "mainframe"] - - -@dataclass(frozen=True) -class EmbedderSpec: - key: str - model_id: str - dimensions: int - query_prefix: str = "" - passage_prefix: str = "" - - -_REGISTRY: dict[str, EmbedderSpec] = { - "minilm": EmbedderSpec("minilm", "all-MiniLM-L6-v2", 384), - "all-minilm-l6-v2": EmbedderSpec( - "all-minilm-l6-v2", "all-MiniLM-L6-v2", 384 - ), - "bge-small": EmbedderSpec( - "bge-small", - "BAAI/bge-small-en-v1.5", - 384, - query_prefix="query: ", - passage_prefix="", - ), - "bge-small-en-v1.5": EmbedderSpec( - "bge-small-en-v1.5", - "BAAI/bge-small-en-v1.5", - 384, - query_prefix="query: ", - passage_prefix="", - ), - "e5-small": EmbedderSpec( - "e5-small", - "intfloat/e5-small-v2", - 384, - query_prefix="query: ", - passage_prefix="passage: ", - ), - "e5-small-v2": EmbedderSpec( - "e5-small-v2", - "intfloat/e5-small-v2", - 384, - query_prefix="query: ", - passage_prefix="passage: ", - ), -} - -DEFAULT_EMBEDDER_KEY = "minilm" - - -def resolve_embedder(name: str | None = None) -> EmbedderSpec: - """Resolve an embedder key from CLI, env, or default.""" - raw = (name or os.environ.get("MINDGRAPH_EMBEDDER") or DEFAULT_EMBEDDER_KEY).strip() - key = raw.lower() - if key not in _REGISTRY: - known = ", ".join(sorted(_REGISTRY)) - raise EmbeddingError(f"Unknown embedder {raw!r}; known keys: {known}") - return _REGISTRY[key] - - -def resolve_embed_template(name: str | None = None) -> EmbedTemplate: - raw = ( - name or os.environ.get("MINDGRAPH_EMBED_TEMPLATE") or "none" - ).strip().lower() - if raw in ("none", ""): - return "none" - if raw == "mainframe": - return "mainframe" - raise EmbeddingError( - f"Unknown embed template {raw!r}; known: none, mainframe" - ) - - -def load_sentence_embedder(spec: EmbedderSpec): - """Load a SentenceTransformer from the local Hugging Face cache only. - - MindGraph expects models to already be cached on disk. Resolving via the - hub client can fail with ``RuntimeError: Cannot send a request, as the - client has been closed`` even when the cache is warm, so we never open a - network metadata request at query/ingest time. - """ - from sentence_transformers import SentenceTransformer - - try: - return SentenceTransformer(spec.model_id, local_files_only=True) - except Exception as exc: - raise EmbeddingError( - f"Failed to load cached embedding model {spec.model_id!r} with " - f"local_files_only=True ({type(exc).__name__}: {exc}). " - "Cache the model once while online, for example:\n" - " python -c \"from sentence_transformers import SentenceTransformer; " - f"SentenceTransformer({spec.model_id!r})\"" - ) from exc - - -def format_query_text(spec: EmbedderSpec, text: str, *, template: EmbedTemplate) -> str: - body = text.strip() - if template == "mainframe": - body = f"[intent=query] {body}" - if spec.query_prefix: - body = f"{spec.query_prefix}{body}" - return body - - -def format_passage_text( - spec: EmbedderSpec, - text: str, - *, - template: EmbedTemplate, - title: str | None = None, - domain: str | None = None, - doc_type: str | None = None, -) -> str: - body = text.strip() - if template == "mainframe": - parts: list[str] = [] - if domain: - parts.append(f"[domain={domain}]") - if doc_type: - parts.append(f"[type={doc_type}]") - if title: - parts.append(title) - prefix = " ".join(parts) - body = f"{prefix} — {body}" if prefix else body - if spec.passage_prefix: - body = f"{spec.passage_prefix}{body}" - return body diff --git a/mindgraph/src/mindgraph/exceptions.py b/mindgraph/src/mindgraph/exceptions.py deleted file mode 100644 index 1fe91e6..0000000 --- a/mindgraph/src/mindgraph/exceptions.py +++ /dev/null @@ -1,26 +0,0 @@ -class MindgraphError(Exception): - """Base exception for all MindGraph errors.""" - - -class DatabaseError(MindgraphError): - """Raised when a database connection, schema, or write operation fails.""" - - -class EmbeddingError(MindgraphError, ValueError): - """Raised when an embedding configuration or runtime cannot be used.""" - - -class IngestionError(MindgraphError): - """Raised when ingesting a file fails. Carries the offending path.""" - - def __init__(self, message: str, path: str | None = None): - super().__init__(message) - self.path = path - - def __str__(self) -> str: - base = super().__str__() - return f"{base} (path={self.path})" if self.path else base - - -class ParseError(IngestionError): - """Raised when a file cannot be parsed (frontmatter, structure, etc.).""" diff --git a/mindgraph/src/mindgraph/idle_lifecycle.py b/mindgraph/src/mindgraph/idle_lifecycle.py deleted file mode 100644 index 23b5843..0000000 --- a/mindgraph/src/mindgraph/idle_lifecycle.py +++ /dev/null @@ -1,106 +0,0 @@ -"""Conservative, opt-in idle lifecycle for the shared HTTP daemon.""" - -from __future__ import annotations - -from contextlib import contextmanager -import os -import signal -import threading -import time -import uuid -from collections.abc import Callable, Iterator - - -class IdleLifecycle: - """Track requests and renewable client leases before allowing idle exit.""" - - def __init__( - self, - idle_seconds: float, - *, - lease_ttl: float = 90.0, - clock: Callable[[], float] = time.monotonic, - request_shutdown: Callable[[], None] | None = None, - ) -> None: - if idle_seconds <= 0 or lease_ttl <= 0: - raise ValueError("idle and lease durations must be positive") - self.idle_seconds = idle_seconds - self.lease_ttl = lease_ttl - self._clock = clock - self._request_shutdown = request_shutdown or ( - lambda: os.kill(os.getpid(), signal.SIGTERM) - ) - self._lock = threading.Lock() - self._in_flight = 0 - self._last_activity = clock() - self._leases: dict[str, float] = {} - self._stopped = threading.Event() - self._thread: threading.Thread | None = None - - @contextmanager - def request(self) -> Iterator[None]: - with self._lock: - self._in_flight += 1 - self._last_activity = self._clock() - try: - yield - finally: - with self._lock: - self._in_flight -= 1 - self._last_activity = self._clock() - - def acquire_lease(self) -> str: - token = uuid.uuid4().hex - self.renew_lease(token) - return token - - def renew_lease(self, token: str) -> bool: - if not token: - return False - with self._lock: - self._leases[token] = self._clock() + self.lease_ttl - self._last_activity = self._clock() - return True - - def release_lease(self, token: str) -> None: - with self._lock: - self._leases.pop(token, None) - self._last_activity = self._clock() - - def snapshot(self) -> dict: - now = self._clock() - with self._lock: - self._leases = {key: expiry for key, expiry in self._leases.items() if expiry > now} - return { - "enabled": True, - "idle_seconds": self.idle_seconds, - "in_flight": self._in_flight, - "active_leases": len(self._leases), - "idle_for_seconds": max(0.0, now - self._last_activity), - } - - def should_shutdown(self) -> bool: - state = self.snapshot() - return ( - state["in_flight"] == 0 - and state["active_leases"] == 0 - and state["idle_for_seconds"] >= self.idle_seconds - ) - - def start(self) -> None: - if self._thread is not None: - return - self._thread = threading.Thread(target=self._monitor, daemon=True) - self._thread.start() - - def stop(self) -> None: - self._stopped.set() - if self._thread is not None: - self._thread.join(timeout=2) - - def _monitor(self) -> None: - interval = min(5.0, max(0.1, self.idle_seconds / 10)) - while not self._stopped.wait(interval): - if self.should_shutdown(): - self._request_shutdown() - return diff --git a/mindgraph/src/mindgraph/intent.py b/mindgraph/src/mindgraph/intent.py deleted file mode 100644 index 7cb687d..0000000 --- a/mindgraph/src/mindgraph/intent.py +++ /dev/null @@ -1,1563 +0,0 @@ -"""Deterministic compilation and read-only traversal for V1 intent graphs. - -Reviewed YAML is the source of truth. SQLite files produced here are generated -control-plane artifacts; they are deliberately separate from MindGraph's -document indexes and contain no embeddings or retrieved content. -""" - -from __future__ import annotations - -import hashlib -import json -import os -import re -import sqlite3 -import tempfile -import unicodedata -from dataclasses import dataclass -from pathlib import Path -from typing import Any, Literal -from urllib.parse import quote, urlsplit - -import yaml -from pydantic import BaseModel, ConfigDict, Field, ValidationError, field_validator - -from mindgraph.exceptions import MindgraphError - - -SCHEMA_VERSION = "1" -NODE_KINDS = frozenset({"goal", "capability", "constraint"}) -RELATIONS = frozenset( - {"decomposes_to", "requires", "next_step", "blocked_by", "routes_to"} -) -ACYCLIC_RELATIONS = frozenset({"decomposes_to", "requires", "next_step"}) -BINDING_SCHEMES = frozenset( - {"knowledge", "project", "retriever", "capability", "policy"} -) -ID_RE = re.compile(r"^[a-z][a-z0-9]*(?:[.-][a-z0-9]+)*$") -VERSION_RE = re.compile(r"^[a-z0-9][a-z0-9]*(?:[.-][a-z0-9]+)*$") -TOKEN_RE = re.compile(r"\w+", re.UNICODE) -REQUIRED_TABLES = frozenset( - { - "schema_meta", - "graph_versions", - "intent_nodes", - "intent_aliases", - "intent_edges", - "intent_bindings", - "intent_rules", - } -) - - -class IntentGraphError(MindgraphError): - """Base intent error with a stable machine-readable code.""" - - def __init__( - self, - code: str, - message: str, - *, - location: str | None = None, - witness: tuple[str, ...] = (), - graph_id: str | None = None, - graph_version: str | None = None, - ) -> None: - super().__init__(message) - self.code = code - self.location = location - self.witness = witness - self.graph_id = graph_id - self.graph_version = graph_version - - def __str__(self) -> str: - details = [] - if self.location: - details.append(f"location={self.location}") - if self.witness: - details.append(f"witness={' -> '.join(self.witness)}") - suffix = f" ({', '.join(details)})" if details else "" - return f"{self.code}: {super().__str__()}{suffix}" - - -class IntentGraphValidationError(IntentGraphError): - """Raised when reviewed source violates the V1 graph contract.""" - - -class IntentVersionConflict(IntentGraphError): - """Raised when compilation would mutate approved version history.""" - - -class IntentStoreError(IntentGraphError): - """Raised when a compiled store is missing or violates its schema.""" - - -class _StrictModel(BaseModel): - model_config = ConfigDict(extra="forbid", frozen=True) - - -def _require_id(value: str) -> str: - value = value.strip() - if not ID_RE.fullmatch(value): - raise ValueError("must be a lowercase dot/hyphen-separated stable ID") - return value - - -def _require_nonempty(value: str) -> str: - value = value.strip() - if not value: - raise ValueError("must not be empty") - return value - - -def _require_version(value: str) -> str: - value = value.strip() - if not VERSION_RE.fullmatch(value): - raise ValueError("must be a lowercase dot/hyphen-separated version") - return value - - -def _require_reference(value: str) -> str: - value = value.strip() - parsed = urlsplit(value) - if not parsed.scheme or not parsed.netloc: - raise ValueError("must be a stable scheme://reference") - return value - - -class GraphMetadata(_StrictModel): - id: str - version: str - status: Literal["approved"] - created_at: str - reviewed_at: str - reviewed_by: tuple[str, ...] - source_refs: tuple[str, ...] - supersedes: str | None = None - - _id = field_validator("id")(_require_id) - _version = field_validator("version")(_require_version) - _times = field_validator("created_at", "reviewed_at")(_require_nonempty) - - @field_validator("reviewed_by") - @classmethod - def validate_reviewers(cls, value: tuple[str, ...]) -> tuple[str, ...]: - cleaned = tuple(sorted({_require_nonempty(item) for item in value})) - if not cleaned: - raise ValueError("must contain at least one reviewer") - return cleaned - - @field_validator("source_refs") - @classmethod - def validate_source_refs(cls, value: tuple[str, ...]) -> tuple[str, ...]: - cleaned = tuple(sorted({_require_reference(item) for item in value})) - if not cleaned: - raise ValueError("must contain at least one source reference") - return cleaned - - @field_validator("supersedes") - @classmethod - def validate_supersedes(cls, value: str | None) -> str | None: - return _require_version(value) if value is not None else None - - -class IntentNodeSource(_StrictModel): - id: str - kind: Literal["goal", "capability", "constraint"] - label: str - aliases: tuple[str, ...] = () - status: Literal["active", "deprecated"] = "active" - source_refs: tuple[str, ...] - - _id = field_validator("id")(_require_id) - _label = field_validator("label")(_require_nonempty) - - @field_validator("aliases") - @classmethod - def validate_aliases(cls, value: tuple[str, ...]) -> tuple[str, ...]: - return tuple(_require_nonempty(item) for item in value) - - @field_validator("source_refs") - @classmethod - def validate_source_refs(cls, value: tuple[str, ...]) -> tuple[str, ...]: - cleaned = tuple(sorted({_require_reference(item) for item in value})) - if not cleaned: - raise ValueError("must contain at least one source reference") - return cleaned - - -class IntentEdgeSource(_StrictModel): - source_id: str - target_id: str - relation: str - status: Literal["active", "deprecated"] = "active" - source_refs: tuple[str, ...] - - _ids = field_validator("source_id", "target_id")(_require_id) - _relation = field_validator("relation")(_require_nonempty) - - @field_validator("source_refs") - @classmethod - def validate_source_refs(cls, value: tuple[str, ...]) -> tuple[str, ...]: - cleaned = tuple(sorted({_require_reference(item) for item in value})) - if not cleaned: - raise ValueError("must contain at least one source reference") - return cleaned - - -class IntentBindingSource(_StrictModel): - id: str - node_id: str - ref: str - required: bool = False - availability: Literal["available", "unavailable"] = "available" - source_refs: tuple[str, ...] - - _ids = field_validator("id", "node_id")(_require_id) - - @field_validator("ref") - @classmethod - def validate_ref(cls, value: str) -> str: - value = _require_reference(value) - if urlsplit(value).scheme not in BINDING_SCHEMES: - raise ValueError( - "binding scheme must be one of " + ", ".join(sorted(BINDING_SCHEMES)) - ) - return value - - @field_validator("source_refs") - @classmethod - def validate_source_refs(cls, value: tuple[str, ...]) -> tuple[str, ...]: - cleaned = tuple(sorted({_require_reference(item) for item in value})) - if not cleaned: - raise ValueError("must contain at least one source reference") - return cleaned - - -class IntentRuleMatch(_StrictModel): - scope: str | None = None - all_terms: tuple[str, ...] = () - any_terms: tuple[str, ...] = () - - @field_validator("scope") - @classmethod - def validate_scope(cls, value: str | None) -> str | None: - return normalize_text(value) if value is not None else None - - @field_validator("all_terms", "any_terms") - @classmethod - def validate_terms(cls, value: tuple[str, ...]) -> tuple[str, ...]: - cleaned = tuple(sorted({normalize_text(item) for item in value})) - if any(not item for item in cleaned): - raise ValueError("rule terms must not be empty") - return cleaned - - -class IntentRuleSource(_StrictModel): - id: str - priority: int = Field(ge=0, le=1000) - goal_id: str - match: IntentRuleMatch - source_refs: tuple[str, ...] - - _ids = field_validator("id", "goal_id")(_require_id) - - @field_validator("source_refs") - @classmethod - def validate_source_refs(cls, value: tuple[str, ...]) -> tuple[str, ...]: - cleaned = tuple(sorted({_require_reference(item) for item in value})) - if not cleaned: - raise ValueError("must contain at least one source reference") - return cleaned - - -class IntentGraphDocument(_StrictModel): - schema_version: Literal["1"] - graph: GraphMetadata - nodes: tuple[IntentNodeSource, ...] - edges: tuple[IntentEdgeSource, ...] = () - bindings: tuple[IntentBindingSource, ...] = () - rules: tuple[IntentRuleSource, ...] = () - - -class VersionHash(_StrictModel): - graph_id: str - version: str - source_hash: str - - -class CompileResult(_StrictModel): - destination: Path - graph_id: str - current_version: str - corpus_hash: str - version_hashes: tuple[VersionHash, ...] - version_count: int - node_count: int - edge_count: int - binding_count: int - rule_count: int - replaced: bool - - -class TraversalLimits(_StrictModel): - max_depth: int = Field(default=2, ge=0, le=32) - max_nodes: int = Field(default=64, ge=1, le=1000) - - -class IntentTraceEdge(_StrictModel): - source_id: str - target_id: str - relation: str - - -class IntentResolution(_StrictModel): - schema_version: Literal["1"] = "1" - graph_id: str - graph_version: str - source_hash: str - outcome: Literal["resolved", "fallback", "refusal"] - resolution_method: Literal["explicit", "alias", "rule", "none"] - matched_goal_ids: tuple[str, ...] = () - prerequisite_goal_ids: tuple[str, ...] = () - intent_path: tuple[str, ...] = () - edge_path: tuple[IntentTraceEdge, ...] = () - capability_hints: tuple[str, ...] = () - rejected_capability_hints: tuple[str, ...] = () - constraint_ids: tuple[str, ...] = () - warnings: tuple[str, ...] = () - truncation_reason: Literal["max_depth", "max_nodes"] | None = None - refusal_reason: str | None = None - - def as_contract_result( - self, result_id: str, behaviors: tuple[str, ...] = () - ) -> dict[str, Any]: - """Return the additive candidate shape used by the Phase 1 evaluator.""" - behavior_set = set(behaviors) - if self.refusal_reason == "intent_no_match": - behavior_set.update({"require explicit scope", "do not query all stores"}) - if self.truncation_reason: - behavior_set.update({"return truncation warning", "preserve visited path"}) - if self.rejected_capability_hints: - behavior_set.add("policy rejected untrusted capability hint") - - payload: dict[str, Any] = { - "id": result_id, - "outcome": self.outcome, - "resolution_method": self.resolution_method, - "matched_goal_ids": list(self.matched_goal_ids), - "prerequisite_goal_ids": list(self.prerequisite_goal_ids), - "intent_path": list(self.intent_path), - "capability_hints": list(self.capability_hints), - "behaviors": sorted(behavior_set), - "fields": sorted(type(self).model_fields), - "graph_id": self.graph_id, - "graph_version": self.graph_version, - } - reason = self.refusal_reason - if self.truncation_reason: - reason = "intent_traversal_truncated" - if reason: - payload["reason"] = reason - return payload - - -@dataclass(frozen=True) -class _ValidatedVersion: - path: Path - document: IntentGraphDocument - canonical_bytes: bytes - source_hash: str - effective_status: str = "approved" - - -@dataclass(frozen=True) -class _ValidatedCorpus: - graph_id: str - current_version: str - corpus_hash: str - versions: tuple[_ValidatedVersion, ...] - - -def normalize_text(value: str) -> str: - """Normalize exact-match text without discarding its word content.""" - normalized = unicodedata.normalize("NFKC", value) - return " ".join(normalized.strip().split()).casefold() - - -def _tokens(value: str) -> frozenset[str]: - return frozenset(TOKEN_RE.findall(normalize_text(value))) - - -def _json(value: Any) -> str: - return json.dumps(value, sort_keys=True, separators=(",", ":"), ensure_ascii=False) - - -def _canonical_payload(document: IntentGraphDocument) -> dict[str, Any]: - payload = document.model_dump(mode="json") - payload["graph"]["reviewed_by"] = sorted(payload["graph"]["reviewed_by"]) - payload["graph"]["source_refs"] = sorted(payload["graph"]["source_refs"]) - for node in payload["nodes"]: - node["aliases"] = sorted(node["aliases"], key=normalize_text) - node["source_refs"] = sorted(node["source_refs"]) - for edge in payload["edges"]: - edge["source_refs"] = sorted(edge["source_refs"]) - for binding in payload["bindings"]: - binding["source_refs"] = sorted(binding["source_refs"]) - for rule in payload["rules"]: - rule["match"]["all_terms"] = sorted(rule["match"]["all_terms"]) - rule["match"]["any_terms"] = sorted(rule["match"]["any_terms"]) - rule["source_refs"] = sorted(rule["source_refs"]) - payload["nodes"] = sorted(payload["nodes"], key=lambda item: item["id"]) - payload["edges"] = sorted( - payload["edges"], - key=lambda item: (item["source_id"], item["target_id"], item["relation"]), - ) - payload["bindings"] = sorted(payload["bindings"], key=lambda item: item["id"]) - payload["rules"] = sorted(payload["rules"], key=lambda item: item["id"]) - return payload - - -def _canonical_bytes(document: IntentGraphDocument) -> bytes: - return (_json(_canonical_payload(document)) + "\n").encode("utf-8") - - -def _validation_error( - code: str, - message: str, - *, - document: IntentGraphDocument | None = None, - location: str | None = None, - witness: tuple[str, ...] = (), -) -> IntentGraphValidationError: - return IntentGraphValidationError( - code, - message, - location=location, - witness=witness, - graph_id=document.graph.id if document else None, - graph_version=document.graph.version if document else None, - ) - - -def _load_document(path: Path) -> IntentGraphDocument: - try: - raw = yaml.safe_load(path.read_text(encoding="utf-8")) - except (OSError, yaml.YAMLError) as exc: - raise IntentGraphValidationError( - "intent_schema_invalid", f"cannot read YAML source: {exc}", location=path.name - ) from exc - if not isinstance(raw, dict): - raise IntentGraphValidationError( - "intent_schema_invalid", "source must be a YAML mapping", location=path.name - ) - try: - return IntentGraphDocument.model_validate(raw) - except ValidationError as exc: - errors = sorted( - exc.errors(), - key=lambda item: (tuple(str(part) for part in item["loc"]), item["type"]), - ) - first = errors[0] - field = ".".join(str(part) for part in first["loc"]) - raise IntentGraphValidationError( - "intent_schema_invalid", - first["msg"], - location=f"{path.name}:{field}", - ) from exc - - -def _duplicate(values: list[str]) -> str | None: - seen: set[str] = set() - for value in values: - if value in seen: - return value - seen.add(value) - return None - - -def _cycle_witness(adjacency: dict[str, tuple[str, ...]]) -> tuple[str, ...] | None: - visited: set[str] = set() - active: list[str] = [] - active_set: set[str] = set() - - def visit(node: str) -> tuple[str, ...] | None: - visited.add(node) - active.append(node) - active_set.add(node) - for target in adjacency.get(node, ()): - if target in active_set: - start = active.index(target) - return tuple(active[start:] + [target]) - if target not in visited: - witness = visit(target) - if witness: - return witness - active.pop() - active_set.remove(node) - return None - - for node in sorted(adjacency): - if node not in visited: - witness = visit(node) - if witness: - return witness - return None - - -def _validate_document(document: IntentGraphDocument, path: Path) -> None: - node_ids = [node.id for node in sorted(document.nodes, key=lambda item: item.id)] - duplicate = _duplicate(node_ids) - if duplicate: - raise _validation_error( - "intent_duplicate_id", - f"duplicate node ID {duplicate}", - document=document, - location=path.name, - witness=(duplicate,), - ) - if not node_ids: - raise _validation_error( - "intent_schema_invalid", - "graph version must contain at least one node", - document=document, - location=path.name, - ) - - node_map = {node.id: node for node in document.nodes} - for node in sorted(document.nodes, key=lambda item: item.id): - if not node.id.startswith(f"{node.kind}."): - raise _validation_error( - "intent_kind_mismatch", - f"node {node.id} must use the {node.kind}. prefix", - document=document, - location=f"{path.name}:nodes.{node.id}", - witness=(node.id,), - ) - if node.kind != "goal" and node.aliases: - raise _validation_error( - "intent_kind_mismatch", - f"only goals may declare aliases: {node.id}", - document=document, - location=f"{path.name}:nodes.{node.id}.aliases", - witness=(node.id,), - ) - - alias_owners: dict[str, str] = {} - for node in sorted(document.nodes, key=lambda item: item.id): - if node.kind != "goal" or node.status != "active": - continue - for alias in (node.label, *node.aliases): - normalized = normalize_text(alias) - owner = alias_owners.get(normalized) - if owner and owner != node.id: - raise _validation_error( - "intent_alias_ambiguous", - f"normalized label/alias {normalized!r} identifies multiple goals", - document=document, - location=f"{path.name}:nodes", - witness=(owner, node.id), - ) - if owner == node.id: - raise _validation_error( - "intent_alias_ambiguous", - f"goal {node.id} repeats normalized label/alias {normalized!r}", - document=document, - location=f"{path.name}:nodes.{node.id}.aliases", - witness=(node.id,), - ) - alias_owners[normalized] = node.id - - edge_keys = [ - (edge.source_id, edge.target_id, edge.relation) - for edge in sorted( - document.edges, - key=lambda item: (item.source_id, item.target_id, item.relation), - ) - ] - duplicate_edge = _duplicate(["\x1f".join(key) for key in edge_keys]) - if duplicate_edge: - raise _validation_error( - "intent_duplicate_id", - "duplicate edge", - document=document, - location=f"{path.name}:edges", - witness=tuple(duplicate_edge.split("\x1f")), - ) - - expected_kinds = { - "decomposes_to": ("goal", "goal"), - "requires": ("goal", "goal"), - "next_step": ("goal", "goal"), - "blocked_by": ("goal", "constraint"), - "routes_to": ("goal", "capability"), - } - for edge in sorted( - document.edges, - key=lambda item: (item.source_id, item.target_id, item.relation), - ): - if edge.relation not in RELATIONS: - raise _validation_error( - "intent_relation_invalid", - f"unknown relation {edge.relation}", - document=document, - location=f"{path.name}:edges", - witness=(edge.source_id, edge.target_id), - ) - missing = [item for item in (edge.source_id, edge.target_id) if item not in node_map] - if missing: - raise _validation_error( - "intent_missing_target", - f"edge references missing node {missing[0]}", - document=document, - location=f"{path.name}:edges", - witness=(edge.source_id, edge.target_id), - ) - actual = (node_map[edge.source_id].kind, node_map[edge.target_id].kind) - if actual != expected_kinds[edge.relation]: - raise _validation_error( - "intent_kind_mismatch", - f"{edge.relation} requires {expected_kinds[edge.relation]}, got {actual}", - document=document, - location=f"{path.name}:edges", - witness=(edge.source_id, edge.target_id), - ) - - for relation in sorted(ACYCLIC_RELATIONS): - adjacency: dict[str, list[str]] = {} - for edge in document.edges: - if edge.status == "active" and edge.relation == relation: - adjacency.setdefault(edge.source_id, []).append(edge.target_id) - ordered = {key: tuple(sorted(value)) for key, value in adjacency.items()} - witness = _cycle_witness(ordered) - if witness: - raise _validation_error( - "intent_cycle_detected", - f"{relation} contains a cycle", - document=document, - location=f"{path.name}:edges", - witness=witness, - ) - - binding_ids = [item.id for item in sorted(document.bindings, key=lambda item: item.id)] - duplicate = _duplicate(binding_ids) - if duplicate: - raise _validation_error( - "intent_duplicate_id", - f"duplicate binding ID {duplicate}", - document=document, - location=f"{path.name}:bindings", - witness=(duplicate,), - ) - for binding in sorted(document.bindings, key=lambda item: item.id): - if binding.node_id not in node_map: - raise _validation_error( - "intent_missing_target", - f"binding references missing node {binding.node_id}", - document=document, - location=f"{path.name}:bindings.{binding.id}", - witness=(binding.node_id,), - ) - if binding.required and binding.availability == "unavailable": - raise _validation_error( - "intent_binding_unavailable", - f"required binding {binding.ref} is unavailable", - document=document, - location=f"{path.name}:bindings.{binding.id}", - witness=(binding.node_id, binding.ref), - ) - - rule_ids = [item.id for item in sorted(document.rules, key=lambda item: item.id)] - duplicate = _duplicate(rule_ids) - if duplicate: - raise _validation_error( - "intent_duplicate_id", - f"duplicate rule ID {duplicate}", - document=document, - location=f"{path.name}:rules", - witness=(duplicate,), - ) - for rule in sorted(document.rules, key=lambda item: item.id): - goal = node_map.get(rule.goal_id) - if goal is None: - raise _validation_error( - "intent_missing_target", - f"rule references missing goal {rule.goal_id}", - document=document, - location=f"{path.name}:rules.{rule.id}", - witness=(rule.goal_id,), - ) - if goal.kind != "goal": - raise _validation_error( - "intent_kind_mismatch", - f"rule target {rule.goal_id} is not a goal", - document=document, - location=f"{path.name}:rules.{rule.id}", - witness=(rule.goal_id,), - ) - match = rule.match - if not match.scope and not match.all_terms and not match.any_terms: - raise _validation_error( - "intent_schema_invalid", - f"rule {rule.id} has no match condition", - document=document, - location=f"{path.name}:rules.{rule.id}.match", - witness=(rule.id,), - ) - - -def validate_intent_corpus(source_dir: Path) -> _ValidatedCorpus: - """Load and validate one append-only graph-family corpus.""" - source_dir = Path(source_dir) - if not source_dir.is_dir(): - raise IntentGraphValidationError( - "intent_schema_invalid", - "source directory does not exist", - location=str(source_dir), - ) - paths = sorted(path for path in source_dir.glob("*.yaml") if path.is_file()) - if not paths: - raise IntentGraphValidationError( - "intent_schema_invalid", - "source corpus contains no direct *.yaml files", - location=str(source_dir), - ) - - versions: list[_ValidatedVersion] = [] - for path in paths: - document = _load_document(path) - _validate_document(document, path) - canonical = _canonical_bytes(document) - versions.append( - _ValidatedVersion( - path=path, - document=document, - canonical_bytes=canonical, - source_hash=hashlib.sha256(canonical).hexdigest(), - ) - ) - - graph_ids = sorted({version.document.graph.id for version in versions}) - if len(graph_ids) != 1: - raise IntentGraphValidationError( - "intent_version_fork", - "a source corpus must contain exactly one graph family", - witness=tuple(graph_ids), - ) - graph_id = graph_ids[0] - - by_version: dict[str, _ValidatedVersion] = {} - for item in sorted(versions, key=lambda version: version.document.graph.version): - version = item.document.graph.version - if version in by_version: - raise _validation_error( - "intent_duplicate_id", - f"duplicate graph version {version}", - document=item.document, - location=item.path.name, - witness=(version,), - ) - by_version[version] = item - - successors: dict[str, list[str]] = {version: [] for version in by_version} - roots: list[str] = [] - for version, item in sorted(by_version.items()): - predecessor = item.document.graph.supersedes - if predecessor is None: - roots.append(version) - continue - if predecessor not in by_version: - raise _validation_error( - "intent_version_missing", - f"version {version} supersedes missing version {predecessor}", - document=item.document, - location=item.path.name, - witness=(predecessor, version), - ) - successors[predecessor].append(version) - - forks = sorted(version for version, items in successors.items() if len(items) > 1) - if len(roots) != 1 or forks: - witness = tuple(sorted(roots + forks)) - raise IntentGraphValidationError( - "intent_version_fork", - "version lineage must be one unbranched chain", - graph_id=graph_id, - witness=witness, - ) - - ordered_names: list[str] = [] - current = roots[0] - seen: set[str] = set() - while current not in seen: - seen.add(current) - ordered_names.append(current) - next_versions = successors[current] - if not next_versions: - break - current = next_versions[0] - if len(seen) != len(by_version): - raise IntentGraphValidationError( - "intent_version_fork", - "version lineage contains a disconnected chain or cycle", - graph_id=graph_id, - witness=tuple(sorted(set(by_version) - seen)), - ) - - ordered: list[_ValidatedVersion] = [] - for name in ordered_names: - item = by_version[name] - ordered.append( - _ValidatedVersion( - path=item.path, - document=item.document, - canonical_bytes=item.canonical_bytes, - source_hash=item.source_hash, - effective_status="approved" if name == ordered_names[-1] else "superseded", - ) - ) - corpus_payload = [ - { - "graph_id": item.document.graph.id, - "version": item.document.graph.version, - "source_hash": item.source_hash, - } - for item in ordered - ] - corpus_hash = hashlib.sha256((_json(corpus_payload) + "\n").encode("utf-8")).hexdigest() - return _ValidatedCorpus( - graph_id=graph_id, - current_version=ordered_names[-1], - corpus_hash=corpus_hash, - versions=tuple(ordered), - ) - - -def _existing_version_hashes(destination: Path) -> tuple[str, dict[str, str]]: - try: - conn = _connect_read_only(destination) - try: - _validate_store(conn) - graph_row = conn.execute( - "SELECT value FROM schema_meta WHERE key = 'graph_id'" - ).fetchone() - rows = conn.execute( - "SELECT version, source_hash FROM graph_versions ORDER BY version" - ).fetchall() - finally: - conn.close() - except (sqlite3.Error, IntentStoreError) as exc: - raise IntentStoreError( - "intent_store_invalid", - f"cannot inspect existing destination: {exc}", - location=str(destination), - ) from exc - if graph_row is None: - raise IntentStoreError( - "intent_store_invalid", - "existing destination has no graph_id", - location=str(destination), - ) - return graph_row["value"], {row["version"]: row["source_hash"] for row in rows} - - -def _check_existing_history(destination: Path, corpus: _ValidatedCorpus) -> None: - if not destination.exists(): - return - graph_id, existing = _existing_version_hashes(destination) - if graph_id != corpus.graph_id: - raise IntentVersionConflict( - "intent_version_conflict", - f"existing graph {graph_id} cannot be replaced by {corpus.graph_id}", - location=str(destination), - graph_id=corpus.graph_id, - ) - proposed = { - item.document.graph.version: item.source_hash for item in corpus.versions - } - for version, source_hash in sorted(existing.items()): - if version not in proposed: - raise IntentVersionConflict( - "intent_version_missing", - f"approved version {version} was removed from the source corpus", - location=str(destination), - witness=(version,), - graph_id=graph_id, - graph_version=version, - ) - if proposed[version] != source_hash: - raise IntentVersionConflict( - "intent_version_conflict", - f"approved version {version} changed source hash", - location=str(destination), - witness=(version,), - graph_id=graph_id, - graph_version=version, - ) - - -SCHEMA_SQL = """ -CREATE TABLE schema_meta ( - key TEXT PRIMARY KEY, - value TEXT NOT NULL -) WITHOUT ROWID; -CREATE TABLE graph_versions ( - graph_id TEXT NOT NULL, - version TEXT NOT NULL, - effective_status TEXT NOT NULL, - schema_version TEXT NOT NULL, - source_hash TEXT NOT NULL, - source_path TEXT NOT NULL, - created_at TEXT NOT NULL, - reviewed_at TEXT NOT NULL, - reviewed_by_json TEXT NOT NULL, - supersedes_version TEXT, - PRIMARY KEY (graph_id, version) -) WITHOUT ROWID; -CREATE TABLE intent_nodes ( - graph_id TEXT NOT NULL, - version TEXT NOT NULL, - node_id TEXT NOT NULL, - kind TEXT NOT NULL, - label TEXT NOT NULL, - normalized_label TEXT NOT NULL, - aliases_json TEXT NOT NULL, - status TEXT NOT NULL, - source_refs_json TEXT NOT NULL, - PRIMARY KEY (graph_id, version, node_id), - FOREIGN KEY (graph_id, version) REFERENCES graph_versions(graph_id, version) -) WITHOUT ROWID; -CREATE TABLE intent_aliases ( - graph_id TEXT NOT NULL, - version TEXT NOT NULL, - normalized_alias TEXT NOT NULL, - node_id TEXT NOT NULL, - alias TEXT NOT NULL, - PRIMARY KEY (graph_id, version, normalized_alias), - FOREIGN KEY (graph_id, version, node_id) - REFERENCES intent_nodes(graph_id, version, node_id) -) WITHOUT ROWID; -CREATE TABLE intent_edges ( - graph_id TEXT NOT NULL, - version TEXT NOT NULL, - source_id TEXT NOT NULL, - target_id TEXT NOT NULL, - relation TEXT NOT NULL, - status TEXT NOT NULL, - source_refs_json TEXT NOT NULL, - PRIMARY KEY (graph_id, version, source_id, target_id, relation), - FOREIGN KEY (graph_id, version, source_id) - REFERENCES intent_nodes(graph_id, version, node_id), - FOREIGN KEY (graph_id, version, target_id) - REFERENCES intent_nodes(graph_id, version, node_id) -) WITHOUT ROWID; -CREATE TABLE intent_bindings ( - graph_id TEXT NOT NULL, - version TEXT NOT NULL, - binding_id TEXT NOT NULL, - node_id TEXT NOT NULL, - ref TEXT NOT NULL, - required INTEGER NOT NULL, - availability TEXT NOT NULL, - source_refs_json TEXT NOT NULL, - PRIMARY KEY (graph_id, version, binding_id), - FOREIGN KEY (graph_id, version, node_id) - REFERENCES intent_nodes(graph_id, version, node_id) -) WITHOUT ROWID; -CREATE TABLE intent_rules ( - graph_id TEXT NOT NULL, - version TEXT NOT NULL, - rule_id TEXT NOT NULL, - priority INTEGER NOT NULL, - goal_id TEXT NOT NULL, - scope TEXT, - all_terms_json TEXT NOT NULL, - any_terms_json TEXT NOT NULL, - source_refs_json TEXT NOT NULL, - PRIMARY KEY (graph_id, version, rule_id), - FOREIGN KEY (graph_id, version, goal_id) - REFERENCES intent_nodes(graph_id, version, node_id) -) WITHOUT ROWID; -""" - - -def _build_database(path: Path, corpus: _ValidatedCorpus) -> None: - conn = sqlite3.connect(str(path)) - conn.row_factory = sqlite3.Row - try: - conn.execute("PRAGMA page_size = 4096") - conn.execute("PRAGMA encoding = 'UTF-8'") - conn.execute("PRAGMA auto_vacuum = NONE") - conn.execute("PRAGMA journal_mode = OFF") - conn.execute("PRAGMA synchronous = OFF") - conn.execute("PRAGMA foreign_keys = ON") - conn.executescript(SCHEMA_SQL) - meta = { - "schema_version": SCHEMA_VERSION, - "graph_id": corpus.graph_id, - "current_version": corpus.current_version, - "corpus_hash": corpus.corpus_hash, - } - conn.executemany( - "INSERT INTO schema_meta(key, value) VALUES (?, ?)", sorted(meta.items()) - ) - for item in corpus.versions: - graph = item.document.graph - conn.execute( - """ - INSERT INTO graph_versions VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?) - """, - ( - graph.id, - graph.version, - item.effective_status, - SCHEMA_VERSION, - item.source_hash, - item.path.name, - graph.created_at, - graph.reviewed_at, - _json(list(graph.reviewed_by)), - graph.supersedes, - ), - ) - for node in sorted(item.document.nodes, key=lambda value: value.id): - aliases = sorted(node.aliases, key=normalize_text) - conn.execute( - "INSERT INTO intent_nodes VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)", - ( - graph.id, - graph.version, - node.id, - node.kind, - node.label, - normalize_text(node.label), - _json(aliases), - node.status, - _json(list(node.source_refs)), - ), - ) - if node.kind == "goal" and node.status == "active": - for alias in (node.label, *aliases): - conn.execute( - "INSERT INTO intent_aliases VALUES (?, ?, ?, ?, ?)", - (graph.id, graph.version, normalize_text(alias), node.id, alias), - ) - for edge in sorted( - item.document.edges, - key=lambda value: (value.source_id, value.target_id, value.relation), - ): - conn.execute( - "INSERT INTO intent_edges VALUES (?, ?, ?, ?, ?, ?, ?)", - ( - graph.id, - graph.version, - edge.source_id, - edge.target_id, - edge.relation, - edge.status, - _json(list(edge.source_refs)), - ), - ) - for binding in sorted(item.document.bindings, key=lambda value: value.id): - conn.execute( - "INSERT INTO intent_bindings VALUES (?, ?, ?, ?, ?, ?, ?, ?)", - ( - graph.id, - graph.version, - binding.id, - binding.node_id, - binding.ref, - int(binding.required), - binding.availability, - _json(list(binding.source_refs)), - ), - ) - for rule in sorted(item.document.rules, key=lambda value: value.id): - conn.execute( - "INSERT INTO intent_rules VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)", - ( - graph.id, - graph.version, - rule.id, - rule.priority, - rule.goal_id, - rule.match.scope, - _json(list(rule.match.all_terms)), - _json(list(rule.match.any_terms)), - _json(list(rule.source_refs)), - ), - ) - conn.commit() - if conn.execute("PRAGMA foreign_key_check").fetchall(): - raise IntentStoreError( - "intent_store_invalid", "compiled store failed foreign_key_check" - ) - integrity = conn.execute("PRAGMA integrity_check").fetchone()[0] - if integrity != "ok": - raise IntentStoreError( - "intent_store_invalid", f"compiled store failed integrity_check: {integrity}" - ) - conn.execute("VACUUM") - integrity = conn.execute("PRAGMA integrity_check").fetchone()[0] - if integrity != "ok": - raise IntentStoreError( - "intent_store_invalid", f"vacuumed store failed integrity_check: {integrity}" - ) - except Exception: - conn.close() - raise - else: - conn.close() - - -def _fsync(path: Path) -> None: - with path.open("rb") as handle: - os.fsync(handle.fileno()) - try: - directory_fd = os.open(path.parent, os.O_RDONLY) - except OSError: - return - try: - os.fsync(directory_fd) - finally: - os.close(directory_fd) - - -def compile_intent_corpus(source_dir: Path, destination: Path) -> CompileResult: - """Validate and atomically compile a reviewed intent corpus.""" - source_dir = Path(source_dir) - destination = Path(destination) - corpus = validate_intent_corpus(source_dir) - _check_existing_history(destination, corpus) - destination.parent.mkdir(parents=True, exist_ok=True) - replaced = destination.exists() - fd, temp_name = tempfile.mkstemp( - prefix=f".{destination.name}.", suffix=".tmp", dir=destination.parent - ) - os.close(fd) - temp_path = Path(temp_name) - temp_path.unlink() - try: - _build_database(temp_path, corpus) - os.chmod(temp_path, 0o600) - check = open_intent_store(temp_path) - check.close() - _fsync(temp_path) - os.replace(temp_path, destination) - _fsync(destination) - finally: - if temp_path.exists(): - temp_path.unlink() - - return CompileResult( - destination=destination, - graph_id=corpus.graph_id, - current_version=corpus.current_version, - corpus_hash=corpus.corpus_hash, - version_hashes=tuple( - VersionHash( - graph_id=item.document.graph.id, - version=item.document.graph.version, - source_hash=item.source_hash, - ) - for item in corpus.versions - ), - version_count=len(corpus.versions), - node_count=sum(len(item.document.nodes) for item in corpus.versions), - edge_count=sum(len(item.document.edges) for item in corpus.versions), - binding_count=sum(len(item.document.bindings) for item in corpus.versions), - rule_count=sum(len(item.document.rules) for item in corpus.versions), - replaced=replaced, - ) - - -def _connect_read_only(path: Path) -> sqlite3.Connection: - resolved = Path(path).expanduser().resolve() - if not resolved.is_file(): - raise IntentStoreError( - "intent_store_invalid", "intent store does not exist", location=str(resolved) - ) - uri = f"file:{quote(str(resolved), safe='/')}?mode=ro&immutable=1" - conn = sqlite3.connect(uri, uri=True) - conn.row_factory = sqlite3.Row - conn.execute("PRAGMA foreign_keys = ON") - conn.execute("PRAGMA query_only = ON") - return conn - - -def _validate_store(conn: sqlite3.Connection) -> None: - tables = { - row["name"] - for row in conn.execute( - "SELECT name FROM sqlite_master WHERE type = 'table'" - ).fetchall() - } - missing = sorted(REQUIRED_TABLES - tables) - if missing: - raise IntentStoreError( - "intent_store_invalid", f"intent store is missing tables: {', '.join(missing)}" - ) - meta = { - row["key"]: row["value"] - for row in conn.execute("SELECT key, value FROM schema_meta").fetchall() - } - required_meta = {"schema_version", "graph_id", "current_version", "corpus_hash"} - if required_meta - set(meta) or meta.get("schema_version") != SCHEMA_VERSION: - raise IntentStoreError( - "intent_store_invalid", "intent store metadata is missing or incompatible" - ) - integrity = conn.execute("PRAGMA integrity_check").fetchone()[0] - if integrity != "ok": - raise IntentStoreError( - "intent_store_invalid", f"intent store failed integrity_check: {integrity}" - ) - if conn.execute("PRAGMA foreign_key_check").fetchall(): - raise IntentStoreError( - "intent_store_invalid", "intent store failed foreign_key_check" - ) - heads = conn.execute( - "SELECT version FROM graph_versions WHERE effective_status = 'approved'" - ).fetchall() - if len(heads) != 1 or heads[0]["version"] != meta["current_version"]: - raise IntentStoreError( - "intent_store_invalid", "intent store must contain exactly one current version" - ) - - -def open_intent_store(path: Path) -> sqlite3.Connection: - """Open and validate a compiled intent store without creating or writing it.""" - conn = _connect_read_only(Path(path)) - try: - _validate_store(conn) - except Exception: - conn.close() - raise - return conn - - -def _metadata(conn: sqlite3.Connection) -> tuple[str, str, str]: - meta = { - row["key"]: row["value"] - for row in conn.execute("SELECT key, value FROM schema_meta").fetchall() - } - version = meta["current_version"] - row = conn.execute( - "SELECT source_hash FROM graph_versions WHERE graph_id = ? AND version = ?", - (meta["graph_id"], version), - ).fetchone() - if row is None: - raise IntentStoreError( - "intent_store_invalid", "current graph version is missing" - ) - return meta["graph_id"], version, row["source_hash"] - - -def _base_resolution( - graph_id: str, - graph_version: str, - source_hash: str, - *, - outcome: Literal["resolved", "fallback", "refusal"], - method: Literal["explicit", "alias", "rule", "none"], - goal_id: str | None = None, - refusal_reason: str | None = None, - path: tuple[str, ...] = (), -) -> IntentResolution: - return IntentResolution( - graph_id=graph_id, - graph_version=graph_version, - source_hash=source_hash, - outcome=outcome, - resolution_method=method, - matched_goal_ids=(goal_id,) if goal_id else (), - intent_path=path, - refusal_reason=refusal_reason, - ) - - -def _resolve_rule( - conn: sqlite3.Connection, - graph_id: str, - version: str, - query_text: str, - scope: str | None, -) -> tuple[str | None, bool]: - query_tokens = _tokens(query_text) - normalized_scope = normalize_text(scope) if scope is not None else None - matches: list[tuple[int, str, str]] = [] - rows = conn.execute( - """ - SELECT rule_id, priority, goal_id, scope, all_terms_json, any_terms_json - FROM intent_rules WHERE graph_id = ? AND version = ? - ORDER BY priority DESC, rule_id ASC - """, - (graph_id, version), - ).fetchall() - for row in rows: - if row["scope"] is not None and row["scope"] != normalized_scope: - continue - all_terms = json.loads(row["all_terms_json"]) - any_terms = json.loads(row["any_terms_json"]) - if any(not _tokens(term).issubset(query_tokens) for term in all_terms): - continue - if any_terms and not any(_tokens(term).issubset(query_tokens) for term in any_terms): - continue - matches.append((row["priority"], row["rule_id"], row["goal_id"])) - if not matches: - return None, False - top_priority = matches[0][0] - top = [item for item in matches if item[0] == top_priority] - goals = sorted({item[2] for item in top}) - if len(goals) > 1: - return None, True - return goals[0], False - - -def _runtime_cycle( - conn: sqlite3.Connection, graph_id: str, version: str, relation: str -) -> tuple[str, ...] | None: - rows = conn.execute( - """ - SELECT source_id, target_id FROM intent_edges - WHERE graph_id = ? AND version = ? AND relation = ? AND status = 'active' - ORDER BY source_id, target_id - """, - (graph_id, version, relation), - ).fetchall() - adjacency: dict[str, list[str]] = {} - for row in rows: - adjacency.setdefault(row["source_id"], []).append(row["target_id"]) - return _cycle_witness( - {source: tuple(sorted(targets)) for source, targets in adjacency.items()} - ) - - -def _traverse_relation( - conn: sqlite3.Connection, - graph_id: str, - version: str, - root: str, - relation: Literal["decomposes_to", "requires", "next_step"], - limits: TraversalLimits, -) -> tuple[ - tuple[str, ...], - tuple[str, ...], - tuple[IntentTraceEdge, ...], - Literal["max_depth", "max_nodes"] | None, -]: - path = [root] - prerequisites: list[str] = [] - edge_path: list[IntentTraceEdge] = [] - visited = {root} - truncation: Literal["max_depth", "max_nodes"] | None = None - - def visit(node: str, depth: int) -> None: - nonlocal truncation - rows = conn.execute( - """ - SELECT target_id FROM intent_edges - WHERE graph_id = ? AND version = ? AND source_id = ? - AND relation = ? AND status = 'active' - ORDER BY target_id - """, - (graph_id, version, node, relation), - ).fetchall() - for row in rows: - target = row["target_id"] - if target in visited: - continue - if depth >= limits.max_depth: - truncation = truncation or "max_depth" - continue - if len(visited) >= limits.max_nodes: - truncation = truncation or "max_nodes" - continue - visited.add(target) - path.append(target) - prerequisites.append(target) - edge_path.append( - IntentTraceEdge(source_id=node, target_id=target, relation=relation) - ) - visit(target, depth + 1) - - visit(root, 0) - return tuple(path), tuple(prerequisites), tuple(edge_path), truncation - - -def _collect_constraints( - conn: sqlite3.Connection, graph_id: str, version: str, goals: tuple[str, ...] -) -> tuple[str, ...]: - constraints: set[str] = set() - for goal in goals: - rows = conn.execute( - """ - SELECT target_id FROM intent_edges - WHERE graph_id = ? AND version = ? AND source_id = ? - AND relation = 'blocked_by' AND status = 'active' - ORDER BY target_id - """, - (graph_id, version, goal), - ).fetchall() - constraints.update(row["target_id"] for row in rows) - return tuple(sorted(constraints)) - - -def _collect_capabilities( - conn: sqlite3.Connection, - graph_id: str, - version: str, - goals: tuple[str, ...], - allowed: frozenset[str], -) -> tuple[tuple[str, ...], tuple[str, ...], tuple[str, ...]]: - hints: set[str] = set() - rejected: set[str] = set() - warnings: set[str] = set() - capability_nodes: set[str] = set() - for goal in goals: - rows = conn.execute( - """ - SELECT target_id FROM intent_edges - WHERE graph_id = ? AND version = ? AND source_id = ? - AND relation = 'routes_to' AND status = 'active' - ORDER BY target_id - """, - (graph_id, version, goal), - ).fetchall() - capability_nodes.update(row["target_id"] for row in rows) - for node in sorted(capability_nodes): - rows = conn.execute( - """ - SELECT ref, availability FROM intent_bindings - WHERE graph_id = ? AND version = ? AND node_id = ? - ORDER BY ref - """, - (graph_id, version, node), - ).fetchall() - for row in rows: - ref = row["ref"] - if row["availability"] != "available": - rejected.add(ref) - warnings.add(f"binding_unavailable:{ref}") - elif ref not in allowed: - rejected.add(ref) - warnings.add(f"capability_not_allowed:{ref}") - else: - hints.add(ref) - return tuple(sorted(hints)), tuple(sorted(rejected)), tuple(sorted(warnings)) - - -def resolve_intent( - conn: sqlite3.Connection, - query_text: str, - *, - intent_id: str | None = None, - scope: str | None = None, - allowed_capability_refs: frozenset[str] = frozenset(), - limits: TraversalLimits = TraversalLimits(), -) -> IntentResolution: - """Resolve one reviewed root goal and return bounded prerequisite context.""" - graph_id, version, source_hash = _metadata(conn) - method: Literal["explicit", "alias", "rule", "none"] = "none" - goal_id: str | None = None - - if intent_id is not None: - method = "explicit" - row = conn.execute( - """ - SELECT node_id FROM intent_nodes - WHERE graph_id = ? AND version = ? AND node_id = ? - AND kind = 'goal' AND status = 'active' - """, - (graph_id, version, intent_id), - ).fetchone() - goal_id = row["node_id"] if row else None - else: - normalized = normalize_text(query_text) - rows = conn.execute( - """ - SELECT node_id FROM intent_aliases - WHERE graph_id = ? AND version = ? AND normalized_alias = ? - ORDER BY node_id - """, - (graph_id, version, normalized), - ).fetchall() - alias_goals = sorted({row["node_id"] for row in rows}) - if len(alias_goals) > 1: - return _base_resolution( - graph_id, - version, - source_hash, - outcome="refusal", - method="none", - refusal_reason="intent_alias_ambiguous", - ) - if alias_goals: - method = "alias" - goal_id = alias_goals[0] - else: - goal_id, ambiguous = _resolve_rule( - conn, graph_id, version, query_text, scope - ) - if ambiguous: - return _base_resolution( - graph_id, - version, - source_hash, - outcome="refusal", - method="rule", - refusal_reason="intent_rule_ambiguous", - ) - if goal_id: - method = "rule" - - if goal_id is None: - return _base_resolution( - graph_id, - version, - source_hash, - outcome="fallback", - method="none" if intent_id is None else "explicit", - refusal_reason="intent_no_match", - ) - - for relation in sorted(ACYCLIC_RELATIONS): - cycle = _runtime_cycle(conn, graph_id, version, relation) - if cycle: - return _base_resolution( - graph_id, - version, - source_hash, - outcome="refusal", - method=method, - goal_id=goal_id, - refusal_reason="intent_cycle_detected", - path=cycle, - ) - - path, prerequisites, edge_path, truncation = _traverse_relation( - conn, graph_id, version, goal_id, "requires", limits - ) - constraints = _collect_constraints(conn, graph_id, version, path) - hints, rejected, capability_warnings = _collect_capabilities( - conn, graph_id, version, path, allowed_capability_refs - ) - warnings = set(capability_warnings) - warnings.update(f"blocked_by:{item}" for item in constraints) - if truncation: - warnings.add("intent_traversal_truncated") - return IntentResolution( - graph_id=graph_id, - graph_version=version, - source_hash=source_hash, - outcome="resolved", - resolution_method=method, - matched_goal_ids=(goal_id,), - prerequisite_goal_ids=prerequisites, - intent_path=path, - edge_path=edge_path, - capability_hints=hints, - rejected_capability_hints=rejected, - constraint_ids=constraints, - warnings=tuple(sorted(warnings)), - truncation_reason=truncation, - ) diff --git a/mindgraph/src/mindgraph/mcp_proxy.py b/mindgraph/src/mindgraph/mcp_proxy.py deleted file mode 100644 index f4193fb..0000000 --- a/mindgraph/src/mindgraph/mcp_proxy.py +++ /dev/null @@ -1,79 +0,0 @@ -"""Official-SDK stdio server proxying to a Streamable HTTP MCP server.""" - -from contextlib import asynccontextmanager -import anyio -import json -import urllib.parse -import urllib.request -from mcp import ClientSession, types -from mcp.client.streamable_http import streamablehttp_client -from mcp.server import Server -from mcp.server.stdio import stdio_server - - -def create_proxy_server(remote: ClientSession) -> Server: - server = Server("mindgraph-mcp-proxy") - - @server.list_tools() - async def list_tools() -> list[types.Tool]: - return (await remote.list_tools()).tools - - @server.call_tool() - async def call_tool(name: str, arguments: dict | None): - return await remote.call_tool(name, arguments) - return server - - -@asynccontextmanager -async def remote_session(url: str): - async with streamablehttp_client(url) as (read, write, _session_id): - async with ClientSession(read, write) as session: - await session.initialize() - yield session - - -def _lease_url(url: str) -> str: - parsed = urllib.parse.urlsplit(url) - return urllib.parse.urlunsplit((parsed.scheme, parsed.netloc, "/lifecycle/lease", "", "")) - - -def _lease_request(url: str, token: str = "", method: str = "POST") -> dict: - request = urllib.request.Request( - _lease_url(url), method=method, - headers={"X-MindGraph-Lease": token} if token else {}, - ) - with urllib.request.urlopen(request, timeout=2) as response: - return json.loads(response.read()) - - -async def _renew_lease(url: str, token: str, interval: float) -> None: - while True: - await anyio.sleep(interval) - await anyio.to_thread.run_sync(_lease_request, url, token) - - -async def run_proxy(url: str, *, lease: bool = False) -> None: - token = "" - if lease: - response = await anyio.to_thread.run_sync(_lease_request, url) - token = response["lease"] - interval = max(1.0, float(response["ttl_seconds"]) / 3) - try: - async with anyio.create_task_group() as tasks: - if token: - tasks.start_soon(_renew_lease, url, token, interval) - async with remote_session(url) as remote: - proxy = create_proxy_server(remote) - async with stdio_server() as (read, write): - await proxy.run(read, write, proxy.create_initialization_options()) - tasks.cancel_scope.cancel() - finally: - if token: - try: - await anyio.to_thread.run_sync(_lease_request, url, token, "DELETE") - except Exception: - pass - - -def run_proxy_sync(url: str, *, lease: bool = False) -> None: - anyio.run(lambda: run_proxy(url, lease=lease)) diff --git a/mindgraph/src/mindgraph/mcp_server.py b/mindgraph/src/mindgraph/mcp_server.py deleted file mode 100644 index 10bf2b3..0000000 --- a/mindgraph/src/mindgraph/mcp_server.py +++ /dev/null @@ -1,430 +0,0 @@ -"""Stdio MCP transport for MindGraph. - -Phase 5 is a transport wrap only. Query behavior stays in `mindgraph.query`. -""" - -import json -import logging -import os -import sqlite3 -import time -from contextlib import contextmanager -from pathlib import Path -from typing import Literal - -from mcp.server.fastmcp import FastMCP -from mcp.types import CallToolResult, TextContent -from starlette.responses import JSONResponse -from starlette.requests import Request - -from mindgraph import db -from mindgraph import query as query_mod -from mindgraph.embedders import EmbedTemplate, EmbedderSpec, format_query_text -from mindgraph import intent as intent_mod -from mindgraph import routing as routing_mod -from mindgraph.exceptions import MindgraphError -from mindgraph.models import QueryResult -from mindgraph.idle_lifecycle import IdleLifecycle - -REQUIRED_TABLES = { - "documents", - "documents_fts", - "chunks", - "vec_chunks", - "edges", -} - -logger = logging.getLogger("mindgraph") - - -class MCPServerStartupError(MindgraphError): - """Raised when the MCP server cannot start cleanly.""" - - -def _find_locked_operational_error(exc: BaseException) -> sqlite3.OperationalError | None: - """Find a SQLite lock error even when a query layer wrapped it.""" - current: BaseException | None = exc - seen: set[int] = set() - while current is not None and id(current) not in seen: - seen.add(id(current)) - if isinstance(current, sqlite3.OperationalError) and "locked" in str( - current - ).lower(): - return current - current = current.__cause__ or current.__context__ - return None - - -def _run_with_lock_retry(operation, *, attempts: int = 5, base_delay: float = 0.5): - """Retry transient SQLite lock errors, including wrapped query errors.""" - for attempt in range(attempts): - try: - return operation() - except Exception as exc: - if _find_locked_operational_error(exc) is None or attempt + 1 >= attempts: - raise - time.sleep(base_delay * (attempt + 1)) - raise AssertionError("unreachable") - - -def open_database(db_path: str) -> sqlite3.Connection: - """Open and validate the single database used by the MCP server.""" - if db_path != ":memory:" and not Path(db_path).exists(): - raise MCPServerStartupError(f"Database does not exist: {db_path}") - - try: - conn = db.get_db(db_path, read_only=True) - _validate_schema(conn, db_path) - return conn - except MindgraphError: - raise - except sqlite3.Error as e: - raise MCPServerStartupError(f"Failed to open database at {db_path}: {e}") from e - - -def open_database_readonly(db_path: str) -> sqlite3.Connection: - """Open and validate one index without allowing persistent writes.""" - expanded = Path(db_path).expanduser().resolve() - if not expanded.exists(): - raise MCPServerStartupError(f"Database does not exist: {expanded}") - try: - conn = db.get_db(str(expanded), read_only=True) - _validate_schema(conn, str(expanded)) - return conn - except MindgraphError: - raise - except sqlite3.Error as exc: - raise MCPServerStartupError( - f"Failed to open database at {expanded}: {exc}" - ) from exc - - -def _validate_schema(conn: sqlite3.Connection, db_path: str) -> None: - rows = conn.execute( - "SELECT name FROM sqlite_master WHERE type IN ('table', 'virtual table')" - ).fetchall() - existing = {row["name"] for row in rows} - missing = sorted(REQUIRED_TABLES - existing) - if missing: - conn.close() - raise MCPServerStartupError( - f"Database at {db_path} is not a MindGraph database; " - f"missing tables: {', '.join(missing)}" - ) - - -def create_server( - conn: sqlite3.Connection, - embedder: query_mod.Embedder, - *, - embedder_spec: EmbedderSpec | None = None, - embed_template: EmbedTemplate = "none", - intent_db_path: str | os.PathLike[str] = "~/.mindgraph/mainframe-intent.sqlite", - log_level: Literal["DEBUG", "INFO"] = "INFO", -) -> FastMCP: - """Create a FastMCP server bound to one DB connection and one embedder.""" - server = FastMCP( - "mindgraph", - instructions=( - "MindGraph retrieves candidate chunks from a local Markdown vault. " - "It does not verify claims. Each result carries semantic_distance " - "(lower is a closer match) and a weak_fit flag: when weak_fit is " - "true, this index has no strong answer for the query and the chunk " - "should be treated as a low-confidence nomination, not an answer. " - "Rows may also carry query_scope_warning when the query appears to " - "ask for inbox, live, or project-status state that belongs in a " - "different lifecycle scope." - ), - log_level=log_level, - ) - - @server.tool( - name="query", - description=( - "Run the MindGraph lexical plus semantic query path, optionally " - "appending outbound graph expansion or association results. " - "Default response is a JSON array of QueryResult nominations. " - "With envelope=true, returns " - "{schema_version, intent_resolution, routing, results} where " - "routing is single-database metadata for this MCP-bound index " - "(not multi-index federation)." - ), - ) - def query_tool( - question: str, - lexical_top_k: int = query_mod.DEFAULT_LEXICAL_TOP_K, - semantic_top_k: int = query_mod.DEFAULT_SEMANTIC_TOP_K, - final_top_k: int = query_mod.DEFAULT_FINAL_TOP_K, - expand: bool = False, - expand_depth: int = query_mod.DEFAULT_EXPAND_DEPTH, - expand_top_k: int = query_mod.DEFAULT_EXPAND_TOP_K, - associate: bool = False, - associate_top_k: int = query_mod.DEFAULT_ASSOCIATE_TOP_K, - associate_seed_k: int = query_mod.DEFAULT_ASSOCIATE_SEED_K, - envelope: bool = False, - ) -> CallToolResult: - formatted_question = question - if embedder_spec is not None: - formatted_question = format_query_text( - embedder_spec, question, template=embed_template - ) - try: - results = _run_with_lock_retry( - lambda: query_mod.run_query( - conn, - formatted_question, - embedder, - lexical_top_k=lexical_top_k, - semantic_top_k=semantic_top_k, - final_top_k=final_top_k, - expand=expand, - expand_depth=expand_depth, - expand_top_k=expand_top_k, - associate=associate, - associate_top_k=associate_top_k, - associate_seed_k=associate_seed_k, - embedder_spec=embedder_spec, - embed_template=embed_template, - ) - ) - if envelope: - payload = _query_envelope( - formatted_question, - results, - intent_db_path=intent_db_path, - ) - return _json_result(payload) - return _json_result([result.model_dump() for result in results]) - except sqlite3.OperationalError as e: - return _tool_error(f"Database error: {e}") - except MindgraphError as e: - return _tool_error(str(e)) - except Exception: - logger.exception("unexpected MCP query tool failure") - raise - - @server.tool( - name="graph_neighbors", - description=( - "List outbound MindGraph edges for a document ID, preserving " - "dangling targets as null target_path values." - ), - ) - def graph_neighbors_tool(doc_id: str) -> CallToolResult: - def lookup_neighbors(): - _ensure_document_exists(conn, doc_id) - return query_mod.list_neighbors(conn, doc_id) - - try: - results = _run_with_lock_retry(lookup_neighbors) - return _json_result([result.model_dump() for result in results]) - except sqlite3.OperationalError as e: - return _tool_error(f"Database error: {e}") - except MindgraphError as e: - return _tool_error(str(e)) - except Exception: - logger.exception("unexpected MCP graph_neighbors tool failure") - raise - - return server - - -def run_stdio(server: FastMCP) -> None: - """Run the server on stdio. Stdout is reserved for MCP protocol frames.""" - server.run("stdio") - - -def create_shared_server( - scopes: dict[str, tuple[sqlite3.Connection, str]], - embedder: query_mod.Embedder, - *, - host: str = "127.0.0.1", - port: int = 8000, - path: str = "/mcp", - embedder_spec: EmbedderSpec | None = None, - embed_template: EmbedTemplate = "none", - lifecycle: IdleLifecycle | None = None, -) -> FastMCP: - """Create a loopback shared server with one explicit scope per call.""" - if host not in {"127.0.0.1", "localhost", "::1"}: - raise MCPServerStartupError("shared MCP daemon host must be loopback") - if not path.startswith("/"): - raise MCPServerStartupError("MCP path must start with '/'") - server = FastMCP( - "mindgraph-shared", - instructions="Select one lifecycle scope. Results are nominations, not verified claims.", - host=host, - port=port, - streamable_http_path=path, - ) - - def selected(scope: str) -> tuple[sqlite3.Connection, str]: - try: - return scopes[scope] - except KeyError as exc: - raise query_mod.QueryError( - f"unknown scope: {scope}; choose one of: {', '.join(sorted(scopes))}" - ) from exc - - @server.tool(name="query") - def scoped_query(question: str, scope: str, final_top_k: int = query_mod.DEFAULT_FINAL_TOP_K) -> CallToolResult: - try: - activity = lifecycle.request() if lifecycle else _null_context() - with activity: - conn, trust_profile = selected(scope) - formatted = ( - format_query_text(embedder_spec, question, template=embed_template) - if embedder_spec is not None else question - ) - rows = _run_with_lock_retry( - lambda: query_mod.run_query(conn, formatted, embedder, final_top_k=final_top_k) - ) - return _json_result({"scope": scope, "trust_profile": trust_profile, - "results": [row.model_dump() for row in rows]}) - except sqlite3.OperationalError as exc: - return _tool_error(f"Database error: {exc}") - except MindgraphError as exc: - return _tool_error(str(exc)) - - @server.tool(name="graph_neighbors") - def scoped_neighbors(doc_id: str, scope: str) -> CallToolResult: - try: - activity = lifecycle.request() if lifecycle else _null_context() - with activity: - conn, trust_profile = selected(scope) - def lookup(): - _ensure_document_exists(conn, doc_id) - return query_mod.list_neighbors(conn, doc_id) - rows = _run_with_lock_retry(lookup) - return _json_result({"scope": scope, "trust_profile": trust_profile, - "results": [row.model_dump() for row in rows]}) - except sqlite3.OperationalError as exc: - return _tool_error(f"Database error: {exc}") - except MindgraphError as exc: - return _tool_error(str(exc)) - - @server.custom_route("/health", methods=["GET"]) - async def health(_request): - return JSONResponse({"status": "ok", "pid": os.getpid(), "scopes": [ - {"scope": name, "trust_profile": trust} - for name, (_conn, trust) in sorted(scopes.items()) - ], "idle_lifecycle": lifecycle.snapshot() if lifecycle else {"enabled": False}}) - if lifecycle: - @server.custom_route("/lifecycle/lease", methods=["POST", "DELETE"]) - async def lease(request: Request): - token = request.headers.get("x-mindgraph-lease", "") - if request.method == "DELETE": - lifecycle.release_lease(token) - return JSONResponse({"status": "released"}) - if token: - lifecycle.renew_lease(token) - else: - token = lifecycle.acquire_lease() - return JSONResponse({"status": "ok", "lease": token, - "ttl_seconds": lifecycle.lease_ttl}) - return server - - -def run_streamable_http(server: FastMCP) -> None: - server.run("streamable-http") - - -@contextmanager -def _null_context(): - yield - - -def _ensure_document_exists(conn: sqlite3.Connection, doc_id: str) -> None: - try: - row = conn.execute( - "SELECT 1 FROM documents WHERE id = ? LIMIT 1", (doc_id,) - ).fetchone() - except sqlite3.Error as e: - raise query_mod.QueryError(f"document lookup failed: {e}") from e - if row is None: - raise query_mod.QueryError(f"unknown doc_id: {doc_id}") - - -def _intent_resolution_payload( - resolution: intent_mod.IntentResolution, -) -> dict: - payload = resolution.model_dump() - payload["method"] = payload.pop("resolution_method") - payload["matched_goals"] = payload.pop("matched_goal_ids") - payload["prerequisite_goals"] = payload.pop("prerequisite_goal_ids") - payload["path"] = payload.pop("intent_path") - return payload - - -def _resolve_intent_payload( - question: str, intent_db_path: str | os.PathLike[str] -) -> tuple[dict | None, str, tuple[str, ...]]: - expanded = Path(os.path.expanduser(os.fspath(intent_db_path))) - if not expanded.exists(): - return None, "intent_store_missing", () - - intent_conn = None - try: - intent_conn = intent_mod.open_intent_store(expanded) - resolution = intent_mod.resolve_intent( - intent_conn, - question, - limits=intent_mod.TraversalLimits(max_depth=2, max_nodes=32), - ) - return ( - _intent_resolution_payload(resolution), - "intent_resolved", - tuple(resolution.warnings), - ) - except Exception as exc: # noqa: BLE001 - MCP envelope reports resolution failures. - logger.warning("MCP intent resolution failed: %s", exc) - return ( - { - "outcome": "fallback", - "method": "none", - "refusal_reason": "intent_resolution_failed", - "error_type": type(exc).__name__, - "message": str(exc), - }, - "intent_resolution_failed", - ("intent_resolution_failed",), - ) - finally: - if intent_conn is not None: - intent_conn.close() - - -def _query_envelope( - question: str, - results: list[QueryResult], - *, - intent_db_path: str | os.PathLike[str], -) -> dict: - intent_payload, reason, warnings = _resolve_intent_payload(question, intent_db_path) - return { - "schema_version": routing_mod.SCHEMA_VERSION, - "intent_resolution": intent_payload, - "routing": { - "mode": "single_database", - "selected_retrievers": ["mcp-bound-db"], - "reason_codes": [reason], - "warnings": list(warnings), - }, - "results": [result.model_dump() for result in results], - } - - -def _json_result(payload) -> CallToolResult: - return CallToolResult( - content=[ - TextContent(type="text", text=json.dumps(payload, indent=2)), - ], - isError=False, - ) - - -def _tool_error(message: str) -> CallToolResult: - return CallToolResult( - content=[TextContent(type="text", text=message)], - isError=True, - ) diff --git a/mindgraph/src/mindgraph/models.py b/mindgraph/src/mindgraph/models.py deleted file mode 100644 index e50318a..0000000 --- a/mindgraph/src/mindgraph/models.py +++ /dev/null @@ -1,95 +0,0 @@ -from typing import Literal - -from pydantic import BaseModel, Field - -Signal = Literal["lexical", "semantic", "fused", "expanded", "associated"] -QueryScopeIntent = Literal["inbox_state", "live_state", "project_status"] - - -class GraphEdge(BaseModel): - source_id: str - target_id: str - relationship_type: str | None = None - - -class ParsedDocument(BaseModel): - id: str - title: str - path: str - content_hash: str - index_id: str | None = None - trust_profile: str | None = None - namespace: str | None = None - source_root: str | None = None - source_path: str | None = None - display_path: str | None = None - metadata: dict = Field(default_factory=dict) - truth_text: str - timeline_text: str | None = None - - -class QueryScopeWarning(BaseModel): - """Query-level warning copied onto result rows. - - The warning means the query appears to ask for lifecycle state that may not - belong in the current DB. It is advisory metadata, not a ranking signal. - """ - - intent: QueryScopeIntent - recommended_trust_profile: str - message: str - - -class QueryResult(BaseModel): - """One ranked retrieval result. See DECISIONS.md § Phase 2 query path. - - `doc_type`/`domain`/`status` are surfaced from the document's frontmatter - (`type`/`domain`/`status`) so a consumer can tell a raw capture from a - curated note without a second lookup. They are null when the source has no - such frontmatter key. They do not affect ranking — see the Phase 7 ADR. - """ - - doc_id: str - chunk_index: int - path: str - title: str - doc_type: str | None = None - domain: str | None = None - status: str | None = None - index_id: str | None = None - trust_profile: str | None = None - namespace: str | None = None - source_root: str | None = None - source_path: str | None = None - display_path: str | None = None - content_hash: str | None = None - eligibility_run_id: str | None = None - signal: Signal - rrf_score: float - lexical_rank: int | None - semantic_rank: int | None - semantic_distance: float | None = None - weak_fit: bool = False - query_scope_warning: QueryScopeWarning | None = None - #: Set when the source document is quarantined or otherwise not citable. - #: Travels on EVERY chunk, which is the whole point: on 2026-08-09, 103 - #: captures with fabricated citations were found in `10_knowledge/`. Each - #: carried a `needs-audit` tag and later a body banner — but a banner only - #: appears in chunk 0, and a frontmatter tag never appears in chunk text at - #: all. A query landing on chunk 3 returned authoritative-looking prose with - #: no indication the source did not exist. A trust flag that is invisible at - #: query time is not a control. - provenance_warning: str | None = None - chunk_text: str - expansion_depth: int = 0 - association_depth: int = 0 - - -class NeighborResult(BaseModel): - """One outbound edge from a source document. Dangling edges have null paths.""" - - source_id: str - target_id: str - relationship_type: str | None - source_path: str | None - target_path: str | None diff --git a/mindgraph/src/mindgraph/parser.py b/mindgraph/src/mindgraph/parser.py deleted file mode 100644 index 7f6945e..0000000 --- a/mindgraph/src/mindgraph/parser.py +++ /dev/null @@ -1,377 +0,0 @@ -import hashlib -import re -from collections.abc import Callable, Iterable -from dataclasses import dataclass, field -from pathlib import Path - -import yaml - -from mindgraph.exceptions import ParseError -from mindgraph.models import GraphEdge, ParsedDocument - -LINK_PATTERN = re.compile(r"\[\[([^\[\]]+?)\]\](?:\s*\(([^)]+)\))?") - -# `---` on its own line followed (possibly across blank lines) by a `## Timeline` heading. -TIMELINE_SPLIT_PATTERN = re.compile( - r"^[ \t]*---[ \t]*\n(?:[ \t]*\n)*[ \t]*##[ \t]+Timeline[ \t]*$", - re.MULTILINE | re.IGNORECASE, -) - -FRONTMATTER_PATTERN = re.compile(r"\A---\s*\n(.*?)\n---\s*\n(.*)\Z", re.DOTALL) -CANONICAL_FILENAME_PREFIX = re.compile(r"^\d{4}-\d{2}-\d{2}__") -WIKILINK_WRAPPER_PATTERN = re.compile(r"^\[\[(.+?)\]\]$") - - -def compute_doc_id(relative_path: str) -> str: - """Stable short hash of a path string. Same input → same ID.""" - return hashlib.sha256(relative_path.encode("utf-8")).hexdigest()[:16] - - -def compute_scoped_doc_id(index_id: str, namespace: str, source_path: str) -> str: - """Stable short hash for a document inside a named index namespace.""" - scoped = f"{index_id}\0{namespace}\0{source_path}" - return compute_doc_id(scoped) - - -def compute_content_hash(body_bytes: bytes) -> str: - return hashlib.sha256(body_bytes).hexdigest() - - -def _normalize_link_target(target: str) -> str: - target = target.strip() - if not target.endswith(".md"): - target = target + ".md" - return target - - -def _normalize_lookup_key(value: str) -> str: - return value.strip().casefold() - - -def canonical_trailing_slug(stem: str) -> str | None: - """Return the trailing slug from a canonical MainFrame filename stem. - - Canonical shape: ``YYYY-MM-DD__domain__type__slug`` (see ADR-033). - """ - if not CANONICAL_FILENAME_PREFIX.match(stem): - return None - parts = stem.split("__") - if len(parts) < 4: - return None - slug = parts[-1].strip() - return slug or None - - -@dataclass -class LinkResolver: - """Resolve wikilink labels against documents in one ingest scope.""" - - paths: set[str] = field(default_factory=set) - ids_by_path: dict[str, str] = field(default_factory=dict) - stems: dict[str, set[str]] = field(default_factory=dict) - titles: dict[str, set[str]] = field(default_factory=dict) - slug_suffixes: dict[str, set[str]] = field(default_factory=dict) - - @classmethod - def from_documents(cls, documents: Iterable[ParsedDocument]) -> "LinkResolver": - resolver = cls() - for doc in documents: - resolver.add_document(doc) - return resolver - - def add_document(self, doc: ParsedDocument) -> None: - self.paths.add(doc.path) - self.ids_by_path[doc.path] = doc.id - stem = Path(doc.path).stem - self.stems.setdefault(_normalize_lookup_key(stem), set()).add(doc.path) - self.titles.setdefault(_normalize_lookup_key(doc.title), set()).add(doc.path) - trailing_slug = canonical_trailing_slug(stem) - if trailing_slug is not None: - self.slug_suffixes.setdefault( - _normalize_lookup_key(trailing_slug), set() - ).add(doc.path) - - def doc_id_for_path(self, path: str) -> str | None: - return self.ids_by_path.get(path) - - def resolve(self, target: str, source_path: str | None = None) -> str | None: - normalized = _normalize_link_target(target) - if normalized in self.paths: - return normalized - - if source_path is not None: - sibling = str(Path(source_path).parent / normalized) - if sibling in self.paths: - return sibling - - if "/" not in normalized: - stem_key = _normalize_lookup_key(Path(normalized).stem) - stem_matches = self.stems.get(stem_key, set()) - if len(stem_matches) == 1: - return next(iter(stem_matches)) - - raw_key = _normalize_lookup_key(target) - title_matches = self.titles.get(raw_key, set()) - if len(title_matches) == 1: - return next(iter(title_matches)) - - if "/" not in normalized: - suffix_key = _normalize_lookup_key(Path(normalized).stem) - suffix_matches = self.slug_suffixes.get(suffix_key, set()) - if len(suffix_matches) == 1: - return next(iter(suffix_matches)) - - return None - - -def parse_frontmatter(text: str) -> tuple[dict, str]: - """Extract YAML frontmatter if present. Returns (metadata, body).""" - match = FRONTMATTER_PATTERN.match(text) - if not match: - return {}, text - raw_yaml, body = match.groups() - try: - metadata = yaml.safe_load(raw_yaml) or {} - except yaml.YAMLError as e: - raise ParseError(f"Malformed YAML frontmatter: {e}") from e - if not isinstance(metadata, dict): - raise ParseError( - f"Frontmatter must be a YAML mapping, got {type(metadata).__name__}" - ) - return metadata, body - - -def split_page_model(body: str) -> tuple[str, str | None]: - """Split body into (truth, timeline) on `---` followed by `## Timeline`. - - Plain `---` horizontal rules elsewhere in the body do not trigger the split. - Returns (body, None) if no Timeline section is present. - """ - match = TIMELINE_SPLIT_PATTERN.search(body) - if not match: - return body.strip(), None - truth = body[: match.start()].strip() - timeline = body[match.end():].strip() - return truth, (timeline or None) - - -def normalize_link_label(raw: str) -> str: - """Normalize a link label from frontmatter or wikilink syntax.""" - target = raw.strip() - if not target: - return "" - wrapper = WIKILINK_WRAPPER_PATTERN.match(target) - if wrapper: - target = wrapper.group(1).strip() - if "|" in target: - target = target.split("|", 1)[0].strip() - return target - - -def extract_metadata_link_targets(metadata: dict) -> list[str]: - """Return deduplicated link targets from frontmatter ``links:``.""" - raw_links = metadata.get("links") - if raw_links is None: - return [] - if isinstance(raw_links, str): - candidates = [raw_links] - elif isinstance(raw_links, list): - candidates = [str(item) for item in raw_links if item is not None] - else: - return [] - - out: list[str] = [] - seen: set[str] = set() - for raw in candidates: - target = normalize_link_label(raw) - if not target: - continue - key = _normalize_lookup_key(target) - if key in seen: - continue - seen.add(key) - out.append(target) - return out - - -def _edge_for_target( - target_raw: str, - source_id: str, - *, - link_resolver: Callable[[str, str | None], str | None] | LinkResolver | None, - source_path: str | None, - relationship_type: str | None = None, -) -> GraphEdge: - resolved_path = None - resolved_id = None - if isinstance(link_resolver, LinkResolver): - resolved_path = link_resolver.resolve(target_raw, source_path) - if resolved_path is not None: - resolved_id = link_resolver.doc_id_for_path(resolved_path) - elif link_resolver is not None: - resolved_path = link_resolver(target_raw, source_path) - target_id = resolved_id or compute_doc_id( - resolved_path or _normalize_link_target(target_raw) - ) - return GraphEdge( - source_id=source_id, - target_id=target_id, - relationship_type=relationship_type, - ) - - -def extract_graph_edges( - text: str, - source_id: str, - *, - link_resolver: Callable[[str, str | None], str | None] | LinkResolver | None = None, - source_path: str | None = None, -) -> list[GraphEdge]: - """Find `[[link]]` and `[[link]] (relationship)` patterns and return edges.""" - edges: list[GraphEdge] = [] - for match in LINK_PATTERN.finditer(text): - target_raw, relationship = match.groups() - edges.append( - _edge_for_target( - target_raw, - source_id, - link_resolver=link_resolver, - source_path=source_path, - relationship_type=relationship.strip() if relationship else None, - ) - ) - return edges - - -def extract_document_graph_edges( - parsed: ParsedDocument, - *, - link_resolver: Callable[[str, str | None], str | None] | LinkResolver | None = None, -) -> list[GraphEdge]: - """Collect graph edges from frontmatter ``links:`` and body wikilinks.""" - by_target: dict[str, GraphEdge] = {} - - for target in extract_metadata_link_targets(parsed.metadata): - edge = _edge_for_target( - target, - parsed.id, - link_resolver=link_resolver, - source_path=parsed.path, - ) - by_target.setdefault(edge.target_id, edge) - - for edge in extract_graph_edges( - parsed.truth_text, - parsed.id, - link_resolver=link_resolver, - source_path=parsed.path, - ): - existing = by_target.get(edge.target_id) - if existing is None: - by_target[edge.target_id] = edge - elif existing.relationship_type is None and edge.relationship_type is not None: - by_target[edge.target_id] = edge - - return list(by_target.values()) - - -# Sentence boundary: whitespace that follows `.`, `!`, or `?`. -_SENTENCE_BOUNDARY = re.compile(r"(?<=[.!?])\s+") - - -def _hard_split(text: str, max_chars: int) -> list[str]: - """Last-resort fixed-width cut for a unit with no usable boundary.""" - return [text[i : i + max_chars] for i in range(0, len(text), max_chars)] - - -def _split_oversized(para: str, max_chars: int) -> list[str]: - """Break a paragraph longer than max_chars into pieces each <= max_chars. - - Prefers line boundaries, then sentence boundaries, and finally a hard - character cut for a single unit that still exceeds max_chars (e.g. a giant - blob with no whitespace, like a raw arXiv capture). Surviving units are - greedily re-packed up to max_chars. - """ - if len(para) <= max_chars: - return [para] - - units: list[str] = [] - for line in para.split("\n"): - if len(line) <= max_chars: - units.append(line) - continue - for sentence in _SENTENCE_BOUNDARY.split(line): - if len(sentence) <= max_chars: - units.append(sentence) - else: - units.extend(_hard_split(sentence, max_chars)) - - pieces: list[str] = [] - current = "" - for unit in units: - if not unit: - continue - if not current: - current = unit - elif len(current) + 1 + len(unit) <= max_chars: - current = current + " " + unit - else: - pieces.append(current) - current = unit - if current: - pieces.append(current) - return pieces - - -def chunk_truth(truth_text: str, max_chars: int = 1000) -> list[str]: - """Pack paragraphs into chunks bounded by max_chars. - - Paragraphs are the natural chunk unit, but a single paragraph longer than - max_chars is first hard-split on sentence/line boundaries so no chunk can - exceed the bound. This protects semantic recall (MiniLM truncates at ~256 - tokens, so an unsplit megachunk is mostly invisible to the embedder) and - avoids dumping an oversized chunk into a `--json`/MCP response. - """ - if not truth_text.strip(): - return [] - paragraphs: list[str] = [] - for raw in re.split(r"\n\s*\n", truth_text): - para = raw.strip() - if para: - paragraphs.extend(_split_oversized(para, max_chars)) - chunks: list[str] = [] - current = "" - for para in paragraphs: - if not current: - current = para - elif len(current) + 2 + len(para) <= max_chars: - current = current + "\n\n" + para - else: - chunks.append(current) - current = para - if current: - chunks.append(current) - return chunks - - -def parse_document(relative_path: str, body_bytes: bytes) -> ParsedDocument: - """Parse a Markdown file's bytes into a validated ParsedDocument.""" - try: - text = body_bytes.decode("utf-8") - except UnicodeDecodeError as e: - raise ParseError(f"File is not valid UTF-8: {e}", path=relative_path) from e - - metadata, body = parse_frontmatter(text) - truth, timeline = split_page_model(body) - - title = metadata.get("title") or Path(relative_path).stem - - return ParsedDocument( - id=compute_doc_id(relative_path), - title=str(title), - path=relative_path, - content_hash=compute_content_hash(body_bytes), - metadata=metadata, - truth_text=truth, - timeline_text=timeline, - ) diff --git a/mindgraph/src/mindgraph/query.py b/mindgraph/src/mindgraph/query.py deleted file mode 100644 index 92056e6..0000000 --- a/mindgraph/src/mindgraph/query.py +++ /dev/null @@ -1,1005 +0,0 @@ -"""Query path: lexical plus semantic retrieval fused with Reciprocal Rank Fusion. - -Design is locked in DECISIONS.md § 2026-05-19 — Phase 2 query path. The fusion -constant RRF_K is canonical (Cormack, Clarke, Buettcher 2009) and not configurable -at runtime by design. -""" - -import json -import logging -import os -import re -import sqlite3 -import struct -from collections.abc import Sequence -from dataclasses import dataclass -from functools import cached_property, lru_cache -from pathlib import Path -from typing import Any, Protocol - -from mindgraph.embedders import EmbedTemplate, EmbedderSpec, format_query_text -from mindgraph.exceptions import DatabaseError, MindgraphError -from mindgraph.models import NeighborResult, QueryResult, QueryScopeWarning - -logger = logging.getLogger(__name__) - -RRF_K = 60 -DEFAULT_LEXICAL_TOP_K = 20 -DEFAULT_SEMANTIC_TOP_K = 20 -DEFAULT_FINAL_TOP_K = 10 -MAX_QUERY_TOP_K = 1000 -MAX_EXPAND_DEPTH = 3 - -# Best-semantic-distance cutoff above which a semantic-only result is flagged -# weak_fit ("this index probably has no answer"). Calibrated against the live -# MainFrame index (Phase 7 ADR): strong hits land near 0.71 (top-5 <= 0.82), -# out-of-scope queries near 1.10, nonsense near 1.22, so 1.0 sits in the gap. -WEAK_FIT_DISTANCE_THRESHOLD = 1.0 - -# Scope-warning vocabulary. -# -# Each entry is a regular-expression *fragment*, not a literal, so `captures?` -# and `state\.md` behave as written. Fragments are joined with `|` and wrapped -# in `\b(...)\b`, matched case-insensitively. -# -# The defaults below describe one vault's lifecycle vocabulary. They are a -# starting point, not a claim about anyone else's notes. Override them with a -# JSON file (see `load_scope_vocabulary`) when your own material uses different -# words for "this is current state, not durable knowledge". - -DEFAULT_INBOX_TERMS = ( - "inbox", "captures?", "routing queue", "ready queue", - "waiting for routing", "00_inbox", "01_ingest", -) -DEFAULT_PROJECT_TERMS = ( - "30_projects", "project status", "active project", "project_state", - "next_action", "next action", "project readme", r"state\.md", - "handoff", "next gate", -) -DEFAULT_FRESHNESS_TERMS = ( - "current", "latest", "today", "this week", "this month", "right now", - "now", "recent", "live", "as of", -) -DEFAULT_LIVE_STATE_TERMS = ( - "job hunt", "finance", "calendar", "workflow metrics", "telemetry", - "live state", "status", "blocked?", "blockers?", "remaining", "next", -) - - -def _compile_terms(terms: Sequence[str]) -> re.Pattern[str] | None: - """Compile alternation fragments into `\\b(a|b|c)\\b`, or None if empty.""" - kept = [t for t in terms if t] - if not kept: - return None - return re.compile(r"\b(" + "|".join(kept) + r")\b", re.IGNORECASE) - - -@dataclass(frozen=True) -class ScopeVocabulary: - """Term fragments driving `classify_query_scope`. - - An empty term list disables that warning branch entirely. - """ - - inbox_terms: tuple[str, ...] = DEFAULT_INBOX_TERMS - project_terms: tuple[str, ...] = DEFAULT_PROJECT_TERMS - freshness_terms: tuple[str, ...] = DEFAULT_FRESHNESS_TERMS - live_state_terms: tuple[str, ...] = DEFAULT_LIVE_STATE_TERMS - - @cached_property - def _patterns(self) -> tuple[re.Pattern[str] | None, ...]: - return ( - _compile_terms(self.inbox_terms), - _compile_terms(self.project_terms), - _compile_terms(self.freshness_terms), - _compile_terms(self.live_state_terms), - ) - - @classmethod - def from_mapping(cls, data: dict) -> "ScopeVocabulary": - """Build from a mapping. Absent keys keep their defaults.""" - known = {f: getattr(cls, f, None) for f in cls.__dataclass_fields__} - unknown = sorted(set(data) - set(known)) - if unknown: - raise QueryError( - f"unknown scope vocabulary key(s): {', '.join(unknown)}; " - f"expected any of: {', '.join(sorted(known))}" - ) - kwargs = {} - for field in cls.__dataclass_fields__: - if field in data: - value = data[field] - if isinstance(value, str) or not isinstance(value, Sequence): - raise QueryError( - f"scope vocabulary key {field!r} must be a list of strings" - ) - kwargs[field] = tuple(str(v) for v in value) - return cls(**kwargs) - - -DEFAULT_SCOPE_VOCABULARY = ScopeVocabulary() - -# Environment variable naming a JSON file that overrides the defaults. -SCOPE_VOCABULARY_ENV = "MINDGRAPH_SCOPE_VOCABULARY" - - -def load_scope_vocabulary(path: str | os.PathLike[str]) -> ScopeVocabulary: - """Load a scope vocabulary from a JSON file. Absent keys keep defaults.""" - resolved = Path(path).expanduser() - try: - data = json.loads(resolved.read_text()) - except FileNotFoundError: - raise QueryError(f"scope vocabulary file not found: {resolved}") - except json.JSONDecodeError as exc: - raise QueryError(f"scope vocabulary file is not valid JSON: {resolved} ({exc})") - if not isinstance(data, dict): - raise QueryError(f"scope vocabulary file must contain a JSON object: {resolved}") - return ScopeVocabulary.from_mapping(data) - - -@lru_cache(maxsize=1) -def _vocabulary_from_env(raw: str | None) -> ScopeVocabulary: - return load_scope_vocabulary(raw) if raw else DEFAULT_SCOPE_VOCABULARY - - -def active_scope_vocabulary() -> ScopeVocabulary: - """The vocabulary in force: the env-var override, else the defaults.""" - return _vocabulary_from_env(os.environ.get(SCOPE_VOCABULARY_ENV)) - -# A word-character run. Everything else — apostrophes, question marks, and the -# rest of `.,;/=%[]<>\|~@#$&!` plus the FTS5 operator symbols `"*():^-` — is a -# token boundary. `\w` is unicode-aware for str patterns, and every `\w+` run is -# a valid FTS5 bareword (ASCII alphanumerics/underscore and/or codepoints >=128), -# so an OR-join of these tokens can never produce a MATCH syntax error. -_FTS5_TOKEN = re.compile(r"\w+") - -# FTS5 reserved operator keywords. FTS5 treats lowercase `and`, `or`, `not`, -# `near` as ordinary tokens, so only the uppercase forms are dropped. -_FTS5_OPERATOR_KEYWORDS = frozenset({"AND", "OR", "NOT", "NEAR"}) -_SHA256_HEX = re.compile(r"^[0-9a-f]{64}$") -_SHA256_TAGGED = re.compile(r"^sha256:[0-9a-f]{64}$") - - -class QueryError(MindgraphError): - """Raised when a query execution fails.""" - - -def _validate_limit(name: str, value: int, *, maximum: int) -> int: - if isinstance(value, bool) or not isinstance(value, int): - raise QueryError(f"{name} must be an integer") - if value < 0 or value > maximum: - raise QueryError(f"{name} must be between 0 and {maximum}") - return value - - -class Embedder(Protocol): - """The minimal interface the query path needs from a sentence embedder. - - Matches the surface of `sentence_transformers.SentenceTransformer.encode` - that the ingest path already depends on. A test embedder can fulfill this - by implementing `encode(texts, convert_to_numpy=True)` returning a 2D array. - """ - - def encode(self, texts, convert_to_numpy=True): ... # pragma: no cover - - -def _encode_without_progress(embedder: Embedder, texts): - """Encode text while suppressing sentence-transformers progress output.""" - try: - return embedder.encode( - texts, convert_to_numpy=True, show_progress_bar=False - ) - except TypeError: - return embedder.encode(texts, convert_to_numpy=True) - - -def sanitize_fts5_query(text: str) -> str: - """Reduce free text to a safe OR-joined FTS5 MATCH expression. - - Uses an allowlist: only word-character runs survive as tokens, so arbitrary - punctuation (apostrophes, question marks, etc.) can never reach the MATCH - parser, and the uppercase operator keywords AND/OR/NOT/NEAR are dropped so - they are not interpreted as operators. The result is an implicit OR over - surviving tokens, per the ADR. Returns an empty string when no tokens - survive (e.g. operator- or punctuation-only input). - """ - tokens = [ - token - for token in _FTS5_TOKEN.findall(text) - if token not in _FTS5_OPERATOR_KEYWORDS - ] - if not tokens: - return "" - return " OR ".join(tokens) - - -def classify_query_scope( - query_text: str, vocabulary: ScopeVocabulary | None = None -) -> QueryScopeWarning | None: - """Return an advisory lifecycle-scope warning for current-state queries. - - This is deliberately a query-intent guardrail, not a no-answer classifier. - A query can have strong lexical and semantic matches while still asking the - wrong database for inbox, live, or project-status state. - - The terms driving it are configurable; see `ScopeVocabulary`. Passing None - uses `active_scope_vocabulary()`, which honours the - `MINDGRAPH_SCOPE_VOCABULARY` environment variable. - """ - vocab = vocabulary if vocabulary is not None else active_scope_vocabulary() - inbox_re, project_re, freshness_re, live_state_re = vocab._patterns - if inbox_re and inbox_re.search(query_text): - return QueryScopeWarning( - intent="inbox_state", - recommended_trust_profile="inbox_or_ingest_queue", - message=( - "Query appears to ask for inbox or routing-queue state. " - "Route it to the live inbox/ingest surface before treating " - "ranked chunks as current-state nominations." - ), - ) - if project_re and project_re.search(query_text): - return QueryScopeWarning( - intent="project_status", - recommended_trust_profile="project_status", - message=( - "Query appears to ask for project status. Route it to a " - "project-status database or inspect the project records before " - "treating ranked chunks as current-state nominations." - ), - ) - if ( - freshness_re - and live_state_re - and freshness_re.search(query_text) - and live_state_re.search(query_text) - ): - return QueryScopeWarning( - intent="live_state", - recommended_trust_profile="time_bound_live_state", - message=( - "Query appears to ask for current or live state. A durable " - "knowledge database can return relevant background notes that " - "are not current-status nominations." - ), - ) - return None - - -def _serialize_embedding(vec) -> bytes: - """Pack a float vector into the bytes format sqlite-vec expects.""" - return struct.pack(f"{len(vec)}f", *vec) - - -def fetch_lexical_ranking( - conn: sqlite3.Connection, - query_text: str, - top_k: int = DEFAULT_LEXICAL_TOP_K, -) -> list[tuple[str, int]]: - """Return (doc_id, rank) pairs from FTS5 BM25 ranking. - - Ranks start at 1. Returns an empty list when sanitization strips all tokens. - """ - _validate_limit("lexical_top_k", top_k, maximum=MAX_QUERY_TOP_K) - if top_k == 0: - return [] - match_expr = sanitize_fts5_query(query_text) - if not match_expr: - return [] - try: - cursor = conn.execute( - """ - SELECT id - FROM documents_fts - WHERE documents_fts MATCH ? - ORDER BY bm25(documents_fts) ASC, id ASC - LIMIT ? - """, - (match_expr, top_k), - ) - return [(row["id"], rank + 1) for rank, row in enumerate(cursor)] - except sqlite3.Error as e: - raise QueryError(f"FTS5 query failed: {e}") from e - - -def fetch_semantic_ranking( - conn: sqlite3.Connection, - query_embedding: list[float], - top_k: int = DEFAULT_SEMANTIC_TOP_K, -) -> list[tuple[str, int, int, float]]: - """Return (doc_id, chunk_index, rank, distance) ordered by vec_chunks distance. - - Promotes to document granularity by keeping the best-ranked chunk per - document, per ADR § retrieval pipeline. Ranks start at 1 and reflect the - chunk-level KNN position before dedup. A document's rank and distance are - those of its first (closest) chunk in the chunk-level ranking. The distance - is carried through fusion for the Phase 7 weak-fit signal. - """ - _validate_limit("semantic_top_k", top_k, maximum=MAX_QUERY_TOP_K) - if top_k == 0: - return [] - try: - # sqlite-vec requires the LIMIT or `k = ?` constraint to be visible on - # the vec0 virtual-table query itself. Wrapping the KNN in a subquery - # keeps the constraint local to vec_chunks, then we join chunks for - # the (doc_id, chunk_index) tuple. - cursor = conn.execute( - """ - SELECT c.doc_id, c.chunk_index, knn.distance - FROM ( - SELECT rowid, distance - FROM vec_chunks - WHERE embedding MATCH ? - AND k = ? - ) AS knn - JOIN chunks c ON c.rowid = knn.rowid - ORDER BY knn.distance ASC, c.doc_id ASC, c.chunk_index ASC - """, - (_serialize_embedding(query_embedding), top_k), - ) - rows = cursor.fetchall() - except sqlite3.Error as e: - raise QueryError(f"vec_chunks query failed: {e}") from e - - seen: dict[str, tuple[int, int, float]] = {} - for rank, row in enumerate(rows, start=1): - doc_id = row["doc_id"] - if doc_id not in seen: - seen[doc_id] = (row["chunk_index"], rank, row["distance"]) - return [ - (doc_id, chunk_index, rank, distance) - for doc_id, (chunk_index, rank, distance) in seen.items() - ] - - -def rrf_fuse( - lexical: list[tuple[str, int]], - semantic: list[tuple[str, int, int, float]], - top_k: int = DEFAULT_FINAL_TOP_K, -) -> list[tuple[str, int | None, float, int | None, int | None, float | None]]: - """Fuse lexical and semantic rankings with RRF at the canonical k = 60. - - Returns (doc_id, chunk_index, rrf_score, lexical_rank, semantic_rank, - semantic_distance). The chunk_index and distance come from the semantic - ranking when present; lexical-only results carry None for both and the - chunk is filled in by the caller with chunk 0 (see ADR § Phase 2 § - lexical-only chunk choice). The distance rides along untouched — it does not - enter the RRF math, which stays rank-based per the locked Phase 2 ADR. Sort - order: rrf_score descending, then doc_id ascending. Deterministic. - """ - _validate_limit("final_top_k", top_k, maximum=MAX_QUERY_TOP_K) - if top_k == 0: - return [] - lex_map = {doc_id: rank for doc_id, rank in lexical} - sem_map = { - doc_id: (chunk_index, rank, distance) - for doc_id, chunk_index, rank, distance in semantic - } - doc_ids = set(lex_map) | set(sem_map) - - scored: list[ - tuple[str, int | None, float, int | None, int | None, float | None] - ] = [] - for doc_id in doc_ids: - lex_rank = lex_map.get(doc_id) - sem_entry = sem_map.get(doc_id) - sem_rank = sem_entry[1] if sem_entry else None - chunk_index = sem_entry[0] if sem_entry else None - sem_distance = sem_entry[2] if sem_entry else None - - score = 0.0 - if lex_rank is not None: - score += 1.0 / (RRF_K + lex_rank) - if sem_rank is not None: - score += 1.0 / (RRF_K + sem_rank) - - scored.append((doc_id, chunk_index, score, lex_rank, sem_rank, sem_distance)) - - scored.sort(key=lambda row: (-row[2], row[0])) - return scored[:top_k] - - -def _attribute_signal(lex_rank: int | None, sem_rank: int | None): - if lex_rank is not None and sem_rank is not None: - return "fused" - if lex_rank is not None: - return "lexical" - return "semantic" - - -def _is_weak_fit( - lex_rank: int | None, sem_rank: int | None, distance: float | None -) -> bool: - """Flag a result the index probably has no real answer for. - - Per the Phase 7 ADR: a result is weak only when it was nominated purely by - semantics (no lexical/fused overlap to corroborate it) and its best chunk - distance is worse than the calibrated threshold. A lexical or fused hit - always carries independent evidence, so it is never weak; expanded graph - results carry no distance and stay False. - """ - return ( - lex_rank is None - and sem_rank is not None - and distance is not None - and distance > WEAK_FIT_DISTANCE_THRESHOLD - ) - - -def _resolve_chunk_text( - conn: sqlite3.Connection, doc_id: str, chunk_index: int -) -> str: - row = conn.execute( - "SELECT text FROM chunks WHERE doc_id = ? AND chunk_index = ?", - (doc_id, chunk_index), - ).fetchone() - return row["text"] if row else "" - - -def _coerce_optional_str(value) -> str | None: - """Normalize a frontmatter scalar to a string, or None when absent/empty.""" - if value is None: - return None - text = str(value).strip() - return text or None - - -#: Frontmatter states that make a document non-citable regardless of how well -#: its text matches a query. -_NON_CITABLE_STATUS = {"quarantined", "retracted", "superseded"} -_NON_CITABLE_TAGS = {"fabricated-citation", "quarantined", "needs-resourcing"} - - -def _provenance_warning(metadata: dict) -> str | None: - """Build a retrieval-time trust warning from a document's frontmatter. - - Returned on every chunk so the warning cannot be outrun by ranking. See - `QueryResult.provenance_warning` for why this exists. - """ - status = str(metadata.get("status") or "").strip().lower() - raw_tags = metadata.get("tags") or [] - if isinstance(raw_tags, str): - raw_tags = [t.strip(" '\"[]") for t in raw_tags.split(",")] - tags = {str(t).strip().lower() for t in raw_tags} - - if status in _NON_CITABLE_STATUS or (tags & _NON_CITABLE_TAGS): - reason = metadata.get("citation_status") or metadata.get("provenance") - detail = f" ({reason})" if reason else "" - return ( - f"NOT CITABLE — this document is {status or 'flagged'}{detail}. " - "Its content may be correct but has no verified source. " - "Do not cite it; re-source any claim before use." - ) - if "needs-audit" in tags or "full-text-pending" in tags: - return ( - "UNVERIFIED — this capture is flagged needs-audit / full-text-pending. " - "Treat as a nomination, not as evidence." - ) - return None - - -def _resolve_document(conn: sqlite3.Connection, doc_id: str) -> dict | None: - """Resolve a document's display + frontmatter fields, or None if missing. - - Returns path, title, provenance, and the `type`/`domain`/`status` - frontmatter values. `domain` is read from its dedicated column; - `doc_type`/`status` come from the stored `metadata_json`. All three are - null when the source omits them. - """ - row = conn.execute( - """ - SELECT - path, title, domain, content_hash, metadata_json, index_id, trust_profile, - namespace, source_root, source_path, display_path - FROM documents - WHERE id = ? - """, - (doc_id,), - ).fetchone() - if row is None: - return None - metadata = {} - if row["metadata_json"]: - try: - parsed = json.loads(row["metadata_json"]) - if isinstance(parsed, dict): - metadata = parsed - except (ValueError, TypeError): - metadata = {} - return { - "path": row["path"], - "title": row["title"], - "content_hash": _coerce_optional_str(row["content_hash"]), - "doc_type": _coerce_optional_str(metadata.get("type")), - "domain": _coerce_optional_str(row["domain"]), - "status": _coerce_optional_str(metadata.get("status")), - "provenance_warning": _provenance_warning(metadata), - "index_id": _coerce_optional_str(row["index_id"]), - "trust_profile": _coerce_optional_str(row["trust_profile"]), - "namespace": _coerce_optional_str(row["namespace"]), - "source_root": _coerce_optional_str(row["source_root"]), - "source_path": _coerce_optional_str(row["source_path"]), - "display_path": _coerce_optional_str(row["display_path"]), - } - - -DEFAULT_EXPAND_DEPTH = 1 -DEFAULT_EXPAND_TOP_K = 20 -DEFAULT_ASSOCIATE_TOP_K = 10 -DEFAULT_ASSOCIATE_SEED_K = 5 - - -def run_query( - conn: sqlite3.Connection, - query_text: str, - embedder: Embedder, - *, - lexical_top_k: int = DEFAULT_LEXICAL_TOP_K, - semantic_top_k: int = DEFAULT_SEMANTIC_TOP_K, - final_top_k: int = DEFAULT_FINAL_TOP_K, - expand: bool = False, - expand_depth: int = DEFAULT_EXPAND_DEPTH, - expand_top_k: int = DEFAULT_EXPAND_TOP_K, - associate: bool = False, - associate_top_k: int = DEFAULT_ASSOCIATE_TOP_K, - associate_seed_k: int = DEFAULT_ASSOCIATE_SEED_K, - embedder_spec: EmbedderSpec | None = None, - embed_template: EmbedTemplate = "none", -) -> list[QueryResult]: - """Run the Phase 2 query pipeline end-to-end and return QueryResult rows. - - Lexical-only results surface chunk 0 by default because there is no semantic - ranking to pick a better chunk from. This is a minor v0.1 simplification - recorded in DECISIONS.md. - - When `expand` is True, walks outbound `[[link]]` edges from the Phase 2 - results to `expand_depth` hops and appends the walked documents with - `signal="expanded"`. See DECISIONS.md § 2026-05-20 — Phase 3 graph expansion. - """ - for name, value in ( - ("lexical_top_k", lexical_top_k), - ("semantic_top_k", semantic_top_k), - ("final_top_k", final_top_k), - ("expand_top_k", expand_top_k), - ("associate_top_k", associate_top_k), - ("associate_seed_k", associate_seed_k), - ): - _validate_limit(name, value, maximum=MAX_QUERY_TOP_K) - _validate_limit("expand_depth", expand_depth, maximum=MAX_EXPAND_DEPTH) - - query_scope_warning = classify_query_scope(query_text) - try: - lexical = fetch_lexical_ranking(conn, query_text, top_k=lexical_top_k) - except QueryError as exc: - # A lexical failure must not sink the whole query: fall back to - # semantic-only retrieval with a logged warning. The sanitizer makes - # this path unreachable for ordinary punctuation, but it guards against - # a malformed FTS index or any future MATCH edge case. - logger.warning( - "lexical ranking failed; degrading to semantic-only: %s", exc - ) - lexical = [] - - if semantic_top_k > 0: - raw = _encode_without_progress(embedder, [query_text]) - query_embedding = [float(x) for x in raw[0]] - semantic = fetch_semantic_ranking( - conn, query_embedding, top_k=semantic_top_k - ) - else: - semantic = [] - - fused = rrf_fuse(lexical, semantic, top_k=final_top_k) - - results: list[QueryResult] = [] - for doc_id, chunk_index, rrf_score, lex_rank, sem_rank, sem_distance in fused: - resolved = _resolve_document(conn, doc_id) - if resolved is None: - # FTS5 row exists but documents row was deleted out from under us. - # Treat as a data-integrity failure rather than silently dropping. - raise QueryError( - f"FTS5 hit for doc_id={doc_id} has no matching documents row" - ) - effective_chunk_index = chunk_index if chunk_index is not None else 0 - chunk_text = _resolve_chunk_text(conn, doc_id, effective_chunk_index) - results.append( - QueryResult( - doc_id=doc_id, - chunk_index=effective_chunk_index, - path=resolved["path"], - title=resolved["title"], - doc_type=resolved["doc_type"], - domain=resolved["domain"], - status=resolved["status"], - provenance_warning=resolved["provenance_warning"], - index_id=resolved["index_id"], - trust_profile=resolved["trust_profile"], - namespace=resolved["namespace"], - source_root=resolved["source_root"], - source_path=resolved["source_path"], - display_path=resolved["display_path"], - content_hash=resolved["content_hash"], - signal=_attribute_signal(lex_rank, sem_rank), - rrf_score=round(rrf_score, 6), - lexical_rank=lex_rank, - semantic_rank=sem_rank, - semantic_distance=( - round(sem_distance, 6) if sem_distance is not None else None - ), - weak_fit=_is_weak_fit(lex_rank, sem_rank, sem_distance), - query_scope_warning=query_scope_warning, - chunk_text=chunk_text, - expansion_depth=0, - ) - ) - - output = list(results) - - if expand: - output.extend( - expand_results( - conn, - results, - depth=expand_depth, - expand_top_k=expand_top_k, - query_scope_warning=query_scope_warning, - ) - ) - - if associate: - output.extend( - associate_results( - conn, - results, - embedder, - top_k=associate_top_k, - seed_k=associate_seed_k, - embedder_spec=embedder_spec, - embed_template=embed_template, - query_scope_warning=query_scope_warning, - ) - ) - - return output - - -def _association_seed_text(conn: sqlite3.Connection, seed: QueryResult) -> str: - chunk = _resolve_chunk_text(conn, seed.doc_id, seed.chunk_index) - title = (seed.title or "").strip() - excerpt = chunk[:500].strip() - if title and excerpt: - return f"{title}\n{excerpt}" - return title or excerpt - - -def associate_results( - conn: sqlite3.Connection, - phase_2_results: list[QueryResult], - embedder: Embedder, - *, - top_k: int = DEFAULT_ASSOCIATE_TOP_K, - seed_k: int = DEFAULT_ASSOCIATE_SEED_K, - embedder_spec: EmbedderSpec | None = None, - embed_template: EmbedTemplate = "none", - query_scope_warning: QueryScopeWarning | None = None, -) -> list[QueryResult]: - """Find embedding-neighbor documents from Phase 2 fused seeds. - - Seeds are the pre-expand Phase 2 rows only (ADR-034). Association is - append-only: rows use signal=associated, rrf_score=0, association_depth=1. - """ - _validate_limit("associate_top_k", top_k, maximum=MAX_QUERY_TOP_K) - _validate_limit("associate_seed_k", seed_k, maximum=MAX_QUERY_TOP_K) - if top_k == 0 or seed_k == 0 or not phase_2_results: - return [] - if query_scope_warning is None: - query_scope_warning = phase_2_results[0].query_scope_warning - - spec = embedder_spec - effective_seed_k = min(seed_k, len(phase_2_results)) if seed_k > 0 else 0 - if effective_seed_k <= 0: - return [] - - seen: set[str] = {r.doc_id for r in phase_2_results} - associated: list[QueryResult] = [] - - for seed in phase_2_results[:effective_seed_k]: - seed_text = _association_seed_text(conn, seed) - if not seed_text: - continue - if spec is not None: - seed_text = format_query_text(spec, seed_text, template=embed_template) - raw = _encode_without_progress(embedder, [seed_text]) - query_embedding = [float(x) for x in raw[0]] - candidate_k = min(MAX_QUERY_TOP_K, top_k + len(seen)) - semantic = fetch_semantic_ranking(conn, query_embedding, top_k=candidate_k) - for doc_id, chunk_index, _rank, distance in semantic: - if doc_id in seen: - continue - resolved = _resolve_document(conn, doc_id) - if resolved is None: - continue - chunk_text = _resolve_chunk_text(conn, doc_id, chunk_index) - associated.append( - QueryResult( - doc_id=doc_id, - chunk_index=chunk_index, - path=resolved["path"], - title=resolved["title"], - doc_type=resolved["doc_type"], - domain=resolved["domain"], - status=resolved["status"], - provenance_warning=resolved["provenance_warning"], - index_id=resolved["index_id"], - trust_profile=resolved["trust_profile"], - namespace=resolved["namespace"], - source_root=resolved["source_root"], - source_path=resolved["source_path"], - display_path=resolved["display_path"], - content_hash=resolved["content_hash"], - signal="associated", - rrf_score=0.0, - lexical_rank=None, - semantic_rank=None, - semantic_distance=round(distance, 6), - weak_fit=_is_weak_fit(None, 1, distance), - query_scope_warning=query_scope_warning, - chunk_text=chunk_text, - expansion_depth=0, - association_depth=1, - ) - ) - seen.add(doc_id) - if len(associated) >= top_k: - break - if len(associated) >= top_k: - break - - associated.sort( - key=lambda r: (r.association_depth, r.semantic_distance or 0.0, r.doc_id) - ) - return associated[:top_k] - - -def expand_results( - conn: sqlite3.Connection, - phase_2_results: list[QueryResult], - *, - depth: int, - expand_top_k: int, - query_scope_warning: QueryScopeWarning | None = None, -) -> list[QueryResult]: - """Walk outbound graph edges from the Phase 2 seeds and return expanded rows. - - Deterministic BFS per DECISIONS.md § 2026-05-20 — Phase 3 graph expansion. - Dangling targets terminate the walk at their depth. Documents already in the - seed set are not re-emitted. Final sort: (expansion_depth, doc_id, chunk_index). - """ - _validate_limit("expand_depth", depth, maximum=MAX_EXPAND_DEPTH) - _validate_limit("expand_top_k", expand_top_k, maximum=MAX_QUERY_TOP_K) - if depth == 0 or expand_top_k == 0 or not phase_2_results: - return [] - if query_scope_warning is None: - query_scope_warning = phase_2_results[0].query_scope_warning - - seen: set[str] = {r.doc_id for r in phase_2_results} - frontier: list[str] = list(seen) - expanded: list[QueryResult] = [] - - for d in range(1, depth + 1): - next_frontier: list[str] = [] - for source_id in frontier: - for edge in list_neighbors(conn, source_id): - target_id = edge.target_id - if edge.target_path is None: - continue - if target_id in seen: - continue - resolved = _resolve_document(conn, target_id) - if resolved is None: - # Defensive: list_neighbors already filtered dangling via - # target_path, so a missing documents row here is a data - # integrity issue worth surfacing. - continue - chunk_text = _resolve_chunk_text(conn, target_id, 0) - expanded.append( - QueryResult( - doc_id=target_id, - chunk_index=0, - path=resolved["path"], - title=resolved["title"], - doc_type=resolved["doc_type"], - domain=resolved["domain"], - status=resolved["status"], - provenance_warning=resolved["provenance_warning"], - index_id=resolved["index_id"], - trust_profile=resolved["trust_profile"], - namespace=resolved["namespace"], - source_root=resolved["source_root"], - source_path=resolved["source_path"], - display_path=resolved["display_path"], - content_hash=resolved["content_hash"], - signal="expanded", - rrf_score=0.0, - lexical_rank=None, - semantic_rank=None, - query_scope_warning=query_scope_warning, - chunk_text=chunk_text, - expansion_depth=d, - ) - ) - seen.add(target_id) - next_frontier.append(target_id) - if not next_frontier: - break - frontier = next_frontier - - expanded.sort(key=lambda r: (r.expansion_depth, r.doc_id, r.chunk_index)) - return expanded[:expand_top_k] - - -def list_neighbors( - conn: sqlite3.Connection, - doc_id: str, -) -> list[NeighborResult]: - """List outbound edges from a document, preserving dangling edges. - - Sort order: target_id ascending, then relationship_type ascending (with - NULL relationship_type sorting first per SQLite default). - """ - try: - cursor = conn.execute( - """ - SELECT - e.source_id, - e.target_id, - e.relationship_type, - src.path AS source_path, - tgt.path AS target_path - FROM edges e - LEFT JOIN documents src ON src.id = e.source_id - LEFT JOIN documents tgt ON tgt.id = e.target_id - WHERE e.source_id = ? - ORDER BY e.target_id ASC, e.relationship_type ASC - """, - (doc_id,), - ) - return [ - NeighborResult( - source_id=row["source_id"], - target_id=row["target_id"], - relationship_type=row["relationship_type"], - source_path=row["source_path"], - target_path=row["target_path"], - ) - for row in cursor - ] - except sqlite3.Error as e: - raise QueryError(f"neighbors lookup failed: {e}") from e - - -def apply_dual_gate_governance( - results: list[QueryResult], - max_seats: int = 3, - max_chars: int = 4000, - quiet_keywords: list[str] | None = None, - *, - eligibility_manifest: dict[str, Any] | None, -) -> list[QueryResult]: - """Return manifest-approved, seat-bounded context without an ungated fallback. - - This is a consumer-side boundary over already-ranked retrieval nominations. - It does not alter ``run_query`` ranking or make a truth/compliance claim. - ``quiet_keywords`` remains a caller-directed ordering input, never a way to - admit a result that lacks C-0 membership. - """ - manifest_run_id, approved = _approved_manifest_records(eligibility_manifest) - if max_seats < 0 or max_chars < 0: - raise QueryError("max_seats and max_chars must be non-negative") - if max_seats == 0: - return [] - - eligible: list[QueryResult] = [] - for result in results: - approved_record = approved.get(result.doc_id) - if approved_record is None: - continue - source_path = _result_source_path(result) - content_hash = _indexed_hash(result.content_hash) - if source_path != approved_record["path"]: - continue - if content_hash != approved_record["sha256"]: - continue - eligible.append(result.model_copy(update={"eligibility_run_id": manifest_run_id})) - - if quiet_keywords and len(eligible) > max_seats: - shortlisted = eligible[:max_seats] - for candidate in eligible[max_seats:]: - candidate_text = (candidate.chunk_text or "").lower() - if any(keyword.lower() in candidate_text for keyword in quiet_keywords): - shortlisted[-1] = candidate - break - else: - shortlisted = eligible[:max_seats] - - governed: list[QueryResult] = [] - current_chars = 0 - for item in shortlisted: - passage = item.chunk_text or "" - if current_chars + len(passage) <= max_chars: - governed.append(item) - current_chars += len(passage) - else: - remaining = max_chars - current_chars - if remaining > 100: - governed.append(item.model_copy(update={"chunk_text": passage[:remaining] + "..."})) - break - return governed - - -def _normalized_path(path: str) -> str: - """Normalize a source path for manifest membership comparison only.""" - normalized = path.replace("\\", "/") - while "//" in normalized: - normalized = normalized.replace("//", "/") - return normalized.rstrip("/") - - -def _result_source_path(result: QueryResult) -> str | None: - if result.source_root and result.source_path: - return _normalized_path(f"{result.source_root}/{result.source_path}") - if result.path: - return _normalized_path(result.path) - return None - - -def _indexed_hash(value: str | None) -> str | None: - if value is None: - return None - normalized = value.strip().lower() - if normalized.startswith("sha256:"): - normalized = normalized.removeprefix("sha256:") - return normalized if _SHA256_HEX.fullmatch(normalized) else None - - -def _approved_manifest_records( - eligibility_manifest: dict[str, Any] | None, -) -> tuple[str, dict[str, dict[str, str]]]: - """Validate the C-0 public boundary contract and index it by document id.""" - if not isinstance(eligibility_manifest, dict): - raise QueryError("governed context requires a C-0 eligibility manifest") - run_id = eligibility_manifest.get("eligibility_run_id") - if not isinstance(run_id, str) or not run_id.strip(): - raise QueryError("C-0 eligibility manifest has no eligibility_run_id") - inventory = eligibility_manifest.get("approved_inventory") - if not isinstance(inventory, list) or not inventory: - raise QueryError("C-0 eligibility manifest approved_inventory is empty") - - records: dict[str, dict[str, str]] = {} - for index, raw_record in enumerate(inventory): - if not isinstance(raw_record, dict): - raise QueryError(f"C-0 approved record {index} is not an object") - doc_id = raw_record.get("doc_id") - path = raw_record.get("path") - sha256 = raw_record.get("sha256") - status = raw_record.get("status") - if not isinstance(doc_id, str) or not doc_id: - raise QueryError(f"C-0 approved record {index} has no doc_id") - if doc_id in records: - raise QueryError(f"C-0 approved inventory has duplicate doc_id '{doc_id}'") - if not isinstance(path, str) or not path: - raise QueryError(f"C-0 approved record '{doc_id}' has no source path") - if not isinstance(sha256, str) or not _SHA256_TAGGED.fullmatch(sha256): - raise QueryError( - f"C-0 approved record '{doc_id}' has malformed sha256; " - "require sha256:<64 lowercase hex>" - ) - if status not in ("Approved", "Effective"): - raise QueryError( - f"C-0 approved record '{doc_id}' has invalid status '{status}'" - ) - records[doc_id] = { - "path": _normalized_path(path), - "sha256": sha256.removeprefix("sha256:"), - } - return run_id.strip(), records diff --git a/mindgraph/src/mindgraph/routing.py b/mindgraph/src/mindgraph/routing.py deleted file mode 100644 index ba8dffc..0000000 --- a/mindgraph/src/mindgraph/routing.py +++ /dev/null @@ -1,901 +0,0 @@ -"""Deterministic intent-aware routing over separately ranked retrievers. - -This module adds an orchestration surface without changing the legacy query or -MCP result lists. MainFrame-specific paths and policy stay with the caller. -""" - -from __future__ import annotations - -import re -import sqlite3 -from collections.abc import Mapping -from dataclasses import dataclass -from pathlib import Path -from types import MappingProxyType -from typing import Any, Literal, Protocol -from urllib.parse import quote, urlsplit - -import sqlite_vec -from pydantic import ( - BaseModel, - ConfigDict, - Field, - field_validator, - model_validator, -) - -from mindgraph.embedders import EmbedTemplate, EmbedderSpec, format_query_text -from mindgraph.intent import ( - IntentGraphError, - IntentResolution, - TraversalLimits, - normalize_text, - resolve_intent, -) -from mindgraph.models import QueryResult -from mindgraph import db as db_mod -from mindgraph import query as query_mod - - -SCHEMA_VERSION = "1" -STABLE_ID_RE = re.compile(r"^[a-z][a-z0-9]*(?:[._-][a-z0-9]+)*$") -TOKEN_RE = re.compile(r"\w+", re.UNICODE) - -RouteMode = Literal["auto", "durable", "projects", "federated"] -DecisionOutcome = Literal["selected", "fallback", "refusal"] -EnvelopeOutcome = Literal["success", "partial", "refusal", "fallback"] - - -class RoutingError(ValueError): - """Invalid routing configuration with a stable machine-readable code.""" - - def __init__(self, code: str, message: str) -> None: - super().__init__(message) - self.code = code - - def __str__(self) -> str: - return f"{self.code}: {super().__str__()}" - - -class _StrictModel(BaseModel): - model_config = ConfigDict(extra="forbid", frozen=True) - - -def _nonempty(value: str) -> str: - value = value.strip() - if not value: - raise ValueError("must not be empty") - return value - - -def _stable_id(value: str) -> str: - value = value.strip() - if not STABLE_ID_RE.fullmatch(value): - raise ValueError( - "must be a lowercase dot/underscore/hyphen-separated stable ID" - ) - return value - - -def _capability_ref(value: str) -> str: - value = value.strip() - parsed = urlsplit(value) - if parsed.scheme not in {"capability", "retriever"} or not parsed.netloc: - raise ValueError("must be a capability:// or retriever:// reference") - return value - - -def _unique_tuple(value: tuple[str, ...], *, label: str) -> tuple[str, ...]: - if len(value) != len(set(value)): - raise ValueError(f"{label} may not contain duplicates") - return value - - -class RouteLimits(_StrictModel): - final_top_k: int = Field(default=10, ge=1, le=1000) - total_timeout_ms: int = Field(default=30000, ge=1, le=600000) - - -class RouteRequest(_StrictModel): - schema_version: Literal["1"] = "1" - query_id: str - query_text: str - mode: RouteMode = "auto" - intent_id: str | None = None - scope: str | None = None - allowed_capability_refs: tuple[str, ...] = () - limits: RouteLimits = RouteLimits() - - _query_id = field_validator("query_id")(_stable_id) - _query_text = field_validator("query_text")(_nonempty) - - @field_validator("intent_id", "scope") - @classmethod - def validate_optional_text(cls, value: str | None) -> str | None: - return _nonempty(value) if value is not None else None - - @field_validator("allowed_capability_refs") - @classmethod - def validate_capabilities(cls, value: tuple[str, ...]) -> tuple[str, ...]: - checked = tuple(_capability_ref(item) for item in value) - return _unique_tuple(checked, label="allowed_capability_refs") - - -class RefusalRule(_StrictModel): - rule_id: str - reason_code: str - all_terms: tuple[str, ...] = () - any_terms: tuple[str, ...] = () - phrases: tuple[str, ...] = () - - _rule_id = field_validator("rule_id")(_stable_id) - _reason_code = field_validator("reason_code")(_stable_id) - - @field_validator("all_terms", "any_terms", "phrases") - @classmethod - def validate_match_values(cls, value: tuple[str, ...]) -> tuple[str, ...]: - normalized = tuple(normalize_text(_nonempty(item)) for item in value) - return _unique_tuple(normalized, label="refusal rule match values") - - @model_validator(mode="after") - def validate_has_match(self) -> RefusalRule: - if not (self.all_terms or self.any_terms or self.phrases): - raise ValueError("refusal rule must declare a match condition") - return self - - def matches(self, query_text: str) -> bool: - normalized = normalize_text(query_text) - token_text = " ".join(TOKEN_RE.findall(normalized)) - - def contains(value: str) -> bool: - needle = " ".join(TOKEN_RE.findall(value)) - return bool(needle) and f" {needle} " in f" {token_text} " - - if any(not contains(term) for term in self.all_terms): - return False - if self.any_terms and not any(contains(term) for term in self.any_terms): - return False - if self.phrases and not any(contains(phrase) for phrase in self.phrases): - return False - return True - - -class RouterPolicy(_StrictModel): - policy_version: str - durable_retriever_id: str - project_retriever_id: str - allow_federated: bool = False - safe_default_retriever_id: str | None = None - refusal_rules: tuple[RefusalRule, ...] = () - - _policy_version = field_validator("policy_version")(_stable_id) - _retriever_ids = field_validator( - "durable_retriever_id", "project_retriever_id" - )(_stable_id) - - @field_validator("safe_default_retriever_id") - @classmethod - def validate_safe_default(cls, value: str | None) -> str | None: - return _stable_id(value) if value is not None else None - - @model_validator(mode="after") - def validate_rules(self) -> RouterPolicy: - rule_ids = tuple(rule.rule_id for rule in self.refusal_rules) - _unique_tuple(rule_ids, label="refusal rule IDs") - return self - - -class RetrieverDescriptor(_StrictModel): - retriever_id: str - capability_ref: str - trust_profile: str - source_surface: str - available: bool = True - unavailable_reason: str | None = None - - _retriever_id = field_validator("retriever_id")(_stable_id) - _capability = field_validator("capability_ref")(_capability_ref) - _labels = field_validator("trust_profile", "source_surface")(_nonempty) - - @field_validator("unavailable_reason") - @classmethod - def validate_unavailable_reason(cls, value: str | None) -> str | None: - return _stable_id(value) if value is not None else None - - @model_validator(mode="after") - def validate_availability(self) -> RetrieverDescriptor: - if self.available and self.unavailable_reason is not None: - raise ValueError("available retriever may not declare unavailable_reason") - if not self.available and self.unavailable_reason is None: - raise ValueError("unavailable retriever must declare unavailable_reason") - return self - - -class RetrieverExecutor(Protocol): - def retrieve(self, query_text: str, *, final_top_k: int) -> list[QueryResult]: ... - - -@dataclass(frozen=True) -class RetrieverRegistration: - descriptor: RetrieverDescriptor - executor: RetrieverExecutor | None = None - - def __post_init__(self) -> None: - if self.descriptor.available and self.executor is None: - raise RoutingError( - "routing_registry_invalid", - f"available retriever {self.descriptor.retriever_id} has no executor", - ) - if not self.descriptor.available and self.executor is not None: - raise RoutingError( - "routing_registry_invalid", - f"unavailable retriever {self.descriptor.retriever_id} has an executor", - ) - - -@dataclass(frozen=True, init=False) -class CapabilityRegistry: - """Immutable deterministic lookup for caller-provided retrievers.""" - - _registrations: tuple[RetrieverRegistration, ...] - _by_id: Mapping[str, RetrieverRegistration] - _by_capability: Mapping[str, RetrieverRegistration] - - def __init__(self, registrations: tuple[RetrieverRegistration, ...]) -> None: - ordered = tuple( - sorted(registrations, key=lambda item: item.descriptor.retriever_id) - ) - ids = tuple(item.descriptor.retriever_id for item in ordered) - capabilities = tuple(item.descriptor.capability_ref for item in ordered) - if len(ids) != len(set(ids)): - raise RoutingError( - "routing_registry_invalid", "duplicate retriever ID in registry" - ) - if len(capabilities) != len(set(capabilities)): - raise RoutingError( - "routing_registry_invalid", "duplicate capability ref in registry" - ) - object.__setattr__(self, "_registrations", ordered) - object.__setattr__( - self, - "_by_id", - MappingProxyType( - {item.descriptor.retriever_id: item for item in ordered} - ), - ) - object.__setattr__( - self, - "_by_capability", - MappingProxyType( - {item.descriptor.capability_ref: item for item in ordered} - ), - ) - - @property - def registrations(self) -> tuple[RetrieverRegistration, ...]: - return self._registrations - - @property - def descriptors(self) -> tuple[RetrieverDescriptor, ...]: - return tuple(item.descriptor for item in self._registrations) - - def by_id(self, retriever_id: str) -> RetrieverRegistration | None: - return self._by_id.get(retriever_id) - - def by_capability(self, capability_ref: str) -> RetrieverRegistration | None: - return self._by_capability.get(capability_ref) - - -class RouteRejection(_StrictModel): - target_id: str - reason_code: str - - _target = field_validator("target_id")(_nonempty) - _reason = field_validator("reason_code")(_stable_id) - - -class RouteDecision(_StrictModel): - schema_version: Literal["1"] = "1" - policy_version: str - outcome: DecisionOutcome - selected_retriever_ids: tuple[str, ...] = () - rejections: tuple[RouteRejection, ...] = () - reason_codes: tuple[str, ...] = () - intent_resolution: IntentResolution - effective_limits: RouteLimits - - _policy_version = field_validator("policy_version")(_stable_id) - - @field_validator("selected_retriever_ids", "reason_codes") - @classmethod - def validate_stable_values(cls, value: tuple[str, ...]) -> tuple[str, ...]: - checked = tuple(_stable_id(item) for item in value) - return _unique_tuple(checked, label="route decision values") - - @model_validator(mode="after") - def validate_decision_state(self) -> RouteDecision: - if not self.reason_codes: - raise ValueError("route decision must contain a reason code") - if self.selected_retriever_ids != tuple(sorted(self.selected_retriever_ids)): - raise ValueError("selected retriever IDs must be sorted") - if self.outcome == "refusal" and self.selected_retriever_ids: - raise ValueError("refusal may not select retrievers") - if self.outcome == "selected" and not self.selected_retriever_ids: - raise ValueError("selected outcome requires at least one retriever") - if self.outcome == "fallback" and len(self.selected_retriever_ids) != 1: - raise ValueError("fallback must select exactly one safe retriever") - rejection_keys = tuple( - (item.target_id, item.reason_code) for item in self.rejections - ) - if len(rejection_keys) != len(set(rejection_keys)): - raise ValueError("route rejections may not contain duplicates") - if rejection_keys != tuple(sorted(rejection_keys)): - raise ValueError("route rejections must be sorted") - return self - - -class RankedNomination(_StrictModel): - local_rank: int = Field(ge=1) - result: QueryResult - - -class RetrieverBatch(_StrictModel): - retriever_id: str - trust_profile: str - source_surface: str - assembly_reason: str - rows: tuple[RankedNomination, ...] = () - - _retriever_id = field_validator("retriever_id")(_stable_id) - _labels = field_validator( - "trust_profile", "source_surface", "assembly_reason" - )(_nonempty) - - @model_validator(mode="after") - def validate_local_ranks(self) -> RetrieverBatch: - ranks = tuple(row.local_rank for row in self.rows) - if ranks != tuple(range(1, len(self.rows) + 1)): - raise ValueError("batch local ranks must be contiguous and one-based") - return self - - -class RetrieverFailure(_StrictModel): - retriever_id: str - reason_code: Literal["retriever_failed"] = "retriever_failed" - error_type: str - message: str - - _retriever_id = field_validator("retriever_id")(_stable_id) - _error = field_validator("error_type", "message")(_nonempty) - - -class RetrievalEnvelope(_StrictModel): - schema_version: Literal["1"] = "1" - query_id: str - outcome: EnvelopeOutcome - decision: RouteDecision - batches: tuple[RetrieverBatch, ...] = () - failures: tuple[RetrieverFailure, ...] = () - partial: bool = False - warnings: tuple[str, ...] = () - - _query_id = field_validator("query_id")(_stable_id) - - @model_validator(mode="after") - def validate_partial_state(self) -> RetrievalEnvelope: - if self.partial != bool(self.failures): - raise ValueError("partial must exactly reflect whether failures exist") - if self.outcome == "partial" and not self.failures: - raise ValueError("partial outcome requires at least one failure") - if self.failures and self.outcome != "partial": - raise ValueError("retriever failures require partial outcome") - batch_ids = tuple(batch.retriever_id for batch in self.batches) - failure_ids = tuple(failure.retriever_id for failure in self.failures) - if batch_ids != tuple(sorted(batch_ids)): - raise ValueError("retriever batches must be sorted") - if failure_ids != tuple(sorted(failure_ids)): - raise ValueError("retriever failures must be sorted") - if set(batch_ids) & set(failure_ids): - raise ValueError("one retriever may not both succeed and fail") - selected = set(self.decision.selected_retriever_ids) - if set(batch_ids) | set(failure_ids) != selected: - raise ValueError("batches and failures must account for selected retrievers") - if self.decision.outcome == "refusal" and self.outcome != "refusal": - raise ValueError("refusal decision requires refusal envelope") - if self.outcome == "success" and self.decision.outcome != "selected": - raise ValueError("success envelope requires selected decision") - if self.outcome == "fallback" and self.decision.outcome != "fallback": - raise ValueError("fallback envelope requires fallback decision") - return self - - def as_contract_result(self, result_id: str) -> dict[str, Any]: - """Return the additive candidate shape used by the Phase 1 evaluator.""" - selected = list(self.decision.selected_retriever_ids) - trust_profiles = sorted({batch.trust_profile for batch in self.batches}) - source_paths: set[str] = set() - for batch in self.batches: - for row in batch.rows: - path = row.result.display_path or row.result.path or row.result.source_path - if path: - source_paths.add(path) - - behaviors: set[str] = set() - if len(self.batches) > 1: - behaviors.update( - { - "group results by retriever and trust profile", - "preserve per-retriever ranks", - "explain why each retriever was called", - } - ) - outcome = self.outcome - if outcome in {"success", "partial"}: - contract_outcome = "success" - else: - contract_outcome = outcome - if ( - contract_outcome == "refusal" - and self.decision.reason_codes - and self.decision.reason_codes[0] - == "raw_transcripts_not_in_default_retrieval" - ): - contract_outcome = "refusal_or_explicit_consent_gate" - fields = ( - ["retriever_id", "trust_profile", "source_surface", "assembly_reason"] - if contract_outcome - not in {"refusal", "refusal_or_explicit_consent_gate", "fallback"} - else [] - ) - payload: dict[str, Any] = { - "id": result_id, - "outcome": contract_outcome, - "selected_retrievers": selected, - "trust_profiles": trust_profiles, - "source_paths": sorted(source_paths), - "behaviors": sorted(behaviors), - "fields": fields, - } - if self.decision.reason_codes: - payload["reason"] = self.decision.reason_codes[0] - return payload - - -def _intent_unavailable_resolution(reason: str = "intent_unavailable") -> IntentResolution: - return IntentResolution( - graph_id="unavailable", - graph_version="unavailable", - source_hash="unavailable", - outcome="fallback", - resolution_method="none", - refusal_reason=reason, - warnings=(reason,), - ) - - -def _validate_policy_registry( - registry: CapabilityRegistry, policy: RouterPolicy -) -> None: - required = {policy.durable_retriever_id, policy.project_retriever_id} - if policy.safe_default_retriever_id is not None: - required.add(policy.safe_default_retriever_id) - missing = sorted(item for item in required if registry.by_id(item) is None) - if missing: - raise RoutingError( - "routing_policy_invalid", - "policy references missing retrievers: " + ", ".join(missing), - ) - - -def _make_decision( - *, - policy: RouterPolicy, - request: RouteRequest, - resolution: IntentResolution, - outcome: DecisionOutcome, - selected: tuple[str, ...] = (), - rejections: tuple[RouteRejection, ...] = (), - reasons: tuple[str, ...] = (), -) -> RouteDecision: - return RouteDecision( - policy_version=policy.policy_version, - outcome=outcome, - selected_retriever_ids=tuple(sorted(selected)), - rejections=tuple(sorted(rejections, key=lambda item: (item.target_id, item.reason_code))), - reason_codes=_unique_tuple(reasons, label="route reason codes"), - intent_resolution=resolution, - effective_limits=request.limits, - ) - - -def _explicit_decision( - request: RouteRequest, - resolution: IntentResolution, - registry: CapabilityRegistry, - policy: RouterPolicy, -) -> RouteDecision: - if request.mode == "federated" and not policy.allow_federated: - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="refusal", - reasons=("federation_not_allowed",), - ) - mode_ids = { - "durable": (policy.durable_retriever_id,), - "projects": (policy.project_retriever_id,), - "federated": tuple( - sorted({policy.durable_retriever_id, policy.project_retriever_id}) - ), - } - reasons = { - "durable": ("explicit_durable_scope",), - "projects": ("explicit_project_scope",), - "federated": ("explicit_federated_scope",), - } - requested = mode_ids[request.mode] - selected: list[str] = [] - rejected: list[RouteRejection] = [] - for retriever_id in requested: - registration = registry.by_id(retriever_id) - assert registration is not None # checked by _validate_policy_registry - if registration.descriptor.available: - selected.append(retriever_id) - else: - rejected.append( - RouteRejection( - target_id=retriever_id, reason_code="retriever_unavailable" - ) - ) - if rejected: - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="refusal", - rejections=tuple(rejected), - reasons=("retriever_unavailable",), - ) - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="selected", - selected=tuple(selected), - reasons=reasons[request.mode], - ) - - -def _safe_default_decision( - request: RouteRequest, - resolution: IntentResolution, - registry: CapabilityRegistry, - policy: RouterPolicy, -) -> RouteDecision | None: - retriever_id = policy.safe_default_retriever_id - if retriever_id is None: - return None - registration = registry.by_id(retriever_id) - assert registration is not None # checked by _validate_policy_registry - if not registration.descriptor.available: - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="refusal", - rejections=( - RouteRejection( - target_id=retriever_id, reason_code="retriever_unavailable" - ), - ), - reasons=("retriever_unavailable",), - ) - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="fallback", - selected=(retriever_id,), - reasons=("safe_default",), - ) - - -def decide_route( - request: RouteRequest, - resolution: IntentResolution, - registry: CapabilityRegistry, - policy: RouterPolicy, -) -> RouteDecision: - """Select retrievers deterministically without executing them.""" - _validate_policy_registry(registry, policy) - for rule in policy.refusal_rules: - if rule.matches(request.query_text): - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="refusal", - reasons=(rule.reason_code,), - ) - - if request.mode != "auto": - return _explicit_decision(request, resolution, registry, policy) - - if resolution.outcome == "refusal": - reason = resolution.refusal_reason or "intent_invalid" - if reason in {"intent_alias_ambiguous", "intent_rule_ambiguous"}: - reason = "intent_ambiguous" - elif reason != "intent_ambiguous": - reason = "intent_invalid" - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="refusal", - reasons=(reason,), - ) - - if resolution.outcome == "fallback": - default = _safe_default_decision(request, resolution, registry, policy) - if default is not None: - return default - reason = ( - "intent_unavailable" - if resolution.refusal_reason == "intent_unavailable" - else "intent_no_match" - ) - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="refusal", - reasons=(reason,), - ) - - allowed = frozenset(request.allowed_capability_refs) - selected: set[str] = set() - rejected: list[RouteRejection] = [] - for capability_ref in resolution.rejected_capability_hints: - reason = "capability_not_allowed" - if f"binding_unavailable:{capability_ref}" in resolution.warnings: - reason = "capability_unavailable" - rejected.append(RouteRejection(target_id=capability_ref, reason_code=reason)) - for capability_ref in resolution.capability_hints: - if capability_ref not in allowed: - rejected.append( - RouteRejection( - target_id=capability_ref, reason_code="capability_not_allowed" - ) - ) - continue - registration = registry.by_capability(capability_ref) - if registration is None: - rejected.append( - RouteRejection( - target_id=capability_ref, reason_code="capability_unavailable" - ) - ) - continue - if not registration.descriptor.available: - rejected.append( - RouteRejection( - target_id=registration.descriptor.retriever_id, - reason_code="retriever_unavailable", - ) - ) - continue - selected.add(registration.descriptor.retriever_id) - - if len(selected) > 1 and not policy.allow_federated: - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="refusal", - rejections=tuple(rejected), - reasons=("federation_not_allowed",), - ) - if selected: - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="selected", - selected=tuple(selected), - rejections=tuple(rejected), - reasons=("intent_capability_hint",), - ) - if rejected: - reason_codes = {item.reason_code for item in rejected} - reason = ( - "capability_unavailable" - if reason_codes & {"capability_unavailable", "retriever_unavailable"} - else "capability_not_allowed" - ) - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="refusal", - rejections=tuple(rejected), - reasons=(reason,), - ) - default = _safe_default_decision(request, resolution, registry, policy) - if default is not None: - return default - return _make_decision( - policy=policy, - request=request, - resolution=resolution, - outcome="refusal", - reasons=("intent_no_match",), - ) - - -def open_query_store_read_only(path: Path) -> sqlite3.Connection: - """Open an existing document index without persisting connection changes.""" - resolved = Path(path).expanduser().resolve() - if not resolved.is_file(): - raise RoutingError( - "retriever_unavailable", f"query store does not exist: {resolved}" - ) - try: - return db_mod.get_db(str(resolved), read_only=True) - except (db_mod.DatabaseError, sqlite3.Error, RuntimeError) as exc: - raise RoutingError( - "retriever_unavailable", f"failed to open query store: {exc}" - ) from exc - - -@dataclass(frozen=True) -class MindGraphQueryRetriever: - descriptor: RetrieverDescriptor - db_path: Path - embedder: Any - embedder_spec: EmbedderSpec | None = None - embed_template: EmbedTemplate = "none" - lexical_top_k: int = query_mod.DEFAULT_LEXICAL_TOP_K - semantic_top_k: int = query_mod.DEFAULT_SEMANTIC_TOP_K - expand: bool = False - expand_depth: int = query_mod.DEFAULT_EXPAND_DEPTH - expand_top_k: int = query_mod.DEFAULT_EXPAND_TOP_K - associate: bool = False - associate_top_k: int = query_mod.DEFAULT_ASSOCIATE_TOP_K - associate_seed_k: int = query_mod.DEFAULT_ASSOCIATE_SEED_K - - def __post_init__(self) -> None: - if not self.descriptor.available: - raise RoutingError( - "routing_registry_invalid", - "MindGraphQueryRetriever requires an available descriptor", - ) - - def retrieve(self, query_text: str, *, final_top_k: int) -> list[QueryResult]: - conn = open_query_store_read_only(self.db_path) - try: - formatted = ( - format_query_text( - self.embedder_spec, query_text, template=self.embed_template - ) - if self.embedder_spec is not None - else query_text - ) - return query_mod.run_query( - conn, - formatted, - self.embedder, - lexical_top_k=self.lexical_top_k, - semantic_top_k=self.semantic_top_k, - final_top_k=final_top_k, - expand=self.expand, - expand_depth=self.expand_depth, - expand_top_k=self.expand_top_k, - associate=self.associate, - associate_top_k=self.associate_top_k, - associate_seed_k=self.associate_seed_k, - embedder_spec=self.embedder_spec, - embed_template=self.embed_template, - ) - finally: - conn.close() - - def registration(self) -> RetrieverRegistration: - return RetrieverRegistration(descriptor=self.descriptor, executor=self) - - -def orchestrate( - request: RouteRequest, - *, - intent_conn: sqlite3.Connection | None, - registry: CapabilityRegistry, - policy: RouterPolicy, -) -> RetrievalEnvelope: - """Resolve intent, decide a route, and return separately ranked batches.""" - if intent_conn is None: - resolution = _intent_unavailable_resolution() - else: - try: - resolution = resolve_intent( - intent_conn, - request.query_text, - intent_id=request.intent_id, - scope=request.scope, - allowed_capability_refs=frozenset(request.allowed_capability_refs), - limits=TraversalLimits(), - ) - except (IntentGraphError, sqlite3.Error): - resolution = _intent_unavailable_resolution() - - decision = decide_route(request, resolution, registry, policy) - if not decision.selected_retriever_ids: - return RetrievalEnvelope( - query_id=request.query_id, - outcome=decision.outcome, - decision=decision, - warnings=tuple(sorted(set(resolution.warnings))), - ) - - batches: list[RetrieverBatch] = [] - failures: list[RetrieverFailure] = [] - for retriever_id in decision.selected_retriever_ids: - registration = registry.by_id(retriever_id) - assert registration is not None and registration.executor is not None - try: - results = registration.executor.retrieve( - request.query_text, final_top_k=request.limits.final_top_k - ) - mismatched = sorted( - { - result.trust_profile - for result in results - if result.trust_profile is not None - and result.trust_profile - != registration.descriptor.trust_profile - } - ) - if mismatched: - raise RoutingError( - "retriever_provenance_mismatch", - f"{retriever_id} returned unexpected trust profiles: " - + ", ".join(mismatched), - ) - rows = tuple( - RankedNomination(local_rank=index, result=result) - for index, result in enumerate(results, start=1) - ) - batches.append( - RetrieverBatch( - retriever_id=retriever_id, - trust_profile=registration.descriptor.trust_profile, - source_surface=registration.descriptor.source_surface, - assembly_reason=decision.reason_codes[0], - rows=rows, - ) - ) - except Exception as exc: # noqa: BLE001 - failures belong in the envelope - failures.append( - RetrieverFailure( - retriever_id=retriever_id, - error_type=type(exc).__name__, - message=str(exc), - ) - ) - - if failures: - outcome: EnvelopeOutcome = "partial" - elif decision.outcome == "fallback": - outcome = "fallback" - else: - outcome = "success" - warnings = set(resolution.warnings) - if failures: - warnings.add("retriever_failed") - return RetrievalEnvelope( - query_id=request.query_id, - outcome=outcome, - decision=decision, - batches=tuple(batches), - failures=tuple(failures), - partial=bool(failures), - warnings=tuple(sorted(warnings)), - ) diff --git a/mindgraph/tests/__init__.py b/mindgraph/tests/__init__.py deleted file mode 100644 index e69de29..0000000 diff --git a/mindgraph/tests/fixtures/intent_graph_cases.yaml b/mindgraph/tests/fixtures/intent_graph_cases.yaml deleted file mode 100644 index 6fbad8a..0000000 --- a/mindgraph/tests/fixtures/intent_graph_cases.yaml +++ /dev/null @@ -1,280 +0,0 @@ -source_ref: &source_ref "plan://plans/retrieval-remodel.md" -reviewer: &reviewer "operator:example" - -documents: - phase1: - schema_version: "1" - graph: - id: "mainframe.core" - version: "2026-06-29.1" - status: "approved" - created_at: "2026-06-29T00:00:00Z" - reviewed_at: "2026-06-29T00:00:00Z" - reviewed_by: [*reviewer] - source_refs: [*source_ref] - supersedes: null - nodes: - - id: "goal.mindgraph-remodel" - kind: "goal" - label: "MindGraph remodel" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "goal.route-contract-evaluation" - kind: "goal" - label: "Route contract evaluation" - aliases: ["Run the route contract checks."] - status: "active" - source_refs: [*source_ref] - - id: "goal.intent-schema-validation" - kind: "goal" - label: "Intent schema validation" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "goal.project-status-review" - kind: "goal" - label: "Project status review" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "goal.deep-root" - kind: "goal" - label: "Deep root" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "goal.deep-1" - kind: "goal" - label: "Deep one" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "goal.deep-2" - kind: "goal" - label: "Deep two" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "goal.deep-3" - kind: "goal" - label: "Deep three" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "capability.fixture-evaluator" - kind: "capability" - label: "Fixture evaluator" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "capability.mindgraph-projects" - kind: "capability" - label: "MindGraph projects retriever" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "capability.gmail" - kind: "capability" - label: "Gmail retriever" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "capability.optional" - kind: "capability" - label: "Optional unavailable capability" - aliases: [] - status: "active" - source_refs: [*source_ref] - - id: "constraint.review-required" - kind: "constraint" - label: "Review required" - aliases: [] - status: "active" - source_refs: [*source_ref] - edges: - - source_id: "goal.mindgraph-remodel" - target_id: "goal.route-contract-evaluation" - relation: "requires" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.route-contract-evaluation" - target_id: "goal.intent-schema-validation" - relation: "requires" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.mindgraph-remodel" - target_id: "capability.mindgraph-projects" - relation: "routes_to" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.mindgraph-remodel" - target_id: "capability.fixture-evaluator" - relation: "routes_to" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.route-contract-evaluation" - target_id: "capability.fixture-evaluator" - relation: "routes_to" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.project-status-review" - target_id: "capability.mindgraph-projects" - relation: "routes_to" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.project-status-review" - target_id: "capability.gmail" - relation: "routes_to" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.project-status-review" - target_id: "capability.optional" - relation: "routes_to" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.mindgraph-remodel" - target_id: "constraint.review-required" - relation: "blocked_by" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.deep-root" - target_id: "goal.deep-1" - relation: "requires" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.deep-1" - target_id: "goal.deep-2" - relation: "requires" - status: "active" - source_refs: [*source_ref] - - source_id: "goal.deep-2" - target_id: "goal.deep-3" - relation: "requires" - status: "active" - source_refs: [*source_ref] - bindings: - - id: "binding.fixture-evaluator" - node_id: "capability.fixture-evaluator" - ref: "capability://fixture-evaluator" - required: true - availability: "available" - source_refs: ["project://mindgraph-eval"] - - id: "binding.mindgraph-projects" - node_id: "capability.mindgraph-projects" - ref: "retriever://mindgraph-projects" - required: true - availability: "available" - source_refs: ["project://mindgraph"] - - id: "binding.gmail" - node_id: "capability.gmail" - ref: "retriever://gmail" - required: false - availability: "available" - source_refs: ["project://mindgraph"] - - id: "binding.optional" - node_id: "capability.optional" - ref: "capability://optional-unavailable" - required: false - availability: "unavailable" - source_refs: ["project://mindgraph"] - rules: - - id: "rule.project-status-review" - priority: 100 - goal_id: "goal.project-status-review" - match: - scope: "project_status" - all_terms: ["project", "status"] - any_terms: ["review", "next"] - source_refs: ["decision://DECISIONS.md"] - - predecessor: - schema_version: "1" - graph: - id: "mainframe.core" - version: "2026-06-28.1" - status: "approved" - created_at: "2026-06-28T00:00:00Z" - reviewed_at: "2026-06-28T00:00:00Z" - reviewed_by: [*reviewer] - source_refs: [*source_ref] - supersedes: null - nodes: - - id: "goal.baseline" - kind: "goal" - label: "Baseline" - aliases: [] - status: "active" - source_refs: [*source_ref] - edges: [] - bindings: [] - rules: [] - -invalid_cases: - duplicate_id: - expected_error: "intent_duplicate_id" - alias_collision: - expected_error: "intent_alias_ambiguous" - missing_target: - expected_error: "intent_missing_target" - invalid_relation: - expected_error: "intent_relation_invalid" - kind_mismatch: - expected_error: "intent_kind_mismatch" - requires_cycle: - expected_error: "intent_cycle_detected" - expected_witness: ["goal.cyclic-a", "goal.cyclic-b", "goal.cyclic-a"] - decomposes_cycle: - expected_error: "intent_cycle_detected" - next_step_cycle: - expected_error: "intent_cycle_detected" - missing_predecessor: - expected_error: "intent_version_missing" - version_fork: - expected_error: "intent_version_fork" - changed_approved_version: - expected_error: "intent_version_conflict" - required_unavailable_binding: - expected_error: "intent_binding_unavailable" - ambiguous_rule: - expected_error: "intent_rule_ambiguous" - -phase1_expected: - intent01_explicit_goal_prerequisites: - outcome: "resolved" - resolution_method: "explicit" - matched_goal_ids: ["goal.mindgraph-remodel"] - prerequisite_goal_ids: - - "goal.route-contract-evaluation" - - "goal.intent-schema-validation" - intent_path: - - "goal.mindgraph-remodel" - - "goal.route-contract-evaluation" - - "goal.intent-schema-validation" - capability_hints: - - "capability://fixture-evaluator" - - "retriever://mindgraph-projects" - intent02_deterministic_alias_resolution: - outcome: "resolved" - resolution_method: "alias" - matched_goal_ids: ["goal.route-contract-evaluation"] - prerequisite_goal_ids: [] - intent_path: ["goal.route-contract-evaluation"] - capability_hints: ["capability://fixture-evaluator"] - intent03_ambiguous_alias_refusal: - outcome: "refusal" - reason: "intent_alias_ambiguous" - intent04_cycle_refusal: - outcome: "refusal" - reason: "intent_cycle_detected" - intent05_missing_goal_fallback: - outcome: "fallback" - reason: "intent_no_match" - intent06_untrusted_route_hint_rejected: - outcome: "resolved" - capability_hints: ["retriever://mindgraph-projects"] - forbidden_capability_hints: ["retriever://gmail"] - intent07_bounded_traversal_truncation: - outcome: "resolved" - reason: "intent_traversal_truncated" - intent_path: ["goal.deep-root", "goal.deep-1", "goal.deep-2"] diff --git a/mindgraph/tests/fixtures/phase3_routing_cases.yaml b/mindgraph/tests/fixtures/phase3_routing_cases.yaml deleted file mode 100644 index 0e4447d..0000000 --- a/mindgraph/tests/fixtures/phase3_routing_cases.yaml +++ /dev/null @@ -1,59 +0,0 @@ -policy: - policy_version: "mainframe.router-v1" - durable_retriever_id: "mindgraph-durable" - project_retriever_id: "mindgraph-projects" - allow_federated: true - safe_default_retriever_id: "mindgraph-durable" - refusal_rules: - - rule_id: "privacy.raw-transcript" - reason_code: "raw_transcripts_not_in_default_retrieval" - phrases: - - "raw session transcript" - - "raw transcript" - -requests: - explicit_durable: - request: - query_id: "route.explicit-durable" - query_text: "Explain option exercise and assignment." - mode: "durable" - expected_outcome: "selected" - expected_retrievers: ["mindgraph-durable"] - expected_reason: "explicit_durable_scope" - - explicit_projects: - request: - query_id: "route.explicit-projects" - query_text: "What is the next approved implementation slice?" - mode: "projects" - expected_outcome: "selected" - expected_retrievers: ["mindgraph-projects"] - expected_reason: "explicit_project_scope" - - explicit_federated: - request: - query_id: "route.explicit-federated" - query_text: "Plan using durable methods and project status." - mode: "federated" - expected_outcome: "selected" - expected_retrievers: ["mindgraph-durable", "mindgraph-projects"] - expected_reason: "explicit_federated_scope" - - automatic_project_hint: - request: - query_id: "route.auto-project" - query_text: "Project status review" - mode: "auto" - allowed_capability_refs: ["retriever://mindgraph-projects"] - expected_outcome: "selected" - expected_retrievers: ["mindgraph-projects"] - expected_reason: "intent_capability_hint" - - raw_transcript_refusal: - request: - query_id: "route.raw-transcript" - query_text: "Search every raw session transcript for the exact prompt." - mode: "durable" - expected_outcome: "refusal" - expected_retrievers: [] - expected_reason: "raw_transcripts_not_in_default_retrieval" diff --git a/mindgraph/tests/test_associate.py b/mindgraph/tests/test_associate.py deleted file mode 100644 index 992506b..0000000 --- a/mindgraph/tests/test_associate.py +++ /dev/null @@ -1,128 +0,0 @@ -"""Phase 9 — semantic association tests (ADR-034).""" - -import pytest - -from mindgraph import cli, db, parser -from mindgraph.query import associate_results, run_query - - -class KeywordEmbedder: - def __init__(self, keyword_to_dim: dict[str, int], dims: int = 384): - self.keyword_to_dim = {k.lower(): v for k, v in keyword_to_dim.items()} - self.dims = dims - - def encode(self, texts, convert_to_numpy=True): - import numpy as np - - out = np.zeros((len(texts), self.dims), dtype=np.float32) - for i, text in enumerate(texts): - lower = text.lower() - for kw, dim in self.keyword_to_dim.items(): - if kw in lower: - out[i, dim] += 1.0 - return out - - -@pytest.fixture -def associate_embedder(): - return KeywordEmbedder({"gamma": 0, "beta": 1, "zzunique": 2}) - - -@pytest.fixture -def associate_db(tmp_path, monkeypatch, associate_embedder): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: associate_embedder) - - notes = tmp_path / "vault" - notes.mkdir() - (notes / "seed.md").write_text( - "# Seed\nzzunique anchor. gamma protocol overview. [[linked]] (refers).\n" - ) - (notes / "neighbor.md").write_text( - "# Neighbor\ngamma protocol details without explicit link.\n" - ) - (notes / "linked.md").write_text("# Linked\nbeta content only.\n") - - db_path = str(tmp_path / "test.sqlite") - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - return db_path - - -class TestAssociate: - def test_associate_finds_semantic_neighbor_not_in_seeds( - self, associate_db, associate_embedder - ): - conn = db.get_db(associate_db) - try: - fused = run_query( - conn, - "zzunique", - associate_embedder, - final_top_k=1, - ) - assert len(fused) == 1 - assert fused[0].path.endswith("seed.md") - - associated = associate_results( - conn, - fused, - associate_embedder, - top_k=5, - seed_k=1, - ) - paths = [r.path for r in associated] - assert any(p.endswith("neighbor.md") for p in paths) - assert all(r.signal == "associated" for r in associated) - assert all(r.association_depth == 1 for r in associated) - assert all(r.rrf_score == 0.0 for r in associated) - finally: - conn.close() - - def test_associate_cli_flag(self, associate_db): - from typer.testing import CliRunner - - runner = CliRunner() - result = runner.invoke( - cli.app, - [ - "query", - "zzunique", - "--db", - associate_db, - "--associate", - "--top-k", - "1", - "--associate-top-k", - "5", - "--json", - ], - ) - assert result.exit_code == 0 - import json - - rows = json.loads(result.stdout) - signals = {row["signal"] for row in rows} - assert "fused" in signals or "lexical" in signals - assert "associated" in signals - - def test_expand_and_associate_are_distinct(self, associate_db, associate_embedder): - conn = db.get_db(associate_db) - try: - rows = run_query( - conn, - "zzunique", - associate_embedder, - final_top_k=1, - expand=True, - expand_top_k=5, - associate=True, - associate_top_k=5, - ) - expanded = [r for r in rows if r.signal == "expanded"] - associated = [r for r in rows if r.signal == "associated"] - assert expanded - assert associated - assert any(r.path.endswith("linked.md") for r in expanded) - assert any(r.path.endswith("neighbor.md") for r in associated) - finally: - conn.close() \ No newline at end of file diff --git a/mindgraph/tests/test_doctor.py b/mindgraph/tests/test_doctor.py deleted file mode 100644 index 51d8ecf..0000000 --- a/mindgraph/tests/test_doctor.py +++ /dev/null @@ -1,162 +0,0 @@ -"""Tests for mindgraph doctor/status and query schema fail-fast (MH01).""" - -from __future__ import annotations - -import json -import os -import sqlite3 -from pathlib import Path - -import pytest -from typer.testing import CliRunner - -from mindgraph import cli, db - - -@pytest.fixture -def runner(): - return CliRunner() - - -def test_inspect_healthy_db(tmp_path): - path = str(tmp_path / "ok.sqlite") - db.init_db(path).close() - report = db.inspect_database(path, role="test", trust_profile="t") - assert report["ok"] is True - assert report["exists"] is True - assert report["tables_missing"] == [] - assert report["counts"]["documents"] == 0 - - -def test_inspect_read_only_wal_db_without_writable_directory(tmp_path): - locked = tmp_path / "locked" - locked.mkdir() - path = locked / "index.sqlite" - db.init_db(str(path)).close() - os.chmod(path, 0o444) - os.chmod(locked, 0o555) - try: - report = db.inspect_database(str(path), role="projects") - assert report["ok"] is True, report - conn = db.get_db(str(path), read_only=True) - try: - assert conn.execute("PRAGMA query_only").fetchone()[0] == 1 - finally: - conn.close() - finally: - os.chmod(locked, 0o755) - os.chmod(path, 0o644) - - -def test_inspect_missing_file(tmp_path): - path = str(tmp_path / "nope.sqlite") - report = db.inspect_database(path) - assert report["ok"] is False - assert "file_missing" in report["issues"] - - -def test_inspect_stub_empty_sqlite(tmp_path): - """Tiny file without MindGraph tables fails schema check.""" - path = tmp_path / "stub.sqlite" - # Create a nearly-empty sqlite (no mindgraph tables) - conn = sqlite3.connect(str(path)) - conn.execute("CREATE TABLE junk (id INTEGER)") - conn.commit() - conn.close() - report = db.inspect_database(str(path)) - assert report["ok"] is False - assert "missing_required_tables" in report["issues"] - assert "documents_fts" in report["tables_missing"] - assert report["likely_stub"] is True - - -def test_validate_query_schema_raises(tmp_path): - from mindgraph.exceptions import DatabaseError - - path = str(tmp_path / "bad.sqlite") - raw = sqlite3.connect(path) - raw.execute("CREATE TABLE junk (id INTEGER)") - raw.commit() - raw.close() - conn = db.get_db(path) - try: - with pytest.raises(DatabaseError, match="missing tables"): - db.validate_query_schema(conn, path) - finally: - conn.close() - - -def test_find_workspace_stubs(tmp_path): - stub = tmp_path / "mainframe.sqlite" - stub.write_bytes(b"\x00" * 100) - hits = db.find_workspace_stub_sqlite([tmp_path]) - assert len(hits) == 1 - assert hits[0]["size_bytes"] == 100 - - -def test_doctor_cli_single_ok(tmp_path, runner): - path = tmp_path / "ok.sqlite" - db.init_db(str(path)).close() - result = runner.invoke(cli.app, ["doctor", "--db", str(path), "--json"]) - assert result.exit_code == 0, result.output - payload = json.loads(result.output) - assert payload["ok"] is True - assert len(payload["databases"]) == 1 - assert payload["databases"][0]["ok"] is True - - -def test_doctor_cli_single_fail(tmp_path, runner): - path = tmp_path / "bad.sqlite" - conn = sqlite3.connect(str(path)) - conn.execute("CREATE TABLE junk (id INTEGER)") - conn.commit() - conn.close() - result = runner.invoke(cli.app, ["doctor", "--db", str(path), "--json"]) - assert result.exit_code == 1 - payload = json.loads(result.output) - assert payload["ok"] is False - - -def test_status_alias(tmp_path, runner): - path = tmp_path / "ok.sqlite" - db.init_db(str(path)).close() - result = runner.invoke(cli.app, ["status", "--db", str(path), "--json"]) - assert result.exit_code == 0, result.output - - -def test_query_fail_fast_on_stub(tmp_path, runner, monkeypatch, caplog): - """Query on incomplete DB exits before embedder load.""" - import logging - - path = tmp_path / "stub.sqlite" - conn = sqlite3.connect(str(path)) - conn.execute("CREATE TABLE junk (id INTEGER)") - conn.commit() - conn.close() - - def boom(*_a, **_k): - raise AssertionError("embedder should not load on schema failure") - - monkeypatch.setattr(cli, "_load_embedder", boom) - with caplog.at_level(logging.ERROR, logger="mindgraph"): - result = runner.invoke(cli.app, ["query", "hello", "--db", str(path)]) - assert result.exit_code == 1 - combined = (result.output or "") + (result.stderr or "") + caplog.text - assert "missing tables" in combined.lower() or "usable MindGraph" in combined - - -def test_doctor_detects_workspace_stub_flag(tmp_path, runner): - path = tmp_path / "ok.sqlite" - db.init_db(str(path)).close() - stub = tmp_path / "mainframe.sqlite" - stub.write_bytes(b"\x00" * 200) - result = runner.invoke( - cli.app, - ["doctor", "--db", str(path), "--workspace", str(tmp_path), "--json"], - ) - # healthy DB → exit 0; stubs listed as non-fatal warnings - assert result.exit_code == 0, result.output - payload = json.loads(result.output) - assert payload["ok"] is True - assert payload["workspace_stubs"] - assert payload["databases"][0]["ok"] is True diff --git a/mindgraph/tests/test_embedders.py b/mindgraph/tests/test_embedders.py deleted file mode 100644 index 5d5932a..0000000 --- a/mindgraph/tests/test_embedders.py +++ /dev/null @@ -1,82 +0,0 @@ -"""Embedder registry and text formatting tests.""" - -import sys -import types - -import pytest - -from mindgraph import embedders -from mindgraph.exceptions import EmbeddingError - - -def test_resolve_default_embedder(): - spec = embedders.resolve_embedder(None) - assert spec.key == "minilm" - assert spec.dimensions == 384 - - -def test_resolve_bge_and_e5_prefixes(): - bge = embedders.resolve_embedder("bge-small") - assert bge.query_prefix == "query: " - e5 = embedders.resolve_embedder("e5-small") - assert e5.passage_prefix == "passage: " - - -def test_unknown_embedder_raises(): - with pytest.raises(ValueError, match="Unknown embedder"): - embedders.resolve_embedder("not-a-model") - - -def test_mainframe_template_prefixes(): - spec = embedders.resolve_embedder("minilm") - query = embedders.format_query_text( - spec, "regulated crosswalk", template="mainframe" - ) - assert query.startswith("[intent=query]") - passage = embedders.format_passage_text( - spec, - "chunk body", - template="mainframe", - title="GxP Note", - domain="regulated-systems", - doc_type="note", - ) - assert "[domain=regulated-systems]" in passage - assert "[type=note]" in passage - assert "GxP Note" in passage - - -def test_load_sentence_embedder_uses_local_files_only(monkeypatch): - """Loader must resolve models from the local cache only (no hub metadata).""" - captured: dict[str, object] = {} - - class FakeSentenceTransformer: - def __init__(self, model_name_or_path, *args, **kwargs): - captured["model_name_or_path"] = model_name_or_path - captured["args"] = args - captured["kwargs"] = kwargs - - fake_module = types.ModuleType("sentence_transformers") - fake_module.SentenceTransformer = FakeSentenceTransformer - monkeypatch.setitem(sys.modules, "sentence_transformers", fake_module) - - spec = embedders.resolve_embedder("minilm") - model = embedders.load_sentence_embedder(spec) - - assert isinstance(model, FakeSentenceTransformer) - assert captured["model_name_or_path"] == "all-MiniLM-L6-v2" - assert captured["kwargs"].get("local_files_only") is True - - -def test_load_sentence_embedder_missing_cache_raises_embedding_error(monkeypatch): - class BoomSentenceTransformer: - def __init__(self, *args, **kwargs): - raise OSError("model not found in local cache") - - fake_module = types.ModuleType("sentence_transformers") - fake_module.SentenceTransformer = BoomSentenceTransformer - monkeypatch.setitem(sys.modules, "sentence_transformers", fake_module) - - spec = embedders.resolve_embedder("minilm") - with pytest.raises(EmbeddingError, match="local_files_only=True"): - embedders.load_sentence_embedder(spec) diff --git a/mindgraph/tests/test_examples.py b/mindgraph/tests/test_examples.py deleted file mode 100644 index 9dee45e..0000000 --- a/mindgraph/tests/test_examples.py +++ /dev/null @@ -1,291 +0,0 @@ -"""Phase 4 — example-vault retrieval smoke (CI-safe). - -Asserts the expected retrieval surface against the committed -`examples/example-vault/` vault using the deterministic `KeywordEmbedder` -stub from `tests/test_query.py`. The real-model captures in the asset README -come from `scripts/run_example_smoke.py`, which uses -`sentence-transformers/all-MiniLM-L6-v2`; this test does not require network -access. - -Topology (locked in DECISIONS.md § 2026-05-21 — Phase 4 public packaging): - - feedback-loops.md → [[reinforcing-loops]] (illustrates), - [[balancing-loops]] (illustrates) - reinforcing-loops.md → [[systems-archetypes]] (related) - balancing-loops.md → [[systems-archetypes]] (related) - systems-archetypes.md → [[unicycle-mental-model]] (cites, dangling) - mental-models-overview.md → [[feedback-loops]] (related) - bounded-rationality.md → [[feedback-loops]] (related) - unrelated-noise.md → (no edges, off-theme) -""" - -import shutil -from pathlib import Path - -import numpy as np -import pytest - -from mindgraph import cli, db, parser -from mindgraph.query import list_neighbors, run_query - - -class KeywordEmbedder: - """Same deterministic embedder pattern as tests/test_query.py and tests/test_expand.py.""" - - def __init__(self, keyword_to_dim: dict[str, int], dims: int = 384): - self.keyword_to_dim = {k.lower(): v for k, v in keyword_to_dim.items()} - self.dims = dims - - def encode(self, texts, convert_to_numpy=True): - out = np.zeros((len(texts), self.dims), dtype=np.float32) - for i, text in enumerate(texts): - lower = text.lower() - for kw, dim in self.keyword_to_dim.items(): - if kw in lower: - out[i, dim] += 1.0 - return out - - -def _doc_id(rel_path: str) -> str: - return parser.compute_doc_id(rel_path) - - -EXAMPLE_VAULT_DIR = ( - Path(__file__).resolve().parent.parent / "examples" / "example-vault" -) - - -@pytest.fixture -def example_embedder(): - """Sparse KeywordEmbedder for the committed example vault. - - Mapping rationale: - - `circular` → dim 0: unique to feedback-loops.md. Lets the expansion test - scope Phase 2 to the seed cleanly via `final_top_k=1`. - - `satisfic` and `heuristic` → dim 1: `satisfic` is a substring of - `satisficing` and `satisficer` in bounded-rationality.md; `heuristic` - appears in no doc. A query for `heuristic` therefore has zero lexical - hits and pulls bounded-rationality on the semantic side only, isolating - the semantic-only path. - - `balancing` → dim 2: hits balancing-loops.md most heavily plus several - other systems docs. Exercises the fused signal on balancing-loops. - """ - return KeywordEmbedder( - { - "circular": 0, - "satisfic": 1, - "heuristic": 1, - "balancing": 2, - } - ) - - -@pytest.fixture -def example_db(tmp_path, monkeypatch, example_embedder): - """Copy the committed example vault into tmp_path and ingest it. - - Copying into tmp_path keeps the test self-contained and matches the - cli._ingest_directory contract (reads from a real directory tree). - """ - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: example_embedder) - - vault_copy = tmp_path / "example-vault" - shutil.copytree(EXAMPLE_VAULT_DIR, vault_copy) - - db_path = str(tmp_path / "test.sqlite") - db.init_db(db_path).close() - cli._ingest_directory(vault_copy, db_path) - return db_path - - -class TestExampleVaultRetrievalPaths: - """One test per retrieval path the example vault is designed to exercise. - - For paths where the KeywordEmbedder behavior on a small vault is fuzzy - (zero-embedding docs come back from sqlite-vec's KNN at distance 1.0), - the test asserts presence rather than a specific signal label. The - semantic-only test has a controlled mapping so it asserts the signal. - """ - - def test_unique_lexical_keyword_finds_mental_models_doc( - self, example_db, example_embedder - ): - """`antinet` is unique to mental-models-overview.md. - - FTS5 matches that one doc only; the keyword is not in the embedder - mapping, so the semantic signal is fuzzy. The test asserts the doc - lands at the top of the fused result rather than asserting a - particular signal label. - """ - conn = db.get_db(example_db) - try: - results = run_query( - conn, "antinet", example_embedder, final_top_k=5 - ) - finally: - conn.close() - - target = _doc_id("mental-models-overview.md") - doc_ids = [r.doc_id for r in results] - assert target in doc_ids - assert results[0].doc_id == target - - def test_semantic_synonym_finds_bounded_rationality( - self, example_db, example_embedder - ): - """`heuristic` has zero lexical hits across the vault. - - The embedder maps `heuristic` and `satisfic` to the same dim. - bounded-rationality.md contains `satisficing` and `satisficer` - (both match the substring `satisfic`), so it is the only doc with - non-zero activation on that dim. The signal must be `semantic` - because the lexical ranking is empty. - """ - conn = db.get_db(example_db) - try: - results = run_query( - conn, "heuristic", example_embedder, final_top_k=5 - ) - finally: - conn.close() - - by_id = {r.doc_id: r for r in results} - target = _doc_id("bounded-rationality.md") - assert target in by_id - assert by_id[target].signal == "semantic" - assert by_id[target].lexical_rank is None - assert by_id[target].semantic_rank is not None - - def test_fused_query_lands_on_balancing_loops( - self, example_db, example_embedder - ): - """`balancing` matches several docs both lexically and semantically. - - balancing-loops.md has the heaviest activation and the most lexical - occurrences, so it carries the fused signal in the result set. - """ - conn = db.get_db(example_db) - try: - results = run_query( - conn, "balancing", example_embedder, final_top_k=5 - ) - finally: - conn.close() - - by_id = {r.doc_id: r for r in results} - target = _doc_id("balancing-loops.md") - assert target in by_id - assert by_id[target].signal == "fused" - assert by_id[target].lexical_rank is not None - assert by_id[target].semantic_rank is not None - - -class TestExampleVaultGraphSurface: - """Asserts the graph topology committed in the example vault.""" - - def test_dangling_edge_preserved_on_systems_archetypes(self, example_db): - """systems-archetypes.md links to unicycle-mental-model.md (not present). - - list_neighbors must return the row with target_path=None so the CLI - can surface broken links rather than silently dropping them. - """ - conn = db.get_db(example_db) - try: - neighbors = list_neighbors( - conn, _doc_id("systems-archetypes.md") - ) - finally: - conn.close() - - dangling = [ - n - for n in neighbors - if n.target_id == _doc_id("unicycle-mental-model.md") - ] - assert len(dangling) == 1 - assert dangling[0].target_path is None - assert dangling[0].relationship_type == "cites" - - def test_seed_has_two_illustrates_edges(self, example_db): - """feedback-loops.md links to both one-hop targets with `illustrates`.""" - conn = db.get_db(example_db) - try: - neighbors = list_neighbors(conn, _doc_id("feedback-loops.md")) - finally: - conn.close() - - target_ids = {n.target_id for n in neighbors} - assert _doc_id("reinforcing-loops.md") in target_ids - assert _doc_id("balancing-loops.md") in target_ids - assert all(n.relationship_type == "illustrates" for n in neighbors) - assert all(n.target_path is not None for n in neighbors) - - -class TestExampleVaultGraphExpansion: - """Exercises the Phase 3 expansion path against the committed vault.""" - - def test_expansion_walks_seed_two_hops_skipping_dangling( - self, example_db, example_embedder - ): - """Seed query `circular` returns feedback-loops at depth 0. - - --depth 2 reaches reinforcing-loops and balancing-loops at depth 1 - (the seed's two outbound edges) and systems-archetypes at depth 2 - (reached through both one-hop docs; the BFS deduplicates). The - dangling unicycle-mental-model target is skipped by the walk. - - final_top_k=1 scopes Phase 2 to the seed only, matching the - convention from tests/test_expand.py. - """ - conn = db.get_db(example_db) - try: - results = run_query( - conn, - "circular", - example_embedder, - final_top_k=1, - expand=True, - expand_depth=2, - ) - finally: - conn.close() - - by_id = {r.doc_id: r for r in results} - - seed = _doc_id("feedback-loops.md") - assert seed in by_id - assert by_id[seed].expansion_depth == 0 - - reinforcing = _doc_id("reinforcing-loops.md") - balancing = _doc_id("balancing-loops.md") - assert by_id[reinforcing].signal == "expanded" - assert by_id[reinforcing].expansion_depth == 1 - assert by_id[balancing].signal == "expanded" - assert by_id[balancing].expansion_depth == 1 - - archetypes = _doc_id("systems-archetypes.md") - assert by_id[archetypes].signal == "expanded" - assert by_id[archetypes].expansion_depth == 2 - - assert _doc_id("unicycle-mental-model.md") not in by_id - - def test_expansion_excludes_unrelated_noise( - self, example_db, example_embedder - ): - """unrelated-noise.md has no inbound edge from any seed-reachable doc.""" - conn = db.get_db(example_db) - try: - results = run_query( - conn, - "circular", - example_embedder, - final_top_k=1, - expand=True, - expand_depth=3, - ) - finally: - conn.close() - - assert _doc_id("unrelated-noise.md") not in { - r.doc_id for r in results - } diff --git a/mindgraph/tests/test_expand.py b/mindgraph/tests/test_expand.py deleted file mode 100644 index 8237a17..0000000 --- a/mindgraph/tests/test_expand.py +++ /dev/null @@ -1,385 +0,0 @@ -"""Phase 3 — graph traversal expansion tests. - -Covers the seven ADR exit-measurement scenarios in DECISIONS.md -§ 2026-05-20 — Phase 3 graph expansion. Fixture topology: - - A.md → [[B]] (cites), [[E]] (kind), [[missing-thing]] (refers, dangling) - B.md → [[C]] (refers) - C.md → (no outbound) - D.md → (no edges, unrelated) - E.md → (no outbound) - -A contains the token `alpha`. A and B both contain the token `beta`, which -exercises the dedup scenario (Phase 2 returns both A and B; the walk from A -must not re-add B). -""" - -import json as jsonlib - -import numpy as np -import pytest -from typer.testing import CliRunner - -from mindgraph import cli, db, parser -from mindgraph.models import QueryResult -from mindgraph.query import expand_results, run_query - - -class KeywordEmbedder: - """Same deterministic embedder pattern as tests/test_query.py.""" - - def __init__(self, keyword_to_dim: dict[str, int], dims: int = 384): - self.keyword_to_dim = {k.lower(): v for k, v in keyword_to_dim.items()} - self.dims = dims - - def encode(self, texts, convert_to_numpy=True): - out = np.zeros((len(texts), self.dims), dtype=np.float32) - for i, text in enumerate(texts): - lower = text.lower() - for kw, dim in self.keyword_to_dim.items(): - if kw in lower: - out[i, dim] += 1.0 - return out - - -def _doc_id(rel_path: str) -> str: - return parser.compute_doc_id(rel_path) - - -@pytest.fixture -def expand_embedder(): - return KeywordEmbedder({"alpha": 0, "beta": 1}) - - -@pytest.fixture -def expand_db(tmp_path, monkeypatch, expand_embedder): - """Ingest the ABCDE fixture vault.""" - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: expand_embedder) - - notes = tmp_path / "vault" - notes.mkdir() - (notes / "A.md").write_text( - "Doc A talks about alpha and beta. " - "Refers to [[B]] (cites), [[E]] (kind), " - "and [[missing-thing]] (refers).\n" - ) - (notes / "B.md").write_text( - "Doc B holds beta content. Links to [[C]] (refers).\n" - ) - (notes / "C.md").write_text("Doc C has no outbound edges.\n") - (notes / "D.md").write_text("Doc D is unrelated.\n") - (notes / "E.md").write_text("Doc E has no outbound edges.\n") - - db_path = str(tmp_path / "test.sqlite") - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - return db_path - - -class TestExpand: - def test_depth_1_walks_one_hop(self, expand_db, expand_embedder): - """ADR scenario 1: A in Phase 2 (depth=0); B and E in expanded (depth=1). - - The dangling [[missing-thing]] target and the unrelated D are absent. - C is not reached at depth=1. - - final_top_k=1 scopes Phase 2 to A only. With a tiny vault, sqlite-vec's - KNN happily returns zero-embedding docs at distance 1.0, so default - top-k would otherwise pull C/D/E into the Phase 2 block as semantic-only - hits and the walk would find nothing new. - """ - conn = db.get_db(expand_db) - try: - results = run_query( - conn, - "alpha", - expand_embedder, - final_top_k=1, - expand=True, - expand_depth=1, - ) - finally: - conn.close() - - by_id = {r.doc_id: r for r in results} - assert _doc_id("A.md") in by_id - assert by_id[_doc_id("A.md")].expansion_depth == 0 - assert by_id[_doc_id("A.md")].signal in ("lexical", "semantic", "fused") - - assert _doc_id("B.md") in by_id - assert by_id[_doc_id("B.md")].signal == "expanded" - assert by_id[_doc_id("B.md")].expansion_depth == 1 - - assert _doc_id("E.md") in by_id - assert by_id[_doc_id("E.md")].signal == "expanded" - assert by_id[_doc_id("E.md")].expansion_depth == 1 - - assert _doc_id("C.md") not in by_id - assert _doc_id("D.md") not in by_id - assert _doc_id("missing-thing.md") not in by_id - - def test_depth_2_walks_two_hops(self, expand_db, expand_embedder): - """ADR scenario 2: depth=2 reaches C through B.""" - conn = db.get_db(expand_db) - try: - results = run_query( - conn, - "alpha", - expand_embedder, - final_top_k=1, - expand=True, - expand_depth=2, - ) - finally: - conn.close() - - by_id = {r.doc_id: r for r in results} - assert by_id[_doc_id("B.md")].expansion_depth == 1 - assert by_id[_doc_id("E.md")].expansion_depth == 1 - assert _doc_id("C.md") in by_id - assert by_id[_doc_id("C.md")].signal == "expanded" - assert by_id[_doc_id("C.md")].expansion_depth == 2 - - def test_no_expand_matches_phase_2(self, expand_db, expand_embedder): - """ADR scenario 3: without --expand and with depth=0 both equal Phase 2.""" - conn = db.get_db(expand_db) - try: - phase_2 = run_query(conn, "alpha", expand_embedder, final_top_k=1) - no_expand = run_query( - conn, - "alpha", - expand_embedder, - final_top_k=1, - expand=False, - expand_depth=1, - ) - zero_depth = run_query( - conn, - "alpha", - expand_embedder, - final_top_k=1, - expand=True, - expand_depth=0, - ) - finally: - conn.close() - - assert [r.doc_id for r in no_expand] == [r.doc_id for r in phase_2] - assert [r.doc_id for r in zero_depth] == [r.doc_id for r in phase_2] - assert all(r.expansion_depth == 0 for r in phase_2) - - def test_dedup_when_phase_2_already_contains_walk_target( - self, expand_db, expand_embedder - ): - """ADR scenario 6: a walk target already in Phase 2 keeps its Phase 2 signal. - - Query 'beta' matches both A and B lexically. With final_top_k=2 both - come back from Phase 2. The walk from A would normally add B at - depth=1, but B is already in seen. Assert no duplicate and B keeps the - Phase 2 signal. - """ - conn = db.get_db(expand_db) - try: - results = run_query( - conn, - "beta", - expand_embedder, - final_top_k=2, - expand=True, - expand_depth=1, - ) - finally: - conn.close() - - b_rows = [r for r in results if r.doc_id == _doc_id("B.md")] - assert len(b_rows) == 1 - assert b_rows[0].signal in ("lexical", "semantic", "fused") - assert b_rows[0].expansion_depth == 0 - - a_rows = [r for r in results if r.doc_id == _doc_id("A.md")] - assert len(a_rows) == 1 - assert a_rows[0].expansion_depth == 0 - - def test_expand_top_k_caps_appended_results( - self, expand_db, expand_embedder - ): - """ADR scenario 7: --expand-top-k 1 with multiple walked neighbors keeps one.""" - conn = db.get_db(expand_db) - try: - results = run_query( - conn, - "alpha", - expand_embedder, - final_top_k=1, - expand=True, - expand_depth=1, - expand_top_k=1, - ) - finally: - conn.close() - - expanded = [r for r in results if r.signal == "expanded"] - assert len(expanded) == 1 - assert expanded[0].expansion_depth == 1 - assert expanded[0].doc_id in {_doc_id("B.md"), _doc_id("E.md")} - - def test_dangling_edge_does_not_appear_as_expanded( - self, expand_db, expand_embedder - ): - """ADR invariant: dangling edges terminate the walk; no QueryResult row.""" - conn = db.get_db(expand_db) - try: - results = run_query( - conn, - "alpha", - expand_embedder, - final_top_k=1, - expand=True, - expand_depth=3, - ) - finally: - conn.close() - - doc_ids = {r.doc_id for r in results} - assert _doc_id("missing-thing.md") not in doc_ids - for r in results: - assert r.path != "" - assert r.path is not None - - def test_unrelated_doc_not_walked(self, expand_db, expand_embedder): - """ADR scenario 1 corollary: D has no inbound edge from any Phase 2 seed.""" - conn = db.get_db(expand_db) - try: - results = run_query( - conn, - "alpha", - expand_embedder, - final_top_k=1, - expand=True, - expand_depth=3, - ) - finally: - conn.close() - - assert _doc_id("D.md") not in {r.doc_id for r in results} - - -class TestExpandResults: - """Direct unit tests for the expand_results primitive.""" - - def test_empty_seed_list_returns_empty(self, expand_db): - conn = db.get_db(expand_db) - try: - assert expand_results(conn, [], depth=2, expand_top_k=20) == [] - finally: - conn.close() - - def test_zero_depth_returns_empty(self, expand_db, expand_embedder): - conn = db.get_db(expand_db) - try: - phase_2 = run_query(conn, "alpha", expand_embedder, final_top_k=1) - expanded = expand_results(conn, phase_2, depth=0, expand_top_k=20) - finally: - conn.close() - assert expanded == [] - - def test_expanded_results_sorted_by_depth_then_doc_id( - self, expand_db, expand_embedder - ): - """Sort key: (expansion_depth ASC, doc_id ASC, chunk_index ASC).""" - conn = db.get_db(expand_db) - try: - phase_2 = run_query(conn, "alpha", expand_embedder, final_top_k=1) - expanded = expand_results(conn, phase_2, depth=2, expand_top_k=20) - finally: - conn.close() - - keys = [(r.expansion_depth, r.doc_id, r.chunk_index) for r in expanded] - assert keys == sorted(keys) - - def test_direct_expand_preserves_seed_scope_warning( - self, expand_db, expand_embedder - ): - conn = db.get_db(expand_db) - try: - phase_2 = run_query( - conn, - "current alpha status this week", - expand_embedder, - final_top_k=1, - ) - expanded = expand_results(conn, phase_2, depth=1, expand_top_k=20) - finally: - conn.close() - - assert expanded - for row in phase_2 + expanded: - assert row.query_scope_warning is not None - assert row.query_scope_warning.intent == "live_state" - - -class TestExpandCLI: - def test_expand_flag_appends_depth_to_header( - self, expand_db, expand_embedder, monkeypatch - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: expand_embedder) - runner = CliRunner() - result = runner.invoke( - cli.app, - [ - "query", - "alpha", - "--db", - expand_db, - "--top-k", - "1", - "--expand", - "--depth", - "1", - ], - ) - assert result.exit_code == 0 - assert "signal=expanded" in result.stdout - assert "depth=1" in result.stdout - - def test_expand_flag_emits_expansion_depth_in_json( - self, expand_db, expand_embedder, monkeypatch - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: expand_embedder) - runner = CliRunner() - result = runner.invoke( - cli.app, - [ - "query", - "alpha", - "--db", - expand_db, - "--top-k", - "1", - "--expand", - "--depth", - "1", - "--json", - ], - ) - assert result.exit_code == 0 - data = jsonlib.loads(result.stdout) - assert isinstance(data, list) - assert all("expansion_depth" in row for row in data) - expanded_rows = [r for r in data if r["signal"] == "expanded"] - assert expanded_rows - for row in expanded_rows: - assert row["expansion_depth"] >= 1 - assert row["lexical_rank"] is None - assert row["semantic_rank"] is None - assert row["rrf_score"] == 0.0 - - def test_depth_above_cap_rejected(self, expand_db, expand_embedder, monkeypatch): - """Hard cap of 3 enforced at the CLI layer.""" - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: expand_embedder) - runner = CliRunner() - result = runner.invoke( - cli.app, - ["query", "alpha", "--db", expand_db, "--expand", "--depth", "4"], - ) - assert result.exit_code != 0 diff --git a/mindgraph/tests/test_governance.py b/mindgraph/tests/test_governance.py deleted file mode 100644 index 370a861..0000000 --- a/mindgraph/tests/test_governance.py +++ /dev/null @@ -1,104 +0,0 @@ -"""C-0 eligibility-manifest boundary tests for governed MindGraph context.""" - -from __future__ import annotations - -import pytest - -from mindgraph.models import QueryResult -from mindgraph.query import QueryError, apply_dual_gate_governance - - -def _result(doc_id: str, *, path: str, content_hash: str, text: str = "passage") -> QueryResult: - return QueryResult( - doc_id=doc_id, - chunk_index=0, - path=path, - title=doc_id, - signal="fused", - rrf_score=1.0, - lexical_rank=1, - semantic_rank=1, - chunk_text=text, - content_hash=content_hash, - ) - - -def _manifest(*records: dict) -> dict: - return {"eligibility_run_id": "c0-test-run", "approved_inventory": list(records)} - - -def _approved(doc_id: str, path: str, content_hash: str) -> dict: - return { - "doc_id": doc_id, - "path": path, - "sha256": f"sha256:{content_hash}", - "status": "Approved", - } - - -def test_manifest_is_required_and_cannot_be_empty() -> None: - result = _result("doc-1", path="/vault/10_knowledge/a.md", content_hash="a" * 64) - with pytest.raises(QueryError, match="requires a C-0"): - apply_dual_gate_governance([result], eligibility_manifest=None) - with pytest.raises(QueryError, match="approved_inventory is empty"): - apply_dual_gate_governance([result], eligibility_manifest=_manifest()) - - -def test_status_alone_never_admits_a_result() -> None: - result = _result("doc-1", path="/vault/10_knowledge/a.md", content_hash="a" * 64) - result = result.model_copy(update={"status": "Approved"}) - assert apply_dual_gate_governance( - [result], - eligibility_manifest=_manifest( - _approved("other", "/vault/10_knowledge/b.md", "b" * 64) - ), - ) == [] - - -@pytest.mark.parametrize( - ("record_path", "record_hash"), - [ - ("/vault/10_knowledge/other.md", "a" * 64), - ("/vault/10_knowledge/a.md", "b" * 64), - ], -) -def test_path_or_hash_mismatch_is_excluded(record_path: str, record_hash: str) -> None: - result = _result("doc-1", path="/vault/10_knowledge/a.md", content_hash="a" * 64) - governed = apply_dual_gate_governance( - [result], - eligibility_manifest=_manifest(_approved("doc-1", record_path, record_hash)), - ) - assert governed == [] - - -def test_exact_manifest_match_preserves_consumed_run_identity() -> None: - result = _result("doc-1", path="/vault/10_knowledge/a.md", content_hash="a" * 64) - governed = apply_dual_gate_governance( - [result], - eligibility_manifest=_manifest( - _approved("doc-1", "/vault/10_knowledge/a.md", "a" * 64) - ), - ) - assert len(governed) == 1 - assert governed[0].eligibility_run_id == "c0-test-run" - - -def test_seat_cap_applies_after_manifest_filtering() -> None: - results = [ - _result( - f"doc-{index}", - path=f"/vault/10_knowledge/{index}.md", - content_hash=f"{index:x}" * 64, - ) - for index in range(1, 5) - ] - inventory = [ - _approved(result.doc_id, result.path, result.content_hash or "") - for result in results - ] - governed = apply_dual_gate_governance( - results, - max_seats=3, - eligibility_manifest=_manifest(*inventory), - ) - assert [result.doc_id for result in governed] == ["doc-1", "doc-2", "doc-3"] diff --git a/mindgraph/tests/test_ingest.py b/mindgraph/tests/test_ingest.py deleted file mode 100644 index c79aa91..0000000 --- a/mindgraph/tests/test_ingest.py +++ /dev/null @@ -1,651 +0,0 @@ -import json - -import numpy as np -import pytest -from typer.testing import CliRunner - -from mindgraph import cli, db, parser -from mindgraph.query import list_neighbors - - -class FakeEmbedder: - """Stand-in for sentence-transformers — returns zero-vectors of the right shape.""" - - def encode(self, texts, convert_to_numpy=True): - return np.zeros((len(texts), 384), dtype=np.float32) - - -@pytest.fixture -def fake_embedder(monkeypatch): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: FakeEmbedder()) - - -@pytest.fixture -def sample_notes(tmp_path): - notes = tmp_path / "notes" - notes.mkdir() - - (notes / "minimal.md").write_text("Just a body, no frontmatter.\n") - - (notes / "with-timeline.md").write_text( - "---\n" - "title: Project Notes\n" - "---\n" - "Project status and goals.\n\n" - "Links to [[people/alice]] (lead).\n\n" - "---\n## Timeline\n- 2026-01-01: kicked off\n" - ) - - people_dir = notes / "people" - people_dir.mkdir() - (people_dir / "alice.md").write_text( - "---\ntitle: Alice\n---\nAlice is a person. Knows [[bob]] (peer).\n" - ) - - return notes - - -@pytest.fixture -def db_path(tmp_path): - return str(tmp_path / "test.sqlite") - - -def test_ingest_end_to_end(sample_notes, db_path, fake_embedder): - db.init_db(db_path).close() - stats = cli._ingest_directory(sample_notes, db_path) - - assert stats["total"] == 3 - assert stats["ingested"] == 3 - assert stats["skipped"] == 0 - assert stats["failed"] == 0 - - conn = db.get_db(db_path) - try: - assert conn.execute("SELECT COUNT(*) FROM documents").fetchone()[0] == 3 - hash_rows = conn.execute("SELECT content_hash FROM documents").fetchall() - assert all(r["content_hash"] for r in hash_rows) - - chunk_count = conn.execute("SELECT COUNT(*) FROM chunks").fetchone()[0] - assert chunk_count >= 3 - vec_count = conn.execute("SELECT COUNT(*) FROM vec_chunks").fetchone()[0] - assert vec_count == chunk_count - - edges = list( - conn.execute( - "SELECT source_id, target_id, relationship_type FROM edges" - ) - ) - assert len(edges) == 2 - rels = {e["relationship_type"] for e in edges} - assert rels == {"lead", "peer"} - resolved_edges = conn.execute( - """ - SELECT COUNT(*) - FROM edges e - JOIN documents d ON d.id = e.target_id - """ - ).fetchone()[0] - assert resolved_edges == 1 - - timeline_row = conn.execute( - "SELECT timeline_text FROM documents WHERE path = ?", - ("with-timeline.md",), - ).fetchone() - assert "kicked off" in timeline_row["timeline_text"] - - minimal_row = conn.execute( - "SELECT timeline_text FROM documents WHERE path = ?", - ("minimal.md",), - ).fetchone() - assert minimal_row["timeline_text"] is None - - fts_count = conn.execute( - "SELECT COUNT(*) FROM documents_fts" - ).fetchone()[0] - assert fts_count == 3 - finally: - conn.close() - - -def test_ingest_serializes_yaml_date_metadata(tmp_path, db_path, fake_embedder): - notes = tmp_path / "notes" - notes.mkdir() - (notes / "dated.md").write_text( - "---\ntitle: Dated Note\nupdated: 2026-06-19\n---\nDate metadata.\n" - ) - - db.init_db(db_path).close() - stats = cli._ingest_directory(notes, db_path) - assert stats["ingested"] == 1 - assert stats["failed"] == 0 - - conn = db.get_db(db_path) - try: - metadata = json.loads( - conn.execute( - "SELECT metadata_json FROM documents WHERE path = ?", - ("dated.md",), - ).fetchone()["metadata_json"] - ) - assert metadata["updated"] == "2026-06-19" - finally: - conn.close() - - -def test_ingest_resolves_neighbors_to_target_path(tmp_path, db_path, fake_embedder): - notes = tmp_path / "notes" - notes.mkdir() - agents = notes / "agents" - agents.mkdir() - ai_business = notes / "ai-business" - ai_business.mkdir() - - (agents / "source.md").write_text( - "Connects to [[same-domain]] and [[cross-domain]].\n" - ) - (agents / "same-domain.md").write_text("Same-domain target.\n") - (ai_business / "cross-domain.md").write_text("Cross-domain target.\n") - - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - - conn = db.get_db(db_path) - try: - neighbors = list_neighbors(conn, parser.compute_doc_id("agents/source.md")) - finally: - conn.close() - - target_paths = {edge.target_path for edge in neighbors} - assert target_paths == {"agents/same-domain.md", "ai-business/cross-domain.md"} - - -def test_reingest_unchanged_source_refreshes_resolved_edges( - tmp_path, db_path, fake_embedder -): - notes = tmp_path / "notes" - notes.mkdir() - (notes / "source.md").write_text("Connects to [[target]] (relates).\n") - - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - - conn = db.get_db(db_path) - try: - before = list_neighbors(conn, parser.compute_doc_id("source.md")) - before_rowids = sorted( - r["rowid"] for r in conn.execute("SELECT rowid FROM chunks") - ) - finally: - conn.close() - - assert before[0].target_path is None - - (notes / "target.md").write_text("Target arrives later.\n") - stats = cli._ingest_directory(notes, db_path) - assert stats["ingested"] == 1 - assert stats["skipped"] == 1 - assert stats["failed"] == 0 - - conn = db.get_db(db_path) - try: - after = list_neighbors(conn, parser.compute_doc_id("source.md")) - after_rowids = sorted( - r["rowid"] for r in conn.execute("SELECT rowid FROM chunks") - ) - finally: - conn.close() - - assert after[0].target_path == "target.md" - assert before_rowids == after_rowids[: len(before_rowids)] - - -def test_skip_path_does_not_rewrite_unchanged_edges( - sample_notes, db_path, fake_embedder, monkeypatch -): - db.init_db(db_path).close() - cli._ingest_directory(sample_notes, db_path) - - calls: list[str] = [] - real_replace = db.replace_edges - - def counting_replace(conn, source_id, edges): - calls.append(source_id) - return real_replace(conn, source_id, edges) - - monkeypatch.setattr(db, "replace_edges", counting_replace) - - stats = cli._ingest_directory(sample_notes, db_path) - assert stats["skipped"] == 3 - # Nothing changed on disk and no link resolution changed, so the skip path - # must not issue a single edge rewrite. - assert calls == [] - - -def test_skip_path_rewrites_edges_when_resolution_changes( - tmp_path, db_path, fake_embedder -): - notes = tmp_path / "notes" - notes.mkdir() - agents = notes / "agents" - agents.mkdir() - # `[[foo]]` is a bare stem. With no sibling agents/foo.md it dangles to - # compute_doc_id("foo.md"). - (agents / "source.md").write_text("Links to [[foo]] (rel).\n") - - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - - conn = db.get_db(db_path) - try: - before = list_neighbors(conn, parser.compute_doc_id("agents/source.md")) - assert before[0].target_id == parser.compute_doc_id("foo.md") - assert before[0].target_path is None - finally: - conn.close() - - # Add the sibling. source.md is byte-identical (skip path), but `[[foo]]` - # now resolves to agents/foo.md — a different target_id — so the skip path - # must rewrite the edge rather than leave the stale dangling target. - (agents / "foo.md").write_text("The foo target.\n") - stats = cli._ingest_directory(notes, db_path) - assert stats["skipped"] == 1 - assert stats["ingested"] == 1 - - conn = db.get_db(db_path) - try: - after = list_neighbors(conn, parser.compute_doc_id("agents/source.md")) - assert after[0].target_id == parser.compute_doc_id("agents/foo.md") - assert after[0].target_path == "agents/foo.md" - finally: - conn.close() - - -def test_reingest_unchanged_is_skipped(sample_notes, db_path, fake_embedder): - db.init_db(db_path).close() - cli._ingest_directory(sample_notes, db_path) - - conn = db.get_db(db_path) - before_rowids = sorted(r["rowid"] for r in conn.execute("SELECT rowid FROM chunks")) - before_total = conn.execute("SELECT COUNT(*) FROM chunks").fetchone()[0] - conn.close() - - stats = cli._ingest_directory(sample_notes, db_path) - assert stats["ingested"] == 0 - assert stats["skipped"] == 3 - assert stats["failed"] == 0 - - conn = db.get_db(db_path) - after_rowids = sorted(r["rowid"] for r in conn.execute("SELECT rowid FROM chunks")) - after_total = conn.execute("SELECT COUNT(*) FROM chunks").fetchone()[0] - conn.close() - - assert before_rowids == after_rowids - assert after_total == before_total - - -def test_reingest_modified_file_refreshes_only_that_file( - sample_notes, db_path, fake_embedder -): - db.init_db(db_path).close() - cli._ingest_directory(sample_notes, db_path) - - conn = db.get_db(db_path) - hashes_before = { - r["path"]: r["content_hash"] - for r in conn.execute("SELECT path, content_hash FROM documents") - } - conn.close() - - (sample_notes / "minimal.md").write_text( - "Completely different content now with [[new/target]] (cites).\n" - ) - - stats = cli._ingest_directory(sample_notes, db_path) - assert stats["ingested"] == 1 - assert stats["skipped"] == 2 - - conn = db.get_db(db_path) - try: - hashes_after = { - r["path"]: r["content_hash"] - for r in conn.execute("SELECT path, content_hash FROM documents") - } - assert hashes_after["minimal.md"] != hashes_before["minimal.md"] - assert hashes_after["with-timeline.md"] == hashes_before["with-timeline.md"] - assert hashes_after["people/alice.md"] == hashes_before["people/alice.md"] - - # The modified file's new edge should be present; old edges from minimal - # (there were none) should still not exist. - edges = list( - conn.execute( - "SELECT relationship_type FROM edges WHERE source_id = ?", - (cli.parser.compute_doc_id("minimal.md"),), - ) - ) - assert [e["relationship_type"] for e in edges] == ["cites"] - finally: - conn.close() - - -def test_ingest_empty_directory(tmp_path, db_path, fake_embedder): - empty = tmp_path / "empty" - empty.mkdir() - db.init_db(db_path).close() - - stats = cli._ingest_directory(empty, db_path) - assert stats == { - "total": 0, - "ingested": 0, - "skipped": 0, - "pruned": 0, - "failed": 0, - } - - -def test_ingest_prunes_database_when_last_source_is_deleted( - tmp_path, db_path, fake_embedder -): - notes = tmp_path / "notes" - notes.mkdir() - only = notes / "only.md" - only.write_text("The final indexed note.\n") - - db.init_db(db_path).close() - assert cli._ingest_directory(notes, db_path)["ingested"] == 1 - only.unlink() - - stats = cli._ingest_directory(notes, db_path) - assert stats["total"] == 0 - assert stats["pruned"] == 1 - conn = db.get_db(db_path) - try: - assert conn.execute("SELECT COUNT(*) FROM documents").fetchone()[0] == 0 - assert conn.execute("SELECT COUNT(*) FROM chunks").fetchone()[0] == 0 - assert conn.execute("SELECT COUNT(*) FROM documents_fts").fetchone()[0] == 0 - finally: - conn.close() - - -def test_ingest_prunes_deleted_file(tmp_path, db_path, fake_embedder): - notes = tmp_path / "notes" - notes.mkdir() - (notes / "keep.md").write_text("This note stays put.\n") - (notes / "remove.md").write_text("This note will be deleted later.\n") - - db.init_db(db_path).close() - stats = cli._ingest_directory(notes, db_path) - assert stats["ingested"] == 2 - assert stats["pruned"] == 0 - - (notes / "remove.md").unlink() - stats = cli._ingest_directory(notes, db_path) - assert stats["pruned"] == 1 - assert stats["skipped"] == 1 # keep.md unchanged - - removed_id = parser.compute_doc_id("remove.md") - conn = db.get_db(db_path) - try: - paths = {r["path"] for r in conn.execute("SELECT path FROM documents")} - assert paths == {"keep.md"} - # All artifacts of the pruned doc are gone, not just the documents row. - for table, col in (("chunks", "doc_id"), ("documents_fts", "id")): - count = conn.execute( - f"SELECT COUNT(*) FROM {table} WHERE {col} = ?", (removed_id,) - ).fetchone()[0] - assert count == 0 - finally: - conn.close() - - -def test_ingest_prunes_renamed_file(tmp_path, db_path, fake_embedder): - notes = tmp_path / "notes" - notes.mkdir() - (notes / "old-name.md").write_text("Stable content that just moves.\n") - - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - - (notes / "old-name.md").rename(notes / "new-name.md") - stats = cli._ingest_directory(notes, db_path) - # A rename is a new doc id at the new path plus an orphan at the old one. - assert stats["pruned"] == 1 - assert stats["ingested"] == 1 - - conn = db.get_db(db_path) - try: - paths = {r["path"] for r in conn.execute("SELECT path FROM documents")} - assert paths == {"new-name.md"} - finally: - conn.close() - - -def test_pruned_doc_leaves_inbound_edges_dangling(tmp_path, db_path, fake_embedder): - notes = tmp_path / "notes" - notes.mkdir() - (notes / "source.md").write_text("Points to [[target]] (cites).\n") - (notes / "target.md").write_text("The target of the link.\n") - - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - - conn = db.get_db(db_path) - try: - before = list_neighbors(conn, parser.compute_doc_id("source.md")) - assert before[0].target_path == "target.md" - finally: - conn.close() - - (notes / "target.md").unlink() - stats = cli._ingest_directory(notes, db_path) - assert stats["pruned"] == 1 - - conn = db.get_db(db_path) - try: - after = list_neighbors(conn, parser.compute_doc_id("source.md")) - # The inbound edge survives but now dangles (target row removed). - assert len(after) == 1 - assert after[0].target_path is None - finally: - conn.close() - - -def test_ingest_many_namespaces_duplicate_project_paths_and_prunes_union( - tmp_path, db_path, fake_embedder -): - projects = tmp_path / "30_projects" - alpha = projects / "alpha" - beta = projects / "beta" - alpha.mkdir(parents=True) - beta.mkdir(parents=True) - (alpha / "README.md").write_text( - "---\ntitle: Alpha Project\n---\nAlpha status links to [[decisions]] (records).\n" - ) - (alpha / "decisions.md").write_text("Alpha decision record.\n") - (alpha / "workbench").mkdir() - (alpha / "workbench" / "README.md").write_text("Noisy nested source.\n") - (beta / "README.md").write_text( - "---\ntitle: Beta Project\n---\nBeta status is separate.\n" - ) - - def scope(slug, root): - return cli.IngestScope( - root=root, - index_id="mainframe-projects", - trust_profile="project_status", - namespace=slug, - source_root=root, - display_prefix=f"30_projects/{slug}", - include_globs=("README.md", "decisions.md"), - exclude_globs=("workbench/*", "workbench/**/*"), - ) - - db.init_db(db_path).close() - stats = cli._ingest_scopes([scope("alpha", alpha), scope("beta", beta)], db_path) - assert stats["total"] == 3 - assert stats["ingested"] == 3 - assert stats["failed"] == 0 - - alpha_readme_id = parser.compute_scoped_doc_id( - "mainframe-projects", "alpha", "README.md" - ) - beta_readme_id = parser.compute_scoped_doc_id( - "mainframe-projects", "beta", "README.md" - ) - assert alpha_readme_id != beta_readme_id - - conn = db.get_db(db_path) - try: - rows = conn.execute( - """ - SELECT id, path, index_id, trust_profile, namespace, source_root, - source_path, display_path - FROM documents - ORDER BY path - """ - ).fetchall() - assert {row["path"] for row in rows} == { - "30_projects/alpha/README.md", - "30_projects/alpha/decisions.md", - "30_projects/beta/README.md", - } - by_id = {row["id"]: row for row in rows} - assert by_id[alpha_readme_id]["namespace"] == "alpha" - assert by_id[alpha_readme_id]["index_id"] == "mainframe-projects" - assert by_id[alpha_readme_id]["trust_profile"] == "project_status" - assert by_id[alpha_readme_id]["source_path"] == "README.md" - assert by_id[alpha_readme_id]["display_path"] == ( - "30_projects/alpha/README.md" - ) - - neighbors = list_neighbors(conn, alpha_readme_id) - assert len(neighbors) == 1 - assert neighbors[0].target_path == "30_projects/alpha/decisions.md" - finally: - conn.close() - - (beta / "README.md").unlink() - stats = cli._ingest_scopes([scope("alpha", alpha), scope("beta", beta)], db_path) - assert stats["pruned"] == 1 - - conn = db.get_db(db_path) - try: - paths = {row["path"] for row in conn.execute("SELECT path FROM documents")} - assert paths == { - "30_projects/alpha/README.md", - "30_projects/alpha/decisions.md", - } - finally: - conn.close() - - -def test_ingest_many_command_loads_manifest(tmp_path, db_path, fake_embedder): - project = tmp_path / "project" - project.mkdir() - (project / "README.md").write_text("Manifest project status.\n") - manifest = tmp_path / "manifest.json" - manifest.write_text( - json.dumps( - { - "index_id": "mainframe-projects", - "trust_profile": "project_status", - "include": ["README.md"], - "scopes": [ - { - "namespace": "manifest-project", - "root": str(project), - "source_root": str(project), - "display_prefix": "30_projects/manifest-project", - } - ], - } - ) - ) - - runner = CliRunner() - result = runner.invoke(cli.app, ["ingest-many", str(manifest), "--db", db_path]) - assert result.exit_code == 0 - - conn = db.get_db(db_path) - try: - row = conn.execute( - "SELECT path, namespace FROM documents" - ).fetchone() - assert row["path"] == "30_projects/manifest-project/README.md" - assert row["namespace"] == "manifest-project" - finally: - conn.close() - - -def test_ingest_many_allow_failures_keeps_partial_index( - tmp_path, db_path, fake_embedder -): - project = tmp_path / "project" - project.mkdir() - (project / "good.md").write_text("Good project context.\n") - (project / "bad.md").write_text("---\nupdated: bad: yaml\n---\nBad.\n") - manifest = tmp_path / "manifest.json" - manifest.write_text( - json.dumps( - { - "include": ["*.md"], - "scopes": [ - { - "namespace": "partial", - "root": str(project), - "display_prefix": "30_projects/partial", - } - ], - } - ) - ) - - runner = CliRunner() - result = runner.invoke( - cli.app, - ["ingest-many", str(manifest), "--db", db_path, "--allow-failures"], - ) - assert result.exit_code == 0 - - conn = db.get_db(db_path) - try: - paths = {row["path"] for row in conn.execute("SELECT path FROM documents")} - assert paths == {"30_projects/partial/good.md"} - finally: - conn.close() - - -def test_ingest_many_allow_failures_does_not_prune_still_present_bad_file( - tmp_path, db_path, fake_embedder -): - project = tmp_path / "project" - project.mkdir() - (project / "good.md").write_text("Good project context.\n") - (project / "fragile.md").write_text("Initially valid project context.\n") - scope = cli.IngestScope( - root=project, - index_id="mainframe-projects", - trust_profile="project_status", - namespace="partial", - source_root=project, - display_prefix="30_projects/partial", - include_globs=("*.md",), - ) - - db.init_db(db_path).close() - stats = cli._ingest_scopes([scope], db_path) - assert stats["ingested"] == 2 - - (project / "fragile.md").write_text("---\nupdated: bad: yaml\n---\nBad.\n") - stats = cli._ingest_scopes([scope], db_path) - assert stats["failed"] == 1 - assert stats["pruned"] == 0 - - conn = db.get_db(db_path) - try: - paths = {row["path"] for row in conn.execute("SELECT path FROM documents")} - assert paths == { - "30_projects/partial/good.md", - "30_projects/partial/fragile.md", - } - finally: - conn.close() diff --git a/mindgraph/tests/test_intent.py b/mindgraph/tests/test_intent.py deleted file mode 100644 index 64b2413..0000000 --- a/mindgraph/tests/test_intent.py +++ /dev/null @@ -1,782 +0,0 @@ -from __future__ import annotations - -import copy -import hashlib -import json -import sqlite3 -from pathlib import Path - -import pytest -import yaml - -from mindgraph.intent import ( - IntentGraphValidationError, - IntentStoreError, - IntentVersionConflict, - TraversalLimits, - compile_intent_corpus, - open_intent_store, - resolve_intent, - validate_intent_corpus, -) - - -FIXTURE_PATH = Path(__file__).parent / "fixtures" / "intent_graph_cases.yaml" - - -def load_catalog() -> dict: - return yaml.safe_load(FIXTURE_PATH.read_text(encoding="utf-8")) - - -@pytest.fixture -def catalog() -> dict: - return load_catalog() - - -def phase1_document(catalog: dict) -> dict: - return copy.deepcopy(catalog["documents"]["phase1"]) - - -def predecessor_document(catalog: dict) -> dict: - return copy.deepcopy(catalog["documents"]["predecessor"]) - - -def write_corpus(root: Path, documents: list[dict]) -> Path: - root.mkdir(parents=True, exist_ok=True) - for index, document in enumerate(documents, start=1): - version = document["graph"]["version"].replace(".", "-") - path = root / f"{index:02d}-{version}.yaml" - path.write_text( - yaml.safe_dump(document, sort_keys=False, allow_unicode=True), - encoding="utf-8", - ) - return root - - -def compile_document(tmp_path: Path, document: dict, name: str = "intent.sqlite"): - source = write_corpus(tmp_path / f"source-{name}", [document]) - destination = tmp_path / name - result = compile_intent_corpus(source, destination) - return destination, result - - -def sha256(path: Path) -> str: - return hashlib.sha256(path.read_bytes()).hexdigest() - - -def expect_error(source: Path, code: str): - with pytest.raises(IntentGraphValidationError) as raised: - validate_intent_corpus(source) - assert raised.value.code == code - return raised.value - - -def goal(node_id: str, label: str | None = None) -> dict: - return { - "id": node_id, - "kind": "goal", - "label": label or node_id, - "aliases": [], - "status": "active", - "source_refs": [ - "plan://plans/retrieval-remodel.md" - ], - } - - -def edge(source: str, target: str, relation: str) -> dict: - return { - "source_id": source, - "target_id": target, - "relation": relation, - "status": "active", - "source_refs": [ - "plan://plans/retrieval-remodel.md" - ], - } - - -def successor_from(document: dict, version: str, supersedes: str) -> dict: - result = copy.deepcopy(document) - result["graph"]["version"] = version - result["graph"]["supersedes"] = supersedes - result["graph"]["created_at"] = "2026-06-29T00:00:00Z" - result["graph"]["reviewed_at"] = "2026-06-29T00:00:00Z" - return result - - -def test_strict_source_schema_rejects_extra_fields(tmp_path, catalog): - document = phase1_document(catalog) - document["unexpected"] = True - error = expect_error(write_corpus(tmp_path / "source", [document]), "intent_schema_invalid") - assert "unexpected" in (error.location or "") - - -def test_valid_compile_has_expected_schema_and_metadata(tmp_path, catalog): - destination, result = compile_document(tmp_path, phase1_document(catalog)) - - assert result.graph_id == "mainframe.core" - assert result.current_version == "2026-06-29.1" - assert result.version_count == 1 - assert result.node_count == 13 - assert result.edge_count == 12 - assert result.binding_count == 4 - assert result.rule_count == 1 - assert destination.stat().st_mode & 0o777 == 0o600 - - conn = open_intent_store(destination) - try: - tables = { - row[0] - for row in conn.execute( - "SELECT name FROM sqlite_master WHERE type = 'table'" - ) - } - assert tables == { - "schema_meta", - "graph_versions", - "intent_nodes", - "intent_aliases", - "intent_edges", - "intent_bindings", - "intent_rules", - } - assert conn.execute("PRAGMA query_only").fetchone()[0] == 1 - assert conn.execute("PRAGMA foreign_key_check").fetchall() == [] - finally: - conn.close() - - -def test_canonical_hash_ignores_nonsemantic_list_and_yaml_order(tmp_path, catalog): - first = phase1_document(catalog) - second = copy.deepcopy(first) - second["nodes"].reverse() - second["edges"].reverse() - second["bindings"].reverse() - second["rules"].reverse() - for node in second["nodes"]: - node["aliases"].reverse() - node["source_refs"].reverse() - - first_result = compile_intent_corpus( - write_corpus(tmp_path / "one", [first]), tmp_path / "one.sqlite" - ) - second_result = compile_intent_corpus( - write_corpus(tmp_path / "two", [second]), tmp_path / "two.sqlite" - ) - assert first_result.corpus_hash == second_result.corpus_hash - assert first_result.version_hashes == second_result.version_hashes - assert (tmp_path / "one.sqlite").read_bytes() == (tmp_path / "two.sqlite").read_bytes() - - -def test_semantic_source_change_changes_hash(tmp_path, catalog): - first = phase1_document(catalog) - second = copy.deepcopy(first) - second["nodes"][0]["label"] = "A changed goal label" - one = compile_intent_corpus( - write_corpus(tmp_path / "one", [first]), tmp_path / "one.sqlite" - ) - two = compile_intent_corpus( - write_corpus(tmp_path / "two", [second]), tmp_path / "two.sqlite" - ) - assert one.corpus_hash != two.corpus_hash - - -def test_invalid_compile_preserves_existing_destination_bytes(tmp_path, catalog): - destination, _ = compile_document(tmp_path, phase1_document(catalog)) - before = destination.read_bytes() - invalid = phase1_document(catalog) - invalid["edges"].append(edge("goal.missing", "goal.deep-root", "requires")) - - with pytest.raises(IntentGraphValidationError): - compile_intent_corpus( - write_corpus(tmp_path / "invalid", [invalid]), destination - ) - assert destination.read_bytes() == before - - -def test_corrupt_existing_store_is_not_replaced(tmp_path, catalog): - document = phase1_document(catalog) - destination, _ = compile_document(tmp_path, document) - conn = sqlite3.connect(destination) - try: - conn.execute( - "UPDATE schema_meta SET value = 'missing-version' " - "WHERE key = 'current_version'" - ) - conn.commit() - finally: - conn.close() - before = destination.read_bytes() - - with pytest.raises(IntentStoreError) as raised: - compile_intent_corpus( - write_corpus(tmp_path / "replacement", [document]), destination - ) - assert raised.value.code == "intent_store_invalid" - assert destination.read_bytes() == before - - -def test_duplicate_alias_missing_target_relation_and_kind_errors(tmp_path, catalog): - cases: list[tuple[str, dict, str]] = [] - - duplicate = phase1_document(catalog) - duplicate["nodes"].append(copy.deepcopy(duplicate["nodes"][0])) - cases.append(("duplicate", duplicate, "intent_duplicate_id")) - - alias = phase1_document(catalog) - alias["nodes"][0]["aliases"] = ["Run the route contract checks."] - cases.append(("alias", alias, "intent_alias_ambiguous")) - - missing = phase1_document(catalog) - missing["edges"].append(edge("goal.mindgraph-remodel", "goal.missing", "requires")) - cases.append(("missing", missing, "intent_missing_target")) - - relation = phase1_document(catalog) - relation["edges"].append(edge("goal.deep-root", "goal.deep-1", "causes")) - cases.append(("relation", relation, "intent_relation_invalid")) - - kind = phase1_document(catalog) - kind["edges"].append( - edge("goal.deep-root", "capability.fixture-evaluator", "requires") - ) - cases.append(("kind", kind, "intent_kind_mismatch")) - - for name, document, code in cases: - expect_error(write_corpus(tmp_path / name, [document]), code) - - -@pytest.mark.parametrize("relation", ["requires", "decomposes_to", "next_step"]) -def test_all_controlled_cycles_are_rejected_with_witness( - tmp_path, catalog, relation -): - document = phase1_document(catalog) - document["nodes"].extend( - [goal("goal.cyclic-a", "Cyclic A"), goal("goal.cyclic-b", "Cyclic B")] - ) - document["edges"].extend( - [ - edge("goal.cyclic-a", "goal.cyclic-b", relation), - edge("goal.cyclic-b", "goal.cyclic-a", relation), - ] - ) - error = expect_error( - write_corpus(tmp_path / relation, [document]), "intent_cycle_detected" - ) - assert error.witness == ( - "goal.cyclic-a", - "goal.cyclic-b", - "goal.cyclic-a", - ) - - -def test_binding_availability_contract(tmp_path, catalog): - optional = phase1_document(catalog) - validate_intent_corpus(write_corpus(tmp_path / "optional", [optional])) - - required = phase1_document(catalog) - required["bindings"][0]["availability"] = "unavailable" - expect_error( - write_corpus(tmp_path / "required", [required]), - "intent_binding_unavailable", - ) - - -def test_version_lineage_append_only_and_effective_status(tmp_path, catalog): - predecessor = predecessor_document(catalog) - source_v1 = write_corpus(tmp_path / "v1", [predecessor]) - destination = tmp_path / "intent.sqlite" - compile_intent_corpus(source_v1, destination) - - successor = successor_from( - phase1_document(catalog), "2026-06-29.1", "2026-06-28.1" - ) - result = compile_intent_corpus( - write_corpus(tmp_path / "v2", [predecessor, successor]), destination - ) - assert result.version_count == 2 - assert result.current_version == "2026-06-29.1" - conn = open_intent_store(destination) - try: - statuses = [ - tuple(row) - for row in conn.execute( - "SELECT version, effective_status FROM graph_versions ORDER BY version" - ) - ] - assert statuses == [ - ("2026-06-28.1", "superseded"), - ("2026-06-29.1", "approved"), - ] - finally: - conn.close() - - -def test_approved_version_mutation_and_removal_are_rejected(tmp_path, catalog): - predecessor = predecessor_document(catalog) - successor = successor_from( - phase1_document(catalog), "2026-06-29.1", "2026-06-28.1" - ) - destination = tmp_path / "intent.sqlite" - compile_intent_corpus( - write_corpus(tmp_path / "complete", [predecessor, successor]), destination - ) - before = destination.read_bytes() - - changed = copy.deepcopy(predecessor) - changed["nodes"][0]["label"] = "Changed approved history" - with pytest.raises(IntentVersionConflict) as conflict: - compile_intent_corpus( - write_corpus(tmp_path / "changed", [changed, successor]), destination - ) - assert conflict.value.code == "intent_version_conflict" - assert destination.read_bytes() == before - - successor_as_root = copy.deepcopy(successor) - successor_as_root["graph"]["supersedes"] = None - with pytest.raises(IntentVersionConflict) as removed: - compile_intent_corpus( - write_corpus(tmp_path / "removed", [successor_as_root]), destination - ) - assert removed.value.code == "intent_version_missing" - assert destination.read_bytes() == before - - -def test_missing_predecessor_and_fork_are_rejected(tmp_path, catalog): - missing = successor_from( - phase1_document(catalog), "2026-06-29.1", "2026-06-27.1" - ) - expect_error( - write_corpus(tmp_path / "missing", [missing]), "intent_version_missing" - ) - - root = predecessor_document(catalog) - first = successor_from(root, "2026-06-29.1", "2026-06-28.1") - second = successor_from(root, "2026-06-29.2", "2026-06-28.1") - expect_error( - write_corpus(tmp_path / "fork", [root, first, second]), - "intent_version_fork", - ) - - -def test_open_is_read_only_and_missing_file_is_not_created(tmp_path, catalog): - missing = tmp_path / "missing.sqlite" - with pytest.raises(IntentStoreError): - open_intent_store(missing) - assert not missing.exists() - - destination, _ = compile_document(tmp_path, phase1_document(catalog)) - conn = open_intent_store(destination) - try: - with pytest.raises(sqlite3.OperationalError): - conn.execute("DELETE FROM intent_nodes") - finally: - conn.close() - - -def test_explicit_resolution_returns_exact_prerequisites_trace_and_hints( - tmp_path, catalog -): - destination, _ = compile_document(tmp_path, phase1_document(catalog)) - conn = open_intent_store(destination) - try: - resolution = resolve_intent( - conn, - "Validate the retrieval remodel before implementation.", - intent_id="goal.mindgraph-remodel", - allowed_capability_refs=frozenset( - { - "retriever://mindgraph-projects", - "capability://fixture-evaluator", - } - ), - ) - finally: - conn.close() - - assert resolution.outcome == "resolved" - assert resolution.resolution_method == "explicit" - assert resolution.intent_path == ( - "goal.mindgraph-remodel", - "goal.route-contract-evaluation", - "goal.intent-schema-validation", - ) - assert resolution.prerequisite_goal_ids == ( - "goal.route-contract-evaluation", - "goal.intent-schema-validation", - ) - assert [edge.model_dump() for edge in resolution.edge_path] == [ - { - "source_id": "goal.mindgraph-remodel", - "target_id": "goal.route-contract-evaluation", - "relation": "requires", - }, - { - "source_id": "goal.route-contract-evaluation", - "target_id": "goal.intent-schema-validation", - "relation": "requires", - }, - ] - assert set(resolution.capability_hints) == { - "retriever://mindgraph-projects", - "capability://fixture-evaluator", - } - assert resolution.constraint_ids == ("constraint.review-required",) - assert "blocked_by:constraint.review-required" in resolution.warnings - - -def test_alias_rule_and_no_match_resolution(tmp_path, catalog): - document = phase1_document(catalog) - document["edges"] = [ - item - for item in document["edges"] - if not ( - item["source_id"] == "goal.route-contract-evaluation" - and item["relation"] == "requires" - ) - ] - destination, _ = compile_document(tmp_path, document) - conn = open_intent_store(destination) - try: - alias = resolve_intent( - conn, - " RUN THE ROUTE CONTRACT CHECKS. ", - allowed_capability_refs=frozenset({"capability://fixture-evaluator"}), - ) - rule = resolve_intent( - conn, - "Review the next project status", - scope="PROJECT_STATUS", - allowed_capability_refs=frozenset({"retriever://mindgraph-projects"}), - ) - missing = resolve_intent(conn, "Help me with an unknown goal") - finally: - conn.close() - - assert alias.resolution_method == "alias" - assert alias.intent_path == ("goal.route-contract-evaluation",) - assert rule.resolution_method == "rule" - assert rule.matched_goal_ids == ("goal.project-status-review",) - assert missing.outcome == "fallback" - assert missing.refusal_reason == "intent_no_match" - assert missing.as_contract_result("missing")["behaviors"] == [ - "do not query all stores", - "require explicit scope", - ] - - -def test_tied_top_rules_for_different_goals_refuse(tmp_path, catalog): - document = phase1_document(catalog) - document["rules"].append( - { - "id": "rule.project-status-remodel", - "priority": 100, - "goal_id": "goal.mindgraph-remodel", - "match": { - "scope": "project_status", - "all_terms": ["project", "status"], - "any_terms": ["review"], - }, - "source_refs": ["decision://DECISIONS.md"], - } - ) - destination, _ = compile_document(tmp_path, document) - conn = open_intent_store(destination) - try: - resolution = resolve_intent( - conn, "Review the project status", scope="project_status" - ) - finally: - conn.close() - assert resolution.outcome == "refusal" - assert resolution.refusal_reason == "intent_rule_ambiguous" - - -def test_capability_allowlist_and_optional_unavailable_binding(tmp_path, catalog): - destination, _ = compile_document(tmp_path, phase1_document(catalog)) - conn = open_intent_store(destination) - try: - resolution = resolve_intent( - conn, - "Use the project review goal", - intent_id="goal.project-status-review", - allowed_capability_refs=frozenset({"retriever://mindgraph-projects"}), - ) - finally: - conn.close() - assert resolution.capability_hints == ("retriever://mindgraph-projects",) - assert set(resolution.rejected_capability_hints) == { - "retriever://gmail", - "capability://optional-unavailable", - } - assert "capability_not_allowed:retriever://gmail" in resolution.warnings - assert ( - "binding_unavailable:capability://optional-unavailable" in resolution.warnings - ) - assert "policy rejected untrusted capability hint" in resolution.as_contract_result( - "intent06" - )["behaviors"] - - -def test_depth_and_node_limits_are_inclusive_and_visible(tmp_path, catalog): - destination, _ = compile_document(tmp_path, phase1_document(catalog)) - conn = open_intent_store(destination) - try: - depth = resolve_intent( - conn, "deep", intent_id="goal.deep-root", limits=TraversalLimits(max_depth=2) - ) - nodes = resolve_intent( - conn, - "deep", - intent_id="goal.deep-root", - limits=TraversalLimits(max_depth=32, max_nodes=2), - ) - finally: - conn.close() - assert depth.intent_path == ("goal.deep-root", "goal.deep-1", "goal.deep-2") - assert depth.truncation_reason == "max_depth" - assert depth.as_contract_result("intent07")["reason"] == "intent_traversal_truncated" - assert nodes.intent_path == ("goal.deep-root", "goal.deep-1") - assert nodes.truncation_reason == "max_nodes" - - -def test_runtime_alias_ambiguity_fails_closed(tmp_path, catalog): - destination, _ = compile_document(tmp_path, phase1_document(catalog)) - conn = sqlite3.connect(destination) - conn.row_factory = sqlite3.Row - try: - conn.execute("ALTER TABLE intent_aliases RENAME TO valid_intent_aliases") - conn.execute( - """ - CREATE TABLE intent_aliases ( - graph_id TEXT, version TEXT, normalized_alias TEXT, - node_id TEXT, alias TEXT - ) - """ - ) - conn.executemany( - "INSERT INTO intent_aliases VALUES (?, ?, ?, ?, ?)", - [ - ( - "mainframe.core", - "2026-06-29.1", - "ambiguous", - "goal.mindgraph-remodel", - "Ambiguous", - ), - ( - "mainframe.core", - "2026-06-29.1", - "ambiguous", - "goal.project-status-review", - "Ambiguous", - ), - ], - ) - conn.commit() - resolution = resolve_intent(conn, "ambiguous") - finally: - conn.close() - assert resolution.outcome == "refusal" - assert resolution.refusal_reason == "intent_alias_ambiguous" - assert resolution.capability_hints == () - - -def test_runtime_cycle_fails_closed_without_capability_hints(tmp_path, catalog): - destination, _ = compile_document(tmp_path, phase1_document(catalog)) - conn = sqlite3.connect(destination) - conn.row_factory = sqlite3.Row - try: - conn.execute("PRAGMA foreign_keys = ON") - conn.execute( - """ - INSERT INTO intent_edges VALUES (?, ?, ?, ?, ?, ?, ?) - """, - ( - "mainframe.core", - "2026-06-29.1", - "goal.intent-schema-validation", - "goal.mindgraph-remodel", - "requires", - "active", - "[]", - ), - ) - conn.commit() - resolution = resolve_intent( - conn, - "cycle", - intent_id="goal.mindgraph-remodel", - allowed_capability_refs=frozenset( - {"retriever://mindgraph-projects", "capability://fixture-evaluator"} - ), - ) - finally: - conn.close() - assert resolution.outcome == "refusal" - assert resolution.refusal_reason == "intent_cycle_detected" - assert resolution.intent_path == ( - "goal.intent-schema-validation", - "goal.mindgraph-remodel", - "goal.route-contract-evaluation", - "goal.intent-schema-validation", - ) - assert resolution.capability_hints == () - - -def _invalid_candidate( - probe_id: str, - error: IntentGraphValidationError, - *, - method: str, - matched: list[str], - path: list[str], -) -> dict: - return { - "id": probe_id, - "outcome": "refusal", - "reason": error.code, - "resolution_method": method, - "matched_goal_ids": matched, - "prerequisite_goal_ids": [], - "intent_path": path, - "capability_hints": [], - "behaviors": [], - "fields": [ - "graph_id", - "graph_version", - "intent_path", - "matched_goal_ids", - "resolution_method", - ], - "graph_id": error.graph_id or "mainframe.core", - "graph_version": error.graph_version or "invalid", - } - - -def build_phase1_candidates(tmp_path: Path, catalog: dict | None = None) -> dict: - """Build actual candidate rows for external Phase 1 evaluator verification.""" - catalog = catalog or load_catalog() - document = phase1_document(catalog) - destination, _ = compile_document(tmp_path, document, "phase1.sqlite") - conn = open_intent_store(destination) - try: - intent01 = resolve_intent( - conn, - "Validate the retrieval remodel before implementation.", - intent_id="goal.mindgraph-remodel", - allowed_capability_refs=frozenset( - { - "retriever://mindgraph-projects", - "capability://fixture-evaluator", - } - ), - ).as_contract_result("intent01_explicit_goal_prerequisites") - intent05 = resolve_intent( - conn, "Handle a goal that is not in the approved graph." - ).as_contract_result("intent05_missing_goal_fallback") - intent06 = resolve_intent( - conn, - "Use a reviewed project goal that contains an untrusted Gmail hint.", - intent_id="goal.project-status-review", - allowed_capability_refs=frozenset({"retriever://mindgraph-projects"}), - ).as_contract_result("intent06_untrusted_route_hint_rejected") - intent07 = resolve_intent( - conn, - "Resolve a deep goal path with the V1 traversal limit.", - intent_id="goal.deep-root", - ).as_contract_result("intent07_bounded_traversal_truncation") - finally: - conn.close() - - alias_document = phase1_document(catalog) - alias_document["edges"] = [ - item - for item in alias_document["edges"] - if not ( - item["source_id"] == "goal.route-contract-evaluation" - and item["relation"] == "requires" - ) - ] - alias_destination, _ = compile_document( - tmp_path, alias_document, "alias.sqlite" - ) - alias_conn = open_intent_store(alias_destination) - try: - intent02 = resolve_intent( - alias_conn, - "Run the route contract checks.", - allowed_capability_refs=frozenset({"capability://fixture-evaluator"}), - ).as_contract_result("intent02_deterministic_alias_resolution") - finally: - alias_conn.close() - - alias_invalid = phase1_document(catalog) - alias_invalid["nodes"][0]["aliases"] = ["Run the route contract checks."] - with pytest.raises(IntentGraphValidationError) as alias_error: - validate_intent_corpus(write_corpus(tmp_path / "ambiguous", [alias_invalid])) - intent03 = _invalid_candidate( - "intent03_ambiguous_alias_refusal", - alias_error.value, - method="none", - matched=[], - path=[], - ) - - cycle_invalid = phase1_document(catalog) - cycle_invalid["nodes"].extend( - [goal("goal.cyclic-a", "Cyclic A"), goal("goal.cyclic-b", "Cyclic B")] - ) - cycle_invalid["edges"].extend( - [ - edge("goal.cyclic-a", "goal.cyclic-b", "requires"), - edge("goal.cyclic-b", "goal.cyclic-a", "requires"), - ] - ) - with pytest.raises(IntentGraphValidationError) as cycle_error: - validate_intent_corpus(write_corpus(tmp_path / "cycle", [cycle_invalid])) - intent04 = _invalid_candidate( - "intent04_cycle_refusal", - cycle_error.value, - method="explicit", - matched=["goal.cyclic-a"], - path=list(cycle_error.value.witness[:-1]), - ) - - return { - "schema_version": "1", - "results": [ - intent01, - intent02, - intent03, - intent04, - intent05, - intent06, - intent07, - ], - } - - -def test_phase1_candidate_adapter_covers_all_seven_cases(tmp_path, catalog): - candidate = build_phase1_candidates(tmp_path, catalog) - results = {item["id"]: item for item in candidate["results"]} - expected = catalog["phase1_expected"] - assert set(results) == set(expected) - - for probe_id, expectation in expected.items(): - observed = results[probe_id] - for key, value in expectation.items(): - if key == "forbidden_capability_hints": - assert not set(value) & set(observed["capability_hints"]) - elif key == "capability_hints": - assert set(observed[key]) == set(value) - else: - assert observed[key] == value - - first = json.dumps(candidate, sort_keys=True, separators=(",", ":")) - second = json.dumps( - build_phase1_candidates(tmp_path / "repeat", catalog), - sort_keys=True, - separators=(",", ":"), - ) - assert first == second diff --git a/mindgraph/tests/test_link_audit.py b/mindgraph/tests/test_link_audit.py deleted file mode 100644 index bdb474b..0000000 --- a/mindgraph/tests/test_link_audit.py +++ /dev/null @@ -1,126 +0,0 @@ -from __future__ import annotations - -import importlib.util -import sys -import tempfile -import unittest -from pathlib import Path - - -MAINFRAME_ROOT = Path(__file__).resolve().parents[2] -SCRIPT = MAINFRAME_ROOT / "scripts" / "mindgraph_link_audit.py" -SPEC = importlib.util.spec_from_file_location("mindgraph_link_audit", SCRIPT) -assert SPEC and SPEC.loader -module = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = module -SPEC.loader.exec_module(module) - - -class LinkAuditTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - (self.root / "10_knowledge" / "alpha").mkdir(parents=True) - - def tearDown(self) -> None: - self.tmp.cleanup() - - def write(self, relative: str, body: str) -> None: - path = self.root / relative - path.parent.mkdir(parents=True, exist_ok=True) - path.write_text(body, encoding="utf-8") - - def test_resolves_supported_typed_link_and_flags_raw_leaf(self) -> None: - self.write( - "10_knowledge/alpha/2026-01-01__alpha__raw__source.md", - "---\ntitle: Source\ndomain: alpha\ntype: raw\nstatus: queued\nsource: manual\ntags: []\n---\n# Source\n", - ) - self.write( - "10_knowledge/alpha/2026-01-02__alpha__note__synthesis.md", - "---\ntitle: Synthesis\ndomain: alpha\ntype: note\nstatus: stable\nsource: manual\ntags: []\n---\n# Synthesis\nEvidence: [[source]] (evidence)\n", - ) - result = module.audit(self.root, self.root / "10_knowledge", "test") - self.assertEqual(result["summary"]["link_classification_counts"], {"resolved": 1}) - self.assertEqual(result["summary"]["document_finding_counts"], {"raw evidence leaf": 1}) - self.assertTrue(result["links"][0]["relationship_supported"]) - self.assertIn("raw evidence leaf", module.markdown_report(result)) - self.assertFalse(result["mutating"]) - - def test_ambiguous_and_external_are_abstentions(self) -> None: - note = "---\ntitle: Note\ndomain: alpha\ntype: note\nstatus: stable\nsource: manual\ntags: []\n---\n# Note\nSee [[same]] and [[30_projects/demo/README.md]].\n" - self.write("10_knowledge/alpha/2026-01-01__alpha__note__one.md", note) - self.write("10_knowledge/alpha/2026-01-02__alpha__note__same.md", "---\ntitle: A\ndomain: alpha\ntype: note\n---\n") - self.write("10_knowledge/beta/2026-01-03__beta__raw__same.md", "---\ntitle: B\ndomain: beta\ntype: raw\n---\n") - result = module.audit(self.root, self.root / "10_knowledge", "test") - counts = result["summary"]["link_classification_counts"] - self.assertEqual(counts["ambiguous"], 1) - self.assertEqual(counts["external/cross-lifecycle"], 1) - - def test_unique_alias_is_resolved_without_rewrite(self) -> None: - self.write( - "10_knowledge/alpha/2026-01-01__alpha__note__one.md", - "---\ntitle: One\ndomain: alpha\ntype: note\n---\n# One\nSee [[renamed-target]].\n", - ) - self.write( - "10_knowledge/alpha/2026-01-02__alpha__note__renamed-target.md", - "---\ntitle: Renamed\ndomain: alpha\ntype: note\n---\n# Renamed\n", - ) - result = module.audit(self.root, self.root / "10_knowledge", "test") - row = next(row for row in result["links"] if row["raw_link_target"] == "renamed-target") - self.assertEqual(row["target_resolution_status"], "resolved") - self.assertIsNone(row["repair_candidate"]) - kind, candidates = module.candidate_paths( - "renamed-target", - {"10_knowledge/alpha/2026-01-02__alpha__note__renamed-target.md"}, - ) - self.assertEqual(kind, "same-canonical-slug") - self.assertEqual(len(candidates), 1) - - def test_same_channel_duplicates_and_metadata_gap_are_reported_without_rewrite(self) -> None: - self.write( - "10_knowledge/alpha/2026-01-01__alpha__note__one.md", - "---\ntitle: One\ntype: note\n---\n# One\nAgain [[target]]. Again [[target]].\n", - ) - self.write( - "10_knowledge/alpha/2026-01-02__alpha__note__target.md", - "---\ntitle: Target\ndomain: alpha\ntype: note\n---\n# Target\n", - ) - before = (self.root / "10_knowledge/alpha/2026-01-01__alpha__note__one.md").read_bytes() - result = module.audit(self.root, self.root / "10_knowledge", "test") - after = (self.root / "10_knowledge/alpha/2026-01-01__alpha__note__one.md").read_bytes() - self.assertEqual(result["summary"]["duplicate_link_count"], 2) - self.assertEqual(result["summary"]["metadata_gap_count"], 1) - self.assertEqual(before, after) - - def test_body_frontmatter_mirror_is_informational_not_duplicate(self) -> None: - self.write( - "10_knowledge/alpha/2026-01-01__alpha__note__one.md", - "---\ntitle: One\ndomain: alpha\ntype: note\nlinks: [target]\n---\n# One\nSee [[target]].\n", - ) - self.write( - "10_knowledge/alpha/2026-01-02__alpha__note__target.md", - "---\ntitle: Target\ndomain: alpha\ntype: note\n---\n# Target\n", - ) - result = module.audit(self.root, self.root / "10_knowledge", "test") - self.assertEqual(result["summary"]["mirror_pair_count"], 1) - self.assertEqual(result["summary"]["duplicate_link_count"], 0) - self.assertTrue(all(row["mirror"] for row in result["links"])) - - def test_valid_curated_dispositions_are_reviewed_not_actionable(self) -> None: - for disposition in ("reviewed-no-link", "standalone"): - self.write( - f"10_knowledge/alpha/2026-01-{len(disposition)}__alpha__note__{disposition}.md", - f"---\ntitle: {disposition}\ndomain: alpha\ntype: note\ngraph_disposition: {disposition}\n---\n# {disposition}\n", - ) - self.write( - "10_knowledge/alpha/2026-01-20__alpha__note__needs-review.md", - "---\ntitle: Needs review\ndomain: alpha\ntype: note\n---\n# Needs review\n", - ) - result = module.audit(self.root, self.root / "10_knowledge", "test") - findings = result["summary"]["document_finding_counts"] - self.assertEqual(findings["curated no-outbound reviewed"], 2) - self.assertEqual(findings["curated no-outbound"], 1) - reviewed = [row for row in result["documents"] if not row["actionable"]] - actionable = [row for row in result["documents"] if row["actionable"]] - self.assertEqual({row["graph_disposition"] for row in reviewed}, {"reviewed-no-link", "standalone"}) - self.assertEqual(len(actionable), 1) diff --git a/mindgraph/tests/test_mcp.py b/mindgraph/tests/test_mcp.py deleted file mode 100644 index 9f94a27..0000000 --- a/mindgraph/tests/test_mcp.py +++ /dev/null @@ -1,346 +0,0 @@ -import json as jsonlib -import sqlite3 -from pathlib import Path - -import pytest -import yaml -from typer.testing import CliRunner - -from mcp.shared.memory import create_connected_server_and_client_session - -from mindgraph import cli, db, parser -from mindgraph import mcp_server -from mindgraph.intent import compile_intent_corpus -from mindgraph.query import list_neighbors, run_query -from tests.test_query import KeywordEmbedder - - -@pytest.fixture -def anyio_backend(): - return "asyncio" - - -def _doc_id(rel_path: str) -> str: - return parser.compute_doc_id(rel_path) - - -def test_lock_retry_handles_wrapped_sqlite_operational_error(monkeypatch): - calls = 0 - sleeps = [] - - def flaky(): - nonlocal calls - calls += 1 - if calls < 3: - try: - raise sqlite3.OperationalError("database is locked") - except sqlite3.OperationalError as exc: - raise mcp_server.query_mod.QueryError("wrapped lock") from exc - return "ok" - - monkeypatch.setattr(mcp_server.time, "sleep", sleeps.append) - assert mcp_server._run_with_lock_retry(flaky) == "ok" - assert calls == 3 - assert sleeps == [0.5, 1.0] - - -@pytest.fixture -def mcp_embedder(): - return KeywordEmbedder( - { - "feedback": 0, - "loops": 0, - "reinforcing": 1, - "balancing": 2, - "archetype": 3, - "compilers": 4, - } - ) - - -@pytest.fixture -def mcp_db(tmp_path, monkeypatch, mcp_embedder): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: mcp_embedder) - - notes = tmp_path / "vault" - notes.mkdir() - (notes / "feedback-loops.md").write_text( - "Feedback loops explain circular system behavior. " - "See [[reinforcing-loops]] (illustrates) and " - "[[balancing-loops]] (illustrates).\n" - ) - (notes / "reinforcing-loops.md").write_text( - "Reinforcing loops amplify change over repeated cycles. " - "Related to [[systems-archetypes]] (related).\n" - ) - (notes / "balancing-loops.md").write_text( - "Balancing loops counter drift and push a system toward a setpoint. " - "Related to [[systems-archetypes]] (related).\n" - ) - (notes / "systems-archetypes.md").write_text( - "A systems archetype is a recurring feedback structure. " - "This note cites [[missing-model]] (cites).\n" - ) - (notes / "unrelated.md").write_text( - "Compilers transform source code into executable forms.\n" - ) - - db_path = str(tmp_path / "test.sqlite") - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - return db_path - - -@pytest.fixture -def mcp_runtime(mcp_db, mcp_embedder): - conn = mcp_server.open_database(mcp_db) - server = mcp_server.create_server(conn, mcp_embedder) - try: - yield server, conn - finally: - conn.close() - - -@pytest.fixture -def mcp_intent_db(tmp_path): - fixture = Path(__file__).parent / "fixtures" / "intent_graph_cases.yaml" - catalog = yaml.safe_load(fixture.read_text(encoding="utf-8")) - source = tmp_path / "intent-source" - source.mkdir() - (source / "mainframe-core.yaml").write_text( - yaml.safe_dump(catalog["documents"]["phase1"], sort_keys=False), - encoding="utf-8", - ) - destination = tmp_path / "intent.sqlite" - compile_intent_corpus(source, destination) - return destination - - -@pytest.fixture -def mcp_runtime_with_intent(mcp_db, mcp_embedder, mcp_intent_db): - conn = mcp_server.open_database(mcp_db) - server = mcp_server.create_server( - conn, - mcp_embedder, - intent_db_path=str(mcp_intent_db), - ) - try: - yield server, conn - finally: - conn.close() - - -def _tool_json(result): - assert result.isError is False - assert len(result.content) == 1 - assert result.content[0].type == "text" - return jsonlib.loads(result.content[0].text) - - -@pytest.mark.anyio -async def test_server_start_happy_path_lists_tools(mcp_runtime): - server, _conn = mcp_runtime - - async with create_connected_server_and_client_session(server) as session: - result = await session.list_tools() - - tool_names = {tool.name for tool in result.tools} - assert {"query", "graph_neighbors"} <= tool_names - - -def test_server_start_missing_db_is_clean_error(tmp_path, mcp_embedder): - missing = tmp_path / "missing.sqlite" - - with pytest.raises(mcp_server.MCPServerStartupError) as exc: - mcp_server.open_database(str(missing)) - - assert "does not exist" in str(exc.value) - assert not missing.exists() - - -@pytest.mark.anyio -async def test_query_tool_shape_matches_cli_json_surface(mcp_runtime, mcp_embedder): - server, conn = mcp_runtime - - async with create_connected_server_and_client_session(server) as session: - result = await session.call_tool( - "query", - { - "question": "feedback loops", - "final_top_k": 3, - }, - ) - - tool_rows = _tool_json(result) - expected = [ - row.model_dump() - for row in run_query( - conn, - "feedback loops", - mcp_embedder, - final_top_k=3, - ) - ] - assert tool_rows == expected - - -@pytest.mark.anyio -async def test_query_tool_envelope_includes_intent_and_routing_metadata( - mcp_runtime_with_intent, -): - server, _conn = mcp_runtime_with_intent - - async with create_connected_server_and_client_session(server) as session: - result = await session.call_tool( - "query", - { - "question": "Run the route contract checks.", - "final_top_k": 3, - "envelope": True, - }, - ) - - payload = _tool_json(result) - assert set(payload) == {"schema_version", "intent_resolution", "routing", "results"} - assert payload["schema_version"] == "1" - assert payload["intent_resolution"]["graph_id"] == "mainframe.core" - assert payload["intent_resolution"]["graph_version"] == "2026-06-29.1" - assert payload["intent_resolution"]["outcome"] == "resolved" - assert payload["intent_resolution"]["method"] == "alias" - assert payload["intent_resolution"]["matched_goals"] == [ - "goal.route-contract-evaluation" - ] - assert payload["routing"]["mode"] == "single_database" - assert payload["routing"]["selected_retrievers"] == ["mcp-bound-db"] - assert payload["routing"]["reason_codes"] == ["intent_resolved"] - assert isinstance(payload["results"], list) - - -@pytest.mark.anyio -async def test_graph_neighbors_tool_shape_matches_cli_json_surface(mcp_runtime): - server, conn = mcp_runtime - doc_id = _doc_id("feedback-loops.md") - - async with create_connected_server_and_client_session(server) as session: - result = await session.call_tool("graph_neighbors", {"doc_id": doc_id}) - - tool_rows = _tool_json(result) - expected = [row.model_dump() for row in list_neighbors(conn, doc_id)] - assert tool_rows == expected - - -@pytest.mark.anyio -async def test_query_tool_routes_expand_parameters(mcp_runtime): - server, _conn = mcp_runtime - - async with create_connected_server_and_client_session(server) as session: - result = await session.call_tool( - "query", - { - "question": "feedback loops", - "final_top_k": 1, - "expand": True, - "expand_depth": 2, - "expand_top_k": 3, - }, - ) - - rows = _tool_json(result) - by_path = {row["path"]: row for row in rows} - assert by_path["feedback-loops.md"]["expansion_depth"] == 0 - assert by_path["reinforcing-loops.md"]["signal"] == "expanded" - assert by_path["reinforcing-loops.md"]["expansion_depth"] == 1 - assert by_path["balancing-loops.md"]["signal"] == "expanded" - assert by_path["balancing-loops.md"]["expansion_depth"] == 1 - assert by_path["systems-archetypes.md"]["signal"] == "expanded" - assert by_path["systems-archetypes.md"]["expansion_depth"] == 2 - - -@pytest.mark.anyio -async def test_query_tool_rejects_negative_expand_limit(mcp_runtime): - server, _conn = mcp_runtime - async with create_connected_server_and_client_session(server) as session: - result = await session.call_tool( - "query", - { - "question": "feedback loops", - "expand": True, - "expand_top_k": -1, - }, - ) - assert result.isError is True - assert "expand_top_k" in result.content[0].text - - -@pytest.mark.anyio -async def test_query_tool_emits_scope_warning(mcp_runtime): - server, _conn = mcp_runtime - - async with create_connected_server_and_client_session(server) as session: - result = await session.call_tool( - "query", - { - "question": "current feedback status this week", - "final_top_k": 3, - }, - ) - - rows = _tool_json(result) - assert rows - assert rows[0]["query_scope_warning"]["intent"] == "live_state" - assert ( - rows[0]["query_scope_warning"]["recommended_trust_profile"] - == "time_bound_live_state" - ) - - -@pytest.mark.anyio -async def test_graph_neighbors_preserves_dangling_target(mcp_runtime): - server, _conn = mcp_runtime - - async with create_connected_server_and_client_session(server) as session: - result = await session.call_tool( - "graph_neighbors", - {"doc_id": _doc_id("systems-archetypes.md")}, - ) - - rows = _tool_json(result) - assert len(rows) == 1 - assert rows[0]["target_id"] == _doc_id("missing-model.md") - assert rows[0]["target_path"] is None - - -@pytest.mark.anyio -async def test_graph_neighbors_unknown_doc_id_is_clean_tool_error(mcp_runtime): - server, _conn = mcp_runtime - - async with create_connected_server_and_client_session(server) as session: - result = await session.call_tool( - "graph_neighbors", - {"doc_id": "not-a-real-doc"}, - ) - - assert result.isError is True - assert len(result.content) == 1 - assert "unknown doc_id" in result.content[0].text - assert "Traceback" not in result.content[0].text - - -def test_serve_mcp_help_is_registered(): - runner = CliRunner() - - result = runner.invoke(cli.app, ["serve-mcp", "--help"]) - - assert result.exit_code == 0 - assert "-db" in result.stdout - - -def test_serve_mcp_missing_db_exits_cleanly(tmp_path): - runner = CliRunner() - missing = tmp_path / "missing.sqlite" - - result = runner.invoke(cli.app, ["serve-mcp", "--db", str(missing)]) - - assert result.exit_code == 1 - assert "does not exist" in result.stderr - assert not Path(missing).exists() diff --git a/mindgraph/tests/test_neighbors.py b/mindgraph/tests/test_neighbors.py deleted file mode 100644 index 38dc52b..0000000 --- a/mindgraph/tests/test_neighbors.py +++ /dev/null @@ -1,157 +0,0 @@ -import json as jsonlib - -import numpy as np -import pytest -from typer.testing import CliRunner - -from mindgraph import cli, db, parser -from mindgraph.models import NeighborResult -from mindgraph.query import list_neighbors - - -class _ZeroEmbedder: - def encode(self, texts, convert_to_numpy=True): - return np.zeros((len(texts), 384), dtype=np.float32) - - -@pytest.fixture -def fake_embedder(monkeypatch): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: _ZeroEmbedder()) - - -@pytest.fixture -def neighbors_db(tmp_path, fake_embedder): - notes = tmp_path / "vault" - notes.mkdir() - (notes / "source.md").write_text( - "Refers to [[target_a]] (cites), [[target_b]], " - "and [[missing/x]] (refers).\n" - ) - (notes / "target_a.md").write_text("Target A.\n") - (notes / "target_b.md").write_text("Target B.\n") - (notes / "no_edges.md").write_text("Pure prose with no outbound links.\n") - - db_path = str(tmp_path / "test.sqlite") - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - return db_path - - -def _doc_id(rel: str) -> str: - return parser.compute_doc_id(rel) - - -class TestListNeighbors: - def test_returns_neighbor_result_objects(self, neighbors_db): - conn = db.get_db(neighbors_db) - try: - edges = list_neighbors(conn, _doc_id("source.md")) - finally: - conn.close() - assert all(isinstance(e, NeighborResult) for e in edges) - - def test_returns_all_outbound_edges_including_dangling(self, neighbors_db): - conn = db.get_db(neighbors_db) - try: - edges = list_neighbors(conn, _doc_id("source.md")) - finally: - conn.close() - target_ids = {e.target_id for e in edges} - assert _doc_id("target_a.md") in target_ids - assert _doc_id("target_b.md") in target_ids - assert _doc_id("missing/x.md") in target_ids - assert len(edges) == 3 - - def test_resolves_source_path_on_every_edge(self, neighbors_db): - conn = db.get_db(neighbors_db) - try: - edges = list_neighbors(conn, _doc_id("source.md")) - finally: - conn.close() - assert all(e.source_path == "source.md" for e in edges) - - def test_dangling_edge_has_null_target_path(self, neighbors_db): - conn = db.get_db(neighbors_db) - try: - edges = list_neighbors(conn, _doc_id("source.md")) - finally: - conn.close() - by_target = {e.target_id: e for e in edges} - dangling = by_target[_doc_id("missing/x.md")] - assert dangling.target_path is None - assert dangling.relationship_type == "refers" - - def test_existing_edges_have_resolved_target_path(self, neighbors_db): - conn = db.get_db(neighbors_db) - try: - edges = list_neighbors(conn, _doc_id("source.md")) - finally: - conn.close() - by_target = {e.target_id: e for e in edges} - assert by_target[_doc_id("target_a.md")].target_path == "target_a.md" - assert by_target[_doc_id("target_b.md")].target_path == "target_b.md" - - def test_relationship_type_preserved_including_null(self, neighbors_db): - conn = db.get_db(neighbors_db) - try: - edges = list_neighbors(conn, _doc_id("source.md")) - finally: - conn.close() - by_target = {e.target_id: e for e in edges} - assert by_target[_doc_id("target_a.md")].relationship_type == "cites" - assert by_target[_doc_id("target_b.md")].relationship_type is None - assert by_target[_doc_id("missing/x.md")].relationship_type == "refers" - - def test_source_with_no_edges_returns_empty(self, neighbors_db): - conn = db.get_db(neighbors_db) - try: - edges = list_neighbors(conn, _doc_id("no_edges.md")) - finally: - conn.close() - assert edges == [] - - def test_sort_order_target_id_ascending(self, neighbors_db): - conn = db.get_db(neighbors_db) - try: - edges = list_neighbors(conn, _doc_id("source.md")) - finally: - conn.close() - target_ids = [e.target_id for e in edges] - assert target_ids == sorted(target_ids) - - -class TestNeighborsCLI: - def test_neighbors_command_text_output(self, neighbors_db): - runner = CliRunner() - result = runner.invoke( - cli.app, ["neighbors", _doc_id("source.md"), "--db", neighbors_db] - ) - assert result.exit_code == 0 - assert "target_path" in result.stdout - - def test_neighbors_command_json_output(self, neighbors_db): - runner = CliRunner() - result = runner.invoke( - cli.app, - ["neighbors", _doc_id("source.md"), "--db", neighbors_db, "--json"], - ) - assert result.exit_code == 0 - data = jsonlib.loads(result.stdout) - assert len(data) == 3 - for row in data: - for field in ( - "source_id", - "target_id", - "relationship_type", - "source_path", - "target_path", - ): - assert field in row - - def test_neighbors_command_reports_no_edges_message(self, neighbors_db): - runner = CliRunner() - result = runner.invoke( - cli.app, ["neighbors", _doc_id("no_edges.md"), "--db", neighbors_db] - ) - assert result.exit_code == 0 - assert "no outbound edges" in result.stdout diff --git a/mindgraph/tests/test_parser.py b/mindgraph/tests/test_parser.py deleted file mode 100644 index 5878251..0000000 --- a/mindgraph/tests/test_parser.py +++ /dev/null @@ -1,406 +0,0 @@ -import pytest - -from mindgraph.exceptions import ParseError -from mindgraph.parser import ( - LinkResolver, - canonical_trailing_slug, - chunk_truth, - compute_doc_id, - extract_document_graph_edges, - extract_graph_edges, - extract_metadata_link_targets, - normalize_link_label, - parse_document, - parse_frontmatter, - split_page_model, -) - - -class TestParseFrontmatter: - def test_with_frontmatter(self): - text = "---\ntitle: Hello\ntags: [a, b]\n---\nBody here" - meta, body = parse_frontmatter(text) - assert meta == {"title": "Hello", "tags": ["a", "b"]} - assert body == "Body here" - - def test_without_frontmatter(self): - text = "Just body, no frontmatter" - meta, body = parse_frontmatter(text) - assert meta == {} - assert body == text - - def test_empty_frontmatter(self): - text = "---\n\n---\nBody" - meta, body = parse_frontmatter(text) - assert meta == {} - assert body == "Body" - - def test_malformed_yaml_raises(self): - text = "---\ntitle: : :\n bad indent\n---\nBody" - with pytest.raises(ParseError): - parse_frontmatter(text) - - def test_non_mapping_frontmatter_raises(self): - text = "---\n- just\n- a\n- list\n---\nBody" - with pytest.raises(ParseError): - parse_frontmatter(text) - - -class TestSplitPageModel: - def test_truth_only(self): - body = "# Heading\n\nSome content here." - truth, timeline = split_page_model(body) - assert truth == "# Heading\n\nSome content here." - assert timeline is None - - def test_truth_and_timeline(self): - body = "Truth content here.\n\n---\n## Timeline\n\n- 2026-01-01: thing\n- 2026-02-01: other" - truth, timeline = split_page_model(body) - assert truth == "Truth content here." - assert timeline == "- 2026-01-01: thing\n- 2026-02-01: other" - - def test_internal_hr_not_split(self): - """A plain `---` HR in the Truth section must not trigger the split.""" - body = "Section A\n\n---\n\nSection B\n\n---\n\nSection C" - truth, timeline = split_page_model(body) - assert truth == body.strip() - assert timeline is None - - def test_internal_hr_then_real_timeline(self): - """Internal HRs come first, then a real `## Timeline` block.""" - body = ( - "Section A\n\n---\n\nSection B\n\n" - "---\n## Timeline\n- event" - ) - truth, timeline = split_page_model(body) - assert truth.startswith("Section A") - assert "Section B" in truth - assert "## Timeline" not in truth - assert timeline == "- event" - - def test_case_insensitive_heading(self): - body = "Truth\n\n---\n## TIMELINE\n- e" - _, timeline = split_page_model(body) - assert timeline == "- e" - - def test_blank_lines_between_dash_and_heading(self): - body = "Truth\n\n---\n\n## Timeline\n- e" - truth, timeline = split_page_model(body) - assert truth == "Truth" - assert timeline == "- e" - - -class TestExtractGraphEdges: - @pytest.fixture - def link_resolver(self): - docs = [ - parse_document( - "agents/source.md", - b"---\ntitle: Source\n---\nBody", - ), - parse_document( - "agents/foo.md", - b"---\ntitle: Agent Foo\n---\nBody", - ), - parse_document( - "ai-business/cross-domain.md", - b"---\ntitle: Cross Domain Note\n---\nBody", - ), - parse_document( - "knowledge-systems/title-target.md", - b"---\ntitle: Display Title\n---\nBody", - ), - parse_document( - "finance/duplicate.md", - b"---\ntitle: Shared Title\n---\nBody", - ), - parse_document( - "agents/duplicate.md", - b"---\ntitle: Shared Title\n---\nBody", - ), - ] - return LinkResolver.from_documents(docs) - - def test_plain_link(self): - edges = extract_graph_edges("see [[people/alice]]", source_id="src") - assert len(edges) == 1 - assert edges[0].source_id == "src" - assert edges[0].target_id == compute_doc_id("people/alice.md") - assert edges[0].relationship_type is None - - def test_link_with_relationship(self): - edges = extract_graph_edges( - "see [[people/alice]] (works_at)", source_id="src" - ) - assert len(edges) == 1 - assert edges[0].relationship_type == "works_at" - - def test_multiple_links_in_paragraph(self): - text = "intro [[a]] middle [[b]] (rel) end [[c/d]]" - edges = extract_graph_edges(text, source_id="src") - assert len(edges) == 3 - assert edges[1].relationship_type == "rel" - assert edges[2].target_id == compute_doc_id("c/d.md") - - def test_nested_brackets_ignored(self): - """A `[[link[with]inner]]` shape should not produce a malformed edge.""" - edges = extract_graph_edges( - "[[normal]] and [[link[inner]brackets]]", source_id="src" - ) - assert len(edges) == 1 - assert edges[0].target_id == compute_doc_id("normal.md") - - def test_link_already_has_md_extension(self): - edges = extract_graph_edges("[[notes/foo.md]]", source_id="src") - assert edges[0].target_id == compute_doc_id("notes/foo.md") - - def test_resolves_explicit_path_target(self, link_resolver): - edges = extract_graph_edges( - "[[agents/foo]] and [[agents/foo.md]]", - source_id="src", - link_resolver=link_resolver, - source_path="agents/source.md", - ) - - assert [e.target_id for e in edges] == [ - compute_doc_id("agents/foo.md"), - compute_doc_id("agents/foo.md"), - ] - - def test_resolves_same_directory_bare_target_first(self, link_resolver): - edges = extract_graph_edges( - "[[foo]]", - source_id="src", - link_resolver=link_resolver, - source_path="agents/source.md", - ) - - assert edges[0].target_id == compute_doc_id("agents/foo.md") - - def test_resolves_unique_cross_domain_stem(self, link_resolver): - edges = extract_graph_edges( - "[[cross-domain]]", - source_id="src", - link_resolver=link_resolver, - source_path="agents/source.md", - ) - - assert edges[0].target_id == compute_doc_id("ai-business/cross-domain.md") - - def test_resolves_unique_title(self, link_resolver): - edges = extract_graph_edges( - "[[Display Title]]", - source_id="src", - link_resolver=link_resolver, - source_path="agents/source.md", - ) - - assert edges[0].target_id == compute_doc_id("knowledge-systems/title-target.md") - - def test_ambiguous_target_remains_dangling(self, link_resolver): - edges = extract_graph_edges( - "[[duplicate]] and [[Shared Title]]", - source_id="src", - link_resolver=link_resolver, - source_path="knowledge-systems/source.md", - ) - - assert [e.target_id for e in edges] == [ - compute_doc_id("duplicate.md"), - compute_doc_id("Shared Title.md"), - ] - - def test_missing_target_remains_dangling(self, link_resolver): - edges = extract_graph_edges( - "[[missing]]", - source_id="src", - link_resolver=link_resolver, - source_path="agents/source.md", - ) - - assert edges[0].target_id == compute_doc_id("missing.md") - - def test_resolves_unique_canonical_trailing_slug(self): - docs = [ - parse_document( - "regulated-systems/2026-06-13__regulated-systems__raw__gxp-pharma-source-catalog.md", - b"---\ntitle: Catalog\n---\nBody", - ), - parse_document( - "regulated-systems/source.md", - b"---\ntitle: Source\n---\nBody", - ), - ] - resolver = LinkResolver.from_documents(docs) - edges = extract_graph_edges( - "[[gxp-pharma-source-catalog]]", - source_id="src", - link_resolver=resolver, - source_path="regulated-systems/source.md", - ) - - assert edges[0].target_id == docs[0].id - - def test_ambiguous_trailing_slug_remains_dangling(self): - docs = [ - parse_document( - "regulated-systems/2026-06-13__regulated-systems__note__shared-slug.md", - b"---\ntitle: One\n---\nBody", - ), - parse_document( - "knowledge-systems/2026-06-14__knowledge-systems__note__shared-slug.md", - b"---\ntitle: Two\n---\nBody", - ), - ] - resolver = LinkResolver.from_documents(docs) - edges = extract_graph_edges( - "[[shared-slug]]", - source_id="src", - link_resolver=resolver, - source_path="regulated-systems/source.md", - ) - - assert edges[0].target_id == compute_doc_id("shared-slug.md") - - -class TestCanonicalTrailingSlug: - def test_extracts_slug_from_canonical_stem(self): - stem = "2026-06-13__regulated-systems__raw__gxp-pharma-source-catalog" - assert canonical_trailing_slug(stem) == "gxp-pharma-source-catalog" - - def test_non_canonical_stem_returns_none(self): - assert canonical_trailing_slug("legacy-short-name") is None - - -class TestMetadataLinks: - def test_extract_metadata_link_targets_dedupes(self): - metadata = { - "links": [ - "hybrid-memory-in-practice", - "[[hybrid-memory-in-practice]]", - "data-integrity-alcoa", - ] - } - assert extract_metadata_link_targets(metadata) == [ - "hybrid-memory-in-practice", - "data-integrity-alcoa", - ] - - def test_normalize_link_label_strips_wrapper_and_alias(self): - assert normalize_link_label("[[target|Display]]") == "target" - - -class TestExtractDocumentGraphEdges: - def test_indexes_frontmatter_and_body_without_duplicates(self): - docs = [ - parse_document( - "knowledge-systems/2026-06-13__knowledge-systems__note__hybrid-memory-in-practice.md", - b"---\ntitle: Hybrid\n---\nBody", - ), - parse_document( - "knowledge-systems/2026-06-04__knowledge-systems__note__knowledge-graph-rag-architecture.md", - b"---\ntitle: KG RAG\n---\nBody", - ), - parse_document( - "knowledge-systems/source.md", - ( - "---\n" - "title: Source\n" - 'links: ["hybrid-memory-in-practice", "knowledge-graph-rag-architecture"]\n' - "---\n" - "See also [[hybrid-memory-in-practice]] (extends).\n" - ).encode("utf-8"), - ), - ] - resolver = LinkResolver.from_documents(docs) - edges = extract_document_graph_edges(docs[2], link_resolver=resolver) - - assert len(edges) == 2 - assert {edge.target_id for edge in edges} == {docs[0].id, docs[1].id} - hybrid_edge = next(edge for edge in edges if edge.target_id == docs[0].id) - assert hybrid_edge.relationship_type == "extends" - - -class TestChunkTruth: - def test_empty(self): - assert chunk_truth("") == [] - assert chunk_truth(" \n\n ") == [] - - def test_single_short_paragraph(self): - assert chunk_truth("hello world") == ["hello world"] - - def test_respects_max_chars(self): - # Two paragraphs that together exceed max_chars must split. - p1 = "a" * 600 - p2 = "b" * 600 - body = f"{p1}\n\n{p2}" - chunks = chunk_truth(body, max_chars=1000) - assert len(chunks) == 2 - assert chunks[0] == p1 - assert chunks[1] == p2 - - def test_packs_paragraphs_until_limit(self): - p1 = "a" * 400 - p2 = "b" * 400 - p3 = "c" * 400 - body = f"{p1}\n\n{p2}\n\n{p3}" - chunks = chunk_truth(body, max_chars=1000) - # First two pack together (400 + 2 + 400 = 802 ≤ 1000), third spills. - assert len(chunks) == 2 - assert "a" * 400 in chunks[0] - assert "b" * 400 in chunks[0] - assert chunks[1] == p3 - - def test_oversized_paragraph_with_no_boundary_is_hard_split(self): - # A paragraph larger than max_chars with no whitespace boundary is - # hard-split so each chunk respects the bound (the old behavior let a - # single ~200KB chunk through, ~99.5% invisible to the embedder). - p = "x" * 2000 - chunks = chunk_truth(p, max_chars=1000) - assert len(chunks) >= 2 - assert all(len(c) <= 1000 for c in chunks) - assert "".join(chunks) == p # no content lost - - def test_oversized_paragraph_splits_on_sentence_boundary(self): - # Two sentences in one paragraph that together exceed max_chars but each - # fit should split at the sentence boundary, not mid-word. - s1 = "A" * 600 + "." - s2 = "B" * 600 + "." - chunks = chunk_truth(f"{s1} {s2}", max_chars=1000) - assert chunks == [s1, s2] - - def test_no_chunk_exceeds_max_chars(self): - # Invariant: whatever the input shape, every chunk respects the bound. - body = "\n\n".join(["word " * 300, "y" * 5000, "short tail."]) - chunks = chunk_truth(body, max_chars=1000) - assert chunks - assert all(len(c) <= 1000 for c in chunks) - - -class TestParseDocument: - def test_end_to_end(self): - body = ( - "---\n" - "title: My Note\n" - "domain: personal\n" - "---\n" - "Truth content with [[people/alice]] (knows).\n\n" - "---\n## Timeline\n- 2026-01-01: created" - ) - doc = parse_document("notes/my-note.md", body.encode("utf-8")) - assert doc.id == compute_doc_id("notes/my-note.md") - assert doc.title == "My Note" - assert doc.metadata["domain"] == "personal" - assert "people/alice" in doc.truth_text - assert doc.timeline_text == "- 2026-01-01: created" - assert doc.content_hash # not empty - - def test_title_falls_back_to_filename(self): - body = b"No frontmatter, just body." - doc = parse_document("notes/example.md", body) - assert doc.title == "example" - assert doc.timeline_text is None - - def test_invalid_utf8_raises(self): - with pytest.raises(ParseError): - parse_document("notes/bad.md", b"\xff\xfe\x00invalid") diff --git a/mindgraph/tests/test_query.py b/mindgraph/tests/test_query.py deleted file mode 100644 index ecf9549..0000000 --- a/mindgraph/tests/test_query.py +++ /dev/null @@ -1,872 +0,0 @@ -import json as jsonlib -import logging -import os -from pathlib import Path - -import numpy as np -import pytest -import yaml -from typer.testing import CliRunner - -from mindgraph import cli, db, parser -from mindgraph.intent import compile_intent_corpus -from mindgraph.models import QueryResult -from mindgraph.query import ( - DEFAULT_SCOPE_VOCABULARY, - RRF_K, - MAX_QUERY_TOP_K, - QueryError, - SCOPE_VOCABULARY_ENV, - ScopeVocabulary, - WEAK_FIT_DISTANCE_THRESHOLD, - _is_weak_fit, - _vocabulary_from_env, - active_scope_vocabulary, - classify_query_scope, - load_scope_vocabulary, - fetch_lexical_ranking, - fetch_semantic_ranking, - rrf_fuse, - run_query, - sanitize_fts5_query, -) - - -# --- Deterministic stub embedder --------------------------------------------- # - - -class KeywordEmbedder: - """Maps keywords (and synonyms) to specific embedding dimensions. - - Each occurrence of a known keyword in a text adds 1.0 to its assigned dim. - Synonyms that share a dim let the test model "semantic similarity without - textual overlap" (e.g. `striped horse` shares a dim with `zebra`, so a doc - that says `striped horse` is semantically close to the query `zebra` even - though FTS5 will not match the literal token). - """ - - def __init__(self, keyword_to_dim: dict[str, int], dims: int = 384): - self.keyword_to_dim = {k.lower(): v for k, v in keyword_to_dim.items()} - self.dims = dims - - def encode(self, texts, convert_to_numpy=True): - out = np.zeros((len(texts), self.dims), dtype=np.float32) - for i, text in enumerate(texts): - lower = text.lower() - for kw, dim in self.keyword_to_dim.items(): - if kw in lower: - out[i, dim] += 1.0 - return out - - -def _doc_id(rel_path: str) -> str: - return parser.compute_doc_id(rel_path) - - -# --- Unit tests: sanitize_fts5_query ----------------------------------------- # - - -class TestSanitizeFTS5Query: - def test_plain_words_or_joined(self): - assert sanitize_fts5_query("cat behavior") == "cat OR behavior" - - def test_strips_uppercase_operator_keywords(self): - assert sanitize_fts5_query("cat AND dog") == "cat OR dog" - assert sanitize_fts5_query("cat OR dog") == "cat OR dog" - assert sanitize_fts5_query("cat NOT dog") == "cat OR dog" - assert sanitize_fts5_query("cat NEAR dog") == "cat OR dog" - - def test_strips_operator_chars(self): - assert sanitize_fts5_query('"cat" *behavior*') == "cat OR behavior" - assert sanitize_fts5_query("foo:bar (baz)") == "foo OR bar OR baz" - assert sanitize_fts5_query("hat^2 -dog") == "hat OR 2 OR dog" - - def test_operator_only_input_returns_empty(self): - assert sanitize_fts5_query("AND OR NOT NEAR") == "" - assert sanitize_fts5_query('""()*:^-') == "" - - def test_empty_input(self): - assert sanitize_fts5_query("") == "" - assert sanitize_fts5_query(" ") == "" - - def test_lowercase_operators_kept_as_tokens(self): - # FTS5 only treats uppercase as operators. - assert sanitize_fts5_query("cat and dog") == "cat OR and OR dog" - - def test_apostrophes_and_question_marks_do_not_survive(self): - # The headline defect: these used to reach the MATCH parser and crash. - assert ( - sanitize_fts5_query("what's blocking the pipeline?") - == "what OR s OR blocking OR the OR pipeline" - ) - - def test_strips_arbitrary_punctuation(self): - # Allowlist: only \w+ runs survive, every other byte is a boundary. - assert ( - sanitize_fts5_query("node.js, react & vue! <tag> a/b=c") - == "node OR js OR react OR vue OR tag OR a OR b OR c" - ) - - def test_punctuation_only_input_returns_empty(self): - assert sanitize_fts5_query("?!.,;:/=%[]<>|~@#$&") == "" - - def test_unicode_tokens_preserved(self): - # \w is unicode-aware, and non-ASCII codepoints are valid FTS5 barewords. - assert sanitize_fts5_query("café señor") == "café OR señor" - - -# --- Unit tests: rrf_fuse ---------------------------------------------------- # - - -class TestRRFFuse: - def test_empty_inputs(self): - assert rrf_fuse([], []) == [] - - def test_lexical_only_doc_has_no_semantic_rank(self): - out = rrf_fuse([("a", 1)], []) - assert len(out) == 1 - doc_id, chunk_index, score, lex_rank, sem_rank, sem_distance = out[0] - assert doc_id == "a" - assert chunk_index is None - assert lex_rank == 1 - assert sem_rank is None - assert sem_distance is None - assert score == pytest.approx(1 / (RRF_K + 1)) - - def test_semantic_only_doc_has_no_lexical_rank(self): - out = rrf_fuse([], [("a", 3, 1, 0.42)]) - assert len(out) == 1 - doc_id, chunk_index, score, lex_rank, sem_rank, sem_distance = out[0] - assert doc_id == "a" - assert chunk_index == 3 - assert lex_rank is None - assert sem_rank == 1 - assert sem_distance == 0.42 - - def test_fused_doc_outranks_solo_doc(self): - # doc 'a' is in both lists; doc 'b' is only in lexical at the same rank. - out = rrf_fuse([("b", 1), ("a", 2)], [("a", 0, 1, 0.42)]) - # 'a' = 1/62 + 1/61 ; 'b' = 1/61. 'a' must be ahead. - assert out[0][0] == "a" - assert out[1][0] == "b" - - def test_tie_break_by_doc_id_ascending(self): - # Both docs tied at lex rank 1, no semantic. RRF scores equal, doc_id wins. - out = rrf_fuse([("zeta", 1)], [("alpha", 0, 1, 0.42)]) - # zeta has lex=1 only, alpha has sem=1 only. Same RRF = 1/61. - # Tie broken by doc_id ascending: alpha < zeta. - assert out[0][0] == "alpha" - assert out[1][0] == "zeta" - - def test_top_k_limit(self): - lex = [(f"d{i:02d}", i + 1) for i in range(20)] - out = rrf_fuse(lex, [], top_k=5) - assert len(out) == 5 - - def test_zero_top_k_returns_empty(self): - assert rrf_fuse([("a", 1)], [], top_k=0) == [] - - def test_score_rounding_happens_at_caller_not_fuse(self): - # rrf_fuse should return raw float; caller may round. - out = rrf_fuse([("a", 1)], []) - # 1 / 61 is not a terminating decimal. - assert out[0][2] != round(out[0][2], 2) or out[0][2] == round(out[0][2], 2) - - @pytest.mark.parametrize("top_k", [-1, MAX_QUERY_TOP_K + 1, True]) - def test_invalid_top_k_is_rejected_instead_of_becoming_unbounded(self, top_k): - with pytest.raises(QueryError, match="final_top_k"): - rrf_fuse([("a", 1)], [], top_k=top_k) - - -# --- Unit tests: weak-fit heuristic ------------------------------------------ # - - -class TestWeakFit: - def test_semantic_only_above_threshold_is_weak(self): - assert _is_weak_fit(None, 1, WEAK_FIT_DISTANCE_THRESHOLD + 0.1) is True - - def test_semantic_only_at_or_below_threshold_is_not_weak(self): - # Strict > threshold: a distance exactly at the cutoff is not weak. - assert _is_weak_fit(None, 1, WEAK_FIT_DISTANCE_THRESHOLD) is False - assert _is_weak_fit(None, 1, WEAK_FIT_DISTANCE_THRESHOLD - 0.1) is False - - def test_fused_result_never_weak(self): - # Lexical corroboration present -> not weak even at a large distance. - assert _is_weak_fit(2, 1, WEAK_FIT_DISTANCE_THRESHOLD + 5) is False - - def test_lexical_only_never_weak(self): - assert _is_weak_fit(1, None, None) is False - - def test_missing_distance_is_not_weak(self): - assert _is_weak_fit(None, 1, None) is False - - -# --- Unit tests: query-scope warning ---------------------------------------- # - - -class TestQueryScopeWarning: - def test_flags_inbox_routing_queries(self): - warning = classify_query_scope("latest inbox captures waiting for routing") - assert warning is not None - assert warning.intent == "inbox_state" - assert warning.recommended_trust_profile == "inbox_or_ingest_queue" - - def test_flags_current_live_state_queries(self): - warning = classify_query_scope("current job hunt status this week") - assert warning is not None - assert warning.intent == "live_state" - assert warning.recommended_trust_profile == "time_bound_live_state" - - def test_flags_project_status_queries(self): - warning = classify_query_scope("agent harness next gate project status") - assert warning is not None - assert warning.intent == "project_status" - assert warning.recommended_trust_profile == "project_status" - - def test_ordinary_durable_knowledge_query_has_no_warning(self): - assert classify_query_scope("agent memory remember cite forget") is None - - -class TestScopeVocabulary: - """The warning terms are configurable; the defaults must not shift.""" - - def test_default_patterns_are_unchanged(self): - expected = ( - r"\b(inbox|captures?|routing queue|ready queue|waiting for routing" - r"|00_inbox|01_ingest)\b", - r"\b(30_projects|project status|active project|project_state" - r"|next_action|next action|project readme|state\.md|handoff|next gate)\b", - r"\b(current|latest|today|this week|this month|right now|now|recent" - r"|live|as of)\b", - r"\b(job hunt|finance|calendar|workflow metrics|telemetry|live state" - r"|status|blocked?|blockers?|remaining|next)\b", - ) - actual = tuple(p.pattern for p in DEFAULT_SCOPE_VOCABULARY._patterns) - assert actual == expected - - def test_custom_terms_replace_defaults_for_that_branch(self): - vocab = ScopeVocabulary(inbox_terms=("unfiled", "to sort")) - warning = classify_query_scope("some unfiled notes", vocab) - assert warning is not None and warning.intent == "inbox_state" - assert classify_query_scope("inbox captures", vocab) is None - - def test_unspecified_branches_keep_defaults(self): - vocab = ScopeVocabulary(inbox_terms=("unfiled",)) - assert vocab.freshness_terms == DEFAULT_SCOPE_VOCABULARY.freshness_terms - assert vocab.project_terms == DEFAULT_SCOPE_VOCABULARY.project_terms - - def test_empty_term_list_disables_that_branch(self): - vocab = ScopeVocabulary(inbox_terms=()) - assert classify_query_scope("inbox captures waiting for routing", vocab) is None - - def test_live_state_needs_both_freshness_and_state_terms(self): - vocab = ScopeVocabulary(live_state_terms=("incident",)) - assert classify_query_scope("latest incident", vocab).intent == "live_state" - assert classify_query_scope("incident", vocab) is None - - def test_from_mapping_partial_override(self): - vocab = ScopeVocabulary.from_mapping({"project_terms": ["standup"]}) - assert vocab.project_terms == ("standup",) - assert vocab.inbox_terms == DEFAULT_SCOPE_VOCABULARY.inbox_terms - - def test_from_mapping_rejects_unknown_key(self): - with pytest.raises(QueryError, match="unknown scope vocabulary key"): - ScopeVocabulary.from_mapping({"bogus_terms": ["x"]}) - - def test_from_mapping_rejects_non_list_value(self): - with pytest.raises(QueryError, match="must be a list of strings"): - ScopeVocabulary.from_mapping({"inbox_terms": "notalist"}) - - def test_load_from_json_file(self, tmp_path): - path = tmp_path / "vocab.json" - path.write_text(jsonlib.dumps({"inbox_terms": ["unfiled"]})) - vocab = load_scope_vocabulary(path) - assert vocab.inbox_terms == ("unfiled",) - - def test_load_missing_file_raises(self, tmp_path): - with pytest.raises(QueryError, match="not found"): - load_scope_vocabulary(tmp_path / "absent.json") - - def test_load_invalid_json_raises(self, tmp_path): - path = tmp_path / "vocab.json" - path.write_text("{not json") - with pytest.raises(QueryError, match="not valid JSON"): - load_scope_vocabulary(path) - - def test_load_non_object_raises(self, tmp_path): - path = tmp_path / "vocab.json" - path.write_text("[1, 2]") - with pytest.raises(QueryError, match="must contain a JSON object"): - load_scope_vocabulary(path) - - def test_env_var_overrides_active_vocabulary(self, tmp_path, monkeypatch): - path = tmp_path / "vocab.json" - path.write_text(jsonlib.dumps({"inbox_terms": ["unfiled"]})) - monkeypatch.setenv(SCOPE_VOCABULARY_ENV, str(path)) - _vocabulary_from_env.cache_clear() - try: - assert active_scope_vocabulary().inbox_terms == ("unfiled",) - assert classify_query_scope("some unfiled notes").intent == "inbox_state" - finally: - _vocabulary_from_env.cache_clear() - - def test_no_env_var_uses_defaults(self, monkeypatch): - monkeypatch.delenv(SCOPE_VOCABULARY_ENV, raising=False) - _vocabulary_from_env.cache_clear() - try: - assert active_scope_vocabulary() == DEFAULT_SCOPE_VOCABULARY - finally: - _vocabulary_from_env.cache_clear() - - -# --- Integration tests: against a small ingested vault ----------------------- # - - -@pytest.fixture -def keyword_embedder(): - # Dim 0: zebra-family. Dim 1: compiler-family. Dim 2: elephant-family. - return KeywordEmbedder( - { - "zebra": 0, - "zebras": 0, - "striped horse": 0, - "savannah": 0, - "compiler": 1, - "programming": 1, - "elephant": 2, - "trunk": 2, - } - ) - - -@pytest.fixture -def vault_db(tmp_path, monkeypatch, keyword_embedder): - """Ingest a small vault designed to exercise each retrieval signal.""" - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: keyword_embedder) - - notes = tmp_path / "vault" - notes.mkdir() - # Lexical-only winner: contains the literal token `zebra`. - (notes / "lex.md").write_text( - "The zebra is the subject of this note. Single mention.\n" - ) - # Semantic-only winner: no literal `zebra` but a synonym `striped horse`. - (notes / "sem.md").write_text( - "About striped horses that live on the savannah grass.\n" - ) - # Fused: contains the literal `zebra` plus extra synonyms. - (notes / "fused.md").write_text( - "The zebra roams the savannah. Striped horse with hooves.\n" - ) - # Unrelated baseline. - (notes / "unrelated.md").write_text( - "About compilers and programming languages.\n" - ) - - db_path = str(tmp_path / "test.sqlite") - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - return db_path - - -class TestFetchLexicalRanking: - def test_returns_only_docs_with_query_token(self, vault_db): - conn = db.get_db(vault_db) - try: - ranking = fetch_lexical_ranking(conn, "zebra", top_k=10) - finally: - conn.close() - doc_ids = [doc_id for doc_id, _ in ranking] - assert _doc_id("lex.md") in doc_ids - assert _doc_id("fused.md") in doc_ids - assert _doc_id("sem.md") not in doc_ids - assert _doc_id("unrelated.md") not in doc_ids - - def test_empty_query_returns_empty_list(self, vault_db): - conn = db.get_db(vault_db) - try: - assert fetch_lexical_ranking(conn, "") == [] - assert fetch_lexical_ranking(conn, "AND OR NOT") == [] - finally: - conn.close() - - def test_punctuation_query_does_not_crash_and_still_matches(self, vault_db): - # Regression for the FTS5 syntax-error defect: a natural-language query - # with an apostrophe and question mark must run, and the surviving - # `zebra` token must still retrieve the zebra docs. - conn = db.get_db(vault_db) - try: - ranking = fetch_lexical_ranking(conn, "what's the zebra?", top_k=10) - finally: - conn.close() - doc_ids = [doc_id for doc_id, _ in ranking] - assert _doc_id("lex.md") in doc_ids - assert _doc_id("fused.md") in doc_ids - - def test_top_k_limits_result_count(self, vault_db): - conn = db.get_db(vault_db) - try: - ranking = fetch_lexical_ranking(conn, "zebra striped", top_k=1) - finally: - conn.close() - assert len(ranking) <= 1 - - def test_ranks_are_one_based_and_strictly_increasing(self, vault_db): - conn = db.get_db(vault_db) - try: - ranking = fetch_lexical_ranking(conn, "zebra", top_k=10) - finally: - conn.close() - ranks = [r for _, r in ranking] - assert ranks - assert ranks[0] == 1 - assert all(b > a for a, b in zip(ranks, ranks[1:])) - - -class TestFetchSemanticRanking: - def test_returns_docs_via_synonym_match(self, vault_db, keyword_embedder): - emb = keyword_embedder.encode(["zebra"])[0].tolist() - conn = db.get_db(vault_db) - try: - ranking = fetch_semantic_ranking(conn, emb, top_k=20) - finally: - conn.close() - doc_ids = [doc_id for doc_id, _, _, _ in ranking] - # sem.md has no `zebra` but is reachable via the synonym embedding. - assert _doc_id("sem.md") in doc_ids - assert _doc_id("lex.md") in doc_ids - assert _doc_id("fused.md") in doc_ids - - def test_at_most_one_chunk_per_document(self, vault_db, keyword_embedder): - emb = keyword_embedder.encode(["zebra"])[0].tolist() - conn = db.get_db(vault_db) - try: - ranking = fetch_semantic_ranking(conn, emb, top_k=50) - finally: - conn.close() - doc_ids = [doc_id for doc_id, _, _, _ in ranking] - assert len(doc_ids) == len(set(doc_ids)) - - def test_zero_top_k_returns_empty(self, vault_db, keyword_embedder): - emb = keyword_embedder.encode(["zebra"])[0].tolist() - conn = db.get_db(vault_db) - try: - assert fetch_semantic_ranking(conn, emb, top_k=0) == [] - finally: - conn.close() - - -class TestRunQuery: - def test_returns_query_result_objects(self, vault_db, keyword_embedder): - conn = db.get_db(vault_db) - try: - results = run_query(conn, "zebra", keyword_embedder) - finally: - conn.close() - assert results - assert all(isinstance(r, QueryResult) for r in results) - assert all(r.signal in ("lexical", "semantic", "fused") for r in results) - - def test_signal_attribution(self, vault_db, keyword_embedder): - conn = db.get_db(vault_db) - try: - results = run_query(conn, "zebra", keyword_embedder, final_top_k=10) - finally: - conn.close() - by_path = {r.path: r for r in results} - # sem.md has no `zebra` but a synonym semantic match. - assert by_path["sem.md"].signal == "semantic" - assert by_path["sem.md"].lexical_rank is None - assert by_path["sem.md"].semantic_rank is not None - # lex.md has both the lexical match and the synonym semantic match. - assert by_path["lex.md"].signal == "fused" - # fused.md is in both rankings. - assert by_path["fused.md"].signal == "fused" - - def test_provenance_fields_populated(self, vault_db, keyword_embedder): - conn = db.get_db(vault_db) - try: - results = run_query(conn, "zebra", keyword_embedder, final_top_k=3) - finally: - conn.close() - assert results - for r in results: - assert r.doc_id - assert r.path - assert r.title - assert r.chunk_text - assert r.chunk_index >= 0 - assert r.rrf_score > 0 - - def test_scoped_ingest_provenance_fields_populated( - self, tmp_path, monkeypatch, keyword_embedder - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: keyword_embedder) - notes = tmp_path / "knowledge" - notes.mkdir() - (notes / "scoped.md").write_text("The zebra appears in durable knowledge.\n") - db_path = str(tmp_path / "scoped.sqlite") - cli._ingest_directory( - notes, - db_path, - index_id="mainframe-knowledge", - trust_profile="durable_knowledge", - namespace="knowledge", - source_root=notes, - display_prefix="10_knowledge", - ) - - conn = db.get_db(db_path) - try: - results = run_query(conn, "zebra", keyword_embedder, final_top_k=1) - finally: - conn.close() - - assert len(results) == 1 - row = results[0] - assert row.path == "10_knowledge/scoped.md" - assert row.display_path == "10_knowledge/scoped.md" - assert row.index_id == "mainframe-knowledge" - assert row.trust_profile == "durable_knowledge" - assert row.namespace == "knowledge" - assert row.source_root == str(notes) - assert row.source_path == "scoped.md" - - def test_final_top_k_limits_output(self, vault_db, keyword_embedder): - conn = db.get_db(vault_db) - try: - results = run_query(conn, "zebra", keyword_embedder, final_top_k=2) - finally: - conn.close() - assert len(results) <= 2 - - def test_operator_only_query_yields_only_semantic_signal( - self, vault_db, keyword_embedder - ): - conn = db.get_db(vault_db) - try: - results = run_query(conn, "AND OR NOT", keyword_embedder) - finally: - conn.close() - # Lexical ranking is empty (sanitized to empty MATCH), so every result - # must carry the `semantic` signal. - for r in results: - assert r.signal == "semantic" - assert r.lexical_rank is None - - def test_degrades_to_semantic_when_lexical_ranking_errors( - self, vault_db, keyword_embedder, monkeypatch - ): - # If FTS5 raises (e.g. a malformed index), run_query must not fail the - # whole call — it degrades to semantic-only with the lexical side empty. - from mindgraph import query as query_mod - - def boom(*args, **kwargs): - raise query_mod.QueryError("FTS5 query failed: simulated") - - monkeypatch.setattr(query_mod, "fetch_lexical_ranking", boom) - conn = db.get_db(vault_db) - try: - results = run_query(conn, "zebra", keyword_embedder) - finally: - conn.close() - assert results # semantic side still produced candidates - for r in results: - assert r.lexical_rank is None - assert r.signal == "semantic" - - def test_rrf_score_rounded_to_six_decimals(self, vault_db, keyword_embedder): - conn = db.get_db(vault_db) - try: - results = run_query(conn, "zebra", keyword_embedder, final_top_k=1) - finally: - conn.close() - assert results - assert results[0].rrf_score == round(results[0].rrf_score, 6) - - def test_results_carry_distance_and_weak_fit_fields( - self, vault_db, keyword_embedder - ): - conn = db.get_db(vault_db) - try: - results = run_query(conn, "zebra", keyword_embedder, final_top_k=10) - finally: - conn.close() - by_path = {r.path: r for r in results} - # A fused result has lexical corroboration, so it is never weak_fit, - # whatever its semantic distance. - assert by_path["fused.md"].signal == "fused" - assert by_path["fused.md"].weak_fit is False - # Raw distance is exposed: a float when there is a semantic rank, else None. - for r in results: - assert isinstance(r.weak_fit, bool) - if r.semantic_rank is not None: - assert isinstance(r.semantic_distance, float) - else: - assert r.semantic_distance is None - - def test_scope_warning_survives_fused_result( - self, vault_db, keyword_embedder - ): - conn = db.get_db(vault_db) - try: - results = run_query( - conn, - "current zebra status this week", - keyword_embedder, - final_top_k=10, - ) - finally: - conn.close() - - by_path = {r.path: r for r in results} - assert by_path["fused.md"].signal == "fused" - assert by_path["fused.md"].weak_fit is False - assert by_path["fused.md"].query_scope_warning is not None - assert by_path["fused.md"].query_scope_warning.intent == "live_state" - for row in results: - assert row.query_scope_warning is not None - assert row.query_scope_warning.intent == "live_state" - - -class TestQueryResultMetadata: - def test_emits_frontmatter_type_domain_status( - self, tmp_path, monkeypatch, keyword_embedder - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: keyword_embedder) - notes = tmp_path / "vault" - notes.mkdir() - (notes / "typed.md").write_text( - "---\ntitle: Typed Note\ntype: note\ndomain: agents\nstatus: stable\n---\n" - "The zebra roams the savannah.\n" - ) - (notes / "bare.md").write_text( - "The zebra appears here with no frontmatter at all.\n" - ) - db_path = str(tmp_path / "meta.sqlite") - db.init_db(db_path).close() - cli._ingest_directory(notes, db_path) - - conn = db.get_db(db_path) - try: - results = run_query(conn, "zebra", keyword_embedder, final_top_k=10) - finally: - conn.close() - - by_path = {r.path: r for r in results} - assert by_path["typed.md"].doc_type == "note" - assert by_path["typed.md"].domain == "agents" - assert by_path["typed.md"].status == "stable" - # A source with no frontmatter carries nulls, not empty strings. - assert by_path["bare.md"].doc_type is None - assert by_path["bare.md"].domain is None - assert by_path["bare.md"].status is None - - -class TestQueryCLI: - def test_query_command_runs_text_output( - self, vault_db, keyword_embedder, monkeypatch - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: keyword_embedder) - runner = CliRunner() - result = runner.invoke( - cli.app, ["query", "zebra", "--db", vault_db, "--top-k", "3"] - ) - assert result.exit_code == 0 - assert "signal=" in result.stdout - assert "rrf_score=" in result.stdout - - def test_query_command_reads_wal_db_from_nonwritable_directory( - self, vault_db, keyword_embedder, monkeypatch - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: keyword_embedder) - db_path = Path(vault_db) - parent = db_path.parent - before = db_path.read_bytes() - os.chmod(db_path, 0o444) - os.chmod(parent, 0o555) - try: - result = CliRunner().invoke( - cli.app, ["query", "zebra", "--db", vault_db, "--json"] - ) - assert result.exit_code == 0, result.output - assert jsonlib.loads(result.stdout) - assert db_path.read_bytes() == before - finally: - os.chmod(parent, 0o755) - os.chmod(db_path, 0o644) - - def test_query_command_reports_embedder_load_failure_cleanly( - self, tmp_path, monkeypatch, caplog - ): - db_path = tmp_path / "query.sqlite" - db.init_db(str(db_path)).close() - - def fail_load(_spec): - raise RuntimeError("offline fixture") - - monkeypatch.setattr(cli.embedders, "load_sentence_embedder", fail_load) - with caplog.at_level(logging.ERROR, logger="mindgraph"): - result = CliRunner().invoke( - cli.app, - ["query", "hello", "--db", str(db_path), "--no-intent"], - ) - assert result.exit_code == 1 - assert not isinstance(result.exception, RuntimeError) - assert "Failed to load embedding model" in caplog.text - - def test_query_command_json_output( - self, vault_db, keyword_embedder, monkeypatch - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: keyword_embedder) - runner = CliRunner() - result = runner.invoke( - cli.app, - ["query", "zebra", "--db", vault_db, "--json", "--top-k", "3"], - ) - assert result.exit_code == 0 - data = jsonlib.loads(result.stdout) - assert isinstance(data, list) - assert len(data) <= 3 - if data: - row = data[0] - for field in ( - "doc_id", - "chunk_index", - "path", - "title", - "doc_type", - "domain", - "status", - "index_id", - "trust_profile", - "namespace", - "source_root", - "source_path", - "display_path", - "signal", - "rrf_score", - "lexical_rank", - "semantic_rank", - "semantic_distance", - "weak_fit", - "query_scope_warning", - "chunk_text", - ): - assert field in row - - def test_query_command_json_envelope_output( - self, tmp_path, vault_db, keyword_embedder, monkeypatch - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: keyword_embedder) - fixture = Path(__file__).parent / "fixtures" / "intent_graph_cases.yaml" - catalog = yaml.safe_load(fixture.read_text(encoding="utf-8")) - source = tmp_path / "intent-source" - source.mkdir() - (source / "mainframe-core.yaml").write_text( - yaml.safe_dump(catalog["documents"]["phase1"], sort_keys=False), - encoding="utf-8", - ) - intent_db = tmp_path / "intent.sqlite" - compile_intent_corpus(source, intent_db) - - runner = CliRunner() - result = runner.invoke( - cli.app, - [ - "query", - "Run the route contract checks.", - "--db", - vault_db, - "--json", - "--envelope", - "--intent-db", - str(intent_db), - "--top-k", - "3", - ], - ) - - assert result.exit_code == 0 - data = jsonlib.loads(result.stdout) - assert set(data) == { - "schema_version", - "intent_resolution", - "routing", - "results", - } - assert data["schema_version"] == "1" - assert data["intent_resolution"]["graph_id"] == "mainframe.core" - assert data["intent_resolution"]["graph_version"] == "2026-06-29.1" - assert data["intent_resolution"]["outcome"] == "resolved" - assert data["intent_resolution"]["method"] == "alias" - assert data["intent_resolution"]["matched_goals"] == [ - "goal.route-contract-evaluation" - ] - assert data["routing"]["mode"] == "single_database" - assert data["routing"]["selected_retrievers"] == ["cli-bound-db"] - assert data["routing"]["reason_codes"] == ["intent_resolved"] - assert isinstance(data["results"], list) - - def test_query_command_text_output_prints_scope_warning( - self, vault_db, keyword_embedder, monkeypatch - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: keyword_embedder) - runner = CliRunner() - result = runner.invoke( - cli.app, - [ - "query", - "current zebra status this week", - "--db", - vault_db, - "--top-k", - "3", - ], - ) - assert result.exit_code == 0 - assert "scope warning:" in result.stdout - assert "recommended_trust_profile=time_bound_live_state" in result.stdout - - def test_query_command_json_output_emits_scope_warning( - self, vault_db, keyword_embedder, monkeypatch - ): - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: keyword_embedder) - runner = CliRunner() - result = runner.invoke( - cli.app, - [ - "query", - "latest inbox captures waiting for routing", - "--db", - vault_db, - "--json", - "--top-k", - "3", - ], - ) - assert result.exit_code == 0 - data = jsonlib.loads(result.stdout) - assert data - assert data[0]["query_scope_warning"]["intent"] == "inbox_state" - - -@pytest.mark.parametrize( - ("kwargs", "field"), - [ - ({"lexical_top_k": -1}, "lexical_top_k"), - ({"semantic_top_k": MAX_QUERY_TOP_K + 1}, "semantic_top_k"), - ({"expand": True, "expand_depth": 4}, "expand_depth"), - ({"expand": True, "expand_top_k": -1}, "expand_top_k"), - ({"associate": True, "associate_seed_k": -1}, "associate_seed_k"), - ], -) -def test_run_query_rejects_invalid_resource_limits_before_work( - vault_db, keyword_embedder, kwargs, field -): - conn = db.get_db(vault_db) - try: - with pytest.raises(QueryError, match=field): - run_query(conn, "zebra", keyword_embedder, **kwargs) - finally: - conn.close() diff --git a/mindgraph/tests/test_routing.py b/mindgraph/tests/test_routing.py deleted file mode 100644 index ce09c8b..0000000 --- a/mindgraph/tests/test_routing.py +++ /dev/null @@ -1,684 +0,0 @@ -from __future__ import annotations - -import hashlib -import json -import sqlite3 -from dataclasses import FrozenInstanceError -from pathlib import Path - -import numpy as np -import pytest -import yaml -from pydantic import ValidationError - -from mindgraph import cli, db -from mindgraph.intent import ( - IntentResolution, - compile_intent_corpus, - open_intent_store, -) -from mindgraph.models import QueryResult -from mindgraph import query as query_mod -from mindgraph.routing import ( - CapabilityRegistry, - MindGraphQueryRetriever, - RefusalRule, - RetrieverDescriptor, - RetrieverRegistration, - RouteRequest, - RouterPolicy, - RoutingError, - decide_route, - open_query_store_read_only, - orchestrate, -) - - -ROUTING_FIXTURE = Path(__file__).parent / "fixtures" / "phase3_routing_cases.yaml" -INTENT_FIXTURE = Path(__file__).parent / "fixtures" / "intent_graph_cases.yaml" - - -def load_routing_fixture() -> dict: - return yaml.safe_load(ROUTING_FIXTURE.read_text(encoding="utf-8")) - - -def make_result( - retriever_id: str, - *, - trust_profile: str, - score: float, - path: str | None = None, -) -> QueryResult: - display_path = path or f"{retriever_id}/result.md" - return QueryResult( - doc_id=f"{retriever_id}-doc", - chunk_index=0, - path=display_path, - title=f"{retriever_id} result", - index_id=retriever_id, - trust_profile=trust_profile, - namespace=retriever_id, - source_root=f"/fixture/{retriever_id}", - source_path="result.md", - display_path=display_path, - signal="fused", - rrf_score=score, - lexical_rank=1, - semantic_rank=1, - semantic_distance=0.25, - chunk_text=f"fixture row from {retriever_id}", - ) - - -class StubRetriever: - def __init__(self, results: list[QueryResult] | None = None) -> None: - self.results = results or [] - self.calls: list[tuple[str, int]] = [] - - def retrieve(self, query_text: str, *, final_top_k: int) -> list[QueryResult]: - self.calls.append((query_text, final_top_k)) - return list(self.results[:final_top_k]) - - -class FailingRetriever: - def __init__(self, message: str) -> None: - self.message = message - self.calls = 0 - - def retrieve(self, query_text: str, *, final_top_k: int) -> list[QueryResult]: - self.calls += 1 - raise RuntimeError(self.message) - - -def descriptor( - retriever_id: str, - capability_ref: str, - trust_profile: str, - *, - available: bool = True, -) -> RetrieverDescriptor: - return RetrieverDescriptor( - retriever_id=retriever_id, - capability_ref=capability_ref, - trust_profile=trust_profile, - source_surface=f"fixture://{retriever_id}", - available=available, - unavailable_reason=None if available else "fixture_unavailable", - ) - - -def make_registry( - *, - durable: object | None = None, - projects: object | None = None, - extras: tuple[RetrieverRegistration, ...] = (), -) -> CapabilityRegistry: - durable_executor = durable if durable is not None else StubRetriever() - project_executor = projects if projects is not None else StubRetriever() - return CapabilityRegistry( - ( - RetrieverRegistration( - descriptor( - "mindgraph-projects", - "retriever://mindgraph-projects", - "project_status", - ), - project_executor, - ), - RetrieverRegistration( - descriptor( - "mindgraph-durable", - "retriever://mindgraph-durable", - "durable_knowledge", - ), - durable_executor, - ), - *extras, - ) - ) - - -@pytest.fixture -def policy() -> RouterPolicy: - return RouterPolicy.model_validate(load_routing_fixture()["policy"]) - - -def unavailable_resolution(reason: str = "intent_unavailable") -> IntentResolution: - return IntentResolution( - graph_id="unavailable", - graph_version="unavailable", - source_hash="unavailable", - outcome="fallback", - resolution_method="none", - refusal_reason=reason, - ) - - -def resolved( - *, - hints: tuple[str, ...] = (), - rejected: tuple[str, ...] = (), - warnings: tuple[str, ...] = (), -) -> IntentResolution: - return IntentResolution( - graph_id="mainframe.core", - graph_version="2026-06-30.1", - source_hash="fixture-hash", - outcome="resolved", - resolution_method="explicit", - matched_goal_ids=("goal.fixture",), - intent_path=("goal.fixture",), - capability_hints=hints, - rejected_capability_hints=rejected, - warnings=warnings, - ) - - -@pytest.fixture -def intent_conn(tmp_path: Path): - catalog = yaml.safe_load(INTENT_FIXTURE.read_text(encoding="utf-8")) - source = tmp_path / "intent-source" - source.mkdir() - (source / "mainframe-core.yaml").write_text( - yaml.safe_dump(catalog["documents"]["phase1"], sort_keys=False), - encoding="utf-8", - ) - destination = tmp_path / "intent.sqlite" - compile_intent_corpus(source, destination) - conn = open_intent_store(destination) - try: - yield conn - finally: - conn.close() - - -class TinyEmbedder: - def encode(self, texts, convert_to_numpy=True, show_progress_bar=False): - output = np.zeros((len(texts), 384), dtype=np.float32) - for index, text in enumerate(texts): - lowered = text.lower() - if "zebra" in lowered or "striped horse" in lowered: - output[index, 0] = 1.0 - if "compiler" in lowered: - output[index, 1] = 1.0 - return output - - -@pytest.fixture -def indexed_db(tmp_path: Path, monkeypatch): - embedder = TinyEmbedder() - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: embedder) - notes = tmp_path / "vault" - notes.mkdir() - (notes / "zebra.md").write_text( - "A zebra is a striped horse on the savannah.\n", encoding="utf-8" - ) - (notes / "compiler.md").write_text( - "A compiler translates programming languages.\n", encoding="utf-8" - ) - destination = tmp_path / "documents.sqlite" - stats = cli._ingest_directory( - notes, - str(destination), - index_id="fixture-durable", - trust_profile="durable_knowledge", - namespace="fixture", - source_root=notes, - display_prefix="10_knowledge/fixture", - ) - assert stats["failed"] == 0 - return destination, embedder - - -def test_wire_contracts_are_strict_and_duplicate_capabilities_fail(): - with pytest.raises(ValidationError): - RouteRequest(query_id="route.strict", query_text="test", surprise=True) - with pytest.raises(ValidationError): - RouteRequest( - query_id="route.duplicate", - query_text="test", - allowed_capability_refs=( - "retriever://mindgraph-durable", - "retriever://mindgraph-durable", - ), - ) - with pytest.raises(ValidationError): - RefusalRule(rule_id="privacy.empty", reason_code="privacy_refusal") - - privacy = RefusalRule( - rule_id="privacy.raw-transcript", - reason_code="privacy_refusal", - phrases=("raw transcript",), - ) - assert privacy.matches("Find the raw transcript, please.") is True - assert privacy.matches("Find the straw transcript, please.") is False - - -def test_registry_rejects_duplicate_ids_capabilities_and_executor_mismatch(): - first = RetrieverRegistration( - descriptor("one", "retriever://one", "durable_knowledge"), StubRetriever() - ) - duplicate_id = RetrieverRegistration( - descriptor("one", "retriever://two", "project_status"), StubRetriever() - ) - with pytest.raises(RoutingError, match="duplicate retriever ID"): - CapabilityRegistry((first, duplicate_id)) - - duplicate_capability = RetrieverRegistration( - descriptor("two", "retriever://one", "project_status"), StubRetriever() - ) - with pytest.raises(RoutingError, match="duplicate capability ref"): - CapabilityRegistry((first, duplicate_capability)) - - with pytest.raises(RoutingError, match="has no executor"): - RetrieverRegistration( - descriptor("missing", "retriever://missing", "project_status") - ) - with pytest.raises(RoutingError, match="has an executor"): - RetrieverRegistration( - descriptor( - "future", "retriever://future", "future_trust", available=False - ), - StubRetriever(), - ) - - registry = CapabilityRegistry((first,)) - with pytest.raises(FrozenInstanceError): - registry._registrations = () - with pytest.raises(TypeError): - registry._by_id["other"] = first - - -@pytest.mark.parametrize( - "case_id", - ["explicit_durable", "explicit_projects", "explicit_federated"], -) -def test_explicit_routes_follow_reviewed_fixture(case_id: str, policy: RouterPolicy): - case = load_routing_fixture()["requests"][case_id] - decision = decide_route( - RouteRequest.model_validate(case["request"]), - unavailable_resolution(), - make_registry(), - policy, - ) - assert decision.outcome == case["expected_outcome"] - assert list(decision.selected_retriever_ids) == case["expected_retrievers"] - assert decision.reason_codes == (case["expected_reason"],) - - -def test_explicit_federation_requires_policy_permission(policy: RouterPolicy): - request = RouteRequest( - query_id="route.federation-denied", - query_text="Use both stores", - mode="federated", - ) - decision = decide_route( - request, - unavailable_resolution(), - make_registry(), - policy.model_copy(update={"allow_federated": False}), - ) - assert decision.outcome == "refusal" - assert decision.selected_retriever_ids == () - assert decision.reason_codes == ("federation_not_allowed",) - - -def test_automatic_intent_hint_selects_only_allowed_project_store( - policy: RouterPolicy, intent_conn -): - case = load_routing_fixture()["requests"]["automatic_project_hint"] - project = StubRetriever( - [ - make_result( - "mindgraph-projects", trust_profile="project_status", score=0.4 - ) - ] - ) - envelope = orchestrate( - RouteRequest.model_validate(case["request"]), - intent_conn=intent_conn, - registry=make_registry(projects=project), - policy=policy, - ) - assert envelope.outcome == "success" - assert envelope.decision.selected_retriever_ids == ("mindgraph-projects",) - assert envelope.decision.reason_codes == (case["expected_reason"],) - assert [batch.retriever_id for batch in envelope.batches] == [ - "mindgraph-projects" - ] - assert project.calls == [(case["request"]["query_text"], 10)] - - -def test_missing_intent_store_preserves_explicit_scope_and_safe_default( - policy: RouterPolicy, -): - durable = StubRetriever( - [make_result("mindgraph-durable", trust_profile="durable_knowledge", score=0.2)] - ) - registry = make_registry(durable=durable) - explicit = orchestrate( - RouteRequest( - query_id="route.no-intent-explicit", - query_text="durable method", - mode="durable", - ), - intent_conn=None, - registry=registry, - policy=policy, - ) - fallback = orchestrate( - RouteRequest( - query_id="route.no-intent-auto", query_text="unclassified", mode="auto" - ), - intent_conn=None, - registry=registry, - policy=policy, - ) - assert explicit.outcome == "success" - assert explicit.decision.reason_codes == ("explicit_durable_scope",) - assert fallback.outcome == "fallback" - assert fallback.decision.selected_retriever_ids == ("mindgraph-durable",) - assert fallback.decision.reason_codes == ("safe_default",) - - -def test_missing_or_unmatched_intent_without_default_refuses(policy: RouterPolicy): - no_default = policy.model_copy(update={"safe_default_retriever_id": None}) - registry = make_registry() - unavailable = decide_route( - RouteRequest(query_id="route.intent-missing", query_text="unknown"), - unavailable_resolution(), - registry, - no_default, - ) - no_match = decide_route( - RouteRequest(query_id="route.intent-no-match", query_text="unknown"), - unavailable_resolution("intent_no_match"), - registry, - no_default, - ) - assert unavailable.reason_codes == ("intent_unavailable",) - assert no_match.reason_codes == ("intent_no_match",) - assert unavailable.selected_retriever_ids == no_match.selected_retriever_ids == () - - -@pytest.mark.parametrize( - ("refusal_reason", "expected"), - [ - ("intent_alias_ambiguous", "intent_ambiguous"), - ("intent_rule_ambiguous", "intent_ambiguous"), - ("intent_cycle_detected", "intent_invalid"), - ("intent_store_invalid", "intent_invalid"), - ], -) -def test_ambiguous_cycle_and_corrupt_intent_never_route( - refusal_reason: str, expected: str, policy: RouterPolicy -): - resolution = IntentResolution( - graph_id="mainframe.core", - graph_version="corrupt", - source_hash="corrupt", - outcome="refusal", - resolution_method="none", - refusal_reason=refusal_reason, - ) - decision = decide_route( - RouteRequest(query_id="route.intent-refusal", query_text="test"), - resolution, - make_registry(), - policy, - ) - assert decision.outcome == "refusal" - assert decision.selected_retriever_ids == () - assert decision.reason_codes == (expected,) - - -def test_disallowed_or_unavailable_capability_is_not_replaced_by_default( - policy: RouterPolicy, -): - registry = make_registry( - extras=( - RetrieverRegistration( - descriptor( - "future-live", - "retriever://future-live", - "time_bound_live_state", - available=False, - ) - ), - ) - ) - disallowed = decide_route( - RouteRequest(query_id="route.disallowed", query_text="test"), - resolved( - rejected=("retriever://mindgraph-projects",), - warnings=( - "capability_not_allowed:retriever://mindgraph-projects", - ), - ), - registry, - policy, - ) - unavailable = decide_route( - RouteRequest( - query_id="route.unavailable", - query_text="test", - allowed_capability_refs=("retriever://future-live",), - ), - resolved(hints=("retriever://future-live",)), - registry, - policy, - ) - assert disallowed.reason_codes == ("capability_not_allowed",) - assert unavailable.reason_codes == ("capability_unavailable",) - assert disallowed.selected_retriever_ids == unavailable.selected_retriever_ids == () - - -def test_raw_transcript_policy_refuses_before_retriever_runs(policy: RouterPolicy): - case = load_routing_fixture()["requests"]["raw_transcript_refusal"] - durable = StubRetriever() - envelope = orchestrate( - RouteRequest.model_validate(case["request"]), - intent_conn=None, - registry=make_registry(durable=durable), - policy=policy, - ) - assert envelope.outcome == "refusal" - assert envelope.decision.reason_codes == (case["expected_reason"],) - assert envelope.batches == () - assert durable.calls == [] - candidate = envelope.as_contract_result("route07_raw_transcript_request") - assert candidate["outcome"] == "refusal_or_explicit_consent_gate" - assert candidate["reason"] == "raw_transcripts_not_in_default_retrieval" - - -def test_grouped_federation_preserves_local_order_without_global_rerank( - policy: RouterPolicy, -): - durable = StubRetriever( - [ - make_result( - "mindgraph-durable", trust_profile="durable_knowledge", score=0.01 - ), - make_result( - "mindgraph-durable-2", - trust_profile="durable_knowledge", - score=0.005, - ), - ] - ) - projects = StubRetriever( - [make_result("mindgraph-projects", trust_profile="project_status", score=0.99)] - ) - envelope = orchestrate( - RouteRequest( - query_id="route.grouped", - query_text="plan across both", - mode="federated", - ), - intent_conn=None, - registry=make_registry(durable=durable, projects=projects), - policy=policy, - ) - assert envelope.outcome == "success" - assert [batch.retriever_id for batch in envelope.batches] == [ - "mindgraph-durable", - "mindgraph-projects", - ] - assert [row.local_rank for row in envelope.batches[0].rows] == [1, 2] - assert envelope.batches[0].rows[0].result.rrf_score == 0.01 - assert envelope.batches[1].rows[0].result.rrf_score == 0.99 - candidate = envelope.as_contract_result("route.grouped") - assert candidate["selected_retrievers"] == [ - "mindgraph-durable", - "mindgraph-projects", - ] - assert "group results by retriever and trust profile" in candidate["behaviors"] - - -def test_partial_failures_are_visible_when_one_or_all_retrievers_fail( - policy: RouterPolicy, -): - durable_result = make_result( - "mindgraph-durable", trust_profile="durable_knowledge", score=0.1 - ) - one_failure = orchestrate( - RouteRequest( - query_id="route.partial-one", query_text="both", mode="federated" - ), - intent_conn=None, - registry=make_registry( - durable=StubRetriever([durable_result]), - projects=FailingRetriever("project fixture failed"), - ), - policy=policy, - ) - all_fail = orchestrate( - RouteRequest( - query_id="route.partial-all", query_text="both", mode="federated" - ), - intent_conn=None, - registry=make_registry( - durable=FailingRetriever("durable fixture failed"), - projects=FailingRetriever("project fixture failed"), - ), - policy=policy, - ) - assert one_failure.outcome == "partial" - assert one_failure.partial is True - assert [batch.retriever_id for batch in one_failure.batches] == [ - "mindgraph-durable" - ] - assert [failure.retriever_id for failure in one_failure.failures] == [ - "mindgraph-projects" - ] - assert all_fail.outcome == "partial" - assert all_fail.batches == () - assert [failure.retriever_id for failure in all_fail.failures] == [ - "mindgraph-durable", - "mindgraph-projects", - ] - - -def test_mismatched_retriever_trust_is_isolated_as_a_failure(policy: RouterPolicy): - mislabeled = StubRetriever( - [ - make_result( - "mindgraph-projects", - trust_profile="durable_knowledge", - score=0.9, - ) - ] - ) - envelope = orchestrate( - RouteRequest( - query_id="route.provenance-mismatch", - query_text="project status", - mode="projects", - ), - intent_conn=None, - registry=make_registry(projects=mislabeled), - policy=policy, - ) - assert envelope.outcome == "partial" - assert envelope.batches == () - assert envelope.failures[0].error_type == "RoutingError" - assert "unexpected trust profiles" in envelope.failures[0].message - - -def test_decisions_and_envelopes_serialize_byte_deterministically(policy: RouterPolicy): - durable = StubRetriever( - [make_result("mindgraph-durable", trust_profile="durable_knowledge", score=0.1)] - ) - request = RouteRequest( - query_id="route.deterministic", query_text="durable", mode="durable" - ) - registry = make_registry(durable=durable) - first = orchestrate( - request, intent_conn=None, registry=registry, policy=policy - ).model_dump_json() - second = orchestrate( - request, intent_conn=None, registry=registry, policy=policy - ).model_dump_json() - assert first.encode("utf-8") == second.encode("utf-8") - assert json.loads(first)["batches"][0]["retriever_id"] == "mindgraph-durable" - - -def test_read_only_adapter_matches_direct_query_and_does_not_change_store(indexed_db): - db_path, embedder = indexed_db - direct_conn = db.get_db(str(db_path)) - try: - direct = query_mod.run_query( - direct_conn, "zebra", embedder, final_top_k=5 - ) - finally: - direct_conn.close() - before = hashlib.sha256(db_path.read_bytes()).hexdigest() - - query_store = open_query_store_read_only(db_path) - try: - assert query_store.execute("PRAGMA query_only").fetchone()[0] == 1 - with pytest.raises(sqlite3.OperationalError): - query_store.execute("CREATE TABLE forbidden_write(id INTEGER)") - finally: - query_store.close() - - retriever = MindGraphQueryRetriever( - descriptor=descriptor( - "mindgraph-durable", - "retriever://mindgraph-durable", - "durable_knowledge", - ), - db_path=db_path, - embedder=embedder, - ) - adapted = retriever.retrieve("zebra", final_top_k=5) - after = hashlib.sha256(db_path.read_bytes()).hexdigest() - - assert [item.model_dump() for item in adapted] == [ - item.model_dump() for item in direct - ] - assert before == after - assert adapted[0].trust_profile == "durable_knowledge" - assert adapted[0].display_path.startswith("10_knowledge/fixture/") - - -def test_policy_must_reference_registered_retrievers(policy: RouterPolicy): - registry = CapabilityRegistry( - ( - RetrieverRegistration( - descriptor( - "mindgraph-durable", - "retriever://mindgraph-durable", - "durable_knowledge", - ), - StubRetriever(), - ), - ) - ) - with pytest.raises(RoutingError, match="policy references missing retrievers"): - decide_route( - RouteRequest(query_id="route.invalid-policy", query_text="test"), - unavailable_resolution(), - registry, - policy, - ) diff --git a/mindgraph/tests/test_shared_mcp.py b/mindgraph/tests/test_shared_mcp.py deleted file mode 100644 index e1dcb4b..0000000 --- a/mindgraph/tests/test_shared_mcp.py +++ /dev/null @@ -1,305 +0,0 @@ -import json -import os -import signal -import sqlite3 -import threading -import time -from unittest.mock import AsyncMock - -import pytest -from mcp.shared.memory import create_connected_server_and_client_session -from httpx import ASGITransport, AsyncClient - -from mindgraph import cli, daemon, idle_lifecycle, mcp_proxy, mcp_server -from mindgraph.exceptions import MindgraphError -from tests.test_mcp import mcp_db, mcp_embedder -from tests.test_query import KeywordEmbedder - - -def _payload(result): - return json.loads(result.content[0].text) - - -def test_streamable_http_defaults_and_configuration(monkeypatch): - conn = sqlite3.connect(":memory:") - server = mcp_server.create_shared_server( - {"knowledge": (conn, "durable_knowledge")}, KeywordEmbedder({}) - ) - assert server.settings.host == "127.0.0.1" - assert server.settings.port == 8000 - assert server.settings.streamable_http_path == "/mcp" - called = [] - monkeypatch.setattr(server, "run", lambda transport: called.append(transport)) - mcp_server.run_streamable_http(server) - assert called == ["streamable-http"] - - -def test_shared_server_rejects_non_loopback(): - with pytest.raises(mcp_server.MCPServerStartupError, match="loopback"): - mcp_server.create_shared_server({}, KeywordEmbedder({}), host="0.0.0.0") - - -@pytest.mark.anyio -async def test_scopes_are_explicit_and_not_blended(tmp_path, monkeypatch): - embedder = KeywordEmbedder({"knowledgeonly": 0, "projectonly": 1}) - monkeypatch.setattr(cli, "_load_embedder", lambda *_a, **_k: embedder) - knowledge_vault = tmp_path / "knowledge-vault" - projects_vault = tmp_path / "projects-vault" - knowledge_vault.mkdir(); projects_vault.mkdir() - (knowledge_vault / "durable-only.md").write_text( - "knowledgeonly durable architecture note\n", encoding="utf-8" - ) - (projects_vault / "project-only.md").write_text( - "projectonly active project status\n", encoding="utf-8" - ) - knowledge_db = tmp_path / "knowledge.sqlite" - projects_db = tmp_path / "projects.sqlite" - cli._ingest_directory(knowledge_vault, str(knowledge_db)) - cli._ingest_directory(projects_vault, str(projects_db)) - first = mcp_server.open_database(str(knowledge_db)) - second = mcp_server.open_database(str(projects_db)) - server = mcp_server.create_shared_server( - {"knowledge": (first, "durable_knowledge"), - "projects": (second, "project_status")}, embedder, - ) - try: - async with create_connected_server_and_client_session(server) as session: - knowledge = _payload(await session.call_tool( - "query", {"question": "knowledgeonly", "scope": "knowledge"})) - projects = _payload(await session.call_tool( - "query", {"question": "projectonly", "scope": "projects"})) - invalid = await session.call_tool( - "query", {"question": "anything", "scope": "both"}) - assert knowledge["scope"] == "knowledge" - assert knowledge["trust_profile"] == "durable_knowledge" - assert projects["scope"] == "projects" - assert projects["trust_profile"] == "project_status" - assert {row["path"] for row in knowledge["results"]} == {"durable-only.md"} - assert {row["path"] for row in projects["results"]} == {"project-only.md"} - assert invalid.isError is True - finally: - first.close(); second.close() - - -def test_open_database_readonly_rejects_writes(mcp_db): - conn = mcp_server.open_database_readonly(mcp_db) - try: - assert conn.execute("PRAGMA query_only").fetchone()[0] == 1 - with pytest.raises(sqlite3.OperationalError, match="readonly"): - conn.execute("DELETE FROM documents") - finally: - conn.close() - - -@pytest.mark.anyio -async def test_proxy_forwards_list_and_call(): - remote = AsyncMock() - remote.list_tools.return_value.tools = [] - remote.call_tool.return_value = {"ok": True} - proxy = mcp_proxy.create_proxy_server(remote) - async with create_connected_server_and_client_session(proxy) as session: - assert (await session.list_tools()).tools == [] - await session.call_tool("query", {"scope": "knowledge", "question": "x"}) - remote.call_tool.assert_awaited_once_with( - "query", {"scope": "knowledge", "question": "x"}) - - -@pytest.mark.anyio -async def test_remote_session_initializes_and_propagates_error(monkeypatch): - initialized = AsyncMock(side_effect=RuntimeError("init failed")) - class FakeClient: - async def __aenter__(self): return self - async def __aexit__(self, *args): return False - initialize = initialized - class Transport: - async def __aenter__(self): return (object(), object(), lambda: None) - async def __aexit__(self, *args): return False - monkeypatch.setattr(mcp_proxy, "streamablehttp_client", lambda _url: Transport()) - monkeypatch.setattr(mcp_proxy, "ClientSession", lambda *_args: FakeClient()) - with pytest.raises(RuntimeError, match="init failed"): - async with mcp_proxy.remote_session("http://127.0.0.1:9/mcp"): - pass - - -def test_daemon_status_stop_and_health_use_isolated_state(tmp_path, monkeypatch): - assert daemon.status(tmp_path) == {"status": "stopped", "pid": None} - daemon.paths(tmp_path)[0].write_text("99999999\n") - assert daemon.stop(tmp_path)["status"] == "stopped" - monkeypatch.setattr(daemon.urllib.request, "urlopen", - lambda *_a, **_k: (_ for _ in ()).throw(OSError("offline"))) - assert daemon.health("http://127.0.0.1:9/health")["status"] == "unhealthy" - - -def test_daemon_start_and_stop_tracks_real_child(tmp_path): - result = daemon.start( - tmp_path, [os.sys.executable, "-c", "import time; time.sleep(30)"]) - try: - assert daemon.status(tmp_path)["status"] == "running" - assert daemon.stop(tmp_path, timeout=2)["status"] == "stopped" - finally: - if daemon.alive(result["pid"]): - os.kill(result["pid"], signal.SIGKILL) - - -def test_daemon_status_reports_healthy_external_supervisor(tmp_path, monkeypatch): - monkeypatch.setattr(daemon, "health", lambda *_a, **_k: { - "status": "ok", "pid": 4242, "scopes": [], - }) - result = daemon.status(tmp_path, "http://127.0.0.1:8000/health") - assert result["status"] == "running" - assert result["supervision"] == "external" - assert result["pid"] == 4242 - - -def test_daemon_stop_refuses_pid_mismatch(tmp_path, monkeypatch): - daemon.paths(tmp_path)[0].write_text("111\n") - monkeypatch.setattr(daemon, "health", lambda *_a, **_k: { - "status": "ok", "pid": 222, "scopes": [], - }) - assert daemon.stop(tmp_path, health_url="http://127.0.0.1:1/health") == { - "status": "refused_pid_mismatch", "pid": 111, "observed_pid": 222, - } - - -def test_daemon_stop_refuses_healthy_listener_without_identity(tmp_path, monkeypatch): - monkeypatch.setattr(daemon, "health", lambda *_a, **_k: { - "status": "ok", "scopes": [], - }) - assert daemon.stop(tmp_path, health_url="http://127.0.0.1:1/health") == { - "status": "refused_unverified_listener", "pid": None, - } - - -def test_daemon_stop_refuses_live_pid_without_start_identity(tmp_path, monkeypatch): - daemon.paths(tmp_path)[0].write_text(f"{os.getpid()}\n") - monkeypatch.setattr(daemon, "health", lambda *_a, **_k: {"status": "unhealthy"}) - assert daemon.stop(tmp_path, health_url="http://127.0.0.1:1/health") == { - "status": "refused_unverified_pid", "pid": os.getpid(), - } - - -def test_idle_lifecycle_never_exits_in_flight_and_honors_lease(): - now = [0.0] - stopped = [] - lifecycle = idle_lifecycle.IdleLifecycle( - 60, lease_ttl=30, clock=lambda: now[0], request_shutdown=lambda: stopped.append(True) - ) - with lifecycle.request(): - now[0] = 120 - assert lifecycle.should_shutdown() is False - token = lifecycle.acquire_lease() - now[0] = 140 - assert lifecycle.should_shutdown() is False - lifecycle.release_lease(token) - now[0] = 201 - assert lifecycle.should_shutdown() is True - - -def test_idle_lifecycle_monitor_requests_shutdown_only_after_grace(): - stopped = threading.Event() - lifecycle = idle_lifecycle.IdleLifecycle(0.15, request_shutdown=stopped.set) - lifecycle.start() - try: - time.sleep(0.05) - assert not stopped.is_set() - assert stopped.wait(1) - finally: - lifecycle.stop() - - -@pytest.mark.anyio -async def test_shared_server_lease_endpoint_tracks_and_releases(): - lifecycle = idle_lifecycle.IdleLifecycle(60) - server = mcp_server.create_shared_server( - {"knowledge": (sqlite3.connect(":memory:"), "durable_knowledge")}, - KeywordEmbedder({}), lifecycle=lifecycle, - ) - async with AsyncClient( - transport=ASGITransport(app=server.streamable_http_app()), - base_url="http://test", - ) as client: - acquired = (await client.post("/lifecycle/lease")).json() - assert lifecycle.snapshot()["active_leases"] == 1 - renewed = (await client.post( - "/lifecycle/lease", headers={"X-MindGraph-Lease": acquired["lease"]} - )).json() - assert renewed["lease"] == acquired["lease"] - assert (await client.delete( - "/lifecycle/lease", headers={"X-MindGraph-Lease": acquired["lease"]} - )).json()["status"] == "released" - assert lifecycle.snapshot()["active_leases"] == 0 - - -def test_daemon_start_serializes_concurrent_callers(tmp_path): - command = [os.sys.executable, "-c", "import time; time.sleep(30)"] - barrier = threading.Barrier(3) - results = [] - def launch(): - barrier.wait() - results.append(daemon.start(tmp_path, command)) - threads = [threading.Thread(target=launch) for _ in range(2)] - for thread in threads: - thread.start() - barrier.wait() - for thread in threads: - thread.join() - try: - assert {result["status"] for result in results} == {"started", "running"} - assert len({result["pid"] for result in results}) == 1 - finally: - assert daemon.stop(tmp_path, timeout=2)["status"] == "stopped" - - -def test_daemon_endpoint_args_rejects_non_loopback(): - assert daemon.daemon_endpoint_args("http://127.0.0.1:8123/custom") == ( - "127.0.0.1", 8123, "/custom" - ) - with pytest.raises(ValueError, match="loopback"): - daemon.daemon_endpoint_args("https://example.com/mcp") - assert daemon.health_url("::1", 8123) == "http://[::1]:8123/health" - - -def test_idle_opt_in_requires_explicit_valid_marker(tmp_path): - assert daemon.idle_opt_in(tmp_path) is None - (tmp_path / "idle-lifecycle.enabled").write_text("900\n") - assert daemon.idle_opt_in(tmp_path) == 900 - (tmp_path / "idle-lifecycle.enabled").write_text("30\n") - assert daemon.idle_opt_in(tmp_path) is None - - -def test_parse_scope_specs_name_only_defaults_trust_to_name(): - assert cli.parse_scope_specs(["notes=/tmp/notes.sqlite"]) == { - "notes": ("/tmp/notes.sqlite", "notes") - } - - -def test_parse_scope_specs_explicit_trust_profile(): - assert cli.parse_scope_specs(["notes:durable_knowledge=/tmp/n.sqlite"]) == { - "notes": ("/tmp/n.sqlite", "durable_knowledge") - } - - -def test_parse_scope_specs_allows_equals_and_colon_in_path(): - parsed = cli.parse_scope_specs(["a=/tmp/x:y=z.sqlite"]) - assert parsed == {"a": ("/tmp/x:y=z.sqlite", "a")} - - -def test_parse_scope_specs_multiple_scopes(): - parsed = cli.parse_scope_specs(["a=/tmp/a.sqlite", "b:vol=/tmp/b.sqlite"]) - assert parsed == {"a": ("/tmp/a.sqlite", "a"), "b": ("/tmp/b.sqlite", "vol")} - - -def test_parse_scope_specs_empty_is_empty_dict(): - assert cli.parse_scope_specs(None) == {} - assert cli.parse_scope_specs([]) == {} - - -@pytest.mark.parametrize("bad", ["notes", "=/tmp/a.sqlite", "notes=", ":t=/tmp/a.sqlite"]) -def test_parse_scope_specs_rejects_malformed(bad): - with pytest.raises(MindgraphError): - cli.parse_scope_specs([bad]) - - -def test_parse_scope_specs_rejects_duplicate_names(): - with pytest.raises(MindgraphError): - cli.parse_scope_specs(["a=/tmp/1.sqlite", "a=/tmp/2.sqlite"]) diff --git a/mindgraph/uv.lock b/mindgraph/uv.lock deleted file mode 100644 index dd802cd..0000000 --- a/mindgraph/uv.lock +++ /dev/null @@ -1,2399 +0,0 @@ -version = 1 -revision = 3 -requires-python = ">=3.10" -resolution-markers = [ - "python_full_version >= '3.14' and sys_platform == 'win32'", - "python_full_version >= '3.14' and sys_platform != 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform == 'win32'", - "python_full_version == '3.11.*' and sys_platform == 'win32'", - "python_full_version < '3.11' and sys_platform == 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform != 'win32'", - "python_full_version == '3.11.*' and sys_platform != 'win32'", - "python_full_version < '3.11' and sys_platform != 'win32'", -] - -[[package]] -name = "annotated-doc" -version = "0.0.4" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/57/ba/046ceea27344560984e26a590f90bc7f4a75b06701f653222458922b558c/annotated_doc-0.0.4.tar.gz", hash = "sha256:fbcda96e87e9c92ad167c2e53839e57503ecfda18804ea28102353485033faa4", size = 7288, upload-time = "2025-11-10T22:07:42.062Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/1e/d3/26bf1008eb3d2daa8ef4cacc7f3bfdc11818d111f7e2d0201bc6e3b49d45/annotated_doc-0.0.4-py3-none-any.whl", hash = "sha256:571ac1dc6991c450b25a9c2d84a3705e2ae7a53467b5d111c24fa8baabbed320", size = 5303, upload-time = "2025-11-10T22:07:40.673Z" }, -] - -[[package]] -name = "annotated-types" -version = "0.7.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/ee/67/531ea369ba64dcff5ec9c3402f9f51bf748cec26dde048a2f973a4eea7f5/annotated_types-0.7.0.tar.gz", hash = "sha256:aff07c09a53a08bc8cfccb9c85b05f1aa9a2a6f23728d790723543408344ce89", size = 16081, upload-time = "2024-05-20T21:33:25.928Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/78/b6/6307fbef88d9b5ee7421e68d78a9f162e0da4900bc5f5793f6d3d0e34fb8/annotated_types-0.7.0-py3-none-any.whl", hash = "sha256:1f02e8b43a8fbbc3f3e0d4f0f4bfc8131bcb4eebe8849b8e5c773f3a1c582a53", size = 13643, upload-time = "2024-05-20T21:33:24.1Z" }, -] - -[[package]] -name = "anyio" -version = "4.14.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "exceptiongroup", marker = "python_full_version < '3.11'" }, - { name = "idna" }, - { name = "typing-extensions", marker = "python_full_version < '3.13'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/1c/b5/001890774a9552aff22502b8da382593109ce0c95314abaebbb116567545/anyio-4.14.0.tar.gz", hash = "sha256:b47c1f9ccf73e67021df785332508f99379c68fa7d0684e8e3492cb1d4b23f89", size = 253586, upload-time = "2026-06-15T22:00:49.021Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/ba/16/9826f089383c593cdfc4a6e5aca94d9e91ae1692c57af82c3b2aa5e810f7/anyio-4.14.0-py3-none-any.whl", hash = "sha256:dd9b7a2a9799ed6552fde617b2c5df02b7fdd7d88392fc48101e51bae46164d9", size = 123506, upload-time = "2026-06-15T22:00:47.595Z" }, -] - -[[package]] -name = "attrs" -version = "26.1.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/9a/8e/82a0fe20a541c03148528be8cac2408564a6c9a0cc7e9171802bc1d26985/attrs-26.1.0.tar.gz", hash = "sha256:d03ceb89cb322a8fd706d4fb91940737b6642aa36998fe130a9bc96c985eff32", size = 952055, upload-time = "2026-03-19T14:22:25.026Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/64/b4/17d4b0b2a2dc85a6df63d1157e028ed19f90d4cd97c36717afef2bc2f395/attrs-26.1.0-py3-none-any.whl", hash = "sha256:c647aa4a12dfbad9333ca4e71fe62ddc36f4e63b2d260a37a8b83d2f043ac309", size = 67548, upload-time = "2026-03-19T14:22:23.645Z" }, -] - -[[package]] -name = "certifi" -version = "2026.6.17" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/c9/c7/424b75da314c1045981bd9777432fad05a9e0c69daa4ed7e308bbaffe405/certifi-2026.6.17.tar.gz", hash = "sha256:024c88eeec92ca068db80f02b8b07c9cef7b9fe261d1d535abfd5abd6f6af432", size = 134594, upload-time = "2026-06-17T10:31:07.894Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/ef/2f/c5464532e965badff2f4c4c1a3a83f5697f0d7c407ed0cda44aaa99bb451/certifi-2026.6.17-py3-none-any.whl", hash = "sha256:2227dcbaafe0d2f59279d1762ddddc37783ed4354594f194ffc31d20f41fc3db", size = 133289, upload-time = "2026-06-17T10:31:06.348Z" }, -] - -[[package]] -name = "cffi" -version = "2.0.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "pycparser", marker = "implementation_name != 'PyPy'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/eb/56/b1ba7935a17738ae8453301356628e8147c79dbb825bcbc73dc7401f9846/cffi-2.0.0.tar.gz", hash = "sha256:44d1b5909021139fe36001ae048dbdde8214afa20200eda0f64c068cac5d5529", size = 523588, upload-time = "2025-09-08T23:24:04.541Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/93/d7/516d984057745a6cd96575eea814fe1edd6646ee6efd552fb7b0921dec83/cffi-2.0.0-cp310-cp310-macosx_10_13_x86_64.whl", hash = "sha256:0cf2d91ecc3fcc0625c2c530fe004f82c110405f101548512cce44322fa8ac44", size = 184283, upload-time = "2025-09-08T23:22:08.01Z" }, - { url = "https://files.pythonhosted.org/packages/9e/84/ad6a0b408daa859246f57c03efd28e5dd1b33c21737c2db84cae8c237aa5/cffi-2.0.0-cp310-cp310-macosx_11_0_arm64.whl", hash = "sha256:f73b96c41e3b2adedc34a7356e64c8eb96e03a3782b535e043a986276ce12a49", size = 180504, upload-time = "2025-09-08T23:22:10.637Z" }, - { url = "https://files.pythonhosted.org/packages/50/bd/b1a6362b80628111e6653c961f987faa55262b4002fcec42308cad1db680/cffi-2.0.0-cp310-cp310-manylinux1_i686.manylinux2014_i686.manylinux_2_17_i686.manylinux_2_5_i686.whl", hash = "sha256:53f77cbe57044e88bbd5ed26ac1d0514d2acf0591dd6bb02a3ae37f76811b80c", size = 208811, upload-time = "2025-09-08T23:22:12.267Z" }, - { url = "https://files.pythonhosted.org/packages/4f/27/6933a8b2562d7bd1fb595074cf99cc81fc3789f6a6c05cdabb46284a3188/cffi-2.0.0-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:3e837e369566884707ddaf85fc1744b47575005c0a229de3327f8f9a20f4efeb", size = 216402, upload-time = "2025-09-08T23:22:13.455Z" }, - { url = "https://files.pythonhosted.org/packages/05/eb/b86f2a2645b62adcfff53b0dd97e8dfafb5c8aa864bd0d9a2c2049a0d551/cffi-2.0.0-cp310-cp310-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:5eda85d6d1879e692d546a078b44251cdd08dd1cfb98dfb77b670c97cee49ea0", size = 203217, upload-time = "2025-09-08T23:22:14.596Z" }, - { url = "https://files.pythonhosted.org/packages/9f/e0/6cbe77a53acf5acc7c08cc186c9928864bd7c005f9efd0d126884858a5fe/cffi-2.0.0-cp310-cp310-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:9332088d75dc3241c702d852d4671613136d90fa6881da7d770a483fd05248b4", size = 203079, upload-time = "2025-09-08T23:22:15.769Z" }, - { url = "https://files.pythonhosted.org/packages/98/29/9b366e70e243eb3d14a5cb488dfd3a0b6b2f1fb001a203f653b93ccfac88/cffi-2.0.0-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:fc7de24befaeae77ba923797c7c87834c73648a05a4bde34b3b7e5588973a453", size = 216475, upload-time = "2025-09-08T23:22:17.427Z" }, - { url = "https://files.pythonhosted.org/packages/21/7a/13b24e70d2f90a322f2900c5d8e1f14fa7e2a6b3332b7309ba7b2ba51a5a/cffi-2.0.0-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:cf364028c016c03078a23b503f02058f1814320a56ad535686f90565636a9495", size = 218829, upload-time = "2025-09-08T23:22:19.069Z" }, - { url = "https://files.pythonhosted.org/packages/60/99/c9dc110974c59cc981b1f5b66e1d8af8af764e00f0293266824d9c4254bc/cffi-2.0.0-cp310-cp310-musllinux_1_2_i686.whl", hash = "sha256:e11e82b744887154b182fd3e7e8512418446501191994dbf9c9fc1f32cc8efd5", size = 211211, upload-time = "2025-09-08T23:22:20.588Z" }, - { url = "https://files.pythonhosted.org/packages/49/72/ff2d12dbf21aca1b32a40ed792ee6b40f6dc3a9cf1644bd7ef6e95e0ac5e/cffi-2.0.0-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:8ea985900c5c95ce9db1745f7933eeef5d314f0565b27625d9a10ec9881e1bfb", size = 218036, upload-time = "2025-09-08T23:22:22.143Z" }, - { url = "https://files.pythonhosted.org/packages/e2/cc/027d7fb82e58c48ea717149b03bcadcbdc293553edb283af792bd4bcbb3f/cffi-2.0.0-cp310-cp310-win32.whl", hash = "sha256:1f72fb8906754ac8a2cc3f9f5aaa298070652a0ffae577e0ea9bd480dc3c931a", size = 172184, upload-time = "2025-09-08T23:22:23.328Z" }, - { url = "https://files.pythonhosted.org/packages/33/fa/072dd15ae27fbb4e06b437eb6e944e75b068deb09e2a2826039e49ee2045/cffi-2.0.0-cp310-cp310-win_amd64.whl", hash = "sha256:b18a3ed7d5b3bd8d9ef7a8cb226502c6bf8308df1525e1cc676c3680e7176739", size = 182790, upload-time = "2025-09-08T23:22:24.752Z" }, - { url = "https://files.pythonhosted.org/packages/12/4a/3dfd5f7850cbf0d06dc84ba9aa00db766b52ca38d8b86e3a38314d52498c/cffi-2.0.0-cp311-cp311-macosx_10_13_x86_64.whl", hash = "sha256:b4c854ef3adc177950a8dfc81a86f5115d2abd545751a304c5bcf2c2c7283cfe", size = 184344, upload-time = "2025-09-08T23:22:26.456Z" }, - { url = "https://files.pythonhosted.org/packages/4f/8b/f0e4c441227ba756aafbe78f117485b25bb26b1c059d01f137fa6d14896b/cffi-2.0.0-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:2de9a304e27f7596cd03d16f1b7c72219bd944e99cc52b84d0145aefb07cbd3c", size = 180560, upload-time = "2025-09-08T23:22:28.197Z" }, - { url = "https://files.pythonhosted.org/packages/b1/b7/1200d354378ef52ec227395d95c2576330fd22a869f7a70e88e1447eb234/cffi-2.0.0-cp311-cp311-manylinux1_i686.manylinux2014_i686.manylinux_2_17_i686.manylinux_2_5_i686.whl", hash = "sha256:baf5215e0ab74c16e2dd324e8ec067ef59e41125d3eade2b863d294fd5035c92", size = 209613, upload-time = "2025-09-08T23:22:29.475Z" }, - { url = "https://files.pythonhosted.org/packages/b8/56/6033f5e86e8cc9bb629f0077ba71679508bdf54a9a5e112a3c0b91870332/cffi-2.0.0-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:730cacb21e1bdff3ce90babf007d0a0917cc3e6492f336c2f0134101e0944f93", size = 216476, upload-time = "2025-09-08T23:22:31.063Z" }, - { url = "https://files.pythonhosted.org/packages/dc/7f/55fecd70f7ece178db2f26128ec41430d8720f2d12ca97bf8f0a628207d5/cffi-2.0.0-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:6824f87845e3396029f3820c206e459ccc91760e8fa24422f8b0c3d1731cbec5", size = 203374, upload-time = "2025-09-08T23:22:32.507Z" }, - { url = "https://files.pythonhosted.org/packages/84/ef/a7b77c8bdc0f77adc3b46888f1ad54be8f3b7821697a7b89126e829e676a/cffi-2.0.0-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:9de40a7b0323d889cf8d23d1ef214f565ab154443c42737dfe52ff82cf857664", size = 202597, upload-time = "2025-09-08T23:22:34.132Z" }, - { url = "https://files.pythonhosted.org/packages/d7/91/500d892b2bf36529a75b77958edfcd5ad8e2ce4064ce2ecfeab2125d72d1/cffi-2.0.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:8941aaadaf67246224cee8c3803777eed332a19d909b47e29c9842ef1e79ac26", size = 215574, upload-time = "2025-09-08T23:22:35.443Z" }, - { url = "https://files.pythonhosted.org/packages/44/64/58f6255b62b101093d5df22dcb752596066c7e89dd725e0afaed242a61be/cffi-2.0.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:a05d0c237b3349096d3981b727493e22147f934b20f6f125a3eba8f994bec4a9", size = 218971, upload-time = "2025-09-08T23:22:36.805Z" }, - { url = "https://files.pythonhosted.org/packages/ab/49/fa72cebe2fd8a55fbe14956f9970fe8eb1ac59e5df042f603ef7c8ba0adc/cffi-2.0.0-cp311-cp311-musllinux_1_2_i686.whl", hash = "sha256:94698a9c5f91f9d138526b48fe26a199609544591f859c870d477351dc7b2414", size = 211972, upload-time = "2025-09-08T23:22:38.436Z" }, - { url = "https://files.pythonhosted.org/packages/0b/28/dd0967a76aab36731b6ebfe64dec4e981aff7e0608f60c2d46b46982607d/cffi-2.0.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:5fed36fccc0612a53f1d4d9a816b50a36702c28a2aa880cb8a122b3466638743", size = 217078, upload-time = "2025-09-08T23:22:39.776Z" }, - { url = "https://files.pythonhosted.org/packages/2b/c0/015b25184413d7ab0a410775fdb4a50fca20f5589b5dab1dbbfa3baad8ce/cffi-2.0.0-cp311-cp311-win32.whl", hash = "sha256:c649e3a33450ec82378822b3dad03cc228b8f5963c0c12fc3b1e0ab940f768a5", size = 172076, upload-time = "2025-09-08T23:22:40.95Z" }, - { url = "https://files.pythonhosted.org/packages/ae/8f/dc5531155e7070361eb1b7e4c1a9d896d0cb21c49f807a6c03fd63fc877e/cffi-2.0.0-cp311-cp311-win_amd64.whl", hash = "sha256:66f011380d0e49ed280c789fbd08ff0d40968ee7b665575489afa95c98196ab5", size = 182820, upload-time = "2025-09-08T23:22:42.463Z" }, - { url = "https://files.pythonhosted.org/packages/95/5c/1b493356429f9aecfd56bc171285a4c4ac8697f76e9bbbbb105e537853a1/cffi-2.0.0-cp311-cp311-win_arm64.whl", hash = "sha256:c6638687455baf640e37344fe26d37c404db8b80d037c3d29f58fe8d1c3b194d", size = 177635, upload-time = "2025-09-08T23:22:43.623Z" }, - { url = "https://files.pythonhosted.org/packages/ea/47/4f61023ea636104d4f16ab488e268b93008c3d0bb76893b1b31db1f96802/cffi-2.0.0-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:6d02d6655b0e54f54c4ef0b94eb6be0607b70853c45ce98bd278dc7de718be5d", size = 185271, upload-time = "2025-09-08T23:22:44.795Z" }, - { url = "https://files.pythonhosted.org/packages/df/a2/781b623f57358e360d62cdd7a8c681f074a71d445418a776eef0aadb4ab4/cffi-2.0.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:8eca2a813c1cb7ad4fb74d368c2ffbbb4789d377ee5bb8df98373c2cc0dee76c", size = 181048, upload-time = "2025-09-08T23:22:45.938Z" }, - { url = "https://files.pythonhosted.org/packages/ff/df/a4f0fbd47331ceeba3d37c2e51e9dfc9722498becbeec2bd8bc856c9538a/cffi-2.0.0-cp312-cp312-manylinux1_i686.manylinux2014_i686.manylinux_2_17_i686.manylinux_2_5_i686.whl", hash = "sha256:21d1152871b019407d8ac3985f6775c079416c282e431a4da6afe7aefd2bccbe", size = 212529, upload-time = "2025-09-08T23:22:47.349Z" }, - { url = "https://files.pythonhosted.org/packages/d5/72/12b5f8d3865bf0f87cf1404d8c374e7487dcf097a1c91c436e72e6badd83/cffi-2.0.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:b21e08af67b8a103c71a250401c78d5e0893beff75e28c53c98f4de42f774062", size = 220097, upload-time = "2025-09-08T23:22:48.677Z" }, - { url = "https://files.pythonhosted.org/packages/c2/95/7a135d52a50dfa7c882ab0ac17e8dc11cec9d55d2c18dda414c051c5e69e/cffi-2.0.0-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:1e3a615586f05fc4065a8b22b8152f0c1b00cdbc60596d187c2a74f9e3036e4e", size = 207983, upload-time = "2025-09-08T23:22:50.06Z" }, - { url = "https://files.pythonhosted.org/packages/3a/c8/15cb9ada8895957ea171c62dc78ff3e99159ee7adb13c0123c001a2546c1/cffi-2.0.0-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:81afed14892743bbe14dacb9e36d9e0e504cd204e0b165062c488942b9718037", size = 206519, upload-time = "2025-09-08T23:22:51.364Z" }, - { url = "https://files.pythonhosted.org/packages/78/2d/7fa73dfa841b5ac06c7b8855cfc18622132e365f5b81d02230333ff26e9e/cffi-2.0.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:3e17ed538242334bf70832644a32a7aae3d83b57567f9fd60a26257e992b79ba", size = 219572, upload-time = "2025-09-08T23:22:52.902Z" }, - { url = "https://files.pythonhosted.org/packages/07/e0/267e57e387b4ca276b90f0434ff88b2c2241ad72b16d31836adddfd6031b/cffi-2.0.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:3925dd22fa2b7699ed2617149842d2e6adde22b262fcbfada50e3d195e4b3a94", size = 222963, upload-time = "2025-09-08T23:22:54.518Z" }, - { url = "https://files.pythonhosted.org/packages/b6/75/1f2747525e06f53efbd878f4d03bac5b859cbc11c633d0fb81432d98a795/cffi-2.0.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:2c8f814d84194c9ea681642fd164267891702542f028a15fc97d4674b6206187", size = 221361, upload-time = "2025-09-08T23:22:55.867Z" }, - { url = "https://files.pythonhosted.org/packages/7b/2b/2b6435f76bfeb6bbf055596976da087377ede68df465419d192acf00c437/cffi-2.0.0-cp312-cp312-win32.whl", hash = "sha256:da902562c3e9c550df360bfa53c035b2f241fed6d9aef119048073680ace4a18", size = 172932, upload-time = "2025-09-08T23:22:57.188Z" }, - { url = "https://files.pythonhosted.org/packages/f8/ed/13bd4418627013bec4ed6e54283b1959cf6db888048c7cf4b4c3b5b36002/cffi-2.0.0-cp312-cp312-win_amd64.whl", hash = "sha256:da68248800ad6320861f129cd9c1bf96ca849a2771a59e0344e88681905916f5", size = 183557, upload-time = "2025-09-08T23:22:58.351Z" }, - { url = "https://files.pythonhosted.org/packages/95/31/9f7f93ad2f8eff1dbc1c3656d7ca5bfd8fb52c9d786b4dcf19b2d02217fa/cffi-2.0.0-cp312-cp312-win_arm64.whl", hash = "sha256:4671d9dd5ec934cb9a73e7ee9676f9362aba54f7f34910956b84d727b0d73fb6", size = 177762, upload-time = "2025-09-08T23:22:59.668Z" }, - { url = "https://files.pythonhosted.org/packages/4b/8d/a0a47a0c9e413a658623d014e91e74a50cdd2c423f7ccfd44086ef767f90/cffi-2.0.0-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:00bdf7acc5f795150faa6957054fbbca2439db2f775ce831222b66f192f03beb", size = 185230, upload-time = "2025-09-08T23:23:00.879Z" }, - { url = "https://files.pythonhosted.org/packages/4a/d2/a6c0296814556c68ee32009d9c2ad4f85f2707cdecfd7727951ec228005d/cffi-2.0.0-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:45d5e886156860dc35862657e1494b9bae8dfa63bf56796f2fb56e1679fc0bca", size = 181043, upload-time = "2025-09-08T23:23:02.231Z" }, - { url = "https://files.pythonhosted.org/packages/b0/1e/d22cc63332bd59b06481ceaac49d6c507598642e2230f201649058a7e704/cffi-2.0.0-cp313-cp313-manylinux1_i686.manylinux2014_i686.manylinux_2_17_i686.manylinux_2_5_i686.whl", hash = "sha256:07b271772c100085dd28b74fa0cd81c8fb1a3ba18b21e03d7c27f3436a10606b", size = 212446, upload-time = "2025-09-08T23:23:03.472Z" }, - { url = "https://files.pythonhosted.org/packages/a9/f5/a2c23eb03b61a0b8747f211eb716446c826ad66818ddc7810cc2cc19b3f2/cffi-2.0.0-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:d48a880098c96020b02d5a1f7d9251308510ce8858940e6fa99ece33f610838b", size = 220101, upload-time = "2025-09-08T23:23:04.792Z" }, - { url = "https://files.pythonhosted.org/packages/f2/7f/e6647792fc5850d634695bc0e6ab4111ae88e89981d35ac269956605feba/cffi-2.0.0-cp313-cp313-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:f93fd8e5c8c0a4aa1f424d6173f14a892044054871c771f8566e4008eaa359d2", size = 207948, upload-time = "2025-09-08T23:23:06.127Z" }, - { url = "https://files.pythonhosted.org/packages/cb/1e/a5a1bd6f1fb30f22573f76533de12a00bf274abcdc55c8edab639078abb6/cffi-2.0.0-cp313-cp313-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:dd4f05f54a52fb558f1ba9f528228066954fee3ebe629fc1660d874d040ae5a3", size = 206422, upload-time = "2025-09-08T23:23:07.753Z" }, - { url = "https://files.pythonhosted.org/packages/98/df/0a1755e750013a2081e863e7cd37e0cdd02664372c754e5560099eb7aa44/cffi-2.0.0-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:c8d3b5532fc71b7a77c09192b4a5a200ea992702734a2e9279a37f2478236f26", size = 219499, upload-time = "2025-09-08T23:23:09.648Z" }, - { url = "https://files.pythonhosted.org/packages/50/e1/a969e687fcf9ea58e6e2a928ad5e2dd88cc12f6f0ab477e9971f2309b57c/cffi-2.0.0-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:d9b29c1f0ae438d5ee9acb31cadee00a58c46cc9c0b2f9038c6b0b3470877a8c", size = 222928, upload-time = "2025-09-08T23:23:10.928Z" }, - { url = "https://files.pythonhosted.org/packages/36/54/0362578dd2c9e557a28ac77698ed67323ed5b9775ca9d3fe73fe191bb5d8/cffi-2.0.0-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:6d50360be4546678fc1b79ffe7a66265e28667840010348dd69a314145807a1b", size = 221302, upload-time = "2025-09-08T23:23:12.42Z" }, - { url = "https://files.pythonhosted.org/packages/eb/6d/bf9bda840d5f1dfdbf0feca87fbdb64a918a69bca42cfa0ba7b137c48cb8/cffi-2.0.0-cp313-cp313-win32.whl", hash = "sha256:74a03b9698e198d47562765773b4a8309919089150a0bb17d829ad7b44b60d27", size = 172909, upload-time = "2025-09-08T23:23:14.32Z" }, - { url = "https://files.pythonhosted.org/packages/37/18/6519e1ee6f5a1e579e04b9ddb6f1676c17368a7aba48299c3759bbc3c8b3/cffi-2.0.0-cp313-cp313-win_amd64.whl", hash = "sha256:19f705ada2530c1167abacb171925dd886168931e0a7b78f5bffcae5c6b5be75", size = 183402, upload-time = "2025-09-08T23:23:15.535Z" }, - { url = "https://files.pythonhosted.org/packages/cb/0e/02ceeec9a7d6ee63bb596121c2c8e9b3a9e150936f4fbef6ca1943e6137c/cffi-2.0.0-cp313-cp313-win_arm64.whl", hash = "sha256:256f80b80ca3853f90c21b23ee78cd008713787b1b1e93eae9f3d6a7134abd91", size = 177780, upload-time = "2025-09-08T23:23:16.761Z" }, - { url = "https://files.pythonhosted.org/packages/92/c4/3ce07396253a83250ee98564f8d7e9789fab8e58858f35d07a9a2c78de9f/cffi-2.0.0-cp314-cp314-macosx_10_13_x86_64.whl", hash = "sha256:fc33c5141b55ed366cfaad382df24fe7dcbc686de5be719b207bb248e3053dc5", size = 185320, upload-time = "2025-09-08T23:23:18.087Z" }, - { url = "https://files.pythonhosted.org/packages/59/dd/27e9fa567a23931c838c6b02d0764611c62290062a6d4e8ff7863daf9730/cffi-2.0.0-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:c654de545946e0db659b3400168c9ad31b5d29593291482c43e3564effbcee13", size = 181487, upload-time = "2025-09-08T23:23:19.622Z" }, - { url = "https://files.pythonhosted.org/packages/d6/43/0e822876f87ea8a4ef95442c3d766a06a51fc5298823f884ef87aaad168c/cffi-2.0.0-cp314-cp314-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:24b6f81f1983e6df8db3adc38562c83f7d4a0c36162885ec7f7b77c7dcbec97b", size = 220049, upload-time = "2025-09-08T23:23:20.853Z" }, - { url = "https://files.pythonhosted.org/packages/b4/89/76799151d9c2d2d1ead63c2429da9ea9d7aac304603de0c6e8764e6e8e70/cffi-2.0.0-cp314-cp314-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:12873ca6cb9b0f0d3a0da705d6086fe911591737a59f28b7936bdfed27c0d47c", size = 207793, upload-time = "2025-09-08T23:23:22.08Z" }, - { url = "https://files.pythonhosted.org/packages/bb/dd/3465b14bb9e24ee24cb88c9e3730f6de63111fffe513492bf8c808a3547e/cffi-2.0.0-cp314-cp314-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:d9b97165e8aed9272a6bb17c01e3cc5871a594a446ebedc996e2397a1c1ea8ef", size = 206300, upload-time = "2025-09-08T23:23:23.314Z" }, - { url = "https://files.pythonhosted.org/packages/47/d9/d83e293854571c877a92da46fdec39158f8d7e68da75bf73581225d28e90/cffi-2.0.0-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:afb8db5439b81cf9c9d0c80404b60c3cc9c3add93e114dcae767f1477cb53775", size = 219244, upload-time = "2025-09-08T23:23:24.541Z" }, - { url = "https://files.pythonhosted.org/packages/2b/0f/1f177e3683aead2bb00f7679a16451d302c436b5cbf2505f0ea8146ef59e/cffi-2.0.0-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:737fe7d37e1a1bffe70bd5754ea763a62a066dc5913ca57e957824b72a85e205", size = 222828, upload-time = "2025-09-08T23:23:26.143Z" }, - { url = "https://files.pythonhosted.org/packages/c6/0f/cafacebd4b040e3119dcb32fed8bdef8dfe94da653155f9d0b9dc660166e/cffi-2.0.0-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:38100abb9d1b1435bc4cc340bb4489635dc2f0da7456590877030c9b3d40b0c1", size = 220926, upload-time = "2025-09-08T23:23:27.873Z" }, - { url = "https://files.pythonhosted.org/packages/3e/aa/df335faa45b395396fcbc03de2dfcab242cd61a9900e914fe682a59170b1/cffi-2.0.0-cp314-cp314-win32.whl", hash = "sha256:087067fa8953339c723661eda6b54bc98c5625757ea62e95eb4898ad5e776e9f", size = 175328, upload-time = "2025-09-08T23:23:44.61Z" }, - { url = "https://files.pythonhosted.org/packages/bb/92/882c2d30831744296ce713f0feb4c1cd30f346ef747b530b5318715cc367/cffi-2.0.0-cp314-cp314-win_amd64.whl", hash = "sha256:203a48d1fb583fc7d78a4c6655692963b860a417c0528492a6bc21f1aaefab25", size = 185650, upload-time = "2025-09-08T23:23:45.848Z" }, - { url = "https://files.pythonhosted.org/packages/9f/2c/98ece204b9d35a7366b5b2c6539c350313ca13932143e79dc133ba757104/cffi-2.0.0-cp314-cp314-win_arm64.whl", hash = "sha256:dbd5c7a25a7cb98f5ca55d258b103a2054f859a46ae11aaf23134f9cc0d356ad", size = 180687, upload-time = "2025-09-08T23:23:47.105Z" }, - { url = "https://files.pythonhosted.org/packages/3e/61/c768e4d548bfa607abcda77423448df8c471f25dbe64fb2ef6d555eae006/cffi-2.0.0-cp314-cp314t-macosx_10_13_x86_64.whl", hash = "sha256:9a67fc9e8eb39039280526379fb3a70023d77caec1852002b4da7e8b270c4dd9", size = 188773, upload-time = "2025-09-08T23:23:29.347Z" }, - { url = "https://files.pythonhosted.org/packages/2c/ea/5f76bce7cf6fcd0ab1a1058b5af899bfbef198bea4d5686da88471ea0336/cffi-2.0.0-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:7a66c7204d8869299919db4d5069a82f1561581af12b11b3c9f48c584eb8743d", size = 185013, upload-time = "2025-09-08T23:23:30.63Z" }, - { url = "https://files.pythonhosted.org/packages/be/b4/c56878d0d1755cf9caa54ba71e5d049479c52f9e4afc230f06822162ab2f/cffi-2.0.0-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:7cc09976e8b56f8cebd752f7113ad07752461f48a58cbba644139015ac24954c", size = 221593, upload-time = "2025-09-08T23:23:31.91Z" }, - { url = "https://files.pythonhosted.org/packages/e0/0d/eb704606dfe8033e7128df5e90fee946bbcb64a04fcdaa97321309004000/cffi-2.0.0-cp314-cp314t-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:92b68146a71df78564e4ef48af17551a5ddd142e5190cdf2c5624d0c3ff5b2e8", size = 209354, upload-time = "2025-09-08T23:23:33.214Z" }, - { url = "https://files.pythonhosted.org/packages/d8/19/3c435d727b368ca475fb8742ab97c9cb13a0de600ce86f62eab7fa3eea60/cffi-2.0.0-cp314-cp314t-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:b1e74d11748e7e98e2f426ab176d4ed720a64412b6a15054378afdb71e0f37dc", size = 208480, upload-time = "2025-09-08T23:23:34.495Z" }, - { url = "https://files.pythonhosted.org/packages/d0/44/681604464ed9541673e486521497406fadcc15b5217c3e326b061696899a/cffi-2.0.0-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:28a3a209b96630bca57cce802da70c266eb08c6e97e5afd61a75611ee6c64592", size = 221584, upload-time = "2025-09-08T23:23:36.096Z" }, - { url = "https://files.pythonhosted.org/packages/25/8e/342a504ff018a2825d395d44d63a767dd8ebc927ebda557fecdaca3ac33a/cffi-2.0.0-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:7553fb2090d71822f02c629afe6042c299edf91ba1bf94951165613553984512", size = 224443, upload-time = "2025-09-08T23:23:37.328Z" }, - { url = "https://files.pythonhosted.org/packages/e1/5e/b666bacbbc60fbf415ba9988324a132c9a7a0448a9a8f125074671c0f2c3/cffi-2.0.0-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:6c6c373cfc5c83a975506110d17457138c8c63016b563cc9ed6e056a82f13ce4", size = 223437, upload-time = "2025-09-08T23:23:38.945Z" }, - { url = "https://files.pythonhosted.org/packages/a0/1d/ec1a60bd1a10daa292d3cd6bb0b359a81607154fb8165f3ec95fe003b85c/cffi-2.0.0-cp314-cp314t-win32.whl", hash = "sha256:1fc9ea04857caf665289b7a75923f2c6ed559b8298a1b8c49e59f7dd95c8481e", size = 180487, upload-time = "2025-09-08T23:23:40.423Z" }, - { url = "https://files.pythonhosted.org/packages/bf/41/4c1168c74fac325c0c8156f04b6749c8b6a8f405bbf91413ba088359f60d/cffi-2.0.0-cp314-cp314t-win_amd64.whl", hash = "sha256:d68b6cef7827e8641e8ef16f4494edda8b36104d79773a334beaa1e3521430f6", size = 191726, upload-time = "2025-09-08T23:23:41.742Z" }, - { url = "https://files.pythonhosted.org/packages/ae/3a/dbeec9d1ee0844c679f6bb5d6ad4e9f198b1224f4e7a32825f47f6192b0c/cffi-2.0.0-cp314-cp314t-win_arm64.whl", hash = "sha256:0a1527a803f0a659de1af2e1fd700213caba79377e27e4693648c2923da066f9", size = 184195, upload-time = "2025-09-08T23:23:43.004Z" }, -] - -[[package]] -name = "click" -version = "8.4.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "colorama", marker = "sys_platform == 'win32'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/9b/98/518d8e5081007684232226f475082b30087d0f585e8457db087298259f49/click-8.4.1.tar.gz", hash = "sha256:918b5633eddf6b41c32d4f454bf0de810065c74e3f7dbf8ee5452f8be88d3e96", size = 353007, upload-time = "2026-05-22T04:08:37.769Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/c7/0d/67e5b4109ea4a837e80daa87c2c696711955e40449a97e8926672534def2/click-8.4.1-py3-none-any.whl", hash = "sha256:482be17c6991b8c19c5429a1e995d9b0efdbb63172824c41f99965dc0ade8ec2", size = 116639, upload-time = "2026-05-22T04:08:35.26Z" }, -] - -[[package]] -name = "colorama" -version = "0.4.6" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/d8/53/6f443c9a4a8358a93a6792e2acffb9d9d5cb0a5cfd8802644b7b1c9a02e4/colorama-0.4.6.tar.gz", hash = "sha256:08695f5cb7ed6e0531a20572697297273c47b8cae5a63ffc6d6ed5c201be6e44", size = 27697, upload-time = "2022-10-25T02:36:22.414Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/d1/d6/3965ed04c63042e047cb6a3e6ed1a63a35087b6a609aa3a15ed8ac56c221/colorama-0.4.6-py2.py3-none-any.whl", hash = "sha256:4f1d9991f5acc0ca119f9d443620b77f9d6b33703e51011c16baf57afb285fc6", size = 25335, upload-time = "2022-10-25T02:36:20.889Z" }, -] - -[[package]] -name = "cryptography" -version = "49.0.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "cffi", marker = "platform_python_implementation != 'PyPy'" }, - { name = "typing-extensions", marker = "python_full_version < '3.11'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/1f/99/d1c90d6041656cc6ee229dc99cd67fd0cd5aec3c5f7d72fffc27cc750054/cryptography-49.0.0.tar.gz", hash = "sha256:f89660a348f4f78a92366240a61404e337586ef7f5909a2fef59ca88ef505493", size = 854345, upload-time = "2026-06-12T20:02:30.512Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/9b/22/adf66990e63584a68dfb50c24f48a125c07b1699899381c8151e63ed458c/cryptography-49.0.0-cp311-abi3-macosx_11_0_arm64.whl", hash = "sha256:966fe0e9c67490071f14c0d2b1cb2dfb3023c5ce39457343931415f08382f2db", size = 4032100, upload-time = "2026-06-12T20:02:32.143Z" }, - { url = "https://files.pythonhosted.org/packages/09/41/3797cfaf69cae04a13ee78ebd83f0678d9c02b4779d21ce24445326f1a69/cryptography-49.0.0-cp311-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:36d1709f992593689b45bda411498d62c6e365f2ca00b84657d4dadd24de16db", size = 4692978, upload-time = "2026-06-12T20:01:21.305Z" }, - { url = "https://files.pythonhosted.org/packages/e6/8b/43011f7ebe515a8aa20d61f290a326cd890c2e738e16e59eaff8d9c3a412/cryptography-49.0.0-cp311-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:0e959b578856a3924bc0cbb710fc12c387b9412a951389f3ca61704a9e25f325", size = 4716422, upload-time = "2026-06-12T20:01:48.566Z" }, - { url = "https://files.pythonhosted.org/packages/4a/91/01ce7303a4579e6d3a6abef01bd322848e9ea7a219adcabc5048b9033571/cryptography-49.0.0-cp311-abi3-manylinux_2_28_aarch64.whl", hash = "sha256:53ecee2e23f7169b6117e99fc8a944e5e50f79e69758a83b52a00cb98ab2b2d2", size = 4700503, upload-time = "2026-06-12T20:02:47.091Z" }, - { url = "https://files.pythonhosted.org/packages/62/99/a2c95cf8293f07491e9e27c20cc4dcd18176d944e674679adeb1d0173fd6/cryptography-49.0.0-cp311-abi3-manylinux_2_28_ppc64le.whl", hash = "sha256:2eda353d8a27bcbcaa4cbed18994a74ab4d19a2ca897db188ea269ab9b71419b", size = 5309779, upload-time = "2026-06-12T20:02:08.987Z" }, - { url = "https://files.pythonhosted.org/packages/20/2c/0622f20ff02b2ef32558733443805dc82fd4c275be01b2d19d14676f3a1b/cryptography-49.0.0-cp311-abi3-manylinux_2_28_x86_64.whl", hash = "sha256:2afe9051da7ae7bd5905da5a949280c7d2bb75682e188f650a9d0f2756b834c6", size = 4749683, upload-time = "2026-06-12T20:02:03.335Z" }, - { url = "https://files.pythonhosted.org/packages/a3/5b/c5246635d5fd3b64e0d45ae10e99fd32fe9676a79915ccfe5a61ba9af1a5/cryptography-49.0.0-cp311-abi3-manylinux_2_31_armv7l.whl", hash = "sha256:0b82e28ee398a386f0807bba7884d30f25218855690f45115831bcce5d90822c", size = 4337874, upload-time = "2026-06-12T20:02:54.323Z" }, - { url = "https://files.pythonhosted.org/packages/6d/88/05563c7fe2e914e87d1a536d06fe83e66b4e1d95cb593e05aea375531da8/cryptography-49.0.0-cp311-abi3-manylinux_2_34_aarch64.whl", hash = "sha256:ccac2bfebc306b862133e3bb71f3f6ee8bb525240089b2d952e4144b3a6d5da7", size = 4700283, upload-time = "2026-06-12T20:01:34.822Z" }, - { url = "https://files.pythonhosted.org/packages/c4/b6/d7696e4e890d6ae1469935164c9e5215c557671cb78d6e3f458ccceaa632/cryptography-49.0.0-cp311-abi3-manylinux_2_34_ppc64le.whl", hash = "sha256:d0527ce944105f257f605a827d6ebead966c752038b6e8656abb9c5edee6fc68", size = 5265844, upload-time = "2026-06-12T20:01:24.09Z" }, - { url = "https://files.pythonhosted.org/packages/a9/3c/f3ad17eecc1a57b0ba236dc01f90e783c51f4a2f35f64777cc4f47a184b2/cryptography-49.0.0-cp311-abi3-manylinux_2_34_x86_64.whl", hash = "sha256:cbc77da8c523d5abd028635ba850a6966fcee2c82e2bf65a41d1d8afe0f98be9", size = 4749290, upload-time = "2026-06-12T20:01:30.848Z" }, - { url = "https://files.pythonhosted.org/packages/4f/01/339573cf1023163a400b0b5d16f6d507de413b9f60be6fd1b77feeaf6737/cryptography-49.0.0-cp311-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:b87e65d263b3e5d3bb92a57e2a6638e2f31110fa7aa890c7b2dbba42248d0a3f", size = 4834612, upload-time = "2026-06-12T20:01:29.246Z" }, - { url = "https://files.pythonhosted.org/packages/71/fd/577302e213a1be9468f92d1afef66fcf1ef83d516819d9992ca547f592bd/cryptography-49.0.0-cp311-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:66ec79c3904820572d7e987abdf304281f141d37ad9a489b8e97066e7b9b6459", size = 4980804, upload-time = "2026-06-12T20:01:42.853Z" }, - { url = "https://files.pythonhosted.org/packages/1f/09/f42b1d190c5ba75f72062a387f8030d1d75f6ab035788f1d9c4b01de6525/cryptography-49.0.0-cp311-abi3-win_amd64.whl", hash = "sha256:e5dfc1e64de5677cec922ffa8da89c546d0415bf6efdf081842e5d44c84e1f0e", size = 3810026, upload-time = "2026-06-12T20:02:39.262Z" }, - { url = "https://files.pythonhosted.org/packages/ec/9e/db72b3ae7fc9cfad53e630e56c6ae83b9b6ff0bf3718ffb8012d20b3aabf/cryptography-49.0.0-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:73a205dce83953d131a4aa1e0fd917a2fd1c5b1eef251e9d7152efefcbf5caf7", size = 4013892, upload-time = "2026-06-12T20:02:10.735Z" }, - { url = "https://files.pythonhosted.org/packages/86/12/c48a424f38db03027be9f7ed5c7dc5de9933dbee992865f98b13727a009d/cryptography-49.0.0-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:196ecd6a36e4e9aa10270393bb98d8df88fccee0bf1e5128b91ae4eb4375896d", size = 4678835, upload-time = "2026-06-12T20:02:48.743Z" }, - { url = "https://files.pythonhosted.org/packages/68/28/8a3ad4653662c93fc44dc4e5d8fd374c25c42e07b34bbfbadf49cf57a5a8/cryptography-49.0.0-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:7abcee80084cda3f7691f3eb1ce480d8df49cec637b429aa35986c1de71738aa", size = 4697239, upload-time = "2026-06-12T20:02:56.03Z" }, - { url = "https://files.pythonhosted.org/packages/a8/b2/2193fc74f81aee4f9b62733133b73b5176718932ed8f2e4b03fa040480a6/cryptography-49.0.0-cp314-cp314t-manylinux_2_28_aarch64.whl", hash = "sha256:4ae387c9cb68ea569ca17e490d66d8142b81c3cc814bf179974b7d146e490bbb", size = 4685593, upload-time = "2026-06-12T20:02:50.666Z" }, - { url = "https://files.pythonhosted.org/packages/47/f1/1d3eaa243bfc5de4a187b22aa8c048b3e4980bfbe830ac46e6bac2e66947/cryptography-49.0.0-cp314-cp314t-manylinux_2_28_ppc64le.whl", hash = "sha256:f37d847238971164fdbc68ade6f6574aecc9c0af714190e2083429ff68f4ce9d", size = 5289961, upload-time = "2026-06-12T20:01:46.468Z" }, - { url = "https://files.pythonhosted.org/packages/58/39/2d51306721330c486495853eda1c567880ff036de15a14c4b74f399934af/cryptography-49.0.0-cp314-cp314t-manylinux_2_28_x86_64.whl", hash = "sha256:c2bc30226390d60ea19d9f82b19db005fe0452154a23c1c410c12ea801e43561", size = 4731145, upload-time = "2026-06-12T20:02:16.832Z" }, - { url = "https://files.pythonhosted.org/packages/17/50/983e838c7fd0d87fd8c969bcdd328edaf5f756e38df5281637424c155873/cryptography-49.0.0-cp314-cp314t-manylinux_2_31_armv7l.whl", hash = "sha256:07cab27cc7b7e0fd28e5e26bb9eeedde5c135c868b46de4a27845abe94af6122", size = 4321719, upload-time = "2026-06-12T20:02:52.611Z" }, - { url = "https://files.pythonhosted.org/packages/a7/f5/8f571d7e27c55bce9f76f026143bcb1e040a4233149ecca0bea5fa5dd5f7/cryptography-49.0.0-cp314-cp314t-manylinux_2_34_aarch64.whl", hash = "sha256:b20133d204d2bb56ba047642199603876c872026ca53e79c35b83772ab2cc505", size = 4685209, upload-time = "2026-06-12T20:02:07.282Z" }, - { url = "https://files.pythonhosted.org/packages/e7/84/0e27016a6fc5a0886f797018b26aa42f40c09a82332bff77822a451deaaa/cryptography-49.0.0-cp314-cp314t-manylinux_2_34_ppc64le.whl", hash = "sha256:b970c6da94d5bb18629db453d14f2a1300f6bf59b61e9b82377931ef95504866", size = 5246285, upload-time = "2026-06-12T20:01:32.439Z" }, - { url = "https://files.pythonhosted.org/packages/11/2d/5e1fb307cb5931881516b464c98774b3f2c36b5d4bb9a2830253cf553cad/cryptography-49.0.0-cp314-cp314t-manylinux_2_34_x86_64.whl", hash = "sha256:d8ecde755e2e91bf773fc94e8c9d730cd7f2007004cb492263a794ec3899a1c8", size = 4730441, upload-time = "2026-06-12T20:02:01.469Z" }, - { url = "https://files.pythonhosted.org/packages/e4/c0/bff5a02ee731d207d6a1ed51732549d8c53d2bc8da1d10ec6f2844201d68/cryptography-49.0.0-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:e3fb64c420688e5319ae25113a354015abbd8dffbfbc41781a1ea66fc7622ac3", size = 4815869, upload-time = "2026-06-12T20:01:36.574Z" }, - { url = "https://files.pythonhosted.org/packages/b9/26/814681d14248d95d73d5c3eea0c39a94eb8302df966f670a2c60de90974b/cryptography-49.0.0-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:32703d93296f5c1f4b53349ad3a250c2cae0fdecd3a3dd5d47e616d8d616af27", size = 4960948, upload-time = "2026-06-12T20:02:18.688Z" }, - { url = "https://files.pythonhosted.org/packages/4c/fe/93ecac273d3738939d023612ad12cca9a3740a5345d69fda04134c43fd96/cryptography-49.0.0-cp314-cp314t-win_amd64.whl", hash = "sha256:33cd0565932807baddb67b96dbee92f2c374b5c89dee09fd74079aeb8c8dba61", size = 3799153, upload-time = "2026-06-12T20:01:39.059Z" }, - { url = "https://files.pythonhosted.org/packages/19/2a/5bb823f5bedcf80718cea7fbc95ec5515cca3769633c4b01a32be7f30e7c/cryptography-49.0.0-cp39-abi3-macosx_11_0_arm64.whl", hash = "sha256:ec5e529fb80935c94fe7b729f9972b50e351a0e6b50aa294fd5cabb109fcc29a", size = 4025947, upload-time = "2026-06-12T20:01:25.745Z" }, - { url = "https://files.pythonhosted.org/packages/3d/df/40577043ca124e17012f408ddddaeb213b856336ac82ddb3bc915f39e29f/cryptography-49.0.0-cp39-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:f78ff2c9ed8dc2d036b0f4d640e22522213d047c1b14e61205a7e55c80a494d4", size = 4692429, upload-time = "2026-06-12T20:01:53.628Z" }, - { url = "https://files.pythonhosted.org/packages/2c/99/2d13299eb3dd27b02dcfaafcc91d6b5cb3329f7cbd6d8f51921acd566c1a/cryptography-49.0.0-cp39-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:35b151772baff2c74cba7fa290ceaff4c3b11c0c881eb93eb5dbc05a7cfbba18", size = 4700968, upload-time = "2026-06-12T20:02:45.383Z" }, - { url = "https://files.pythonhosted.org/packages/a5/4d/9c0cd02f95e2602dd5e563da149ee0830abef3537be8b34dc56281ebe27a/cryptography-49.0.0-cp39-abi3-manylinux_2_28_aarch64.whl", hash = "sha256:0f21641cf4b30fca7aee061ced0ec7ad7b073518088b7c9969a297c0ae796c69", size = 4697758, upload-time = "2026-06-12T20:01:41.13Z" }, - { url = "https://files.pythonhosted.org/packages/24/01/186c825898477d77e2324d5360fefe622ff1d8d1963ec0554e2cada8ec77/cryptography-49.0.0-cp39-abi3-manylinux_2_28_ppc64le.whl", hash = "sha256:9e82dcc8e56052715fb18b2429e3bca4823b1629136a2084fc45a9a5cecb9b64", size = 5298863, upload-time = "2026-06-12T20:02:24.579Z" }, - { url = "https://files.pythonhosted.org/packages/b8/7b/62cbbab75d0659865bf0273790031544a0b16c8072d258f9428dcd8190dc/cryptography-49.0.0-cp39-abi3-manylinux_2_28_x86_64.whl", hash = "sha256:6f2debedf9ca60cf1d5bd466475638af5130f89965605cd818484d19987d3a21", size = 4735983, upload-time = "2026-06-12T20:01:50.14Z" }, - { url = "https://files.pythonhosted.org/packages/6c/72/3e798c064bc39e471008075d0f9bc9daf77a80879c092e4a8e170c585ed4/cryptography-49.0.0-cp39-abi3-manylinux_2_31_armv7l.whl", hash = "sha256:8c25ceb16df5b9435f3f6a9829204985b0e0cbee3b48aacd432c7d2c850b44d9", size = 4334173, upload-time = "2026-06-12T20:01:44.743Z" }, - { url = "https://files.pythonhosted.org/packages/f0/ee/6fca21d1ac73e06f8bef71940abfd4d2f6472b4bca284d770f32bd4086f6/cryptography-49.0.0-cp39-abi3-manylinux_2_34_aarch64.whl", hash = "sha256:28d8b15e6275f12c8a207dc309dfa957903c927d08d0cc937ee3f63f200693cc", size = 4697298, upload-time = "2026-06-12T20:02:20.918Z" }, - { url = "https://files.pythonhosted.org/packages/67/d0/a5fcd3515f0bae49a7b6d0413cc1bdccdcc1fc0047037a0d480642cdc5d6/cryptography-49.0.0-cp39-abi3-manylinux_2_34_ppc64le.whl", hash = "sha256:6fc361c34fb6aac015ce19435876635e5c6d21db31998b0920f675f131e043b8", size = 5254338, upload-time = "2026-06-12T20:02:22.737Z" }, - { url = "https://files.pythonhosted.org/packages/a0/84/84fe36f19caf857d61cb7fc9c63035a47ffabd84ea12d1d393148efa3615/cryptography-49.0.0-cp39-abi3-manylinux_2_34_x86_64.whl", hash = "sha256:2400ef9c9e2299a25614eb1dea3db54a69b1349efd043bfac9c67630d136df36", size = 4735650, upload-time = "2026-06-12T20:02:41.389Z" }, - { url = "https://files.pythonhosted.org/packages/6c/a0/db537264e234f7273a73ec020873d6d6b39dfd8a53db78b550ca8320440e/cryptography-49.0.0-cp39-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:67e1d20ad9ef3a563c59ef22e7a8a0b8210bd26604369ea4a30a7c66aefe504e", size = 4834820, upload-time = "2026-06-12T20:01:51.847Z" }, - { url = "https://files.pythonhosted.org/packages/93/77/8df9eb486495979bccecd1062e2eaf435250e84437040295b57d09048b0b/cryptography-49.0.0-cp39-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:42b0684e0e40cf26122427802486f6d93aea593612603a94fbf260c7eb1e9c1b", size = 4967968, upload-time = "2026-06-12T20:02:12.524Z" }, - { url = "https://files.pythonhosted.org/packages/c2/e6/f60198ea8d9dfa15fff9ed4ca02ce362f6eadd9ba757dcc50634c4257b63/cryptography-49.0.0-cp39-abi3-win_amd64.whl", hash = "sha256:026ac7423e6fa66872d3bf889be5974507da3944f866f704fa200eadacd00001", size = 3785547, upload-time = "2026-06-12T20:02:26.847Z" }, - { url = "https://files.pythonhosted.org/packages/63/d3/4a83af35d65e3fad632c926fad684c193ea4398569ccb0bbbc7fe8f5dc9a/cryptography-49.0.0-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:fc1e275c2f1d97b1a6450b8b0ea3ebfa6e087a611c2b26cb2404d48588abab7b", size = 3993685, upload-time = "2026-06-12T20:02:14.883Z" }, - { url = "https://files.pythonhosted.org/packages/d6/a7/f9dac0ab7f80368c56993a7bf638ef9935f825c91902798481fac0898138/cryptography-49.0.0-pp311-pypy311_pp73-manylinux_2_28_aarch64.whl", hash = "sha256:c83782480a4a9da4d0feb51950131ba32e12e70813848b3343f6e18c28a66838", size = 4676239, upload-time = "2026-06-12T20:02:28.793Z" }, - { url = "https://files.pythonhosted.org/packages/d7/70/2ba3769dd0ae167e2f33dfa9592d45db6ff9a61d62ca1a5b3d1bdd09068f/cryptography-49.0.0-pp311-pypy311_pp73-manylinux_2_28_x86_64.whl", hash = "sha256:b39efa323140595abd3ecca8529d321ae50f55f3aa3ba9cc81ea56a6011953d5", size = 4715584, upload-time = "2026-06-12T20:01:27.495Z" }, - { url = "https://files.pythonhosted.org/packages/94/64/2923570ac1c0bd3a737aa366ac3abbbbde273042308b8cde95e2364a6e6a/cryptography-49.0.0-pp311-pypy311_pp73-manylinux_2_34_aarch64.whl", hash = "sha256:b47db11c2c3525083296069b98ac5221907455e989ae0c2e3008bde851921615", size = 4675885, upload-time = "2026-06-12T20:01:55.49Z" }, - { url = "https://files.pythonhosted.org/packages/ab/f8/614dc7e051418cfe53d55173c1e24c6b0085e89996fe90508c2fdf769aef/cryptography-49.0.0-pp311-pypy311_pp73-manylinux_2_34_x86_64.whl", hash = "sha256:084ef1af862eb07ec46d25f68689f2102a9fc0e05ce7b80f14f5fe51e4eef0f6", size = 4715449, upload-time = "2026-06-12T20:02:05.469Z" }, - { url = "https://files.pythonhosted.org/packages/aa/50/a9caea39ad19c431c1a3f8a31114df65b260cdfe67786b6c7e7c040c4c44/cryptography-49.0.0-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:be9fcb48a55f023493482827d4f459bd263cc20efde64f204b97c123201850c6", size = 3783731, upload-time = "2026-06-12T20:02:43.319Z" }, -] - -[[package]] -name = "cuda-bindings" -version = "13.3.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "cuda-pathfinder", marker = "sys_platform != 'win32'" }, -] -wheels = [ - { url = "https://files.pythonhosted.org/packages/a9/21/8464d133752951c154feafb3b65c297e7d80f301183d220bec4c830f1441/cuda_bindings-13.3.1-cp310-cp310-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:120fcc53d57903df529c3486962c56528cba5b7d6c57c99537320ed9922c8b86", size = 6073403, upload-time = "2026-05-29T23:11:36.22Z" }, - { url = "https://files.pythonhosted.org/packages/a8/1f/5ef51f5fbaa5d4d3201bb3d7555af028ec1aa4416275ccbf73c9e34e3d2d/cuda_bindings-13.3.1-cp310-cp310-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:9851b0caa8bfd3bc6fa054eaf57bea7c8e9c3a62db2d2621224677f49f3c53d0", size = 6675244, upload-time = "2026-05-29T23:11:38.664Z" }, - { url = "https://files.pythonhosted.org/packages/51/6b/457ca12dad3ee9bfcc9a545cfd6b64b359ba49de40f776f6e028e678f262/cuda_bindings-13.3.1-cp311-cp311-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:c5879712accf6e14bb01aa5e67440eb84998b8d104b509cc7a6dc0b8f656a474", size = 6053539, upload-time = "2026-05-29T23:11:43.19Z" }, - { url = "https://files.pythonhosted.org/packages/95/7a/c5e3c34a409b148f5c0f5a4ea374158f95d488862c1dffedf9aa5c639df9/cuda_bindings-13.3.1-cp311-cp311-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:04436a9364059c84b8f9636f359eccda1cf814341f5b670c71d80d2f79dbc708", size = 6674166, upload-time = "2026-05-29T23:11:45.478Z" }, - { url = "https://files.pythonhosted.org/packages/ce/67/5e7dba1ba576dd73da5dee894ca076ca5e959450dfff66d6d510a255d1f7/cuda_bindings-13.3.1-cp312-cp312-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:c7855c4868aabc0cfae28abbe83d56734bdfbd08f08fc234ac1912a12858bf49", size = 6025351, upload-time = "2026-05-29T23:11:49.685Z" }, - { url = "https://files.pythonhosted.org/packages/39/2a/6d2e9047d1fb243dbaa364b01e0297534b9ed7fd27dba1c9f361519cf69b/cuda_bindings-13.3.1-cp312-cp312-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:e32d08f71ebcdf00f0f41eab2eb37e8da94c8ed411cc9f7f7a019ce6b34abe3a", size = 6657965, upload-time = "2026-05-29T23:11:52.227Z" }, - { url = "https://files.pythonhosted.org/packages/cc/6e/2394f8163360f8391f8f1b7e72d300a82724edb81a7b7084c799fbd4c91f/cuda_bindings-13.3.1-cp313-cp313-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:9efb21c1ee64981e184b9e0ba5eb3179e5ba3d4b51665a6cb52b8ef3d01a7cbf", size = 5920504, upload-time = "2026-05-29T23:11:56.883Z" }, - { url = "https://files.pythonhosted.org/packages/34/c2/ef9b6a63f7dc432712a462c816662e662e00d38caa9b861c8c2588195d03/cuda_bindings-13.3.1-cp313-cp313-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:2732904099e0a4d4db774a5fc6d91ee95fae065b4d2ecabb4968c5fe2406c9d7", size = 6476660, upload-time = "2026-05-29T23:11:59.188Z" }, - { url = "https://files.pythonhosted.org/packages/b1/81/bff68ce829999c1e4209c761bbf903b1c06ec570416ddb25020864ad5907/cuda_bindings-13.3.1-cp314-cp314-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:1ab2f74ed65bfef4163ba07a8db16f1085e0729291db12a2423aff84ee8278b8", size = 6013639, upload-time = "2026-05-29T23:12:03.509Z" }, - { url = "https://files.pythonhosted.org/packages/d4/e0/c8a1f0c8f9ffdea4f5fe6dbab89b326cef4d85caf489dad39e209da89416/cuda_bindings-13.3.1-cp314-cp314-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:efd4c814d311ec08c981f6dded1dbe7d4b371067ee4f6c14cccec4bde9590f80", size = 6534419, upload-time = "2026-05-29T23:12:05.633Z" }, - { url = "https://files.pythonhosted.org/packages/52/b8/83b1f563925b290f2d11a01a77a84013ba56052fe3653a5bef3ccfbb43d6/cuda_bindings-13.3.1-cp314-cp314t-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:c3c772dfff49681541d59630c90f858e173ac926b9c593a2b7123f2a1043cc76", size = 5809771, upload-time = "2026-05-29T23:12:10.422Z" }, - { url = "https://files.pythonhosted.org/packages/12/20/e79b4bfe98f075195afb6343d41c498f9dbd2d161d7021d4d28bceb83581/cuda_bindings-13.3.1-cp314-cp314t-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:36febb7c1079d68a981dbbd8d5a67235b399802b82075c9388624719607e52b9", size = 6358584, upload-time = "2026-05-29T23:12:12.767Z" }, -] - -[[package]] -name = "cuda-pathfinder" -version = "1.5.5" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/11/c8/26f2e4aae92f11522a96043892ba39a90eac610d5242523aa863212bc1c7/cuda_pathfinder-1.5.5-py3-none-any.whl", hash = "sha256:0228c023f95d1480f143ef5c8922d27a2ab052087a942e81dc289c9eb8f91689", size = 51671, upload-time = "2026-05-27T01:21:25.413Z" }, -] - -[[package]] -name = "cuda-toolkit" -version = "13.0.2" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/57/b2/453099f5f3b698d7d0eab38916aac44c7f76229f451709e2eb9db6615dcd/cuda_toolkit-13.0.2-py2.py3-none-any.whl", hash = "sha256:b198824cf2f54003f50d64ada3a0f184b42ca0846c1c94192fa269ecd97a66eb", size = 2364, upload-time = "2025-12-19T23:24:07.328Z" }, -] - -[package.optional-dependencies] -cudart = [ - { name = "nvidia-cuda-runtime", marker = "sys_platform == 'linux'" }, -] -cufft = [ - { name = "nvidia-cufft", marker = "sys_platform == 'linux'" }, -] -cufile = [ - { name = "nvidia-cufile", marker = "sys_platform == 'linux'" }, -] -cupti = [ - { name = "nvidia-cuda-cupti", marker = "sys_platform == 'linux'" }, -] -curand = [ - { name = "nvidia-curand", marker = "sys_platform == 'linux'" }, -] -cusolver = [ - { name = "nvidia-cusolver", marker = "sys_platform == 'linux'" }, -] -cusparse = [ - { name = "nvidia-cusparse", marker = "sys_platform == 'linux'" }, -] -nvjitlink = [ - { name = "nvidia-nvjitlink", marker = "sys_platform == 'linux'" }, -] -nvrtc = [ - { name = "nvidia-cuda-nvrtc", marker = "sys_platform == 'linux'" }, -] -nvtx = [ - { name = "nvidia-nvtx", marker = "sys_platform == 'linux'" }, -] - -[[package]] -name = "exceptiongroup" -version = "1.3.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "typing-extensions", marker = "python_full_version < '3.11'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/50/79/66800aadf48771f6b62f7eb014e352e5d06856655206165d775e675a02c9/exceptiongroup-1.3.1.tar.gz", hash = "sha256:8b412432c6055b0b7d14c310000ae93352ed6754f70fa8f7c34141f91c4e3219", size = 30371, upload-time = "2025-11-21T23:01:54.787Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/8a/0e/97c33bf5009bdbac74fd2beace167cab3f978feb69cc36f1ef79360d6c4e/exceptiongroup-1.3.1-py3-none-any.whl", hash = "sha256:a7a39a3bd276781e98394987d3a5701d0c4edffb633bb7a5144577f82c773598", size = 16740, upload-time = "2025-11-21T23:01:53.443Z" }, -] - -[[package]] -name = "filelock" -version = "3.29.4" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/e6/dc/be6cbe99670cd6e4ad387123647cb08e0c32975e223f82551e914c5568a6/filelock-3.29.4.tar.gz", hash = "sha256:10cdb3656fc44541cdf30652a93fb10ec6b05325620eb316bd26893e4201538a", size = 63028, upload-time = "2026-06-13T16:12:00.744Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/13/37/a065dc3bd6e49423a6532c642ca7378d3f467b1ef44c2800c937af7f9739/filelock-3.29.4-py3-none-any.whl", hash = "sha256:dac1648087d5115554850d113e7dd8c83ab2d38e3435dde2d4f163847e57b767", size = 42757, upload-time = "2026-06-13T16:11:59.582Z" }, -] - -[[package]] -name = "fsspec" -version = "2026.6.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/10/a1/ae4e3e5003468d6391d2c77b6fa1cd73bd5d13511d81c642d7b28ac90ed4/fsspec-2026.6.0.tar.gz", hash = "sha256:f5bac145310fe30e16e1471bd6840b2d990d609e872251d7e674241822abf01a", size = 313646, upload-time = "2026-06-16T01:57:28.105Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/e5/22/4222d7ddf3da30f363edaa98e329c2bce6c65497c9cb2810931c8b2c0fbc/fsspec-2026.6.0-py3-none-any.whl", hash = "sha256:02e0b71817df9b2169dc30a16832045764def1191b43dcff5bb85bdee212d2a1", size = 203949, upload-time = "2026-06-16T01:57:26.358Z" }, -] - -[[package]] -name = "h11" -version = "0.16.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/01/ee/02a2c011bdab74c6fb3c75474d40b3052059d95df7e73351460c8588d963/h11-0.16.0.tar.gz", hash = "sha256:4e35b956cf45792e4caa5885e69fba00bdbc6ffafbfa020300e549b208ee5ff1", size = 101250, upload-time = "2025-04-24T03:35:25.427Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/04/4b/29cac41a4d98d144bf5f6d33995617b185d14b22401f75ca86f384e87ff1/h11-0.16.0-py3-none-any.whl", hash = "sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86", size = 37515, upload-time = "2025-04-24T03:35:24.344Z" }, -] - -[[package]] -name = "hf-xet" -version = "1.5.1" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/4b/2d/57fd21d84d93efb4bd0b962383790e19dd1bc053501b4264c97903b4e83e/hf_xet-1.5.1.tar.gz", hash = "sha256:51ef4500dab3764b41135ee1381a4b62ce56fc54d4c92b719b59e597d6df5bf6", size = 876636, upload-time = "2026-06-08T23:02:53.897Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/64/ee/dd9ba7beae1005e54131b7d45263cc74c8a066d47d354e6d58ae9445a388/hf_xet-1.5.1-cp313-cp313t-macosx_10_12_x86_64.whl", hash = "sha256:dbf48c0d02cf0b2e568944330c60d9120c272dabe013bd892d48e25bc6797577", size = 4069485, upload-time = "2026-06-08T23:02:13.193Z" }, - { url = "https://files.pythonhosted.org/packages/b6/bc/9cae6cfeb4e03070874e73e5c97c66eb90369d3206b6a2b1ef5f96520888/hf_xet-1.5.1-cp313-cp313t-macosx_11_0_arm64.whl", hash = "sha256:e78e4e5192ad2b674c2e1160b651cb9134db974f8ae1835bdfbfb0166b894a43", size = 3838493, upload-time = "2026-06-08T23:02:15.282Z" }, - { url = "https://files.pythonhosted.org/packages/ba/b4/d5c01e0eb6d9f2ca2dacd84d0d1b71e6cfbb2ef3208c968528e010e9b3d7/hf_xet-1.5.1-cp313-cp313t-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:6f7a04a8ad962422e225bc49fbbac99dc1806764b1f3e54dbd154bffa7593947", size = 4505658, upload-time = "2026-06-08T23:02:17.196Z" }, - { url = "https://files.pythonhosted.org/packages/76/c5/29a7598c0c6383c523dc22186d577f4e04267a626cd95ae60f67c00bfe66/hf_xet-1.5.1-cp313-cp313t-manylinux_2_28_aarch64.whl", hash = "sha256:d48199c2bf4f8df0adc55d31d1368b6ec0e4d4f45bc86b08038089c23db0bed8", size = 4292822, upload-time = "2026-06-08T23:02:18.608Z" }, - { url = "https://files.pythonhosted.org/packages/04/9a/dceaf6ca69390126b86ea825fb354b93d01163199070b7bd849225de9468/hf_xet-1.5.1-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:97f212a88d14bbf573619a74b7fecb238de77d08fc702e54dec6f78276ca3283", size = 4491255, upload-time = "2026-06-08T23:02:20.124Z" }, - { url = "https://files.pythonhosted.org/packages/48/a7/e5a7afaacf6c1791fdbeeac42951fb81c3d2bc482992b115dedcc86d963e/hf_xet-1.5.1-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:f61e3665892a6c8c5e765395838b8ddf36185da835253d4bc4509a81e49fb342", size = 4711062, upload-time = "2026-06-08T23:02:21.863Z" }, - { url = "https://files.pythonhosted.org/packages/53/49/2802f8433c9742ce281bddc1e65c02c32268ca3098d66828b05e12e45ee2/hf_xet-1.5.1-cp313-cp313t-win_amd64.whl", hash = "sha256:f4ad3ebd4c32dd2b27099d69dc7b2df821e30767e46fb6ee6a0713778243b8ff", size = 4017205, upload-time = "2026-06-08T23:02:23.495Z" }, - { url = "https://files.pythonhosted.org/packages/9e/5a/50c71195b9fb883659f596e7252faf4c18c58e753a9013bdbf9bac5d2250/hf_xet-1.5.1-cp313-cp313t-win_arm64.whl", hash = "sha256:8298485c1e36e7e67cbd01eeb1376619b7af43d4f1ec245caae306f890a8a32d", size = 3845426, upload-time = "2026-06-08T23:02:25.124Z" }, - { url = "https://files.pythonhosted.org/packages/05/24/5e0c28f80371c17d49fed004597d9d132cb75c1f6f53db2cb95f459d2312/hf_xet-1.5.1-cp314-cp314t-macosx_10_12_x86_64.whl", hash = "sha256:3474760d10e3bb6f92ff3f024fcb00c0b3e4001e9b035c7483e49a5dd17aa70f", size = 4069676, upload-time = "2026-06-08T23:02:26.759Z" }, - { url = "https://files.pythonhosted.org/packages/d2/17/261ba565b6a4d960fb478f61fdf919c0be5824645aaf1c319eca660c1611/hf_xet-1.5.1-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:6762d89b9e3267dfd502b29b2a327b4525f33b17e7b509a78d94e2151a30ce30", size = 3838509, upload-time = "2026-06-08T23:02:28.573Z" }, - { url = "https://files.pythonhosted.org/packages/4e/44/7ffdc2e184b0d41fc0f683ba3936ef669ab63cf242cf36ef50e57d683668/hf_xet-1.5.1-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:bf67e6ed10260cef62e852789dc91ebb03f382d5bdc4b1dbeb64763ea275e7d6", size = 4505881, upload-time = "2026-06-08T23:02:30.257Z" }, - { url = "https://files.pythonhosted.org/packages/63/b6/788060d5aa4d5e671f1a31bf69624c314eb2d8babab3aa562f9e5d53444e/hf_xet-1.5.1-cp314-cp314t-manylinux_2_28_aarch64.whl", hash = "sha256:c6b6cd08ca095058780b50b8ce4d6cbf6787bcf27841705d58a9d32246e3e47a", size = 4292995, upload-time = "2026-06-08T23:02:31.993Z" }, - { url = "https://files.pythonhosted.org/packages/22/93/c5540cbd6b55529b7dc42f6734e88cebee21aefbea34128b66229df56c57/hf_xet-1.5.1-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:e1af0de8ca6f190d4294a28b88023db64a1e2d1d719cab044baf75bec569e7a9", size = 4491570, upload-time = "2026-06-08T23:02:33.86Z" }, - { url = "https://files.pythonhosted.org/packages/03/f3/9d8ceab30f44f36c1679b1b8683054c71a0dadc787dbf07421891742d3ca/hf_xet-1.5.1-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:4f561cbbb92f80960772059864b7fb07eae879adde1b2e781ec6f86f6ac26c59", size = 4711565, upload-time = "2026-06-08T23:02:35.454Z" }, - { url = "https://files.pythonhosted.org/packages/cd/54/27ed9a5e2cc583b4df82f75a03a4df8dbf55f5a9fa1f47f1fadfb20dbeac/hf_xet-1.5.1-cp314-cp314t-win_amd64.whl", hash = "sha256:e7dbb40617410f432182d918e37c12303fe6700fd6aa6c5964e30a535a4461d6", size = 4017343, upload-time = "2026-06-08T23:02:37.14Z" }, - { url = "https://files.pythonhosted.org/packages/ae/12/ecb2fc8d45e767580e3a37faa97cb895608b614965567efb4f18cff67e27/hf_xet-1.5.1-cp314-cp314t-win_arm64.whl", hash = "sha256:6071d5ccb4d8d2cbd5fea5cc798da4f0ba3f44e25369591c4e89a4987050e61d", size = 3845716, upload-time = "2026-06-08T23:02:39.073Z" }, - { url = "https://files.pythonhosted.org/packages/7a/d8/5e54cf37434759d1f4f2ba9b66077ff9d4c4e1f37b6bd7975da5c40d94ab/hf_xet-1.5.1-cp37-abi3-macosx_10_12_x86_64.whl", hash = "sha256:6abd35c3221eff63836618ddfb954dcf84798603f71d8e33e3ed7b04acfdbe6e", size = 4077794, upload-time = "2026-06-08T23:02:40.656Z" }, - { url = "https://files.pythonhosted.org/packages/35/94/4b2ecfbad8f8b04701a23aefb62f540b9137d058b7e1dbef16a32676f0e9/hf_xet-1.5.1-cp37-abi3-macosx_11_0_arm64.whl", hash = "sha256:94e761bbd266bf4c03cee73753916062665ce8365aa40ed321f45afcb934b41e", size = 3845354, upload-time = "2026-06-08T23:02:42.702Z" }, - { url = "https://files.pythonhosted.org/packages/de/cc/f99f4bc7295023d7bd9ebbfd51f75cc530ca262c1227666268b8208f4b77/hf_xet-1.5.1-cp37-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:892e3a3a3aecc12aded8b93cf4f9cd059282c7de0732f7d55026f3abdf474350", size = 4514864, upload-time = "2026-06-08T23:02:44.497Z" }, - { url = "https://files.pythonhosted.org/packages/cd/6e/21f7e5a2381278bd3b7b7a5a4d90038518bb6308a0c1daf5d9f8268bb178/hf_xet-1.5.1-cp37-abi3-manylinux_2_28_aarch64.whl", hash = "sha256:a93df2039190502835b1db8cd7e178b0b7b889fe9ab51299d5ced26e0dd879a4", size = 4303784, upload-time = "2026-06-08T23:02:46.203Z" }, - { url = "https://files.pythonhosted.org/packages/35/0e/f992bb6927ac1cb30ef74e62268f551f338bc32b2191f7c96a44c6f7283e/hf_xet-1.5.1-cp37-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:0c97106032ef70467b4f6bc2d0ccc266d7613ee076afc56516c502f87ce1c4a6", size = 4500703, upload-time = "2026-06-08T23:02:47.628Z" }, - { url = "https://files.pythonhosted.org/packages/fb/d1/90a498d05447980b977b1669246eeeeae4cfb0ea3e7a286eaba627f91bf9/hf_xet-1.5.1-cp37-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:6208adb15d192b90e4c2ad2a27ed864359b2cb0f2494eb6d7c7f3699ac02e2bf", size = 4719498, upload-time = "2026-06-08T23:02:49.268Z" }, - { url = "https://files.pythonhosted.org/packages/6d/b6/20f99cfe97cc663a711f7b33cc21d4793e51968e9a26125b4afcd77315ba/hf_xet-1.5.1-cp37-abi3-win_amd64.whl", hash = "sha256:f7b3002f95d1c13e24bcb4537baa8f0eb3838957067c91bb4959bc004a6435f5", size = 4026419, upload-time = "2026-06-08T23:02:50.829Z" }, - { url = "https://files.pythonhosted.org/packages/f9/fa/77453694888f03e5a8c8852d1514a0894d8e81c622d39edbaf308ea0dcf4/hf_xet-1.5.1-cp37-abi3-win_arm64.whl", hash = "sha256:93d090b57b211133f6c0dab0205ef5cb6d89162979ba75a74845045cc3063b8e", size = 3855178, upload-time = "2026-06-08T23:02:52.452Z" }, -] - -[[package]] -name = "httpcore" -version = "1.0.9" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "certifi" }, - { name = "h11" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/06/94/82699a10bca87a5556c9c59b5963f2d039dbd239f25bc2a63907a05a14cb/httpcore-1.0.9.tar.gz", hash = "sha256:6e34463af53fd2ab5d807f399a9b45ea31c3dfa2276f15a2c3f00afff6e176e8", size = 85484, upload-time = "2025-04-24T22:06:22.219Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/7e/f5/f66802a942d491edb555dd61e3a9961140fd64c90bce1eafd741609d334d/httpcore-1.0.9-py3-none-any.whl", hash = "sha256:2d400746a40668fc9dec9810239072b40b4484b640a8c38fd654a024c7a1bf55", size = 78784, upload-time = "2025-04-24T22:06:20.566Z" }, -] - -[[package]] -name = "httpx" -version = "0.28.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "anyio" }, - { name = "certifi" }, - { name = "httpcore" }, - { name = "idna" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/b1/df/48c586a5fe32a0f01324ee087459e112ebb7224f646c0b5023f5e79e9956/httpx-0.28.1.tar.gz", hash = "sha256:75e98c5f16b0f35b567856f597f06ff2270a374470a5c2392242528e3e3e42fc", size = 141406, upload-time = "2024-12-06T15:37:23.222Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/2a/39/e50c7c3a983047577ee07d2a9e53faf5a69493943ec3f6a384bdc792deb2/httpx-0.28.1-py3-none-any.whl", hash = "sha256:d909fcccc110f8c7faf814ca82a9a4d816bc5a6dbfea25d6591d6985b8ba59ad", size = 73517, upload-time = "2024-12-06T15:37:21.509Z" }, -] - -[[package]] -name = "httpx-sse" -version = "0.4.3" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/0f/4c/751061ffa58615a32c31b2d82e8482be8dd4a89154f003147acee90f2be9/httpx_sse-0.4.3.tar.gz", hash = "sha256:9b1ed0127459a66014aec3c56bebd93da3c1bc8bb6618c8082039a44889a755d", size = 15943, upload-time = "2025-10-10T21:48:22.271Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/d2/fd/6668e5aec43ab844de6fc74927e155a3b37bf40d7c3790e49fc0406b6578/httpx_sse-0.4.3-py3-none-any.whl", hash = "sha256:0ac1c9fe3c0afad2e0ebb25a934a59f4c7823b60792691f779fad2c5568830fc", size = 8960, upload-time = "2025-10-10T21:48:21.158Z" }, -] - -[[package]] -name = "huggingface-hub" -version = "1.20.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "click" }, - { name = "filelock" }, - { name = "fsspec" }, - { name = "hf-xet", marker = "platform_machine == 'AMD64' or platform_machine == 'aarch64' or platform_machine == 'amd64' or platform_machine == 'arm64' or platform_machine == 'x86_64'" }, - { name = "httpx" }, - { name = "packaging" }, - { name = "pyyaml" }, - { name = "tqdm" }, - { name = "typer" }, - { name = "typing-extensions" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/e6/7e/fad82ad491b226e832d2da90a1a59f36acd4526cda8c726f639834754aa4/huggingface_hub-1.20.1.tar.gz", hash = "sha256:9f6d63bfbeab2d2a8357200a9bc4f18cd2c8bfac9579f792f5922e77bf6471d0", size = 859910, upload-time = "2026-06-18T22:06:53.348Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/8e/b5/ff8516e74b459da3dce9567540c39f2d305ee7a2655109f6802873ff1588/huggingface_hub-1.20.1-py3-none-any.whl", hash = "sha256:274448a45c1ba6f112fe2fb168ead05574c654faa156904157a84085cfae14bd", size = 719837, upload-time = "2026-06-18T22:06:51.486Z" }, -] - -[[package]] -name = "idna" -version = "3.18" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/cd/63/9496c57188a2ee585e0f1db071d75089a11e98aa86eb99d9d7618fc1edce/idna-3.18.tar.gz", hash = "sha256:ffb385a7e039654cef1ab9ef32c6fafe283c0c0467bba1d9029738ce4a14a848", size = 196711, upload-time = "2026-06-02T14:34:07.794Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/1e/5e/d4e9f1a599fb8e573b7b87160658329fbf28d19eac2718f51fc3def3aa5a/idna-3.18-py3-none-any.whl", hash = "sha256:7f952cbe720b688055e3f87de14f5c3e5fdaa8bc3928985c4077ca689de849a2", size = 65455, upload-time = "2026-06-02T14:34:06.319Z" }, -] - -[[package]] -name = "iniconfig" -version = "2.3.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/72/34/14ca021ce8e5dfedc35312d08ba8bf51fdd999c576889fc2c24cb97f4f10/iniconfig-2.3.0.tar.gz", hash = "sha256:c76315c77db068650d49c5b56314774a7804df16fee4402c1f19d6d15d8c4730", size = 20503, upload-time = "2025-10-18T21:55:43.219Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/cb/b1/3846dd7f199d53cb17f49cba7e651e9ce294d8497c8c150530ed11865bb8/iniconfig-2.3.0-py3-none-any.whl", hash = "sha256:f631c04d2c48c52b84d0d0549c99ff3859c98df65b3101406327ecc7d53fbf12", size = 7484, upload-time = "2025-10-18T21:55:41.639Z" }, -] - -[[package]] -name = "jinja2" -version = "3.1.6" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "markupsafe" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/df/bf/f7da0350254c0ed7c72f3e33cef02e048281fec7ecec5f032d4aac52226b/jinja2-3.1.6.tar.gz", hash = "sha256:0137fb05990d35f1275a587e9aee6d56da821fc83491a0fb838183be43f66d6d", size = 245115, upload-time = "2025-03-05T20:05:02.478Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/62/a1/3d680cbfd5f4b8f15abc1d571870c5fc3e594bb582bc3b64ea099db13e56/jinja2-3.1.6-py3-none-any.whl", hash = "sha256:85ece4451f492d0c13c5dd7c13a64681a86afae63a5f347908daf103ce6d2f67", size = 134899, upload-time = "2025-03-05T20:05:00.369Z" }, -] - -[[package]] -name = "joblib" -version = "1.5.3" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/41/f2/d34e8b3a08a9cc79a50b2208a93dce981fe615b64d5a4d4abee421d898df/joblib-1.5.3.tar.gz", hash = "sha256:8561a3269e6801106863fd0d6d84bb737be9e7631e33aaed3fb9ce5953688da3", size = 331603, upload-time = "2025-12-15T08:41:46.427Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/7b/91/984aca2ec129e2757d1e4e3c81c3fcda9d0f85b74670a094cc443d9ee949/joblib-1.5.3-py3-none-any.whl", hash = "sha256:5fc3c5039fc5ca8c0276333a188bbd59d6b7ab37fe6632daa76bc7f9ec18e713", size = 309071, upload-time = "2025-12-15T08:41:44.973Z" }, -] - -[[package]] -name = "jsonschema" -version = "4.26.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "attrs" }, - { name = "jsonschema-specifications" }, - { name = "referencing" }, - { name = "rpds-py", version = "0.30.0", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, - { name = "rpds-py", version = "2026.5.1", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.11'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/b3/fc/e067678238fa451312d4c62bf6e6cf5ec56375422aee02f9cb5f909b3047/jsonschema-4.26.0.tar.gz", hash = "sha256:0c26707e2efad8aa1bfc5b7ce170f3fccc2e4918ff85989ba9ffa9facb2be326", size = 366583, upload-time = "2026-01-07T13:41:07.246Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/69/90/f63fb5873511e014207a475e2bb4e8b2e570d655b00ac19a9a0ca0a385ee/jsonschema-4.26.0-py3-none-any.whl", hash = "sha256:d489f15263b8d200f8387e64b4c3a75f06629559fb73deb8fdfb525f2dab50ce", size = 90630, upload-time = "2026-01-07T13:41:05.306Z" }, -] - -[[package]] -name = "jsonschema-specifications" -version = "2025.9.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "referencing" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/19/74/a633ee74eb36c44aa6d1095e7cc5569bebf04342ee146178e2d36600708b/jsonschema_specifications-2025.9.1.tar.gz", hash = "sha256:b540987f239e745613c7a9176f3edb72b832a4ac465cf02712288397832b5e8d", size = 32855, upload-time = "2025-09-08T01:34:59.186Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/41/45/1a4ed80516f02155c51f51e8cedb3c1902296743db0bbc66608a0db2814f/jsonschema_specifications-2025.9.1-py3-none-any.whl", hash = "sha256:98802fee3a11ee76ecaca44429fda8a41bff98b00a0f2838151b113f210cc6fe", size = 18437, upload-time = "2025-09-08T01:34:57.871Z" }, -] - -[[package]] -name = "markdown-it-py" -version = "4.2.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "mdurl" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/06/ff/7841249c247aa650a76b9ee4bbaeae59370dc8bfd2f6c01f3630c35eb134/markdown_it_py-4.2.0.tar.gz", hash = "sha256:04a21681d6fbb623de53f6f364d352309d4094dd4194040a10fd51833e418d49", size = 82454, upload-time = "2026-05-07T12:08:28.36Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/b3/81/4da04ced5a082363ecfa159c010d200ecbd959ae410c10c0264a38cac0f5/markdown_it_py-4.2.0-py3-none-any.whl", hash = "sha256:9f7ebbcd14fe59494226453aed97c1070d83f8d24b6fc3a3bcf9a38092641c4a", size = 91687, upload-time = "2026-05-07T12:08:27.182Z" }, -] - -[[package]] -name = "markupsafe" -version = "3.0.3" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/7e/99/7690b6d4034fffd95959cbe0c02de8deb3098cc577c67bb6a24fe5d7caa7/markupsafe-3.0.3.tar.gz", hash = "sha256:722695808f4b6457b320fdc131280796bdceb04ab50fe1795cd540799ebe1698", size = 80313, upload-time = "2025-09-27T18:37:40.426Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/e8/4b/3541d44f3937ba468b75da9eebcae497dcf67adb65caa16760b0a6807ebb/markupsafe-3.0.3-cp310-cp310-macosx_10_9_x86_64.whl", hash = "sha256:2f981d352f04553a7171b8e44369f2af4055f888dfb147d55e42d29e29e74559", size = 11631, upload-time = "2025-09-27T18:36:05.558Z" }, - { url = "https://files.pythonhosted.org/packages/98/1b/fbd8eed11021cabd9226c37342fa6ca4e8a98d8188a8d9b66740494960e4/markupsafe-3.0.3-cp310-cp310-macosx_11_0_arm64.whl", hash = "sha256:e1c1493fb6e50ab01d20a22826e57520f1284df32f2d8601fdd90b6304601419", size = 12057, upload-time = "2025-09-27T18:36:07.165Z" }, - { url = "https://files.pythonhosted.org/packages/40/01/e560d658dc0bb8ab762670ece35281dec7b6c1b33f5fbc09ebb57a185519/markupsafe-3.0.3-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:1ba88449deb3de88bd40044603fafffb7bc2b055d626a330323a9ed736661695", size = 22050, upload-time = "2025-09-27T18:36:08.005Z" }, - { url = "https://files.pythonhosted.org/packages/af/cd/ce6e848bbf2c32314c9b237839119c5a564a59725b53157c856e90937b7a/markupsafe-3.0.3-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:f42d0984e947b8adf7dd6dde396e720934d12c506ce84eea8476409563607591", size = 20681, upload-time = "2025-09-27T18:36:08.881Z" }, - { url = "https://files.pythonhosted.org/packages/c9/2a/b5c12c809f1c3045c4d580b035a743d12fcde53cf685dbc44660826308da/markupsafe-3.0.3-cp310-cp310-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:c0c0b3ade1c0b13b936d7970b1d37a57acde9199dc2aecc4c336773e1d86049c", size = 20705, upload-time = "2025-09-27T18:36:10.131Z" }, - { url = "https://files.pythonhosted.org/packages/cf/e3/9427a68c82728d0a88c50f890d0fc072a1484de2f3ac1ad0bfc1a7214fd5/markupsafe-3.0.3-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:0303439a41979d9e74d18ff5e2dd8c43ed6c6001fd40e5bf2e43f7bd9bbc523f", size = 21524, upload-time = "2025-09-27T18:36:11.324Z" }, - { url = "https://files.pythonhosted.org/packages/bc/36/23578f29e9e582a4d0278e009b38081dbe363c5e7165113fad546918a232/markupsafe-3.0.3-cp310-cp310-musllinux_1_2_riscv64.whl", hash = "sha256:d2ee202e79d8ed691ceebae8e0486bd9a2cd4794cec4824e1c99b6f5009502f6", size = 20282, upload-time = "2025-09-27T18:36:12.573Z" }, - { url = "https://files.pythonhosted.org/packages/56/21/dca11354e756ebd03e036bd8ad58d6d7168c80ce1fe5e75218e4945cbab7/markupsafe-3.0.3-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:177b5253b2834fe3678cb4a5f0059808258584c559193998be2601324fdeafb1", size = 20745, upload-time = "2025-09-27T18:36:13.504Z" }, - { url = "https://files.pythonhosted.org/packages/87/99/faba9369a7ad6e4d10b6a5fbf71fa2a188fe4a593b15f0963b73859a1bbd/markupsafe-3.0.3-cp310-cp310-win32.whl", hash = "sha256:2a15a08b17dd94c53a1da0438822d70ebcd13f8c3a95abe3a9ef9f11a94830aa", size = 14571, upload-time = "2025-09-27T18:36:14.779Z" }, - { url = "https://files.pythonhosted.org/packages/d6/25/55dc3ab959917602c96985cb1253efaa4ff42f71194bddeb61eb7278b8be/markupsafe-3.0.3-cp310-cp310-win_amd64.whl", hash = "sha256:c4ffb7ebf07cfe8931028e3e4c85f0357459a3f9f9490886198848f4fa002ec8", size = 15056, upload-time = "2025-09-27T18:36:16.125Z" }, - { url = "https://files.pythonhosted.org/packages/d0/9e/0a02226640c255d1da0b8d12e24ac2aa6734da68bff14c05dd53b94a0fc3/markupsafe-3.0.3-cp310-cp310-win_arm64.whl", hash = "sha256:e2103a929dfa2fcaf9bb4e7c091983a49c9ac3b19c9061b6d5427dd7d14d81a1", size = 13932, upload-time = "2025-09-27T18:36:17.311Z" }, - { url = "https://files.pythonhosted.org/packages/08/db/fefacb2136439fc8dd20e797950e749aa1f4997ed584c62cfb8ef7c2be0e/markupsafe-3.0.3-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:1cc7ea17a6824959616c525620e387f6dd30fec8cb44f649e31712db02123dad", size = 11631, upload-time = "2025-09-27T18:36:18.185Z" }, - { url = "https://files.pythonhosted.org/packages/e1/2e/5898933336b61975ce9dc04decbc0a7f2fee78c30353c5efba7f2d6ff27a/markupsafe-3.0.3-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:4bd4cd07944443f5a265608cc6aab442e4f74dff8088b0dfc8238647b8f6ae9a", size = 12058, upload-time = "2025-09-27T18:36:19.444Z" }, - { url = "https://files.pythonhosted.org/packages/1d/09/adf2df3699d87d1d8184038df46a9c80d78c0148492323f4693df54e17bb/markupsafe-3.0.3-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:6b5420a1d9450023228968e7e6a9ce57f65d148ab56d2313fcd589eee96a7a50", size = 24287, upload-time = "2025-09-27T18:36:20.768Z" }, - { url = "https://files.pythonhosted.org/packages/30/ac/0273f6fcb5f42e314c6d8cd99effae6a5354604d461b8d392b5ec9530a54/markupsafe-3.0.3-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:0bf2a864d67e76e5c9a34dc26ec616a66b9888e25e7b9460e1c76d3293bd9dbf", size = 22940, upload-time = "2025-09-27T18:36:22.249Z" }, - { url = "https://files.pythonhosted.org/packages/19/ae/31c1be199ef767124c042c6c3e904da327a2f7f0cd63a0337e1eca2967a8/markupsafe-3.0.3-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:bc51efed119bc9cfdf792cdeaa4d67e8f6fcccab66ed4bfdd6bde3e59bfcbb2f", size = 21887, upload-time = "2025-09-27T18:36:23.535Z" }, - { url = "https://files.pythonhosted.org/packages/b2/76/7edcab99d5349a4532a459e1fe64f0b0467a3365056ae550d3bcf3f79e1e/markupsafe-3.0.3-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:068f375c472b3e7acbe2d5318dea141359e6900156b5b2ba06a30b169086b91a", size = 23692, upload-time = "2025-09-27T18:36:24.823Z" }, - { url = "https://files.pythonhosted.org/packages/a4/28/6e74cdd26d7514849143d69f0bf2399f929c37dc2b31e6829fd2045b2765/markupsafe-3.0.3-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:7be7b61bb172e1ed687f1754f8e7484f1c8019780f6f6b0786e76bb01c2ae115", size = 21471, upload-time = "2025-09-27T18:36:25.95Z" }, - { url = "https://files.pythonhosted.org/packages/62/7e/a145f36a5c2945673e590850a6f8014318d5577ed7e5920a4b3448e0865d/markupsafe-3.0.3-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:f9e130248f4462aaa8e2552d547f36ddadbeaa573879158d721bbd33dfe4743a", size = 22923, upload-time = "2025-09-27T18:36:27.109Z" }, - { url = "https://files.pythonhosted.org/packages/0f/62/d9c46a7f5c9adbeeeda52f5b8d802e1094e9717705a645efc71b0913a0a8/markupsafe-3.0.3-cp311-cp311-win32.whl", hash = "sha256:0db14f5dafddbb6d9208827849fad01f1a2609380add406671a26386cdf15a19", size = 14572, upload-time = "2025-09-27T18:36:28.045Z" }, - { url = "https://files.pythonhosted.org/packages/83/8a/4414c03d3f891739326e1783338e48fb49781cc915b2e0ee052aa490d586/markupsafe-3.0.3-cp311-cp311-win_amd64.whl", hash = "sha256:de8a88e63464af587c950061a5e6a67d3632e36df62b986892331d4620a35c01", size = 15077, upload-time = "2025-09-27T18:36:29.025Z" }, - { url = "https://files.pythonhosted.org/packages/35/73/893072b42e6862f319b5207adc9ae06070f095b358655f077f69a35601f0/markupsafe-3.0.3-cp311-cp311-win_arm64.whl", hash = "sha256:3b562dd9e9ea93f13d53989d23a7e775fdfd1066c33494ff43f5418bc8c58a5c", size = 13876, upload-time = "2025-09-27T18:36:29.954Z" }, - { url = "https://files.pythonhosted.org/packages/5a/72/147da192e38635ada20e0a2e1a51cf8823d2119ce8883f7053879c2199b5/markupsafe-3.0.3-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:d53197da72cc091b024dd97249dfc7794d6a56530370992a5e1a08983ad9230e", size = 11615, upload-time = "2025-09-27T18:36:30.854Z" }, - { url = "https://files.pythonhosted.org/packages/9a/81/7e4e08678a1f98521201c3079f77db69fb552acd56067661f8c2f534a718/markupsafe-3.0.3-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:1872df69a4de6aead3491198eaf13810b565bdbeec3ae2dc8780f14458ec73ce", size = 12020, upload-time = "2025-09-27T18:36:31.971Z" }, - { url = "https://files.pythonhosted.org/packages/1e/2c/799f4742efc39633a1b54a92eec4082e4f815314869865d876824c257c1e/markupsafe-3.0.3-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:3a7e8ae81ae39e62a41ec302f972ba6ae23a5c5396c8e60113e9066ef893da0d", size = 24332, upload-time = "2025-09-27T18:36:32.813Z" }, - { url = "https://files.pythonhosted.org/packages/3c/2e/8d0c2ab90a8c1d9a24f0399058ab8519a3279d1bd4289511d74e909f060e/markupsafe-3.0.3-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:d6dd0be5b5b189d31db7cda48b91d7e0a9795f31430b7f271219ab30f1d3ac9d", size = 22947, upload-time = "2025-09-27T18:36:33.86Z" }, - { url = "https://files.pythonhosted.org/packages/2c/54/887f3092a85238093a0b2154bd629c89444f395618842e8b0c41783898ea/markupsafe-3.0.3-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:94c6f0bb423f739146aec64595853541634bde58b2135f27f61c1ffd1cd4d16a", size = 21962, upload-time = "2025-09-27T18:36:35.099Z" }, - { url = "https://files.pythonhosted.org/packages/c9/2f/336b8c7b6f4a4d95e91119dc8521402461b74a485558d8f238a68312f11c/markupsafe-3.0.3-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:be8813b57049a7dc738189df53d69395eba14fb99345e0a5994914a3864c8a4b", size = 23760, upload-time = "2025-09-27T18:36:36.001Z" }, - { url = "https://files.pythonhosted.org/packages/32/43/67935f2b7e4982ffb50a4d169b724d74b62a3964bc1a9a527f5ac4f1ee2b/markupsafe-3.0.3-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:83891d0e9fb81a825d9a6d61e3f07550ca70a076484292a70fde82c4b807286f", size = 21529, upload-time = "2025-09-27T18:36:36.906Z" }, - { url = "https://files.pythonhosted.org/packages/89/e0/4486f11e51bbba8b0c041098859e869e304d1c261e59244baa3d295d47b7/markupsafe-3.0.3-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:77f0643abe7495da77fb436f50f8dab76dbc6e5fd25d39589a0f1fe6548bfa2b", size = 23015, upload-time = "2025-09-27T18:36:37.868Z" }, - { url = "https://files.pythonhosted.org/packages/2f/e1/78ee7a023dac597a5825441ebd17170785a9dab23de95d2c7508ade94e0e/markupsafe-3.0.3-cp312-cp312-win32.whl", hash = "sha256:d88b440e37a16e651bda4c7c2b930eb586fd15ca7406cb39e211fcff3bf3017d", size = 14540, upload-time = "2025-09-27T18:36:38.761Z" }, - { url = "https://files.pythonhosted.org/packages/aa/5b/bec5aa9bbbb2c946ca2733ef9c4ca91c91b6a24580193e891b5f7dbe8e1e/markupsafe-3.0.3-cp312-cp312-win_amd64.whl", hash = "sha256:26a5784ded40c9e318cfc2bdb30fe164bdb8665ded9cd64d500a34fb42067b1c", size = 15105, upload-time = "2025-09-27T18:36:39.701Z" }, - { url = "https://files.pythonhosted.org/packages/e5/f1/216fc1bbfd74011693a4fd837e7026152e89c4bcf3e77b6692fba9923123/markupsafe-3.0.3-cp312-cp312-win_arm64.whl", hash = "sha256:35add3b638a5d900e807944a078b51922212fb3dedb01633a8defc4b01a3c85f", size = 13906, upload-time = "2025-09-27T18:36:40.689Z" }, - { url = "https://files.pythonhosted.org/packages/38/2f/907b9c7bbba283e68f20259574b13d005c121a0fa4c175f9bed27c4597ff/markupsafe-3.0.3-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:e1cf1972137e83c5d4c136c43ced9ac51d0e124706ee1c8aa8532c1287fa8795", size = 11622, upload-time = "2025-09-27T18:36:41.777Z" }, - { url = "https://files.pythonhosted.org/packages/9c/d9/5f7756922cdd676869eca1c4e3c0cd0df60ed30199ffd775e319089cb3ed/markupsafe-3.0.3-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:116bb52f642a37c115f517494ea5feb03889e04df47eeff5b130b1808ce7c219", size = 12029, upload-time = "2025-09-27T18:36:43.257Z" }, - { url = "https://files.pythonhosted.org/packages/00/07/575a68c754943058c78f30db02ee03a64b3c638586fba6a6dd56830b30a3/markupsafe-3.0.3-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:133a43e73a802c5562be9bbcd03d090aa5a1fe899db609c29e8c8d815c5f6de6", size = 24374, upload-time = "2025-09-27T18:36:44.508Z" }, - { url = "https://files.pythonhosted.org/packages/a9/21/9b05698b46f218fc0e118e1f8168395c65c8a2c750ae2bab54fc4bd4e0e8/markupsafe-3.0.3-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:ccfcd093f13f0f0b7fdd0f198b90053bf7b2f02a3927a30e63f3ccc9df56b676", size = 22980, upload-time = "2025-09-27T18:36:45.385Z" }, - { url = "https://files.pythonhosted.org/packages/7f/71/544260864f893f18b6827315b988c146b559391e6e7e8f7252839b1b846a/markupsafe-3.0.3-cp313-cp313-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:509fa21c6deb7a7a273d629cf5ec029bc209d1a51178615ddf718f5918992ab9", size = 21990, upload-time = "2025-09-27T18:36:46.916Z" }, - { url = "https://files.pythonhosted.org/packages/c2/28/b50fc2f74d1ad761af2f5dcce7492648b983d00a65b8c0e0cb457c82ebbe/markupsafe-3.0.3-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:a4afe79fb3de0b7097d81da19090f4df4f8d3a2b3adaa8764138aac2e44f3af1", size = 23784, upload-time = "2025-09-27T18:36:47.884Z" }, - { url = "https://files.pythonhosted.org/packages/ed/76/104b2aa106a208da8b17a2fb72e033a5a9d7073c68f7e508b94916ed47a9/markupsafe-3.0.3-cp313-cp313-musllinux_1_2_riscv64.whl", hash = "sha256:795e7751525cae078558e679d646ae45574b47ed6e7771863fcc079a6171a0fc", size = 21588, upload-time = "2025-09-27T18:36:48.82Z" }, - { url = "https://files.pythonhosted.org/packages/b5/99/16a5eb2d140087ebd97180d95249b00a03aa87e29cc224056274f2e45fd6/markupsafe-3.0.3-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:8485f406a96febb5140bfeca44a73e3ce5116b2501ac54fe953e488fb1d03b12", size = 23041, upload-time = "2025-09-27T18:36:49.797Z" }, - { url = "https://files.pythonhosted.org/packages/19/bc/e7140ed90c5d61d77cea142eed9f9c303f4c4806f60a1044c13e3f1471d0/markupsafe-3.0.3-cp313-cp313-win32.whl", hash = "sha256:bdd37121970bfd8be76c5fb069c7751683bdf373db1ed6c010162b2a130248ed", size = 14543, upload-time = "2025-09-27T18:36:51.584Z" }, - { url = "https://files.pythonhosted.org/packages/05/73/c4abe620b841b6b791f2edc248f556900667a5a1cf023a6646967ae98335/markupsafe-3.0.3-cp313-cp313-win_amd64.whl", hash = "sha256:9a1abfdc021a164803f4d485104931fb8f8c1efd55bc6b748d2f5774e78b62c5", size = 15113, upload-time = "2025-09-27T18:36:52.537Z" }, - { url = "https://files.pythonhosted.org/packages/f0/3a/fa34a0f7cfef23cf9500d68cb7c32dd64ffd58a12b09225fb03dd37d5b80/markupsafe-3.0.3-cp313-cp313-win_arm64.whl", hash = "sha256:7e68f88e5b8799aa49c85cd116c932a1ac15caaa3f5db09087854d218359e485", size = 13911, upload-time = "2025-09-27T18:36:53.513Z" }, - { url = "https://files.pythonhosted.org/packages/e4/d7/e05cd7efe43a88a17a37b3ae96e79a19e846f3f456fe79c57ca61356ef01/markupsafe-3.0.3-cp313-cp313t-macosx_10_13_x86_64.whl", hash = "sha256:218551f6df4868a8d527e3062d0fb968682fe92054e89978594c28e642c43a73", size = 11658, upload-time = "2025-09-27T18:36:54.819Z" }, - { url = "https://files.pythonhosted.org/packages/99/9e/e412117548182ce2148bdeacdda3bb494260c0b0184360fe0d56389b523b/markupsafe-3.0.3-cp313-cp313t-macosx_11_0_arm64.whl", hash = "sha256:3524b778fe5cfb3452a09d31e7b5adefeea8c5be1d43c4f810ba09f2ceb29d37", size = 12066, upload-time = "2025-09-27T18:36:55.714Z" }, - { url = "https://files.pythonhosted.org/packages/bc/e6/fa0ffcda717ef64a5108eaa7b4f5ed28d56122c9a6d70ab8b72f9f715c80/markupsafe-3.0.3-cp313-cp313t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:4e885a3d1efa2eadc93c894a21770e4bc67899e3543680313b09f139e149ab19", size = 25639, upload-time = "2025-09-27T18:36:56.908Z" }, - { url = "https://files.pythonhosted.org/packages/96/ec/2102e881fe9d25fc16cb4b25d5f5cde50970967ffa5dddafdb771237062d/markupsafe-3.0.3-cp313-cp313t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:8709b08f4a89aa7586de0aadc8da56180242ee0ada3999749b183aa23df95025", size = 23569, upload-time = "2025-09-27T18:36:57.913Z" }, - { url = "https://files.pythonhosted.org/packages/4b/30/6f2fce1f1f205fc9323255b216ca8a235b15860c34b6798f810f05828e32/markupsafe-3.0.3-cp313-cp313t-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:b8512a91625c9b3da6f127803b166b629725e68af71f8184ae7e7d54686a56d6", size = 23284, upload-time = "2025-09-27T18:36:58.833Z" }, - { url = "https://files.pythonhosted.org/packages/58/47/4a0ccea4ab9f5dcb6f79c0236d954acb382202721e704223a8aafa38b5c8/markupsafe-3.0.3-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:9b79b7a16f7fedff2495d684f2b59b0457c3b493778c9eed31111be64d58279f", size = 24801, upload-time = "2025-09-27T18:36:59.739Z" }, - { url = "https://files.pythonhosted.org/packages/6a/70/3780e9b72180b6fecb83a4814d84c3bf4b4ae4bf0b19c27196104149734c/markupsafe-3.0.3-cp313-cp313t-musllinux_1_2_riscv64.whl", hash = "sha256:12c63dfb4a98206f045aa9563db46507995f7ef6d83b2f68eda65c307c6829eb", size = 22769, upload-time = "2025-09-27T18:37:00.719Z" }, - { url = "https://files.pythonhosted.org/packages/98/c5/c03c7f4125180fc215220c035beac6b9cb684bc7a067c84fc69414d315f5/markupsafe-3.0.3-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:8f71bc33915be5186016f675cd83a1e08523649b0e33efdb898db577ef5bb009", size = 23642, upload-time = "2025-09-27T18:37:01.673Z" }, - { url = "https://files.pythonhosted.org/packages/80/d6/2d1b89f6ca4bff1036499b1e29a1d02d282259f3681540e16563f27ebc23/markupsafe-3.0.3-cp313-cp313t-win32.whl", hash = "sha256:69c0b73548bc525c8cb9a251cddf1931d1db4d2258e9599c28c07ef3580ef354", size = 14612, upload-time = "2025-09-27T18:37:02.639Z" }, - { url = "https://files.pythonhosted.org/packages/2b/98/e48a4bfba0a0ffcf9925fe2d69240bfaa19c6f7507b8cd09c70684a53c1e/markupsafe-3.0.3-cp313-cp313t-win_amd64.whl", hash = "sha256:1b4b79e8ebf6b55351f0d91fe80f893b4743f104bff22e90697db1590e47a218", size = 15200, upload-time = "2025-09-27T18:37:03.582Z" }, - { url = "https://files.pythonhosted.org/packages/0e/72/e3cc540f351f316e9ed0f092757459afbc595824ca724cbc5a5d4263713f/markupsafe-3.0.3-cp313-cp313t-win_arm64.whl", hash = "sha256:ad2cf8aa28b8c020ab2fc8287b0f823d0a7d8630784c31e9ee5edea20f406287", size = 13973, upload-time = "2025-09-27T18:37:04.929Z" }, - { url = "https://files.pythonhosted.org/packages/33/8a/8e42d4838cd89b7dde187011e97fe6c3af66d8c044997d2183fbd6d31352/markupsafe-3.0.3-cp314-cp314-macosx_10_13_x86_64.whl", hash = "sha256:eaa9599de571d72e2daf60164784109f19978b327a3910d3e9de8c97b5b70cfe", size = 11619, upload-time = "2025-09-27T18:37:06.342Z" }, - { url = "https://files.pythonhosted.org/packages/b5/64/7660f8a4a8e53c924d0fa05dc3a55c9cee10bbd82b11c5afb27d44b096ce/markupsafe-3.0.3-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:c47a551199eb8eb2121d4f0f15ae0f923d31350ab9280078d1e5f12b249e0026", size = 12029, upload-time = "2025-09-27T18:37:07.213Z" }, - { url = "https://files.pythonhosted.org/packages/da/ef/e648bfd021127bef5fa12e1720ffed0c6cbb8310c8d9bea7266337ff06de/markupsafe-3.0.3-cp314-cp314-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:f34c41761022dd093b4b6896d4810782ffbabe30f2d443ff5f083e0cbbb8c737", size = 24408, upload-time = "2025-09-27T18:37:09.572Z" }, - { url = "https://files.pythonhosted.org/packages/41/3c/a36c2450754618e62008bf7435ccb0f88053e07592e6028a34776213d877/markupsafe-3.0.3-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:457a69a9577064c05a97c41f4e65148652db078a3a509039e64d3467b9e7ef97", size = 23005, upload-time = "2025-09-27T18:37:10.58Z" }, - { url = "https://files.pythonhosted.org/packages/bc/20/b7fdf89a8456b099837cd1dc21974632a02a999ec9bf7ca3e490aacd98e7/markupsafe-3.0.3-cp314-cp314-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:e8afc3f2ccfa24215f8cb28dcf43f0113ac3c37c2f0f0806d8c70e4228c5cf4d", size = 22048, upload-time = "2025-09-27T18:37:11.547Z" }, - { url = "https://files.pythonhosted.org/packages/9a/a7/591f592afdc734f47db08a75793a55d7fbcc6902a723ae4cfbab61010cc5/markupsafe-3.0.3-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:ec15a59cf5af7be74194f7ab02d0f59a62bdcf1a537677ce67a2537c9b87fcda", size = 23821, upload-time = "2025-09-27T18:37:12.48Z" }, - { url = "https://files.pythonhosted.org/packages/7d/33/45b24e4f44195b26521bc6f1a82197118f74df348556594bd2262bda1038/markupsafe-3.0.3-cp314-cp314-musllinux_1_2_riscv64.whl", hash = "sha256:0eb9ff8191e8498cca014656ae6b8d61f39da5f95b488805da4bb029cccbfbaf", size = 21606, upload-time = "2025-09-27T18:37:13.485Z" }, - { url = "https://files.pythonhosted.org/packages/ff/0e/53dfaca23a69fbfbbf17a4b64072090e70717344c52eaaaa9c5ddff1e5f0/markupsafe-3.0.3-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:2713baf880df847f2bece4230d4d094280f4e67b1e813eec43b4c0e144a34ffe", size = 23043, upload-time = "2025-09-27T18:37:14.408Z" }, - { url = "https://files.pythonhosted.org/packages/46/11/f333a06fc16236d5238bfe74daccbca41459dcd8d1fa952e8fbd5dccfb70/markupsafe-3.0.3-cp314-cp314-win32.whl", hash = "sha256:729586769a26dbceff69f7a7dbbf59ab6572b99d94576a5592625d5b411576b9", size = 14747, upload-time = "2025-09-27T18:37:15.36Z" }, - { url = "https://files.pythonhosted.org/packages/28/52/182836104b33b444e400b14f797212f720cbc9ed6ba34c800639d154e821/markupsafe-3.0.3-cp314-cp314-win_amd64.whl", hash = "sha256:bdc919ead48f234740ad807933cdf545180bfbe9342c2bb451556db2ed958581", size = 15341, upload-time = "2025-09-27T18:37:16.496Z" }, - { url = "https://files.pythonhosted.org/packages/6f/18/acf23e91bd94fd7b3031558b1f013adfa21a8e407a3fdb32745538730382/markupsafe-3.0.3-cp314-cp314-win_arm64.whl", hash = "sha256:5a7d5dc5140555cf21a6fefbdbf8723f06fcd2f63ef108f2854de715e4422cb4", size = 14073, upload-time = "2025-09-27T18:37:17.476Z" }, - { url = "https://files.pythonhosted.org/packages/3c/f0/57689aa4076e1b43b15fdfa646b04653969d50cf30c32a102762be2485da/markupsafe-3.0.3-cp314-cp314t-macosx_10_13_x86_64.whl", hash = "sha256:1353ef0c1b138e1907ae78e2f6c63ff67501122006b0f9abad68fda5f4ffc6ab", size = 11661, upload-time = "2025-09-27T18:37:18.453Z" }, - { url = "https://files.pythonhosted.org/packages/89/c3/2e67a7ca217c6912985ec766c6393b636fb0c2344443ff9d91404dc4c79f/markupsafe-3.0.3-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:1085e7fbddd3be5f89cc898938f42c0b3c711fdcb37d75221de2666af647c175", size = 12069, upload-time = "2025-09-27T18:37:19.332Z" }, - { url = "https://files.pythonhosted.org/packages/f0/00/be561dce4e6ca66b15276e184ce4b8aec61fe83662cce2f7d72bd3249d28/markupsafe-3.0.3-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:1b52b4fb9df4eb9ae465f8d0c228a00624de2334f216f178a995ccdcf82c4634", size = 25670, upload-time = "2025-09-27T18:37:20.245Z" }, - { url = "https://files.pythonhosted.org/packages/50/09/c419f6f5a92e5fadde27efd190eca90f05e1261b10dbd8cbcb39cd8ea1dc/markupsafe-3.0.3-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:fed51ac40f757d41b7c48425901843666a6677e3e8eb0abcff09e4ba6e664f50", size = 23598, upload-time = "2025-09-27T18:37:21.177Z" }, - { url = "https://files.pythonhosted.org/packages/22/44/a0681611106e0b2921b3033fc19bc53323e0b50bc70cffdd19f7d679bb66/markupsafe-3.0.3-cp314-cp314t-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:f190daf01f13c72eac4efd5c430a8de82489d9cff23c364c3ea822545032993e", size = 23261, upload-time = "2025-09-27T18:37:22.167Z" }, - { url = "https://files.pythonhosted.org/packages/5f/57/1b0b3f100259dc9fffe780cfb60d4be71375510e435efec3d116b6436d43/markupsafe-3.0.3-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:e56b7d45a839a697b5eb268c82a71bd8c7f6c94d6fd50c3d577fa39a9f1409f5", size = 24835, upload-time = "2025-09-27T18:37:23.296Z" }, - { url = "https://files.pythonhosted.org/packages/26/6a/4bf6d0c97c4920f1597cc14dd720705eca0bf7c787aebc6bb4d1bead5388/markupsafe-3.0.3-cp314-cp314t-musllinux_1_2_riscv64.whl", hash = "sha256:f3e98bb3798ead92273dc0e5fd0f31ade220f59a266ffd8a4f6065e0a3ce0523", size = 22733, upload-time = "2025-09-27T18:37:24.237Z" }, - { url = "https://files.pythonhosted.org/packages/14/c7/ca723101509b518797fedc2fdf79ba57f886b4aca8a7d31857ba3ee8281f/markupsafe-3.0.3-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:5678211cb9333a6468fb8d8be0305520aa073f50d17f089b5b4b477ea6e67fdc", size = 23672, upload-time = "2025-09-27T18:37:25.271Z" }, - { url = "https://files.pythonhosted.org/packages/fb/df/5bd7a48c256faecd1d36edc13133e51397e41b73bb77e1a69deab746ebac/markupsafe-3.0.3-cp314-cp314t-win32.whl", hash = "sha256:915c04ba3851909ce68ccc2b8e2cd691618c4dc4c4232fb7982bca3f41fd8c3d", size = 14819, upload-time = "2025-09-27T18:37:26.285Z" }, - { url = "https://files.pythonhosted.org/packages/1a/8a/0402ba61a2f16038b48b39bccca271134be00c5c9f0f623208399333c448/markupsafe-3.0.3-cp314-cp314t-win_amd64.whl", hash = "sha256:4faffd047e07c38848ce017e8725090413cd80cbc23d86e55c587bf979e579c9", size = 15426, upload-time = "2025-09-27T18:37:27.316Z" }, - { url = "https://files.pythonhosted.org/packages/70/bc/6f1c2f612465f5fa89b95bead1f44dcb607670fd42891d8fdcd5d039f4f4/markupsafe-3.0.3-cp314-cp314t-win_arm64.whl", hash = "sha256:32001d6a8fc98c8cb5c947787c5d08b0a50663d139f1305bac5885d98d9b40fa", size = 14146, upload-time = "2025-09-27T18:37:28.327Z" }, -] - -[[package]] -name = "mcp" -version = "1.28.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "anyio" }, - { name = "httpx" }, - { name = "httpx-sse" }, - { name = "jsonschema" }, - { name = "pydantic" }, - { name = "pydantic-settings" }, - { name = "pyjwt", extra = ["crypto"] }, - { name = "python-multipart" }, - { name = "pywin32", marker = "sys_platform == 'win32'" }, - { name = "sse-starlette" }, - { name = "starlette" }, - { name = "typing-extensions" }, - { name = "typing-inspection" }, - { name = "uvicorn", marker = "sys_platform != 'emscripten'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/c1/ee/94c6c50ffc5b5cf4737052275d11b57367f32d1a8516e31dcd60591b3916/mcp-1.28.0.tar.gz", hash = "sha256:559d3f9943674cafbe5744c5d3794f3237e8b47f9bbc58e20c0fad680d8487c2", size = 636040, upload-time = "2026-06-16T21:37:17.996Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/2e/e1/4c1dc1fbb688641a712d34650c3d58bbbdcb314ddb75bc5817bbf33515a4/mcp-1.28.0-py3-none-any.whl", hash = "sha256:9c1e7cf3a9125557e418ecd4fed8e9adddce81b0dfdae4d6601d700f5beb71a4", size = 221959, upload-time = "2026-06-16T21:37:16.579Z" }, -] - -[[package]] -name = "mdurl" -version = "0.1.2" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/d6/54/cfe61301667036ec958cb99bd3efefba235e65cdeb9c84d24a8293ba1d90/mdurl-0.1.2.tar.gz", hash = "sha256:bb413d29f5eea38f31dd4754dd7377d4465116fb207585f97bf925588687c1ba", size = 8729, upload-time = "2022-08-14T12:40:10.846Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/b3/38/89ba8ad64ae25be8de66a6d463314cf1eb366222074cfda9ee839c56a4b4/mdurl-0.1.2-py3-none-any.whl", hash = "sha256:84008a41e51615a49fc9966191ff91509e3c40b939176e643fd50a5c2196b8f8", size = 9979, upload-time = "2022-08-14T12:40:09.779Z" }, -] - -[[package]] -name = "mindgraph" -version = "0.2.0" -source = { editable = "." } -dependencies = [ - { name = "mcp" }, - { name = "pydantic" }, - { name = "pyyaml" }, - { name = "sentence-transformers" }, - { name = "sqlite-vec" }, - { name = "typer" }, -] - -[package.optional-dependencies] -dev = [ - { name = "pytest" }, -] - -[package.metadata] -requires-dist = [ - { name = "mcp", specifier = ">=1.0.0" }, - { name = "pydantic", specifier = ">=2.0.0" }, - { name = "pytest", marker = "extra == 'dev'", specifier = ">=8.0.0" }, - { name = "pyyaml", specifier = ">=6.0.1" }, - { name = "sentence-transformers", specifier = ">=3.0.0" }, - { name = "sqlite-vec", specifier = ">=0.1.6" }, - { name = "typer", specifier = ">=0.12.0" }, -] -provides-extras = ["dev"] - -[[package]] -name = "mpmath" -version = "1.3.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/e0/47/dd32fa426cc72114383ac549964eecb20ecfd886d1e5ccf5340b55b02f57/mpmath-1.3.0.tar.gz", hash = "sha256:7a28eb2a9774d00c7bc92411c19a89209d5da7c4c9a9e227be8330a23a25b91f", size = 508106, upload-time = "2023-03-07T16:47:11.061Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/43/e3/7d92a15f894aa0c9c4b49b8ee9ac9850d6e63b03c9c32c0367a13ae62209/mpmath-1.3.0-py3-none-any.whl", hash = "sha256:a0b2b9fe80bbcd81a6647ff13108738cfb482d481d826cc0e02f5b35e5c88d2c", size = 536198, upload-time = "2023-03-07T16:47:09.197Z" }, -] - -[[package]] -name = "narwhals" -version = "2.22.1" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/62/3c/c4ef2164a71c1a63d7f1ae411c4082c5fa872405106db60a4b7114989ad7/narwhals-2.22.1.tar.gz", hash = "sha256:d62920805a0a43b7ff8b54b0c0d3142d796f8a9301836ada37e573d6a33cbcd9", size = 647493, upload-time = "2026-06-05T12:34:34.051Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/48/ca/36339329c4604adbcc99c899b7eb1ce1a555c499b6a6860757dc9bfed36d/narwhals-2.22.1-py3-none-any.whl", hash = "sha256:60567d774edf77db53906f89d9fbd164e66e56d66d388e1e6990f17ac33cfb53", size = 454815, upload-time = "2026-06-05T12:34:32.289Z" }, -] - -[[package]] -name = "networkx" -version = "3.4.2" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version < '3.11' and sys_platform == 'win32'", - "python_full_version < '3.11' and sys_platform != 'win32'", -] -sdist = { url = "https://files.pythonhosted.org/packages/fd/1d/06475e1cd5264c0b870ea2cc6fdb3e37177c1e565c43f56ff17a10e3937f/networkx-3.4.2.tar.gz", hash = "sha256:307c3669428c5362aab27c8a1260aa8f47c4e91d3891f48be0141738d8d053e1", size = 2151368, upload-time = "2024-10-21T12:39:38.695Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/b9/54/dd730b32ea14ea797530a4479b2ed46a6fb250f682a9cfb997e968bf0261/networkx-3.4.2-py3-none-any.whl", hash = "sha256:df5d4365b724cf81b8c6a7312509d0c22386097011ad1abe274afd5e9d3bbc5f", size = 1723263, upload-time = "2024-10-21T12:39:36.247Z" }, -] - -[[package]] -name = "networkx" -version = "3.6.1" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version >= '3.14' and sys_platform == 'win32'", - "python_full_version >= '3.14' and sys_platform != 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform == 'win32'", - "python_full_version == '3.11.*' and sys_platform == 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform != 'win32'", - "python_full_version == '3.11.*' and sys_platform != 'win32'", -] -sdist = { url = "https://files.pythonhosted.org/packages/6a/51/63fe664f3908c97be9d2e4f1158eb633317598cfa6e1fc14af5383f17512/networkx-3.6.1.tar.gz", hash = "sha256:26b7c357accc0c8cde558ad486283728b65b6a95d85ee1cd66bafab4c8168509", size = 2517025, upload-time = "2025-12-08T17:02:39.908Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/9e/c9/b2622292ea83fbb4ec318f5b9ab867d0a28ab43c5717bb85b0a5f6b3b0a4/networkx-3.6.1-py3-none-any.whl", hash = "sha256:d47fbf302e7d9cbbb9e2555a0d267983d2aa476bac30e90dfbe5669bd57f3762", size = 2068504, upload-time = "2025-12-08T17:02:38.159Z" }, -] - -[[package]] -name = "numpy" -version = "2.2.6" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version < '3.11' and sys_platform == 'win32'", - "python_full_version < '3.11' and sys_platform != 'win32'", -] -sdist = { url = "https://files.pythonhosted.org/packages/76/21/7d2a95e4bba9dc13d043ee156a356c0a8f0c6309dff6b21b4d71a073b8a8/numpy-2.2.6.tar.gz", hash = "sha256:e29554e2bef54a90aa5cc07da6ce955accb83f21ab5de01a62c8478897b264fd", size = 20276440, upload-time = "2025-05-17T22:38:04.611Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/9a/3e/ed6db5be21ce87955c0cbd3009f2803f59fa08df21b5df06862e2d8e2bdd/numpy-2.2.6-cp310-cp310-macosx_10_9_x86_64.whl", hash = "sha256:b412caa66f72040e6d268491a59f2c43bf03eb6c96dd8f0307829feb7fa2b6fb", size = 21165245, upload-time = "2025-05-17T21:27:58.555Z" }, - { url = "https://files.pythonhosted.org/packages/22/c2/4b9221495b2a132cc9d2eb862e21d42a009f5a60e45fc44b00118c174bff/numpy-2.2.6-cp310-cp310-macosx_11_0_arm64.whl", hash = "sha256:8e41fd67c52b86603a91c1a505ebaef50b3314de0213461c7a6e99c9a3beff90", size = 14360048, upload-time = "2025-05-17T21:28:21.406Z" }, - { url = "https://files.pythonhosted.org/packages/fd/77/dc2fcfc66943c6410e2bf598062f5959372735ffda175b39906d54f02349/numpy-2.2.6-cp310-cp310-macosx_14_0_arm64.whl", hash = "sha256:37e990a01ae6ec7fe7fa1c26c55ecb672dd98b19c3d0e1d1f326fa13cb38d163", size = 5340542, upload-time = "2025-05-17T21:28:30.931Z" }, - { url = "https://files.pythonhosted.org/packages/7a/4f/1cb5fdc353a5f5cc7feb692db9b8ec2c3d6405453f982435efc52561df58/numpy-2.2.6-cp310-cp310-macosx_14_0_x86_64.whl", hash = "sha256:5a6429d4be8ca66d889b7cf70f536a397dc45ba6faeb5f8c5427935d9592e9cf", size = 6878301, upload-time = "2025-05-17T21:28:41.613Z" }, - { url = "https://files.pythonhosted.org/packages/eb/17/96a3acd228cec142fcb8723bd3cc39c2a474f7dcf0a5d16731980bcafa95/numpy-2.2.6-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:efd28d4e9cd7d7a8d39074a4d44c63eda73401580c5c76acda2ce969e0a38e83", size = 14297320, upload-time = "2025-05-17T21:29:02.78Z" }, - { url = "https://files.pythonhosted.org/packages/b4/63/3de6a34ad7ad6646ac7d2f55ebc6ad439dbbf9c4370017c50cf403fb19b5/numpy-2.2.6-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:fc7b73d02efb0e18c000e9ad8b83480dfcd5dfd11065997ed4c6747470ae8915", size = 16801050, upload-time = "2025-05-17T21:29:27.675Z" }, - { url = "https://files.pythonhosted.org/packages/07/b6/89d837eddef52b3d0cec5c6ba0456c1bf1b9ef6a6672fc2b7873c3ec4e2e/numpy-2.2.6-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:74d4531beb257d2c3f4b261bfb0fc09e0f9ebb8842d82a7b4209415896adc680", size = 15807034, upload-time = "2025-05-17T21:29:51.102Z" }, - { url = "https://files.pythonhosted.org/packages/01/c8/dc6ae86e3c61cfec1f178e5c9f7858584049b6093f843bca541f94120920/numpy-2.2.6-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:8fc377d995680230e83241d8a96def29f204b5782f371c532579b4f20607a289", size = 18614185, upload-time = "2025-05-17T21:30:18.703Z" }, - { url = "https://files.pythonhosted.org/packages/5b/c5/0064b1b7e7c89137b471ccec1fd2282fceaae0ab3a9550f2568782d80357/numpy-2.2.6-cp310-cp310-win32.whl", hash = "sha256:b093dd74e50a8cba3e873868d9e93a85b78e0daf2e98c6797566ad8044e8363d", size = 6527149, upload-time = "2025-05-17T21:30:29.788Z" }, - { url = "https://files.pythonhosted.org/packages/a3/dd/4b822569d6b96c39d1215dbae0582fd99954dcbcf0c1a13c61783feaca3f/numpy-2.2.6-cp310-cp310-win_amd64.whl", hash = "sha256:f0fd6321b839904e15c46e0d257fdd101dd7f530fe03fd6359c1ea63738703f3", size = 12904620, upload-time = "2025-05-17T21:30:48.994Z" }, - { url = "https://files.pythonhosted.org/packages/da/a8/4f83e2aa666a9fbf56d6118faaaf5f1974d456b1823fda0a176eff722839/numpy-2.2.6-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:f9f1adb22318e121c5c69a09142811a201ef17ab257a1e66ca3025065b7f53ae", size = 21176963, upload-time = "2025-05-17T21:31:19.36Z" }, - { url = "https://files.pythonhosted.org/packages/b3/2b/64e1affc7972decb74c9e29e5649fac940514910960ba25cd9af4488b66c/numpy-2.2.6-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:c820a93b0255bc360f53eca31a0e676fd1101f673dda8da93454a12e23fc5f7a", size = 14406743, upload-time = "2025-05-17T21:31:41.087Z" }, - { url = "https://files.pythonhosted.org/packages/4a/9f/0121e375000b5e50ffdd8b25bf78d8e1a5aa4cca3f185d41265198c7b834/numpy-2.2.6-cp311-cp311-macosx_14_0_arm64.whl", hash = "sha256:3d70692235e759f260c3d837193090014aebdf026dfd167834bcba43e30c2a42", size = 5352616, upload-time = "2025-05-17T21:31:50.072Z" }, - { url = "https://files.pythonhosted.org/packages/31/0d/b48c405c91693635fbe2dcd7bc84a33a602add5f63286e024d3b6741411c/numpy-2.2.6-cp311-cp311-macosx_14_0_x86_64.whl", hash = "sha256:481b49095335f8eed42e39e8041327c05b0f6f4780488f61286ed3c01368d491", size = 6889579, upload-time = "2025-05-17T21:32:01.712Z" }, - { url = "https://files.pythonhosted.org/packages/52/b8/7f0554d49b565d0171eab6e99001846882000883998e7b7d9f0d98b1f934/numpy-2.2.6-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:b64d8d4d17135e00c8e346e0a738deb17e754230d7e0810ac5012750bbd85a5a", size = 14312005, upload-time = "2025-05-17T21:32:23.332Z" }, - { url = "https://files.pythonhosted.org/packages/b3/dd/2238b898e51bd6d389b7389ffb20d7f4c10066d80351187ec8e303a5a475/numpy-2.2.6-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:ba10f8411898fc418a521833e014a77d3ca01c15b0c6cdcce6a0d2897e6dbbdf", size = 16821570, upload-time = "2025-05-17T21:32:47.991Z" }, - { url = "https://files.pythonhosted.org/packages/83/6c/44d0325722cf644f191042bf47eedad61c1e6df2432ed65cbe28509d404e/numpy-2.2.6-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:bd48227a919f1bafbdda0583705e547892342c26fb127219d60a5c36882609d1", size = 15818548, upload-time = "2025-05-17T21:33:11.728Z" }, - { url = "https://files.pythonhosted.org/packages/ae/9d/81e8216030ce66be25279098789b665d49ff19eef08bfa8cb96d4957f422/numpy-2.2.6-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:9551a499bf125c1d4f9e250377c1ee2eddd02e01eac6644c080162c0c51778ab", size = 18620521, upload-time = "2025-05-17T21:33:39.139Z" }, - { url = "https://files.pythonhosted.org/packages/6a/fd/e19617b9530b031db51b0926eed5345ce8ddc669bb3bc0044b23e275ebe8/numpy-2.2.6-cp311-cp311-win32.whl", hash = "sha256:0678000bb9ac1475cd454c6b8c799206af8107e310843532b04d49649c717a47", size = 6525866, upload-time = "2025-05-17T21:33:50.273Z" }, - { url = "https://files.pythonhosted.org/packages/31/0a/f354fb7176b81747d870f7991dc763e157a934c717b67b58456bc63da3df/numpy-2.2.6-cp311-cp311-win_amd64.whl", hash = "sha256:e8213002e427c69c45a52bbd94163084025f533a55a59d6f9c5b820774ef3303", size = 12907455, upload-time = "2025-05-17T21:34:09.135Z" }, - { url = "https://files.pythonhosted.org/packages/82/5d/c00588b6cf18e1da539b45d3598d3557084990dcc4331960c15ee776ee41/numpy-2.2.6-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:41c5a21f4a04fa86436124d388f6ed60a9343a6f767fced1a8a71c3fbca038ff", size = 20875348, upload-time = "2025-05-17T21:34:39.648Z" }, - { url = "https://files.pythonhosted.org/packages/66/ee/560deadcdde6c2f90200450d5938f63a34b37e27ebff162810f716f6a230/numpy-2.2.6-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:de749064336d37e340f640b05f24e9e3dd678c57318c7289d222a8a2f543e90c", size = 14119362, upload-time = "2025-05-17T21:35:01.241Z" }, - { url = "https://files.pythonhosted.org/packages/3c/65/4baa99f1c53b30adf0acd9a5519078871ddde8d2339dc5a7fde80d9d87da/numpy-2.2.6-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:894b3a42502226a1cac872f840030665f33326fc3dac8e57c607905773cdcde3", size = 5084103, upload-time = "2025-05-17T21:35:10.622Z" }, - { url = "https://files.pythonhosted.org/packages/cc/89/e5a34c071a0570cc40c9a54eb472d113eea6d002e9ae12bb3a8407fb912e/numpy-2.2.6-cp312-cp312-macosx_14_0_x86_64.whl", hash = "sha256:71594f7c51a18e728451bb50cc60a3ce4e6538822731b2933209a1f3614e9282", size = 6625382, upload-time = "2025-05-17T21:35:21.414Z" }, - { url = "https://files.pythonhosted.org/packages/f8/35/8c80729f1ff76b3921d5c9487c7ac3de9b2a103b1cd05e905b3090513510/numpy-2.2.6-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:f2618db89be1b4e05f7a1a847a9c1c0abd63e63a1607d892dd54668dd92faf87", size = 14018462, upload-time = "2025-05-17T21:35:42.174Z" }, - { url = "https://files.pythonhosted.org/packages/8c/3d/1e1db36cfd41f895d266b103df00ca5b3cbe965184df824dec5c08c6b803/numpy-2.2.6-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:fd83c01228a688733f1ded5201c678f0c53ecc1006ffbc404db9f7a899ac6249", size = 16527618, upload-time = "2025-05-17T21:36:06.711Z" }, - { url = "https://files.pythonhosted.org/packages/61/c6/03ed30992602c85aa3cd95b9070a514f8b3c33e31124694438d88809ae36/numpy-2.2.6-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:37c0ca431f82cd5fa716eca9506aefcabc247fb27ba69c5062a6d3ade8cf8f49", size = 15505511, upload-time = "2025-05-17T21:36:29.965Z" }, - { url = "https://files.pythonhosted.org/packages/b7/25/5761d832a81df431e260719ec45de696414266613c9ee268394dd5ad8236/numpy-2.2.6-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:fe27749d33bb772c80dcd84ae7e8df2adc920ae8297400dabec45f0dedb3f6de", size = 18313783, upload-time = "2025-05-17T21:36:56.883Z" }, - { url = "https://files.pythonhosted.org/packages/57/0a/72d5a3527c5ebffcd47bde9162c39fae1f90138c961e5296491ce778e682/numpy-2.2.6-cp312-cp312-win32.whl", hash = "sha256:4eeaae00d789f66c7a25ac5f34b71a7035bb474e679f410e5e1a94deb24cf2d4", size = 6246506, upload-time = "2025-05-17T21:37:07.368Z" }, - { url = "https://files.pythonhosted.org/packages/36/fa/8c9210162ca1b88529ab76b41ba02d433fd54fecaf6feb70ef9f124683f1/numpy-2.2.6-cp312-cp312-win_amd64.whl", hash = "sha256:c1f9540be57940698ed329904db803cf7a402f3fc200bfe599334c9bd84a40b2", size = 12614190, upload-time = "2025-05-17T21:37:26.213Z" }, - { url = "https://files.pythonhosted.org/packages/f9/5c/6657823f4f594f72b5471f1db1ab12e26e890bb2e41897522d134d2a3e81/numpy-2.2.6-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:0811bb762109d9708cca4d0b13c4f67146e3c3b7cf8d34018c722adb2d957c84", size = 20867828, upload-time = "2025-05-17T21:37:56.699Z" }, - { url = "https://files.pythonhosted.org/packages/dc/9e/14520dc3dadf3c803473bd07e9b2bd1b69bc583cb2497b47000fed2fa92f/numpy-2.2.6-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:287cc3162b6f01463ccd86be154f284d0893d2b3ed7292439ea97eafa8170e0b", size = 14143006, upload-time = "2025-05-17T21:38:18.291Z" }, - { url = "https://files.pythonhosted.org/packages/4f/06/7e96c57d90bebdce9918412087fc22ca9851cceaf5567a45c1f404480e9e/numpy-2.2.6-cp313-cp313-macosx_14_0_arm64.whl", hash = "sha256:f1372f041402e37e5e633e586f62aa53de2eac8d98cbfb822806ce4bbefcb74d", size = 5076765, upload-time = "2025-05-17T21:38:27.319Z" }, - { url = "https://files.pythonhosted.org/packages/73/ed/63d920c23b4289fdac96ddbdd6132e9427790977d5457cd132f18e76eae0/numpy-2.2.6-cp313-cp313-macosx_14_0_x86_64.whl", hash = "sha256:55a4d33fa519660d69614a9fad433be87e5252f4b03850642f88993f7b2ca566", size = 6617736, upload-time = "2025-05-17T21:38:38.141Z" }, - { url = "https://files.pythonhosted.org/packages/85/c5/e19c8f99d83fd377ec8c7e0cf627a8049746da54afc24ef0a0cb73d5dfb5/numpy-2.2.6-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:f92729c95468a2f4f15e9bb94c432a9229d0d50de67304399627a943201baa2f", size = 14010719, upload-time = "2025-05-17T21:38:58.433Z" }, - { url = "https://files.pythonhosted.org/packages/19/49/4df9123aafa7b539317bf6d342cb6d227e49f7a35b99c287a6109b13dd93/numpy-2.2.6-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:1bc23a79bfabc5d056d106f9befb8d50c31ced2fbc70eedb8155aec74a45798f", size = 16526072, upload-time = "2025-05-17T21:39:22.638Z" }, - { url = "https://files.pythonhosted.org/packages/b2/6c/04b5f47f4f32f7c2b0e7260442a8cbcf8168b0e1a41ff1495da42f42a14f/numpy-2.2.6-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:e3143e4451880bed956e706a3220b4e5cf6172ef05fcc397f6f36a550b1dd868", size = 15503213, upload-time = "2025-05-17T21:39:45.865Z" }, - { url = "https://files.pythonhosted.org/packages/17/0a/5cd92e352c1307640d5b6fec1b2ffb06cd0dabe7d7b8227f97933d378422/numpy-2.2.6-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:b4f13750ce79751586ae2eb824ba7e1e8dba64784086c98cdbbcc6a42112ce0d", size = 18316632, upload-time = "2025-05-17T21:40:13.331Z" }, - { url = "https://files.pythonhosted.org/packages/f0/3b/5cba2b1d88760ef86596ad0f3d484b1cbff7c115ae2429678465057c5155/numpy-2.2.6-cp313-cp313-win32.whl", hash = "sha256:5beb72339d9d4fa36522fc63802f469b13cdbe4fdab4a288f0c441b74272ebfd", size = 6244532, upload-time = "2025-05-17T21:43:46.099Z" }, - { url = "https://files.pythonhosted.org/packages/cb/3b/d58c12eafcb298d4e6d0d40216866ab15f59e55d148a5658bb3132311fcf/numpy-2.2.6-cp313-cp313-win_amd64.whl", hash = "sha256:b0544343a702fa80c95ad5d3d608ea3599dd54d4632df855e4c8d24eb6ecfa1c", size = 12610885, upload-time = "2025-05-17T21:44:05.145Z" }, - { url = "https://files.pythonhosted.org/packages/6b/9e/4bf918b818e516322db999ac25d00c75788ddfd2d2ade4fa66f1f38097e1/numpy-2.2.6-cp313-cp313t-macosx_10_13_x86_64.whl", hash = "sha256:0bca768cd85ae743b2affdc762d617eddf3bcf8724435498a1e80132d04879e6", size = 20963467, upload-time = "2025-05-17T21:40:44Z" }, - { url = "https://files.pythonhosted.org/packages/61/66/d2de6b291507517ff2e438e13ff7b1e2cdbdb7cb40b3ed475377aece69f9/numpy-2.2.6-cp313-cp313t-macosx_11_0_arm64.whl", hash = "sha256:fc0c5673685c508a142ca65209b4e79ed6740a4ed6b2267dbba90f34b0b3cfda", size = 14225144, upload-time = "2025-05-17T21:41:05.695Z" }, - { url = "https://files.pythonhosted.org/packages/e4/25/480387655407ead912e28ba3a820bc69af9adf13bcbe40b299d454ec011f/numpy-2.2.6-cp313-cp313t-macosx_14_0_arm64.whl", hash = "sha256:5bd4fc3ac8926b3819797a7c0e2631eb889b4118a9898c84f585a54d475b7e40", size = 5200217, upload-time = "2025-05-17T21:41:15.903Z" }, - { url = "https://files.pythonhosted.org/packages/aa/4a/6e313b5108f53dcbf3aca0c0f3e9c92f4c10ce57a0a721851f9785872895/numpy-2.2.6-cp313-cp313t-macosx_14_0_x86_64.whl", hash = "sha256:fee4236c876c4e8369388054d02d0e9bb84821feb1a64dd59e137e6511a551f8", size = 6712014, upload-time = "2025-05-17T21:41:27.321Z" }, - { url = "https://files.pythonhosted.org/packages/b7/30/172c2d5c4be71fdf476e9de553443cf8e25feddbe185e0bd88b096915bcc/numpy-2.2.6-cp313-cp313t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:e1dda9c7e08dc141e0247a5b8f49cf05984955246a327d4c48bda16821947b2f", size = 14077935, upload-time = "2025-05-17T21:41:49.738Z" }, - { url = "https://files.pythonhosted.org/packages/12/fb/9e743f8d4e4d3c710902cf87af3512082ae3d43b945d5d16563f26ec251d/numpy-2.2.6-cp313-cp313t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:f447e6acb680fd307f40d3da4852208af94afdfab89cf850986c3ca00562f4fa", size = 16600122, upload-time = "2025-05-17T21:42:14.046Z" }, - { url = "https://files.pythonhosted.org/packages/12/75/ee20da0e58d3a66f204f38916757e01e33a9737d0b22373b3eb5a27358f9/numpy-2.2.6-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:389d771b1623ec92636b0786bc4ae56abafad4a4c513d36a55dce14bd9ce8571", size = 15586143, upload-time = "2025-05-17T21:42:37.464Z" }, - { url = "https://files.pythonhosted.org/packages/76/95/bef5b37f29fc5e739947e9ce5179ad402875633308504a52d188302319c8/numpy-2.2.6-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:8e9ace4a37db23421249ed236fdcdd457d671e25146786dfc96835cd951aa7c1", size = 18385260, upload-time = "2025-05-17T21:43:05.189Z" }, - { url = "https://files.pythonhosted.org/packages/09/04/f2f83279d287407cf36a7a8053a5abe7be3622a4363337338f2585e4afda/numpy-2.2.6-cp313-cp313t-win32.whl", hash = "sha256:038613e9fb8c72b0a41f025a7e4c3f0b7a1b5d768ece4796b674c8f3fe13efff", size = 6377225, upload-time = "2025-05-17T21:43:16.254Z" }, - { url = "https://files.pythonhosted.org/packages/67/0e/35082d13c09c02c011cf21570543d202ad929d961c02a147493cb0c2bdf5/numpy-2.2.6-cp313-cp313t-win_amd64.whl", hash = "sha256:6031dd6dfecc0cf9f668681a37648373bddd6421fff6c66ec1624eed0180ee06", size = 12771374, upload-time = "2025-05-17T21:43:35.479Z" }, - { url = "https://files.pythonhosted.org/packages/9e/3b/d94a75f4dbf1ef5d321523ecac21ef23a3cd2ac8b78ae2aac40873590229/numpy-2.2.6-pp310-pypy310_pp73-macosx_10_15_x86_64.whl", hash = "sha256:0b605b275d7bd0c640cad4e5d30fa701a8d59302e127e5f79138ad62762c3e3d", size = 21040391, upload-time = "2025-05-17T21:44:35.948Z" }, - { url = "https://files.pythonhosted.org/packages/17/f4/09b2fa1b58f0fb4f7c7963a1649c64c4d315752240377ed74d9cd878f7b5/numpy-2.2.6-pp310-pypy310_pp73-macosx_14_0_x86_64.whl", hash = "sha256:7befc596a7dc9da8a337f79802ee8adb30a552a94f792b9c9d18c840055907db", size = 6786754, upload-time = "2025-05-17T21:44:47.446Z" }, - { url = "https://files.pythonhosted.org/packages/af/30/feba75f143bdc868a1cc3f44ccfa6c4b9ec522b36458e738cd00f67b573f/numpy-2.2.6-pp310-pypy310_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:ce47521a4754c8f4593837384bd3424880629f718d87c5d44f8ed763edd63543", size = 16643476, upload-time = "2025-05-17T21:45:11.871Z" }, - { url = "https://files.pythonhosted.org/packages/37/48/ac2a9584402fb6c0cd5b5d1a91dcf176b15760130dd386bbafdbfe3640bf/numpy-2.2.6-pp310-pypy310_pp73-win_amd64.whl", hash = "sha256:d042d24c90c41b54fd506da306759e06e568864df8ec17ccc17e9e884634fd00", size = 12812666, upload-time = "2025-05-17T21:45:31.426Z" }, -] - -[[package]] -name = "numpy" -version = "2.4.6" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version >= '3.14' and sys_platform == 'win32'", - "python_full_version >= '3.14' and sys_platform != 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform == 'win32'", - "python_full_version == '3.11.*' and sys_platform == 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform != 'win32'", - "python_full_version == '3.11.*' and sys_platform != 'win32'", -] -sdist = { url = "https://files.pythonhosted.org/packages/d0/ad/fed0499ce6a338d2a03ebae59cd15093910c8875328855781952abf6c2fe/numpy-2.4.6.tar.gz", hash = "sha256:f3a3570c4a2a16746ac2c31a7c7c7b0c186b95ce902e33db6f28094ed7387dda", size = 20735807, upload-time = "2026-05-18T23:37:14.07Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/b3/49/ec46835a70be8fa6446c495126ac84fdb28cb2558e1620ffb87a10c8b64c/numpy-2.4.6-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:0280e0356c0829a18d9de1cb7eee50ec22ca639878d7240307ca0943d73cd2c4", size = 16969194, upload-time = "2026-05-18T23:33:13.503Z" }, - { url = "https://files.pythonhosted.org/packages/0e/0d/f5957185c0ee2f3e12f78715aa9e3b353fd83633316c8532b38faa37e3f6/numpy-2.4.6-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:110f8b71aacb688ec69062bb7f6938a0f8acb01b7c1c4beb453c65b6d234584d", size = 14964111, upload-time = "2026-05-18T23:33:17.795Z" }, - { url = "https://files.pythonhosted.org/packages/ad/40/40a40ee0ddf7ceb782c49af278894b686e586d65d8c1889c8b5da01a3d7d/numpy-2.4.6-cp311-cp311-macosx_14_0_arm64.whl", hash = "sha256:4cfe66903cc32a9921a6733d96b19bb6abf310397581bbad89c228f5abaf0ee8", size = 5469159, upload-time = "2026-05-18T23:33:20.654Z" }, - { url = "https://files.pythonhosted.org/packages/63/13/f9a8046535cb21deae82f8d03de9617e08882d274fad2539630761888228/numpy-2.4.6-cp311-cp311-macosx_14_0_x86_64.whl", hash = "sha256:8155154c7c691289fe18f510b5d4657c68c67989f293f0535a91360392ff6538", size = 6798936, upload-time = "2026-05-18T23:33:22.987Z" }, - { url = "https://files.pythonhosted.org/packages/33/a8/6fa8c1a345a8c85dbb21932c447bee07c30a2c2a3f31e369c0a84b300147/numpy-2.4.6-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:0ab0a9c4ffb1a6d95ef519fe4247dba8eb6b18ad93999f76b7f657039acabd47", size = 15966692, upload-time = "2026-05-18T23:33:26.62Z" }, - { url = "https://files.pythonhosted.org/packages/02/03/74fe2a4cb3817d94d86402f2506554130a2f01414e299b5a843e5a8a957f/numpy-2.4.6-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:89cd468399cfd2504718f0ba50e410dca55a170b61a02ad92bb18c8a65186e93", size = 16918164, upload-time = "2026-05-18T23:33:29.955Z" }, - { url = "https://files.pythonhosted.org/packages/c5/80/3615be3313f7e7696609bc194b9f0101da809df79e859bdb84e0cd043f46/numpy-2.4.6-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:c2d37ab77531417474168eb79d6d80b14f821a966818505d03013d0833edb7a8", size = 17322877, upload-time = "2026-05-18T23:33:34.724Z" }, - { url = "https://files.pythonhosted.org/packages/ca/ac/a691e0fe2675e370d0e08ff905adc49a1c8830e8cae03efe4477e92cd55d/numpy-2.4.6-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:f407cb6b8e9d6d8c626bc73c945db1706035af8fd632295547bf1c9e46d092d6", size = 18651487, upload-time = "2026-05-18T23:33:38.217Z" }, - { url = "https://files.pythonhosted.org/packages/15/a7/9bc1cd626d7bf6869bfedf27b91b6ab5dd607758bf8e959d6fa80c6a59cb/numpy-2.4.6-cp311-cp311-win32.whl", hash = "sha256:ddea102b48f9e339f3948bf22040944184627a30fdf7f858667673b9c5f033c8", size = 6233945, upload-time = "2026-05-18T23:33:41.331Z" }, - { url = "https://files.pythonhosted.org/packages/c5/31/7fc6239c12bce7e931463251cca4426c465e1876ba3cc785402ef4dd8f4e/numpy-2.4.6-cp311-cp311-win_amd64.whl", hash = "sha256:1e254a00cdf42b1e4d5b3d68d33af63268d41340d8885df2ab6470f2e1500147", size = 12608406, upload-time = "2026-05-18T23:33:44.131Z" }, - { url = "https://files.pythonhosted.org/packages/27/83/140f85a466595a16382996a1bf06b2b54bcd597488921b0c9daaeeda72af/numpy-2.4.6-cp311-cp311-win_arm64.whl", hash = "sha256:ed9749eef4cbd126da3dc1d6bcb3a57f5eb7ac6a6484146bdbf743f552dfc577", size = 10479528, upload-time = "2026-05-18T23:33:50.725Z" }, - { url = "https://files.pythonhosted.org/packages/95/2a/3d7b5ac8aac24feaf9ad7ed58f45b0bbc06d37e4338ae84c9f2298b570f9/numpy-2.4.6-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:001fbb8e08d942dd57599e781f2472269ee7f2755fae407b4f67b2f0b17da3f1", size = 16689119, upload-time = "2026-05-18T23:33:54.065Z" }, - { url = "https://files.pythonhosted.org/packages/ea/12/92c4c131527599e8288d6918e888d88726f84d805d784b771f32408aeaef/numpy-2.4.6-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:ebfb099f8dcf083deef3ac1ca4c1503f387cf76296fcb3816b66f5ecb5f54fdb", size = 14699246, upload-time = "2026-05-18T23:33:57.621Z" }, - { url = "https://files.pythonhosted.org/packages/ad/fe/c0a6b7b2ca128a8fb228575147073b660656734b8ebe4d76c8fd748dcc79/numpy-2.4.6-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:3213d622a0283a39a93d188f3cf72b26862df52fbb4ca3697f51705016523d41", size = 5204410, upload-time = "2026-05-18T23:34:00.302Z" }, - { url = "https://files.pythonhosted.org/packages/f3/d4/9770d14ba719432bb90a421bfd443872ed0f70f7264b64bec12ea363d5fd/numpy-2.4.6-cp312-cp312-macosx_14_0_x86_64.whl", hash = "sha256:357cc07a6d7b0b182ff02249616a03742827ebb1277546b5c7cd7f7620a45698", size = 6551240, upload-time = "2026-05-18T23:34:02.852Z" }, - { url = "https://files.pythonhosted.org/packages/c9/c6/50a46a6205feba2343f1d6d17438107c5dc491ed1c736e6ea68689fd906b/numpy-2.4.6-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5f9fb9157b4ce2971008323afe46053787b526ef624fea915b261468a8421a0f", size = 15671012, upload-time = "2026-05-18T23:34:05.485Z" }, - { url = "https://files.pythonhosted.org/packages/99/60/14115e6364fa676c5397c2ad3004e527e9aa487abf5d0706ec81bbd08529/numpy-2.4.6-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:90f9849678c75fe7afa2d348ac842c168b0a4d3d61919687216dfc547976d853", size = 16645538, upload-time = "2026-05-18T23:34:09.265Z" }, - { url = "https://files.pythonhosted.org/packages/ae/c5/693cbe59e57db94d2231fa519ca3978dc9e19da5a8f088588f5c6e947ff2/numpy-2.4.6-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:c1a2af6c6ef86344a6b0db6b97834208bf598db514f2b155042439b62605601a", size = 17020706, upload-time = "2026-05-18T23:34:13.053Z" }, - { url = "https://files.pythonhosted.org/packages/ef/fc/85b7c4eff9b4966ade25c2273cf7e7012e92366c032058653934b37de044/numpy-2.4.6-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:e5805d5a22fd19c8ccff10a9561f9df94436b0545619ea579db2d3c35294bce2", size = 18368541, upload-time = "2026-05-18T23:34:17.024Z" }, - { url = "https://files.pythonhosted.org/packages/f6/81/e1b27545deedce7f4a0b348618c6b62d74e36a4dc9ccd42f3eb2f85eee32/numpy-2.4.6-cp312-cp312-win32.whl", hash = "sha256:e3eeb0aabd6bd5ce64faae67e9935203a6991b4bc2a485a767fbafb2c5125f45", size = 5962825, upload-time = "2026-05-18T23:34:20.3Z" }, - { url = "https://files.pythonhosted.org/packages/ab/ca/feab00bd44aa5fe1ad2c18f08b4d3bb92e26484b0b1d1443897809ed528c/numpy-2.4.6-cp312-cp312-win_amd64.whl", hash = "sha256:d8e8286dd7cea7895157318d1b91cdacac64c479f3cbc8dce548331728484751", size = 12321687, upload-time = "2026-05-18T23:34:23.095Z" }, - { url = "https://files.pythonhosted.org/packages/63/cf/5a6d34850a39d1093558564f77ee8e8e0bee5061151b8f05a55711001ec7/numpy-2.4.6-cp312-cp312-win_arm64.whl", hash = "sha256:4081eb135ac24158bd51cdfbef16f1c64df7063b1143f24731387137c092bec8", size = 10221482, upload-time = "2026-05-18T23:34:25.876Z" }, - { url = "https://files.pythonhosted.org/packages/fb/82/bdab26d7438c6791ca31b7c024ca37c1eab8b726ba236129005cd4a06e45/numpy-2.4.6-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:511dbaf848decaaaf4b4ca48032619fb3138710c4bf7da7617765edad1ef96b0", size = 16684648, upload-time = "2026-05-18T23:34:29.41Z" }, - { url = "https://files.pythonhosted.org/packages/1b/30/a80189bcc7f5e4258b3fbc3968d909d1756f54d023299ecc39ad6fdb9ef8/numpy-2.4.6-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:bf162abab1c1a736333192707cef898e735a5ca00f38f27eeedf44b39d9e85eb", size = 14693902, upload-time = "2026-05-18T23:34:33.013Z" }, - { url = "https://files.pythonhosted.org/packages/97/12/70b5d0d7c15e1ebb8a6a84a8caa1d19e181d84fb58bb6d70aca29099dec1/numpy-2.4.6-cp313-cp313-macosx_14_0_arm64.whl", hash = "sha256:043191bfa8eab18c776647b62723ac9dddece59743b13f49b2016094129c2b3f", size = 5198992, upload-time = "2026-05-18T23:34:36.132Z" }, - { url = "https://files.pythonhosted.org/packages/ba/8c/ebd2a8f8a83541f8d38cc5667e8c2b69cecfd30da6e45693e8158857d44b/numpy-2.4.6-cp313-cp313-macosx_14_0_x86_64.whl", hash = "sha256:6180d8b35af935aed8ece3a85e0a43f87393ae0ac87c8d2c8bd2c993f7270ef3", size = 6546944, upload-time = "2026-05-18T23:34:38.484Z" }, - { url = "https://files.pythonhosted.org/packages/bb/c5/7b863a97a91671a0338f4253bd3b5a3d3852f0692dae91711c9f4a10e787/numpy-2.4.6-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:72fbe16c6fac95aedf5937fa873445cec2110be35d8a4e9433d7501fd98dae6b", size = 15669392, upload-time = "2026-05-18T23:34:41.257Z" }, - { url = "https://files.pythonhosted.org/packages/a5/9d/3584b9984ca4c047aea75214ce1a4c4c73d849bd71b604264b7f5653f8a8/numpy-2.4.6-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:a7830bab239b79cda9c08c2da014761cafb48da6150e1da17ac06283f43b6089", size = 16633220, upload-time = "2026-05-18T23:34:45.075Z" }, - { url = "https://files.pythonhosted.org/packages/05/ae/7c67fba23bd98caec7c99261f3a16072ade14813486b0282cb29846de832/numpy-2.4.6-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:ef4aea96ce4d3b074422cb4f2f64e216bf9e213004bb58ecfdf50ea02ea8eb9a", size = 17020800, upload-time = "2026-05-18T23:34:49.065Z" }, - { url = "https://files.pythonhosted.org/packages/d9/5d/3b6725cb31d983c5e66916f5d36f6d7e5521129e4c4404d64f918292a5b6/numpy-2.4.6-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:dfa20cc6ca228e6b155b11da03825975ce66aea520985dbbddf0f2a5a495c605", size = 18357600, upload-time = "2026-05-18T23:34:52.709Z" }, - { url = "https://files.pythonhosted.org/packages/f7/da/2ccc6c2fe8898dee01d90c75c5f5f914a23daf99e3e0f59516a08760c8b5/numpy-2.4.6-cp313-cp313-win32.whl", hash = "sha256:56b39e5e0622a09a25bf5baf62f4bcf0cb8a41ae6e2819cf49bbc5a74c083f91", size = 5961134, upload-time = "2026-05-18T23:34:55.618Z" }, - { url = "https://files.pythonhosted.org/packages/b5/cd/9cc4dc876fb065d5c220aae4d5e14826b2715331bb7618ce1fb07a679d99/numpy-2.4.6-cp313-cp313-win_amd64.whl", hash = "sha256:c4fc99836233ea196540b17ab0983aff60ed07941751930f5f4d05bc3b3b7359", size = 12318598, upload-time = "2026-05-18T23:34:58.928Z" }, - { url = "https://files.pythonhosted.org/packages/39/1e/c0bcba1f8694116485fe28fd1be698c278fcda4141c5b0e53a2aed8b12a8/numpy-2.4.6-cp313-cp313-win_arm64.whl", hash = "sha256:a7c711e21628b52034bb5ab8d1bce291f752fcc5e92accc615778acee1ff4778", size = 10222272, upload-time = "2026-05-18T23:35:02.167Z" }, - { url = "https://files.pythonhosted.org/packages/63/6d/cc5619247c8f4204e507f5883528372e4ac4bb189e579fb859a12e480b1f/numpy-2.4.6-cp313-cp313t-macosx_11_0_arm64.whl", hash = "sha256:112b06a867b235ef466ed3508ddf0238050df9c727cafb5301ac385b899189a1", size = 14821197, upload-time = "2026-05-18T23:35:05.468Z" }, - { url = "https://files.pythonhosted.org/packages/00/58/f1c39161c87d9e9bed660f1ed4bafc0e403d5ec9650b6dd77aead07d489b/numpy-2.4.6-cp313-cp313t-macosx_14_0_arm64.whl", hash = "sha256:eaf7fa2de5c0be8ae6ff8e9bea2ccd725e980541244521d8d4b5f3354a27babe", size = 5326287, upload-time = "2026-05-18T23:35:08.693Z" }, - { url = "https://files.pythonhosted.org/packages/af/57/3917ab0fd97f271a8694513581b8a36c655f111c446852c302f04ccdb6fc/numpy-2.4.6-cp313-cp313t-macosx_14_0_x86_64.whl", hash = "sha256:7265a2f3d436e54ef9f2b52b5c937e6be778781bd97a590319d7348f1c1ca997", size = 6646763, upload-time = "2026-05-18T23:35:11.459Z" }, - { url = "https://files.pythonhosted.org/packages/eb/0f/037e64c494b67581ae18193d770adef354c41f3f2c8ebf865602d949bf8f/numpy-2.4.6-cp313-cp313t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:f74a575920ab21fe304421a3fc28793d82e299cae9eccb37084e9fc7f3617c20", size = 15728070, upload-time = "2026-05-18T23:35:14.79Z" }, - { url = "https://files.pythonhosted.org/packages/21/a6/5d2bae9c9542eb4df16dc9c46dc79c186e9bad53805dfa5399a6023c6db0/numpy-2.4.6-cp313-cp313t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:ede83e07a75dd06bc501566c1eca2afc0d61677c1472ac9ad93fdee6e638a48d", size = 16681752, upload-time = "2026-05-18T23:35:18.836Z" }, - { url = "https://files.pythonhosted.org/packages/92/14/23d1dfb410ae362cd59ce53e936b1513d545eb40db3949ced632e19a459e/numpy-2.4.6-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:68bb27509ac1b9a3443094260f6326150663b06abe40b73a2f81160623da5b67", size = 17086024, upload-time = "2026-05-18T23:35:22.52Z" }, - { url = "https://files.pythonhosted.org/packages/4b/6e/23595a2c642cdf3bc567877064bdd7f91c8b0038a4453cf2daf7248eafe9/numpy-2.4.6-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:a0df0043bdb289bde1f62da130d20df23d58b45429f752bc7a8fc5325a225ecd", size = 18403398, upload-time = "2026-05-18T23:35:26.398Z" }, - { url = "https://files.pythonhosted.org/packages/8a/90/0ac3bc947217e66dec77e7cbc6a1979d1af70b6461b82f620d3bccd5e4c8/numpy-2.4.6-cp313-cp313t-win32.whl", hash = "sha256:29a287e0cf63ff528da061de6b9f64a4618da591ca1046aafc54062e40ca7eab", size = 6084971, upload-time = "2026-05-18T23:35:29.387Z" }, - { url = "https://files.pythonhosted.org/packages/77/71/5673e351671a1d2bd6063b91b44f70c0affea7d1516fa7a6572941ba4aa1/numpy-2.4.6-cp313-cp313t-win_amd64.whl", hash = "sha256:25c692919ac5a01f170a3bfcd62d745b24fd095c353d50812637d6fcab442e75", size = 12458532, upload-time = "2026-05-18T23:35:32.175Z" }, - { url = "https://files.pythonhosted.org/packages/3f/88/19d3503c5046e688f049274b27a3ef3d771152fa80d3ba3d01a3dff61abe/numpy-2.4.6-cp313-cp313t-win_arm64.whl", hash = "sha256:1e978ec1e8bd0e0e4de6bb75de9d30cbb74db6b6a2bb727618613703ca0167dd", size = 10291881, upload-time = "2026-05-18T23:35:35.465Z" }, - { url = "https://files.pythonhosted.org/packages/f8/91/3ab2044d05fd16d343c5ac2e69b127f1b2854040dd20b193257c78028bd3/numpy-2.4.6-cp314-cp314-macosx_10_15_x86_64.whl", hash = "sha256:06ca2f61ec4385a07a6977c55ba998a4466c123642b4a32694d3128fce18c079", size = 16683458, upload-time = "2026-05-18T23:35:38.353Z" }, - { url = "https://files.pythonhosted.org/packages/8e/62/764ce66fa4147ae6d73071a3abf804ffe606f174618697c571acdf26a7c9/numpy-2.4.6-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:38efbc8de75c7a0fc1ac190162d892787f3f47b57cc291231aafee36b80982b7", size = 14704559, upload-time = "2026-05-18T23:35:42.14Z" }, - { url = "https://files.pythonhosted.org/packages/60/61/23f27c172f022e04025b7dc2367f4d63c1a398120607ec896228649a6f48/numpy-2.4.6-cp314-cp314-macosx_14_0_arm64.whl", hash = "sha256:d581b735e177fdcdce6fed8e7e8880a3fb6ee4e3653a3ac6af01c6f4c03effc5", size = 5209716, upload-time = "2026-05-18T23:35:45.377Z" }, - { url = "https://files.pythonhosted.org/packages/03/71/21cf70dc6ea3e3acb95fc53a265b2fc248b981f0194ceb5b475271b8809d/numpy-2.4.6-cp314-cp314-macosx_14_0_x86_64.whl", hash = "sha256:0a041d3d761dc3c35cc56ce0351506a02bcbc25f7b169f652435141a17db9096", size = 6543947, upload-time = "2026-05-18T23:35:47.926Z" }, - { url = "https://files.pythonhosted.org/packages/d5/91/64288395ee1799bd2e0b04a305dce9666da90c961e1f3fe982a05ee1c036/numpy-2.4.6-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:40fdc1ae7125e518ea98e53e69a4ebc27e1fd50510c47b7ea130cf21e5e1d42b", size = 15685197, upload-time = "2026-05-18T23:35:50.863Z" }, - { url = "https://files.pythonhosted.org/packages/f3/eb/ebffaa97dc55502df69584a8f0dcf07f69a3e0b3e2323670a2722db9aa39/numpy-2.4.6-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:a2c306dea656c12c68f51f4cea133cbe78ca7435eb28c735eac1d3ebe73be6e8", size = 16638245, upload-time = "2026-05-18T23:35:54.752Z" }, - { url = "https://files.pythonhosted.org/packages/b8/0b/54f9da33128d7e350fab89c7455902eeae70349ee52bddb448dc4a576f45/numpy-2.4.6-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:33111801a01c12a8a1e3721f0a9232f8cfc8ae2c6b7098167e6f623c6073f402", size = 17036587, upload-time = "2026-05-18T23:35:58.355Z" }, - { url = "https://files.pythonhosted.org/packages/b6/f0/fdebc1052db1cc37c64beb22072d67cd6d1c71adca1299f53dec2b5e20d3/numpy-2.4.6-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:ae506e6902902557576a26ff33eda8695e7ecb3cb36c3b573a0765dee114ebdb", size = 18363226, upload-time = "2026-05-18T23:36:02.845Z" }, - { url = "https://files.pythonhosted.org/packages/aa/b4/298628d98c72b57e57f7165ae6a481a1deaf6f3c28262a6e4c739c275930/numpy-2.4.6-cp314-cp314-win32.whl", hash = "sha256:aaf159caa35993cb1f56fb9b8e4610d35758e7ca005412eb1daa856a78c9c4b1", size = 6010196, upload-time = "2026-05-18T23:36:05.92Z" }, - { url = "https://files.pythonhosted.org/packages/df/ac/46de6dda46478f7942f839e094970be2d4a861e005c4b3bf07c92e291a09/numpy-2.4.6-cp314-cp314-win_amd64.whl", hash = "sha256:b507f5c4c1d508876d1819b6bf9a49d365b96320b5d4993426b33a23ca4b8261", size = 12450334, upload-time = "2026-05-18T23:36:09.107Z" }, - { url = "https://files.pythonhosted.org/packages/78/92/b8b798ac784102c0da830d2257d59358e3d3d90d1e2b3f2575dad976c5cf/numpy-2.4.6-cp314-cp314-win_arm64.whl", hash = "sha256:6f41ae150c4e32db4f3310cdaf64b1593a03dbabe29eec77fc9b50fe64061df6", size = 10495678, upload-time = "2026-05-18T23:36:12.766Z" }, - { url = "https://files.pythonhosted.org/packages/30/34/ec28d1aa8115971537c01469ab2011ee96827930f0a124de1000cc2a7ed7/numpy-2.4.6-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:ece3d2cfe132e7d51f44a832b303895e6f2d499c5e74dfbdb06ee246147a304a", size = 14823672, upload-time = "2026-05-18T23:36:16.473Z" }, - { url = "https://files.pythonhosted.org/packages/16/bd/f6d1fede4e54e8042a7ff97bb495510f3c220f94bcd9e8b228e87c92cc0d/numpy-2.4.6-cp314-cp314t-macosx_14_0_arm64.whl", hash = "sha256:e3e5193ef5a3dc73bceee50f7fdc2c90dbb76c42df8d8fae3d1067a583df579e", size = 5328731, upload-time = "2026-05-18T23:36:19.767Z" }, - { url = "https://files.pythonhosted.org/packages/f4/f0/e105b9e2fd728a9910103884decd6951d9dd73896b914a98d9a231de02ee/numpy-2.4.6-cp314-cp314t-macosx_14_0_x86_64.whl", hash = "sha256:17f9ade344e7d9b464a084d69bcf18fc691cb1db67c62ed80820bf4926d78f0e", size = 6649805, upload-time = "2026-05-18T23:36:22.266Z" }, - { url = "https://files.pythonhosted.org/packages/82/dd/1206a7ca6ab15e3f02069707ca96222e202af681bb73756da7527f3cb837/numpy-2.4.6-cp314-cp314t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:9cd5ffd25db4e7ba6a375693b3fc0fc1791ec636c17db3720da19bde7180ec43", size = 15730496, upload-time = "2026-05-18T23:36:25.713Z" }, - { url = "https://files.pythonhosted.org/packages/51/e7/38d3ea825dcab85a591734decb2f6c67caa7c8367d374df1a1c3842f9b07/numpy-2.4.6-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:7d92c3819208a60205a12a245c91ad70cb0a85336659b19b834205573ac8456e", size = 16679616, upload-time = "2026-05-18T23:36:29.652Z" }, - { url = "https://files.pythonhosted.org/packages/93/b7/caabfdf53edf663e0b4eb74d7d405d83baef09eb5e83bcd32d601d72b93e/numpy-2.4.6-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:e85b752a1e912b70eaad4fafbd4d1238007ab221de2009b9a2f5ae7461239895", size = 17085145, upload-time = "2026-05-18T23:36:33.449Z" }, - { url = "https://files.pythonhosted.org/packages/f9/45/68d7c33a6bcf3e5aa3bdbd57a367e6f615286dfd6482f97e8ffeb734306e/numpy-2.4.6-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:29cb7f67d10b479ff07c17d33e39f78c07f71c40ef30d63c153d340e96cd3fb4", size = 18403813, upload-time = "2026-05-18T23:36:37.369Z" }, - { url = "https://files.pythonhosted.org/packages/9c/50/0753655aa844c99cd9e018aacf76f130f1bd81d881bb74bc0aef5d73a8ba/numpy-2.4.6-cp314-cp314t-win32.whl", hash = "sha256:260a5d70215b61ab4fadf5c7baacd64821842975eea312125ed3c39a6391b063", size = 6156982, upload-time = "2026-05-18T23:36:40.817Z" }, - { url = "https://files.pythonhosted.org/packages/b2/d4/7c67becf668f973cb490cec3e98dfd799d866f9c989a54d355672cfa0db6/numpy-2.4.6-cp314-cp314t-win_amd64.whl", hash = "sha256:81a1cca95ed5bb92aa8b10dd2cdc9a0d3853a50fad926c28b5d7e8ea54389627", size = 12638908, upload-time = "2026-05-18T23:36:43.996Z" }, - { url = "https://files.pythonhosted.org/packages/43/bb/e1c71a4295b1b1d1393d50dbb4f2a36283c6859d9d3892e84f00ec5a91d5/numpy-2.4.6-cp314-cp314t-win_arm64.whl", hash = "sha256:0c9136e14ed34a9e343a31c533d78a9813a69a3148332bce5e9821cb2f996e66", size = 10565867, upload-time = "2026-05-18T23:36:47.114Z" }, - { url = "https://files.pythonhosted.org/packages/de/12/b422cc84439adc0d00de605bf4a308890ae5c26f2c71fbd73e5d08fbb0dd/numpy-2.4.6-pp311-pypy311_pp73-macosx_10_15_x86_64.whl", hash = "sha256:55cced7c52e981362f708ad635198e97a752dfba412cc03c23bbf3bd8d5cd662", size = 16847511, upload-time = "2026-05-18T23:36:50.673Z" }, - { url = "https://files.pythonhosted.org/packages/44/53/f481bef68011740f8849418d82db07230e825013f31f4eef5ba5b805316a/numpy-2.4.6-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:d6da64deb6b8ed903e7560180a92f2d804ee1ba5eeb849ac2748b8c1aba1f6d7", size = 14889064, upload-time = "2026-05-18T23:36:53.879Z" }, - { url = "https://files.pythonhosted.org/packages/7f/57/42ed575c10ced8af951d426bc4e1f8aff16fd851db33f067036215a7f860/numpy-2.4.6-pp311-pypy311_pp73-macosx_14_0_arm64.whl", hash = "sha256:68a5124b13fa6cc2086764a20005d30bc0548146f7f5322f02fce212ca14317f", size = 5394157, upload-time = "2026-05-18T23:36:57.194Z" }, - { url = "https://files.pythonhosted.org/packages/6a/ef/f66cc724fcc36c1e364c67f51ae9146090b8b584f27d58b97fdae3edd737/numpy-2.4.6-pp311-pypy311_pp73-macosx_14_0_x86_64.whl", hash = "sha256:948424b06129ce883307e8cff868c31396d8dc7630a59c61d70d98dbe70f222c", size = 6708728, upload-time = "2026-05-18T23:36:59.575Z" }, - { url = "https://files.pythonhosted.org/packages/1a/9c/c531f2293b91265d8b48e9b329f54fdd7ffae73cb4134ea10cca4237e9cc/numpy-2.4.6-pp311-pypy311_pp73-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5dbbdb29840ca3d91ee0fece42fc29278886d908280bfec0a5846c6f901a3eb0", size = 15798374, upload-time = "2026-05-18T23:37:02.674Z" }, - { url = "https://files.pythonhosted.org/packages/1a/b0/413077f6b1153ed3cba361401c6783bbad6114804a000cc22eb71c13e190/numpy-2.4.6-pp311-pypy311_pp73-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:8ad03c0965fb3c692200e74d458ca28c1dbb4ce96f9a479a8aa041ad5fabca02", size = 16747286, upload-time = "2026-05-18T23:37:06.327Z" }, - { url = "https://files.pythonhosted.org/packages/15/ce/e5ec180bc41812edcd8daeb8639d205622c0e8c02259d8ab25a0201b3c2a/numpy-2.4.6-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:2803abfebfc990042cd494d8ce2d5f82e9d847af6d35ec486923aa19dbad5e73", size = 12504263, upload-time = "2026-05-18T23:37:09.715Z" }, -] - -[[package]] -name = "nvidia-cublas" -version = "13.1.1.3" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "nvidia-cuda-nvrtc", marker = "sys_platform != 'win32'" }, -] -wheels = [ - { url = "https://files.pythonhosted.org/packages/a7/a1/0bd24ee8c8d03adac032fd2909426a00c88f8c57961b1277ded97f91119f/nvidia_cublas-13.1.1.3-py3-none-manylinux_2_27_aarch64.whl", hash = "sha256:b7a210458267ac818974c53038fbec2e969d5c99f305ab15c72522fa9f001dd5", size = 542848918, upload-time = "2026-04-08T18:46:22.985Z" }, - { url = "https://files.pythonhosted.org/packages/3b/cd/154ca20c38269e05eff77c1464e6c1da89f50a6390b565e9d82e06bc11e1/nvidia_cublas-13.1.1.3-py3-none-manylinux_2_27_x86_64.whl", hash = "sha256:37936a16db8fe4ac1f065c2139360608a543a09275cb1a1af612e08cfa065436", size = 423138758, upload-time = "2026-04-08T18:46:58.655Z" }, -] - -[[package]] -name = "nvidia-cuda-cupti" -version = "13.0.85" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/2a/2a/80353b103fc20ce05ef51e928daed4b6015db4aaa9162ed0997090fe2250/nvidia_cuda_cupti-13.0.85-py3-none-manylinux_2_25_aarch64.whl", hash = "sha256:796bd679890ee55fb14a94629b698b6db54bcfd833d391d5e94017dd9d7d3151", size = 10310827, upload-time = "2025-09-04T08:26:42.012Z" }, - { url = "https://files.pythonhosted.org/packages/33/6d/737d164b4837a9bbd202f5ae3078975f0525a55730fe871d8ed4e3b952b0/nvidia_cuda_cupti-13.0.85-py3-none-manylinux_2_25_x86_64.whl", hash = "sha256:4eb01c08e859bf924d222250d2e8f8b8ff6d3db4721288cf35d14252a4d933c8", size = 10715597, upload-time = "2025-09-04T08:26:51.312Z" }, -] - -[[package]] -name = "nvidia-cuda-nvrtc" -version = "13.0.88" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/c3/68/483a78f5e8f31b08fb1bb671559968c0ca3a065ac7acabfc7cee55214fd6/nvidia_cuda_nvrtc-13.0.88-py3-none-manylinux2010_x86_64.manylinux_2_12_x86_64.whl", hash = "sha256:ad9b6d2ead2435f11cbb6868809d2adeeee302e9bb94bcf0539c7a40d80e8575", size = 90215200, upload-time = "2025-09-04T08:28:44.204Z" }, - { url = "https://files.pythonhosted.org/packages/b7/dc/6bb80850e0b7edd6588d560758f17e0550893a1feaf436807d64d2da040f/nvidia_cuda_nvrtc-13.0.88-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:d27f20a0ca67a4bb34268a5e951033496c5b74870b868bacd046b1b8e0c3267b", size = 43015449, upload-time = "2025-09-04T08:28:20.239Z" }, -] - -[[package]] -name = "nvidia-cuda-runtime" -version = "13.0.96" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/87/4f/17d7b9b8e285199c58ce28e31b5c5bbaa4d8271af06a89b6405258245de2/nvidia_cuda_runtime-13.0.96-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:ef9bcbe90493a2b9d810e43d249adb3d02e98dd30200d86607d8d02687c43f55", size = 2261060, upload-time = "2025-10-09T08:55:15.78Z" }, - { url = "https://files.pythonhosted.org/packages/2e/24/d1558f3b68b1d26e706813b1d10aa1d785e4698c425af8db8edc3dced472/nvidia_cuda_runtime-13.0.96-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:7f82250d7782aa23b6cfe765ecc7db554bd3c2870c43f3d1821f1d18aebf0548", size = 2243632, upload-time = "2025-10-09T08:55:36.117Z" }, -] - -[[package]] -name = "nvidia-cudnn-cu13" -version = "9.20.0.48" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "nvidia-cublas", marker = "sys_platform != 'win32'" }, -] -wheels = [ - { url = "https://files.pythonhosted.org/packages/56/c5/83384d846b2fd17c44bd499b36c75a45ed4f095fbbb2252294e89cea5c5c/nvidia_cudnn_cu13-9.20.0.48-py3-none-manylinux_2_27_aarch64.whl", hash = "sha256:e31454ae00094b0c55319d9d15b6fa2fc50a9e1c0f5c8c80fb75258234e731e1", size = 444574296, upload-time = "2026-03-09T19:28:27.751Z" }, - { url = "https://files.pythonhosted.org/packages/6e/5e/edb9c0ae051602c3ccaffe424256463636d639e27d7f302dde9975ef9e7a/nvidia_cudnn_cu13-9.20.0.48-py3-none-manylinux_2_27_x86_64.whl", hash = "sha256:0c45dd8eeb50b603f07995b1b300c62ffe6a1980482b82b3bcf94a4ca9d49304", size = 366173588, upload-time = "2026-03-09T19:29:34.474Z" }, -] - -[[package]] -name = "nvidia-cufft" -version = "12.0.0.61" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "nvidia-nvjitlink", marker = "sys_platform != 'win32'" }, -] -wheels = [ - { url = "https://files.pythonhosted.org/packages/8b/ae/f417a75c0259e85c1d2f83ca4e960289a5f814ed0cea74d18c353d3e989d/nvidia_cufft-12.0.0.61-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:2708c852ef8cd89d1d2068bdbece0aa188813a0c934db3779b9b1faa8442e5f5", size = 214053554, upload-time = "2025-09-04T08:31:38.196Z" }, - { url = "https://files.pythonhosted.org/packages/a8/2f/7b57e29836ea8714f81e9898409196f47d772d5ddedddf1592eadb8ab743/nvidia_cufft-12.0.0.61-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:6c44f692dce8fd5ffd3e3df134b6cdb9c2f72d99cf40b62c32dde45eea9ddad3", size = 214085489, upload-time = "2025-09-04T08:31:56.044Z" }, -] - -[[package]] -name = "nvidia-cufile" -version = "1.15.1.6" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/3f/70/4f193de89a48b71714e74602ee14d04e4019ad36a5a9f20c425776e72cd6/nvidia_cufile-1.15.1.6-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:08a3ecefae5a01c7f5117351c64f17c7c62efa5fffdbe24fc7d298da19cd0b44", size = 1223672, upload-time = "2025-09-04T08:32:22.779Z" }, - { url = "https://files.pythonhosted.org/packages/ab/73/cc4a14c9813a8a0d509417cf5f4bdaba76e924d58beb9864f5a7baceefbf/nvidia_cufile-1.15.1.6-py3-none-manylinux_2_27_aarch64.whl", hash = "sha256:bdc0deedc61f548bddf7733bdc216456c2fdb101d020e1ab4b88d232d5e2f6d1", size = 1136992, upload-time = "2025-09-04T08:32:14.119Z" }, -] - -[[package]] -name = "nvidia-curand" -version = "10.4.0.35" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/1e/72/7c2ae24fb6b63a32e6ae5d241cc65263ea18d08802aaae087d9f013335a2/nvidia_curand-10.4.0.35-py3-none-manylinux_2_27_aarch64.whl", hash = "sha256:133df5a7509c3e292aaa2b477afd0194f06ce4ea24d714d616ff36439cee349a", size = 61962106, upload-time = "2025-08-04T10:21:41.128Z" }, - { url = "https://files.pythonhosted.org/packages/a5/9f/be0a41ca4a4917abf5cb9ae0daff1a6060cc5de950aec0396de9f3b52bc5/nvidia_curand-10.4.0.35-py3-none-manylinux_2_27_x86_64.whl", hash = "sha256:1aee33a5da6e1db083fe2b90082def8915f30f3248d5896bcec36a579d941bfc", size = 59544258, upload-time = "2025-08-04T10:22:03.992Z" }, -] - -[[package]] -name = "nvidia-cusolver" -version = "12.0.4.66" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "nvidia-cublas", marker = "sys_platform != 'win32'" }, - { name = "nvidia-cusparse", marker = "sys_platform != 'win32'" }, - { name = "nvidia-nvjitlink", marker = "sys_platform != 'win32'" }, -] -wheels = [ - { url = "https://files.pythonhosted.org/packages/c8/c3/b30c9e935fc01e3da443ec0116ed1b2a009bb867f5324d3f2d7e533e776b/nvidia_cusolver-12.0.4.66-py3-none-manylinux_2_27_aarch64.whl", hash = "sha256:02c2457eaa9e39de20f880f4bd8820e6a1cfb9f9a34f820eb12a155aa5bc92d2", size = 223467760, upload-time = "2025-09-04T08:33:04.222Z" }, - { url = "https://files.pythonhosted.org/packages/5f/67/cba3777620cdacb99102da4042883709c41c709f4b6323c10781a9c3aa34/nvidia_cusolver-12.0.4.66-py3-none-manylinux_2_27_x86_64.whl", hash = "sha256:0a759da5dea5c0ea10fd307de75cdeb59e7ea4fcb8add0924859b944babf1112", size = 200941980, upload-time = "2025-09-04T08:33:22.767Z" }, -] - -[[package]] -name = "nvidia-cusparse" -version = "12.6.3.3" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "nvidia-nvjitlink", marker = "sys_platform != 'win32'" }, -] -wheels = [ - { url = "https://files.pythonhosted.org/packages/f8/94/5c26f33738ae35276672f12615a64bd008ed5be6d1ebcb23579285d960a9/nvidia_cusparse-12.6.3.3-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:80bcc4662f23f1054ee334a15c72b8940402975e0eab63178fc7e670aa59472c", size = 162155568, upload-time = "2025-09-04T08:33:42.864Z" }, - { url = "https://files.pythonhosted.org/packages/fa/18/623c77619c31d62efd55302939756966f3ecc8d724a14dab2b75f1508850/nvidia_cusparse-12.6.3.3-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:2b3c89c88d01ee0e477cb7f82ef60a11a4bcd57b6b87c33f789350b59759360b", size = 145942937, upload-time = "2025-09-04T08:33:58.029Z" }, -] - -[[package]] -name = "nvidia-cusparselt-cu13" -version = "0.8.1" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/46/e1/cdc1797eadf82d3a9a575a19b33fdc871a97edbec42c00b5b5e914f4aff4/nvidia_cusparselt_cu13-0.8.1-py3-none-manylinux2014_aarch64.whl", hash = "sha256:4dca476c50bf4780d46cd0bfbd82e2bc10a08e4fef7950917ce8d7578d22a23f", size = 221051344, upload-time = "2025-09-05T18:49:51.289Z" }, - { url = "https://files.pythonhosted.org/packages/34/7d/2661f2fb3ac4302f3a246f5fc030213ac60c1fe0bce84f9783dbd831dbb7/nvidia_cusparselt_cu13-0.8.1-py3-none-manylinux2014_x86_64.whl", hash = "sha256:786ce87568c303fadb5afcc7102d454cd3040d75f6f8626f5db460d1871f4dd0", size = 170148586, upload-time = "2025-09-05T18:50:50.248Z" }, -] - -[[package]] -name = "nvidia-nccl-cu13" -version = "2.29.7" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/72/0d/daf50d44177ee0cbc7ff0a0c91eb5ff676c82be42f9a970bc7597f440c3a/nvidia_nccl_cu13-2.29.7-py3-none-manylinux_2_18_aarch64.whl", hash = "sha256:674a12383e3c38a1bcccae7d4f3633b37852230b6047883cb2f4c2d1b36d9bf5", size = 206014712, upload-time = "2026-03-03T05:34:20.843Z" }, - { url = "https://files.pythonhosted.org/packages/67/f4/58e4e91b6919367c7aafb8e36fce9aad1a3047e536bf7e2fd560927d3a4c/nvidia_nccl_cu13-2.29.7-py3-none-manylinux_2_18_x86_64.whl", hash = "sha256:edd81538446786ec3b73972543e53bb43bcaf0bfc8ef76cb679fcc390ffe136d", size = 205976000, upload-time = "2026-03-03T05:36:24.472Z" }, -] - -[[package]] -name = "nvidia-nvjitlink" -version = "13.0.88" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/56/7a/123e033aaff487c77107195fa5a2b8686795ca537935a24efae476c41f05/nvidia_nvjitlink-13.0.88-py3-none-manylinux2010_x86_64.manylinux_2_12_x86_64.whl", hash = "sha256:13a74f429e23b921c1109976abefacc69835f2f433ebd323d3946e11d804e47b", size = 40713933, upload-time = "2025-09-04T08:35:43.553Z" }, - { url = "https://files.pythonhosted.org/packages/ab/2c/93c5250e64df4f894f1cbb397c6fd71f79813f9fd79d7cd61de3f97b3c2d/nvidia_nvjitlink-13.0.88-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:e931536ccc7d467a98ba1d8b89ff7fa7f1fa3b13f2b0069118cd7f47bff07d0c", size = 38768748, upload-time = "2025-09-04T08:35:20.008Z" }, -] - -[[package]] -name = "nvidia-nvshmem-cu13" -version = "3.4.5" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/dc/0f/05cc9c720236dcd2db9c1ab97fff629e96821be2e63103569da0c9b72f19/nvidia_nvshmem_cu13-3.4.5-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:6dc2a197f38e5d0376ad52cd1a2a3617d3cdc150fd5966f4aee9bcebb1d68fe9", size = 60215947, upload-time = "2025-09-06T00:32:20.022Z" }, - { url = "https://files.pythonhosted.org/packages/3c/35/a9bf80a609e74e3b000fef598933235c908fcefcef9026042b8e6dfde2a9/nvidia_nvshmem_cu13-3.4.5-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:290f0a2ee94c9f3687a02502f3b9299a9f9fe826e6d0287ee18482e78d495b80", size = 60412546, upload-time = "2025-09-06T00:32:41.564Z" }, -] - -[[package]] -name = "nvidia-nvtx" -version = "13.0.85" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/c2/f3/d86c845465a2723ad7e1e5c36dcd75ddb82898b3f53be47ebd429fb2fa5d/nvidia_nvtx-13.0.85-py3-none-manylinux1_x86_64.manylinux_2_5_x86_64.whl", hash = "sha256:4936d1d6780fbe68db454f5e72a42ff64d1fd6397df9f363ae786930fd5c1cd4", size = 148047, upload-time = "2025-09-04T08:29:01.761Z" }, - { url = "https://files.pythonhosted.org/packages/a8/64/3708a90d1ebe202ffdeb7185f878a3c84d15c2b2c31858da2ce0583e2def/nvidia_nvtx-13.0.85-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:cb7780edb6b14107373c835bf8b72e7a178bac7367e23da7acb108f973f157a6", size = 148878, upload-time = "2025-09-04T08:28:53.627Z" }, -] - -[[package]] -name = "packaging" -version = "26.2" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/d7/f1/e7a6dd94a8d4a5626c03e4e99c87f241ba9e350cd9e6d75123f992427270/packaging-26.2.tar.gz", hash = "sha256:ff452ff5a3e828ce110190feff1178bb1f2ea2281fa2075aadb987c2fb221661", size = 228134, upload-time = "2026-04-24T20:15:23.917Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/df/b2/87e62e8c3e2f4b32e5fe99e0b86d576da1312593b39f47d8ceef365e95ed/packaging-26.2-py3-none-any.whl", hash = "sha256:5fc45236b9446107ff2415ce77c807cee2862cb6fac22b8a73826d0693b0980e", size = 100195, upload-time = "2026-04-24T20:15:22.081Z" }, -] - -[[package]] -name = "pluggy" -version = "1.6.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/f9/e2/3e91f31a7d2b083fe6ef3fa267035b518369d9511ffab804f839851d2779/pluggy-1.6.0.tar.gz", hash = "sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3", size = 69412, upload-time = "2025-05-15T12:30:07.975Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/54/20/4d324d65cc6d9205fabedc306948156824eb9f0ee1633355a8f7ec5c66bf/pluggy-1.6.0-py3-none-any.whl", hash = "sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746", size = 20538, upload-time = "2025-05-15T12:30:06.134Z" }, -] - -[[package]] -name = "pycparser" -version = "3.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/1b/7d/92392ff7815c21062bea51aa7b87d45576f649f16458d78b7cf94b9ab2e6/pycparser-3.0.tar.gz", hash = "sha256:600f49d217304a5902ac3c37e1281c9fe94e4d0489de643a9504c5cdfdfc6b29", size = 103492, upload-time = "2026-01-21T14:26:51.89Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/0c/c3/44f3fbbfa403ea2a7c779186dc20772604442dde72947e7d01069cbe98e3/pycparser-3.0-py3-none-any.whl", hash = "sha256:b727414169a36b7d524c1c3e31839a521725078d7b2ff038656844266160a992", size = 48172, upload-time = "2026-01-21T14:26:50.693Z" }, -] - -[[package]] -name = "pydantic" -version = "2.13.4" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "annotated-types" }, - { name = "pydantic-core" }, - { name = "typing-extensions" }, - { name = "typing-inspection" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/18/a5/b60d21ac674192f8ab0ba4e9fd860690f9b4a6e51ca5df118733b487d8d6/pydantic-2.13.4.tar.gz", hash = "sha256:c40756b57adaa8b1efeeced5c196f3f3b7c435f90e84ea7f443901bec8099ef6", size = 844775, upload-time = "2026-05-06T13:43:05.343Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/fd/7b/122376b1fd3c62c1ed9dc80c931ace4844b3c55407b6fb2d199377c9736f/pydantic-2.13.4-py3-none-any.whl", hash = "sha256:45a282cde31d808236fd7ea9d919b128653c8b38b393d1c4ab335c62924d9aba", size = 472262, upload-time = "2026-05-06T13:43:02.641Z" }, -] - -[[package]] -name = "pydantic-core" -version = "2.46.4" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "typing-extensions" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/9d/56/921726b776ace8d8f5db44c4ef961006580d91dc52b803c489fafd1aa249/pydantic_core-2.46.4.tar.gz", hash = "sha256:62f875393d7f270851f20523dd2e29f082bcc82292d66db2b64ea71f64b6e1c1", size = 471464, upload-time = "2026-05-06T13:37:06.98Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/e7/08/f1ba952f1c8ae5581c70fa9c6da89f247b83e3dd8c09c035d5d7931fc23d/pydantic_core-2.46.4-cp310-cp310-macosx_10_12_x86_64.whl", hash = "sha256:a396dcc17e5a0b164dbe026896245a4fa9ff402edca1dff0be3d53a517f74de4", size = 2113146, upload-time = "2026-05-06T13:37:36.537Z" }, - { url = "https://files.pythonhosted.org/packages/56/c6/65f646c7ff09bd257f660434adb45c4dfcbbcebcc030562fecf6f5bf887d/pydantic_core-2.46.4-cp310-cp310-macosx_11_0_arm64.whl", hash = "sha256:da4b951fe36dc7c3a1ccb4e3cd1747c3542b8c9ceede8fc86cae054e764485f5", size = 1949769, upload-time = "2026-05-06T13:37:46.365Z" }, - { url = "https://files.pythonhosted.org/packages/64/ba/bfb1d928fd5b49e1258935ff104ae356e9fd89384a55bf9f847e9193ad40/pydantic_core-2.46.4-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:bb63e0198ca18aad131c089b9204c23079c3afa95487e561f4c522d519e55aba", size = 1974958, upload-time = "2026-05-06T13:37:28.611Z" }, - { url = "https://files.pythonhosted.org/packages/4e/74/76223bfb117b64af743c9b6670d1364516f5c0604f96b48f3272f6af6cc6/pydantic_core-2.46.4-cp310-cp310-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:f47286a97f0bc9b8859519809077b91b2cefe4ae47fcbf5e466a009c1c5d742b", size = 2042118, upload-time = "2026-05-06T13:36:55.216Z" }, - { url = "https://files.pythonhosted.org/packages/cb/7b/848732968bc8f48f3187542f08358b9d842db564147b256669426ebb1652/pydantic_core-2.46.4-cp310-cp310-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:905a0ed8ea6f2d61c1738835f99b699348d7857379083e5fc497fa0c967a407c", size = 2222876, upload-time = "2026-05-06T13:38:25.455Z" }, - { url = "https://files.pythonhosted.org/packages/b5/2f/e90b63ee2e14bd8d3db8f705a6d75d64e6ee1b7c2c8833747ce706e1e0ce/pydantic_core-2.46.4-cp310-cp310-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:ea793e075b70290d89d8142074262885d3f7da19634845135751bd6344f73b50", size = 2286703, upload-time = "2026-05-06T13:37:53.304Z" }, - { url = "https://files.pythonhosted.org/packages/ba/1e/acc4d70f88a0a277e4a1fa77ebb985ceabaf900430f875bf9338e11c9420/pydantic_core-2.46.4-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:395aebd9183f9d112f569aeb5b2214d1a10a33bec8456447f7fbdfa51d38d4cd", size = 2092042, upload-time = "2026-05-06T13:38:46.981Z" }, - { url = "https://files.pythonhosted.org/packages/a9/da/0a422b57bf8504102bf3c4ccea9c41bab5a5cee6a54650acf8faf67f5a24/pydantic_core-2.46.4-cp310-cp310-manylinux_2_31_riscv64.whl", hash = "sha256:b078afbc25f3a1436c7a1d2cd3e322497ee99615ba97c563566fdf46aff1ee01", size = 2117231, upload-time = "2026-05-06T13:39:23.146Z" }, - { url = "https://files.pythonhosted.org/packages/bd/2a/2ac13c3af305843e23c5078c53d135656b3f05a2fd78cb7bbbb12e97b473/pydantic_core-2.46.4-cp310-cp310-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:f747929cf940cddb5b3668a390056ddd5ba2e5010615ea2dcf4f9c4f3ab8791d", size = 2168388, upload-time = "2026-05-06T13:40:08.06Z" }, - { url = "https://files.pythonhosted.org/packages/72/04/2beacf7e1607e93eefe4aed1b4709f079b905fb77530179d4f7c71745f22/pydantic_core-2.46.4-cp310-cp310-musllinux_1_1_aarch64.whl", hash = "sha256:daa27d92c36f24388fe3ad306b174781c747627f134452e4f128ea00ce1fe8c4", size = 2184769, upload-time = "2026-05-06T13:38:13.901Z" }, - { url = "https://files.pythonhosted.org/packages/9e/29/d2b9fd9f539133548eaf622c06a4ce176cb46ac59f32d0359c4abc0de047/pydantic_core-2.46.4-cp310-cp310-musllinux_1_1_armv7l.whl", hash = "sha256:19e51f073cd3df251856a8a4189fbdf1de4012c3ebacfb1884f94f1eb406079f", size = 2319312, upload-time = "2026-05-06T13:39:08.24Z" }, - { url = "https://files.pythonhosted.org/packages/7c/af/0f7a5b85fec6075bea96e3ef9187de38fccced0de92c1e7feda8d5cc7bb9/pydantic_core-2.46.4-cp310-cp310-musllinux_1_1_x86_64.whl", hash = "sha256:c1747f85cee84c26985853c6f3d9bd3e75da5212912443fa111c113b9c246f39", size = 2361817, upload-time = "2026-05-06T13:38:43.2Z" }, - { url = "https://files.pythonhosted.org/packages/25/a4/73363fec545fd3ec025490bdda2743c56d0dd5b6266b1a53bbe9e4265375/pydantic_core-2.46.4-cp310-cp310-win32.whl", hash = "sha256:2f84c03c8607173d16b5a854ec68a2f9079ae03237a54fb506d13af47e1d018d", size = 1987085, upload-time = "2026-05-06T13:39:25.497Z" }, - { url = "https://files.pythonhosted.org/packages/01/aa/62f082da2c91fac1c234bc9ee0066257ce83f0604abd72e4c9d5991f2d84/pydantic_core-2.46.4-cp310-cp310-win_amd64.whl", hash = "sha256:8358a950c8909158e3df31538a7e4edc2d7265a7c54b47f0864d9e5bae9dcebf", size = 2074311, upload-time = "2026-05-06T13:39:59.922Z" }, - { url = "https://files.pythonhosted.org/packages/5c/fa/6d7708d2cfc1a832acb6aeb0cd16e801902df8a0f583bb3b4b527fde022e/pydantic_core-2.46.4-cp311-cp311-macosx_10_12_x86_64.whl", hash = "sha256:0e96592440881c74a213e5ad528e2b24d3d4f940de2766bed9010ab1d9e51594", size = 2111872, upload-time = "2026-05-06T13:40:27.596Z" }, - { url = "https://files.pythonhosted.org/packages/ae/6f/aa064a3e74b5745afbdf250594f38e7ead05e2d651bcb35994b9417a0d4d/pydantic_core-2.46.4-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:e0d65b8c354be7fb5f720c3caa8bc940bc2d20ce749c8e06135f07f8ed95dd7c", size = 1948255, upload-time = "2026-05-06T13:39:12.574Z" }, - { url = "https://files.pythonhosted.org/packages/43/3a/41114a9f7569b84b4d84e7a018c57c56347dac30c0d4a872946ec4e36c46/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:7bfb192b3f4b9e8a89b6277b6ce787564f62cfd272055f6e685726b111dc7826", size = 1972827, upload-time = "2026-05-06T13:38:19.841Z" }, - { url = "https://files.pythonhosted.org/packages/ef/25/1ab42e8048fe551934d9884e8d64daa7e990ad386f310a15981aeb6a5b08/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:9037063db01f09b09e237c282b6792bd4da634b5402c4e7f0c61effed7701a04", size = 2041051, upload-time = "2026-05-06T13:38:10.447Z" }, - { url = "https://files.pythonhosted.org/packages/94/c2/1a934597ddf08da410385b3b7aae91956a5a76c635effef456074fad7e88/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:fc010ab034c8c7452522748bf937df58020d256ccae0874463d1f4d01758af8e", size = 2221314, upload-time = "2026-05-06T13:40:13.089Z" }, - { url = "https://files.pythonhosted.org/packages/02/6d/9e8ad178c9c4df27ad3c8f25d1fe2a7ab0d2ba0559fad4aee5d3d1f16771/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:8c5dac79fa1614d1e06ca695109c6105923bd9c7d1d6c918d4e637b7e6b32fd3", size = 2285146, upload-time = "2026-05-06T13:38:59.224Z" }, - { url = "https://files.pythonhosted.org/packages/80/50/540cd3aeefc041beb111125c4bff779831a2111fc6b15a9138cda277d32c/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:f9fa868638bf362d3d138ea55829cefb3d5f4b0d7f142234382a15e2485dbec4", size = 2089685, upload-time = "2026-05-06T13:38:17.762Z" }, - { url = "https://files.pythonhosted.org/packages/6b/a4/b440ad35f05f6a38f89fa0f149accb3f0e02be94ca5e15f3c449a61b4bc9/pydantic_core-2.46.4-cp311-cp311-manylinux_2_31_riscv64.whl", hash = "sha256:17299feefe090f2caa5b8e37222bb5f663e4935a8bfa6931d4102e5df1a9f398", size = 2115420, upload-time = "2026-05-06T13:37:58.195Z" }, - { url = "https://files.pythonhosted.org/packages/99/61/de4f55db8dfd57bfdfa9a12ec90fe1b57c4f41062f7ca86f08586b3e0ac0/pydantic_core-2.46.4-cp311-cp311-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:4c63ebc82684aa89d9a3bcbd13d515b3be44250dc68dd3bd81526c1cb31286c3", size = 2165122, upload-time = "2026-05-06T13:37:01.167Z" }, - { url = "https://files.pythonhosted.org/packages/f7/52/7c529d7bdb2d1068bd52f51fe32572c8301f9a4febf1948f10639f1436f5/pydantic_core-2.46.4-cp311-cp311-musllinux_1_1_aarch64.whl", hash = "sha256:aaa2a54443eff1950ba5ddc6b6ccda0d9c84a364276a62f969bdf2a390650848", size = 2182573, upload-time = "2026-05-06T13:38:45.04Z" }, - { url = "https://files.pythonhosted.org/packages/37/b3/7c40325848ba78247f2812dcf9c7274e38cd801820ca6dd9fe63bcfb0eb4/pydantic_core-2.46.4-cp311-cp311-musllinux_1_1_armv7l.whl", hash = "sha256:18e5ceec2ab67e6d5f1a9085e5a24c9c4e2ac4545730bfe668680bca05e555f3", size = 2317139, upload-time = "2026-05-06T13:37:15.539Z" }, - { url = "https://files.pythonhosted.org/packages/d9/37/f913f81a657c865b75da6c0dbed79876073c2a43b5bd9edbe8da785e4d49/pydantic_core-2.46.4-cp311-cp311-musllinux_1_1_x86_64.whl", hash = "sha256:a0f62d0a58f4e7da165457e995725421e0064f2255d8eccebc49f41bbc23b109", size = 2360433, upload-time = "2026-05-06T13:37:30.099Z" }, - { url = "https://files.pythonhosted.org/packages/c4/67/6acaa1be2567f9256b056d8477158cac7240813956ce86e49deae8e173b4/pydantic_core-2.46.4-cp311-cp311-win32.whl", hash = "sha256:041bde0a48fd37cf71cab1c9d56d3e8625a3793fef1f7dd232b3ff37e978ecda", size = 1985513, upload-time = "2026-05-06T13:38:15.669Z" }, - { url = "https://files.pythonhosted.org/packages/aa/e6/c505f83dfeda9a2e5c995cfd872949e4d05e12f7feb3dca72f633daefa94/pydantic_core-2.46.4-cp311-cp311-win_amd64.whl", hash = "sha256:6f2eeda33a839975441c86a4119e1383c50b47faf0cbb5176985565c6bb02c33", size = 2071114, upload-time = "2026-05-06T13:40:35.416Z" }, - { url = "https://files.pythonhosted.org/packages/0f/da/7a263a96d965d9d0df5e8de8a475f33495451117035b09acb110288c381f/pydantic_core-2.46.4-cp311-cp311-win_arm64.whl", hash = "sha256:14f4c5d6db102bd796a627bbb3a17b4cf4574b9ae861d8b7c9a9661c6dd3362d", size = 2044298, upload-time = "2026-05-06T13:38:29.754Z" }, - { url = "https://files.pythonhosted.org/packages/ce/8c/af022f0af448d7747c5154288d46b5f2bc5f17366eaa0e23e9aa04d59f3b/pydantic_core-2.46.4-cp312-cp312-macosx_10_12_x86_64.whl", hash = "sha256:3245406455a5d98187ec35530fd772b1d799b26667980872c8d4614991e2c4a2", size = 2106158, upload-time = "2026-05-06T13:38:57.215Z" }, - { url = "https://files.pythonhosted.org/packages/19/95/6195171e385007300f0f5574592e467c568becce2d937a0b6804f218bc49/pydantic_core-2.46.4-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:962ccbab7b642487b1d8b7df90ef677e03134cf1fd8880bf698649b22a69371f", size = 1951724, upload-time = "2026-05-06T13:37:02.697Z" }, - { url = "https://files.pythonhosted.org/packages/8e/bc/f47d1ff9cbb1620e1b5b697eef06010035735f07820180e74178226b27b3/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:8233f2947cf85404441fd7e0085f53b10c93e0ee78611099b5c7237e36aacbf7", size = 1975742, upload-time = "2026-05-06T13:37:09.448Z" }, - { url = "https://files.pythonhosted.org/packages/5b/11/9b9a5b0306345664a2da6410877af6e8082481b5884b3ddd78d47c6013ce/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:3a233125ac121aa3ffba9a2b59edfc4a985a76092dc8279586ab4b71390875e7", size = 2052418, upload-time = "2026-05-06T13:37:38.234Z" }, - { url = "https://files.pythonhosted.org/packages/f1/b7/a65fec226f5d78fc39f4a13c4cc0c768c22b113438f60c14adc9d2865038/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:5b712b53160b79a5850310b912a5ef8e57e56947c8ad690c227f5c9d7e561712", size = 2232274, upload-time = "2026-05-06T13:38:27.753Z" }, - { url = "https://files.pythonhosted.org/packages/68/f0/92039db98b907ef49269a8271f67db9cb78ae2fc68062ef7e4e77adb5f61/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:9401557acd873c3a7f3eb9383edef8ac4968f9510e340f4808d427e75667e7b4", size = 2309940, upload-time = "2026-05-06T13:38:05.353Z" }, - { url = "https://files.pythonhosted.org/packages/5f/97/2aab507d3d00ca626e8e57c1eac6a79e4e5fbcc63eb99733ff55d1717f65/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:926c9541b14b12b1681dca8a0b75feb510b06c6341b70a8e500c2fdcff837cce", size = 2094516, upload-time = "2026-05-06T13:39:10.577Z" }, - { url = "https://files.pythonhosted.org/packages/22/37/a8aca44d40d737dde2bc05b3c6c07dff0de07ce6f82e9f3167aeaf4d5dea/pydantic_core-2.46.4-cp312-cp312-manylinux_2_31_riscv64.whl", hash = "sha256:56cb4851bcaf3d117eddcef4fe66afd750a50274b0da8e22be256d10e5611987", size = 2136854, upload-time = "2026-05-06T13:40:22.59Z" }, - { url = "https://files.pythonhosted.org/packages/24/99/fcef1b79238c06a8cbec70819ac722ba76e02bc8ada9b0fd66eba40da01b/pydantic_core-2.46.4-cp312-cp312-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:c68fcd102d71ea85c5b2dfac3f4f8476eff42a9e078fd5faefff6d145063536b", size = 2180306, upload-time = "2026-05-06T13:40:10.666Z" }, - { url = "https://files.pythonhosted.org/packages/ae/6c/fc44000918855b42779d007ae63b0532794739027b2f417321cddbc44f6a/pydantic_core-2.46.4-cp312-cp312-musllinux_1_1_aarch64.whl", hash = "sha256:b2f69dec1725e79a012d920df1707de5caf7ed5e08f3be4435e25803efc47458", size = 2190044, upload-time = "2026-05-06T13:40:43.231Z" }, - { url = "https://files.pythonhosted.org/packages/6b/65/d9cadc9f1920d7a127ad2edba16c1db7916e59719285cd6c94600b0080ba/pydantic_core-2.46.4-cp312-cp312-musllinux_1_1_armv7l.whl", hash = "sha256:8d0820e8192167f80d88d64038e609c31452eeca865b4e1d9950a27a4609b00b", size = 2329133, upload-time = "2026-05-06T13:39:57.365Z" }, - { url = "https://files.pythonhosted.org/packages/d0/cf/c873d91679f3a30bcf5e7ac280ce5573483e72295307685120d0d5ad3416/pydantic_core-2.46.4-cp312-cp312-musllinux_1_1_x86_64.whl", hash = "sha256:fbdb89b3e1c94a30cc5edfce477c6e6a5dc4d8f84665b455c27582f211a1c72c", size = 2374464, upload-time = "2026-05-06T13:38:06.976Z" }, - { url = "https://files.pythonhosted.org/packages/47/bd/6f2fc8188f31bf10590f1e98e7b306336161fac930a8c514cd7bd828c7dc/pydantic_core-2.46.4-cp312-cp312-win32.whl", hash = "sha256:9aa768456404a8bf48a4406685ac2bec8e72b62c69313734fa3b73cf33b3a894", size = 1974823, upload-time = "2026-05-06T13:40:47.985Z" }, - { url = "https://files.pythonhosted.org/packages/40/8c/985c1d41ea1107c2534abd9870e4ed5c8e7669b5c308297835c001e7a1c4/pydantic_core-2.46.4-cp312-cp312-win_amd64.whl", hash = "sha256:e9c26f834c65f5752f3f06cb08cb86a913ceb7274d0db6e267808a708b46bc89", size = 2072919, upload-time = "2026-05-06T13:39:21.153Z" }, - { url = "https://files.pythonhosted.org/packages/c4/ba/f463d006e0c47373ca7ec5e1a261c59dc01ef4d62b2657af925fb0deee3a/pydantic_core-2.46.4-cp312-cp312-win_arm64.whl", hash = "sha256:4fc73cb559bdb54b1134a706a2802a4cddd27a0633f5abb7e53056268751ac6a", size = 2027604, upload-time = "2026-05-06T13:39:03.753Z" }, - { url = "https://files.pythonhosted.org/packages/51/a2/5d30b469c5267a17b39dec53208222f76a8d351dfac4af661888c5aee77d/pydantic_core-2.46.4-cp313-cp313-macosx_10_12_x86_64.whl", hash = "sha256:5d5902252db0d3cedf8d4a1bc68f70eeb430f7e4c7104c8c476753519b423008", size = 2106306, upload-time = "2026-05-06T13:37:48.029Z" }, - { url = "https://files.pythonhosted.org/packages/c1/81/4fa520eaffa8bd7d1525e644cd6d39e7d60b1592bc5b516693c7340b50f1/pydantic_core-2.46.4-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:c94f0688e7b8d0a67abf40e57a7eaaecd17cc9586706a31b76c031f63df052b4", size = 1951906, upload-time = "2026-05-06T13:37:17.012Z" }, - { url = "https://files.pythonhosted.org/packages/03/d5/fd02da45b659668b05923b17ba3a0100a0a3d5541e3bd8fcc4ecb711309e/pydantic_core-2.46.4-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:f027324c56cd5406ca49c124b0db10e56c69064fec039acc571c29020cc87c76", size = 1976802, upload-time = "2026-05-06T13:37:35.113Z" }, - { url = "https://files.pythonhosted.org/packages/21/f2/95727e1368be3d3ed485eaab7adbd7dda408f33f7a36e8b48e0144002b91/pydantic_core-2.46.4-cp313-cp313-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:e739fee756ba1010f8bcccb534252e85a35fe45ae92c295a06059ce58b74ccd3", size = 2052446, upload-time = "2026-05-06T13:37:12.313Z" }, - { url = "https://files.pythonhosted.org/packages/9c/86/5d99feea3f77c7234b8718075b23db11532773c1a0dbd9b9490215dc2eeb/pydantic_core-2.46.4-cp313-cp313-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:9d56801be94b86a9da183e5f3766e6310752b99ff647e38b09a9500d88e46e76", size = 2232757, upload-time = "2026-05-06T13:39:01.149Z" }, - { url = "https://files.pythonhosted.org/packages/d2/3a/508ac615935ef7588cf6d9e9b91309fdc2da751af865e02a9098de88258c/pydantic_core-2.46.4-cp313-cp313-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:2412e734dcb48da14d4e4006b82b46b74f2518b8a26ee7e58c6844a6cd6d03c4", size = 2309275, upload-time = "2026-05-06T13:37:41.406Z" }, - { url = "https://files.pythonhosted.org/packages/07/f8/41db9de19d7987d6b04715a02b3b40aea467000275d9d758ffaa31af7d50/pydantic_core-2.46.4-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:9551187363ffc0de2a00b2e47c25aeaeb1020b69b668762966df15fc5659dd5a", size = 2094467, upload-time = "2026-05-06T13:39:18.847Z" }, - { url = "https://files.pythonhosted.org/packages/2c/e2/f35033184cb11d0052daf4416e8e10a502ea2ac006fc4f459aee872727d1/pydantic_core-2.46.4-cp313-cp313-manylinux_2_31_riscv64.whl", hash = "sha256:0186750b482eefa11d7f435892b09c5c606193ef3375bcf94aa00ae6bfb66262", size = 2134417, upload-time = "2026-05-06T13:40:17.944Z" }, - { url = "https://files.pythonhosted.org/packages/7e/7b/6ceeb1cc90e193862f444ebe373d8fdf613f0a82572dde03fb10734c6c71/pydantic_core-2.46.4-cp313-cp313-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:5855698a4856556d86e8e6cd8434bc3ac0314ee8e12089ae0e143f64c6256e4e", size = 2179782, upload-time = "2026-05-06T13:40:32.618Z" }, - { url = "https://files.pythonhosted.org/packages/5a/f2/c8d7773ede6af08036423a00ae0ceffce266c3c52a096c435d68c896083f/pydantic_core-2.46.4-cp313-cp313-musllinux_1_1_aarch64.whl", hash = "sha256:cbaf13819775b7f769bf4a1f066cb6df7a28d4480081a589828ef190226881cd", size = 2188782, upload-time = "2026-05-06T13:36:51.018Z" }, - { url = "https://files.pythonhosted.org/packages/59/31/0c864784e31f09f05cdd87606f08923b9c9e7f6e51dd27f20f62f975ce9f/pydantic_core-2.46.4-cp313-cp313-musllinux_1_1_armv7l.whl", hash = "sha256:633147d34cf4550417f12e2b1a0383973bdf5cdfde212cb09e9a581cf10820be", size = 2328334, upload-time = "2026-05-06T13:40:37.764Z" }, - { url = "https://files.pythonhosted.org/packages/c2/eb/4f6c8a41efa30baa755590f4141abf3a8c370fab610915733e74134a7270/pydantic_core-2.46.4-cp313-cp313-musllinux_1_1_x86_64.whl", hash = "sha256:82cf5301172168103724d49a1444d3378cb20cdee30b116a1bd6031236298a5d", size = 2372986, upload-time = "2026-05-06T13:39:34.152Z" }, - { url = "https://files.pythonhosted.org/packages/5b/24/b375a480d53113860c299764bfe9f349a3dc9108b3adc0d7f0d786492ebf/pydantic_core-2.46.4-cp313-cp313-win32.whl", hash = "sha256:9fa8ae11da9e2b3126c6426f147e0fba88d96d65921799bb30c6abd1cb2c97fb", size = 1973693, upload-time = "2026-05-06T13:37:55.072Z" }, - { url = "https://files.pythonhosted.org/packages/7e/e8/cff247591966f2d22ec8c003cd7587e27b7ba7b81ab2fb888e3ab75dc285/pydantic_core-2.46.4-cp313-cp313-win_amd64.whl", hash = "sha256:6b3ace8194b0e5204818c92802dcdca7fc6d88aabbb799d7c795540d9cd6d292", size = 2071819, upload-time = "2026-05-06T13:38:49.139Z" }, - { url = "https://files.pythonhosted.org/packages/c6/1a/f4aee670d5670e9e148e0c82c7db98d780be566c6e6a97ee8035528ca0b3/pydantic_core-2.46.4-cp313-cp313-win_arm64.whl", hash = "sha256:184c081504d17f1c1066e430e117142b2c77d9448a97f7b65c6ac9fd9aee238d", size = 2027411, upload-time = "2026-05-06T13:40:45.796Z" }, - { url = "https://files.pythonhosted.org/packages/8d/74/228a26ddad29c6672b805d9fd78e8d251cd04004fa7eed0e622096cd0250/pydantic_core-2.46.4-cp314-cp314-macosx_10_12_x86_64.whl", hash = "sha256:428e04521a40150c85216fc8b85e8d39fece235a9cf5e383761238c7fa9b96fb", size = 2102079, upload-time = "2026-05-06T13:38:41.019Z" }, - { url = "https://files.pythonhosted.org/packages/ad/1f/8970b150a4b4365623ae00fc88603491f763c627311ae8031e3111356d6e/pydantic_core-2.46.4-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:23ace664830ee0bfe014a0c7bc248b1f7f25ed7ad103852c317624a1083af462", size = 1952179, upload-time = "2026-05-06T13:36:59.812Z" }, - { url = "https://files.pythonhosted.org/packages/95/30/5211a831ae054928054b2f79731661087a2bc5c01e825c672b3a4a8f1b3e/pydantic_core-2.46.4-cp314-cp314-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:ce5c1d2a8b27468f433ca974829c44060b8097eedc39933e3c206a90ee49c4a9", size = 1978926, upload-time = "2026-05-06T13:37:39.933Z" }, - { url = "https://files.pythonhosted.org/packages/57/e9/689668733b1eb67adeef047db3c2e8788fcf65a7fd9c9e2b46b7744fe245/pydantic_core-2.46.4-cp314-cp314-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:7283d57845ecf5a163403eb0702dfc220cc4fbdd18919cb5ccea4f95ee1cdab4", size = 2046785, upload-time = "2026-05-06T13:38:01.995Z" }, - { url = "https://files.pythonhosted.org/packages/60/d9/6715260422ff50a2109878fd24d948a6c3446bb2664f34ee78cd972b3acd/pydantic_core-2.46.4-cp314-cp314-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:8daafc69c93ee8a0204506a3b6b30f586ef54028f52aeeeb5c4cfc5184fd5914", size = 2228733, upload-time = "2026-05-06T13:40:50.371Z" }, - { url = "https://files.pythonhosted.org/packages/18/ae/fdb2f64316afca925640f8e70bb1a564b0ec2721c1389e25b8eb4bf9a299/pydantic_core-2.46.4-cp314-cp314-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:cd2213145bcc2ba85884d0ac63d222fece9209678f77b9b4d76f054c561adb28", size = 2307534, upload-time = "2026-05-06T13:37:21.531Z" }, - { url = "https://files.pythonhosted.org/packages/89/1d/8eff589b45bb8190a9d12c49cfad0f176a5cbd1534908a6b5125e2886239/pydantic_core-2.46.4-cp314-cp314-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:7a5f930472650a82629163023e630d160863fce524c616f4e5186e5de9d9a49b", size = 2099732, upload-time = "2026-05-06T13:39:31.942Z" }, - { url = "https://files.pythonhosted.org/packages/06/d5/ee5a3366637fee41dee51a1fc91562dcf12ddbc68fda34e6b253da2324bb/pydantic_core-2.46.4-cp314-cp314-manylinux_2_31_riscv64.whl", hash = "sha256:c1b3f518abeca3aa13c712fd202306e145abf59a18b094a6bafb2d2bbf59192c", size = 2129627, upload-time = "2026-05-06T13:37:25.033Z" }, - { url = "https://files.pythonhosted.org/packages/94/33/2414be571d2c6a6c4d08be21f9292b6d3fdb08949a97b6dfe985017821db/pydantic_core-2.46.4-cp314-cp314-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:1a7dd0b3ee80d90150e3495a3a13ac34dbcbfd4f012996a6a1d8900e91b5c0fb", size = 2179141, upload-time = "2026-05-06T13:37:14.046Z" }, - { url = "https://files.pythonhosted.org/packages/7b/79/7daa95be995be0eecc4cf75064cb33f9bbbfe3fe0158caf2f0d4a996a5c7/pydantic_core-2.46.4-cp314-cp314-musllinux_1_1_aarch64.whl", hash = "sha256:3fb702cd90b0446a3a1c5e470bfa0dd23c0233b676a9099ddcc964fa6ca13898", size = 2184325, upload-time = "2026-05-06T13:36:53.615Z" }, - { url = "https://files.pythonhosted.org/packages/9f/cb/d0a382f5c0de8a222dc61c65348e0ce831b1f68e0a018450d31c2cace3a5/pydantic_core-2.46.4-cp314-cp314-musllinux_1_1_armv7l.whl", hash = "sha256:b8458003118a712e66286df6a707db01c52c0f52f7db8e4a38f0da1d3b94fc4e", size = 2323990, upload-time = "2026-05-06T13:40:29.971Z" }, - { url = "https://files.pythonhosted.org/packages/05/db/d9ba624cc4a5aced1598e88c04fdbd8310c8a69b9d38b9a3d39ce3a61ed7/pydantic_core-2.46.4-cp314-cp314-musllinux_1_1_x86_64.whl", hash = "sha256:372429a130e469c9cd698925ce5fc50940b7a1336b0d82038e63d5bbc4edc519", size = 2369978, upload-time = "2026-05-06T13:37:23.027Z" }, - { url = "https://files.pythonhosted.org/packages/f2/20/d15df15ba918c423461905802bfd2981c3af0bfa0e40d05e13edbfa48bc3/pydantic_core-2.46.4-cp314-cp314-win32.whl", hash = "sha256:85bb3611ff1802f3ee7fdd7dbff26b56f343fb432d57a4728fdd49b6ef35e2f4", size = 1966354, upload-time = "2026-05-06T13:38:03.499Z" }, - { url = "https://files.pythonhosted.org/packages/fc/b6/6b8de4c0a7d7ab3004c439c80c5c1e0a3e8d78bbae19379b01960383d9e5/pydantic_core-2.46.4-cp314-cp314-win_amd64.whl", hash = "sha256:811ff8e9c313ab425368bcbb36e5c4ebd7108c2bbf4e4089cfbb0b01eff63fac", size = 2072238, upload-time = "2026-05-06T13:39:40.807Z" }, - { url = "https://files.pythonhosted.org/packages/32/36/51eb763beec1f4cf59b1db243a7dcc39cbb41230f050a09b9d69faaf0a48/pydantic_core-2.46.4-cp314-cp314-win_arm64.whl", hash = "sha256:bfec22eab3c8cc2ceec0248aec886624116dc079afa027ecc8ad4a7e62010f8a", size = 2018251, upload-time = "2026-05-06T13:37:26.72Z" }, - { url = "https://files.pythonhosted.org/packages/e8/91/855af51d625b23aa987116a19e231d2aaef9c4a415273ddc189b79a45fee/pydantic_core-2.46.4-cp314-cp314t-macosx_10_12_x86_64.whl", hash = "sha256:af8244b2bef6aaad6d92cda81372de7f8c8d36c9f0c3ea36e827c60e7d9467a0", size = 2099593, upload-time = "2026-05-06T13:39:47.682Z" }, - { url = "https://files.pythonhosted.org/packages/fb/1b/8784a54c65edb5f49f0a14d6977cf1b209bba85a4c77445b255c2de58ab3/pydantic_core-2.46.4-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:5a4330cdbc57162e4b3aa303f588ba752257694c9c9be3e7ebb11b4aca659b5d", size = 1935226, upload-time = "2026-05-06T13:40:40.428Z" }, - { url = "https://files.pythonhosted.org/packages/e8/e7/1955d28d1afc56dd4b3ad7cc0cf39df1b9852964cf16e5d13912756d6d6b/pydantic_core-2.46.4-cp314-cp314t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:29c61fc04a3d840155ff08e475a04809278972fe6aef51e2720554e96367e34b", size = 1974605, upload-time = "2026-05-06T13:37:32.029Z" }, - { url = "https://files.pythonhosted.org/packages/93/e2/3fedbf0ba7a22850e6e9fd78117f1c0f10f950182344d8a6c535d468fdd8/pydantic_core-2.46.4-cp314-cp314t-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:c50f2528cf200c5eed56faf3f4e22fcd5f38c157a8b78576e6ba3168ec35f000", size = 2030777, upload-time = "2026-05-06T13:38:55.239Z" }, - { url = "https://files.pythonhosted.org/packages/f8/61/46be275fcaaba0b4f5b9669dd852267ce1ff616592dccf7a7845588df091/pydantic_core-2.46.4-cp314-cp314t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:0cbe8b01f948de4286c74cdd6c667aceb38f5c1e26f0693b3983d9d74887c65e", size = 2236641, upload-time = "2026-05-06T13:37:08.096Z" }, - { url = "https://files.pythonhosted.org/packages/60/db/12e93e46a8bac9988be3c016860f83293daea8c716c029c9ace279036f2f/pydantic_core-2.46.4-cp314-cp314t-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:617d7e2ca7dcb8c5cf6bcb8c59b8832c94b36196bbf1cbd1bfb56ed341905edd", size = 2286404, upload-time = "2026-05-06T13:40:20.221Z" }, - { url = "https://files.pythonhosted.org/packages/e2/4a/4d8b19008f38d31c53b8219cfedc2e3d5de5fe99d90076b7e767de29274f/pydantic_core-2.46.4-cp314-cp314t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:7027560ee92211647d0d34e3f7cd6f50da56399d26a9c8ad0da286d3869a53f3", size = 2109219, upload-time = "2026-05-06T13:38:12.153Z" }, - { url = "https://files.pythonhosted.org/packages/88/70/3cbc40978fefb7bb09c6708d40d4ad1a5d70fd7213c3d17f971de868ec1f/pydantic_core-2.46.4-cp314-cp314t-manylinux_2_31_riscv64.whl", hash = "sha256:f99626688942fb746e545232e7726926f3be91b5975f8b55327665fafda991c7", size = 2110594, upload-time = "2026-05-06T13:40:02.971Z" }, - { url = "https://files.pythonhosted.org/packages/9d/20/b8d36736216e29491125531685b2f9e61aa5b4b2599893f8268551da3338/pydantic_core-2.46.4-cp314-cp314t-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:fc3e9034a63de20e15e8ade85358bc6efc614008cab72898b4b4952bea0509ff", size = 2159542, upload-time = "2026-05-06T13:39:27.506Z" }, - { url = "https://files.pythonhosted.org/packages/1d/a2/367df868eb584dacf6bf82a389272406d7178e301c4ac82545ab98bc2dd9/pydantic_core-2.46.4-cp314-cp314t-musllinux_1_1_aarch64.whl", hash = "sha256:97e7cf2be5c77b7d1a9713a05605d49460d02c6078d38d8bef3cbe323c548424", size = 2168146, upload-time = "2026-05-06T13:38:31.93Z" }, - { url = "https://files.pythonhosted.org/packages/c1/b8/4460f77f7e201893f649a29ab355dddd3beee8a97bcb1a320db414f9a06e/pydantic_core-2.46.4-cp314-cp314t-musllinux_1_1_armv7l.whl", hash = "sha256:3bf92c5d0e00fefaab325a4d27828fe6b6e2a21848686b5b60d2d9eeb09d76c6", size = 2306309, upload-time = "2026-05-06T13:37:44.717Z" }, - { url = "https://files.pythonhosted.org/packages/64/c4/be2639293acd87dc8ddbcec41a73cee9b2ebf996fe6d892a1a74e88ad3f7/pydantic_core-2.46.4-cp314-cp314t-musllinux_1_1_x86_64.whl", hash = "sha256:3ecbc122d18468d06ca279dc26a8c2e2d5acb10943bb35e36ae92096dc3b5565", size = 2369736, upload-time = "2026-05-06T13:37:05.645Z" }, - { url = "https://files.pythonhosted.org/packages/30/a6/9f9f380dbb301f67023bf8f707aaa75daadf84f7152d95c410fd7e81d994/pydantic_core-2.46.4-cp314-cp314t-win32.whl", hash = "sha256:e846ae7835bf0703ae43f534ab79a867146dadd59dc9ca5c8b53d5c8f7c9ef02", size = 1955575, upload-time = "2026-05-06T13:38:51.116Z" }, - { url = "https://files.pythonhosted.org/packages/40/1f/f1eb9eb350e795d1af8586289746f5c5677d16043040d63710e22abc43c9/pydantic_core-2.46.4-cp314-cp314t-win_amd64.whl", hash = "sha256:2108ba5c1c1eca18030634489dc544844144ee36357f2f9f780b93e7ddbb44b5", size = 2051624, upload-time = "2026-05-06T13:38:21.672Z" }, - { url = "https://files.pythonhosted.org/packages/f6/d2/42dd53d0a85c27606f316d3aa5d2869c4e8470a5ed6dec30e4a1abe19192/pydantic_core-2.46.4-cp314-cp314t-win_arm64.whl", hash = "sha256:4fcbe087dbc2068af7eda3aa87634eba216dbda64d1ae73c8684b621d33f6596", size = 2017325, upload-time = "2026-05-06T13:40:52.723Z" }, - { url = "https://files.pythonhosted.org/packages/ee/a4/73995fd4ebbb46ba0ee51e6fa049b8f02c40daebb762208feda8a6b7894d/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-macosx_10_12_x86_64.whl", hash = "sha256:14d4edf427bdcf950a8a02d7cb44a08614388dd6e1bdcbf4f67504fa7887da9c", size = 2111589, upload-time = "2026-05-06T13:37:10.817Z" }, - { url = "https://files.pythonhosted.org/packages/fb/7f/f37d3a5e8bfcc2e403f5c57a730f2d815693fb42119e8ea48b3789335af1/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-macosx_11_0_arm64.whl", hash = "sha256:0ce40cd7b21210e99342afafbd4d0f76d784eb5b1d60f3bdc566be4983c6c73b", size = 1944552, upload-time = "2026-05-06T13:36:56.717Z" }, - { url = "https://files.pythonhosted.org/packages/15/3c/d7eb777b3ff43e8433a4efb39a17aa8fd98a4ee8561a24a67ef5db07b2d6/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:90884113d8b48f760e9587002789ddd741e76ab9f89518cd1e43b1f1a52ec44b", size = 1982984, upload-time = "2026-05-06T13:39:06.207Z" }, - { url = "https://files.pythonhosted.org/packages/63/87/70b9f40170a81afd55ca26c9b2acb25c20d64bcfbf888fafecb3ba077d4c/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:66ce7632c22d837c95301830e111ad0128a32b8207533b60896a96c4915192ea", size = 2138417, upload-time = "2026-05-06T13:39:45.476Z" }, - { url = "https://files.pythonhosted.org/packages/9d/1d/8987ad40f65ae1432753072f214fb5c74fe47ffbd0698bb9cbbb585664f8/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-macosx_10_12_x86_64.whl", hash = "sha256:1d8ba486450b14f3b1d63bc521d410ec7565e52f887b9fb671791886436a42f7", size = 2095527, upload-time = "2026-05-06T13:39:52.283Z" }, - { url = "https://files.pythonhosted.org/packages/64/d3/84c282a7eee1d3ac4c0377546ef5a1ea436ce26840d9ac3b7ed54a377507/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-macosx_11_0_arm64.whl", hash = "sha256:3009f12e4e90b7f88b4f9adb1b0c4a3d58fe7820f3238c190047209d148026df", size = 1936024, upload-time = "2026-05-06T13:40:15.671Z" }, - { url = "https://files.pythonhosted.org/packages/d7/ca/eac61596cdeb4d7e174d3dc0bd8a6238f14f75f97a24e7b7db4c7e7340a0/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:ad785e92e6dc634c21555edc8bd6b64957ab844541bcb96a1366c202951ae526", size = 1990696, upload-time = "2026-05-06T13:38:34.717Z" }, - { url = "https://files.pythonhosted.org/packages/fa/c3/7c8b240552251faf6b3a957db200fcfbbcec36763c050428b601e0c9b83b/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:00c603d540afdd6b80eb39f078f33ebd46211f02f33e34a32d9f053bba711de0", size = 2147590, upload-time = "2026-05-06T13:39:29.883Z" }, - { url = "https://files.pythonhosted.org/packages/11/cb/428de0385b6c8d44b716feba566abfacfbd23ee3c4439faa789a1456242f/pydantic_core-2.46.4-pp311-pypy311_pp73-macosx_10_12_x86_64.whl", hash = "sha256:0c563b08bca408dc7f65f700633d8442fffb2421fc47b8101377e9fd65051ff0", size = 2112782, upload-time = "2026-05-06T13:37:04.016Z" }, - { url = "https://files.pythonhosted.org/packages/0b/b5/6a17bdadd0fc1f170adfd05a20d37c832f52b117b4d9131da1f41bb097ce/pydantic_core-2.46.4-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:db06ffe51636ffe9ca531fe9023dd64bdd794be8754cb5df57c5498ae5b518a7", size = 1952146, upload-time = "2026-05-06T13:39:43.092Z" }, - { url = "https://files.pythonhosted.org/packages/2a/dc/03734d80e362cd43ef65428e9de77c730ce7f2f11c60d2b1e1b39f0fbf99/pydantic_core-2.46.4-pp311-pypy311_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:133878133d271ade3d41d1bfb2a45ec38dbdbda40bc065921c6b04e4630127e2", size = 2134492, upload-time = "2026-05-06T13:36:58.124Z" }, - { url = "https://files.pythonhosted.org/packages/de/df/5e5ffc085ed07cc22d298134d3d911c63e91f6a0eb91fe646750a3209910/pydantic_core-2.46.4-pp311-pypy311_pp73-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:9bc519fbf2b7578398853d815009ae5e4d4603d12f4e3f91da8c06852d3da3e9", size = 2156604, upload-time = "2026-05-06T13:37:49.88Z" }, - { url = "https://files.pythonhosted.org/packages/81/44/6e112a4253e56f5705467cbab7ab5e91ee7398ba3d56d358635958893d3e/pydantic_core-2.46.4-pp311-pypy311_pp73-musllinux_1_1_aarch64.whl", hash = "sha256:c7a7bd4e39e8e4c12c39cd480356842b6a8a06e41b23a55a5e3e191718838ddf", size = 2183828, upload-time = "2026-05-06T13:37:43.053Z" }, - { url = "https://files.pythonhosted.org/packages/ac/ad/5565071e937d8e752842ac241463944c9eb14c87e2d269f2658a5bd05e98/pydantic_core-2.46.4-pp311-pypy311_pp73-musllinux_1_1_armv7l.whl", hash = "sha256:d396ec2b979760aaf3218e76c24e65bd0aca24983298653b3a9d7a45f9e47b30", size = 2310000, upload-time = "2026-05-06T13:37:56.694Z" }, - { url = "https://files.pythonhosted.org/packages/4f/c3/66883a5cec183e7fba4d024b4cbbe61851a63750ef606b0afecc46d1f2bf/pydantic_core-2.46.4-pp311-pypy311_pp73-musllinux_1_1_x86_64.whl", hash = "sha256:86e1a4418c6cd97d60c95c71164158eaf7324fae7b0923264016baa993eba6fc", size = 2361286, upload-time = "2026-05-06T13:40:05.667Z" }, - { url = "https://files.pythonhosted.org/packages/4b/2d/69abac8f838090bbecd5df894befb2c2619e7996a98ddb949db9f3b93225/pydantic_core-2.46.4-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:d51026d73fcfd93610abc7b27789c26b313920fcfb20e27462d74a7f8b06e983", size = 2193071, upload-time = "2026-05-06T13:38:08.682Z" }, -] - -[[package]] -name = "pydantic-settings" -version = "2.14.2" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "pydantic" }, - { name = "python-dotenv" }, - { name = "typing-inspection" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/5c/b5/8f48e906c3e0205276e8bd8cb7512217a87b2685304d64be27cad5b3019f/pydantic_settings-2.14.2.tar.gz", hash = "sha256:c19dd64b19097f1de80184f0cc7b0272a13ae6e170cbf240a3e27e381ed14a5f", size = 237700, upload-time = "2026-06-19T13:44:56.324Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/77/c1/6e422f34e569cf8e18df68d1939c81c099d2b61e4f7d9621c8a77560799c/pydantic_settings-2.14.2-py3-none-any.whl", hash = "sha256:a20c97b37910b6550d5ea50fbcc2d4187defe58cd57070b73863d069419c9440", size = 61715, upload-time = "2026-06-19T13:44:55.02Z" }, -] - -[[package]] -name = "pygments" -version = "2.20.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/c3/b2/bc9c9196916376152d655522fdcebac55e66de6603a76a02bca1b6414f6c/pygments-2.20.0.tar.gz", hash = "sha256:6757cd03768053ff99f3039c1a36d6c0aa0b263438fcab17520b30a303a82b5f", size = 4955991, upload-time = "2026-03-29T13:29:33.898Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/f4/7e/a72dd26f3b0f4f2bf1dd8923c85f7ceb43172af56d63c7383eb62b332364/pygments-2.20.0-py3-none-any.whl", hash = "sha256:81a9e26dd42fd28a23a2d169d86d7ac03b46e2f8b59ed4698fb4785f946d0176", size = 1231151, upload-time = "2026-03-29T13:29:30.038Z" }, -] - -[[package]] -name = "pyjwt" -version = "2.13.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "typing-extensions", marker = "python_full_version < '3.11'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/3b/81/58d0ac84e1ef3a3843791d6954d94c0b33d526c75eeb1efbce9d0a4c4077/pyjwt-2.13.0.tar.gz", hash = "sha256:41571c89ca91598c79e8ef18a2d07367d4810fbbd6f637794879baf1b7703423", size = 107515, upload-time = "2026-05-21T19:54:36.618Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/a3/5e/ecf12fdb62546d64385c158514e9b2b671f7832108ef2ecd2020ce0af2d1/pyjwt-2.13.0-py3-none-any.whl", hash = "sha256:66adcc2aff09b3f1bbd95fc1e1577df8ac8723c978552fd43304c8a290ac5728", size = 31274, upload-time = "2026-05-21T19:54:35.362Z" }, -] - -[package.optional-dependencies] -crypto = [ - { name = "cryptography" }, -] - -[[package]] -name = "pytest" -version = "9.1.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "colorama", marker = "sys_platform == 'win32'" }, - { name = "exceptiongroup", marker = "python_full_version < '3.11'" }, - { name = "iniconfig" }, - { name = "packaging" }, - { name = "pluggy" }, - { name = "pygments" }, - { name = "tomli", marker = "python_full_version < '3.11'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/e4/47/b9efed96c114afcfa3c9d3fe98a76a1d14c74a9e266d397cf6eb64be5e01/pytest-9.1.1.tar.gz", hash = "sha256:1088fbde8f2b49d95a549a195707afa7a76a3ce9bcadc26b6d71f0ffda5fe313", size = 1636369, upload-time = "2026-06-19T10:58:32.857Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/24/25/1de2678b631f5a49215c6c96fff41ba892b0a34df68d6d80292b1b48aa7f/pytest-9.1.1-py3-none-any.whl", hash = "sha256:37a86b45efb9a47a61a36449063e8e18d0cab3161329fc099eb21783169c4f0c", size = 386536, upload-time = "2026-06-19T10:58:31.347Z" }, -] - -[[package]] -name = "python-dotenv" -version = "1.2.2" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/82/ed/0301aeeac3e5353ef3d94b6ec08bbcabd04a72018415dcb29e588514bba8/python_dotenv-1.2.2.tar.gz", hash = "sha256:2c371a91fbd7ba082c2c1dc1f8bf89ca22564a087c2c287cd9b662adde799cf3", size = 50135, upload-time = "2026-03-01T16:00:26.196Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/0b/d7/1959b9648791274998a9c3526f6d0ec8fd2233e4d4acce81bbae76b44b2a/python_dotenv-1.2.2-py3-none-any.whl", hash = "sha256:1d8214789a24de455a8b8bd8ae6fe3c6b69a5e3d64aa8a8e5d68e694bbcb285a", size = 22101, upload-time = "2026-03-01T16:00:25.09Z" }, -] - -[[package]] -name = "python-multipart" -version = "0.0.32" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/5b/42/55c32bb9b12693c092ad250a0e82edb5b31ddeda6eb772de5f308b3804ad/python_multipart-0.0.32.tar.gz", hash = "sha256:be54b7f3fa167bb83e4fcd936b887b708f4e57fe75911c02aebf53efaf8d938e", size = 46881, upload-time = "2026-06-04T16:18:58.647Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/e1/04/e8135ebd1ad02c56ec633277529b2602ff99ff634be76cdba5744cf554fd/python_multipart-0.0.32-py3-none-any.whl", hash = "sha256:ff6d3f776f16878c894e52e107296ffc890e913c611b1a4ec6c44e2821fe2e23", size = 30042, upload-time = "2026-06-04T16:18:57.319Z" }, -] - -[[package]] -name = "pywin32" -version = "312" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/fe/1b/9cfdeac80ee45bebbbcb31f1b7b99a0d81a1c72de48d837be984e0e88b1d/pywin32-312-cp310-cp310-win32.whl", hash = "sha256:772235332b5d1024c696f11cea1ae4be7930f0a8b894bb43db14e3f435f1ff7e", size = 6361387, upload-time = "2026-06-04T07:49:14.329Z" }, - { url = "https://files.pythonhosted.org/packages/33/b1/7afc96d041d982c27bc2df6f853d43f01fd273e3d39d04be3647ddeb533d/pywin32-312-cp310-cp310-win_amd64.whl", hash = "sha256:5dbc35d2b5320dc07f25fa31269cfb767471002b17de5eb067d03da68c7cb2db", size = 6926780, upload-time = "2026-06-04T07:49:16.881Z" }, - { url = "https://files.pythonhosted.org/packages/ce/3a/4140da9ad54108e517f4a16b2d83da3033e08662144623e1239587cb7db6/pywin32-312-cp310-cp310-win_arm64.whl", hash = "sha256:3020656e34f1cf7faeb7bccd2b84653a607c6ff0c55ada85e6487d61716deabd", size = 4307203, upload-time = "2026-06-04T07:49:18.993Z" }, - { url = "https://files.pythonhosted.org/packages/1f/f5/10a6e845a00fc5e7afd0a988b744f403d4d57162a28d160a093c4d9322f0/pywin32-312-cp311-cp311-win32.whl", hash = "sha256:17948aeadbdb091f0ced6ef0841620794e68327b94ee415571c1203594b7215c", size = 6362659, upload-time = "2026-06-04T07:49:21.349Z" }, - { url = "https://files.pythonhosted.org/packages/35/c4/dcd2d62b5944b6d5db53413a5899016ccd57ffcb7278f3f81655d25d2027/pywin32-312-cp311-cp311-win_amd64.whl", hash = "sha256:d11417d84412f859b722fad0841b3614459ed0047f7542d8362e77884f6b6e8a", size = 6928825, upload-time = "2026-06-04T07:49:23.934Z" }, - { url = "https://files.pythonhosted.org/packages/b7/56/3cbb433fe4501cdba2eb9040f56a4e1a8243faa4186b25295564d1a7a79d/pywin32-312-cp311-cp311-win_arm64.whl", hash = "sha256:b2200a054ca6d6625c4842fc56a4976a4b47f96b73dbe5538c3f813a80359f47", size = 6721875, upload-time = "2026-06-04T07:49:26.416Z" }, - { url = "https://files.pythonhosted.org/packages/83/ff/32aa7d2ed0ab12b323aaa64f9b75e6ad4f8fd09f9ccfc28c79414d46838d/pywin32-312-cp312-cp312-win32.whl", hash = "sha256:dab4f65ac9c4e48400a2a0530c46c3c579cd5905ecd11b80692373915269208b", size = 6371877, upload-time = "2026-06-04T07:49:28.836Z" }, - { url = "https://files.pythonhosted.org/packages/03/d9/77040d3b43df3f3be32ea289433d660d2727f5ba327bc73be835127d9d60/pywin32-312-cp312-cp312-win_amd64.whl", hash = "sha256:b457f6d628a47e8a7346ce22acb7e1a46a4a78b52e1d17e1af56871bd19a93bc", size = 6914841, upload-time = "2026-06-04T07:49:31.85Z" }, - { url = "https://files.pythonhosted.org/packages/e3/cc/7b1ec671775756020a0ee7f4feeaf3c568f0ab86bd3900088cf986937a92/pywin32-312-cp312-cp312-win_arm64.whl", hash = "sha256:6017c58e12f6809fbb0555b75df144c2922a9ffd18e4b9b5afa863b6c1a9d950", size = 6727901, upload-time = "2026-06-04T07:49:34.244Z" }, - { url = "https://files.pythonhosted.org/packages/2d/41/12fbfd7f36ed2146d8bc9de96c2741296bf0d490b98508496cff322e274c/pywin32-312-cp313-cp313-win32.whl", hash = "sha256:7a27df850933d16a8eabfbaeb73d52b273e2da667f80d70b01a89d1f6828d02c", size = 6370184, upload-time = "2026-06-04T07:49:36.253Z" }, - { url = "https://files.pythonhosted.org/packages/ba/db/36a78e3403099d31d9746d13fdcde5accc43c1155f375a34d15983a479a7/pywin32-312-cp313-cp313-win_amd64.whl", hash = "sha256:c53e878d15a1c44788082bfe712a905433473aa38f86375b7cf8b45e3acbaaf9", size = 6914298, upload-time = "2026-06-04T07:49:38.876Z" }, - { url = "https://files.pythonhosted.org/packages/84/37/c1697194092b76de9ed47ca124323f02c57ffc8a45c06f88a3d5acaf01eb/pywin32-312-cp313-cp313-win_arm64.whl", hash = "sha256:59aba5d5940842075343a5ddc6b11f1cdf0d1567fe745290359dfbcc7c2eb831", size = 6727640, upload-time = "2026-06-04T07:49:41.083Z" }, - { url = "https://files.pythonhosted.org/packages/fc/2b/1f3cded5822fd49c02f40544cbb5f58c7cfd6b1694869fd476cb6170ee97/pywin32-312-cp314-cp314-win32.whl", hash = "sha256:a77a90fbb6881238d2ca9c6fd797b25817f3768fe78d214a90137ff055a75f5b", size = 6468928, upload-time = "2026-06-04T07:49:43.188Z" }, - { url = "https://files.pythonhosted.org/packages/21/82/3bf86d2e2808902013132e1ce905a7da0da53790f3836c64bf44d55e24f3/pywin32-312-cp314-cp314-win_amd64.whl", hash = "sha256:a4dd3a848290ef724347b19f301045831d8e802fa4464f491b98b1e0a081432e", size = 7024157, upload-time = "2026-06-04T07:49:45.34Z" }, - { url = "https://files.pythonhosted.org/packages/a4/0e/73f6d6800b4f27655abd9e9f6aaeaefcddb2b946e4674efa2bab184a7f7b/pywin32-312-cp314-cp314-win_arm64.whl", hash = "sha256:9fce94568364e0155e6dfb781ac5d95903be8baf28670632beab1b523f300daa", size = 6839598, upload-time = "2026-06-04T07:49:47.613Z" }, - { url = "https://files.pythonhosted.org/packages/eb/61/caa39686032d2ebdd04ff0ab5cbe163126c0066d98e00c9018646e42393b/pywin32-312-cp315-cp315-win32.whl", hash = "sha256:5c1fbe4a937a73ae9297384a3da38518cbc694c68ad8a809b2e19acd350f03ed", size = 6471159, upload-time = "2026-06-04T07:49:50.035Z" }, - { url = "https://files.pythonhosted.org/packages/0f/cd/7e1de64a4a6f69c04214169657ccab0d93a670ea50e35eb8f489d7378249/pywin32-312-cp315-cp315-win_amd64.whl", hash = "sha256:c2f03a0f73f804a13c2735b99392b0cd426bb4f2c4d0178e5ac966a0f21618d5", size = 7025293, upload-time = "2026-06-04T07:49:54.857Z" }, - { url = "https://files.pythonhosted.org/packages/23/ed/4532e9388e65fa16b46776ef47ad631a64eda1631884488af707666350ed/pywin32-312-cp315-cp315-win_arm64.whl", hash = "sha256:a8597d28f267b39074aef51fa593530082b39cbe5a074226096857b1fed2dfb9", size = 6840337, upload-time = "2026-06-04T07:49:57.531Z" }, -] - -[[package]] -name = "pyyaml" -version = "6.0.3" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/05/8e/961c0007c59b8dd7729d542c61a4d537767a59645b82a0b521206e1e25c2/pyyaml-6.0.3.tar.gz", hash = "sha256:d76623373421df22fb4cf8817020cbb7ef15c725b9d5e45f17e189bfc384190f", size = 130960, upload-time = "2025-09-25T21:33:16.546Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/f4/a0/39350dd17dd6d6c6507025c0e53aef67a9293a6d37d3511f23ea510d5800/pyyaml-6.0.3-cp310-cp310-macosx_10_13_x86_64.whl", hash = "sha256:214ed4befebe12df36bcc8bc2b64b396ca31be9304b8f59e25c11cf94a4c033b", size = 184227, upload-time = "2025-09-25T21:31:46.04Z" }, - { url = "https://files.pythonhosted.org/packages/05/14/52d505b5c59ce73244f59c7a50ecf47093ce4765f116cdb98286a71eeca2/pyyaml-6.0.3-cp310-cp310-macosx_11_0_arm64.whl", hash = "sha256:02ea2dfa234451bbb8772601d7b8e426c2bfa197136796224e50e35a78777956", size = 174019, upload-time = "2025-09-25T21:31:47.706Z" }, - { url = "https://files.pythonhosted.org/packages/43/f7/0e6a5ae5599c838c696adb4e6330a59f463265bfa1e116cfd1fbb0abaaae/pyyaml-6.0.3-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:b30236e45cf30d2b8e7b3e85881719e98507abed1011bf463a8fa23e9c3e98a8", size = 740646, upload-time = "2025-09-25T21:31:49.21Z" }, - { url = "https://files.pythonhosted.org/packages/2f/3a/61b9db1d28f00f8fd0ae760459a5c4bf1b941baf714e207b6eb0657d2578/pyyaml-6.0.3-cp310-cp310-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:66291b10affd76d76f54fad28e22e51719ef9ba22b29e1d7d03d6777a9174198", size = 840793, upload-time = "2025-09-25T21:31:50.735Z" }, - { url = "https://files.pythonhosted.org/packages/7a/1e/7acc4f0e74c4b3d9531e24739e0ab832a5edf40e64fbae1a9c01941cabd7/pyyaml-6.0.3-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:9c7708761fccb9397fe64bbc0395abcae8c4bf7b0eac081e12b809bf47700d0b", size = 770293, upload-time = "2025-09-25T21:31:51.828Z" }, - { url = "https://files.pythonhosted.org/packages/8b/ef/abd085f06853af0cd59fa5f913d61a8eab65d7639ff2a658d18a25d6a89d/pyyaml-6.0.3-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:418cf3f2111bc80e0933b2cd8cd04f286338bb88bdc7bc8e6dd775ebde60b5e0", size = 732872, upload-time = "2025-09-25T21:31:53.282Z" }, - { url = "https://files.pythonhosted.org/packages/1f/15/2bc9c8faf6450a8b3c9fc5448ed869c599c0a74ba2669772b1f3a0040180/pyyaml-6.0.3-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:5e0b74767e5f8c593e8c9b5912019159ed0533c70051e9cce3e8b6aa699fcd69", size = 758828, upload-time = "2025-09-25T21:31:54.807Z" }, - { url = "https://files.pythonhosted.org/packages/a3/00/531e92e88c00f4333ce359e50c19b8d1de9fe8d581b1534e35ccfbc5f393/pyyaml-6.0.3-cp310-cp310-win32.whl", hash = "sha256:28c8d926f98f432f88adc23edf2e6d4921ac26fb084b028c733d01868d19007e", size = 142415, upload-time = "2025-09-25T21:31:55.885Z" }, - { url = "https://files.pythonhosted.org/packages/2a/fa/926c003379b19fca39dd4634818b00dec6c62d87faf628d1394e137354d4/pyyaml-6.0.3-cp310-cp310-win_amd64.whl", hash = "sha256:bdb2c67c6c1390b63c6ff89f210c8fd09d9a1217a465701eac7316313c915e4c", size = 158561, upload-time = "2025-09-25T21:31:57.406Z" }, - { url = "https://files.pythonhosted.org/packages/6d/16/a95b6757765b7b031c9374925bb718d55e0a9ba8a1b6a12d25962ea44347/pyyaml-6.0.3-cp311-cp311-macosx_10_13_x86_64.whl", hash = "sha256:44edc647873928551a01e7a563d7452ccdebee747728c1080d881d68af7b997e", size = 185826, upload-time = "2025-09-25T21:31:58.655Z" }, - { url = "https://files.pythonhosted.org/packages/16/19/13de8e4377ed53079ee996e1ab0a9c33ec2faf808a4647b7b4c0d46dd239/pyyaml-6.0.3-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:652cb6edd41e718550aad172851962662ff2681490a8a711af6a4d288dd96824", size = 175577, upload-time = "2025-09-25T21:32:00.088Z" }, - { url = "https://files.pythonhosted.org/packages/0c/62/d2eb46264d4b157dae1275b573017abec435397aa59cbcdab6fc978a8af4/pyyaml-6.0.3-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:10892704fc220243f5305762e276552a0395f7beb4dbf9b14ec8fd43b57f126c", size = 775556, upload-time = "2025-09-25T21:32:01.31Z" }, - { url = "https://files.pythonhosted.org/packages/10/cb/16c3f2cf3266edd25aaa00d6c4350381c8b012ed6f5276675b9eba8d9ff4/pyyaml-6.0.3-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:850774a7879607d3a6f50d36d04f00ee69e7fc816450e5f7e58d7f17f1ae5c00", size = 882114, upload-time = "2025-09-25T21:32:03.376Z" }, - { url = "https://files.pythonhosted.org/packages/71/60/917329f640924b18ff085ab889a11c763e0b573da888e8404ff486657602/pyyaml-6.0.3-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:b8bb0864c5a28024fac8a632c443c87c5aa6f215c0b126c449ae1a150412f31d", size = 806638, upload-time = "2025-09-25T21:32:04.553Z" }, - { url = "https://files.pythonhosted.org/packages/dd/6f/529b0f316a9fd167281a6c3826b5583e6192dba792dd55e3203d3f8e655a/pyyaml-6.0.3-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:1d37d57ad971609cf3c53ba6a7e365e40660e3be0e5175fa9f2365a379d6095a", size = 767463, upload-time = "2025-09-25T21:32:06.152Z" }, - { url = "https://files.pythonhosted.org/packages/f2/6a/b627b4e0c1dd03718543519ffb2f1deea4a1e6d42fbab8021936a4d22589/pyyaml-6.0.3-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:37503bfbfc9d2c40b344d06b2199cf0e96e97957ab1c1b546fd4f87e53e5d3e4", size = 794986, upload-time = "2025-09-25T21:32:07.367Z" }, - { url = "https://files.pythonhosted.org/packages/45/91/47a6e1c42d9ee337c4839208f30d9f09caa9f720ec7582917b264defc875/pyyaml-6.0.3-cp311-cp311-win32.whl", hash = "sha256:8098f252adfa6c80ab48096053f512f2321f0b998f98150cea9bd23d83e1467b", size = 142543, upload-time = "2025-09-25T21:32:08.95Z" }, - { url = "https://files.pythonhosted.org/packages/da/e3/ea007450a105ae919a72393cb06f122f288ef60bba2dc64b26e2646fa315/pyyaml-6.0.3-cp311-cp311-win_amd64.whl", hash = "sha256:9f3bfb4965eb874431221a3ff3fdcddc7e74e3b07799e0e84ca4a0f867d449bf", size = 158763, upload-time = "2025-09-25T21:32:09.96Z" }, - { url = "https://files.pythonhosted.org/packages/d1/33/422b98d2195232ca1826284a76852ad5a86fe23e31b009c9886b2d0fb8b2/pyyaml-6.0.3-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:7f047e29dcae44602496db43be01ad42fc6f1cc0d8cd6c83d342306c32270196", size = 182063, upload-time = "2025-09-25T21:32:11.445Z" }, - { url = "https://files.pythonhosted.org/packages/89/a0/6cf41a19a1f2f3feab0e9c0b74134aa2ce6849093d5517a0c550fe37a648/pyyaml-6.0.3-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:fc09d0aa354569bc501d4e787133afc08552722d3ab34836a80547331bb5d4a0", size = 173973, upload-time = "2025-09-25T21:32:12.492Z" }, - { url = "https://files.pythonhosted.org/packages/ed/23/7a778b6bd0b9a8039df8b1b1d80e2e2ad78aa04171592c8a5c43a56a6af4/pyyaml-6.0.3-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:9149cad251584d5fb4981be1ecde53a1ca46c891a79788c0df828d2f166bda28", size = 775116, upload-time = "2025-09-25T21:32:13.652Z" }, - { url = "https://files.pythonhosted.org/packages/65/30/d7353c338e12baef4ecc1b09e877c1970bd3382789c159b4f89d6a70dc09/pyyaml-6.0.3-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:5fdec68f91a0c6739b380c83b951e2c72ac0197ace422360e6d5a959d8d97b2c", size = 844011, upload-time = "2025-09-25T21:32:15.21Z" }, - { url = "https://files.pythonhosted.org/packages/8b/9d/b3589d3877982d4f2329302ef98a8026e7f4443c765c46cfecc8858c6b4b/pyyaml-6.0.3-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:ba1cc08a7ccde2d2ec775841541641e4548226580ab850948cbfda66a1befcdc", size = 807870, upload-time = "2025-09-25T21:32:16.431Z" }, - { url = "https://files.pythonhosted.org/packages/05/c0/b3be26a015601b822b97d9149ff8cb5ead58c66f981e04fedf4e762f4bd4/pyyaml-6.0.3-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:8dc52c23056b9ddd46818a57b78404882310fb473d63f17b07d5c40421e47f8e", size = 761089, upload-time = "2025-09-25T21:32:17.56Z" }, - { url = "https://files.pythonhosted.org/packages/be/8e/98435a21d1d4b46590d5459a22d88128103f8da4c2d4cb8f14f2a96504e1/pyyaml-6.0.3-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:41715c910c881bc081f1e8872880d3c650acf13dfa8214bad49ed4cede7c34ea", size = 790181, upload-time = "2025-09-25T21:32:18.834Z" }, - { url = "https://files.pythonhosted.org/packages/74/93/7baea19427dcfbe1e5a372d81473250b379f04b1bd3c4c5ff825e2327202/pyyaml-6.0.3-cp312-cp312-win32.whl", hash = "sha256:96b533f0e99f6579b3d4d4995707cf36df9100d67e0c8303a0c55b27b5f99bc5", size = 137658, upload-time = "2025-09-25T21:32:20.209Z" }, - { url = "https://files.pythonhosted.org/packages/86/bf/899e81e4cce32febab4fb42bb97dcdf66bc135272882d1987881a4b519e9/pyyaml-6.0.3-cp312-cp312-win_amd64.whl", hash = "sha256:5fcd34e47f6e0b794d17de1b4ff496c00986e1c83f7ab2fb8fcfe9616ff7477b", size = 154003, upload-time = "2025-09-25T21:32:21.167Z" }, - { url = "https://files.pythonhosted.org/packages/1a/08/67bd04656199bbb51dbed1439b7f27601dfb576fb864099c7ef0c3e55531/pyyaml-6.0.3-cp312-cp312-win_arm64.whl", hash = "sha256:64386e5e707d03a7e172c0701abfb7e10f0fb753ee1d773128192742712a98fd", size = 140344, upload-time = "2025-09-25T21:32:22.617Z" }, - { url = "https://files.pythonhosted.org/packages/d1/11/0fd08f8192109f7169db964b5707a2f1e8b745d4e239b784a5a1dd80d1db/pyyaml-6.0.3-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:8da9669d359f02c0b91ccc01cac4a67f16afec0dac22c2ad09f46bee0697eba8", size = 181669, upload-time = "2025-09-25T21:32:23.673Z" }, - { url = "https://files.pythonhosted.org/packages/b1/16/95309993f1d3748cd644e02e38b75d50cbc0d9561d21f390a76242ce073f/pyyaml-6.0.3-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:2283a07e2c21a2aa78d9c4442724ec1eb15f5e42a723b99cb3d822d48f5f7ad1", size = 173252, upload-time = "2025-09-25T21:32:25.149Z" }, - { url = "https://files.pythonhosted.org/packages/50/31/b20f376d3f810b9b2371e72ef5adb33879b25edb7a6d072cb7ca0c486398/pyyaml-6.0.3-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:ee2922902c45ae8ccada2c5b501ab86c36525b883eff4255313a253a3160861c", size = 767081, upload-time = "2025-09-25T21:32:26.575Z" }, - { url = "https://files.pythonhosted.org/packages/49/1e/a55ca81e949270d5d4432fbbd19dfea5321eda7c41a849d443dc92fd1ff7/pyyaml-6.0.3-cp313-cp313-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:a33284e20b78bd4a18c8c2282d549d10bc8408a2a7ff57653c0cf0b9be0afce5", size = 841159, upload-time = "2025-09-25T21:32:27.727Z" }, - { url = "https://files.pythonhosted.org/packages/74/27/e5b8f34d02d9995b80abcef563ea1f8b56d20134d8f4e5e81733b1feceb2/pyyaml-6.0.3-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:0f29edc409a6392443abf94b9cf89ce99889a1dd5376d94316ae5145dfedd5d6", size = 801626, upload-time = "2025-09-25T21:32:28.878Z" }, - { url = "https://files.pythonhosted.org/packages/f9/11/ba845c23988798f40e52ba45f34849aa8a1f2d4af4b798588010792ebad6/pyyaml-6.0.3-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:f7057c9a337546edc7973c0d3ba84ddcdf0daa14533c2065749c9075001090e6", size = 753613, upload-time = "2025-09-25T21:32:30.178Z" }, - { url = "https://files.pythonhosted.org/packages/3d/e0/7966e1a7bfc0a45bf0a7fb6b98ea03fc9b8d84fa7f2229e9659680b69ee3/pyyaml-6.0.3-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:eda16858a3cab07b80edaf74336ece1f986ba330fdb8ee0d6c0d68fe82bc96be", size = 794115, upload-time = "2025-09-25T21:32:31.353Z" }, - { url = "https://files.pythonhosted.org/packages/de/94/980b50a6531b3019e45ddeada0626d45fa85cbe22300844a7983285bed3b/pyyaml-6.0.3-cp313-cp313-win32.whl", hash = "sha256:d0eae10f8159e8fdad514efdc92d74fd8d682c933a6dd088030f3834bc8e6b26", size = 137427, upload-time = "2025-09-25T21:32:32.58Z" }, - { url = "https://files.pythonhosted.org/packages/97/c9/39d5b874e8b28845e4ec2202b5da735d0199dbe5b8fb85f91398814a9a46/pyyaml-6.0.3-cp313-cp313-win_amd64.whl", hash = "sha256:79005a0d97d5ddabfeeea4cf676af11e647e41d81c9a7722a193022accdb6b7c", size = 154090, upload-time = "2025-09-25T21:32:33.659Z" }, - { url = "https://files.pythonhosted.org/packages/73/e8/2bdf3ca2090f68bb3d75b44da7bbc71843b19c9f2b9cb9b0f4ab7a5a4329/pyyaml-6.0.3-cp313-cp313-win_arm64.whl", hash = "sha256:5498cd1645aa724a7c71c8f378eb29ebe23da2fc0d7a08071d89469bf1d2defb", size = 140246, upload-time = "2025-09-25T21:32:34.663Z" }, - { url = "https://files.pythonhosted.org/packages/9d/8c/f4bd7f6465179953d3ac9bc44ac1a8a3e6122cf8ada906b4f96c60172d43/pyyaml-6.0.3-cp314-cp314-macosx_10_13_x86_64.whl", hash = "sha256:8d1fab6bb153a416f9aeb4b8763bc0f22a5586065f86f7664fc23339fc1c1fac", size = 181814, upload-time = "2025-09-25T21:32:35.712Z" }, - { url = "https://files.pythonhosted.org/packages/bd/9c/4d95bb87eb2063d20db7b60faa3840c1b18025517ae857371c4dd55a6b3a/pyyaml-6.0.3-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:34d5fcd24b8445fadc33f9cf348c1047101756fd760b4dacb5c3e99755703310", size = 173809, upload-time = "2025-09-25T21:32:36.789Z" }, - { url = "https://files.pythonhosted.org/packages/92/b5/47e807c2623074914e29dabd16cbbdd4bf5e9b2db9f8090fa64411fc5382/pyyaml-6.0.3-cp314-cp314-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:501a031947e3a9025ed4405a168e6ef5ae3126c59f90ce0cd6f2bfc477be31b7", size = 766454, upload-time = "2025-09-25T21:32:37.966Z" }, - { url = "https://files.pythonhosted.org/packages/02/9e/e5e9b168be58564121efb3de6859c452fccde0ab093d8438905899a3a483/pyyaml-6.0.3-cp314-cp314-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:b3bc83488de33889877a0f2543ade9f70c67d66d9ebb4ac959502e12de895788", size = 836355, upload-time = "2025-09-25T21:32:39.178Z" }, - { url = "https://files.pythonhosted.org/packages/88/f9/16491d7ed2a919954993e48aa941b200f38040928474c9e85ea9e64222c3/pyyaml-6.0.3-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:c458b6d084f9b935061bc36216e8a69a7e293a2f1e68bf956dcd9e6cbcd143f5", size = 794175, upload-time = "2025-09-25T21:32:40.865Z" }, - { url = "https://files.pythonhosted.org/packages/dd/3f/5989debef34dc6397317802b527dbbafb2b4760878a53d4166579111411e/pyyaml-6.0.3-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:7c6610def4f163542a622a73fb39f534f8c101d690126992300bf3207eab9764", size = 755228, upload-time = "2025-09-25T21:32:42.084Z" }, - { url = "https://files.pythonhosted.org/packages/d7/ce/af88a49043cd2e265be63d083fc75b27b6ed062f5f9fd6cdc223ad62f03e/pyyaml-6.0.3-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:5190d403f121660ce8d1d2c1bb2ef1bd05b5f68533fc5c2ea899bd15f4399b35", size = 789194, upload-time = "2025-09-25T21:32:43.362Z" }, - { url = "https://files.pythonhosted.org/packages/23/20/bb6982b26a40bb43951265ba29d4c246ef0ff59c9fdcdf0ed04e0687de4d/pyyaml-6.0.3-cp314-cp314-win_amd64.whl", hash = "sha256:4a2e8cebe2ff6ab7d1050ecd59c25d4c8bd7e6f400f5f82b96557ac0abafd0ac", size = 156429, upload-time = "2025-09-25T21:32:57.844Z" }, - { url = "https://files.pythonhosted.org/packages/f4/f4/a4541072bb9422c8a883ab55255f918fa378ecf083f5b85e87fc2b4eda1b/pyyaml-6.0.3-cp314-cp314-win_arm64.whl", hash = "sha256:93dda82c9c22deb0a405ea4dc5f2d0cda384168e466364dec6255b293923b2f3", size = 143912, upload-time = "2025-09-25T21:32:59.247Z" }, - { url = "https://files.pythonhosted.org/packages/7c/f9/07dd09ae774e4616edf6cda684ee78f97777bdd15847253637a6f052a62f/pyyaml-6.0.3-cp314-cp314t-macosx_10_13_x86_64.whl", hash = "sha256:02893d100e99e03eda1c8fd5c441d8c60103fd175728e23e431db1b589cf5ab3", size = 189108, upload-time = "2025-09-25T21:32:44.377Z" }, - { url = "https://files.pythonhosted.org/packages/4e/78/8d08c9fb7ce09ad8c38ad533c1191cf27f7ae1effe5bb9400a46d9437fcf/pyyaml-6.0.3-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:c1ff362665ae507275af2853520967820d9124984e0f7466736aea23d8611fba", size = 183641, upload-time = "2025-09-25T21:32:45.407Z" }, - { url = "https://files.pythonhosted.org/packages/7b/5b/3babb19104a46945cf816d047db2788bcaf8c94527a805610b0289a01c6b/pyyaml-6.0.3-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:6adc77889b628398debc7b65c073bcb99c4a0237b248cacaf3fe8a557563ef6c", size = 831901, upload-time = "2025-09-25T21:32:48.83Z" }, - { url = "https://files.pythonhosted.org/packages/8b/cc/dff0684d8dc44da4d22a13f35f073d558c268780ce3c6ba1b87055bb0b87/pyyaml-6.0.3-cp314-cp314t-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:a80cb027f6b349846a3bf6d73b5e95e782175e52f22108cfa17876aaeff93702", size = 861132, upload-time = "2025-09-25T21:32:50.149Z" }, - { url = "https://files.pythonhosted.org/packages/b1/5e/f77dc6b9036943e285ba76b49e118d9ea929885becb0a29ba8a7c75e29fe/pyyaml-6.0.3-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:00c4bdeba853cc34e7dd471f16b4114f4162dc03e6b7afcc2128711f0eca823c", size = 839261, upload-time = "2025-09-25T21:32:51.808Z" }, - { url = "https://files.pythonhosted.org/packages/ce/88/a9db1376aa2a228197c58b37302f284b5617f56a5d959fd1763fb1675ce6/pyyaml-6.0.3-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:66e1674c3ef6f541c35191caae2d429b967b99e02040f5ba928632d9a7f0f065", size = 805272, upload-time = "2025-09-25T21:32:52.941Z" }, - { url = "https://files.pythonhosted.org/packages/da/92/1446574745d74df0c92e6aa4a7b0b3130706a4142b2d1a5869f2eaa423c6/pyyaml-6.0.3-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:16249ee61e95f858e83976573de0f5b2893b3677ba71c9dd36b9cf8be9ac6d65", size = 829923, upload-time = "2025-09-25T21:32:54.537Z" }, - { url = "https://files.pythonhosted.org/packages/f0/7a/1c7270340330e575b92f397352af856a8c06f230aa3e76f86b39d01b416a/pyyaml-6.0.3-cp314-cp314t-win_amd64.whl", hash = "sha256:4ad1906908f2f5ae4e5a8ddfce73c320c2a1429ec52eafd27138b7f1cbe341c9", size = 174062, upload-time = "2025-09-25T21:32:55.767Z" }, - { url = "https://files.pythonhosted.org/packages/f1/12/de94a39c2ef588c7e6455cfbe7343d3b2dc9d6b6b2f40c4c6565744c873d/pyyaml-6.0.3-cp314-cp314t-win_arm64.whl", hash = "sha256:ebc55a14a21cb14062aa4162f906cd962b28e2e9ea38f9b4391244cd8de4ae0b", size = 149341, upload-time = "2025-09-25T21:32:56.828Z" }, -] - -[[package]] -name = "referencing" -version = "0.37.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "attrs" }, - { name = "rpds-py", version = "0.30.0", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, - { name = "rpds-py", version = "2026.5.1", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.11'" }, - { name = "typing-extensions", marker = "python_full_version < '3.13'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/22/f5/df4e9027acead3ecc63e50fe1e36aca1523e1719559c499951bb4b53188f/referencing-0.37.0.tar.gz", hash = "sha256:44aefc3142c5b842538163acb373e24cce6632bd54bdb01b21ad5863489f50d8", size = 78036, upload-time = "2025-10-13T15:30:48.871Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/2c/58/ca301544e1fa93ed4f80d724bf5b194f6e4b945841c5bfd555878eea9fcb/referencing-0.37.0-py3-none-any.whl", hash = "sha256:381329a9f99628c9069361716891d34ad94af76e461dcb0335825aecc7692231", size = 26766, upload-time = "2025-10-13T15:30:47.625Z" }, -] - -[[package]] -name = "regex" -version = "2026.5.9" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/dc/0e/49aee608ad09480e7fd276898c99ec6192985fa331abe4eb3a986094490b/regex-2026.5.9.tar.gz", hash = "sha256:a8234aa23ec39894bfe4a3f1b85616a7032481964a13ac6fc9f10de4f6fca270", size = 416074, upload-time = "2026-05-09T23:15:19.37Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/fe/ed/0ad2c8edf634918eb4484365d3819fa7bd7f58daf807fe7fb21812c316e5/regex-2026.5.9-cp310-cp310-macosx_10_9_universal2.whl", hash = "sha256:a9e1328e17c84c1a5d22ec9f785ecef4a967fab9a42b6a8dc3bcbebd0a0c9e44", size = 489438, upload-time = "2026-05-09T23:11:29.374Z" }, - { url = "https://files.pythonhosted.org/packages/89/a9/4ed972ad263963b860b7c3e86e0e1bcc791def47b43b8c8efe57e710f139/regex-2026.5.9-cp310-cp310-macosx_10_9_x86_64.whl", hash = "sha256:bfe1ce50cbfb569d74e1e4337da6468961f31dbea55fd85aa5de59c0947a805a", size = 291270, upload-time = "2026-05-09T23:11:33.254Z" }, - { url = "https://files.pythonhosted.org/packages/16/81/075930d9fa28c4ea1f53398dd015ee7c882f623539759113cda1257f4b82/regex-2026.5.9-cp310-cp310-macosx_11_0_arm64.whl", hash = "sha256:15ee42209947f4ca045412eae98416317238163618ace2a8e54f99586a466733", size = 289198, upload-time = "2026-05-09T23:11:35.769Z" }, - { url = "https://files.pythonhosted.org/packages/d4/c8/5cdfbf0b5dc6599e1b6131eff43262e5275d4ec3469ce10216061659aadb/regex-2026.5.9-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:b4bb445ff3f725f59df8f6014edb547ee928ec7023a774f6a39a3f953038cbb2", size = 784765, upload-time = "2026-05-09T23:11:37.689Z" }, - { url = "https://files.pythonhosted.org/packages/cd/ca/ae5fd6edc59b7f84b904b31d6ec39a860cbcecd10f64bd5a062ca83a4864/regex-2026.5.9-cp310-cp310-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:446ddd671e43ab535810c4b21cff7104945c701d4a14d1e6d1cd6f4e445a8bea", size = 852115, upload-time = "2026-05-09T23:11:39.973Z" }, - { url = "https://files.pythonhosted.org/packages/f6/ce/a91cf555afb51f3b74a182e24ba073b91ea7bb64592fc4b315c111bb19fd/regex-2026.5.9-cp310-cp310-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:7b92817338591505f282cf3864c145244b1edcf5381d237038df955001091538", size = 899503, upload-time = "2026-05-09T23:11:42.48Z" }, - { url = "https://files.pythonhosted.org/packages/55/7f/725a0a2b245a4cf0c4bab29d0e97c74285d94136a65d1b55a6459a583502/regex-2026.5.9-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:d6b8a143aca6c39b446ea8092cde25cc8fe9304d4f5fecfbc1a9dbb0282703c2", size = 794093, upload-time = "2026-05-09T23:11:44.681Z" }, - { url = "https://files.pythonhosted.org/packages/e3/2a/996efbd59ce6b5d4a09e3af6180ceb62af171f4a9a6fb557d2f0ae0d462b/regex-2026.5.9-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:0f03aa6898aaaac4592479821df16e68e8d0e29e903e65d8f2dfb2f19028a989", size = 786234, upload-time = "2026-05-09T23:11:46.882Z" }, - { url = "https://files.pythonhosted.org/packages/4b/0a/8731e8b8806174c9cdd5903f80a14990331c1f42fc4209b540952e9e010d/regex-2026.5.9-cp310-cp310-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:ed457d8e98ae812ed7732bef7bf78de78e834eae0372a74e23ca90ef21d910f9", size = 769895, upload-time = "2026-05-09T23:11:49.324Z" }, - { url = "https://files.pythonhosted.org/packages/9a/0b/932473194bd563f342a412ae2ffbbd6da608306a2bc4e99249a41c2b0b92/regex-2026.5.9-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:71b61c5bfe1c806332defc42ad6c780b3c55f661986d7f40283a3a88274b4c00", size = 774991, upload-time = "2026-05-09T23:11:51.261Z" }, - { url = "https://files.pythonhosted.org/packages/98/80/9523d196010031df25f7177ee0a467efbee436324038e5d99def17a57515/regex-2026.5.9-cp310-cp310-musllinux_1_2_ppc64le.whl", hash = "sha256:3b1e39888c5e0c7d92cea4fc777396c4a90363b05de75d02eb459a4752200808", size = 848790, upload-time = "2026-05-09T23:11:53.232Z" }, - { url = "https://files.pythonhosted.org/packages/3c/07/56987b35e89edf47e4a38cf2845aeee476bfa688a6bdbd3e820cda461dc1/regex-2026.5.9-cp310-cp310-musllinux_1_2_riscv64.whl", hash = "sha256:6ba42b2e7e7f46cf68cc6a5ca36fa07959f9bbd9c6bdcc47b6ee76549a590248", size = 757679, upload-time = "2026-05-09T23:11:55.82Z" }, - { url = "https://files.pythonhosted.org/packages/04/2a/ff713fff0c566507c06a4ce2dc0ae8e7eeebc88811a95fc81cf1e7d534dd/regex-2026.5.9-cp310-cp310-musllinux_1_2_s390x.whl", hash = "sha256:c010eb8caca74bdb40c07498d7ece26b4428fd3f04aa8a72c9ac6f79e8faaac6", size = 837116, upload-time = "2026-05-09T23:11:57.934Z" }, - { url = "https://files.pythonhosted.org/packages/77/90/df6d982b03e3614785c6937ba51b57f6733d97d2ee1c9bc7531dbfab3a54/regex-2026.5.9-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:a6a563446a41adc451393dc6b8e6ad87979efaee3c8738690a8d1b08ebead1b4", size = 782081, upload-time = "2026-05-09T23:11:59.607Z" }, - { url = "https://files.pythonhosted.org/packages/c7/8a/4e88a5f7c3e98489aac4dd23142723d907b2a595b4a6abcbacabefeded09/regex-2026.5.9-cp310-cp310-win32.whl", hash = "sha256:954cc214c04663ee6d266fc61739cad83054683048de65c5bd1d640ad28098ac", size = 266247, upload-time = "2026-05-09T23:12:01.116Z" }, - { url = "https://files.pythonhosted.org/packages/6a/40/4b224cb0582b2dca1786726e6cdabe26abbf757d7f6718332f186da155d2/regex-2026.5.9-cp310-cp310-win_amd64.whl", hash = "sha256:b310768746dd314ea6e2ff4cc89ef215426813396ff4e94ee8e6f7096c8b6e03", size = 278416, upload-time = "2026-05-09T23:12:03.2Z" }, - { url = "https://files.pythonhosted.org/packages/12/4d/014fbe803204cab0947ee428f09f658a29632053dde1d3c6176bb4f0fd4c/regex-2026.5.9-cp310-cp310-win_arm64.whl", hash = "sha256:19c16ceb4a267a8789e25733e583983eeab9f0f8664e66b0bd1c5d21f14c2d4b", size = 270413, upload-time = "2026-05-09T23:12:04.649Z" }, - { url = "https://files.pythonhosted.org/packages/c2/dc/c1f2df4027e82fc54b5a473e4b250f5139faca49a0fbe29a48668d228f34/regex-2026.5.9-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:ccf5249114cc3e772ecdd88a98a86eca0fd74c61ce32a94743758c083fc05d48", size = 489445, upload-time = "2026-05-09T23:12:06.111Z" }, - { url = "https://files.pythonhosted.org/packages/03/d2/59f01110660081cce9c0bc30ebd0b5ee250dacf658e3248ed92f01e0e8ee/regex-2026.5.9-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:46f1326ca6e65b0879d23ca302c0f2415aad42ff0309b9c818e7949fe19a41d8", size = 291271, upload-time = "2026-05-09T23:12:07.731Z" }, - { url = "https://files.pythonhosted.org/packages/58/b6/14b2c84ff90ddb370c81d27503f4a0fcf071496416f4855f6cc8c5d81c35/regex-2026.5.9-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:ef31cbfe458e21c6122ba8150ff060e0c7789ed0d26eb423f25472584920b555", size = 289212, upload-time = "2026-05-09T23:12:09.266Z" }, - { url = "https://files.pythonhosted.org/packages/03/d0/4db86529117320de0c84afd90e70bb47434625875e34fcef9d8c127c5b16/regex-2026.5.9-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:992604d02e6d9c6d786c24a706a71ecffe1020fc1ef264044474cd81fa2c3919", size = 792310, upload-time = "2026-05-09T23:12:11.416Z" }, - { url = "https://files.pythonhosted.org/packages/07/78/fe4800cd322f862ecffd2d553409b20d80650e5ed71b9d178f853d020b82/regex-2026.5.9-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:c9411dd64ca95477225734a93dfc8583b51916b8d5942f99d6cac21e09965451", size = 861721, upload-time = "2026-05-09T23:12:13.681Z" }, - { url = "https://files.pythonhosted.org/packages/b5/d0/b3618a895dd8feb897c61bb2954edd265e1767d82a01d53065d5871127a3/regex-2026.5.9-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:3dd4a3ff360dfb836fecdb93a4598f9d6e2ac81e3e397125145c6221bf58cf4c", size = 906460, upload-time = "2026-05-09T23:12:15.443Z" }, - { url = "https://files.pythonhosted.org/packages/33/6f/1481597e859ef19508b345eec4afd1416ed6e6b459c75a64026ef193aecf/regex-2026.5.9-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:2a661a7d270a61f7cf460caee8b9fa2d5ef9e5c681234bcb9e0fe14f488e7dfc", size = 799843, upload-time = "2026-05-09T23:12:16.892Z" }, - { url = "https://files.pythonhosted.org/packages/73/59/955734c803f59108deccba3597ae440c76b62a652733c0006e6243758420/regex-2026.5.9-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:f079e50a0d3cc3cd5091fa9ff45869a2e6b2cd35895731edafb0327901a8d86d", size = 773610, upload-time = "2026-05-09T23:12:19.127Z" }, - { url = "https://files.pythonhosted.org/packages/68/8f/70c04a236d651c81881dac42ef8538bddda6121434509d0a22d9e601503b/regex-2026.5.9-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:4ebe8f0b5ec5a5024dc4a4c59f444c4e9afc5f2abdbb8962065b75d27fb971f9", size = 781645, upload-time = "2026-05-09T23:12:20.806Z" }, - { url = "https://files.pythonhosted.org/packages/1d/96/05c7434d88185e5d27fe54aeb74df86bd77cd79f52f0b4eae54faa8fea70/regex-2026.5.9-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:97cf3bc1b7d7d2306772ec07366c80d9df00ff79e79cea32898883a646d2fae2", size = 854473, upload-time = "2026-05-09T23:12:22.465Z" }, - { url = "https://files.pythonhosted.org/packages/4e/c1/6e3d8202d981f3117004bf341ee74893ba4ba8a9fbaf4b94615846550a08/regex-2026.5.9-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:0f9eede6a5cbdc02d4978090186390936e1776a7d1359b21e41014c609880bcf", size = 763311, upload-time = "2026-05-09T23:12:24.351Z" }, - { url = "https://files.pythonhosted.org/packages/93/c7/e7737f1526b3fb32bd4c337fd6c71c3ebb5c8296fc34d11197e0955d2e35/regex-2026.5.9-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:01f0f5f55f4b64dacec85dc116d3c05fd23ad3ff037bbc73a2085775953c2611", size = 844593, upload-time = "2026-05-09T23:12:26.341Z" }, - { url = "https://files.pythonhosted.org/packages/a5/27/0daffb1a535bb39f422c3d200f4ab023c71110ad66a32b366bee708baba0/regex-2026.5.9-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:1268eddd8486dc561d08eee1156e40aa3a8fe10f4bdec8fa653b455fcbffd12c", size = 789167, upload-time = "2026-05-09T23:12:27.975Z" }, - { url = "https://files.pythonhosted.org/packages/ce/fc/294fe4fac4f2ed67207b17471815870c1c45b3a489e08e0ac96daea16ef6/regex-2026.5.9-cp311-cp311-win32.whl", hash = "sha256:8676474c07469d6f33dd1085ca2cd45f65785f32518f2b20e36d9953ca07f994", size = 266249, upload-time = "2026-05-09T23:12:30.141Z" }, - { url = "https://files.pythonhosted.org/packages/d0/b0/8dce459f6245bcf8f6e9f23ac9569f1a0f15c131cc0745e82b43226204cf/regex-2026.5.9-cp311-cp311-win_amd64.whl", hash = "sha256:246de9d60aa3f8538b519834dd95cbf276ea263d6a7bd5a3666dc3fa0230505b", size = 278423, upload-time = "2026-05-09T23:12:31.676Z" }, - { url = "https://files.pythonhosted.org/packages/db/8d/f9aeff6ad63a3ef720386f2907e6d34a35a510a6e498ebad28b0fb3f6ab6/regex-2026.5.9-cp311-cp311-win_arm64.whl", hash = "sha256:d726ca3f0d76969bf1e8e477d160d3d666bbf999f6860bd314889e5345782046", size = 270420, upload-time = "2026-05-09T23:12:33.194Z" }, - { url = "https://files.pythonhosted.org/packages/50/9b/6550044bc44e17c84d312c031c2ec42fbdb6a4ec4e29093be3a172d08772/regex-2026.5.9-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:57eeeb05db7979413dec5438f2db21d7ecbba787cde7a711df1a6f6df672aa06", size = 490451, upload-time = "2026-05-09T23:12:34.72Z" }, - { url = "https://files.pythonhosted.org/packages/1e/95/fc7ba4303b5a0f92446a12ee6778ef2c6c799233f5060042a31bf390cfe9/regex-2026.5.9-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:398c521292f4c7fb807001dcd54694d3a1fcafc179a36ad9cc56f98df85930b6", size = 292112, upload-time = "2026-05-09T23:12:36.285Z" }, - { url = "https://files.pythonhosted.org/packages/54/4b/ee27938d1b2c443e89a9a10e00d2d19aa5ee300cd3d61140644e93bb083e/regex-2026.5.9-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:f7a7c26137296beba7784de6eba69c6a93a63ccebc385e4962fe67e267a91225", size = 289599, upload-time = "2026-05-09T23:12:38.089Z" }, - { url = "https://files.pythonhosted.org/packages/d8/dd/ba103dc19614e25f3880800ca67ce093d6e21b325d72b8383c7bf906e9fa/regex-2026.5.9-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:6441cc660d76107934a09c22167200839a0e89604a6297f78a974e66e931d2c0", size = 796732, upload-time = "2026-05-09T23:12:40.062Z" }, - { url = "https://files.pythonhosted.org/packages/cf/e7/f035b4fd858b050b0080bf302968dc0f59ba34e391872d54936758e6844e/regex-2026.5.9-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:91328f1c23d47595ca3ef0a7557fa129c5a23404b775c770697d2f35b33e0107", size = 865440, upload-time = "2026-05-09T23:12:42.059Z" }, - { url = "https://files.pythonhosted.org/packages/0a/51/8cd301ecc899aea28124357f729f4272f44de7806fc7ca02490bfbe253e8/regex-2026.5.9-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:93a7860539414dddaefba2b40f8771765ae17949d4c7182b876ce429e11a8309", size = 912329, upload-time = "2026-05-09T23:12:44.373Z" }, - { url = "https://files.pythonhosted.org/packages/cc/1e/3fbe2fa1e8cebd62f3bb7d3321cff1640aca2e240b51d9bd624aad949260/regex-2026.5.9-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:dd2810d22146b6d838acc5ec15602cb6b47920aa4e33015df3868eedfd20bab8", size = 801239, upload-time = "2026-05-09T23:12:46.268Z" }, - { url = "https://files.pythonhosted.org/packages/17/2f/6f6008682bf2cf98040a0d3153a8e557b6ab728d7713d045cee4ce544ab8/regex-2026.5.9-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:daff2bdbaf1d23e52fdff7c0b7bc2048b68f978df6a4d107ac981f94caef2e66", size = 777054, upload-time = "2026-05-09T23:12:48.051Z" }, - { url = "https://files.pythonhosted.org/packages/19/2b/eee0d20a6842ba04df4b8847a920b57ef56853f14ef85405473e586b605a/regex-2026.5.9-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:4eeb011098fcb77af513dcef521a3dbecbf8849b1e38940759d293b7a93f5026", size = 785098, upload-time = "2026-05-09T23:12:49.851Z" }, - { url = "https://files.pythonhosted.org/packages/4a/98/6fc1e6410feefb92159edaed5041992bfe390e8d26c721865434acbca558/regex-2026.5.9-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:ea9c8ecfa1b73c73b626534d6626e5340d429630943672b8480724f44e84b962", size = 860095, upload-time = "2026-05-09T23:12:51.666Z" }, - { url = "https://files.pythonhosted.org/packages/18/a3/bd855e0f2cb1a978ecf6fa6bb69632dd9c3f6ea3b81cde62fde14c9daec7/regex-2026.5.9-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:cd2846168eb9ee3c513902bc8225409cb1caab31d04728b145171fa1625d9621", size = 765762, upload-time = "2026-05-09T23:12:53.413Z" }, - { url = "https://files.pythonhosted.org/packages/dc/66/0ae8c092e60b14c79d24f8e0b7f0aea5bfbffdcab00b5483d13404d3c3a5/regex-2026.5.9-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:39617fb0cde9c0e6306dc70e3bfc096f3da793219879f7ae7aa341a69fbdcf6d", size = 852100, upload-time = "2026-05-09T23:12:55.256Z" }, - { url = "https://files.pythonhosted.org/packages/21/de/8dfde60fc1b21c946a893ba273403b72617edb261370cb1087099a83f088/regex-2026.5.9-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:fd03c4f0e33280d15cae17159b899245d6b7c53d21def19b263b39655061f5ce", size = 789479, upload-time = "2026-05-09T23:12:57.573Z" }, - { url = "https://files.pythonhosted.org/packages/c3/1c/bdcc98f9a4af4fdd166c74941174619ccff4726d3ce32faa8e9a2ecd38dd/regex-2026.5.9-cp312-cp312-win32.whl", hash = "sha256:164eba9b755ea6f244b0d881196fbc1fac09714e9782c9e2732b813142033c8e", size = 266699, upload-time = "2026-05-09T23:12:59.14Z" }, - { url = "https://files.pythonhosted.org/packages/78/87/240d36864f9e48ace85f72e79ced97ceb7f27ce87739a947dcb834b4e6bc/regex-2026.5.9-cp312-cp312-win_amd64.whl", hash = "sha256:86f40a5d6444db30a125c9c9177e6b25dad981cbc37451fd838f145e6edac92e", size = 277783, upload-time = "2026-05-09T23:13:00.789Z" }, - { url = "https://files.pythonhosted.org/packages/4f/b5/7b30f312b0669dff5beebe5b0989dc2d1a312b1a44fab852199c387a5b96/regex-2026.5.9-cp312-cp312-win_arm64.whl", hash = "sha256:96f5f58b54a063d7ea9dca08e1cf57bfe10499c4d579ee672da284f57f5f0070", size = 270513, upload-time = "2026-05-09T23:13:02.426Z" }, - { url = "https://files.pythonhosted.org/packages/aa/da/797e91ecec6f84135da778ddce78c20e0af5d2a15c26f87a81bc3eadb6db/regex-2026.5.9-cp313-cp313-macosx_10_13_universal2.whl", hash = "sha256:d626b84406444b165fc0ba981604edea39f0588ff1f92baa23fe50799ea9afdb", size = 490303, upload-time = "2026-05-09T23:13:04.382Z" }, - { url = "https://files.pythonhosted.org/packages/44/da/bf30abaaa737b58f4a4b8c4a03659e02fd92092c822e0197ed9e0daab917/regex-2026.5.9-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:d7bdc0ab8f3dd7e1b4f9ab88634e13374669db86bb3c72e8292f07ae313f539f", size = 292019, upload-time = "2026-05-09T23:13:06.022Z" }, - { url = "https://files.pythonhosted.org/packages/2d/e7/d0eaf5713828417b9e5648cf81fa9bacd4961f6ab98c380c2034f8716e35/regex-2026.5.9-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:a8820737949116ffff55fe18f9fc644530063ba6ebfcb8314239416e78f1347c", size = 289468, upload-time = "2026-05-09T23:13:08.214Z" }, - { url = "https://files.pythonhosted.org/packages/d3/9b/b3fdd62b003baa1a9b593cd8c8699c9651c2e80cc21a5c715707983c42d7/regex-2026.5.9-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:aa0fbdbac82cb3e4450d0ccde7d7a35607f4cb2dd9fba4b8b69bfaf8c9fa6aed", size = 796749, upload-time = "2026-05-09T23:13:10.573Z" }, - { url = "https://files.pythonhosted.org/packages/d4/30/66ab84588765f5b4b271a9ca09ef7ce2b87caa95176ec3d2ad65d7bc4902/regex-2026.5.9-cp313-cp313-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:57e8915c7986aa33d25e4d3629cef711cd2863f2961b10409f0c04cb8b7d9020", size = 865445, upload-time = "2026-05-09T23:13:12.523Z" }, - { url = "https://files.pythonhosted.org/packages/1a/89/f05169e8588aac365f35ffc7f3bc3184f095ef4cfded7cfaa3c7fd5dbd89/regex-2026.5.9-cp313-cp313-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:508f56a89ba9cb26e4168cbc37dbd60a28d82430a9e18ad1d25fe0883c314ca2", size = 912322, upload-time = "2026-05-09T23:13:14.281Z" }, - { url = "https://files.pythonhosted.org/packages/30/e1/c93444052cf41581f3c884ab3fb5823daf0992f11cd4388d4275ca610558/regex-2026.5.9-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:b6d189041f15691cfa2b6c4290448ec221244d225b3f5fe9e7771b34ffcdf6e2", size = 801269, upload-time = "2026-05-09T23:13:16.569Z" }, - { url = "https://files.pythonhosted.org/packages/50/fe/0cf96b882f540e62e8b9956599798203d599c44cf4c77917ca27400ff69b/regex-2026.5.9-cp313-cp313-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:e82db382b44d0111b22601c509c89f64434816c9e0eef9d1989cda8cc6ff1c04", size = 777085, upload-time = "2026-05-09T23:13:18.675Z" }, - { url = "https://files.pythonhosted.org/packages/23/5c/d78d4924e7fc875557b9e9b768423925fdfaac5549d06da7810019a9bd26/regex-2026.5.9-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:2acfb48634f64996b57f90f39afa692ff362162722581921fe92239a59960f3c", size = 785153, upload-time = "2026-05-09T23:13:20.525Z" }, - { url = "https://files.pythonhosted.org/packages/bf/e0/5214774090e7b4524dcea3e3c4aa74141d43043f8beb49c1599db1c8b53a/regex-2026.5.9-cp313-cp313-musllinux_1_2_ppc64le.whl", hash = "sha256:d29eebfc9525db68cad3c97eedd7f754fa265aa5cd0cf4f863b2421e1b48fc9f", size = 860164, upload-time = "2026-05-09T23:13:22.263Z" }, - { url = "https://files.pythonhosted.org/packages/6e/e1/4a57a83350319b1271f0d7a249b8672513ed928b237a741631270de6caea/regex-2026.5.9-cp313-cp313-musllinux_1_2_riscv64.whl", hash = "sha256:debb893095e944091c16e641a6e33c1b0f4cb61ab945ec5afbf53ce7068834d8", size = 765731, upload-time = "2026-05-09T23:13:24.277Z" }, - { url = "https://files.pythonhosted.org/packages/12/f4/499e74a20c156fc75836ee04a72a38d1a063978f600937f9760467beb1b0/regex-2026.5.9-cp313-cp313-musllinux_1_2_s390x.whl", hash = "sha256:d659eee77986549c9ea45b861c7567e44d6287c3dc9a4565478853f7b9fe2ff6", size = 852062, upload-time = "2026-05-09T23:13:26.125Z" }, - { url = "https://files.pythonhosted.org/packages/5b/92/7eebc0d0a01e78629695f342ba17e0deaff8fb45e79cc0d7b98287da6e3e/regex-2026.5.9-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:2efa205e6d98b24d1f3ab395c11aa15cdf10935bca283d0285e0499c284fba21", size = 789577, upload-time = "2026-05-09T23:13:27.814Z" }, - { url = "https://files.pythonhosted.org/packages/05/a4/018e71f7d2ad48c1ebe6d3ae0026f9b7cb4802fd15c7cc02fdf724355102/regex-2026.5.9-cp313-cp313-win32.whl", hash = "sha256:f3844f134e834076677dd369976e9f5068679fcb8e50102fdf6b7ac96a3ec127", size = 266691, upload-time = "2026-05-09T23:13:29.549Z" }, - { url = "https://files.pythonhosted.org/packages/e6/1d/861a93719fb9ee7dbfc3761b3797b7a3e112a5d42c6129459d2d741be9b5/regex-2026.5.9-cp313-cp313-win_amd64.whl", hash = "sha256:3527bb4942d2c14552155406cdedd906567456821848aed1cb4933a391bf5eca", size = 277747, upload-time = "2026-05-09T23:13:31.859Z" }, - { url = "https://files.pythonhosted.org/packages/d9/c6/0a2436ae4da1ba76e51cb98943c6838a9a721faa40ebe2dce07694ae34e3/regex-2026.5.9-cp313-cp313-win_arm64.whl", hash = "sha256:56a33f191f17d8c417f99945ebdc1e691d3af9605d86ec68c7e54a57e3e17af6", size = 270500, upload-time = "2026-05-09T23:13:33.525Z" }, - { url = "https://files.pythonhosted.org/packages/e8/e9/d21346f7b60ed58789371358ed66b09d00f832e1bd7c06e55d9da5679882/regex-2026.5.9-cp313-cp313t-macosx_10_13_universal2.whl", hash = "sha256:01f28d868834624c934b8d2e0aa1c8341337e37831f4a012f18a5afcba4cbaf3", size = 494172, upload-time = "2026-05-09T23:13:35.935Z" }, - { url = "https://files.pythonhosted.org/packages/c4/43/fd1177a2032037c681baecdb3422ee4e1424aec4e4f470ef47793d325274/regex-2026.5.9-cp313-cp313t-macosx_10_13_x86_64.whl", hash = "sha256:48036f6374aaa79eb3b754ec29c61d1c6b1606749d705a13f8854fa2539671f6", size = 293952, upload-time = "2026-05-09T23:13:38.307Z" }, - { url = "https://files.pythonhosted.org/packages/f2/7d/9fbf919768368d3f8a4f6c692cf2aa61e482b2b81ec6a298ace4cbf02480/regex-2026.5.9-cp313-cp313t-macosx_11_0_arm64.whl", hash = "sha256:b96350aa424e79d4fd6b567b344dcbe2b2d6bfc48dfe7717587e1fa6d43da6ff", size = 292314, upload-time = "2026-05-09T23:13:40.353Z" }, - { url = "https://files.pythonhosted.org/packages/e2/6c/e41bfeecb589716843e7c4df09ba46ff2a42961457afece19059d85caeef/regex-2026.5.9-cp313-cp313t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:8f3af7a4903c5c04a11a196a5aa75cdd7dd3f8508132f9fb3259d9f5908e3b88", size = 811681, upload-time = "2026-05-09T23:13:42.543Z" }, - { url = "https://files.pythonhosted.org/packages/87/83/a5c1c525fba0aa656e88ad0face0b1829788ef4c2fb6b26df58aa1151b84/regex-2026.5.9-cp313-cp313t-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:7e87577720152d2caae19fe2baaf1f8d5ca12091e9e229f03915c37d1e4b9178", size = 871135, upload-time = "2026-05-09T23:13:44.326Z" }, - { url = "https://files.pythonhosted.org/packages/18/d4/80882e799e440dd878b0979cbebf8fa4d54624a332c83037c7a701649e3f/regex-2026.5.9-cp313-cp313t-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:c8b9b9d294cfea3cd19c718ade7cc93492b2c4991abd9a68d0b3477ae6d8e100", size = 917265, upload-time = "2026-05-09T23:13:47.295Z" }, - { url = "https://files.pythonhosted.org/packages/ae/ff/8db60211e2286e396aad7dc7725356c502bff0901ea05bd6cdc2e1a042b9/regex-2026.5.9-cp313-cp313t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:728d8bfd28a8845c8b6bc5dc7ce010453d206396786c0765c2740cb65f37791e", size = 816311, upload-time = "2026-05-09T23:13:49.885Z" }, - { url = "https://files.pythonhosted.org/packages/4c/47/742ef579c61730f8d268e5cf1f9ce0e37e2ea041ad0f5644724f2378e463/regex-2026.5.9-cp313-cp313t-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:7e30b874d341fac767d7df5a0870540541c2c054b80cfaac116e8d367a8a7ff2", size = 785498, upload-time = "2026-05-09T23:13:52.25Z" }, - { url = "https://files.pythonhosted.org/packages/7f/ab/cb0999802dcb0fb95b1ab005e8d4163d8afdd67efc2cb6b6630ac13f8cb1/regex-2026.5.9-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:fd190e88a895a8901325fad284a3f74ea52b1da8525b76cc811fa9b1edf0ce2b", size = 801348, upload-time = "2026-05-09T23:13:54.127Z" }, - { url = "https://files.pythonhosted.org/packages/7d/62/8ca59a24c55bc34d166eefaf3717bd77772f329fdbf984d86581e0a3571c/regex-2026.5.9-cp313-cp313t-musllinux_1_2_ppc64le.whl", hash = "sha256:8e76e8161ad00694cfce6767d5dea860c6391ac5b83e5c3a39661e696f11fc7e", size = 866493, upload-time = "2026-05-09T23:13:56.067Z" }, - { url = "https://files.pythonhosted.org/packages/8d/3d/30f2ae62cef3278bb5bb821f467277a55fb73f01032cf85997e15e8289a8/regex-2026.5.9-cp313-cp313t-musllinux_1_2_riscv64.whl", hash = "sha256:ddda5340e6c01a293027dd46232fa79eaff1b48058ce7a98f572b6445b088041", size = 772811, upload-time = "2026-05-09T23:13:57.867Z" }, - { url = "https://files.pythonhosted.org/packages/d8/ae/7d2089bcd78ad0c0161bc684339df50032acb438a7bd3305e7ddb1193cec/regex-2026.5.9-cp313-cp313t-musllinux_1_2_s390x.whl", hash = "sha256:205109e96b3cf5adf8f4cd62bedde9487feb282b9497a3535451e5a24cd706a0", size = 856584, upload-time = "2026-05-09T23:13:59.679Z" }, - { url = "https://files.pythonhosted.org/packages/a9/29/92ff47f75990131ea4f24ba17819e5a9d141e10819807e09addd73409af6/regex-2026.5.9-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:dfbe4579b9f08036aa7d101d1835437a20783574ac66327e6b29b4018a138081", size = 803453, upload-time = "2026-05-09T23:14:01.978Z" }, - { url = "https://files.pythonhosted.org/packages/04/99/eff29f1037dcab36702c9ee5d6858cf1ce2336ea8ea2987f64245b99ea5e/regex-2026.5.9-cp313-cp313t-win32.whl", hash = "sha256:ed2c9e8068b614c574d8d30e543d617cf5379b0535d46f97ef00e904745a08b5", size = 269951, upload-time = "2026-05-09T23:14:03.661Z" }, - { url = "https://files.pythonhosted.org/packages/0e/9d/8870b8981d27b22cda77bb26a5ac7ebfa9c7d9e0dea195a834a82380e748/regex-2026.5.9-cp313-cp313t-win_amd64.whl", hash = "sha256:b46b0f094dc1d3b90356c85a0bd2c9bafc4a6a190b9d6f8ddd5a033b6e088ed4", size = 281240, upload-time = "2026-05-09T23:14:05.56Z" }, - { url = "https://files.pythonhosted.org/packages/72/b1/3379415e8f135c13ac551353397cc4fe97b4978f3cac73c5fcbcded548b8/regex-2026.5.9-cp313-cp313t-win_arm64.whl", hash = "sha256:872acc074bd29ffc9913ecdfedf6ea77502312ca44a4aa0d3779089c6069d8de", size = 272383, upload-time = "2026-05-09T23:14:07.843Z" }, - { url = "https://files.pythonhosted.org/packages/13/3e/9c3cd292d8808b3645a2ce517e200179b6d0e903f176300bd8b542e14de5/regex-2026.5.9-cp314-cp314-macosx_10_13_universal2.whl", hash = "sha256:1bd7587a2948b4085195d5a3374eaf4a425dc3e55784c038175355ecf3bbbf8a", size = 490376, upload-time = "2026-05-09T23:14:09.64Z" }, - { url = "https://files.pythonhosted.org/packages/60/70/d43ee8a2ca0a8b68d167f21658b85520ac0574617c7f320367c5047f7556/regex-2026.5.9-cp314-cp314-macosx_10_13_x86_64.whl", hash = "sha256:dea2e88e1cce4522496cce630e11e67b98b7076620bc4336c3f674bc21a375f4", size = 291964, upload-time = "2026-05-09T23:14:11.424Z" }, - { url = "https://files.pythonhosted.org/packages/21/91/9d50b433828d8e74196904e168a43abf1e6e88b2a15d47ed742456720c37/regex-2026.5.9-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:2099f7e7ff7b6aa3192312650a56e91cc091e49d50b04e4f6f8b6e28b3b27f1c", size = 289682, upload-time = "2026-05-09T23:14:13.123Z" }, - { url = "https://files.pythonhosted.org/packages/3e/d2/b835e3cafbb9d977736912436259ff551d60919f7d7b3d37d46659c63564/regex-2026.5.9-cp314-cp314-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:ecd353045824e4477562a2ac718c25799cdaaa41f7aa925a806a8a3e6848a5b9", size = 796996, upload-time = "2026-05-09T23:14:14.923Z" }, - { url = "https://files.pythonhosted.org/packages/2c/a6/9f992d00019166b9de01c546dd4549bc679f2a68df11b877740b0760b7c2/regex-2026.5.9-cp314-cp314-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:65c8c8c37377794bd5b2f3ebe51919042bf17aec802e23c833d89782ed0c78af", size = 866089, upload-time = "2026-05-09T23:14:17.757Z" }, - { url = "https://files.pythonhosted.org/packages/e0/08/4d32af657e049b19cb62b02e46e38fe1518797bfb2203ee93a510b21b0dc/regex-2026.5.9-cp314-cp314-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:5b73ab8afcf66c622db143d1c6fda4e58e4d537ee4f125229ad47b1ab80f34c0", size = 911530, upload-time = "2026-05-09T23:14:20.353Z" }, - { url = "https://files.pythonhosted.org/packages/d9/27/2af43dd1dc201d1fecefda64a45f4ad0995855b92724f795a777b402ee69/regex-2026.5.9-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:0de5cf193997384ed2ca6f1cd4f78055b255d93d82d5a8cd6ba0d11c10b167e4", size = 800643, upload-time = "2026-05-09T23:14:22.265Z" }, - { url = "https://files.pythonhosted.org/packages/a4/dd/23a249047013b5321d4a60c4d2437462086f601b061776a525e5fba2a59f/regex-2026.5.9-cp314-cp314-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:d641a8c9a61618047796d572a39a79b26167b0411d2c3031937b2fe2d081e2cf", size = 777223, upload-time = "2026-05-09T23:14:24.179Z" }, - { url = "https://files.pythonhosted.org/packages/94/6a/e85ed9538cd19586d0465076a4578a12e093ce776d15f3f8ce92733a8dd6/regex-2026.5.9-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:24b2355ef5cc9aa5b8f07d17704face1c166fdcc2290fa7bd6e6c925655a8346", size = 785760, upload-time = "2026-05-09T23:14:26.065Z" }, - { url = "https://files.pythonhosted.org/packages/2a/c4/f25473209438638e947c55f9156fd8f236f74169229028cc99116380868e/regex-2026.5.9-cp314-cp314-musllinux_1_2_ppc64le.whl", hash = "sha256:a24852d3c29ad9e47593593d8a247c44ccc3d0548ef12c822d6ed0810affe676", size = 860891, upload-time = "2026-05-09T23:14:28.17Z" }, - { url = "https://files.pythonhosted.org/packages/f9/f7/f4f86e3c74419c37370e91f150ae0c2ef7d34b2e0e4cdd5da046a02e4022/regex-2026.5.9-cp314-cp314-musllinux_1_2_riscv64.whl", hash = "sha256:916714069da19329ef7de197dcbc77bb3104145c7c2c864dbfbe318f46b88b14", size = 765891, upload-time = "2026-05-09T23:14:30.06Z" }, - { url = "https://files.pythonhosted.org/packages/26/70/704d8e13765939146b1cd0ef4e2feb71d7929727d2290f026eed10095955/regex-2026.5.9-cp314-cp314-musllinux_1_2_s390x.whl", hash = "sha256:fa411799ca8da32a8d38d020a88faa5b6f91657d284761352940ecf9f7c3bbdd", size = 851380, upload-time = "2026-05-09T23:14:32.123Z" }, - { url = "https://files.pythonhosted.org/packages/26/29/1a13582a8460038edc38e49f64ceb0dd7c60f5caba77571f4bf6601965d9/regex-2026.5.9-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:1e6da47d679b7010ef27556b6e0f99771b744936db1792a10ceac6547ae1503e", size = 789350, upload-time = "2026-05-09T23:14:34.799Z" }, - { url = "https://files.pythonhosted.org/packages/73/56/3dcafe34fc72e271d62ad9a291801e88a1457bb251c132f15fcc2e5aad1a/regex-2026.5.9-cp314-cp314-win32.whl", hash = "sha256:98bd73080e8756255137e1bd3f3f00295bbc5aa383c0e0f973920e9134d7c4ad", size = 272130, upload-time = "2026-05-09T23:14:36.729Z" }, - { url = "https://files.pythonhosted.org/packages/d0/9c/02eebf0be95efe416c664db7fb8b6b05b7a0b06a7544f2884f2558b0526f/regex-2026.5.9-cp314-cp314-win_amd64.whl", hash = "sha256:ff8d372ac2acdc048d1c19916f27ee61bc5722728458ba6ca5052f2c72d51763", size = 280999, upload-time = "2026-05-09T23:14:39.126Z" }, - { url = "https://files.pythonhosted.org/packages/70/5a/1dd1abee76cb7a846a0bcf42fdc87e5720c3c33c24f3e37814310a513d9f/regex-2026.5.9-cp314-cp314-win_arm64.whl", hash = "sha256:e1d93bf647916292e8edcec150c07ddf3dc50179ccaf770c04a7f9e452155372", size = 273500, upload-time = "2026-05-09T23:14:41.059Z" }, - { url = "https://files.pythonhosted.org/packages/86/c1/c5f619b0057a7965cb78ec559c1d7a45ce8c99a35bea95483d64959a93d9/regex-2026.5.9-cp314-cp314t-macosx_10_13_universal2.whl", hash = "sha256:83d0ee4a57d1c87cb549e195ec300b8f0ec3a82eba66d835e4e2ed8634fe4499", size = 494269, upload-time = "2026-05-09T23:14:42.869Z" }, - { url = "https://files.pythonhosted.org/packages/05/2c/5d01f1aee33de4bbe60c8452945bfc8477ca7c5ae4450f6bfe711036cb36/regex-2026.5.9-cp314-cp314t-macosx_10_13_x86_64.whl", hash = "sha256:d3d7eb5c9a7f6df82ed3cfac9beb93882a5cbcb5b8b157b56cb2b3b276574ac1", size = 293954, upload-time = "2026-05-09T23:14:44.822Z" }, - { url = "https://files.pythonhosted.org/packages/7a/fe/e8988b2ae2108c6ef71bd4aa8d87fbe257976dd0810e826cd75f701c68b6/regex-2026.5.9-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:075160bf16658e16d35233300b8453aac25de4cbea808d22348b6979668e924d", size = 292405, upload-time = "2026-05-09T23:14:47.211Z" }, - { url = "https://files.pythonhosted.org/packages/79/34/d2b0937faa7859263f7f0a3c6b103a1296306be6952dc173d0154e9a2f49/regex-2026.5.9-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:45375819235558a4ff1c4971dc32881f022613abdb180128f5cb4768c1765a1c", size = 811855, upload-time = "2026-05-09T23:14:49.21Z" }, - { url = "https://files.pythonhosted.org/packages/80/fe/daf53a47457a8486db66c66c01ceb9c2303eecee3f87197f1e77eb1a736d/regex-2026.5.9-cp314-cp314t-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:ead4b163ac30a29574510cd4b3e2e985ac5290c05fc7095557d6a5f403fc31b5", size = 871189, upload-time = "2026-05-09T23:14:51.555Z" }, - { url = "https://files.pythonhosted.org/packages/1c/75/058fc4470cbfbf57d800aff1a0022b929a3f9fa553ee10a0cdf2070eb31f/regex-2026.5.9-cp314-cp314t-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:8c6e4218fbdfbcd4f6c19efca40930d24a621bf4b48cb76bc6640543bd28ef20", size = 917485, upload-time = "2026-05-09T23:14:53.633Z" }, - { url = "https://files.pythonhosted.org/packages/88/e7/179cfda3a28bc843b5c6cfe7f79f23489c791ed95f151083803660878432/regex-2026.5.9-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:6351571c8a42b505eb555c0dc47d740d0fb66977dc142919eea6f4325b7c56a0", size = 816369, upload-time = "2026-05-09T23:14:56.198Z" }, - { url = "https://files.pythonhosted.org/packages/41/90/6f0cc422071688266d344fca8462d787cba0a2c144acb25721f9a61ec265/regex-2026.5.9-cp314-cp314t-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:002205cafd2a9e78c6290c7d1df277bf3277b3b7a30e0b4bb0dac2e2e3f7cb2d", size = 785869, upload-time = "2026-05-09T23:14:58.602Z" }, - { url = "https://files.pythonhosted.org/packages/02/67/a31f1760f09c27b251ef39e9beb541f462cf977381d067faa764c2c0e393/regex-2026.5.9-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:8abd33fef90b2a9efac5557d6033ca82d1195ed3a15fea5af15ba7b463c6a63b", size = 801427, upload-time = "2026-05-09T23:15:00.642Z" }, - { url = "https://files.pythonhosted.org/packages/e3/c4/1a80654597b6bc1e1ea0494824c31200e8a956abe290afae9b19a166a148/regex-2026.5.9-cp314-cp314t-musllinux_1_2_ppc64le.whl", hash = "sha256:31037c82eccb44b7ea2e9e221d7c01429430e989a1f4b91ea5a855f6017b509a", size = 866482, upload-time = "2026-05-09T23:15:03.384Z" }, - { url = "https://files.pythonhosted.org/packages/d1/11/960724e06482c08466ff5611e242e86f80062949cdf6b4b9cc317b9dd93d/regex-2026.5.9-cp314-cp314t-musllinux_1_2_riscv64.whl", hash = "sha256:5604dfd046dc37eca90250fc3be938b076c8059fa772ac0ed6f499b0f0fb0415", size = 773022, upload-time = "2026-05-09T23:15:05.625Z" }, - { url = "https://files.pythonhosted.org/packages/50/a8/a9979c3e7918280e93159ebcab5ef1a65116dd4f3bd6091be0eae4a126e8/regex-2026.5.9-cp314-cp314t-musllinux_1_2_s390x.whl", hash = "sha256:0e1b1b4e496afbb24f4a62aba855ee4f88f25578927697b340702e48c9ee6bc2", size = 856642, upload-time = "2026-05-09T23:15:07.966Z" }, - { url = "https://files.pythonhosted.org/packages/fe/d4/a9b732f2f0072c0ab12227483abb24fffcb9f73f8a2b203df0a6d0434735/regex-2026.5.9-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:be3372b9df6ddecff6486d37e19095a7b4973137caf5512407a89f4455361f41", size = 803552, upload-time = "2026-05-09T23:15:10.215Z" }, - { url = "https://files.pythonhosted.org/packages/d5/fe/1b3113817447a1d4155e4ac76d2e072f42c0bcba2f43fa8a0e756ea2cd91/regex-2026.5.9-cp314-cp314t-win32.whl", hash = "sha256:3ddd90103f9e5c471c49c7852ecc1fe27c7e45eb99e977aefe7caa4e779f4f58", size = 275746, upload-time = "2026-05-09T23:15:12.609Z" }, - { url = "https://files.pythonhosted.org/packages/92/73/93d42045302636c91f2e5ef588b65b84b01428f28ec77de256b1dfdfbe5c/regex-2026.5.9-cp314-cp314t-win_amd64.whl", hash = "sha256:ca518ed29c46eecba6010b15f1b9a479314d2de409536e71b6a13aa04e3b8a77", size = 285685, upload-time = "2026-05-09T23:15:15.086Z" }, - { url = "https://files.pythonhosted.org/packages/da/80/35b4c33c804a165a7f55289afda3ea9e3eb6d15800341a2d66455c0f1f30/regex-2026.5.9-cp314-cp314t-win_arm64.whl", hash = "sha256:5e41809d2683fcde7d5a8c87a6567ba1fb1ce0de9f31bff578de00a4b2d76daa", size = 275713, upload-time = "2026-05-09T23:15:16.98Z" }, -] - -[[package]] -name = "rich" -version = "15.0.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "markdown-it-py" }, - { name = "pygments" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/c0/8f/0722ca900cc807c13a6a0c696dacf35430f72e0ec571c4275d2371fca3e9/rich-15.0.0.tar.gz", hash = "sha256:edd07a4824c6b40189fb7ac9bc4c52536e9780fbbfbddf6f1e2502c31b068c36", size = 230680, upload-time = "2026-04-12T08:24:00.75Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/82/3b/64d4899d73f91ba49a8c18a8ff3f0ea8f1c1d75481760df8c68ef5235bf5/rich-15.0.0-py3-none-any.whl", hash = "sha256:33bd4ef74232fb73fe9279a257718407f169c09b78a87ad3d296f548e27de0bb", size = 310654, upload-time = "2026-04-12T08:24:02.83Z" }, -] - -[[package]] -name = "rpds-py" -version = "0.30.0" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version < '3.11' and sys_platform == 'win32'", - "python_full_version < '3.11' and sys_platform != 'win32'", -] -sdist = { url = "https://files.pythonhosted.org/packages/20/af/3f2f423103f1113b36230496629986e0ef7e199d2aa8392452b484b38ced/rpds_py-0.30.0.tar.gz", hash = "sha256:dd8ff7cf90014af0c0f787eea34794ebf6415242ee1d6fa91eaba725cc441e84", size = 69469, upload-time = "2025-11-30T20:24:38.837Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/06/0c/0c411a0ec64ccb6d104dcabe0e713e05e153a9a2c3c2bd2b32ce412166fe/rpds_py-0.30.0-cp310-cp310-macosx_10_12_x86_64.whl", hash = "sha256:679ae98e00c0e8d68a7fda324e16b90fd5260945b45d3b824c892cec9eea3288", size = 370490, upload-time = "2025-11-30T20:21:33.256Z" }, - { url = "https://files.pythonhosted.org/packages/19/6a/4ba3d0fb7297ebae71171822554abe48d7cab29c28b8f9f2c04b79988c05/rpds_py-0.30.0-cp310-cp310-macosx_11_0_arm64.whl", hash = "sha256:4cc2206b76b4f576934f0ed374b10d7ca5f457858b157ca52064bdfc26b9fc00", size = 359751, upload-time = "2025-11-30T20:21:34.591Z" }, - { url = "https://files.pythonhosted.org/packages/cd/7c/e4933565ef7f7a0818985d87c15d9d273f1a649afa6a52ea35ad011195ea/rpds_py-0.30.0-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:389a2d49eded1896c3d48b0136ead37c48e221b391c052fba3f4055c367f60a6", size = 389696, upload-time = "2025-11-30T20:21:36.122Z" }, - { url = "https://files.pythonhosted.org/packages/5e/01/6271a2511ad0815f00f7ed4390cf2567bec1d4b1da39e2c27a41e6e3b4de/rpds_py-0.30.0-cp310-cp310-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:32c8528634e1bf7121f3de08fa85b138f4e0dc47657866630611b03967f041d7", size = 403136, upload-time = "2025-11-30T20:21:37.728Z" }, - { url = "https://files.pythonhosted.org/packages/55/64/c857eb7cd7541e9b4eee9d49c196e833128a55b89a9850a9c9ac33ccf897/rpds_py-0.30.0-cp310-cp310-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:f207f69853edd6f6700b86efb84999651baf3789e78a466431df1331608e5324", size = 524699, upload-time = "2025-11-30T20:21:38.92Z" }, - { url = "https://files.pythonhosted.org/packages/9c/ed/94816543404078af9ab26159c44f9e98e20fe47e2126d5d32c9d9948d10a/rpds_py-0.30.0-cp310-cp310-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:67b02ec25ba7a9e8fa74c63b6ca44cf5707f2fbfadae3ee8e7494297d56aa9df", size = 412022, upload-time = "2025-11-30T20:21:40.407Z" }, - { url = "https://files.pythonhosted.org/packages/61/b5/707f6cf0066a6412aacc11d17920ea2e19e5b2f04081c64526eb35b5c6e7/rpds_py-0.30.0-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:0c0e95f6819a19965ff420f65578bacb0b00f251fefe2c8b23347c37174271f3", size = 390522, upload-time = "2025-11-30T20:21:42.17Z" }, - { url = "https://files.pythonhosted.org/packages/13/4e/57a85fda37a229ff4226f8cbcf09f2a455d1ed20e802ce5b2b4a7f5ed053/rpds_py-0.30.0-cp310-cp310-manylinux_2_31_riscv64.whl", hash = "sha256:a452763cc5198f2f98898eb98f7569649fe5da666c2dc6b5ddb10fde5a574221", size = 404579, upload-time = "2025-11-30T20:21:43.769Z" }, - { url = "https://files.pythonhosted.org/packages/f9/da/c9339293513ec680a721e0e16bf2bac3db6e5d7e922488de471308349bba/rpds_py-0.30.0-cp310-cp310-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:e0b65193a413ccc930671c55153a03ee57cecb49e6227204b04fae512eb657a7", size = 421305, upload-time = "2025-11-30T20:21:44.994Z" }, - { url = "https://files.pythonhosted.org/packages/f9/be/522cb84751114f4ad9d822ff5a1aa3c98006341895d5f084779b99596e5c/rpds_py-0.30.0-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:858738e9c32147f78b3ac24dc0edb6610000e56dc0f700fd5f651d0a0f0eb9ff", size = 572503, upload-time = "2025-11-30T20:21:46.91Z" }, - { url = "https://files.pythonhosted.org/packages/a2/9b/de879f7e7ceddc973ea6e4629e9b380213a6938a249e94b0cdbcc325bb66/rpds_py-0.30.0-cp310-cp310-musllinux_1_2_i686.whl", hash = "sha256:da279aa314f00acbb803da1e76fa18666778e8a8f83484fba94526da5de2cba7", size = 598322, upload-time = "2025-11-30T20:21:48.709Z" }, - { url = "https://files.pythonhosted.org/packages/48/ac/f01fc22efec3f37d8a914fc1b2fb9bcafd56a299edbe96406f3053edea5a/rpds_py-0.30.0-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:7c64d38fb49b6cdeda16ab49e35fe0da2e1e9b34bc38bd78386530f218b37139", size = 560792, upload-time = "2025-11-30T20:21:50.024Z" }, - { url = "https://files.pythonhosted.org/packages/e2/da/4e2b19d0f131f35b6146425f846563d0ce036763e38913d917187307a671/rpds_py-0.30.0-cp310-cp310-win32.whl", hash = "sha256:6de2a32a1665b93233cde140ff8b3467bdb9e2af2b91079f0333a0974d12d464", size = 221901, upload-time = "2025-11-30T20:21:51.32Z" }, - { url = "https://files.pythonhosted.org/packages/96/cb/156d7a5cf4f78a7cc571465d8aec7a3c447c94f6749c5123f08438bcf7bc/rpds_py-0.30.0-cp310-cp310-win_amd64.whl", hash = "sha256:1726859cd0de969f88dc8673bdd954185b9104e05806be64bcd87badbe313169", size = 235823, upload-time = "2025-11-30T20:21:52.505Z" }, - { url = "https://files.pythonhosted.org/packages/4d/6e/f964e88b3d2abee2a82c1ac8366da848fce1c6d834dc2132c3fda3970290/rpds_py-0.30.0-cp311-cp311-macosx_10_12_x86_64.whl", hash = "sha256:a2bffea6a4ca9f01b3f8e548302470306689684e61602aa3d141e34da06cf425", size = 370157, upload-time = "2025-11-30T20:21:53.789Z" }, - { url = "https://files.pythonhosted.org/packages/94/ba/24e5ebb7c1c82e74c4e4f33b2112a5573ddc703915b13a073737b59b86e0/rpds_py-0.30.0-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:dc4f992dfe1e2bc3ebc7444f6c7051b4bc13cd8e33e43511e8ffd13bf407010d", size = 359676, upload-time = "2025-11-30T20:21:55.475Z" }, - { url = "https://files.pythonhosted.org/packages/84/86/04dbba1b087227747d64d80c3b74df946b986c57af0a9f0c98726d4d7a3b/rpds_py-0.30.0-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:422c3cb9856d80b09d30d2eb255d0754b23e090034e1deb4083f8004bd0761e4", size = 389938, upload-time = "2025-11-30T20:21:57.079Z" }, - { url = "https://files.pythonhosted.org/packages/42/bb/1463f0b1722b7f45431bdd468301991d1328b16cffe0b1c2918eba2c4eee/rpds_py-0.30.0-cp311-cp311-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:07ae8a593e1c3c6b82ca3292efbe73c30b61332fd612e05abee07c79359f292f", size = 402932, upload-time = "2025-11-30T20:21:58.47Z" }, - { url = "https://files.pythonhosted.org/packages/99/ee/2520700a5c1f2d76631f948b0736cdf9b0acb25abd0ca8e889b5c62ac2e3/rpds_py-0.30.0-cp311-cp311-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:12f90dd7557b6bd57f40abe7747e81e0c0b119bef015ea7726e69fe550e394a4", size = 525830, upload-time = "2025-11-30T20:21:59.699Z" }, - { url = "https://files.pythonhosted.org/packages/e0/ad/bd0331f740f5705cc555a5e17fdf334671262160270962e69a2bdef3bf76/rpds_py-0.30.0-cp311-cp311-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:99b47d6ad9a6da00bec6aabe5a6279ecd3c06a329d4aa4771034a21e335c3a97", size = 412033, upload-time = "2025-11-30T20:22:00.991Z" }, - { url = "https://files.pythonhosted.org/packages/f8/1e/372195d326549bb51f0ba0f2ecb9874579906b97e08880e7a65c3bef1a99/rpds_py-0.30.0-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:33f559f3104504506a44bb666b93a33f5d33133765b0c216a5bf2f1e1503af89", size = 390828, upload-time = "2025-11-30T20:22:02.723Z" }, - { url = "https://files.pythonhosted.org/packages/ab/2b/d88bb33294e3e0c76bc8f351a3721212713629ffca1700fa94979cb3eae8/rpds_py-0.30.0-cp311-cp311-manylinux_2_31_riscv64.whl", hash = "sha256:946fe926af6e44f3697abbc305ea168c2c31d3e3ef1058cf68f379bf0335a78d", size = 404683, upload-time = "2025-11-30T20:22:04.367Z" }, - { url = "https://files.pythonhosted.org/packages/50/32/c759a8d42bcb5289c1fac697cd92f6fe01a018dd937e62ae77e0e7f15702/rpds_py-0.30.0-cp311-cp311-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:495aeca4b93d465efde585977365187149e75383ad2684f81519f504f5c13038", size = 421583, upload-time = "2025-11-30T20:22:05.814Z" }, - { url = "https://files.pythonhosted.org/packages/2b/81/e729761dbd55ddf5d84ec4ff1f47857f4374b0f19bdabfcf929164da3e24/rpds_py-0.30.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:d9a0ca5da0386dee0655b4ccdf46119df60e0f10da268d04fe7cc87886872ba7", size = 572496, upload-time = "2025-11-30T20:22:07.713Z" }, - { url = "https://files.pythonhosted.org/packages/14/f6/69066a924c3557c9c30baa6ec3a0aa07526305684c6f86c696b08860726c/rpds_py-0.30.0-cp311-cp311-musllinux_1_2_i686.whl", hash = "sha256:8d6d1cc13664ec13c1b84241204ff3b12f9bb82464b8ad6e7a5d3486975c2eed", size = 598669, upload-time = "2025-11-30T20:22:09.312Z" }, - { url = "https://files.pythonhosted.org/packages/5f/48/905896b1eb8a05630d20333d1d8ffd162394127b74ce0b0784ae04498d32/rpds_py-0.30.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:3896fa1be39912cf0757753826bc8bdc8ca331a28a7c4ae46b7a21280b06bb85", size = 561011, upload-time = "2025-11-30T20:22:11.309Z" }, - { url = "https://files.pythonhosted.org/packages/22/16/cd3027c7e279d22e5eb431dd3c0fbc677bed58797fe7581e148f3f68818b/rpds_py-0.30.0-cp311-cp311-win32.whl", hash = "sha256:55f66022632205940f1827effeff17c4fa7ae1953d2b74a8581baaefb7d16f8c", size = 221406, upload-time = "2025-11-30T20:22:13.101Z" }, - { url = "https://files.pythonhosted.org/packages/fa/5b/e7b7aa136f28462b344e652ee010d4de26ee9fd16f1bfd5811f5153ccf89/rpds_py-0.30.0-cp311-cp311-win_amd64.whl", hash = "sha256:a51033ff701fca756439d641c0ad09a41d9242fa69121c7d8769604a0a629825", size = 236024, upload-time = "2025-11-30T20:22:14.853Z" }, - { url = "https://files.pythonhosted.org/packages/14/a6/364bba985e4c13658edb156640608f2c9e1d3ea3c81b27aa9d889fff0e31/rpds_py-0.30.0-cp311-cp311-win_arm64.whl", hash = "sha256:47b0ef6231c58f506ef0b74d44e330405caa8428e770fec25329ed2cb971a229", size = 229069, upload-time = "2025-11-30T20:22:16.577Z" }, - { url = "https://files.pythonhosted.org/packages/03/e7/98a2f4ac921d82f33e03f3835f5bf3a4a40aa1bfdc57975e74a97b2b4bdd/rpds_py-0.30.0-cp312-cp312-macosx_10_12_x86_64.whl", hash = "sha256:a161f20d9a43006833cd7068375a94d035714d73a172b681d8881820600abfad", size = 375086, upload-time = "2025-11-30T20:22:17.93Z" }, - { url = "https://files.pythonhosted.org/packages/4d/a1/bca7fd3d452b272e13335db8d6b0b3ecde0f90ad6f16f3328c6fb150c889/rpds_py-0.30.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:6abc8880d9d036ecaafe709079969f56e876fcf107f7a8e9920ba6d5a3878d05", size = 359053, upload-time = "2025-11-30T20:22:19.297Z" }, - { url = "https://files.pythonhosted.org/packages/65/1c/ae157e83a6357eceff62ba7e52113e3ec4834a84cfe07fa4b0757a7d105f/rpds_py-0.30.0-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:ca28829ae5f5d569bb62a79512c842a03a12576375d5ece7d2cadf8abe96ec28", size = 390763, upload-time = "2025-11-30T20:22:21.661Z" }, - { url = "https://files.pythonhosted.org/packages/d4/36/eb2eb8515e2ad24c0bd43c3ee9cd74c33f7ca6430755ccdb240fd3144c44/rpds_py-0.30.0-cp312-cp312-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:a1010ed9524c73b94d15919ca4d41d8780980e1765babf85f9a2f90d247153dd", size = 408951, upload-time = "2025-11-30T20:22:23.408Z" }, - { url = "https://files.pythonhosted.org/packages/d6/65/ad8dc1784a331fabbd740ef6f71ce2198c7ed0890dab595adb9ea2d775a1/rpds_py-0.30.0-cp312-cp312-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:f8d1736cfb49381ba528cd5baa46f82fdc65c06e843dab24dd70b63d09121b3f", size = 514622, upload-time = "2025-11-30T20:22:25.16Z" }, - { url = "https://files.pythonhosted.org/packages/63/8e/0cfa7ae158e15e143fe03993b5bcd743a59f541f5952e1546b1ac1b5fd45/rpds_py-0.30.0-cp312-cp312-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:d948b135c4693daff7bc2dcfc4ec57237a29bd37e60c2fabf5aff2bbacf3e2f1", size = 414492, upload-time = "2025-11-30T20:22:26.505Z" }, - { url = "https://files.pythonhosted.org/packages/60/1b/6f8f29f3f995c7ffdde46a626ddccd7c63aefc0efae881dc13b6e5d5bb16/rpds_py-0.30.0-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:47f236970bccb2233267d89173d3ad2703cd36a0e2a6e92d0560d333871a3d23", size = 394080, upload-time = "2025-11-30T20:22:27.934Z" }, - { url = "https://files.pythonhosted.org/packages/6d/d5/a266341051a7a3ca2f4b750a3aa4abc986378431fc2da508c5034d081b70/rpds_py-0.30.0-cp312-cp312-manylinux_2_31_riscv64.whl", hash = "sha256:2e6ecb5a5bcacf59c3f912155044479af1d0b6681280048b338b28e364aca1f6", size = 408680, upload-time = "2025-11-30T20:22:29.341Z" }, - { url = "https://files.pythonhosted.org/packages/10/3b/71b725851df9ab7a7a4e33cf36d241933da66040d195a84781f49c50490c/rpds_py-0.30.0-cp312-cp312-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:a8fa71a2e078c527c3e9dc9fc5a98c9db40bcc8a92b4e8858e36d329f8684b51", size = 423589, upload-time = "2025-11-30T20:22:31.469Z" }, - { url = "https://files.pythonhosted.org/packages/00/2b/e59e58c544dc9bd8bd8384ecdb8ea91f6727f0e37a7131baeff8d6f51661/rpds_py-0.30.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:73c67f2db7bc334e518d097c6d1e6fed021bbc9b7d678d6cc433478365d1d5f5", size = 573289, upload-time = "2025-11-30T20:22:32.997Z" }, - { url = "https://files.pythonhosted.org/packages/da/3e/a18e6f5b460893172a7d6a680e86d3b6bc87a54c1f0b03446a3c8c7b588f/rpds_py-0.30.0-cp312-cp312-musllinux_1_2_i686.whl", hash = "sha256:5ba103fb455be00f3b1c2076c9d4264bfcb037c976167a6047ed82f23153f02e", size = 599737, upload-time = "2025-11-30T20:22:34.419Z" }, - { url = "https://files.pythonhosted.org/packages/5c/e2/714694e4b87b85a18e2c243614974413c60aa107fd815b8cbc42b873d1d7/rpds_py-0.30.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:7cee9c752c0364588353e627da8a7e808a66873672bcb5f52890c33fd965b394", size = 563120, upload-time = "2025-11-30T20:22:35.903Z" }, - { url = "https://files.pythonhosted.org/packages/6f/ab/d5d5e3bcedb0a77f4f613706b750e50a5a3ba1c15ccd3665ecc636c968fd/rpds_py-0.30.0-cp312-cp312-win32.whl", hash = "sha256:1ab5b83dbcf55acc8b08fc62b796ef672c457b17dbd7820a11d6c52c06839bdf", size = 223782, upload-time = "2025-11-30T20:22:37.271Z" }, - { url = "https://files.pythonhosted.org/packages/39/3b/f786af9957306fdc38a74cef405b7b93180f481fb48453a114bb6465744a/rpds_py-0.30.0-cp312-cp312-win_amd64.whl", hash = "sha256:a090322ca841abd453d43456ac34db46e8b05fd9b3b4ac0c78bcde8b089f959b", size = 240463, upload-time = "2025-11-30T20:22:39.021Z" }, - { url = "https://files.pythonhosted.org/packages/f3/d2/b91dc748126c1559042cfe41990deb92c4ee3e2b415f6b5234969ffaf0cc/rpds_py-0.30.0-cp312-cp312-win_arm64.whl", hash = "sha256:669b1805bd639dd2989b281be2cfd951c6121b65e729d9b843e9639ef1fd555e", size = 230868, upload-time = "2025-11-30T20:22:40.493Z" }, - { url = "https://files.pythonhosted.org/packages/ed/dc/d61221eb88ff410de3c49143407f6f3147acf2538c86f2ab7ce65ae7d5f9/rpds_py-0.30.0-cp313-cp313-macosx_10_12_x86_64.whl", hash = "sha256:f83424d738204d9770830d35290ff3273fbb02b41f919870479fab14b9d303b2", size = 374887, upload-time = "2025-11-30T20:22:41.812Z" }, - { url = "https://files.pythonhosted.org/packages/fd/32/55fb50ae104061dbc564ef15cc43c013dc4a9f4527a1f4d99baddf56fe5f/rpds_py-0.30.0-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:e7536cd91353c5273434b4e003cbda89034d67e7710eab8761fd918ec6c69cf8", size = 358904, upload-time = "2025-11-30T20:22:43.479Z" }, - { url = "https://files.pythonhosted.org/packages/58/70/faed8186300e3b9bdd138d0273109784eea2396c68458ed580f885dfe7ad/rpds_py-0.30.0-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:2771c6c15973347f50fece41fc447c054b7ac2ae0502388ce3b6738cd366e3d4", size = 389945, upload-time = "2025-11-30T20:22:44.819Z" }, - { url = "https://files.pythonhosted.org/packages/bd/a8/073cac3ed2c6387df38f71296d002ab43496a96b92c823e76f46b8af0543/rpds_py-0.30.0-cp313-cp313-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:0a59119fc6e3f460315fe9d08149f8102aa322299deaa5cab5b40092345c2136", size = 407783, upload-time = "2025-11-30T20:22:46.103Z" }, - { url = "https://files.pythonhosted.org/packages/77/57/5999eb8c58671f1c11eba084115e77a8899d6e694d2a18f69f0ba471ec8b/rpds_py-0.30.0-cp313-cp313-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:76fec018282b4ead0364022e3c54b60bf368b9d926877957a8624b58419169b7", size = 515021, upload-time = "2025-11-30T20:22:47.458Z" }, - { url = "https://files.pythonhosted.org/packages/e0/af/5ab4833eadc36c0a8ed2bc5c0de0493c04f6c06de223170bd0798ff98ced/rpds_py-0.30.0-cp313-cp313-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:692bef75a5525db97318e8cd061542b5a79812d711ea03dbc1f6f8dbb0c5f0d2", size = 414589, upload-time = "2025-11-30T20:22:48.872Z" }, - { url = "https://files.pythonhosted.org/packages/b7/de/f7192e12b21b9e9a68a6d0f249b4af3fdcdff8418be0767a627564afa1f1/rpds_py-0.30.0-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:9027da1ce107104c50c81383cae773ef5c24d296dd11c99e2629dbd7967a20c6", size = 394025, upload-time = "2025-11-30T20:22:50.196Z" }, - { url = "https://files.pythonhosted.org/packages/91/c4/fc70cd0249496493500e7cc2de87504f5aa6509de1e88623431fec76d4b6/rpds_py-0.30.0-cp313-cp313-manylinux_2_31_riscv64.whl", hash = "sha256:9cf69cdda1f5968a30a359aba2f7f9aa648a9ce4b580d6826437f2b291cfc86e", size = 408895, upload-time = "2025-11-30T20:22:51.87Z" }, - { url = "https://files.pythonhosted.org/packages/58/95/d9275b05ab96556fefff73a385813eb66032e4c99f411d0795372d9abcea/rpds_py-0.30.0-cp313-cp313-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:a4796a717bf12b9da9d3ad002519a86063dcac8988b030e405704ef7d74d2d9d", size = 422799, upload-time = "2025-11-30T20:22:53.341Z" }, - { url = "https://files.pythonhosted.org/packages/06/c1/3088fc04b6624eb12a57eb814f0d4997a44b0d208d6cace713033ff1a6ba/rpds_py-0.30.0-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:5d4c2aa7c50ad4728a094ebd5eb46c452e9cb7edbfdb18f9e1221f597a73e1e7", size = 572731, upload-time = "2025-11-30T20:22:54.778Z" }, - { url = "https://files.pythonhosted.org/packages/d8/42/c612a833183b39774e8ac8fecae81263a68b9583ee343db33ab571a7ce55/rpds_py-0.30.0-cp313-cp313-musllinux_1_2_i686.whl", hash = "sha256:ba81a9203d07805435eb06f536d95a266c21e5b2dfbf6517748ca40c98d19e31", size = 599027, upload-time = "2025-11-30T20:22:56.212Z" }, - { url = "https://files.pythonhosted.org/packages/5f/60/525a50f45b01d70005403ae0e25f43c0384369ad24ffe46e8d9068b50086/rpds_py-0.30.0-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:945dccface01af02675628334f7cf49c2af4c1c904748efc5cf7bbdf0b579f95", size = 563020, upload-time = "2025-11-30T20:22:58.2Z" }, - { url = "https://files.pythonhosted.org/packages/0b/5d/47c4655e9bcd5ca907148535c10e7d489044243cc9941c16ed7cd53be91d/rpds_py-0.30.0-cp313-cp313-win32.whl", hash = "sha256:b40fb160a2db369a194cb27943582b38f79fc4887291417685f3ad693c5a1d5d", size = 223139, upload-time = "2025-11-30T20:23:00.209Z" }, - { url = "https://files.pythonhosted.org/packages/f2/e1/485132437d20aa4d3e1d8b3fb5a5e65aa8139f1e097080c2a8443201742c/rpds_py-0.30.0-cp313-cp313-win_amd64.whl", hash = "sha256:806f36b1b605e2d6a72716f321f20036b9489d29c51c91f4dd29a3e3afb73b15", size = 240224, upload-time = "2025-11-30T20:23:02.008Z" }, - { url = "https://files.pythonhosted.org/packages/24/95/ffd128ed1146a153d928617b0ef673960130be0009c77d8fbf0abe306713/rpds_py-0.30.0-cp313-cp313-win_arm64.whl", hash = "sha256:d96c2086587c7c30d44f31f42eae4eac89b60dabbac18c7669be3700f13c3ce1", size = 230645, upload-time = "2025-11-30T20:23:03.43Z" }, - { url = "https://files.pythonhosted.org/packages/ff/1b/b10de890a0def2a319a2626334a7f0ae388215eb60914dbac8a3bae54435/rpds_py-0.30.0-cp313-cp313t-macosx_10_12_x86_64.whl", hash = "sha256:eb0b93f2e5c2189ee831ee43f156ed34e2a89a78a66b98cadad955972548be5a", size = 364443, upload-time = "2025-11-30T20:23:04.878Z" }, - { url = "https://files.pythonhosted.org/packages/0d/bf/27e39f5971dc4f305a4fb9c672ca06f290f7c4e261c568f3dea16a410d47/rpds_py-0.30.0-cp313-cp313t-macosx_11_0_arm64.whl", hash = "sha256:922e10f31f303c7c920da8981051ff6d8c1a56207dbdf330d9047f6d30b70e5e", size = 353375, upload-time = "2025-11-30T20:23:06.342Z" }, - { url = "https://files.pythonhosted.org/packages/40/58/442ada3bba6e8e6615fc00483135c14a7538d2ffac30e2d933ccf6852232/rpds_py-0.30.0-cp313-cp313t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:cdc62c8286ba9bf7f47befdcea13ea0e26bf294bda99758fd90535cbaf408000", size = 383850, upload-time = "2025-11-30T20:23:07.825Z" }, - { url = "https://files.pythonhosted.org/packages/14/14/f59b0127409a33c6ef6f5c1ebd5ad8e32d7861c9c7adfa9a624fc3889f6c/rpds_py-0.30.0-cp313-cp313t-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:47f9a91efc418b54fb8190a6b4aa7813a23fb79c51f4bb84e418f5476c38b8db", size = 392812, upload-time = "2025-11-30T20:23:09.228Z" }, - { url = "https://files.pythonhosted.org/packages/b3/66/e0be3e162ac299b3a22527e8913767d869e6cc75c46bd844aa43fb81ab62/rpds_py-0.30.0-cp313-cp313t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:1f3587eb9b17f3789ad50824084fa6f81921bbf9a795826570bda82cb3ed91f2", size = 517841, upload-time = "2025-11-30T20:23:11.186Z" }, - { url = "https://files.pythonhosted.org/packages/3d/55/fa3b9cf31d0c963ecf1ba777f7cf4b2a2c976795ac430d24a1f43d25a6ba/rpds_py-0.30.0-cp313-cp313t-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:39c02563fc592411c2c61d26b6c5fe1e51eaa44a75aa2c8735ca88b0d9599daa", size = 408149, upload-time = "2025-11-30T20:23:12.864Z" }, - { url = "https://files.pythonhosted.org/packages/60/ca/780cf3b1a32b18c0f05c441958d3758f02544f1d613abf9488cd78876378/rpds_py-0.30.0-cp313-cp313t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:51a1234d8febafdfd33a42d97da7a43f5dcb120c1060e352a3fbc0c6d36e2083", size = 383843, upload-time = "2025-11-30T20:23:14.638Z" }, - { url = "https://files.pythonhosted.org/packages/82/86/d5f2e04f2aa6247c613da0c1dd87fcd08fa17107e858193566048a1e2f0a/rpds_py-0.30.0-cp313-cp313t-manylinux_2_31_riscv64.whl", hash = "sha256:eb2c4071ab598733724c08221091e8d80e89064cd472819285a9ab0f24bcedb9", size = 396507, upload-time = "2025-11-30T20:23:16.105Z" }, - { url = "https://files.pythonhosted.org/packages/4b/9a/453255d2f769fe44e07ea9785c8347edaf867f7026872e76c1ad9f7bed92/rpds_py-0.30.0-cp313-cp313t-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:6bdfdb946967d816e6adf9a3d8201bfad269c67efe6cefd7093ef959683c8de0", size = 414949, upload-time = "2025-11-30T20:23:17.539Z" }, - { url = "https://files.pythonhosted.org/packages/a3/31/622a86cdc0c45d6df0e9ccb6becdba5074735e7033c20e401a6d9d0e2ca0/rpds_py-0.30.0-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:c77afbd5f5250bf27bf516c7c4a016813eb2d3e116139aed0096940c5982da94", size = 565790, upload-time = "2025-11-30T20:23:19.029Z" }, - { url = "https://files.pythonhosted.org/packages/1c/5d/15bbf0fb4a3f58a3b1c67855ec1efcc4ceaef4e86644665fff03e1b66d8d/rpds_py-0.30.0-cp313-cp313t-musllinux_1_2_i686.whl", hash = "sha256:61046904275472a76c8c90c9ccee9013d70a6d0f73eecefd38c1ae7c39045a08", size = 590217, upload-time = "2025-11-30T20:23:20.885Z" }, - { url = "https://files.pythonhosted.org/packages/6d/61/21b8c41f68e60c8cc3b2e25644f0e3681926020f11d06ab0b78e3c6bbff1/rpds_py-0.30.0-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:4c5f36a861bc4b7da6516dbdf302c55313afa09b81931e8280361a4f6c9a2d27", size = 555806, upload-time = "2025-11-30T20:23:22.488Z" }, - { url = "https://files.pythonhosted.org/packages/f9/39/7e067bb06c31de48de3eb200f9fc7c58982a4d3db44b07e73963e10d3be9/rpds_py-0.30.0-cp313-cp313t-win32.whl", hash = "sha256:3d4a69de7a3e50ffc214ae16d79d8fbb0922972da0356dcf4d0fdca2878559c6", size = 211341, upload-time = "2025-11-30T20:23:24.449Z" }, - { url = "https://files.pythonhosted.org/packages/0a/4d/222ef0b46443cf4cf46764d9c630f3fe4abaa7245be9417e56e9f52b8f65/rpds_py-0.30.0-cp313-cp313t-win_amd64.whl", hash = "sha256:f14fc5df50a716f7ece6a80b6c78bb35ea2ca47c499e422aa4463455dd96d56d", size = 225768, upload-time = "2025-11-30T20:23:25.908Z" }, - { url = "https://files.pythonhosted.org/packages/86/81/dad16382ebbd3d0e0328776d8fd7ca94220e4fa0798d1dc5e7da48cb3201/rpds_py-0.30.0-cp314-cp314-macosx_10_12_x86_64.whl", hash = "sha256:68f19c879420aa08f61203801423f6cd5ac5f0ac4ac82a2368a9fcd6a9a075e0", size = 362099, upload-time = "2025-11-30T20:23:27.316Z" }, - { url = "https://files.pythonhosted.org/packages/2b/60/19f7884db5d5603edf3c6bce35408f45ad3e97e10007df0e17dd57af18f8/rpds_py-0.30.0-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:ec7c4490c672c1a0389d319b3a9cfcd098dcdc4783991553c332a15acf7249be", size = 353192, upload-time = "2025-11-30T20:23:29.151Z" }, - { url = "https://files.pythonhosted.org/packages/bf/c4/76eb0e1e72d1a9c4703c69607cec123c29028bff28ce41588792417098ac/rpds_py-0.30.0-cp314-cp314-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:f251c812357a3fed308d684a5079ddfb9d933860fc6de89f2b7ab00da481e65f", size = 384080, upload-time = "2025-11-30T20:23:30.785Z" }, - { url = "https://files.pythonhosted.org/packages/72/87/87ea665e92f3298d1b26d78814721dc39ed8d2c74b86e83348d6b48a6f31/rpds_py-0.30.0-cp314-cp314-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:ac98b175585ecf4c0348fd7b29c3864bda53b805c773cbf7bfdaffc8070c976f", size = 394841, upload-time = "2025-11-30T20:23:32.209Z" }, - { url = "https://files.pythonhosted.org/packages/77/ad/7783a89ca0587c15dcbf139b4a8364a872a25f861bdb88ed99f9b0dec985/rpds_py-0.30.0-cp314-cp314-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:3e62880792319dbeb7eb866547f2e35973289e7d5696c6e295476448f5b63c87", size = 516670, upload-time = "2025-11-30T20:23:33.742Z" }, - { url = "https://files.pythonhosted.org/packages/5b/3c/2882bdac942bd2172f3da574eab16f309ae10a3925644e969536553cb4ee/rpds_py-0.30.0-cp314-cp314-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:4e7fc54e0900ab35d041b0601431b0a0eb495f0851a0639b6ef90f7741b39a18", size = 408005, upload-time = "2025-11-30T20:23:35.253Z" }, - { url = "https://files.pythonhosted.org/packages/ce/81/9a91c0111ce1758c92516a3e44776920b579d9a7c09b2b06b642d4de3f0f/rpds_py-0.30.0-cp314-cp314-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:47e77dc9822d3ad616c3d5759ea5631a75e5809d5a28707744ef79d7a1bcfcad", size = 382112, upload-time = "2025-11-30T20:23:36.842Z" }, - { url = "https://files.pythonhosted.org/packages/cf/8e/1da49d4a107027e5fbc64daeab96a0706361a2918da10cb41769244b805d/rpds_py-0.30.0-cp314-cp314-manylinux_2_31_riscv64.whl", hash = "sha256:b4dc1a6ff022ff85ecafef7979a2c6eb423430e05f1165d6688234e62ba99a07", size = 399049, upload-time = "2025-11-30T20:23:38.343Z" }, - { url = "https://files.pythonhosted.org/packages/df/5a/7ee239b1aa48a127570ec03becbb29c9d5a9eb092febbd1699d567cae859/rpds_py-0.30.0-cp314-cp314-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:4559c972db3a360808309e06a74628b95eaccbf961c335c8fe0d590cf587456f", size = 415661, upload-time = "2025-11-30T20:23:40.263Z" }, - { url = "https://files.pythonhosted.org/packages/70/ea/caa143cf6b772f823bc7929a45da1fa83569ee49b11d18d0ada7f5ee6fd6/rpds_py-0.30.0-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:0ed177ed9bded28f8deb6ab40c183cd1192aa0de40c12f38be4d59cd33cb5c65", size = 565606, upload-time = "2025-11-30T20:23:42.186Z" }, - { url = "https://files.pythonhosted.org/packages/64/91/ac20ba2d69303f961ad8cf55bf7dbdb4763f627291ba3d0d7d67333cced9/rpds_py-0.30.0-cp314-cp314-musllinux_1_2_i686.whl", hash = "sha256:ad1fa8db769b76ea911cb4e10f049d80bf518c104f15b3edb2371cc65375c46f", size = 591126, upload-time = "2025-11-30T20:23:44.086Z" }, - { url = "https://files.pythonhosted.org/packages/21/20/7ff5f3c8b00c8a95f75985128c26ba44503fb35b8e0259d812766ea966c7/rpds_py-0.30.0-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:46e83c697b1f1c72b50e5ee5adb4353eef7406fb3f2043d64c33f20ad1c2fc53", size = 553371, upload-time = "2025-11-30T20:23:46.004Z" }, - { url = "https://files.pythonhosted.org/packages/72/c7/81dadd7b27c8ee391c132a6b192111ca58d866577ce2d9b0ca157552cce0/rpds_py-0.30.0-cp314-cp314-win32.whl", hash = "sha256:ee454b2a007d57363c2dfd5b6ca4a5d7e2c518938f8ed3b706e37e5d470801ed", size = 215298, upload-time = "2025-11-30T20:23:47.696Z" }, - { url = "https://files.pythonhosted.org/packages/3e/d2/1aaac33287e8cfb07aab2e6b8ac1deca62f6f65411344f1433c55e6f3eb8/rpds_py-0.30.0-cp314-cp314-win_amd64.whl", hash = "sha256:95f0802447ac2d10bcc69f6dc28fe95fdf17940367b21d34e34c737870758950", size = 228604, upload-time = "2025-11-30T20:23:49.501Z" }, - { url = "https://files.pythonhosted.org/packages/e8/95/ab005315818cc519ad074cb7784dae60d939163108bd2b394e60dc7b5461/rpds_py-0.30.0-cp314-cp314-win_arm64.whl", hash = "sha256:613aa4771c99f03346e54c3f038e4cc574ac09a3ddfb0e8878487335e96dead6", size = 222391, upload-time = "2025-11-30T20:23:50.96Z" }, - { url = "https://files.pythonhosted.org/packages/9e/68/154fe0194d83b973cdedcdcc88947a2752411165930182ae41d983dcefa6/rpds_py-0.30.0-cp314-cp314t-macosx_10_12_x86_64.whl", hash = "sha256:7e6ecfcb62edfd632e56983964e6884851786443739dbfe3582947e87274f7cb", size = 364868, upload-time = "2025-11-30T20:23:52.494Z" }, - { url = "https://files.pythonhosted.org/packages/83/69/8bbc8b07ec854d92a8b75668c24d2abcb1719ebf890f5604c61c9369a16f/rpds_py-0.30.0-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:a1d0bc22a7cdc173fedebb73ef81e07faef93692b8c1ad3733b67e31e1b6e1b8", size = 353747, upload-time = "2025-11-30T20:23:54.036Z" }, - { url = "https://files.pythonhosted.org/packages/ab/00/ba2e50183dbd9abcce9497fa5149c62b4ff3e22d338a30d690f9af970561/rpds_py-0.30.0-cp314-cp314t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:0d08f00679177226c4cb8c5265012eea897c8ca3b93f429e546600c971bcbae7", size = 383795, upload-time = "2025-11-30T20:23:55.556Z" }, - { url = "https://files.pythonhosted.org/packages/05/6f/86f0272b84926bcb0e4c972262f54223e8ecc556b3224d281e6598fc9268/rpds_py-0.30.0-cp314-cp314t-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:5965af57d5848192c13534f90f9dd16464f3c37aaf166cc1da1cae1fd5a34898", size = 393330, upload-time = "2025-11-30T20:23:57.033Z" }, - { url = "https://files.pythonhosted.org/packages/cb/e9/0e02bb2e6dc63d212641da45df2b0bf29699d01715913e0d0f017ee29438/rpds_py-0.30.0-cp314-cp314t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:9a4e86e34e9ab6b667c27f3211ca48f73dba7cd3d90f8d5b11be56e5dbc3fb4e", size = 518194, upload-time = "2025-11-30T20:23:58.637Z" }, - { url = "https://files.pythonhosted.org/packages/ee/ca/be7bca14cf21513bdf9c0606aba17d1f389ea2b6987035eb4f62bd923f25/rpds_py-0.30.0-cp314-cp314t-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:e5d3e6b26f2c785d65cc25ef1e5267ccbe1b069c5c21b8cc724efee290554419", size = 408340, upload-time = "2025-11-30T20:24:00.2Z" }, - { url = "https://files.pythonhosted.org/packages/c2/c7/736e00ebf39ed81d75544c0da6ef7b0998f8201b369acf842f9a90dc8fce/rpds_py-0.30.0-cp314-cp314t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:626a7433c34566535b6e56a1b39a7b17ba961e97ce3b80ec62e6f1312c025551", size = 383765, upload-time = "2025-11-30T20:24:01.759Z" }, - { url = "https://files.pythonhosted.org/packages/4a/3f/da50dfde9956aaf365c4adc9533b100008ed31aea635f2b8d7b627e25b49/rpds_py-0.30.0-cp314-cp314t-manylinux_2_31_riscv64.whl", hash = "sha256:acd7eb3f4471577b9b5a41baf02a978e8bdeb08b4b355273994f8b87032000a8", size = 396834, upload-time = "2025-11-30T20:24:03.687Z" }, - { url = "https://files.pythonhosted.org/packages/4e/00/34bcc2565b6020eab2623349efbdec810676ad571995911f1abdae62a3a0/rpds_py-0.30.0-cp314-cp314t-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:fe5fa731a1fa8a0a56b0977413f8cacac1768dad38d16b3a296712709476fbd5", size = 415470, upload-time = "2025-11-30T20:24:05.232Z" }, - { url = "https://files.pythonhosted.org/packages/8c/28/882e72b5b3e6f718d5453bd4d0d9cf8df36fddeb4ddbbab17869d5868616/rpds_py-0.30.0-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:74a3243a411126362712ee1524dfc90c650a503502f135d54d1b352bd01f2404", size = 565630, upload-time = "2025-11-30T20:24:06.878Z" }, - { url = "https://files.pythonhosted.org/packages/3b/97/04a65539c17692de5b85c6e293520fd01317fd878ea1995f0367d4532fb1/rpds_py-0.30.0-cp314-cp314t-musllinux_1_2_i686.whl", hash = "sha256:3e8eeb0544f2eb0d2581774be4c3410356eba189529a6b3e36bbbf9696175856", size = 591148, upload-time = "2025-11-30T20:24:08.445Z" }, - { url = "https://files.pythonhosted.org/packages/85/70/92482ccffb96f5441aab93e26c4d66489eb599efdcf96fad90c14bbfb976/rpds_py-0.30.0-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:dbd936cde57abfee19ab3213cf9c26be06d60750e60a8e4dd85d1ab12c8b1f40", size = 556030, upload-time = "2025-11-30T20:24:10.956Z" }, - { url = "https://files.pythonhosted.org/packages/20/53/7c7e784abfa500a2b6b583b147ee4bb5a2b3747a9166bab52fec4b5b5e7d/rpds_py-0.30.0-cp314-cp314t-win32.whl", hash = "sha256:dc824125c72246d924f7f796b4f63c1e9dc810c7d9e2355864b3c3a73d59ade0", size = 211570, upload-time = "2025-11-30T20:24:12.735Z" }, - { url = "https://files.pythonhosted.org/packages/d0/02/fa464cdfbe6b26e0600b62c528b72d8608f5cc49f96b8d6e38c95d60c676/rpds_py-0.30.0-cp314-cp314t-win_amd64.whl", hash = "sha256:27f4b0e92de5bfbc6f86e43959e6edd1425c33b5e69aab0984a72047f2bcf1e3", size = 226532, upload-time = "2025-11-30T20:24:14.634Z" }, - { url = "https://files.pythonhosted.org/packages/69/71/3f34339ee70521864411f8b6992e7ab13ac30d8e4e3309e07c7361767d91/rpds_py-0.30.0-pp311-pypy311_pp73-macosx_10_12_x86_64.whl", hash = "sha256:c2262bdba0ad4fc6fb5545660673925c2d2a5d9e2e0fb603aad545427be0fc58", size = 372292, upload-time = "2025-11-30T20:24:16.537Z" }, - { url = "https://files.pythonhosted.org/packages/57/09/f183df9b8f2d66720d2ef71075c59f7e1b336bec7ee4c48f0a2b06857653/rpds_py-0.30.0-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:ee6af14263f25eedc3bb918a3c04245106a42dfd4f5c2285ea6f997b1fc3f89a", size = 362128, upload-time = "2025-11-30T20:24:18.086Z" }, - { url = "https://files.pythonhosted.org/packages/7a/68/5c2594e937253457342e078f0cc1ded3dd7b2ad59afdbf2d354869110a02/rpds_py-0.30.0-pp311-pypy311_pp73-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:3adbb8179ce342d235c31ab8ec511e66c73faa27a47e076ccc92421add53e2bb", size = 391542, upload-time = "2025-11-30T20:24:20.092Z" }, - { url = "https://files.pythonhosted.org/packages/49/5c/31ef1afd70b4b4fbdb2800249f34c57c64beb687495b10aec0365f53dfc4/rpds_py-0.30.0-pp311-pypy311_pp73-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:250fa00e9543ac9b97ac258bd37367ff5256666122c2d0f2bc97577c60a1818c", size = 404004, upload-time = "2025-11-30T20:24:22.231Z" }, - { url = "https://files.pythonhosted.org/packages/e3/63/0cfbea38d05756f3440ce6534d51a491d26176ac045e2707adc99bb6e60a/rpds_py-0.30.0-pp311-pypy311_pp73-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:9854cf4f488b3d57b9aaeb105f06d78e5529d3145b1e4a41750167e8c213c6d3", size = 527063, upload-time = "2025-11-30T20:24:24.302Z" }, - { url = "https://files.pythonhosted.org/packages/42/e6/01e1f72a2456678b0f618fc9a1a13f882061690893c192fcad9f2926553a/rpds_py-0.30.0-pp311-pypy311_pp73-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:993914b8e560023bc0a8bf742c5f303551992dcb85e247b1e5c7f4a7d145bda5", size = 413099, upload-time = "2025-11-30T20:24:25.916Z" }, - { url = "https://files.pythonhosted.org/packages/b8/25/8df56677f209003dcbb180765520c544525e3ef21ea72279c98b9aa7c7fb/rpds_py-0.30.0-pp311-pypy311_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:58edca431fb9b29950807e301826586e5bbf24163677732429770a697ffe6738", size = 392177, upload-time = "2025-11-30T20:24:27.834Z" }, - { url = "https://files.pythonhosted.org/packages/4a/b4/0a771378c5f16f8115f796d1f437950158679bcd2a7c68cf251cfb00ed5b/rpds_py-0.30.0-pp311-pypy311_pp73-manylinux_2_31_riscv64.whl", hash = "sha256:dea5b552272a944763b34394d04577cf0f9bd013207bc32323b5a89a53cf9c2f", size = 406015, upload-time = "2025-11-30T20:24:29.457Z" }, - { url = "https://files.pythonhosted.org/packages/36/d8/456dbba0af75049dc6f63ff295a2f92766b9d521fa00de67a2bd6427d57a/rpds_py-0.30.0-pp311-pypy311_pp73-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:ba3af48635eb83d03f6c9735dfb21785303e73d22ad03d489e88adae6eab8877", size = 423736, upload-time = "2025-11-30T20:24:31.22Z" }, - { url = "https://files.pythonhosted.org/packages/13/64/b4d76f227d5c45a7e0b796c674fd81b0a6c4fbd48dc29271857d8219571c/rpds_py-0.30.0-pp311-pypy311_pp73-musllinux_1_2_aarch64.whl", hash = "sha256:dff13836529b921e22f15cb099751209a60009731a68519630a24d61f0b1b30a", size = 573981, upload-time = "2025-11-30T20:24:32.934Z" }, - { url = "https://files.pythonhosted.org/packages/20/91/092bacadeda3edf92bf743cc96a7be133e13a39cdbfd7b5082e7ab638406/rpds_py-0.30.0-pp311-pypy311_pp73-musllinux_1_2_i686.whl", hash = "sha256:1b151685b23929ab7beec71080a8889d4d6d9fa9a983d213f07121205d48e2c4", size = 599782, upload-time = "2025-11-30T20:24:35.169Z" }, - { url = "https://files.pythonhosted.org/packages/d1/b7/b95708304cd49b7b6f82fdd039f1748b66ec2b21d6a45180910802f1abf1/rpds_py-0.30.0-pp311-pypy311_pp73-musllinux_1_2_x86_64.whl", hash = "sha256:ac37f9f516c51e5753f27dfdef11a88330f04de2d564be3991384b2f3535d02e", size = 562191, upload-time = "2025-11-30T20:24:36.853Z" }, -] - -[[package]] -name = "rpds-py" -version = "2026.5.1" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version >= '3.14' and sys_platform == 'win32'", - "python_full_version >= '3.14' and sys_platform != 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform == 'win32'", - "python_full_version == '3.11.*' and sys_platform == 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform != 'win32'", - "python_full_version == '3.11.*' and sys_platform != 'win32'", -] -sdist = { url = "https://files.pythonhosted.org/packages/2e/43/25a8dcd3feedd735039a8f0b5b7e3b118232b5eae288c4fd9ab200d41094/rpds_py-2026.5.1.tar.gz", hash = "sha256:07b24fea40541e28570e5b795a4a38fbdcd12550c06bd0748005ecc8116ca256", size = 64459, upload-time = "2026-05-28T12:02:13.232Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/4f/a0/acf8b6fc20bfdcd3a45bd3f57680fb198e157b7e997b9123b10763798bd2/rpds_py-2026.5.1-cp311-cp311-macosx_10_12_x86_64.whl", hash = "sha256:3397a5ed7174dc2786bb214030232fc36fe8e5584fec43a9952cc542b1a12036", size = 355609, upload-time = "2026-05-28T11:58:50.78Z" }, - { url = "https://files.pythonhosted.org/packages/b6/95/f8203fd997484b1690a6869cd0e503b6c3c6be55b0ecc36d1a491fe742f0/rpds_py-2026.5.1-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:99ab6ba7bfa2cb0f96a04e3652355bf04e3f51aceb1e943b8541dab7ba4828cc", size = 348460, upload-time = "2026-05-28T11:58:52.374Z" }, - { url = "https://files.pythonhosted.org/packages/33/8c/b47326ad2f0be545a5e5c1a55937a12afaea7d392ba2837bb9680f57e6c9/rpds_py-2026.5.1-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:d0efbe45632665e53e3db8fe1e5692db58fc5cb9bab4459d570b83efefe11164", size = 381031, upload-time = "2026-05-28T11:58:53.775Z" }, - { url = "https://files.pythonhosted.org/packages/22/0b/e83bbd97ffac6f6389b605cd4e1c8ac5761dc7e977769c9255d8c5adb7bd/rpds_py-2026.5.1-cp311-cp311-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:01d17b29c0c23d82b1f4751147ec49cf451f1fc2554eb9ef5f957e55d2656ead", size = 387121, upload-time = "2026-05-28T11:58:55.243Z" }, - { url = "https://files.pythonhosted.org/packages/fd/0e/d285d1bc8864245919c61e1ca82263e4a66d337759c3a4cef72766ff9afc/rpds_py-2026.5.1-cp311-cp311-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:7559f72b94ae52659086c595dfa017cde03155f7832071d30959049052cb3ece", size = 501026, upload-time = "2026-05-28T11:58:56.788Z" }, - { url = "https://files.pythonhosted.org/packages/86/06/ccb2109a1e543437b5e43816f2b43b9554cc6783145528a4e3711e05c011/rpds_py-2026.5.1-cp311-cp311-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:9e25b7088f9ccbfc0dfcaa52bf969300ca229e10ecf758974ebcbb080a4b37bb", size = 391865, upload-time = "2026-05-28T11:58:58.298Z" }, - { url = "https://files.pythonhosted.org/packages/3d/33/237173db1cfef10105b3839a24de00eb8d2a523711add4632447cdf0aedd/rpds_py-2026.5.1-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:613fc4ee9eaef26dc5840666214dd6fbcebcf32f46e76f4abc473059f4e13dda", size = 378012, upload-time = "2026-05-28T11:58:59.589Z" }, - { url = "https://files.pythonhosted.org/packages/97/64/1eae54e34d5161f9969295e80bd6b62a55f2b6ac5f2a5b60d02c2140e758/rpds_py-2026.5.1-cp311-cp311-manylinux_2_31_riscv64.whl", hash = "sha256:85264a90ff4c05c1568dd65f5921c837614b67c60358fb4c17df3b7f2e90690a", size = 391111, upload-time = "2026-05-28T11:59:01.104Z" }, - { url = "https://files.pythonhosted.org/packages/d8/34/5bb334a5a0f65d77869217c4654f34c78a7d11b93938a3c076a2edeafc52/rpds_py-2026.5.1-cp311-cp311-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:fe71bca7d547acb17027c7fd1624ff8aae623499c498d3e7011182c4de5c25e0", size = 409225, upload-time = "2026-05-28T11:59:02.433Z" }, - { url = "https://files.pythonhosted.org/packages/16/0f/007ec21283b5b040b4ec3bd95e0402591e22bfa7d5c93dfe01c465c2d2d7/rpds_py-2026.5.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:a05fa4f41f37ec97c9c260441a940450a192f78d774d2b097eee1379f1e1246a", size = 556487, upload-time = "2026-05-28T11:59:04.012Z" }, - { url = "https://files.pythonhosted.org/packages/ff/10/5437c94508169b6b22d8418fef7a66e9ffb5f3b9e9c94460f2eedafe06ff/rpds_py-2026.5.1-cp311-cp311-musllinux_1_2_i686.whl", hash = "sha256:df1d2a1996755b24b9ecee92cb4d36c28f86f464a6a173349c26bab41e94b8c2", size = 620798, upload-time = "2026-05-28T11:59:05.485Z" }, - { url = "https://files.pythonhosted.org/packages/e0/d5/9937dce4d6bda74157b954e7d1460db05a22f5929dccfeeba1ed27a93df0/rpds_py-2026.5.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:8895840ac4809e5f60c88fd07617cd71326e73d6e5a8aa783c5c0f7c24985de2", size = 584053, upload-time = "2026-05-28T11:59:06.837Z" }, - { url = "https://files.pythonhosted.org/packages/6c/31/750617dd0ae1752471bf43f9e41d263398fae7cde7849d23b8574a70e617/rpds_py-2026.5.1-cp311-cp311-win32.whl", hash = "sha256:3684a59b158a7683aaeb8e25352e9a9dd2122cec78f2d8530266e4f91b4c7b3f", size = 214390, upload-time = "2026-05-28T11:59:08.402Z" }, - { url = "https://files.pythonhosted.org/packages/3c/bb/3dcab0e1d9516303f2eb672a5d6f62eca5a69e2886301e9c8c54b520c39b/rpds_py-2026.5.1-cp311-cp311-win_amd64.whl", hash = "sha256:7bd530e6a530bb3ea892f194fafa455f3516ac25ecf7143fd33c09be62b0470a", size = 231097, upload-time = "2026-05-28T11:59:09.786Z" }, - { url = "https://files.pythonhosted.org/packages/49/d6/c6bbf5cb1cf12b9732df8074b57f6ef8341ba884c95d40632ae8bddb44e4/rpds_py-2026.5.1-cp311-cp311-win_arm64.whl", hash = "sha256:0a5ae4dbe43c1076983b72616496919872ae7bbe7a1e21cc48336bc3154d130b", size = 226361, upload-time = "2026-05-28T11:59:11.079Z" }, - { url = "https://files.pythonhosted.org/packages/d4/e7/a78582dc57caa592dcc7d4fb69b61390561e908eb3d2f5df5928a8e354c0/rpds_py-2026.5.1-cp312-cp312-macosx_10_12_x86_64.whl", hash = "sha256:3abe24a66e57adcfa645d718063a5fa5103ecc71ddbf26d78af8f9368018ff1d", size = 353040, upload-time = "2026-05-28T11:59:12.531Z" }, - { url = "https://files.pythonhosted.org/packages/a3/43/35e3f136343aef451e545ce8c38d36c2f93c0ed88703db8b64ba2b205c68/rpds_py-2026.5.1-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:58b1d94308ddf0b1982f61f2eb54bf92997c9ece8a8093ef014250f4a517906c", size = 345775, upload-time = "2026-05-28T11:59:13.827Z" }, - { url = "https://files.pythonhosted.org/packages/20/e1/0f2160c5982d3157734d5cb3ed63d8b2d583a73c9864f77b666449f32cf8/rpds_py-2026.5.1-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:0fa92420128dadce7f54bd73ba1825a273e9268fe9e35dbf7e6362890efa4e08", size = 376329, upload-time = "2026-05-28T11:59:15.271Z" }, - { url = "https://files.pythonhosted.org/packages/d0/11/ee0ba42aff83bf4effdbc576673c6be64c5e173978c3f6d537e94482f77d/rpds_py-2026.5.1-cp312-cp312-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:ca653c6546386227cd9800d1bef6a348099acf8db4250341da6d90f663d6dfcb", size = 383539, upload-time = "2026-05-28T11:59:16.665Z" }, - { url = "https://files.pythonhosted.org/packages/11/df/d94aa6a499d4ac40afe2d7620f2c597fd3c0f182e854ad7cf3f596a81cb6/rpds_py-2026.5.1-cp312-cp312-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:66c93681c4729e4e3ecba31b8179fae083ff3118841672835140338b4b9867c1", size = 494674, upload-time = "2026-05-28T11:59:17.991Z" }, - { url = "https://files.pythonhosted.org/packages/1f/75/33d30f43bb2f458de11979486a591b1bf6e5651765ed1704c6197c2dc773/rpds_py-2026.5.1-cp312-cp312-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:40ff257542e04796880e011e15cd4dc21c2599975df2aaa8f2c8495ca574e1a5", size = 389268, upload-time = "2026-05-28T11:59:19.434Z" }, - { url = "https://files.pythonhosted.org/packages/f4/1e/2c9096fc19d5fd084b0184ca2b651e659aa0a37e6fdbecf6ece47f147fe1/rpds_py-2026.5.1-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:b6825cc329b290e93c5f6a9be2393118a763f6ccf6abd83704e0c102ca583644", size = 376280, upload-time = "2026-05-28T11:59:21Z" }, - { url = "https://files.pythonhosted.org/packages/b9/e5/61ec9f8be8211ea7f48448195549e4aaf02004083475493b0e137702ecb2/rpds_py-2026.5.1-cp312-cp312-manylinux_2_31_riscv64.whl", hash = "sha256:de42116e69cb53b911cc34aee5ab98f36c597b822545045d49e938818b99e5e4", size = 387233, upload-time = "2026-05-28T11:59:22.454Z" }, - { url = "https://files.pythonhosted.org/packages/0d/ca/bcec1005c4f4a234f92a29078631fee49206c7265ccae966f18fd332e80e/rpds_py-2026.5.1-cp312-cp312-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:c0f920015df2a504bebaba6d4c31ccf3fcf942f92655c086da30b671aad19aa6", size = 405009, upload-time = "2026-05-28T11:59:23.845Z" }, - { url = "https://files.pythonhosted.org/packages/72/e6/4d5718c5cf26c522dc7c9999e238da1e77380b81d0c5d1df11e271ddfeb1/rpds_py-2026.5.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:0408a24e44feb919423dc6d9da677cb5cddb894d2ca9e763967d156d9c60fab4", size = 553113, upload-time = "2026-05-28T11:59:25.184Z" }, - { url = "https://files.pythonhosted.org/packages/d4/25/2ee807bdb3e1f0b7eddf7782acd5665a8b5205a331a7d7244a52c4812fd9/rpds_py-2026.5.1-cp312-cp312-musllinux_1_2_i686.whl", hash = "sha256:cea68bcd53467561ae2f96a6bdad1544299ba97b5b0ddcd5ac3d376e5c781c24", size = 618838, upload-time = "2026-05-28T11:59:26.749Z" }, - { url = "https://files.pythonhosted.org/packages/6a/c1/7d4c26f167f8c41501cc073d30ee22082b16ce358cf5b00ec97cbc7804ea/rpds_py-2026.5.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:4be8b1d2a705cc37d08256004e1d07de143fa0075c8e85a3df020b776f62b732", size = 582436, upload-time = "2026-05-28T11:59:28.11Z" }, - { url = "https://files.pythonhosted.org/packages/04/1d/9d12b0a337bab46f4769f8857f4007e3b2d639e14f9a44a0efe157696e64/rpds_py-2026.5.1-cp312-cp312-win32.whl", hash = "sha256:6736718bd4fc49cbcb538ba30516fdbef161522acefb739657d48b97bd864fed", size = 212734, upload-time = "2026-05-28T11:59:29.689Z" }, - { url = "https://files.pythonhosted.org/packages/c5/93/e4116f2de7f56bc7406a76033dc501811ddeb22b7f056b92d632871ebb0c/rpds_py-2026.5.1-cp312-cp312-win_amd64.whl", hash = "sha256:0a7d1eec967df0e9b22614a5e177622e0c89611d03727fa0cb48e45028907870", size = 229045, upload-time = "2026-05-28T11:59:31.033Z" }, - { url = "https://files.pythonhosted.org/packages/cb/53/6c3419d85eb2ec5938a37627c585b42d76a63bb731d6e42ed4b079ebf486/rpds_py-2026.5.1-cp312-cp312-win_arm64.whl", hash = "sha256:1841d067089e117142d79b98aa0df2f08b52f2ecc1819dd2700636c0db74a473", size = 223967, upload-time = "2026-05-28T11:59:32.318Z" }, - { url = "https://files.pythonhosted.org/packages/6c/32/14c961ad295f490eb0849ada8b79683e93a59b9de3afdd983eaf55fa6867/rpds_py-2026.5.1-cp313-cp313-macosx_10_12_x86_64.whl", hash = "sha256:efef4ac29c6ff495531eb17ee705b62841ecaa291b7c7077e848ea03e237164d", size = 352787, upload-time = "2026-05-28T11:59:33.655Z" }, - { url = "https://files.pythonhosted.org/packages/ca/bb/d1b85117967c11191441a7274ae616c65d93901d082c588f89a50a8da5ae/rpds_py-2026.5.1-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:c39f5b67a8a2e67179ada2a954227d670fe65fa9098457f698f56ddf248709b3", size = 345179, upload-time = "2026-05-28T11:59:35Z" }, - { url = "https://files.pythonhosted.org/packages/7c/46/d84105f062e626a1b233f863907288a4708c2d833b8b4c6fb2764bc080c0/rpds_py-2026.5.1-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:b5c30f3f04eef4fbd362226a6f31d7c8895ca4fbb6e0b790f6890a98d8da8559", size = 376173, upload-time = "2026-05-28T11:59:36.43Z" }, - { url = "https://files.pythonhosted.org/packages/e2/ae/469d7959ce5b1201e1de135dc735b86db3b35dd0d1734f6a44246d5f061c/rpds_py-2026.5.1-cp313-cp313-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:277f6c82f0580848796c7ecc8a7173aa3bfb928e4ff831261c2f60a81dc270db", size = 383162, upload-time = "2026-05-28T11:59:37.995Z" }, - { url = "https://files.pythonhosted.org/packages/dc/a2/57853d31a1116a561aa072794602ad3f6341e18d70a8523f1bd5b9fc1e5a/rpds_py-2026.5.1-cp313-cp313-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:63c2c4c213f1a4e3f3de28ecab029dbdee976324e729c0d7a55211be72576b02", size = 495093, upload-time = "2026-05-28T11:59:39.453Z" }, - { url = "https://files.pythonhosted.org/packages/99/63/3a8eabcad9314b7daf5c65f451d2c33d989235cd8a5762186cf2c3f5a4f8/rpds_py-2026.5.1-cp313-cp313-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:3350ec808fb538fe71a1f94dfaa0e29c598dfad805ce49f0caec5ae3183c652b", size = 389829, upload-time = "2026-05-28T11:59:40.896Z" }, - { url = "https://files.pythonhosted.org/packages/4b/25/05678d97fc25e2622df14dc530fb82023174ecfff6733991ed0d78f167bd/rpds_py-2026.5.1-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:b1b964e3ab599e718dc46c018d104b1ebc007cbc6567d827c94a687fca56d77e", size = 374786, upload-time = "2026-05-28T11:59:42.626Z" }, - { url = "https://files.pythonhosted.org/packages/88/d1/8c90b6431e80a3b91b284a5c7c8c0c4f9c006444d90477a740d6e0f9c694/rpds_py-2026.5.1-cp313-cp313-manylinux_2_31_riscv64.whl", hash = "sha256:19cb09fab7b7fc96b2a6e28f2e34b72a3705ff27b37edb77455316e5d3f3dc9b", size = 386920, upload-time = "2026-05-28T11:59:44.124Z" }, - { url = "https://files.pythonhosted.org/packages/ff/99/4638f672ab356682d633ee0da9255f5b67ce6efd0b85eb94ad3e255e65a5/rpds_py-2026.5.1-cp313-cp313-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:abe76bcdba31e576cb83eeb8797aa0d882b738fef6dc65d0601fc753806a5b46", size = 405059, upload-time = "2026-05-28T11:59:47.177Z" }, - { url = "https://files.pythonhosted.org/packages/66/3f/3546524b6eb4cc2e1f363a3d638fa52f6c24faae3500c25fb488b02f1740/rpds_py-2026.5.1-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:8bff7073db3899158fff55ebf57b113a67030af26f80a18978f9f0aa60250ddf", size = 553030, upload-time = "2026-05-28T11:59:48.603Z" }, - { url = "https://files.pythonhosted.org/packages/c6/c3/7b3388c796fcf471bd17194242d4dc1a7608567c0fa422bcc1c5e79f9c1e/rpds_py-2026.5.1-cp313-cp313-musllinux_1_2_i686.whl", hash = "sha256:8ba264fa49be666cd9cc56bf34ec7002fb3d27a4aee5bcb4d43d0d18feb1bb6f", size = 618975, upload-time = "2026-05-28T11:59:50.314Z" }, - { url = "https://files.pythonhosted.org/packages/61/1e/a3cb07f2795075d1d88efddae2f541359fde5f08c81ee114c29c2949c90a/rpds_py-2026.5.1-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:4860b603ddda0475a8885499b3729e90229d480105b42651962a5397d995fa89", size = 581178, upload-time = "2026-05-28T11:59:51.673Z" }, - { url = "https://files.pythonhosted.org/packages/a1/74/e758c03a5ef46f04c37f2651a2893db846d569ba8a7bca469d4b58939bcd/rpds_py-2026.5.1-cp313-cp313-win32.whl", hash = "sha256:7944270ae71383f6e2657dd7d5ce4eeb4ac2d0059a6738f0510583d462ab4842", size = 212481, upload-time = "2026-05-28T11:59:53.148Z" }, - { url = "https://files.pythonhosted.org/packages/70/ec/a2aca432db9c7359b40fa393eeeaa0d166c2f70175be956e75fa24197c44/rpds_py-2026.5.1-cp313-cp313-win_amd64.whl", hash = "sha256:88647f43a73c4e01be19b04ceef0c8d3a1958153604d13c773becd8016f2a0cf", size = 228519, upload-time = "2026-05-28T11:59:54.505Z" }, - { url = "https://files.pythonhosted.org/packages/29/60/a73bfdd45b096574556acf303bbd9fa9eed36ca8a818b514e2a5d5fe2b9d/rpds_py-2026.5.1-cp313-cp313-win_arm64.whl", hash = "sha256:453895624ecf7db7063b1004e44037522bbaef9ff6a945e59bc71662d7a03abd", size = 223446, upload-time = "2026-05-28T11:59:56.081Z" }, - { url = "https://files.pythonhosted.org/packages/18/e2/408105fd611823f00882aea810f3989a30d26b1bab8b6beb20f98c724e0e/rpds_py-2026.5.1-cp313-cp313t-macosx_10_12_x86_64.whl", hash = "sha256:b4e4bc98639ec915f512fde3aa7a95e0041d95d9c3cc86eea841fa63cb1e8600", size = 355287, upload-time = "2026-05-28T11:59:57.448Z" }, - { url = "https://files.pythonhosted.org/packages/8d/58/5c4a43436843c90d0f6d19f82c200c80e3843ca9fa07b237623327f6d384/rpds_py-2026.5.1-cp313-cp313t-macosx_11_0_arm64.whl", hash = "sha256:cacedb7a6e167680acba45ad5716e89067d225dc80da0d7040cae8c81d4572fa", size = 347033, upload-time = "2026-05-28T11:59:58.881Z" }, - { url = "https://files.pythonhosted.org/packages/fb/c2/1a71acdacaf4e259b10278fb87b039ded3cf80041bcd89dd8a3ea702ded6/rpds_py-2026.5.1-cp313-cp313t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:68700371c5d7ae1412862ddfa719090925c93ecf351c566d66f09d04b136ea00", size = 376891, upload-time = "2026-05-28T12:00:00.516Z" }, - { url = "https://files.pythonhosted.org/packages/c2/c8/535f3d9b65addd8e28aa87b83c6e526799c3717a88273db8ea795beeef7a/rpds_py-2026.5.1-cp313-cp313t-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:296c799becfa849c779c8725494fe9ed94959ed886787df4364b058465bad7f0", size = 385646, upload-time = "2026-05-28T12:00:02.394Z" }, - { url = "https://files.pythonhosted.org/packages/1c/91/dc033f313345c354ade914dbe73cdb90b615a4409ea02430d5356794f3d8/rpds_py-2026.5.1-cp313-cp313t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:d3858b908218ee108d0bbfb2095ccc237648053c9bf98affad7cb079acaf1d97", size = 498830, upload-time = "2026-05-28T12:00:04.189Z" }, - { url = "https://files.pythonhosted.org/packages/27/fc/90fcbea459dbb8ddc18a2e0fd1de9412b48bc84ffff2db771cf714bacfd6/rpds_py-2026.5.1-cp313-cp313t-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:4fb8d2e7cb2f850b169806d61d1b991738acec96500a75c30f49caf064ce7cef", size = 392830, upload-time = "2026-05-28T12:00:05.797Z" }, - { url = "https://files.pythonhosted.org/packages/b2/1d/46cd11a228c9750684a798d98f878be6f614aa762438da7378f035e79e35/rpds_py-2026.5.1-cp313-cp313t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:27b74c10ed6a8f190f4287f53bcfea348b92a84a9c9f70d30183d1e6172d580d", size = 379613, upload-time = "2026-05-28T12:00:07.433Z" }, - { url = "https://files.pythonhosted.org/packages/24/4a/d9b0c6af3a1de03eb93741bbe8be2bdce84d8fda8224f3005451d86df389/rpds_py-2026.5.1-cp313-cp313t-manylinux_2_31_riscv64.whl", hash = "sha256:b9a6528956191c48c52294a592dbd4a8386d7048bdb25c0efcb6b966466c6d83", size = 388183, upload-time = "2026-05-28T12:00:09.227Z" }, - { url = "https://files.pythonhosted.org/packages/c5/b4/db7aaabdda6d020afc87d981bcc2f57a434c7dec60ecfc2ab3dd50b20351/rpds_py-2026.5.1-cp313-cp313t-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:af03e34e860047bc7a352b842856fcf78798fbb81132cc98bd2f907ab4eb9cd2", size = 408578, upload-time = "2026-05-28T12:00:10.779Z" }, - { url = "https://files.pythonhosted.org/packages/08/d6/070f6a41cbb343e2ac4171859bf3f3623e0ab002f72619d6d505313ec2de/rpds_py-2026.5.1-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:fea6e836d10abbe191d557d33bd58bd5987725fe63aa1eefe557d230209855bd", size = 553573, upload-time = "2026-05-28T12:00:12.443Z" }, - { url = "https://files.pythonhosted.org/packages/75/ab/1a71ea3589c4345dac0a0518f0e6a031cb42689277851b683c46d27463a5/rpds_py-2026.5.1-cp313-cp313t-musllinux_1_2_i686.whl", hash = "sha256:fc0c0f878ea770a0a8a462456c5ad36fc9fe6358e6b76fdadc7f17575e0b8bf1", size = 620861, upload-time = "2026-05-28T12:00:14.09Z" }, - { url = "https://files.pythonhosted.org/packages/8a/22/9bf80a56069c0c443fcfefac639a86a744550a2898817a6dfd3e26654924/rpds_py-2026.5.1-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:e0b360f316d966b048b085857630b3cc51f3db2f07b06f440eac8f695374d1e3", size = 585633, upload-time = "2026-05-28T12:00:15.66Z" }, - { url = "https://files.pythonhosted.org/packages/da/68/3b2c0a75c9e04125696f84ebdbbf304acf5a40b58ba4481cdb98a922c3ba/rpds_py-2026.5.1-cp313-cp313t-win32.whl", hash = "sha256:a2999883eedf72fdfb7520b92c7d4ec2572a71ff40239377aa604cc529eecafc", size = 210074, upload-time = "2026-05-28T12:00:17.291Z" }, - { url = "https://files.pythonhosted.org/packages/e7/8b/609157d5a25d37d4f29f92840ba531f416907c34ae5c5739dd21fc2bef98/rpds_py-2026.5.1-cp313-cp313t-win_amd64.whl", hash = "sha256:e07be2a9d7122bd6e82dea89814ef8dc893feb1aae97fec1630f3263bbb30e55", size = 228635, upload-time = "2026-05-28T12:00:18.73Z" }, - { url = "https://files.pythonhosted.org/packages/d4/6f/19c1918a4b590d8de87e712e4abe4b3875771eff60216fb6153cf6665c68/rpds_py-2026.5.1-cp314-cp314-macosx_10_12_x86_64.whl", hash = "sha256:1f2c391c3059798093b65df23aca2cac150460ae9c630d99dec83d703d9485b9", size = 349756, upload-time = "2026-05-28T12:00:20.217Z" }, - { url = "https://files.pythonhosted.org/packages/e5/60/a06fe7da34eca79dacbf958a2ba0c6eea85bc2b29de20080bf40f72f66fa/rpds_py-2026.5.1-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:413b424f7c4ee65ab5e5be91f5731be0f8b41a1ee2b12dfe810d716312e95a78", size = 343831, upload-time = "2026-05-28T12:00:21.711Z" }, - { url = "https://files.pythonhosted.org/packages/bf/ec/b2333b97b90e2a6ef6ca8ad386ee284968e74bcfe113b3f1a8d9036429a9/rpds_py-2026.5.1-cp314-cp314-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:2c595a1d9255dce0599e13130d1440ab2506654f2b50294226ee06402f8fef63", size = 375127, upload-time = "2026-05-28T12:00:23.326Z" }, - { url = "https://files.pythonhosted.org/packages/14/7f/e00aae54067f2b488c4637961d5f58204d470795fc791085fa3f15060d2e/rpds_py-2026.5.1-cp314-cp314-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:1c27c5f6102eac8c03e7595a00827a53b271ba40a53b59ff8709170e0855ea4a", size = 379034, upload-time = "2026-05-28T12:00:24.89Z" }, - { url = "https://files.pythonhosted.org/packages/be/cc/423999bbb8ae8dc93c77fc1d5e984ade5eb89d237d3bb884ccfa72ae2890/rpds_py-2026.5.1-cp314-cp314-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:6c7fcf61d44cacecaf3aea542b0e053db77972a4573e7ceda16fb2b399161195", size = 490823, upload-time = "2026-05-28T12:00:26.676Z" }, - { url = "https://files.pythonhosted.org/packages/0f/aa/c671bf660f12e68d3c52ff86c7066ed1372df5a0f4f2ff584e419b8207e7/rpds_py-2026.5.1-cp314-cp314-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:2c817a189d4ee14290420e5ff051e4dd6baa13f3edf84685071dee07a6d538ee", size = 388144, upload-time = "2026-05-28T12:00:28.577Z" }, - { url = "https://files.pythonhosted.org/packages/19/c8/d63bb75b68afe77b229e3021c6031bcaf01da5db5b0e69d0d10f9ba679a7/rpds_py-2026.5.1-cp314-cp314-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:21846aac0ed2e0589f38c12dc44e77bb64e494b771eadbcf169cba00566ba7ba", size = 371959, upload-time = "2026-05-28T12:00:30.304Z" }, - { url = "https://files.pythonhosted.org/packages/82/35/c51122014d8274ff37dc606d60049c3db7d83da02b5b282511e5a906a9a6/rpds_py-2026.5.1-cp314-cp314-manylinux_2_31_riscv64.whl", hash = "sha256:b317c87a13f769a4e787819bd508aaa5d69aa09b0880de9af6d3a8a54571cdec", size = 383558, upload-time = "2026-05-28T12:00:31.764Z" }, - { url = "https://files.pythonhosted.org/packages/e3/f9/2790cb99c136a5363acdeacf5c27c56f3de0d4118a1f48fca83404c99c89/rpds_py-2026.5.1-cp314-cp314-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:ce87129d9f2c14fa6c4a8601fb80eb4488c80d38a20cd13758ef11123e14995d", size = 402789, upload-time = "2026-05-28T12:00:33.247Z" }, - { url = "https://files.pythonhosted.org/packages/e5/1b/e4fb584f8c75d35c38150ff6a332cda949e6f97acba1f4fd123b14ab56fe/rpds_py-2026.5.1-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:9cdddb6c1207d284d94fd1530adf57fbd797fe7c4b8704ba85f49414f2557e7d", size = 551405, upload-time = "2026-05-28T12:00:34.819Z" }, - { url = "https://files.pythonhosted.org/packages/d8/f7/a6731b4216cb3793ea1af5391da240f5683dacc0d13e034fe5fc3503f240/rpds_py-2026.5.1-cp314-cp314-musllinux_1_2_i686.whl", hash = "sha256:4e237e139f94d3c036fd28eb9f564c99055476ff4ff05cd42be55ce349b5aa02", size = 616975, upload-time = "2026-05-28T12:00:36.268Z" }, - { url = "https://files.pythonhosted.org/packages/2c/ea/2e051a81d95d8e63f4b35a1c463a87e8766bc3d083c067c5dfb6bf220747/rpds_py-2026.5.1-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:ed0954b524873214369184a9c82b0eaa45a3fbb9a798cd95b17e0d98499e7ea0", size = 578701, upload-time = "2026-05-28T12:00:37.82Z" }, - { url = "https://files.pythonhosted.org/packages/65/56/b5f6fdb2083e32bca8a8993d89e70db114b4756c9e2c38421328126689d2/rpds_py-2026.5.1-cp314-cp314-win32.whl", hash = "sha256:2d88621d6a7d4dfa633d21abe90f280bb205274e16b1d1e61c6ad4640b2453b7", size = 209806, upload-time = "2026-05-28T12:00:39.492Z" }, - { url = "https://files.pythonhosted.org/packages/fb/80/65a5aa96c155e611d1ed844e4e1f57f3e36b021f396d9f8585d756e6b90d/rpds_py-2026.5.1-cp314-cp314-win_amd64.whl", hash = "sha256:cef8ac28d26f4dda3533060c20fbf80a325458fa9fd23ea72a73cdfa8e978838", size = 225985, upload-time = "2026-05-28T12:00:40.94Z" }, - { url = "https://files.pythonhosted.org/packages/27/7c/ad185212e87b05f196daef92bc5f3caf07298eb47c295b5585c3dd3093ac/rpds_py-2026.5.1-cp314-cp314-win_arm64.whl", hash = "sha256:eaaea962c68cdc68d4a533ba985ab8e9484277910bbfaa2ab3ef7732667bfed8", size = 221219, upload-time = "2026-05-28T12:00:43.15Z" }, - { url = "https://files.pythonhosted.org/packages/23/58/e14ae18759020334646b031e708ab4158d653a938822bfb7b95ef2e93aa3/rpds_py-2026.5.1-cp314-cp314t-macosx_10_12_x86_64.whl", hash = "sha256:21942f52dbbd5f8758bf021213d28bd45c39e873e65e2407faf5f1846f5761ad", size = 352148, upload-time = "2026-05-28T12:00:44.638Z" }, - { url = "https://files.pythonhosted.org/packages/31/9b/5f4a1e2f960bca3ac5d052b139dd31eed97b259f9d909173821760d542e8/rpds_py-2026.5.1-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:f414556f6e3958300ff941e40c9f97e3dc9774ddd1b3434c475d73dd354bbed3", size = 345196, upload-time = "2026-05-28T12:00:46.14Z" }, - { url = "https://files.pythonhosted.org/packages/1a/71/1d9574d6a2fa20ab60eaa55c7467f5aa20cbc770f341a05f09c0876f59e2/rpds_py-2026.5.1-cp314-cp314t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:ef1013a8625c74043210190b246f5b1551e09757c1f356c6e4160ef96c5bc081", size = 374981, upload-time = "2026-05-28T12:00:47.531Z" }, - { url = "https://files.pythonhosted.org/packages/0c/9a/37e99f4915a80aa71670263c1267f7ae0af95f53a3f61e6c3bdc016d4515/rpds_py-2026.5.1-cp314-cp314t-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:cc68e231a77a5f0d774ae278a1f8e55c0456501820847c1e4efb3829f3441df6", size = 379961, upload-time = "2026-05-28T12:00:49.216Z" }, - { url = "https://files.pythonhosted.org/packages/a8/ff/6e73f74b89d2e0715e0fc86b7dde893f9a61ae2f9b256ff3bdfe41ac4e94/rpds_py-2026.5.1-cp314-cp314t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:9baffb505aff33acc69b422a19f77806680f3c8632227d79f48de8a810d1c2c5", size = 495965, upload-time = "2026-05-28T12:00:51.111Z" }, - { url = "https://files.pythonhosted.org/packages/ea/e0/425faba25f59d74d4638b267f7c7a80e8649d2ef4db10a19b0c4a71e6e6f/rpds_py-2026.5.1-cp314-cp314t-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:b8d2f912928d426e8cfa396f7f3f8d29a59e6689c86dcca3c420730c1096322b", size = 389526, upload-time = "2026-05-28T12:00:52.77Z" }, - { url = "https://files.pythonhosted.org/packages/c6/76/7a41960e3fddae47fab43a28684d5da981401dffd88253de0944148654cb/rpds_py-2026.5.1-cp314-cp314t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:90f628283be835db980c941767d41c9a27b5239e54ba0a9c1335247e82406964", size = 376190, upload-time = "2026-05-28T12:00:54.215Z" }, - { url = "https://files.pythonhosted.org/packages/27/60/5f38dc70824fc6951b51d35377e577a3a3a4c81a6769cc5a2de25ebe0ad1/rpds_py-2026.5.1-cp314-cp314t-manylinux_2_31_riscv64.whl", hash = "sha256:1ebb2f0ab7e16132995a72de805170e0203df0c3dd22e1ef1cd1fdd90bd7a131", size = 383921, upload-time = "2026-05-28T12:00:55.673Z" }, - { url = "https://files.pythonhosted.org/packages/60/1a/d60a38caa1505f4b9483c3fbbde12c94e1079154f4f401a6da96f7e77621/rpds_py-2026.5.1-cp314-cp314t-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:f3df3d16ded76f1f8c9cdebd0e1ea55fdf4c23b812de189814da7cf229c22a81", size = 404766, upload-time = "2026-05-28T12:00:57.518Z" }, - { url = "https://files.pythonhosted.org/packages/87/ff/602fd3f174d6425f0bce05ad0dfbec0e96b38d0f7d08a79af5aa20083885/rpds_py-2026.5.1-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:9af8905b8f854990e40d5206aa5ac58d9b0fe0b7f351ff2bb086c20f6c8c6a47", size = 551343, upload-time = "2026-05-28T12:00:58.978Z" }, - { url = "https://files.pythonhosted.org/packages/b8/c1/1be13327acdbead3eca1fde03b6a34dbb011f1e864e217f0d32cc1779a7f/rpds_py-2026.5.1-cp314-cp314t-musllinux_1_2_i686.whl", hash = "sha256:036a36a87fb1cd3b214d11c4b3c4f7d2ddad933625dca1c900b56a057c07740a", size = 618502, upload-time = "2026-05-28T12:01:00.656Z" }, - { url = "https://files.pythonhosted.org/packages/f3/d7/afb49b49d7f2be8b7ba1a9f0977fa5168003437b93086726f066544e8351/rpds_py-2026.5.1-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:62ae3853454fe9ef283a03c96c2d835d39e84b14643a9d62c82ef0fb87d702ca", size = 581916, upload-time = "2026-05-28T12:01:02.22Z" }, - { url = "https://files.pythonhosted.org/packages/25/d1/dbef8c1f8a10f07beb62b5f054e20099fd9924b3ec001b8f0b6ac7813a85/rpds_py-2026.5.1-cp314-cp314t-win32.whl", hash = "sha256:6c3d771a46ec18b12af06ce36243a9a80b07a5d0515236332d90863ca8bb326a", size = 207855, upload-time = "2026-05-28T12:01:03.821Z" }, - { url = "https://files.pythonhosted.org/packages/2a/72/bfa4e61ab8e7dc1c8adf397e05e6cbdd4239357bd72b248d3de662f23915/rpds_py-2026.5.1-cp314-cp314t-win_amd64.whl", hash = "sha256:c93c629be4636cf54337bd5f06c104d55e42ced54d681f6fe21ae510a65116f6", size = 225422, upload-time = "2026-05-28T12:01:05.194Z" }, - { url = "https://files.pythonhosted.org/packages/27/3a/7b5da92b640f67b6717ccafc83cdd06bfa7ff2395c3685c68922bb54d703/rpds_py-2026.5.1-cp315-cp315-macosx_10_12_x86_64.whl", hash = "sha256:3574b55c604b8f75dacb007136508bbc0db406e626301778096a133327e7f2fb", size = 349576, upload-time = "2026-05-28T12:01:06.722Z" }, - { url = "https://files.pythonhosted.org/packages/d7/8a/2aafd7ad355a1bd48ca76e2262b74b15e6432b5a1efe150efd4d779cd55d/rpds_py-2026.5.1-cp315-cp315-macosx_11_0_arm64.whl", hash = "sha256:94068eb3ae6d43f5a786b7db96a406a34e6d5c24489feef32fd6e8946ea7b291", size = 343640, upload-time = "2026-05-28T12:01:08.441Z" }, - { url = "https://files.pythonhosted.org/packages/f7/7d/6c9523c1abbe840a1b7fba3c516d48e1d3487cc80fea4366c4071cf56784/rpds_py-2026.5.1-cp315-cp315-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:f3a5b10e8ce894825f380a8f1b6444cf73c294dfea62afbb2d13e3a9e630cec1", size = 375322, upload-time = "2026-05-28T12:01:09.934Z" }, - { url = "https://files.pythonhosted.org/packages/5a/5d/0b7b03fb1dc509321f01de3149784ab773e34c8573022029af8076afcb9c/rpds_py-2026.5.1-cp315-cp315-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:fc09f82e63d4bcd58149572f857a431bae851dc747e313c3b5bdf7abb907fda8", size = 379066, upload-time = "2026-05-28T12:01:11.48Z" }, - { url = "https://files.pythonhosted.org/packages/d7/e2/8ef6012999ebf1cb1c22f876d9ce5e63d960fd4631d2af3202d3f480aa25/rpds_py-2026.5.1-cp315-cp315-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:e10464d17df3b582745c25cec695cb9558bca2cb6ddb631aee1787fc72c767b2", size = 494586, upload-time = "2026-05-28T12:01:13.051Z" }, - { url = "https://files.pythonhosted.org/packages/80/af/1eeb029bec67582c226b7809172207cd005073af4ebd906e65ff494f4983/rpds_py-2026.5.1-cp315-cp315-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:ba05adbf15d994c38ec0b7ab32e858e5110c21e9009a00a86545fd220f84e038", size = 388415, upload-time = "2026-05-28T12:01:14.631Z" }, - { url = "https://files.pythonhosted.org/packages/18/23/ffbe10711c4d766c1cab0557d6906c074f795814863c67b351355d29354a/rpds_py-2026.5.1-cp315-cp315-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:77c004fdc7b891967106f78ddfd7b076bfe6813c6139c6fff6aed3bcaa960b26", size = 372427, upload-time = "2026-05-28T12:01:16.153Z" }, - { url = "https://files.pythonhosted.org/packages/bd/3a/30ba4a6ad457e5b070c18d742a33fb77d8d922b565cc881f8a5313d63bfe/rpds_py-2026.5.1-cp315-cp315-manylinux_2_31_riscv64.whl", hash = "sha256:83bcf894486c9d78dd290d3c0124ff6dd8875d3025e2090a8ec49fcc37c55fdd", size = 383615, upload-time = "2026-05-28T12:01:17.809Z" }, - { url = "https://files.pythonhosted.org/packages/d3/69/62e242b53ce39c0814bd24e1a6e6eba6c92be716277745f317f9540a2e7b/rpds_py-2026.5.1-cp315-cp315-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:c3df104083952a0e0c6f10de33e440eabe98fb6317d23e1a58c68f6df08d01b9", size = 402786, upload-time = "2026-05-28T12:01:19.419Z" }, - { url = "https://files.pythonhosted.org/packages/38/c1/a770b9c186928a1ed0f7e6d7ae50e7f3950ed23e3f9e366dbc8e38cb55de/rpds_py-2026.5.1-cp315-cp315-musllinux_1_2_aarch64.whl", hash = "sha256:980450826cf22e133c57e0835070bdd0dd3f73b9b708c3ce223def2cb9469e14", size = 551583, upload-time = "2026-05-28T12:01:21.013Z" }, - { url = "https://files.pythonhosted.org/packages/21/7c/68e8579b95375b70d2a963103c42e705856cdb98569258bd807f4423891c/rpds_py-2026.5.1-cp315-cp315-musllinux_1_2_i686.whl", hash = "sha256:205dde846f24332ab0c1188699a043b8d165b79bb84529ce272c45048ff6be01", size = 616941, upload-time = "2026-05-28T12:01:22.548Z" }, - { url = "https://files.pythonhosted.org/packages/70/a1/a6135aed5730ff03ab957182259987ac11e55fb392a28dc6f0592048a280/rpds_py-2026.5.1-cp315-cp315-musllinux_1_2_x86_64.whl", hash = "sha256:3966b82dd563176396df030f3dd52a6e54cb69b718e95e78bd555ed3d1e0185d", size = 578349, upload-time = "2026-05-28T12:01:24.118Z" }, - { url = "https://files.pythonhosted.org/packages/09/6e/f24201a76a84e6c49d0bdfdfcb735210e21701e9b21c5bfc0ba497dd62f6/rpds_py-2026.5.1-cp315-cp315-win32.whl", hash = "sha256:7818f8d0a415be74d2be3590b0a1c1f463a642f4d0217e7d10602dceef5b79aa", size = 209922, upload-time = "2026-05-28T12:01:25.522Z" }, - { url = "https://files.pythonhosted.org/packages/9e/e4/966bc240bb0485fc265278f6de44d05834bf0b3618886e0b22e33d54c49a/rpds_py-2026.5.1-cp315-cp315-win_amd64.whl", hash = "sha256:b3cc20c0d800af78fd0fac68086e28c1856cec51ea528bb81ea851aa40d39325", size = 226003, upload-time = "2026-05-28T12:01:27.062Z" }, - { url = "https://files.pythonhosted.org/packages/5c/5c/a15a59269cd5e74472734516c73795c15eccfc841b3d4b0228c3f53f19d0/rpds_py-2026.5.1-cp315-cp315-win_arm64.whl", hash = "sha256:3609e9939a8a76cd904cf98a3f1f13b5dc7e150adeaee89e0ea09652ea213e16", size = 221245, upload-time = "2026-05-28T12:01:28.51Z" }, - { url = "https://files.pythonhosted.org/packages/e0/22/135ce03804e179a71ceb13be095deda4a279bc88f7a6b8fa161c5ad44e12/rpds_py-2026.5.1-cp315-cp315t-macosx_10_12_x86_64.whl", hash = "sha256:5d333a7127d4b307601ac37792bee01bb95c867cbfacf21b6375b804d6bbd723", size = 352015, upload-time = "2026-05-28T12:01:30.214Z" }, - { url = "https://files.pythonhosted.org/packages/3b/5f/f1f6d2652eb9d848f6eb369d8db83a2da6249bb49ad2c2a48f45d54538d3/rpds_py-2026.5.1-cp315-cp315t-macosx_11_0_arm64.whl", hash = "sha256:b5f077b44a4f7808520f66dae234988d867deb9aed9be5da057ce9ba831b2a41", size = 345016, upload-time = "2026-05-28T12:01:31.656Z" }, - { url = "https://files.pythonhosted.org/packages/88/66/b74182775691ea2290c99e52ac8d5db844e56fbec90ce421f107658c8314/rpds_py-2026.5.1-cp315-cp315t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:55d8f9b7b78c9538fc9e04e82ec0e888ff0c3cffcfad152c77e57cd09351a98a", size = 374775, upload-time = "2026-05-28T12:01:33.136Z" }, - { url = "https://files.pythonhosted.org/packages/ff/8f/15e5a61d9f0a43902d36561d4f07cae6ae9f4716be825159fd72717f33af/rpds_py-2026.5.1-cp315-cp315t-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:e3a8ae58895ac107ed934a6bf51e5846f95c53b9b940c2c6d310838fd5846358", size = 380270, upload-time = "2026-05-28T12:01:34.574Z" }, - { url = "https://files.pythonhosted.org/packages/02/c3/f859b12763a80540cdf2af0f15b19904cf756a71d7bdd3f82ff3e5b1bbf9/rpds_py-2026.5.1-cp315-cp315t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:0957cf3c2b8632ec7aaebffebea8005b353cc2a237b6e2ae3c2cac0820704cfb", size = 495285, upload-time = "2026-05-28T12:01:36.127Z" }, - { url = "https://files.pythonhosted.org/packages/1c/c7/ff27c2ac8411d30b03b1829fd88cae8dad1a4d0da48dd25e57c4038042e6/rpds_py-2026.5.1-cp315-cp315t-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:c396c1304de421050b3681ea70f371874b54d41b0151e96109758144c231e30b", size = 389581, upload-time = "2026-05-28T12:01:37.635Z" }, - { url = "https://files.pythonhosted.org/packages/6e/67/fe92ee32a6cc05c77228a2f8b1762e7124f386ec20ff83d0757b762d58d0/rpds_py-2026.5.1-cp315-cp315t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:aad1bff7f666b9598e573815affd666aac6a13a585dde336f843e33350c7fadc", size = 376041, upload-time = "2026-05-28T12:01:39.307Z" }, - { url = "https://files.pythonhosted.org/packages/f8/91/b4d6685c27aba55bd82f25b278be8237038117d05f9659a6213ad3408130/rpds_py-2026.5.1-cp315-cp315t-manylinux_2_31_riscv64.whl", hash = "sha256:656a042550878f12d45752452d47094b7cfe5ad1e9d7b87b5a22ad3ae5ff8015", size = 383946, upload-time = "2026-05-28T12:01:41.043Z" }, - { url = "https://files.pythonhosted.org/packages/bd/79/2c1d832a53c8e0f8e98fc970ec257b950fecd4f62be2ab7182b500a0cbc8/rpds_py-2026.5.1-cp315-cp315t-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:73c4bd4f70294737b5206a3e8e30ccadbf8a60301831c8ea23eec5dbeea1ecfa", size = 405526, upload-time = "2026-05-28T12:01:43.032Z" }, - { url = "https://files.pythonhosted.org/packages/78/c4/c98117b03c6a8581ab2c2dfccfe9a5ad82bd8128a3c28b46a6ad2d97c393/rpds_py-2026.5.1-cp315-cp315t-musllinux_1_2_aarch64.whl", hash = "sha256:43bca78665423cabae77146f2fe7ce55272b6c8d55d82cca83effd42c7e13972", size = 551165, upload-time = "2026-05-28T12:01:44.648Z" }, - { url = "https://files.pythonhosted.org/packages/3b/c1/bc479ca069200af730881b1bd525e3114b2b391a351509fcb1b772f28086/rpds_py-2026.5.1-cp315-cp315t-musllinux_1_2_i686.whl", hash = "sha256:42d0f20e85e549c870749d0e247f0c10d318a45b7e9676d575d2dcb04a1b2e66", size = 618778, upload-time = "2026-05-28T12:01:46.337Z" }, - { url = "https://files.pythonhosted.org/packages/77/65/38ab2f90df44c2febfb63cc10ced40763d9b4bc94d173e734528663fe7f5/rpds_py-2026.5.1-cp315-cp315t-musllinux_1_2_x86_64.whl", hash = "sha256:b1be5c35683684d5331b93600c210e8367c254683d8a6df6bd21bd2da3a334fb", size = 581839, upload-time = "2026-05-28T12:01:48.109Z" }, - { url = "https://files.pythonhosted.org/packages/15/2d/ce1f605fe036aadd460e5822e578c6c7ec3a860936cca37d6e0f299daa77/rpds_py-2026.5.1-cp315-cp315t-win32.whl", hash = "sha256:75808f6c38ce7749bb68cc2770161aae5045e6c6f6781a9782e74b93304399df", size = 207866, upload-time = "2026-05-28T12:01:49.648Z" }, - { url = "https://files.pythonhosted.org/packages/79/cb/966040123eb102371559746908ef2c9471f4d43e17ec9a645a2258dab64b/rpds_py-2026.5.1-cp315-cp315t-win_amd64.whl", hash = "sha256:90bd6630002a1c7f09e7843dd79f0d24f3d2897cc25a753480917865d14f15b3", size = 225441, upload-time = "2026-05-28T12:01:51.408Z" }, - { url = "https://files.pythonhosted.org/packages/42/56/3fe0fb34820ff667be791b3a3c22b85e8bcba54e9c832f47438c191fa7be/rpds_py-2026.5.1-pp311-pypy311_pp73-macosx_10_12_x86_64.whl", hash = "sha256:edf2765d84e42447f112ad877af8fe1db0089aaec5b28e88d6eab45e7fe99cea", size = 357151, upload-time = "2026-05-28T12:01:53.43Z" }, - { url = "https://files.pythonhosted.org/packages/8b/f2/3eb9ccdb9f143b8c9b003978898cb497f942a324c077401e6b8834238e63/rpds_py-2026.5.1-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:ad3773236e95f7f33991eb125224b7da66f206504d032a253a02da7e134519fb", size = 350195, upload-time = "2026-05-28T12:01:54.901Z" }, - { url = "https://files.pythonhosted.org/packages/a7/24/dbda232bc4f3ed732120692ab0d2c8402cb020516556d8bee622dcef2413/rpds_py-2026.5.1-pp311-pypy311_pp73-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:a04df86b3f0fade39ec8fd0e0aab089b1da9fbd2b48df778a57ef96f5e7d38df", size = 381850, upload-time = "2026-05-28T12:01:56.601Z" }, - { url = "https://files.pythonhosted.org/packages/40/30/32e769839a358f78810c234f160f2cc21d1e4e47e1c0e0e0d535be5a0219/rpds_py-2026.5.1-pp311-pypy311_pp73-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:6142dbd80c4df62a5d899f0d616d417f84e0bc8d32526c8e5589019d75d028a7", size = 387899, upload-time = "2026-05-28T12:01:58.212Z" }, - { url = "https://files.pythonhosted.org/packages/ab/86/ec84d243aadb3b34b71dd26a010d0930b2d284ff5fc9a69fec53810ee6fd/rpds_py-2026.5.1-pp311-pypy311_pp73-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:0b35217adefe87f2fe4db7e9766cabe84744bfe9616d9667be18988928c7f2dc", size = 501618, upload-time = "2026-05-28T12:01:59.888Z" }, - { url = "https://files.pythonhosted.org/packages/74/25/b60e52686bbff777a64f9e4f4d3dd57980dc846913777177a2c92e4937aa/rpds_py-2026.5.1-pp311-pypy311_pp73-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:b95d5e11fc712b752081183a55a244c03cd00570489edd7014d8899f8ceb8162", size = 394003, upload-time = "2026-05-28T12:02:01.482Z" }, - { url = "https://files.pythonhosted.org/packages/9b/c7/b3a6a588cc2219510ef3f42e207483a93950bedd1e3a0fd4015c95cff9e5/rpds_py-2026.5.1-pp311-pypy311_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:141c9498daf2ace9eda35d2b0e376f9ea8b058d84f2aef4f96fccfd449a2f251", size = 379778, upload-time = "2026-05-28T12:02:03.197Z" }, - { url = "https://files.pythonhosted.org/packages/31/00/c7dba3fc8a3da8cb3f6db1eb3386be4d79c2e97c6890d20eb9ac66ae8c43/rpds_py-2026.5.1-pp311-pypy311_pp73-manylinux_2_31_riscv64.whl", hash = "sha256:6f249f8b860a200ad35193af961183ebe9132710484e6f6ce0cf89fd83c63a9a", size = 392359, upload-time = "2026-05-28T12:02:04.817Z" }, - { url = "https://files.pythonhosted.org/packages/93/dd/472ba494c70753f93745992c99855bee0636daf74e6984e5e003f150316f/rpds_py-2026.5.1-pp311-pypy311_pp73-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:e4abbf391a70be864920858bf360f4fb380577c9a0f732438a1996726e2c195b", size = 412820, upload-time = "2026-05-28T12:02:06.401Z" }, - { url = "https://files.pythonhosted.org/packages/1d/6f/93831a3bfe789542ed0c1d0d74b78b440f055d6dc3ea4640eba2d95e6e23/rpds_py-2026.5.1-pp311-pypy311_pp73-musllinux_1_2_aarch64.whl", hash = "sha256:c74005a7bb87752acf351c93897ec63ad77a07a0da7ecad9c050e32e7286ba34", size = 557243, upload-time = "2026-05-28T12:02:08.013Z" }, - { url = "https://files.pythonhosted.org/packages/1f/ff/0b3d604614ffc77522c6b288fdbce68957eb583da1002aa65ba38ac0ee40/rpds_py-2026.5.1-pp311-pypy311_pp73-musllinux_1_2_i686.whl", hash = "sha256:8213afbe8a3a906fb9acb2014423fe3359ee783d0bf90995f70623a3217bfa6c", size = 623541, upload-time = "2026-05-28T12:02:09.661Z" }, - { url = "https://files.pythonhosted.org/packages/ea/ea/e7b0251441da9adfeaebcf29601d10f2a1455fcf0772fae9e7e19032bd96/rpds_py-2026.5.1-pp311-pypy311_pp73-musllinux_1_2_x86_64.whl", hash = "sha256:8c43a8a973270fd173bf48cdf80bbe66312421cba68d40845034f174f2389049", size = 586326, upload-time = "2026-05-28T12:02:11.47Z" }, -] - -[[package]] -name = "safetensors" -version = "0.8.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/45/06/f955dbbb1859e3bd23c8ac6141af5106e7ad5fedec4a3a6e3d60f94b7001/safetensors-0.8.0.tar.gz", hash = "sha256:fabaf3e0f18a6618d9b36560682562157f77c2b71fcffc7b432be2baed9d753d", size = 325846, upload-time = "2026-06-09T07:52:25.563Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/39/a0/f718cda65b05407d228f97602cf60dca269c979867aa5beb25410de26cd3/safetensors-0.8.0-cp310-abi3-macosx_10_12_x86_64.whl", hash = "sha256:c554f85858e05226d3c2828e32395e677434685d6d94594a41643361c5e837f0", size = 473568, upload-time = "2026-06-09T07:52:18.829Z" }, - { url = "https://files.pythonhosted.org/packages/f5/b1/fa7c600e7dceae12e9606c7578cbc9ff1e1ed55844883ee5c92205e86226/safetensors-0.8.0-cp310-abi3-macosx_11_0_arm64.whl", hash = "sha256:c80201d22cbf405b80647a60ada77bba06c8fba2da2743ba1e89cdcc39a81f25", size = 484562, upload-time = "2026-06-09T07:52:17.518Z" }, - { url = "https://files.pythonhosted.org/packages/09/7d/65a7de0af421317bb36a067241e4235fff194eed60b961ed6d3f59a3fc60/safetensors-0.8.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:7a46e5ff292c356d6991e60942ba7f79817682d3a2cef0702136448cb9c4d235", size = 502844, upload-time = "2026-06-09T07:52:07.624Z" }, - { url = "https://files.pythonhosted.org/packages/91/4f/3175c9d75634e0e0dda0082794193521035edd7c70a6f212bf33ca06ddf4/safetensors-0.8.0-cp310-abi3-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:4124502b78f03534117c848f87a39b8f31e577b15eff423bf8bfb95f2a8c30d0", size = 511823, upload-time = "2026-06-09T07:52:09.565Z" }, - { url = "https://files.pythonhosted.org/packages/20/87/846c289e7aa2299eff406335717cf43ce8777194ece8aad75772e0411615/safetensors-0.8.0-cp310-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:7bc0a787ba8a35be368ee3574edfa2b1ad389eebd0a72e482ae275490e3f6c98", size = 633461, upload-time = "2026-06-09T07:52:11.128Z" }, - { url = "https://files.pythonhosted.org/packages/76/22/8d64d9df2c45d5ded401df889d0ad90882804ca172d79ec4f0df8f727fe0/safetensors-0.8.0-cp310-abi3-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:040070828e36dc8e122178bbbd5830ff9e97920affb84cbe0f46442497bed358", size = 545148, upload-time = "2026-06-09T07:52:13.603Z" }, - { url = "https://files.pythonhosted.org/packages/28/50/f203ff3a3ddfe19308efc83c5a3a29ed02bf786732ec35e68bf9162f3365/safetensors-0.8.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:fd6f3f93c9a0a7cc2788ee63fb763353d4bd2e89b0751bc78fcf7dda00bea774", size = 516040, upload-time = "2026-06-09T07:52:16.29Z" }, - { url = "https://files.pythonhosted.org/packages/46/fb/cdaed17ceb2948784fd9c36b6fd3e951b608547cea81a48e8ee6f8cfdfcb/safetensors-0.8.0-cp310-abi3-manylinux_2_31_riscv64.whl", hash = "sha256:fcdd41ec4628fee5799f807c73c353629130fbd942aa23d83c623dd6c9d52d78", size = 513832, upload-time = "2026-06-09T07:52:12.37Z" }, - { url = "https://files.pythonhosted.org/packages/0d/49/1e15de264dcc3b77943d2d0c56a95809956883b1c2d6d585c792523f180b/safetensors-0.8.0-cp310-abi3-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:8e9f537aa183a38ace122d27303dcd986b26bd2a7591f9181d7f0c396f4677ca", size = 559930, upload-time = "2026-06-09T07:52:14.743Z" }, - { url = "https://files.pythonhosted.org/packages/2a/43/bf38443278eab4b1be1fce2931e2b012ad9cb7df52ada751d0aab8f7659a/safetensors-0.8.0-cp310-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:87eec7ffed2b809f05a398a8becb7d013f19f7837cd15d9748580d6cf30dbaf4", size = 678670, upload-time = "2026-06-09T07:52:20.032Z" }, - { url = "https://files.pythonhosted.org/packages/72/e3/68cd3fa5b48488e84add63e04cb12f3bc28ae4638c06d4508c6e88823d0e/safetensors-0.8.0-cp310-abi3-musllinux_1_2_armv7l.whl", hash = "sha256:4a95ae2b05d7726d751da4ebf626a2ca782b706e101bd894c95bc2450b1cffcc", size = 786679, upload-time = "2026-06-09T07:52:21.322Z" }, - { url = "https://files.pythonhosted.org/packages/29/4b/1c19c509d56e01f4fbb3d0a2e597450f6cc04d1d56cf52defb0a62dfd715/safetensors-0.8.0-cp310-abi3-musllinux_1_2_i686.whl", hash = "sha256:3ae091f16662658bdc019a4ff6cb4c085bb7d725eb5978b183ffd265863b6d2d", size = 765683, upload-time = "2026-06-09T07:52:22.594Z" }, - { url = "https://files.pythonhosted.org/packages/27/43/41c1621732edd934d868a00d1b891584c892a7b62a9aab82ea5a0a5623ee/safetensors-0.8.0-cp310-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:8e080062fcde23be189565e1c3305d16751a218ecf9412c8601e64204eb6f846", size = 722361, upload-time = "2026-06-09T07:52:23.924Z" }, - { url = "https://files.pythonhosted.org/packages/8e/3f/73ccf82579412b4a71c4ca673f10b5f1f888d7cf5af7fe24f27d30307be4/safetensors-0.8.0-cp310-abi3-win32.whl", hash = "sha256:2ddf52eac562eda224f99acfa7889d02968c1fd59a5b011ae7d8137c37e9c02d", size = 342401, upload-time = "2026-06-09T07:52:28.895Z" }, - { url = "https://files.pythonhosted.org/packages/1b/6d/3fba214c1e5e0f69991677ec3bc17023f0421776975e1de0c682dca475e2/safetensors-0.8.0-cp310-abi3-win_amd64.whl", hash = "sha256:096ec1a98435df7beb08853bb5aa9081a84f23d0adc67ed1a0a10550f608373f", size = 355540, upload-time = "2026-06-09T07:52:27.832Z" }, - { url = "https://files.pythonhosted.org/packages/8d/fc/7eedc3510d97878876e32774eebbeb61c43f148a96e915c84229a3e967aa/safetensors-0.8.0-cp310-abi3-win_arm64.whl", hash = "sha256:f7838e5135a406ad3e02efdcb8cf2e5397d368b0154537c4fec682dbc544d452", size = 340500, upload-time = "2026-06-09T07:52:26.745Z" }, -] - -[[package]] -name = "scikit-learn" -version = "1.7.2" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version < '3.11' and sys_platform == 'win32'", - "python_full_version < '3.11' and sys_platform != 'win32'", -] -dependencies = [ - { name = "joblib", marker = "python_full_version < '3.11'" }, - { name = "numpy", version = "2.2.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, - { name = "scipy", version = "1.15.3", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, - { name = "threadpoolctl", marker = "python_full_version < '3.11'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/98/c2/a7855e41c9d285dfe86dc50b250978105dce513d6e459ea66a6aeb0e1e0c/scikit_learn-1.7.2.tar.gz", hash = "sha256:20e9e49ecd130598f1ca38a1d85090e1a600147b9c02fa6f15d69cb53d968fda", size = 7193136, upload-time = "2025-09-09T08:21:29.075Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/ba/3e/daed796fd69cce768b8788401cc464ea90b306fb196ae1ffed0b98182859/scikit_learn-1.7.2-cp310-cp310-macosx_10_9_x86_64.whl", hash = "sha256:6b33579c10a3081d076ab403df4a4190da4f4432d443521674637677dc91e61f", size = 9336221, upload-time = "2025-09-09T08:20:19.328Z" }, - { url = "https://files.pythonhosted.org/packages/1c/ce/af9d99533b24c55ff4e18d9b7b4d9919bbc6cd8f22fe7a7be01519a347d5/scikit_learn-1.7.2-cp310-cp310-macosx_12_0_arm64.whl", hash = "sha256:36749fb62b3d961b1ce4fedf08fa57a1986cd409eff2d783bca5d4b9b5fce51c", size = 8653834, upload-time = "2025-09-09T08:20:22.073Z" }, - { url = "https://files.pythonhosted.org/packages/58/0e/8c2a03d518fb6bd0b6b0d4b114c63d5f1db01ff0f9925d8eb10960d01c01/scikit_learn-1.7.2-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:7a58814265dfc52b3295b1900cfb5701589d30a8bb026c7540f1e9d3499d5ec8", size = 9660938, upload-time = "2025-09-09T08:20:24.327Z" }, - { url = "https://files.pythonhosted.org/packages/2b/75/4311605069b5d220e7cf5adabb38535bd96f0079313cdbb04b291479b22a/scikit_learn-1.7.2-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:4a847fea807e278f821a0406ca01e387f97653e284ecbd9750e3ee7c90347f18", size = 9477818, upload-time = "2025-09-09T08:20:26.845Z" }, - { url = "https://files.pythonhosted.org/packages/7f/9b/87961813c34adbca21a6b3f6b2bea344c43b30217a6d24cc437c6147f3e8/scikit_learn-1.7.2-cp310-cp310-win_amd64.whl", hash = "sha256:ca250e6836d10e6f402436d6463d6c0e4d8e0234cfb6a9a47835bd392b852ce5", size = 8886969, upload-time = "2025-09-09T08:20:29.329Z" }, - { url = "https://files.pythonhosted.org/packages/43/83/564e141eef908a5863a54da8ca342a137f45a0bfb71d1d79704c9894c9d1/scikit_learn-1.7.2-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:c7509693451651cd7361d30ce4e86a1347493554f172b1c72a39300fa2aea79e", size = 9331967, upload-time = "2025-09-09T08:20:32.421Z" }, - { url = "https://files.pythonhosted.org/packages/18/d6/ba863a4171ac9d7314c4d3fc251f015704a2caeee41ced89f321c049ed83/scikit_learn-1.7.2-cp311-cp311-macosx_12_0_arm64.whl", hash = "sha256:0486c8f827c2e7b64837c731c8feff72c0bd2b998067a8a9cbc10643c31f0fe1", size = 8648645, upload-time = "2025-09-09T08:20:34.436Z" }, - { url = "https://files.pythonhosted.org/packages/ef/0e/97dbca66347b8cf0ea8b529e6bb9367e337ba2e8be0ef5c1a545232abfde/scikit_learn-1.7.2-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:89877e19a80c7b11a2891a27c21c4894fb18e2c2e077815bcade10d34287b20d", size = 9715424, upload-time = "2025-09-09T08:20:36.776Z" }, - { url = "https://files.pythonhosted.org/packages/f7/32/1f3b22e3207e1d2c883a7e09abb956362e7d1bd2f14458c7de258a26ac15/scikit_learn-1.7.2-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:8da8bf89d4d79aaec192d2bda62f9b56ae4e5b4ef93b6a56b5de4977e375c1f1", size = 9509234, upload-time = "2025-09-09T08:20:38.957Z" }, - { url = "https://files.pythonhosted.org/packages/9f/71/34ddbd21f1da67c7a768146968b4d0220ee6831e4bcbad3e03dd3eae88b6/scikit_learn-1.7.2-cp311-cp311-win_amd64.whl", hash = "sha256:9b7ed8d58725030568523e937c43e56bc01cadb478fc43c042a9aca1dacb3ba1", size = 8894244, upload-time = "2025-09-09T08:20:41.166Z" }, - { url = "https://files.pythonhosted.org/packages/a7/aa/3996e2196075689afb9fce0410ebdb4a09099d7964d061d7213700204409/scikit_learn-1.7.2-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:8d91a97fa2b706943822398ab943cde71858a50245e31bc71dba62aab1d60a96", size = 9259818, upload-time = "2025-09-09T08:20:43.19Z" }, - { url = "https://files.pythonhosted.org/packages/43/5d/779320063e88af9c4a7c2cf463ff11c21ac9c8bd730c4a294b0000b666c9/scikit_learn-1.7.2-cp312-cp312-macosx_12_0_arm64.whl", hash = "sha256:acbc0f5fd2edd3432a22c69bed78e837c70cf896cd7993d71d51ba6708507476", size = 8636997, upload-time = "2025-09-09T08:20:45.468Z" }, - { url = "https://files.pythonhosted.org/packages/5c/d0/0c577d9325b05594fdd33aa970bf53fb673f051a45496842caee13cfd7fe/scikit_learn-1.7.2-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:e5bf3d930aee75a65478df91ac1225ff89cd28e9ac7bd1196853a9229b6adb0b", size = 9478381, upload-time = "2025-09-09T08:20:47.982Z" }, - { url = "https://files.pythonhosted.org/packages/82/70/8bf44b933837ba8494ca0fc9a9ab60f1c13b062ad0197f60a56e2fc4c43e/scikit_learn-1.7.2-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:b4d6e9deed1a47aca9fe2f267ab8e8fe82ee20b4526b2c0cd9e135cea10feb44", size = 9300296, upload-time = "2025-09-09T08:20:50.366Z" }, - { url = "https://files.pythonhosted.org/packages/c6/99/ed35197a158f1fdc2fe7c3680e9c70d0128f662e1fee4ed495f4b5e13db0/scikit_learn-1.7.2-cp312-cp312-win_amd64.whl", hash = "sha256:6088aa475f0785e01bcf8529f55280a3d7d298679f50c0bb70a2364a82d0b290", size = 8731256, upload-time = "2025-09-09T08:20:52.627Z" }, - { url = "https://files.pythonhosted.org/packages/ae/93/a3038cb0293037fd335f77f31fe053b89c72f17b1c8908c576c29d953e84/scikit_learn-1.7.2-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:0b7dacaa05e5d76759fb071558a8b5130f4845166d88654a0f9bdf3eb57851b7", size = 9212382, upload-time = "2025-09-09T08:20:54.731Z" }, - { url = "https://files.pythonhosted.org/packages/40/dd/9a88879b0c1104259136146e4742026b52df8540c39fec21a6383f8292c7/scikit_learn-1.7.2-cp313-cp313-macosx_12_0_arm64.whl", hash = "sha256:abebbd61ad9e1deed54cca45caea8ad5f79e1b93173dece40bb8e0c658dbe6fe", size = 8592042, upload-time = "2025-09-09T08:20:57.313Z" }, - { url = "https://files.pythonhosted.org/packages/46/af/c5e286471b7d10871b811b72ae794ac5fe2989c0a2df07f0ec723030f5f5/scikit_learn-1.7.2-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:502c18e39849c0ea1a5d681af1dbcf15f6cce601aebb657aabbfe84133c1907f", size = 9434180, upload-time = "2025-09-09T08:20:59.671Z" }, - { url = "https://files.pythonhosted.org/packages/f1/fd/df59faa53312d585023b2da27e866524ffb8faf87a68516c23896c718320/scikit_learn-1.7.2-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:7a4c328a71785382fe3fe676a9ecf2c86189249beff90bf85e22bdb7efaf9ae0", size = 9283660, upload-time = "2025-09-09T08:21:01.71Z" }, - { url = "https://files.pythonhosted.org/packages/a7/c7/03000262759d7b6f38c836ff9d512f438a70d8a8ddae68ee80de72dcfb63/scikit_learn-1.7.2-cp313-cp313-win_amd64.whl", hash = "sha256:63a9afd6f7b229aad94618c01c252ce9e6fa97918c5ca19c9a17a087d819440c", size = 8702057, upload-time = "2025-09-09T08:21:04.234Z" }, - { url = "https://files.pythonhosted.org/packages/55/87/ef5eb1f267084532c8e4aef98a28b6ffe7425acbfd64b5e2f2e066bc29b3/scikit_learn-1.7.2-cp313-cp313t-macosx_10_13_x86_64.whl", hash = "sha256:9acb6c5e867447b4e1390930e3944a005e2cb115922e693c08a323421a6966e8", size = 9558731, upload-time = "2025-09-09T08:21:06.381Z" }, - { url = "https://files.pythonhosted.org/packages/93/f8/6c1e3fc14b10118068d7938878a9f3f4e6d7b74a8ddb1e5bed65159ccda8/scikit_learn-1.7.2-cp313-cp313t-macosx_12_0_arm64.whl", hash = "sha256:2a41e2a0ef45063e654152ec9d8bcfc39f7afce35b08902bfe290c2498a67a6a", size = 9038852, upload-time = "2025-09-09T08:21:08.628Z" }, - { url = "https://files.pythonhosted.org/packages/83/87/066cafc896ee540c34becf95d30375fe5cbe93c3b75a0ee9aa852cd60021/scikit_learn-1.7.2-cp313-cp313t-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:98335fb98509b73385b3ab2bd0639b1f610541d3988ee675c670371d6a87aa7c", size = 9527094, upload-time = "2025-09-09T08:21:11.486Z" }, - { url = "https://files.pythonhosted.org/packages/9c/2b/4903e1ccafa1f6453b1ab78413938c8800633988c838aa0be386cbb33072/scikit_learn-1.7.2-cp313-cp313t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:191e5550980d45449126e23ed1d5e9e24b2c68329ee1f691a3987476e115e09c", size = 9367436, upload-time = "2025-09-09T08:21:13.602Z" }, - { url = "https://files.pythonhosted.org/packages/b5/aa/8444be3cfb10451617ff9d177b3c190288f4563e6c50ff02728be67ad094/scikit_learn-1.7.2-cp313-cp313t-win_amd64.whl", hash = "sha256:57dc4deb1d3762c75d685507fbd0bc17160144b2f2ba4ccea5dc285ab0d0e973", size = 9275749, upload-time = "2025-09-09T08:21:15.96Z" }, - { url = "https://files.pythonhosted.org/packages/d9/82/dee5acf66837852e8e68df6d8d3a6cb22d3df997b733b032f513d95205b7/scikit_learn-1.7.2-cp314-cp314-macosx_10_13_x86_64.whl", hash = "sha256:fa8f63940e29c82d1e67a45d5297bdebbcb585f5a5a50c4914cc2e852ab77f33", size = 9208906, upload-time = "2025-09-09T08:21:18.557Z" }, - { url = "https://files.pythonhosted.org/packages/3c/30/9029e54e17b87cb7d50d51a5926429c683d5b4c1732f0507a6c3bed9bf65/scikit_learn-1.7.2-cp314-cp314-macosx_12_0_arm64.whl", hash = "sha256:f95dc55b7902b91331fa4e5845dd5bde0580c9cd9612b1b2791b7e80c3d32615", size = 8627836, upload-time = "2025-09-09T08:21:20.695Z" }, - { url = "https://files.pythonhosted.org/packages/60/18/4a52c635c71b536879f4b971c2cedf32c35ee78f48367885ed8025d1f7ee/scikit_learn-1.7.2-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:9656e4a53e54578ad10a434dc1f993330568cfee176dff07112b8785fb413106", size = 9426236, upload-time = "2025-09-09T08:21:22.645Z" }, - { url = "https://files.pythonhosted.org/packages/99/7e/290362f6ab582128c53445458a5befd471ed1ea37953d5bcf80604619250/scikit_learn-1.7.2-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:96dc05a854add0e50d3f47a1ef21a10a595016da5b007c7d9cd9d0bffd1fcc61", size = 9312593, upload-time = "2025-09-09T08:21:24.65Z" }, - { url = "https://files.pythonhosted.org/packages/8e/87/24f541b6d62b1794939ae6422f8023703bbf6900378b2b34e0b4384dfefd/scikit_learn-1.7.2-cp314-cp314-win_amd64.whl", hash = "sha256:bb24510ed3f9f61476181e4db51ce801e2ba37541def12dc9333b946fc7a9cf8", size = 8820007, upload-time = "2025-09-09T08:21:26.713Z" }, -] - -[[package]] -name = "scikit-learn" -version = "1.9.0" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version >= '3.14' and sys_platform == 'win32'", - "python_full_version >= '3.14' and sys_platform != 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform == 'win32'", - "python_full_version == '3.11.*' and sys_platform == 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform != 'win32'", - "python_full_version == '3.11.*' and sys_platform != 'win32'", -] -dependencies = [ - { name = "joblib", marker = "python_full_version >= '3.11'" }, - { name = "narwhals", marker = "python_full_version >= '3.11'" }, - { name = "numpy", version = "2.4.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.11'" }, - { name = "scipy", version = "1.17.1", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version == '3.11.*'" }, - { name = "scipy", version = "1.18.0", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.12'" }, - { name = "threadpoolctl", marker = "python_full_version >= '3.11'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/fa/6f/37092bdb25f712817231799fc5674d8e704066a8a70c1d2d40517e18b4ab/scikit_learn-1.9.0.tar.gz", hash = "sha256:8833266989d3a5110178a9fae30783675460724d0e1efb13b14901d2c660c557", size = 7750767, upload-time = "2026-06-02T11:54:32.706Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/f5/be/e844fd9586e66540a15b71924d17a6cbc1bb749e81ddd0a796bcdba4c055/scikit_learn-1.9.0-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:9db6f4d34e68c8899e4cab27fdf8eafe6ed21f2ba52ceb25ea250cd237f8e47b", size = 8789686, upload-time = "2026-06-02T11:53:05.439Z" }, - { url = "https://files.pythonhosted.org/packages/42/e2/ff880f62677a17d035817d543cb0fc8727d01eccbee81c5f7fc733a9d856/scikit_learn-1.9.0-cp311-cp311-macosx_12_0_arm64.whl", hash = "sha256:f401448645a3e7bc115aa3c094097865155b34bff1cba8101857d9104e99074c", size = 8256782, upload-time = "2026-06-02T11:53:08.904Z" }, - { url = "https://files.pythonhosted.org/packages/25/64/eb40435e1a508ab1b4e284ce43ae80f6a162e5be5e38ed5a6fab467a9ea4/scikit_learn-1.9.0-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:fd3a8ef0c758555a3b23c03adaa858af32f7736785ded50ad5991f59c4ed03fa", size = 8992419, upload-time = "2026-06-02T11:53:11.551Z" }, - { url = "https://files.pythonhosted.org/packages/8d/da/4810a28e473185429e45a57eebcc91fc991b33d889cc0676063e671db03d/scikit_learn-1.9.0-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:f7e254636164090da847715a27f8e5478feb98c40a9e0ee90cbd277de9e5ceb8", size = 9281411, upload-time = "2026-06-02T11:53:15.063Z" }, - { url = "https://files.pythonhosted.org/packages/3b/67/be3d369f40d8178ba3bd86635d132e08cb5329b023e4669d9426d84bc007/scikit_learn-1.9.0-cp311-cp311-win_amd64.whl", hash = "sha256:5dc1818c77575d149e25fce9ef82dd7b7263ae372f03494158668ad632a69759", size = 8272736, upload-time = "2026-06-02T11:53:18.108Z" }, - { url = "https://files.pythonhosted.org/packages/37/79/a733f02dc2118da7e77a134b34f39f40201a353311b011d20859d2db3556/scikit_learn-1.9.0-cp311-cp311-win_arm64.whl", hash = "sha256:366652351f092b219c248f1e72821e841960a63d8f358f1dcfd54dc1cbdbbc28", size = 7919564, upload-time = "2026-06-02T11:53:21.2Z" }, - { url = "https://files.pythonhosted.org/packages/ac/20/75f915ff375d6249e6550ac740fdbbd66159a068fd3af1400ff62036b07a/scikit_learn-1.9.0-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:2bd41b0d201bc81575531b96b713d3eb5e5f50fb0b82101ff0f92294fdc236ac", size = 8741122, upload-time = "2026-06-02T11:53:24.08Z" }, - { url = "https://files.pythonhosted.org/packages/cc/d5/2b5148f2279196775e1db2aeb85d14b70ac80e7e32b3b28e7ebeafb0901d/scikit_learn-1.9.0-cp312-cp312-macosx_12_0_arm64.whl", hash = "sha256:5be45aa4a42a68a533913a6ed736cf309de2226411c79ef8d609a5456f1939b1", size = 8261512, upload-time = "2026-06-02T11:53:27.183Z" }, - { url = "https://files.pythonhosted.org/packages/a0/ee/5adbc77656b71f9456a2f5a7a9fdb4bcf9207a6b962889f1c2f9323afa4e/scikit_learn-1.9.0-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5e50ed4da51974e86e940690e9a3d82e729b62b5a49f7c9bac534d515d39d86f", size = 8837603, upload-time = "2026-06-02T11:53:30.328Z" }, - { url = "https://files.pythonhosted.org/packages/6c/c2/63fdda36c56437eeb44aaf9493c8bcd62ce230ab1598924fc626ffbfa943/scikit_learn-1.9.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:056c92bb67ad4c28463c2f2653d9701449201e7e7a9e94e321be0f71c4fef2b8", size = 9132097, upload-time = "2026-06-02T11:53:33.456Z" }, - { url = "https://files.pythonhosted.org/packages/83/a4/c8e67227c680e2259c8864ae72ff48b06e16a6f51253a22167aa02a8aa4e/scikit_learn-1.9.0-cp312-cp312-win_amd64.whl", hash = "sha256:4306775fad04cc4b472a1b15af1ae9cede1540fbfcc17fbce3767cd8dc7ae283", size = 8211173, upload-time = "2026-06-02T11:53:36.602Z" }, - { url = "https://files.pythonhosted.org/packages/cf/fd/3c0863792e98e67e9184aa4029288a175935eb65443afcd30d4f143450cf/scikit_learn-1.9.0-cp312-cp312-win_arm64.whl", hash = "sha256:26e22435f63bcdcf396b574273f29f13dd531f5ea035801f5be10ba1540a4e60", size = 7867451, upload-time = "2026-06-02T11:53:39.075Z" }, - { url = "https://files.pythonhosted.org/packages/3c/01/cf3310626b6d48d3e9be69a1223f9180360b5e6edb045f50fade723ce494/scikit_learn-1.9.0-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:80746d63bd4b6eaca54d36fe5feaf4d28bb38dc6f9470f81c7cad7c40155f119", size = 8705188, upload-time = "2026-06-02T11:53:41.964Z" }, - { url = "https://files.pythonhosted.org/packages/3e/04/5acd7ae280c5f93b6ac5ef6cdec14eef4c8d1cd91d85b3292989c94d96b1/scikit_learn-1.9.0-cp313-cp313-macosx_12_0_arm64.whl", hash = "sha256:5b934c45c252844a91d69fda3a34cff5e7307e1db10d77cb10a3980312c74713", size = 8228299, upload-time = "2026-06-02T11:53:44.817Z" }, - { url = "https://files.pythonhosted.org/packages/0c/39/ffe829a5b8ecb40a518724a997794657fdc354ada5e8fe8e64d998c0bac9/scikit_learn-1.9.0-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:38c3dcb9a1ffb85505ec53d54c7b4aea0cff70050425a7760c2af661ac85df05", size = 8789690, upload-time = "2026-06-02T11:53:47.461Z" }, - { url = "https://files.pythonhosted.org/packages/1f/88/8dab5de10c638c083772a6be83a3d8106ced492f74a928c8693638e5bb50/scikit_learn-1.9.0-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:da76d09304a4706db7cc1e3ebaa3b6b98a67365cc11d2996c4f1e58ba47df714", size = 9087723, upload-time = "2026-06-02T11:53:50.702Z" }, - { url = "https://files.pythonhosted.org/packages/20/3f/7917ca72464038f6240ec70c29f94862d08a34a74291ae4d4ec5eb8186a0/scikit_learn-1.9.0-cp313-cp313-win_amd64.whl", hash = "sha256:5808d98f15c6bf6d9d96d2348c1997392a5888ce7097e664105f930c4bca1277", size = 8184330, upload-time = "2026-06-02T11:53:53.396Z" }, - { url = "https://files.pythonhosted.org/packages/78/c7/15739eb2f61fda3c54639e9942414e5a19ad8a8d1f5a3266afad7cb7df80/scikit_learn-1.9.0-cp313-cp313-win_arm64.whl", hash = "sha256:d77f54c017633791bc0225a43e2f8d03745fdcfe4880268fcc4df15f505dec2e", size = 7840653, upload-time = "2026-06-02T11:53:56.035Z" }, - { url = "https://files.pythonhosted.org/packages/f4/7d/c9a35cf59b20a86fec24d306f1547b78dec194b08d367ce2a3e4854169d9/scikit_learn-1.9.0-cp314-cp314-macosx_10_15_x86_64.whl", hash = "sha256:9656acd4e93f74e0b66c8a36c88830a99252dfa900044d36bc2212ae89a47162", size = 8713289, upload-time = "2026-06-02T11:53:58.788Z" }, - { url = "https://files.pythonhosted.org/packages/3c/a7/552a7821597c632b907f7bfe8f36f9f572777af8ef8a48353041cf8e091a/scikit_learn-1.9.0-cp314-cp314-macosx_12_0_arm64.whl", hash = "sha256:24360002ae845e7866522b0a5bbf690802e7bc388cac8663502e78aa98598aa2", size = 8245141, upload-time = "2026-06-02T11:54:01.694Z" }, - { url = "https://files.pythonhosted.org/packages/7d/79/f4a0c4fe9711154cddabf913471153af79056382ddc612cfe5ee0ff4b72e/scikit_learn-1.9.0-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5162ad10a418c8a282dde04c9aa06965de3e9a65f33c1440c0ae69bb1a09d913", size = 8847671, upload-time = "2026-06-02T11:54:04.448Z" }, - { url = "https://files.pythonhosted.org/packages/f0/af/4d72d9e475ac83719160c662619e4bf7b95c19507cd582e7d0167a3c3dae/scikit_learn-1.9.0-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:1fea2cc5677ab49d6f5bade978c866da44957b712d92e9635e8b4f723013c3cb", size = 9118104, upload-time = "2026-06-02T11:54:07.205Z" }, - { url = "https://files.pythonhosted.org/packages/a2/d5/6a58eea2cb9abbb9b3f2bb8b2cfb3243d1152d69f442d256c7af71304769/scikit_learn-1.9.0-cp314-cp314-win_amd64.whl", hash = "sha256:64fa347efc1c839c487433e40c5144d38c336e8a2b59c81aa8660373945c2673", size = 8290674, upload-time = "2026-06-02T11:54:10.087Z" }, - { url = "https://files.pythonhosted.org/packages/65/5b/d4c879cf358f1187141cf90ced473f087183489090244f50c124a2ee478b/scikit_learn-1.9.0-cp314-cp314-win_arm64.whl", hash = "sha256:1b944b6db288f6b926e3650026ddafb988929de95d11fc2cc5fa117773c9ba42", size = 7978807, upload-time = "2026-06-02T11:54:12.769Z" }, - { url = "https://files.pythonhosted.org/packages/8a/43/bfae3121ec67ae09150d453c442c7c1cc166e9aefe056e6ab3b7728a5cfc/scikit_learn-1.9.0-cp314-cp314t-macosx_10_15_x86_64.whl", hash = "sha256:4ccacf04ca5f4b492158a5f28afe0ace43f81b2571e4b9a66d34848b46128949", size = 9031941, upload-time = "2026-06-02T11:54:15.436Z" }, - { url = "https://files.pythonhosted.org/packages/75/b0/20a4546eb17f3b25d3c66df15810411c14ed5065bcfab50b53c96fb627b2/scikit_learn-1.9.0-cp314-cp314t-macosx_12_0_arm64.whl", hash = "sha256:ee1a8db2c18c08e34c7412d4b10be1cac214cd4ea7dc9715a6a327eb49a37c96", size = 8613528, upload-time = "2026-06-02T11:54:18.842Z" }, - { url = "https://files.pythonhosted.org/packages/18/3c/e440e039bb82cd19004edaaad00acbde0fb9b461083c3ecf37941c557312/scikit_learn-1.9.0-cp314-cp314t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:147e9329ef0e39f75d4cffa02b2aa48d827832684926cd5210d9a2cb5c57246b", size = 8855050, upload-time = "2026-06-02T11:54:21.699Z" }, - { url = "https://files.pythonhosted.org/packages/43/26/b341b8dab5998da6270a3a42c2152c578501354d36f944b5856757035ef8/scikit_learn-1.9.0-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:5bad8f8b9950321b54c965fdcbac6c6c55e79e16646b49977bcf3668d3870a1a", size = 9097190, upload-time = "2026-06-02T11:54:24.454Z" }, - { url = "https://files.pythonhosted.org/packages/fb/de/b650b4d69b84468cfa2e28a3ff7b8103743029e6446ce1a97fe060ef688c/scikit_learn-1.9.0-cp314-cp314t-win_amd64.whl", hash = "sha256:78fc56eafd4edb9575d2d8950d1dd152061abb573341a1cb7e099fc40f6c6666", size = 8963204, upload-time = "2026-06-02T11:54:27.428Z" }, - { url = "https://files.pythonhosted.org/packages/ee/f3/ff83d76d7418112e5a61326443cdda87be3545dd8d6599c95b2481a4419e/scikit_learn-1.9.0-cp314-cp314t-win_arm64.whl", hash = "sha256:051075bda8b7aab87b1906ab3d4740a1e1224a19d7b3781a576736edc94e76aa", size = 8222661, upload-time = "2026-06-02T11:54:30.192Z" }, -] - -[[package]] -name = "scipy" -version = "1.15.3" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version < '3.11' and sys_platform == 'win32'", - "python_full_version < '3.11' and sys_platform != 'win32'", -] -dependencies = [ - { name = "numpy", version = "2.2.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/0f/37/6964b830433e654ec7485e45a00fc9a27cf868d622838f6b6d9c5ec0d532/scipy-1.15.3.tar.gz", hash = "sha256:eae3cf522bc7df64b42cad3925c876e1b0b6c35c1337c93e12c0f366f55b0eaf", size = 59419214, upload-time = "2025-05-08T16:13:05.955Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/78/2f/4966032c5f8cc7e6a60f1b2e0ad686293b9474b65246b0c642e3ef3badd0/scipy-1.15.3-cp310-cp310-macosx_10_13_x86_64.whl", hash = "sha256:a345928c86d535060c9c2b25e71e87c39ab2f22fc96e9636bd74d1dbf9de448c", size = 38702770, upload-time = "2025-05-08T16:04:20.849Z" }, - { url = "https://files.pythonhosted.org/packages/a0/6e/0c3bf90fae0e910c274db43304ebe25a6b391327f3f10b5dcc638c090795/scipy-1.15.3-cp310-cp310-macosx_12_0_arm64.whl", hash = "sha256:ad3432cb0f9ed87477a8d97f03b763fd1d57709f1bbde3c9369b1dff5503b253", size = 30094511, upload-time = "2025-05-08T16:04:27.103Z" }, - { url = "https://files.pythonhosted.org/packages/ea/b1/4deb37252311c1acff7f101f6453f0440794f51b6eacb1aad4459a134081/scipy-1.15.3-cp310-cp310-macosx_14_0_arm64.whl", hash = "sha256:aef683a9ae6eb00728a542b796f52a5477b78252edede72b8327a886ab63293f", size = 22368151, upload-time = "2025-05-08T16:04:31.731Z" }, - { url = "https://files.pythonhosted.org/packages/38/7d/f457626e3cd3c29b3a49ca115a304cebb8cc6f31b04678f03b216899d3c6/scipy-1.15.3-cp310-cp310-macosx_14_0_x86_64.whl", hash = "sha256:1c832e1bd78dea67d5c16f786681b28dd695a8cb1fb90af2e27580d3d0967e92", size = 25121732, upload-time = "2025-05-08T16:04:36.596Z" }, - { url = "https://files.pythonhosted.org/packages/db/0a/92b1de4a7adc7a15dcf5bddc6e191f6f29ee663b30511ce20467ef9b82e4/scipy-1.15.3-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:263961f658ce2165bbd7b99fa5135195c3a12d9bef045345016b8b50c315cb82", size = 35547617, upload-time = "2025-05-08T16:04:43.546Z" }, - { url = "https://files.pythonhosted.org/packages/8e/6d/41991e503e51fc1134502694c5fa7a1671501a17ffa12716a4a9151af3df/scipy-1.15.3-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:9e2abc762b0811e09a0d3258abee2d98e0c703eee49464ce0069590846f31d40", size = 37662964, upload-time = "2025-05-08T16:04:49.431Z" }, - { url = "https://files.pythonhosted.org/packages/25/e1/3df8f83cb15f3500478c889be8fb18700813b95e9e087328230b98d547ff/scipy-1.15.3-cp310-cp310-musllinux_1_2_aarch64.whl", hash = "sha256:ed7284b21a7a0c8f1b6e5977ac05396c0d008b89e05498c8b7e8f4a1423bba0e", size = 37238749, upload-time = "2025-05-08T16:04:55.215Z" }, - { url = "https://files.pythonhosted.org/packages/93/3e/b3257cf446f2a3533ed7809757039016b74cd6f38271de91682aa844cfc5/scipy-1.15.3-cp310-cp310-musllinux_1_2_x86_64.whl", hash = "sha256:5380741e53df2c566f4d234b100a484b420af85deb39ea35a1cc1be84ff53a5c", size = 40022383, upload-time = "2025-05-08T16:05:01.914Z" }, - { url = "https://files.pythonhosted.org/packages/d1/84/55bc4881973d3f79b479a5a2e2df61c8c9a04fcb986a213ac9c02cfb659b/scipy-1.15.3-cp310-cp310-win_amd64.whl", hash = "sha256:9d61e97b186a57350f6d6fd72640f9e99d5a4a2b8fbf4b9ee9a841eab327dc13", size = 41259201, upload-time = "2025-05-08T16:05:08.166Z" }, - { url = "https://files.pythonhosted.org/packages/96/ab/5cc9f80f28f6a7dff646c5756e559823614a42b1939d86dd0ed550470210/scipy-1.15.3-cp311-cp311-macosx_10_13_x86_64.whl", hash = "sha256:993439ce220d25e3696d1b23b233dd010169b62f6456488567e830654ee37a6b", size = 38714255, upload-time = "2025-05-08T16:05:14.596Z" }, - { url = "https://files.pythonhosted.org/packages/4a/4a/66ba30abe5ad1a3ad15bfb0b59d22174012e8056ff448cb1644deccbfed2/scipy-1.15.3-cp311-cp311-macosx_12_0_arm64.whl", hash = "sha256:34716e281f181a02341ddeaad584205bd2fd3c242063bd3423d61ac259ca7eba", size = 30111035, upload-time = "2025-05-08T16:05:20.152Z" }, - { url = "https://files.pythonhosted.org/packages/4b/fa/a7e5b95afd80d24313307f03624acc65801846fa75599034f8ceb9e2cbf6/scipy-1.15.3-cp311-cp311-macosx_14_0_arm64.whl", hash = "sha256:3b0334816afb8b91dab859281b1b9786934392aa3d527cd847e41bb6f45bee65", size = 22384499, upload-time = "2025-05-08T16:05:24.494Z" }, - { url = "https://files.pythonhosted.org/packages/17/99/f3aaddccf3588bb4aea70ba35328c204cadd89517a1612ecfda5b2dd9d7a/scipy-1.15.3-cp311-cp311-macosx_14_0_x86_64.whl", hash = "sha256:6db907c7368e3092e24919b5e31c76998b0ce1684d51a90943cb0ed1b4ffd6c1", size = 25152602, upload-time = "2025-05-08T16:05:29.313Z" }, - { url = "https://files.pythonhosted.org/packages/56/c5/1032cdb565f146109212153339f9cb8b993701e9fe56b1c97699eee12586/scipy-1.15.3-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:721d6b4ef5dc82ca8968c25b111e307083d7ca9091bc38163fb89243e85e3889", size = 35503415, upload-time = "2025-05-08T16:05:34.699Z" }, - { url = "https://files.pythonhosted.org/packages/bd/37/89f19c8c05505d0601ed5650156e50eb881ae3918786c8fd7262b4ee66d3/scipy-1.15.3-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:39cb9c62e471b1bb3750066ecc3a3f3052b37751c7c3dfd0fd7e48900ed52982", size = 37652622, upload-time = "2025-05-08T16:05:40.762Z" }, - { url = "https://files.pythonhosted.org/packages/7e/31/be59513aa9695519b18e1851bb9e487de66f2d31f835201f1b42f5d4d475/scipy-1.15.3-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:795c46999bae845966368a3c013e0e00947932d68e235702b5c3f6ea799aa8c9", size = 37244796, upload-time = "2025-05-08T16:05:48.119Z" }, - { url = "https://files.pythonhosted.org/packages/10/c0/4f5f3eeccc235632aab79b27a74a9130c6c35df358129f7ac8b29f562ac7/scipy-1.15.3-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:18aaacb735ab38b38db42cb01f6b92a2d0d4b6aabefeb07f02849e47f8fb3594", size = 40047684, upload-time = "2025-05-08T16:05:54.22Z" }, - { url = "https://files.pythonhosted.org/packages/ab/a7/0ddaf514ce8a8714f6ed243a2b391b41dbb65251affe21ee3077ec45ea9a/scipy-1.15.3-cp311-cp311-win_amd64.whl", hash = "sha256:ae48a786a28412d744c62fd7816a4118ef97e5be0bee968ce8f0a2fba7acf3bb", size = 41246504, upload-time = "2025-05-08T16:06:00.437Z" }, - { url = "https://files.pythonhosted.org/packages/37/4b/683aa044c4162e10ed7a7ea30527f2cbd92e6999c10a8ed8edb253836e9c/scipy-1.15.3-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:6ac6310fdbfb7aa6612408bd2f07295bcbd3fda00d2d702178434751fe48e019", size = 38766735, upload-time = "2025-05-08T16:06:06.471Z" }, - { url = "https://files.pythonhosted.org/packages/7b/7e/f30be3d03de07f25dc0ec926d1681fed5c732d759ac8f51079708c79e680/scipy-1.15.3-cp312-cp312-macosx_12_0_arm64.whl", hash = "sha256:185cd3d6d05ca4b44a8f1595af87f9c372bb6acf9c808e99aa3e9aa03bd98cf6", size = 30173284, upload-time = "2025-05-08T16:06:11.686Z" }, - { url = "https://files.pythonhosted.org/packages/07/9c/0ddb0d0abdabe0d181c1793db51f02cd59e4901da6f9f7848e1f96759f0d/scipy-1.15.3-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:05dc6abcd105e1a29f95eada46d4a3f251743cfd7d3ae8ddb4088047f24ea477", size = 22446958, upload-time = "2025-05-08T16:06:15.97Z" }, - { url = "https://files.pythonhosted.org/packages/af/43/0bce905a965f36c58ff80d8bea33f1f9351b05fad4beaad4eae34699b7a1/scipy-1.15.3-cp312-cp312-macosx_14_0_x86_64.whl", hash = "sha256:06efcba926324df1696931a57a176c80848ccd67ce6ad020c810736bfd58eb1c", size = 25242454, upload-time = "2025-05-08T16:06:20.394Z" }, - { url = "https://files.pythonhosted.org/packages/56/30/a6f08f84ee5b7b28b4c597aca4cbe545535c39fe911845a96414700b64ba/scipy-1.15.3-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:c05045d8b9bfd807ee1b9f38761993297b10b245f012b11b13b91ba8945f7e45", size = 35210199, upload-time = "2025-05-08T16:06:26.159Z" }, - { url = "https://files.pythonhosted.org/packages/0b/1f/03f52c282437a168ee2c7c14a1a0d0781a9a4a8962d84ac05c06b4c5b555/scipy-1.15.3-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:271e3713e645149ea5ea3e97b57fdab61ce61333f97cfae392c28ba786f9bb49", size = 37309455, upload-time = "2025-05-08T16:06:32.778Z" }, - { url = "https://files.pythonhosted.org/packages/89/b1/fbb53137f42c4bf630b1ffdfc2151a62d1d1b903b249f030d2b1c0280af8/scipy-1.15.3-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:6cfd56fc1a8e53f6e89ba3a7a7251f7396412d655bca2aa5611c8ec9a6784a1e", size = 36885140, upload-time = "2025-05-08T16:06:39.249Z" }, - { url = "https://files.pythonhosted.org/packages/2e/2e/025e39e339f5090df1ff266d021892694dbb7e63568edcfe43f892fa381d/scipy-1.15.3-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:0ff17c0bb1cb32952c09217d8d1eed9b53d1463e5f1dd6052c7857f83127d539", size = 39710549, upload-time = "2025-05-08T16:06:45.729Z" }, - { url = "https://files.pythonhosted.org/packages/e6/eb/3bf6ea8ab7f1503dca3a10df2e4b9c3f6b3316df07f6c0ded94b281c7101/scipy-1.15.3-cp312-cp312-win_amd64.whl", hash = "sha256:52092bc0472cfd17df49ff17e70624345efece4e1a12b23783a1ac59a1b728ed", size = 40966184, upload-time = "2025-05-08T16:06:52.623Z" }, - { url = "https://files.pythonhosted.org/packages/73/18/ec27848c9baae6e0d6573eda6e01a602e5649ee72c27c3a8aad673ebecfd/scipy-1.15.3-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:2c620736bcc334782e24d173c0fdbb7590a0a436d2fdf39310a8902505008759", size = 38728256, upload-time = "2025-05-08T16:06:58.696Z" }, - { url = "https://files.pythonhosted.org/packages/74/cd/1aef2184948728b4b6e21267d53b3339762c285a46a274ebb7863c9e4742/scipy-1.15.3-cp313-cp313-macosx_12_0_arm64.whl", hash = "sha256:7e11270a000969409d37ed399585ee530b9ef6aa99d50c019de4cb01e8e54e62", size = 30109540, upload-time = "2025-05-08T16:07:04.209Z" }, - { url = "https://files.pythonhosted.org/packages/5b/d8/59e452c0a255ec352bd0a833537a3bc1bfb679944c4938ab375b0a6b3a3e/scipy-1.15.3-cp313-cp313-macosx_14_0_arm64.whl", hash = "sha256:8c9ed3ba2c8a2ce098163a9bdb26f891746d02136995df25227a20e71c396ebb", size = 22383115, upload-time = "2025-05-08T16:07:08.998Z" }, - { url = "https://files.pythonhosted.org/packages/08/f5/456f56bbbfccf696263b47095291040655e3cbaf05d063bdc7c7517f32ac/scipy-1.15.3-cp313-cp313-macosx_14_0_x86_64.whl", hash = "sha256:0bdd905264c0c9cfa74a4772cdb2070171790381a5c4d312c973382fc6eaf730", size = 25163884, upload-time = "2025-05-08T16:07:14.091Z" }, - { url = "https://files.pythonhosted.org/packages/a2/66/a9618b6a435a0f0c0b8a6d0a2efb32d4ec5a85f023c2b79d39512040355b/scipy-1.15.3-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:79167bba085c31f38603e11a267d862957cbb3ce018d8b38f79ac043bc92d825", size = 35174018, upload-time = "2025-05-08T16:07:19.427Z" }, - { url = "https://files.pythonhosted.org/packages/b5/09/c5b6734a50ad4882432b6bb7c02baf757f5b2f256041da5df242e2d7e6b6/scipy-1.15.3-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:c9deabd6d547aee2c9a81dee6cc96c6d7e9a9b1953f74850c179f91fdc729cb7", size = 37269716, upload-time = "2025-05-08T16:07:25.712Z" }, - { url = "https://files.pythonhosted.org/packages/77/0a/eac00ff741f23bcabd352731ed9b8995a0a60ef57f5fd788d611d43d69a1/scipy-1.15.3-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:dde4fc32993071ac0c7dd2d82569e544f0bdaff66269cb475e0f369adad13f11", size = 36872342, upload-time = "2025-05-08T16:07:31.468Z" }, - { url = "https://files.pythonhosted.org/packages/fe/54/4379be86dd74b6ad81551689107360d9a3e18f24d20767a2d5b9253a3f0a/scipy-1.15.3-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:f77f853d584e72e874d87357ad70f44b437331507d1c311457bed8ed2b956126", size = 39670869, upload-time = "2025-05-08T16:07:38.002Z" }, - { url = "https://files.pythonhosted.org/packages/87/2e/892ad2862ba54f084ffe8cc4a22667eaf9c2bcec6d2bff1d15713c6c0703/scipy-1.15.3-cp313-cp313-win_amd64.whl", hash = "sha256:b90ab29d0c37ec9bf55424c064312930ca5f4bde15ee8619ee44e69319aab163", size = 40988851, upload-time = "2025-05-08T16:08:33.671Z" }, - { url = "https://files.pythonhosted.org/packages/1b/e9/7a879c137f7e55b30d75d90ce3eb468197646bc7b443ac036ae3fe109055/scipy-1.15.3-cp313-cp313t-macosx_10_13_x86_64.whl", hash = "sha256:3ac07623267feb3ae308487c260ac684b32ea35fd81e12845039952f558047b8", size = 38863011, upload-time = "2025-05-08T16:07:44.039Z" }, - { url = "https://files.pythonhosted.org/packages/51/d1/226a806bbd69f62ce5ef5f3ffadc35286e9fbc802f606a07eb83bf2359de/scipy-1.15.3-cp313-cp313t-macosx_12_0_arm64.whl", hash = "sha256:6487aa99c2a3d509a5227d9a5e889ff05830a06b2ce08ec30df6d79db5fcd5c5", size = 30266407, upload-time = "2025-05-08T16:07:49.891Z" }, - { url = "https://files.pythonhosted.org/packages/e5/9b/f32d1d6093ab9eeabbd839b0f7619c62e46cc4b7b6dbf05b6e615bbd4400/scipy-1.15.3-cp313-cp313t-macosx_14_0_arm64.whl", hash = "sha256:50f9e62461c95d933d5c5ef4a1f2ebf9a2b4e83b0db374cb3f1de104d935922e", size = 22540030, upload-time = "2025-05-08T16:07:54.121Z" }, - { url = "https://files.pythonhosted.org/packages/e7/29/c278f699b095c1a884f29fda126340fcc201461ee8bfea5c8bdb1c7c958b/scipy-1.15.3-cp313-cp313t-macosx_14_0_x86_64.whl", hash = "sha256:14ed70039d182f411ffc74789a16df3835e05dc469b898233a245cdfd7f162cb", size = 25218709, upload-time = "2025-05-08T16:07:58.506Z" }, - { url = "https://files.pythonhosted.org/packages/24/18/9e5374b617aba742a990581373cd6b68a2945d65cc588482749ef2e64467/scipy-1.15.3-cp313-cp313t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:0a769105537aa07a69468a0eefcd121be52006db61cdd8cac8a0e68980bbb723", size = 34809045, upload-time = "2025-05-08T16:08:03.929Z" }, - { url = "https://files.pythonhosted.org/packages/e1/fe/9c4361e7ba2927074360856db6135ef4904d505e9b3afbbcb073c4008328/scipy-1.15.3-cp313-cp313t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:9db984639887e3dffb3928d118145ffe40eff2fa40cb241a306ec57c219ebbbb", size = 36703062, upload-time = "2025-05-08T16:08:09.558Z" }, - { url = "https://files.pythonhosted.org/packages/b7/8e/038ccfe29d272b30086b25a4960f757f97122cb2ec42e62b460d02fe98e9/scipy-1.15.3-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:40e54d5c7e7ebf1aa596c374c49fa3135f04648a0caabcb66c52884b943f02b4", size = 36393132, upload-time = "2025-05-08T16:08:15.34Z" }, - { url = "https://files.pythonhosted.org/packages/10/7e/5c12285452970be5bdbe8352c619250b97ebf7917d7a9a9e96b8a8140f17/scipy-1.15.3-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:5e721fed53187e71d0ccf382b6bf977644c533e506c4d33c3fb24de89f5c3ed5", size = 38979503, upload-time = "2025-05-08T16:08:21.513Z" }, - { url = "https://files.pythonhosted.org/packages/81/06/0a5e5349474e1cbc5757975b21bd4fad0e72ebf138c5592f191646154e06/scipy-1.15.3-cp313-cp313t-win_amd64.whl", hash = "sha256:76ad1fb5f8752eabf0fa02e4cc0336b4e8f021e2d5f061ed37d6d264db35e3ca", size = 40308097, upload-time = "2025-05-08T16:08:27.627Z" }, -] - -[[package]] -name = "scipy" -version = "1.17.1" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version == '3.11.*' and sys_platform == 'win32'", - "python_full_version == '3.11.*' and sys_platform != 'win32'", -] -dependencies = [ - { name = "numpy", version = "2.4.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version == '3.11.*'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/7a/97/5a3609c4f8d58b039179648e62dd220f89864f56f7357f5d4f45c29eb2cc/scipy-1.17.1.tar.gz", hash = "sha256:95d8e012d8cb8816c226aef832200b1d45109ed4464303e997c5b13122b297c0", size = 30573822, upload-time = "2026-02-23T00:26:24.851Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/df/75/b4ce781849931fef6fd529afa6b63711d5a733065722d0c3e2724af9e40a/scipy-1.17.1-cp311-cp311-macosx_10_14_x86_64.whl", hash = "sha256:1f95b894f13729334fb990162e911c9e5dc1ab390c58aa6cbecb389c5b5e28ec", size = 31613675, upload-time = "2026-02-23T00:16:00.13Z" }, - { url = "https://files.pythonhosted.org/packages/f7/58/bccc2861b305abdd1b8663d6130c0b3d7cc22e8d86663edbc8401bfd40d4/scipy-1.17.1-cp311-cp311-macosx_12_0_arm64.whl", hash = "sha256:e18f12c6b0bc5a592ed23d3f7b891f68fd7f8241d69b7883769eb5d5dfb52696", size = 28162057, upload-time = "2026-02-23T00:16:09.456Z" }, - { url = "https://files.pythonhosted.org/packages/6d/ee/18146b7757ed4976276b9c9819108adbc73c5aad636e5353e20746b73069/scipy-1.17.1-cp311-cp311-macosx_14_0_arm64.whl", hash = "sha256:a3472cfbca0a54177d0faa68f697d8ba4c80bbdc19908c3465556d9f7efce9ee", size = 20334032, upload-time = "2026-02-23T00:16:17.358Z" }, - { url = "https://files.pythonhosted.org/packages/ec/e6/cef1cf3557f0c54954198554a10016b6a03b2ec9e22a4e1df734936bd99c/scipy-1.17.1-cp311-cp311-macosx_14_0_x86_64.whl", hash = "sha256:766e0dc5a616d026a3a1cffa379af959671729083882f50307e18175797b3dfd", size = 22709533, upload-time = "2026-02-23T00:16:25.791Z" }, - { url = "https://files.pythonhosted.org/packages/4d/60/8804678875fc59362b0fb759ab3ecce1f09c10a735680318ac30da8cd76b/scipy-1.17.1-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:744b2bf3640d907b79f3fd7874efe432d1cf171ee721243e350f55234b4cec4c", size = 33062057, upload-time = "2026-02-23T00:16:36.931Z" }, - { url = "https://files.pythonhosted.org/packages/09/7d/af933f0f6e0767995b4e2d705a0665e454d1c19402aa7e895de3951ebb04/scipy-1.17.1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:43af8d1f3bea642559019edfe64e9b11192a8978efbd1539d7bc2aaa23d92de4", size = 35349300, upload-time = "2026-02-23T00:16:49.108Z" }, - { url = "https://files.pythonhosted.org/packages/b4/3d/7ccbbdcbb54c8fdc20d3b6930137c782a163fa626f0aef920349873421ba/scipy-1.17.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:cd96a1898c0a47be4520327e01f874acfd61fb48a9420f8aa9f6483412ffa444", size = 35127333, upload-time = "2026-02-23T00:17:01.293Z" }, - { url = "https://files.pythonhosted.org/packages/e8/19/f926cb11c42b15ba08e3a71e376d816ac08614f769b4f47e06c3580c836a/scipy-1.17.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:4eb6c25dd62ee8d5edf68a8e1c171dd71c292fdae95d8aeb3dd7d7de4c364082", size = 37741314, upload-time = "2026-02-23T00:17:12.576Z" }, - { url = "https://files.pythonhosted.org/packages/95/da/0d1df507cf574b3f224ccc3d45244c9a1d732c81dcb26b1e8a766ae271a8/scipy-1.17.1-cp311-cp311-win_amd64.whl", hash = "sha256:d30e57c72013c2a4fe441c2fcb8e77b14e152ad48b5464858e07e2ad9fbfceff", size = 36607512, upload-time = "2026-02-23T00:17:23.424Z" }, - { url = "https://files.pythonhosted.org/packages/68/7f/bdd79ceaad24b671543ffe0ef61ed8e659440eb683b66f033454dcee90eb/scipy-1.17.1-cp311-cp311-win_arm64.whl", hash = "sha256:9ecb4efb1cd6e8c4afea0daa91a87fbddbce1b99d2895d151596716c0b2e859d", size = 24599248, upload-time = "2026-02-23T00:17:34.561Z" }, - { url = "https://files.pythonhosted.org/packages/35/48/b992b488d6f299dbe3f11a20b24d3dda3d46f1a635ede1c46b5b17a7b163/scipy-1.17.1-cp312-cp312-macosx_10_14_x86_64.whl", hash = "sha256:35c3a56d2ef83efc372eaec584314bd0ef2e2f0d2adb21c55e6ad5b344c0dcb8", size = 31610954, upload-time = "2026-02-23T00:17:49.855Z" }, - { url = "https://files.pythonhosted.org/packages/b2/02/cf107b01494c19dc100f1d0b7ac3cc08666e96ba2d64db7626066cee895e/scipy-1.17.1-cp312-cp312-macosx_12_0_arm64.whl", hash = "sha256:fcb310ddb270a06114bb64bbe53c94926b943f5b7f0842194d585c65eb4edd76", size = 28172662, upload-time = "2026-02-23T00:18:01.64Z" }, - { url = "https://files.pythonhosted.org/packages/cf/a9/599c28631bad314d219cf9ffd40e985b24d603fc8a2f4ccc5ae8419a535b/scipy-1.17.1-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:cc90d2e9c7e5c7f1a482c9875007c095c3194b1cfedca3c2f3291cdc2bc7c086", size = 20344366, upload-time = "2026-02-23T00:18:12.015Z" }, - { url = "https://files.pythonhosted.org/packages/35/f5/906eda513271c8deb5af284e5ef0206d17a96239af79f9fa0aebfe0e36b4/scipy-1.17.1-cp312-cp312-macosx_14_0_x86_64.whl", hash = "sha256:c80be5ede8f3f8eded4eff73cc99a25c388ce98e555b17d31da05287015ffa5b", size = 22704017, upload-time = "2026-02-23T00:18:21.502Z" }, - { url = "https://files.pythonhosted.org/packages/da/34/16f10e3042d2f1d6b66e0428308ab52224b6a23049cb2f5c1756f713815f/scipy-1.17.1-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:e19ebea31758fac5893a2ac360fedd00116cbb7628e650842a6691ba7ca28a21", size = 32927842, upload-time = "2026-02-23T00:18:35.367Z" }, - { url = "https://files.pythonhosted.org/packages/01/8e/1e35281b8ab6d5d72ebe9911edcdffa3f36b04ed9d51dec6dd140396e220/scipy-1.17.1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:02ae3b274fde71c5e92ac4d54bc06c42d80e399fec704383dcd99b301df37458", size = 35235890, upload-time = "2026-02-23T00:18:49.188Z" }, - { url = "https://files.pythonhosted.org/packages/c5/5c/9d7f4c88bea6e0d5a4f1bc0506a53a00e9fcb198de372bfe4d3652cef482/scipy-1.17.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:8a604bae87c6195d8b1045eddece0514d041604b14f2727bbc2b3020172045eb", size = 35003557, upload-time = "2026-02-23T00:18:54.74Z" }, - { url = "https://files.pythonhosted.org/packages/65/94/7698add8f276dbab7a9de9fb6b0e02fc13ee61d51c7c3f85ac28b65e1239/scipy-1.17.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:f590cd684941912d10becc07325a3eeb77886fe981415660d9265c4c418d0bea", size = 37625856, upload-time = "2026-02-23T00:19:00.307Z" }, - { url = "https://files.pythonhosted.org/packages/a2/84/dc08d77fbf3d87d3ee27f6a0c6dcce1de5829a64f2eae85a0ecc1f0daa73/scipy-1.17.1-cp312-cp312-win_amd64.whl", hash = "sha256:41b71f4a3a4cab9d366cd9065b288efc4d4f3c0b37a91a8e0947fb5bd7f31d87", size = 36549682, upload-time = "2026-02-23T00:19:07.67Z" }, - { url = "https://files.pythonhosted.org/packages/bc/98/fe9ae9ffb3b54b62559f52dedaebe204b408db8109a8c66fdd04869e6424/scipy-1.17.1-cp312-cp312-win_arm64.whl", hash = "sha256:f4115102802df98b2b0db3cce5cb9b92572633a1197c77b7553e5203f284a5b3", size = 24547340, upload-time = "2026-02-23T00:19:12.024Z" }, - { url = "https://files.pythonhosted.org/packages/76/27/07ee1b57b65e92645f219b37148a7e7928b82e2b5dbeccecb4dff7c64f0b/scipy-1.17.1-cp313-cp313-macosx_10_14_x86_64.whl", hash = "sha256:5e3c5c011904115f88a39308379c17f91546f77c1667cea98739fe0fccea804c", size = 31590199, upload-time = "2026-02-23T00:19:17.192Z" }, - { url = "https://files.pythonhosted.org/packages/ec/ae/db19f8ab842e9b724bf5dbb7db29302a91f1e55bc4d04b1025d6d605a2c5/scipy-1.17.1-cp313-cp313-macosx_12_0_arm64.whl", hash = "sha256:6fac755ca3d2c3edcb22f479fceaa241704111414831ddd3bc6056e18516892f", size = 28154001, upload-time = "2026-02-23T00:19:22.241Z" }, - { url = "https://files.pythonhosted.org/packages/5b/58/3ce96251560107b381cbd6e8413c483bbb1228a6b919fa8652b0d4090e7f/scipy-1.17.1-cp313-cp313-macosx_14_0_arm64.whl", hash = "sha256:7ff200bf9d24f2e4d5dc6ee8c3ac64d739d3a89e2326ba68aaf6c4a2b838fd7d", size = 20325719, upload-time = "2026-02-23T00:19:26.329Z" }, - { url = "https://files.pythonhosted.org/packages/b2/83/15087d945e0e4d48ce2377498abf5ad171ae013232ae31d06f336e64c999/scipy-1.17.1-cp313-cp313-macosx_14_0_x86_64.whl", hash = "sha256:4b400bdc6f79fa02a4d86640310dde87a21fba0c979efff5248908c6f15fad1b", size = 22683595, upload-time = "2026-02-23T00:19:30.304Z" }, - { url = "https://files.pythonhosted.org/packages/b4/e0/e58fbde4a1a594c8be8114eb4aac1a55bcd6587047efc18a61eb1f5c0d30/scipy-1.17.1-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:2b64ca7d4aee0102a97f3ba22124052b4bd2152522355073580bf4845e2550b6", size = 32896429, upload-time = "2026-02-23T00:19:35.536Z" }, - { url = "https://files.pythonhosted.org/packages/f5/5f/f17563f28ff03c7b6799c50d01d5d856a1d55f2676f537ca8d28c7f627cd/scipy-1.17.1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:581b2264fc0aa555f3f435a5944da7504ea3a065d7029ad60e7c3d1ae09c5464", size = 35203952, upload-time = "2026-02-23T00:19:42.259Z" }, - { url = "https://files.pythonhosted.org/packages/8d/a5/9afd17de24f657fdfe4df9a3f1ea049b39aef7c06000c13db1530d81ccca/scipy-1.17.1-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:beeda3d4ae615106d7094f7e7cef6218392e4465cc95d25f900bebabfded0950", size = 34979063, upload-time = "2026-02-23T00:19:47.547Z" }, - { url = "https://files.pythonhosted.org/packages/8b/13/88b1d2384b424bf7c924f2038c1c409f8d88bb2a8d49d097861dd64a57b2/scipy-1.17.1-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:6609bc224e9568f65064cfa72edc0f24ee6655b47575954ec6339534b2798369", size = 37598449, upload-time = "2026-02-23T00:19:53.238Z" }, - { url = "https://files.pythonhosted.org/packages/35/e5/d6d0e51fc888f692a35134336866341c08655d92614f492c6860dc45bb2c/scipy-1.17.1-cp313-cp313-win_amd64.whl", hash = "sha256:37425bc9175607b0268f493d79a292c39f9d001a357bebb6b88fdfaff13f6448", size = 36510943, upload-time = "2026-02-23T00:20:50.89Z" }, - { url = "https://files.pythonhosted.org/packages/2a/fd/3be73c564e2a01e690e19cc618811540ba5354c67c8680dce3281123fb79/scipy-1.17.1-cp313-cp313-win_arm64.whl", hash = "sha256:5cf36e801231b6a2059bf354720274b7558746f3b1a4efb43fcf557ccd484a87", size = 24545621, upload-time = "2026-02-23T00:20:55.871Z" }, - { url = "https://files.pythonhosted.org/packages/6f/6b/17787db8b8114933a66f9dcc479a8272e4b4da75fe03b0c282f7b0ade8cd/scipy-1.17.1-cp313-cp313t-macosx_10_14_x86_64.whl", hash = "sha256:d59c30000a16d8edc7e64152e30220bfbd724c9bbb08368c054e24c651314f0a", size = 31936708, upload-time = "2026-02-23T00:19:58.694Z" }, - { url = "https://files.pythonhosted.org/packages/38/2e/524405c2b6392765ab1e2b722a41d5da33dc5c7b7278184a8ad29b6cb206/scipy-1.17.1-cp313-cp313t-macosx_12_0_arm64.whl", hash = "sha256:010f4333c96c9bb1a4516269e33cb5917b08ef2166d5556ca2fd9f082a9e6ea0", size = 28570135, upload-time = "2026-02-23T00:20:03.934Z" }, - { url = "https://files.pythonhosted.org/packages/fd/c3/5bd7199f4ea8556c0c8e39f04ccb014ac37d1468e6cfa6a95c6b3562b76e/scipy-1.17.1-cp313-cp313t-macosx_14_0_arm64.whl", hash = "sha256:2ceb2d3e01c5f1d83c4189737a42d9cb2fc38a6eeed225e7515eef71ad301dce", size = 20741977, upload-time = "2026-02-23T00:20:07.935Z" }, - { url = "https://files.pythonhosted.org/packages/d9/b8/8ccd9b766ad14c78386599708eb745f6b44f08400a5fd0ade7cf89b6fc93/scipy-1.17.1-cp313-cp313t-macosx_14_0_x86_64.whl", hash = "sha256:844e165636711ef41f80b4103ed234181646b98a53c8f05da12ca5ca289134f6", size = 23029601, upload-time = "2026-02-23T00:20:12.161Z" }, - { url = "https://files.pythonhosted.org/packages/6d/a0/3cb6f4d2fb3e17428ad2880333cac878909ad1a89f678527b5328b93c1d4/scipy-1.17.1-cp313-cp313t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:158dd96d2207e21c966063e1635b1063cd7787b627b6f07305315dd73d9c679e", size = 33019667, upload-time = "2026-02-23T00:20:17.208Z" }, - { url = "https://files.pythonhosted.org/packages/f3/c3/2d834a5ac7bf3a0c806ad1508efc02dda3c8c61472a56132d7894c312dea/scipy-1.17.1-cp313-cp313t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:74cbb80d93260fe2ffa334efa24cb8f2f0f622a9b9febf8b483c0b865bfb3475", size = 35264159, upload-time = "2026-02-23T00:20:23.087Z" }, - { url = "https://files.pythonhosted.org/packages/4d/77/d3ed4becfdbd217c52062fafe35a72388d1bd82c2d0ba5ca19d6fcc93e11/scipy-1.17.1-cp313-cp313t-musllinux_1_2_aarch64.whl", hash = "sha256:dbc12c9f3d185f5c737d801da555fb74b3dcfa1a50b66a1a93e09190f41fab50", size = 35102771, upload-time = "2026-02-23T00:20:28.636Z" }, - { url = "https://files.pythonhosted.org/packages/bd/12/d19da97efde68ca1ee5538bb261d5d2c062f0c055575128f11a2730e3ac1/scipy-1.17.1-cp313-cp313t-musllinux_1_2_x86_64.whl", hash = "sha256:94055a11dfebe37c656e70317e1996dc197e1a15bbcc351bcdd4610e128fe1ca", size = 37665910, upload-time = "2026-02-23T00:20:34.743Z" }, - { url = "https://files.pythonhosted.org/packages/06/1c/1172a88d507a4baaf72c5a09bb6c018fe2ae0ab622e5830b703a46cc9e44/scipy-1.17.1-cp313-cp313t-win_amd64.whl", hash = "sha256:e30bdeaa5deed6bc27b4cc490823cd0347d7dae09119b8803ae576ea0ce52e4c", size = 36562980, upload-time = "2026-02-23T00:20:40.575Z" }, - { url = "https://files.pythonhosted.org/packages/70/b0/eb757336e5a76dfa7911f63252e3b7d1de00935d7705cf772db5b45ec238/scipy-1.17.1-cp313-cp313t-win_arm64.whl", hash = "sha256:a720477885a9d2411f94a93d16f9d89bad0f28ca23c3f8daa521e2dcc3f44d49", size = 24856543, upload-time = "2026-02-23T00:20:45.313Z" }, - { url = "https://files.pythonhosted.org/packages/cf/83/333afb452af6f0fd70414dc04f898647ee1423979ce02efa75c3b0f2c28e/scipy-1.17.1-cp314-cp314-macosx_10_14_x86_64.whl", hash = "sha256:a48a72c77a310327f6a3a920092fa2b8fd03d7deaa60f093038f22d98e096717", size = 31584510, upload-time = "2026-02-23T00:21:01.015Z" }, - { url = "https://files.pythonhosted.org/packages/ed/a6/d05a85fd51daeb2e4ea71d102f15b34fedca8e931af02594193ae4fd25f7/scipy-1.17.1-cp314-cp314-macosx_12_0_arm64.whl", hash = "sha256:45abad819184f07240d8a696117a7aacd39787af9e0b719d00285549ed19a1e9", size = 28170131, upload-time = "2026-02-23T00:21:05.888Z" }, - { url = "https://files.pythonhosted.org/packages/db/7b/8624a203326675d7746a254083a187398090a179335b2e4a20e2ddc46e83/scipy-1.17.1-cp314-cp314-macosx_14_0_arm64.whl", hash = "sha256:3fd1fcdab3ea951b610dc4cef356d416d5802991e7e32b5254828d342f7b7e0b", size = 20342032, upload-time = "2026-02-23T00:21:09.904Z" }, - { url = "https://files.pythonhosted.org/packages/c9/35/2c342897c00775d688d8ff3987aced3426858fd89d5a0e26e020b660b301/scipy-1.17.1-cp314-cp314-macosx_14_0_x86_64.whl", hash = "sha256:7bdf2da170b67fdf10bca777614b1c7d96ae3ca5794fd9587dce41eb2966e866", size = 22678766, upload-time = "2026-02-23T00:21:14.313Z" }, - { url = "https://files.pythonhosted.org/packages/ef/f2/7cdb8eb308a1a6ae1e19f945913c82c23c0c442a462a46480ce487fdc0ac/scipy-1.17.1-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:adb2642e060a6549c343603a3851ba76ef0b74cc8c079a9a58121c7ec9fe2350", size = 32957007, upload-time = "2026-02-23T00:21:19.663Z" }, - { url = "https://files.pythonhosted.org/packages/0b/2e/7eea398450457ecb54e18e9d10110993fa65561c4f3add5e8eccd2b9cd41/scipy-1.17.1-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:eee2cfda04c00a857206a4330f0c5e3e56535494e30ca445eb19ec624ae75118", size = 35221333, upload-time = "2026-02-23T00:21:25.278Z" }, - { url = "https://files.pythonhosted.org/packages/d9/77/5b8509d03b77f093a0d52e606d3c4f79e8b06d1d38c441dacb1e26cacf46/scipy-1.17.1-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:d2650c1fb97e184d12d8ba010493ee7b322864f7d3d00d3f9bb97d9c21de4068", size = 35042066, upload-time = "2026-02-23T00:21:31.358Z" }, - { url = "https://files.pythonhosted.org/packages/f9/df/18f80fb99df40b4070328d5ae5c596f2f00fffb50167e31439e932f29e7d/scipy-1.17.1-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:08b900519463543aa604a06bec02461558a6e1cef8fdbb8098f77a48a83c8118", size = 37612763, upload-time = "2026-02-23T00:21:37.247Z" }, - { url = "https://files.pythonhosted.org/packages/4b/39/f0e8ea762a764a9dc52aa7dabcfad51a354819de1f0d4652b6a1122424d6/scipy-1.17.1-cp314-cp314-win_amd64.whl", hash = "sha256:3877ac408e14da24a6196de0ddcace62092bfc12a83823e92e49e40747e52c19", size = 37290984, upload-time = "2026-02-23T00:22:35.023Z" }, - { url = "https://files.pythonhosted.org/packages/7c/56/fe201e3b0f93d1a8bcf75d3379affd228a63d7e2d80ab45467a74b494947/scipy-1.17.1-cp314-cp314-win_arm64.whl", hash = "sha256:f8885db0bc2bffa59d5c1b72fad7a6a92d3e80e7257f967dd81abb553a90d293", size = 25192877, upload-time = "2026-02-23T00:22:39.798Z" }, - { url = "https://files.pythonhosted.org/packages/96/ad/f8c414e121f82e02d76f310f16db9899c4fcde36710329502a6b2a3c0392/scipy-1.17.1-cp314-cp314t-macosx_10_14_x86_64.whl", hash = "sha256:1cc682cea2ae55524432f3cdff9e9a3be743d52a7443d0cba9017c23c87ae2f6", size = 31949750, upload-time = "2026-02-23T00:21:42.289Z" }, - { url = "https://files.pythonhosted.org/packages/7c/b0/c741e8865d61b67c81e255f4f0a832846c064e426636cd7de84e74d209be/scipy-1.17.1-cp314-cp314t-macosx_12_0_arm64.whl", hash = "sha256:2040ad4d1795a0ae89bfc7e8429677f365d45aa9fd5e4587cf1ea737f927b4a1", size = 28585858, upload-time = "2026-02-23T00:21:47.706Z" }, - { url = "https://files.pythonhosted.org/packages/ed/1b/3985219c6177866628fa7c2595bfd23f193ceebbe472c98a08824b9466ff/scipy-1.17.1-cp314-cp314t-macosx_14_0_arm64.whl", hash = "sha256:131f5aaea57602008f9822e2115029b55d4b5f7c070287699fe45c661d051e39", size = 20757723, upload-time = "2026-02-23T00:21:52.039Z" }, - { url = "https://files.pythonhosted.org/packages/c0/19/2a04aa25050d656d6f7b9e7b685cc83d6957fb101665bfd9369ca6534563/scipy-1.17.1-cp314-cp314t-macosx_14_0_x86_64.whl", hash = "sha256:9cdc1a2fcfd5c52cfb3045feb399f7b3ce822abdde3a193a6b9a60b3cb5854ca", size = 23043098, upload-time = "2026-02-23T00:21:56.185Z" }, - { url = "https://files.pythonhosted.org/packages/86/f1/3383beb9b5d0dbddd030335bf8a8b32d4317185efe495374f134d8be6cce/scipy-1.17.1-cp314-cp314t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:6e3dcd57ab780c741fde8dc68619de988b966db759a3c3152e8e9142c26295ad", size = 33030397, upload-time = "2026-02-23T00:22:01.404Z" }, - { url = "https://files.pythonhosted.org/packages/41/68/8f21e8a65a5a03f25a79165ec9d2b28c00e66dc80546cf5eb803aeeff35b/scipy-1.17.1-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:a9956e4d4f4a301ebf6cde39850333a6b6110799d470dbbb1e25326ac447f52a", size = 35281163, upload-time = "2026-02-23T00:22:07.024Z" }, - { url = "https://files.pythonhosted.org/packages/84/8d/c8a5e19479554007a5632ed7529e665c315ae7492b4f946b0deb39870e39/scipy-1.17.1-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:a4328d245944d09fd639771de275701ccadf5f781ba0ff092ad141e017eccda4", size = 35116291, upload-time = "2026-02-23T00:22:12.585Z" }, - { url = "https://files.pythonhosted.org/packages/52/52/e57eceff0e342a1f50e274264ed47497b59e6a4e3118808ee58ddda7b74a/scipy-1.17.1-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:a77cbd07b940d326d39a1d1b37817e2ee4d79cb30e7338f3d0cddffae70fcaa2", size = 37682317, upload-time = "2026-02-23T00:22:18.513Z" }, - { url = "https://files.pythonhosted.org/packages/11/2f/b29eafe4a3fbc3d6de9662b36e028d5f039e72d345e05c250e121a230dd4/scipy-1.17.1-cp314-cp314t-win_amd64.whl", hash = "sha256:eb092099205ef62cd1782b006658db09e2fed75bffcae7cc0d44052d8aa0f484", size = 37345327, upload-time = "2026-02-23T00:22:24.442Z" }, - { url = "https://files.pythonhosted.org/packages/07/39/338d9219c4e87f3e708f18857ecd24d22a0c3094752393319553096b98af/scipy-1.17.1-cp314-cp314t-win_arm64.whl", hash = "sha256:200e1050faffacc162be6a486a984a0497866ec54149a01270adc8a59b7c7d21", size = 25489165, upload-time = "2026-02-23T00:22:29.563Z" }, -] - -[[package]] -name = "scipy" -version = "1.18.0" -source = { registry = "https://pypi.org/simple" } -resolution-markers = [ - "python_full_version >= '3.14' and sys_platform == 'win32'", - "python_full_version >= '3.14' and sys_platform != 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform == 'win32'", - "python_full_version >= '3.12' and python_full_version < '3.14' and sys_platform != 'win32'", -] -dependencies = [ - { name = "numpy", version = "2.4.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.12'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/a7/25/c2700dfaf6442b4effaa91af24ebce5dc9d31bb4a69706313aae70d72cd0/scipy-1.18.0.tar.gz", hash = "sha256:67b2ad2ad54c72ca6d04975a9b2df8c3638c34ddd5b28738e94fc2b57929d378", size = 30774447, upload-time = "2026-06-19T15:01:43.456Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/6a/19/ca10ead60b0acc80b2b833c2c4a4f2ff753d0f58b811f70d911c7e94a25c/scipy-1.18.0-cp312-cp312-macosx_10_15_x86_64.whl", hash = "sha256:7bd21faaf5a1a3b2eff922d02db5f191b99a6518db9078a8fb23169f6d22259a", size = 31056519, upload-time = "2026-06-19T14:59:45.203Z" }, - { url = "https://files.pythonhosted.org/packages/96/72/1e6442a00cd2924d361aa1b642ab6373ec35c6fabf311a760be9f76e0f13/scipy-1.18.0-cp312-cp312-macosx_12_0_arm64.whl", hash = "sha256:265915e79107de9f946b855e50d7470d5893ec3f54b342e1aa6201cbdcd8bb6b", size = 28681889, upload-time = "2026-06-19T14:59:48.103Z" }, - { url = "https://files.pythonhosted.org/packages/9b/2d/11dd93d21e147a73ba22bd75c0b9208d3a2e0ec76d53170ce7d9029b1015/scipy-1.18.0-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:9ab7b758be6940954a713ee466e2043e9f6e2ed965c1fce5c91039f4be3d90a9", size = 20423580, upload-time = "2026-06-19T14:59:50.665Z" }, - { url = "https://files.pythonhosted.org/packages/9c/01/93552f75e0d2a7dd115a45e59209c51e8d514daff02fc887d2623be06fe1/scipy-1.18.0-cp312-cp312-macosx_14_0_x86_64.whl", hash = "sha256:97b6cddaaee0a779ef6b5ca83c9604b27cc16b2b8fc22c142652df8793319fb8", size = 23054441, upload-time = "2026-06-19T14:59:53.564Z" }, - { url = "https://files.pythonhosted.org/packages/3c/23/21f5e703643d66f21faa6b4c73195bfcad70c55efcb4f1ab327cd7c4101a/scipy-1.18.0-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:52a96e21517c7292375c0e27dd796a811f03fcea5fd4d108fdfea8145dcf17ab", size = 33968720, upload-time = "2026-06-19T14:59:56.415Z" }, - { url = "https://files.pythonhosted.org/packages/dd/aa/1b939f6c67ed68635bb538e6752d3dacc02f66535182e939a89581a44e9c/scipy-1.18.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:1f55797419e16e7f30cf88ffb3113ce0467f00cfe3f70d5c281730b21769bfc2", size = 35287115, upload-time = "2026-06-19T14:59:59.411Z" }, - { url = "https://files.pythonhosted.org/packages/b6/ff/eec46be7e9234208f801062b53e1983085eddebd693f6c9bfb03b459830d/scipy-1.18.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:ad033410e2e0672ffdc1042110cef20e1c46f8fd0616cee1d44d8d58fad8fc11", size = 35577989, upload-time = "2026-06-19T15:00:02.235Z" }, - { url = "https://files.pythonhosted.org/packages/84/ca/210d4759c7210bb7d269437421959b39a33434e2776b60c5cb8a763bb30a/scipy-1.18.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:4a55985d54c769c872e64b7f4c8a81cc30ef700cc04296abbbf3705439c126de", size = 37421717, upload-time = "2026-06-19T15:00:05.102Z" }, - { url = "https://files.pythonhosted.org/packages/2b/54/9a9edb45345bd6744da5ddfb6628e5d5185920494c6a67ec45b6381004cb/scipy-1.18.0-cp312-cp312-win_amd64.whl", hash = "sha256:71ccc8faa2dd16ac310233203474a8b5cb67f10dedd54a3116d34943f4b19132", size = 36597428, upload-time = "2026-06-19T15:00:08.112Z" }, - { url = "https://files.pythonhosted.org/packages/99/0e/33f32a2a58987e26aec0f7df252cbbad1e90ae77bdbc76f40dd4ed0cf0ea/scipy-1.18.0-cp312-cp312-win_arm64.whl", hash = "sha256:d88363fd9d8fbd3511bd273f1a49efb2a540773ddf92a91d57498ce7dd7f3e76", size = 24351481, upload-time = "2026-06-19T15:00:11.103Z" }, - { url = "https://files.pythonhosted.org/packages/05/52/9c0136c2de7ae0779b7b366447766cec6d9f0702c56bb8ffeb04c8fd3af4/scipy-1.18.0-cp313-cp313-macosx_10_15_x86_64.whl", hash = "sha256:09143f676d157d9f546d663504ef9c1becb819824f1afc018814176411942446", size = 31036107, upload-time = "2026-06-19T15:00:14.03Z" }, - { url = "https://files.pythonhosted.org/packages/02/73/0291a64843270f4efb86cdcf2ee0f2048631b65ec6b405398b2b4dbf11bf/scipy-1.18.0-cp313-cp313-macosx_12_0_arm64.whl", hash = "sha256:5efe260f69417b97ddae455bfb5a95e8359f7f66ad7fa9522a60feb66f169520", size = 28663303, upload-time = "2026-06-19T15:00:16.819Z" }, - { url = "https://files.pythonhosted.org/packages/d3/0f/10ffa0b697a572f4e0d48b92a88895d366422f019f723e7e14a84c050dac/scipy-1.18.0-cp313-cp313-macosx_14_0_arm64.whl", hash = "sha256:68363b7eaacd8b5dd426df56d782cc156468ac79a127a1b87ca597d6e2e82197", size = 20404960, upload-time = "2026-06-19T15:00:19.635Z" }, - { url = "https://files.pythonhosted.org/packages/7e/d2/e896cea21ba8edd6c81d4c55b1ffcc717e79698dcbebf9641b4cfb4c6622/scipy-1.18.0-cp313-cp313-macosx_14_0_x86_64.whl", hash = "sha256:c5557d8be5da8e41353fcd4d21491fdbab83b062fc579e94dc09a7c8ab4f669b", size = 23034074, upload-time = "2026-06-19T15:00:22.107Z" }, - { url = "https://files.pythonhosted.org/packages/ea/b2/e83ea34279a52c03374477c74006256ec78df65fc877baa4617d6de1d202/scipy-1.18.0-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:0d13bca67c096d89fb95ced0d8921807300fce0275643aef9533cc63a0773468", size = 33942038, upload-time = "2026-06-19T15:00:24.964Z" }, - { url = "https://files.pythonhosted.org/packages/f6/af/e8fe5fb136f51e2b01678b92cb4106d10d8cd68ec147ead2e7cb0ac75398/scipy-1.18.0-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:a46f9273dbd0eb1cefba61c9b8648b4dfe3cbc14a080176f9a73e44b8336dc7f", size = 35266390, upload-time = "2026-06-19T15:00:28.059Z" }, - { url = "https://files.pythonhosted.org/packages/3a/49/2c5cbb907b56695fc67517811d1db234dfd83381a84814ec220aded2794d/scipy-1.18.0-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:5aba46108853ddfc77906b6557aac839d2b52e900c1d72a1180adaaab58d265f", size = 35551324, upload-time = "2026-06-19T15:00:31.014Z" }, - { url = "https://files.pythonhosted.org/packages/bb/73/eda39f7a2d306ff0ffc574afd13c0bbb6d10a603d9a413998ee269487a80/scipy-1.18.0-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:b6f758e35f12757b5d95c00bc6de2438e229c2664b7a92e96f205959d9f2dfa4", size = 37404785, upload-time = "2026-06-19T15:00:34.072Z" }, - { url = "https://files.pythonhosted.org/packages/b7/d2/ae881ee28d014f38e0ccbfd974a06a919ba9af34f1f74bf42b5301891d63/scipy-1.18.0-cp313-cp313-win_amd64.whl", hash = "sha256:1afac4a847207c7ff8efd321734a50b06d0280b3b2a2c0fc2f413101747ad7c7", size = 36554943, upload-time = "2026-06-19T15:00:36.903Z" }, - { url = "https://files.pythonhosted.org/packages/70/3a/21154e2d54eb3639c6bf4dbae2e531c68356bfe95990daa30df33b30d556/scipy-1.18.0-cp313-cp313-win_arm64.whl", hash = "sha256:c5dbddf60e58c2312316d097271a8e73d40eaf2eabfa4d95ed7d3695bbf2ce7b", size = 24350911, upload-time = "2026-06-19T15:00:40.062Z" }, - { url = "https://files.pythonhosted.org/packages/78/b5/915a19b3de2f7430062b509653563db1633ddbb6f021b06731521115d4e2/scipy-1.18.0-cp314-cp314-macosx_10_15_x86_64.whl", hash = "sha256:4c256ee70c0d1a8a2ace807e199ccd4e3f57037433842abb3fb36bc17eaa9578", size = 31036253, upload-time = "2026-06-19T15:00:43.216Z" }, - { url = "https://files.pythonhosted.org/packages/d7/88/b72def7262e150d16be13fca37a96481138d624e700340bc3362a7588929/scipy-1.18.0-cp314-cp314-macosx_12_0_arm64.whl", hash = "sha256:2ef3abc54a4ffc53765374b0d5728532dfdd2585ed23f6b11c206a1f0b1b9af8", size = 28673758, upload-time = "2026-06-19T15:00:46.663Z" }, - { url = "https://files.pythonhosted.org/packages/91/02/2e636a61a525632c373cf6a9c24442a3ffb79e364d38e98b32042964ac32/scipy-1.18.0-cp314-cp314-macosx_14_0_arm64.whl", hash = "sha256:f2a6af57bd9e4a75d70e4117e78a1bbee84f79ae3fbb6d0111005d6ebcc4cb8d", size = 20415514, upload-time = "2026-06-19T15:00:49.399Z" }, - { url = "https://files.pythonhosted.org/packages/c9/b6/2135974442f6aba159d9d39d774a1c8cb19947016725d69fecc685df45bf/scipy-1.18.0-cp314-cp314-macosx_14_0_x86_64.whl", hash = "sha256:3f1ac564d3bf6c03d861d2cd87a1bea0da2887136f7fb1bf519c05a8971452d6", size = 23034398, upload-time = "2026-06-19T15:00:51.941Z" }, - { url = "https://files.pythonhosted.org/packages/f6/e6/ba89ec5abf6ee9257c0d1ec985573f3ae32742c24bc03e016388a40b1b15/scipy-1.18.0-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:40395a5fcd1abee49a5c7aaa98c29db393eedc835138560a588c47ec16156690", size = 33998032, upload-time = "2026-06-19T15:00:54.838Z" }, - { url = "https://files.pythonhosted.org/packages/7f/c4/bc41eb19b0fd0db868f4132920879019318d80cc522ad8f2bca4611af808/scipy-1.18.0-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:8ca01e8ae69f1b18e9a58d91afead31be3cef0dd905a10249dac559ee15460a0", size = 35283333, upload-time = "2026-06-19T15:00:58.152Z" }, - { url = "https://files.pythonhosted.org/packages/53/a4/cbdeef6eb3830a8462a9d4ada814de5fc984345cc9ecf17cbec51a036f1e/scipy-1.18.0-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:7a7f3b01647384dbc3a711e8c6778e0aabbe93959249fef5c7393396bcac0867", size = 35610216, upload-time = "2026-06-19T15:01:01.155Z" }, - { url = "https://files.pythonhosted.org/packages/80/4d/b2b82502b65f661d1b789c1665dcdf315d5f12194e06fc0b37946294ebae/scipy-1.18.0-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:6aa94e78ec192a30063a5e72e561c28af769dc311190b24fe91774eff1969709", size = 37418960, upload-time = "2026-06-19T15:01:04.155Z" }, - { url = "https://files.pythonhosted.org/packages/93/3e/902d836831474b0ab5a37d16404f7bc5fafd9efba632890e271ba952635f/scipy-1.18.0-cp314-cp314-win_amd64.whl", hash = "sha256:2d8bbdc6c817f5b4006a54d799d4f5bab6f910193cbb9a1ff310833d4d270f61", size = 37288845, upload-time = "2026-06-19T15:01:07.822Z" }, - { url = "https://files.pythonhosted.org/packages/b6/43/8d73b337a3bdb14daa0314f0434210747c02d79d729ce1777574a817dcf6/scipy-1.18.0-cp314-cp314-win_arm64.whl", hash = "sha256:18e9575f1569b2c54174e6159d32942e03731177f63dce7975f0a0c88d102f5b", size = 24988971, upload-time = "2026-06-19T15:01:11.076Z" }, - { url = "https://files.pythonhosted.org/packages/b4/b4/f11918b0508a2787031a0499a03fbe3546f3bb5ca05d01038c45b278c09a/scipy-1.18.0-cp314-cp314t-macosx_10_15_x86_64.whl", hash = "sha256:f351e0dd702687d12a402b867a1b4146a256923e1c38317cbc472f6372b94707", size = 31399325, upload-time = "2026-06-19T15:01:13.723Z" }, - { url = "https://files.pythonhosted.org/packages/7b/d1/1f287b57c0ff0ee5185dff3946d92c8017d39b0e431f0ae79a3ff1859512/scipy-1.18.0-cp314-cp314t-macosx_12_0_arm64.whl", hash = "sha256:7c7a51b33ce387193c97f228320cf8e87361daa1bba750638677729598b3e677", size = 29092110, upload-time = "2026-06-19T15:01:16.908Z" }, - { url = "https://files.pythonhosted.org/packages/ff/1a/7b74eb6c392fdcb27d414c0e7558a6d0231eb3b6d73571f479bb81ea8794/scipy-1.18.0-cp314-cp314t-macosx_14_0_arm64.whl", hash = "sha256:84031d7b052a54fae2f8632e0ec802073d385476eb9a63079bce6e23ef9283d4", size = 20833811, upload-time = "2026-06-19T15:01:20.488Z" }, - { url = "https://files.pythonhosted.org/packages/7c/ad/f3941716320a7b9cb4d68734a903b45fe16eff5fb7da7e16f2e619304979/scipy-1.18.0-cp314-cp314t-macosx_14_0_x86_64.whl", hash = "sha256:56abf29a7c067dde59be8b9a22d606a4ea1b2f2a4b756d9d903c62818f5dacce", size = 23396644, upload-time = "2026-06-19T15:01:23.364Z" }, - { url = "https://files.pythonhosted.org/packages/22/22/1446b62ffe07f9719b7d9b1b6a4e05a772833ae8f441fe4c22c34c9b250f/scipy-1.18.0-cp314-cp314t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:1ad44305cfa24b1ba5803cbbebf033590ccbac1aa5d612d727b785325ab408b0", size = 34079318, upload-time = "2026-06-19T15:01:26.002Z" }, - { url = "https://files.pythonhosted.org/packages/56/3b/b87da667098bb470fa30c7011b0ba351ee976dd395c78798c66e941665a3/scipy-1.18.0-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:945c1761b93f38d7f99ae81ae80c63e621471608c7eeead563f6df025585cd58", size = 35324320, upload-time = "2026-06-19T15:01:28.881Z" }, - { url = "https://files.pythonhosted.org/packages/f8/a1/c7932f91909759b0267f75fdea34e91309f96b895757534b76a90b6b4344/scipy-1.18.0-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:1a4441f15d620578772a49e5ab48c0ee1f7a0220e387110283062729136b2553", size = 35699541, upload-time = "2026-06-19T15:01:31.968Z" }, - { url = "https://files.pythonhosted.org/packages/f7/86/5185061a1fcc41d18c5dc2463969b3a3964b31d9ac67b2fb05d4c7ff7670/scipy-1.18.0-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:9aac6192fac56bf2ca534389d24623f07b39ff83317d58287285e7fbd622ff76", size = 37472480, upload-time = "2026-06-19T15:01:35.136Z" }, - { url = "https://files.pythonhosted.org/packages/31/8e/f04c68e39919a010d34f2ee1367fd705b0a25a02f609d755f0bfbc0a15fc/scipy-1.18.0-cp314-cp314t-win_amd64.whl", hash = "sha256:e40baea28ae7f5475c779741e2d90b1247c78531207b49c7030e698ff81cee3f", size = 37365390, upload-time = "2026-06-19T15:01:38.091Z" }, - { url = "https://files.pythonhosted.org/packages/d5/19/969dc072906c84dd0a3b05dcf57ea750936087d7873549e408b35cfc3f97/scipy-1.18.0-cp314-cp314t-win_arm64.whl", hash = "sha256:368e0a705903c466aa5f08eefb39e6b1b6b2d659e7352a31fd9e2438365be0f8", size = 25279661, upload-time = "2026-06-19T15:01:40.817Z" }, -] - -[[package]] -name = "sentence-transformers" -version = "5.6.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "huggingface-hub" }, - { name = "numpy", version = "2.2.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, - { name = "numpy", version = "2.4.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.11'" }, - { name = "scikit-learn", version = "1.7.2", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, - { name = "scikit-learn", version = "1.9.0", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.11'" }, - { name = "scipy", version = "1.15.3", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, - { name = "scipy", version = "1.17.1", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version == '3.11.*'" }, - { name = "scipy", version = "1.18.0", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.12'" }, - { name = "torch" }, - { name = "tqdm" }, - { name = "transformers" }, - { name = "typing-extensions" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/f9/56/d2cb00765a6b15c994a7fccf20f9032f16e8193ca49147cb5155166ad744/sentence_transformers-5.6.0.tar.gz", hash = "sha256:0e7164d051e416c1853ade7c274ff52af3f9da0f4be7f0b83d734c27699e1057", size = 453194, upload-time = "2026-06-16T14:01:56.42Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/76/c1/dc1582b79e9a2eb0cddf9559cd9bcdff084f541d6fe881fdd9d98630dba7/sentence_transformers-5.6.0-py3-none-any.whl", hash = "sha256:d2075b5e687a1611005e20ab04a6846994d51adfcf39610aed066af3c0c0b81f", size = 596411, upload-time = "2026-06-16T14:01:55.103Z" }, -] - -[[package]] -name = "setuptools" -version = "81.0.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/0d/1c/73e719955c59b8e424d015ab450f51c0af856ae46ea2da83eba51cc88de1/setuptools-81.0.0.tar.gz", hash = "sha256:487b53915f52501f0a79ccfd0c02c165ffe06631443a886740b91af4b7a5845a", size = 1198299, upload-time = "2026-02-06T21:10:39.601Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/e1/e3/c164c88b2e5ce7b24d667b9bd83589cf4f3520d97cad01534cd3c4f55fdb/setuptools-81.0.0-py3-none-any.whl", hash = "sha256:fdd925d5c5d9f62e4b74b30d6dd7828ce236fd6ed998a08d81de62ce5a6310d6", size = 1062021, upload-time = "2026-02-06T21:10:37.175Z" }, -] - -[[package]] -name = "shellingham" -version = "1.5.4" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/58/15/8b3609fd3830ef7b27b655beb4b4e9c62313a4e8da8c676e142cc210d58e/shellingham-1.5.4.tar.gz", hash = "sha256:8dbca0739d487e5bd35ab3ca4b36e11c4078f3a234bfce294b0a0291363404de", size = 10310, upload-time = "2023-10-24T04:13:40.426Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/e0/f9/0595336914c5619e5f28a1fb793285925a8cd4b432c9da0a987836c7f822/shellingham-1.5.4-py2.py3-none-any.whl", hash = "sha256:7ecfff8f2fd72616f7481040475a65b2bf8af90a56c89140852d1120324e8686", size = 9755, upload-time = "2023-10-24T04:13:38.866Z" }, -] - -[[package]] -name = "sqlite-vec" -version = "0.1.9" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/68/85/9fad0045d8e7c8df3e0fa5a56c630e8e15ad6e5ca2e6106fceb666aa6638/sqlite_vec-0.1.9-py3-none-macosx_10_6_x86_64.whl", hash = "sha256:1b62a7f0a060d9475575d4e599bbf94a13d85af896bc1ce86ee80d1b5b48e5fb", size = 131171, upload-time = "2026-03-31T08:02:31.717Z" }, - { url = "https://files.pythonhosted.org/packages/a4/3d/3677e0cd2f92e5ebc43cd29fbf565b75582bff1ccfa0b8327c7508e1084f/sqlite_vec-0.1.9-py3-none-macosx_11_0_arm64.whl", hash = "sha256:1d52e30513bae4cc9778ddbf6145610434081be4c3afe57cd877893bad9f6b6c", size = 165434, upload-time = "2026-03-31T08:02:32.712Z" }, - { url = "https://files.pythonhosted.org/packages/00/d4/f2b936d3bdc38eadcbd2a87875815db36430fab0363182ba5d12cd8e0b51/sqlite_vec-0.1.9-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:4e921e592f24a5f9a18f590b6ddd530eb637e2d474e3b1972f9bbeb773aa3cb9", size = 160076, upload-time = "2026-03-31T08:02:33.796Z" }, - { url = "https://files.pythonhosted.org/packages/6f/ad/6afd073b0f817b3e03f9e37ad626ae341805891f23c74b5292818f49ac63/sqlite_vec-0.1.9-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.manylinux1_x86_64.whl", hash = "sha256:1515727990b49e79bcaf75fdee2ffc7d461f8b66905013231251f1c8938e7786", size = 163388, upload-time = "2026-03-31T08:02:34.888Z" }, - { url = "https://files.pythonhosted.org/packages/42/89/81b2907cda14e566b9bf215e2ad82fc9b349edf07d2010756ffdb902f328/sqlite_vec-0.1.9-py3-none-win_amd64.whl", hash = "sha256:4a28dc12fa4b53d7b1dced22da2488fade444e96b5d16fd2d698cd670675cf32", size = 292804, upload-time = "2026-03-31T08:02:36.035Z" }, -] - -[[package]] -name = "sse-starlette" -version = "3.4.4" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "anyio" }, - { name = "starlette" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/f7/2b/58abc2d1fd397e7dde08e947e05c884d8ef2f78d5e2588c17a12d42d6994/sse_starlette-3.4.4.tar.gz", hash = "sha256:07e0fa0460138baf25cdd5fb28683472c3995dc1642225191b3832d62526bcb0", size = 31819, upload-time = "2026-05-12T17:37:17.019Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/dc/67/805710444ea8cc75fbf70b920ed431a560c4bf9c57f7d5a3117213189399/sse_starlette-3.4.4-py3-none-any.whl", hash = "sha256:3f4dd50d8aed2771a091f3a83000323fc3844541c16b4fe585ae2420cc6df973", size = 16514, upload-time = "2026-05-12T17:37:15.601Z" }, -] - -[[package]] -name = "starlette" -version = "1.3.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "anyio" }, - { name = "typing-extensions", marker = "python_full_version < '3.13'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/eb/e3/7c1dc7381d9f8ab7d854328ebfa884e62cb3f3d8549ddfd37c7814f42afa/starlette-1.3.1.tar.gz", hash = "sha256:05d0213193f2fbaae60e2ecb593b4add4262ad4e46536b54abe36f11a71724e0", size = 2703240, upload-time = "2026-06-12T09:23:11.602Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/ec/bb/2799cc2ede3ed41131f8975621e7213dfc7ef4acbbaadfa440f32500c370/starlette-1.3.1-py3-none-any.whl", hash = "sha256:c7372aae11c3c3f26a42df7bd626cec2f47d03483d261d369516a615a53714c6", size = 73632, upload-time = "2026-06-12T09:23:10.017Z" }, -] - -[[package]] -name = "sympy" -version = "1.14.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "mpmath" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/83/d3/803453b36afefb7c2bb238361cd4ae6125a569b4db67cd9e79846ba2d68c/sympy-1.14.0.tar.gz", hash = "sha256:d3d3fe8df1e5a0b42f0e7bdf50541697dbe7d23746e894990c030e2b05e72517", size = 7793921, upload-time = "2025-04-27T18:05:01.611Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/a2/09/77d55d46fd61b4a135c444fc97158ef34a095e5681d0a6c10b75bf356191/sympy-1.14.0-py3-none-any.whl", hash = "sha256:e091cc3e99d2141a0ba2847328f5479b05d94a6635cb96148ccb3f34671bd8f5", size = 6299353, upload-time = "2025-04-27T18:04:59.103Z" }, -] - -[[package]] -name = "threadpoolctl" -version = "3.6.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/b7/4d/08c89e34946fce2aec4fbb45c9016efd5f4d7f24af8e5d93296e935631d8/threadpoolctl-3.6.0.tar.gz", hash = "sha256:8ab8b4aa3491d812b623328249fab5302a68d2d71745c8a4c719a2fcaba9f44e", size = 21274, upload-time = "2025-03-13T13:49:23.031Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/32/d5/f9a850d79b0851d1d4ef6456097579a9005b31fea68726a4ae5f2d82ddd9/threadpoolctl-3.6.0-py3-none-any.whl", hash = "sha256:43a0b8fd5a2928500110039e43a5eed8480b918967083ea48dc3ab9f13c4a7fb", size = 18638, upload-time = "2025-03-13T13:49:21.846Z" }, -] - -[[package]] -name = "tokenizers" -version = "0.22.2" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "huggingface-hub" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/73/6f/f80cfef4a312e1fb34baf7d85c72d4411afde10978d4657f8cdd811d3ccc/tokenizers-0.22.2.tar.gz", hash = "sha256:473b83b915e547aa366d1eee11806deaf419e17be16310ac0a14077f1e28f917", size = 372115, upload-time = "2026-01-05T10:45:15.988Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/92/97/5dbfabf04c7e348e655e907ed27913e03db0923abb5dfdd120d7b25630e1/tokenizers-0.22.2-cp39-abi3-macosx_10_12_x86_64.whl", hash = "sha256:544dd704ae7238755d790de45ba8da072e9af3eea688f698b137915ae959281c", size = 3100275, upload-time = "2026-01-05T10:41:02.158Z" }, - { url = "https://files.pythonhosted.org/packages/2e/47/174dca0502ef88b28f1c9e06b73ce33500eedfac7a7692108aec220464e7/tokenizers-0.22.2-cp39-abi3-macosx_11_0_arm64.whl", hash = "sha256:1e418a55456beedca4621dbab65a318981467a2b188e982a23e117f115ce5001", size = 2981472, upload-time = "2026-01-05T10:41:00.276Z" }, - { url = "https://files.pythonhosted.org/packages/d6/84/7990e799f1309a8b87af6b948f31edaa12a3ed22d11b352eaf4f4b2e5753/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:2249487018adec45d6e3554c71d46eb39fa8ea67156c640f7513eb26f318cec7", size = 3290736, upload-time = "2026-01-05T10:40:32.165Z" }, - { url = "https://files.pythonhosted.org/packages/78/59/09d0d9ba94dcd5f4f1368d4858d24546b4bdc0231c2354aa31d6199f0399/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:25b85325d0815e86e0bac263506dd114578953b7b53d7de09a6485e4a160a7dd", size = 3168835, upload-time = "2026-01-05T10:40:38.847Z" }, - { url = "https://files.pythonhosted.org/packages/47/50/b3ebb4243e7160bda8d34b731e54dd8ab8b133e50775872e7a434e524c28/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_i686.manylinux2014_i686.whl", hash = "sha256:bfb88f22a209ff7b40a576d5324bf8286b519d7358663db21d6246fb17eea2d5", size = 3521673, upload-time = "2026-01-05T10:40:56.614Z" }, - { url = "https://files.pythonhosted.org/packages/e0/fa/89f4cb9e08df770b57adb96f8cbb7e22695a4cb6c2bd5f0c4f0ebcf33b66/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:1c774b1276f71e1ef716e5486f21e76333464f47bece56bbd554485982a9e03e", size = 3724818, upload-time = "2026-01-05T10:40:44.507Z" }, - { url = "https://files.pythonhosted.org/packages/64/04/ca2363f0bfbe3b3d36e95bf67e56a4c88c8e3362b658e616d1ac185d47f2/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:df6c4265b289083bf710dff49bc51ef252f9d5be33a45ee2bed151114a56207b", size = 3379195, upload-time = "2026-01-05T10:40:51.139Z" }, - { url = "https://files.pythonhosted.org/packages/2e/76/932be4b50ef6ccedf9d3c6639b056a967a86258c6d9200643f01269211ca/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:369cc9fc8cc10cb24143873a0d95438bb8ee257bb80c71989e3ee290e8d72c67", size = 3274982, upload-time = "2026-01-05T10:40:58.331Z" }, - { url = "https://files.pythonhosted.org/packages/1d/28/5f9f5a4cc211b69e89420980e483831bcc29dade307955cc9dc858a40f01/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:29c30b83d8dcd061078b05ae0cb94d3c710555fbb44861139f9f83dcca3dc3e4", size = 9478245, upload-time = "2026-01-05T10:41:04.053Z" }, - { url = "https://files.pythonhosted.org/packages/6c/fb/66e2da4704d6aadebf8cb39f1d6d1957df667ab24cff2326b77cda0dcb85/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_armv7l.whl", hash = "sha256:37ae80a28c1d3265bb1f22464c856bd23c02a05bb211e56d0c5301a435be6c1a", size = 9560069, upload-time = "2026-01-05T10:45:10.673Z" }, - { url = "https://files.pythonhosted.org/packages/16/04/fed398b05caa87ce9b1a1bb5166645e38196081b225059a6edaff6440fac/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_i686.whl", hash = "sha256:791135ee325f2336f498590eb2f11dc5c295232f288e75c99a36c5dbce63088a", size = 9899263, upload-time = "2026-01-05T10:45:12.559Z" }, - { url = "https://files.pythonhosted.org/packages/05/a1/d62dfe7376beaaf1394917e0f8e93ee5f67fea8fcf4107501db35996586b/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:38337540fbbddff8e999d59970f3c6f35a82de10053206a7562f1ea02d046fa5", size = 10033429, upload-time = "2026-01-05T10:45:14.333Z" }, - { url = "https://files.pythonhosted.org/packages/fd/18/a545c4ea42af3df6effd7d13d250ba77a0a86fb20393143bbb9a92e434d4/tokenizers-0.22.2-cp39-abi3-win32.whl", hash = "sha256:a6bf3f88c554a2b653af81f3204491c818ae2ac6fbc09e76ef4773351292bc92", size = 2502363, upload-time = "2026-01-05T10:45:20.593Z" }, - { url = "https://files.pythonhosted.org/packages/65/71/0670843133a43d43070abeb1949abfdef12a86d490bea9cd9e18e37c5ff7/tokenizers-0.22.2-cp39-abi3-win_amd64.whl", hash = "sha256:c9ea31edff2968b44a88f97d784c2f16dc0729b8b143ed004699ebca91f05c48", size = 2747786, upload-time = "2026-01-05T10:45:18.411Z" }, - { url = "https://files.pythonhosted.org/packages/72/f4/0de46cfa12cdcbcd464cc59fde36912af405696f687e53a091fb432f694c/tokenizers-0.22.2-cp39-abi3-win_arm64.whl", hash = "sha256:9ce725d22864a1e965217204946f830c37876eee3b2ba6fc6255e8e903d5fcbc", size = 2612133, upload-time = "2026-01-05T10:45:17.232Z" }, - { url = "https://files.pythonhosted.org/packages/84/04/655b79dbcc9b3ac5f1479f18e931a344af67e5b7d3b251d2dcdcd7558592/tokenizers-0.22.2-pp310-pypy310_pp73-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:753d47ebd4542742ef9261d9da92cd545b2cacbb48349a1225466745bb866ec4", size = 3282301, upload-time = "2026-01-05T10:40:34.858Z" }, - { url = "https://files.pythonhosted.org/packages/46/cd/e4851401f3d8f6f45d8480262ab6a5c8cb9c4302a790a35aa14eeed6d2fd/tokenizers-0.22.2-pp310-pypy310_pp73-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:e10bf9113d209be7cd046d40fbabbaf3278ff6d18eb4da4c500443185dc1896c", size = 3161308, upload-time = "2026-01-05T10:40:40.737Z" }, - { url = "https://files.pythonhosted.org/packages/6f/6e/55553992a89982cd12d4a66dddb5e02126c58677ea3931efcbe601d419db/tokenizers-0.22.2-pp310-pypy310_pp73-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:64d94e84f6660764e64e7e0b22baa72f6cd942279fdbb21d46abd70d179f0195", size = 3718964, upload-time = "2026-01-05T10:40:46.56Z" }, - { url = "https://files.pythonhosted.org/packages/59/8c/b1c87148aa15e099243ec9f0cf9d0e970cc2234c3257d558c25a2c5304e6/tokenizers-0.22.2-pp310-pypy310_pp73-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:f01a9c019878532f98927d2bacb79bbb404b43d3437455522a00a30718cdedb5", size = 3373542, upload-time = "2026-01-05T10:40:52.803Z" }, -] - -[[package]] -name = "tomli" -version = "2.4.1" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/22/de/48c59722572767841493b26183a0d1cc411d54fd759c5607c4590b6563a6/tomli-2.4.1.tar.gz", hash = "sha256:7c7e1a961a0b2f2472c1ac5b69affa0ae1132c39adcb67aba98568702b9cc23f", size = 17543, upload-time = "2026-03-25T20:22:03.828Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/f4/11/db3d5885d8528263d8adc260bb2d28ebf1270b96e98f0e0268d32b8d9900/tomli-2.4.1-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:f8f0fc26ec2cc2b965b7a3b87cd19c5c6b8c5e5f436b984e85f486d652285c30", size = 154704, upload-time = "2026-03-25T20:21:10.473Z" }, - { url = "https://files.pythonhosted.org/packages/6d/f7/675db52c7e46064a9aa928885a9b20f4124ecb9bc2e1ce74c9106648d202/tomli-2.4.1-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:4ab97e64ccda8756376892c53a72bd1f964e519c77236368527f758fbc36a53a", size = 149454, upload-time = "2026-03-25T20:21:12.036Z" }, - { url = "https://files.pythonhosted.org/packages/61/71/81c50943cf953efa35bce7646caab3cf457a7d8c030b27cfb40d7235f9ee/tomli-2.4.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:96481a5786729fd470164b47cdb3e0e58062a496f455ee41b4403be77cb5a076", size = 237561, upload-time = "2026-03-25T20:21:13.098Z" }, - { url = "https://files.pythonhosted.org/packages/48/c1/f41d9cb618acccca7df82aaf682f9b49013c9397212cb9f53219e3abac37/tomli-2.4.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:5a881ab208c0baf688221f8cecc5401bd291d67e38a1ac884d6736cbcd8247e9", size = 243824, upload-time = "2026-03-25T20:21:14.569Z" }, - { url = "https://files.pythonhosted.org/packages/22/e4/5a816ecdd1f8ca51fb756ef684b90f2780afc52fc67f987e3c61d800a46d/tomli-2.4.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:47149d5bd38761ac8be13a84864bf0b7b70bc051806bc3669ab1cbc56216b23c", size = 242227, upload-time = "2026-03-25T20:21:15.712Z" }, - { url = "https://files.pythonhosted.org/packages/6b/49/2b2a0ef529aa6eec245d25f0c703e020a73955ad7edf73e7f54ddc608aa5/tomli-2.4.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:ec9bfaf3ad2df51ace80688143a6a4ebc09a248f6ff781a9945e51937008fcbc", size = 247859, upload-time = "2026-03-25T20:21:17.001Z" }, - { url = "https://files.pythonhosted.org/packages/83/bd/6c1a630eaca337e1e78c5903104f831bda934c426f9231429396ce3c3467/tomli-2.4.1-cp311-cp311-win32.whl", hash = "sha256:ff2983983d34813c1aeb0fa89091e76c3a22889ee83ab27c5eeb45100560c049", size = 97204, upload-time = "2026-03-25T20:21:18.079Z" }, - { url = "https://files.pythonhosted.org/packages/42/59/71461df1a885647e10b6bb7802d0b8e66480c61f3f43079e0dcd315b3954/tomli-2.4.1-cp311-cp311-win_amd64.whl", hash = "sha256:5ee18d9ebdb417e384b58fe414e8d6af9f4e7a0ae761519fb50f721de398dd4e", size = 108084, upload-time = "2026-03-25T20:21:18.978Z" }, - { url = "https://files.pythonhosted.org/packages/b8/83/dceca96142499c069475b790e7913b1044c1a4337e700751f48ed723f883/tomli-2.4.1-cp311-cp311-win_arm64.whl", hash = "sha256:c2541745709bad0264b7d4705ad453b76ccd191e64aa6f0fc66b69a293a45ece", size = 95285, upload-time = "2026-03-25T20:21:20.309Z" }, - { url = "https://files.pythonhosted.org/packages/c1/ba/42f134a3fe2b370f555f44b1d72feebb94debcab01676bf918d0cb70e9aa/tomli-2.4.1-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:c742f741d58a28940ce01d58f0ab2ea3ced8b12402f162f4d534dfe18ba1cd6a", size = 155924, upload-time = "2026-03-25T20:21:21.626Z" }, - { url = "https://files.pythonhosted.org/packages/dc/c7/62d7a17c26487ade21c5422b646110f2162f1fcc95980ef7f63e73c68f14/tomli-2.4.1-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:7f86fd587c4ed9dd76f318225e7d9b29cfc5a9d43de44e5754db8d1128487085", size = 150018, upload-time = "2026-03-25T20:21:23.002Z" }, - { url = "https://files.pythonhosted.org/packages/5c/05/79d13d7c15f13bdef410bdd49a6485b1c37d28968314eabee452c22a7fda/tomli-2.4.1-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:ff18e6a727ee0ab0388507b89d1bc6a22b138d1e2fa56d1ad494586d61d2eae9", size = 244948, upload-time = "2026-03-25T20:21:24.04Z" }, - { url = "https://files.pythonhosted.org/packages/10/90/d62ce007a1c80d0b2c93e02cab211224756240884751b94ca72df8a875ca/tomli-2.4.1-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:136443dbd7e1dee43c68ac2694fde36b2849865fa258d39bf822c10e8068eac5", size = 253341, upload-time = "2026-03-25T20:21:25.177Z" }, - { url = "https://files.pythonhosted.org/packages/1a/7e/caf6496d60152ad4ed09282c1885cca4eea150bfd007da84aea07bcc0a3e/tomli-2.4.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:5e262d41726bc187e69af7825504c933b6794dc3fbd5945e41a79bb14c31f585", size = 248159, upload-time = "2026-03-25T20:21:26.364Z" }, - { url = "https://files.pythonhosted.org/packages/99/e7/c6f69c3120de34bbd882c6fba7975f3d7a746e9218e56ab46a1bc4b42552/tomli-2.4.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:5cb41aa38891e073ee49d55fbc7839cfdb2bc0e600add13874d048c94aadddd1", size = 253290, upload-time = "2026-03-25T20:21:27.46Z" }, - { url = "https://files.pythonhosted.org/packages/d6/2f/4a3c322f22c5c66c4b836ec58211641a4067364f5dcdd7b974b4c5da300c/tomli-2.4.1-cp312-cp312-win32.whl", hash = "sha256:da25dc3563bff5965356133435b757a795a17b17d01dbc0f42fb32447ddfd917", size = 98141, upload-time = "2026-03-25T20:21:28.492Z" }, - { url = "https://files.pythonhosted.org/packages/24/22/4daacd05391b92c55759d55eaee21e1dfaea86ce5c571f10083360adf534/tomli-2.4.1-cp312-cp312-win_amd64.whl", hash = "sha256:52c8ef851d9a240f11a88c003eacb03c31fc1c9c4ec64a99a0f922b93874fda9", size = 108847, upload-time = "2026-03-25T20:21:29.386Z" }, - { url = "https://files.pythonhosted.org/packages/68/fd/70e768887666ddd9e9f5d85129e84910f2db2796f9096aa02b721a53098d/tomli-2.4.1-cp312-cp312-win_arm64.whl", hash = "sha256:f758f1b9299d059cc3f6546ae2af89670cb1c4d48ea29c3cacc4fe7de3058257", size = 95088, upload-time = "2026-03-25T20:21:30.677Z" }, - { url = "https://files.pythonhosted.org/packages/07/06/b823a7e818c756d9a7123ba2cda7d07bc2dd32835648d1a7b7b7a05d848d/tomli-2.4.1-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:36d2bd2ad5fb9eaddba5226aa02c8ec3fa4f192631e347b3ed28186d43be6b54", size = 155866, upload-time = "2026-03-25T20:21:31.65Z" }, - { url = "https://files.pythonhosted.org/packages/14/6f/12645cf7f08e1a20c7eb8c297c6f11d31c1b50f316a7e7e1e1de6e2e7b7e/tomli-2.4.1-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:eb0dc4e38e6a1fd579e5d50369aa2e10acfc9cace504579b2faabb478e76941a", size = 149887, upload-time = "2026-03-25T20:21:33.028Z" }, - { url = "https://files.pythonhosted.org/packages/5c/e0/90637574e5e7212c09099c67ad349b04ec4d6020324539297b634a0192b0/tomli-2.4.1-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:c7f2c7f2b9ca6bdeef8f0fa897f8e05085923eb091721675170254cbc5b02897", size = 243704, upload-time = "2026-03-25T20:21:34.51Z" }, - { url = "https://files.pythonhosted.org/packages/10/8f/d3ddb16c5a4befdf31a23307f72828686ab2096f068eaf56631e136c1fdd/tomli-2.4.1-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:f3c6818a1a86dd6dca7ddcaaf76947d5ba31aecc28cb1b67009a5877c9a64f3f", size = 251628, upload-time = "2026-03-25T20:21:36.012Z" }, - { url = "https://files.pythonhosted.org/packages/e3/f1/dbeeb9116715abee2485bf0a12d07a8f31af94d71608c171c45f64c0469d/tomli-2.4.1-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:d312ef37c91508b0ab2cee7da26ec0b3ed2f03ce12bd87a588d771ae15dcf82d", size = 247180, upload-time = "2026-03-25T20:21:37.136Z" }, - { url = "https://files.pythonhosted.org/packages/d3/74/16336ffd19ed4da28a70959f92f506233bd7cfc2332b20bdb01591e8b1d1/tomli-2.4.1-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:51529d40e3ca50046d7606fa99ce3956a617f9b36380da3b7f0dd3dd28e68cb5", size = 251674, upload-time = "2026-03-25T20:21:38.298Z" }, - { url = "https://files.pythonhosted.org/packages/16/f9/229fa3434c590ddf6c0aa9af64d3af4b752540686cace29e6281e3458469/tomli-2.4.1-cp313-cp313-win32.whl", hash = "sha256:2190f2e9dd7508d2a90ded5ed369255980a1bcdd58e52f7fe24b8162bf9fedbd", size = 97976, upload-time = "2026-03-25T20:21:39.316Z" }, - { url = "https://files.pythonhosted.org/packages/6a/1e/71dfd96bcc1c775420cb8befe7a9d35f2e5b1309798f009dca17b7708c1e/tomli-2.4.1-cp313-cp313-win_amd64.whl", hash = "sha256:8d65a2fbf9d2f8352685bc1364177ee3923d6baf5e7f43ea4959d7d8bc326a36", size = 108755, upload-time = "2026-03-25T20:21:40.248Z" }, - { url = "https://files.pythonhosted.org/packages/83/7a/d34f422a021d62420b78f5c538e5b102f62bea616d1d75a13f0a88acb04a/tomli-2.4.1-cp313-cp313-win_arm64.whl", hash = "sha256:4b605484e43cdc43f0954ddae319fb75f04cc10dd80d830540060ee7cd0243cd", size = 95265, upload-time = "2026-03-25T20:21:41.219Z" }, - { url = "https://files.pythonhosted.org/packages/3c/fb/9a5c8d27dbab540869f7c1f8eb0abb3244189ce780ba9cd73f3770662072/tomli-2.4.1-cp314-cp314-macosx_10_15_x86_64.whl", hash = "sha256:fd0409a3653af6c147209d267a0e4243f0ae46b011aa978b1080359fddc9b6cf", size = 155726, upload-time = "2026-03-25T20:21:42.23Z" }, - { url = "https://files.pythonhosted.org/packages/62/05/d2f816630cc771ad836af54f5001f47a6f611d2d39535364f148b6a92d6b/tomli-2.4.1-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:a120733b01c45e9a0c34aeef92bf0cf1d56cfe81ed9d47d562f9ed591a9828ac", size = 149859, upload-time = "2026-03-25T20:21:43.386Z" }, - { url = "https://files.pythonhosted.org/packages/ce/48/66341bdb858ad9bd0ceab5a86f90eddab127cf8b046418009f2125630ecb/tomli-2.4.1-cp314-cp314-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:559db847dc486944896521f68d8190be1c9e719fced785720d2216fe7022b662", size = 244713, upload-time = "2026-03-25T20:21:44.474Z" }, - { url = "https://files.pythonhosted.org/packages/df/6d/c5fad00d82b3c7a3ab6189bd4b10e60466f22cfe8a08a9394185c8a8111c/tomli-2.4.1-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:01f520d4f53ef97964a240a035ec2a869fe1a37dde002b57ebc4417a27ccd853", size = 252084, upload-time = "2026-03-25T20:21:45.62Z" }, - { url = "https://files.pythonhosted.org/packages/00/71/3a69e86f3eafe8c7a59d008d245888051005bd657760e96d5fbfb0b740c2/tomli-2.4.1-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:7f94b27a62cfad8496c8d2513e1a222dd446f095fca8987fceef261225538a15", size = 247973, upload-time = "2026-03-25T20:21:46.937Z" }, - { url = "https://files.pythonhosted.org/packages/67/50/361e986652847fec4bd5e4a0208752fbe64689c603c7ae5ea7cb16b1c0ca/tomli-2.4.1-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:ede3e6487c5ef5d28634ba3f31f989030ad6af71edfb0055cbbd14189ff240ba", size = 256223, upload-time = "2026-03-25T20:21:48.467Z" }, - { url = "https://files.pythonhosted.org/packages/8c/9a/b4173689a9203472e5467217e0154b00e260621caa227b6fa01feab16998/tomli-2.4.1-cp314-cp314-win32.whl", hash = "sha256:3d48a93ee1c9b79c04bb38772ee1b64dcf18ff43085896ea460ca8dec96f35f6", size = 98973, upload-time = "2026-03-25T20:21:49.526Z" }, - { url = "https://files.pythonhosted.org/packages/14/58/640ac93bf230cd27d002462c9af0d837779f8773bc03dee06b5835208214/tomli-2.4.1-cp314-cp314-win_amd64.whl", hash = "sha256:88dceee75c2c63af144e456745e10101eb67361050196b0b6af5d717254dddf7", size = 109082, upload-time = "2026-03-25T20:21:50.506Z" }, - { url = "https://files.pythonhosted.org/packages/d5/2f/702d5e05b227401c1068f0d386d79a589bb12bf64c3d2c72ce0631e3bc49/tomli-2.4.1-cp314-cp314-win_arm64.whl", hash = "sha256:b8c198f8c1805dc42708689ed6864951fd2494f924149d3e4bce7710f8eb5232", size = 96490, upload-time = "2026-03-25T20:21:51.474Z" }, - { url = "https://files.pythonhosted.org/packages/45/4b/b877b05c8ba62927d9865dd980e34a755de541eb65fffba52b4cc495d4d2/tomli-2.4.1-cp314-cp314t-macosx_10_15_x86_64.whl", hash = "sha256:d4d8fe59808a54658fcc0160ecfb1b30f9089906c50b23bcb4c69eddc19ec2b4", size = 164263, upload-time = "2026-03-25T20:21:52.543Z" }, - { url = "https://files.pythonhosted.org/packages/24/79/6ab420d37a270b89f7195dec5448f79400d9e9c1826df982f3f8e97b24fd/tomli-2.4.1-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:7008df2e7655c495dd12d2a4ad038ff878d4ca4b81fccaf82b714e07eae4402c", size = 160736, upload-time = "2026-03-25T20:21:53.674Z" }, - { url = "https://files.pythonhosted.org/packages/02/e0/3630057d8eb170310785723ed5adcdfb7d50cb7e6455f85ba8a3deed642b/tomli-2.4.1-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:1d8591993e228b0c930c4bb0db464bdad97b3289fb981255d6c9a41aedc84b2d", size = 270717, upload-time = "2026-03-25T20:21:55.129Z" }, - { url = "https://files.pythonhosted.org/packages/7a/b4/1613716072e544d1a7891f548d8f9ec6ce2faf42ca65acae01d76ea06bb0/tomli-2.4.1-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:734e20b57ba95624ecf1841e72b53f6e186355e216e5412de414e3c51e5e3c41", size = 278461, upload-time = "2026-03-25T20:21:56.228Z" }, - { url = "https://files.pythonhosted.org/packages/05/38/30f541baf6a3f6df77b3df16b01ba319221389e2da59427e221ef417ac0c/tomli-2.4.1-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:8a650c2dbafa08d42e51ba0b62740dae4ecb9338eefa093aa5c78ceb546fcd5c", size = 274855, upload-time = "2026-03-25T20:21:57.653Z" }, - { url = "https://files.pythonhosted.org/packages/77/a3/ec9dd4fd2c38e98de34223b995a3b34813e6bdadf86c75314c928350ed14/tomli-2.4.1-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:504aa796fe0569bb43171066009ead363de03675276d2d121ac1a4572397870f", size = 283144, upload-time = "2026-03-25T20:21:59.089Z" }, - { url = "https://files.pythonhosted.org/packages/ef/be/605a6261cac79fba2ec0c9827e986e00323a1945700969b8ee0b30d85453/tomli-2.4.1-cp314-cp314t-win32.whl", hash = "sha256:b1d22e6e9387bf4739fbe23bfa80e93f6b0373a7f1b96c6227c32bef95a4d7a8", size = 108683, upload-time = "2026-03-25T20:22:00.214Z" }, - { url = "https://files.pythonhosted.org/packages/12/64/da524626d3b9cc40c168a13da8335fe1c51be12c0a63685cc6db7308daae/tomli-2.4.1-cp314-cp314t-win_amd64.whl", hash = "sha256:2c1c351919aca02858f740c6d33adea0c5deea37f9ecca1cc1ef9e884a619d26", size = 121196, upload-time = "2026-03-25T20:22:01.169Z" }, - { url = "https://files.pythonhosted.org/packages/5a/cd/e80b62269fc78fc36c9af5a6b89c835baa8af28ff5ad28c7028d60860320/tomli-2.4.1-cp314-cp314t-win_arm64.whl", hash = "sha256:eab21f45c7f66c13f2a9e0e1535309cee140182a9cdae1e041d02e47291e8396", size = 100393, upload-time = "2026-03-25T20:22:02.137Z" }, - { url = "https://files.pythonhosted.org/packages/7b/61/cceae43728b7de99d9b847560c262873a1f6c98202171fd5ed62640b494b/tomli-2.4.1-py3-none-any.whl", hash = "sha256:0d85819802132122da43cb86656f8d1f8c6587d54ae7dcaf30e90533028b49fe", size = 14583, upload-time = "2026-03-25T20:22:03.012Z" }, -] - -[[package]] -name = "torch" -version = "2.12.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "cuda-bindings", marker = "sys_platform == 'linux'" }, - { name = "cuda-toolkit", extra = ["cudart", "cufft", "cufile", "cupti", "curand", "cusolver", "cusparse", "nvjitlink", "nvrtc", "nvtx"], marker = "sys_platform == 'linux'" }, - { name = "filelock" }, - { name = "fsspec" }, - { name = "jinja2" }, - { name = "networkx", version = "3.4.2", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, - { name = "networkx", version = "3.6.1", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.11'" }, - { name = "nvidia-cublas", marker = "sys_platform == 'linux'" }, - { name = "nvidia-cudnn-cu13", marker = "sys_platform == 'linux'" }, - { name = "nvidia-cusparselt-cu13", marker = "sys_platform == 'linux'" }, - { name = "nvidia-nccl-cu13", marker = "sys_platform == 'linux'" }, - { name = "nvidia-nvshmem-cu13", marker = "sys_platform == 'linux'" }, - { name = "setuptools" }, - { name = "sympy" }, - { name = "triton", marker = "sys_platform == 'linux'" }, - { name = "typing-extensions" }, -] -wheels = [ - { url = "https://files.pythonhosted.org/packages/db/ed/ff0c4f8cef63977a646dc80e40c05cae873f4097b12dc87e1cd7e1cecf42/torch-2.12.1-cp310-cp310-macosx_14_0_arm64.whl", hash = "sha256:ec56e82be6a8b0c036771a77f7d32ad3c299770571af9815b3dafe61434389d5", size = 87967927, upload-time = "2026-06-17T21:08:43.16Z" }, - { url = "https://files.pythonhosted.org/packages/85/1b/c8ecf60c9dba535f9ea341c359c600c0bd877a7ca14b3296f13316321847/torch-2.12.1-cp310-cp310-manylinux_2_28_aarch64.whl", hash = "sha256:42cd7339bf266f14944710e8274be63e7e012bb937834a8d85a8327a9860eba6", size = 426366829, upload-time = "2026-06-17T21:07:18.574Z" }, - { url = "https://files.pythonhosted.org/packages/ab/d6/73d4a3f27e00526e98086f3a64ab609af1345cca62367749fbc3c8e4b83c/torch-2.12.1-cp310-cp310-manylinux_2_28_x86_64.whl", hash = "sha256:a7817f0f89a796d9de239d06f69faf5d7e19a6a5db6710a5ead777c912f9f50a", size = 532144834, upload-time = "2026-06-17T21:08:00.633Z" }, - { url = "https://files.pythonhosted.org/packages/e3/51/4010c8fa6f9d1f42c054a321970ca95ec58e4e4494f5b53a34c3f3c9e310/torch-2.12.1-cp310-cp310-win_amd64.whl", hash = "sha256:2af3d9cc866e0a15ae7635ff0a9c61d6624a353ad657f5bcd8d86c26cdc64693", size = 122949863, upload-time = "2026-06-17T21:08:39.016Z" }, - { url = "https://files.pythonhosted.org/packages/59/38/7028d3be540f1dcdf41660a2b01d0c51d2cb73915fe370d84e4d277a6d47/torch-2.12.1-cp311-cp311-macosx_14_0_arm64.whl", hash = "sha256:ef81f503912effea2ce3d9b12a2e3a6ed488943e91271c90c7a829f60baf6aa2", size = 87975425, upload-time = "2026-06-17T21:08:34.094Z" }, - { url = "https://files.pythonhosted.org/packages/5a/e3/750b3e3548635ceac03ba255daa26dbc7ed66ca3484dc4b4d955ab7f4501/torch-2.12.1-cp311-cp311-manylinux_2_28_aarch64.whl", hash = "sha256:107df6888624bdea41508f9aeb6149d9333c737a5530ceecb56c904e811369ae", size = 426379894, upload-time = "2026-06-17T21:06:55.077Z" }, - { url = "https://files.pythonhosted.org/packages/dc/ca/ed24783da629ff3e640ba3f70a7639e9045d3d88b93ee6bc47b8a28a1f2c/torch-2.12.1-cp311-cp311-manylinux_2_28_x86_64.whl", hash = "sha256:6e29e7e74d05bda7d955c75e99459f878ebd970ef851b4057edbd3b34a5eb4a3", size = 532169264, upload-time = "2026-06-17T21:08:17.65Z" }, - { url = "https://files.pythonhosted.org/packages/46/61/c63f0158446f3a98ea672b004d761b848911eba567ea4a624c7db5aadc04/torch-2.12.1-cp311-cp311-win_amd64.whl", hash = "sha256:a513506cfda3c1c78dabeb6574c1597538c0254b3d39af174dde35d8177f4ce3", size = 122953086, upload-time = "2026-06-17T21:08:27.69Z" }, - { url = "https://files.pythonhosted.org/packages/f0/54/efb7ebca77970012b0cc21687a55d70eb2ba514b2c2b8e18d9fb1222f3be/torch-2.12.1-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:d2dd0f2c5f7ccbddaf34cade0deaf476808368f902b9cdb7f36a2ab42301bc0e", size = 87991951, upload-time = "2026-06-17T21:07:49.309Z" }, - { url = "https://files.pythonhosted.org/packages/1e/00/4210d76ca7424981f04033ebe7e48816ab83287a62538747a58825db770c/torch-2.12.1-cp312-cp312-manylinux_2_28_aarch64.whl", hash = "sha256:2de4e19b88a481482c6c75291f2d6a52eda3ce51f311b29aa9b68499c830c07c", size = 426382721, upload-time = "2026-06-17T21:06:41.842Z" }, - { url = "https://files.pythonhosted.org/packages/76/1f/bc9f5a5aa569307076365f25afcebacb22e9c754b1bcfbaaa146627c7fda/torch-2.12.1-cp312-cp312-manylinux_2_28_x86_64.whl", hash = "sha256:649e4ced014ba646f76f8cb9c9726735a6323eb321b7919f942790a923f90921", size = 532261322, upload-time = "2026-06-17T21:06:06.673Z" }, - { url = "https://files.pythonhosted.org/packages/9e/49/c549461daa008159d006a76a991fbc2f26fa8bac27a4030c858463dcb20f/torch-2.12.1-cp312-cp312-win_amd64.whl", hash = "sha256:e86550597877fb272ddc52db2f85b82cb601ea7bd932576a0340152cae2200b3", size = 122988095, upload-time = "2026-06-17T21:07:44.9Z" }, - { url = "https://files.pythonhosted.org/packages/ff/4a/0300261818e1560d72cc160ac826005507e8b7ca0a35788b591436d05b4a/torch-2.12.1-cp313-cp313-macosx_14_0_arm64.whl", hash = "sha256:c75e93173c700bccd6bfcc4a9d19ce242ab6dacd1f1781483027a16239b9e650", size = 87992358, upload-time = "2026-06-17T21:07:40.299Z" }, - { url = "https://files.pythonhosted.org/packages/30/a7/874a5ca05e8f159211dca7921060f7057acc1adb26431e119fd150623efc/torch-2.12.1-cp313-cp313-manylinux_2_28_aarch64.whl", hash = "sha256:fcb61ccd20784b62bdd78ec84238a5cfb383b4994902e03bac95505ab360884c", size = 426386134, upload-time = "2026-06-17T21:07:31.481Z" }, - { url = "https://files.pythonhosted.org/packages/e1/75/20bb8fe9c1ad6538cce8cd0391b51927ae5af0b17ed1eab44b8824465dc1/torch-2.12.1-cp313-cp313-manylinux_2_28_x86_64.whl", hash = "sha256:f4afc8083dff08719edbea346644476e3cec0cf40ebe256be0ee5d5b7c7e8c0d", size = 532268019, upload-time = "2026-06-17T21:05:37.925Z" }, - { url = "https://files.pythonhosted.org/packages/d1/fa/824ddb662af55b2eabc0dbb7b57c7c0b1bcd93693754a2b8509ec4d16490/torch-2.12.1-cp313-cp313-win_amd64.whl", hash = "sha256:f92609e3b3ce72f25e2eb780d043ced2480c1a86c47c852604fc7a9108648386", size = 122987777, upload-time = "2026-06-17T21:07:09.49Z" }, - { url = "https://files.pythonhosted.org/packages/63/b7/1b49fe7086ea36839cc80abc43174c43d0ab6f676c0891c871c162f44fe3/torch-2.12.1-cp314-cp314-macosx_14_0_arm64.whl", hash = "sha256:e9b6f7d2dd66ea87a3ae620069d31335d594c06effb1a383bdd21cfe61e44ece", size = 88010025, upload-time = "2026-06-17T21:07:03.934Z" }, - { url = "https://files.pythonhosted.org/packages/d7/06/5b44063a6545036dcc680d2d303b137d9176cfb2cc1e1863e3ef94abeb52/torch-2.12.1-cp314-cp314-manylinux_2_28_aarch64.whl", hash = "sha256:7973ccd3d2cd35c74449213f7bded199bec6c6247e705cbeda7407af79703d91", size = 426392891, upload-time = "2026-06-17T21:05:52.261Z" }, - { url = "https://files.pythonhosted.org/packages/f8/dd/c9ce9a4b0eb3c5bb92d9ea56766e2c22559f0b45171149188494edcce80f/torch-2.12.1-cp314-cp314-manylinux_2_28_x86_64.whl", hash = "sha256:c64ac4aac16be5e296dcd912305605804b203333c690bf98c55bc09494ee92ad", size = 532272494, upload-time = "2026-06-17T21:06:22.72Z" }, - { url = "https://files.pythonhosted.org/packages/21/7c/f3a601fc1b1f663ff269bfe553654e638651939aa6563e8daa7167c33098/torch-2.12.1-cp314-cp314-win_amd64.whl", hash = "sha256:f6dc4caf7eb4adb38a2d9f536b51db56310fdd1254e69a2d96767e1367c892b3", size = 122987254, upload-time = "2026-06-17T21:06:33.199Z" }, - { url = "https://files.pythonhosted.org/packages/e6/8c/b8087556cf81ddd808dbeb34afb8396d7ae7a1694ab489f08b1a0004e7d0/torch-2.12.1-cp314-cp314t-macosx_14_0_arm64.whl", hash = "sha256:2afbb2bdaa8a95040e733f05492ddf133c3967c9b7ce0abd218d704b6cab437d", size = 88303173, upload-time = "2026-06-17T21:05:06.603Z" }, - { url = "https://files.pythonhosted.org/packages/4a/07/fe09d1699fbed2afa10ebc692ff2b99d113f2605b6748cea633989e2789a/torch-2.12.1-cp314-cp314t-manylinux_2_28_aarch64.whl", hash = "sha256:97eba061fcb042fed191400b15568990073d67eaacaa6ee9b7ca01dd8b790fe9", size = 426404009, upload-time = "2026-06-17T21:04:57.557Z" }, - { url = "https://files.pythonhosted.org/packages/2e/f7/0ce4f6c1962c60ded7270e0a9eb560fb615c92b89d332cf9e3dff36d5ecc/torch-2.12.1-cp314-cp314t-manylinux_2_28_x86_64.whl", hash = "sha256:3867b861391701012adb2df93360efb88494dca245a185e3bb7624495cfe3f33", size = 532184292, upload-time = "2026-06-17T21:05:17.526Z" }, - { url = "https://files.pythonhosted.org/packages/70/db/e384c12aba30320ca92aaaf557456cbcb26f04b4df307728bb8f019f5000/torch-2.12.1-cp314-cp314t-win_amd64.whl", hash = "sha256:dd15595f8fc764cffde8c6361a3beb6ef69a028c851b1b3e70e077f615980d4e", size = 123231142, upload-time = "2026-06-17T21:05:27.061Z" }, -] - -[[package]] -name = "tqdm" -version = "4.68.3" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "colorama", marker = "sys_platform == 'win32'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/87/d7/0535a28b1f5f24f6612fb3ff1e89fb1a8d160fee0f976e0aa6803862134b/tqdm-4.68.3.tar.gz", hash = "sha256:00dfa48452b6b6cfae3dd9885636c23d3422d1ec97c66d96818cbd5e0821d482", size = 170596, upload-time = "2026-06-17T07:36:52.105Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/d8/8e/bb97bb0c71802080bfc8952937d174e49cfc50de5c951dd47b2496f0dcdb/tqdm-4.68.3-py3-none-any.whl", hash = "sha256:39832cc2def2789a6f29df83f172db7416cea70052c0907a57801c5f2fdccb03", size = 78337, upload-time = "2026-06-17T07:36:50.132Z" }, -] - -[[package]] -name = "transformers" -version = "5.12.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "huggingface-hub" }, - { name = "numpy", version = "2.2.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.11'" }, - { name = "numpy", version = "2.4.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.11'" }, - { name = "packaging" }, - { name = "pyyaml" }, - { name = "regex" }, - { name = "safetensors" }, - { name = "tokenizers" }, - { name = "tqdm" }, - { name = "typer" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/aa/7c/8240f612819718100a9346dc28dea6a11370c3ca9c8c6eabadd3dea4ef29/transformers-5.12.1.tar.gz", hash = "sha256:679ee731c8225347889ad4fb3b2c926a62e9da3b7d284e9d12c791da7272466b", size = 8924054, upload-time = "2026-06-15T17:27:50.604Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/df/56/bbd60dd8668055803bf8ba55a81f9b8a8b31497f620109a9671d26a2076d/transformers-5.12.1-py3-none-any.whl", hash = "sha256:2a5e109d2021265df7098ffbb738295acaf5ad256f12cbc586db2ea4dcbb1a8a", size = 11150587, upload-time = "2026-06-15T17:27:46.679Z" }, -] - -[[package]] -name = "triton" -version = "3.7.1" -source = { registry = "https://pypi.org/simple" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/ec/ea/629cc37436ca5df93ce98956d09cd2ca1498bfee8ef4972d2fe48b9f958c/triton-3.7.1-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:3daf64305d6cea88d3334c65ebc9bcd0c64c9564a977084366aa768d57cbcf64", size = 184551013, upload-time = "2026-06-17T20:03:37.551Z" }, - { url = "https://files.pythonhosted.org/packages/15/76/c79c34311625227a288df3e483fc5cdf3d596624cbd4b4758c4cbdc14af3/triton-3.7.1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:ee89fbf782ec2ad50391dd1cf26cbea4f4467154c37f4773026da8fc31c0f58e", size = 197596267, upload-time = "2026-06-17T19:53:06.898Z" }, - { url = "https://files.pythonhosted.org/packages/7b/f9/19d842d06a08559534fa1eaab6ca551b1bcf40f06620bddec1babaa2772d/triton-3.7.1-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:d4a0e1cd4c4a76370ed74a8432a53cea28716827d19e40ffc732233e35ceb3f6", size = 184664887, upload-time = "2026-06-17T20:03:42.913Z" }, - { url = "https://files.pythonhosted.org/packages/cd/5e/fce69606f7f240297f163e25539906732b199530d486ce67ae319877e821/triton-3.7.1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:6744957e9fd610a29680ec2346057d0c86948ed3812468670719f391e94b44a5", size = 197701306, upload-time = "2026-06-17T19:53:13.673Z" }, - { url = "https://files.pythonhosted.org/packages/94/fa/f856e24deb462d5f18bd4b5a746957862ab9b6ee5834bda60605ec348366/triton-3.7.1-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:9497f2e696ee368862a181a90b2dcc03ca978cc4f602abd67c7d81022a6988e1", size = 184692359, upload-time = "2026-06-17T20:03:48.288Z" }, - { url = "https://files.pythonhosted.org/packages/c4/6f/fb96d15db6f36d6eae4cafb998c2e0353bf59d7c4ea1662d7497f269134a/triton-3.7.1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:7e40869937a68206ec70d7f25bb7ec6433cb083f9135e1f36dbd318dc449a728", size = 197719725, upload-time = "2026-06-17T19:53:20.419Z" }, - { url = "https://files.pythonhosted.org/packages/00/42/c5089d4d9327fcd1e862c599cc2927f39418f84dd11a84cb2ccff9d4787a/triton-3.7.1-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:cdbfc09d9ec58bc5e68321525653220de7515c199e7a8097a97c85e62b52cd0a", size = 184694629, upload-time = "2026-06-17T20:03:53.444Z" }, - { url = "https://files.pythonhosted.org/packages/07/42/2c3ac59253ae8892b6f307875263dd23dc875cdf732d3aea40d6d41fb7cb/triton-3.7.1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:58c0e131da05134a2a4788ccbcc0c1105cf0f54c8e98f19e34cd465396dc15eb", size = 197729241, upload-time = "2026-06-17T19:53:27.801Z" }, - { url = "https://files.pythonhosted.org/packages/40/71/e01aa7ad573883ed9456f130226babdec70b005e098c4d6226a6238e761b/triton-3.7.1-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:fe4ea396a06171f1f1f58cbd39c70b09294398f7dd7c620939bab54ad6f934fa", size = 184705764, upload-time = "2026-06-17T20:03:59.064Z" }, - { url = "https://files.pythonhosted.org/packages/a4/09/5683146fda6a2b569deb78ccfd8fbfea8bfe55f726b081c0a6bb18dd6f28/triton-3.7.1-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:2020153b08280415ec0da6607834e79166442147e78e144df06b508c75b186d2", size = 197729537, upload-time = "2026-06-17T19:53:35.516Z" }, - { url = "https://files.pythonhosted.org/packages/e9/f8/448220c3092019f9fdfab39ec47985968181d67da34b44f6a7f6280a5cbb/triton-3.7.1-cp314-cp314t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:c58e4c61f0c73b5dba3b5d19b4a7093c32f90dc18b2a7f121a7c16ccd31107b7", size = 184814760, upload-time = "2026-06-17T20:04:04.984Z" }, - { url = "https://files.pythonhosted.org/packages/f0/ac/229b7d4589d2e5937310e72c6d46e89599d16a4a12b479ffa1499fee8eb8/triton-3.7.1-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:10ba85fa2cca4a2fbdeb36bf1cb082f2c252bda55bf9fccd74f65ec5bc647e68", size = 197824404, upload-time = "2026-06-17T19:53:42.772Z" }, -] - -[[package]] -name = "typer" -version = "0.25.1" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "annotated-doc" }, - { name = "click" }, - { name = "rich" }, - { name = "shellingham" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/e4/51/9aed62104cea109b820bbd6c14245af756112017d309da813ef107d42e7e/typer-0.25.1.tar.gz", hash = "sha256:9616eb8853a09ffeabab1698952f33c6f29ffdbceb4eaeecf571880e8d7664cc", size = 122276, upload-time = "2026-04-30T19:32:16.964Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/3f/f9/2b3ff4e56e5fa7debfaf9eb135d0da96f3e9a1d5b27222223c7296336e5f/typer-0.25.1-py3-none-any.whl", hash = "sha256:75caa44ed46a03fb2dab8808753ffacdbfea88495e74c85a28c5eefcf5f39c89", size = 58409, upload-time = "2026-04-30T19:32:18.271Z" }, -] - -[[package]] -name = "typing-extensions" -version = "4.15.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/72/94/1a15dd82efb362ac84269196e94cf00f187f7ed21c242792a923cdb1c61f/typing_extensions-4.15.0.tar.gz", hash = "sha256:0cea48d173cc12fa28ecabc3b837ea3cf6f38c6d1136f85cbaaf598984861466", size = 109391, upload-time = "2025-08-25T13:49:26.313Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/18/67/36e9267722cc04a6b9f15c7f3441c2363321a3ea07da7ae0c0707beb2a9c/typing_extensions-4.15.0-py3-none-any.whl", hash = "sha256:f0fa19c6845758ab08074a0cfa8b7aecb71c999ca73d62883bc25cc018c4e548", size = 44614, upload-time = "2025-08-25T13:49:24.86Z" }, -] - -[[package]] -name = "typing-inspection" -version = "0.4.2" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "typing-extensions" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/55/e3/70399cb7dd41c10ac53367ae42139cf4b1ca5f36bb3dc6c9d33acdb43655/typing_inspection-0.4.2.tar.gz", hash = "sha256:ba561c48a67c5958007083d386c3295464928b01faa735ab8547c5692e87f464", size = 75949, upload-time = "2025-10-01T02:14:41.687Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/dc/9b/47798a6c91d8bdb567fe2698fe81e0c6b7cb7ef4d13da4114b41d239f65d/typing_inspection-0.4.2-py3-none-any.whl", hash = "sha256:4ed1cacbdc298c220f1bd249ed5287caa16f34d44ef4e9c3d0cbad5b521545e7", size = 14611, upload-time = "2025-10-01T02:14:40.154Z" }, -] - -[[package]] -name = "uvicorn" -version = "0.49.0" -source = { registry = "https://pypi.org/simple" } -dependencies = [ - { name = "click" }, - { name = "h11" }, - { name = "typing-extensions", marker = "python_full_version < '3.11'" }, -] -sdist = { url = "https://files.pythonhosted.org/packages/c4/1f/fa18009dea8469069cca78a4e877a008ab78f08b064bfc9ab891579077ff/uvicorn-0.49.0.tar.gz", hash = "sha256:ebf4271aa580d9de97f93192d4595176df6e91f9aae919ca73e4fc07df1e66a3", size = 91284, upload-time = "2026-06-03T22:01:30.448Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/88/fa/e1388bbcf24ef3274f45c0c1c7b501fd14971037c1b6ee23610553307497/uvicorn-0.49.0-py3-none-any.whl", hash = "sha256:ba3d14c3ee7e41c6c654c46c9eb489d33213cdd30aa1696eab1374337c13f68f", size = 71376, upload-time = "2026-06-03T22:01:29.037Z" }, -] diff --git a/scripts/eval_schedule.py b/scripts/eval_schedule.py deleted file mode 100644 index 935bc15..0000000 --- a/scripts/eval_schedule.py +++ /dev/null @@ -1,1170 +0,0 @@ -#!/usr/bin/env python3 -"""Run MainFrame eval suites on a fixed cadence and record schedule manifests.""" - -from __future__ import annotations - -import argparse -import json -import os -import re -import shutil -import subprocess -import sys -import time -from dataclasses import asdict, dataclass, field -from datetime import date, datetime, timezone -from pathlib import Path -from typing import Any - -ROOT = Path(__file__).resolve().parents[1] -REGISTRY_DIR = ROOT / "20_live" / "eval-registry" -SCHEDULE_LOG = REGISTRY_DIR / "schedule-runs.jsonl" -PROCESS_EVAL_OUTPUTS = ROOT / "30_projects" / "mainframe-process-eval" / "outputs" -MINDGRAPH_PROBE = ROOT / "30_projects" / "mindgraph-eval" / "scripts" / "retrieval_quality_probe.py" -MINDGRAPH_LIVE_ENVELOPE = ( - ROOT / "30_projects" / "mindgraph-eval" / "scripts" / "live_envelope_probe.py" -) -MINDGRAPH_GRAPH_HEALTH = ( - ROOT / "30_projects" / "mindgraph-eval" / "scripts" / "graph_health_trend.py" -) -MINDGRAPH_PROJECT = ROOT / "mindgraph" -DEFAULT_DB = Path.home() / ".mindgraph" / "mainframe.sqlite" -DEFAULT_INTENT_DB = Path.home() / ".mindgraph" / "mainframe-intent.sqlite" -EVAL_SNAPSHOT_DIR = ROOT / "30_projects" / "mindgraph-eval" / "snapshots" -EVAL_SNAPSHOT_PATHS = ( - EVAL_SNAPSHOT_DIR / "frozen-mainframe.sqlite", - EVAL_SNAPSHOT_DIR / "frozen-mainframe-projects.sqlite", - EVAL_SNAPSHOT_DIR / "frozen-mainframe-intent.sqlite", -) -LOG_DIR = REGISTRY_DIR / "logs" -# launchd stderr/stdout outside Desktop/Documents — those trees need Full Disk Access. -LAUNCHD_LOG_DIR = Path.home() / "Library" / "Logs" / "MainFrame" / "eval-schedule" -DAEMON_LABEL_DAILY = "com.mainframe.eval-schedule.daily" -DAEMON_LABEL_WEEKLY = "com.mainframe.eval-schedule.weekly" -SCHEDULED_PROBE_QUERY_IDS = ( - "q04_memory_cite_forget,q07_ai_detection,q11_negative_live_state,q12_negative_inbox" -) -# macOS TCC blocks background agents from Desktop/Documents/Downloads without FDA. -_TCC_PROTECTED_SEGMENTS = frozenset({"Desktop", "Documents", "Downloads"}) -_TCC_ERR_MARKERS = ( - "Operation not permitted", - "getcwd: cannot access parent directories", -) -STEP_PROCESS_IDS = { - "ingest_minion_dry_run": ["cli-ingest-minion", "workflow-ingest-minion"], - "mindgraph_refresh_dry_run": ["cli-mindgraph-refresh", "workflow-mindgraph-refresh"], - "sync_project_index_check": ["cli-sync-project-index"], - "eval_registry_status": ["cli-eval-registry"], - "unittest_suite": ["workflow-process-evaluation"], - "workflow_report_7d": ["cli-workflow-report", "workflow-workflow-telemetry"], - "mindgraph_retrieval_probe": ["workflow-mindgraph-refresh"], - "mindgraph_live_envelope": ["workflow-mindgraph-refresh"], - "mindgraph_graph_health": ["workflow-mindgraph-refresh"], - "eval_registry_harvest": ["cli-eval-registry"], -} -WEEKLY_STALE_DAYS = 8 -DAILY_STALE_DAYS = 2 -OPERATOR_CARD = REGISTRY_DIR / "OPERATOR.md" - - -@dataclass -class StepResult: - name: str - command: list[str] - exit_code: int - duration_seconds: float - process_ids: list[str] = field(default_factory=list) - stdout_tail: str = "" - stderr_tail: str = "" - skipped: bool = False - skip_reason: str | None = None - operator_gated: bool = False - - -@dataclass -class ScheduleRun: - run_id: str - cadence: str - started_at: str - finished_at: str - git_sha: str | None - all_passed: bool - steps: list[StepResult] = field(default_factory=list) - # Unit 2.4: provenance for control honesty (launchd vs manual) - trigger: str = "manual" - - -@dataclass -class ScheduleHealth: - launchd_daily: bool - launchd_weekly: bool - last_weekly: dict[str, Any] | None - last_daily: dict[str, Any] | None - weekly_age_days: float | None - daily_age_days: float | None - problems: list[str] = field(default_factory=list) - degraded: list[str] = field(default_factory=list) - - @property - def ok(self) -> bool: - return not self.problems and not self.degraded - - @property - def status_label(self) -> str: - if self.problems: - return "FAIL" - if self.degraded: - return "DEGRADED" - return "OK" - - -def repo_root() -> Path: - return ROOT - - -def git_short_sha(root: Path) -> str | None: - try: - res = subprocess.run( - ["git", "-C", str(root), "rev-parse", "--short", "HEAD"], - capture_output=True, - text=True, - timeout=10, - ) - if res.returncode == 0: - return res.stdout.strip() or None - except (OSError, subprocess.TimeoutExpired): - return None - return None - - -def tail_text(text: str, max_lines: int = 12) -> str: - lines = [ln for ln in text.splitlines() if ln.strip()] - if len(lines) <= max_lines: - return "\n".join(lines) - return "\n".join(lines[-max_lines:]) - - -def run_step(name: str, command: list[str], *, cwd: Path, timeout: int | None = None) -> StepResult: - started = time.monotonic() - try: - res = subprocess.run( - command, - cwd=cwd, - capture_output=True, - text=True, - timeout=timeout, - ) - duration = time.monotonic() - started - return StepResult( - name=name, - command=command, - exit_code=res.returncode, - duration_seconds=round(duration, 3), - process_ids=STEP_PROCESS_IDS.get(name, []), - stdout_tail=tail_text(res.stdout), - stderr_tail=tail_text(res.stderr), - ) - except subprocess.TimeoutExpired as exc: - duration = time.monotonic() - started - stdout = exc.stdout.decode("utf-8", errors="replace") if isinstance(exc.stdout, bytes) else (exc.stdout or "") - stderr = exc.stderr.decode("utf-8", errors="replace") if isinstance(exc.stderr, bytes) else (exc.stderr or "") - return StepResult( - name=name, - command=command, - exit_code=124, - duration_seconds=round(duration, 3), - process_ids=STEP_PROCESS_IDS.get(name, []), - stdout_tail=tail_text(stdout), - stderr_tail=tail_text(stderr or "timeout"), - ) - except OSError as exc: - duration = time.monotonic() - started - return StepResult( - name=name, - command=command, - exit_code=127, - duration_seconds=round(duration, 3), - process_ids=STEP_PROCESS_IDS.get(name, []), - stderr_tail=str(exc), - ) - - -def skipped_step(name: str, reason: str, *, operator_gated: bool = False) -> StepResult: - return StepResult( - name=name, - command=[], - exit_code=0, - duration_seconds=0.0, - process_ids=STEP_PROCESS_IDS.get(name, []), - skipped=True, - skip_reason=reason, - operator_gated=operator_gated, - ) - - -def missing_eval_snapshots() -> list[Path]: - """Return frozen evaluation inputs absent from the local snapshot set. - - Refreshing a snapshot deliberately rebases evaluation evidence, so a - missing snapshot remains an explicit operator gate rather than falling - back to the installed indexes. - """ - return [path for path in EVAL_SNAPSHOT_PATHS if not path.is_file()] - - -def parse_harvest_stats(stdout: str) -> dict[str, int]: - stats = {"errors": 0, "runs_new": 0, "skipped": 0} - match = re.search( - r"runs_new=(\d+).*?skipped=(\d+).*?errors=(\d+)", - stdout.replace("\n", " "), - ) - if match: - stats["runs_new"] = int(match.group(1)) - stats["skipped"] = int(match.group(2)) - stats["errors"] = int(match.group(3)) - return stats - - -def make_run_id(cadence: str) -> str: - stamp = datetime.now().strftime("%Y-%m-%dT%H%M%S") - return f"{stamp}-scheduled-{cadence}" - - -def mindgraph_uv_python_cmd(script: Path, *script_args: str) -> list[str] | None: - """Run a MindGraph-eval script with the mindgraph project's deps (PyYAML, etc.). - - Operator card and live_envelope_probe docs use - ``uv run --project mindgraph python …``. Bare ``sys.executable`` lacks - PyYAML on the MainFrame host Python and fails the weekly canary in <1s. - """ - uv = shutil.which("uv") - if not uv or not MINDGRAPH_PROJECT.is_dir(): - return None - return [ - uv, - "run", - "--project", - str(MINDGRAPH_PROJECT), - "python", - str(script), - *script_args, - ] - - -def steps_for_cadence( - cadence: str, - *, - skip_tests: bool, - skip_probe: bool, - full_probe: bool, - run_id: str, -) -> list[tuple[str, list[str] | None, str | None]]: - """Return (name, command-or-None, skip_reason) tuples.""" - bin_dir = repo_root() / "bin" - py = sys.executable - root = repo_root() - - out: list[tuple[str, list[str] | None, str | None]] = [] - - def add(name: str, cmd: list[str]) -> None: - out.append((name, cmd, None)) - - add("ingest_minion_dry_run", [str(bin_dir / "ingest-minion"), "run", "--dry-run"]) - add("mindgraph_refresh_dry_run", [str(bin_dir / "mindgraph-refresh"), "--dry-run"]) - add("sync_project_index_check", [str(bin_dir / "sync-project-index"), "--check"]) - add("eval_registry_status", [str(bin_dir / "eval-registry"), "status"]) - - if cadence in {"weekly", "monthly"} and not skip_tests: - add("unittest_suite", [py, "-m", "unittest", "discover", "-s", "tests"]) - - if cadence == "monthly": - add("workflow_report_7d", [str(bin_dir / "workflow-report"), "--days", "7", "--json"]) - - if cadence in {"weekly", "monthly"} and not skip_probe: - if MINDGRAPH_PROBE.exists() and DEFAULT_DB.exists(): - probe_run_id = f"{run_id}-probe" - probe_cmd = [ - py, - str(MINDGRAPH_PROBE), - "--run-id", - probe_run_id, - "--registry", - ] - if not full_probe: - probe_cmd.extend( - ["--fused-only", "--query-ids", SCHEDULED_PROBE_QUERY_IDS] - ) - add("mindgraph_retrieval_probe", probe_cmd) - else: - reason = "probe script or MindGraph DB missing" - out.append(("mindgraph_retrieval_probe", None, reason)) - - # Weekly envelope check reads frozen evaluation inputs only. Refreshing - # those snapshots is an operator-gated rebase, never a live fallback. - missing_snapshots = missing_eval_snapshots() - if MINDGRAPH_LIVE_ENVELOPE.exists() and not missing_snapshots: - env_run_id = f"{run_id}-live-envelope" - live_cmd = mindgraph_uv_python_cmd( - MINDGRAPH_LIVE_ENVELOPE, "--run-id", env_run_id - ) - if live_cmd is not None: - add("mindgraph_live_envelope", live_cmd) - else: - out.append( - ( - "mindgraph_live_envelope", - None, - "uv or mindgraph/ project missing (live envelope needs PyYAML)", - ) - ) - else: - missing_names = ", ".join(path.name for path in missing_snapshots) - if not missing_names: - missing_names = "live envelope script" - out.append( - ( - "mindgraph_live_envelope", - None, - "OPERATOR GATE: frozen evaluation snapshot missing " - f"({missing_names}); approve refresh_eval_snapshots.py before rerunning", - ) - ) - - if MINDGRAPH_GRAPH_HEALTH.exists(): - add("mindgraph_graph_health", [py, str(MINDGRAPH_GRAPH_HEALTH)]) - else: - out.append(("mindgraph_graph_health", None, "graph health script missing")) - - return out - - -def resolve_trigger(explicit: str | None = None) -> str: - """Record how a suite was started (Unit 2.4 provenance).""" - if explicit in {"launchd", "manual", "unknown"}: - return explicit - env = (os.environ.get("MAINFRAME_EVAL_TRIGGER") or "").strip().lower() - if env in {"launchd", "manual", "unknown"}: - return env - return "manual" - - -def execute_suite( - cadence: str, - *, - skip_tests: bool = False, - skip_probe: bool = False, - full_probe: bool = False, - trigger: str | None = None, -) -> ScheduleRun: - root = repo_root() - started_at = datetime.now(timezone.utc).isoformat() - run_id = make_run_id(cadence) - steps: list[StepResult] = [] - - for name, command, skip_reason in steps_for_cadence( - cadence, - skip_tests=skip_tests, - skip_probe=skip_probe, - full_probe=full_probe, - run_id=run_id, - ): - if command is None: - reason = skip_reason or "skipped" - steps.append( - skipped_step( - name, - reason, - operator_gated=reason.startswith("OPERATOR GATE:"), - ) - ) - continue - if name == "unittest_suite": - timeout = 1800 - elif name in {"mindgraph_retrieval_probe", "mindgraph_live_envelope"}: - timeout = 1200 - else: - timeout = 600 - steps.append(run_step(name, command, cwd=root, timeout=timeout)) - - finished_at = datetime.now(timezone.utc).isoformat() - all_passed = all( - s.exit_code == 0 and not s.operator_gated - for s in steps - if not s.skipped or s.operator_gated - ) - return ScheduleRun( - run_id=run_id, - cadence=cadence, - started_at=started_at, - finished_at=finished_at, - git_sha=git_short_sha(root), - all_passed=all_passed, - steps=steps, - trigger=resolve_trigger(trigger), - ) - - -def append_schedule_log(run: ScheduleRun, *, dry_run: bool) -> None: - if dry_run: - return - REGISTRY_DIR.mkdir(parents=True, exist_ok=True) - payload = { - "run_id": run.run_id, - "cadence": run.cadence, - "started_at": run.started_at, - "finished_at": run.finished_at, - "git_sha": run.git_sha, - "all_passed": run.all_passed, - "trigger": run.trigger, - "steps": [asdict(s) for s in run.steps], - } - with SCHEDULE_LOG.open("a", encoding="utf-8") as fh: - fh.write(json.dumps(payload, ensure_ascii=False) + "\n") - - -def write_process_eval_output(run: ScheduleRun, *, dry_run: bool) -> Path | None: - if dry_run or run.cadence == "daily": - return None - - PROCESS_EVAL_OUTPUTS.mkdir(parents=True, exist_ok=True) - out_path = PROCESS_EVAL_OUTPUTS / f"{run.run_id}.md" - # Portfolio triage may exit 2 when severity is high but the card was written; - # that is allowed for all_passed and must not be labeled a suite failure. - def _step_failed(s: StepResult) -> bool: - if s.skipped or s.operator_gated: - return False - if s.name == "eval_portfolio_triage" and s.exit_code == 2: - return False - return s.exit_code != 0 - - failed = [s.name for s in run.steps if _step_failed(s)] - triage_sev = [ - s.name - for s in run.steps - if s.name == "eval_portfolio_triage" and s.exit_code == 2 and not s.skipped - ] - skipped = [s.name for s in run.steps if s.skipped and not s.operator_gated] - gated = [s.name for s in run.steps if s.operator_gated] - passed_count = sum( - 1 - for s in run.steps - if not s.skipped - and (s.exit_code == 0 or (s.name == "eval_portfolio_triage" and s.exit_code == 2)) - ) - step_total = sum(1 for s in run.steps if not s.skipped or s.operator_gated) - - answer = ( - f"Scheduled {run.cadence} suite {'passed' if run.all_passed else 'failed'}: " - f"{passed_count}/{step_total} steps green." - ) - if failed: - answer += f" Failed: {', '.join(failed)}." - if triage_sev: - answer += " Triage exit 2 (severity signal only; card written)." - if skipped: - answer += f" Skipped: {', '.join(skipped)}." - if gated: - answer += f" Operator-gated: {', '.join(gated)}." - - metrics = [ - { - "name": "scheduled_steps_passed", - "slice": run.cadence, - "value": passed_count, - "n": step_total, - "unit": "count", - }, - { - "name": "scheduled_all_passed", - "slice": run.cadence, - "value": 1 if run.all_passed else 0, - "n": 1, - "unit": "binary", - }, - ] - - irregularities: list[dict[str, Any]] = [] - for step in run.steps: - if step.operator_gated: - irregularities.append( - { - "id": f"operator-gated-{step.name}", - "severity": "medium", - "category": "operator_gate", - "observation": step.skip_reason or "operator decision required", - "context": run.run_id, - "resolved": False, - } - ) - elif step.skipped: - irregularities.append( - { - "id": f"skipped-{step.name}", - "severity": "info", - "category": "intentional_skip", - "observation": step.skip_reason or "step skipped", - "context": run.run_id, - "resolved": True, - } - ) - elif step.exit_code != 0: - if step.name == "eval_portfolio_triage" and step.exit_code == 2: - irregularities.append( - { - "id": f"triage-severity-{step.name}", - "severity": "info", - "category": "triage_severity", - "observation": ( - f"{step.name} exit_code=2 (portfolio severity high; " - "action card written; not a suite failure)" - ), - "context": step.stderr_tail or step.stdout_tail or "", - "resolved": True, - } - ) - else: - irregularities.append( - { - "id": f"failed-{step.name}", - "severity": "medium" if run.cadence == "weekly" else "high", - "category": "step_failure", - "observation": f"{step.name} exit_code={step.exit_code}", - "context": step.stderr_tail or step.stdout_tail or "", - "resolved": False, - } - ) - elif step.name == "eval_registry_harvest": - stats = parse_harvest_stats(step.stdout_tail) - if stats["errors"]: - irregularities.append( - { - "id": "harvest-parse-errors", - "severity": "low", - "category": "registry_hygiene", - "observation": f"eval-registry harvest reported errors={stats['errors']}", - "context": "outputs missing metric extract or non-harvestable artifacts", - "resolved": False, - } - ) - - yaml_metrics = "\n".join( - f" - name: {m['name']}\n slice: {m['slice']}\n value: {m['value']}\n n: {m['n']}\n unit: {m['unit']}" - for m in metrics - ) - if irregularities: - yaml_irreg = "\n".join( - " - id: {id}\n severity: {severity}\n category: {category}\n observation: \"{observation}\"\n context: \"{context}\"\n resolved: {resolved}".format( - id=i["id"], - severity=i["severity"], - category=i["category"], - observation=i["observation"].replace('"', "'"), - context=(i.get("context") or "").replace('"', "'")[:200], - resolved=str(i["resolved"]).lower(), - ) - for i in irregularities - ) - else: - yaml_irreg = " []" - - step_rows = "\n".join( - f"| {s.name} | {'gate' if s.operator_gated else ('skip' if s.skipped else s.exit_code)} | {s.duration_seconds} | {', '.join(s.process_ids) or 'none'} |" - for s in run.steps - ) - - body = f"""--- -title: "Scheduled process evaluation — {run.cadence} — {date.today().isoformat()}" -domain: "knowledge-systems" -type: "project" -status: "active" -study_type: "observational" -eval_run_id: "{run.run_id}" -protocol_ref: ".context/workflows/eval-schedule.md" -decision_sentence: "If scheduled steps fail twice in a row, fix the failing bin before changing agent workflows." -project: "mainframe-process-eval" -tags: ["evaluation", "eval-registry", "eval-profile", "scheduled"] -updated: "{date.today().isoformat()}" -source: "bin/eval-schedule" ---- - -# Scheduled process evaluation — {run.cadence} — {date.today().isoformat()} - -## Answer - -{answer} - -## Checks run - -| Step | Exit | Duration (s) | Process IDs | -| --- | ---: | ---: | --- | -{step_rows} - -## Metric extract (eval-registry) - -```yaml -registry: - project: mainframe-process-eval - run_id: {run.run_id} - study_type: observational - protocol_ref: .context/workflows/eval-schedule.md - date: {date.today().isoformat()} - decision_sentence: "If scheduled steps fail twice in a row, fix the failing bin before changing agent workflows." - artifact_path: outputs/{run.run_id}.md - raw_path: null - environment: - git_sha: {run.git_sha or "null"} - cadence: {run.cadence} - decision_use: regression_only -metrics: -{yaml_metrics} -irregularities: -{yaml_irreg} -``` -""" - out_path.write_text(body, encoding="utf-8") - return out_path - - -def print_run_summary(run: ScheduleRun) -> None: - print( - f"eval-schedule: cadence={run.cadence} run_id={run.run_id} " - f"all_passed={run.all_passed} trigger={run.trigger}" - ) - for step in run.steps: - if step.operator_gated: - print(f" - {step.name}: OPERATOR GATE ({step.skip_reason})") - elif step.skipped: - print(f" - {step.name}: SKIP ({step.skip_reason})") - else: - print(f" - {step.name}: exit={step.exit_code} duration={step.duration_seconds}s") - - -def plist_path(label: str) -> Path: - return Path.home() / "Library" / "LaunchAgents" / f"{label}.plist" - - -def python_for_launchd() -> str: - """Absolute interpreter for LaunchAgents (avoid /usr/bin/env + bash wrappers).""" - return sys.executable - - -def root_needs_full_disk_access(root: Path) -> bool: - """True when root sits under a macOS TCC-protected user folder.""" - try: - parts = set(root.expanduser().resolve().parts) - except OSError: - parts = set(root.parts) - return bool(parts & _TCC_PROTECTED_SEGMENTS) - - -def detect_launchd_tcc_blocks(*, log_dirs: list[Path] | None = None) -> list[str]: - """Scan launchd err logs for Desktop/TCC denials. - - Default: only ~/Library/Logs/MainFrame/eval-schedule (post-2026-07-23 plists). - Legacy Desktop-tree logs are ignored so historical denials do not forever-red - the check after the agent is fixed. - """ - dirs = log_dirs if log_dirs is not None else [LAUNCHD_LOG_DIR] - hits: list[str] = [] - for directory in dirs: - if not directory.is_dir(): - continue - for path in sorted(directory.glob("*.err")): - try: - text = path.read_text(encoding="utf-8", errors="replace") - except OSError: - continue - # Prefer the tail — recent failures matter more than ancient noise. - tail = "\n".join(text.splitlines()[-40:]) - if any(marker in tail for marker in _TCC_ERR_MARKERS): - hits.append(str(path)) - return hits - - -def build_plist(label: str, cadence: str, root: Path) -> str: - """Build a LaunchAgent plist that minimizes Desktop TCC friction. - - - Invoke python + scripts/eval_schedule.py directly (no bash wrapper). - - WorkingDirectory = $HOME (not the Desktop tree). - - Logs under ~/Library/Logs/MainFrame/eval-schedule/. - - Still requires Full Disk Access when root is under Desktop. - """ - LAUNCHD_LOG_DIR.mkdir(parents=True, exist_ok=True) - log_out = LAUNCHD_LOG_DIR / f"{cadence}.log" - log_err = LAUNCHD_LOG_DIR / f"{cadence}.err" - python = python_for_launchd() - module = (root / "scripts" / "eval_schedule.py").resolve() - home = str(Path.home()) - if cadence == "daily": - calendar = """ - <key>StartCalendarInterval</key> - <dict> - <key>Hour</key> - <integer>6</integer> - <key>Minute</key> - <integer>15</integer> - </dict>""" - else: - calendar = """ - <key>StartCalendarInterval</key> - <dict> - <key>Weekday</key> - <integer>0</integer> - <key>Hour</key> - <integer>7</integer> - <key>Minute</key> - <integer>30</integer> - </dict>""" - - return f"""<?xml version="1.0" encoding="UTF-8"?> -<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd"> -<plist version="1.0"> -<dict> - <key>Label</key> - <string>{label}</string> - <key>ProgramArguments</key> - <array> - <string>{python}</string> - <string>{module}</string> - <string>run</string> - <string>--cadence</string> - <string>{cadence}</string> - </array> - <key>WorkingDirectory</key> - <string>{home}</string>{calendar} - <key>StandardOutPath</key> - <string>{log_out}</string> - <key>StandardErrorPath</key> - <string>{log_err}</string> - <key>EnvironmentVariables</key> - <dict> - <key>PATH</key> - <string>{home}/.local/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin</string> - <key>MAINFRAME_EVAL_TRIGGER</key> - <string>launchd</string> - <key>MAINFRAME_ROOT</key> - <string>{root.resolve()}</string> - <key>PYTHONPATH</key> - <string>{(root / "scripts").resolve()}</string> - <key>HOME</key> - <string>{home}</string> - </dict> -</dict> -</plist> -""" - - -def launchctl_load(plist: Path) -> tuple[int, str]: - subprocess.run(["launchctl", "unload", str(plist)], capture_output=True) - res = subprocess.run(["launchctl", "load", "-w", str(plist)], capture_output=True, text=True) - return res.returncode, (res.stderr or res.stdout or "").strip() - - -def cmd_run(args: argparse.Namespace) -> int: - run = execute_suite( - args.cadence, - skip_tests=args.skip_tests, - skip_probe=args.skip_probe, - full_probe=args.full_probe, - trigger=getattr(args, "trigger", None), - ) - out = write_process_eval_output(run, dry_run=args.dry_run) - if out: - print(f"process_eval_output: {out.relative_to(ROOT)}") - - if not args.dry_run and args.cadence in {"weekly", "monthly"}: - harvest = run_step( - "eval_registry_harvest", - [str(ROOT / "bin" / "eval-registry"), "harvest"], - cwd=ROOT, - ) - run.steps.append(harvest) - stats = parse_harvest_stats(harvest.stdout_tail) - if harvest.exit_code != 0: - run.all_passed = False - elif stats["errors"]: - print( - f" - harvest warning: errors={stats['errors']} " - f"(non-fatal — fix outputs missing metric extract)" - ) - - # Action layer: suite already ran + harvested; portfolio triage so all - # MainFrame evals are actionable (project-experiment-loop). - triage_bin = ROOT / "bin" / "project-experiment-loop" - if triage_bin.exists(): - triage = run_step( - "eval_portfolio_triage", - [str(triage_bin), "triage", "--scope", "portfolio"], - cwd=ROOT, - timeout=180, - ) - run.steps.append(triage) - if triage.exit_code not in (0, 2): - print(f" - triage warning: exit={triage.exit_code}") - else: - print( - " - eval action card: " - f"{(REGISTRY_DIR / 'last-eval-action.md').relative_to(ROOT)}" - ) - - if out and not args.dry_run and run.cadence in {"weekly", "monthly"}: - out = write_process_eval_output(run, dry_run=False) - - print_run_summary(run) - append_schedule_log(run, dry_run=args.dry_run) - return 0 if run.all_passed else 1 - - -def cmd_install(args: argparse.Namespace) -> int: - root = repo_root() - LOG_DIR.mkdir(parents=True, exist_ok=True) - LAUNCHD_LOG_DIR.mkdir(parents=True, exist_ok=True) - REGISTRY_DIR.mkdir(parents=True, exist_ok=True) - ensure_operator_card() - plist_dir = Path.home() / "Library" / "LaunchAgents" - plist_dir.mkdir(parents=True, exist_ok=True) - - targets: list[tuple[str, str]] = [] - if args.cadence in {"daily", "both"}: - targets.append((DAEMON_LABEL_DAILY, "daily")) - if args.cadence in {"weekly", "both"}: - targets.append((DAEMON_LABEL_WEEKLY, "weekly")) - - for label, cadence in targets: - path = plist_path(label) - path.write_text(build_plist(label, cadence, root), encoding="utf-8") - code, msg = launchctl_load(path) - if code == 0: - print(f"installed: {label} -> {path}") - else: - print(f"install warning ({label}): {msg}", file=sys.stderr) - return 1 - - py = python_for_launchd() - print(f"launchd python: {py}") - print(f"launchd logs: {LAUNCHD_LOG_DIR}") - if root_needs_full_disk_access(root): - print( - "\nNote: MainFrame is under Desktop/Documents/Downloads (TCC-protected).\n" - " Agents invoke Python directly (not bash). If kickstart fails with\n" - f" Operation not permitted, grant Full Disk Access to:\n • {py}\n" - " Prove: launchctl kickstart -k gui/$(id -u)/com.mainframe.eval-schedule.weekly\n" - " Longer-term option: keep the repo outside Desktop (e.g. ~/MainFrame).\n", - file=sys.stderr, - ) - tcc_hits = detect_launchd_tcc_blocks() - if tcc_hits: - print( - "Recent launchd TCC denials still present in:\n - " - + "\n - ".join(tcc_hits[:4]), - file=sys.stderr, - ) - print( - "Prove service: launchctl kickstart -k gui/$(id -u)/com.mainframe.eval-schedule.weekly" - ) - return 0 - - -def cmd_uninstall(args: argparse.Namespace) -> int: - labels: list[str] = [] - if args.cadence in {"daily", "both"}: - labels.append(DAEMON_LABEL_DAILY) - if args.cadence in {"weekly", "both"}: - labels.append(DAEMON_LABEL_WEEKLY) - - for label in labels: - path = plist_path(label) - if path.exists(): - subprocess.run(["launchctl", "unload", str(path)], capture_output=True) - path.unlink() - print(f"uninstalled: {label}") - else: - print(f"not installed: {label}") - return 0 - - -def _parse_iso(ts: str | None) -> datetime | None: - if not ts: - return None - try: - return datetime.fromisoformat(ts.replace("Z", "+00:00")) - except ValueError: - return None - - -def load_schedule_runs() -> list[dict[str, Any]]: - if not SCHEDULE_LOG.exists(): - return [] - runs: list[dict[str, Any]] = [] - for line in SCHEDULE_LOG.read_text(encoding="utf-8").splitlines(): - if line.strip(): - runs.append(json.loads(line)) - return runs - - -def last_run_for_cadence(runs: list[dict[str, Any]], cadence: str) -> dict[str, Any] | None: - matches = [r for r in runs if r.get("cadence") == cadence] - if not matches: - return None - return max(matches, key=lambda r: r.get("finished_at") or r.get("started_at") or "") - - -def failed_step_summary(run: dict[str, Any] | None) -> str: - """Summarize failed or operator-gated child steps from a schedule receipt.""" - if not run: - return "" - labels: list[str] = [] - for step in run.get("steps") or []: - name = step.get("name") or "unnamed" - if step.get("operator_gated"): - labels.append(f"{name} (operator gate)") - elif not step.get("skipped") and step.get("exit_code") not in (0, None): - labels.append(f"{name} (exit {step.get('exit_code')})") - return ", ".join(labels) - - -def assess_schedule_health(*, now: datetime | None = None) -> ScheduleHealth: - now = now or datetime.now(timezone.utc) - runs = load_schedule_runs() - last_weekly = last_run_for_cadence(runs, "weekly") - last_daily = last_run_for_cadence(runs, "daily") - - weekly_finished = _parse_iso(last_weekly.get("finished_at") if last_weekly else None) - daily_finished = _parse_iso(last_daily.get("finished_at") if last_daily else None) - weekly_age = ( - (now - weekly_finished).total_seconds() / 86400 if weekly_finished else None - ) - daily_age = (now - daily_finished).total_seconds() / 86400 if daily_finished else None - - launchd_daily = plist_path(DAEMON_LABEL_DAILY).exists() - launchd_weekly = plist_path(DAEMON_LABEL_WEEKLY).exists() - problems: list[str] = [] - degraded: list[str] = [] - - if not launchd_weekly: - problems.append("launchd weekly agent not installed — run: bin/eval-schedule install --cadence both") - if not launchd_daily: - problems.append("launchd daily agent not installed — run: bin/eval-schedule install --cadence both") - - # TCC / Full Disk Access — hard-fail only on concrete denial evidence in - # current launchd logs. A recent launchd-triggered weekly proves the agent - # can read the tree (bash-wrapper denials are historical). - tcc_logs = detect_launchd_tcc_blocks() - if tcc_logs: - problems.append( - "launchd blocked by macOS TCC (Operation not permitted on Desktop path) — " - f"grant Full Disk Access to {python_for_launchd()}, then: " - "bin/eval-schedule install --cadence both && " - "launchctl kickstart -k gui/$(id -u)/com.mainframe.eval-schedule.weekly " - f"(see {tcc_logs[0]})" - ) - elif ( - root_needs_full_disk_access(ROOT) - and (launchd_weekly or launchd_daily) - and not ( - last_weekly - and last_weekly.get("trigger") == "launchd" - and weekly_age is not None - and weekly_age <= WEEKLY_STALE_DAYS - ) - ): - degraded.append( - "MainFrame is under Desktop/Documents/Downloads; if launchd fails with " - f"Operation not permitted, grant Full Disk Access to {python_for_launchd()}" - ) - - if last_weekly is None: - problems.append("no weekly eval run logged — run: bin/eval-schedule run --cadence weekly") - elif weekly_age is not None and weekly_age > WEEKLY_STALE_DAYS: - problems.append( - f"weekly eval stale ({weekly_age:.1f}d) — run: bin/eval-schedule run --cadence weekly" - ) - elif last_weekly and not last_weekly.get("all_passed"): - details = failed_step_summary(last_weekly) - suffix = f"; failed steps: {details}" if details else "" - problems.append( - f"last weekly eval failed ({last_weekly.get('run_id')}){suffix} — inspect 20_live/eval-registry/logs/" - ) - if daily_age is not None and daily_age > DAILY_STALE_DAYS: - problems.append(f"daily eval stale ({daily_age:.1f}d) — check launchd logs") - - # Unit 2.4: recency alone is insufficient — require launchd provenance - if last_weekly and last_weekly.get("all_passed") and ( - weekly_age is None or weekly_age <= WEEKLY_STALE_DAYS - ): - trigger = last_weekly.get("trigger") - if trigger != "launchd": - degraded.append( - "last successful weekly run lacks launchd provenance " - f"(trigger={trigger!r}); manual/unknown runs cannot prove the service " - "(Unit 2.4 control-honesty)" - ) - - return ScheduleHealth( - launchd_daily=launchd_daily, - launchd_weekly=launchd_weekly, - last_weekly=last_weekly, - last_daily=last_daily, - weekly_age_days=weekly_age, - daily_age_days=daily_age, - problems=problems, - degraded=degraded, - ) - - -def ensure_operator_card() -> None: - REGISTRY_DIR.mkdir(parents=True, exist_ok=True) - if OPERATOR_CARD.exists(): - return - OPERATOR_CARD.write_text( - """# Eval registry — operator card - -**Do not let scheduled evals fade.** This surface is the live reminder. - -## Weekly ritual (~15–20 min, after Sunday 07:30 run) - -1. `bin/eval-schedule check` — must exit 0 -2. `bin/eval-schedule status` — last weekly green? -3. Read latest `30_projects/mainframe-process-eval/outputs/*-scheduled-weekly.md` -4. Read MindGraph canaries: retrieval probe, live-envelope, graph-health-trend under `30_projects/mindgraph-eval/outputs/` -5. `bin/eval-registry status` — new metrics or irregularities? -6. Pick **one** improvement slice; rerun `bin/eval-schedule run --cadence weekly` after the fix - -### MindGraph canaries (weekly) - -- `mindgraph_retrieval_probe` — q04/q07/q11/q12 fused -- `mindgraph_live_envelope` — installed intent graph (also **required after every intent install**) -- `mindgraph_graph_health` — density trend - -Post-intent-install: run `live_envelope_probe.py` until exit 0 before calling the install done. - -## Commands - -```bash -bin/eval-schedule check -bin/eval-schedule status -bin/eval-schedule run --cadence weekly -bin/eval-schedule run --cadence weekly --full-probe -bin/eval-schedule install --cadence both # macOS launchd -``` - -## Surfaces that nag you - -- `bin/session-close --check` — warns when eval schedule is unhealthy -- `bin/session-open` — prints eval health on session start -- Lane **EV01** — `30_projects/research-lanes-strategy/lanes/scheduled-process-evaluation/` -- Workflow — `.context/workflows/eval-schedule.md` - -## Logs - -- Manifest: `20_live/eval-registry/schedule-runs.jsonl` -- launchd: `20_live/eval-registry/logs/{daily,weekly}.{log,err}` -""", - encoding="utf-8", - ) - - -def cmd_status(args: argparse.Namespace) -> int: - ensure_operator_card() - health = assess_schedule_health() - runs = load_schedule_runs() - print(f"schedule_log: {SCHEDULE_LOG}") - print(f"log_dir: {LOG_DIR}") - print(f"operator_card: {OPERATOR_CARD}") - print(f"schedule_runs: {len(runs)}") - if health.last_weekly: - weekly_age = ( - f"{health.weekly_age_days:.1f}" - if health.weekly_age_days is not None - else "n/a" - ) - print( - f"last_weekly: {health.last_weekly.get('run_id')} " - f"all_passed={health.last_weekly.get('all_passed')} age_days={weekly_age}" - ) - else: - print("last_weekly: none") - if health.last_daily: - daily_age = ( - f"{health.daily_age_days:.1f}" if health.daily_age_days is not None else "n/a" - ) - print( - f"last_daily: {health.last_daily.get('run_id')} " - f"all_passed={health.last_daily.get('all_passed')} age_days={daily_age}" - ) - else: - print("last_daily: none") - for label in (DAEMON_LABEL_DAILY, DAEMON_LABEL_WEEKLY): - path = plist_path(label) - print(f"launchd {label}: {'installed' if path.exists() else 'not installed'} ({path})") - if health.problems: - print("problems:") - for problem in health.problems: - print(f" - {problem}") - return 0 - - -def cmd_check(args: argparse.Namespace) -> int: - ensure_operator_card() - health = assess_schedule_health() - if health.ok: - print("eval-schedule check: OK") - if health.last_weekly and health.weekly_age_days is not None: - trigger = health.last_weekly.get("trigger") or "unknown" - print( - f" weekly: {health.last_weekly.get('run_id')} " - f"({health.weekly_age_days:.1f}d ago) trigger={trigger}" - ) - return 0 - label = health.status_label - n = len(health.problems) + len(health.degraded) - print(f"eval-schedule check: {label} ({n} issue(s))") - for problem in health.problems: - print(f" - [problem] {problem}") - for item in health.degraded: - print(f" - [degraded] {item}") - print(f"operator_card: {OPERATOR_CARD}") - return 1 - - -def main() -> int: - parser = argparse.ArgumentParser(description="MainFrame scheduled evaluation runner") - sub = parser.add_subparsers(dest="command", required=True) - - run = sub.add_parser("run", help="Execute a scheduled eval suite") - run.add_argument("--cadence", choices=["daily", "weekly", "monthly"], default="weekly") - run.add_argument("--dry-run", action="store_true", help="Run suite but do not write logs/outputs") - run.add_argument("--skip-tests", action="store_true") - run.add_argument("--skip-probe", action="store_true") - run.add_argument( - "--full-probe", - action="store_true", - help="Run full 12-query fused+expanded probe instead of scheduled 4-query fused regression", - ) - run.add_argument( - "--trigger", - choices=["launchd", "manual", "unknown"], - default=None, - help="Provenance label for this run (default: MAINFRAME_EVAL_TRIGGER or manual)", - ) - run.set_defaults(func=cmd_run) - - inst = sub.add_parser("install", help="Install launchd LaunchAgents (macOS)") - inst.add_argument("--cadence", choices=["daily", "weekly", "both"], default="both") - inst.set_defaults(func=cmd_install) - - uninst = sub.add_parser("uninstall", help="Remove launchd LaunchAgents") - uninst.add_argument("--cadence", choices=["daily", "weekly", "both"], default="both") - uninst.set_defaults(func=cmd_uninstall) - - stat = sub.add_parser("status", help="Show schedule log and launchd state") - stat.set_defaults(func=cmd_status) - - chk = sub.add_parser("check", help="Exit 1 if launchd missing or eval runs are stale/failed") - chk.set_defaults(func=cmd_check) - - args = parser.parse_args() - return args.func(args) - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/fetch_source_text.py b/scripts/fetch_source_text.py deleted file mode 100755 index dba3cd6..0000000 --- a/scripts/fetch_source_text.py +++ /dev/null @@ -1,832 +0,0 @@ -#!/usr/bin/env python3 -"""Fetch open-access full text (or best available excerpt) for source-literature stubs. - -Scans type: raw stubs under 10_knowledge/, resolves DOI/PMID/PDF/HTML sources, -and appends a ## Full text extract section following the business-operations audit -pattern. Does not paste paywalled PDFs. - -Usage: - bin/fetch-source-text --dry-run - bin/fetch-source-text --apply --subset healthcare-practice - bin/fetch-source-text --apply --file path/to/stub.md - bin/fetch-source-text --json --subset healthcare-practice -""" - -from __future__ import annotations - -import argparse -import fcntl -import json -import os -import re -import shutil -import stat -import subprocess -import sys -import tempfile -import textwrap -import urllib.error -import urllib.parse -import urllib.request -import xml.etree.ElementTree as ET -from dataclasses import asdict, dataclass, field -from datetime import date -from pathlib import Path -from typing import Any - -ROOT = Path(__file__).resolve().parents[1] -KNOWLEDGE = ROOT / "10_knowledge" -USER_AGENT = "MainFrame-fetch-source-text/1.0 (mailto:mainframe@local)" -EXCERPT_MAX = 6000 -HTTP_TIMEOUT = 45 -HTTP_MAX_BYTES = 25 * 1024 * 1024 -LOCK_RETRIES = 4 - - -class PathScopeError(ValueError): - """Raised when a requested stub is outside the durable knowledge tree.""" - - -class DownloadTooLargeError(urllib.error.URLError): - """Raised when a remote response exceeds the configured download bound.""" - - -class ConcurrentUpdateError(RuntimeError): - """Raised when a stub changes while an atomic update is being prepared.""" - - -@dataclass -class FetchResult: - path: str - status: str # fetched | pending | skipped | na | error - method: str = "" - access_url: str = "" - message: str = "" - excerpt_chars: int = 0 - section: str = "" - - -def parse_frontmatter(text: str) -> tuple[dict[str, str], str]: - if not text.startswith("---"): - return {}, text - end = text.find("\n---", 3) - if end == -1: - return {}, text - header = text[3:end] - body = text[end + 5 :] - fm: dict[str, str] = {} - for raw_line in header.splitlines(): - line = raw_line.strip() - if not line or line.startswith("#"): - continue - if ":" not in line: - continue - key, val = line.split(":", 1) - fm[key.strip().lower()] = val.strip().strip('"').strip("'") - return fm, body - - -def parse_tags(tag_field: str) -> list[str]: - if not tag_field: - return [] - tag_field = tag_field.strip() - if tag_field.startswith("["): - tag_field = tag_field.strip("[]").strip() - return [t.strip().strip('"').strip("'") for t in tag_field.split(",") if t.strip()] - - -def normalize_source_url(source: str) -> str: - """Pick the first fetchable URL from messy source fields.""" - source = source.strip().strip('"').strip("'") - if not source: - return "" - urls = re.findall(r"https?://[^\s\"'<>]+", source) - if urls: - return urls[0].rstrip(").,;") - if source.startswith("/"): - arxiv = re.search(r"/abs/([\d.]+)", source) - if arxiv: - return f"https://arxiv.org/abs/{arxiv.group(1)}" - if source.startswith("arxiv:"): - return f"https://arxiv.org/abs/{source.split(':', 1)[1].strip()}" - if source.startswith("10."): - return f"https://doi.org/{source}" - return source.split(";")[0].split()[0].strip() - - -def extract_identifiers(fm: dict[str, str]) -> tuple[str | None, str | None]: - doi = fm.get("doi") - pmid = fm.get("pmid") - source = normalize_source_url(fm.get("source", "")) - if not doi: - m = re.search(r"10\.\d{4,9}/[^\s\"'>]+", source) - if m: - doi = m.group(0).rstrip(").,;") - if not pmid: - m = re.search(r"pubmed\.ncbi\.nlm\.nih\.gov/(\d+)", source, re.I) - if m: - pmid = m.group(1) - return doi, pmid - - -def http_get( - url: str, - accept: str = "*/*", - *, - max_bytes: int = HTTP_MAX_BYTES, -) -> bytes: - url = normalize_source_url(url) if " " in url or ";" in url else url - if not url.startswith(("http://", "https://")): - raise urllib.error.URLError(f"unsupported URL: {url!r}") - if max_bytes < 1: - raise ValueError("max_bytes must be positive") - req = urllib.request.Request( - url, - headers={"User-Agent": USER_AGENT, "Accept": accept}, - ) - with urllib.request.urlopen(req, timeout=HTTP_TIMEOUT) as resp: - content_length = resp.headers.get("Content-Length") - if content_length: - try: - declared_bytes = int(content_length) - except (TypeError, ValueError): - declared_bytes = -1 - if declared_bytes > max_bytes: - raise DownloadTooLargeError( - f"response declares {declared_bytes} bytes; limit is {max_bytes}" - ) - data = resp.read(max_bytes + 1) - if len(data) > max_bytes: - raise DownloadTooLargeError( - f"response exceeded download limit of {max_bytes} bytes" - ) - return data - - -def http_get_json(url: str) -> Any: - raw = http_get(url, accept="application/json") - return json.loads(raw.decode("utf-8", errors="replace")) - - -def europepmc_search(doi: str | None, pmid: str | None) -> dict[str, Any] | None: - if pmid: - query = f"EXT_ID:{pmid}" - elif doi: - query = f"DOI:{doi}" - else: - return None - url = ( - "https://www.ebi.ac.uk/europepmc/webservices/rest/search?" - + urllib.parse.urlencode({"query": query, "format": "json", "pageSize": 1}) - ) - try: - data = http_get_json(url) - except (urllib.error.URLError, urllib.error.HTTPError, json.JSONDecodeError): - return None - results = data.get("resultList", {}).get("result") or [] - return results[0] if results else None - - -def openalex_abstract(doi: str) -> str: - url = f"https://api.openalex.org/works/https://doi.org/{urllib.parse.quote(doi)}" - try: - data = http_get_json(url) - except (urllib.error.URLError, urllib.error.HTTPError, json.JSONDecodeError): - return "" - inv = data.get("abstract_inverted_index") - if not inv: - return "" - positions: list[tuple[int, str]] = [] - for word, idxs in inv.items(): - for i in idxs: - positions.append((i, word)) - positions.sort() - return " ".join(w for _, w in positions) - - -def crossref_abstract(doi: str) -> str: - url = f"https://api.crossref.org/works/{urllib.parse.quote(doi)}" - try: - data = http_get_json(url) - except (urllib.error.URLError, urllib.error.HTTPError, json.JSONDecodeError): - return "" - msg = data.get("message", {}) - abstract = msg.get("abstract") or "" - abstract = re.sub(r"<[^>]+>", "", abstract) - return " ".join(abstract.split()) - - -def unpaywall_best_pdf(doi: str, email: str | None) -> str | None: - if not email: - return None - url = ( - f"https://api.unpaywall.org/v2/{urllib.parse.quote(doi)}" - f"?email={urllib.parse.quote(email)}" - ) - try: - data = http_get_json(url) - except (urllib.error.URLError, urllib.error.HTTPError, json.JSONDecodeError): - return None - if not data.get("is_oa"): - return None - loc = data.get("best_oa_location") or {} - return loc.get("url_for_pdf") or loc.get("url") - - -def jats_xml_to_text(xml_bytes: bytes) -> tuple[str, str]: - """Return (abstract, body_text) from JATS XML.""" - try: - root = ET.fromstring(xml_bytes) - except ET.ParseError: - return "", "" - - def collect(tag: str) -> str: - parts: list[str] = [] - for el in root.iter(tag): - text = "".join(el.itertext()).strip() - if text: - parts.append(text) - return "\n\n".join(parts) - - abstract = collect("abstract") - body = collect("body") - if not body: - # Fallback: all paragraphs outside front matter - paras = [] - for el in root.iter("p"): - text = "".join(el.itertext()).strip() - if text: - paras.append(text) - body = "\n\n".join(paras) - return abstract, body - - -def fetch_europepmc_fulltext(pmcid: str) -> tuple[str, str, str]: - """Returns method, access_url, combined_text.""" - access = f"https://europepmc.org/article/PMC/{pmcid.replace('PMC', '')}" - url = f"https://www.ebi.ac.uk/europepmc/webservices/rest/{pmcid}/fullTextXML" - try: - xml_bytes = http_get(url, accept="application/xml") - except (urllib.error.URLError, urllib.error.HTTPError): - return "europepmc-blocked", access, "" - abstract, body = jats_xml_to_text(xml_bytes) - chunks = [] - if abstract: - chunks.append(abstract) - if body: - chunks.append(body) - return "europepmc-xml", access, "\n\n".join(chunks) - - -def pdf_to_text(pdf_bytes: bytes) -> str: - pdftotext = shutil.which("pdftotext") - if not pdftotext: - return "" - with tempfile.NamedTemporaryFile(suffix=".pdf", delete=True) as pdf_tmp: - pdf_tmp.write(pdf_bytes) - pdf_tmp.flush() - with tempfile.NamedTemporaryFile(suffix=".txt", delete=True) as txt_tmp: - proc = subprocess.run( - [pdftotext, "-layout", pdf_tmp.name, txt_tmp.name], - capture_output=True, - text=True, - check=False, - ) - if proc.returncode != 0: - return "" - return Path(txt_tmp.name).read_text(encoding="utf-8", errors="replace") - - -def fetch_direct_pdf(url: str) -> tuple[str, str, str]: - try: - data = http_get(url) - except (urllib.error.URLError, urllib.error.HTTPError) as exc: - return "direct-pdf-blocked", url, f"HTTP error: {exc}" - if not data.startswith(b"%PDF"): - return "direct-pdf-blocked", url, "Response was not a PDF" - text = pdf_to_text(data) - if not text.strip(): - return "direct-pdf-empty", url, "PDF downloaded; text extraction failed or empty" - return "direct-pdf", url, text - - -def strip_html(html: str) -> str: - html = re.sub(r"(?is)<(script|style|noscript).*?>.*?</\1>", " ", html) - html = re.sub(r"(?is)<br\s*/?>", "\n", html) - html = re.sub(r"(?is)</p>", "\n\n", html) - html = re.sub(r"<[^>]+>", " ", html) - html = re.sub(r"[ \t]+\n", "\n", html) - html = re.sub(r"\n{3,}", "\n\n", html) - html = re.sub(r" +", " ", html) - return html.strip() - - -def fetch_html(url: str) -> tuple[str, str, str]: - try: - raw = http_get(url, accept="text/html") - except (urllib.error.URLError, urllib.error.HTTPError) as exc: - return "direct-html-blocked", url, f"HTTP error: {exc}" - text = strip_html(raw.decode("utf-8", errors="replace")) - if len(text) < 200: - return "direct-html-empty", url, text - return "direct-html", url, text - - -def truncate_excerpt(text: str, limit: int = EXCERPT_MAX) -> str: - text = text.strip() - if len(text) <= limit: - return text - return text[:limit].rstrip() + f"\n\n[… truncated at {limit} chars — full text available at Access URL]" - - -def has_fulltext_section(body: str) -> bool: - return bool(re.search(r"^##\s+Full text extract\b", body, re.M)) - - -def remove_fulltext_section(body: str) -> str: - return re.sub(r"\n?## Full text extract\b[\s\S]*\Z", "", body.rstrip()) + "\n" - - -def channel_source_na(source: str) -> bool: - host = urllib.parse.urlparse(source).netloc.lower() - return any( - x in host - for x in ( - "youtube.com", - "youtu.be", - "reddit.com", - "twitter.com", - "x.com", - ) - ) - - -def build_section( - *, - today: str, - method: str, - access_url: str, - abstract: str, - body_text: str, - verdict: str, -) -> str: - lines = [f"## Full text extract (fetch {today})", ""] - if access_url: - lines.append(f"**Access:** {access_url}") - lines.append(f"**Fetch method:** {method}") - lines.append("") - if abstract: - lines.append("**Abstract (fetched):**") - lines.append("") - lines.append(truncate_excerpt(abstract, 2500)) - lines.append("") - if body_text: - lines.append("**Body excerpt:**") - lines.append("") - lines.append(truncate_excerpt(body_text)) - lines.append("") - lines.append(f"**Audit verdict:** {verdict}") - lines.append("") - return "\n".join(lines) - - -def update_fulltext_assessment(body: str, assessment: str) -> str: - if re.search(r"^\-\s+\*\*Full text:\*\*", body, re.M): - return re.sub( - r"^\-\s+\*\*Full text:\*\*.*$", - f"- **Full text:** {assessment}", - body, - count=1, - flags=re.M, - ) - # Insert before Dedup note inside Source assessment if present - m = re.search( - r"(## Source assessment\s*\n)(.*?)(^\-\s+\*\*Dedup note:\*\*)", - body, - re.M | re.S, - ) - if m: - prefix, middle, dedup = m.group(1), m.group(2), m.group(3) - if "**Full text:**" not in middle: - insert = f"- **Full text:** {assessment}\n" - return body.replace( - m.group(0), - prefix + middle + insert + dedup, - 1, - ) - return body - - -def update_tags_in_frontmatter(header: str, add: list[str], remove_prefixes: tuple[str, ...]) -> str: - # Non-greedy: header may contain other bracketed YAML values (e.g. author arrays). - m = re.search(r"^tags:\s*\[(.*?)\]\s*$", header, re.M) - if not m: - return header - tags = parse_tags(m.group(1)) - tags = [t for t in tags if not any(t.startswith(p) for p in remove_prefixes)] - for t in add: - if t not in tags: - tags.append(t) - new_inner = ", ".join(f'"{t}"' for t in tags) - return header[: m.start()] + f"tags: [{new_inner}]" + header[m.end() :] - - -def fetch_for_stub(path: Path, *, today: str, unpaywall_email: str | None, force: bool) -> FetchResult: - rel = str(path.relative_to(ROOT)) - try: - text = path.read_text(encoding="utf-8") - except OSError as exc: - return FetchResult(rel, "error", message=str(exc)) - - fm, body = parse_frontmatter(text) - if fm.get("type", "").lower() != "raw": - return FetchResult(rel, "skipped", message="not type: raw") - - if has_fulltext_section(body) and not force: - return FetchResult(rel, "skipped", message="already has Full text extract") - - source = normalize_source_url(fm.get("source", "")) - doi, pmid = extract_identifiers(fm) - abstract = "" - body_text = "" - method = "" - access_url = source - verdict = "" - status = "pending" - - if channel_source_na(source): - section = build_section( - today=today, - method="channel-source", - access_url=source, - abstract="", - body_text="", - verdict="Full text **not applicable** — channel/community source; use manual monitor notes.", - ) - return FetchResult(rel, "na", method="channel-source", access_url=source, section=section) - - # Europe PMC for scholarly IDs - epmc = europepmc_search(doi, pmid) if (doi or pmid) else None - if epmc: - abstract = (epmc.get("abstractText") or "").strip() - pmcid = epmc.get("pmcid") - is_oa = (epmc.get("isOpenAccess") or "").upper() == "Y" - if pmcid and is_oa: - method, access_url, full = fetch_europepmc_fulltext(pmcid) - if full: - body_text = full - status = "fetched" - verdict = "Full text **fetch verified** (Europe PMC OA XML). Claims still need human appraisal." - else: - status = "pending" - verdict = "**full-text-pending** — indexed in Europe PMC but XML fetch blocked." - elif abstract: - status = "pending" - method = "europepmc-abstract" - access_url = f"https://doi.org/{doi}" if doi else source - verdict = "**full-text-pending** — abstract only (not open access in Europe PMC)." - if not abstract and doi: - abstract = openalex_abstract(doi) or crossref_abstract(doi) - - # Unpaywall PDF fallback for DOI - if status != "fetched" and doi: - pdf_url = unpaywall_best_pdf(doi, unpaywall_email) - if pdf_url: - m, u, txt = fetch_direct_pdf(pdf_url) - if txt and not txt.startswith("HTTP") and not txt.startswith("Response"): - method, access_url, body_text = m, u, txt - status = "fetched" - verdict = "Full text **fetch verified** (Unpaywall OA PDF). Claims still need human appraisal." - - # Direct PDF URL in source field - if status != "fetched" and source.lower().endswith(".pdf"): - m, u, txt = fetch_direct_pdf(source) - if m == "direct-pdf": - method, access_url, body_text = m, u, txt - status = "fetched" - verdict = "Full text **fetch verified** (direct PDF). Claims still need human appraisal." - elif not verdict: - method, access_url = m, u - status = "pending" - verdict = f"**full-text-pending** — {txt}" - - # HTML institutional / guidance pages - if status != "fetched" and source.startswith("http") and not source.lower().endswith(".pdf"): - if not method or method in ("europepmc-abstract", ""): - m, u, txt = fetch_html(source) - if m == "direct-html" and len(txt) > 400: - method, access_url, body_text = m, u, txt - status = "fetched" - verdict = "Full text **fetch verified** (HTML page). Claims still need human appraisal." - elif not verdict and txt: - method, access_url = m, u - abstract = abstract or truncate_excerpt(txt, 1500) - status = "pending" - verdict = "**full-text-pending** — partial HTML only." - - if not verdict: - if abstract: - status = "pending" - method = method or "abstract-only" - verdict = "**full-text-pending** — abstract/metadata only; no OA full text found." - else: - status = "pending" - method = method or "none" - verdict = "**full-text-pending** — no automated full text or abstract found." - - if not abstract and doi: - abstract = openalex_abstract(doi) or crossref_abstract(doi) - - section = build_section( - today=today, - method=method or "none", - access_url=access_url, - abstract=abstract, - body_text=body_text, - verdict=verdict, - ) - return FetchResult( - rel, - status, - method=method, - access_url=access_url, - excerpt_chars=len(body_text), - section=section, - ) - - -def apply_result(path: Path, original: str, result: FetchResult, today: str) -> str: - fm, body = parse_frontmatter(original) - header_end = original.find("\n---", 3) - header = original[3:header_end] - - body = remove_fulltext_section(body) - body = update_fulltext_assessment( - body, - "verified (automated fetch)" if result.status == "fetched" else ( - "not applicable" if result.status == "na" else "pending (automated fetch)" - ), - ) - body = body.rstrip() + "\n\n" + result.section - - tag_add: list[str] = [] - if result.status == "fetched": - tag_add.append(f"full-text-fetched-{today}") - elif result.status == "na": - tag_add.append("full-text-na") - else: - tag_add.append("full-text-pending") - - header = update_tags_in_frontmatter( - header, - tag_add, - ("full-text-fetched-", "full-text-pending", "full-text-na"), - ) - return f"---\n{header}\n---\n{body}" - - -def resolve_knowledge_file(path: Path) -> Path: - """Resolve an existing Markdown stub and keep it inside 10_knowledge/.""" - knowledge_root = KNOWLEDGE.resolve() - try: - resolved = path.resolve(strict=True) - except OSError as exc: - raise PathScopeError(f"stub does not exist: {path}") from exc - try: - resolved.relative_to(knowledge_root) - except ValueError as exc: - raise PathScopeError( - f"stub must be inside {knowledge_root}: {path}" - ) from exc - if not resolved.is_file(): - raise PathScopeError(f"stub is not a regular file: {path}") - if resolved.suffix.lower() != ".md": - raise PathScopeError(f"stub must be a Markdown file: {path}") - return resolved - - -def resolve_knowledge_directory(path: Path) -> Path: - """Resolve a requested scan root without permitting traversal escapes.""" - knowledge_root = KNOWLEDGE.resolve() - resolved = path.resolve() - try: - resolved.relative_to(knowledge_root) - except ValueError as exc: - raise PathScopeError( - f"subset must stay inside {knowledge_root}: {path}" - ) from exc - return resolved - - -def _same_file(left: os.stat_result, right: os.stat_result) -> bool: - return (left.st_dev, left.st_ino) == (right.st_dev, right.st_ino) - - -def _same_version(left: os.stat_result, right: os.stat_result) -> bool: - return ( - left.st_dev, - left.st_ino, - left.st_size, - left.st_mtime_ns, - left.st_ctime_ns, - ) == ( - right.st_dev, - right.st_ino, - right.st_size, - right.st_mtime_ns, - right.st_ctime_ns, - ) - - -def atomic_replace_text( - path: Path, - text: str, - *, - expected: os.stat_result, -) -> None: - """Durably replace a file, refusing a known competing modification.""" - temp_path: Path | None = None - try: - with tempfile.NamedTemporaryFile( - mode="w", - encoding="utf-8", - dir=path.parent, - prefix=f".{path.name}.", - suffix=".tmp", - delete=False, - ) as tmp: - temp_path = Path(tmp.name) - tmp.write(text) - tmp.flush() - os.fsync(tmp.fileno()) - os.chmod(temp_path, stat.S_IMODE(expected.st_mode)) - - current = os.stat(path, follow_symlinks=False) - if not _same_version(expected, current): - raise ConcurrentUpdateError(f"stub changed before replace: {path}") - - os.replace(temp_path, path) - temp_path = None - - try: - parent_fd = os.open(path.parent, os.O_RDONLY) - try: - os.fsync(parent_fd) - finally: - os.close(parent_fd) - except OSError: - # Some filesystems do not support directory fsync. The file replace - # is still atomic; only the extra crash-durability barrier is absent. - pass - finally: - if temp_path is not None: - try: - temp_path.unlink() - except FileNotFoundError: - pass - - -def apply_result_atomically( - path: Path, - result: FetchResult, - *, - today: str, - force: bool, -) -> bool: - """Apply a fetched section once, preserving edits from competing runs.""" - path = resolve_knowledge_file(path) - nofollow = getattr(os, "O_NOFOLLOW", 0) - - for _ in range(LOCK_RETRIES): - fd = os.open(path, os.O_RDONLY | nofollow) - try: - fcntl.flock(fd, fcntl.LOCK_EX) - locked_stat = os.fstat(fd) - path_stat = os.stat(path, follow_symlinks=False) - if not _same_file(locked_stat, path_stat): - continue - - # Re-resolve after taking the lock so a swapped symlink or parent - # cannot redirect the update beyond the knowledge tree. - if resolve_knowledge_file(path) != path: - raise ConcurrentUpdateError(f"stub path changed before apply: {path}") - - with os.fdopen(os.dup(fd), "r", encoding="utf-8") as current_file: - original = current_file.read() - current_fm, current_body = parse_frontmatter(original) - if current_fm.get("type", "").lower() != "raw": - return False - if has_fulltext_section(current_body) and not force: - return False - - updated = apply_result(path, original, result, today) - atomic_replace_text(path, updated, expected=locked_stat) - return True - finally: - fcntl.flock(fd, fcntl.LOCK_UN) - os.close(fd) - - raise ConcurrentUpdateError(f"stub kept changing while acquiring lock: {path}") - - -def find_stubs(subset: str | None, file_arg: str | None) -> list[Path]: - if file_arg: - p = Path(file_arg) - if not p.is_absolute(): - p = ROOT / p - return [resolve_knowledge_file(p)] - base = resolve_knowledge_directory(KNOWLEDGE / subset if subset else KNOWLEDGE) - if not base.is_dir(): - return [] - out: list[Path] = [] - seen: set[Path] = set() - for p in sorted(base.rglob("*.md")): - if p.name.lower() == "index.md": - continue - if "__raw__" in p.name or p.parent.name == "raw": - try: - resolved = resolve_knowledge_file(p) - except PathScopeError: - continue - if resolved not in seen: - seen.add(resolved) - out.append(resolved) - return out - - -def main(argv: list[str] | None = None) -> int: - parser = argparse.ArgumentParser(description="Fetch full text for source-literature stubs.") - parser.add_argument("--dry-run", action="store_true", help="Report only; do not write files") - parser.add_argument("--apply", action="store_true", help="Append extracts to stub files") - parser.add_argument("--json", action="store_true", help="Emit machine-readable report") - parser.add_argument("--subset", metavar="DOMAIN", help="Limit to 10_knowledge/<domain>/") - parser.add_argument("--file", metavar="PATH", help="Single stub file") - parser.add_argument("--force", action="store_true", help="Re-fetch even if extract exists") - args = parser.parse_args(argv) - - if not args.dry_run and not args.apply and not args.json: - args.dry_run = True - - today = date.today().isoformat() - unpaywall_email = ( - __import__("os").environ.get("UNPAYWALL_EMAIL") - or __import__("os").environ.get("MAINFRAME_UNPAYWALL_EMAIL") - ) - - try: - stubs = find_stubs(args.subset, args.file) - except PathScopeError as exc: - parser.error(str(exc)) - results: list[FetchResult] = [] - for path in stubs: - try: - result = fetch_for_stub( - path, today=today, unpaywall_email=unpaywall_email, force=args.force - ) - except Exception as exc: # noqa: BLE001 — batch must survive one bad stub - result = FetchResult( - str(path.relative_to(ROOT)), - "error", - message=str(exc), - ) - results.append(result) - if args.apply and result.section and result.status in ("fetched", "pending", "na"): - try: - applied = apply_result_atomically( - path, - result, - today=today, - force=args.force, - ) - except (OSError, PathScopeError, ConcurrentUpdateError) as exc: - result.status = "error" - result.section = "" - result.message = f"apply failed: {exc}" - else: - if not applied: - result.status = "skipped" - result.section = "" - result.message = "stub changed before apply; left unchanged" - - if args.json: - print(json.dumps([asdict(r) for r in results], indent=2)) - else: - for r in results: - flag = r.status.upper() - extra = f" ({r.method}, {r.excerpt_chars} chars)" if r.method else "" - print(f"[{flag}] {r.path}{extra}") - if r.message: - print(f" {r.message}") - - fetched = sum(1 for r in results if r.status == "fetched") - pending = sum(1 for r in results if r.status == "pending") - print( - f"\nSummary: {len(results)} scanned, {fetched} fetched, {pending} pending, " - f"{sum(1 for r in results if r.status == 'skipped')} skipped", - file=sys.stderr, - ) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/generate-client-hooks.py b/scripts/generate-client-hooks.py deleted file mode 100755 index 47540e8..0000000 --- a/scripts/generate-client-hooks.py +++ /dev/null @@ -1,93 +0,0 @@ -#!/usr/bin/env python3 -""" -Stub for unifying client hook configs (Claude, Codex, Antigravity, Aider watcher). - -Part of Fix 5 (client config fragmentation). - -In a full impl, this would read a single source (e.g. .context/client-hooks.json or yaml) -and emit the per-client files: - .claude/settings.json (or hooks section) - .codex/hooks.json - .antigravity/hooks.json - etc. - -For now: a thin generator stub that documents the pattern and can be expanded. -Run with --dry-run to preview; --apply to write (respecting local overrides). - -This keeps the public surface clean while reducing duplication of hook wiring -(bin/workflow-event calls remain the single backend). -""" - -from __future__ import annotations - -import argparse -import json -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] - - -def load_source() -> dict: - # Placeholder: in real version load from .context/client-hooks-source.json - # or derive from the existing settings we maintain in .claude/ etc. - # For demo, hardcode a minimal unified view. Use a shell-safe template that - # the real hook files expand with their ${VAR} prefixes. - return { - "events": [ - "SessionStart", "UserPromptSubmit", "PreToolUse", "PostToolUse", - "PostToolUseFailure", "Stop", "SubagentStart", "SubagentStop", - "PermissionRequest", "PreCompact", "PostCompact" - ], - # MAINFRAME_ROOT / client project-dir env only — never hard-code a personal Desktop path. - "command_template": "${MAINFRAME_ROOT:-${CLAUDE_PROJECT_DIR}}/bin/workflow-event --client {client}", - "timeout": 5, - } - - -def generate_for_client(client: str, source: dict) -> dict: - # Very simplified emitter. Real version would match the exact layout of - # .claude/settings.json hooks, .codex/hooks.json, etc. - cmd = source["command_template"].format(client=client) - hooks = {} - for ev in source["events"]: - hooks[ev] = [{"type": "command", "command": cmd, "timeout": source["timeout"]}] - return {"hooks": hooks} - - -def main() -> int: - parser = argparse.ArgumentParser() - parser.add_argument("--dry-run", action="store_true") - parser.add_argument("--apply", action="store_true") - parser.add_argument("--client", choices=["claude", "codex", "antigravity"], default="claude") - parser.add_argument("--pixel", default=None, help="Unique pixel/agent ID for trackable hooks in the local agents pixel agent tracker (e.g. main-claude-pixel)") - args = parser.parse_args() - - source = load_source() - cfg = generate_for_client(args.client, source) - - # Inject pixel into all hook commands for unique tracking (like web_panel/human clients) - if args.pixel: - pixel_arg = f" --pixel {args.pixel}" - for event_hooks in cfg.get("hooks", {}).values(): - for h in event_hooks: - for hook in h.get("hooks", []): - if "command" in hook: - hook["command"] += pixel_arg - - if args.dry_run: - print(json.dumps(cfg, indent=2)) - return 0 - - if args.apply: - # In real: write to the correct dotfile location for the client. - # e.g. for claude: (ROOT / ".claude" / "settings.json") but preserve other keys. - print(f"Would write generated hooks for {args.client} with pixel={args.pixel} (stub).") - print("Expand this script to do real merge/write while keeping manual local overrides.") - return 0 - - print("Use --dry-run or --apply --pixel ID. This is a stub for centralizing unique/trackable hook definitions.") - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/lifecycle_identity.py b/scripts/lifecycle_identity.py new file mode 100644 index 0000000..c6498ed --- /dev/null +++ b/scripts/lifecycle_identity.py @@ -0,0 +1,500 @@ +"""Direct lifecycle identity and cross-root record resolution. + +MainFrame has two lifecycle roots whose records share one slug namespace. This +module is deliberately small and dependency-free so command-line readers, +writers, inventory checks, and tests can use the same authority rules. + +MindGraph, generated indexes, and Workstation projections may validate the +result returned here, but they are never used to choose a path. +""" + +from __future__ import annotations + +import ast +import hashlib +import json +import os +import re +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any, Iterable + + +SLUG_RE = re.compile(r"^[A-Za-z0-9][A-Za-z0-9._-]*$") +LIFECYCLE_ROOTS = { + "project": "30_projects", + "operation": "40_operations", +} +VALID_RECORD_TYPES = frozenset({"project", "operation"}) + + +class IdentityError(RuntimeError): + """Base class for fail-closed identity failures.""" + + +class MissingIdentity(IdentityError): + """No direct authority exists for the requested slug.""" + + +class DuplicateIdentity(IdentityError): + """The slug exists in more than one lifecycle root.""" + + +class InvalidIdentity(IdentityError): + """The direct authority exists but is unsafe or internally inconsistent.""" + + +class FrontmatterError(InvalidIdentity): + """Lifecycle frontmatter is ambiguous or malformed.""" + + +def _coerce_value(value: str) -> Any: + value = value.strip() + if not value: + return "" + if value[0:1] in {"\"", "'"}: + try: + return ast.literal_eval(value) + except (ValueError, SyntaxError): + return value.strip("\"'") + if value.startswith(("[", "{", "(")): + try: + return ast.literal_eval(value) + except (ValueError, SyntaxError): + return value + lowered = value.lower() + if lowered in {"true", "false"}: + return lowered == "true" + if lowered in {"null", "none", "~"}: + return None + return value + + +def parse_frontmatter(text: str, *, strict: bool = True) -> dict[str, Any]: + """Parse the simple metadata subset used by lifecycle README files. + + MainFrame's README frontmatter is intentionally Markdown-first. The + resolver only needs scalar identity fields, and refusing malformed lines is + safer than silently importing a full YAML parser into every CLI. + """ + + if not text.startswith("---"): + return {} + lines = text.splitlines() + if not lines or lines[0].strip() != "---": + return {} + end = next((i for i, line in enumerate(lines[1:], start=1) if line.strip() == "---"), None) + if end is None: + if strict: + raise FrontmatterError("frontmatter opening delimiter has no closing delimiter") + return {} + result: dict[str, Any] = {} + for line in lines[1:end]: + stripped = line.strip() + if not stripped or stripped.startswith("#"): + continue + if line[:1].isspace(): + if strict: + raise FrontmatterError("indented/nested frontmatter declarations are not supported") + continue + if ":" not in line: + if strict: + raise FrontmatterError(f"frontmatter line is not a key/value declaration: {line!r}") + continue + key, value = line.split(":", 1) + key = key.strip() + if not key: + if strict: + raise FrontmatterError("frontmatter contains an empty key") + continue + if not re.fullmatch(r"[A-Za-z_][A-Za-z0-9_-]*", key): + if strict: + raise FrontmatterError(f"frontmatter key is invalid: {key!r}") + continue + if key in result: + raise FrontmatterError(f"duplicate frontmatter key: {key}") + result[key] = _coerce_value(value) + return result + + +def sha256_file(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as handle: + for chunk in iter(lambda: handle.read(1 << 20), b""): + digest.update(chunk) + return digest.hexdigest() + + +def _root_value(mainframe_root: Path, name: str, override: Path | str | None) -> Path: + if override is not None: + value = Path(override).expanduser() + else: + env_name = "MAINFRAME_PROJECTS_ROOT" if name == "projects" else "MAINFRAME_OPERATIONS_ROOT" + env_value = os.environ.get(env_name) + value = Path(env_value).expanduser() if env_value else mainframe_root / { + "projects": "30_projects", + "operations": "40_operations", + }[name] + if not value.is_absolute(): + value = mainframe_root / value + # Keep the lexical path here. Resolving before the scan would erase the + # fact that a lifecycle root itself is symlinked, which is a direct- + # authority violation that must fail closed. + return value.absolute() + + +def lifecycle_roots( + mainframe_root: Path | str, + *, + projects_root: Path | str | None = None, + operations_root: Path | str | None = None, +) -> dict[str, Path]: + root = Path(mainframe_root).expanduser().resolve() + return { + "projects": _root_value(root, "projects", projects_root), + "operations": _root_value(root, "operations", operations_root), + } + + +def display_path(mainframe_root: Path | str, path: Path | str) -> str: + root = Path(mainframe_root).expanduser().resolve() + candidate = Path(path).expanduser().resolve() + try: + return candidate.relative_to(root).as_posix() + except ValueError: + return candidate.as_posix() + + +def _safe_child(parent: Path, child: Path) -> tuple[bool, str | None]: + if child.is_symlink(): + return False, "symlinked lifecycle entity is not an authority" + try: + resolved_parent = parent.resolve() + resolved_child = child.resolve() + except OSError as exc: + return False, f"cannot resolve lifecycle entity: {exc}" + if not resolved_child.is_relative_to(resolved_parent): + return False, "lifecycle entity escapes its root" + if not child.is_dir(): + return False, "lifecycle entity is not a directory" + return True, None + + +def _effective_record_type(metadata: dict[str, Any], root_kind: str) -> str | None: + keys = [key for key in ("record_type", "type") if key in metadata] + if keys: + values: list[str | None] = [] + for key in keys: + raw = metadata.get(key) + if not isinstance(raw, str): + values.append(None) + continue + normalized = raw.strip().strip("\"'").lower() + if normalized == "operation": + values.append("operation") + elif normalized in {"project", "program", "evaluation", "lifecycle"}: + values.append("project") + else: + # An explicit but unknown declaration must not silently fall + # back to the containing folder's historical default. + values.append(None) + if len(values) == 2 and values[0] != values[1]: + return None + return values[0] + # Existing project READMEs predate record_type. A project-root record is + # compatible as a project; an operation-root record must opt in explicitly. + return "project" if root_kind == "projects" else None + + +def _lifecycle_state(metadata: dict[str, Any]) -> tuple[str | None, str | None]: + project_state = metadata.get("project_state") + lifecycle_state = metadata.get("lifecycle_state") + if project_state is not None and lifecycle_state is not None: + if str(project_state).strip() != str(lifecycle_state).strip(): + return None, "project_state and lifecycle_state conflict" + value = lifecycle_state if lifecycle_state is not None else project_state + if value is None: + value = metadata.get("status") + return (str(value).strip() if value is not None else None), None + + +@dataclass(frozen=True) +class LifecycleRecord: + slug: str + path: Path + root_kind: str + record_type: str + metadata: dict[str, Any] + lifecycle_state: str | None + readme_path: Path + readme_sha256: str + issues: tuple[str, ...] = field(default_factory=tuple) + coordination_path: Path | None = None + coordination_sha256: str | None = None + coordination_metadata: dict[str, Any] = field(default_factory=dict) + wip_class: str | None = None + state_source: str = "README.md" + + @property + def root_name(self) -> str: + return "30_projects" if self.root_kind == "projects" else "40_operations" + + def relative_path(self, mainframe_root: Path | str) -> str: + return display_path(mainframe_root, self.path) + + def manifest_entry(self, mainframe_root: Path | str) -> dict[str, Any]: + result = { + "id": self.slug, + "record_type": self.record_type, + "path": self.relative_path(mainframe_root), + "lifecycle_state": self.lifecycle_state, + "readme_sha256": self.readme_sha256, + } + if self.wip_class is not None: + result["wip_class"] = self.wip_class + return result + + +@dataclass +class IdentityScan: + records: list[LifecycleRecord] = field(default_factory=list) + issues: list[str] = field(default_factory=list) + root_issues: list[str] = field(default_factory=list) + + @property + def by_slug(self) -> dict[str, list[LifecycleRecord]]: + result: dict[str, list[LifecycleRecord]] = {} + for record in self.records: + result.setdefault(record.slug, []).append(record) + return result + + +def _record_for_child(child: Path, root_kind: str) -> LifecycleRecord | None: + slug = child.name + issues: list[str] = [] + ok, issue = _safe_child(child.parent, child) + if not ok: + issues.append(issue or "unsafe lifecycle entity") + if not SLUG_RE.fullmatch(slug): + issues.append("invalid lifecycle slug") + readme = child / "README.md" + if readme.is_symlink(): + issues.append("README.md is symlinked") + if not readme.is_file(): + issues.append("README.md missing or not a regular file") + metadata: dict[str, Any] = {} + readme_hash = "" + else: + try: + resolved_readme = readme.resolve() + if not resolved_readme.is_relative_to(child.resolve()): + issues.append("README.md escapes lifecycle entity") + raw = readme.read_bytes() + metadata = parse_frontmatter(raw.decode("utf-8"), strict=True) + readme_hash = hashlib.sha256(raw).hexdigest() + except FrontmatterError as exc: + metadata = {} + readme_hash = "" + issues.append(str(exc)) + except (OSError, UnicodeError) as exc: + metadata = {} + readme_hash = "" + issues.append(f"README.md unreadable: {type(exc).__name__}: {exc}") + + coordination_path = child / "PROJECT.md" + coordination_metadata: dict[str, Any] = {} + coordination_hash: str | None = None + if coordination_path.exists() or coordination_path.is_symlink(): + if coordination_path.is_symlink(): + issues.append("PROJECT.md is symlinked") + elif not coordination_path.is_file(): + issues.append("PROJECT.md is not a regular file") + else: + try: + coordination_raw = coordination_path.read_bytes() + coordination_hash = hashlib.sha256(coordination_raw).hexdigest() + coordination_text = coordination_raw.decode("utf-8") + # PROJECT.md is an optional public-mirror coordination surface. + # A plain Markdown file is valid when it carries no frontmatter; + # an attempted frontmatter block must still be unambiguous. + coordination_metadata = parse_frontmatter(coordination_text, strict=True) + except FrontmatterError as exc: + issues.append(f"PROJECT.md: {exc}") + except (OSError, UnicodeError) as exc: + issues.append(f"PROJECT.md unreadable: {type(exc).__name__}: {exc}") + + record_type = _effective_record_type(metadata, root_kind) + if record_type not in VALID_RECORD_TYPES: + issues.append("record_type is missing or invalid") + record_type = "invalid" + readme_state, state_issue = _lifecycle_state(metadata) + if state_issue: + issues.append(state_issue) + coordination_state, coordination_state_issue = _lifecycle_state(coordination_metadata) + if coordination_state_issue: + issues.append(f"PROJECT.md: {coordination_state_issue}") + if coordination_state is not None: + if readme_state is not None and readme_state != coordination_state: + issues.append("README lifecycle state differs from PROJECT.md owner") + state = coordination_state + state_source = "PROJECT.md" + else: + state = readme_state + state_source = "README.md" + readme_wip = metadata.get("wip_class") + coordination_wip = coordination_metadata.get("wip_class") + if readme_wip is not None and coordination_wip is not None and str(readme_wip) != str(coordination_wip): + issues.append("README wip_class differs from PROJECT.md owner") + wip_raw = coordination_wip if coordination_wip is not None else readme_wip + wip_class = str(wip_raw).strip() if wip_raw is not None else None + if wip_class is not None and wip_class not in {"product", "eval", "anchor"}: + issues.append(f"invalid wip_class: {wip_class}") + project_record_type = _effective_record_type(coordination_metadata, root_kind) + if "record_type" in coordination_metadata and project_record_type != record_type: + issues.append("README and PROJECT.md record_type declarations differ") + if root_kind == "operations" and metadata.get("record_type") != "operation": + issues.append("operation-root README must declare record_type: operation") + return LifecycleRecord( + slug=slug, + path=child, + root_kind=root_kind, + record_type=record_type, + metadata=metadata, + lifecycle_state=state, + readme_path=readme, + readme_sha256=readme_hash, + issues=tuple(issues), + coordination_path=coordination_path if coordination_path.is_file() else None, + coordination_sha256=coordination_hash, + coordination_metadata=coordination_metadata, + wip_class=wip_class, + state_source=state_source, + ) + + +def scan_lifecycle_records( + mainframe_root: Path | str, + *, + projects_root: Path | str | None = None, + operations_root: Path | str | None = None, +) -> IdentityScan: + root = Path(mainframe_root).expanduser().resolve() + roots = lifecycle_roots(root, projects_root=projects_root, operations_root=operations_root) + result = IdentityScan() + for root_kind, root_path in (("projects", roots["projects"]), ("operations", roots["operations"])): + if root_path.is_symlink(): + issue = f"{root_kind} lifecycle root is symlinked: {root_path}" + result.issues.append(issue) + result.root_issues.append(issue) + continue + if not root_path.exists(): + continue + if not root_path.is_dir(): + issue = f"{root_kind} lifecycle root is not a directory: {root_path}" + result.issues.append(issue) + result.root_issues.append(issue) + continue + try: + children: Iterable[Path] = sorted(root_path.iterdir(), key=lambda p: p.name) + except OSError as exc: + issue = f"cannot enumerate {root_kind} lifecycle root: {exc}" + result.issues.append(issue) + result.root_issues.append(issue) + continue + for child in children: + if child.name.startswith("."): + continue + if not child.is_dir() and not child.is_symlink(): + continue + record = _record_for_child(child, root_kind) + if record is None: + continue + result.records.append(record) + result.issues.extend( + f"{record.root_name}/{record.slug}: {issue}" for issue in record.issues + ) + for slug, matches in result.by_slug.items(): + if len(matches) > 1: + result.issues.append( + f"duplicate lifecycle identity {slug}: " + + ", ".join(sorted(r.relative_path(root) for r in matches)) + ) + return result + + +def resolve_record( + mainframe_root: Path | str, + slug: str, + *, + expected_record_type: str | None = None, + projects_root: Path | str | None = None, + operations_root: Path | str | None = None, +) -> LifecycleRecord: + if not SLUG_RE.fullmatch(slug): + raise InvalidIdentity(f"invalid lifecycle slug: {slug!r}") + root = Path(mainframe_root).expanduser().resolve() + scan = scan_lifecycle_records( + root, + projects_root=projects_root, + operations_root=operations_root, + ) + if scan.root_issues: + raise InvalidIdentity("cannot prove cross-root lifecycle uniqueness: " + "; ".join(scan.root_issues)) + matches = scan.by_slug.get(slug, []) + if len(matches) == 0: + raise MissingIdentity(f"no direct lifecycle authority for {slug!r}") + if len(matches) > 1: + locations = ", ".join(sorted(record.relative_path(root) for record in matches)) + raise DuplicateIdentity(f"duplicate lifecycle identity {slug!r}: {locations}") + record = matches[0] + if record.issues: + raise InvalidIdentity(f"invalid lifecycle authority {record.relative_path(root)}: " + "; ".join(record.issues)) + if expected_record_type and record.record_type != expected_record_type: + raise InvalidIdentity( + f"{slug!r} has record_type={record.record_type!r}, expected {expected_record_type!r}" + ) + return record + + +def resolve_path( + mainframe_root: Path | str, + slug: str, + *, + expected_record_type: str | None = None, + projects_root: Path | str | None = None, + operations_root: Path | str | None = None, +) -> Path: + return resolve_record( + mainframe_root, + slug, + expected_record_type=expected_record_type, + projects_root=projects_root, + operations_root=operations_root, + ).path + + +def stable_identity_snapshot(mainframe_root: Path | str, slug: str) -> dict[str, Any]: + """Return a compact direct-authority snapshot for a writer precondition.""" + + record = resolve_record(mainframe_root, slug) + return { + "id": record.slug, + "record_type": record.record_type, + "path": record.relative_path(mainframe_root), + "readme_sha256": record.readme_sha256, + "coordination_sha256": record.coordination_sha256, + "lifecycle_state": record.lifecycle_state, + "state_source": record.state_source, + "wip_class": record.wip_class, + } + + +def write_json_atomic(path: Path, payload: dict[str, Any]) -> None: + """Write a small receipt without replacing an existing receipt silently.""" + + path.parent.mkdir(parents=True, exist_ok=True) + if path.exists(): + raise FileExistsError(path) + temporary = path.with_name(f".{path.name}.{os.getpid()}.tmp") + temporary.write_text(json.dumps(payload, indent=2, sort_keys=True) + "\n", encoding="utf-8") + os.replace(temporary, path) diff --git a/scripts/lifecycle_runtime.py b/scripts/lifecycle_runtime.py new file mode 100644 index 0000000..491268a --- /dev/null +++ b/scripts/lifecycle_runtime.py @@ -0,0 +1,53 @@ +"""Generic lease-aware lifecycle writers. + +Readers resolve a slug from direct lifecycle authority at use time. Writers +acquire the shared migration lease before resolving and hold it until the +filesystem write is complete. This module is MainFrame-owned substrate: it +must not import process-evaluation internals. +""" + +from __future__ import annotations + +from contextlib import contextmanager +from pathlib import Path +from typing import Iterator + +from lifecycle_identity import LifecycleRecord, resolve_record +from migration_lease import MigrationLease, shared_writer_lease + + +def resolve_lifecycle( + root: Path | str, + slug: str, + *, + expected_record_type: str | None = None, +) -> LifecycleRecord: + return resolve_record(root, slug, expected_record_type=expected_record_type) + + +@contextmanager +def lifecycle_writer( + root: Path | str, + slug: str, + holder: str, + *, + expected_record_type: str | None = None, + pause_token: str | None = None, + state_root: Path | str | None = None, +) -> Iterator[tuple[LifecycleRecord, MigrationLease]]: + """Acquire the writer fence before resolving a lifecycle path.""" + + root_path = Path(root).expanduser().resolve() + with shared_writer_lease( + root_path, + holder, + state_root=state_root, + pause_token=pause_token, + pause_scope=slug, + ) as lease: + record = resolve_lifecycle( + root_path, + slug, + expected_record_type=expected_record_type, + ) + yield record, lease diff --git a/scripts/mainframe_doctor/__init__.py b/scripts/mainframe_doctor/__init__.py deleted file mode 100644 index 04ab66b..0000000 --- a/scripts/mainframe_doctor/__init__.py +++ /dev/null @@ -1,3 +0,0 @@ -"""MainFrame doctor — read-only health aggregator (Unit 1.3 shell).""" - -__version__ = "0.1.0" diff --git a/scripts/mainframe_doctor/__main__.py b/scripts/mainframe_doctor/__main__.py deleted file mode 100644 index d692423..0000000 --- a/scripts/mainframe_doctor/__main__.py +++ /dev/null @@ -1,80 +0,0 @@ -"""CLI entry: python -m mainframe_doctor""" - -from __future__ import annotations - -import argparse -import json -import sys -from pathlib import Path - - -def main(argv: list[str] | None = None) -> int: - # Ensure scripts/ is on path when executed as module from bin wrapper - scripts_dir = Path(__file__).resolve().parents[1] - if str(scripts_dir) not in sys.path: - sys.path.insert(0, str(scripts_dir)) - - from mainframe_doctor.runner import format_human, run_doctor - - parser = argparse.ArgumentParser( - prog="mainframe-doctor", - description="Read-only MainFrame health aggregator (vector, not one boolean).", - ) - profile = parser.add_mutually_exclusive_group() - profile.add_argument("--quick", action="store_true", help="Session-start orientation profile (default)") - profile.add_argument("--deep", action="store_true", help="Full catalogue profile") - parser.add_argument("--component", metavar="SUBSYSTEM", help="Run one subsystem only") - parser.add_argument("--json", action="store_true", help="JSON only on stdout") - parser.add_argument( - "--fixture", - type=Path, - help="Fixture JSON file or directory (fixture.json); no live mutation", - ) - parser.add_argument( - "--root", - type=Path, - default=None, - help="MainFrame root (default: detect from this install)", - ) - parser.add_argument( - "--catalogue", - type=Path, - default=None, - help="Override catalogue path", - ) - parser.add_argument( - "--invariants", - type=Path, - default=None, - help="Override required-invariants path", - ) - args = parser.parse_args(argv) - - if args.root is not None: - root = args.root.resolve() - else: - # bin/mainframe-doctor → repo root; or scripts/mainframe_doctor → parents[2] - here = Path(__file__).resolve() - root = here.parents[2] - - profile_name = "deep" if args.deep else "quick" - report, code = run_doctor( - root=root, - profile=profile_name, - component=args.component, - fixture_path=args.fixture.resolve() if args.fixture else None, - catalogue_path=args.catalogue.resolve() if args.catalogue else None, - invariants_path=args.invariants.resolve() if args.invariants else None, - ) - - if args.json: - # JSON only on stdout - sys.stdout.write(json.dumps(report.to_dict(), indent=2, sort_keys=False) + "\n") - else: - sys.stdout.write(format_human(report)) - - return code - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/mainframe_doctor/catalogue.py b/scripts/mainframe_doctor/catalogue.py deleted file mode 100644 index e065cf2..0000000 --- a/scripts/mainframe_doctor/catalogue.py +++ /dev/null @@ -1,123 +0,0 @@ -"""Load and validate doctor catalogue against required-invariants manifest.""" - -from __future__ import annotations - -import json -from dataclasses import dataclass -from pathlib import Path -from typing import Any - - -REQUIRED_CHECK_FIELDS = ( - "id", - "owner", - "subsystem", - "layer", - "provider", - "required", - "mutates", - "isolation", - "timeout_seconds", - "skip_policy", - "pass_condition", - "remediation", - "safe_fix_available", - "contract_version", -) - - -@dataclass -class CatalogueLoad: - catalogue: dict[str, Any] - required_ids: list[str] - checks_by_id: dict[str, dict[str, Any]] - errors: list[str] - - @property - def ok(self) -> bool: - return not self.errors - - -def load_json(path: Path) -> Any: - return json.loads(path.read_text(encoding="utf-8")) - - -def validate_catalogue( - catalogue: dict[str, Any], - required_ids: list[str], -) -> list[str]: - errors: list[str] = [] - checks = catalogue.get("checks") - if not isinstance(checks, list) or not checks: - return ["catalogue.checks must be a non-empty list"] - - seen: set[str] = set() - for i, check in enumerate(checks): - if not isinstance(check, dict): - errors.append(f"checks[{i}] is not an object") - continue - for field in REQUIRED_CHECK_FIELDS: - if field not in check: - errors.append(f"checks[{i}] missing field {field}") - cid = check.get("id") - if not isinstance(cid, str) or not cid: - errors.append(f"checks[{i}] missing id") - continue - if cid in seen: - errors.append(f"duplicate check id: {cid}") - seen.add(cid) - if check.get("mutates") is True: - errors.append(f"{cid}: mutates=true not allowed in doctor catalogue") - provider = check.get("provider") - if not provider: - errors.append(f"{cid}: provider required") - timeout_seconds = check.get("timeout_seconds") - if ( - isinstance(timeout_seconds, bool) - or not isinstance(timeout_seconds, int) - or timeout_seconds <= 0 - ): - errors.append(f"{cid}: timeout_seconds must be a positive integer") - # Reporter-as-check guard: providers named *report* without threshold forbidden - if isinstance(provider, str) and provider.endswith("_report") and not check.get("threshold"): - errors.append(f"{cid}: reporter provider requires threshold adapter") - - # Completeness: every required invariant id must appear in catalogue - missing = [rid for rid in required_ids if rid not in seen] - for rid in missing: - errors.append(f"catalogue missing required invariant id: {rid}") - - # Extra catalogue ids are allowed (forward-looking), but warn-as-error for empty owner - for cid, check in ((c.get("id"), c) for c in checks if isinstance(c, dict)): - if cid and not check.get("owner"): - errors.append(f"{cid}: owner required") - - return errors - - -def load_catalogue_pair( - catalogue_path: Path, - invariants_path: Path, -) -> CatalogueLoad: - errors: list[str] = [] - try: - catalogue = load_json(catalogue_path) - except Exception as exc: # noqa: BLE001 — surface as config failure - return CatalogueLoad({}, [], {}, [f"catalogue load failed: {exc}"]) - try: - inv = load_json(invariants_path) - except Exception as exc: # noqa: BLE001 - return CatalogueLoad(catalogue, [], {}, [f"required-invariants load failed: {exc}"]) - - required_ids = inv.get("required_check_ids") - if not isinstance(required_ids, list) or not all(isinstance(x, str) for x in required_ids): - errors.append("required-invariants.required_check_ids must be a string list") - required_ids = [] - - errors.extend(validate_catalogue(catalogue, required_ids)) - checks_by_id = { - c["id"]: c - for c in catalogue.get("checks", []) - if isinstance(c, dict) and isinstance(c.get("id"), str) - } - return CatalogueLoad(catalogue, list(required_ids), checks_by_id, errors) diff --git a/scripts/mainframe_doctor/providers.py b/scripts/mainframe_doctor/providers.py deleted file mode 100644 index d7341fa..0000000 --- a/scripts/mainframe_doctor/providers.py +++ /dev/null @@ -1,1140 +0,0 @@ -"""Non-mutating check providers for mainframe-doctor.""" - -from __future__ import annotations - -import hashlib -import json -import re -import signal -import sqlite3 -import stat -import subprocess -import sys -import time -from pathlib import Path -from typing import Any, Callable -from urllib.parse import quote - -from mainframe_doctor.schema import CheckResult, CheckStatus, Severity, redact_secret_shaped, utc_now_rfc3339 - - -ProviderFn = Callable[[dict[str, Any], dict[str, Any]], CheckResult] - - -class ProviderTimeoutError(TimeoutError): - """Raised when a doctor provider exceeds its catalogue time budget.""" - - -def _result( - check: dict[str, Any], - status: CheckStatus, - *, - observed: str, - message: str, - severity: Severity | None = None, - evidence_refs: list[str] | None = None, - duration_ms: int = 0, -) -> CheckResult: - sev: Severity - if severity is not None: - sev = severity - elif status == "fail": - sev = "high" - elif status in ("unknown", "stale"): - sev = "medium" - elif status == "warn": - sev = "low" - else: - sev = "info" - return CheckResult( - id=check["id"], - subsystem=check.get("subsystem", "unknown"), - layer=check.get("layer", "unknown"), - status=status, - severity=sev, - required=bool(check.get("required", True)), - expected=str(check.get("pass_condition", "")), - observed=redact_secret_shaped(observed)[:500], - observed_at=utc_now_rfc3339(), - freshness_seconds=check.get("freshness_seconds"), - authority=str(check.get("owner", "")), - evidence_refs=evidence_refs or [], - message=redact_secret_shaped(message)[:500], - remediation=str(check.get("remediation", "")), - safe_fix_available=bool(check.get("safe_fix_available", False)), - duration_ms=duration_ms, - proves=str(check.get("pass_condition", "")), - does_not_prove="Outcome value, income benefit, or unreproduced live claims", - ) - - -def provider_unimplemented(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - # Optional checks with skip_policy=skip may skip rather than force unknown - if not check.get("required", True) and str(check.get("skip_policy", "")).lower() == "skip": - return _result( - check, - "skip", - observed="provider not implemented; optional skip_policy=skip", - message=f"{check['id']}: skipped (unimplemented optional)", - severity="info", - ) - return _result( - check, - "unknown", - observed="provider not implemented in Unit 1.3 shell", - message=f"{check['id']}: explicit unknown/unimplemented", - severity="medium", - ) - - -def provider_fixture_override(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult | None: - """If fixture defines this check id, return that result.""" - fixture = ctx.get("fixture") or {} - overrides = fixture.get("checks") or {} - if check["id"] not in overrides: - return None - o = overrides[check["id"]] - status = o.get("status", "unknown") - if status not in ("pass", "warn", "fail", "stale", "unknown", "skip"): - status = "unknown" - return _result( - check, - status, - observed=str(o.get("observed", "fixture override")), - message=str(o.get("message", "from fixture")), - severity=o.get("severity"), - evidence_refs=list(o.get("evidence_refs") or ["fixture"]), - ) - - -def provider_auth_focus(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - root: Path = ctx["root"] - try: - from focus_authority import load_focus - - loaded = load_focus(root, now=ctx.get("now")) - except Exception as exc: # noqa: BLE001 - return _result( - check, - "unknown", - observed=type(exc).__name__, - message="focus loader failed", - ) - if loaded.path is None: - return _result( - check, - "fail", - observed="20_live/focus/current.{yaml,json} absent", - message="structured focus authority not present (ADR-044 / MPE-024)", - severity="high", - evidence_refs=["20_live/focus/"], - ) - if not loaded.ok: - return _result( - check, - "fail", - observed="; ".join(loaded.errors)[:300], - message="focus authority failed validation", - severity="high", - evidence_refs=[str(loaded.path.name)], - ) - if loaded.warnings: - return _result( - check, - "stale" if any("stale" in w or "past" in w for w in loaded.warnings) else "warn", - observed=f"project={loaded.primary_project}; " + "; ".join(loaded.warnings), - message="focus present with warnings", - severity="low", - evidence_refs=[str(loaded.path.name)], - ) - return _result( - check, - "pass", - observed=f"project={loaded.primary_project} revision={loaded.revision}", - message="structured focus authority parses within review window", - evidence_refs=[str(loaded.path.name)], - ) - - -def provider_sched_service(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - """SCHED-001: launchd present + last run healthy with provenance preference.""" - root: Path = ctx["root"] - try: - sys.path.insert(0, str(root / "scripts")) - import eval_schedule as es # type: ignore - - health = es.assess_schedule_health() - except Exception as exc: # noqa: BLE001 - return _result(check, "unknown", observed=type(exc).__name__, message="schedule assess failed") - if health.problems: - return _result( - check, - "fail", - observed="; ".join(health.problems)[:300], - message="scheduler hard problems present", - severity="high", - ) - if health.degraded: - return _result( - check, - "fail", - observed="; ".join(health.degraded)[:300], - message="scheduler degraded (no launchd-proven run)", - severity="high", - ) - return _result( - check, - "pass", - observed=( - f"weekly={health.last_weekly and health.last_weekly.get('run_id')} " - f"trigger={health.last_weekly and health.last_weekly.get('trigger')}" - ), - message="scheduler service + provenance ok", - ) - - -def provider_sched_provenance(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - """SCHED-002: latest scheduled success came from launchd, not manual substitute.""" - root: Path = ctx["root"] - try: - sys.path.insert(0, str(root / "scripts")) - import eval_schedule as es # type: ignore - - health = es.assess_schedule_health() - except Exception as exc: # noqa: BLE001 - return _result(check, "unknown", observed=type(exc).__name__, message="schedule assess failed") - last = health.last_weekly - if not last: - return _result(check, "fail", observed="no weekly run", message="no scheduled weekly run") - trigger = last.get("trigger") - if trigger == "launchd" and last.get("all_passed"): - return _result( - check, - "pass", - observed=f"run_id={last.get('run_id')} trigger=launchd", - message="latest weekly success has launchd provenance", - ) - return _result( - check, - "fail", - observed=f"run_id={last.get('run_id')} trigger={trigger!r} all_passed={last.get('all_passed')}", - message="latest weekly is not a launchd-proven success", - severity="high", - ) - - -def provider_task_manifest_quarantine(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - root: Path = ctx["root"] - path = root / "30_projects" / "tasks_manifest.json" - if not path.exists(): - return _result(check, "fail", observed="tasks_manifest.json missing", message="no task projection") - try: - payload = json.loads(path.read_text(encoding="utf-8")) - except Exception as exc: # noqa: BLE001 - return _result(check, "fail", observed=type(exc).__name__, message="manifest unreadable") - q = payload.get("quarantine") if isinstance(payload, dict) else None - if not isinstance(q, dict) or q.get("status") != "quarantined": - return _result( - check, - "fail", - observed="quarantine.status missing or not quarantined", - message="task projection not quarantined — may be mistaken for executable work", - severity="critical", - evidence_refs=["30_projects/tasks_manifest.json"], - ) - tasks = payload.get("tasks") or [] - executable_true = sum(1 for t in tasks if isinstance(t, dict) and t.get("executable") is True) - if executable_true: - return _result( - check, - "fail", - observed=f"executable=true count={executable_true}", - message="quarantined manifest still marks rows executable", - severity="high", - ) - return _result( - check, - "pass", - observed=f"quarantined n={len(tasks)} baseline={str(q.get('baseline_sha256') or '')[:12]}…", - message="task projection quarantined; not executable authority", - evidence_refs=["30_projects/tasks_manifest.json"], - ) - - -def _detect_project_from_state(state_path: Path) -> str | None: - if not state_path.exists(): - return None - found = False - for line in state_path.read_text(encoding="utf-8", errors="replace").splitlines(): - if line.strip().lower() == "## active project": - found = True - continue - if found: - value = line.strip() - if value and not value.startswith("#"): - return value - if value.startswith("#"): - break - return None - - -def provider_session_project_path(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - root: Path = ctx["root"] - state = root / "STATE.md" - project = _detect_project_from_state(state) - if not project: - return _result( - check, - "fail", - observed="no Active Project in STATE.md", - message="session selected project missing", - severity="high", - ) - # slug path: first token if simple, else whole line as slug folder attempt - slug = project.strip() - # compound narrative → almost certainly invalid as single path segment - if " + " in slug or "(" in slug: - bad = root / "30_projects" / re.sub(r"[^\w.\-]+", "-", slug.lower()).strip("-") / "README.md" - return _result( - check, - "fail", - observed=f"compound STATE project label; path would be missing ({bad.name}…)", - message="session-open false-green class: compound focus cannot resolve project path", - severity="critical", - evidence_refs=["STATE.md"], - ) - readme = root / "30_projects" / slug / "README.md" - if readme.exists(): - return _result( - check, - "pass", - observed=f"30_projects/{slug}/README.md exists", - message="selected project path resolves", - evidence_refs=[f"30_projects/{slug}/README.md"], - ) - return _result( - check, - "fail", - observed=f"30_projects/{slug}/README.md missing", - message="selected project path does not exist", - severity="critical", - evidence_refs=["STATE.md"], - ) - - -def provider_project_index_check(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - root: Path = ctx["root"] - script = root / "bin" / "sync-project-index" - if not script.exists(): - return _result(check, "unknown", observed="bin/sync-project-index missing", message="tool absent") - t0 = time.perf_counter() - try: - proc = subprocess.run( - [str(script), "--check"], - cwd=str(root), - capture_output=True, - text=True, - timeout=min(int(check.get("timeout_seconds") or 10), 30), - check=False, - ) - except Exception as exc: # noqa: BLE001 - return _result(check, "unknown", observed=str(type(exc).__name__), message="index check failed to run") - ms = int((time.perf_counter() - t0) * 1000) - out = (proc.stdout or "") + (proc.stderr or "") - if proc.returncode == 0 and "current" in out.lower(): - return _result( - check, - "pass", - observed="sync-project-index --check current", - message="generated index matches authorities (parser rules)", - duration_ms=ms, - ) - return _result( - check, - "fail", - observed=f"exit={proc.returncode}; {(out.strip().splitlines() or [''])[0][:200]}", - message="project index stale or invalid", - severity="high", - duration_ms=ms, - ) - - -def provider_eval_registry_strict(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - root: Path = ctx["root"] - script = root / "bin" / "eval-registry" - if not script.exists(): - return _result(check, "unknown", observed="bin/eval-registry missing", message="tool absent") - t0 = time.perf_counter() - try: - proc = subprocess.run( - [str(script), "check", "--strict"], - cwd=str(root), - capture_output=True, - text=True, - timeout=min(int(check.get("timeout_seconds") or 15), 60), - check=False, - ) - except Exception as exc: # noqa: BLE001 - return _result(check, "unknown", observed=str(type(exc).__name__), message="registry check failed to run") - ms = int((time.perf_counter() - t0) * 1000) - out = (proc.stdout or "") + (proc.stderr or "") - # count problems without dumping paths that might be sensitive — keep count only - m = re.search(r"(\d+)\s+problem", out) - n = int(m.group(1)) if m else (0 if proc.returncode == 0 else -1) - if proc.returncode == 0: - return _result( - check, - "pass", - observed="0 strict problems", - message="eval-registry strict clean", - duration_ms=ms, - ) - return _result( - check, - "fail", - observed=f"strict problems={n if n >= 0 else 'unknown'}; exit={proc.returncode}", - message="eval-registry strict failures remain", - severity="medium", - duration_ms=ms, - ) - - -def _open_sqlite_read_only(path: Path) -> sqlite3.Connection: - """Open a SQLite snapshot without requiring write access to its directory.""" - resolved = path.expanduser().resolve() - base_uri = f"file:{quote(str(resolved), safe='/')}?mode=ro" - last_error: Exception | None = None - for immutable in (False, True): - con: sqlite3.Connection | None = None - uri = f"{base_uri}&immutable=1" if immutable else base_uri - try: - con = sqlite3.connect(uri, uri=True) - con.execute("PRAGMA query_only = ON") - con.execute("SELECT 1 FROM sqlite_master LIMIT 1").fetchone() - return con - except sqlite3.Error as exc: - last_error = exc - if con is not None: - con.close() - assert last_error is not None - raise last_error - - -def provider_mg_db_presence(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - home = Path.home() / ".mindgraph" - knowledge = home / "mainframe.sqlite" - projects = home / "mainframe-projects.sqlite" - missing = [p.name for p in (knowledge, projects) if not p.exists() or p.stat().st_size < 10_000] - if missing: - return _result( - check, - "fail", - observed=f"missing or stub-like: {', '.join(missing)}", - message="installed MindGraph DBs not healthy by size/presence", - severity="high", - ) - # Schema-sniff both canonical stores. Presence of one healthy DB must not - # make the pair look green. - need = {"documents", "documents_fts", "chunks", "vec_chunks", "edges"} - for label, path in (("knowledge", knowledge), ("projects", projects)): - try: - con = _open_sqlite_read_only(path) - try: - tables = { - r[0] - for r in con.execute( - "SELECT name FROM sqlite_master " - "WHERE type IN ('table', 'virtual table')" - ) - } - finally: - con.close() - except Exception as exc: # noqa: BLE001 - return _result( - check, - "unknown", - observed=f"{label} open failed: {type(exc).__name__}", - message=f"could not inspect {label} DB", - ) - missing_tables = sorted(need - tables) - if missing_tables: - return _result( - check, - "fail", - observed=f"{label} tables missing {missing_tables}", - message=f"{label} schema incomplete", - severity="high", - ) - return _result( - check, - "pass", - observed="~/.mindgraph mainframe + projects DB schemas valid", - message="both DB presence/schema sniffs ok (not freshness/coverage)", - ) - - -def provider_mg_manifest_coverage(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - root: Path = ctx["root"] - man_path = root / "30_projects" / "mindgraph-projects.json" - projects_dir = root / "30_projects" - if not man_path.exists(): - return _result(check, "fail", observed="mindgraph-projects.json missing", message="no coverage policy file") - try: - man = json.loads(man_path.read_text(encoding="utf-8")) - names = man.get("projects") or [] - except Exception as exc: # noqa: BLE001 - return _result(check, "fail", observed=type(exc).__name__, message="manifest unreadable") - real = sorted( - p.name - for p in projects_dir.iterdir() - if p.is_dir() and not p.name.startswith(".") and (p / "README.md").exists() - ) - set_m, set_r = set(names), set(real) - missing_dirs = sorted(set_m - set_r) - omitted = sorted(set_r - set_m) - # active omitted is worse - active_omitted = [] - for slug in omitted: - text = (projects_dir / slug / "README.md").read_text(encoding="utf-8", errors="replace")[:1500] - if re.search(r'^project_state:\s*["\']?active', text, re.M): - active_omitted.append(slug) - if missing_dirs or active_omitted or omitted: - return _result( - check, - "fail", - observed=( - f"manifest={len(names)} real={len(real)} " - f"missing_dirs={len(missing_dirs)} omitted={len(omitted)} " - f"active_omitted={active_omitted}" - ), - message="project MindGraph coverage incomplete or stale", - severity="high", - evidence_refs=["30_projects/mindgraph-projects.json"], - ) - - installed = Path.home() / ".mindgraph" / "mainframe-projects.sqlite" - if not installed.exists(): - return _result( - check, - "fail", - observed="installed projects DB missing", - message="manifest is complete but installed project index is absent", - severity="high", - evidence_refs=["30_projects/mindgraph-projects.json"], - ) - try: - con = _open_sqlite_read_only(installed) - try: - installed_namespaces = { - str(row[0]) - for row in con.execute( - "SELECT DISTINCT namespace FROM documents " - "WHERE namespace IS NOT NULL AND namespace != ''" - ) - } - finally: - con.close() - except Exception as exc: # noqa: BLE001 - return _result( - check, - "unknown", - observed=f"installed projects DB unreadable: {type(exc).__name__}", - message="could not verify installed project namespaces", - severity="high", - evidence_refs=["~/.mindgraph/mainframe-projects.sqlite"], - ) - - expected_namespaces = set(names) - missing_namespaces = sorted(expected_namespaces - installed_namespaces) - extra_namespaces = sorted(installed_namespaces - expected_namespaces) - if missing_namespaces or extra_namespaces: - return _result( - check, - "fail", - observed=( - f"manifest_namespaces={len(expected_namespaces)} " - f"installed_namespaces={len(installed_namespaces)} " - f"missing={missing_namespaces} extra={extra_namespaces}" - ), - message="installed project MindGraph namespaces are stale", - severity="high", - evidence_refs=[ - "30_projects/mindgraph-projects.json", - "~/.mindgraph/mainframe-projects.sqlite", - ], - ) - return _result( - check, - "pass", - observed=f"manifest and installed DB cover all {len(real)} project namespaces", - message="manifest and installed namespace coverage ok", - ) - - -def provider_cli_mutation_help_safety(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - """Source inspection only — never invoke --help on unsafe wrappers.""" - root: Path = ctx["root"] - path = root / "bin" / "mindgraph-refresh-projects" - if not path.exists(): - return _result(check, "unknown", observed="wrapper missing", message="cannot assess") - src = path.read_text(encoding="utf-8", errors="replace") - has_help = bool( - re.search(r"--help\|-h|-h\|--help", src) - or re.search(r"usage\(\)|Show this help", src) - or ("--help)" in src and "case" in src) - ) - has_unknown_fail = bool( - re.search(r"unknown option|unknown flag", src, re.I) - or ('-*)' in src and "exit 2" in src) - ) - requires_apply = "--apply" in src and "refusing to mutate without --apply" in src - # Legacy unsafe: only leading --dry-run accepted, no help, falls through to ingest - legacy_unsafe = ( - re.search(r'\$\{1:-\}.*"--dry-run"', src) is not None - and not has_help - and "--apply" not in src - ) - if legacy_unsafe: - return _result( - check, - "fail", - observed="no --help handler; non-dry-run first args fall through to ingest", - message="bin/mindgraph-refresh-projects violates help/unknown fail-closed (static proof)", - severity="critical", - evidence_refs=["bin/mindgraph-refresh-projects"], - ) - if has_help and has_unknown_fail and requires_apply: - return _result( - check, - "pass", - observed="help + unknown fail-closed + --apply mutation boundary present", - message="CLI help safety contract present (static Unit 2.1)", - evidence_refs=["bin/mindgraph-refresh-projects"], - ) - if has_help and has_unknown_fail: - return _result( - check, - "warn", - observed="help/unknown present; --apply boundary unclear", - message="partial CLI safety", - severity="low", - ) - return _result( - check, - "fail", - observed=f"has_help={has_help} unknown_fail={has_unknown_fail} apply={requires_apply}", - message="CLI help/unknown/apply contract incomplete", - severity="high", - evidence_refs=["bin/mindgraph-refresh-projects"], - ) - - -def provider_ws_db_presence(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - root: Path = ctx["root"] - db = root / "20_live" / "workstation" / "workstation.sqlite" - if not db.exists(): - return _result(check, "fail", observed="workstation.sqlite missing", message="storage preflight fail") - try: - con = sqlite3.connect(f"file:{db}?mode=ro", uri=True) - tables = {r[0] for r in con.execute("SELECT name FROM sqlite_master WHERE type='table'")} - con.close() - except Exception as exc: # noqa: BLE001 - return _result(check, "fail", observed=type(exc).__name__, message="cannot open workstation db") - if "tasks" not in tables and "runs" not in tables: - return _result(check, "fail", observed=f"tables={sorted(tables)[:12]}", message="unexpected schema") - return _result( - check, - "pass", - observed=f"db size={db.stat().st_size} tables={len(tables)}", - message="workstation storage present (not operational fullness)", - ) - - -def provider_ws_operational_rows(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - root: Path = ctx["root"] - db = root / "20_live" / "workstation" / "workstation.sqlite" - if not db.exists(): - return _result(check, "unknown", observed="db missing", message="skip operational row probe") - try: - con = sqlite3.connect(f"file:{db}?mode=ro", uri=True) - runs = con.execute("SELECT COUNT(*) FROM runs").fetchone()[0] if _has_table(con, "runs") else 0 - approvals = ( - con.execute("SELECT COUNT(*) FROM approvals").fetchone()[0] if _has_table(con, "approvals") else 0 - ) - artifacts = ( - con.execute("SELECT COUNT(*) FROM artifacts").fetchone()[0] if _has_table(con, "artifacts") else 0 - ) - tasks = con.execute("SELECT COUNT(*) FROM tasks").fetchone()[0] if _has_table(con, "tasks") else 0 - con.close() - except Exception as exc: # noqa: BLE001 - return _result(check, "unknown", observed=type(exc).__name__, message="row probe failed") - if runs == 0 and approvals == 0 and artifacts == 0 and tasks > 0: - return _result( - check, - "fail", - observed=f"tasks={tasks} runs={runs} approvals={approvals} artifacts={artifacts}", - message="projection full of tasks but zero operational runs/approvals/artifacts", - severity="high", - ) - if runs == 0 and tasks == 0: - return _result( - check, - "warn", - observed="empty operational and task tables", - message="workstation empty", - ) - return _result( - check, - "pass", - observed=f"tasks={tasks} runs={runs} approvals={approvals} artifacts={artifacts}", - message="operational rows present or explicitly empty with no false-full tasks", - ) - - -def _has_table(con: sqlite3.Connection, name: str) -> bool: - row = con.execute( - "SELECT 1 FROM sqlite_master WHERE type='table' AND name=?", - (name,), - ).fetchone() - return row is not None - - -def _parse_readme_frontmatter(path: Path) -> dict[str, str]: - if not path.exists(): - return {} - text = path.read_text(encoding="utf-8", errors="replace") - if not text.startswith("---\n"): - return {} - end = text.find("\n---", 4) - if end == -1: - return {} - meta: dict[str, str] = {} - for line in text[4:end].splitlines(): - if ":" not in line: - continue - key, value = line.split(":", 1) - meta[key.strip()] = value.strip().strip("\"'") - return meta - - -_VALID_PROJECT_STATES = { - "active", - "paused", - "planned", - "blocked", - "suspended", - "shipped", - "trashed", -} - - -def provider_session_phase_alignment(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - """SESSION-002: selected project state and re-entry pointer agree.""" - root: Path = ctx["root"] - project: str | None = None - source = "none" - try: - from focus_authority import load_focus - - loaded = load_focus(root) - if loaded.ok and loaded.primary_project: - project = loaded.primary_project - source = "focus" - except Exception: - loaded = None # type: ignore[assignment] - - if not project: - project = _detect_project_from_state(root / "STATE.md") - source = "STATE.md" if project else source - - if not project: - return _result( - check, - "fail", - observed="no focus primary or STATE Active Project", - message="session project selection missing for phase alignment", - severity="high", - ) - - if " + " in project or "(" in project: - return _result( - check, - "fail", - observed=f"compound project label from {source}", - message="phase alignment cannot use compound focus labels", - severity="high", - evidence_refs=["STATE.md"], - ) - - readme = root / "30_projects" / project / "README.md" - if not readme.exists(): - return _result( - check, - "fail", - observed=f"30_projects/{project}/README.md missing", - message="selected project path does not exist for phase alignment", - severity="critical", - evidence_refs=[f"30_projects/{project}/"], - ) - - meta = _parse_readme_frontmatter(readme) - state = (meta.get("project_state") or meta.get("status") or "").strip() - next_action = (meta.get("next_action") or "").strip() - if not state: - return _result( - check, - "fail", - observed=f"{project}: missing project_state", - message="selected project lacks lifecycle state", - severity="high", - evidence_refs=[f"30_projects/{project}/README.md"], - ) - if state not in _VALID_PROJECT_STATES: - return _result( - check, - "fail", - observed=f"{project}: invalid project_state={state!r}", - message="selected project state not in ADR-041 vocabulary", - severity="high", - evidence_refs=[f"30_projects/{project}/README.md"], - ) - - needs_reentry = state in ("active", "paused", "blocked", "suspended") - if needs_reentry and not next_action: - return _result( - check, - "fail", - observed=f"{project}: state={state} without next_action", - message="active/paused/blocked/suspended project missing re-entry pointer", - severity="high", - evidence_refs=[f"30_projects/{project}/README.md"], - ) - - return _result( - check, - "pass", - observed=f"source={source} project={project} state={state} next_action={'set' if next_action else 'n/a'}", - message="selected project state and re-entry pointer agree", - evidence_refs=[f"30_projects/{project}/README.md"], - ) - - -def provider_structure_bounds(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - """STRUCT-001: required structural contracts and private/public bounds present.""" - root: Path = ctx["root"] - required = [ - "AGENTS.md", - "HARNESS.md", - "DECISIONS.md", - "EPISTEMIC_STANCE.md", - "20_live/AGENTS.md", - "30_projects/AGENTS.md", - ".context/primitives.md", - ] - missing = [rel for rel in required if not (root / rel).exists()] - if missing: - return _result( - check, - "fail", - observed=f"missing={missing[:6]}", - message="required structural contracts missing", - severity="high", - evidence_refs=missing[:4], - ) - - # Private live zone must not be a symlink out of tree into a public export root. - live = root / "20_live" - issues: list[str] = [] - if live.is_symlink(): - issues.append("20_live is symlink") - public = root / "30_projects" / "mainframe-public-portfolio" - if public.exists(): - # flag only direct 20_live path embeds in public README (not deep scan) - pub_readme = public / "README.md" - if pub_readme.exists(): - blob = pub_readme.read_text(encoding="utf-8", errors="replace") - if re.search(r"(?m)20_live/|/Users/.*/20_live", blob): - issues.append("public portfolio README references 20_live paths") - - if issues: - return _result( - check, - "fail", - observed="; ".join(issues), - message="private/public structural boundary issues", - severity="high", - evidence_refs=["20_live/", "30_projects/mainframe-public-portfolio/README.md"], - ) - - return _result( - check, - "pass", - observed=f"structural contracts present n={len(required)}; private bounds ok", - message="structural links and private/public boundaries valid (contract presence)", - evidence_refs=["AGENTS.md", "20_live/AGENTS.md"], - ) - - -def _telemetry_compute_hash(event_data: dict[str, Any], prev_hash: str) -> str: - event_str = json.dumps(event_data, sort_keys=True, separators=(",", ":")) - return hashlib.sha256((event_str + prev_hash).encode("utf-8")).hexdigest()[:16] - - -def provider_tel_hash_integrity(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - """TEL-001: recent event store parses and intra-file hash chains hold.""" - root: Path = ctx["root"] - events_dir = root / "20_live" / "workflow-metrics" / "events" - if not events_dir.is_dir(): - return _result( - check, - "fail", - observed="20_live/workflow-metrics/events missing", - message="telemetry event store absent", - severity="high", - evidence_refs=["20_live/workflow-metrics/events"], - ) - - files = sorted(events_dir.glob("*.jsonl")) - if not files: - return _result( - check, - "fail", - observed="no event jsonl files", - message="telemetry event store empty", - severity="high", - ) - - # Operational health: verify the newest day file fully (bounded), not all history. - newest = files[-1] - parsed = 0 - chain_ok = 0 - chain_bad = 0 - parse_errors = 0 - prev_hash: str | None = None - try: - for line in newest.read_text(encoding="utf-8", errors="replace").splitlines(): - if not line.strip(): - continue - try: - event = json.loads(line) - except json.JSONDecodeError: - parse_errors += 1 - prev_hash = None - continue - if not isinstance(event, dict): - parse_errors += 1 - prev_hash = None - continue - parsed += 1 - h = event.get("hash_chain") - if not isinstance(h, str) or not h: - chain_bad += 1 - prev_hash = None - continue - body = {k: v for k, v in event.items() if k != "hash_chain"} - if prev_hash is None: - # First event or post-error resync — accept as chain anchor. - prev_hash = h - continue - expected = _telemetry_compute_hash(body, prev_hash) - if expected == h: - chain_ok += 1 - else: - chain_bad += 1 - prev_hash = h - except OSError as exc: - return _result( - check, - "unknown", - observed=type(exc).__name__, - message="could not read telemetry event file", - ) - - observed = ( - f"file={newest.name} parsed={parsed} chain_ok={chain_ok} " - f"chain_bad={chain_bad} parse_errors={parse_errors}" - ) - if parse_errors: - return _result( - check, - "fail", - observed=observed, - message="telemetry event store has unreadable records", - severity="high", - evidence_refs=[f"20_live/workflow-metrics/events/{newest.name}"], - ) - if parsed == 0: - return _result( - check, - "warn", - observed=observed, - message="newest telemetry day file is empty", - severity="low", - ) - if chain_bad: - return _result( - check, - "fail", - observed=observed, - message="telemetry hash-chain breaks in newest day file", - severity="high", - evidence_refs=[f"20_live/workflow-metrics/events/{newest.name}"], - ) - return _result( - check, - "pass", - observed=observed, - message="newest telemetry day parses with intact hash chain", - evidence_refs=[f"20_live/workflow-metrics/events/{newest.name}"], - ) - - -def provider_sec_secret_store(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - """Presence/mode only — never print secret values.""" - root: Path = ctx["root"] - # operator disposition can force pass via fixture or marker file - disposition = root / "20_live" / "security" / "SEC-001-disposition.md" - mcp = root / ".mcp.json" - issues: list[str] = [] - if mcp.exists(): - mode = mcp.stat().st_mode - if mode & stat.S_IROTH: - issues.append("mcp_world_readable") - if mode & stat.S_IRGRP: - issues.append("mcp_group_readable") - try: - data = json.loads(mcp.read_text(encoding="utf-8")) - blob = json.dumps(data) - # detect secret-shaped keys without values - if re.search(r"(?i)client_secret|api_key|refresh_token", blob): - # check if values look like placeholders - def walk(o: Any, path: str = "") -> list[str]: - hits: list[str] = [] - if isinstance(o, dict): - for k, v in o.items(): - p = f"{path}.{k}" if path else k - if re.search(r"(?i)secret|token|password|api_key", k) and isinstance(v, str): - if v and not v.startswith("${") and "REDACTED" not in v and len(v) > 8: - hits.append(p) - hits.extend(walk(v, p)) - elif isinstance(o, list): - for i, v in enumerate(o[:20]): - hits.extend(walk(v, f"{path}[{i}]")) - return hits - - embedded = walk(data) - if embedded: - issues.append(f"embedded_secret_fields={len(embedded)}") - except Exception: - issues.append("mcp_unreadable") - else: - issues.append("mcp_absent") - - # disposition file can record accepted residual risk - accepted = False - if disposition.exists(): - text = disposition.read_text(encoding="utf-8", errors="replace") - if re.search(r"(?i)status:\s*accepted|disposition:\s*accepted|residual.?risk.?accepted", text): - accepted = True - - if not issues: - return _result(check, "pass", observed="no world-readable secret-shaped stores found", message="SEC-001 clean") - if accepted: - return _result( - check, - "warn", - observed=f"issues={issues}; operator disposition accepted", - message="SEC-001 residual risk accepted by operator disposition", - severity="low", - evidence_refs=["20_live/security/SEC-001-disposition.md"], - ) - return _result( - check, - "fail", - observed=f"issues={issues}", - message="secret store mode or embedding not approved", - severity="high", - evidence_refs=[".mcp.json"], - ) - - -PROVIDERS: dict[str, ProviderFn] = { - "unimplemented": provider_unimplemented, - "auth_focus": provider_auth_focus, - "session_project_path": provider_session_project_path, - "session_phase_alignment": provider_session_phase_alignment, - "project_index_check": provider_project_index_check, - "eval_registry_strict": provider_eval_registry_strict, - "mg_db_presence": provider_mg_db_presence, - "mg_manifest_coverage": provider_mg_manifest_coverage, - "cli_mutation_help_safety": provider_cli_mutation_help_safety, - "ws_db_presence": provider_ws_db_presence, - "ws_operational_rows": provider_ws_operational_rows, - "structure_bounds": provider_structure_bounds, - "tel_hash_integrity": provider_tel_hash_integrity, - "sec_secret_store": provider_sec_secret_store, - "task_manifest_quarantine": provider_task_manifest_quarantine, - "sched_service": provider_sched_service, - "sched_provenance": provider_sched_provenance, -} - - -def run_provider(check: dict[str, Any], ctx: dict[str, Any]) -> CheckResult: - # Fixture overrides always win when present - overridden = provider_fixture_override(check, ctx) - if overridden is not None: - return overridden - - name = check.get("provider") or "unimplemented" - fn = PROVIDERS.get(str(name), provider_unimplemented) - if check.get("mutates") is True: - return _result( - check, - "fail", - observed="mutates=true", - message="doctor refused mutating provider", - severity="critical", - ) - t0 = time.perf_counter() - timeout_seconds = int(check.get("timeout_seconds") or 10) - previous_handler = signal.getsignal(signal.SIGALRM) - previous_timer = signal.getitimer(signal.ITIMER_REAL) - - def timeout_handler(_signum: int, _frame: object) -> None: - raise ProviderTimeoutError - - try: - signal.signal(signal.SIGALRM, timeout_handler) - signal.setitimer(signal.ITIMER_REAL, timeout_seconds) - try: - result = fn(check, ctx) - finally: - signal.setitimer(signal.ITIMER_REAL, 0) - signal.signal(signal.SIGALRM, previous_handler) - previous_delay, previous_interval = previous_timer - if previous_delay > 0: - elapsed = time.perf_counter() - t0 - signal.setitimer( - signal.ITIMER_REAL, - max(previous_delay - elapsed, 0.000001), - previous_interval, - ) - except ProviderTimeoutError: - result = _result( - check, - "unknown", - observed=f"timeout_seconds={timeout_seconds}", - message=f"provider timed out after {timeout_seconds}s", - severity="high", - ) - except Exception as exc: # noqa: BLE001 — provider fault → unknown/fail - result = _result( - check, - "unknown", - observed=type(exc).__name__, - message=f"provider exception: {type(exc).__name__}", - severity="high", - ) - if result.duration_ms == 0: - result.duration_ms = int((time.perf_counter() - t0) * 1000) - return result diff --git a/scripts/mainframe_doctor/runner.py b/scripts/mainframe_doctor/runner.py deleted file mode 100644 index 407389d..0000000 --- a/scripts/mainframe_doctor/runner.py +++ /dev/null @@ -1,211 +0,0 @@ -"""Doctor runner: catalogue → providers → aggregate.""" - -from __future__ import annotations - -import json -import time -from pathlib import Path -from typing import Any - -from mainframe_doctor.catalogue import load_catalogue_pair -from mainframe_doctor.providers import run_provider -from mainframe_doctor.schema import ( - DoctorReport, - aggregate_health, - exit_code, - summarize, - utc_now_rfc3339, -) -from mainframe_doctor import __version__ - - -QUICK_SUBSYSTEMS = { - "authority", - "session", - "projects", - "scheduler", - "mindgraph", - "telemetry", - "structure", - "cli", - "security", - "workstation", - "evaluation", -} - -# quick profile: prefer these ids if present -QUICK_IDS = { - "AUTH-001", - "SESSION-001", - "SESSION-002", - "PROJECT-002", - "SCHED-001", - "EVAL-002", - "MG-001", - "MG-003", - "TEL-001", - "WS-001", - "WS-003", - "CLI-001", - "SEC-001", - "STRUCT-001", - "TASK-001", -} - - -def default_paths(root: Path) -> tuple[Path, Path]: - return ( - root / ".context" / "doctor" / "catalogue.json", - root / ".context" / "doctor" / "required-invariants.json", - ) - - -def load_fixture(path: Path | None) -> dict[str, Any]: - if path is None: - return {} - if path.is_dir(): - f = path / "fixture.json" - else: - f = path - if not f.exists(): - raise FileNotFoundError(f"fixture not found: {f}") - return json.loads(f.read_text(encoding="utf-8")) - - -def select_checks( - checks: list[dict[str, Any]], - *, - profile: str, - component: str | None, -) -> list[dict[str, Any]]: - if component: - return [c for c in checks if c.get("subsystem") == component] - if profile == "quick": - selected = [c for c in checks if c.get("id") in QUICK_IDS] - if selected: - return selected - return [c for c in checks if c.get("subsystem") in QUICK_SUBSYSTEMS] - # deep = all - return list(checks) - - -def run_doctor( - *, - root: Path, - profile: str = "quick", - component: str | None = None, - fixture_path: Path | None = None, - catalogue_path: Path | None = None, - invariants_path: Path | None = None, -) -> tuple[DoctorReport, int]: - t0 = time.perf_counter() - cat_path, inv_path = default_paths(root) - if catalogue_path: - cat_path = catalogue_path - if invariants_path: - inv_path = invariants_path - - loaded = load_catalogue_pair(cat_path, inv_path) - if not loaded.ok: - report = DoctorReport( - schema_version=1, - doctor_version=__version__, - profile=profile, - health="unknown", - checked_at=utc_now_rfc3339(), - duration_ms=int((time.perf_counter() - t0) * 1000), - authority_revision=None, - summary={"pass": 0, "warn": 0, "fail": 0, "stale": 0, "unknown": 0, "skip": 0}, - checks=[], - catalogue_version=None, - mode="fixture" if fixture_path else "live", - internal_error="; ".join(loaded.errors)[:1000], - ) - return report, exit_code("unknown", internal_error=True) - - try: - fixture = load_fixture(fixture_path) - except Exception as exc: # noqa: BLE001 - report = DoctorReport( - schema_version=1, - doctor_version=__version__, - profile=profile, - health="unknown", - checked_at=utc_now_rfc3339(), - duration_ms=int((time.perf_counter() - t0) * 1000), - authority_revision=None, - summary={"pass": 0, "warn": 0, "fail": 0, "stale": 0, "unknown": 0, "skip": 0}, - checks=[], - catalogue_version=str(loaded.catalogue.get("catalogue_version")), - mode="fixture", - internal_error=f"fixture load failed: {exc}", - ) - return report, exit_code("unknown", internal_error=True) - - # Fixture may point at a synthetic root - run_root = root - if fixture.get("root"): - run_root = Path(fixture["root"]) - if not run_root.is_absolute(): - # relative to fixture file dir - base = fixture_path if fixture_path and fixture_path.is_dir() else (fixture_path.parent if fixture_path else root) - run_root = (base / fixture["root"]).resolve() - - ctx: dict[str, Any] = { - "root": run_root, - "fixture": fixture, - "profile": profile, - } - - all_checks = [c for c in loaded.catalogue.get("checks", []) if isinstance(c, dict)] - selected = select_checks(all_checks, profile=profile, component=component) - - results = [] - for check in selected: - results.append(run_provider(check, ctx)) - - health = aggregate_health(results) - report = DoctorReport( - schema_version=1, - doctor_version=__version__, - profile=profile if not component else f"component:{component}", - health=health, - checked_at=utc_now_rfc3339(), - duration_ms=int((time.perf_counter() - t0) * 1000), - authority_revision=fixture.get("authority_revision"), - summary=summarize(results), - checks=results, - catalogue_version=str(loaded.catalogue.get("catalogue_version")), - mode="fixture" if fixture_path else "live", - internal_error=None, - ) - return report, exit_code(health, internal_error=False) - - -def format_human(report: DoctorReport) -> str: - lines = [ - f"mainframe-doctor {report.doctor_version} profile={report.profile} mode={report.mode}", - f"health={report.health} checked_at={report.checked_at} duration_ms={report.duration_ms}", - f"summary: {report.summary}", - ] - if report.internal_error: - lines.append(f"INTERNAL ERROR: {report.internal_error}") - return "\n".join(lines) + "\n" - - # order by severity interest: fail, unknown, stale, warn, skip, pass - order = {"fail": 0, "unknown": 1, "stale": 2, "warn": 3, "skip": 4, "pass": 5} - checks = sorted(report.checks, key=lambda c: (order.get(c.status, 9), c.id)) - lines.append("") - for c in checks: - if c.status == "pass" and report.profile.startswith("quick"): - continue # keep quick human output short; still in JSON - flag = c.status.upper() - req = "req" if c.required else "opt" - lines.append(f" [{flag}] {c.id} ({req}/{c.severity}) {c.message}") - if c.status != "pass" and c.remediation: - lines.append(f" remediation: {c.remediation}") - # always show non-pass count reminder - non_pass = [c for c in report.checks if c.status != "pass"] - lines.append("") - lines.append(f"{len(non_pass)} non-pass / {len(report.checks)} checks shown partially; use --json for full vector") - return "\n".join(lines) + "\n" diff --git a/scripts/mainframe_doctor/schema.py b/scripts/mainframe_doctor/schema.py deleted file mode 100644 index 1e3ab0e..0000000 --- a/scripts/mainframe_doctor/schema.py +++ /dev/null @@ -1,134 +0,0 @@ -"""Result schema and aggregation for mainframe-doctor.""" - -from __future__ import annotations - -from dataclasses import asdict, dataclass, field -from datetime import datetime, timezone -from typing import Any, Literal - -CheckStatus = Literal["pass", "warn", "fail", "stale", "unknown", "skip"] -Health = Literal["healthy", "degraded", "unhealthy", "unknown"] -Severity = Literal["info", "low", "medium", "high", "critical"] - -STATUS_ORDER = ("pass", "warn", "fail", "stale", "unknown", "skip") - - -@dataclass -class CheckResult: - id: str - subsystem: str - layer: str - status: CheckStatus - severity: Severity - required: bool - expected: str - observed: str - observed_at: str | None - freshness_seconds: int | None - authority: str - evidence_refs: list[str] = field(default_factory=list) - message: str = "" - remediation: str = "" - safe_fix_available: bool = False - duration_ms: int = 0 - proves: str = "" - does_not_prove: str = "" - - def to_dict(self) -> dict[str, Any]: - return asdict(self) - - -@dataclass -class DoctorReport: - schema_version: int - doctor_version: str - profile: str - health: Health - checked_at: str - duration_ms: int - authority_revision: str | None - summary: dict[str, int] - checks: list[CheckResult] - catalogue_version: str | None = None - mode: str = "live" - internal_error: str | None = None - - def to_dict(self) -> dict[str, Any]: - d = asdict(self) - return d - - -def utc_now_rfc3339() -> str: - return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") - - -def summarize(checks: list[CheckResult]) -> dict[str, int]: - summary = {k: 0 for k in STATUS_ORDER} - for c in checks: - summary[c.status] = summary.get(c.status, 0) + 1 - return summary - - -def aggregate_health(checks: list[CheckResult]) -> Health: - """Aggregate check statuses into overall health. - - Rules (doctor contract): - - healthy: all required checks pass (optional may warn/skip) - - unhealthy: any required failure or critical defect - - unknown: a critical required check is unknown, or insufficient evidence - - degraded: no critical failure, but required warn/stale or accepted optional fail - """ - if not checks: - return "unknown" - - required = [c for c in checks if c.required] - optional = [c for c in checks if not c.required] - - # Critical severity fail on any check → unhealthy - for c in checks: - if c.status == "fail" and c.severity == "critical": - return "unhealthy" - - for c in required: - if c.status == "fail": - return "unhealthy" - - # Required unknown / skip-policy-fail → unknown (cannot claim healthy) - for c in required: - if c.status == "unknown": - return "unknown" - if c.status == "skip": - # skip_policy fail for required is treated as unknown, not pass - return "unknown" - - for c in required: - if c.status in ("warn", "stale"): - return "degraded" - - for c in optional: - if c.status == "fail": - return "degraded" - - for c in required: - if c.status != "pass": - return "unknown" - - return "healthy" - - -def exit_code(health: Health, *, internal_error: bool = False) -> int: - if internal_error: - return 2 - if health == "healthy": - return 0 - return 1 - - -def redact_secret_shaped(text: str) -> str: - """Redact long token-like substrings; never echo secret values.""" - import re - - # crude: long base64/hex-ish runs - text = re.sub(r"(?i)(secret|token|password|api[_-]?key)\s*[:=]\s*\S+", r"\1=[REDACTED]", text) - text = re.sub(r"\b[A-Za-z0-9_\-]{40,}\b", "[REDACTED_LONG]", text) - return text diff --git a/scripts/mainframe_paths.py b/scripts/mainframe_paths.py new file mode 100644 index 0000000..e44bc38 --- /dev/null +++ b/scripts/mainframe_paths.py @@ -0,0 +1,43 @@ +"""Path display helpers shared by bin/ tools. + +## Why this module exists + +`Path.relative_to` raises `ValueError` on any path outside the repo. Three tools +have now shipped that crash independently: + + 2026-08-11 bin/capture-validate crashed on any path outside the tree, and + exited 1 — which was also its "found + errors" code, so a crash was + indistinguishable from a result + 2026-08-23 bin/contract-lint same crash, linting a file by absolute + path from outside the repo + 2026-08-23 bin/doctor-remediate same crash, writing a receipt whose + snapshot directory was a temp dir + +`bin/papercut harvest`'s rule is that one papercut is noise and the same +papercut three times is a defect with an address. This is the address. A fourth +private `try/except ValueError` would have been the wrong fix. + +Import from a bin/ script the way the repo already does it:: + + sys.path.insert(0, str(ROOT)) + from scripts.mainframe_paths import rel_to_root +""" + +from __future__ import annotations + +from pathlib import Path + + +def rel_to_root(path: Path | str, root: Path) -> str: + """Repo-relative when possible, absolute otherwise. Never raises. + + Display only. Do not use the result to reopen the file — an absolute return + value means the path was outside `root`, and re-joining it to `root` would + silently point somewhere else. + """ + p = Path(path) + try: + return str(p.resolve().relative_to(root.resolve())) + except (ValueError, OSError): + return str(p) diff --git a/scripts/migration_lease.py b/scripts/migration_lease.py new file mode 100644 index 0000000..e01debc --- /dev/null +++ b/scripts/migration_lease.py @@ -0,0 +1,422 @@ +"""External migration epoch/lease used by lifecycle writers. + +The lock and state live under ``20_live/system-health`` (or an injected shadow +root), never under either lifecycle root. Supported writers acquire a shared +lease immediately before resolving a path and hold it until their write is +complete. The migration runner acquires an exclusive lease, so a writer that +resolved the old path earlier cannot recreate it after the rename. +""" + +from __future__ import annotations + +import fcntl +import hashlib +import json +import os +import uuid +from contextlib import contextmanager +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any + + +class LeaseError(RuntimeError): + """Base class for lease failures.""" + + +class LeaseBusy(LeaseError): + """A compatible writer or migration lease is currently held.""" + + +class LeaseUnavailable(LeaseError): + """The lease could not be established safely.""" + + +PAUSE_STATE_FILENAME = "writer-pause.json" + + +def utc_now() -> str: + return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + +def default_state_root(mainframe_root: Path | str) -> Path: + override = os.environ.get("MAINFRAME_MIGRATION_STATE_ROOT") + if override: + candidate = Path(override).expanduser() + if not candidate.is_absolute(): + candidate = Path(mainframe_root).expanduser().resolve() / candidate + return candidate.resolve() + return Path(mainframe_root).expanduser().resolve() / "20_live" / "system-health" / "mpe-migration" + + +def _read_json(path: Path) -> dict[str, Any]: + try: + payload = json.loads(path.read_text(encoding="utf-8")) + except FileNotFoundError: + return {} + except (OSError, json.JSONDecodeError) as exc: + raise LeaseUnavailable(f"lease state unreadable: {type(exc).__name__}: {exc}") from exc + if not isinstance(payload, dict): + raise LeaseUnavailable("lease state must be a JSON object") + return payload + + +def _write_state(path: Path, payload: dict[str, Any]) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + temporary = path.with_name(f".{path.name}.{os.getpid()}.{uuid.uuid4().hex}.tmp") + encoded = (json.dumps(payload, indent=2, sort_keys=True) + "\n").encode("utf-8") + try: + descriptor = os.open(temporary, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600) + with os.fdopen(descriptor, "wb") as handle: + handle.write(encoded) + handle.flush() + os.fsync(handle.fileno()) + os.replace(temporary, path) + _fsync_directory(path.parent) + finally: + try: + temporary.unlink() + except FileNotFoundError: + pass + + +def _fsync_directory(path: Path) -> None: + descriptor = os.open(path, os.O_RDONLY) + try: + os.fsync(descriptor) + finally: + os.close(descriptor) + + +def _remove_durable(path: Path) -> None: + try: + path.unlink() + except FileNotFoundError: + return + _fsync_directory(path.parent) + + +@contextmanager +def _audit_guard(state_root: Path): + """Serialize the audit projection without weakening the migration lock.""" + + path = state_root / "audit.lock" + handle = path.open("a+") + try: + fcntl.flock(handle.fileno(), fcntl.LOCK_EX) + yield + finally: + try: + fcntl.flock(handle.fileno(), fcntl.LOCK_UN) + finally: + handle.close() + + +def _pid_alive(pid: Any) -> bool: + if not isinstance(pid, int) or pid <= 0: + return False + try: + os.kill(pid, 0) + except ProcessLookupError: + return False + except PermissionError: + return True + except OSError: + return False + return True + + +@dataclass +class MigrationLease: + mainframe_root: Path + mode: str + holder: str + state_root: Path | None = None + nonblocking: bool = True + pause_token: str | None = None + pause_scope: str | None = None + _lock_handle: Any = None + _epoch: str | None = None + _state_path: Path | None = None + _holder_path: Path | None = None + + def __post_init__(self) -> None: + if self.mode not in {"shared", "exclusive"}: + raise ValueError("lease mode must be shared or exclusive") + self.mainframe_root = Path(self.mainframe_root).expanduser().resolve() + self.state_root = (self.state_root or default_state_root(self.mainframe_root)).expanduser().resolve() + + @property + def epoch(self) -> str: + if not self._epoch: + raise LeaseError("lease has not been acquired") + return self._epoch + + @property + def lock_path(self) -> Path: + assert self.state_root is not None + return self.state_root / "migration.lock" + + @property + def state_path(self) -> Path: + assert self.state_root is not None + return self.state_root / "lease-state.json" + + @property + def holders_dir(self) -> Path: + assert self.state_root is not None + return self.state_root / "holders" + + @property + def audit_lock_path(self) -> Path: + assert self.state_root is not None + return self.state_root / "audit.lock" + + @property + def writes_path(self) -> Path: + assert self.state_root is not None + return self.state_root / "write-events.jsonl" + + @property + def pause_state_path(self) -> Path: + assert self.state_root is not None + return self.state_root / PAUSE_STATE_FILENAME + + def _check_writer_pause(self) -> None: + """Fail closed while a migration has paused normal writers. + + The migration runner uses an exclusive lease and is therefore allowed + to create or advance the pause state. Shared lifecycle writers must + present the exact one-shot token while the state is paused; a missing, + malformed, or unknown state never becomes an accidental allow. + """ + + if self.mode != "shared": + return + pause = read_writer_pause_state(self.state_root) + if not pause: + return + status = pause.get("status") + scope = pause.get("slug") + if self.pause_scope and scope and scope != self.pause_scope: + return + if status == "resume-permitted": + return + if status != "paused": + raise LeaseUnavailable(f"writer pause state has unsupported status: {status!r}") + expected = pause.get("pause_token") + if not isinstance(expected, str) or not expected: + raise LeaseUnavailable("paused writer state has no valid pause token") + if self.pause_token != expected: + raise LeaseBusy("normal lifecycle writers are persistently paused by an active migration") + + def acquire(self) -> "MigrationLease": + assert self.state_root is not None + self.state_root.mkdir(parents=True, exist_ok=True) + self._state_path = self.state_path + self._lock_handle = self.lock_path.open("a+") + operation = fcntl.LOCK_SH if self.mode == "shared" else fcntl.LOCK_EX + if self.nonblocking: + operation |= fcntl.LOCK_NB + try: + fcntl.flock(self._lock_handle.fileno(), operation) + except BlockingIOError as exc: + self._lock_handle.close() + self._lock_handle = None + raise LeaseBusy(f"migration lease busy: {self.mode}") from exc + except OSError as exc: + self._lock_handle.close() + self._lock_handle = None + raise LeaseUnavailable(f"cannot acquire migration lease: {exc}") from exc + + try: + self._check_writer_pause() + self._epoch = str(uuid.uuid4()) + self.holders_dir.mkdir(parents=True, exist_ok=True) + with _audit_guard(self.state_root): + prior = _read_json(self.state_path) + stale_state = bool(prior and not _pid_alive(prior.get("owner_pid"))) + # Holder files are the concurrent truth. The summary JSON is + # an auditable projection, never the lock authority. + active: list[dict[str, Any]] = [] + for holder_path in sorted(self.holders_dir.glob("*.json")): + holder_payload = _read_json(holder_path) + if not _pid_alive(holder_payload.get("owner_pid")): + _remove_durable(holder_path) + continue + active.append(holder_payload) + self._holder_path = self.holders_dir / f"{self._epoch}.json" + holder = { + "schema_version": 2, + "mode": self.mode, + "holder": self.holder, + "owner_pid": os.getpid(), + "epoch": self._epoch, + "started_at": utc_now(), + } + _write_state(self._holder_path, holder) + active.append(holder) + summary = { + "schema_version": 2, + "mode": self.mode, + "holder": self.holder, + "owner_pid": os.getpid(), + "epoch": self._epoch, + "started_at": holder["started_at"], + "recovered_stale_state": stale_state, + "active_holders": sorted(active, key=lambda item: str(item.get("epoch"))), + } + _write_state(self.state_path, summary) + except LeaseError: + self._release_fd_only() + raise + except Exception as exc: # noqa: BLE001 + # Metadata failures must never leak an acquired migration fd. + self._cleanup_holder_best_effort() + self._release_fd_only() + raise LeaseUnavailable(f"cannot record migration lease safely: {type(exc).__name__}: {exc}") from exc + return self + + def _cleanup_holder_best_effort(self) -> None: + if self._holder_path is None: + return + try: + with _audit_guard(self.state_root): + _remove_durable(self._holder_path) + except Exception: + pass + self._holder_path = None + + def _release_fd_only(self) -> None: + if self._lock_handle is None: + return + try: + fcntl.flock(self._lock_handle.fileno(), fcntl.LOCK_UN) + finally: + self._lock_handle.close() + self._lock_handle = None + self._state_path = None + self._epoch = None + + def release(self) -> None: + if self._lock_handle is None: + return + audit_error: Exception | None = None + try: + try: + with _audit_guard(self.state_root): + _remove_durable(self._holder_path) if self._holder_path else None + state = _read_json(self.state_path) + active = [ + item for item in state.get("active_holders", []) + if item.get("epoch") != self._epoch + ] + state.update({ + "mode": "released" if not active else "shared", + "released_at": utc_now(), + "active_holders": active, + }) + _write_state(self.state_path, state) + except Exception as exc: # noqa: BLE001 + audit_error = exc + fcntl.flock(self._lock_handle.fileno(), fcntl.LOCK_UN) + finally: + self._lock_handle.close() + self._lock_handle = None + self._state_path = None + self._holder_path = None + self._epoch = None + if audit_error is not None: + raise LeaseUnavailable(f"cannot finalize migration lease audit: {type(audit_error).__name__}: {audit_error}") from audit_error + + def record_write(self, path: Path | str, *, action: str) -> dict[str, Any]: + if self._lock_handle is None: + raise LeaseError("cannot record a write without an acquired lease") + target = Path(path).expanduser() + try: + relative = target.resolve().relative_to(self.mainframe_root) + displayed = relative.as_posix() + except ValueError: + displayed = str(target) + event = { + "schema_version": 1, + "event_id": str(uuid.uuid4()), + "epoch": self.epoch, + "holder": self.holder, + "pid": os.getpid(), + "at": utc_now(), + "action": action, + "path": displayed, + "lease_mode": self.mode, + } + self.writes_path.parent.mkdir(parents=True, exist_ok=True) + with self.writes_path.open("a", encoding="utf-8") as handle: + handle.write(json.dumps(event, sort_keys=True) + "\n") + handle.flush() + os.fsync(handle.fileno()) + return event + + def __enter__(self) -> "MigrationLease": + return self.acquire() + + def __exit__(self, exc_type: Any, exc: Any, tb: Any) -> None: + self.release() + + +def shared_writer_lease( + mainframe_root: Path | str, + holder: str, + *, + state_root: Path | str | None = None, + pause_token: str | None = None, + pause_scope: str | None = None, +) -> MigrationLease: + return MigrationLease( + Path(mainframe_root), + "shared", + holder, + Path(state_root) if state_root is not None else None, + pause_token=pause_token, + pause_scope=pause_scope, + ) + + +def exclusive_migration_lease( + mainframe_root: Path | str, + holder: str = "mpe-migration-runner", + *, + state_root: Path | str | None = None, +) -> MigrationLease: + return MigrationLease( + Path(mainframe_root), + "exclusive", + holder, + Path(state_root) if state_root is not None else None, + ) + + +def state_hash(state_root: Path | str) -> str | None: + path = Path(state_root).expanduser().resolve() / "lease-state.json" + if not path.is_file(): + return None + return hashlib.sha256(path.read_bytes()).hexdigest() + + +def read_writer_pause_state(state_root: Path | str) -> dict[str, Any]: + """Read the durable writer pause projection. + + A missing file means no migration pause is active. Any other read or JSON + failure raises ``LeaseUnavailable`` so consumers cannot silently resume + during an unreadable intermediate state. + """ + + path = Path(state_root).expanduser().resolve() / PAUSE_STATE_FILENAME + return _read_json(path) + + +def pause_state_hash(state_root: Path | str) -> str | None: + path = Path(state_root).expanduser().resolve() / PAUSE_STATE_FILENAME + if not path.is_file(): + return None + return hashlib.sha256(path.read_bytes()).hexdigest() diff --git a/scripts/mindgraph_link_audit.py b/scripts/mindgraph_link_audit.py deleted file mode 100644 index f38598f..0000000 --- a/scripts/mindgraph_link_audit.py +++ /dev/null @@ -1,439 +0,0 @@ -#!/usr/bin/env python3 -"""Advisory, source-authoritative MainFrame wikilink audit. - -Profile: deterministic-operation / root / generated-state. -Authority: Markdown source files in the requested scope and the canonical -MindGraph parser/resolver. This command never opens or mutates SQLite and -never rewrites source notes. It is deliberately separate from the -mindgraph-eval frozen-snapshot inventory. - -The audit reports link-level classifications (resolved, dangling/unresolved, -ambiguous, external/cross-lifecycle, intentional body/frontmatter mirror, -same-channel duplicate) and document-level findings (raw evidence leaf, -curated no-outbound, reviewed curated disposition, metadata domain/type gaps). -A unique candidate is recorded as an exact-safe repair candidate for review; it -is never selected or written automatically. -""" - -from __future__ import annotations - -import argparse -import hashlib -import json -import re -import sys -from collections import Counter -from datetime import UTC, datetime -from pathlib import Path -from typing import Any - - -ROOT = Path(__file__).resolve().parents[1] -MINDGRAPH_SRC = ROOT / "mindgraph" / "src" -if str(MINDGRAPH_SRC) not in sys.path: - sys.path.insert(0, str(MINDGRAPH_SRC)) - -from mindgraph.parser import ( # noqa: E402 - LINK_PATTERN, - LinkResolver, - extract_metadata_link_targets, - parse_document, -) - - -SUPPORTED_RELATIONSHIPS = frozenset( - {"evidence", "extends", "contrasts", "implements", "navigation"} -) -VALID_GRAPH_DISPOSITIONS = frozenset({"reviewed-no-link", "standalone"}) -CANONICAL_PREFIX = re.compile(r"^\d{4}-\d{2}-\d{2}__") - - -def normalize_target(target: str) -> str: - target = target.strip() - return target if target.endswith(".md") else f"{target}.md" - - -def lifecycle_classification(target: str) -> str | None: - lowered = target.strip().casefold() - if lowered.startswith(("http://", "https://", "doi:")): - return "external/cross-lifecycle" - if lowered.startswith(("00_inbox/", "01_ingest/", "20_live/", "30_projects/", "90_archive/")): - return "external/cross-lifecycle" - return None - - -def candidate_paths(target: str, paths: set[str]) -> tuple[str, list[str]]: - """Return conservative candidates using the resolver's supported aliases.""" - normalized = normalize_target(target) - name = Path(normalized).name.casefold() - stem = Path(normalized).stem.casefold() - same_name = sorted(path for path in paths if Path(path).name.casefold() == name) - if same_name: - return "same-filename", same_name - same_stem = sorted(path for path in paths if Path(path).stem.casefold() == stem) - if same_stem: - return "same-stem", same_stem - if "__" not in stem: - same_slug = sorted( - path - for path in paths - if CANONICAL_PREFIX.match(Path(path).stem) - and Path(path).stem.rsplit("__", 1)[-1].casefold() == stem - ) - if same_slug: - return "same-canonical-slug", same_slug - return "none", [] - - -def _context(text: str, start: int, end: int) -> str: - left = max(0, start - 160) - right = min(len(text), end + 160) - return " ".join(text[left:right].split()) - - -def _display_path(scope_root: Path, root: Path, relative: str) -> str: - return (scope_root.relative_to(root).as_posix().rstrip("/") + "/" + relative).lstrip("/") - - -def _source_occurrences(parsed: Any, text: str) -> list[dict[str, Any]]: - occurrences: list[dict[str, Any]] = [] - for target in extract_metadata_link_targets(parsed.metadata): - occurrences.append( - { - "raw_link_target": target, - "location": "frontmatter", - "relationship": None, - "context": f"links: {target}", - } - ) - for match in LINK_PATTERN.finditer(parsed.truth_text): - relationship = match.group(2).strip() if match.group(2) else None - occurrences.append( - { - "raw_link_target": match.group(1).strip(), - "location": "body", - "relationship": relationship, - "context": _context(parsed.truth_text, match.start(), match.end()), - } - ) - return occurrences - - -def audit(root: Path, scope: Path, run_id: str) -> dict[str, Any]: - files = sorted(path for path in scope.rglob("*.md") if path.is_file()) - parsed_docs: list[tuple[Path, Any, str]] = [] - parse_errors: list[dict[str, str]] = [] - digest = hashlib.sha256() - for source in files: - relative = source.relative_to(scope).as_posix() - digest.update(relative.encode("utf-8")) - try: - body = source.read_bytes() - digest.update(body) - display = _display_path(scope, root, relative) - parsed = parse_document(relative, body).model_copy( - update={"path": display, "display_path": display, "source_path": relative} - ) - parsed_docs.append((source, parsed, display)) - except Exception as exc: # keep the audit advisory and complete the queue - parse_errors.append({"source_path": str(source.relative_to(root)), "error": str(exc)}) - - resolver = LinkResolver.from_documents(parsed for _, parsed, _ in parsed_docs) - paths = {display for _, _, display in parsed_docs} - link_rows: list[dict[str, Any]] = [] - document_rows: list[dict[str, Any]] = [] - metadata_gap_rows: list[dict[str, Any]] = [] - - for source, parsed, display in parsed_docs: - text = source.read_text(encoding="utf-8") - occurrences = _source_occurrences(parsed, text) - resolved_targets: set[str] = set() - occurrence_keys: Counter[str] = Counter() - source_link_rows: list[dict[str, Any]] = [] - for occurrence in occurrences: - target = occurrence["raw_link_target"] - lifecycle = lifecycle_classification(target) - resolved_path = resolver.resolve(target, display) if lifecycle is None else None - kind, candidates = candidate_paths(target, paths) if resolved_path is None else ("resolved", [resolved_path]) - if resolved_path: - classification = "resolved" - resolved_targets.add(resolved_path) - occurrence_key = resolved_path.casefold() - elif lifecycle: - classification = lifecycle - occurrence_key = target.casefold() - elif len(candidates) > 1: - classification = "ambiguous" - occurrence_key = target.casefold() - else: - classification = "dangling/unresolved" - occurrence_key = target.casefold() - occurrence_keys[occurrence_key] += 1 - row = { - "source_path": display, - "source_domain": parsed.metadata.get("domain") or "unset", - "source_type": parsed.metadata.get("type") or "unset", - "source_title": parsed.title, - "raw_link_target": target, - "location": occurrence["location"], - "relationship": occurrence["relationship"], - "relationship_supported": occurrence["relationship"] in SUPPORTED_RELATIONSHIPS - if occurrence["relationship"] - else None, - "target_resolution_status": classification, - "resolved_target_path": resolved_path, - "candidate_kind": kind, - "candidate_paths": candidates, - "repair_candidate": ( - "exact-safe-repair-candidate" if len(candidates) == 1 and not lifecycle and not resolved_path else None - ), - "context": occurrence["context"], - } - source_link_rows.append(row) - rows_by_key: dict[str, list[dict[str, Any]]] = {} - for row in source_link_rows: - key = row["resolved_target_path"] or row["raw_link_target"].casefold() - rows_by_key.setdefault(key, []).append(row) - for rows in rows_by_key.values(): - body_rows = [row for row in rows if row["location"] == "body"] - frontmatter_rows = [row for row in rows if row["location"] == "frontmatter"] - mirror_pairs = 0 - if rows[0]["target_resolution_status"] == "resolved": - mirror_pairs = min(len(body_rows), len(frontmatter_rows)) - for row in body_rows[:mirror_pairs] + frontmatter_rows[:mirror_pairs]: - row["mirror"] = True - row["duplicate"] = False - row["classifications"] = [row["target_resolution_status"], "mirror"] - for channel_rows, channel_total in ( - (body_rows[mirror_pairs:], len(body_rows)), - (frontmatter_rows[mirror_pairs:], len(frontmatter_rows)), - ): - is_duplicate = channel_total > 1 - for row in channel_rows: - row["mirror"] = False - row["duplicate"] = is_duplicate - row["classifications"] = [row["target_resolution_status"]] - if is_duplicate: - row["classifications"].append("duplicate") - if mirror_pairs == 0: - for row in rows: - row.setdefault("mirror", False) - row.setdefault("duplicate", len(rows) > 1) - row.setdefault("classifications", [row["target_resolution_status"]]) - link_rows.extend(source_link_rows) - - metadata = parsed.metadata - missing = [field for field in ("domain", "type") if not metadata.get(field)] - if missing: - gap = { - "source_path": display, - "classification": "metadata domain/type gaps", - "missing_fields": missing, - "context": "frontmatter metadata", - } - metadata_gap_rows.append(gap) - - item_type = str(metadata.get("type") or "unset").casefold() - if item_type == "raw" and not resolved_targets: - document_rows.append( - { - "source_path": display, - "source_domain": metadata.get("domain") or "unset", - "source_type": "raw", - "title": parsed.title, - "classification": "raw evidence leaf", - "resolved_outbound_count": 0, - "raw_link_count": len(occurrences), - "context": "raw document with zero resolved authored outbound links", - } - ) - if item_type == "note" and not resolved_targets: - disposition = str(metadata.get("graph_disposition") or "").strip().casefold() - disposition_valid = disposition in VALID_GRAPH_DISPOSITIONS - document_rows.append( - { - "source_path": display, - "source_domain": metadata.get("domain") or "unset", - "source_type": "note", - "title": parsed.title, - "classification": ( - "curated no-outbound reviewed" - if disposition_valid - else "curated no-outbound" - ), - "actionable": not disposition_valid, - "resolved_outbound_count": 0, - "raw_link_count": len(occurrences), - "graph_disposition": disposition or "missing-review-disposition", - "disposition_valid": disposition_valid, - "context": "curated note with zero resolved authored outbound links", - } - ) - - link_counts = Counter() - for row in link_rows: - link_counts.update(row["classifications"]) - duplicate_count = link_counts.get("duplicate", 0) - mirror_count = link_counts.get("mirror", 0) - disposition_counts = Counter( - row["graph_disposition"] - for row in document_rows - if row["classification"] == "curated no-outbound reviewed" - ) - doc_counts = Counter(row["classification"] for row in document_rows) - if metadata_gap_rows: - doc_counts["metadata domain/type gaps"] = len(metadata_gap_rows) - summary = { - "run_id": run_id, - "root": str(root), - "scope": str(scope), - "source_digest": digest.hexdigest(), - "source_file_count": len(files), - "parsed_file_count": len(parsed_docs), - "parse_error_count": len(parse_errors), - "link_count": len(link_rows), - "link_classification_counts": dict(sorted(link_counts.items())), - "duplicate_link_count": duplicate_count, - "mirror_occurrence_count": mirror_count, - "mirror_pair_count": mirror_count // 2, - "document_finding_counts": dict(sorted(doc_counts.items())), - "metadata_gap_count": len(metadata_gap_rows), - "curated_reviewed_disposition_counts": dict(sorted(disposition_counts.items())), - "supported_relationship_vocabulary": sorted(SUPPORTED_RELATIONSHIPS), - } - action_queue = { - "auto_detectable": [ - "review dangling/unresolved targets and exact-safe candidates", - "review ambiguous targets without fuzzy selection", - "review same-channel duplicate authored occurrences or duplicate edge semantics", - "review metadata domain/type gaps", - "review curated no-outbound notes with missing or invalid graph_disposition", - ], - "informational": [ - "body/frontmatter mirror pairs are intentional ingest mirrors and are not duplicate findings", - "raw evidence leaves (informational): raw documents with zero resolved authored outbound links", - "curated no-outbound notes with graph_disposition reviewed-no-link or standalone are reviewed/informational", - "external/cross-lifecycle references remain intentional unless a bridge policy is approved", - ], - "human_decision_required": [ - "source renames, identity conflicts, and any source rewrite or link repair", - ], - } - return { - "status": "advisory", - "mutating": False, - "snapshot_mode": False, - "summary": summary, - "action_queue": action_queue, - "links": link_rows, - "documents": document_rows, - "metadata_gaps": metadata_gap_rows, - "parse_errors": parse_errors, - } - - -def markdown_report(data: dict[str, Any]) -> str: - summary = data["summary"] - lines = [ - f"# MindGraph advisory link audit — {summary['run_id']}", - "", - "Source-authoritative, read-only preflight. No SQLite or source-note writes are performed; findings do not fail ingest.", - "", - f"- Scope: `{summary['scope']}`", - f"- Source digest: `{summary['source_digest']}`", - f"- Files / authored link occurrences: **{summary['source_file_count']} / {summary['link_count']}**", - f"- Status: **{data['status']}**; mutating: **{data['mutating']}**", - "", - "## Link classifications", - "", - "| Classification | Count |", - "| --- | ---: |", - ] - for key, value in summary["link_classification_counts"].items(): - lines.append(f"| `{key}` | {value} |") - lines.extend( - [ - f"| `mirror` (body/frontmatter ingest pairs) | {summary['mirror_occurrence_count']} occurrences / {summary['mirror_pair_count']} pairs |", - f"| `duplicate` (same-channel authored occurrences) | {summary['duplicate_link_count']} |", - "", - "## Document findings", - "", - "| Finding | Count |", - "| --- | ---: |", - ] - ) - for key, value in summary["document_finding_counts"].items(): - lines.append(f"| `{key}` | {value} |") - lines.extend(["", "## Action queue", ""]) - for bucket, entries in data["action_queue"].items(): - lines.append(f"### {bucket.replace('_', ' ').title()}") - lines.append("") - lines.extend(f"- {entry}" for entry in entries) - lines.append("") - if data["documents"]: - lines.extend(["## Document queue", "", "| Source | Type | Finding | Actionable | Disposition | Context |", "| --- | --- | --- | --- | --- | --- |"]) - for row in data["documents"]: - disposition = row.get("graph_disposition", "-") - lines.append( - f"| `{row['source_path']}` | `{row['source_type']}` | `{row['classification']}` | `{row.get('actionable', False)}` | `{disposition}` | {row['context']} |" - ) - if data["metadata_gaps"]: - lines.extend(["", "## Metadata gaps", ""]) - for row in data["metadata_gaps"]: - lines.append(f"- `{row['source_path']}` — missing `{', '.join(row['missing_fields'])}`") - unresolved = [row for row in data["links"] if row["target_resolution_status"] != "resolved"] - if unresolved: - lines.extend(["", "## Unresolved and candidate links", "", "| Source | Raw target | Status | Candidates | Context |", "| --- | --- | --- | --- | --- |"]) - for row in unresolved: - candidates = ", ".join(row["candidate_paths"]) or "-" - context = row["context"].replace("|", "\\|") - lines.append(f"| `{row['source_path']}` | `{row['raw_link_target']}` | `{row['target_resolution_status']}` | `{candidates}` | {context} |") - lines.extend(["", "## Non-mutating contract", "", "This audit only reads source Markdown, uses the canonical parser/resolver semantics, and emits advisory JSON/Markdown. It does not rewrite notes, delete links, mutate SQLite, fuzzy-select targets, or act as a blocking gate.", ""]) - return "\n".join(lines) - - -def build_parser() -> argparse.ArgumentParser: - parser = argparse.ArgumentParser( - prog="bin/mindgraph-audit-links", - description="Run the advisory, read-only MainFrame wikilink audit.", - epilog=( - "Recommended sequence: pre-refresh audit -> source review/edit -> " - "durable refresh -> post-refresh audit/probe. Findings are advisory " - "and never a blocking gate." - ), - ) - parser.add_argument("--root", default=str(ROOT), help="MainFrame root") - parser.add_argument("--scope", default="10_knowledge", help="scope directory, relative to --root") - parser.add_argument("--run-id", default=datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")) - parser.add_argument("--output-dir", help="optional directory for JSON and Markdown reports") - parser.add_argument("--json", action="store_true", help="print the complete JSON report") - parser.add_argument("--dry-run", action="store_true", help="explicitly document read-only mode") - return parser - - -def main(argv: list[str] | None = None) -> int: - args = build_parser().parse_args(argv) - root = Path(args.root).resolve() - scope_arg = Path(args.scope) - scope = (root / scope_arg if not scope_arg.is_absolute() else scope_arg).resolve() - if not scope.is_dir(): - print(f"audit scope does not exist: {scope}", file=sys.stderr) - return 2 - data = audit(root, scope, args.run_id) - human = markdown_report(data) - if args.output_dir: - output_dir = Path(args.output_dir).resolve() - output_dir.mkdir(parents=True, exist_ok=True) - (output_dir / f"{args.run_id}.json").write_text(json.dumps(data, indent=2, sort_keys=True) + "\n", encoding="utf-8") - (output_dir / f"{args.run_id}.md").write_text(human, encoding="utf-8") - print(f"json: {output_dir / f'{args.run_id}.json'}", file=sys.stderr) - print(f"report: {output_dir / f'{args.run_id}.md'}", file=sys.stderr) - if args.json: - print(json.dumps(data, indent=2, sort_keys=True)) - else: - print(human) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/mindgraph_projects_apply.py b/scripts/mindgraph_projects_apply.py deleted file mode 100644 index 46652ca..0000000 --- a/scripts/mindgraph_projects_apply.py +++ /dev/null @@ -1,416 +0,0 @@ -"""Staged apply for the projects MindGraph index (ADR-045). - -plan → coverage of manifest vs on-disk projects; dry-run scopes -stage → ingest into a staging DB; verify namespaces; write receipt -promote → backup installed DB; replace from a green stage receipt -status → compare installed DB namespaces to manifest - -Never mutates the installed DB without --promote. -Never ingests 20_live telemetry. -""" - -from __future__ import annotations - -import argparse -import hashlib -import json -import os -import shutil -import sqlite3 -import subprocess -import sys -from datetime import datetime, timezone -from pathlib import Path -from typing import Any -from urllib.parse import quote - - -ROOT = Path(__file__).resolve().parents[1] -DEFAULT_MANIFEST = ROOT / "30_projects" / "mindgraph-projects.json" -DEEP_MANIFEST = ROOT / "30_projects" / "mindgraph-projects-deep.json" -DEFAULT_INSTALLED = Path.home() / ".mindgraph" / "mainframe-projects.sqlite" -STAGING_DIR = Path.home() / ".mindgraph" / "staging" -RECEIPT_DIR = ROOT / "20_live" / "system-health" / "mindgraph-projects-apply" -REFRESH = ROOT / "bin" / "mindgraph-refresh-projects" - - -def utc_now() -> str: - return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") - - -def sha256_file(path: Path) -> str | None: - if not path.exists() or not path.is_file(): - return None - h = hashlib.sha256() - with path.open("rb") as f: - for chunk in iter(lambda: f.read(1 << 20), b""): - h.update(chunk) - return h.hexdigest() - - -def load_manifest_projects(manifest_path: Path) -> list[str]: - payload = json.loads(manifest_path.read_text(encoding="utf-8")) - projects = payload.get("projects") - if not isinstance(projects, list) or not all(isinstance(p, str) for p in projects): - raise SystemExit(f"manifest.projects must be a string list: {manifest_path}") - return list(projects) - - -def real_project_dirs(projects_dir: Path) -> list[str]: - return sorted( - p.name - for p in projects_dir.iterdir() - if p.is_dir() and not p.name.startswith(".") and (p / "README.md").exists() - ) - - -def coverage_report(manifest_projects: list[str], real_dirs: list[str]) -> dict[str, Any]: - set_m, set_r = set(manifest_projects), set(real_dirs) - return { - "manifest_count": len(manifest_projects), - "real_count": len(real_dirs), - "missing_dirs": sorted(set_m - set_r), - "omitted_dirs": sorted(set_r - set_m), - "intersection": sorted(set_m & set_r), - "complete": set_m <= set_r and len(set_m) == len(set_r), - } - - -def namespaces_from_db(db_path: Path) -> dict[str, int]: - if not db_path.exists(): - return {} - resolved = db_path.expanduser().resolve() - uri = f"file:{quote(str(resolved), safe='/')}?mode=ro&immutable=1" - con = sqlite3.connect(uri, uri=True) - try: - con.execute("PRAGMA query_only = ON") - rows = con.execute( - "SELECT namespace, COUNT(*) FROM documents GROUP BY namespace ORDER BY 1" - ).fetchall() - finally: - con.close() - return {str(n): int(c) for n, c in rows if n} - - -def verify_stage( - db_path: Path, - expected_namespaces: list[str], -) -> dict[str, Any]: - observed = namespaces_from_db(db_path) - expected = set(expected_namespaces) - present = set(observed) - missing = sorted(expected - present) - extra = sorted(present - expected) - return { - "expected_count": len(expected), - "observed_count": len(present), - "doc_counts": observed, - "missing_namespaces": missing, - "extra_namespaces": extra, - "ok": not missing and db_path.exists() and db_path.stat().st_size > 10_000, - } - - -def run_refresh_apply( - db_path: Path, - *, - full: bool = False, - deep: bool = False, - manifest: Path | None = None, -) -> subprocess.CompletedProcess[str]: - env = os.environ.copy() - env["MINDGRAPH_DB_PATH"] = str(db_path) - if manifest is not None: - env["MAINFRAME_MINDGRAPH_PROJECTS_MANIFEST"] = str(manifest) - cmd = [str(REFRESH), "--apply"] - if full: - cmd.append("--full") - if deep and manifest is None: - cmd.append("--deep") - return subprocess.run( - cmd, - cwd=str(ROOT), - capture_output=True, - text=True, - env=env, - check=False, - ) - - -def resolve_manifest(args: argparse.Namespace) -> Path: - if getattr(args, "manifest", None) and args.manifest != str(DEFAULT_MANIFEST): - return Path(args.manifest).expanduser() - if getattr(args, "deep", False): - return DEEP_MANIFEST - return Path(args.manifest).expanduser() - - -def cmd_plan(args: argparse.Namespace) -> int: - manifest = resolve_manifest(args) - projects = load_manifest_projects(manifest) - real = real_project_dirs(ROOT / "30_projects") - cov = coverage_report(projects, real) - profile = json.loads(manifest.read_text(encoding="utf-8")).get("profile", "unknown") - print("mindgraph-projects-apply plan") - print(f" profile: {profile}") - print(f" manifest: {manifest}") - print(f" installed: {args.installed}") - print(f" coverage_complete: {cov['complete']}") - print(f" manifest_count: {cov['manifest_count']} real_count: {cov['real_count']}") - includes = json.loads(manifest.read_text(encoding="utf-8")).get("include", []) - print(f" include: {includes}") - if cov["missing_dirs"]: - print(f" missing_dirs: {cov['missing_dirs']}") - if cov["omitted_dirs"]: - print(f" omitted_dirs: {cov['omitted_dirs']}") - # dry-run wrapper for scopes - env = os.environ.copy() - env["MINDGRAPH_DB_PATH"] = str(Path(args.installed).expanduser()) - env["MAINFRAME_MINDGRAPH_PROJECTS_MANIFEST"] = str(manifest) - dry_cmd = [str(REFRESH), "--dry-run"] - if args.full: - dry_cmd.append("--full") - # Manifest path is forced via MAINFRAME_MINDGRAPH_PROJECTS_MANIFEST - proc = subprocess.run( - dry_cmd, - cwd=str(ROOT), - capture_output=True, - text=True, - env=env, - check=False, - ) - print("--- dry-run ---") - sys.stdout.write(proc.stdout or "") - if proc.stderr: - sys.stderr.write(proc.stderr) - if not cov["complete"]: - print("plan: coverage incomplete — fix manifest before stage", file=sys.stderr) - return 1 - if proc.returncode != 0: - return proc.returncode - print("plan: OK — ready to --stage") - return 0 - - -def cmd_stage(args: argparse.Namespace) -> int: - manifest = resolve_manifest(args) - projects = load_manifest_projects(manifest) - real = real_project_dirs(ROOT / "30_projects") - cov = coverage_report(projects, real) - if not cov["complete"] and not args.allow_incomplete_coverage: - print("error: manifest coverage incomplete; refuse stage", file=sys.stderr) - print(json.dumps(cov, indent=2), file=sys.stderr) - return 1 - - STAGING_DIR.mkdir(parents=True, exist_ok=True) - RECEIPT_DIR.mkdir(parents=True, exist_ok=True) - stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ") - profile = json.loads(manifest.read_text(encoding="utf-8")).get("profile", "default") - stage_db = STAGING_DIR / f"mainframe-projects-{profile}-{stamp}.sqlite" - if stage_db.exists(): - stage_db.unlink() - - print(f"stage: profile={profile} ingest → {stage_db}") - proc = run_refresh_apply( - stage_db, full=args.full, deep=args.deep, manifest=manifest - ) - sys.stdout.write(proc.stdout or "") - if proc.stderr: - sys.stderr.write(proc.stderr) - if proc.returncode != 0: - print(f"error: stage ingest exit {proc.returncode}", file=sys.stderr) - return proc.returncode - - # Expected namespaces: scopes that actually exist on disk from manifest - expected = [p for p in projects if (ROOT / "30_projects" / p).is_dir()] - verify = verify_stage(stage_db, expected) - receipt = { - "schema_version": 1, - "kind": "mindgraph-projects-stage", - "as_of": utc_now(), - "stamp": stamp, - "manifest": str(manifest), - "manifest_profile": profile, - "manifest_sha256": sha256_file(manifest), - "include": json.loads(manifest.read_text(encoding="utf-8")).get("include"), - "stage_db": str(stage_db), - "stage_db_sha256": sha256_file(stage_db), - "stage_db_size": stage_db.stat().st_size if stage_db.exists() else 0, - "coverage": cov, - "verify": verify, - "green": bool(verify.get("ok")), - "policy_ref": ".context/live-retention.md", - "adr": "ADR-045", - } - receipt_path = RECEIPT_DIR / f"{stamp}-stage.json" - receipt_path.write_text(json.dumps(receipt, indent=2) + "\n", encoding="utf-8") - print(f"stage receipt: {receipt_path.relative_to(ROOT)}") - print(f" namespaces: {verify['observed_count']}/{verify['expected_count']}") - if verify["missing_namespaces"]: - print(f" missing: {verify['missing_namespaces']}") - if verify["extra_namespaces"]: - print(f" extra: {verify['extra_namespaces']}") - if not receipt["green"]: - print("stage: NOT GREEN — do not promote", file=sys.stderr) - return 1 - print("stage: GREEN — promote with:") - print(f" bin/mindgraph-projects-apply --promote --receipt {receipt_path}") - return 0 - - -def cmd_promote(args: argparse.Namespace) -> int: - receipt_path = Path(args.receipt).expanduser() - if not receipt_path.is_absolute(): - # allow relative to ROOT - cand = ROOT / receipt_path - if cand.exists(): - receipt_path = cand - if not receipt_path.exists(): - print(f"error: receipt not found: {receipt_path}", file=sys.stderr) - return 1 - receipt = json.loads(receipt_path.read_text(encoding="utf-8")) - if not receipt.get("green"): - print("error: refuse promote of non-green receipt", file=sys.stderr) - return 1 - stage_db = Path(receipt["stage_db"]) - if not stage_db.exists(): - print(f"error: stage DB missing: {stage_db}", file=sys.stderr) - return 1 - current_hash = sha256_file(stage_db) - if receipt.get("stage_db_sha256") and current_hash != receipt["stage_db_sha256"]: - print("error: stage DB hash mismatch vs receipt (stage changed)", file=sys.stderr) - return 1 - - installed = Path(args.installed).expanduser() - installed.parent.mkdir(parents=True, exist_ok=True) - backup = None - if installed.exists(): - stamp = receipt.get("stamp") or datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ") - backup = installed.with_name(f"mainframe-projects.pre-promote-{stamp}.sqlite") - shutil.copy2(installed, backup) - print(f"promote: backed up → {backup}") - - # Atomic-ish replace: copy to temp beside target then replace - tmp_target = installed.with_suffix(installed.suffix + ".promoting") - shutil.copy2(stage_db, tmp_target) - os.replace(tmp_target, installed) - - manifest_path = Path(receipt["manifest"]) - if manifest_path.exists(): - expected = [ - p - for p in load_manifest_projects(manifest_path) - if (ROOT / "30_projects" / p).is_dir() - ] - else: - expected = list((receipt.get("verify") or {}).get("doc_counts") or {}) - post = verify_stage(installed, expected) - - promote_receipt = { - "schema_version": 1, - "kind": "mindgraph-projects-promote", - "as_of": utc_now(), - "from_stage_receipt": str(receipt_path), - "installed": str(installed), - "installed_sha256": sha256_file(installed), - "backup": str(backup) if backup else None, - "verify": post, - "green": bool(post.get("ok")), - "adr": "ADR-045", - } - RECEIPT_DIR.mkdir(parents=True, exist_ok=True) - out = RECEIPT_DIR / f"{receipt.get('stamp', 'promote')}-promote.json" - out.write_text(json.dumps(promote_receipt, indent=2) + "\n", encoding="utf-8") - print(f"promote receipt: {out.relative_to(ROOT)}") - if not promote_receipt["green"]: - print("promote: installed but verify failed — inspect backup", file=sys.stderr) - return 1 - print("promote: OK") - return 0 - - -def cmd_status(args: argparse.Namespace) -> int: - manifest = resolve_manifest(args) - installed = Path(args.installed).expanduser() - projects = load_manifest_projects(manifest) - real = real_project_dirs(ROOT / "30_projects") - cov = coverage_report(projects, real) - expected = [p for p in projects if (ROOT / "30_projects" / p).is_dir()] - verify = verify_stage(installed, expected) if installed.exists() else { - "ok": False, - "missing_namespaces": expected, - "observed_count": 0, - "expected_count": len(expected), - "doc_counts": {}, - "extra_namespaces": [], - } - profile = json.loads(manifest.read_text(encoding="utf-8")).get("profile", "unknown") - print("mindgraph-projects-apply status") - print(f" profile: {profile}") - print(f" manifest: {manifest}") - print(f" manifest_coverage_complete: {cov['complete']}") - print(f" installed_exists: {installed.exists()} size={installed.stat().st_size if installed.exists() else 0}") - print(f" installed_namespaces: {verify.get('observed_count')}/{verify.get('expected_count')}") - if verify.get("missing_namespaces"): - print(f" missing_in_db: {verify['missing_namespaces']}") - if verify.get("extra_namespaces"): - print(f" extra_in_db: {verify['extra_namespaces']}") - print(f" installed_green: {verify.get('ok')}") - return 0 if cov["complete"] and verify.get("ok") else 1 - - -def main(argv: list[str] | None = None) -> int: - parser = argparse.ArgumentParser( - description="Staged projects MindGraph apply (plan → stage → promote)." - ) - parser.add_argument( - "--manifest", - default=str(DEFAULT_MANIFEST), - help="projects manifest JSON (default: lean mindgraph-projects.json)", - ) - parser.add_argument( - "--deep", - action="store_true", - help="use mindgraph-projects-deep.json (includes outputs/**)", - ) - parser.add_argument( - "--installed", - default=str(DEFAULT_INSTALLED), - help="installed projects DB path", - ) - parser.add_argument("--full", action="store_true", help="pass --full to refresh wrapper") - mode = parser.add_mutually_exclusive_group(required=True) - mode.add_argument("--plan", action="store_true", help="coverage + dry-run only") - mode.add_argument("--stage", action="store_true", help="ingest into staging DB + receipt") - mode.add_argument("--promote", action="store_true", help="promote green stage to installed") - mode.add_argument("--status", action="store_true", help="compare installed DB to manifest") - parser.add_argument( - "--receipt", - help="stage receipt path (required for --promote)", - ) - parser.add_argument( - "--allow-incomplete-coverage", - action="store_true", - help="allow stage when manifest ≠ disk (not recommended)", - ) - args = parser.parse_args(argv) - # If user passed both --deep and explicit default path, deep wins via resolve_manifest - if args.deep and args.manifest == str(DEFAULT_MANIFEST): - pass # resolve_manifest uses DEEP_MANIFEST - - if args.plan: - return cmd_plan(args) - if args.stage: - return cmd_stage(args) - if args.promote: - if not args.receipt: - print("error: --promote requires --receipt", file=sys.stderr) - return 2 - return cmd_promote(args) - if args.status: - return cmd_status(args) - return 2 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/scripts/work_inventory.py b/scripts/work_inventory.py new file mode 100644 index 0000000..2ccd2a0 --- /dev/null +++ b/scripts/work_inventory.py @@ -0,0 +1,659 @@ +"""Typed cross-lifecycle inventory and projects-index coverage checks. + +The inventory is a derived projection. It validates direct lifecycle +authority from :mod:`lifecycle_identity`; it never resolves a record for a +reader or writer. The same typed member map is used by the MindGraph stage, +promotion, and status paths so a namespace count cannot masquerade as green +coverage. +""" + +from __future__ import annotations + +import fnmatch +import hashlib +import json +import sqlite3 +from dataclasses import dataclass +from pathlib import Path +from typing import Any, Iterable +from urllib.parse import quote + +from lifecycle_identity import ( + IdentityScan, + LifecycleRecord, + display_path, + scan_lifecycle_records, +) + + +INDEX_ID = "mainframe-projects" +TRUST_PROFILE = "project_status" +DEFAULT_INCLUDE = ( + "README.md", + "AGENTS.md", + "log.md", + "decisions.md", + "methodology-approach.md", + "plans/*.md", + "plans/**/*.md", +) +DEFAULT_EXCLUDE = ( + ".git/*", + ".git/**/*", + ".venv/*", + ".venv/**/*", + "__pycache__/*", + "__pycache__/**/*", + ".pytest_cache/*", + ".pytest_cache/**/*", + "workbench/*", + "workbench/**/*", + "raw-materials/*", + "raw-materials/**/*", + "outputs/*", + "outputs/**/*", + "*sealed*", + "*blind-sheet-*", + "*items-blind*", +) + + +@dataclass(frozen=True) +class ManifestMember: + slug: str + record_type: str + path: str + lifecycle_state: str | None = None + readme_sha256: str | None = None + + def as_dict(self) -> dict[str, Any]: + result: dict[str, Any] = { + "id": self.slug, + "record_type": self.record_type, + "path": self.path, + } + if self.lifecycle_state is not None: + result["lifecycle_state"] = self.lifecycle_state + if self.readme_sha256 is not None: + result["readme_sha256"] = self.readme_sha256 + return result + + +@dataclass(frozen=True) +class ManifestExclusion: + """A direct lifecycle record intentionally outside an index profile.""" + + slug: str + record_type: str + path: str + reason: str + + def as_dict(self) -> dict[str, Any]: + return { + "id": self.slug, + "record_type": self.record_type, + "path": self.path, + "reason": self.reason, + } + + +def sha256_file(path: Path) -> str | None: + if not path.is_file(): + return None + digest = hashlib.sha256() + with path.open("rb") as handle: + for chunk in iter(lambda: handle.read(1 << 20), b""): + digest.update(chunk) + return digest.hexdigest() + + +def load_manifest(path: Path) -> dict[str, Any]: + payload = json.loads(path.read_text(encoding="utf-8")) + if not isinstance(payload, dict): + raise ValueError(f"manifest must be an object: {path}") + return payload + + +def _member_from_payload(item: Any, index: int) -> ManifestMember: + if not isinstance(item, dict): + raise ValueError(f"manifest member #{index} must be an object") + slug = item.get("id") or item.get("slug") + record_type = item.get("record_type") + path = item.get("path") + if not isinstance(slug, str) or not slug: + raise ValueError(f"manifest member #{index} has no id") + if not isinstance(record_type, str) or record_type not in {"project", "operation"}: + raise ValueError(f"manifest member {slug!r} has invalid record_type") + if not isinstance(path, str) or not path: + raise ValueError(f"manifest member {slug!r} has no path") + return ManifestMember( + slug=slug, + record_type=record_type, + path=path.rstrip("/"), + lifecycle_state=(str(item["lifecycle_state"]) if item.get("lifecycle_state") is not None else None), + readme_sha256=(str(item["readme_sha256"]) if item.get("readme_sha256") is not None else None), + ) + + +def _exclusion_from_payload(item: Any, index: int) -> ManifestExclusion: + if not isinstance(item, dict): + raise ValueError(f"manifest exclusion #{index} must be an object") + slug = item.get("id") or item.get("slug") + record_type = item.get("record_type") + path = item.get("path") + reason = item.get("reason") + if not isinstance(slug, str) or not slug: + raise ValueError(f"manifest exclusion #{index} has no id") + if not isinstance(record_type, str) or record_type not in {"project", "operation"}: + raise ValueError(f"manifest exclusion {slug!r} has invalid record_type") + if not isinstance(path, str) or not path: + raise ValueError(f"manifest exclusion {slug!r} has no path") + path = path.rstrip("/") + if ( + not path + or path.startswith("/") + or "\\" in path + or any(part in {"", ".", ".."} for part in path.split("/")) + or not (path.startswith("30_projects/") or path.startswith("40_operations/")) + ): + raise ValueError(f"manifest exclusion {slug!r} path is outside lifecycle roots: {path}") + if not isinstance(reason, str) or not reason.strip(): + raise ValueError(f"manifest exclusion {slug!r} has no reason") + return ManifestExclusion( + slug=slug, + record_type=record_type, + path=path, + reason=reason.strip(), + ) + + +def manifest_exclusions(payload: dict[str, Any]) -> tuple[list[ManifestExclusion], list[str]]: + """Parse explicit profile exclusions without treating them as members.""" + + issues: list[str] = [] + raw_exclusions = payload.get("exclusions", []) + if raw_exclusions is None: + raw_exclusions = [] + if not isinstance(raw_exclusions, list): + return [], ["manifest.exclusions must be a list"] + exclusions: list[ManifestExclusion] = [] + seen_ids: set[str] = set() + seen_paths: set[str] = set() + for index, item in enumerate(raw_exclusions, start=1): + try: + exclusion = _exclusion_from_payload(item, index) + except ValueError as exc: + issues.append(str(exc)) + continue + if exclusion.slug in seen_ids: + issues.append(f"duplicate manifest exclusion identity: {exclusion.slug}") + if exclusion.path in seen_paths: + issues.append(f"duplicate manifest exclusion path: {exclusion.path}") + seen_ids.add(exclusion.slug) + seen_paths.add(exclusion.path) + exclusions.append(exclusion) + return exclusions, issues + + +def manifest_members(payload: dict[str, Any]) -> tuple[list[ManifestMember], list[str]]: + """Return typed members plus structural manifest issues.""" + + issues: list[str] = [] + raw_members = payload.get("members") + if not isinstance(raw_members, list): + issues.append("manifest.members is missing; exact typed coverage is unavailable") + # Legacy manifests can still be read for a diagnostic, but they are not + # eligible for a green migration stage. + raw_projects = payload.get("projects") + if isinstance(raw_projects, list): + raw_members = [ + { + "id": slug, + "record_type": "project", + "path": f"30_projects/{slug}", + } + for slug in raw_projects + if isinstance(slug, str) + ] + else: + raw_members = [] + members: list[ManifestMember] = [] + seen_ids: set[str] = set() + seen_paths: set[str] = set() + for index, item in enumerate(raw_members, start=1): + try: + member = _member_from_payload(item, index) + except ValueError as exc: + issues.append(str(exc)) + continue + if member.slug in seen_ids: + issues.append(f"duplicate manifest identity: {member.slug}") + if member.path in seen_paths: + issues.append(f"duplicate manifest path: {member.path}") + seen_ids.add(member.slug) + seen_paths.add(member.path) + if not (member.path.startswith("30_projects/") or member.path.startswith("40_operations/")): + issues.append(f"manifest member {member.slug} path is outside lifecycle roots: {member.path}") + members.append(member) + projects = payload.get("projects") + if isinstance(projects, list): + project_ids = [str(value) for value in projects] + # The legacy list is physically rooted in 30_projects. Typed + # operation members under 40_operations remain in `members` but are + # intentionally absent from this compatibility projection. + member_ids = [ + member.slug for member in members if member.path.startswith("30_projects/") + ] + if project_ids != member_ids: + issues.append( + "manifest.projects does not exactly preserve 30_projects member order/identity" + ) + return members, issues + + +def _record_map(scan: IdentityScan) -> dict[str, LifecycleRecord]: + result: dict[str, LifecycleRecord] = {} + for record in scan.records: + result.setdefault(record.slug, record) + return result + + +def validate_manifest( + mainframe_root: Path | str, + manifest_path: Path | str, + *, + projects_root: Path | str | None = None, + operations_root: Path | str | None = None, +) -> dict[str, Any]: + root = Path(mainframe_root).expanduser().resolve() + manifest = Path(manifest_path).expanduser().resolve() + issues: list[str] = [] + try: + payload = load_manifest(manifest) + members, manifest_issues = manifest_members(payload) + issues.extend(manifest_issues) + exclusions, exclusion_issues = manifest_exclusions(payload) + issues.extend(exclusion_issues) + except (OSError, ValueError, json.JSONDecodeError) as exc: + return { + "ok": False, + "manifest": str(manifest), + "manifest_sha256": sha256_file(manifest), + "issues": [f"manifest unreadable: {type(exc).__name__}: {exc}"], + "members": [], + "exclusions": [], + "records": [], + } + scan = scan_lifecycle_records( + root, + projects_root=projects_root, + operations_root=operations_root, + ) + issues.extend(scan.issues) + observed = _record_map(scan) + expected = {member.slug: member for member in members} + excluded = {exclusion.slug: exclusion for exclusion in exclusions} + member_paths = {member.path: member.slug for member in members} + for exclusion in exclusions: + if exclusion.slug in expected: + issues.append(f"manifest exclusion overlaps member identity: {exclusion.slug}") + if exclusion.path in member_paths: + issues.append( + f"manifest exclusion overlaps member path: {exclusion.path}" + ) + missing = sorted(set(expected) - set(observed)) + excluded_missing = sorted(set(excluded) - set(observed)) + omitted = sorted(set(observed) - set(expected) - set(excluded)) + path_conflicts: list[str] = [] + type_conflicts: list[str] = [] + excluded_path_conflicts: list[str] = [] + excluded_type_conflicts: list[str] = [] + state_conflicts: list[str] = [] + readme_conflicts: list[str] = [] + for slug in sorted(set(expected) & set(observed)): + member = expected[slug] + record = observed[slug] + observed_path = record.relative_path(root) + if observed_path != member.path: + path_conflicts.append(f"{slug}: manifest={member.path} direct={observed_path}") + if record.record_type != member.record_type: + type_conflicts.append(f"{slug}: manifest={member.record_type} direct={record.record_type}") + if member.lifecycle_state is not None and record.lifecycle_state != member.lifecycle_state: + state_conflicts.append( + f"{slug}: manifest={member.lifecycle_state} direct={record.lifecycle_state}" + ) + if member.readme_sha256 and record.readme_sha256 != member.readme_sha256: + readme_conflicts.append(slug) + for slug in sorted(set(excluded) & set(observed)): + exclusion = excluded[slug] + record = observed[slug] + observed_path = record.relative_path(root) + if observed_path != exclusion.path: + excluded_path_conflicts.append( + f"{slug}: exclusion={exclusion.path} direct={observed_path}" + ) + if record.record_type != exclusion.record_type: + excluded_type_conflicts.append( + f"{slug}: exclusion={exclusion.record_type} direct={record.record_type}" + ) + if missing: + issues.append(f"manifest members missing from direct roots: {missing}") + if excluded_missing: + issues.append(f"manifest exclusions missing from direct roots: {excluded_missing}") + if omitted: + issues.append(f"direct records omitted from manifest: {omitted}") + if path_conflicts: + issues.extend(path_conflicts) + if type_conflicts: + issues.extend(type_conflicts) + if excluded_path_conflicts: + issues.extend(excluded_path_conflicts) + if excluded_type_conflicts: + issues.extend(excluded_type_conflicts) + if state_conflicts: + issues.extend(state_conflicts) + if readme_conflicts: + issues.append(f"manifest README hashes changed: {readme_conflicts}") + return { + "ok": not issues, + "manifest": str(manifest), + "manifest_sha256": sha256_file(manifest), + "schema_version": payload.get("schema_version"), + "profile": payload.get("profile"), + "members": [member.as_dict() for member in members], + "exclusions": [exclusion.as_dict() for exclusion in exclusions], + "records": [record.manifest_entry(root) for record in sorted(scan.records, key=lambda r: r.slug)], + "manifest_count": len(members), + "real_count": len(scan.records), + "missing_members": missing, + "excluded_records": sorted(set(excluded) & set(observed)), + "excluded_missing": excluded_missing, + "excluded_count": len(exclusions), + "omitted_records": omitted, + "path_conflicts": path_conflicts, + "type_conflicts": type_conflicts, + "excluded_path_conflicts": excluded_path_conflicts, + "excluded_type_conflicts": excluded_type_conflicts, + "state_conflicts": state_conflicts, + "readme_conflicts": readme_conflicts, + "issues": issues, + } + + +def coverage_report(manifest_projects: list[str], real_dirs: list[str]) -> dict[str, Any]: + """Compatibility report for callers that only have old slug lists.""" + + expected = list(manifest_projects) + real = list(real_dirs) + set_m, set_r = set(expected), set(real) + return { + "manifest_count": len(expected), + "real_count": len(real), + "missing_dirs": sorted(set_m - set_r), + "omitted_dirs": sorted(set_r - set_m), + "intersection": sorted(set_m & set_r), + "complete": set_m == set_r and len(expected) == len(real), + } + + +def real_record_slugs(mainframe_root: Path | str) -> list[str]: + scan = scan_lifecycle_records(mainframe_root) + return sorted(record.slug for record in scan.records if not record.issues) + + +def _matches(path: str, patterns: Iterable[str]) -> bool: + return any(fnmatch.fnmatch(path, pattern) for pattern in patterns) + + +def selected_source_files( + mainframe_root: Path | str, + payload: dict[str, Any], + members: Iterable[ManifestMember] | None = None, +) -> list[dict[str, Any]]: + """Build the exact Markdown file set the projects ingester should see.""" + + root = Path(mainframe_root).expanduser().resolve() + members = list(members if members is not None else manifest_members(payload)[0]) + include = tuple(payload.get("include") or DEFAULT_INCLUDE) + exclude = tuple(payload.get("exclude") or DEFAULT_EXCLUDE) + rows: list[dict[str, Any]] = [] + for member in sorted(members, key=lambda item: item.slug): + record_root = (root / member.path).resolve() + if not record_root.is_dir() or not record_root.is_relative_to(root): + continue + for path in sorted(record_root.rglob("*.md")): + if path.is_symlink() or not path.is_file(): + continue + rel = path.relative_to(record_root).as_posix() + if include and not _matches(rel, include): + continue + if exclude and _matches(rel, exclude): + continue + source_path = rel + scoped = f"{INDEX_ID}\0{member.slug}\0{source_path}" + doc_id = hashlib.sha256(scoped.encode("utf-8")).hexdigest()[:16] + rows.append( + { + "id": doc_id, + "namespace": member.slug, + "index_id": INDEX_ID, + "trust_profile": TRUST_PROFILE, + "source_root": str(record_root), + "source_path": source_path, + "display_path": f"{member.path}/{source_path}", + "content_hash": sha256_file(path), + "path": str(path), + "record_type": member.record_type, + } + ) + return rows + + +def _open_readonly(db_path: Path) -> sqlite3.Connection: + resolved = db_path.expanduser().resolve() + # `immutable=1` is only truthful for a separately frozen copy. The live + # MindGraph database may have a WAL or a concurrent writer, so use a real + # read-only transaction snapshot instead. + uri = f"file:{quote(str(resolved), safe='/')}?mode=ro" + con = sqlite3.connect(uri, uri=True) + con.row_factory = sqlite3.Row + con.execute("PRAGMA query_only=ON") + con.execute("BEGIN") + return con + + +def document_map_hash(rows: Iterable[dict[str, Any]], *, include_record_type: bool = True) -> str: + keys = ( + "id", + "namespace", + "index_id", + "trust_profile", + "source_root", + "source_path", + "display_path", + "content_hash", + ) + if include_record_type: + keys = (*keys, "record_type") + canonical = [ + { + key: row.get(key) + for key in keys + } + for row in sorted(rows, key=lambda item: (str(item.get("namespace")), str(item.get("source_path")))) + ] + return hashlib.sha256(json.dumps(canonical, sort_keys=True, separators=(",", ":")).encode()).hexdigest() + + +def verify_typed_document_map( + db_path: Path | str, + mainframe_root: Path | str, + manifest_path: Path | str, + *, + projects_root: Path | str | None = None, + operations_root: Path | str | None = None, +) -> dict[str, Any]: + """Verify exact typed rows, paths, IDs, and content hashes in a DB.""" + + root = Path(mainframe_root).expanduser().resolve() + manifest = Path(manifest_path).expanduser().resolve() + coverage = validate_manifest( + root, + manifest, + projects_root=projects_root, + operations_root=operations_root, + ) + result: dict[str, Any] = { + "ok": False, + "db": str(Path(db_path).expanduser().resolve()), + "manifest": str(manifest), + "coverage_ok": coverage.get("ok", False), + "coverage": coverage, + "expected_count": 0, + "observed_count": 0, + "missing_documents": [], + "extra_documents": [], + "identity_mismatches": [], + "path_mismatches": [], + "content_mismatches": [], + "provenance_mismatches": [], + "duplicate_database_ids": [], + "record_type_source": "unknown", + "record_type_direct_check": False, + "record_type_derived_check": False, + "record_type_mismatches": [], + "record_type_unknown_documents": [], + "document_map_ok": False, + "read_snapshot": None, + "expected_document_map_hash": None, + "observed_document_map_hash": None, + "expected_typed_document_map_hash": None, + "observed_typed_document_map_hash": None, + } + if not coverage.get("ok"): + return result + payload = load_manifest(manifest) + members, _ = manifest_members(payload) + expected_rows = selected_source_files(root, payload, members) + result["expected_count"] = len(expected_rows) + result["expected_document_map_hash"] = document_map_hash(expected_rows, include_record_type=False) + result["expected_typed_document_map_hash"] = document_map_hash(expected_rows, include_record_type=True) + db = Path(db_path).expanduser().resolve() + if not db.is_file(): + result["extra_documents"] = ["<database missing>"] + return result + try: + con = _open_readonly(db) + try: + table = {row["name"] for row in con.execute("SELECT name FROM sqlite_master WHERE type='table'")} + required = {"documents"} + missing_tables = sorted(required - table) + if missing_tables: + result["extra_documents"] = [f"missing tables: {missing_tables}"] + return result + columns = {row[1] for row in con.execute("PRAGMA table_info(documents)")} + has_record_type = "record_type" in columns + result["record_type_source"] = "documents.record_type" if has_record_type else "derived_from_manifest" + selected = "id, namespace, index_id, trust_profile, source_root, source_path, display_path, content_hash" + if has_record_type: + selected += ", record_type" + rows = [dict(row) for row in con.execute(f"SELECT {selected} FROM documents")] + result["read_snapshot"] = { + "mode": "read_only_transaction", + "journal_mode": str(con.execute("PRAGMA journal_mode").fetchone()[0]), + "query_only": True, + } + finally: + con.close() + except (OSError, sqlite3.Error) as exc: + result["extra_documents"] = [f"database unreadable: {type(exc).__name__}: {exc}"] + return result + result["observed_count"] = len(rows) + expected_by_id = {row["id"]: row for row in expected_rows} + observed_by_id: dict[str, dict[str, Any]] = {} + duplicate_ids: list[str] = [] + for row in rows: + doc_id = str(row.get("id")) + if doc_id in observed_by_id: + duplicate_ids.append(doc_id) + observed_by_id[doc_id] = row + result["duplicate_database_ids"] = sorted(set(duplicate_ids)) + result["missing_documents"] = sorted(set(expected_by_id) - set(observed_by_id)) + result["extra_documents"] = sorted(set(observed_by_id) - set(expected_by_id)) + for doc_id in sorted(set(expected_by_id) & set(observed_by_id)): + expected = expected_by_id[doc_id] + observed = observed_by_id[doc_id] + if observed.get("namespace") != expected["namespace"]: + result["identity_mismatches"].append(doc_id) + for key in ("index_id", "trust_profile", "source_root", "source_path", "display_path"): + if observed.get(key) != expected[key]: + result["path_mismatches" if key in {"source_root", "source_path", "display_path"} else "provenance_mismatches"].append( + {"id": doc_id, "field": key, "expected": expected[key], "observed": observed.get(key)} + ) + if observed.get("content_hash") != expected["content_hash"]: + result["content_mismatches"].append( + {"id": doc_id, "expected": expected["content_hash"], "observed": observed.get("content_hash")} + ) + result["document_map_ok"] = bool( + not result["missing_documents"] + and not result["extra_documents"] + and not result["identity_mismatches"] + and not result["path_mismatches"] + and not result["content_mismatches"] + and not result["provenance_mismatches"] + and not result["duplicate_database_ids"] + ) + if result["record_type_source"] == "documents.record_type": + result["record_type_direct_check"] = True + for row in rows: + doc_id = str(row.get("id")) + expected = expected_by_id.get(doc_id) + observed_type = row.get("record_type") + if expected is None: + continue + if observed_type is None or observed_type == "": + result["record_type_unknown_documents"].append(doc_id) + elif observed_type != expected.get("record_type"): + result["record_type_mismatches"].append( + {"id": doc_id, "expected": expected.get("record_type"), "observed": observed_type} + ) + else: + # The current schema has no typed lifecycle column. Do not inject the + # manifest's expected type into observed rows or into the observed DB + # hash and call that a DB check. The manifest's classification is + # still authoritative because validate_manifest proved its typed + # members against the direct lifecycle README authorities. + result["record_type_direct_check"] = False + result["record_type_derived_check"] = bool(result["coverage_ok"]) + result["record_type_unknown_documents"] = sorted(set(result["record_type_unknown_documents"])) + result["record_type_mismatches"] = sorted(result["record_type_mismatches"], key=lambda item: str(item.get("id"))) + observed_rows = [{**row} for row in rows] + result["observed_document_map_hash"] = document_map_hash(observed_rows, include_record_type=False) + if result["record_type_source"] == "documents.record_type": + result["observed_typed_document_map_hash"] = document_map_hash(observed_rows, include_record_type=True) + result["ok"] = bool( + result["coverage_ok"] + and result["document_map_ok"] + and ( + result["record_type_direct_check"] + or result["record_type_derived_check"] + ) + and not result["record_type_mismatches"] + and not result["record_type_unknown_documents"] + ) + return result + + +def inventory_payload(mainframe_root: Path | str) -> dict[str, Any]: + root = Path(mainframe_root).expanduser().resolve() + scan = scan_lifecycle_records(root) + records = [record.manifest_entry(root) for record in sorted(scan.records, key=lambda r: r.slug)] + return { + "schema_version": 1, + "kind": "mainframe-work-inventory", + "root": str(root), + "records": records, + "count": len(records), + "issues": scan.issues, + "ok": not scan.issues, + } diff --git a/tests/conftest.py b/tests/conftest.py new file mode 100644 index 0000000..f22faf7 --- /dev/null +++ b/tests/conftest.py @@ -0,0 +1,28 @@ +"""Pytest conftest hook for MainFrame Brain sensory afference.""" +from __future__ import annotations + +import os +import subprocess +import sys +from pathlib import Path + + +def pytest_sessionfinish(session, exitstatus): + """Notify MainFrame Brain of test completion. + Executes in <1ms directly in Python; completely fails open so it never breaks tests. + """ + if os.environ.get("MAINFRAME_BRAIN_NO_HOOK") == "1": + return + + try: + root = Path(__file__).resolve().parents[1] + brain_path = str(root / "30_projects/mainframe-brain") + if brain_path not in sys.path: + sys.path.insert(0, brain_path) + from engine import storage + total = getattr(session, "testscollected", 0) + status_label = "passed" if exitstatus == 0 else f"failed (exit code {exitstatus})" + summary = f"Test suite {status_label}: {total} tests collected" + storage.add_event(source="test", summary=summary) + except Exception: + pass diff --git a/tests/test_aider_watcher.py b/tests/test_aider_watcher.py deleted file mode 100644 index 4bfaa09..0000000 --- a/tests/test_aider_watcher.py +++ /dev/null @@ -1,150 +0,0 @@ -from __future__ import annotations - -import importlib.util -import tempfile -import unittest -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -WATCHER_PATH = ROOT / "bin" / "aider-watcher" -LOADER = SourceFileLoader("aider_watcher", str(WATCHER_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -aider_watcher = importlib.util.module_from_spec(SPEC) -SPEC.loader.exec_module(aider_watcher) - - -class AiderWatcherParserTests(unittest.TestCase): - def parse(self, text: str) -> list[dict[str, object]]: - events: list[dict[str, object]] = [] - path = ROOT / "30_projects" / "example" / ".aider.chat.history.md" - state = aider_watcher.new_state() - aider_watcher.process_bytes( - text.encode(), - path, - state, - events.append, - final=True, - ) - aider_watcher.end_session(path, state, events.append, "test_eof") - return events - - def test_one_start_per_session_and_next_session_closes_previous(self) -> None: - events = self.parse( - """ -# aider chat started at 2026-06-13 10:00:00 -> Model: ollama_chat/qwen2.5-coder:14b with diff edit format -#### Fix one warning. -# aider chat started at 2026-06-13 10:01:00 -> Model: ollama_chat/qwen2.5-coder:14b with diff edit format -#### Fix the second warning. -""" - ) - names = [event["hook_event_name"] for event in events] - self.assertEqual(names.count("SessionStart"), 2) - self.assertEqual(names.count("UserPromptSubmit"), 2) - self.assertEqual(names.count("SessionEnd"), 2) - session_ids = { - event["session_id"] - for event in events - if event["hook_event_name"] == "SessionStart" - } - self.assertEqual(len(session_ids), 2) - - def test_records_coarse_diagnostics_without_model_output(self) -> None: - transcript = """ -# aider chat started at 2026-06-13 10:00:00 -> Model: ollama_chat/qwen2.5-coder:14b with diff edit format -> Your estimated chat context of 47,226 tokens exceeds the 24,576 token limit! -> The LLM did not conform to the edit format. -> # 1 SEARCH/REPLACE block failed to match! -""" - events = self.parse(transcript) - diagnostics = [ - event["diagnostic"] - for event in events - if event["hook_event_name"] == "Diagnostic" - ] - self.assertEqual( - diagnostics, - ["context_limit_exceeded", "edit_format_failure"], - ) - self.assertNotIn("47,226", repr(events)) - - def test_denied_shell_command_is_not_recorded_as_execution(self) -> None: - events = self.parse( - """ -# aider chat started at 2026-06-13 10:00:00 -> Model: ollama_chat/qwen2.5-coder:14b with diff edit format -> git am ../private.patch -> Run shell commands? (Y)es/(N)o [Yes]: n -""" - ) - self.assertFalse( - any(event.get("tool_name") == "Bash" for event in events) - ) - self.assertTrue( - any( - event.get("diagnostic") == "shell_command_denied" - for event in events - ) - ) - - def test_applied_edits_and_direct_tests_are_observed(self) -> None: - events = self.parse( - """ -# aider chat started at 2026-06-13 10:00:00 -> Model: ollama_chat/qwen2.5-coder:14b with diff edit format -# The other 2 SEARCH/REPLACE blocks were applied successfully. -/run pytest -q -""" - ) - edit = next(event for event in events if event.get("tool_name") == "Edit") - test = next(event for event in events if event.get("tool_name") == "Bash") - self.assertEqual(edit["observed_count"], 2) - self.assertEqual(test["tool_purpose"], "verification") - - def test_only_verification_shaped_package_commands_are_marked(self) -> None: - self.assertFalse(aider_watcher.is_test_command("npm install")) - self.assertTrue(aider_watcher.is_test_command("npm test")) - self.assertTrue( - aider_watcher.is_test_command("python3 -m unittest discover") - ) - - -class AiderWatcherFileTests(unittest.TestCase): - def test_new_history_can_be_read_from_the_beginning(self) -> None: - events: list[dict[str, object]] = [] - with tempfile.TemporaryDirectory() as tmp: - path = Path(tmp) / ".aider.chat.history.md" - path.write_text( - "# aider chat started at 2026-06-13 10:00:00\n" - "> Model: local/model with diff edit format\n", - encoding="utf-8", - ) - state = aider_watcher.new_state() - aider_watcher.read_growth(path, state, events.append, final=True) - - self.assertEqual( - [event["hook_event_name"] for event in events], - ["SessionStart"], - ) - - def test_existing_history_offset_does_not_replay_old_content(self) -> None: - events: list[dict[str, object]] = [] - with tempfile.TemporaryDirectory() as tmp: - path = Path(tmp) / ".aider.chat.history.md" - path.write_text( - "# aider chat started at 2026-06-13 10:00:00\n", - encoding="utf-8", - ) - state = aider_watcher.new_state(size=path.stat().st_size) - aider_watcher.read_growth(path, state, events.append) - - self.assertEqual(events, []) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_audit_sweep.py b/tests/test_audit_sweep.py deleted file mode 100644 index 1877b3e..0000000 --- a/tests/test_audit_sweep.py +++ /dev/null @@ -1,131 +0,0 @@ -"""Focused tests for bin/audit-sweep candidate discovery and CLI validation.""" - -from __future__ import annotations - -import importlib.util -import json -import os -import subprocess -import sys -import tempfile -import unittest -from datetime import date, datetime, time, timedelta -from importlib.machinery import SourceFileLoader -from pathlib import Path -from unittest.mock import patch - - -ROOT = Path(__file__).resolve().parents[1] -SWEEP_PATH = ROOT / "bin" / "audit-sweep" -LOADER = SourceFileLoader("audit_sweep", str(SWEEP_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -audit_sweep = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = audit_sweep -SPEC.loader.exec_module(audit_sweep) - - -class AuditSweepTests(unittest.TestCase): - def test_parse_frontmatter_tags_supports_inline_and_block_lists(self) -> None: - inline = """--- -tags: ["x-capture", "needs-audit", "clippings"] ---- -""" - block = """--- -tags: -- needs-audit -- raw -- pdf ---- -""" - - self.assertEqual( - audit_sweep.parse_simple_frontmatter_tags(inline), - ["x-capture", "needs-audit", "clippings"], - ) - self.assertEqual( - audit_sweep.parse_simple_frontmatter_tags(block), - ["needs-audit", "raw", "pdf"], - ) - self.assertEqual( - audit_sweep.parse_simple_frontmatter_tags("just body"), - [], - ) - - def test_max_days_limits_samples_but_not_explicit_audit_signals(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - knowledge = root / "10_knowledge" - domain = knowledge / "example" - domain.mkdir(parents=True) - - at_boundary = domain / "2026-01-01__example__raw__boundary.md" - too_old = domain / "2026-01-01__example__raw__old.md" - explicit = domain / "2026-01-01__example__note__audit.md" - at_boundary.write_text("---\ntags: [raw]\n---\n", encoding="utf-8") - too_old.write_text("---\ntags: [raw]\n---\n", encoding="utf-8") - explicit.write_text( - "---\ntags: [needs-audit]\n---\n", - encoding="utf-8", - ) - - self._set_age(at_boundary, 7) - self._set_age(too_old, 8) - self._set_age(explicit, 30) - - with ( - patch.object(audit_sweep, "ROOT", root), - patch.object(audit_sweep, "KNOWLEDGE", knowledge), - ): - candidates = audit_sweep.collect_candidates("example", max_days=7) - - by_name = {candidate["name"]: candidate for candidate in candidates} - self.assertEqual(set(by_name), {at_boundary.name, explicit.name}) - self.assertFalse(by_name[at_boundary.name]["has_needs_audit"]) - self.assertTrue(by_name[explicit.name]["has_needs_audit"]) - - def test_cli_rejects_negative_max_days(self) -> None: - result = subprocess.run( - [str(SWEEP_PATH), "--dry-run", "--json", "--max-days", "-1"], - capture_output=True, - text=True, - timeout=30, - ) - - self.assertEqual(result.returncode, 2) - self.assertIn("must be zero or greater", result.stderr) - self.assertEqual(result.stdout, "") - - def test_dry_run_json_reports_subset_and_counts(self) -> None: - result = subprocess.run( - [ - str(SWEEP_PATH), - "--dry-run", - "--json", - "--subset", - "regulated-systems", - "--max-days", - "7", - ], - capture_output=True, - text=True, - timeout=30, - ) - - self.assertEqual(result.returncode, 0, msg=result.stderr or result.stdout) - data = json.loads(result.stdout) - self.assertEqual(data["subset"], "regulated-systems") - self.assertGreaterEqual(data["candidate_count"], 0) - self.assertGreaterEqual(data["explicit_needs_audit"], 0) - - @staticmethod - def _set_age(path: Path, days: int) -> None: - modified = datetime.combine( - date.today() - timedelta(days=days), - time(hour=12), - ).timestamp() - os.utime(path, (modified, modified)) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_capture_validate.py b/tests/test_capture_validate.py deleted file mode 100644 index 89d875c..0000000 --- a/tests/test_capture_validate.py +++ /dev/null @@ -1,52 +0,0 @@ -from __future__ import annotations - -import importlib.util -import sys -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -LOADER = SourceFileLoader("capture_validate", str(ROOT / "bin" / "capture-validate")) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -capture_validate = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = capture_validate -SPEC.loader.exec_module(capture_validate) - - -def write_capture(tmp_path: Path, source: str) -> Path: - path = tmp_path / "capture.md" - path.write_text( - "---\n" - "title: \"Test capture\"\n" - "type: raw\n" - f"source: \"{source}\"\n" - "---\n\nBody.\n", - encoding="utf-8", - ) - return path - - -def test_bare_domain_source_is_a_warning(tmp_path: Path) -> None: - findings = capture_validate.validate(write_capture(tmp_path, "https://ollama.com")) - - assert any( - finding.rule == "R7" and finding.severity == "warn" for finding in findings - ) - - -def test_specific_document_source_is_not_a_bare_domain(tmp_path: Path) -> None: - findings = capture_validate.validate( - write_capture(tmp_path, "https://ollama.com/library/llama3") - ) - - assert not any(finding.rule == "R7" for finding in findings) - - -def test_query_or_fragment_does_not_count_as_a_bare_domain(tmp_path: Path) -> None: - findings = capture_validate.validate( - write_capture(tmp_path, "https://example.org/?page=about") - ) - - assert not any(finding.rule == "R7" for finding in findings) diff --git a/tests/test_eval_schedule.py b/tests/test_eval_schedule.py deleted file mode 100644 index 9cc91d6..0000000 --- a/tests/test_eval_schedule.py +++ /dev/null @@ -1,344 +0,0 @@ -"""Tests for scripts/eval_schedule.py.""" - -from __future__ import annotations - -import json -import tempfile -import unittest -from datetime import datetime, timezone -from pathlib import Path -from unittest import mock - -import scripts.eval_schedule as es - - -class EvalScheduleTests(unittest.TestCase): - def test_steps_for_daily_excludes_tests_and_probe(self) -> None: - steps = es.steps_for_cadence( - "daily", skip_tests=False, skip_probe=False, full_probe=False, run_id="test-run" - ) - names = [s[0] for s in steps] - self.assertIn("ingest_minion_dry_run", names) - self.assertNotIn("unittest_suite", names) - self.assertNotIn("eval_registry_harvest", names) - - def test_steps_for_weekly_includes_tests_not_mid_harvest(self) -> None: - steps = es.steps_for_cadence( - "weekly", skip_tests=False, skip_probe=True, full_probe=False, run_id="test-run" - ) - names = [s[0] for s in steps] - self.assertIn("unittest_suite", names) - self.assertNotIn("eval_registry_harvest", names) - self.assertNotIn("mindgraph_retrieval_probe", names) - - def test_scheduled_probe_uses_fused_subset(self) -> None: - steps = es.steps_for_cadence( - "weekly", skip_tests=True, skip_probe=False, full_probe=False, run_id="2026-06-22T120000-scheduled-weekly" - ) - probe = next(s for s in steps if s[0] == "mindgraph_retrieval_probe") - cmd = probe[1] or [] - self.assertIn("--fused-only", cmd) - self.assertIn("--query-ids", cmd) - self.assertIn("--registry", cmd) - - def test_live_envelope_uses_mindgraph_uv_env(self) -> None: - """Live envelope imports PyYAML; bare host python fails (weekly 2026-07-12).""" - steps = es.steps_for_cadence( - "weekly", - skip_tests=True, - skip_probe=False, - full_probe=False, - run_id="2026-07-12T235432-scheduled-weekly", - ) - step = next(s for s in steps if s[0] == "mindgraph_live_envelope") - cmd = step[1] - if cmd is None: - # uv or mindgraph/ absent in this environment — skip is allowed. - self.assertIsNotNone(step[2]) - return - self.assertEqual(cmd[0:4], [cmd[0], "run", "--project", str(es.MINDGRAPH_PROJECT)]) - self.assertIn("uv", Path(cmd[0]).name) - self.assertEqual(cmd[4], "python") - self.assertEqual(cmd[5], str(es.MINDGRAPH_LIVE_ENVELOPE)) - self.assertIn("--run-id", cmd) - self.assertIn("2026-07-12T235432-scheduled-weekly-live-envelope", cmd) - - def test_live_envelope_missing_frozen_snapshot_is_operator_gate(self) -> None: - with mock.patch.object( - es, "missing_eval_snapshots", return_value=[Path("/tmp/frozen-mainframe-intent.sqlite")] - ): - steps = es.steps_for_cadence( - "weekly", skip_tests=True, skip_probe=False, full_probe=False, run_id="test-run" - ) - step = next(s for s in steps if s[0] == "mindgraph_live_envelope") - self.assertIsNone(step[1]) - self.assertTrue((step[2] or "").startswith("OPERATOR GATE:")) - - def test_failed_step_summary_names_failed_and_operator_gated_steps(self) -> None: - summary = es.failed_step_summary( - {"steps": [ - {"name": "ok", "exit_code": 0}, - {"name": "probe", "exit_code": 2}, - {"name": "snapshot", "operator_gated": True, "skipped": True}, - ]} - ) - self.assertEqual(summary, "probe (exit 2), snapshot (operator gate)") - - def test_mindgraph_uv_python_cmd_shape(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - project = Path(tmp) - with mock.patch.object(es.shutil, "which", return_value="/usr/bin/uv"): - with mock.patch.object(es, "MINDGRAPH_PROJECT", project): - cmd = es.mindgraph_uv_python_cmd( - Path("/tmp/probe.py"), "--run-id", "rid-1" - ) - self.assertEqual( - cmd, - [ - "/usr/bin/uv", - "run", - "--project", - str(project), - "python", - "/tmp/probe.py", - "--run-id", - "rid-1", - ], - ) - - def test_mindgraph_uv_python_cmd_missing_uv(self) -> None: - with mock.patch.object(es.shutil, "which", return_value=None): - self.assertIsNone( - es.mindgraph_uv_python_cmd(Path("/tmp/probe.py"), "--run-id", "x") - ) - - def test_parse_harvest_stats(self) -> None: - stats = es.parse_harvest_stats( - "harvest: projects=5 files_scanned=42 runs_new=1 metrics=4 irregularities=0 skipped=37 errors=2" - ) - self.assertEqual(stats["errors"], 2) - self.assertEqual(stats["runs_new"], 1) - - def test_run_step_attaches_process_ids(self) -> None: - completed = mock.Mock(returncode=0, stdout="", stderr="") - with mock.patch.object(es.subprocess, "run", return_value=completed): - step = es.run_step( - "workflow_report_7d", - ["bin/workflow-report", "--days", "7", "--json"], - cwd=Path("/tmp"), - ) - - self.assertEqual( - step.process_ids, - ["cli-workflow-report", "workflow-workflow-telemetry"], - ) - - def test_write_process_eval_output_skips_daily(self) -> None: - run = es.ScheduleRun( - run_id="2026-06-22-scheduled-daily", - cadence="daily", - started_at="t0", - finished_at="t1", - git_sha="abc", - all_passed=True, - steps=[es.StepResult("ingest_minion_dry_run", ["bin"], 0, 1.0)], - ) - with tempfile.TemporaryDirectory() as tmp: - with mock.patch.object(es, "PROCESS_EVAL_OUTPUTS", Path(tmp)): - out = es.write_process_eval_output(run, dry_run=False) - self.assertIsNone(out) - - def test_write_process_eval_output_weekly(self) -> None: - run = es.ScheduleRun( - run_id="2026-06-22-scheduled-weekly", - cadence="weekly", - started_at="t0", - finished_at="t1", - git_sha="abc", - all_passed=True, - steps=[ - es.StepResult("unittest_suite", ["py"], 0, 2.0), - es.StepResult("eval_registry_harvest", ["bin"], 0, 0.5), - ], - ) - with tempfile.TemporaryDirectory() as tmp: - with mock.patch.object(es, "PROCESS_EVAL_OUTPUTS", Path(tmp)): - out = es.write_process_eval_output(run, dry_run=False) - self.assertIsNotNone(out) - text = out.read_text(encoding="utf-8") - self.assertIn("## Metric extract (eval-registry)", text) - self.assertIn("scheduled_steps_passed", text) - - def test_append_schedule_log(self) -> None: - run = es.ScheduleRun( - run_id="2026-06-22-scheduled-daily", - cadence="daily", - started_at="t0", - finished_at="t1", - git_sha=None, - all_passed=True, - steps=[], - ) - with tempfile.TemporaryDirectory() as tmp: - log = Path(tmp) / "schedule-runs.jsonl" - with mock.patch.object(es, "SCHEDULE_LOG", log), mock.patch.object(es, "REGISTRY_DIR", Path(tmp)): - es.append_schedule_log(run, dry_run=False) - lines = log.read_text(encoding="utf-8").strip().splitlines() - self.assertEqual(len(lines), 1) - payload = json.loads(lines[0]) - self.assertEqual(payload["run_id"], run.run_id) - - def test_build_plist_contains_cadence(self) -> None: - plist = es.build_plist(es.DAEMON_LABEL_WEEKLY, "weekly", Path("/tmp/MainFrame")) - self.assertIn(es.DAEMON_LABEL_WEEKLY, plist) - self.assertIn("weekly", plist) - self.assertIn("StartCalendarInterval", plist) - self.assertIn("MAINFRAME_EVAL_TRIGGER", plist) - self.assertIn("launchd", plist) - self.assertIn("scripts/eval_schedule.py", plist) - self.assertIn("MAINFRAME_ROOT", plist) - # WorkingDirectory should not force Desktop chdir - self.assertIn(str(Path.home()), plist) - - def test_detect_launchd_tcc_blocks(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - log_dir = Path(tmp) - err = log_dir / "weekly.err" - err.write_text( - "bash: /path/to/MainFrame/bin/eval-schedule: Operation not permitted\n", - encoding="utf-8", - ) - hits = es.detect_launchd_tcc_blocks(log_dirs=[log_dir]) - self.assertEqual(hits, [str(err)]) - - def test_root_needs_full_disk_access(self) -> None: - self.assertTrue(es.root_needs_full_disk_access(Path("/Users/x/Desktop/MainFrame"))) - self.assertFalse(es.root_needs_full_disk_access(Path("/Users/x/src/MainFrame"))) - - def test_assess_schedule_health_stale_weekly(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - log = Path(tmp) / "schedule-runs.jsonl" - old = { - "run_id": "old-weekly", - "cadence": "weekly", - "finished_at": "2020-01-01T00:00:00+00:00", - "all_passed": True, - "trigger": "launchd", - } - log.write_text(json.dumps(old) + "\n", encoding="utf-8") - with mock.patch.object(es, "SCHEDULE_LOG", log): - with mock.patch.object(es, "plist_path", return_value=Path(tmp) / "missing.plist"): - with mock.patch.object(es, "detect_launchd_tcc_blocks", return_value=[]): - with mock.patch.object( - es, "root_needs_full_disk_access", return_value=False - ): - health = es.assess_schedule_health( - now=datetime(2026, 6, 22, tzinfo=timezone.utc) - ) - self.assertFalse(health.ok) - self.assertTrue(any("stale" in p for p in health.problems)) - - def test_assess_schedule_health_ok_when_recent_launchd(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - log = Path(tmp) / "schedule-runs.jsonl" - recent = { - "run_id": "new-weekly", - "cadence": "weekly", - "finished_at": "2026-06-22T10:00:00+00:00", - "all_passed": True, - "trigger": "launchd", - } - log.write_text(json.dumps(recent) + "\n", encoding="utf-8") - weekly_plist = Path(tmp) / "weekly.plist" - daily_plist = Path(tmp) / "daily.plist" - weekly_plist.write_text("plist", encoding="utf-8") - daily_plist.write_text("plist", encoding="utf-8") - - def fake_plist(label: str) -> Path: - if label == es.DAEMON_LABEL_WEEKLY: - return weekly_plist - return daily_plist - - with mock.patch.object(es, "SCHEDULE_LOG", log): - with mock.patch.object(es, "plist_path", side_effect=fake_plist): - with mock.patch.object(es, "detect_launchd_tcc_blocks", return_value=[]): - with mock.patch.object( - es, "root_needs_full_disk_access", return_value=False - ): - health = es.assess_schedule_health( - now=datetime(2026, 6, 22, 12, 0, tzinfo=timezone.utc) - ) - self.assertTrue(health.ok) - self.assertEqual(health.status_label, "OK") - - def test_assess_schedule_health_degraded_without_launchd_provenance(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - log = Path(tmp) / "schedule-runs.jsonl" - recent = { - "run_id": "manual-weekly", - "cadence": "weekly", - "finished_at": "2026-06-22T10:00:00+00:00", - "all_passed": True, - # no trigger field → treated as missing launchd provenance - } - log.write_text(json.dumps(recent) + "\n", encoding="utf-8") - weekly_plist = Path(tmp) / "weekly.plist" - daily_plist = Path(tmp) / "daily.plist" - weekly_plist.write_text("plist", encoding="utf-8") - daily_plist.write_text("plist", encoding="utf-8") - - def fake_plist(label: str) -> Path: - if label == es.DAEMON_LABEL_WEEKLY: - return weekly_plist - return daily_plist - - with mock.patch.object(es, "SCHEDULE_LOG", log): - with mock.patch.object(es, "plist_path", side_effect=fake_plist): - with mock.patch.object(es, "detect_launchd_tcc_blocks", return_value=[]): - with mock.patch.object( - es, "root_needs_full_disk_access", return_value=False - ): - health = es.assess_schedule_health( - now=datetime(2026, 6, 22, 12, 0, tzinfo=timezone.utc) - ) - self.assertFalse(health.ok) - self.assertEqual(health.status_label, "DEGRADED") - self.assertTrue(any("provenance" in d for d in health.degraded)) - - def test_assess_schedule_health_tcc_denial_is_problem(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - log = Path(tmp) / "schedule-runs.jsonl" - recent = { - "run_id": "new-weekly", - "cadence": "weekly", - "finished_at": "2026-06-22T10:00:00+00:00", - "all_passed": True, - "trigger": "launchd", - } - log.write_text(json.dumps(recent) + "\n", encoding="utf-8") - weekly_plist = Path(tmp) / "weekly.plist" - daily_plist = Path(tmp) / "daily.plist" - weekly_plist.write_text("plist", encoding="utf-8") - daily_plist.write_text("plist", encoding="utf-8") - - def fake_plist(label: str) -> Path: - if label == es.DAEMON_LABEL_WEEKLY: - return weekly_plist - return daily_plist - - with mock.patch.object(es, "SCHEDULE_LOG", log): - with mock.patch.object(es, "plist_path", side_effect=fake_plist): - with mock.patch.object( - es, - "detect_launchd_tcc_blocks", - return_value=["/tmp/weekly.err"], - ): - health = es.assess_schedule_health( - now=datetime(2026, 6, 22, 12, 0, tzinfo=timezone.utc) - ) - self.assertFalse(health.ok) - self.assertTrue(any("TCC" in p for p in health.problems)) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_extract_knowledge.py b/tests/test_extract_knowledge.py deleted file mode 100644 index 5aed8d6..0000000 --- a/tests/test_extract_knowledge.py +++ /dev/null @@ -1,283 +0,0 @@ -from __future__ import annotations - -import importlib.util -import json -import sys -import tempfile -import unittest -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -SCRIPT_PATH = ROOT / "bin" / "extract-knowledge" -LOADER = SourceFileLoader("extract_knowledge", str(SCRIPT_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -mod = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = mod -SPEC.loader.exec_module(mod) - -run_extract = mod.run_extract -slugify = mod.slugify - - -def make_project(root: Path, slug: str = "test-proj", - log: bool = True, decisions: bool = True) -> None: - proj = root / "30_projects" / slug - proj.mkdir(parents=True) - (proj / "README.md").write_text("# Test Project\n", encoding="utf-8") - if log: - (proj / "log.md").write_text("# Log\n", encoding="utf-8") - if decisions: - (proj / "decisions.md").write_text("# Decisions\n", encoding="utf-8") - - -def make_domain(root: Path, domain: str = "ai-systems") -> None: - (root / "10_knowledge" / domain).mkdir(parents=True) - - -class SlugifyTests(unittest.TestCase): - def test_basic(self) -> None: - self.assertEqual(slugify("My Great Insight"), "my-great-insight") - - def test_special_chars(self) -> None: - self.assertEqual(slugify("Test: A (thing)!"), "test-a-thing") - - def test_leading_trailing(self) -> None: - self.assertEqual(slugify(" --hello-- "), "hello") - - -class ExtractKnowledgeTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - - def tearDown(self) -> None: - self.tmp.cleanup() - - def test_all_prerequisites_met(self) -> None: - make_project(self.root) - make_domain(self.root) - result = run_extract(self.root, "test-proj", "ai-systems", - "Test Note", ["extracted"]) - self.assertTrue(result.ok) - self.assertFalse(result.written) - self.assertIn("ai-systems", result.target_path) - - def test_missing_project_dir(self) -> None: - make_domain(self.root) - result = run_extract(self.root, "nonexistent", "ai-systems", - "Test Note", ["extracted"]) - self.assertFalse(result.ok) - failed = [p for p in result.prerequisites if not p.exists and p.required] - self.assertTrue(len(failed) >= 1) - - def test_missing_readme(self) -> None: - proj = self.root / "30_projects" / "test-proj" - proj.mkdir(parents=True) - make_domain(self.root) - result = run_extract(self.root, "test-proj", "ai-systems", - "Test Note", ["extracted"]) - self.assertFalse(result.ok) - readme_prereq = [p for p in result.prerequisites - if p.name == "project README"][0] - self.assertFalse(readme_prereq.exists) - self.assertTrue(readme_prereq.required) - - def test_missing_log_warns_but_ok(self) -> None: - make_project(self.root, log=False) - make_domain(self.root) - result = run_extract(self.root, "test-proj", "ai-systems", - "Test Note", ["extracted"]) - self.assertTrue(result.ok) - log_prereq = [p for p in result.prerequisites - if p.name == "project log"][0] - self.assertFalse(log_prereq.exists) - self.assertFalse(log_prereq.required) - - def test_missing_decisions_warns_but_ok(self) -> None: - make_project(self.root, decisions=False) - make_domain(self.root) - result = run_extract(self.root, "test-proj", "ai-systems", - "Test Note", ["extracted"]) - self.assertTrue(result.ok) - dec_prereq = [p for p in result.prerequisites - if p.name == "project decisions"][0] - self.assertFalse(dec_prereq.exists) - self.assertFalse(dec_prereq.required) - - def test_unknown_domain(self) -> None: - make_project(self.root) - result = run_extract(self.root, "test-proj", "unknown-domain", - "Test Note", ["extracted"]) - self.assertFalse(result.ok) - domain_prereq = [p for p in result.prerequisites - if p.name == "knowledge domain"][0] - self.assertFalse(domain_prereq.exists) - - def test_rejects_path_bearing_project_and_domain_values(self) -> None: - make_project(self.root) - make_domain(self.root) - cases = ( - ("../test-proj", "ai-systems"), - ("test-proj", "../ai-systems"), - ("test-proj", "/tmp/ai-systems"), - ) - for project, domain in cases: - with self.subTest(project=project, domain=domain): - result = run_extract( - self.root, - project, - domain, - "Test Note", - ["extracted"], - write=True, - ) - self.assertFalse(result.ok) - self.assertIsNotNone(result.input_error) - self.assertIsNone(result.target_path) - self.assertFalse(result.written) - - def test_rejects_empty_or_multiline_titles(self) -> None: - make_project(self.root) - make_domain(self.root) - for title in ("", "---", "Injected\nstatus: stable"): - with self.subTest(title=title): - result = run_extract( - self.root, - "test-proj", - "ai-systems", - title, - ["extracted"], - write=True, - ) - self.assertFalse(result.ok) - self.assertEqual( - result.input_error, "title must be a non-empty single-line title" - ) - self.assertFalse(result.written) - - def test_rejects_symlinked_domain_that_escapes_knowledge_root(self) -> None: - make_project(self.root) - outside = self.root / "outside" - outside.mkdir() - knowledge = self.root / "10_knowledge" - knowledge.mkdir() - (knowledge / "escaped").symlink_to(outside, target_is_directory=True) - - result = run_extract( - self.root, - "test-proj", - "escaped", - "Test Note", - ["extracted"], - write=True, - ) - - self.assertFalse(result.ok) - self.assertFalse(result.written) - self.assertEqual(list(outside.iterdir()), []) - - def test_broken_target_symlink_cannot_redirect_write(self) -> None: - make_project(self.root) - make_domain(self.root) - check = run_extract( - self.root, "test-proj", "ai-systems", "Test Note", ["extracted"] - ) - outside = self.root / "outside-note.md" - target = self.root / check.target_path - target.symlink_to(outside) - - result = run_extract( - self.root, - "test-proj", - "ai-systems", - "Test Note", - ["extracted"], - write=True, - ) - - self.assertFalse(result.ok) - self.assertTrue(result.collision) - self.assertFalse(result.written) - self.assertFalse(outside.exists()) - - def test_write_creates_scaffold(self) -> None: - make_project(self.root) - make_domain(self.root) - result = run_extract(self.root, "test-proj", "ai-systems", - "Test Note", ["extracted"], write=True) - self.assertTrue(result.ok) - self.assertTrue(result.written) - target = self.root / result.target_path - self.assertTrue(target.exists()) - content = target.read_text(encoding="utf-8") - self.assertIn('title: "Test Note"', content) - self.assertIn('domain: "ai-systems"', content) - self.assertIn('type: "note"', content) - self.assertIn('status: "queued"', content) - self.assertIn('source: "30_projects/test-proj/README.md"', content) - self.assertIn("# Test Note", content) - - def test_collision_blocks_write(self) -> None: - make_project(self.root) - make_domain(self.root) - result1 = run_extract(self.root, "test-proj", "ai-systems", - "Test Note", ["extracted"], write=True) - self.assertTrue(result1.written) - - result2 = run_extract(self.root, "test-proj", "ai-systems", - "Test Note", ["extracted"], write=True) - self.assertFalse(result2.ok) - self.assertTrue(result2.collision) - self.assertFalse(result2.written) - - def test_title_slugification_in_filename(self) -> None: - make_project(self.root) - make_domain(self.root) - result = run_extract(self.root, "test-proj", "ai-systems", - "My Great Insight", ["extracted"]) - self.assertIn("my-great-insight", result.target_path) - - def test_custom_tags(self) -> None: - make_project(self.root) - make_domain(self.root) - result = run_extract(self.root, "test-proj", "ai-systems", - "Test Note", ["foo", "bar"], write=True) - target = self.root / result.target_path - content = target.read_text(encoding="utf-8") - self.assertIn('["foo", "bar"]', content) - - def test_scaffold_json_escapes_metadata_values(self) -> None: - make_project(self.root) - make_domain(self.root) - title = 'Quoted "title"' - tags = ['tag"\nstatus: stable'] - result = run_extract( - self.root, - "test-proj", - "ai-systems", - title, - tags, - write=True, - ) - - content = (self.root / result.target_path).read_text(encoding="utf-8") - self.assertIn(f"title: {json.dumps(title)}", content) - self.assertIn(f"tags: {json.dumps(tags)}", content) - self.assertEqual(content.count("\nstatus:"), 1) - - def test_check_mode_does_not_write(self) -> None: - make_project(self.root) - make_domain(self.root) - result = run_extract(self.root, "test-proj", "ai-systems", - "Test Note", ["extracted"], write=False) - self.assertTrue(result.ok) - self.assertFalse(result.written) - target = self.root / result.target_path - self.assertFalse(target.exists()) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_fetch_source_text.py b/tests/test_fetch_source_text.py deleted file mode 100644 index f201138..0000000 --- a/tests/test_fetch_source_text.py +++ /dev/null @@ -1,71 +0,0 @@ -"""Tests for scripts/fetch_source_text.py pure helpers.""" - -from __future__ import annotations - -import unittest - -from scripts.fetch_source_text import ( - build_section, - extract_identifiers, - has_fulltext_section, - jats_xml_to_text, - parse_frontmatter, - strip_html, - truncate_excerpt, -) - - -class TestFetchSourceTextHelpers(unittest.TestCase): - def test_parse_frontmatter(self) -> None: - text = """--- -title: "Example" -doi: "10.1000/xyz" -type: raw ---- -# Body -""" - fm, body = parse_frontmatter(text) - self.assertEqual(fm.get("doi"), "10.1000/xyz") - self.assertIn("# Body", body) - - def test_extract_identifiers_from_source(self) -> None: - fm = {"source": "https://doi.org/10.1177/10870547231161533"} - doi, pmid = extract_identifiers(fm) - self.assertEqual(doi, "10.1177/10870547231161533") - self.assertIsNone(pmid) - - def test_has_fulltext_section(self) -> None: - self.assertTrue(has_fulltext_section("## Full text extract (fetch 2026-06-21)\n")) - self.assertFalse(has_fulltext_section("## Bibliographic record\n")) - - def test_jats_xml_to_text_minimal(self) -> None: - xml = b"""<article><body><p>First paragraph.</p><p>Second.</p></body></article>""" - abstract, body = jats_xml_to_text(xml) - self.assertIn("First paragraph", body) - self.assertIn("Second", body) - - def test_strip_html(self) -> None: - html = "<html><body><p>Hello <b>world</b></p></body></html>" - self.assertIn("Hello world", strip_html(html)) - - def test_truncate_excerpt(self) -> None: - long = "a" * 100 - out = truncate_excerpt(long, limit=50) - self.assertIn("truncated", out) - self.assertLess(len(out), 120) - - def test_build_section_contains_verdict(self) -> None: - section = build_section( - today="2026-06-21", - method="europepmc-xml", - access_url="https://example.com", - abstract="Short abstract.", - body_text="Body here.", - verdict="Full text **fetch verified**.", - ) - self.assertIn("## Full text extract", section) - self.assertIn("**Audit verdict:**", section) - - -if __name__ == "__main__": - unittest.main() \ No newline at end of file diff --git a/tests/test_focus_authority.py b/tests/test_focus_authority.py deleted file mode 100644 index 2e94f8b..0000000 --- a/tests/test_focus_authority.py +++ /dev/null @@ -1,127 +0,0 @@ -"""MPE-024 — structured focus authority loader and session-open preference.""" - -from __future__ import annotations - -import importlib.util -import sys -import tempfile -import unittest -from importlib.machinery import SourceFileLoader -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] -sys.path.insert(0, str(ROOT / "scripts")) - -from focus_authority import load_focus, parse_focus_yaml # noqa: E402 - -SCRIPT_PATH = ROOT / "bin" / "session-open" -LOADER = SourceFileLoader("session_open_focus", str(SCRIPT_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -mod = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = mod -SPEC.loader.exec_module(mod) -build_context = mod.build_context - - -SAMPLE = """\ -schema_version: 1 -decision_id: focus-test-01 -revision: abcd1234 -as_of: 2026-07-15T00:00:00Z -review_by: 2099-01-01T00:00:00Z -selected_by: operator -primary: - project: demo-proj - desired_outcome: ship - success_boundary: done - why_now: now - confidence: high - evidence_refs: - - a.md -supporting_slots: [] -snoozed: [] -candidate_snapshot: null -""" - - -class FocusParseTests(unittest.TestCase): - def test_parse_primary_project(self) -> None: - data = parse_focus_yaml(SAMPLE) - self.assertEqual(data["decision_id"], "focus-test-01") - self.assertEqual(data["primary"]["project"], "demo-proj") - self.assertEqual(data["primary"]["evidence_refs"], ["a.md"]) - - def test_parse_list_mapping_preserves_sibling_fields(self) -> None: - data = parse_focus_yaml( - """\ -supporting_slots: - - project: support-one - desired_outcome: unblock primary - evidence_refs: - - evidence-a.md - - evidence-b.md - - project: support-two - desired_outcome: preserve follow-up -""" - ) - - self.assertEqual( - data["supporting_slots"], - [ - { - "project": "support-one", - "desired_outcome": "unblock primary", - "evidence_refs": ["evidence-a.md", "evidence-b.md"], - }, - { - "project": "support-two", - "desired_outcome": "preserve follow-up", - }, - ], - ) - - def test_load_focus_requires_project_readme(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - focus = root / "20_live" / "focus" - focus.mkdir(parents=True) - (focus / "current.yaml").write_text(SAMPLE, encoding="utf-8") - loaded = load_focus(root) - self.assertFalse(loaded.ok) - self.assertTrue(any("missing" in e for e in loaded.errors)) - - proj = root / "30_projects" / "demo-proj" - proj.mkdir(parents=True) - (proj / "README.md").write_text("# Demo\n", encoding="utf-8") - loaded2 = load_focus(root) - self.assertTrue(loaded2.ok, loaded2.errors) - self.assertEqual(loaded2.primary_project, "demo-proj") - - -class SessionOpenFocusTests(unittest.TestCase): - def test_prefers_focus_over_state(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - (root / "AGENTS.md").write_text("# a\n", encoding="utf-8") - (root / "STATE.md").write_text( - "# S\n\n## Active Project\n\nother-project\n", - encoding="utf-8", - ) - focus = root / "20_live" / "focus" - focus.mkdir(parents=True) - (focus / "current.yaml").write_text(SAMPLE, encoding="utf-8") - for slug in ("demo-proj", "other-project"): - d = root / "30_projects" / slug - d.mkdir(parents=True) - (d / "README.md").write_text(f"# {slug}\n", encoding="utf-8") - - result = build_context(root) - self.assertEqual(result.project, "demo-proj") - self.assertEqual(result.project_source, "focus") - self.assertTrue(result.ok) - self.assertEqual(result.focus_revision, "abcd1234") - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_generate_project_tasks.py b/tests/test_generate_project_tasks.py deleted file mode 100644 index 215c6c3..0000000 --- a/tests/test_generate_project_tasks.py +++ /dev/null @@ -1,158 +0,0 @@ -from __future__ import annotations - -import importlib.util -import json -import sys -import tempfile -import unittest -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -SCRIPT = ROOT / "bin" / "generate-project-tasks" -LOADER = SourceFileLoader("generate_project_tasks", str(SCRIPT)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -generate_tasks = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = generate_tasks -SPEC.loader.exec_module(generate_tasks) - - -class GenerateProjectTasksTests(unittest.TestCase): - def test_phase_tasks_and_compiled_packets_are_both_preserved(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = projects / "demo" - project.mkdir() - (project / "README.md").write_text( - "---\n" - 'title: "Demo"\n' - 'status: "active"\n' - 'project_state: "active"\n' - 'next_action: "Keep going"\n' - "---\n", - encoding="utf-8", - ) - - plans_dir = project / "plans" - plans_dir.mkdir() - (plans_dir / "phase-1-test.md").write_text( - "# Phase 1 Test\n\n" - "status: active\n\n" - "## Unit Plan\n\n" - "### Unit 1 — Foundation\n\n" - "- [ ] Task checkbox item\n", - encoding="utf-8", - ) - - packet_manifest = projects / "task_packets_manifest.json" - packet_manifest.write_text( - json.dumps( - { - "packets": [ - { - "id": "packet-demo-fix", - "task_id": "fix", - "project_slug": "demo", - "title": "Packet task", - "status": "ready", - "task_kind": "code", - "agent_profile": "local", - "goal": "Fix it", - "next_action": "Delegate it", - "packet_path": "30_projects/demo/plans/task-packets/fix.md", - } - ] - } - ), - encoding="utf-8", - ) - - tasks = generate_tasks.compile_tasks(projects, packet_manifest) - - self.assertEqual([task["id"] for task in tasks], ["task-demo-phase-1-test-l9", "packet-demo-fix"]) - - extracted_phase_task = tasks[0] - self.assertEqual(extracted_phase_task["status"], "active") - self.assertEqual(extracted_phase_task["title"], "[Unit 1 — Foundation] Task checkbox item") - self.assertEqual(extracted_phase_task["mainframe_target"], "30_projects/demo/plans/phase-1-test.md#L9") - - packet = tasks[1] - self.assertEqual(packet["status"], "active") - self.assertEqual(packet["mainframe_target"], packet["packet_path"]) - self.assertEqual(extracted_phase_task.get("authority_class"), "plan_item") - self.assertFalse(extracted_phase_task.get("executable")) - self.assertEqual(packet.get("authority_class"), "packet_candidate") - self.assertFalse(packet.get("executable")) - - envelope = generate_tasks.build_quarantine_envelope(tasks) - self.assertEqual(envelope["quarantine"]["status"], "quarantined") - self.assertTrue(generate_tasks.is_quarantined_manifest(envelope)) - - def test_plain_bullets_and_decision_log_lines_are_not_tasks(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = projects / "demo" - project.mkdir() - (project / "README.md").write_text( - "---\n" - 'title: "Demo"\n' - 'status: "active"\n' - "---\n", - encoding="utf-8", - ) - - plans_dir = project / "plans" - plans_dir.mkdir() - adr_blob = "**Promoted 2026-05-12:** ADR-011 (four-state enum " + "x" * 2000 + ")" - (plans_dir / "master-plan.md").write_text( - "# Master Plan\n\n" - "status: active\n\n" - "## Execution order\n\n" - f"- {adr_blob}\n" - "- A plain status note that is not a commitment\n" - f"- [ ] {adr_blob}\n" - "- [ ] A real checkbox task\n", - encoding="utf-8", - ) - - tasks = generate_tasks.compile_tasks( - projects, projects / "task_packets_manifest.json" - ) - - titles = [task["title"] for task in tasks] - self.assertEqual(len(tasks), 1) - self.assertIn("A real checkbox task", titles[0]) - for title in titles: - self.assertLessEqual(len(title), 240) - self.assertNotIn("Promoted", title) - - def test_overlong_checkbox_titles_are_truncated_with_ellipsis(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = projects / "demo" - project.mkdir() - (project / "README.md").write_text("---\ntitle: \"Demo\"\n---\n", encoding="utf-8") - - plans_dir = project / "plans" - plans_dir.mkdir() - long_line = "Implement the thing " * 30 - (plans_dir / "phase-1.md").write_text( - "# Phase 1\n\nstatus: active\n\n## Unit plan\n\n" - f"- [ ] {long_line}\n", - encoding="utf-8", - ) - - tasks = generate_tasks.compile_tasks( - projects, projects / "task_packets_manifest.json" - ) - - self.assertEqual(len(tasks), 1) - title = tasks[0]["title"] - self.assertTrue(title.endswith("…"), title[-20:]) - self.assertLessEqual(len(title), generate_tasks.MAX_TASK_TITLE_LENGTH + 40) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_ingest_minion.py b/tests/test_ingest_minion.py deleted file mode 100644 index f56fa7e..0000000 --- a/tests/test_ingest_minion.py +++ /dev/null @@ -1,658 +0,0 @@ -from __future__ import annotations - -import importlib.util -import sys -import tempfile -import unittest -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -MINION_PATH = ROOT / "01_ingest" / "minion.py" -SPEC = importlib.util.spec_from_file_location("ingest_minion", MINION_PATH) -assert SPEC and SPEC.loader -minion_module = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = minion_module -SPEC.loader.exec_module(minion_module) - -IngestMinion = minion_module.IngestMinion -FrontmatterError = minion_module.FrontmatterError -validate_strict = minion_module.validate_strict - - -def valid_note(domain: str = "ai-systems", item_type: str = "note", status: str = "queued") -> str: - return "\n".join( - [ - "---", - 'title: "Example Note"', - f'domain: "{domain}"', - f'type: "{item_type}"', - f'status: "{status}"', - 'source: "manual test source"', - 'tags: ["test", "ingest"]', - "links: []", - "---", - "", - "# Example Note", - "", - ] - ) - - -def pdf_bytes_with_metadata( - *, - title: str = "Metadata Title", - author: str = "Ada Lovelace", - subject: str = "Metadata subject", - keywords: str = "agentic systems, pdf metadata", - creation_date: str = "D:20260528112233-04'00'", -) -> bytes: - return "\n".join( - [ - "%PDF-1.4", - "1 0 obj", - "<<", - f"/Title ({title})", - f"/Author ({author})", - f"/Subject ({subject})", - f"/Keywords ({keywords})", - f"/CreationDate ({creation_date})", - ">>", - "endobj", - "trailer", - "<< /Info 1 0 R >>", - "%%EOF", - ] - ).encode("utf-8") - - -class FrontmatterValidationTests(unittest.TestCase): - def test_strict_validation_rejects_blank_tag_values(self) -> None: - metadata = { - "title": "Example Note", - "domain": "ai-systems", - "type": "note", - "status": "queued", - "source": "manual test source", - "tags": ["test", " "], - } - - with self.assertRaisesRegex(FrontmatterError, "only non-empty strings"): - validate_strict(metadata) - - def test_strip_quotes_accepts_backticks_and_trims_wrappers(self) -> None: - self.assertEqual(minion_module.strip_quotes("` Example title `"), "Example title") - self.assertEqual(minion_module.strip_quotes("' Example source '"), "Example source") - - -class IngestMinionTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - for directory in ( - "00_inbox", - "01_ingest/queue", - "01_ingest/ready", - "01_ingest/rejected", - "10_knowledge/ai-systems/raw", - "10_knowledge/productivity-systems/raw", - "10_knowledge/robotics/raw", - ): - (self.root / directory).mkdir(parents=True, exist_ok=True) - self.minion = IngestMinion(self.root) - - def tearDown(self) -> None: - self.tmp.cleanup() - - def test_reserved_operational_files_are_not_ingested(self) -> None: - """A contract file in 00_inbox governs the folder; it is not a capture. - - Before 2026-08-10 the scanner returned every non-dotfile, so an - `AGENTS.md` written into `00_inbox/` to govern that folder was - normalized and moved to `01_ingest/ready/` on the first apply. The - contract silently migrated out of the folder it governed. The scanner - had no concept of a non-capture file, and treated presence in a - directory as intent to ingest. - """ - for name in ("AGENTS.md", "README.md", "index.md", "agents.md"): - (self.root / "00_inbox" / name).write_text( - valid_note(), encoding="utf-8" - ) - capture = self.root / "00_inbox" / "real-capture.md" - capture.write_text(valid_note(), encoding="utf-8") - - found = {p.name for p in self.minion._files_in(self.root / "00_inbox")} - - self.assertEqual(found, {"real-capture.md"}) - for name in ("AGENTS.md", "README.md", "index.md", "agents.md"): - self.assertTrue( - (self.root / "00_inbox" / name).exists(), - f"{name} must stay where it was written", - ) - - def test_routes_valid_markdown_from_inbox_through_queue(self) -> None: - source = self.root / "00_inbox" / "2026-05-23__ai-systems__note__example.md" - source.write_text(valid_note(), encoding="utf-8") - - result = self.minion.run(apply=True) - - target = self.root / "10_knowledge" / "ai-systems" / source.name - self.assertTrue(result.ok) - self.assertFalse(source.exists()) - self.assertFalse((self.root / "01_ingest" / "queue" / source.name).exists()) - self.assertTrue(target.exists()) - self.assertEqual(target.read_text(encoding="utf-8"), valid_note()) - self.assertIn("stage", [event.kind for event in result.events]) - self.assertIn("route", [event.kind for event in result.events]) - - def test_dry_run_does_not_move_files(self) -> None: - source = self.root / "00_inbox" / "2026-05-23__ai-systems__note__example.md" - source.write_text(valid_note(), encoding="utf-8") - - result = self.minion.run(apply=False) - - target = self.root / "10_knowledge" / "ai-systems" / source.name - self.assertTrue(result.ok) - self.assertTrue(source.exists()) - self.assertFalse(target.exists()) - self.assertIn("stage", [event.kind for event in result.events]) - self.assertIn("route", [event.kind for event in result.events]) - - def test_extracted_status_routes_directly_to_queue(self) -> None: - source = self.root / "00_inbox" / "2026-05-24__ai-systems__note__agent-enriched.md" - source.write_text(valid_note(status="extracted"), encoding="utf-8") - - result = self.minion.run(apply=True) - - target = self.root / "10_knowledge" / "ai-systems" / source.name - self.assertTrue(result.ok) - self.assertTrue(target.exists()) - self.assertNotIn("normalize", [event.kind for event in result.events]) - self.assertIn("stage", [event.kind for event in result.events]) - self.assertIn("route", [event.kind for event in result.events]) - - def test_inbox_empty_tags_requires_enrichment_instead_of_routing(self) -> None: - source = self.root / "00_inbox" / "2026-05-24__ai-systems__note__untagged.md" - source.write_text( - valid_note().replace('tags: ["test", "ingest"]', "tags: []"), - encoding="utf-8", - ) - - result = self.minion.run(apply=True) - - ready_target = self.root / "01_ingest" / "ready" / source.name - durable_target = self.root / "10_knowledge" / "ai-systems" / source.name - self.assertTrue(result.ok) - self.assertTrue(ready_target.exists()) - self.assertFalse(durable_target.exists()) - self.assertIn("tags: []", ready_target.read_text(encoding="utf-8")) - - def test_inbox_no_frontmatter_is_normalized_to_ready(self) -> None: - source = self.root / "00_inbox" / "raw-clipping.md" - source.write_text("# Loose Thought\n\nA half-formed idea.\n", encoding="utf-8") - - result = self.minion.run(apply=True) - - ready_target = self.root / "01_ingest" / "ready" / source.name - rejected = self.root / "01_ingest" / "rejected" / source.name - self.assertTrue(result.ok) - self.assertFalse(source.exists()) - self.assertFalse(rejected.exists()) - self.assertTrue(ready_target.exists()) - - body = ready_target.read_text(encoding="utf-8") - self.assertIn('title: "Loose Thought"', body) - self.assertIn('type: "note"', body) - self.assertIn('status: "skimmed"', body) - self.assertIn('domain: ""', body) - self.assertIn(f'source: "00_inbox/{source.name}"', body) - self.assertIn("tags: []", body) - self.assertIn("links: []", body) - self.assertIn("# Loose Thought", body) - self.assertIn("normalize", [event.kind for event in result.events]) - - def test_inbox_partial_frontmatter_is_normalized_to_ready(self) -> None: - source = self.root / "00_inbox" / "partial.md" - source.write_text( - "\n".join( - [ - "---", - 'title: "Existing Title"', - 'tags: ["draft"]', - "---", - "", - "Body paragraph.", - "", - ] - ), - encoding="utf-8", - ) - - result = self.minion.run(apply=True) - - ready_target = self.root / "01_ingest" / "ready" / source.name - self.assertTrue(result.ok) - self.assertTrue(ready_target.exists()) - - body = ready_target.read_text(encoding="utf-8") - # Preserves what the author wrote - self.assertIn('title: "Existing Title"', body) - self.assertIn('tags: ["draft"]', body) - # Fills missing required keys with defaults - self.assertIn('type: "note"', body) - self.assertIn('status: "skimmed"', body) - self.assertIn('domain: ""', body) - self.assertIn(f'source: "00_inbox/{source.name}"', body) - self.assertIn("links: []", body) - self.assertIn("normalize", [event.kind for event in result.events]) - - def test_inbox_block_list_frontmatter_is_normalized_to_ready(self) -> None: - source = self.root / "00_inbox" / "clipping.md" - source.write_text( - "\n".join( - [ - "---", - 'title: "Existing Clipping"', - "author:", - ' - "[[@source-author]]"', - "tags:", - ' - "clippings"', - "---", - "", - "Body paragraph.", - "", - ] - ), - encoding="utf-8", - ) - - result = self.minion.run(apply=True) - - ready_target = self.root / "01_ingest" / "ready" / source.name - self.assertTrue(result.ok) - self.assertTrue(ready_target.exists()) - - body = ready_target.read_text(encoding="utf-8") - self.assertIn('title: "Existing Clipping"', body) - self.assertIn('tags: ["clippings"]', body) - self.assertIn('author: ["[[@source-author]]"]', body) - self.assertIn('status: "skimmed"', body) - self.assertIn("Body paragraph.", body) - - def test_inbox_malformed_frontmatter_stages_to_ready_with_warning(self) -> None: - source = self.root / "00_inbox" / "bad-frontmatter.md" - source.write_text( - "\n".join( - [ - "---", - 'title: "Broken Capture"', - "not valid yaml structure", - "---", - "", - "Captured body.", - "", - ] - ), - encoding="utf-8", - ) - - result = self.minion.run(apply=True) - - ready_target = self.root / "01_ingest" / "ready" / source.name - rejected = self.root / "01_ingest" / "rejected" / source.name - self.assertTrue(result.ok) - self.assertFalse(rejected.exists()) - self.assertTrue(ready_target.exists()) - self.assertEqual(result.events[0].kind, "normalize") - self.assertEqual(result.events[0].severity, "warning") - self.assertIn("frontmatter needs agent repair", result.events[0].message) - - body = ready_target.read_text(encoding="utf-8") - self.assertIn('title: "Bad Frontmatter"', body) - self.assertIn("---\ntitle: \"Broken Capture\"\nnot valid yaml structure\n---", body) - self.assertIn("Captured body.", body) - - def test_inbox_unknown_domain_stages_to_ready_with_warning(self) -> None: - source = self.root / "00_inbox" / "2026-05-23__new-domain__note__example.md" - source.write_text(valid_note(domain="new-domain"), encoding="utf-8") - - result = self.minion.run(apply=True) - - ready_target = self.root / "01_ingest" / "ready" / source.name - rejected = self.root / "01_ingest" / "rejected" / source.name - target = self.root / "10_knowledge" / "new-domain" / source.name - self.assertTrue(result.ok) - self.assertFalse(rejected.exists()) - self.assertFalse(target.exists()) - self.assertTrue(ready_target.exists()) - self.assertEqual(result.events[0].kind, "normalize") - self.assertEqual(result.events[0].severity, "warning") - self.assertIn("not established", result.events[0].message) - - body = ready_target.read_text(encoding="utf-8") - self.assertIn('domain: "new-domain"', body) - self.assertIn('status: "skimmed"', body) - - def test_wikilinks_inside_code_spans_are_ignored(self) -> None: - source = self.root / "00_inbox" / "discusses-wikilinks.md" - source.write_text( - "\n".join( - [ - "# Note that talks about wikilinks", - "", - "Body refers to [[real-target]] as a real connection.", - "", - "But `[[inline-syntax-example]]` is just talking about the syntax.", - "", - "And the fenced block below is illustrative:", - "", - "```", - "tags: [\"author:[[@somebody]]\"]", - "see also [[fenced-example]]", - "```", - "", - "End paragraph also has [[real-target]] again (dedup).", - "", - ] - ), - encoding="utf-8", - ) - - result = self.minion.run(apply=True) - - ready_target = self.root / "01_ingest" / "ready" / source.name - self.assertTrue(result.ok) - self.assertTrue(ready_target.exists()) - - body = ready_target.read_text(encoding="utf-8") - self.assertIn('links: ["real-target"]', body) - # Body itself is preserved verbatim — the code spans/blocks still appear. - self.assertIn("`[[inline-syntax-example]]`", body) - self.assertIn("[[fenced-example]]", body) - - def test_inbox_wikilinks_extracted_to_links_field(self) -> None: - source = self.root / "00_inbox" / "with-links.md" - source.write_text( - "\n".join( - [ - "# Note With Links", - "", - "Refers to [[mindgraph-design]] and [[robot-cell|the new cell]].", - "", - "Also [[mindgraph-design]] again to test dedup.", - "", - ] - ), - encoding="utf-8", - ) - - result = self.minion.run(apply=True) - - ready_target = self.root / "01_ingest" / "ready" / source.name - self.assertTrue(result.ok) - self.assertTrue(ready_target.exists()) - - body = ready_target.read_text(encoding="utf-8") - self.assertIn('links: ["mindgraph-design", "robot-cell"]', body) - # The body itself is preserved verbatim (still contains the raw wikilinks). - self.assertIn("[[mindgraph-design]]", body) - self.assertIn("[[robot-cell|the new cell]]", body) - - def test_queue_missing_frontmatter_is_rejected(self) -> None: - source = self.root / "01_ingest" / "queue" / "missing-frontmatter.md" - source.write_text("# Missing frontmatter\n", encoding="utf-8") - - result = self.minion.run(apply=True) - - rejected = self.root / "01_ingest" / "rejected" / source.name - self.assertFalse(result.ok) - self.assertFalse(source.exists()) - self.assertTrue(rejected.exists()) - self.assertEqual(result.events[0].kind, "reject") - - def test_unknown_domain_is_rejected(self) -> None: - source = self.root / "01_ingest" / "queue" / "2026-05-23__unknown__note__example.md" - source.write_text(valid_note(domain="unknown"), encoding="utf-8") - - result = self.minion.run(apply=True) - - rejected = self.root / "01_ingest" / "rejected" / source.name - self.assertFalse(result.ok) - self.assertFalse(source.exists()) - self.assertTrue(rejected.exists()) - self.assertIn("unknown knowledge domain", result.events[0].message) - - def test_convention_named_pdf_creates_raw_file_and_stub(self) -> None: - source = self.root / "00_inbox" / "2026-05-21__ai-systems__raw__sample-paper.pdf" - source.write_bytes(b"%PDF-1.4\n") - - result = self.minion.run(apply=True) - - raw = self.root / "10_knowledge" / "ai-systems" / "raw" / source.name - stub = ( - self.root - / "10_knowledge" - / "ai-systems" - / "2026-05-21__ai-systems__raw__sample-paper.md" - ) - self.assertTrue(result.ok) - self.assertFalse(source.exists()) - self.assertTrue(raw.exists()) - self.assertTrue(stub.exists()) - stub_text = stub.read_text(encoding="utf-8") - self.assertIn('title: "Sample Paper"', stub_text) - self.assertIn('domain: "ai-systems"', stub_text) - self.assertIn('source: "./raw/2026-05-21__ai-systems__raw__sample-paper.pdf"', stub_text) - self.assertIn('tags: ["pdf", "evidence"]', stub_text) - self.assertIn('source_type: "pdf"', stub_text) - - def test_pdf_metadata_enriches_raw_stub(self) -> None: - source = self.root / "00_inbox" / "2026-05-21__ai-systems__raw__sample-paper.pdf" - source.write_bytes(pdf_bytes_with_metadata()) - - result = self.minion.run(apply=True) - - stub = ( - self.root - / "10_knowledge" - / "ai-systems" - / "2026-05-21__ai-systems__raw__sample-paper.md" - ) - self.assertTrue(result.ok) - stub_text = stub.read_text(encoding="utf-8") - self.assertIn('title: "Metadata Title"', stub_text) - self.assertIn('author: ["Ada Lovelace"]', stub_text) - self.assertIn('description: "Metadata subject"', stub_text) - self.assertIn('keywords: ["agentic systems", "pdf metadata"]', stub_text) - self.assertIn('created: "2026-05-28"', stub_text) - - def test_malformed_pdf_octal_metadata_does_not_abort_routing(self) -> None: - source = self.root / "00_inbox" / "2026-05-21__ai-systems__raw__sample-paper.pdf" - source.write_bytes(pdf_bytes_with_metadata(title=r"Broken\777Title")) - - result = self.minion.run(apply=True) - - stub = ( - self.root - / "10_knowledge" - / "ai-systems" - / "2026-05-21__ai-systems__raw__sample-paper.md" - ) - self.assertTrue(result.ok) - self.assertTrue(stub.exists()) - stub_text = stub.read_text(encoding="utf-8") - self.assertIn('title: "Sample Paper"', stub_text) - self.assertIn('author: ["Ada Lovelace"]', stub_text) - - def test_inbox_unconvention_named_pdf_suggests_without_rejecting(self) -> None: - source = self.root / "00_inbox" / "Prompting in the Wild.pdf" - source.write_bytes(pdf_bytes_with_metadata(title="Prompting Study")) - - result = self.minion.run(apply=True) - - rejected = self.root / "01_ingest" / "rejected" / source.name - queued = self.root / "01_ingest" / "queue" / source.name - self.assertTrue(result.ok) - self.assertTrue(source.exists()) - self.assertFalse(rejected.exists()) - self.assertFalse(queued.exists()) - self.assertEqual(result.events[0].kind, "suggest") - self.assertEqual(result.events[0].severity, "warning") - self.assertIn("__<domain>__raw__prompting-in-the-wild.pdf", result.events[0].message) - self.assertIn("metadata found: title: Prompting Study", result.events[0].message) - - def test_inbox_pdf_unknown_domain_suggests_without_rejecting(self) -> None: - source = self.root / "00_inbox" / "2026-05-21__new-domain__raw__sample-paper.pdf" - source.write_bytes(b"%PDF-1.4\n") - - result = self.minion.run(apply=True) - - rejected = self.root / "01_ingest" / "rejected" / source.name - queued = self.root / "01_ingest" / "queue" / source.name - self.assertTrue(result.ok) - self.assertTrue(source.exists()) - self.assertFalse(rejected.exists()) - self.assertFalse(queued.exists()) - self.assertEqual(result.events[0].kind, "suggest") - self.assertEqual(result.events[0].severity, "warning") - self.assertIn("no matching 10_knowledge domain exists", result.events[0].message) - - def test_destination_collision_blocks_without_overwrite(self) -> None: - source = self.root / "01_ingest" / "queue" / "2026-05-23__ai-systems__note__example.md" - target = self.root / "10_knowledge" / "ai-systems" / source.name - source.write_text(valid_note(), encoding="utf-8") - target.write_text("existing content\n", encoding="utf-8") - - result = self.minion.run(apply=True) - - self.assertFalse(result.ok) - self.assertTrue(source.exists()) - self.assertEqual(target.read_text(encoding="utf-8"), "existing content\n") - self.assertEqual(result.events[0].kind, "blocked") - -extract_wikilinks = minion_module.extract_wikilinks -_is_plausible_wikilink = minion_module._is_plausible_wikilink - - -class WikilinkExtractionTests(unittest.TestCase): - """Regression tests for extract_wikilinks — especially unfenced code.""" - - def test_unfenced_python_double_bracket_indexing_is_not_a_wikilink(self) -> None: - """Repro case: df[['Open', 'High', 'Low', 'Close', 'Volume']]""" - body = ( - "import pandas as pd\n" - "df = pd.read_csv('data.csv')\n" - "ohlcv = df[['Open', 'High', 'Low', 'Close', 'Volume']]\n" - "filtered = df[['Close']]\n" - ) - self.assertEqual(extract_wikilinks(body), []) - - def test_multiline_bracket_expression_is_not_a_wikilink(self) -> None: - body = "result = df[['Open',\n 'High',\n 'Low']]\n" - self.assertEqual(extract_wikilinks(body), []) - - def test_real_wikilinks_still_extracted_alongside_unfenced_code(self) -> None: - body = ( - "See [[strategy-hunter]] for details.\n" - "\n" - "ohlcv = df[['Open', 'High', 'Low', 'Close', 'Volume']]\n" - "\n" - "Also [[backtest-results|results]].\n" - ) - self.assertEqual(extract_wikilinks(body), ["strategy-hunter", "backtest-results"]) - - def test_fenced_code_block_still_stripped(self) -> None: - body = ( - "Some text\n" - "\n" - "```python\n" - "ohlcv = df[['Open', 'High', 'Low', 'Close', 'Volume']]\n" - "link = [[fenced-link]]\n" - "```\n" - "\n" - "End with [[real-target]].\n" - ) - self.assertEqual(extract_wikilinks(body), ["real-target"]) - - def test_inline_code_span_still_stripped(self) -> None: - body = "Use `[[syntax-example]]` to link. Real: [[actual-link]].\n" - self.assertEqual(extract_wikilinks(body), ["actual-link"]) - - def test_target_with_quotes_is_implausible(self) -> None: - self.assertFalse(_is_plausible_wikilink("'Open', 'High'")) - - def test_target_with_comma_is_implausible(self) -> None: - self.assertFalse(_is_plausible_wikilink("a, b")) - - def test_target_with_newline_is_implausible(self) -> None: - self.assertFalse(_is_plausible_wikilink("open\nhigh")) - - def test_normal_wikilink_target_is_plausible(self) -> None: - self.assertTrue(_is_plausible_wikilink("strategy-hunter")) - self.assertTrue(_is_plausible_wikilink("@author-name")) - self.assertTrue(_is_plausible_wikilink("My Note Title")) - - def test_end_to_end_strategy_hunter_repro(self) -> None: - """Full repro of the original bug: a raw Python script saved as markdown.""" - body = "\n".join([ - "# Strategy Hunter Backtest Script", - "", - "import pandas as pd", - "import numpy as np", - "", - "df = pd.read_csv('data.csv')", - "ohlcv = df[['Open', 'High', 'Low', 'Close', 'Volume']]", - "signals = df[['Signal', 'Position']]", - "", - "def backtest(data):", - " returns = data[['Close']].pct_change()", - " return returns", - "", - ]) - self.assertEqual(extract_wikilinks(body), []) - - -class WikilinkExtractionIntegrationTests(unittest.TestCase): - """Integration: unfenced code in inbox markdown doesn't pollute links.""" - - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - for directory in ( - "00_inbox", - "01_ingest/queue", - "01_ingest/ready", - "01_ingest/rejected", - "10_knowledge/finance/raw", - ): - (self.root / directory).mkdir(parents=True, exist_ok=True) - self.minion = IngestMinion(self.root) - - def tearDown(self) -> None: - self.tmp.cleanup() - - def test_inbox_python_script_as_markdown_produces_empty_links(self) -> None: - source = self.root / "00_inbox" / "Back test - strategy hunter.md" - source.write_text( - "\n".join([ - "import pandas as pd", - "df = pd.read_csv('data.csv')", - "ohlcv = df[['Open', 'High', 'Low', 'Close', 'Volume']]", - "filtered = df[['Close']]", - "", - ]), - encoding="utf-8", - ) - - result = self.minion.run(apply=True) - - ready_target = self.root / "01_ingest" / "ready" / source.name - self.assertTrue(result.ok) - self.assertTrue(ready_target.exists()) - - body = ready_target.read_text(encoding="utf-8") - self.assertIn("links: []", body) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_ingest_status.py b/tests/test_ingest_status.py deleted file mode 100644 index 0150a33..0000000 --- a/tests/test_ingest_status.py +++ /dev/null @@ -1,208 +0,0 @@ -from __future__ import annotations - -import importlib.util -import os -import sys -import tempfile -import unittest -from datetime import date, datetime, timedelta -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -SCRIPT_PATH = ROOT / "bin" / "ingest-status" -LOADER = SourceFileLoader("ingest_status", str(SCRIPT_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -ingest_status = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = ingest_status -SPEC.loader.exec_module(ingest_status) - - -TODAY = date(2026, 6, 9) - -MANIFEST_HEADER = ( - "batch_id,source,relative_path,size_bytes,sha256,created_at,modified_at," - "extension,embedded_dates,category,proposed_destination,review_status," - "duplicate_group\n" -) -LEDGER_HEADER = ( - "recorded_at,batch_id,relative_path,sha256,disposition,destination," - "evidence,notes\n" -) - - -def set_age(path: Path, days: int) -> None: - stamp = datetime.combine( - TODAY - timedelta(days=days), datetime.min.time() - ).timestamp() - os.utime(path, (stamp, stamp)) - - -class IngestStatusTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - for lane in ("00_inbox", "01_ingest/ready", "01_ingest/queue", - "01_ingest/rejected"): - (self.root / lane).mkdir(parents=True) - self.batches = ( - self.root / "30_projects/second-brain-migration/raw-materials/batches" - ) - self.batches.mkdir(parents=True) - - def tearDown(self) -> None: - self.tmp.cleanup() - - def write_inbox(self, name: str, body: str, age_days: int) -> Path: - path = self.root / "00_inbox" / name - path.write_text(body, encoding="utf-8") - set_age(path, age_days) - return path - - def register_batch(self, batch_id: str, files: dict[str, Path], - ledger_rows: list[tuple[str, str]] | None = None) -> None: - batch_dir = self.batches / batch_id - batch_dir.mkdir() - manifest_lines = [MANIFEST_HEADER] - for name, path in files.items(): - digest = ingest_status.sha256_of(path) - manifest_lines.append( - f"{batch_id},test drop,{name},1,{digest},,,.md,,cat,dest,unreviewed,\n" - ) - (batch_dir / "manifest.csv").write_text( - "".join(manifest_lines), encoding="utf-8" - ) - if ledger_rows is not None: - ledger_lines = [LEDGER_HEADER] - for name, disposition in ledger_rows: - digest = ( - ingest_status.sha256_of(files[name]) if name in files else "" - ) - ledger_lines.append( - f"2026-06-06T12:00:00,{batch_id},{name},{digest}," - f"{disposition},,evidence,\n" - ) - (batch_dir / "disposition-ledger.csv").write_text( - "".join(ledger_lines), encoding="utf-8" - ) - - def test_lane_age_buckets_and_hidden_files(self) -> None: - self.write_inbox("fresh.md", "fresh", age_days=2) - self.write_inbox("aging.md", "aging", age_days=14) - self.write_inbox("stale.md", "stale", age_days=90) - hidden = self.write_inbox(".DS_Store", "junk", age_days=400) - ready = self.root / "01_ingest/ready/note.md" - ready.write_text("ready", encoding="utf-8") - set_age(ready, 5) - (self.root / "00_inbox" / "subdir").mkdir() - - status = ingest_status.build_status(self.root, today=TODAY) - lanes = {lane["name"]: lane for lane in status["lanes"]} - - inbox = lanes["00_inbox"] - self.assertEqual(inbox["files"], 3) # hidden file skipped - self.assertEqual(inbox["directories"], 1) - self.assertEqual(inbox["age_days_0_7"], 1) - self.assertEqual(inbox["age_days_8_30"], 1) - self.assertEqual(inbox["age_days_31_plus"], 1) - self.assertEqual(inbox["oldest_file_age_days"], 90) - self.assertEqual(lanes["01_ingest/ready"]["files"], 1) - self.assertEqual(lanes["01_ingest/queue"]["files"], 0) - self.assertTrue(hidden.exists()) - - def test_inbox_composition_splits_registered_from_organic(self) -> None: - registered = self.write_inbox("migrated.md", "old note", age_days=60) - self.write_inbox("organic.md", "fresh capture", age_days=1) - self.register_batch( - "2026-06-05-001", - {"migrated.md": registered}, - ledger_rows=[("migrated.md", "unresolved")], - ) - - status = ingest_status.build_status(self.root, today=TODAY) - composition = status["inbox_composition"] - - self.assertEqual(composition["batch_registered"], 1) - self.assertEqual(composition["organic"], 1) - self.assertEqual(composition["batches_matched"], ["2026-06-05-001"]) - self.assertEqual( - composition["registered_by_disposition"], - [{"name": "unresolved", "count": 1}], - ) - - def test_last_ledger_row_wins(self) -> None: - kept = self.write_inbox("kept.md", "kept body", age_days=10) - self.register_batch( - "2026-06-05-001", - {"kept.md": kept}, - ledger_rows=[ - ("kept.md", "unresolved"), - ("kept.md", "retained"), - ], - ) - - status = ingest_status.build_status(self.root, today=TODAY) - composition = status["inbox_composition"] - batch = status["batches"][0] - - self.assertEqual( - composition["registered_by_disposition"], - [{"name": "retained", "count": 1}], - ) - self.assertEqual( - batch["dispositions"], [{"name": "retained", "count": 1}] - ) - self.assertEqual(batch["ledger_rows"], 2) - - def test_renamed_registered_file_matches_by_hash(self) -> None: - original = self.write_inbox("original-name.md", "same bytes", age_days=30) - self.register_batch( - "2026-06-05-001", - {"original-name.md": original}, - ledger_rows=[("original-name.md", "parked")], - ) - original.rename(self.root / "00_inbox" / "renamed.md") - set_age(self.root / "00_inbox" / "renamed.md", 30) - - status = ingest_status.build_status(self.root, today=TODAY) - composition = status["inbox_composition"] - - self.assertEqual(composition["batch_registered"], 1) - self.assertEqual(composition["organic"], 0) - self.assertEqual( - composition["registered_by_disposition"], - [{"name": "parked", "count": 1}], - ) - - def test_batch_without_ledger_and_invalid_dir(self) -> None: - registered = self.write_inbox("noledger.md", "body", age_days=5) - self.register_batch("2026-06-05-001", {"noledger.md": registered}) - (self.batches / "stray-dir").mkdir() - - status = ingest_status.build_status(self.root, today=TODAY) - composition = status["inbox_composition"] - batch = status["batches"][0] - - self.assertFalse(batch["has_ledger"]) - self.assertEqual(batch["dispositions"], []) - self.assertEqual( - composition["registered_by_disposition"], - [{"name": "unresolved", "count": 1}], - ) - self.assertEqual(status["invalid_batch_dirs"], ["stray-dir"]) - - def test_missing_batches_dir_means_all_organic(self) -> None: - self.write_inbox("capture.md", "body", age_days=1) - - status = ingest_status.build_status(self.root, today=TODAY) - composition = status["inbox_composition"] - - self.assertEqual(composition["batch_registered"], 0) - self.assertEqual(composition["organic"], 1) - self.assertEqual(status["batches"], []) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_lane_intake.py b/tests/test_lane_intake.py deleted file mode 100644 index 5a167d8..0000000 --- a/tests/test_lane_intake.py +++ /dev/null @@ -1,191 +0,0 @@ -"""Tests for bin/lane-intake list filters and priority/status matching.""" - -from __future__ import annotations - -import importlib.util -import sys -import tempfile -import unittest -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -SCRIPT_PATH = ROOT / "bin" / "lane-intake" -LOADER = SourceFileLoader("lane_intake", str(SCRIPT_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -mod = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = mod -SPEC.loader.exec_module(mod) - - -class PriorityMatchTests(unittest.TestCase): - def test_p0_matches_immediate(self) -> None: - self.assertTrue(mod._priority_matches("immediate", "P0")) - self.assertTrue(mod._priority_matches("immediate", "p0")) - - def test_p0_matches_p0(self) -> None: - self.assertTrue(mod._priority_matches("P0", "P0")) - self.assertTrue(mod._priority_matches("P0 critical", "P0")) - - def test_p0_rejects_next(self) -> None: - self.assertFalse(mod._priority_matches("next", "P0")) - self.assertFalse(mod._priority_matches("P1", "P0")) - - def test_p1_matches_next_and_high(self) -> None: - self.assertTrue(mod._priority_matches("next", "P1")) - self.assertTrue(mod._priority_matches("high", "P1")) - self.assertTrue(mod._priority_matches("P1", "P1")) - - def test_p4_matches_monitor_alongside(self) -> None: - self.assertTrue(mod._priority_matches("monitor", "P4")) - self.assertTrue(mod._priority_matches("alongside", "P4")) - self.assertTrue(mod._priority_matches("deferred", "P4")) - - def test_empty_filter_matches_all(self) -> None: - self.assertTrue(mod._priority_matches("immediate", "")) - self.assertTrue(mod._priority_matches("", "")) - - -class StatusMatchTests(unittest.TestCase): - def test_exact_active(self) -> None: - self.assertTrue(mod._status_matches("active", "active")) - - def test_prefix_em_dash(self) -> None: - self.assertTrue( - mod._status_matches("active — specialization pass done", "active") - ) - - def test_prefix_hyphen(self) -> None: - self.assertTrue(mod._status_matches("active - foundation", "active")) - - def test_complete_not_active(self) -> None: - self.assertFalse(mod._status_matches("complete", "active")) - self.assertFalse(mod._status_matches("parked", "active")) - - def test_case_insensitive(self) -> None: - self.assertTrue(mod._status_matches("Active", "ACTIVE")) - - -class ListLanesFixtureTests(unittest.TestCase): - """list_lanes against a temporary lanes tree.""" - - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - self.lanes = self.root / "lanes" - self.completed = self.lanes / "completed" - self.lanes.mkdir(parents=True) - self.completed.mkdir(parents=True) - - self._write_lane( - self.lanes / "alpha-lane", - lane_id="A1", - status="active", - priority="immediate", - domain="agents", - ) - self._write_lane( - self.lanes / "beta-lane", - lane_id="B1", - status="active — specialization pass done", - priority="P0", - domain="agents", - ) - self._write_lane( - self.lanes / "gamma-lane", - lane_id="G1", - status="active", - priority="next", - domain="finance", - ) - self._write_lane( - self.completed / "done-lane", - lane_id="D1", - status="complete", - priority="P0", - domain="agents", - ) - - self._orig_lanes = mod.LANES_DIR - self._orig_completed = mod.COMPLETED_DIR - mod.LANES_DIR = self.lanes - mod.COMPLETED_DIR = self.completed - - def tearDown(self) -> None: - mod.LANES_DIR = self._orig_lanes - mod.COMPLETED_DIR = self._orig_completed - self.tmp.cleanup() - - def _write_lane( - self, - path: Path, - *, - lane_id: str, - status: str, - priority: str, - domain: str, - ) -> None: - path.mkdir(parents=True, exist_ok=True) - (path / "README.md").write_text( - f"""--- -title: "Lane: {path.name}" -status: "{status}" -lane_id: "{lane_id}" -knowledge_domain: "{domain}" -priority: "{priority}" ---- - -# {path.name} -""", - encoding="utf-8", - ) - - def test_p0_active_includes_immediate_and_p0(self) -> None: - rows = mod.list_lanes(status_filter="active", priority_filter="P0") - slugs = {r["slug"] for r in rows} - self.assertIn("alpha-lane", slugs) - self.assertIn("beta-lane", slugs) - self.assertNotIn("gamma-lane", slugs) - self.assertNotIn("done-lane", slugs) - - def test_status_prefix_matches_decorated_active(self) -> None: - rows = mod.list_lanes(status_filter="active") - slugs = {r["slug"] for r in rows} - self.assertIn("beta-lane", slugs) - self.assertEqual(len(rows), 3) - - def test_p1_matches_next(self) -> None: - rows = mod.list_lanes(status_filter="active", priority_filter="P1") - slugs = {r["slug"] for r in rows} - self.assertEqual(slugs, {"gamma-lane"}) - - def test_complete_p0(self) -> None: - rows = mod.list_lanes(status_filter="complete", priority_filter="P0") - slugs = {r["slug"] for r in rows} - self.assertEqual(slugs, {"done-lane"}) - - def test_format_json(self) -> None: - rows = mod.list_lanes(status_filter="active", priority_filter="P0") - out = mod.format_list(rows, fmt="json") - self.assertIn("alpha-lane", out) - self.assertTrue(out.strip().startswith("[")) - - -class ParseFrontmatterTests(unittest.TestCase): - def test_basic(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - p = Path(tmp) / "README.md" - p.write_text( - '---\nstatus: "active"\npriority: "immediate"\nlane_id: "X1"\n---\n\n# Hi\n', - encoding="utf-8", - ) - fm = mod.parse_frontmatter(p) - self.assertEqual(fm.get("status"), "active") - self.assertEqual(fm.get("priority"), "immediate") - self.assertEqual(fm.get("lane_id"), "X1") - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_lifecycle_identity.py b/tests/test_lifecycle_identity.py new file mode 100644 index 0000000..8b83b4d --- /dev/null +++ b/tests/test_lifecycle_identity.py @@ -0,0 +1,328 @@ +from __future__ import annotations + +import json +import sqlite3 +import sys +import tempfile +import unittest +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] +sys.path.insert(0, str(ROOT / "scripts")) + +from lifecycle_identity import ( # noqa: E402 + DuplicateIdentity, + FrontmatterError, + InvalidIdentity, + MissingIdentity, + resolve_record, + scan_lifecycle_records, +) +from work_inventory import ( # noqa: E402 + document_map_hash, + manifest_members, + selected_source_files, + validate_manifest, + verify_typed_document_map, +) + + +def write_record(root: Path, slug: str, *, record_type: str = "project", state: str = "active") -> Path: + path = root / slug + path.mkdir(parents=True, exist_ok=True) + path.joinpath("README.md").write_text( + "---\n" + f"record_type: {record_type}\n" + f"project_state: {state}\n" + f"status: {state}\n" + "---\n" + f"# {slug}\n", + encoding="utf-8", + ) + path.joinpath("plans").mkdir() + path.joinpath("plans", "plan.md").write_text("# plan\n", encoding="utf-8") + path.joinpath("log.md").write_text("# log\n", encoding="utf-8") + return path + + +def typed_manifest(root: Path) -> Path: + scan = scan_lifecycle_records(root) + assert not scan.issues, scan.issues + members = [record.manifest_entry(root) for record in sorted(scan.records, key=lambda item: item.slug)] + path = root / "30_projects" / "manifest.json" + path.write_text( + json.dumps( + { + "schema_version": 2, + "profile": "test", + "projects": [ + member["id"] + for member in members + if member["path"].startswith("30_projects/") + ], + "members": members, + "include": ["README.md", "plans/*.md", "log.md"], + "exclude": ["outputs/*"], + }, + indent=2, + ), + encoding="utf-8", + ) + return path + + +class LifecycleIdentityTests(unittest.TestCase): + def test_resolves_project_and_operation_across_shared_namespace(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + write_record(root / "30_projects", "project-a") + write_record(root / "40_operations", "operation-a", record_type="operation") + + project = resolve_record(root, "project-a", expected_record_type="project") + operation = resolve_record(root, "operation-a", expected_record_type="operation") + + self.assertEqual(project.root_name, "30_projects") + self.assertEqual(operation.root_name, "40_operations") + self.assertEqual(operation.relative_path(root), "40_operations/operation-a") + + def test_duplicate_identity_fails_closed(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + write_record(root / "30_projects", "same") + write_record(root / "40_operations", "same", record_type="operation") + with self.assertRaises(DuplicateIdentity): + resolve_record(root, "same") + + def test_missing_and_invalid_operation_identity_fail_closed(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + with self.assertRaises(MissingIdentity): + resolve_record(root, "ghost") + write_record(root / "40_operations", "bad") + with self.assertRaises(InvalidIdentity): + resolve_record(root, "bad") + + def test_symlinked_root_and_child_are_not_authority(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + real = root / "real-projects" + write_record(real, "project-a") + (root / "30_projects").symlink_to(real, target_is_directory=True) + self.assertFalse(scan_lifecycle_records(root).records) + self.assertTrue(scan_lifecycle_records(root).issues) + + root.unlink() if root.is_symlink() else None + + def test_readme_state_conflict_is_reported(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + path = write_record(root / "30_projects", "conflict") + readme = path / "README.md" + readme.write_text( + "---\nrecord_type: project\nproject_state: active\nlifecycle_state: paused\n---\n", + encoding="utf-8", + ) + with self.assertRaises(InvalidIdentity): + resolve_record(root, "conflict") + + def test_explicit_invalid_record_type_does_not_fall_back_to_project_root(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + path = write_record(root / "30_projects", "bad-type") + (path / "README.md").write_text( + "---\nrecord_type: definitely-not-a-type\nproject_state: active\n---\n", + encoding="utf-8", + ) + with self.assertRaises(InvalidIdentity): + resolve_record(root, "bad-type") + + def test_malformed_duplicate_and_nested_frontmatter_fail_closed(self) -> None: + with self.assertRaises(FrontmatterError): + from lifecycle_identity import parse_frontmatter + parse_frontmatter("---\nrecord_type: project\nrecord_type: operation\n---\n") + with self.assertRaises(FrontmatterError): + from lifecycle_identity import parse_frontmatter + parse_frontmatter("---\nidentity:\n record_type: project\n---\n") + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + path = write_record(root / "30_projects", "malformed") + (path / "README.md").write_text("---\nrecord_type: project\n project_state: active\n---\n", encoding="utf-8") + with self.assertRaises(InvalidIdentity): + resolve_record(root, "malformed") + + def test_uninspectable_sibling_root_blocks_cross_root_resolution(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + write_record(root / "30_projects", "project-a") + real_operations = root / "real-operations" + write_record(real_operations, "operation-a", record_type="operation") + (root / "40_operations").symlink_to(real_operations, target_is_directory=True) + with self.assertRaises(InvalidIdentity): + resolve_record(root, "project-a") + + def test_project_md_owns_lifecycle_state_without_copying_private_fields_to_manifest(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + path = write_record(root / "30_projects", "public-mirror") + (path / "README.md").write_text("---\nrecord_type: project\n---\n# public mirror\n", encoding="utf-8") + (path / "PROJECT.md").write_text( + "---\nproject_state: paused\nwip_class: eval\nnext_action: private reentry\n---\n", + encoding="utf-8", + ) + record = resolve_record(root, "public-mirror") + self.assertEqual(record.lifecycle_state, "paused") + self.assertEqual(record.state_source, "PROJECT.md") + self.assertEqual(record.wip_class, "eval") + entry = record.manifest_entry(root) + self.assertNotIn("next_action", entry) + + +class TypedInventoryTests(unittest.TestCase): + def test_manifest_exactly_covers_both_lifecycle_roots(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + write_record(root / "30_projects", "project-a") + write_record(root / "40_operations", "operation-a", record_type="operation") + manifest = typed_manifest(root) + report = validate_manifest(root, manifest) + self.assertTrue(report["ok"], report) + self.assertEqual(report["manifest_count"], 2) + + payload = json.loads(manifest.read_text(encoding="utf-8")) + self.assertEqual(payload["projects"], ["project-a"]) + self.assertEqual( + [member["id"] for member in payload["members"]], + ["operation-a", "project-a"], + ) + + def test_explicit_exclusion_closes_coverage_without_ingesting_record(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + write_record(root / "30_projects", "project-a") + write_record(root / "30_projects", "excluded-project") + manifest = typed_manifest(root) + payload = json.loads(manifest.read_text(encoding="utf-8")) + payload["members"] = [ + member for member in payload["members"] if member["id"] == "project-a" + ] + payload["projects"] = ["project-a"] + payload["exclusions"] = [ + { + "id": "excluded-project", + "record_type": "project", + "path": "30_projects/excluded-project", + "reason": "independent repository is outside this profile", + } + ] + manifest.write_text(json.dumps(payload), encoding="utf-8") + + report = validate_manifest(root, manifest) + self.assertTrue(report["ok"], report) + self.assertEqual(report["omitted_records"], []) + self.assertEqual(report["excluded_records"], ["excluded-project"]) + self.assertEqual(report["excluded_count"], 1) + members, issues = manifest_members(payload) + self.assertEqual(issues, []) + rows = selected_source_files(root, payload, members) + self.assertEqual({row["namespace"] for row in rows}, {"project-a"}) + + def test_manifest_omission_and_type_mismatch_are_distinct(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + write_record(root / "30_projects", "project-a") + write_record(root / "40_operations", "operation-a", record_type="operation") + manifest = typed_manifest(root) + payload = json.loads(manifest.read_text(encoding="utf-8")) + payload["members"] = [ + member for member in payload["members"] if member["id"] == "project-a" + ] + payload["projects"] = ["project-a"] + manifest.write_text(json.dumps(payload), encoding="utf-8") + report = validate_manifest(root, manifest) + self.assertFalse(report["ok"]) + self.assertEqual(report["omitted_records"], ["operation-a"]) + + payload = json.loads(manifest.read_text(encoding="utf-8")) + payload["members"][0]["record_type"] = "operation" + manifest.write_text(json.dumps(payload), encoding="utf-8") + report = validate_manifest(root, manifest) + self.assertIn("project-a: manifest=operation direct=project", report["type_conflicts"]) + + def test_exact_database_document_map_detects_stale_projection(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + write_record(root / "30_projects", "project-a") + manifest = typed_manifest(root) + payload = json.loads(manifest.read_text(encoding="utf-8")) + members, issues = manifest_members(payload) + self.assertEqual(issues, []) + expected = selected_source_files(root, payload, members) + self.assertTrue(expected) + + db = root / "projection.sqlite" + con = sqlite3.connect(db) + con.execute( + "CREATE TABLE documents (id TEXT, namespace TEXT, index_id TEXT, trust_profile TEXT, " + "source_root TEXT, source_path TEXT, display_path TEXT, content_hash TEXT, record_type TEXT)" + ) + for row in expected: + con.execute( + "INSERT INTO documents VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)", + tuple(row[key] for key in ( + "id", "namespace", "index_id", "trust_profile", "source_root", + "source_path", "display_path", "content_hash", "record_type", + )), + ) + con.commit() + con.close() + + green = verify_typed_document_map(db, root, manifest) + self.assertTrue(green["ok"], green) + self.assertEqual(green["expected_typed_document_map_hash"], document_map_hash(expected)) + self.assertEqual(green["expected_document_map_hash"], document_map_hash(expected, include_record_type=False)) + + (root / "30_projects" / "project-a" / "plans" / "plan.md").write_text( + "# changed\n", encoding="utf-8" + ) + stale = verify_typed_document_map(db, root, manifest) + self.assertFalse(stale["ok"]) + self.assertTrue(stale["content_mismatches"]) + + untyped_db = root / "untyped.sqlite" + con = sqlite3.connect(untyped_db) + con.execute( + "CREATE TABLE documents (id TEXT, namespace TEXT, index_id TEXT, trust_profile TEXT, " + "source_root TEXT, source_path TEXT, display_path TEXT, content_hash TEXT)" + ) + current_expected = selected_source_files(root, payload, members) + for row in current_expected: + con.execute( + "INSERT INTO documents VALUES (?, ?, ?, ?, ?, ?, ?, ?)", + tuple(row[key] for key in ( + "id", "namespace", "index_id", "trust_profile", "source_root", + "source_path", "display_path", "content_hash", + )), + ) + con.commit() + con.close() + untyped = verify_typed_document_map(untyped_db, root, manifest) + self.assertTrue(untyped["document_map_ok"], untyped) + self.assertFalse(untyped["record_type_direct_check"]) + self.assertTrue(untyped["record_type_derived_check"]) + self.assertEqual(untyped["record_type_source"], "derived_from_manifest") + self.assertEqual(untyped["record_type_unknown_documents"], []) + self.assertTrue(untyped["ok"]) + + def test_legacy_manifest_is_diagnostic_only(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + write_record(root / "30_projects", "project-a") + manifest = root / "legacy.json" + manifest.write_text(json.dumps({"projects": ["project-a"]}), encoding="utf-8") + report = validate_manifest(root, manifest) + self.assertFalse(report["ok"]) + self.assertIn("manifest.members is missing", report["issues"][0]) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_lifecycle_substrate_isolation.py b/tests/test_lifecycle_substrate_isolation.py new file mode 100644 index 0000000..74ed22e --- /dev/null +++ b/tests/test_lifecycle_substrate_isolation.py @@ -0,0 +1,53 @@ +"""Generic lifecycle substrate must not import MPE internals.""" + +from __future__ import annotations + +import ast +import unittest +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] +SCRIPTS = ROOT / "scripts" +GENERIC = ( + "lifecycle_identity.py", + "lifecycle_runtime.py", + "migration_lease.py", + "work_inventory.py", +) + + +def _imports(path: Path) -> set[str]: + tree = ast.parse(path.read_text(encoding="utf-8")) + names: set[str] = set() + for node in ast.walk(tree): + if isinstance(node, ast.Import): + names.update(alias.name.split(".")[0] for alias in node.names) + elif isinstance(node, ast.ImportFrom) and node.module: + names.add(node.module.split(".")[0]) + return names + + +class LifecycleSubstrateIsolationTests(unittest.TestCase): + def test_generic_modules_do_not_import_mpe_engine(self) -> None: + banned = { + "runtime", + "process_eval", + "migration", + "writer_smoke", + "mpe_runtime", + "mpe_migration", + } + for name in GENERIC: + imported = _imports(SCRIPTS / name) + overlap = imported & banned + self.assertFalse(overlap, f"{name} imports {overlap}") + + def test_generic_modules_do_not_hardcode_operation_src(self) -> None: + needle = "40_operations/mainframe-process-eval/src" + for name in GENERIC: + text = (SCRIPTS / name).read_text(encoding="utf-8") + self.assertNotIn(needle, text, name) + + def test_lifecycle_identity_has_no_mpe_helper(self) -> None: + text = (SCRIPTS / "lifecycle_identity.py").read_text(encoding="utf-8") + self.assertNotIn("def resolve_mpe", text) diff --git a/tests/test_loop_catalogue.py b/tests/test_loop_catalogue.py deleted file mode 100644 index fc787ad..0000000 --- a/tests/test_loop_catalogue.py +++ /dev/null @@ -1,137 +0,0 @@ -"""Loop-eval catalogue schema, scores, tier gates, surface paths.""" - -from __future__ import annotations - -import importlib.util -import sys -import unittest -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] -WB = ROOT / "30_projects" / "mainframe-process-eval" / "workbench" / "loop-eval" -SCRIPT = WB / "scripts" / "score_catalogue.py" -CATALOGUE = WB / "catalogue" / "entries.json" -SCHEMA = WB / "schema" / "loop-entry.schema.json" - - -def _load_score_mod(): - spec = importlib.util.spec_from_file_location("score_catalogue", SCRIPT) - assert spec and spec.loader - mod = importlib.util.module_from_spec(spec) - sys.modules[spec.name] = mod - spec.loader.exec_module(mod) - return mod - - -class LoopCatalogueTests(unittest.TestCase): - @classmethod - def setUpClass(cls) -> None: - cls.mod = _load_score_mod() - cls.data = cls.mod.load_catalogue(CATALOGUE) - - def test_workbench_files_exist(self) -> None: - for rel in ( - "README.md", - "THRESHOLDS.md", - "catalogue/entries.json", - "catalogue/entries.yaml", - "catalogue/INDEX.md", - "catalogue/01-well-defined.md", - "catalogue/02-underspecified.md", - "catalogue/03-promotion-candidates.md", - "compositions/README.md", - "schema/loop-entry.schema.json", - "scripts/score_catalogue.py", - ): - path = WB / rel - self.assertTrue(path.is_file(), msg=f"missing {rel}") - - def test_catalogue_validates_clean(self) -> None: - problems = self.mod.validate_catalogue(self.data, root=ROOT) - self.assertEqual(problems, [], msg="\n".join(problems)) - - def test_unique_ids(self) -> None: - ids = [e["id"] for e in self.data["entries"]] - self.assertEqual(len(ids), len(set(ids))) - - def test_scores_match_checklist(self) -> None: - for e in self.data["entries"]: - computed = self.mod.score_entry(e) - self.assertEqual( - computed, - sum(1 for v in e["checklist"].values() if v is True), - msg=e["id"], - ) - self.assertGreaterEqual(computed, 0) - self.assertLessEqual(computed, 10) - - def test_tier_a_entries_meet_gates(self) -> None: - for e in self.data["entries"]: - if e.get("tier") != "well_defined": - continue - self.assertTrue( - self.mod.tier_a_gates_ok(e), - msg=f"{e['id']} fails tier A gates (score={self.mod.score_entry(e)})", - ) - - def test_at_least_three_well_defined(self) -> None: - n = sum(1 for e in self.data["entries"] if e.get("tier") == "well_defined") - self.assertGreaterEqual(n, 3) - - def test_solidified_loops_are_well_defined(self) -> None: - """Former tier-B loops must stay A after 2026-07-23 solidification.""" - required = { - "loop-process-evaluation", - "loop-income-wave1", - "loop-repo-radar-harvest", - "loop-session-lifecycle", - } - by_id = {e["id"]: e for e in self.data["entries"]} - for eid in required: - self.assertIn(eid, by_id) - self.assertEqual(by_id[eid]["tier"], "well_defined", msg=eid) - self.assertTrue(self.mod.tier_a_gates_ok(by_id[eid]), msg=eid) - - def test_process_eval_bin_help(self) -> None: - import subprocess - - proc = subprocess.run( - [str(ROOT / "bin" / "process-eval"), "--help"], - cwd=str(ROOT), - capture_output=True, - text=True, - check=False, - ) - self.assertEqual(proc.returncode, 0, msg=proc.stderr) - self.assertIn("preflight", proc.stdout) - - def test_required_tiers_present(self) -> None: - """A and C must exist; B may be empty after solidification sweeps.""" - tiers = {e.get("tier") for e in self.data["entries"]} - self.assertIn("well_defined", tiers) - self.assertIn("candidate", tiers) - # underspecified optional - for t in tiers: - self.assertIn(t, {"well_defined", "underspecified", "candidate"}) - - def test_schema_file_lists_checklist_keys(self) -> None: - text = SCHEMA.read_text(encoding="utf-8") - for key in self.mod.CHECKLIST_KEYS: - self.assertIn(key, text) - - def test_cli_score_script_exits_zero(self) -> None: - import subprocess - - proc = subprocess.run( - [sys.executable, str(SCRIPT)], - cwd=str(ROOT), - capture_output=True, - text=True, - check=False, - ) - self.assertEqual(proc.returncode, 0, msg=proc.stdout + proc.stderr) - self.assertIn("OK", proc.stdout) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_mainframe_doctor.py b/tests/test_mainframe_doctor.py deleted file mode 100644 index e88e14a..0000000 --- a/tests/test_mainframe_doctor.py +++ /dev/null @@ -1,541 +0,0 @@ -"""Unit 1.3 — mainframe-doctor catalogue, aggregation, fixture shell.""" - -from __future__ import annotations - -import json -import shutil -import sqlite3 -import stat -import sys -import tempfile -import time -import unittest -from datetime import datetime, timezone -from pathlib import Path -from unittest import mock - -ROOT = Path(__file__).resolve().parents[1] -SCRIPTS = ROOT / "scripts" -sys.path.insert(0, str(SCRIPTS)) - -from mainframe_doctor.catalogue import load_catalogue_pair, validate_catalogue # noqa: E402 -from mainframe_doctor import providers # noqa: E402 -from mainframe_doctor.runner import run_doctor # noqa: E402 -from mainframe_doctor.schema import ( # noqa: E402 - CheckResult, - aggregate_health, - exit_code, - redact_secret_shaped, -) - - -CAT = ROOT / ".context" / "doctor" / "catalogue.json" -INV = ROOT / ".context" / "doctor" / "required-invariants.json" -FIXTURES = ROOT / "tests" / "fixtures" / "mainframe_doctor" - - -class AggregationTests(unittest.TestCase): - def _c( - self, - status: str, - *, - required: bool = True, - severity: str = "high", - cid: str = "X", - ) -> CheckResult: - return CheckResult( - id=cid, - subsystem="t", - layer="t", - status=status, # type: ignore[arg-type] - severity=severity, # type: ignore[arg-type] - required=required, - expected="e", - observed="o", - observed_at=None, - freshness_seconds=0, - authority="test", - ) - - def test_required_unknown_not_healthy(self) -> None: - h = aggregate_health([self._c("pass"), self._c("unknown", cid="U")]) - self.assertEqual(h, "unknown") - self.assertEqual(exit_code(h), 1) - - def test_required_fail_unhealthy(self) -> None: - h = aggregate_health([self._c("pass"), self._c("fail", cid="F")]) - self.assertEqual(h, "unhealthy") - - def test_optional_skip_does_not_block_healthy(self) -> None: - h = aggregate_health( - [ - self._c("pass", cid="A"), - self._c("skip", required=False, cid="B"), - ] - ) - self.assertEqual(h, "healthy") - self.assertEqual(exit_code(h), 0) - - def test_required_skip_is_unknown(self) -> None: - h = aggregate_health([self._c("pass"), self._c("skip", required=True, cid="S")]) - self.assertEqual(h, "unknown") - - def test_warn_degraded(self) -> None: - h = aggregate_health([self._c("pass"), self._c("warn", severity="low", cid="W")]) - self.assertEqual(h, "degraded") - - def test_internal_error_exit_2(self) -> None: - self.assertEqual(exit_code("unknown", internal_error=True), 2) - - def test_redact_secret_shaped(self) -> None: - s = redact_secret_shaped("client_secret=abcDEF1234567890supersecretvalue") - self.assertNotIn("supersecretvalue", s) - self.assertIn("REDACTED", s) - - -class CatalogueTests(unittest.TestCase): - def test_live_catalogue_complete_against_invariants(self) -> None: - loaded = load_catalogue_pair(CAT, INV) - self.assertTrue(loaded.ok, msg=loaded.errors) - self.assertGreaterEqual(len(loaded.checks_by_id), len(loaded.required_ids)) - - def test_missing_required_id_fails_completeness(self) -> None: - cat = json.loads(CAT.read_text(encoding="utf-8")) - cat["checks"] = [c for c in cat["checks"] if c["id"] != "CLI-001"] - errors = validate_catalogue(cat, json.loads(INV.read_text())["required_check_ids"]) - self.assertTrue(any("CLI-001" in e for e in errors)) - - def test_duplicate_id_fails(self) -> None: - cat = json.loads(CAT.read_text(encoding="utf-8")) - cat["checks"] = cat["checks"] + [cat["checks"][0]] - errors = validate_catalogue(cat, []) - self.assertTrue(any("duplicate" in e for e in errors)) - - def test_mutating_provider_rejected_in_catalogue(self) -> None: - cat = { - "checks": [ - { - "id": "X-001", - "owner": "t", - "subsystem": "t", - "layer": "t", - "provider": "x", - "required": True, - "mutates": True, - "isolation": "live", - "timeout_seconds": 1, - "skip_policy": "fail", - "pass_condition": "x", - "remediation": "x", - "safe_fix_available": False, - "contract_version": 1, - } - ] - } - errors = validate_catalogue(cat, ["X-001"]) - self.assertTrue(any("mutates" in e for e in errors)) - - def test_nonpositive_provider_timeout_is_rejected(self) -> None: - cat = json.loads(CAT.read_text(encoding="utf-8")) - cat["checks"][0]["timeout_seconds"] = 0 - errors = validate_catalogue(cat, []) - self.assertTrue(any("positive integer" in error for error in errors)) - - -class ProviderTimeoutTests(unittest.TestCase): - def test_slow_provider_becomes_unknown_at_its_catalogue_deadline(self) -> None: - check = { - "id": "SLOW-001", - "owner": "test", - "subsystem": "test", - "layer": "unit", - "provider": "slow_test_provider", - "required": True, - "mutates": False, - "timeout_seconds": 1, - "pass_condition": "returns before deadline", - "remediation": "fix slow provider", - } - - def slow_provider(_check: dict, _ctx: dict) -> CheckResult: - time.sleep(10) - raise AssertionError("deadline was not enforced") - - started = time.perf_counter() - with mock.patch.dict( - providers.PROVIDERS, - {"slow_test_provider": slow_provider}, - ): - result = providers.run_provider(check, {"root": ROOT}) - elapsed = time.perf_counter() - started - - self.assertLess(elapsed, 2.0) - self.assertEqual(result.status, "unknown") - self.assertIn("timed out after 1s", result.message) - - -class MindGraphProviderTests(unittest.TestCase): - def _check(self, check_id: str) -> dict: - catalogue = json.loads(CAT.read_text(encoding="utf-8")) - return next(c for c in catalogue["checks"] if c["id"] == check_id) - - def _write_db(self, path: Path, *, missing: set[str] | None = None) -> None: - path.parent.mkdir(parents=True, exist_ok=True) - missing = missing or set() - con = sqlite3.connect(path) - for table in ("documents", "documents_fts", "chunks", "vec_chunks", "edges"): - if table not in missing: - con.execute(f"CREATE TABLE {table} (id TEXT, namespace TEXT)") - con.commit() - con.close() - - def test_mg_001_checks_projects_schema_not_only_knowledge(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - home = Path(tmp) - mg = home / ".mindgraph" - self._write_db(mg / "mainframe.sqlite") - self._write_db(mg / "mainframe-projects.sqlite", missing={"edges"}) - with mock.patch.object(providers.Path, "home", return_value=home): - result = providers.provider_mg_db_presence( - self._check("MG-001"), {"root": ROOT} - ) - self.assertEqual(result.status, "fail") - self.assertIn("projects tables missing", result.observed) - - def test_mg_003_fails_when_installed_namespace_is_stale(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - temp = Path(tmp) - root = temp / "root" - projects = root / "30_projects" - for slug in ("alpha", "beta"): - project = projects / slug - project.mkdir(parents=True) - (project / "README.md").write_text("# project\n", encoding="utf-8") - (projects / "mindgraph-projects.json").write_text( - json.dumps({"projects": ["alpha", "beta"]}), encoding="utf-8" - ) - - home = temp / "home" - db_path = home / ".mindgraph" / "mainframe-projects.sqlite" - db_path.parent.mkdir(parents=True) - con = sqlite3.connect(db_path) - con.execute("CREATE TABLE documents (namespace TEXT)") - con.execute("INSERT INTO documents VALUES ('alpha')") - con.commit() - con.close() - - with mock.patch.object(providers.Path, "home", return_value=home): - result = providers.provider_mg_manifest_coverage( - self._check("MG-003"), {"root": root} - ) - self.assertEqual(result.status, "fail") - self.assertIn("beta", result.observed) - self.assertIn("stale", result.message) - - -class FixtureShellTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - # minimal catalogue copy - doc = self.root / ".context" / "doctor" - doc.mkdir(parents=True) - shutil.copy(CAT, doc / "catalogue.json") - shutil.copy(INV, doc / "required-invariants.json") - - def tearDown(self) -> None: - self.tmp.cleanup() - - def _write_fixture(self, name: str, payload: dict) -> Path: - d = self.root / "fixtures" / name - d.mkdir(parents=True) - (d / "fixture.json").write_text(json.dumps(payload), encoding="utf-8") - return d - - def test_all_pass_fixture_healthy(self) -> None: - cat = json.loads(CAT.read_text(encoding="utf-8")) - checks = { - c["id"]: {"status": "pass", "observed": "ok", "message": "fixture pass"} - for c in cat["checks"] - if c.get("required") - } - # optional can skip - for c in cat["checks"]: - if not c.get("required"): - checks[c["id"]] = {"status": "skip", "observed": "opt", "message": "skip"} - fix = self._write_fixture("all-pass", {"checks": checks}) - report, code = run_doctor( - root=self.root, - profile="deep", - fixture_path=fix, - ) - self.assertIsNone(report.internal_error) - self.assertEqual(report.health, "healthy") - self.assertEqual(code, 0) - - def test_required_unknown_fixture_not_healthy(self) -> None: - cat = json.loads(CAT.read_text(encoding="utf-8")) - checks = { - c["id"]: {"status": "pass", "observed": "ok", "message": "p"} - for c in cat["checks"] - } - checks["SESSION-001"] = { - "status": "unknown", - "observed": "missing evidence", - "message": "unknown", - } - fix = self._write_fixture("req-unknown", {"checks": checks}) - report, code = run_doctor(root=self.root, profile="deep", fixture_path=fix) - self.assertEqual(report.health, "unknown") - self.assertEqual(code, 1) - - def test_compound_session_fixture_fails(self) -> None: - # synthetic root with compound STATE - synth = self.root / "synth" - synth.mkdir() - (synth / "STATE.md").write_text( - "# S\n\n## Active Project\n\nfoo (bar) + baz (qux)\n", - encoding="utf-8", - ) - (synth / "AGENTS.md").write_text("# a\n", encoding="utf-8") - (synth / "30_projects").mkdir() - # only override non-session checks to pass; let session provider run - cat = json.loads(CAT.read_text(encoding="utf-8")) - checks = {} - for c in cat["checks"]: - if c["id"] == "SESSION-001": - continue # live provider against fixture root - checks[c["id"]] = {"status": "pass", "observed": "ok", "message": "p"} - fix = self._write_fixture( - "compound-session", - {"root": str(synth), "checks": checks}, - ) - report, code = run_doctor(root=self.root, profile="deep", fixture_path=fix) - by_id = {c.id: c for c in report.checks} - self.assertEqual(by_id["SESSION-001"].status, "fail") - self.assertIn(report.health, ("unhealthy", "unknown", "degraded")) - self.assertEqual(code, 1) - - def test_cli_help_unsafe_source_fixture(self) -> None: - synth = self.root / "cli-root" - bin_dir = synth / "bin" - bin_dir.mkdir(parents=True) - # minimal unsafe wrapper (legacy pre-Unit-2.1) - (bin_dir / "mindgraph-refresh-projects").write_text( - '#!/bin/bash\n' - 'if [[ "${1:-}" == "--dry-run" ]]; then exit 0; fi\n' - 'echo ingest\n', - encoding="utf-8", - ) - cat = json.loads(CAT.read_text(encoding="utf-8")) - checks = { - c["id"]: {"status": "pass", "observed": "ok", "message": "p"} - for c in cat["checks"] - if c["id"] != "CLI-001" - } - fix = self._write_fixture( - "cli-unsafe", - {"root": str(synth), "checks": checks}, - ) - report, _code = run_doctor(root=self.root, profile="deep", fixture_path=fix) - by_id = {c.id: c for c in report.checks} - self.assertEqual(by_id["CLI-001"].status, "fail") - - def test_cli_help_safe_source_fixture(self) -> None: - synth = self.root / "cli-safe" - bin_dir = synth / "bin" - bin_dir.mkdir(parents=True) - (bin_dir / "mindgraph-refresh-projects").write_text( - "#!/bin/bash\n" - "usage() { echo Show this help; }\n" - 'while [[ $# -gt 0 ]]; do case "$1" in\n' - " -h|--help) usage; exit 0 ;;\n" - " --dry-run) shift ;;\n" - " --apply) shift ;;\n" - ' -*) echo "error: unknown option: $1"; exit 2 ;;\n' - "esac; done\n" - 'if [[ "$DRY_RUN" -eq 0 ]]; then echo "refusing to mutate without --apply"; exit 2; fi\n', - encoding="utf-8", - ) - cat = json.loads(CAT.read_text(encoding="utf-8")) - checks = { - c["id"]: {"status": "pass", "observed": "ok", "message": "p"} - for c in cat["checks"] - if c["id"] != "CLI-001" - } - fix = self._write_fixture( - "cli-safe", - {"root": str(synth), "checks": checks}, - ) - report, _code = run_doctor(root=self.root, profile="deep", fixture_path=fix) - by_id = {c.id: c for c in report.checks} - self.assertEqual(by_id["CLI-001"].status, "pass") - - def test_workstation_zero_ops_fixture(self) -> None: - import sqlite3 - - synth = self.root / "ws-root" - db_path = synth / "20_live" / "workstation" - db_path.mkdir(parents=True) - db = db_path / "workstation.sqlite" - con = sqlite3.connect(db) - con.execute("CREATE TABLE tasks (id INTEGER)") - con.execute("CREATE TABLE runs (id INTEGER)") - con.execute("CREATE TABLE approvals (id INTEGER)") - con.execute("CREATE TABLE artifacts (id INTEGER)") - for _ in range(5): - con.execute("INSERT INTO tasks VALUES (1)") - con.commit() - con.close() - cat = json.loads(CAT.read_text(encoding="utf-8")) - checks = { - c["id"]: {"status": "pass", "observed": "ok", "message": "p"} - for c in cat["checks"] - if c["id"] not in ("WS-001", "WS-003") - } - fix = self._write_fixture( - "ws-empty-ops", - {"root": str(synth), "checks": checks}, - ) - report, _code = run_doctor(root=self.root, profile="deep", fixture_path=fix) - by_id = {c.id: c for c in report.checks} - self.assertEqual(by_id["WS-001"].status, "pass") - self.assertEqual(by_id["WS-003"].status, "fail") - - def test_provider_exception_becomes_unknown(self) -> None: - # force broken root for mg provider only — use override for exception sim via bad fixture status - cat = json.loads(CAT.read_text(encoding="utf-8")) - checks = { - c["id"]: {"status": "pass", "observed": "ok", "message": "p"} - for c in cat["checks"] - } - checks["TEL-001"] = { - "status": "unknown", - "observed": "RuntimeError", - "message": "provider exception: RuntimeError", - } - fix = self._write_fixture("prov-exc", {"checks": checks}) - report, code = run_doctor(root=self.root, profile="deep", fixture_path=fix) - self.assertEqual(report.health, "unknown") - self.assertEqual(code, 1) - - def test_bad_catalogue_exit_2(self) -> None: - bad = self.root / ".context" / "doctor" / "catalogue.json" - bad.write_text('{"checks":[]}', encoding="utf-8") - report, code = run_doctor(root=self.root, profile="quick") - self.assertEqual(code, 2) - self.assertIsNotNone(report.internal_error) - - def test_json_serializable(self) -> None: - cat = json.loads(CAT.read_text(encoding="utf-8")) - checks = { - c["id"]: {"status": "fail", "observed": "x", "message": "y"} - for c in cat["checks"] - } - fix = self._write_fixture("ser", {"checks": checks}) - report, _code = run_doctor(root=self.root, profile="quick", fixture_path=fix) - blob = json.dumps(report.to_dict()) - self.assertIn("schema_version", blob) - self.assertNotIn("supersecret", blob) - - -class LiveSmokeTests(unittest.TestCase): - """Live probes — doctor shell must not claim healthy while system degraded.""" - - def test_live_quick_not_healthy_or_internal_ok(self) -> None: - report, code = run_doctor(root=ROOT, profile="quick") - self.assertIsNone(report.internal_error, msg=report.internal_error) - # SEC-001 residual disposition keeps live vector from claiming healthy. - self.assertIn(report.health, ("unknown", "unhealthy", "degraded")) - self.assertEqual(code, 1) - ids = {c.id for c in report.checks} - self.assertIn("SESSION-001", ids) - self.assertIn("CLI-001", ids) - - def test_live_cli_001_passes_after_unit_2_1(self) -> None: - report, _code = run_doctor(root=ROOT, profile="deep") - by_id = {c.id: c for c in report.checks} - self.assertEqual(by_id["CLI-001"].status, "pass") - - def test_live_auth_001_reports_state_and_task_001_passes(self) -> None: - report, _code = run_doctor(root=ROOT, profile="quick") - by_id = {c.id: c for c in report.checks} - self.assertIn(by_id["AUTH-001"].status, ("pass", "warn", "stale"), by_id["AUTH-001"].message) - self.assertEqual(by_id["TASK-001"].status, "pass", by_id["TASK-001"].message) - - def test_focus_review_staleness_is_deterministic(self) -> None: - check = next( - c for c in json.loads(CAT.read_text(encoding="utf-8"))["checks"] if c["id"] == "AUTH-001" - ) - result = providers.provider_auth_focus( - check, - {"root": ROOT, "now": datetime(2026, 8, 2, tzinfo=timezone.utc)}, - ) - self.assertEqual(result.status, "stale") - - def test_live_quick_has_no_unimplemented_unknowns(self) -> None: - report, _code = run_doctor(root=ROOT, profile="quick") - unknowns = [c for c in report.checks if c.status == "unknown"] - self.assertEqual(unknowns, [], msg=[(c.id, c.message) for c in unknowns]) - by_id = {c.id: c for c in report.checks} - for cid in ("SESSION-002", "STRUCT-001", "TEL-001"): - self.assertIn(cid, by_id) - self.assertNotEqual(by_id[cid].status, "unknown", by_id[cid].message) - # May be degraded (SEC-001) or unhealthy if a required live surface fails; - # the contract is no silent unimplemented unknowns in the quick profile. - self.assertNotEqual(report.health, "unknown") - - -class NewProviderUnitTests(unittest.TestCase): - def _check(self, cid: str = "X") -> dict: - return { - "id": cid, - "subsystem": "t", - "layer": "t", - "required": True, - "pass_condition": "p", - "remediation": "r", - "owner": "t", - } - - def test_session_phase_alignment_missing_project(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - (root / "STATE.md").write_text("# no active project\n", encoding="utf-8") - result = providers.provider_session_phase_alignment(self._check(), {"root": root}) - self.assertEqual(result.status, "fail") - - def test_structure_bounds_missing_contract(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - result = providers.provider_structure_bounds(self._check(), {"root": root}) - self.assertEqual(result.status, "fail") - self.assertIn("missing", result.observed) - - def test_tel_hash_integrity_missing_store(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - result = providers.provider_tel_hash_integrity(self._check(), {"root": root}) - self.assertEqual(result.status, "fail") - - def test_tel_hash_integrity_intact_chain(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - events = root / "20_live" / "workflow-metrics" / "events" - events.mkdir(parents=True) - prev = "genesis-seed" - lines = [] - for i in range(3): - body = {"event": "test", "n": i, "as_of": f"2026-07-23T00:00:0{i}Z"} - h = providers._telemetry_compute_hash(body, prev) - row = dict(body) - row["hash_chain"] = h - lines.append(json.dumps(row, sort_keys=True, separators=(",", ":"))) - prev = h - (events / "2026-07-23.jsonl").write_text("\n".join(lines) + "\n", encoding="utf-8") - result = providers.provider_tel_hash_integrity(self._check(), {"root": root}) - self.assertEqual(result.status, "pass", result.observed) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_migration_lease.py b/tests/test_migration_lease.py new file mode 100644 index 0000000..1211364 --- /dev/null +++ b/tests/test_migration_lease.py @@ -0,0 +1,128 @@ +from __future__ import annotations + +import json +import os +import subprocess +import sys +import tempfile +import time +import unittest +from unittest import mock +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] +sys.path.insert(0, str(ROOT / "scripts")) + +from migration_lease import ( # noqa: E402 + LeaseBusy, + LeaseUnavailable, + MigrationLease, + exclusive_migration_lease, + shared_writer_lease, +) + + +class MigrationLeaseTests(unittest.TestCase): + def test_exclusive_lease_fences_shared_writer(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + state = root / "state" + with exclusive_migration_lease(root, state_root=state): + writer = shared_writer_lease(root, "old-path-writer", state_root=state) + with self.assertRaises(LeaseBusy): + writer.acquire() + self.assertFalse((root / "30_projects" / "mainframe-process-eval").exists()) + + def test_stale_state_is_recovered_under_the_lock(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + state = root / "state" + state.mkdir() + (state / "lease-state.json").write_text( + json.dumps( + { + "mode": "exclusive", + "owner_pid": 999999999, + "epoch": "stale", + } + ), + encoding="utf-8", + ) + lease = MigrationLease(root, "shared", "recovery-test", state_root=state) + lease.acquire() + try: + current = json.loads((state / "lease-state.json").read_text(encoding="utf-8")) + self.assertTrue(current["recovered_stale_state"]) + event = lease.record_write(root / "receipt.json", action="test") + self.assertEqual(event["holder"], "recovery-test") + finally: + lease.release() + + def test_writer_must_hold_lease_to_record_write(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + lease = shared_writer_lease(root, "writer", state_root=root / "state") + with self.assertRaises(Exception): + lease.record_write(root / "old-path", action="write") + + def test_concurrent_shared_holders_have_distinct_audit_records(self) -> None: + worker = """ +import sys +from pathlib import Path +sys.path.insert(0, sys.argv[2]) +from migration_lease import shared_writer_lease +root = Path(sys.argv[1]); state = Path(sys.argv[3]) +with shared_writer_lease(root, sys.argv[4], state_root=state) as lease: + print(lease.epoch, flush=True) + sys.stdin.read() +""" + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + state = root / "state" + env = os.environ.copy() + env["PYTHONPATH"] = str(ROOT / "scripts") + processes = [ + subprocess.Popen( + [sys.executable, "-c", worker, str(root), str(ROOT / "scripts"), str(state), f"holder-{i}"], + stdin=subprocess.PIPE, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + text=True, + env=env, + ) + for i in range(2) + ] + epochs = [process.stdout.readline().strip() for process in processes] + self.assertEqual(len(set(epochs)), 2) + holders = list((state / "holders").glob("*.json")) + self.assertEqual(len(holders), 2) + summary = json.loads((state / "lease-state.json").read_text(encoding="utf-8")) + self.assertEqual(len(summary["active_holders"]), 2) + for process in processes: + process.stdin.write("release\n") + process.stdin.close() + process.wait(timeout=10) + self.assertFalse(list((state / "holders").glob("*.json"))) + + def test_corrupt_or_failed_audit_metadata_does_not_leak_lock_fd(self) -> None: + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + state = root / "state" + state.mkdir() + (state / "lease-state.json").write_text("{not-json", encoding="utf-8") + lease = MigrationLease(root, "shared", "corrupt", state_root=state) + with self.assertRaises(LeaseUnavailable): + lease.acquire() + self.assertIsNone(lease._lock_handle) + (state / "lease-state.json").unlink() + with mock.patch("migration_lease._write_state", side_effect=OSError("disk full")): + failing = MigrationLease(root, "shared", "write-failure", state_root=state) + with self.assertRaises(LeaseUnavailable): + failing.acquire() + self.assertIsNone(failing._lock_handle) + with shared_writer_lease(root, "after-failure", state_root=state): + self.assertTrue((state / "lease-state.json").exists()) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_mindgraph_projects_apply.py b/tests/test_mindgraph_projects_apply.py deleted file mode 100644 index 63de4f6..0000000 --- a/tests/test_mindgraph_projects_apply.py +++ /dev/null @@ -1,119 +0,0 @@ -"""Tests for staged projects MindGraph apply (ADR-045).""" - -from __future__ import annotations - -import json -import sqlite3 -import sys -import tempfile -import unittest -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] -sys.path.insert(0, str(ROOT / "scripts")) - -import mindgraph_projects_apply as mpa # noqa: E402 - - -class CoverageTests(unittest.TestCase): - def test_coverage_complete(self) -> None: - cov = mpa.coverage_report(["a", "b"], ["a", "b"]) - self.assertTrue(cov["complete"]) - self.assertEqual(cov["missing_dirs"], []) - self.assertEqual(cov["omitted_dirs"], []) - - def test_coverage_missing_and_omitted(self) -> None: - cov = mpa.coverage_report(["a", "ghost"], ["a", "extra"]) - self.assertFalse(cov["complete"]) - self.assertEqual(cov["missing_dirs"], ["ghost"]) - self.assertEqual(cov["omitted_dirs"], ["extra"]) - - -class NamespaceVerifyTests(unittest.TestCase): - def test_verify_stage_ok(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - db = Path(tmp) / "t.sqlite" - con = sqlite3.connect(db) - con.execute( - "CREATE TABLE documents (id TEXT, namespace TEXT, title TEXT)" - ) - for ns in ("alpha", "beta"): - con.execute( - "INSERT INTO documents VALUES (?, ?, ?)", - (f"id-{ns}", ns, "t"), - ) - con.commit() - con.close() - # pad size > 10k for ok check - with db.open("ab") as f: - f.write(b"x" * 12_000) - v = mpa.verify_stage(db, ["alpha", "beta"]) - self.assertTrue(v["ok"], v) - self.assertEqual(v["missing_namespaces"], []) - - def test_verify_stage_missing(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - db = Path(tmp) / "t.sqlite" - con = sqlite3.connect(db) - con.execute( - "CREATE TABLE documents (id TEXT, namespace TEXT, title TEXT)" - ) - con.execute("INSERT INTO documents VALUES ('1', 'alpha', 't')") - con.commit() - con.close() - with db.open("ab") as f: - f.write(b"x" * 12_000) - v = mpa.verify_stage(db, ["alpha", "beta"]) - self.assertFalse(v["ok"]) - self.assertEqual(v["missing_namespaces"], ["beta"]) - - -class ManifestLoadTests(unittest.TestCase): - def test_live_manifest_loads(self) -> None: - projects = mpa.load_manifest_projects(mpa.DEFAULT_MANIFEST) - self.assertGreaterEqual(len(projects), 20) - real = mpa.real_project_dirs(ROOT / "30_projects") - cov = mpa.coverage_report(projects, real) - self.assertTrue(cov["complete"], cov) - - def test_default_manifest_excludes_outputs(self) -> None: - payload = json.loads(mpa.DEFAULT_MANIFEST.read_text(encoding="utf-8")) - self.assertEqual(payload.get("profile"), "default") - includes = payload.get("include") or [] - self.assertFalse(any("outputs" in p for p in includes)) - excludes = payload.get("exclude") or [] - self.assertTrue(any(p.startswith("outputs") for p in excludes)) - - def test_deep_manifest_includes_outputs(self) -> None: - payload = json.loads(mpa.DEEP_MANIFEST.read_text(encoding="utf-8")) - self.assertEqual(payload.get("profile"), "deep") - includes = payload.get("include") or [] - self.assertTrue(any("outputs" in p for p in includes)) - - def test_resolve_manifest_deep(self) -> None: - from types import SimpleNamespace - - args = SimpleNamespace(manifest=str(mpa.DEFAULT_MANIFEST), deep=True) - self.assertEqual(mpa.resolve_manifest(args), mpa.DEEP_MANIFEST) - args2 = SimpleNamespace(manifest=str(mpa.DEFAULT_MANIFEST), deep=False) - self.assertEqual(mpa.resolve_manifest(args2), mpa.DEFAULT_MANIFEST) - - -class CliHelpTests(unittest.TestCase): - def test_bin_help(self) -> None: - import subprocess - - proc = subprocess.run( - [str(ROOT / "bin" / "mindgraph-projects-apply"), "--help"], - capture_output=True, - text=True, - check=False, - ) - self.assertEqual(proc.returncode, 0) - self.assertIn("--plan", proc.stdout) - self.assertIn("--stage", proc.stdout) - self.assertIn("--promote", proc.stdout) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_mindgraph_refresh_projects.py b/tests/test_mindgraph_refresh_projects.py deleted file mode 100644 index d31250e..0000000 --- a/tests/test_mindgraph_refresh_projects.py +++ /dev/null @@ -1,121 +0,0 @@ -"""Unit 2.1 — mindgraph-refresh-projects fail-closed CLI contract.""" - -from __future__ import annotations - -import hashlib -import os -import subprocess -import tempfile -import unittest -from pathlib import Path - -ROOT = Path(__file__).resolve().parents[1] -SCRIPT = ROOT / "bin" / "mindgraph-refresh-projects" - - -def _run(args: list[str], *, env: dict | None = None, cwd: Path | None = None) -> subprocess.CompletedProcess[str]: - e = os.environ.copy() - if env: - e.update(env) - return subprocess.run( - [str(SCRIPT), *args], - cwd=str(cwd or ROOT), - capture_output=True, - text=True, - env=e, - check=False, - ) - - -def _sha256(path: Path) -> str: - h = hashlib.sha256() - with path.open("rb") as f: - for chunk in iter(lambda: f.read(1 << 20), b""): - h.update(chunk) - return h.hexdigest() - - -class MindgraphRefreshProjectsCLITests(unittest.TestCase): - def test_help_exits_zero_and_mentions_apply(self) -> None: - proc = _run(["--help"]) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn("--apply", proc.stdout) - self.assertIn("--dry-run", proc.stdout) - self.assertIn("--deep", proc.stdout) - self.assertNotIn("ingest-many", proc.stdout.split("Safety:")[0] if "Safety:" in proc.stdout else "") - - def test_dry_run_lean_excludes_outputs_include(self) -> None: - proc = _run(["--dry-run"]) - self.assertEqual(proc.returncode, 0, proc.stderr + proc.stdout) - self.assertIn("lean/default", proc.stdout) - # Includes block should not list outputs globs for default - after = proc.stdout.split("Includes:")[-1] - self.assertNotIn("outputs/*.md", after) - self.assertNotIn("outputs/**/*.md", after) - - def test_dry_run_deep_includes_outputs(self) -> None: - proc = _run(["--dry-run", "--deep"]) - self.assertEqual(proc.returncode, 0, proc.stderr + proc.stdout) - self.assertIn("(deep)", proc.stdout) - after = proc.stdout.split("Includes:")[-1] - self.assertIn("outputs", after) - - def test_help_flag_order_independent(self) -> None: - # --help wins as first matched option when alone; with other flags, - # any --help in argv should still be non-mutating if processed in loop. - proc = _run(["--full", "--help"]) - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertIn("Usage:", proc.stdout) - - def test_unknown_flag_fails_closed(self) -> None: - proc = _run(["--not-a-real-flag"]) - self.assertEqual(proc.returncode, 2) - self.assertIn("unknown option", proc.stderr.lower()) - - def test_positional_arg_fails_closed(self) -> None: - proc = _run(["surprise"]) - self.assertEqual(proc.returncode, 2) - self.assertIn("unexpected argument", proc.stderr.lower()) - - def test_bare_invocation_refuses_without_apply(self) -> None: - with tempfile.TemporaryDirectory() as td: - sentinel = Path(td) / "sentinel.sqlite" - sentinel.write_bytes(b"before") - before = _sha256(sentinel) - proc = _run([], env={"MINDGRAPH_DB_PATH": str(sentinel)}) - self.assertEqual(proc.returncode, 2) - self.assertIn("refusing to mutate without --apply", proc.stderr) - self.assertEqual(_sha256(sentinel), before) - self.assertEqual(sentinel.read_bytes(), b"before") - - def test_dry_run_does_not_change_sentinel_db_bytes(self) -> None: - with tempfile.TemporaryDirectory() as td: - sentinel = Path(td) / "sentinel.sqlite" - payload = b"x" * 4096 - sentinel.write_bytes(payload) - before = _sha256(sentinel) - # flag order: --full before --dry-run - proc = _run( - ["--full", "--dry-run"], - env={"MINDGRAPH_DB_PATH": str(sentinel)}, - ) - self.assertEqual(proc.returncode, 0, proc.stderr + proc.stdout) - self.assertIn("dry-run", proc.stdout.lower()) - self.assertEqual(sentinel.read_bytes(), payload) - self.assertEqual(_sha256(sentinel), before) - - def test_dry_run_apply_together_rejected(self) -> None: - proc = _run(["--dry-run", "--apply"]) - self.assertEqual(proc.returncode, 2) - self.assertIn("refuse", proc.stderr.lower()) - - def test_help_does_not_require_db(self) -> None: - with tempfile.TemporaryDirectory() as td: - missing = Path(td) / "nope.sqlite" - proc = _run(["--help"], env={"MINDGRAPH_DB_PATH": str(missing)}) - self.assertEqual(proc.returncode, 0) - self.assertFalse(missing.exists()) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_prep_ingest.py b/tests/test_prep_ingest.py deleted file mode 100644 index bfabfc7..0000000 --- a/tests/test_prep_ingest.py +++ /dev/null @@ -1,225 +0,0 @@ -from __future__ import annotations - -import importlib.util -import sys -import tempfile -import unittest -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -PREP_PATH = ROOT / "01_ingest" / "prep_ingest.py" -SPEC = importlib.util.spec_from_file_location("prep_ingest", PREP_PATH) -assert SPEC and SPEC.loader -prep_module = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = prep_module -SPEC.loader.exec_module(prep_module) - -PrepIngest = prep_module.PrepIngest - - -def enriched_note( - domain: str = "ai-systems", - item_type: str = "note", - status: str = "extracted", - title: str = "Enriched Note", -) -> str: - return "\n".join( - [ - "---", - f'title: "{title}"', - f'domain: "{domain}"', - f'type: "{item_type}"', - f'status: "{status}"', - 'source: "00_inbox/raw-clipping.md"', - 'tags: ["draft", "ai"]', - 'links: ["related-note"]', - "---", - "", - "# Enriched Note", - "", - "Body content.", - "", - ] - ) - - -class PrepIngestTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - for directory in ( - "01_ingest/queue", - "01_ingest/ready", - "10_knowledge/ai-systems", - "10_knowledge/productivity-systems", - "10_knowledge/robotics", - ): - (self.root / directory).mkdir(parents=True, exist_ok=True) - self.prep = PrepIngest(self.root) - - def tearDown(self) -> None: - self.tmp.cleanup() - - def _write_ready(self, name: str, content: str) -> Path: - path = self.root / "01_ingest" / "ready" / name - path.write_text(content, encoding="utf-8") - return path - - def test_promotes_valid_extracted_file_to_queue(self) -> None: - source = self._write_ready( - "2026-05-27__ai-systems__note__enriched-clipping.md", - enriched_note(), - ) - - result = self.prep.run(apply=True) - - target = self.root / "01_ingest" / "queue" / source.name - self.assertTrue(result.ok) - self.assertFalse(source.exists()) - self.assertTrue(target.exists()) - self.assertIn("promote", [event.kind for event in result.events]) - - def test_dry_run_does_not_move_files(self) -> None: - source = self._write_ready( - "2026-05-27__ai-systems__note__enriched-clipping.md", - enriched_note(), - ) - - result = self.prep.run(apply=False) - - target = self.root / "01_ingest" / "queue" / source.name - self.assertTrue(result.ok) - self.assertTrue(source.exists()) - self.assertFalse(target.exists()) - self.assertIn("promote", [event.kind for event in result.events]) - - def test_blocks_partial_frontmatter(self) -> None: - source = self._write_ready( - "2026-05-27__ai-systems__note__partial.md", - "\n".join( - [ - "---", - 'title: "Partial"', - 'domain: "ai-systems"', - 'type: "note"', - 'status: "extracted"', - "---", - "", - ] - ), - ) - - result = self.prep.run(apply=True) - - target = self.root / "01_ingest" / "queue" / source.name - self.assertFalse(result.ok) - self.assertTrue(source.exists()) - self.assertFalse(target.exists()) - self.assertEqual(result.events[0].kind, "blocked") - self.assertIn("frontmatter not ready", result.events[0].message) - - def test_skip_invalid_demotes_validation_block_to_warning(self) -> None: - source = self._write_ready( - "2026-05-27__ai-systems__note__partial.md", - "\n".join( - [ - "---", - 'title: "Partial"', - 'domain: "ai-systems"', - 'type: "note"', - 'status: "extracted"', - "---", - "", - ] - ), - ) - - result = self.prep.run(apply=False, skip_invalid=True) - - self.assertTrue(result.ok) - self.assertTrue(source.exists()) - self.assertEqual(result.events[0].severity, "warning") - - def test_root_option_is_preserved_for_cli_compatibility(self) -> None: - args = prep_module.build_parser().parse_args( - ["--root", str(self.root), "run", "--dry-run"] - ) - - self.assertEqual(args.root, str(self.root)) - - def test_blocks_non_extracted_status(self) -> None: - source = self._write_ready( - "2026-05-27__ai-systems__note__premature.md", - enriched_note(status="skimmed"), - ) - - result = self.prep.run(apply=True) - - self.assertFalse(result.ok) - self.assertTrue(source.exists()) - self.assertEqual(result.events[0].kind, "blocked") - self.assertIn("status must be 'extracted'", result.events[0].message) - - def test_blocks_malformed_filename(self) -> None: - source = self._write_ready( - "just-a-slug.md", - enriched_note(), - ) - - result = self.prep.run(apply=True) - - self.assertFalse(result.ok) - self.assertTrue(source.exists()) - self.assertEqual(result.events[0].kind, "blocked") - self.assertIn("filename must match", result.events[0].message) - - def test_blocks_filename_domain_mismatch(self) -> None: - source = self._write_ready( - "2026-05-27__robotics__note__wrong-domain.md", - enriched_note(domain="ai-systems"), - ) - - result = self.prep.run(apply=True) - - self.assertFalse(result.ok) - self.assertTrue(source.exists()) - self.assertIn("filename domain", result.events[0].message) - - def test_blocks_unknown_domain(self) -> None: - source = self._write_ready( - "2026-05-27__mythics__note__unknown.md", - enriched_note(domain="mythics"), - ) - - result = self.prep.run(apply=True) - - self.assertFalse(result.ok) - self.assertTrue(source.exists()) - self.assertIn("unknown knowledge domain", result.events[0].message) - - def test_blocks_destination_collision(self) -> None: - source = self._write_ready( - "2026-05-27__ai-systems__note__enriched-clipping.md", - enriched_note(), - ) - target = self.root / "01_ingest" / "queue" / source.name - target.write_text("preexisting\n", encoding="utf-8") - - result = self.prep.run(apply=True) - - self.assertFalse(result.ok) - self.assertTrue(source.exists()) - self.assertEqual(target.read_text(encoding="utf-8"), "preexisting\n") - self.assertEqual(result.events[0].kind, "blocked") - self.assertIn("queue destination already exists", result.events[0].message) - - def test_empty_ready_directory_is_ok(self) -> None: - result = self.prep.run(apply=True) - - self.assertTrue(result.ok) - self.assertEqual(result.events, []) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_process_eval_shim.py b/tests/test_process_eval_shim.py new file mode 100644 index 0000000..a11166c --- /dev/null +++ b/tests/test_process_eval_shim.py @@ -0,0 +1,32 @@ +"""Root shim for bin/process-eval delegates to the operation-owned engine.""" + +from __future__ import annotations + +import subprocess +import sys +import unittest +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] + + +class ProcessEvalShimTests(unittest.TestCase): + def test_help_runs(self) -> None: + result = subprocess.run( + [sys.executable, str(ROOT / "bin" / "process-eval"), "--help"], + cwd=ROOT, + capture_output=True, + text=True, + ) + self.assertEqual(result.returncode, 0, result.stderr) + self.assertIn("process-evaluation", result.stdout.lower() + result.stderr.lower()) + + def test_status_resolves_tracked_operation_authority(self) -> None: + result = subprocess.run( + [sys.executable, str(ROOT / "bin" / "process-eval"), "status", "--json"], + cwd=ROOT, + capture_output=True, + text=True, + ) + self.assertEqual(result.returncode, 0, result.stderr) + self.assertIn("40_operations/mainframe-process-eval", result.stdout) diff --git a/tests/test_public_mpe_synthetic.py b/tests/test_public_mpe_synthetic.py new file mode 100644 index 0000000..804f320 --- /dev/null +++ b/tests/test_public_mpe_synthetic.py @@ -0,0 +1,288 @@ +"""Synthetic MainFrame fixture: MPE/identity without private evidence.""" + +from __future__ import annotations + +import json +import shutil +import subprocess +import sys +import tempfile +import unittest +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] + + +def _write_record(path: Path, record_type: str, slug: str) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text( + "\n".join( + [ + "---", + f'title: "{slug}"', + f"type: {record_type}", + 'status: "active"', + "project_state: active", + f"record_type: {record_type}", + 'updated: "2026-09-10"', + 'source: "synthetic"', + "---", + "", + f"# {slug}", + "", + "Synthetic fixture. Not a private MainFrame record.", + "", + ] + ), + encoding="utf-8", + ) + + +class SyntheticMpeFeasibility(unittest.TestCase): + def _assert_output_not_exported(self, path: Path) -> None: + manifest_path = ROOT / ".context" / "public-export" / "manifest.json" + if manifest_path.is_file(): + manifest = json.loads(manifest_path.read_text(encoding="utf-8")) + exported = set(manifest.get("includes", [])) + exported.update(item["dest"] for item in manifest.get("transforms", [])) + self.assertNotIn(path.as_posix(), exported) + self.assertFalse(any("/outputs/" in item for item in exported)) + return + self.assertIn("outputs", path.parts) + shipped = ROOT / "40_operations" / "mainframe-process-eval" / "outputs" + self.assertFalse(shipped.exists()) + def test_identity_project_vs_operation_and_failures(self) -> None: + sys.path.insert(0, str(ROOT / "scripts")) + from lifecycle_identity import ( + DuplicateIdentity, + MissingIdentity, + resolve_record, + ) + + with tempfile.TemporaryDirectory() as td: + root = Path(td) + _write_record( + root / "30_projects" / "example-project" / "README.md", + "project", + "example-project", + ) + _write_record( + root / "40_operations" / "example-operation" / "README.md", + "operation", + "example-operation", + ) + project = resolve_record(root, "example-project") + operation = resolve_record(root, "example-operation") + self.assertEqual(project.record_type, "project") + self.assertEqual(operation.record_type, "operation") + with self.assertRaises(MissingIdentity): + resolve_record(root, "missing-slug") + _write_record( + root / "30_projects" / "example-operation" / "README.md", + "project", + "example-operation", + ) + with self.assertRaises(DuplicateIdentity): + resolve_record(root, "example-operation") + + def test_process_eval_status_on_synthetic_public_layout(self) -> None: + with tempfile.TemporaryDirectory() as td: + root = Path(td) + src = root / "40_operations" / "mainframe-process-eval" / "src" + src.mkdir(parents=True) + for name in ("process_eval.py", "runtime.py", "__init__.py"): + shutil.copy2( + ROOT / "40_operations" / "mainframe-process-eval" / "src" / name, + src / name, + ) + (root / "bin").mkdir(parents=True, exist_ok=True) + shutil.copy2(ROOT / "bin" / "process-eval", root / "bin" / "process-eval") + (root / "scripts").mkdir() + for name in ( + "lifecycle_identity.py", + "lifecycle_runtime.py", + "migration_lease.py", + ): + shutil.copy2(ROOT / "scripts" / name, root / "scripts" / name) + _write_record( + root / "40_operations" / "mainframe-process-eval" / "README.md", + "operation", + "mainframe-process-eval", + ) + workflow = root / ".context" / "workflows" + workflow.mkdir(parents=True) + (workflow / "process-evaluation.md").write_text( + "# synthetic process evaluation\n", encoding="utf-8" + ) + proc = subprocess.run( + [sys.executable, str(root / "bin" / "process-eval"), "status", "--json"], + cwd=root, + capture_output=True, + text=True, + check=False, + ) + self.assertEqual(proc.returncode, 0, proc.stderr) + payload = json.loads(proc.stdout) + self.assertEqual( + payload["project"], "40_operations/mainframe-process-eval" + ) + self.assertIsNone(payload["latest_output"]) + self.assertIsNone(payload["loop_eval"]) + + def test_real_loop_eval_scorer_on_synthetic_catalogue(self) -> None: + scorer = ( + ROOT + / "40_operations" + / "mainframe-process-eval" + / "workbench" + / "loop-eval" + / "scripts" + / "score_catalogue.py" + ) + catalogue = ( + ROOT + / "examples" + / "demo-mainframe" + / "40_operations" + / "demo-eval" + / "catalogue" + / "entries.json" + ) + self.assertTrue(catalogue.exists()) + self.assertIn("SYNTHETIC", catalogue.read_text(encoding="utf-8")) + + with tempfile.TemporaryDirectory() as td: + out = Path(td) / "outputs" / "synthetic-loop-eval.json" + proc = subprocess.run( + [ + sys.executable, + str(scorer), + "--catalogue", + str(catalogue), + "--json", + "--out", + str(out), + ], + cwd=ROOT, + capture_output=True, + text=True, + check=False, + ) + self.assertEqual(proc.returncode, 0, proc.stderr + proc.stdout) + payload = json.loads(out.read_text(encoding="utf-8")) + self.assertTrue(payload["ok"], payload) + scores = {row["id"]: row["score"] for row in payload["summary"]["scores"]} + self.assertEqual(scores["loop-synthetic-capture"], 8) + self.assertEqual(scores["loop-synthetic-thin"], 4) + self.assertEqual(payload["summary"]["n"], 2) + self.assertFalse( + payload["summary"]["by_tier_counts"]["well_defined"] == 0 + ) + + self._assert_output_not_exported(out) + + def test_loop_eval_scorer_is_not_a_noop(self) -> None: + scorer = ( + ROOT + / "40_operations" + / "mainframe-process-eval" + / "workbench" + / "loop-eval" + / "scripts" + / "score_catalogue.py" + ) + with tempfile.TemporaryDirectory() as td: + broken = Path(td) / "broken.json" + broken.write_text( + json.dumps( + { + "synthetic": True, + "entries": [ + { + "id": "loop-broken", + "title": "Broken", + "tier": "well_defined", + "kind": "loop", + "question": "x", + "unit_of_work": "x", + "status": "active", + "last_reviewed": "2026-09-10", + "checklist": {"unit_of_work": True}, + } + ], + } + ), + encoding="utf-8", + ) + proc = subprocess.run( + [sys.executable, str(scorer), "--catalogue", str(broken), "--json"], + cwd=ROOT, + capture_output=True, + text=True, + check=False, + ) + self.assertNotEqual(proc.returncode, 0, proc.stdout) + payload = json.loads(proc.stdout) + self.assertFalse(payload["ok"]) + self.assertTrue(payload["problems"]) + + def test_close_writes_local_only_synthetic_artifact(self) -> None: + with tempfile.TemporaryDirectory() as td: + root = Path(td) + src = root / "40_operations" / "mainframe-process-eval" / "src" + src.mkdir(parents=True) + for name in ("process_eval.py", "runtime.py", "__init__.py"): + shutil.copy2( + ROOT / "40_operations" / "mainframe-process-eval" / "src" / name, + src / name, + ) + (root / "bin").mkdir() + shutil.copy2(ROOT / "bin" / "process-eval", root / "bin" / "process-eval") + (root / "scripts").mkdir() + for name in ( + "lifecycle_identity.py", + "lifecycle_runtime.py", + "migration_lease.py", + ): + shutil.copy2(ROOT / "scripts" / name, root / "scripts" / name) + _write_record( + root / "40_operations" / "mainframe-process-eval" / "README.md", + "operation", + "mainframe-process-eval", + ) + wf = root / ".context" / "workflows" + wf.mkdir(parents=True) + (wf / "process-evaluation.md").write_text("# synthetic\n", encoding="utf-8") + proc = subprocess.run( + [ + sys.executable, + str(root / "bin" / "process-eval"), + "close", + "--run-id", + "2026-09-10-synthetic", + "--question", + "SYNTHETIC: does close write a local-only artifact?", + "--write", + ], + cwd=root, + capture_output=True, + text=True, + check=False, + ) + self.assertEqual(proc.returncode, 0, proc.stderr + proc.stdout) + written = ( + root + / "40_operations" + / "mainframe-process-eval" + / "outputs" + / "2026-09-10-synthetic-process-eval-close.md" + ) + self.assertTrue(written.is_file(), proc.stdout) + text = written.read_text(encoding="utf-8") + self.assertIn("SYNTHETIC", text) + self.assertIn("same_slice", text) + self._assert_output_not_exported(written) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_research_lane_loop.py b/tests/test_research_lane_loop.py deleted file mode 100644 index cd18ca1..0000000 --- a/tests/test_research_lane_loop.py +++ /dev/null @@ -1,274 +0,0 @@ -"""Tests for bin/research-lane-loop preflight classification and index audit.""" - -from __future__ import annotations - -import importlib.util -import sys -import tempfile -import unittest -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -SCRIPT_PATH = ROOT / "bin" / "research-lane-loop" -LOADER = SourceFileLoader("research_lane_loop", str(SCRIPT_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -mod = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = mod -SPEC.loader.exec_module(mod) - - -class ClassifyResumeTests(unittest.TestCase): - def test_class_a_fresh(self) -> None: - c, _ = mod.classify_resume([], [], [], []) - self.assertEqual(c, "A") - - def test_class_b_inbox(self) -> None: - c, reason = mod.classify_resume(["00_inbox/x.md"], [], [], []) - self.assertEqual(c, "B") - self.assertIn("inbox", reason.lower()) - - def test_class_c_raws_no_notes(self) -> None: - raws = [f"r{i}" for i in range(3)] - c, _ = mod.classify_resume([], raws, [], []) - self.assertEqual(c, "C") - - def test_class_d_stale_when_knowledge_exists(self) -> None: - stale = [ - mod.StaleRef( - tracker_path="00_inbox/foo.md", - basename="foo.md", - knowledge_path="10_knowledge/agents/foo.md", - still_in_inbox=False, - ) - ] - c, reason = mod.classify_resume([], ["10_knowledge/agents/foo.md"], [], stale) - self.assertEqual(c, "D") - self.assertIn("00_inbox", reason) - - def test_class_b_wins_over_empty_knowledge(self) -> None: - c, _ = mod.classify_resume(["00_inbox/x.md"], ["raw1"], ["note1"], []) - self.assertEqual(c, "B") - - def test_class_c_sparse_existing_corpus(self) -> None: - """1 raw + 1 note is not 'fresh' — avoid re-discovering from zero.""" - c, reason = mod.classify_resume([], ["raw1"], ["note1"], []) - self.assertEqual(c, "C") - self.assertIn("coverage", reason.lower()) - - def test_sparse_corpus_reason_carries_no_source_quota(self) -> None: - """The resume reason must never imply a target number of sources. - - Until 2026-08-09 this path said "gap-fill" and the loop emitted - "need ~3-4 raws" whenever a lane held fewer than three. An agent that - searched honestly and found one source had no compliant way to say so, - because no field meant "I looked and found nothing" — so it produced - three that looked right. 107 captures citing papers that do not exist - followed, across four months and fourteen domains. - - The guarantee is not a wording preference. It is that the loop asks a - coverage question rather than naming a count, so that recording a gap - stays an available and honest answer. Asserted on the shape of the - reason string rather than its exact text, so a rewrite that keeps the - guarantee does not fail and a quota that returns does. - """ - for raws, notes in (([], []), (["r1"], []), (["r1"], ["n1"]), - (["r1", "r2"], ["n1"])): - _, reason = mod.classify_resume([], raws, notes, []) - low = reason.lower() - self.assertNotRegex( - low, - r"(need|fill|at least|minimum|target|quota)\b[^.]{0,24}\b\d+", - f"resume reason reintroduces a source quota: {reason!r}", - ) - self.assertNotIn("gap-fill", low) - - -class StaleScanTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - self.knowledge = self.root / "10_knowledge" / "agents" - self.knowledge.mkdir(parents=True) - self.inbox = self.root / "00_inbox" - self.inbox.mkdir() - (self.knowledge / "2026-01-01__agents__raw__sample.md").write_text("# raw\n") - - self._orig_root = mod.ROOT - self._orig_know = mod.KNOWLEDGE - self._orig_inbox = mod.INBOX - mod.ROOT = self.root - mod.KNOWLEDGE = self.root / "10_knowledge" - mod.INBOX = self.inbox - - def tearDown(self) -> None: - mod.ROOT = self._orig_root - mod.KNOWLEDGE = self._orig_know - mod.INBOX = self._orig_inbox - self.tmp.cleanup() - - def test_detects_fixable_stale_path(self) -> None: - text = ( - "| 2026-01-01 | `00_inbox/2026-01-01__agents__raw__sample.md` | raw |\n" - ) - refs = mod.scan_stale_inbox_refs(text, "agents") - self.assertEqual(len(refs), 1) - self.assertEqual( - refs[0].knowledge_path, - "10_knowledge/agents/2026-01-01__agents__raw__sample.md", - ) - self.assertFalse(refs[0].still_in_inbox) - - def test_no_stale_when_only_knowledge_paths(self) -> None: - text = "`10_knowledge/agents/2026-01-01__agents__raw__sample.md`" - refs = mod.scan_stale_inbox_refs(text, "agents") - self.assertEqual(refs, []) - - -class CaptureTagsTests(unittest.TestCase): - def test_parse_yaml_list(self) -> None: - fm = { - "capture_tags": '["research-lane", "lane-mh01", "mindgraph"]', - } - tags = mod.parse_capture_tags(fm, "") - self.assertIn("lane-mh01", tags) - self.assertIn("research-lane", tags) - - def test_fallback_lane_tag_from_body(self) -> None: - tags = mod.parse_capture_tags({}, "see lane-c28 work") - self.assertIn("lane-c28", tags) - - def test_paths_from_capture_index(self) -> None: - text = ( - "| d | `10_knowledge/loop-engineering/2026-07-10__x__raw__y.md` |\n" - "| e | 10_knowledge/ai-business/z.md |\n" - ) - paths = mod.paths_from_capture_index(text) - self.assertIn( - "10_knowledge/loop-engineering/2026-07-10__x__raw__y.md", paths - ) - self.assertIn("10_knowledge/ai-business/z.md", paths) - - -class AuditRepairTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - self.lanes = self.root / "30_projects" / "research-lanes-strategy" / "lanes" - self.completed = self.lanes / "completed" - self.lanes.mkdir(parents=True) - self.completed.mkdir(parents=True) - self.knowledge = self.root / "10_knowledge" / "agents" - self.knowledge.mkdir(parents=True) - self.inbox = self.root / "00_inbox" - self.inbox.mkdir() - - base = "2026-01-01__agents__raw__widget.md" - (self.knowledge / base).write_text("# raw\n") - - lane = self.lanes / "widget-lane" - lane.mkdir() - self.readme = lane / "README.md" - self.readme.write_text( - f"""--- -title: "Lane: Widget" -status: "active" -lane_id: "W1" -knowledge_domain: "agents" -priority: "immediate" -capture_tags: ["research-lane", "lane-w1"] ---- - -# Widget - -| Date | Path | -|------|------| -| 2026-01-01 | `00_inbox/{base}` | -""", - encoding="utf-8", - ) - - # Point module paths + lane_intake dirs at fixture - self._paths = { - "ROOT": mod.ROOT, - "KNOWLEDGE": mod.KNOWLEDGE, - "INBOX": mod.INBOX, - "LANES_DIR": mod.LANES_DIR, - "COMPLETED_DIR": mod.COMPLETED_DIR, - "li_LANES": mod.lane_intake.LANES_DIR, - "li_COMPLETED": mod.lane_intake.COMPLETED_DIR, - } - mod.ROOT = self.root - mod.KNOWLEDGE = self.root / "10_knowledge" - mod.INBOX = self.inbox - mod.LANES_DIR = self.lanes - mod.COMPLETED_DIR = self.completed - mod.lane_intake.LANES_DIR = self.lanes - mod.lane_intake.COMPLETED_DIR = self.completed - - def tearDown(self) -> None: - mod.ROOT = self._paths["ROOT"] - mod.KNOWLEDGE = self._paths["KNOWLEDGE"] - mod.INBOX = self._paths["INBOX"] - mod.LANES_DIR = self._paths["LANES_DIR"] - mod.COMPLETED_DIR = self._paths["COMPLETED_DIR"] - mod.lane_intake.LANES_DIR = self._paths["li_LANES"] - mod.lane_intake.COMPLETED_DIR = self._paths["li_COMPLETED"] - self.tmp.cleanup() - - def test_audit_dry_run_finds_stale(self) -> None: - result = mod.audit_indexes("active", repair=False) - self.assertEqual(result["lanes_with_stale_refs"], 1) - self.assertEqual(result["repaired"], []) - self.assertIn("00_inbox/", self.readme.read_text()) - - def test_preflight_does_not_rewrite_stale_path(self) -> None: - mod.preflight_slug("widget-lane") - - self.assertIn("00_inbox/", self.readme.read_text()) - - def test_audit_repair_rewrites_path(self) -> None: - result = mod.audit_indexes("active", repair=True) - self.assertEqual(result["repaired"], ["widget-lane"]) - text = self.readme.read_text() - self.assertIn("10_knowledge/agents/2026-01-01__agents__raw__widget.md", text) - self.assertNotIn("00_inbox/2026-01-01__agents__raw__widget.md", text) - - -if __name__ == "__main__": - unittest.main() - - -class PhaseAndBlockerTests(unittest.TestCase): - def test_parse_phase_status_ignores_capture_index(self) -> None: - text = """ -## Captured knowledge - -| Date | Path | Type | -|------|------|------| -| 2026-07-13 | `foo.md` | raw | - -## Phase status - -| Phase | Status | Notes | -|-------|--------|-------| -| Foundation | ✅ done | x | -| Taxonomy | ✅ done | y | -| Application | next | z | -""" - ps = mod.parse_phase_status(text) - self.assertEqual(ps.get("Foundation"), "done") - self.assertEqual(ps.get("Taxonomy"), "done") - self.assertEqual(ps.get("Application"), "next") - self.assertNotIn("Date", ps) - - def test_infer_next_phase_and_blocker(self) -> None: - ps = {"Foundation": "done", "Taxonomy": "done", "Application": "next"} - nxt = mod.infer_next_phase(ps) - self.assertEqual(nxt, "Application") - blocker = mod.infer_blocker( - "C", nxt, "10_knowledge/x/, 30_projects/local-agent/", raw_count=5, note_count=2 - ) - self.assertEqual(blocker, "project") diff --git a/tests/test_second_brain_batch.py b/tests/test_second_brain_batch.py deleted file mode 100644 index 4b7d4f2..0000000 --- a/tests/test_second_brain_batch.py +++ /dev/null @@ -1,291 +0,0 @@ -import csv -import hashlib -import importlib.util -import importlib.machinery -import subprocess -import tempfile -import unittest -from pathlib import Path - - -SCRIPT = Path(__file__).resolve().parents[1] / "bin" / "second-brain-batch" - - -class SecondBrainBatchTests(unittest.TestCase): - def setUp(self): - self.temp_dir = tempfile.TemporaryDirectory() - self.root = Path(self.temp_dir.name) - (self.root / "bin").mkdir() - register = ( - self.root - / "30_projects" - / "second-brain-migration" - / "raw-materials" - / "batch-register.md" - ) - register.parent.mkdir(parents=True) - register.write_text( - "# Batch Register\n\n" - "| Batch | Registered | Source | Files | Status | Manifest |\n" - "|---|---|---|---:|---|---|\n", - encoding="utf-8", - ) - script_text = SCRIPT.read_text(encoding="utf-8").replace( - "ROOT = Path(__file__).resolve().parents[1]", - f"ROOT = Path({str(self.root)!r})", - ) - self.script = self.root / "bin" / "second-brain-batch" - self.script.write_text(script_text, encoding="utf-8") - self.script.chmod(0o755) - loader = importlib.machinery.SourceFileLoader( - f"second_brain_batch_{id(self)}", str(self.script) - ) - spec = importlib.util.spec_from_loader( - f"second_brain_batch_{id(self)}", loader - ) - assert spec and spec.loader - self.module = importlib.util.module_from_spec(spec) - spec.loader.exec_module(self.module) - - def tearDown(self): - self.temp_dir.cleanup() - - def run_script(self, *args): - return subprocess.run( - [str(self.script), *args], - text=True, - capture_output=True, - check=False, - ) - - def hash_file(self, path: Path) -> str: - digest = hashlib.sha256() - digest.update(path.read_bytes()) - return digest.hexdigest() - - def test_registers_every_file_copies_snapshot_and_initializes_ledger(self): - inbox = self.root / "drop" - (inbox / "nested").mkdir(parents=True) - (inbox / "cc-state.md").write_text( - "Last Updated: 2026-04-01\n", encoding="utf-8" - ) - (inbox / "copy.md").write_text("same\n", encoding="utf-8") - (inbox / "nested" / "copy.md").write_text("same\n", encoding="utf-8") - - result = self.run_script( - "--batch-id", - "2026-06-05-001", - "--source", - str(inbox), - "--source-label", - "old-vault export", - ) - - self.assertEqual(result.returncode, 0, result.stderr) - batch = ( - self.root - / "30_projects" - / "second-brain-migration" - / "raw-materials" - / "batches" - / "2026-06-05-001" - ) - with (batch / "manifest.csv").open(encoding="utf-8") as handle: - rows = list(csv.DictReader(handle)) - with (batch / "disposition-ledger.csv").open(encoding="utf-8") as handle: - ledger_rows = list(csv.DictReader(handle)) - self.assertEqual(len(rows), 3) - self.assertEqual(len(ledger_rows), 3) - duplicate_rows = [row for row in rows if row["duplicate_group"]] - self.assertEqual(len(duplicate_rows), 2) - self.assertEqual( - duplicate_rows[0]["duplicate_group"], - duplicate_rows[1]["duplicate_group"], - ) - self.assertEqual(rows[0]["review_status"], "unreviewed") - self.assertTrue( - (batch / "source-files" / "cc-state.md").exists(), - "registered batch should preserve a recoverable source snapshot", - ) - self.assertTrue( - (batch / "source-files" / "nested" / "copy.md").exists(), - "nested files should keep their relative paths inside the snapshot", - ) - self.assertTrue( - all(row["disposition"] == "unresolved" for row in ledger_rows), - "registration should initialize every file with an unresolved ledger row", - ) - self.assertEqual( - ledger_rows[0]["evidence"], - "raw-materials/batches/2026-06-05-001/manifest.csv", - ) - - def test_snapshot_recovers_original_bytes_after_source_changes(self): - inbox = self.root / "drop" - inbox.mkdir() - source = inbox / "note.md" - source.write_text("original\n", encoding="utf-8") - - result = self.run_script( - "--batch-id", - "2026-06-05-001", - "--source", - str(inbox), - ) - - self.assertEqual(result.returncode, 0, result.stderr) - batch = ( - self.root - / "30_projects" - / "second-brain-migration" - / "raw-materials" - / "batches" - / "2026-06-05-001" - ) - snapshot = batch / "source-files" / "note.md" - with (batch / "manifest.csv").open(encoding="utf-8") as handle: - manifest_row = next(csv.DictReader(handle)) - - source.write_text("changed later\n", encoding="utf-8") - - self.assertEqual(snapshot.read_text(encoding="utf-8"), "original\n") - self.assertEqual(self.hash_file(snapshot), manifest_row["sha256"]) - self.assertNotEqual(self.hash_file(source), manifest_row["sha256"]) - - def test_registration_does_not_mutate_source_files(self): - inbox = self.root / "drop" - inbox.mkdir() - source = inbox / "note.md" - original = "unchanged body\n" - source.write_text(original, encoding="utf-8") - - result = self.run_script( - "--batch-id", - "2026-06-05-001", - "--source", - str(inbox), - ) - - self.assertEqual(result.returncode, 0, result.stderr) - self.assertEqual(source.read_text(encoding="utf-8"), original) - - def test_ledger_append_keeps_prior_rows(self): - inbox = self.root / "drop" - inbox.mkdir() - (inbox / "note.md").write_text("note\n", encoding="utf-8") - - batch = self.module.register_batch( - "2026-06-05-001", inbox.resolve(), str(inbox.resolve()) - ) - ledger = batch / "disposition-ledger.csv" - self.module.append_disposition_entries( - ledger, - [ - { - "recorded_at": "2026-06-06T12:00:00-04:00", - "batch_id": "2026-06-05-001", - "relative_path": "note.md", - "sha256": self.hash_file(batch / "source-files" / "note.md"), - "disposition": "promoted", - "destination": "10_knowledge/finance/example.md", - "evidence": "manual verification", - "notes": "test append", - } - ], - ) - - with ledger.open(encoding="utf-8") as handle: - rows = list(csv.DictReader(handle)) - self.assertEqual(len(rows), 2) - self.assertEqual(rows[0]["disposition"], "unresolved") - self.assertEqual(rows[1]["disposition"], "promoted") - - def test_refuses_to_overwrite_existing_batch(self): - inbox = self.root / "drop" - inbox.mkdir() - (inbox / "note.md").write_text("note\n", encoding="utf-8") - args = ("--batch-id", "2026-06-05-001", "--source", str(inbox)) - - self.assertEqual(self.run_script(*args).returncode, 0) - second = self.run_script(*args) - - self.assertNotEqual(second.returncode, 0) - self.assertIn("already registered", second.stderr) - - def test_registered_batch_is_rejected_before_creating_output(self): - inbox = self.root / "drop" - inbox.mkdir() - register = ( - self.root - / "30_projects" - / "second-brain-migration" - / "raw-materials" - / "batch-register.md" - ) - register.write_text( - register.read_text(encoding="utf-8") - + "| `2026-06-05-001` | 2026-06-05 | `prior` | 1 | registered | `prior.csv` |\n", - encoding="utf-8", - ) - - result = self.run_script( - "--batch-id", - "2026-06-05-001", - "--source", - str(inbox), - ) - - batch = ( - self.root - / "30_projects" - / "second-brain-migration" - / "raw-materials" - / "batches" - / "2026-06-05-001" - ) - self.assertNotEqual(result.returncode, 0) - self.assertIn("already registered", result.stderr) - self.assertFalse(batch.exists()) - - def test_rejects_preexisting_batch_directory(self): - inbox = self.root / "drop" - inbox.mkdir() - (inbox / "note.md").write_text("note\n", encoding="utf-8") - batch = ( - self.root - / "30_projects" - / "second-brain-migration" - / "raw-materials" - / "batches" - / "2026-06-05-001" - ) - batch.mkdir(parents=True) - - result = self.run_script( - "--batch-id", - "2026-06-05-001", - "--source", - str(inbox), - ) - - self.assertNotEqual(result.returncode, 0) - self.assertIn("already exists", result.stderr) - - def test_rejects_batch_id_that_can_escape_batch_directory(self): - inbox = self.root / "drop" - inbox.mkdir() - - result = self.run_script( - "--batch-id", - "../../outside", - "--source", - str(inbox), - ) - - self.assertNotEqual(result.returncode, 0) - self.assertIn("YYYY-MM-DD-NNN", result.stderr) - self.assertFalse((self.root / "outside").exists()) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_session_close.py b/tests/test_session_close.py deleted file mode 100644 index 68ebfdc..0000000 --- a/tests/test_session_close.py +++ /dev/null @@ -1,336 +0,0 @@ -from __future__ import annotations - -import hashlib -import importlib.util -import json -import os -import subprocess -import sys -import tempfile -import unittest -from datetime import date, datetime, timedelta, timezone -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -SCRIPT_PATH = ROOT / "bin" / "session-close" -LOADER = SourceFileLoader("session_close", str(SCRIPT_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -mod = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = mod -SPEC.loader.exec_module(mod) - -check_session = mod.check_session - - -def make_runner(changed_files: list[str] | None = None, - fail_commands: set[str] | None = None): - changed = changed_files or [] - failures = fail_commands or set() - - def run_command(cmd: list[str]) -> subprocess.CompletedProcess[str]: - cmd_str = " ".join(cmd) - if any(f in cmd_str for f in failures): - return subprocess.CompletedProcess(cmd, 1, "", "error") - - if "git" in cmd and "diff" in cmd and "--name-only" in cmd: - return subprocess.CompletedProcess(cmd, 0, "\n".join(changed), "") - if "git" in cmd and "status" in cmd: - porcelain = "\n".join(f" M {f}" for f in changed) - return subprocess.CompletedProcess(cmd, 0, porcelain, "") - - if "sync-project-index" in cmd_str and "--check" in cmd_str: - project_meta_changed = any( - f.startswith("30_projects/") and f.endswith("/README.md") - for f in changed - ) - return subprocess.CompletedProcess(cmd, 1 if project_meta_changed else 0, "", "") - - if "eval-schedule" in cmd_str and "check" in cmd_str: - return subprocess.CompletedProcess(cmd, 0, "eval-schedule check: OK\n", "") - - return subprocess.CompletedProcess(cmd, 0, "", "") - - return run_command - - -STATE_MD = """\ -# STATE - -**Status**: Active - -## Active Project -test-project - -## Next Actions -- things -""" - - -class SessionCloseTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - (self.root / "STATE.md").write_text(STATE_MD, encoding="utf-8") - (self.root / "bin").mkdir() - for script in ("sync-project-index", "mindgraph-refresh", "workflow-report"): - (self.root / "bin" / script).write_text("#!/bin/sh\n", encoding="utf-8") - - def tearDown(self) -> None: - self.tmp.cleanup() - - def test_clean_state_no_changes(self) -> None: - runner = make_runner([]) - result = check_session(self.root, run_command=runner) - self.assertTrue(result.ok) - auto_needed = [a for a in result.actions - if a.kind == "auto" and a.needed and a.name != "handoff-digest"] - self.assertEqual(len(auto_needed), 0) - - def test_detects_project_metadata_change(self) -> None: - runner = make_runner(["30_projects/foo/README.md"]) - result = check_session(self.root, run_command=runner) - sync = [a for a in result.actions if a.name == "sync-project-index"][0] - self.assertTrue(sync.needed) - self.assertEqual(sync.kind, "auto") - - def test_detects_knowledge_change(self) -> None: - runner = make_runner(["10_knowledge/ai-systems/note.md"]) - result = check_session(self.root, run_command=runner) - refresh = [a for a in result.actions if a.name == "mindgraph-refresh"][0] - self.assertTrue(refresh.needed) - self.assertEqual(refresh.kind, "auto") - - def test_no_knowledge_change_not_needed(self) -> None: - runner = make_runner(["AGENTS.md"]) - result = check_session(self.root, run_command=runner) - refresh = [a for a in result.actions if a.name == "mindgraph-refresh"][0] - self.assertFalse(refresh.needed) - - def test_missing_state_md(self) -> None: - (self.root / "STATE.md").unlink() - runner = make_runner([]) - result = check_session(self.root, run_command=runner) - warn = [a for a in result.actions if a.name == "state-md"][0] - self.assertEqual(warn.kind, "warn") - self.assertTrue(warn.needed) - - def test_telemetry_available(self) -> None: - events_dir = self.root / "20_live" / "workflow-metrics" / "events" - events_dir.mkdir(parents=True) - today = date.today().isoformat() - (events_dir / f"{today}.jsonl").write_text("{}\n", encoding="utf-8") - - runner = make_runner([]) - result = check_session(self.root, run_command=runner) - report = [a for a in result.actions if a.name == "workflow-report"][0] - self.assertTrue(report.needed) - self.assertEqual(report.kind, "auto") - - def test_no_telemetry(self) -> None: - runner = make_runner([]) - result = check_session(self.root, run_command=runner) - report = [a for a in result.actions if a.name == "workflow-report"][0] - self.assertFalse(report.needed) - - def test_eval_schedule_check_always_considered(self) -> None: - runner = make_runner([]) - result = check_session(self.root, run_command=runner) - names = [a.name for a in result.actions] - self.assertIn("eval-schedule", names) - ev = [a for a in result.actions if a.name == "eval-schedule"][0] - self.assertEqual(ev.kind, "warn") - - def test_apply_runs_sync_project_index(self) -> None: - runner = make_runner(["30_projects/foo/README.md"]) - result = check_session(self.root, run_command=runner, apply=True) - sync = [a for a in result.actions if a.name == "sync-project-index"][0] - self.assertTrue(sync.ran) - self.assertTrue(sync.success) - - def test_apply_runs_mindgraph_refresh(self) -> None: - runner = make_runner(["10_knowledge/ai-systems/note.md"]) - result = check_session(self.root, run_command=runner, apply=True) - refresh = [a for a in result.actions if a.name == "mindgraph-refresh"][0] - self.assertTrue(refresh.ran) - self.assertTrue(refresh.success) - - def test_apply_reports_failure(self) -> None: - runner = make_runner( - ["30_projects/foo/README.md"], - fail_commands={"sync-project-index"}, - ) - result = check_session(self.root, run_command=runner, apply=True) - sync = [a for a in result.actions if a.name == "sync-project-index"][0] - self.assertTrue(sync.ran) - self.assertFalse(sync.success) - self.assertFalse(result.ok) - - def test_manual_reminders_always_present(self) -> None: - runner = make_runner([]) - result = check_session(self.root, run_command=runner) - manual = [a for a in result.actions if a.kind == "manual"] - names = {a.name for a in manual} - self.assertIn("state-md-narrative", names) - self.assertIn("decisions-review", names) - - def test_dirty_tree_warning(self) -> None: - runner = make_runner(["AGENTS.md", "STATE.md", "README.md"]) - result = check_session(self.root, run_command=runner) - warn = [a for a in result.actions if a.name == "working-tree"][0] - self.assertEqual(warn.kind, "warn") - self.assertIn("changed file(s)", warn.reason) - - def test_check_not_needed_returns_ok(self) -> None: - runner = make_runner([]) - result = check_session(self.root, run_command=runner) - self.assertTrue(result.ok) - auto_pending = [a for a in result.actions - if a.kind == "auto" and a.needed and not a.ran and a.name != "handoff-digest"] - self.assertEqual(len(auto_pending), 0) - - -class CheckpointTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.root = Path(self.tmp.name) - (self.root / "STATE.md").write_text(STATE_MD, encoding="utf-8") - (self.root / "20_live").mkdir() - projects = self.root / "30_projects" - (projects / "foo").mkdir(parents=True) - (projects / "foo" / "notes.md").write_text("recent\n", encoding="utf-8") - (projects / "old-proj").mkdir() - old_file = projects / "old-proj" / "stale.md" - old_file.write_text("old\n", encoding="utf-8") - old_ts = (datetime.now() - timedelta(days=30)).timestamp() - os.utime(old_file, (old_ts, old_ts)) - os.utime(projects / "old-proj", (old_ts, old_ts)) - - def tearDown(self) -> None: - self.tmp.cleanup() - - def draft_path(self) -> Path: - return self.root / "20_live" / "last-handoff-draft.md" - - def feed_path(self) -> Path: - return self.root / "20_live" / "workstation" / "session-close-feed.jsonl" - - def test_checkpoint_creates_draft_and_feed(self) -> None: - hook = {"session_hash": "abc123", "hook_event_name": "PreCompact", - "trigger": "auto"} - record, draft = mod.run_checkpoint( - self.root, run_command=make_runner([]), hook=hook) - text = draft.read_text(encoding="utf-8") - self.assertIn(mod.CHECKPOINT_HEADING, text) - self.assertIn("foo (files", text) - self.assertIn("PreCompact (auto)", text) - self.assertEqual(record["kind"], "checkpoint") - lines = self.feed_path().read_text(encoding="utf-8").strip().splitlines() - self.assertEqual(len(lines), 1) - feed_record = json.loads(lines[0]) - self.assertEqual(feed_record["kind"], "checkpoint") - self.assertEqual(feed_record["session_hash"], "abc123") - self.assertEqual(feed_record["source"], "PreCompact") - - def test_derive_orders_and_windows(self) -> None: - derived = mod.derive_active_projects(self.root, make_runner([])) - self.assertEqual([d["project"] for d in derived], ["foo"]) - - def test_drift_flag(self) -> None: - record = mod.build_checkpoint(self.root, make_runner([])) - self.assertTrue(record["drift"]) - match_dir = self.root / "30_projects" / "test-project" - match_dir.mkdir() - (match_dir / "work.md").write_text("hot\n", encoding="utf-8") - record = mod.build_checkpoint(self.root, make_runner([])) - self.assertFalse(record["drift"]) - - def test_snapshot_cap(self) -> None: - for _ in range(mod.SNAPSHOT_CAP + 2): - mod.run_checkpoint(self.root, run_command=make_runner([])) - text = self.draft_path().read_text(encoding="utf-8") - self.assertEqual(text.count("\n### "), mod.SNAPSHOT_CAP) - self.assertEqual(text.count("\nDerived active: "), 1) - - def test_checkpoint_preserves_digest_scaffold(self) -> None: - self.draft_path().write_text( - "# Session Handoff Draft (auto-generated)\n\n" - "Ingest status: XYZ-sentinel\n", - encoding="utf-8", - ) - mod.run_checkpoint(self.root, run_command=make_runner([])) - text = self.draft_path().read_text(encoding="utf-8") - self.assertIn("Ingest status: XYZ-sentinel", text) - self.assertLess(text.index("XYZ-sentinel"), - text.index(mod.CHECKPOINT_HEADING)) - - def test_digest_rewrite_preserves_checkpoints(self) -> None: - (self.root / "bin").mkdir() - mod.run_checkpoint(self.root, run_command=make_runner([])) - result = check_session(self.root, run_command=make_runner([]), apply=True) - digest = [a for a in result.actions if a.name == "handoff-digest"][0] - self.assertTrue(digest.success) - text = self.draft_path().read_text(encoding="utf-8") - self.assertIn(mod.CHECKPOINT_HEADING, text) - self.assertEqual(text.count("\n### "), 1) - self.assertIn("## What changed / remains / blockers / next", text) - - def test_check_feed_record_shape(self) -> None: - result = check_session(self.root, run_command=make_runner(["AGENTS.md"])) - hook = {"session_hash": "xyz", "hook_event_name": "SessionEnd", - "reason": "logout"} - record = mod.check_feed_record(result, hook, "close-check") - self.assertEqual(record["kind"], "close-check") - self.assertEqual(record["source"], "SessionEnd") - self.assertEqual(record["reason"], "logout") - self.assertIn("working-tree", record["warnings"]) - self.assertIn("handoff-digest", record["pending_auto"]) - self.assertFalse(record["ok"]) - names = [a["name"] for a in record["actions"]] - self.assertIn("state-md-narrative", names) - feed = mod.append_feed(self.root, record) - parsed = json.loads(feed.read_text(encoding="utf-8").strip()) - self.assertEqual(parsed["kind"], "close-check") - - def test_parse_hook_stdin(self) -> None: - payload = json.dumps({ - "session_id": "abc", - "hook_event_name": "PreCompact", - "trigger": "auto", - "prompt": "verbatim content must never be copied", - }) - hook = mod.parse_hook_stdin(payload) - expected = hashlib.sha256(b"abc").hexdigest()[:16] - self.assertEqual(hook["session_hash"], expected) - self.assertEqual(hook["hook_event_name"], "PreCompact") - self.assertEqual(hook["trigger"], "auto") - self.assertNotIn("prompt", hook) - self.assertNotIn("session_id", hook) - self.assertEqual(mod.parse_hook_stdin("not json"), {}) - self.assertEqual(mod.parse_hook_stdin(""), {}) - - def test_eval_weekly_status_staleness(self) -> None: - runs_path = self.root / "20_live" / "eval-registry" / "schedule-runs.jsonl" - runs_path.parent.mkdir(parents=True) - stale_ts = (datetime.now(timezone.utc) - timedelta(days=10)).isoformat() - runs_path.write_text(json.dumps({ - "run_id": "old-weekly", "cadence": "weekly", - "finished_at": stale_ts, "all_passed": True, - }) + "\n", encoding="utf-8") - status = mod.eval_weekly_status(self.root) - self.assertTrue(status["stale"]) - fresh_ts = datetime.now(timezone.utc).isoformat() - with runs_path.open("a", encoding="utf-8") as handle: - handle.write(json.dumps({ - "run_id": "fresh-weekly", "cadence": "weekly", - "finished_at": fresh_ts, "all_passed": True, - }) + "\n") - status = mod.eval_weekly_status(self.root) - self.assertFalse(status["stale"]) - self.assertEqual(status["run_id"], "fresh-weekly") - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_session_open.py b/tests/test_session_open.py index 0952a96..8eb8fb1 100644 --- a/tests/test_session_open.py +++ b/tests/test_session_open.py @@ -74,7 +74,14 @@ def setUp(self) -> None: self.tmp = tempfile.TemporaryDirectory() self.root = Path(self.tmp.name) (self.root / "AGENTS.md").write_text("# Agents\n", encoding="utf-8") + (self.root / "HARNESS.md").write_text("# Harness\n", encoding="utf-8") (self.root / "STATE.md").write_text(STATE_WITHOUT_PROJECT, encoding="utf-8") + (self.root / "30_projects").mkdir() + (self.root / "30_projects/AGENTS.md").write_text("# Projects\n", encoding="utf-8") + workflow = self.root / ".context/workflows/project-resume-and-candidate-lifecycle.md" + workflow.parent.mkdir(parents=True) + workflow.write_text("# Project reconstruction\n", encoding="utf-8") + (workflow.parent / "session-open.md").write_text("# Session route\n", encoding="utf-8") def tearDown(self) -> None: self.tmp.cleanup() @@ -84,9 +91,10 @@ def test_minimal_open_no_project(self) -> None: self.assertTrue(result.ok) self.assertIsNone(result.project) real_entries = [e for e in result.entries if not e.path.startswith("(")] - self.assertEqual(len(real_entries), 2) + self.assertEqual(len(real_entries), 4) self.assertEqual(real_entries[0].path, "AGENTS.md") - self.assertEqual(real_entries[1].path, "STATE.md") + self.assertEqual(real_entries[1].path, "HARNESS.md") + self.assertEqual(real_entries[2].path, "STATE.md") def test_open_with_project_flag(self) -> None: proj_dir = self.root / "30_projects" / "foo" @@ -106,7 +114,7 @@ def test_auto_detect_project_from_state(self) -> None: proj_dir.mkdir(parents=True) (proj_dir / "README.md").write_text("# Test\n", encoding="utf-8") - result = build_context(self.root) + result = build_context(self.root, intent="resume") self.assertEqual(result.project, "test-project") self.assertEqual(result.project_source, "state") paths = [e.path for e in result.entries] @@ -132,6 +140,80 @@ def test_missing_state_is_not_ok(self) -> None: result = build_context(self.root) self.assertFalse(result.ok) + def test_missing_harness_cannot_report_valid_context(self) -> None: + (self.root / "HARNESS.md").unlink() + result = build_context(self.root) + self.assertFalse(result.ok) + self.assertTrue(result.degraded) + self.assertEqual(result.missing_required, ["HARNESS.md"]) + + def test_missing_project_contracts_cannot_report_valid_context(self) -> None: + project = self.root / "30_projects/foo" + project.mkdir() + (project / "README.md").write_text("# Foo\n", encoding="utf-8") + (self.root / "30_projects/AGENTS.md").unlink() + (self.root / ".context/workflows/project-resume-and-candidate-lifecycle.md").unlink() + result = build_context(self.root, project="foo") + self.assertFalse(result.ok) + self.assertIn("30_projects/AGENTS.md", result.missing_required) + self.assertIn(".context/workflows/project-resume-and-candidate-lifecycle.md", + result.missing_required) + + def test_task_path_loads_only_its_ancestor_contracts_before_project_content(self) -> None: + project = self.root / "30_projects/foo" + target = project / "workbench/src/module.py" + target.parent.mkdir(parents=True) + target.write_text("pass\n", encoding="utf-8") + (project / "README.md").write_text("# Public readme\n", encoding="utf-8") + (project / "PROJECT.md").write_text("project_state: paused\n", encoding="utf-8") + for directory in (project, project / "workbench", project / "raw-materials", + self.root / "30_projects/unrelated"): + directory.mkdir(exist_ok=True) + (directory / "AGENTS.md").write_text("# Local rule\n", encoding="utf-8") + + result = build_context(self.root, project="foo", task_path=str(target.relative_to(self.root))) + self.assertTrue(result.ok) + paths = [e.path for e in result.entries] + expected = ["AGENTS.md", "HARNESS.md", "30_projects/AGENTS.md", + "30_projects/foo/AGENTS.md", "30_projects/foo/workbench/AGENTS.md", + "30_projects/foo/README.md", "30_projects/foo/PROJECT.md"] + self.assertEqual([p for p in paths if p in expected], expected) + self.assertNotIn("30_projects/foo/raw-materials/AGENTS.md", paths) + self.assertNotIn("30_projects/unrelated/AGENTS.md", paths) + + project_only = build_context(self.root, project="foo") + self.assertNotIn("30_projects/foo/workbench/AGENTS.md", + [e.path for e in project_only.entries]) + + def test_invalid_task_paths_fail_without_loading_other_contracts(self) -> None: + project = self.root / "30_projects/foo" + project.mkdir() + (project / "README.md").write_text("# Foo\n", encoding="utf-8") + outside = self.root / "private-other" + outside.mkdir() + (outside / "AGENTS.md").write_text("# Other contract\n", encoding="utf-8") + (project / "outside-link").symlink_to(outside, target_is_directory=True) + for path in ("private-other", "30_projects/foo/missing", "30_projects/foo/outside-link"): + with self.subTest(path=path): + result = build_context(self.root, project="foo", task_path=path) + self.assertFalse(result.ok) + self.assertTrue(result.path_error) + self.assertNotIn("private-other/AGENTS.md", [e.path for e in result.entries]) + + def test_task_path_requires_explicit_project(self) -> None: + result = build_context(self.root, task_path=".") + self.assertFalse(result.ok) + self.assertIn("explicit --project", result.path_error) + + def test_directory_named_contract_is_a_missing_required_file(self) -> None: + project = self.root / "30_projects/foo" + project.mkdir() + (project / "README.md").write_text("# Foo\n", encoding="utf-8") + (project / "AGENTS.md").mkdir() + result = build_context(self.root, project="foo") + self.assertFalse(result.ok) + self.assertIn("30_projects/foo/AGENTS.md", result.missing_required) + def test_phase_plan_discovery(self) -> None: (self.root / "STATE.md").write_text(STATE_WITH_PROJECT, encoding="utf-8") proj_dir = self.root / "30_projects" / "test-project" @@ -143,11 +225,11 @@ def test_phase_plan_discovery(self) -> None: encoding="utf-8", ) - result = build_context(self.root) + result = build_context(self.root, intent="resume", task="first") paths = [e.path for e in result.entries] self.assertTrue(any("01-first.md" in p for p in paths)) plan_entry = [e for e in result.entries if "01-first.md" in e.path][0] - self.assertEqual(plan_entry.note, "active phase plan") + self.assertIn("candidate", plan_entry.note) def test_skips_non_active_phase_plan(self) -> None: (self.root / "STATE.md").write_text(STATE_WITH_PROJECT, encoding="utf-8") @@ -160,7 +242,7 @@ def test_skips_non_active_phase_plan(self) -> None: encoding="utf-8", ) - result = build_context(self.root) + result = build_context(self.root, intent="resume") paths = [e.path for e in result.entries] self.assertFalse(any("01-done.md" in p for p in paths)) @@ -208,7 +290,7 @@ def test_compound_state_label_not_ok(self) -> None: "agent-tracker-eval (Focus Board) + mainframe-process-eval (canaries)\n", encoding="utf-8", ) - result = build_context(self.root) + result = build_context(self.root, intent="resume") self.assertFalse(result.ok) self.assertIsNotNone(result.project_error) self.assertTrue(result.degraded) @@ -223,7 +305,7 @@ def test_missing_project_readme_not_ok(self) -> None: "# S\n\n## Active Project\n\nmissing-slug\n", encoding="utf-8", ) - result = build_context(self.root) + result = build_context(self.root, intent="resume") self.assertFalse(result.ok) self.assertEqual(result.project, "missing-slug") self.assertFalse(result.project_exists) diff --git a/tests/test_session_open_routes.py b/tests/test_session_open_routes.py new file mode 100644 index 0000000..0f03774 --- /dev/null +++ b/tests/test_session_open_routes.py @@ -0,0 +1,184 @@ +"""Exercise real context files, bounded delivery and disconnected CLI behavior.""" + +from __future__ import annotations + +import importlib.util +import json +import shutil +import subprocess +import sys +from importlib.machinery import SourceFileLoader +from pathlib import Path + +import pytest + + +ROOT = Path(__file__).resolve().parents[1] +loader = SourceFileLoader("session_open_routes", str(ROOT / "bin/session-open")) +spec = importlib.util.spec_from_loader(loader.name, loader) +mod = importlib.util.module_from_spec(spec) +sys.modules[spec.name] = mod +spec.loader.exec_module(mod) + + +@pytest.fixture +def workspace(tmp_path): + files = { + "AGENTS.md": "# Rules\nPreserve local work.\n", + "HARNESS.md": "# Harness\n\n## Session orientation\nArrival rules.\n\n## Execution\nFull action rules.\n", + "STATE.md": "# State\n\n## Active Project\nexample\n", + ".context/workflows/session-open.md": "# Read the routed context completely.\n", + ".context/workflows/project-resume-and-candidate-lifecycle.md": "# Reconstruct source and Git.\n", + "30_projects/AGENTS.md": "# Project lifecycle rules\n", + "30_projects/example/README.md": "# Example\n", + "30_projects/example/plans/01-older.md": '---\ntitle: "Brainstorming"\nstatus: active\n---\nOld unrelated plan.\n', + "30_projects/example/plans/99-onboarding.md": '---\ntitle: "Onboarding"\nstatus: active\n---\nCurrent relevant plan.\n', + } + for relative, content in files.items(): + path = tmp_path / relative + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(content) + return tmp_path + + +def content_for(result, root, path): + return "".join(mod.read_batch(result, root, b["id"])["content"] + for b in result.read_batches if b["path"] == path) + + +def test_arrival_does_not_enter_recorded_project_or_require_its_contracts(workspace): + (workspace / "30_projects/AGENTS.md").unlink() + result = mod.build_context(workspace) + assert result.ok + assert result.intent == "arrival" + assert result.project is None + assert not any(e.path.startswith("30_projects/") for e in result.entries) + assert "Arrival rules." in content_for(result, workspace, "HARNESS.md") + assert "Full action rules." not in content_for(result, workspace, "HARNESS.md") + assert result.context_status == "unread" + + +def test_named_resume_retains_full_contracts_and_nominates_task_plan(workspace): + result = mod.build_context(workspace, project="example", task="pick up onboarding work") + assert result.ok + assert result.intent == "resume" + assert "Full action rules." in content_for(result, workspace, "HARNESS.md") + assert result.plan.endswith("99-onboarding.md") + assert result.plan_status == "task_match_candidate" + assert any(e.path == "30_projects/AGENTS.md" and e.required for e in result.entries) + + +def test_tied_or_missing_task_match_does_not_choose_alphabetical_plan(workspace): + plans = workspace / "30_projects/example/plans" + (plans / "98-onboarding.md").write_text('---\ntitle: "Onboarding two"\nstatus: active\n---\n') + for task in (None, "onboarding", "unmatched request"): + result = mod.build_context(workspace, project="example", task=task) + assert result.ok + assert result.plan is None + assert len(result.plan_candidates) == 3 + explicit = mod.build_context(workspace, project="example", plan="30_projects/example/plans/98-onboarding.md") + assert explicit.plan_status == "explicit" + assert next(e for e in explicit.entries if e.path == explicit.plan).required + + +def test_invalid_and_escaping_plan_paths_remain_incomplete(workspace, tmp_path): + outside = tmp_path / "outside.md" + outside.write_text("Outside plan\n") + link = workspace / "30_projects/example/plans/link.md" + link.symlink_to(outside) + for plan in ("outside.md", "30_projects/example/plans/missing.md", str(link)): + result = mod.build_context(workspace, project="example", plan=plan) + assert not result.ok + assert result.context_status == "incomplete" + assert result.plan_error + + +def test_arrival_rejects_project_arguments_and_resume_needs_a_project(workspace): + result = mod.build_context(workspace, project="example", intent="arrival") + assert not result.ok and result.intent_error + (workspace / "STATE.md").write_text("# No focus\n") + result = mod.build_context(workspace, intent="resume") + assert not result.ok and result.project_error + + +def test_plan_directory_cannot_redirect_to_another_project(workspace): + plans = workspace / "30_projects/example/plans" + other = workspace / "30_projects/other/plans" + other.parent.mkdir() + plans.rename(other) + plans.symlink_to(other, target_is_directory=True) + result = mod.build_context(workspace, project="example", task="onboarding") + assert not result.ok + assert result.plan_error + assert result.plan_candidates == [] + + +def test_missing_workflow_and_binary_contract_are_explicit(workspace): + (workspace / ".context/workflows/session-open.md").unlink() + (workspace / "AGENTS.md").write_bytes(b"\xff\xfe") + result = mod.build_context(workspace) + assert not result.ok + assert result.context_status == "incomplete" + assert ".context/workflows/session-open.md" in result.missing_required + assert any("AGENTS.md" in error for error in result.read_errors) + + +@pytest.mark.parametrize("content", ["Read this complete rule.\n" * 1500, "\U0001f331\u754c" * 5000]) +def test_content_batches_are_lossless_bounded_and_utf8_safe(workspace, content): + (workspace / "AGENTS.md").write_text(content) + result = mod.build_context(workspace) + assert result.ok + assert content_for(result, workspace, "AGENTS.md") == content + assert all(0 < b["bytes"] <= mod.BATCH_BYTES for b in result.read_batches) + assert len([b for b in result.read_batches if b["path"] == "AGENTS.md"]) > 1 + + +def test_changed_source_is_rejected_after_a_batch_manifest(workspace): + result = mod.build_context(workspace) + (workspace / "AGENTS.md").write_text("Changed rules\n") + with pytest.raises(ValueError, match="source changed"): + mod.read_batch(result, workspace, 1) + + +def test_latest_coordination_is_bounded_without_losing_current_entry(workspace): + log = workspace / "30_projects/example/log.md" + log.write_text("# Log\n\n## Today\nCurrent authority.\n" + "A long current entry.\n" * 900 + "\n## Yesterday\nHistorical detail.\n") + result = mod.build_context(workspace, project="example") + delivered = content_for(result, workspace, "30_projects/example/log.md") + assert "Current authority." in delivered + assert delivered.count("A long current entry.") == 900 + assert "Historical detail." not in delivered + + +def test_real_cli_survives_missing_scheduler_and_serves_one_batch(workspace): + for relative in ( + "bin/session-open", + "scripts/focus_authority.py", + "scripts/lifecycle_identity.py", + ): + target = workspace / relative + target.parent.mkdir(parents=True, exist_ok=True) + shutil.copy2(ROOT / relative, target) + (workspace / "AGENTS.md").write_text("Long rule.\n" * 2000) + command = [sys.executable, str(workspace / "bin/session-open")] + listing = subprocess.run(command + ["--json"], cwd=workspace, capture_output=True, text=True) + assert listing.returncode == 0, listing.stderr + payload = json.loads(listing.stdout) + assert payload["eval_schedule_ok"] is None + assert "unavailable" in payload["eval_schedule_summary"] + assert payload["reading_verified"] is False + assert payload["context_status"] == "unread" + batch = subprocess.run(command + ["--read-batch", "1", "--json"], cwd=workspace, capture_output=True, text=True) + assert batch.returncode == 0, batch.stderr + delivered = json.loads(batch.stdout) + assert len(delivered["batch"]["content"].encode()) <= mod.BATCH_BYTES + assert "read_batches" not in delivered + wrong_hash = subprocess.run(command + ["--read-batch", "1", "--expect-hash", "wrong", "--json"], cwd=workspace, capture_output=True, text=True) + assert wrong_hash.returncode == 1 + assert json.loads(wrong_hash.stdout)["context_status"] == "incomplete" + invalid = subprocess.run(command + ["--read-batch", "0", "--json"], cwd=workspace, capture_output=True, text=True) + assert invalid.returncode == 1 + printed = subprocess.run(command + ["--print-contents"], cwd=workspace, capture_output=True, text=True) + assert printed.returncode == 0 + assert "Only batch 1 printed" in printed.stdout + assert printed.stdout.count("Long rule.") < 2000 diff --git a/tests/test_sync_project_index.py b/tests/test_sync_project_index.py deleted file mode 100644 index 3f71d59..0000000 --- a/tests/test_sync_project_index.py +++ /dev/null @@ -1,241 +0,0 @@ -from __future__ import annotations - -import importlib.util -import os -import subprocess -import sys -import tempfile -import unittest -from datetime import datetime, timedelta -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -SCRIPT_PATH = ROOT / "bin" / "sync-project-index" -LOADER = SourceFileLoader("sync_project_index", str(SCRIPT_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -mod = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = mod -SPEC.loader.exec_module(mod) - - -README_TEMPLATE = """\ ---- -title: "{title}" -project_state: "{state}" -goal: "Do the thing" -next_action: "{next_action}" -updated: "2026-06-01" -{wip_class_line}--- - -# {title} -""" - - -class ValidateProjectsTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.projects = Path(self.tmp.name) / "30_projects" - self.projects.mkdir() - - def tearDown(self) -> None: - self.tmp.cleanup() - - def add_project( - self, - name: str, - state: str = "active", - next_action: str = "Do the next step", - age_days: float = 0.0, - frontmatter: bool = True, - wip_class: str | None = None, - ) -> Path: - project = self.projects / name - project.mkdir() - readme = project / "README.md" - if frontmatter: - wip_line = f'wip_class: "{wip_class}"\n' if wip_class else "" - readme.write_text( - README_TEMPLATE.format( - title=name, - state=state, - next_action=next_action, - wip_class_line=wip_line, - ), - encoding="utf-8", - ) - else: - readme.write_text(f"# {name}\n\nNo frontmatter here.\n", - encoding="utf-8") - if age_days: - old_ts = (datetime.now() - timedelta(days=age_days)).timestamp() - os.utime(readme, (old_ts, old_ts)) - os.utime(project, (old_ts, old_ts)) - return project - - def problems(self, **kwargs) -> list[str]: - return mod.validate_projects(self.projects, **kwargs) - - def test_fresh_active_passes(self) -> None: - self.add_project("hot-project") - self.assertEqual(self.problems(), []) - - def test_idle_active_fails_loudly(self) -> None: - self.add_project("stale-project", age_days=20) - problems = self.problems() - self.assertEqual(len(problems), 1) - self.assertIn("stale-project: active but idle", problems[0]) - self.assertIn("pause it", problems[0]) - - def test_nested_repo_commit_counts_as_evidence(self) -> None: - project = self.add_project("repo-project", age_days=20) - subprocess.run(["git", "-C", str(project), "init", "-q", "-b", "main"], - check=True) - subprocess.run(["git", "-C", str(project), "add", "-A"], check=True) - subprocess.run( - ["git", "-C", str(project), "-c", "user.email=t@t", "-c", - "user.name=t", "commit", "-q", "-m", "fresh work"], - check=True) - # File mtimes are 20d old, but the commit is from just now. - self.assertEqual(self.problems(), []) - - def test_missing_frontmatter_flagged(self) -> None: - self.add_project("bare-project", frontmatter=False) - problems = self.problems() - self.assertEqual(len(problems), 1) - self.assertIn("bare-project: missing project_state", problems[0]) - - def test_missing_readme_flagged(self) -> None: - (self.projects / "empty-project").mkdir() - problems = self.problems() - self.assertIn("empty-project: missing README.md", problems) - - def test_unknown_state_flagged(self) -> None: - self.add_project("weird-project", state="wip") - problems = self.problems() - self.assertEqual(len(problems), 1) - self.assertIn("unknown project_state 'wip'", problems[0]) - - def test_paused_requires_reentry_pointer(self) -> None: - self.add_project("shelved-ok", state="paused", - next_action="Reread plans/reentry.md") - self.add_project("shelved-bare", state="paused", next_action="") - problems = self.problems() - self.assertEqual(len(problems), 1) - self.assertIn("shelved-bare: paused without next_action", problems[0]) - - def test_active_requires_next_action(self) -> None: - self.add_project("aimless", next_action="") - problems = self.problems() - self.assertEqual(len(problems), 1) - self.assertIn("aimless: active without next_action", problems[0]) - - def test_product_wip_cap_breach(self) -> None: - for i in range(6): - self.add_project(f"proj-{i}", wip_class="product") - problems = self.problems() - self.assertEqual(len(problems), 1) - self.assertIn("product WIP cap breach: 6 active product projects (cap 5)", problems[0]) - for i in range(6): - self.assertIn(f"proj-{i}", problems[0]) - - def test_product_wip_cap_boundary_passes(self) -> None: - for i in range(5): - self.add_project(f"proj-{i}", wip_class="product") - self.assertEqual(self.problems(), []) - - def test_eval_actives_do_not_consume_product_seats(self) -> None: - for i in range(5): - self.add_project(f"product-{i}", wip_class="product") - for i in range(5): - self.add_project(f"metric-{i}-eval") # suffix heuristic → eval - self.assertEqual(self.problems(), []) - - def test_total_active_cap_breach(self) -> None: - for i in range(5): - self.add_project(f"product-{i}", wip_class="product") - for i in range(6): - self.add_project(f"suite-{i}-eval") - problems = self.problems() - self.assertTrue(any("total active cap breach: 11" in p for p in problems), problems) - - def test_default_eval_slug_set(self) -> None: - for i in range(5): - self.add_project(f"product-{i}", wip_class="product") - # claim-audit-lab is in DEFAULT_EVAL_SLUGS (no -eval suffix) - self.add_project("claim-audit-lab") - self.assertEqual(self.problems(), []) - - def test_explicit_wip_class_overrides_suffix(self) -> None: - # Force a *-eval slug into product pool — counts toward product cap. - for i in range(5): - self.add_project(f"other-{i}", wip_class="product") - self.add_project("forced-eval", wip_class="product") - problems = self.problems() - self.assertTrue(any("product WIP cap breach" in p for p in problems), problems) - - def test_anchor_does_not_consume_product_or_total_seats(self) -> None: - for i in range(5): - self.add_project(f"product-{i}", wip_class="product") - for i in range(5): - self.add_project(f"suite-{i}-eval") - self.add_project("income-engine") # default anchor - self.assertEqual(self.problems(), []) - - def test_non_active_states_have_no_evidence_rule(self) -> None: - self.add_project("old-shipped", state="shipped", age_days=90) - self.add_project("old-planned", state="planned", age_days=90) - self.add_project("old-suspended", state="suspended", age_days=90) - self.assertEqual(self.problems(), []) - - def test_only_scopes_to_single_project(self) -> None: - self.add_project("healthy", state="shipped") - self.add_project("broken", frontmatter=False) - self.assertEqual(self.problems(only="healthy"), []) - problems = self.problems(only="broken") - self.assertEqual(len(problems), 1) - self.assertIn("broken: missing project_state frontmatter", problems[0]) - - def test_only_missing_directory_is_a_problem(self) -> None: - problems = self.problems(only="no-such-project") - self.assertEqual( - problems, ["no-such-project: project directory does not exist"]) - - def test_only_skips_wip_cap(self) -> None: - for i in range(6): - self.add_project(f"proj-{i}") - self.assertEqual(self.problems(only="proj-0"), []) - - -class RenderTests(unittest.TestCase): - def setUp(self) -> None: - self.tmp = tempfile.TemporaryDirectory() - self.projects = Path(self.tmp.name) / "30_projects" - self.projects.mkdir() - - def tearDown(self) -> None: - self.tmp.cleanup() - - def test_render_includes_evidence_column(self) -> None: - project = self.projects / "alpha" - project.mkdir() - (project / "README.md").write_text( - README_TEMPLATE.format( - title="Alpha", - state="active", - next_action="Next", - wip_class_line="", - ), - encoding="utf-8", - ) - output = mod.render(self.projects) - self.assertIn("| Project | State | Goal | Next action | Updated | Evidence |", - output) - today = datetime.now().date().isoformat() - self.assertIn(f"| {today} |", output) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_task_packet.py b/tests/test_task_packet.py deleted file mode 100644 index 4242e65..0000000 --- a/tests/test_task_packet.py +++ /dev/null @@ -1,537 +0,0 @@ -from __future__ import annotations - -import importlib.util -import json -import subprocess -import sys -import tempfile -import unittest -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -SCRIPT = ROOT / "bin" / "task-packet" -LOADER = SourceFileLoader("task_packet", str(SCRIPT)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -task_packet = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = task_packet -SPEC.loader.exec_module(task_packet) - - -def packet_text(**overrides: object) -> str: - metadata: dict[str, object] = { - "packet_version": 1, - "task_id": "fix-one", - "project_slug": "demo", - "title": "Fix one thing", - "status": "ready", - "task_kind": "code", - "agent_profile": "local-test", - "workdir": "workbench", - "timeout_seconds": 60, - "editable_files": ["src/app.py"], - "create_files": [], - "read_only_files": ["AGENTS.md"], - "verification_commands": ["python3 -m unittest -q"], - "mindgraph_mode": "off", - "knowledge_queries": [], - } - metadata.update(overrides) - lines = ["---"] - for key, value in metadata.items(): - if isinstance(value, list): - rendered = json.dumps(value) - elif isinstance(value, int): - rendered = str(value) - else: - rendered = json.dumps(value) - lines.append(f"{key}: {rendered}") - lines.extend(["---", "", "# Task Packet", ""]) - for section in task_packet.REQUIRED_SECTIONS: - lines.extend([f"## {section}", "", f"{section} content.", ""]) - return "\n".join(lines) - - -class TaskPacketTests(unittest.TestCase): - def make_project(self, root: Path) -> Path: - project = root / "demo" - packet_dir = project / "plans" / "task-packets" - (project / "workbench" / "src").mkdir(parents=True) - (project / "workbench" / "AGENTS.md").write_text("", encoding="utf-8") - packet_dir.mkdir(parents=True) - return project - - def test_valid_ready_packet(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text(packet_text(), encoding="utf-8") - packet = task_packet.load_packet(path) - - errors = task_packet.validate_packet( - packet, - projects_dir=projects, - require_ready=True, - ) - - self.assertEqual(errors, []) - - def test_inspect_emits_validated_canonical_summary(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - projects = root / "30_projects" - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text(packet_text(), encoding="utf-8") - - completed = subprocess.run( - [ - sys.executable, - str(SCRIPT), - "inspect", - str(path), - "--projects-dir", - str(projects), - "--require-ready", - ], - check=True, - capture_output=True, - text=True, - ) - summary = json.loads(completed.stdout) - - self.assertEqual(summary["task_id"], "fix-one") - self.assertEqual(summary["project_slug"], "demo") - self.assertEqual(summary["status"], "ready") - self.assertEqual( - summary["packet_path"], - "30_projects/demo/plans/task-packets/fix-one.md", - ) - self.assertRegex(summary["contract_sha256"], r"^[a-f0-9]{64}$") - self.assertRegex(summary["source_sha256"], r"^[a-f0-9]{64}$") - self.assertEqual(summary["workdir"], "workbench") - - def test_optional_query_pass_section_allowed(self) -> None: - text = packet_text() + ( - "## MindGraph Query Pass\n\n" - "intent: demo\n" - "doctor: Overall: OK\n" - "knowledge_query: demo knowledge\n" - "projects_query: demo project\n" - "knowledge (durable_knowledge):\n" - "- path: 10_knowledge/x.md\n" - " reason: test\n" - "projects (project_status):\n" - "- path: 30_projects/demo/README.md\n" - " reason: test\n" - "weak_or_excluded:\n" - "- none\n" - "source_inspection_required:\n" - "- none\n" - "do_not_merge_without_trust_labels: true\n" - ) - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text(text, encoding="utf-8") - packet = task_packet.load_packet(path) - errors = task_packet.validate_packet(packet, projects_dir=projects) - self.assertEqual(errors, []) - - def test_curated_mode_rejects_replace_placeholders_in_query_pass(self) -> None: - text = packet_text( - mindgraph_mode="curated", - knowledge_queries=["demo query"], - ) + ( - "## MindGraph Query Pass\n\n" - "intent: REPLACE — decision this retrieval supports\n" - "doctor: REPLACE\n" - "knowledge_query: REPLACE\n" - ) - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text(text, encoding="utf-8") - packet = task_packet.load_packet(path) - errors = task_packet.validate_packet(packet, projects_dir=projects) - self.assertTrue(any("REPLACE" in e for e in errors)) - - def test_rejects_unsafe_paths_overlap_and_shell_operators(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text( - packet_text( - editable_files=["../secret", "AGENTS.md"], - read_only_files=["AGENTS.md"], - verification_commands=["pytest && rm -rf /"], - ), - encoding="utf-8", - ) - packet = task_packet.load_packet(path) - errors = task_packet.validate_packet(packet, projects_dir=projects) - - self.assertTrue(any("unsafe path" in error for error in errors)) - self.assertTrue(any("overlap" in error for error in errors)) - self.assertTrue(any("shell operator" in error for error in errors)) - - def test_rejects_ancestor_descendant_scope_overlap(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text( - packet_text( - editable_files=["src", "src/app.py"], - create_files=["generated/report.json"], - read_only_files=["src/secrets.txt", "generated"], - ), - encoding="utf-8", - ) - - errors = task_packet.validate_packet( - task_packet.load_packet(path), projects_dir=projects - ) - - self.assertTrue( - any("editable_files contains overlapping paths" in error for error in errors) - ) - self.assertTrue( - any("editable_files and read_only_files" in error for error in errors) - ) - self.assertTrue( - any("create_files and read_only_files" in error for error in errors) - ) - - def test_rejects_noncanonical_relative_path_aliases(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text( - packet_text( - editable_files=["src//app.py", "src/./other.py"], - read_only_files=["src\\secrets.txt"], - ), - encoding="utf-8", - ) - - errors = task_packet.validate_packet( - task_packet.load_packet(path), projects_dir=projects - ) - - self.assertEqual(sum("unsafe path" in error for error in errors), 3) - - def test_rejects_missing_section_and_nonexistent_workdir(self) -> None: - text = packet_text().replace( - "## Stop conditions\n\nStop conditions content.\n", - "", - ) - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text( - text.replace('workdir: "workbench"', 'workdir: "missing"'), - encoding="utf-8", - ) - packet = task_packet.load_packet(path) - errors = task_packet.validate_packet(packet, projects_dir=projects) - - self.assertTrue(any("Stop conditions" in error for error in errors)) - self.assertTrue(any("workdir does not exist" in error for error in errors)) - - def test_rejects_absolute_paths_missing_verification_and_draft_execution(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text( - packet_text( - status="draft", - editable_files=["/tmp/escape.py"], - verification_commands=[], - ), - encoding="utf-8", - ) - packet = task_packet.load_packet(path) - errors = task_packet.validate_packet( - packet, - projects_dir=projects, - require_ready=True, - ) - - self.assertTrue(any("relative paths" in error for error in errors)) - self.assertTrue(any("verification command" in error for error in errors)) - self.assertTrue(any("status ready" in error for error in errors)) - - def test_accepts_optional_routing_metadata(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text( - packet_text( - task_category="mechanical-edit", - executor="local", - harness_recommendation="H1-packet", - allow_fusion_plan=False, - needs_deliberation=False, - ), - encoding="utf-8", - ) - packet = task_packet.load_packet(path) - errors = task_packet.validate_packet(packet, projects_dir=projects) - - self.assertEqual(errors, []) - summary = task_packet.packet_summary(packet, root=projects.parent) - self.assertEqual(summary["task_category"], "mechanical-edit") - self.assertEqual(summary["harness_recommendation"], "H1-packet") - - def test_rejects_invalid_routing_metadata(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text( - packet_text( - task_category="not-a-category", - executor="hybrid", - harness_recommendation="H9-fusion", - ), - encoding="utf-8", - ) - packet = task_packet.load_packet(path) - errors = task_packet.validate_packet(packet, projects_dir=projects) - - self.assertTrue(any("task_category" in error for error in errors)) - self.assertTrue(any("executor" in error for error in errors)) - self.assertTrue(any("harness_recommendation" in error for error in errors)) - - def test_infer_routing_hints_for_multi_file_code(self) -> None: - hints = task_packet.infer_routing_hints( - { - "task_kind": "code", - "editable_files": ["a.py", "b.py"], - "create_files": [], - } - ) - self.assertEqual(hints["task_category"], "multi-file-coordination") - self.assertEqual(hints["harness_recommendation"], "H2-repair") - self.assertEqual(hints["agent_profile"], "local-qwen25-coder-14b") - - def test_rejects_unknown_metadata_and_duplicate_paths(self) -> None: - text = packet_text( - editable_files=["src/app.py", "src/app.py"], - ).replace( - 'title: "Fix one thing"', - 'title: "Fix one thing"\nmisspelled_timeout: 99', - ) - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text(text, encoding="utf-8") - packet = task_packet.load_packet(path) - errors = task_packet.validate_packet(packet, projects_dir=projects) - - self.assertTrue(any("unknown frontmatter key" in error for error in errors)) - self.assertTrue(any("duplicate paths" in error for error in errors)) - - def test_rejects_unknown_and_duplicate_sections(self) -> None: - unknown = packet_text() + "\n## Surprise\n\nNot part of the contract.\n" - duplicate = packet_text() + "\n## Goal\n\nA second goal.\n" - with tempfile.TemporaryDirectory() as tmp: - projects = Path(tmp) - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text(unknown, encoding="utf-8") - errors = task_packet.validate_packet( - task_packet.load_packet(path), - projects_dir=projects, - ) - path.write_text(duplicate, encoding="utf-8") - with self.assertRaises(task_packet.PacketError): - task_packet.load_packet(path) - - self.assertTrue(any("unknown packet section" in error for error in errors)) - - def test_compile_writes_deterministic_manifest(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - projects = root / "30_projects" - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - path.write_text(packet_text(), encoding="utf-8") - output = projects / "task_packets_manifest.json" - - first = task_packet.compile_packets( - projects_dir=projects, - output=output, - ) - first_bytes = output.read_bytes() - second = task_packet.compile_packets( - projects_dir=projects, - output=output, - ) - - self.assertEqual(first, second) - self.assertEqual(first_bytes, json.dumps(second, indent=2, sort_keys=True).encode() + b"\n") - self.assertEqual(first["packets"][0]["task_id"], "fix-one") - self.assertEqual( - first["packets"][0]["packet_path"], - "30_projects/demo/plans/task-packets/fix-one.md", - ) - - def test_compile_does_not_reuse_predictable_temporary_path(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - projects = root / "30_projects" - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - output = projects / "task_packets_manifest.json" - predictable_temp = output.with_suffix(output.suffix + ".tmp") - path.write_text(packet_text(), encoding="utf-8") - predictable_temp.write_text("sentinel", encoding="utf-8") - - payload = task_packet.compile_packets( - projects_dir=projects, output=output - ) - - self.assertEqual( - predictable_temp.read_text(encoding="utf-8"), "sentinel" - ) - self.assertEqual(payload["invalid"], []) - - def test_invalid_compile_preserves_last_good_manifest(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - projects = root / "30_projects" - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - output = projects / "task_packets_manifest.json" - path.write_text(packet_text(), encoding="utf-8") - task_packet.compile_packets(projects_dir=projects, output=output) - last_good = output.read_bytes() - path.write_text(packet_text(verification_commands=[]), encoding="utf-8") - - payload = task_packet.compile_packets( - projects_dir=projects, - output=output, - ) - output_bytes = output.read_bytes() - - self.assertTrue(payload["invalid"]) - self.assertEqual(output_bytes, last_good) - - def test_corrupt_manifest_fails_closed_and_is_preserved(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - projects = root / "30_projects" - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - output = projects / "task_packets_manifest.json" - path.write_text(packet_text(), encoding="utf-8") - corrupt = b'{"packet_version": 1, "packets": [' - output.write_bytes(corrupt) - - payload = task_packet.compile_packets( - projects_dir=projects, output=output - ) - - self.assertTrue(payload["invalid"]) - self.assertIn( - "invalid existing manifest JSON", - payload["invalid"][0]["errors"][0], - ) - self.assertEqual(output.read_bytes(), corrupt) - - def test_ready_manifest_entry_without_contract_hash_fails_closed(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - projects = root / "30_projects" - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - output = projects / "task_packets_manifest.json" - path.write_text(packet_text(), encoding="utf-8") - corrupt = json.dumps( - { - "packet_version": 1, - "packets": [ - { - "id": "packet-demo-fix-one", - "status": "ready", - } - ], - "invalid": [], - } - ).encode("utf-8") - output.write_bytes(corrupt) - - payload = task_packet.compile_packets( - projects_dir=projects, output=output - ) - - self.assertTrue(payload["invalid"]) - self.assertIn( - "no valid contract hash", payload["invalid"][0]["errors"][0] - ) - self.assertEqual(output.read_bytes(), corrupt) - - def test_compiled_ready_contract_is_immutable(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - projects = root / "30_projects" - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - output = projects / "task_packets_manifest.json" - path.write_text(packet_text(), encoding="utf-8") - task_packet.compile_packets(projects_dir=projects, output=output) - last_good = output.read_bytes() - - path.write_text( - packet_text(title="Silently changed contract"), - encoding="utf-8", - ) - payload = task_packet.compile_packets( - projects_dir=projects, - output=output, - ) - - preserved = output.read_bytes() - - self.assertTrue(payload["invalid"]) - self.assertIn("immutable", payload["invalid"][0]["errors"][0]) - self.assertEqual(preserved, last_good) - - def test_compiled_ready_contract_may_not_be_deleted(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - root = Path(tmp) - projects = root / "30_projects" - project = self.make_project(projects) - path = project / "plans" / "task-packets" / "fix-one.md" - output = projects / "task_packets_manifest.json" - path.write_text(packet_text(), encoding="utf-8") - task_packet.compile_packets(projects_dir=projects, output=output) - last_good = output.read_bytes() - - path.unlink() - payload = task_packet.compile_packets( - projects_dir=projects, - output=output, - ) - preserved = output.read_bytes() - - self.assertTrue(payload["invalid"]) - self.assertIn("may not be deleted", payload["invalid"][0]["errors"][0]) - self.assertEqual(preserved, last_good) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_workflow_event.py b/tests/test_workflow_event.py deleted file mode 100644 index 738fb7a..0000000 --- a/tests/test_workflow_event.py +++ /dev/null @@ -1,802 +0,0 @@ -from __future__ import annotations - -from datetime import datetime -import importlib.util -import json -import multiprocessing -import os -import shutil -import subprocess -import sys -import tempfile -import time -import unittest -from importlib.machinery import SourceFileLoader -from io import StringIO -from pathlib import Path -from unittest import mock - - - -ROOT = Path(__file__).resolve().parents[1] -EVENT_PATH = ROOT / "bin" / "workflow-event" -# workflow-event has no .py suffix, so spec_from_file_location refuses to -# infer a loader. Use SourceFileLoader explicitly. -LOADER = SourceFileLoader("workflow_event", str(EVENT_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -workflow_event = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = workflow_event -SPEC.loader.exec_module(workflow_event) - - -class CommandHeadRedactionTests(unittest.TestCase): - """The command_head field must never carry a filesystem basename. - - A leak surfaced during the initial publish-readiness review: when a - Bash invocation began with an env assignment like ``SANDBOX=/tmp/foo`` - the old logic took ``Path(...).name`` of that token and wrote ``foo`` - into telemetry. These tests pin the post-fix behaviour. - """ - - def test_plain_command_keeps_bare_name(self) -> None: - self.assertEqual(workflow_event.command_head("git status"), "git") - self.assertEqual(workflow_event.command_head("python3 -m unittest"), "python3") - self.assertEqual(workflow_event.command_head("ls"), "ls") - - def test_env_assignment_prefix_is_skipped(self) -> None: - self.assertEqual( - workflow_event.command_head("SANDBOX=/tmp/secret-project python3 run.py"), - "python3", - ) - self.assertEqual( - workflow_event.command_head("FOO=bar BAZ=qux make build"), - "make", - ) - - def test_path_shaped_head_collapses(self) -> None: - self.assertEqual( - workflow_event.command_head("/opt/tools/run.sh --flag"), - "<path>", - ) - self.assertEqual(workflow_event.command_head("./bin/private-helper"), "<path>") - - def test_only_env_assignments_returns_none(self) -> None: - self.assertIsNone(workflow_event.command_head("FOO=bar")) - - def test_empty_command_returns_none(self) -> None: - self.assertIsNone(workflow_event.command_head("")) - - def test_safe_head_punctuation_is_preserved(self) -> None: - self.assertEqual(workflow_event.command_head("python3.14 -V"), "python3.14") - self.assertEqual(workflow_event.command_head("npm install"), "npm") - - -class BashSummaryTests(unittest.TestCase): - def test_bash_summary_hashes_full_command(self) -> None: - summary = workflow_event.bash_summary( - {"command": "SANDBOX=/tmp/x python3 run.py", "timeout": 60} - ) - self.assertEqual(summary["command_head"], "python3") - self.assertEqual(summary["timeout"], 60) - self.assertIsNotNone(summary["command_hash"]) - # The hash is deterministic for the same input. - again = workflow_event.bash_summary( - {"command": "SANDBOX=/tmp/x python3 run.py", "timeout": 60} - ) - self.assertEqual(summary["command_hash"], again["command_hash"]) - - def test_bash_summary_handles_empty_command(self) -> None: - summary = workflow_event.bash_summary({}) - self.assertIsNone(summary["command_head"]) - self.assertIsNone(summary["command_hash"]) - - def test_bash_summary_handles_codex_cmd_field(self) -> None: - summary = workflow_event.bash_summary({"cmd": "git status --short"}) - self.assertEqual(summary["command_head"], "git") - self.assertIsNotNone(summary["command_hash"]) - - def test_grok_terminal_tool_uses_redacted_command_summary(self) -> None: - summary = workflow_event.input_summary( - "run_terminal_command", - {"command": "PRIVATE=/tmp/path npm test"}, - ) - self.assertEqual(summary["command_head"], "npm") - self.assertNotIn("/tmp/path", json.dumps(summary)) - - -class ProcessIdTests(unittest.TestCase): - def test_process_ids_keep_catalogue_shaped_values(self) -> None: - ids = workflow_event.process_ids( - { - "process_id": "cli-workflow-report", - "process_ids": ["workflow-workflow-telemetry", "cli-workflow-report"], - } - ) - self.assertEqual( - ids, - ["cli-workflow-report", "workflow-workflow-telemetry"], - ) - - def test_process_ids_reject_paths_and_free_text(self) -> None: - ids = workflow_event.process_ids( - { - "process_id": "/tmp/example-mainframe/bin/workflow-report", - "catalogue_id": "ran private command", - "catalogue_ids": ["client-name-sensitive-case"], - } - ) - self.assertEqual(ids, []) - - def test_process_ids_keep_only_catalogue_prefixes(self) -> None: - ids = workflow_event.process_ids( - { - "process_ids": [ - "cli-workflow-report", - "workflow-workflow-telemetry", - "skill-source-literature", - "agent-ingest-agent", - "script-eval-schedule-implementation", - "ingest-minion-module", - "private-case-slug", - ], - } - ) - self.assertEqual( - ids, - [ - "cli-workflow-report", - "workflow-workflow-telemetry", - "skill-source-literature", - "agent-ingest-agent", - "script-eval-schedule-implementation", - "ingest-minion-module", - ], - ) - - def test_process_ids_can_come_from_environment(self) -> None: - with mock.patch.dict( - "os.environ", - {"MAINFRAME_PROCESS_ID": "cli-workflow-event,workflow-workflow-telemetry"}, - clear=False, - ): - ids = workflow_event.process_ids({}) - self.assertEqual( - ids, - ["cli-workflow-event", "workflow-workflow-telemetry"], - ) - - -class PathClassTests(unittest.TestCase): - def test_path_inside_repo_zone(self) -> None: - cls = workflow_event.path_class(str(ROOT / "10_knowledge" / "index.md")) - self.assertEqual(cls["path_zone"], "10_knowledge") - self.assertEqual(cls["extension"], ".md") - - def test_path_outside_repo_marked_external(self) -> None: - cls = workflow_event.path_class("/etc/hosts") - self.assertEqual(cls["path_zone"], "external") - - -class ResponseSuccessTests(unittest.TestCase): - """Codex reports tool outcomes inside PostToolUse's tool_response.""" - - def test_exit_code_zero_is_success(self) -> None: - self.assertTrue(workflow_event.response_success({"exit_code": 0})) - - def test_nonzero_exit_code_is_failure(self) -> None: - self.assertFalse(workflow_event.response_success({"exit_code": 2})) - self.assertFalse(workflow_event.response_success({"returncode": 1})) - - def test_exit_code_nested_in_metadata(self) -> None: - self.assertFalse( - workflow_event.response_success({"metadata": {"exitCode": 3}}) - ) - - def test_explicit_success_flag(self) -> None: - self.assertFalse(workflow_event.response_success({"success": False})) - self.assertTrue(workflow_event.response_success({"success": True})) - - def test_bool_exit_code_is_not_an_exit_code(self) -> None: - # True is an int in Python; it must not be read as exit code 1. - self.assertIsNone(workflow_event.response_success({"exit_code": True})) - - def test_unrecognized_shapes_return_none(self) -> None: - self.assertIsNone(workflow_event.response_success("plain output")) - self.assertIsNone(workflow_event.response_success(None)) - self.assertIsNone(workflow_event.response_success({"output": "ok"})) - - -class HookPayloadNormalizationTests(unittest.TestCase): - def test_grok_camel_case_payload_is_normalized(self) -> None: - normalized = workflow_event.normalize_hook_payload( - { - "hookEventName": "pre_tool_use", - "sessionId": "grok-session", - "workspaceRoot": str(ROOT), - "toolName": "run_terminal_command", - "toolInput": {"command": "python3 -m unittest"}, - "toolUseId": "tool-1", - "modelId": "grok-build", - } - ) - - self.assertEqual(normalized["hook_event_name"], "PreToolUse") - self.assertEqual(normalized["session_id"], "grok-session") - self.assertEqual(normalized["cwd"], str(ROOT)) - self.assertEqual(normalized["tool_name"], "run_terminal_command") - self.assertEqual(normalized["tool_input"]["command"], "python3 -m unittest") - self.assertEqual(normalized["tool_use_id"], "tool-1") - self.assertEqual(normalized["model"], "grok-build") - - def test_grok_event_aliases_cover_session_and_failure_events(self) -> None: - self.assertEqual( - workflow_event.normalize_hook_payload( - {"hookEventName": "session_start"} - )["hook_event_name"], - "SessionStart", - ) - self.assertEqual( - workflow_event.normalize_hook_payload( - {"hookEventName": "post_tool_use_failure"} - )["hook_event_name"], - "PostToolUseFailure", - ) - self.assertEqual( - workflow_event.normalize_hook_payload( - {"hookEventName": "session_end"} - )["hook_event_name"], - "SessionEnd", - ) - - -class SafeEventNewSignalsTests(unittest.TestCase): - def test_permission_request_sets_permission_kind(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "PermissionRequest", - "tool_name": "Bash", - "tool_input": {"command": "rm -rf build"}, - "session_id": "s1", - }, - client="codex", - ) - self.assertEqual(event["notification_kind"], "permission") - self.assertNotIn("rm -rf", json.dumps(event)) - - def test_post_tool_use_derives_failure_from_response(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "PostToolUse", - "tool_name": "Bash", - "tool_input": {"command": "make test"}, - "tool_response": {"exit_code": 1}, - "session_id": "s1", - }, - client="codex", - ) - self.assertFalse(event["success"]) - - def test_post_tool_use_without_outcome_stays_success(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "PostToolUse", - "tool_name": "Read", - "tool_input": {"file_path": "README.md"}, - "tool_response": "text", - "session_id": "s1", - }, - client="claude", - ) - self.assertTrue(event["success"]) - - def test_post_tool_use_failure_event_stays_failure(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "PostToolUseFailure", - "tool_name": "Bash", - "tool_input": {"command": "ls"}, - "tool_response": {"exit_code": 0}, - "session_id": "s1", - }, - client="claude", - ) - self.assertFalse(event["success"]) - - def test_grok_permission_denied_is_blocked(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "PermissionDenied", - "tool_name": "run_terminal_command", - "tool_input": {"command": "git push"}, - "session_id": "s1", - }, - client="grok", - ) - self.assertFalse(event["success"]) - - def test_compact_events_record_trigger_only(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "PreCompact", - "trigger": "auto", - "custom_instructions": "keep the secret plan", - "session_id": "s1", - }, - client="claude", - ) - self.assertEqual(event["trigger"], "auto") - self.assertNotIn("secret plan", json.dumps(event)) - - def test_aider_diagnostic_uses_allowlisted_code(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "Diagnostic", - "diagnostic": "context_limit_exceeded", - "detail": "private model output", - "session_id": "s1", - }, - client="local", - ) - self.assertEqual(event["diagnostic"], "context_limit_exceeded") - self.assertNotIn("private model output", json.dumps(event)) - - def test_verification_purpose_and_observed_count_are_preserved(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "PostToolUse", - "tool_name": "Bash", - "tool_input": {"command": "pytest -q"}, - "tool_purpose": "verification", - "observed_count": 2, - "session_id": "s1", - }, - client="local", - ) - self.assertEqual(event["tool_purpose"], "verification") - self.assertEqual(event["observed_count"], 2) - - def test_aider_source_is_preserved_on_tool_events(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "PostToolUse", - "source": "aider", - "tool_name": "Edit", - "session_id": "s1", - }, - client="local", - ) - self.assertEqual(event["source"], "aider") - - def test_safe_event_preserves_valid_process_ids_only(self) -> None: - event = workflow_event.safe_event( - { - "hook_event_name": "PostToolUse", - "tool_name": "Bash", - "tool_input": {"command": "bin/workflow-report --days 1"}, - "process_ids": [ - "cli-workflow-report", - "/tmp/example-mainframe/bin/workflow-report", - "workflow-workflow-telemetry", - ], - "session_id": "s1", - }, - client="codex", - ) - self.assertEqual(event["process_id"], "cli-workflow-report") - self.assertEqual( - event["process_ids"], - ["cli-workflow-report", "workflow-workflow-telemetry"], - ) - self.assertNotIn("/tmp/example-user", json.dumps(event)) - - -class WriteEventTests(unittest.TestCase): - """End-to-end: stdin JSON -> redacted JSONL on disk.""" - - def test_writes_redacted_event(self) -> None: - payload = { - "hook_event_name": "PostToolUse", - "tool_name": "Bash", - "tool_input": { - "command": "SANDBOX=/tmp/fixture python3 minion.py --root /tmp/fixture", - }, - "tool_use_id": "abc123", - "session_id": "session-xyz", - "duration_ms": 42, - "cwd": str(ROOT), - } - with tempfile.TemporaryDirectory() as tmp: - env = {"WORKFLOW_METRICS_DIR": tmp} - with mock.patch.dict("os.environ", env, clear=False): - with mock.patch.object(sys, "stdin", StringIO(json.dumps(payload))): - rc = workflow_event.main() - self.assertEqual(rc, 0) - files = list(Path(tmp).glob("*.jsonl")) - self.assertEqual(len(files), 1) - line = files[0].read_text(encoding="utf-8").strip() - record = json.loads(line) - self.assertEqual(record["tool_name"], "Bash") - self.assertEqual(record["input_summary"]["command_head"], "python3") - self.assertNotIn("fixture", line, "basename leaked into telemetry line") - self.assertNotIn("/tmp/fixture", line, "raw path leaked into telemetry line") - self.assertTrue(record["success"]) - - def test_writes_codex_exec_command_summary(self) -> None: - payload = { - "hook_event_name": "PreToolUse", - "tool_name": "functions.exec_command", - "tool_input": {"cmd": "python3 -m unittest"}, - "tool_use_id": "abc123", - "session_id": "session-xyz", - "cwd": str(ROOT), - } - with tempfile.TemporaryDirectory() as tmp: - env = {"WORKFLOW_METRICS_DIR": tmp} - with mock.patch.dict("os.environ", env, clear=False): - with mock.patch.object(sys, "stdin", StringIO(json.dumps(payload))): - rc = workflow_event.main() - self.assertEqual(rc, 0) - files = list(Path(tmp).glob("*.jsonl")) - self.assertEqual(len(files), 1) - line = files[0].read_text(encoding="utf-8").strip() - record = json.loads(line) - self.assertEqual(record["tool_name"], "functions.exec_command") - self.assertEqual(record["input_summary"]["command_head"], "python3") - self.assertIsNotNone(record["input_summary"]["command_hash"]) - - def test_concurrent_local_appends_preserve_one_linear_hash_chain(self) -> None: - worker_count = 16 - ctx = multiprocessing.get_context("fork") - start = ctx.Event() - ready = ctx.Queue() - results = ctx.Queue() - - with tempfile.TemporaryDirectory() as tmp: - log_path = Path(tmp) / "2026-07-19.jsonl" - - def append_worker(index: int) -> None: - event = { - "as_of": "2026-07-19", - "event": "PreToolUse", - "logged_at": "2026-07-19T12:00:00-04:00", - "sequence": index, - } - ready.put(index) - start.wait() - results.put( - workflow_event.post_or_append_event( - event, - log_path, - allow_post=False, - ) - ) - - workers = [ - ctx.Process(target=append_worker, args=(index,)) - for index in range(worker_count) - ] - for worker in workers: - worker.start() - for _ in workers: - ready.get(timeout=5) - start.set() - for worker in workers: - worker.join(timeout=5) - self.assertFalse(worker.is_alive(), "concurrent writer did not finish") - self.assertEqual(worker.exitcode, 0) - - self.assertTrue(all(results.get(timeout=1) for _ in workers)) - records = [ - json.loads(line) - for line in log_path.read_text(encoding="utf-8").splitlines() - if line.strip() - ] - - self.assertEqual(len(records), worker_count) - self.assertEqual({record["sequence"] for record in records}, set(range(worker_count))) - previous = "genesis-seed" - for record in records: - stored = record.pop("hash_chain") - self.assertEqual(stored, workflow_event.compute_hash(record, previous)) - previous = stored - - def test_local_append_does_not_chain_past_a_corrupt_tail(self) -> None: - event = {"as_of": "2026-07-19", "event": "PreToolUse"} - with tempfile.TemporaryDirectory() as tmp: - log_path = Path(tmp) / "2026-07-19.jsonl" - corrupt = b'{"event":"broken"\n' - log_path.write_bytes(corrupt) - with mock.patch.object(sys, "stderr", StringIO()): - success = workflow_event.post_or_append_event( - event, - log_path, - allow_post=False, - ) - self.assertFalse(success) - self.assertEqual(log_path.read_bytes(), corrupt) - - def test_local_append_rejects_an_oversized_record(self) -> None: - event = { - "as_of": "2026-07-19", - "event": "PreToolUse", - "unexpected": "x" * workflow_event.MAX_EVENT_LINE_BYTES, - } - with tempfile.TemporaryDirectory() as tmp: - log_path = Path(tmp) / "2026-07-19.jsonl" - with mock.patch.object(sys, "stderr", StringIO()): - success = workflow_event.post_or_append_event( - event, - log_path, - allow_post=False, - ) - self.assertFalse(success) - self.assertEqual(log_path.read_bytes(), b"") - - def test_oversized_hook_payload_is_skipped_without_creating_a_log(self) -> None: - payload = "x" * (workflow_event.MAX_HOOK_PAYLOAD_CHARS + 1) - with tempfile.TemporaryDirectory() as tmp: - metrics = Path(tmp) / "metrics" - with ( - mock.patch.dict( - "os.environ", - {"WORKFLOW_METRICS_DIR": str(metrics)}, - clear=False, - ), - mock.patch.object(sys, "stdin", StringIO(payload)), - mock.patch.object(sys, "stderr", StringIO()), - ): - rc = workflow_event.main() - self.assertEqual(rc, 0, "telemetry rejection must not block the client") - self.assertFalse(metrics.exists()) - - def test_permission_request_is_nonblocking_telemetry_only(self) -> None: - payload = { - "hook_event_name": "PermissionRequest", - "tool_name": "Bash", - "tool_input": {"command": "git push"}, - "session_id": "permission-session", - "cwd": str(ROOT), - } - with tempfile.TemporaryDirectory() as tmp: - temp_root = Path(tmp) - copied = temp_root / "bin" / "workflow-event" - copied.parent.mkdir(parents=True) - shutil.copy2(EVENT_PATH, copied) - metrics = temp_root / "metrics" - env = os.environ.copy() - env["WORKFLOW_METRICS_DIR"] = str(metrics) - env.pop("PYTEST_CURRENT_TEST", None) - started = time.perf_counter() - proc = subprocess.run( - [sys.executable, str(copied), "--client", "codex"], - input=json.dumps(payload), - text=True, - capture_output=True, - timeout=2, - env=env, - check=False, - ) - elapsed = time.perf_counter() - started - self.assertEqual(proc.returncode, 0, proc.stderr) - self.assertLess(elapsed, 1.0) - self.assertFalse( - (temp_root / "20_live" / "workflow-metrics" / "approvals").exists() - ) - record = json.loads(next(metrics.glob("*.jsonl")).read_text()) - self.assertEqual(record["event"], "PermissionRequest") - self.assertEqual(record["notification_kind"], "permission") - - def test_antigravity_agent_env_override(self) -> None: - payload = { - "hook_event_name": "PreToolUse", - "tool_name": "Read", - "tool_input": {"file_path": "README.md"}, - "session_id": "session-xyz", - "cwd": str(ROOT), - } - with tempfile.TemporaryDirectory() as tmp: - env = {"WORKFLOW_METRICS_DIR": tmp, "ANTIGRAVITY_AGENT": "1"} - with mock.patch.dict("os.environ", env, clear=True): - with mock.patch.object(sys, "stdin", StringIO(json.dumps(payload))): - rc = workflow_event.main() - self.assertEqual(rc, 0) - files = list(Path(tmp).glob("*.jsonl")) - self.assertEqual(len(files), 1) - line = files[0].read_text(encoding="utf-8").strip() - record = json.loads(line) - self.assertEqual(record["client"], "antigravity") - - def test_explicit_client_wins_over_parent_inference(self) -> None: - payload = { - "hook_event_name": "SessionStart", - "session_id": "session-xyz", - "cwd": str(ROOT), - } - with tempfile.TemporaryDirectory() as tmp: - env = {"WORKFLOW_METRICS_DIR": tmp, "ANTIGRAVITY_AGENT": "1"} - with mock.patch.dict("os.environ", env, clear=True): - with mock.patch.object( - workflow_event, - "get_client_override", - return_value="antigravity", - ): - with mock.patch.object( - sys, - "argv", - ["workflow-event", "--client", "local"], - ): - with mock.patch.object( - sys, "stdin", StringIO(json.dumps(payload)) - ): - rc = workflow_event.main() - self.assertEqual(rc, 0) - record = json.loads( - next(Path(tmp).glob("*.jsonl")).read_text(encoding="utf-8") - ) - self.assertEqual(record["client"], "local") - - def test_writes_redacted_grok_hook_event(self) -> None: - payload = { - "hookEventName": "pre_tool_use", - "sessionId": "grok-session", - "workspaceRoot": str(ROOT), - "toolName": "run_terminal_command", - "toolInput": { - "command": "SECRET=/tmp/private python3 -m unittest", - }, - "toolUseId": "tool-1", - } - with tempfile.TemporaryDirectory() as tmp: - env = {"WORKFLOW_METRICS_DIR": tmp} - with mock.patch.dict("os.environ", env, clear=False): - with mock.patch.object( - sys, - "argv", - ["workflow-event", "--client", "grok"], - ): - with mock.patch.object( - sys, - "stdin", - StringIO(json.dumps(payload)), - ): - rc = workflow_event.main() - - record = json.loads( - next(Path(tmp).glob("*.jsonl")).read_text(encoding="utf-8") - ) - - self.assertEqual(rc, 0) - self.assertEqual(record["client"], "grok") - self.assertEqual(record["event"], "PreToolUse") - self.assertEqual(record["input_summary"]["command_head"], "python3") - self.assertNotIn("/tmp/private", json.dumps(record)) - - -class HooksConfigurationTests(unittest.TestCase): - def test_grok_hooks_valid_json(self) -> None: - hooks_path = ROOT / ".grok" / "hooks" / "mainframe-telemetry.json" - self.assertTrue(hooks_path.exists()) - with hooks_path.open("r", encoding="utf-8") as handle: - data = json.load(handle) - self.assertIn("SessionStart", data["hooks"]) - self.assertIn("SessionEnd", data["hooks"]) - commands = [ - hook.get("command", "") - for groups in data["hooks"].values() - for group in groups - for hook in group.get("hooks", []) - ] - self.assertTrue( - all("--client grok --pixel main-grok-build" in command for command in commands) - ) - - def test_tracked_claude_hooks_are_valid_and_nonblocking(self) -> None: - hooks_path = ROOT / ".claude" / "settings.json" - self.assertTrue(hooks_path.exists()) - with hooks_path.open("r", encoding="utf-8") as f: - data = json.load(f) - self.assertIn("hooks", data) - self.assertIn("SessionEnd", data["hooks"]) - permission_hooks = [ - hook - for group in data["hooks"]["PermissionRequest"] - for hook in group.get("hooks", []) - ] - self.assertTrue(permission_hooks) - self.assertTrue( - all("--client claude" in hook.get("command", "") for hook in permission_hooks) - ) - self.assertTrue(all(hook.get("timeout", 0) <= 5 for hook in permission_hooks)) - - -class DetectLoopTests(unittest.TestCase): - def test_detect_loop_no_file(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - p = Path(tmp) / "nonexistent.jsonl" - res = workflow_event.detect_loop(p, "session-1") - self.assertEqual(res, {"is_loop": False, "loop_count": 0}) - - def test_detect_loop_no_loop(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - p = Path(tmp) / "telemetry.jsonl" - events = [ - {"session_hash": "session-1", "success": True, "event": "PostToolUse"}, - {"session_hash": "session-1", "success": True, "event": "PostToolUse"}, - ] - with p.open("w", encoding="utf-8") as f: - for ev in events: - f.write(json.dumps(ev) + "\n") - res = workflow_event.detect_loop(p, "session-1") - self.assertEqual(res, {"is_loop": False, "loop_count": 0}) - - def test_detect_loop_with_loop(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - p = Path(tmp) / "telemetry.jsonl" - events = [ - {"session_hash": "session-1", "success": True, "event": "PostToolUse"}, - {"session_hash": "session-1", "success": False, "event": "PostToolUse"}, - {"session_hash": "session-1", "success": False, "event": "PostToolUseFailure"}, - ] - with p.open("w", encoding="utf-8") as f: - for ev in events: - f.write(json.dumps(ev) + "\n") - res = workflow_event.detect_loop(p, "session-1") - self.assertEqual(res, {"is_loop": True, "loop_count": 2}) - - -class CalculateDurationTests(unittest.TestCase): - def test_calculate_duration_no_file(self) -> None: - p = Path("/nonexistent/file.jsonl") - res = workflow_event.calculate_duration_from_pre_tool( - p, "session-1", "tool-1", datetime.now() - ) - self.assertIsNone(res) - - def test_calculate_duration_no_match(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - p = Path(tmp) / "telemetry.jsonl" - events = [ - { - "session_hash": "session-1", - "tool_use_hash": "tool-2", - "event": "PreToolUse", - "logged_at": "2026-06-18T00:00:00-04:00", - } - ] - with p.open("w", encoding="utf-8") as f: - for ev in events: - f.write(json.dumps(ev) + "\n") - res = workflow_event.calculate_duration_from_pre_tool( - p, - "session-1", - "tool-1", - datetime.fromisoformat("2026-06-18T00:00:05-04:00"), - ) - self.assertIsNone(res) - - def test_calculate_duration_success(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - p = Path(tmp) / "telemetry.jsonl" - events = [ - { - "session_hash": "session-1", - "tool_use_hash": "tool-1", - "event": "PreToolUse", - "logged_at": "2026-06-18T00:00:00-04:00", - } - ] - with p.open("w", encoding="utf-8") as f: - for ev in events: - f.write(json.dumps(ev) + "\n") - res = workflow_event.calculate_duration_from_pre_tool( - p, - "session-1", - "tool-1", - datetime.fromisoformat("2026-06-18T00:00:05-04:00"), - ) - self.assertEqual(res, 5000) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_workflow_report.py b/tests/test_workflow_report.py deleted file mode 100644 index c872ca0..0000000 --- a/tests/test_workflow_report.py +++ /dev/null @@ -1,354 +0,0 @@ -from __future__ import annotations - -import importlib.util -import json -import sys -import tempfile -import unittest -from importlib.machinery import SourceFileLoader -from pathlib import Path - - -ROOT = Path(__file__).resolve().parents[1] -REPORT_PATH = ROOT / "bin" / "workflow-report" -LOADER = SourceFileLoader("workflow_report", str(REPORT_PATH)) -SPEC = importlib.util.spec_from_loader(LOADER.name, LOADER) -assert SPEC and SPEC.loader -workflow_report = importlib.util.module_from_spec(SPEC) -sys.modules[SPEC.name] = workflow_report -SPEC.loader.exec_module(workflow_report) - - -class WorkflowReportTests(unittest.TestCase): - def test_build_summary_reports_coverage_and_pairing(self) -> None: - events = [ - { - "event": "SessionStart", - "session_hash": "a", - "client": "codex", - }, - { - "event": "PreToolUse", - "session_hash": "a", - "client": "codex", - "tool_name": "Bash", - }, - { - "event": "PostToolUse", - "session_hash": "a", - "client": "codex", - "tool_name": "Bash", - "success": True, - "duration_ms": 20, - }, - { - "event": "SessionEnd", - "session_hash": "a", - "client": "codex", - }, - { - "event": "Diagnostic", - "session_hash": "a", - "client": "codex", - "diagnostic": "context_limit_exceeded", - }, - { - "event": "SessionStart", - "session_hash": "b", - "client": None, - }, - { - "event": "PreToolUse", - "session_hash": "b", - "client": None, - "tool_name": "Read", - }, - ] - - summary = workflow_report.build_summary(events, days=7) - quality = summary["telemetry_quality"] - - self.assertEqual(summary["events"], 7) - self.assertEqual(summary["sessions"], 2) - self.assertEqual(summary["tool_failures"], 0) - self.assertEqual(quality["client_tagged_events"], 5) - self.assertEqual(quality["duration_coverage_pct"], 100.0) - self.assertEqual(quality["session_close_coverage_pct"], 50.0) - self.assertEqual(quality["unbalanced_tool_sessions"], 1) - self.assertEqual(quality["unmatched_pre_tool_events"], 1) - self.assertEqual( - summary["diagnostics"], - [{"name": "context_limit_exceeded", "count": 1}], - ) - - def test_build_summary_segments_quality_by_client(self) -> None: - events = [ - {"event": "SessionStart", "session_hash": "a", "client": "codex"}, - { - "event": "PreToolUse", - "session_hash": "a", - "client": "codex", - "tool_name": "Bash", - }, - { - "event": "PostToolUse", - "session_hash": "a", - "client": "codex", - "tool_name": "Bash", - "success": False, - "duration_ms": 12, - }, - {"event": "SessionStart", "session_hash": "b", "client": "claude"}, - { - "event": "PostToolUse", - "session_hash": "b", - "client": "claude", - "tool_name": "Read", - "success": True, - }, - {"event": "SessionEnd", "session_hash": "b", "client": "claude"}, - ] - - summary = workflow_report.build_summary(events, days=7) - by_client = { - item["name"]: item - for item in summary["telemetry_quality"]["by_client"] - } - - self.assertEqual(set(by_client), {"codex", "claude"}) - - codex = by_client["codex"] - self.assertEqual(codex["events"], 3) - self.assertEqual(codex["sessions"], 1) - self.assertEqual(codex["sessions_with_start"], 1) - self.assertEqual(codex["sessions_with_end"], 0) - self.assertEqual(codex["session_close_coverage_pct"], 0.0) - self.assertEqual(codex["duration_coverage_pct"], 100.0) - self.assertEqual(codex["tool_failures"], 1) - self.assertEqual( - codex["event_types"], - ["PostToolUse", "PreToolUse", "SessionStart"], - ) - - claude = by_client["claude"] - self.assertEqual(claude["session_close_coverage_pct"], 100.0) - self.assertEqual(claude["duration_coverage_pct"], 0.0) - self.assertEqual(claude["tool_failures"], 0) - - # Per-client list is sorted by event volume, descending. - names = [ - item["name"] - for item in summary["telemetry_quality"]["by_client"] - ] - self.assertEqual(names, ["codex", "claude"]) - - def test_load_events_skips_malformed_and_non_object_json_values(self) -> None: - with tempfile.TemporaryDirectory() as tmp: - log_dir = Path(tmp) - path = log_dir / "2099-01-01.jsonl" - path.write_text( - "\n".join( - [ - json.dumps({"event": "SessionStart"}), - "not-json", - "[]", - "null", - '"string"', - "42", - ] - ) - + "\n", - encoding="utf-8", - ) - events = workflow_report.load_events(log_dir, days=30000) - - self.assertEqual(events, [{"event": "SessionStart"}]) - - def test_empty_summary_has_no_division_errors(self) -> None: - summary = workflow_report.build_summary([], days=1) - quality = summary["telemetry_quality"] - - self.assertEqual(summary["events"], 0) - self.assertIsNone(quality["client_tag_coverage_pct"]) - self.assertIsNone(quality["duration_coverage_pct"]) - self.assertIsNone(quality["session_close_coverage_pct"]) - self.assertEqual(quality["by_client"], []) - - def test_aider_outcomes_do_not_count_as_pairing_defects(self) -> None: - summary = workflow_report.build_summary( - [ - { - "event": "PostToolUse", - "session_hash": "a", - "client": "local", - "source": "aider", - "tool_name": "Edit", - "success": True, - } - ], - days=1, - ) - quality = summary["telemetry_quality"] - self.assertEqual(quality["unmatched_outcome_events"], 0) - self.assertEqual(quality["outcome_only_source_events"], 1) - - def test_process_rollup_is_opt_in(self) -> None: - events = [ - { - "event": "PostToolUse", - "session_hash": "a", - "client": "codex", - "tool_name": "Bash", - "success": True, - "process_id": "cli-workflow-report", - } - ] - - default_summary = workflow_report.build_summary(events, days=1) - self.assertNotIn("processes", default_summary) - self.assertEqual( - default_summary["telemetry_quality"]["process_tag_coverage_pct"], - 100.0, - ) - - process_summary = workflow_report.build_summary( - events, - days=1, - by_process=True, - ) - self.assertEqual( - process_summary["processes"], - [ - { - "process_id": "cli-workflow-report", - "events": 1, - "sessions": 1, - "tool_failures": 0, - "event_types": ["PostToolUse"], - "clients": [{"name": "codex", "count": 1}], - } - ], - ) - - def test_process_rollup_counts_multi_process_events_once_per_process(self) -> None: - summary = workflow_report.build_summary( - [ - { - "event": "PostToolUse", - "session_hash": "a", - "client": "codex", - "success": False, - "process_ids": [ - "cli-workflow-event", - "workflow-workflow-telemetry", - ], - } - ], - days=1, - by_process=True, - ) - by_process = { - item["process_id"]: item - for item in summary["processes"] - } - - self.assertEqual( - set(by_process), - {"cli-workflow-event", "workflow-workflow-telemetry"}, - ) - self.assertEqual(by_process["cli-workflow-event"]["tool_failures"], 1) - self.assertEqual(by_process["workflow-workflow-telemetry"]["events"], 1) - - def test_process_rollup_ignores_unsafe_process_ids(self) -> None: - summary = workflow_report.build_summary( - [ - { - "event": "PostToolUse", - "session_hash": "a", - "client": "codex", - "process_ids": [ - "cli-workflow-report", - "/tmp/example-mainframe/bin/workflow-report", - "ran private command", - "client-name-sensitive-case", - ], - } - ], - days=1, - by_process=True, - ) - - self.assertEqual( - summary["processes"], - [ - { - "process_id": "cli-workflow-report", - "events": 1, - "sessions": 1, - "tool_failures": 0, - "event_types": ["PostToolUse"], - "clients": [{"name": "codex", "count": 1}], - } - ], - ) - - def test_input_signals_are_opt_in_aggregate_counts(self) -> None: - events = [ - { - "event": "PreToolUse", - "tool_name": "Bash", - "input_summary": { - "command_head": "git", - "command_hash": "not-reported", - }, - }, - { - "event": "PreToolUse", - "tool_name": "Read", - "input_summary": { - "path_zone": "30_projects", - "extension": ".md", - }, - }, - { - "event": "PreToolUse", - "tool_name": "Read", - "input_summary": { - "path_zone": "30_projects", - "extension": ".py", - }, - }, - { - "event": "PreToolUse", - "tool_name": "Other", - "input_summary": { - "input_hash": "not-reported", - }, - }, - ] - - default_summary = workflow_report.build_summary(events, days=1) - self.assertNotIn("input_signals", default_summary) - - summary = workflow_report.build_summary( - events, - days=1, - input_signals=True, - ) - - self.assertEqual( - summary["input_signals"], - { - "redacted_input_events": 4, - "command_heads": [{"name": "git", "count": 1}], - "path_zones": [{"name": "30_projects", "count": 2}], - "file_extensions": [ - {"name": ".md", "count": 1}, - {"name": ".py", "count": 1}, - ], - }, - ) - - -if __name__ == "__main__": - unittest.main()