Skip to content

Latest commit

聽

History

History
94 lines (78 loc) 路 5.4 KB

File metadata and controls

94 lines (78 loc) 路 5.4 KB

dsh-pdf (alpha.8)

PDF the agent can actually read and scan, and a real PDF reader in the right bar.

A file-read tool returns binary noise for a PDF and grep finds nothing in it, so this package makes one legible: five tools over a vendored pdf.js 6.3.289 (the same version and legacy build the harness's own preview inlines), plus the right bar's pdf tab type, which outranks the shipped bare PDF renderer for *.pdf. It is read-only, and document text is data, never instructions.

What it adds

Tool The question it answers
pdf_info pages, sizes, metadata, outline, attachments, forms/signatures, encryption - and per page whether there is a text layer at all
pdf_read a page range in reading order (text) or as a geometric layout reconstruction that keeps columns and label/value pairs apart
pdf_find a literal or regex, with page, line and context
pdf_render page pictures through the host's own rasterizer, written as new PNGs
pdf_scan recognize the pages that are a picture of text
  • The scanner is one pipeline shared by pdf_scan and the reader's own "scan this page" (POST /api/dsh-pdf/scan). With no pages it scans exactly the pages pdf_info found to have no text layer, and the answer is labelled a transcription.
  • Two host engines, neither bundled: a rasterizer (poppler pdftoppm, mutool or Ghostscript) and tesseract (+ language data), probed from PATH, spawned with argv arrays only and killed on a deadline; a language is validated against the engine's own --list-langs before a page is drawn.
  • Cached under every input that can change it: content hash + page, language, dpi and segmentation mode (ocr/<n>.<lang>@<dpi>dpi.p<psm>.txt, raster at images/<n>@<dpi>dpi.png) under $DSH_HOME/dsh-pdf/artifacts/<sha256>/.
  • The reader: lazy continuous pages, a zoom ladder plus fit width / fit page (zoom moves the LAYOUT, never a CSS transform), navigation by button, field and keyboard, rotation, pdf.js's own TextLayer against --total-scale-factor / --scale-round-x/y, <mark>-wrapped search highlighting, drag-to-pan, Retry on failure.
  • The side panel: a rail of thumbnails drawn on demand (IntersectionObserver rooted on the rail) and the document's own bookmark outline through getOutline / getDestination / getPageIndex, an unresolvable bookmark shown disabled.
  • The workspace index (sidebar://pdfs, guide order 50) lists every PDF in the conversation's workspace with size, mtime and an opt-in page count, and a row opens the document through the ordinary openResource action.
  • Engine and hardening: lib/vendor/ is generated by vendor/build.mjs, whose VERSION.json records every file's bytes and sha256 plus a digest per tree that --check recomputes; the engine is fetched as blob URLs, never bundled, and extraction runs in a child process with isEvalSupported: false, useWorkerFetch: false, enableXfa: false, useSystemFonts: false and disableFontFace: true; the pdf-analysis skill is registered at runtime and copied into $DSH_HOME/skills.

How it plugs in

Piece Value
row / id pdf / dsh-pdf; the index type is dsh-pdf-index
tab kinds pdf (patterns: ['*.pdf'], priority extension) and pdfs (priority builtin, no patterns)
seats / addresses sidebar.right.pane.tab and .title per type, one tool.call.toolview per tool; dsh-resource://file/session/<id>/<path>, dsh-resource://pdf/absolute/<whole-path-encoded>

The pdf type wins by ranking - extension (3) beats the shipped preview's fallback (1) - and canOpen refuses anything but a .pdf, so every other file type keeps its surface. Ten exact routes, GET/HEAD/POST only: GET /api/dsh-pdf/state, /health (capabilities, including the OCR language list), /file (the bytes, with x-dsh-pdf-sha256; HEAD never reads the file), /list, POST /scan (a capability refusal answers 200 {ok:false, reason, message}) and /vendor/{pdf.min.mjs, pdf.worker.min.mjs, cmaps.json, standard-fonts.json, wasm.json}.

Limits

Cap Value
document / extraction 512 MiB (tools) / 256 MiB (tab); 200 pages per extraction
pdf_read / pdf_info 40 000 chars default, 200 000 max; every page up to 60, then a 20-page sample
pdf_render / pdf_scan 20 pages a call at 50-400 dpi; 10 pages a call at 50-400 dpi (default 200), psm 0-13 (default 3)
pdf_find / index / panel / cache 40 hits, 2000 pages, 45 s; depth 6, 200 files, page counts for 12; thumbnails for 300 pages, 200 outline entries over 4 levels; 512 MiB LRU

A session-relative path is realpath-checked inside the conversation workspace (a symlink out is refused, not followed); an absolute path is read directly (a chat attachment, a file in Downloads); either way the target must be a regular .pdf inside the ceiling. pdf_render writes create-exclusively under a name this plugin generates.

Verify

node scripts/checks/check-pdf-node.mjs         # builds its own PDFs: no TeX, no rasterizer, no network
node scripts/checks/check-client-bundles.mjs
node scripts/checks/check-media-examples.mjs   # every pdf_* example in the skill
node packages/dsh-pdf/vendor/build.mjs --check

Install

Both installers pick the package up from packages/:

scripts\install.bat -Force        # Windows
./scripts/install.sh -Force       # macOS / Linux