v8.0.0 — Spec-Driven Agent, modular backend, and a security pass over generated code - #604
Merged
Merged
Conversation
1. push-to-github 404'd for every worker run.
restore_persisted() runs once, in the lifespan. That was fine while the
process that produced a run also served its download and its push. Since
generation moved to besser-wme-smartgen it is not: push-to-github and
import-github-run stay on the backend because they need its process-local
OAuth session, so every worker run is written to the shared volume AFTER
the backend's one-shot scan. Proven live: worker had 1 run on disk, the
backend registry had 0 entries.
get() now falls back to a TARGETED rescan for the single run_id it was
asked about, sharing the manifest parse with restore_persisted. A cached
hit still touches no disk, a malformed id never reaches os.listdir, and an
expired run stays gone.
An earlier commit claimed this was fixed and verified. It was not: I had
checked that the backend could SEE the files, which is a different thing
from the endpoint resolving the run.
2. A modelled default could escape into an executable os.system() call.
backend/templates/router.py.j2 renders param.default_value in four
executable places and two docstrings. The executable ones were hardened
first; python_default never raises for a str, so generation SUCCEEDED and a
value containing a triple-double-quote closed the generated docstring and
dropped statements at body indentation -- in a router /deploy-app executes.
Verified: reverting the fix makes the new test report the escape at a
concrete line.
That was the third site of this one sink in this branch, after the
SQLAlchemy macro and its association-class copy.
The same lines also passed param[1], the RENDERED ANNOTATION ('Any', or an
FK class name) rather than the modelled type, so python_default fell to its
quoted fallback and emitted '5' for an int default. The tuple now carries
param.type.name and both uses read it.
3. A symlink in a run workspace could exfiltrate the backend's secrets.
github_service collected files with rglob + is_file(), which FOLLOWS
symlinks. The tree being pushed is written by the worker, which runs
model-authored code with shell tools enabled; the backend doing the push
holds SMTP_PASSWORD, GITHUB_CLIENT_SECRET, GPG_KEY and the rest. `ln -s
/proc/self/environ leak.txt` would have been read through and committed to
the user's repository. :ro blocks writes, not traversal.
Symlinks are skipped and every survivor is re-checked with commonpath.
This was not exploitable while (1) was broken -- which is exactly why it
ships in the same commit as the fix that would have armed it.
The keyless picker listed BESSER_FREE_LLM_FALLBACK_MODEL as a choosable model, which put the self-hosted box directly in front of users. That box serves ONE request at a time: a parallel batch measured on 2026-09-11 drove two of five runs to zero turns, and the two concurrent runs today showed the same shape -- one finished in 203s while the other queued behind it for 1342s and died on the runtime cap. It is still configured as the fallback and still runs automatically when the primary keyless endpoint fails -- which is not hypothetical, poolside /laguna-s-2.1-free is returning "Upstream model provider is temporarily unavailable" right now. Only its appearance in the picker is removed. The two config tests that asserted it WAS offered are updated rather than relaxed: they now assert it is absent from the choices AND that free_fallback_model() still returns it, so hiding it from the picker cannot silently disable the failover.
…nds on it _get_pricing asks _is_free_local_model first, and a True answer returns _ZERO_PRICING. Every call then costs $0, so max_cost_usd can never trip: a run continues to the turn or time cap while the provider bills normally and the run card reads $0.00 throughout. The rule matched FAMILY NAMES -- qwen, deepseek, llama, gemma, mistral-small and so on. That was right while open-weight implied self-hosted. It stopped being right when Command Code began reselling those same families: Qwen/Qwen3.8-Max, Qwen/Qwen3.8-Flash, deepseek/deepseek-v4-pro and deepseek/deepseek-v4-flash are all billed models on the plan, and all four priced at $0. Selecting one as a paid option would have spent real money with the cap inert. The decision now keys on how a model is IDENTIFIED rather than its family: 1. explicit free marker (:free / -free) -> free 2. Ollama name:tag with no vendor namespace -> self-hosted, free 3. vendor/model (a gateway id) -> billed 4. otherwise the family list, which still covers tagless self-hosted ids Verified against the real ids in play: qwen3-coder:30b, LongCat-2.0:free, laguna-s-2.1-free and ling-3.0-flash-sante:free stay free; Qwen3.8-Max, deepseek-v4-pro, Kimi-K3, GLM-5.3, MiniMax-M3 and gemini-3.8-flash become billed. The gateway models fall to the gpt-4o default, which is a guess -- but a guess that keeps the cap working, where $0 switched it off.
…nu item" This reverts commit cdd2c0e.
BESSER_FREE_LLM_PILOT_MODEL names the keyless model a facilitated pilot session should start on; /spec-driven/config exposes it as free_tier.pilot_model (null when unset, so pilots and the public get the same default). free_pilot_model() refuses an id the server does not actually offer. Advertising a default we would decline to honour is worse than having none: the client sends the pre-selected id as llm_model, an unknown id is pinned back to the default, and the pilot would silently run on the public model while the UI showed something else. Server-side because the choice is volatile -- it changed three times in one afternoon -- so it should be an env edit and a container restart. Bumps the frontend submodule for the matching client change.
/spec-driven/config is served by the worker, which uses an explicit env allowlist rather than env_file. A variable missing from that list reads as unset, so the pilot default silently did nothing -- the third feature this allowlist has swallowed, after the alt-model list and run telemetry. The allowlist is still right (it cut the generation path from nine secrets to three) but it fails silently, so every new BESSER_* var the spec-driven path reads has to be added here deliberately.
Server side of surviving a TLS-inspecting corporate proxy (Netskope on
LIST laptops). The proxy buffers a response body before releasing it, and
an SSE body never ends, so the run stream is the one endpoint it breaks —
REST and the agent WebSocket both work through it.
GET /spec-driven/runs/{id}/events.json returns the same durable events as
a short, terminating JSON response the proxy can release normally. It
needed almost no new machinery: the event store already assigns sequence
numbers before any subscriber sees a frame, so polling and streaming share
one cursor and one replay semantic. Payloads are decoded so one client
reducer serves both transports; `hasMore` lets a poller drain a backlog
without sleeping through its interval.
POST /spec-driven/generate now honours an Idempotency-Key. Starting a run
is not idempotent, which is why the client never retried a start that
failed at the transport layer — and behind such a proxy that is exactly the
request that fails, stranding the user with a run they cannot rejoin. The
key is claimed before the run starts, so a retry arriving mid-startup waits
for that run rather than opening a second one, and is released if startup
fails. In-process and best-effort: a missed hit costs one duplicate run,
never corruption, and the durable store stays the source of truth.
Also fixes the free tier 400-ing on gpt-5.6. reasoning_effort='none' is
required by the official OpenAI API for tools on those models, but a
gateway re-serving them validates the value against its own enum, and
'none' is not in it — api.commandcode.ai answers `Invalid option: expected
one of "low"|"medium"|"high"|"xhigh"|"max"`, which aborted the tool-driven
Phase 2 loop and shipped a deterministic-only bundle flagged incomplete.
Verified against gpt-5.6-luna there: 'none' -> 400, while omitted / 'low' /
'medium' each returned a proper tool call. So the flag is now sent only on
the official endpoint and omitted on any custom base_url.
…njection Turn cap 120 -> 150. A turn is a unit of round-trip, not of work: a model that batches tool calls does several actions per turn, while one that emits a single call per turn needs roughly 4x the turns for the same result. Two measured runs bracket it — 30 calls in 14 turns (2.1/turn, finished at 17% of budget) against 2026-09-11's 80 turns / 80 calls / 18 files, which consumed the whole budget for less output. The ceiling has to fit the worst ratio or a non-batching model hits it mid-build having done nothing wrong. Cost and runtime are still checked at every turn boundary, so a higher turn ceiling cannot raise spend past the cost cap. Docs updated for the new ceiling, and two pre-existing wrong numbers in the same list corrected: BESSER_LLM_DEFAULT_MAX_COST_USD is 5.0 and BESSER_LLM_DEFAULT_MAX_RUNTIME_SECONDS is 1200, not the documented 1.0/600. runs.rst already had the right values, so the two pages contradicted. Corporate CA injection is now opt-in behind --build-arg TRUST_EXTRA_CAS=1 (Dockerfile). deploy.sh builds from the local working tree, so the unconditional COPY would have baked whatever .crt sits in ca-certs-extra/ into the image shipped to the shared host — making the deployed backend trust a corporate CA for every outbound TLS call. Defaulting to 0 keeps production clean by construction rather than by remembering to empty a directory. The certs themselves stay gitignored; only .gitkeep is tracked.
…es to 4
Turn default 80 -> 120. The cap was never the spend guard: cost and runtime
are checked at every turn boundary before the next billable call, so a
higher turn ceiling cannot raise what a run spends. What it does fix is the
model-dependence — a turn is a unit of round-trip, not of work. Measured:
one run did 30 tool calls in 14 turns, while 2026-09-11 did 80 calls in 80
turns for 18 files. At 1 call/turn, 80 turns strangles a model that has done
nothing wrong.
Truncation retries 2 -> 4. This budget is a per-run TOTAL that deliberately
never resets after a clean turn, and it ends Phase 2 outright. At 2, a run
that truncated once early, adapted, produced a hundred good turns and then
truncated once more was killed one bad turn from done — while the stated
rationale ("beyond that it is not adapting") had been false for a hundred
turns. Raising the turn budget makes that sharper, so the two move together.
_MAX_PARALLEL_WORKERS stays at 4, deliberately. The deploy host is 2 vCPU /
1.9GB RAM with ~450MB free, and LLM_MAX_CONCURRENT_RUNS is 10 — already up
to 40 concurrent tool executions on two cores. The CPU-bound tools cannot
outrun the cores, so raising it buys contention, not throughput.
preview.py no longer hardcodes 80 with a comment claiming it "matches
LLMOrchestrator.MAX_TURNS"; it reads LLM_DEFAULT_MAX_TURNS, so the estimate
cannot silently drift from the budget again. Docs updated for both values,
including the truncation budget's non-resetting semantics, and a stale
comment claiming the cost default is $1 (it is $5) removed.
The Django template collapsed newlines in two places, so any realistic
model produced a models.py that fails ast.parse. Pre-existing on master,
not a regression from this branch.
Enumerations: `{%- for %}` / `{%- endfor %}` strip the preceding newline
and trim_blocks strips the following one, so every literal and every enum
class ran together on one line:
class LoanStatus(models.TextChoices): CLOSED = 'CLOSED', ...class MemberStatus(...)
Relationship fields: `{%- endif -%}` ate the newline after a field's
closing paren, so two consecutive ForeignKeys on one class joined into
`null=True) member = models.ForeignKey(` — a statement-level join, not
keyword arguments sharing a line. Same tag in the ManyToMany and OneToOne
blocks, so all three are fixed. This one is independent of the enum defect
and reproduces on an enum-free model.
Whitespace control only; trim_blocks/lstrip_blocks are untouched because
every other block in the template depends on them.
The generator's own subprocesses also leaked bytecode into the download:
`manage.py startapp` imports the settings it just wrote, so CPython left
new_project/new_project/__pycache__/*.pyc in the packaged tree. They now
run with PYTHONDONTWRITEBYTECODE, so nothing is written to clean up.
Existing Django compile tests all used enum-free, single-FK models, which
is why this shipped. The new tests use both shapes and assert ast.parse.
The comments accumulated across the spec-driven work had grown into 15-20 line blocks retelling a debugging story around a single fact. Trimmed to the fact: ~200 fewer comment lines across 23 files, the largest share in llm/orchestrator.py, and duplicated rationale collapsed to one location instead of being restated at the constant, the docstring and the call site. Kept, deliberately: every dated measurement (the 62-turn livelock, the 40%-bookkeeping batch, the 128k prompt truncated to 65_536), the security rationale for a default being off, and the "why the obvious simpler version is wrong" warnings — those are what stop a future cleanup reintroducing the bug. Comments and docstrings only. Verified by parsing each changed file at HEAD and in the working tree, stripping docstrings, and comparing ast.dump(): identical for all 23.
…e comments A corporate CA is needed to FETCH packages when building from a laptop behind TLS inspection, but the deployed backend must not trust it for its own outbound calls. The final layer removes it and asserts it is gone — the grep fails the build rather than shipping an image that trusts the proxy. It runs unconditionally, so the guarantee holds whether or not TRUST_EXTRA_CAS was passed. That closes the gap the opt-in ARG alone left: without it, building behind the proxy meant choosing between a failed build and a shipped corporate CA. Comments trimmed 95 -> 72 lines. Instructions are unchanged apart from the new layer; the facts worth keeping (why 3.11+, why the toolchains must be on PATH, why ruff is pinned here) are kept, the narrative around them is not.
…ad Tailwind Two template-level defects, both deterministic, both invisible to the current validation because it checks files in isolation and nothing exercises a write or a stylesheet. 1. Every POST returned 500. pydantic_classes_template deliberately drops the surrogate id and the audit timestamps from Create -- the ORM stamps them, and its comment says leaving them in "breaks the generated app". router.py.j2 dropped only the id and read createdAt/updatedAt off the payload regardless, so FastAPI raised AttributeError on every create. Verified live on a generated hotel app: 7 of 9 entities returned 500 while the run was reported a clean success. The router now shares one _client_supplied macro that mirrors the pydantic rule, so the two cannot drift again. 2. The React generator declared tailwindcss ^4.1.14 and never wired it: no @tailwindcss/vite plugin, no @import, no config, and postcss+autoprefixer alongside -- the v3 toolchain for a v4 package. The generator's own components use zero utility classes so deterministic output looked fine, but an LLM reads package.json, concludes Tailwind is available, and writes utility classes that are inert. Observed: a nav rendered as a browser-default bulleted list because "flex gap-1 py-3" did nothing. Wiring it turned the same unchanged LLM output into a styled app. "type": "module" is required with it: @tailwindcss/vite is ESM-only and Vite otherwise tries to require() it and refuses to start. Found by booting the app -- a template test asserting the import line exists would have passed.
…nces No relationship could be set through the API of any generated app whose model uses inheritance. Observed live on a hotel app: Person carried a string uuid PK, Guest and Employee inherited it (joined-table -- their id IS a ForeignKey to person.id), but every FK pointing at Guest or Employee came out Mapped_[int], and the Pydantic field came out `employee: int`. POST /booking/ with a real id returned 422 "Input should be a valid integer". The schema was internally inconsistent too, not just the API layer. SQLite creates those tables regardless, so boot checks, imports and ast.parse all passed and the run was reported a clean success. Cause: the PK-type resolver looked only at a class's OWN attributes. A subclass has no id attribute of its own, so it was absent from the map and every caller fell back to the 'int' default. It now follows parents(). There were THREE copies of this logic. Both docstrings already described this exact failure -- "a Mapped_[int] FK pointing at a String PK breaks joins at runtime", "a guest: int Pydantic field 422-rejects every real id" -- and neither implementation handled inheritance. pk_types.pk_python_types now delegates to structural_utils.get_pk_py_types so there is one implementation, and the pydantic template stops hardcoding `int` for association ends.
Five defects, each traced to a live run and each covered by a test that was verified to fail against the pre-fix code. Neither entity could be created. A 1:1 association emitted a mandatory relationship field on BOTH create schemas, because the FK-owning and non-owning branches of the Pydantic template were byte-identical. BookingCreate demanded a guest, GuestCreate demanded a booking. The router never read the non-owning field either - it validates and assigns only the FK-owning end - so the client was forced to send an id that was then silently discarded. The N:1 branch already emitted nothing there; the 1:1 branch now matches it. Two existing tests asserted the discarded field and have been updated to assert its absence. Two tables each holding a NOT NULL foreign key to the other cannot be populated: POST /guest/ died on "NOT NULL constraint failed" before anything existed to point at. A cycle needs two associations with opposite FK ownership - one association yields only one FK column - so one edge has to give. get_deferred_fk_associations finds the cycles and breaks one edge each, chosen by association name so the choice is stable across regenerations. That FK becomes nullable with use_alter, its relationships get post_update, and the create schema and router stop demanding it. The three-step create flow the REST API actually performs now completes; before, the first step failed. tsc ran without node_modules, so every package import was unresolvable and the type errors cascaded from there. A frontend that boots and renders was reported as "3 blocker-level issues remain", headlined by TS2688 on vite/client. We never run npm install during validation - it needs network the host may not have - so instead the findings are demoted to advisory when node_modules is absent. One exception stays a blocker: a TS2307 naming a RELATIVE path, which is a genuine missing file either way. Every non-FastAPI app was reported as "cannot start". django and flask were simply absent from a hardcoded allowlist, as were argparse, sqlite3, glob and ~180 other stdlib modules. The stdlib half now comes from sys.stdlib_module_names, the framework families are listed, and - the part that generalises to all 20 generators - the app's own manifest is read. A package the app declares is a package the app has. Also fixed: a directory of .py files is importable from its PARENT, not only as a sibling of the importing file, so `from core.config import settings` in backend/routers/x.py no longer reports missing. On a django/flask/stdlib/import-core sample: 11 false blockers, now 0. The three genuine-defect cases still report. The missing-frontend blocker never fired on "build a web application": the trailing \b in \b(web ?app|...)\b cannot match when "app" continues into "lication". Nine of ten sweep runs shipped backend-only and reported success. "spa" is deliberately NOT an alternative - a hotel has one. Also: the gap analyser compared the generated code against the MODEL rather than against the REQUEST. On run ca48a6dd it returned four cosmetic tasks for a spec with twelve real gaps - two state machines, five methods, four OCL constraints and an association-class attribute were all invisible to it. It is the only stage that holds the request and the model side by side, so what it misses is lost for the rest of the run. It now performs an explicit two-pass diff, request against model first, and the system prompt states that the model is a lossy capture of the request rather than the authority. And MethodButtonProps lacked the `id` the page builder puts on every component it renders, which was 8 x TS2322 on every entity page.
Three changes, all from measuring a live run against the app it produced. A create schema and its router can disagree, and every disagreement is a guaranteed 500. On 2026-09-17 the LLM rewrote PersonCreate to add an email validator and dropped lastName; the deterministic router still read guest_data.lastName, so person, guest AND employee returned 500 - every way to get a row into the system. Phase 3 ran, found 29 issues and zero blockers. The new check is structural: for every X_data.<field> a router reads, the field must be declared on XCreate or one of its Create bases. Three findings on that app, none on a clean scaffold. It is the third instance of this shape today, after createdAt/updatedAt and a 1:1 relationship field, so it is a class of bug, not an accident. The agent could not see the file it most needed to edit. A generated routers/booking.py runs to 34,884 characters against a 20,000-character read limit - 59% visible - and write_file then asked it to reproduce the whole thing. Modeled-method bodies now render into their own <entity>_methods.py: booking.py drops to 24,911 with a self-contained 10,247-character methods module beside it. Routes are untouched - same paths, same tags, a second APIRouter that main_api includes - verified by diffing the OpenAPI surface before and after: 62 paths, identical, boots clean. Method bodies are the hand-written half of the generation gap, so giving them their own file also stops a botched rewrite taking 24k of CRUD with it. MAX_FILE_READ goes to 60,000. The largest generated file is ~8,700 tokens against an 80,000-token compaction threshold, so reading it whole costs about 11% of the budget - far cheaper than editing it blind. Also: the default runtime cap moves from 20 to 40 minutes. Both qwen runs today overran 1200s (1206s and 1243s) and lost Phase 3 entirely, one of them mid-write, shipping a module Python cannot parse.
The code default moved to 2400s but compose pinned it with
${BESSER_LLM_DEFAULT_MAX_RUNTIME_SECONDS:-1200}, so the container kept
running at 20 minutes. Three qwen runs in a row overran it - 1206s, 1243s
and 1215s - and all three skipped Phase 3 entirely. That is the stage
that runs the data-contract checks, the frontend contract and the bounded
fix loop, so every check we add is inert until the cap stops cutting the
run short.
… model reaches the run The server-paid sponsored tier accepted provider="sponsored" from anyone and built a client on the server's own endpoint and bearer: an open proxy for the org key, bounded only by the per-run cost cap. It now requires the demo secret (BESSER_DEMO_TOKEN, carried by demo links as ?demo=<token>). BESSER_FREE_LLM_PILOT_MODEL never took effect: 17 of 17 pilot runs across 9 participants went out on the public default because only the client resolved it and the no-popup free tier stores nothing. The server now resolves it.
The threshold could only clamp down from a flat 80k, so a 1M-window model used ~8% of its context and compaction fired into a read/compact/re-read spiral. Resolution: measured overrides (always win) -> free and sponsored catalogs (GET /models, once per process, never fails a run) -> known hosted families -> the unchanged default. Advertised windows yield (window - reserve) * 0.8, capped by BESSER_LLM_MAX_COMPACT_THRESHOLD (200k) and, when set, by BESSER_LLM_COMPACT_THRESHOLD. Invariant test: no known model lands materially below today's default unless a measured row says so.
instructions[:500] in the generator selector never saw a stack clause that conventionally comes last: a 4,622-char spec ending "Frontend -> React" was scaffolded with no frontend. Three such clips were fixed separately today, each believed to be the last. user_request() is now the one sanctioned way to shorten the request, and a test fails on any bound that is not declared.
Live run: "Add authentication" for a request that never mentioned it (the condition trailed twelve vivid nouns); two enumerations proposed that the model already carried (name matched, identical literal sets ignored); the same task listed three times, worked through per copy. The gate now leads the invention list; an add-enumeration task whose literal set the model already has is dropped; tasks are deduplicated before the cap; matching is on meaning and member sets, not exact names. The six declared-but-unimplemented method tasks from the same run are pinned as kept.
…rom Windows hosts Five byte-identical 14-line 501 stubs per router made any quoted stub body ambiguous for modify_file; the stub comment, stdout capture and 500 detail now name their method (26 -> 3 duplicated 4-line windows, all pure try/except). React templates copied verbatim shipped CRLF from a Windows checkout and text-mode writes produced CRLF workspaces; copies and renders now write LF.
…ndbox The import smoke check, the constructibility and API probes, tsc, cargo check, npm build and the write-time import probe run in bubblewrap with no network; npm install and the pip dependency check run confined with network. On Linux a check whose sandbox cannot start is reported unverified instead of running unconfined. /usr/local and /root are read-only with a per-run HOME, and the telemetry folder is masked from runs.
…njection BPMN B-UML import now runs through the safe loader with a BPMN-only allowlist. GUI export escapes screen names in comments, and the code builders escape every character splitlines() treats as a line break. Rx/Ry/Rz gates, function-gate names with symbols, quantum input values and agent condition names now export to importable code; GUI form labels and the submit button survive re-import, and bind_data_source keeps field order. Adds round-trip tests for BPMN, quantum and GUI forms.
Templates iterate unordered sets by name and sort_by_timestamp breaks ties by name, so two runs of the same model give byte-identical output. Django models.py with several modeled methods is valid Python again, inferred method bodies include inherited attributes, and update_settings errors propagate. Class rename no longer creates duplicate association-end names, validate() reports them, and role names pluralise with English rules. pk_types no longer swallows errors; AssociationClass accepts timestamp and metadata.
The Dockerfile decodes certificate subjects and fails the build if a proxy CA remains in the system store or the worker's JDK keystore, and mounts ca-certs-extra instead of copying it into a layer (requires BuildKit). The local compose file gets the sandbox security_opt used in production. The serializer reads the real metamodel attributes, so a component's bound entity reaches the prompt; refusal tests fail instead of skipping; swallowed check errors are logged; dead helpers removed; ClaudeLLMClient exported.
Covers the events.json endpoint, the on-by-default smoke-check and requirements-ledger flags, sandboxed validators and installs, link-free packaging, the BuildKit requirement, the shell-tool policy, and per-generator 8.0.0 breaking changes.
Architecture tours and endpoint lists now point to the docs pages that cover them; kept facts were re-checked against the code and stale statements corrected.
…tput Auth and retry decisions use the SDK status code, so a 400 that mentions a token count is no longer reported as an invalid key and a provider failure after files were written delivers them as incomplete. Timeouts go straight to the fallback model when one is configured; SDK retries no longer stack with ours. Paid routes are priced from the vendored table, so the cost cap covers pinned ids, and fallback tokens are billed at the fallback rate. The download zip drops artifact files, the run's own key is redacted from every frame, and unparseable tool arguments reach the executor as an explicit error.
Tools, walks, copies and named-file reads and writes in the harness skip or replace symlinks and special files, so a link planted by a sandboxed command cannot surface or overwrite files outside the workspace. Captured command output is capped at 8 MB per stream; line-number stripping only removes the read tool's own NNN| prefix; the probe restarts its server when the sandbox has torn it down; missing tool arguments return a clear error.
A stuck-edit stop in Phase 3 no longer marks a later verified repair as incomplete, and a truncated reply is retried with less output instead of ending the phase. The Dockerfile check resolves COPY sources against the build context and only reports; a separate step restores a FastAPI requirements.txt. The Phase 2 prompt now includes the runbook when shell tools are on, the planner keeps the scaffold's stack, validator crashes are reported as not run, and Phase 3 compacts its history.
A request for an app, web app, UI, website or dashboard now gets the FastAPI scaffold plus an LLM-authored React frontend when no GUI model exists; explicit API-only requests stay backend-only. One shared rule feeds the Phase 2 frontend rule, the checklist task, the Phase 3 missing-frontend blocker and requirement scoping. UI requirements on a run that neither has nor asked for a frontend are reported out of scope, extraction no longer invents requirements for a vague request, and a Phase 3 round that clears no blocker is no longer counted as progress.
SandboxUnavailable now reads as a generic "shell sandbox is unavailable" message; the cause and the operator override are logged once at error level and never reach a tool result, validator finding or event.
Free-tier models that draw on provider credits are priced at list rate so the per-run cap still applies, while the user is shown no cost; -free ids and the self-hosted fallback stay at zero. The free route plans with its own model instead of gpt-4o-mini, which the gateway does not serve. The Anthropic default is claude-sonnet-5. A done event with incomplete=false no longer marks the stored run partial.
Generated zips, GitHub deploys and Spec-Driven downloads carry a BESSER_GENERATION.md with the BESSER version, generator and its options (no timestamp, so output stays reproducible); single .py/.sql downloads and B-UML exports start with a version comment; GET /besser_api/ reports the version. The Docker images now copy setup.cfg so the installed package reports its real version. Closes #617.
A caller that omits env= no longer passes the worker's provider keys to sandboxed generated code; every current caller already passed the safe environment. The production compose gives the Spec-Driven worker systempaths=unconfined and no-new-privileges, which bubblewrap needs on Amazon Linux to mount its private /proc.
The ruff check now excludes the harness's own .besser_* files as well as the snapshot, in one --exclude list (with an absolute target a separate --extend-exclude silently replaced it). In live run 7d188d29 all 20 lint warnings were in the runbook's .besser_probe.py. The scaffold README takes its title from the backend's FastAPI title (the model name) instead of the run's temp folder name.
- BESSER_GENERATION.md is written without following a link planted at that name in the run workspace. - The sandbox hides the incident log folder from runs, like telemetry: it names other runs, and a run id is enough to fetch their output. - Deleting a run also removes the sandbox $HOME beside it (cargo, npm and pip caches) instead of waiting for the 24 h sweep. - Every 5xx is retryable again: with the SDK retry off, Anthropic's 529 and a Cloudflare tunnel's 52x were failing on the first attempt and never reached the fallback. - The tools that walk the workspace skip special files, so a FIFO can't hang them.
…s code Default element timestamps now strictly follow creation order, so constructor parameters, columns and enumeration literals keep declaration order in every generator (literals had become name-sorted in Django, Pydantic and SQLAlchemy). The remaining set-order loops are sorted, including inherited association ends, the REST API's associations and the whole Java generator. An inherited to-one end now gets a default, so a subclass __init__ is no longer a SyntaxError. A Django one-to-many self-association emits a single foreign key; renaming a class checks the opposite class's subclasses for duplicate end names; role names pluralise quizzes and analyses; an agent condition without code exports loadable B-UML.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
v8.0.0 — the Spec-Driven Agent arrives in BESSER
BESSER's deterministic generators are what make this possible, and they stay the
foundation. This release builds an agentic layer directly on top of them.
The Spec-Driven Agent is a hybrid generator, not a wrapper around a chat model. A
deterministic BESSER generator produces the scaffold from your model; an LLM customizes
it with a bounded tool surface; a validation gate checks the result back against the
model it came from and fixes what it can. The model never starts from a blank directory,
and nothing ships without being checked — the generators supply the correctness the
agent is measured against.
It is a peer of the deterministic generators — documented, tested, and configured like
one — not an add-on.
202 commits since
development. Full notes:docs/source/releases/v8/v8.0.0.rst.How it works
When no generator fits, Phase 0.5 emits stack metadata and the build starts from
scratch instead.
generators as tools, and a
task_listchecklist that gates the end of the turn.after two without progress, rolling back to the pre-Phase-3 snapshot if the fixes make
blockers worse.
Guards throughout: turn cap, cost and runtime caps checked at turn boundaries, truncation
recovery, per-file modify-loop detection, parallel tool execution grouped by write path.
What else is new
that apply to models, with
blocker/warning/styleseverities. Includes amodel-derived data contract (prompt rules + per-write lint + Phase 3 gate), a frontend
contract check (blank-on-load routers, dead submit handlers), endpoint coherence, and a
per-entity acceptance matrix.
assigned before any subscriber sees a frame, replay via
?after=/Last-Event-ID,resume from a per-turn checkpoint. Closing the tab no longer kills the run.
chain; BYOK across Anthropic, OpenAI and Mistral. Run budgets default to $5 / 20 min.
incrementally rather than regenerating over it.
association classes, and server-owned
id/createdAt/updatedAtexcluded fromCreate schemas.
Security fixes in this PR
defaultValue.default_valuereaches the metamodelunvalidated from request JSON, and the SQLAlchemy template wrote it unquoted:
default=__import__("os").getcwd().SQLGeneratorexecutes the generated module in asubprocess to dump DDL, so that ran as the backend user. The
strbranch wasinjectable too — the value sat inside a quote it could close. The same raw
interpolation existed in the pydantic, python and django templates. Now rendered
through
besser/generators/default_literals.py, which coerces to the declared type andemits
repr().list; the handler table still held
run_commandandexecute_typeddid no membershipcheck, so a model naming a tool it was never offered still ran it. On an
OpenAI-compatible endpoint the tool name comes verbatim from model output. Now refused
at dispatch.
Correctness fixes in this PR
default=str, turning aToolUseBlockinto itsrepr()and destroying theidthat the next message'stool_resultreferenced. Every resume from a checkpoint saved mid-tool-turn wasrejected by the provider. Blocks now serialize to their wire shape.
_restore_snapshotdeleted everything andthen copied the snapshot back, swallowing failures. It now moves the tree aside and
discards it only after the restore succeeds, and never reverts the append-only trace or
the checkpoint.
[]on timeout,launch failure, or error — which the verdict renders as "0 blockers". ruff ignored its
return code entirely (
--exit-zeromeans a non-zero code is ruff itself failing); tscignored its return code unless it could parse an
errorline, so a frontend whosetoolchain would not start passed. Each now emits a visible warning saying the check was
skipped.
project_diris just<output_dir>/<project_name>and wasrmtree'd unconditionally.Breaking changes
Seven things callers will notice. Upgrade notes are in the release file.
allow_shell_tools=False; refused at dispatchenable_toolchain_validation=Falsedefault_valueInvalidDefaultValueErrorat generation timeDjangoGenerator.generate()main.pysmart-generate/vibe/besser_api/spec-driven/*DEFAULT_SQL_DIALECTstandard(not inVALID_DBMS)sqliteVerification
pytest tests/ --ignore=tests/generators/nn— 2585 passed, 17 skipped, 3 xfailedruff check besser/ --select F841,F401,F541,F811,E711,E721,E731,E741 --ignore E501— cleanbash docs/check-docs-warnings.sh— clean (only the two allowlisted systemic categories)anthropic/mistralaimade unimportable, mirroring CI's dependency set —2585 passed
37 new tests accompany the fixes above, each written to fail against the previous
behaviour:
test_default_value_injection.py,test_rollback_safety.py,test_validation_honesty.py,test_django_output_safety.py, and the dispatch-gate casesin
test_shell_tool_gate.py.Companion PR
The frontend submodule pointer moves to
feature/smart-generator@7f529ab2(BESSER-PEARL/BESSER-Web-Modeling-Editor#193). Merge the frontend PR first. That repo
has no CI, so the branch was built locally: webapp 5476 modules transformed, server
webpack exit 0.