Skip to content

v8.0.0 — Spec-Driven Agent, modular backend, and a security pass over generated code - #604

Merged
ArmenSl merged 462 commits into
developmentfrom
feature/smart-generator
Sep 30, 2026
Merged

ArmenSl merged 462 commits into
developmentfrom
feature/smart-generator

Conversation

@ArmenSl

@ArmenSl ArmenSl commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

v8.0.0 — the Spec-Driven Agent arrives in BESSER

BESSER's deterministic generators are what make this possible, and they stay the
foundation. This release builds an agentic layer directly on top of them.

The Spec-Driven Agent is a hybrid generator, not a wrapper around a chat model. A
deterministic BESSER generator produces the scaffold from your model; an LLM customizes
it with a bounded tool surface; a validation gate checks the result back against the
model it came from and fixes what it can. The model never starts from a blank directory,
and nothing ships without being checked — the generators supply the correctness the
agent is measured against.

It is a peer of the deterministic generators — documented, tested, and configured like
one — not an add-on.

202 commits since development. Full notes: docs/source/releases/v8/v8.0.0.rst.

How it works

  • Phase 1 — a deterministic scaffold from the BESSER generator that fits the request.
    When no generator fits, Phase 0.5 emits stack metadata and the build starts from
    scratch instead.
  • Phase 1.5 — validate the scaffold before a single token is spent customizing it.
  • Phase 2 — the LLM customization loop: files, model queries, the other BESSER
    generators as tools, and a task_list checklist that gates the end of the turn.
  • Phase 3 — validate, then a bounded auto-fix loop: up to five rounds, giving up
    after two without progress, rolling back to the pre-Phase-3 snapshot if the fixes make
    blockers worse.

Guards throughout: turn cap, cost and runtime caps checked at turn boundaries, truncation
recovery, per-file modify-loop detection, parallel tool execution grouped by write path.

What else is new

  • Validation over generated code — a fourth validation layer, distinct from the three
    that apply to models, with blocker / warning / style severities. Includes a
    model-derived data contract (prompt rules + per-write lint + Phase 3 gate), a frontend
    contract check (blank-on-load routers, dead submit handlers), endpoint coherence, and a
    per-entity acceptance matrix.
  • Durable runs — a run is owned by the server. SQLite event store, sequence numbers
    assigned before any subscriber sees a frame, replay via ?after= / Last-Event-ID,
    resume from a per-turn checkpoint. Closing the tab no longer kills the run.
  • Free / sponsored / BYOK tiers — a keyless free tier with a cloud-first fallback
    chain; BYOK across Anthropic, OpenAI and Mistral. Run budgets default to $5 / 20 min.
  • GitHub round-trip — push a generated project, import it back, modify it
    incrementally rather than regenerating over it.
  • Modular FastAPI backend — per-entity routers, typed primary keys end to end,
    association classes, and server-owned id / createdAt / updatedAt excluded from
    Create schemas.

Security fixes in this PR

  • Code injection via defaultValue. default_value reaches the metamodel
    unvalidated from request JSON, and the SQLAlchemy template wrote it unquoted:
    default=__import__("os").getcwd(). SQLGenerator executes the generated module in a
    subprocess to dump DDL, so that ran as the backend user. The str branch was
    injectable too — the value sat inside a quote it could close. The same raw
    interpolation existed in the pydantic, python and django templates. Now rendered
    through besser/generators/default_literals.py, which coerces to the declared type and
    emits repr().
  • Shell tools were hidden, not withheld. The gate only filtered the advertised tool
    list; the handler table still held run_command and execute_typed did no membership
    check, so a model naming a tool it was never offered still ran it. On an
    OpenAI-compatible endpoint the tool name comes verbatim from model output. Now refused
    at dispatch.

Correctness fixes in this PR

  • Resume was broken. The checkpoint serialized messages with default=str, turning a
    ToolUseBlock into its repr() and destroying the id that the next message's
    tool_result referenced. Every resume from a checkpoint saved mid-tool-turn was
    rejected by the provider. Blocks now serialize to their wire shape.
  • Rollback could destroy the workspace. _restore_snapshot deleted everything and
    then copied the snapshot back, swallowing failures. It now moves the tree aside and
    discards it only after the restore succeeds, and never reverts the append-only trace or
    the checkpoint.
  • Validation reported clean when it hadn't run. Collectors returned [] on timeout,
    launch failure, or error — which the verdict renders as "0 blockers". ruff ignored its
    return code entirely (--exit-zero means a non-zero code is ruff itself failing); tsc
    ignored its return code unless it could parse an error line, so a frontend whose
    toolchain would not start passed. Each now emits a visible warning saying the check was
    skipped.
  • DjangoGenerator deleted user content. project_dir is just
    <output_dir>/<project_name> and was rmtree'd unconditionally.

Breaking changes

Seven things callers will notice. Upgrade notes are in the release file.

Change Before After
Agent shell tools on by default allow_shell_tools=False; refused at dispatch
Toolchain validation on enable_toolchain_validation=False
Non-literal default_value interpolated raw into source InvalidDefaultValueError at generation time
DjangoGenerator.generate() printed the exception, returned normally propagates; refuses to delete non-generated output
FastAPI backend generator one main.py modular, per-entity routers
Agent REST paths smart-generate / vibe /besser_api/spec-driven/*
DEFAULT_SQL_DIALECT standard (not in VALID_DBMS) sqlite

Verification

  • pytest tests/ --ignore=tests/generators/nn — 2585 passed, 17 skipped, 3 xfailed
  • ruff check besser/ --select F841,F401,F541,F811,E711,E721,E731,E741 --ignore E501 — clean
  • bash docs/check-docs-warnings.sh — clean (only the two allowlisted systemic categories)
  • Re-run with anthropic / mistralai made unimportable, mirroring CI's dependency set —
    2585 passed

37 new tests accompany the fixes above, each written to fail against the previous
behaviour: test_default_value_injection.py, test_rollback_safety.py,
test_validation_honesty.py, test_django_output_safety.py, and the dispatch-gate cases
in test_shell_tool_gate.py.

Companion PR

The frontend submodule pointer moves to feature/smart-generator @ 7f529ab2
(BESSER-PEARL/BESSER-Web-Modeling-Editor#193). Merge the frontend PR first. That repo
has no CI, so the branch was built locally: webapp 5476 modules transformed, server
webpack exit 0.

1. push-to-github 404'd for every worker run.

   restore_persisted() runs once, in the lifespan. That was fine while the
   process that produced a run also served its download and its push. Since
   generation moved to besser-wme-smartgen it is not: push-to-github and
   import-github-run stay on the backend because they need its process-local
   OAuth session, so every worker run is written to the shared volume AFTER
   the backend's one-shot scan. Proven live: worker had 1 run on disk, the
   backend registry had 0 entries.

   get() now falls back to a TARGETED rescan for the single run_id it was
   asked about, sharing the manifest parse with restore_persisted. A cached
   hit still touches no disk, a malformed id never reaches os.listdir, and an
   expired run stays gone.

   An earlier commit claimed this was fixed and verified. It was not: I had
   checked that the backend could SEE the files, which is a different thing
   from the endpoint resolving the run.

2. A modelled default could escape into an executable os.system() call.

   backend/templates/router.py.j2 renders param.default_value in four
   executable places and two docstrings. The executable ones were hardened
   first; python_default never raises for a str, so generation SUCCEEDED and a
   value containing a triple-double-quote closed the generated docstring and
   dropped statements at body indentation -- in a router /deploy-app executes.
   Verified: reverting the fix makes the new test report the escape at a
   concrete line.

   That was the third site of this one sink in this branch, after the
   SQLAlchemy macro and its association-class copy.

   The same lines also passed param[1], the RENDERED ANNOTATION ('Any', or an
   FK class name) rather than the modelled type, so python_default fell to its
   quoted fallback and emitted '5' for an int default. The tuple now carries
   param.type.name and both uses read it.

3. A symlink in a run workspace could exfiltrate the backend's secrets.

   github_service collected files with rglob + is_file(), which FOLLOWS
   symlinks. The tree being pushed is written by the worker, which runs
   model-authored code with shell tools enabled; the backend doing the push
   holds SMTP_PASSWORD, GITHUB_CLIENT_SECRET, GPG_KEY and the rest. `ln -s
   /proc/self/environ leak.txt` would have been read through and committed to
   the user's repository. :ro blocks writes, not traversal.

   Symlinks are skipped and every survivor is re-checked with commonpath.

   This was not exploitable while (1) was broken -- which is exactly why it
   ships in the same commit as the fix that would have armed it.
The keyless picker listed BESSER_FREE_LLM_FALLBACK_MODEL as a choosable
model, which put the self-hosted box directly in front of users. That box
serves ONE request at a time: a parallel batch measured on 2026-09-11
drove two of five runs to zero turns, and the two concurrent runs today
showed the same shape -- one finished in 203s while the other queued
behind it for 1342s and died on the runtime cap.

It is still configured as the fallback and still runs automatically when
the primary keyless endpoint fails -- which is not hypothetical, poolside
/laguna-s-2.1-free is returning "Upstream model provider is temporarily
unavailable" right now. Only its appearance in the picker is removed.

The two config tests that asserted it WAS offered are updated rather than
relaxed: they now assert it is absent from the choices AND that
free_fallback_model() still returns it, so hiding it from the picker
cannot silently disable the failover.
…nds on it

_get_pricing asks _is_free_local_model first, and a True answer returns
_ZERO_PRICING. Every call then costs $0, so max_cost_usd can never trip: a
run continues to the turn or time cap while the provider bills normally and
the run card reads $0.00 throughout.

The rule matched FAMILY NAMES -- qwen, deepseek, llama, gemma, mistral-small
and so on. That was right while open-weight implied self-hosted. It stopped
being right when Command Code began reselling those same families:
Qwen/Qwen3.8-Max, Qwen/Qwen3.8-Flash, deepseek/deepseek-v4-pro and
deepseek/deepseek-v4-flash are all billed models on the plan, and all four
priced at $0. Selecting one as a paid option would have spent real money
with the cap inert.

The decision now keys on how a model is IDENTIFIED rather than its family:

  1. explicit free marker (:free / -free)      -> free
  2. Ollama name:tag with no vendor namespace  -> self-hosted, free
  3. vendor/model (a gateway id)               -> billed
  4. otherwise the family list, which still covers tagless self-hosted ids

Verified against the real ids in play: qwen3-coder:30b, LongCat-2.0:free,
laguna-s-2.1-free and ling-3.0-flash-sante:free stay free; Qwen3.8-Max,
deepseek-v4-pro, Kimi-K3, GLM-5.3, MiniMax-M3 and gemini-3.8-flash become
billed. The gateway models fall to the gpt-4o default, which is a guess --
but a guess that keeps the cap working, where $0 switched it off.
BESSER_FREE_LLM_PILOT_MODEL names the keyless model a facilitated pilot
session should start on; /spec-driven/config exposes it as
free_tier.pilot_model (null when unset, so pilots and the public get the
same default).

free_pilot_model() refuses an id the server does not actually offer.
Advertising a default we would decline to honour is worse than having
none: the client sends the pre-selected id as llm_model, an unknown id is
pinned back to the default, and the pilot would silently run on the public
model while the UI showed something else.

Server-side because the choice is volatile -- it changed three times in
one afternoon -- so it should be an env edit and a container restart.

Bumps the frontend submodule for the matching client change.
/spec-driven/config is served by the worker, which uses an explicit env
allowlist rather than env_file. A variable missing from that list reads as
unset, so the pilot default silently did nothing -- the third feature this
allowlist has swallowed, after the alt-model list and run telemetry.

The allowlist is still right (it cut the generation path from nine secrets
to three) but it fails silently, so every new BESSER_* var the spec-driven
path reads has to be added here deliberately.
Server side of surviving a TLS-inspecting corporate proxy (Netskope on
LIST laptops). The proxy buffers a response body before releasing it, and
an SSE body never ends, so the run stream is the one endpoint it breaks —
REST and the agent WebSocket both work through it.

GET /spec-driven/runs/{id}/events.json returns the same durable events as
a short, terminating JSON response the proxy can release normally. It
needed almost no new machinery: the event store already assigns sequence
numbers before any subscriber sees a frame, so polling and streaming share
one cursor and one replay semantic. Payloads are decoded so one client
reducer serves both transports; `hasMore` lets a poller drain a backlog
without sleeping through its interval.

POST /spec-driven/generate now honours an Idempotency-Key. Starting a run
is not idempotent, which is why the client never retried a start that
failed at the transport layer — and behind such a proxy that is exactly the
request that fails, stranding the user with a run they cannot rejoin. The
key is claimed before the run starts, so a retry arriving mid-startup waits
for that run rather than opening a second one, and is released if startup
fails. In-process and best-effort: a missed hit costs one duplicate run,
never corruption, and the durable store stays the source of truth.

Also fixes the free tier 400-ing on gpt-5.6. reasoning_effort='none' is
required by the official OpenAI API for tools on those models, but a
gateway re-serving them validates the value against its own enum, and
'none' is not in it — api.commandcode.ai answers `Invalid option: expected
one of "low"|"medium"|"high"|"xhigh"|"max"`, which aborted the tool-driven
Phase 2 loop and shipped a deterministic-only bundle flagged incomplete.
Verified against gpt-5.6-luna there: 'none' -> 400, while omitted / 'low' /
'medium' each returned a proper tool call. So the flag is now sent only on
the official endpoint and omitted on any custom base_url.
…njection

Turn cap 120 -> 150. A turn is a unit of round-trip, not of work: a model
that batches tool calls does several actions per turn, while one that emits
a single call per turn needs roughly 4x the turns for the same result. Two
measured runs bracket it — 30 calls in 14 turns (2.1/turn, finished at 17%
of budget) against 2026-09-11's 80 turns / 80 calls / 18 files, which
consumed the whole budget for less output. The ceiling has to fit the worst
ratio or a non-batching model hits it mid-build having done nothing wrong.
Cost and runtime are still checked at every turn boundary, so a higher turn
ceiling cannot raise spend past the cost cap.

Docs updated for the new ceiling, and two pre-existing wrong numbers in the
same list corrected: BESSER_LLM_DEFAULT_MAX_COST_USD is 5.0 and
BESSER_LLM_DEFAULT_MAX_RUNTIME_SECONDS is 1200, not the documented 1.0/600.
runs.rst already had the right values, so the two pages contradicted.

Corporate CA injection is now opt-in behind --build-arg TRUST_EXTRA_CAS=1
(Dockerfile). deploy.sh builds from the local working tree, so the
unconditional COPY would have baked whatever .crt sits in ca-certs-extra/
into the image shipped to the shared host — making the deployed backend
trust a corporate CA for every outbound TLS call. Defaulting to 0 keeps
production clean by construction rather than by remembering to empty a
directory. The certs themselves stay gitignored; only .gitkeep is tracked.
…es to 4

Turn default 80 -> 120. The cap was never the spend guard: cost and runtime
are checked at every turn boundary before the next billable call, so a
higher turn ceiling cannot raise what a run spends. What it does fix is the
model-dependence — a turn is a unit of round-trip, not of work. Measured:
one run did 30 tool calls in 14 turns, while 2026-09-11 did 80 calls in 80
turns for 18 files. At 1 call/turn, 80 turns strangles a model that has done
nothing wrong.

Truncation retries 2 -> 4. This budget is a per-run TOTAL that deliberately
never resets after a clean turn, and it ends Phase 2 outright. At 2, a run
that truncated once early, adapted, produced a hundred good turns and then
truncated once more was killed one bad turn from done — while the stated
rationale ("beyond that it is not adapting") had been false for a hundred
turns. Raising the turn budget makes that sharper, so the two move together.

_MAX_PARALLEL_WORKERS stays at 4, deliberately. The deploy host is 2 vCPU /
1.9GB RAM with ~450MB free, and LLM_MAX_CONCURRENT_RUNS is 10 — already up
to 40 concurrent tool executions on two cores. The CPU-bound tools cannot
outrun the cores, so raising it buys contention, not throughput.

preview.py no longer hardcodes 80 with a comment claiming it "matches
LLMOrchestrator.MAX_TURNS"; it reads LLM_DEFAULT_MAX_TURNS, so the estimate
cannot silently drift from the budget again. Docs updated for both values,
including the truncation budget's non-resetting semantics, and a stale
comment claiming the cost default is $1 (it is $5) removed.
The Django template collapsed newlines in two places, so any realistic
model produced a models.py that fails ast.parse. Pre-existing on master,
not a regression from this branch.

Enumerations: `{%- for %}` / `{%- endfor %}` strip the preceding newline
and trim_blocks strips the following one, so every literal and every enum
class ran together on one line:

    class LoanStatus(models.TextChoices):    CLOSED = 'CLOSED', ...class MemberStatus(...)

Relationship fields: `{%- endif -%}` ate the newline after a field's
closing paren, so two consecutive ForeignKeys on one class joined into
`null=True)    member = models.ForeignKey(` — a statement-level join, not
keyword arguments sharing a line. Same tag in the ManyToMany and OneToOne
blocks, so all three are fixed. This one is independent of the enum defect
and reproduces on an enum-free model.

Whitespace control only; trim_blocks/lstrip_blocks are untouched because
every other block in the template depends on them.

The generator's own subprocesses also leaked bytecode into the download:
`manage.py startapp` imports the settings it just wrote, so CPython left
new_project/new_project/__pycache__/*.pyc in the packaged tree. They now
run with PYTHONDONTWRITEBYTECODE, so nothing is written to clean up.

Existing Django compile tests all used enum-free, single-FK models, which
is why this shipped. The new tests use both shapes and assert ast.parse.
The comments accumulated across the spec-driven work had grown into
15-20 line blocks retelling a debugging story around a single fact. Trimmed
to the fact: ~200 fewer comment lines across 23 files, the largest share in
llm/orchestrator.py, and duplicated rationale collapsed to one location
instead of being restated at the constant, the docstring and the call site.

Kept, deliberately: every dated measurement (the 62-turn livelock, the
40%-bookkeeping batch, the 128k prompt truncated to 65_536), the security
rationale for a default being off, and the "why the obvious simpler version
is wrong" warnings — those are what stop a future cleanup reintroducing the
bug.

Comments and docstrings only. Verified by parsing each changed file at HEAD
and in the working tree, stripping docstrings, and comparing ast.dump():
identical for all 23.
…e comments

A corporate CA is needed to FETCH packages when building from a laptop behind
TLS inspection, but the deployed backend must not trust it for its own
outbound calls. The final layer removes it and asserts it is gone — the grep
fails the build rather than shipping an image that trusts the proxy. It runs
unconditionally, so the guarantee holds whether or not TRUST_EXTRA_CAS was
passed.

That closes the gap the opt-in ARG alone left: without it, building behind the
proxy meant choosing between a failed build and a shipped corporate CA.

Comments trimmed 95 -> 72 lines. Instructions are unchanged apart from the new
layer; the facts worth keeping (why 3.11+, why the toolchains must be on PATH,
why ruff is pinned here) are kept, the narrative around them is not.
…ad Tailwind

Two template-level defects, both deterministic, both invisible to the current
validation because it checks files in isolation and nothing exercises a write
or a stylesheet.

1. Every POST returned 500. pydantic_classes_template deliberately drops the
   surrogate id and the audit timestamps from Create -- the ORM stamps them,
   and its comment says leaving them in "breaks the generated app". router.py.j2
   dropped only the id and read createdAt/updatedAt off the payload regardless,
   so FastAPI raised AttributeError on every create. Verified live on a
   generated hotel app: 7 of 9 entities returned 500 while the run was reported
   a clean success. The router now shares one _client_supplied macro that
   mirrors the pydantic rule, so the two cannot drift again.

2. The React generator declared tailwindcss ^4.1.14 and never wired it: no
   @tailwindcss/vite plugin, no @import, no config, and postcss+autoprefixer
   alongside -- the v3 toolchain for a v4 package. The generator's own
   components use zero utility classes so deterministic output looked fine, but
   an LLM reads package.json, concludes Tailwind is available, and writes
   utility classes that are inert. Observed: a nav rendered as a browser-default
   bulleted list because "flex gap-1 py-3" did nothing. Wiring it turned the
   same unchanged LLM output into a styled app.

   "type": "module" is required with it: @tailwindcss/vite is ESM-only and Vite
   otherwise tries to require() it and refuses to start. Found by booting the
   app -- a template test asserting the import line exists would have passed.
…nces

No relationship could be set through the API of any generated app whose model
uses inheritance. Observed live on a hotel app: Person carried a string uuid
PK, Guest and Employee inherited it (joined-table -- their id IS a ForeignKey
to person.id), but every FK pointing at Guest or Employee came out
Mapped_[int], and the Pydantic field came out `employee: int`. POST /booking/
with a real id returned 422 "Input should be a valid integer".

The schema was internally inconsistent too, not just the API layer. SQLite
creates those tables regardless, so boot checks, imports and ast.parse all
passed and the run was reported a clean success.

Cause: the PK-type resolver looked only at a class's OWN attributes. A
subclass has no id attribute of its own, so it was absent from the map and
every caller fell back to the 'int' default. It now follows parents().

There were THREE copies of this logic. Both docstrings already described this
exact failure -- "a Mapped_[int] FK pointing at a String PK breaks joins at
runtime", "a guest: int Pydantic field 422-rejects every real id" -- and
neither implementation handled inheritance. pk_types.pk_python_types now
delegates to structural_utils.get_pk_py_types so there is one implementation,
and the pydantic template stops hardcoding `int` for association ends.
Five defects, each traced to a live run and each covered by a test that
was verified to fail against the pre-fix code.

Neither entity could be created. A 1:1 association emitted a mandatory
relationship field on BOTH create schemas, because the FK-owning and
non-owning branches of the Pydantic template were byte-identical.
BookingCreate demanded a guest, GuestCreate demanded a booking. The
router never read the non-owning field either - it validates and assigns
only the FK-owning end - so the client was forced to send an id that was
then silently discarded. The N:1 branch already emitted nothing there;
the 1:1 branch now matches it. Two existing tests asserted the discarded
field and have been updated to assert its absence.

Two tables each holding a NOT NULL foreign key to the other cannot be
populated: POST /guest/ died on "NOT NULL constraint failed" before
anything existed to point at. A cycle needs two associations with
opposite FK ownership - one association yields only one FK column - so
one edge has to give. get_deferred_fk_associations finds the cycles and
breaks one edge each, chosen by association name so the choice is stable
across regenerations. That FK becomes nullable with use_alter, its
relationships get post_update, and the create schema and router stop
demanding it. The three-step create flow the REST API actually performs
now completes; before, the first step failed.

tsc ran without node_modules, so every package import was unresolvable
and the type errors cascaded from there. A frontend that boots and
renders was reported as "3 blocker-level issues remain", headlined by
TS2688 on vite/client. We never run npm install during validation - it
needs network the host may not have - so instead the findings are
demoted to advisory when node_modules is absent. One exception stays a
blocker: a TS2307 naming a RELATIVE path, which is a genuine missing
file either way.

Every non-FastAPI app was reported as "cannot start". django and flask
were simply absent from a hardcoded allowlist, as were argparse,
sqlite3, glob and ~180 other stdlib modules. The stdlib half now comes
from sys.stdlib_module_names, the framework families are listed, and -
the part that generalises to all 20 generators - the app's own manifest
is read. A package the app declares is a package the app has. Also
fixed: a directory of .py files is importable from its PARENT, not only
as a sibling of the importing file, so `from core.config import
settings` in backend/routers/x.py no longer reports missing. On a
django/flask/stdlib/import-core sample: 11 false blockers, now 0. The
three genuine-defect cases still report.

The missing-frontend blocker never fired on "build a web application":
the trailing \b in \b(web ?app|...)\b cannot match when "app" continues
into "lication". Nine of ten sweep runs shipped backend-only and
reported success. "spa" is deliberately NOT an alternative - a hotel
has one.

Also: the gap analyser compared the generated code against the MODEL
rather than against the REQUEST. On run ca48a6dd it returned four
cosmetic tasks for a spec with twelve real gaps - two state machines,
five methods, four OCL constraints and an association-class attribute
were all invisible to it. It is the only stage that holds the request
and the model side by side, so what it misses is lost for the rest of
the run. It now performs an explicit two-pass diff, request against
model first, and the system prompt states that the model is a lossy
capture of the request rather than the authority.

And MethodButtonProps lacked the `id` the page builder puts on every
component it renders, which was 8 x TS2322 on every entity page.
Three changes, all from measuring a live run against the app it produced.

A create schema and its router can disagree, and every disagreement is a
guaranteed 500. On 2026-09-17 the LLM rewrote PersonCreate to add an email
validator and dropped lastName; the deterministic router still read
guest_data.lastName, so person, guest AND employee returned 500 - every
way to get a row into the system. Phase 3 ran, found 29 issues and zero
blockers. The new check is structural: for every X_data.<field> a router
reads, the field must be declared on XCreate or one of its Create bases.
Three findings on that app, none on a clean scaffold. It is the third
instance of this shape today, after createdAt/updatedAt and a 1:1
relationship field, so it is a class of bug, not an accident.

The agent could not see the file it most needed to edit. A generated
routers/booking.py runs to 34,884 characters against a 20,000-character
read limit - 59% visible - and write_file then asked it to reproduce the
whole thing. Modeled-method bodies now render into their own
<entity>_methods.py: booking.py drops to 24,911 with a self-contained
10,247-character methods module beside it. Routes are untouched - same
paths, same tags, a second APIRouter that main_api includes - verified by
diffing the OpenAPI surface before and after: 62 paths, identical, boots
clean. Method bodies are the hand-written half of the generation gap, so
giving them their own file also stops a botched rewrite taking 24k of CRUD
with it.

MAX_FILE_READ goes to 60,000. The largest generated file is ~8,700 tokens
against an 80,000-token compaction threshold, so reading it whole costs
about 11% of the budget - far cheaper than editing it blind.

Also: the default runtime cap moves from 20 to 40 minutes. Both qwen runs
today overran 1200s (1206s and 1243s) and lost Phase 3 entirely, one of
them mid-write, shipping a module Python cannot parse.
The code default moved to 2400s but compose pinned it with
${BESSER_LLM_DEFAULT_MAX_RUNTIME_SECONDS:-1200}, so the container kept
running at 20 minutes. Three qwen runs in a row overran it - 1206s, 1243s
and 1215s - and all three skipped Phase 3 entirely. That is the stage
that runs the data-contract checks, the frontend contract and the bounded
fix loop, so every check we add is inert until the cap stops cutting the
run short.
… model reaches the run

The server-paid sponsored tier accepted provider="sponsored" from anyone and
built a client on the server's own endpoint and bearer: an open proxy for the
org key, bounded only by the per-run cost cap. It now requires the demo secret
(BESSER_DEMO_TOKEN, carried by demo links as ?demo=<token>).

BESSER_FREE_LLM_PILOT_MODEL never took effect: 17 of 17 pilot runs across 9
participants went out on the public default because only the client resolved
it and the no-popup free tier stores nothing. The server now resolves it.
The threshold could only clamp down from a flat 80k, so a 1M-window model
used ~8% of its context and compaction fired into a read/compact/re-read
spiral. Resolution: measured overrides (always win) -> free and sponsored
catalogs (GET /models, once per process, never fails a run) -> known hosted
families -> the unchanged default. Advertised windows yield
(window - reserve) * 0.8, capped by BESSER_LLM_MAX_COMPACT_THRESHOLD (200k)
and, when set, by BESSER_LLM_COMPACT_THRESHOLD. Invariant test: no known
model lands materially below today's default unless a measured row says so.
instructions[:500] in the generator selector never saw a stack clause that
conventionally comes last: a 4,622-char spec ending "Frontend -> React" was
scaffolded with no frontend. Three such clips were fixed separately today,
each believed to be the last. user_request() is now the one sanctioned way to
shorten the request, and a test fails on any bound that is not declared.
Live run: "Add authentication" for a request that never mentioned it (the
condition trailed twelve vivid nouns); two enumerations proposed that the model
already carried (name matched, identical literal sets ignored); the same task
listed three times, worked through per copy. The gate now leads the invention
list; an add-enumeration task whose literal set the model already has is
dropped; tasks are deduplicated before the cap; matching is on meaning and
member sets, not exact names. The six declared-but-unimplemented method tasks
from the same run are pinned as kept.
…rom Windows hosts

Five byte-identical 14-line 501 stubs per router made any quoted stub body
ambiguous for modify_file; the stub comment, stdout capture and 500 detail now
name their method (26 -> 3 duplicated 4-line windows, all pure try/except).
React templates copied verbatim shipped CRLF from a Windows checkout and
text-mode writes produced CRLF workspaces; copies and renders now write LF.
…ndbox

The import smoke check, the constructibility and API probes, tsc, cargo
check, npm build and the write-time import probe run in bubblewrap with no
network; npm install and the pip dependency check run confined with network.
On Linux a check whose sandbox cannot start is reported unverified instead
of running unconfined. /usr/local and /root are read-only with a per-run
HOME, and the telemetry folder is masked from runs.
…njection

BPMN B-UML import now runs through the safe loader with a BPMN-only
allowlist. GUI export escapes screen names in comments, and the code
builders escape every character splitlines() treats as a line break.
Rx/Ry/Rz gates, function-gate names with symbols, quantum input values and
agent condition names now export to importable code; GUI form labels and
the submit button survive re-import, and bind_data_source keeps field
order. Adds round-trip tests for BPMN, quantum and GUI forms.
Templates iterate unordered sets by name and sort_by_timestamp breaks ties
by name, so two runs of the same model give byte-identical output. Django
models.py with several modeled methods is valid Python again, inferred
method bodies include inherited attributes, and update_settings errors
propagate. Class rename no longer creates duplicate association-end names,
validate() reports them, and role names pluralise with English rules.
pk_types no longer swallows errors; AssociationClass accepts timestamp and
metadata.
The Dockerfile decodes certificate subjects and fails the build if a proxy
CA remains in the system store or the worker's JDK keystore, and mounts
ca-certs-extra instead of copying it into a layer (requires BuildKit).
The local compose file gets the sandbox security_opt used in production.
The serializer reads the real metamodel attributes, so a component's bound
entity reaches the prompt; refusal tests fail instead of skipping; swallowed
check errors are logged; dead helpers removed; ClaudeLLMClient exported.
Covers the events.json endpoint, the on-by-default smoke-check and
requirements-ledger flags, sandboxed validators and installs, link-free
packaging, the BuildKit requirement, the shell-tool policy, and per-generator
8.0.0 breaking changes.
Architecture tours and endpoint lists now point to the docs pages that
cover them; kept facts were re-checked against the code and stale
statements corrected.
…tput

Auth and retry decisions use the SDK status code, so a 400 that mentions
a token count is no longer reported as an invalid key and a provider
failure after files were written delivers them as incomplete. Timeouts go
straight to the fallback model when one is configured; SDK retries no
longer stack with ours. Paid routes are priced from the vendored table,
so the cost cap covers pinned ids, and fallback tokens are billed at the
fallback rate. The download zip drops artifact files, the run's own key is
redacted from every frame, and unparseable tool arguments reach the
executor as an explicit error.
Tools, walks, copies and named-file reads and writes in the harness skip
or replace symlinks and special files, so a link planted by a sandboxed
command cannot surface or overwrite files outside the workspace. Captured
command output is capped at 8 MB per stream; line-number stripping only
removes the read tool's own NNN| prefix; the probe restarts its server
when the sandbox has torn it down; missing tool arguments return a clear
error.
A stuck-edit stop in Phase 3 no longer marks a later verified repair as
incomplete, and a truncated reply is retried with less output instead of
ending the phase. The Dockerfile check resolves COPY sources against the
build context and only reports; a separate step restores a FastAPI
requirements.txt. The Phase 2 prompt now includes the runbook when shell
tools are on, the planner keeps the scaffold's stack, validator crashes are
reported as not run, and Phase 3 compacts its history.
A request for an app, web app, UI, website or dashboard now gets the
FastAPI scaffold plus an LLM-authored React frontend when no GUI model
exists; explicit API-only requests stay backend-only. One shared rule
feeds the Phase 2 frontend rule, the checklist task, the Phase 3
missing-frontend blocker and requirement scoping. UI requirements on a
run that neither has nor asked for a frontend are reported out of scope,
extraction no longer invents requirements for a vague request, and a
Phase 3 round that clears no blocker is no longer counted as progress.
SandboxUnavailable now reads as a generic "shell sandbox is unavailable"
message; the cause and the operator override are logged once at error
level and never reach a tool result, validator finding or event.
Free-tier models that draw on provider credits are priced at list rate so
the per-run cap still applies, while the user is shown no cost; -free ids
and the self-hosted fallback stay at zero. The free route plans with its
own model instead of gpt-4o-mini, which the gateway does not serve. The
Anthropic default is claude-sonnet-5. A done event with incomplete=false
no longer marks the stored run partial.
Generated zips, GitHub deploys and Spec-Driven downloads carry a
BESSER_GENERATION.md with the BESSER version, generator and its options
(no timestamp, so output stays reproducible); single .py/.sql downloads
and B-UML exports start with a version comment; GET /besser_api/ reports
the version. The Docker images now copy setup.cfg so the installed
package reports its real version. Closes #617.
A caller that omits env= no longer passes the worker's provider keys to
sandboxed generated code; every current caller already passed the safe
environment. The production compose gives the Spec-Driven worker
systempaths=unconfined and no-new-privileges, which bubblewrap needs on
Amazon Linux to mount its private /proc.
The ruff check now excludes the harness's own .besser_* files as well as
the snapshot, in one --exclude list (with an absolute target a separate
--extend-exclude silently replaced it). In live run 7d188d29 all 20 lint
warnings were in the runbook's .besser_probe.py. The scaffold README takes
its title from the backend's FastAPI title (the model name) instead of the
run's temp folder name.
- BESSER_GENERATION.md is written without following a link planted at that
  name in the run workspace.
- The sandbox hides the incident log folder from runs, like telemetry: it
  names other runs, and a run id is enough to fetch their output.
- Deleting a run also removes the sandbox $HOME beside it (cargo, npm and
  pip caches) instead of waiting for the 24 h sweep.
- Every 5xx is retryable again: with the SDK retry off, Anthropic's 529 and
  a Cloudflare tunnel's 52x were failing on the first attempt and never
  reached the fallback.
- The tools that walk the workspace skip special files, so a FIFO can't
  hang them.
…s code

Default element timestamps now strictly follow creation order, so
constructor parameters, columns and enumeration literals keep declaration
order in every generator (literals had become name-sorted in Django,
Pydantic and SQLAlchemy). The remaining set-order loops are sorted,
including inherited association ends, the REST API's associations and the
whole Java generator. An inherited to-one end now gets a default, so a
subclass __init__ is no longer a SyntaxError. A Django one-to-many
self-association emits a single foreign key; renaming a class checks the
opposite class's subclasses for duplicate end names; role names pluralise
quizzes and analyses; an agent condition without code exports loadable
B-UML.
@ArmenSl
ArmenSl marked this pull request as ready for review September 30, 2026 09:28
@ArmenSl
ArmenSl merged commit ac81f2a into development Sep 30, 2026
6 checks passed
@ArmenSl
ArmenSl deleted the feature/smart-generator branch September 30, 2026 14:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants