release: v7.18.0 — live agent simulation and methods implemented by a neural network - #616
Merged
Merged
Conversation
… and align frontend auth flow
- secure backend agent-test routes in `besser/utilities/web_modeling_editor/backend/routers/agent_test_router.py`
- add optional auth gate via `AGENT_TEST_REQUIRE_AUTH` (GitHub session validation)
- add in-memory sliding-window rate limiter (`AGENT_TEST_RATE_LIMIT_WINDOW_SECONDS`, `AGENT_TEST_RATE_LIMIT_MAX_REQUESTS`)
- validate session IDs as UUIDs on REST + WS paths
- enforce WS auth using `github_session` query param and return explicit WS close codes for auth/input failures
- harden sandbox API/session boundaries in `besser/utilities/web_modeling_editor/agent-sandbox/agent_sandbox_api.py`
- validate `session_id` on create/delete/ws
- add startup permissions bootstrap for `/tmp/sessions` (`root:root`, `0711`)
- strengthen sandbox process + filesystem isolation in `besser/utilities/web_modeling_editor/agent-sandbox/session_manager.py`
- add thread-safe state management with `RLock`
- add safe session path resolution and canonical UUID normalization
- add per-session UID allocation (`AGENT_SANDBOX_SESSION_UID_BASE`) and release lifecycle
- drop each agent subprocess privileges to its dedicated UID/GID (`setuid`/`setgid`)
- enforce strict per-session FS permissions (`0700` dirs, `0600` files), isolated `HOME`/`TMPDIR`
- replace inherited env with strict allowlist to avoid container secret leakage
- keep existing resource limits and fail fast if privilege drop cannot be applied
- update sandbox bootstrap/runtime files
- `besser/utilities/web_modeling_editor/agent-sandbox/entrypoint.sh`: initialize `/tmp/sessions` with restrictive permissions
- `besser/utilities/web_modeling_editor/agent-sandbox/Dockerfile`:
- preload NLTK tokenizer data (`punkt`, `punkt_tab`) under `/opt/nltk_data`
- expose `NLTK_DATA` for runtime consistency
- adjust runtime model to support per-session ownership/privilege-drop workflow
- fix docker compose integration in `docker-compose.yml`
- set backend sandbox URL to service DNS (`http://besser-agent-sandbox:8001`)
- add `AGENT_TEST_REQUIRE_AUTH` env passthrough
- add compose-effective sandbox limits (`cpus`, `mem_limit`, `pids_limit`)
… feature/agent-live-testing
…of agents with custom Python code bodies
… prompts, store_in_session, send_reply across all LLM actions
…de generation Previously, when a `when_file_received` transition had a list of allowed file types, the BAF template and code builder both treated `allowed_types` as a plain string, collapsing any list to its repr instead of emitting a proper Python list literal. This caused broken generated code for multi-type file transitions.
- BAFGenerator creates __init__.py in the guis/ output directory so it is importable as a package - Agent simulator router collects generated GUI files into support_files and restores gui_models onto the loaded agent after import
Enforce that custom code bodies contain exactly one top-level function definition with 'session' as the first parameter. This prevents module-level code injection via exec_module() in the backend process. A second validation pass (simulation=True) runs just before exec_module() and additionally blocks high-risk imports and dangerous built-in calls (exec, eval, open, etc.) for the simulation path only. Standard code generation is unrestricted beyond the structural rules. New file: services/validators/python_code_validator.py New exception: CodeValidationError in services/exceptions.py
…/limits endpoint Blocker 2: agent_simulator_router.py ran exec_module() in the backend container, which holds .env secrets, GitHub OAuth tokens, and Docker access. Tool.code from diagram JSON was emitted verbatim by the builder and was not covered by validate_custom_code_action (state-body CustomCodeAction only), creating an unvalidated RCE path. The validator itself is a bypassable AST denylist and must never be the isolation boundary. Fix: - Remove spec.loader.exec_module(); use the in-memory agent_model from process_agent_diagram directly (it is already a complete Agent object). - Extend the pre-exec validation to also cover Tool.code and fallback bodies (both were skipped before). - Add a comment explaining why exec belongs only in the sandbox. Blocker 8: the frontend's fetchLimitsThunk called GET /simulation/limits which had no backend route, killing the quota/api-key mode on arrival. Fix: add GET /limits endpoint that reads AGENT_SIMULATOR_MEMORY_MB, AGENT_SIMULATOR_CPU_CORES, AGENT_SIMULATOR_DISK_MB, AGENT_SIMULATOR_SESSION_LIFETIME_SECONDS, and AGENT_SIMULATOR_QUOTA_ENABLED env vars and returns them as a typed response. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…rt is lossless The code builder never emitted Agent.gui_models, so every exec'd module started with an empty dict. Four call sites worked around this by hand-patching resolved_agent.gui_models = getattr(agent_model, ...) after exec. Export a BUML .py file, import it back, and every GUI screen definition vanished silently. Fix: - Add `import json` to agent_model_builder.py (module-level) and to the generated file header. - Serialize gui_models at the end of the generated file as agent.gui_models = json.loads(<json_string>), mirroring how the diagram processor populates it. - Remove the four hand-patches in conversion_router.py, generation_router.py, agent_simulator_router.py (already removed in the blocker-2 commit), and agent_generation_utils.py — the exec'd module's agent now carries gui_models by itself. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
….message_id Two data-loss bugs on the BUML→JSON export path: 1. when_form_submitted had no branch in agent_diagram_converter.py; the call chain fell through to the default and was exported as a plain auto transition. On re-import the form transition silently became go_to(), losing the FormSubmitMatcher entirely. Fix: add elif chain_attr == "when_form_submitted" branch that extracts the form_id keyword and writes it as predefined_block["formGuiId"], matching the key the processor reads. 2. GUIEvent(message_id=...) was never written back to the JSON. The processor read transition_payload["guiEventGuiId"] but transition_payload was built from only "event" and "conditions" — guiEventGuiId was never copied from custom_block, so the lookup always returned "". Fix (converter): extract message_id from the AST GUIEvent() call and write it as custom_block["guiEventGuiId"]. Fix (processor): also copy guiEventGuiId from custom_block (or the top-level relationship dict) into transition_payload so the field survives even when the JSON was produced by non-converter paths. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
A method can now declare MethodImplementationType.NEURAL_NETWORK and reference an NN model through Method.neural_network. The editor stores the link as neuralNetworkId; both converters and the BUML code builder carry it. The Backend generator (and so the WebApp generator) turns such a method into a class-level endpoint that runs the network: it emits the PyTorch module under neural_networks/, a weights folder, nn_runtime.py and the torch requirement. The endpoint answers 503 until the trained weights are in place. Project generation links methods to the project's NNDiagrams, and project-mode backend generation now uses the active ClassDiagram. Refs #570
…ation feat: let a method be implemented by a neural network
JPA entities need an @id, and the generator rejected every class without an is_id attribute. None of the editor's templates mark one, so generating a Spring backend from them always failed with a 400. A class hierarchy that marks no identifier now gets an `id: int` on its root class, which becomes an auto-generated Integer @id; an existing attribute named `id` is promoted instead. Generation works on a copy, so the input model is not modified.
…n and new parameters
Review follow-up for #587, following the Spec-Driven Agent's sandbox (#604). - Each agent session runs in bubblewrap: its own user, PID, IPC and UTS namespaces, a private /proc, a read-only filesystem and only its own work dir writable. Sibling sessions and /proc/1/environ are out of view. It fails closed when the sandbox cannot start (AGENT_SIMULATOR_SANDBOX=off is a single-tenant opt-out). - A session's whole process group is killed on stop or expiry, and its UID and port only return to the pool once nothing runs under them. Popen drops supplementary groups; an rlimit that cannot be set fails the session instead of logging a warning. - The file listing no longer follows symlinks (lstat, O_NOFOLLOW, realpath check), so agent code cannot read files outside its work dir through it. - Every simulator endpoint requires X-Agent-Simulator-Token; without AGENT_SIMULATOR_API_TOKEN configured the simulator refuses requests. Agent websockets bind to 127.0.0.1. - Compose: no env_file (explicit env list), cap_drop ALL plus the five caps the UID drop needs, no-new-privileges, pid/memory/cpu limits, tmpfs for sessions, a healthcheck, and its own network the backend joins. The image is built from this repo, pinned, Python 3.12, and deploy-wme.yml builds and deploys it. - The directory is renamed agent_simulator so it is an importable package. - New shared besser.utilities.path_utils.normalize_relative_path. - Tests for the sandbox, session manager and API; docs page for the service and its environment variables.
…d sessions to their owner Review follow-up for #587. - Endpoints use @handle_endpoint_errors and the custom exceptions; request and response models live in models/, settings in constants.py, each read once from the environment. The duplicate env-var aliases are gone. - CodeValidationError is a ValidationError, so invalid custom code gives a 400 with its message on /generate-output, /export-buml and the simulation routes instead of a 500. - A session belongs to the actor that created it: another user gets 404 on its files and delete and 4404 on its websocket. One session per actor by default, so a single user cannot take every simulator slot. - The browser websocket authenticates with a first frame instead of a query-string token. Every call to the simulator sends X-Agent-Simulator-Token. - The rate limiter prunes idle keys and is bounded; the single-process assumption is documented. - AGENT_SIMULATOR_RESTRICT_CUSTOM_CODE defaults to true. The code linter is documented as early feedback, not a security boundary, and tools get their own lint (they have no `session` parameter). - The GitHub-session check is shared from routers/auth.py. - The default simulator URL matches the compose service name. - Router and validator tests; the simulation endpoints are documented.
…mport Review follow-up for #587. - Agent.gui_models holds GUIModel objects (validated, add_gui_model), not raw GrapesJS JSON. The processor builds them from the AgentGUI components; Agent.validate() reports a GUIReplyAction whose gui_id is not registered. - BUML to JSON now handles GUIReplyAction and gui_models, so exporting and re-importing an agent keeps its GUI replies and GUI definitions. - The builder writes GUIEvent(message_id=...), no longer drops the filter, and no longer emits `from __future__ import annotations`, which broke project exports containing an agent (SyntaxError mid-file). - sanitize_text no longer escapes quotes and backslashes into model strings; escaping happens once, at code emission. Before, every round trip doubled them. - input_prompt_mode is validated (custom requires a prompt); the new WebCrawlLLMReply parameters come after llm_name so positional calls keep working; send_reply and the other new fields are documented and shown in __repr__. - Direct attribute access instead of getattr guards; the body and fallback body share one helper per reply type. - Round-trip tests for every new field, all three component layouts (components, elements, legacy agentComponents), project export and re-import, and text with quotes and backslashes.
…ates Review follow-up for #587. - use_ui is False only in test mode. The PR hard-coded use_ui=False for a model-declared WebSocket platform, which removed the UI from every existing generated agent. - The session-variable interpolation (pasted about 16 times) is one Jinja macro, and the body and fallback share one action macro. Generated code for existing agents is byte-identical apart from the use_ui fix. - guis/<id>.py is written by gui_model_to_code plus the agent_gui.py.j2 template instead of inline f-strings, and the generator no longer imports the web editor backend. A GUIReplyAction whose GUI is missing, or two gui ids with the same module name, raise instead of generating a gui_model = None stub. - allowed_types go through tojson; direct attribute access instead of getattr guards; workspace paths use normalize_relative_path and a path escaping the output dir is an error. - Tests for every new template branch, each generated file compiled. - baf.rst: the GUI example is complete and runs; test_mode and the new reply options are documented.
… itself CodeQL (py/clear-text-logging-sensitive-data) read the env-var name constant passed to the logger as token data. The messages now state the variable name literally; no value was ever logged.
- `from besser.generators.agents.baf_generator import BAFGenerator` as a first import raised a circular ImportError (the generator imported the web editor backend at module level, which imports the generator registry). The backend import now happens where the personalized JSON export uses it. The conftest stubs that hid this in tests are removed, and a fresh-interpreter test covers both import paths. - Generated agents called use_websocket_platform twice: a standalone platform loop rendered before the config/fallback block, and the found_platform flag did not survive the Jinja loop. The platform is now created exactly once. Both issues already existed on development.
- agent.rst: RAG examples use valid names (no spaces); hybrid BM25 is marked Python-API only; DBReply lists its accepted values; the Greetings example validates; every fragment runs in page order. - baf.rst: the generated file is <agent name>.py; every example runs. - web_editor_backend.rst / agent_simulator.rst: linked both ways; /limits nulls, WS close 1011 and HTTP 429 at capacity documented; LLM keys also reach the agent through config.yaml. - web_editor.rst: Agent Simulation section linking the backend pages and the editor user guide. - CLAUDE.md: agent_simulator_router.py and the simulator service.
…n the JSON structure for class diagrams
Feature/agent live testing
… neural network - Bump setup.cfg 7.17.2 -> 7.18.0 and add v7.18.0 release notes. - Frontend submodule bump to develop f9baafad: the agent live testing editor work (#184: simulation page, agent components and runtime pages, component migration, assistant transition fixes) on top of the NN method option (#199). - Backend: agent live testing (#587: sandboxed agent simulator service, session-variable and GUI replies, BAF generator fixes), methods implemented by a neural network (#613, closes #570) and the Spring generator's generated id for classes without one.
2 of 4 tasks
…ntation Review follow-up for #615. - BinaryAssociation checks both navigability rules in one place and names the association and end in its errors. DomainModel.validate() reports them too, so ends made invalid after construction are caught. - The class-diagram processor accepts only a real boolean for `navigable`; anything else falls back to the legacy default with a warning. The correction branches describe what they do, and aggregation follows the same path. - BUML to JSON emits plain associations as ClassBidirectional with explicit per-end navigable and keeps the drawn orientation. The navigability-based end swap flipped one-way associations and put their layout on the wrong class. - Tests for the metamodel rules, JSON -> BUML -> JSON round trips of every navigability combination, legacy payloads and the correction warnings, and the BUML code builder; structural.rst documents navigability.
Navigability constraints on binary associations
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Minor release: live agent simulation in the web editor, and methods implemented by a neural network. Per-end navigability by @ivan-alfonso (#615, editor #200). The agent live-testing work was contributed by @mgv99 (#587, editor #184, modeling-agent #19) and reviewed and completed by the BESSER team. Methods backed by a neural network close #570.
This release adds a new deployable service, the agent simulator. See Deployment below before upgrading a hosted instance.
Highlights
/proc, read-only filesystem except its work dir) and fails closed. The simulator container has no host port, an explicit env allowlist and dropped capabilities, and requiresX-Agent-Simulator-Token. Custom code is refused by default, GitHub sign-in is required, and each user gets one session at a time.store_in_session,send_reply, custom input prompts,GUIReplyAction,GUIEventandwhen_form_submitted;Agent.gui_modelsholdsGUIModelobjects.MethodImplementationType.NEURAL_NETWORK): the Backend and Web App generators emit the PyTorch module, annn_runtime.pyand an endpoint that runs the network (503 until weights are provided).BinaryAssociationenforces at least one navigable end and a navigable composition part end (also invalidate()), and old diagrams, templates and assistant output are migrated on load.idon the hierarchy root; the input model is untouched.Fixes
from __future__mid-file).BAFGeneratorfailed to import on its own (circular import).add_intent_training_phrasewas ignored.Deployment
besser-wme-agent-simulatorservice in both compose files, built by the deploy workflow; the backend joinsagent_simulator_network.AGENT_SIMULATOR_API_TOKENfor both backend and simulator.seccomp=unconfined/apparmor=unconfined(already in compose) and unprivileged user namespaces on the host.developincludes Add @property decorator to associations and generalizations methods #19).Test plan
docker composewith a real agent simulated from the editor