Skip to content

Feature/agent live testing - #587

Merged
ArmenSl merged 49 commits into
developmentfrom
feature/agent-live-testing
Sep 24, 2026
Merged

ArmenSl merged 49 commits into
developmentfrom
feature/agent-live-testing

Conversation

@mgv99

@mgv99 mgv99 commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR introduces an Agent Simulator — an isolated execution service for live-testing BESSER-generated agents directly from the web modeling editor — along with the backend routing layer to drive it. It also extends the agent metamodel and code generator with session variable interpolation, GUI reply actions, and improved LLM/RAG configurability.


  1. New Service: besser/utilities/web_modeling_editor/agent-simulator/

A standalone FastAPI microservice that runs generated agents in isolation.

agent_simulator_api.py — REST + WebSocket API for the simulator container:

  • POST /sessions — validates the session UUID, receives agent code and config, delegates to SessionManager, returns {session_id, port}
  • GET /sessions/{id}/files — returns file tree + content (≤512 KB per file) from the session work directory
  • DELETE /sessions/{id} — terminates the session and frees all resources
  • GET /health — liveness check with active session count
  • WS /sessions/{id}/ws — bidirectional relay; parses stdout for [] Running body lines and emits structured state_change events; all other lines become {type:"stdout", line:"..."} events

session_manager.py — Secure subprocess and resource management:

  • Allocates a dedicated Unix UID per session (SESSION_UID_BASE + offset) from a pre-built pool; the subprocess drops to that UID via os.setuid/os.setgid before executing
  • Assigns a port from a configurable pool (default start: 7700), injects it into config.yaml under platforms.websocket.port
  • Applies resource.setrlimit caps: 4 GB virtual memory, 120 s CPU, 100 MB file size, 64 processes, 1024 file descriptors (all env-configurable)
  • Environment passed to the subprocess is a strict allowlist (PATH, HOME, optional API keys only)
  • cleanup_expired() runs every 60 s, terminating sessions that exceeded lifetime or whose subprocess exited
  • Path traversal prevention on all file writes via _normalize_session_relative_path()

Dockerfile — Python 3.11-slim image; pre-downloads NLTK tokenizers (punkt, punkt_tab) at build time; installs CPU-only PyTorch; creates /tmp/sessions with restricted permissions; exposes port 8001

README.md — Full architecture documentation, Docker instructions, and reference for all environment variables (AGENT_SIMULATOR_MAX_SESSIONS, AGENT_SIMULATOR_SESSION_LIFETIME_SECONDS, AGENT_SIMULATOR_PORT_POOL_START, AGENT_SIMULATOR_SESSION_UID_BASE, RLIMIT_*)


  1. New Backend Router: agent_simulator_router.py

Acts as a secure middle layer between the frontend and the simulator service (http://besser-agent-simulator:8001):

┌──────────────────────────────────────────────┬─────────────────────────────────────────────────────────────┐
│ Endpoint │ Purpose │
├──────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
│ POST /besser_api/simulation/sessions │ Generates agent code, validates it, and creates a simulator │
│ │ session │
├──────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
│ POST /besser_api/simulation/validate │ Runs code generation + validation without starting a │
│ │ session; returns agent code and event list │
├──────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
│ GET │ Proxies the file listing from the simulator │
│ /besser_api/simulation/sessions/{id}/files │ │
├──────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
│ DELETE /besser_api/simulation/sessions/{id} │ Terminates the session │
├──────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
│ WS /besser_api/simulation/{id}/ws │ Two-coroutine relay bridging the frontend WebSocket to the │
│ │ simulator │
└──────────────────────────────────────────────┴─────────────────────────────────────────────────────────────┘

Authentication: Optional GitHub session enforcement via X-GitHub-Session header (query param for WS); returns 401 when required but absent/expired.

Rate limiting: Sliding-window limiter, 12 requests / 60 s per actor (GitHub session ID or client IP); 429 on breach.

Custom code restriction: Optional flag (AGENT_SIMULATOR_RESTRICT_CUSTOM_CODE) blocks agents containing CustomCodeAction or replyType: "code".


  1. Agent Metamodel Extensions: besser/BUML/metamodel/state_machine/agent.py

New session-variable interpolation across all reply action types:

  • AgentReply, WebSocketReplyMarkdown/HTML/Speech — use_session_vars: bool = False; {key} placeholders replaced at runtime from session.get("key")
  • LLMReply, LLMChatReply — system_prompt_use_session_vars, store_in_session, send_reply
  • RAGReply — input_prompt_mode, custom_input_prompt, custom_input_prompt_use_session_vars, prompt_use_session_vars, store_in_session, send_reply
  • LLMReply / RAGReply — input_prompt_mode ('last_user_message' default) + custom_input_prompt
  • WebCrawlLLMReply — system_message_prefix_use_session_vars, store_in_session, send_reply

New classes:

  • GUIReplyAction(Action) — sends a BESSER GUI model as a chat reply; params: gui_id, persist, width, is_form
  • GUIEvent(Event) — triggered by user interaction with a GUI component; optionally filtered by message_id
  • FormSubmitMatcher(Condition) — matches GUI form submission events; optionally filtered by form_id

Agent class: new gui_models: dict[str, dict] attribute mapping GUI IDs to GrapesJS model dicts

State: new when_form_submitted(form_id?) builder method


  1. BAF Generator Changes: besser/generators/agents/baf_generator.py
  • New test_mode: bool = False parameter; when True, pre-creates declared workspace directories before running the agent
  • New workspace_rel_dir() helper sanitizes paths (strips drive letters, prevents traversal)
  • New extract_braced_vars() helper for template string variable extraction
  • GUI generation block: collects unique GUIReplyAction instances, creates guis/ package, calls process_gui_diagram + gui_model_to_code per GUI, falls back to gui_model = None stub on failure

  1. Jinja2 Template Changes: baf_agent_template.py.j2
  • Streamlit platform conditionally sets use_ui=False when test_mode=True
  • Logging changed from basicConfig to logger.setLevel(logging.INFO)
  • RAG load_pdfs wrapped in os.path.exists guard
  • GUI imports emitted from collected GUIReplyAction instances
  • States now explicitly emit initial=False for non-initial states
  • All action types emit session-var interpolation, store_in_session, and send_reply logic
  • GUIReplyAction renders platform.reply_gui(session, <gui_var>)
  • FileTypeMatcher: fixed list rendering for allowed_types
  • New FormSubmitMatcher transition branch
  • New GUIEvent transition branch

  1. Agent Model Builder Changes: agent_model_builder.py
  • Added from future import annotations to generated headers
  • Added GUIReplyAction, GUIEvent to generated imports
  • set_default_llm now always emits even if the default matches the first LLM
  • Serialization updated for all new reply action parameters
  • New GUIReplyAction serialization block
  • GUIEvent + FormSubmitMatcher transition serialization
  • FileTypeMatcher list handling fixed

mgv99 added 29 commits June 4, 2026 14:35
… and align frontend auth flow

- secure backend agent-test routes in `besser/utilities/web_modeling_editor/backend/routers/agent_test_router.py`
  - add optional auth gate via `AGENT_TEST_REQUIRE_AUTH` (GitHub session validation)
  - add in-memory sliding-window rate limiter (`AGENT_TEST_RATE_LIMIT_WINDOW_SECONDS`, `AGENT_TEST_RATE_LIMIT_MAX_REQUESTS`)
  - validate session IDs as UUIDs on REST + WS paths
  - enforce WS auth using `github_session` query param and return explicit WS close codes for auth/input failures

- harden sandbox API/session boundaries in `besser/utilities/web_modeling_editor/agent-sandbox/agent_sandbox_api.py`
  - validate `session_id` on create/delete/ws
  - add startup permissions bootstrap for `/tmp/sessions` (`root:root`, `0711`)

- strengthen sandbox process + filesystem isolation in `besser/utilities/web_modeling_editor/agent-sandbox/session_manager.py`
  - add thread-safe state management with `RLock`
  - add safe session path resolution and canonical UUID normalization
  - add per-session UID allocation (`AGENT_SANDBOX_SESSION_UID_BASE`) and release lifecycle
  - drop each agent subprocess privileges to its dedicated UID/GID (`setuid`/`setgid`)
  - enforce strict per-session FS permissions (`0700` dirs, `0600` files), isolated `HOME`/`TMPDIR`
  - replace inherited env with strict allowlist to avoid container secret leakage
  - keep existing resource limits and fail fast if privilege drop cannot be applied

- update sandbox bootstrap/runtime files
  - `besser/utilities/web_modeling_editor/agent-sandbox/entrypoint.sh`: initialize `/tmp/sessions` with restrictive permissions
  - `besser/utilities/web_modeling_editor/agent-sandbox/Dockerfile`:
    - preload NLTK tokenizer data (`punkt`, `punkt_tab`) under `/opt/nltk_data`
    - expose `NLTK_DATA` for runtime consistency
    - adjust runtime model to support per-session ownership/privilege-drop workflow

- fix docker compose integration in `docker-compose.yml`
  - set backend sandbox URL to service DNS (`http://besser-agent-sandbox:8001`)
  - add `AGENT_TEST_REQUIRE_AUTH` env passthrough
  - add compose-effective sandbox limits (`cpus`, `mem_limit`, `pids_limit`)
… prompts, store_in_session, send_reply across all LLM actions
…de generation

Previously, when a `when_file_received` transition had a list of allowed
file types, the BAF template and code builder both treated `allowed_types`
as a plain string, collapsing any list to its repr instead of emitting a
proper Python list literal. This caused broken generated code for multi-type file transitions.
- BAFGenerator creates __init__.py in the guis/ output directory so it is importable as a package

- Agent simulator router collects generated GUI files into support_files and restores gui_models onto the loaded agent after import
Enforce that custom code bodies contain exactly one top-level function
definition with 'session' as the first parameter. This prevents
module-level code injection via exec_module() in the backend process.

A second validation pass (simulation=True) runs just before exec_module()
and additionally blocks high-risk imports and dangerous built-in calls
(exec, eval, open, etc.) for the simulation path only. Standard code
generation is unrestricted beyond the structural rules.

New file: services/validators/python_code_validator.py
New exception: CodeValidationError in services/exceptions.py
@mgv99
mgv99 requested a review from ArmenSl September 2, 2026 18:01
mgv99 and others added 6 commits September 14, 2026 13:06
…ression

Blocker 1: session-var field names (store_in_session, form_id, message_id)
were interpolated raw into Python string literals.  The only upstream
escaping, sanitize_text, escaped single quotes but NOT backslashes, so a
value ending \' could close the literal and inject arbitrary Python into
every generated agent.
- Fix sanitize_text to escape \ before ' (text_parser.py).
- Use | string | tojson for all store_in_session, form_id, and message_id
  occurrences in the template (mirrors state.name | string | tojson already
  present at line 602).

Blocker 6: line 71 called logger.setLevel() on a name the template never
defines — logger resolved only through a `from …base_events import *` leak.
Every regenerated agent silently lost its handler and format; the call would
break outright if BAF adds __all__.
- Emit logging.basicConfig(...) to install a handler and logger = logging.getLogger(__name__).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…circular dep

baf_generator.py:22's top-level import of process_gui_diagram created the
cycle baf_generator → services.converters → config.generators → baf_generator.
This caused pytest tests/generators/agents/ to error at collection on this
branch (while the full-suite run stayed green because the backend imports
first and caches the modules).

Fix: remove the module-level import; add the deferred import immediately
before the first call site inside generate(), which runs only at generation
time, never at import time.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…/limits endpoint

Blocker 2: agent_simulator_router.py ran exec_module() in the backend
container, which holds .env secrets, GitHub OAuth tokens, and Docker
access.  Tool.code from diagram JSON was emitted verbatim by the builder
and was not covered by validate_custom_code_action (state-body
CustomCodeAction only), creating an unvalidated RCE path.  The validator
itself is a bypassable AST denylist and must never be the isolation
boundary.

Fix:
- Remove spec.loader.exec_module(); use the in-memory agent_model from
  process_agent_diagram directly (it is already a complete Agent object).
- Extend the pre-exec validation to also cover Tool.code and fallback
  bodies (both were skipped before).
- Add a comment explaining why exec belongs only in the sandbox.

Blocker 8: the frontend's fetchLimitsThunk called GET /simulation/limits
which had no backend route, killing the quota/api-key mode on arrival.

Fix: add GET /limits endpoint that reads AGENT_SIMULATOR_MEMORY_MB,
AGENT_SIMULATOR_CPU_CORES, AGENT_SIMULATOR_DISK_MB,
AGENT_SIMULATOR_SESSION_LIFETIME_SECONDS, and AGENT_SIMULATOR_QUOTA_ENABLED
env vars and returns them as a typed response.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…rt is lossless

The code builder never emitted Agent.gui_models, so every exec'd module
started with an empty dict.  Four call sites worked around this by
hand-patching resolved_agent.gui_models = getattr(agent_model, ...) after
exec.  Export a BUML .py file, import it back, and every GUI screen
definition vanished silently.

Fix:
- Add `import json` to agent_model_builder.py (module-level) and to the
  generated file header.
- Serialize gui_models at the end of the generated file as
  agent.gui_models = json.loads(<json_string>), mirroring how the diagram
  processor populates it.
- Remove the four hand-patches in conversion_router.py, generation_router.py,
  agent_simulator_router.py (already removed in the blocker-2 commit), and
  agent_generation_utils.py — the exec'd module's agent now carries
  gui_models by itself.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
….message_id

Two data-loss bugs on the BUML→JSON export path:

1. when_form_submitted had no branch in agent_diagram_converter.py; the
   call chain fell through to the default and was exported as a plain auto
   transition.  On re-import the form transition silently became go_to(),
   losing the FormSubmitMatcher entirely.
   Fix: add elif chain_attr == "when_form_submitted" branch that extracts
   the form_id keyword and writes it as predefined_block["formGuiId"],
   matching the key the processor reads.

2. GUIEvent(message_id=...) was never written back to the JSON.  The
   processor read transition_payload["guiEventGuiId"] but transition_payload
   was built from only "event" and "conditions" — guiEventGuiId was never
   copied from custom_block, so the lookup always returned "".
   Fix (converter): extract message_id from the AST GUIEvent() call and
   write it as custom_block["guiEventGuiId"].
   Fix (processor): also copy guiEventGuiId from custom_block (or the
   top-level relationship dict) into transition_payload so the field
   survives even when the JSON was produced by non-converter paths.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Copilot stopped work on behalf of mgv99 due to an error September 14, 2026 11:15
mgv99 and others added 8 commits September 14, 2026 13:16
Review follow-up for #587, following the Spec-Driven Agent's sandbox (#604).

- Each agent session runs in bubblewrap: its own user, PID, IPC and UTS
  namespaces, a private /proc, a read-only filesystem and only its own work
  dir writable. Sibling sessions and /proc/1/environ are out of view. It
  fails closed when the sandbox cannot start (AGENT_SIMULATOR_SANDBOX=off
  is a single-tenant opt-out).
- A session's whole process group is killed on stop or expiry, and its UID
  and port only return to the pool once nothing runs under them. Popen
  drops supplementary groups; an rlimit that cannot be set fails the
  session instead of logging a warning.
- The file listing no longer follows symlinks (lstat, O_NOFOLLOW, realpath
  check), so agent code cannot read files outside its work dir through it.
- Every simulator endpoint requires X-Agent-Simulator-Token; without
  AGENT_SIMULATOR_API_TOKEN configured the simulator refuses requests.
  Agent websockets bind to 127.0.0.1.
- Compose: no env_file (explicit env list), cap_drop ALL plus the five caps
  the UID drop needs, no-new-privileges, pid/memory/cpu limits, tmpfs for
  sessions, a healthcheck, and its own network the backend joins. The
  image is built from this repo, pinned, Python 3.12, and deploy-wme.yml
  builds and deploys it.
- The directory is renamed agent_simulator so it is an importable package.
- New shared besser.utilities.path_utils.normalize_relative_path.
- Tests for the sandbox, session manager and API; docs page for the
  service and its environment variables.
…d sessions to their owner

Review follow-up for #587.

- Endpoints use @handle_endpoint_errors and the custom exceptions; request
  and response models live in models/, settings in constants.py, each read
  once from the environment. The duplicate env-var aliases are gone.
- CodeValidationError is a ValidationError, so invalid custom code gives a
  400 with its message on /generate-output, /export-buml and the
  simulation routes instead of a 500.
- A session belongs to the actor that created it: another user gets 404 on
  its files and delete and 4404 on its websocket. One session per actor by
  default, so a single user cannot take every simulator slot.
- The browser websocket authenticates with a first frame instead of a
  query-string token. Every call to the simulator sends
  X-Agent-Simulator-Token.
- The rate limiter prunes idle keys and is bounded; the single-process
  assumption is documented.
- AGENT_SIMULATOR_RESTRICT_CUSTOM_CODE defaults to true. The code linter is
  documented as early feedback, not a security boundary, and tools get
  their own lint (they have no `session` parameter).
- The GitHub-session check is shared from routers/auth.py.
- The default simulator URL matches the compose service name.
- Router and validator tests; the simulation endpoints are documented.
…mport

Review follow-up for #587.

- Agent.gui_models holds GUIModel objects (validated, add_gui_model),
  not raw GrapesJS JSON. The processor builds them from the AgentGUI
  components; Agent.validate() reports a GUIReplyAction whose gui_id is not
  registered.
- BUML to JSON now handles GUIReplyAction and gui_models, so exporting and
  re-importing an agent keeps its GUI replies and GUI definitions.
- The builder writes GUIEvent(message_id=...), no longer drops the filter,
  and no longer emits `from __future__ import annotations`, which broke
  project exports containing an agent (SyntaxError mid-file).
- sanitize_text no longer escapes quotes and backslashes into model
  strings; escaping happens once, at code emission. Before, every round
  trip doubled them.
- input_prompt_mode is validated (custom requires a prompt); the new
  WebCrawlLLMReply parameters come after llm_name so positional calls keep
  working; send_reply and the other new fields are documented and shown in
  __repr__.
- Direct attribute access instead of getattr guards; the body and
  fallback body share one helper per reply type.
- Round-trip tests for every new field, all three component layouts
  (components, elements, legacy agentComponents), project export and
  re-import, and text with quotes and backslashes.
…ates

Review follow-up for #587.

- use_ui is False only in test mode. The PR hard-coded use_ui=False for a
  model-declared WebSocket platform, which removed the UI from every
  existing generated agent.
- The session-variable interpolation (pasted about 16 times) is one Jinja
  macro, and the body and fallback share one action macro. Generated code
  for existing agents is byte-identical apart from the use_ui fix.
- guis/<id>.py is written by gui_model_to_code plus the agent_gui.py.j2
  template instead of inline f-strings, and the generator no longer
  imports the web editor backend. A GUIReplyAction whose GUI is missing,
  or two gui ids with the same module name, raise instead of generating a
  gui_model = None stub.
- allowed_types go through tojson; direct attribute access instead of
  getattr guards; workspace paths use normalize_relative_path and a path
  escaping the output dir is an error.
- Tests for every new template branch, each generated file compiled.
- baf.rst: the GUI example is complete and runs; test_mode and the new
  reply options are documented.
… itself

CodeQL (py/clear-text-logging-sensitive-data) read the env-var name
constant passed to the logger as token data. The messages now state the
variable name literally; no value was ever logged.
@ArmenSl
ArmenSl requested a balanced review from Copilot September 23, 2026 15:48

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

- `from besser.generators.agents.baf_generator import BAFGenerator` as a
  first import raised a circular ImportError (the generator imported the
  web editor backend at module level, which imports the generator
  registry). The backend import now happens where the personalized JSON
  export uses it. The conftest stubs that hid this in tests are removed,
  and a fresh-interpreter test covers both import paths.
- Generated agents called use_websocket_platform twice: a standalone
  platform loop rendered before the config/fallback block, and the
  found_platform flag did not survive the Jinja loop. The platform is now
  created exactly once. Both issues already existed on development.
- agent.rst: RAG examples use valid names (no spaces); hybrid BM25 is
  marked Python-API only; DBReply lists its accepted values; the
  Greetings example validates; every fragment runs in page order.
- baf.rst: the generated file is <agent name>.py; every example runs.
- web_editor_backend.rst / agent_simulator.rst: linked both ways; /limits
  nulls, WS close 1011 and HTTP 429 at capacity documented; LLM keys also
  reach the agent through config.yaml.
- web_editor.rst: Agent Simulation section linking the backend pages and
  the editor user guide.
- CLAUDE.md: agent_simulator_router.py and the simulator service.
@ArmenSl
ArmenSl merged commit 51ae322 into development Sep 24, 2026
5 checks passed
ArmenSl added a commit that referenced this pull request Sep 24, 2026
… neural network

- Bump setup.cfg 7.17.2 -> 7.18.0 and add v7.18.0 release notes.
- Frontend submodule bump to develop f9baafad: the agent live testing
  editor work (#184: simulation page, agent components and runtime pages,
  component migration, assistant transition fixes) on top of the NN method
  option (#199).
- Backend: agent live testing (#587: sandboxed agent simulator service,
  session-variable and GUI replies, BAF generator fixes), methods
  implemented by a neural network (#613, closes #570) and the Spring
  generator's generated id for classes without one.
@ArmenSl
ArmenSl deleted the feature/agent-live-testing branch September 24, 2026 11:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants