Verified task planning from compiled equipment manuals, E5 retrieval, and guarded Dobot execution for robotic industrial panel operation.
MaCoPlanner conditions a vision-language planner on task-aligned manual knowledge, retrieves relevant procedural constraints, checks candidate action sequences, and performs symbolic rollout with a device-specific safety state machine. A failed candidate is returned to the model for targeted repair. If the repair budget is exhausted, the task is rejected and no non-compliant plan is exposed as an executable result.
MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation
Guipeng Xin, Jiahe Xu, Mohammad Deghat, Chenhui Wan, Jie Liu, Youmin Hu, and Zhongxu Hu
Abstract
Robotic industrial panel operation requires not only accurate control localization but also compliance with operating procedures, safety rules, and device-state constraints distributed across heterogeneous manuals. This study presents MaCoPlanner, a task-planning framework built on knowledge compiled from equipment manuals that converts equipment manuals into a typed intermediate representation, retrieves task- and state-relevant evidence, and uses it to support plan generation. Before actuation, candidate plans are symbolically rolled out and checked against procedural and state-transition constraints; detected violations are localized and returned for targeted repair, while unresolved plans are rejected. A separate execution interface grounds verified symbolic actions to physical controls and updates the device state. Under an independent evaluation oracle, MaCoPlanner achieves a final violation rate of 2.7%, and 26.3% of the runs in the repair analysis are rejected after exhausting the refinement budget. Compared with Raw-Manual, task success increases from 62.8% to 84.4% on Level-2 tasks and from 25.9% to 43.2% on Level-3 tasks. Experiments on a controller-panel simulator without an attached industrial load further demonstrate integrated execution feasibility under representative interaction conditions, without claiming industrial deployment readiness.
- One leak-free planning API and CLI shared by engine, generator, and VFD profiles.
- Device-specific primitive schemas, rule checks, and symbolic state machines.
- 120 benchmark tasks, compiled knowledge, and formal-rule records per device.
- A versioned 400-case retrieval benchmark and 1,440 structured planning evaluations.
- E5 dense retrieval by default; TF-IDF and Haystack BM25 remain explicit alternatives.
- A strict execution adapter for move, press, turn, and observe primitives; dry-run by default.
- Offline verification for explicit plans, batch inference, packaging metadata, GitHub CI, and tests.
Raw manuals, historical versions, ROS build trees, vendor drivers, hardware calibration, raw experiment logs, model weights/caches, transcripts, and unrelated generated outputs are intentionally not part of this release. See the release audit.
MaCoPlanner requires Python 3.10 or newer.
python -m venv .venv
python -m pip install -U pip
python -m pip install -e .To use Haystack BM25 instead of the built-in TF-IDF retriever:
python -m pip install -e ".[haystack]"Install E5 dense retrieval or OpenCV-based execution when needed:
python -m pip install -e ".[e5]"
python -m pip install -e ".[execution]"Set your API key in the process environment. .env is ignored by Git, but this package does not
silently load it.
# Linux/macOS
export OPENAI_API_KEY="..."
# Windows PowerShell
$env:OPENAI_API_KEY = "..."List bundled tasks:
macoplanner list-tasks --device engineVerify an explicit plan without making an API call:
macoplanner verify \
--device engine \
--task-id L1-01 \
--plan "move_to_pose('POSE_DRIVE_DIR'); knob_turn_with_ag95('POSE_DRIVE_DIR', target_state='FWD')" \
--prettyGenerate and verify a plan:
macoplanner plan \
--device engine \
--task-id L2-01 \
--model gpt-4o \
--rag-backend e5 \
--max-refinements 3 \
--pretty--max-refinements 3 means one initial candidate followed by at most three repair calls. The
canonical profile rejects silent changes to the model, E5 backend, retrieval K, or repair budget.
Use --allow-custom-config to run an explicitly customized experiment.
The bundled panel image is used unless --panel-image supplies a local file or HTTP(S) URL.
OPENAI_BASE_URL or --base-url may be used with a compatible endpoint.
Inspect retrieval without calling a planning model:
macoplanner retrieve \
--device engine \
--query "Switch drive direction to forward safely" \
--backend e5 \
--top-k 8 \
--prettyThe first online E5 run may download
intfloat/e5-base-v2 into the normal Hugging Face
cache; model files are not stored in this repository. Use --local-files-only to prohibit
downloads. E5's input and language limits are documented in
docs/retrieval.md.
Inspect the active runtime settings:
macoplanner settings --prettyVerify and dry-run a physical action sequence:
macoplanner execute \
--device generator \
--task-id L1-01 \
--plan-file examples/generator_l1_01.plan.txt \
--control-map examples/control_map.example.json \
--prettyThis example never contacts ROS2 or hardware. Live motion requires a separately installed Dobot
driver, an installation-specific calibrated control map, --live, and
--acknowledge-hardware-risk. See docs/execution.md.
Run selected task IDs and write auditable JSON Lines output:
macoplanner batch \
--device vfd \
--task-id L1-01 \
--task-id L2-01 \
--output results/vfd.jsonlOmit --task-id to run all bundled tasks. The output is excluded by .gitignore.
Research ablations are available through --ablation:
off: full retrieval, verification, iterative repair, and rejection.rag_nl_only: one retrieval-conditioned generation; output remains marked unverified unless it passes the checks.ltl_full_only: one generation followed by checks without iterative repair.no_safety: generation only; output is always marked unverified and is never symbolically rolled out by the public pipeline.
The reproducibility benchmarks bundled with the package can be discovered and fully checked offline:
macoplanner benchmark list --pretty
macoplanner benchmark validate --prettyThe validator checks all 400 retrieval cases, 1,440 per-case planning records, and recomputes the 12 aggregate rows. See the benchmark data card for schemas, provenance, known task snapshot differences, and limitations.
from macoplanner import plan_task, retrieve_rules, verify_plan
result = plan_task("generator", "L1-01", model="gpt-4o")
if result.accepted:
print(result.verified_plan)
else:
print(result.status, result.violations)
offline = verify_plan(
"engine",
"L1-01",
"move_to_pose('POSE_DRIVE_DIR'); "
"knob_turn_with_ag95('POSE_DRIVE_DIR', target_state='FWD')",
)
hits = retrieve_rules("vfd", "start only after the fault is cleared", backend="e5")PlanResult.status is verified, rejected, or unverified. Only verified results populate
verified_plan. attempts counts all planner calls, while refinements excludes the initial
candidate. Benchmark reference plans are never used as model fallbacks.
MaCoPlanner/
├── src/macoplanner/
│ ├── pipeline.py # shared leak-free orchestration API
│ ├── cli.py # planning, retrieval, execution, and data commands
│ ├── settings.py # immutable runtime configuration
│ ├── runtime_helpers.py # deterministic configuration helpers
│ ├── devices/ # action schemas, checks, and FSMs
│ ├── retrieval/ # E5 index, typed fusion, and evidence gate
│ ├── execution/ # strict actions, control maps, Dobot and vision adapters
│ ├── benchmarks/ # loaders, validator, 400 retrieval cases, 1,440 evaluations
│ └── resources/ # tasks, rules, compiled knowledge, panel images
├── examples/ # non-calibrated control-map template
├── docs/ # retrieval, execution, benchmark, and release-audit notes
├── .github/ # CI and contribution templates
├── tests/ # offline tests
├── pyproject.toml
├── LICENSE
└── CITATION.cff
Every device resource directory contains:
tasks.json: an array withtask_id,task_description,state_before, and optional reference metadata;knowledge.json: an array of{task_id, knowledge}records;rules.json: a mapping from rule ID to{description, ltl};panel.jpgorpanel.png: the representative panel image.
Use --tasks, --rules, and --knowledge to override bundled files. Custom tasks must use the
selected device's canonical state keys and action vocabulary. Invalid action lines are rejected by
the offline verifier instead of being ignored.
The planning default is the intfloat/e5-base-v2 dense retriever. Sampling controls are supplied by
the configured API endpoint rather than silently hard-coded in device profiles. API model aliases
can change over time; record the exact provider snapshot, endpoint, sampling controls, dependency
lock file, retrieval backend, and JSONL output for a publication run. Canonical runtime values live
in macoplanner.settings.
MaCoPlanner verifies encoded task-level rules only. It is not a certified industrial safety system and does not cover collision avoidance, force limits, perception/grounding errors, plant interlocks, or complete execution-time fault handling. The live adapter is an integration example, not a safety controller. Do not connect it to energized or safety-critical equipment without independent safety engineering, guarding, interlocks, and qualified supervision. See SECURITY.md, NOTICE, and NOTICE.md.
We thank the authors of the following related planning and safety methods, which helped shape the research context and comparative evaluation of MaCoPlanner:
- LLM³: Large Language Model-based Task and Motion Planning with Motion Failure Reasoning (official code);
- ISR-LLM: Iterative Self-Refined Large Language Model for Long-Horizon Sequential Task Planning (official code); and
- SafePlan: Leveraging Formal Logic and Chain-of-Thought Reasoning for Enhanced Safety in LLM-based Robotic Task Planning.
The optional dense retrieval backend uses
intfloat/e5-base-v2; we thank its authors and
maintainers for releasing the model and documentation.
These links acknowledge related work; no source code from these repositories is vendored into MaCoPlanner. Please consult each project for its own license and citation requirements.
If MaCoPlanner is useful in your research, please cite the paper:
@article{xin2026macoplanner,
title = {MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation},
author = {Xin, Guipeng and Xu, Jiahe and Deghat, Mohammad and Wan, Chenhui and Liu, Jie and Hu, Youmin and Hu, Zhongxu},
year = {2026},
note = {SSRN preprint},
doi = {10.2139/ssrn.6799077},
url = {https://ssrn.com/abstract=6799077}
}Machine-readable citation metadata is available in CITATION.cff.
For questions, contact Guipeng Xin at
xinguipeng@hust.edu.cn.
python -m pip install -e ".[dev]"
python -m unittest discover -s tests -v
ruff check src tests
python -m build
twine check dist/*See CONTRIBUTING.md. MaCoPlanner is released under the MIT License.



