RISEBench++
Xue Yang†, Peiyuan Zhang*, Yilun Zhu*, Qihao Yang*, Mingxin Liu, Xiangyu Zhao, Ziqian Fan, Zhaokai Wang, Yan Li, Yifan Yang, Xu Yang, Xiaosong Jia, Yue Zhou, Zhihang Zhong, Junchi Yan
* Core student contributors · † Corresponding author
Shanghai Jiao Tong University · Southeast University · South China University of Technology · Fudan University · East China Normal University
If you find our work helpful, please consider giving us a ⭐ or citation 😊
In this work, we introduce RISEBench++, an extension of RISEBench for evaluating Reasoning-Informed viSual Editing (RISE). RISEBench++ focuses on six key reasoning types: Temporal, Causal, Spatial, Logical, Counterfactual, and Hybrid Reasoning.
RISEBench++ comprises 1,000 human-annotated test cases, provided in English and Chinese, with a hierarchical taxonomy spanning six reasoning dimensions, 12 subcategories, and 65 fine-grained task types. It supports single-image, multi-image, and multi-turn editing. Temporal, causal, spatial, and logical reasoning each contain 200 cases; counterfactual and hybrid reasoning each contain 100 cases. Hybrid cases contain three to five successive editing instructions.
To comprehensively assess model performance across diverse task types, we define three key evaluation dimensions: Instruction Reasoning, Appearance Consistency, and Visual Plausibility.
Besides, we design a robust LMM-as-a-Judge evaluation pipeline with dimension-specific reference evidence and scoring rubrics. Gemini-3-Flash evaluates instruction reasoning and appearance consistency, while Gemini-3.1-Flash-Lite evaluates visual plausibility. Logical tasks are not evaluated for visual plausibility; counterfactual tasks are assessed for perceptual quality without penalizing intentionally impossible physical outcomes.
RISEBench++ aims to support comprehensive, reliable, and fine-grained evaluation of reasoning-aware visual editing across diverse models and input settings.
We also introduce RISE-Agent, a training-free framework that combines reasoning-driven planning, tool-augmented execution, and verifier-guided refinement. The planner infers the intended target state using visual reasoning, external knowledge, and computational tools; the executor realizes the edit through a generative model or precise programmatic operations; and the verifier checks the result and guides refinement or replanning. Together, they form a closed loop for reasoning-informed visual editing.
We evaluate 58 visual editing approaches, including 19 closed-source models, 34 open-source models, and 5 agentic methods. The following tables report results from the accompanying manuscript. Accuracy measures the percentage of cases that receive full marks on every applicable evaluation dimension.
All accuracy values are percentages. Single, Multi, and Total refer to single-image cases, multi-image cases, and the full benchmark, respectively. Models without multi-image support are evaluated only on single-image cases; — means not evaluated. Bold and underlined values follow the manuscript's best and second-best markings within each model group.
| Model | Multi- image | Temporal | Causal | Spatial | Logical | Counter factual | Hybrid | Single | Multi | Total |
|---|---|---|---|---|---|---|---|---|---|---|
| Closed-source Models | ||||||||||
| GPT-Image-2.5 Sunburst | Yes | 67.5 | 69.0 | 64.0 | 27.0 | 72.0 | 39.0 | 56.9 | 52.6 | 56.6 |
| Nano Banana Pro | Yes | 63.0 | 71.0 | 51.5 | 36.0 | 65.0 | 38.0 | 55.4 | 44.9 | 54.6 |
| Nano Banana 2 | Yes | 67.5 | 67.0 | 45.0 | 28.5 | 62.0 | 30.0 | 51.3 | 44.9 | 50.8 |
| GPT-Image-2 | Yes | 59.5 | 64.5 | 60.5 | 15.0 | 58.0 | 21.0 | 48.2 | 43.6 | 47.8 |
| GPT-Image-2.5 Flare | Yes | 61.0 | 63.0 | 56.5 | 11.0 | 57.0 | 20.0 | 46.6 | 38.5 | 46.0 |
| Luma Uni-1 | Yes | 61.0 | 57.0 | 50.5 | 18.0 | 55.0 | 24.0 | 45.7 | 39.7 | 45.2 |
| Qwen-Image-3.0 | Yes | 58.5 | 47.5 | 41.0 | 36.0 | 49.0 | 34.0 | 44.5 | 50.0 | 44.9 |
| Nano Banana 2 Lite | Yes | 60.5 | 61.5 | 46.5 | 15.0 | 62.0 | 12.0 | 44.5 | 39.7 | 44.1 |
| FLUX.2 Max | Yes | 57.0 | 57.0 | 42.5 | 12.5 | 61.0 | 18.0 | 41.5 | 43.6 | 41.7 |
| GPT-Image-1.5 | Yes | 57.5 | 62.5 | 39.5 | 8.0 | 59.0 | 12.0 | 41.2 | 33.3 | 40.6 |
| FLUX.2 Pro | Yes | 53.5 | 48.5 | 36.5 | 10.5 | 61.0 | 18.0 | 37.3 | 42.3 | 37.7 |
| Nano Banana | Yes | 52.0 | 52.5 | 39.5 | 11.0 | 47.0 | 11.0 | 36.8 | 37.2 | 36.8 |
| HunyuanImage-3.5 | Yes | 43.5 | 46.0 | 43.5 | 14.0 | 37.0 | 17.0 | 34.5 | 38.5 | 34.8 |
| Wan2.7-Image-Pro | Yes | 41.5 | 52.0 | 32.5 | 10.5 | 49.0 | 13.0 | 36.3 | 0.0 | 33.5 |
| GPT-Image-1 | Yes | 51.0 | 49.0 | 29.5 | 5.5 | 41.0 | 7.0 | 32.3 | 25.6 | 31.8 |
| Qwen-Image-2.0 | Yes | 44.0 | 47.0 | 36.0 | 6.5 | 36.0 | 0.0 | 30.2 | 32.1 | 30.3 |
| Seedream 4.5 | Yes | 42.5 | 42.0 | 26.5 | 7.5 | 36.0 | 6.0 | 28.0 | 26.9 | 27.9 |
| Seedream 5.0 | Yes | 42.5 | 43.0 | 26.0 | 5.5 | 30.0 | 3.0 | 27.2 | 20.5 | 26.7 |
| Step Image Edit 2 | No | 18.5 | 22.9 | 26.4 | 1.7 | 10.0 | 1.0 | 14.9 | — | — |
| Open-source Models | ||||||||||
| HunyuanImage-3.0-Instruct | Yes | 30.5 | 37.5 | 29.5 | 7.0 | 21.0 | 0.0 | 22.5 | 29.5 | 23.0 |
| Qwen-Image-Edit-2511 | Yes | 22.0 | 24.5 | 23.5 | 4.5 | 17.0 | 5.0 | 16.9 | 19.2 | 17.1 |
| Qwen-Image-2.1 | Yes | 21.5 | 25.5 | 24.5 | 2.0 | 16.0 | 0.0 | 15.0 | 32.1 | 16.3 |
| FireRed-Image-Edit-1.0 | Yes | 22.0 | 18.0 | 28.0 | 2.0 | 21.0 | 2.0 | 16.1 | 19.2 | 16.3 |
| FireRed-Image-Edit-1.1 | Yes | 22.5 | 23.5 | 22.5 | 1.5 | 16.0 | 2.0 | 14.8 | 28.2 | 15.8 |
| Qwen-Image-Edit-2509 | Yes | 21.0 | 24.5 | 20.0 | 0.5 | 16.0 | 0.0 | 15.0 | 12.8 | 14.8 |
| HiDream-O1-Image | Yes | 27.5 | 18.0 | 20.5 | 3.0 | 5.0 | 0.0 | 13.6 | 23.1 | 14.3 |
| FLUX.2 Dev | Yes | 19.5 | 27.5 | 17.0 | 1.0 | 10.0 | 1.0 | 12.8 | 29.5 | 14.1 |
| FLUX.2 Klein 9B | Yes | 16.0 | 16.5 | 22.0 | 0.5 | 17.0 | 2.0 | 11.4 | 30.8 | 12.9 |
| JoyAI-Image-Edit-Plus | Yes | 16.5 | 16.0 | 20.0 | 4.0 | 13.0 | 3.0 | 11.7 | 26.9 | 12.9 |
| SenseNova-U1-A3B-MoT-SFT | Yes | 20.5 | 25.5 | 2.5 | 3.0 | 17.0 | 0.0 | 12.9 | 1.3 | 12.0 |
| UniPic 3.0 | Yes | 18.5 | 17.5 | 18.0 | 0.5 | 9.0 | 1.0 | 11.1 | 21.8 | 11.9 |
| BAGEL (w/ CoT) | Yes | 18.5 | 16.0 | 16.0 | 1.5 | 10.0 | 0.0 | 10.4 | 23.1 | 11.4 |
| SenseNova-U1-8B-MoT-SFT | Yes | 16.5 | 22.5 | 7.5 | 3.5 | 13.0 | 0.0 | 12.3 | 0.0 | 11.3 |
| FLUX.2 Klein 4B | Yes | 15.5 | 14.5 | 17.0 | 2.5 | 8.0 | 1.0 | 8.7 | 35.9 | 10.8 |
| UniWorld-V2 | Yes | 10.0 | 18.0 | 12.0 | 0.5 | 13.0 | 1.0 | 9.9 | 5.1 | 9.5 |
| BAGEL | Yes | 6.5 | 6.0 | 8.0 | 1.0 | 2.0 | 0.0 | 3.1 | 20.5 | 4.5 |
| OmniGen2 | Yes | 7.0 | 6.0 | 5.5 | 0.0 | 8.0 | 0.0 | 3.7 | 14.1 | 4.5 |
| OmniGen | Yes | 4.5 | 4.0 | 5.0 | 1.5 | 1.0 | 0.0 | 2.7 | 7.7 | 3.1 |
| Emu2-Gen | Yes | 4.0 | 1.5 | 0.5 | 0.5 | 1.0 | 0.0 | 1.4 | 1.3 | 1.4 |
| JoyAI-Image-Edit | No | 17.9 | 24.0 | 22.0 | 2.2 | 16.0 | 2.0 | 15.0 | — | — |
| Step1X-Edit (thinking) | No | 19.6 | 27.1 | 14.3 | 0.6 | 16.0 | 2.0 | 14.1 | — | — |
| Boogu-Image-0.1-Edit | No | 15.5 | 21.9 | 17.6 | 6.1 | 11.0 | 3.0 | 13.6 | — | — |
| Step1X-Edit (thinking + reflection) | No | 21.4 | 22.9 | 9.9 | 1.7 | 19.0 | 1.0 | 13.1 | — | — |
| LongCat-Image-Edit | No | 10.1 | 20.8 | 15.4 | 1.1 | 9.0 | 0.0 | 10.4 | — | — |
| Step1X-Edit (base) | No | 13.1 | 18.8 | 10.4 | 1.1 | 14.0 | 1.0 | 10.2 | — | — |
| InternVL-U (w/ CoT) | No | 8.9 | 14.1 | 4.9 | 1.1 | 16.0 | 0.0 | 7.5 | — | — |
| FLUX.1 Kontext Dev | No | 5.4 | 9.4 | 12.1 | 1.7 | 7.0 | 0.0 | 6.4 | — | — |
| LLaDA-Image | No | 7.1 | 4.7 | 3.8 | 0.0 | 4.0 | 0.0 | 3.5 | — | — |
| Ovis-U1 | No | 4.8 | 6.3 | 2.2 | 1.7 | 2.0 | 0.0 | 3.1 | — | — |
| InternVL-U | No | 3.6 | 2.6 | 3.8 | 0.6 | 5.0 | 0.0 | 2.6 | — | — |
| UniPic2-SD3.5M-Kontext-2B | No | 1.2 | 2.1 | 0.5 | 0.0 | 3.0 | 0.0 | 1.1 | — | — |
| Lumina-DiMOO | No | 0.6 | 1.6 | 0.0 | 0.6 | 4.0 | 0.0 | 1.0 | — | — |
| UniWorld-V1 | No | 1.2 | 1.0 | 0.5 | 0.6 | 1.0 | 0.0 | 0.8 | — | — |
| Agentic Methods | ||||||||||
| RISE-Agent (Ours) | Yes | 58.5 | 58.0 | 37.5 | 39.0 | 66.0 | 28.0 | 48.3 | 44.9 | 48.0 |
| MindBrush | No | 35.1 | 33.3 | 16.5 | 8.9 | 32.0 | 0.0 | 21.8 | — | — |
| InterleaveThinker | No | 38.7 | 30.7 | 15.9 | 1.7 | 26.0 | 8.0 | 20.6 | — | — |
| IntentEdit | No | 23.2 | 24.0 | 11.5 | 1.1 | 14.0 | 0.0 | 13.2 | — | — |
| ImAgent | No | 5.4 | 4.7 | 4.4 | 0.6 | 5.0 | 0.0 | 3.5 | — | — |
Scores are reported on a 100-point scale. Models without multi-image support use the single-image subset. Visual plausibility excludes logical tasks. These are dimension scores, not task accuracy.
| Model | Multi-image | 🧠 Instruction Reasoning | 🪞 Appearance Consistency | 👁️ Visual Plausibility |
|---|---|---|---|---|
| Closed-source Models | ||||
| GPT-Image-2.5 Sunburst | Yes | 68.1 | 96.3 | 98.0 |
| Nano Banana Pro | Yes | 69.8 | 87.7 | 96.8 |
| Nano Banana 2 | Yes | 68.5 | 91.1 | 97.2 |
| GPT-Image-2 | Yes | 61.5 | 95.2 | 97.6 |
| GPT-Image-2.5 Flare | Yes | 59.8 | 93.2 | 98.3 |
| Luma Uni-1 | Yes | 62.2 | 91.0 | 94.7 |
| Qwen-Image-3.0 | Yes | 68.4 | 87.3 | 95.9 |
| Nano Banana 2 Lite | Yes | 60.8 | 89.7 | 96.3 |
| FLUX.2 Max | Yes | 61.5 | 79.2 | 95.5 |
| GPT-Image-1.5 | Yes | 59.9 | 81.0 | 96.5 |
| FLUX.2 Pro | Yes | 57.5 | 74.7 | 94.9 |
| Nano Banana | Yes | 52.9 | 84.1 | 95.8 |
| HunyuanImage-3.5 | Yes | 51.3 | 93.4 | 95.2 |
| Wan2.7-Image-Pro | Yes | 48.8 | 80.1 | 93.3 |
| GPT-Image-1 | Yes | 53.4 | 73.6 | 95.6 |
| Qwen-Image-2.0 | Yes | 49.6 | 79.7 | 91.5 |
| Seedream 4.5 | Yes | 49.8 | 76.0 | 94.5 |
| Seedream 5.0 | Yes | 47.8 | 71.7 | 94.3 |
| Step Image Edit 2 | No | 31.0 | 68.3 | 86.3 |
| Open-source Models | ||||
| HunyuanImage-3.0-Instruct | Yes | 38.2 | 75.8 | 88.5 |
| Qwen-Image-Edit-2511 | Yes | 36.0 | 64.0 | 90.7 |
| Qwen-Image-2.1 | Yes | 31.6 | 77.6 | 86.2 |
| FireRed-Image-Edit-1.0 | Yes | 32.2 | 64.0 | 93.7 |
| FireRed-Image-Edit-1.1 | Yes | 32.8 | 66.2 | 93.0 |
| Qwen-Image-Edit-2509 | Yes | 29.8 | 61.3 | 85.7 |
| HiDream-O1-Image | Yes | 35.2 | 57.3 | 87.3 |
| FLUX.2 Dev | Yes | 31.1 | 57.7 | 91.9 |
| FLUX.2 Klein 9B | Yes | 28.9 | 67.2 | 92.3 |
| JoyAI-Image-Edit-Plus | Yes | 26.7 | 76.9 | 92.1 |
| SenseNova-U1-A3B-MoT-SFT | Yes | 38.0 | 53.0 | 58.7 |
| UniPic 3.0 | Yes | 28.2 | 60.7 | 83.4 |
| BAGEL (w/ CoT) | Yes | 33.3 | 61.2 | 75.0 |
| SenseNova-U1-8B-MoT-SFT | Yes | 36.0 | 58.1 | 60.0 |
| FLUX.2 Klein 4B | Yes | 25.6 | 71.0 | 90.3 |
| UniWorld-V2 | Yes | 30.7 | 51.6 | 84.7 |
| BAGEL | Yes | 23.1 | 47.7 | 76.1 |
| OmniGen2 | Yes | 15.4 | 58.4 | 80.7 |
| OmniGen | Yes | 13.6 | 44.5 | 61.2 |
| Emu2-Gen | Yes | 10.7 | 25.5 | 62.5 |
| JoyAI-Image-Edit | No | 31.6 | 71.9 | 90.5 |
| Step1X-Edit (thinking) | No | 28.5 | 70.3 | 89.9 |
| Boogu-Image-0.1-Edit | No | 27.5 | 79.8 | 91.0 |
| Step1X-Edit (thinking + reflection) | No | 30.2 | 68.0 | 89.4 |
| LongCat-Image-Edit | No | 25.7 | 62.0 | 83.9 |
| Step1X-Edit (base) | No | 26.0 | 66.1 | 87.1 |
| InternVL-U (w/ CoT) | No | 27.8 | 52.1 | 71.1 |
| FLUX.1 Kontext Dev | No | 18.6 | 72.8 | 91.0 |
| LLaDA-Image | No | 21.7 | 26.7 | 78.0 |
| Ovis-U1 | No | 20.5 | 32.1 | 71.2 |
| InternVL-U | No | 21.0 | 39.7 | 63.4 |
| UniPic2-SD3.5M-Kontext-2B | No | 16.2 | 34.4 | 54.2 |
| Lumina-DiMOO | No | 11.2 | 46.8 | 63.9 |
| UniWorld-V1 | No | 15.8 | 23.3 | 58.3 |
| Agentic Methods | ||||
| RISE-Agent (Ours) | Yes | 67.6 | 90.6 | 92.9 |
| MindBrush | No | 49.2 | 56.4 | 81.1 |
| InterleaveThinker | No | 40.0 | 66.4 | 87.3 |
| IntentEdit | No | 41.0 | 58.4 | 78.7 |
| ImAgent | No | 24.6 | 29.2 | 65.4 |
Data: This repository includes only a demo subset; please download the full RISEBench++ dataset from our Hugging Face collection for full-benchmark evaluation.
The bundled data/overall_data.json contains 60 sample cases, with 10 cases per reasoning dimension, numbered 1–10 within each dimension. This subset is provided for trying the generation and evaluation workflow; the tables above report full-benchmark results.
Run the following commands from the repository root to install the dependencies:
cd 'RISEBench++'
python -m pip install "numpy<2" pandas Pillow openai tqdm xlsxwriter openpyxlAll data/... and outputs/... paths below are relative to RISEBench++/.
Input images for the six categories are located in the data directory. Each sample contains an instruction and an image list. Use all listed input images, in order, to generate the edited output. For hybrid cases, execute the instructions sequentially and save the final output image. Reference answers are used only for evaluation.
The sample data include English and Chinese instructions. The evaluator reads instruction and reference_txt; for Chinese evaluation, prepare a separate manifest using their Chinese counterparts.
Output File Structure:
Generated outputs should be saved in the following directory structure:
outputs/{MODEL_NAME}/images/{CATEGORY}/{INDEX_NAME}.{FORMAT}
{MODEL_NAME}: The name of the model being evaluated.{CATEGORY}: The sample's reasoning category; usehybrid_reasoningfor multi-turn cases.{INDEX_NAME}: The sample'sindexinoverall_data.json.{FORMAT}:png,jpg, orjpeg.
For example:
outputs/MODEL_NAME/images/temporal_reasoning/temporal_reasoning_1.png
For hybrid cases:
outputs/MODEL_NAME/images/hybrid_reasoning/hybrid_reasoning_1.png
From RISEBench++/, install the agent dependencies with python -m pip install -r rise-agent/requirements.txt. Before running, edit rise-agent/agent.cfg: set api_base to your OpenAI-compatible API URL, planner_model and verifier_model to model IDs available at that endpoint, and model_path to your local FLUX.2-klein-9B directory. The default data, input_dir, and output_dir paths can be changed there as needed. Set API_KEY (for the planner/verifier API) and TAVILY_API_KEY (for web search) in rise-agent/run_agent.sh, replacing the placeholder values.
Then run:
bash rise-agent/run_agent.shTo try one sample first, use bash rise-agent/run_agent.sh --limit 1. Outputs are written to outputs/rise-agent/images/ by default.
Once the outputs are generated and saved in the specified format, evaluate them using gemini_eval.py. The default judges are google/gemini-3-flash-preview for instruction reasoning and appearance consistency, and google/gemini-3.1-flash-lite for visual plausibility.
Set the API key and base URL for an OpenAI-compatible endpoint that serves these Gemini models:
export OPENAI_API_KEY="YOUR_API_KEY"
export OPENAI_BASE_URL="https://YOUR_API_HOST/v1"Replace both placeholders with your provider's credentials and endpoint. If the provider uses different model identifiers, update the two judge model strings in gemini_eval.py.
python gemini_eval.py \
--data data/overall_data.json \
--input data \
--output outputs/MODEL_NAME \
--model MODEL_NAME \
--nproc 10--model names the model being evaluated. --nproc controls concurrency. Use separate output directories or --prefix values for different languages and subsets to keep cached results separate.
Three result files will be generated in outputs/MODEL_NAME/:
MODEL_NAME_judge.csv: Overall and fine-grained evaluation scores and accuracy. Accuracy is stored as a fraction; multiply by 100 to report a percentage.MODEL_NAME_judge.xlsx: Sample-level judge responses and scores.MODEL_NAME.pkl: Cached judge responses for resuming evaluations.
Rerunning the same command reuses cached sample IDs. Missing output images receive the minimum score and are also cached; use a new prefix or remove the corresponding cache before evaluating newly supplied outputs.
We exhibit representative model outputs below, including examples illustrating reasoning accuracy and consistency across editing turns. Download the model outputs from Hugging Face.
If you find RISEBench++ or RISE-Agent useful, please cite our paper:
@misc{yang2026reasoninginformedvisualediting,
title={Reasoning-Informed Visual Editing},
author={Xue Yang and Peiyuan Zhang and Yilun Zhu and Qihao Yang and Mingxin Liu and Xiangyu Zhao and Ziqian Fan and Zhaokai Wang and Yan Li and Yifan Yang and Xu Yang and Xiaosong Jia and Yue Zhou and Zhihang Zhong and Junchi Yan},
year={2026},
eprint={2610.12343},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.12343}
}



