Skip to content

Latest commit

 

History

History
320 lines (250 loc) · 88.6 KB

File metadata and controls

320 lines (250 loc) · 88.6 KB

← Overview · RISEBench

Reasoning-Informed Visual Editing

RISEBench++

Xue Yang†, Peiyuan Zhang*, Yilun Zhu*, Qihao Yang*, Mingxin Liu, Xiangyu Zhao, Ziqian Fan, Zhaokai Wang, Yan Li, Yifan Yang, Xu Yang, Xiaosong Jia, Yue Zhou, Zhihang Zhong, Junchi Yan

* Core student contributors · † Corresponding author

Shanghai Jiao Tong University · Southeast University · South China University of Technology · Fudan University · East China Normal University

arXiv Paper Hugging Face Data Hugging Face Model Outputs Hugging Face Paper Leaderboard

If you find our work helpful, please consider giving us a ⭐ or citation 😊

Overview and example tasks of RISEBench++

📖 Introduction

RISEBench++ task taxonomy and data distribution

In this work, we introduce RISEBench++, an extension of RISEBench for evaluating Reasoning-Informed viSual Editing (RISE). RISEBench++ focuses on six key reasoning types: Temporal, Causal, Spatial, Logical, Counterfactual, and Hybrid Reasoning.

RISEBench++ comprises 1,000 human-annotated test cases, provided in English and Chinese, with a hierarchical taxonomy spanning six reasoning dimensions, 12 subcategories, and 65 fine-grained task types. It supports single-image, multi-image, and multi-turn editing. Temporal, causal, spatial, and logical reasoning each contain 200 cases; counterfactual and hybrid reasoning each contain 100 cases. Hybrid cases contain three to five successive editing instructions.

To comprehensively assess model performance across diverse task types, we define three key evaluation dimensions: Instruction Reasoning, Appearance Consistency, and Visual Plausibility.

Besides, we design a robust LMM-as-a-Judge evaluation pipeline with dimension-specific reference evidence and scoring rubrics. Gemini-3-Flash evaluates instruction reasoning and appearance consistency, while Gemini-3.1-Flash-Lite evaluates visual plausibility. Logical tasks are not evaluated for visual plausibility; counterfactual tasks are assessed for perceptual quality without penalizing intentionally impossible physical outcomes.

RISEBench++ aims to support comprehensive, reliable, and fine-grained evaluation of reasoning-aware visual editing across diverse models and input settings.

RISEBench++ evaluation metrics and judge pipeline

We also introduce RISE-Agent, a training-free framework that combines reasoning-driven planning, tool-augmented execution, and verifier-guided refinement. The planner infers the intended target state using visual reasoning, external knowledge, and computational tools; the executor realizes the edit through a generative model or precise programmatic operations; and the verifier checks the result and guides refinement or replanning. Together, they form a closed loop for reasoning-informed visual editing.

Overview of RISE-Agent: reasoning-driven planning, tool-augmented execution, and verifier-guided refinement

🔥 Benchmark Performance

We evaluate 58 visual editing approaches, including 19 closed-source models, 34 open-source models, and 5 agentic methods. The following tables report results from the accompanying manuscript. Accuracy measures the percentage of cases that receive full marks on every applicable evaluation dimension.

📊 Overall performance on RISEBench++ (English)

All accuracy values are percentages. Single, Multi, and Total refer to single-image cases, multi-image cases, and the full benchmark, respectively. Models without multi-image support are evaluated only on single-image cases; — means not evaluated. Bold and underlined values follow the manuscript's best and second-best markings within each model group.

ModelMulti-
image
TemporalCausalSpatialLogicalCounter
factual
HybridSingleMultiTotal
Closed-source Models
GPT-Image-2.5 SunburstYes67.569.064.027.072.039.056.952.656.6
Nano Banana ProYes63.071.051.536.065.038.055.444.954.6
Nano Banana 2Yes67.567.045.028.562.030.051.344.950.8
GPT-Image-2Yes59.564.560.515.058.021.048.243.647.8
GPT-Image-2.5 FlareYes61.063.056.511.057.020.046.638.546.0
Luma Uni-1Yes61.057.050.518.055.024.045.739.745.2
Qwen-Image-3.0Yes58.547.541.036.049.034.044.550.044.9
Nano Banana 2 LiteYes60.561.546.515.062.012.044.539.744.1
FLUX.2 MaxYes57.057.042.512.561.018.041.543.641.7
GPT-Image-1.5Yes57.562.539.58.059.012.041.233.340.6
FLUX.2 ProYes53.548.536.510.561.018.037.342.337.7
Nano BananaYes52.052.539.511.047.011.036.837.236.8
HunyuanImage-3.5Yes43.546.043.514.037.017.034.538.534.8
Wan2.7-Image-ProYes41.552.032.510.549.013.036.30.033.5
GPT-Image-1Yes51.049.029.55.541.07.032.325.631.8
Qwen-Image-2.0Yes44.047.036.06.536.00.030.232.130.3
Seedream 4.5Yes42.542.026.57.536.06.028.026.927.9
Seedream 5.0Yes42.543.026.05.530.03.027.220.526.7
Step Image Edit 2No18.522.926.41.710.01.014.9——
Open-source Models
HunyuanImage-3.0-InstructYes30.537.529.57.021.00.022.529.523.0
Qwen-Image-Edit-2511Yes22.024.523.54.517.05.016.919.217.1
Qwen-Image-2.1Yes21.525.524.52.016.00.015.032.116.3
FireRed-Image-Edit-1.0Yes22.018.028.02.021.02.016.119.216.3
FireRed-Image-Edit-1.1Yes22.523.522.51.516.02.014.828.215.8
Qwen-Image-Edit-2509Yes21.024.520.00.516.00.015.012.814.8
HiDream-O1-ImageYes27.518.020.53.05.00.013.623.114.3
FLUX.2 DevYes19.527.517.01.010.01.012.829.514.1
FLUX.2 Klein 9BYes16.016.522.00.517.02.011.430.812.9
JoyAI-Image-Edit-PlusYes16.516.020.04.013.03.011.726.912.9
SenseNova-U1-A3B-MoT-SFTYes20.525.52.53.017.00.012.91.312.0
UniPic 3.0Yes18.517.518.00.59.01.011.121.811.9
BAGEL (w/ CoT)Yes18.516.016.01.510.00.010.423.111.4
SenseNova-U1-8B-MoT-SFTYes16.522.57.53.513.00.012.30.011.3
FLUX.2 Klein 4BYes15.514.517.02.58.01.08.735.910.8
UniWorld-V2Yes10.018.012.00.513.01.09.95.19.5
BAGELYes6.56.08.01.02.00.03.120.54.5
OmniGen2Yes7.06.05.50.08.00.03.714.14.5
OmniGenYes4.54.05.01.51.00.02.77.73.1
Emu2-GenYes4.01.50.50.51.00.01.41.31.4
JoyAI-Image-EditNo17.924.022.02.216.02.015.0——
Step1X-Edit (thinking)No19.627.114.30.616.02.014.1——
Boogu-Image-0.1-EditNo15.521.917.66.111.03.013.6——
Step1X-Edit (thinking + reflection)No21.422.99.91.719.01.013.1——
LongCat-Image-EditNo10.120.815.41.19.00.010.4——
Step1X-Edit (base)No13.118.810.41.114.01.010.2——
InternVL-U (w/ CoT)No8.914.14.91.116.00.07.5——
FLUX.1 Kontext DevNo5.49.412.11.77.00.06.4——
LLaDA-ImageNo7.14.73.80.04.00.03.5——
Ovis-U1No4.86.32.21.72.00.03.1——
InternVL-UNo3.62.63.80.65.00.02.6——
UniPic2-SD3.5M-Kontext-2BNo1.22.10.50.03.00.01.1——
Lumina-DiMOONo0.61.60.00.64.00.01.0——
UniWorld-V1No1.21.00.50.61.00.00.8——
Agentic Methods
RISE-Agent (Ours)Yes58.558.037.539.066.028.048.344.948.0
MindBrushNo35.133.316.58.932.00.021.8——
InterleaveThinkerNo38.730.715.91.726.08.020.6——
IntentEditNo23.224.011.51.114.00.013.2——
ImAgentNo5.44.74.40.65.00.03.5——

🎨 Comparison across models on three evaluation sub-dimensions

Scores are reported on a 100-point scale. Models without multi-image support use the single-image subset. Visual plausibility excludes logical tasks. These are dimension scores, not task accuracy.

ModelMulti-image🧠 Instruction Reasoning🪞 Appearance Consistency👁️ Visual Plausibility
Closed-source Models
GPT-Image-2.5 SunburstYes68.196.398.0
Nano Banana ProYes69.887.796.8
Nano Banana 2Yes68.591.197.2
GPT-Image-2Yes61.595.297.6
GPT-Image-2.5 FlareYes59.893.298.3
Luma Uni-1Yes62.291.094.7
Qwen-Image-3.0Yes68.487.395.9
Nano Banana 2 LiteYes60.889.796.3
FLUX.2 MaxYes61.579.295.5
GPT-Image-1.5Yes59.981.096.5
FLUX.2 ProYes57.574.794.9
Nano BananaYes52.984.195.8
HunyuanImage-3.5Yes51.393.495.2
Wan2.7-Image-ProYes48.880.193.3
GPT-Image-1Yes53.473.695.6
Qwen-Image-2.0Yes49.679.791.5
Seedream 4.5Yes49.876.094.5
Seedream 5.0Yes47.871.794.3
Step Image Edit 2No31.068.386.3
Open-source Models
HunyuanImage-3.0-InstructYes38.275.888.5
Qwen-Image-Edit-2511Yes36.064.090.7
Qwen-Image-2.1Yes31.677.686.2
FireRed-Image-Edit-1.0Yes32.264.093.7
FireRed-Image-Edit-1.1Yes32.866.293.0
Qwen-Image-Edit-2509Yes29.861.385.7
HiDream-O1-ImageYes35.257.387.3
FLUX.2 DevYes31.157.791.9
FLUX.2 Klein 9BYes28.967.292.3
JoyAI-Image-Edit-PlusYes26.776.992.1
SenseNova-U1-A3B-MoT-SFTYes38.053.058.7
UniPic 3.0Yes28.260.783.4
BAGEL (w/ CoT)Yes33.361.275.0
SenseNova-U1-8B-MoT-SFTYes36.058.160.0
FLUX.2 Klein 4BYes25.671.090.3
UniWorld-V2Yes30.751.684.7
BAGELYes23.147.776.1
OmniGen2Yes15.458.480.7
OmniGenYes13.644.561.2
Emu2-GenYes10.725.562.5
JoyAI-Image-EditNo31.671.990.5
Step1X-Edit (thinking)No28.570.389.9
Boogu-Image-0.1-EditNo27.579.891.0
Step1X-Edit (thinking + reflection)No30.268.089.4
LongCat-Image-EditNo25.762.083.9
Step1X-Edit (base)No26.066.187.1
InternVL-U (w/ CoT)No27.852.171.1
FLUX.1 Kontext DevNo18.672.891.0
LLaDA-ImageNo21.726.778.0
Ovis-U1No20.532.171.2
InternVL-UNo21.039.763.4
UniPic2-SD3.5M-Kontext-2BNo16.234.454.2
Lumina-DiMOONo11.246.863.9
UniWorld-V1No15.823.358.3
Agentic Methods
RISE-Agent (Ours)Yes67.690.692.9
MindBrushNo49.256.481.1
InterleaveThinkerNo40.066.487.3
IntentEditNo41.058.478.7
ImAgentNo24.629.265.4

Data: This repository includes only a demo subset; please download the full RISEBench++ dataset from our Hugging Face collection for full-benchmark evaluation.

🛠️ Quick Start

The bundled data/overall_data.json contains 60 sample cases, with 10 cases per reasoning dimension, numbered 1–10 within each dimension. This subset is provided for trying the generation and evaluation workflow; the tables above report full-benchmark results.

Run the following commands from the repository root to install the dependencies:

cd 'RISEBench++'
python -m pip install "numpy<2" pandas Pillow openai tqdm xlsxwriter openpyxl

All data/... and outputs/... paths below are relative to RISEBench++/.

1. Output Generation

Input images for the six categories are located in the data directory. Each sample contains an instruction and an image list. Use all listed input images, in order, to generate the edited output. For hybrid cases, execute the instructions sequentially and save the final output image. Reference answers are used only for evaluation.

The sample data include English and Chinese instructions. The evaluator reads instruction and reference_txt; for Chinese evaluation, prepare a separate manifest using their Chinese counterparts.

Output File Structure:

Generated outputs should be saved in the following directory structure:

outputs/{MODEL_NAME}/images/{CATEGORY}/{INDEX_NAME}.{FORMAT}

  • {MODEL_NAME}: The name of the model being evaluated.
  • {CATEGORY}: The sample's reasoning category; use hybrid_reasoning for multi-turn cases.
  • {INDEX_NAME}: The sample's index in overall_data.json.
  • {FORMAT}: png, jpg, or jpeg.

For example: outputs/MODEL_NAME/images/temporal_reasoning/temporal_reasoning_1.png

For hybrid cases: outputs/MODEL_NAME/images/hybrid_reasoning/hybrid_reasoning_1.png

Run RISE-Agent

From RISEBench++/, install the agent dependencies with python -m pip install -r rise-agent/requirements.txt. Before running, edit rise-agent/agent.cfg: set api_base to your OpenAI-compatible API URL, planner_model and verifier_model to model IDs available at that endpoint, and model_path to your local FLUX.2-klein-9B directory. The default data, input_dir, and output_dir paths can be changed there as needed. Set API_KEY (for the planner/verifier API) and TAVILY_API_KEY (for web search) in rise-agent/run_agent.sh, replacing the placeholder values.

Then run:

bash rise-agent/run_agent.sh

To try one sample first, use bash rise-agent/run_agent.sh --limit 1. Outputs are written to outputs/rise-agent/images/ by default.

2. Evaluation By LMM Judges

Once the outputs are generated and saved in the specified format, evaluate them using gemini_eval.py. The default judges are google/gemini-3-flash-preview for instruction reasoning and appearance consistency, and google/gemini-3.1-flash-lite for visual plausibility.

Step 1: Configure API Settings

Set the API key and base URL for an OpenAI-compatible endpoint that serves these Gemini models:

export OPENAI_API_KEY="YOUR_API_KEY"
export OPENAI_BASE_URL="https://YOUR_API_HOST/v1"

Replace both placeholders with your provider's credentials and endpoint. If the provider uses different model identifiers, update the two judge model strings in gemini_eval.py.

Step 2: Run the Evaluation Script

python gemini_eval.py \
  --data data/overall_data.json \
  --input data \
  --output outputs/MODEL_NAME \
  --model MODEL_NAME \
  --nproc 10

--model names the model being evaluated. --nproc controls concurrency. Use separate output directories or --prefix values for different languages and subsets to keep cached results separate.

Step 3: Review the Results

Three result files will be generated in outputs/MODEL_NAME/:

  1. MODEL_NAME_judge.csv: Overall and fine-grained evaluation scores and accuracy. Accuracy is stored as a fraction; multiply by 100 to report a percentage.
  2. MODEL_NAME_judge.xlsx: Sample-level judge responses and scores.
  3. MODEL_NAME.pkl: Cached judge responses for resuming evaluations.

Rerunning the same command reuses cached sample IDs. Missing output images receive the minimum score and are also cached; use a new prefix or remove the corresponding cache before evaluating newly supplied outputs.

🔥 Outputs of Current Models

We exhibit representative model outputs below, including examples illustrating reasoning accuracy and consistency across editing turns. Download the model outputs from Hugging Face.

Representative visual editing outputs on RISEBench++

Citation

If you find RISEBench++ or RISE-Agent useful, please cite our paper:

@misc{yang2026reasoninginformedvisualediting,
  title={Reasoning-Informed Visual Editing},
  author={Xue Yang and Peiyuan Zhang and Yilun Zhu and Qihao Yang and Mingxin Liu and Xiangyu Zhao and Ziqian Fan and Zhaokai Wang and Yan Li and Yifan Yang and Xu Yang and Xiaosong Jia and Yue Zhou and Zhihang Zhong and Junchi Yan},
  year={2026},
  eprint={2610.12343},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.12343}
}