Image-based modeling
Construct engineering artifacts from visual references and drawings.
Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages.
We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks.
We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.
Six task types, spanning the complete engineering design loop.
Construct engineering artifacts from visual references and drawings.
Choose suitable applications from a permitted set of engineering tools.
Complete a specified engineering objective within one designated application.
Meet feasibility constraints and optimize a continuous design-quality objective.
Transfer intermediate artifacts and preserve dependencies across applications.
Choose tools and execution strategies without a predefined software workflow.
Experts define representative engineering objectives, initial workspaces, deliverables, and acceptance criteria. AI assists task specification and verifier implementation. The final task packages are executed and reviewed in their designated software environments.
Agents work through GUI or CLI interfaces. Domain-specific verifiers reopen their deliverables and check the engineering properties that matter.
Start from a clean, task-specific workspace. Use visual observations and mouse or keyboard actions, or terminal commands and scripts.
Submit native models, simulation outputs, circuits, building models, or 3D scenes, including intermediate files when the workflow requires them.
Recompute geometry, topology, physical quantities, rule compliance, and consistency across workflow stages against task-specific criteria.
EngiScore = 100 × mean task score. Binary tasks receive 1 only when every required criterion passes. Quantitative design tasks receive a continuous score between 0 and 1 after passing feasibility checks; infeasible outputs receive 0.
Seven frontier models evaluated on the same stratified subset of 300 tasks: 152 CLI and 148 GUI tasks across 73 software–category strata.
Highest overall EngiScore
Claude Opus 5 leads the main evaluation.
Multi-software success
6 of 168 attempts succeed across all seven models.
Tasks with zero scores for every model
Out of the 300 tasks in the main evaluation.
| Model | ||||||
|---|---|---|---|---|---|---|
| 01 | 44.3 | 47.9 | 40.5 | 70.2 | 2.4 | 19.02 |
| 02 | 38.0 | 47.1 | 28.6 | 54.9 | 1.0 | 9.73 |
| 03 | 25.8 | 34.7 | 16.6 | 108.1 | 2.6 | 6.52 |
| 04 | 25.5 | 48.3 | 2.0 | 116.2 | 9.9 | 0.81 |
| 05 | 25.3 | 36.1 | 14.2 | 110.8 | 0.7 | 2.11 |
| 06 | 24.1 | 31.3 | 16.7 | 92.5 | 0.9 | 5.90 |
| 07 | 16.1 | 30.8 | 1.0 | 120.5 | 1.3 | 0.45 |
EngiScore: higher is better. Steps and API cost: per-task means over all 300 tasks, including unsuccessful runs. Tokens (K): output tokens per decision step, averaged over the 300 tasks and reported in thousands. Lower Steps, Tokens, and Cost are preferred. Click a column heading to sort.
CLI and GUI scores are measured on different task subsets, not a paired comparison of the same tasks. Scores are reported as in the supplied manuscript.
Controlled ablations vary visual inputs, accessibility information, screenshot resolution, and interaction history.
For GPT-5.6 Sol, GUI EngiScore rises from 15.0 to 31.7 when the reference drawing is shown directly in the task message. Improvements are not uniform across individual tasks.
Accessibility trees improve Kimi, change GPT little, and lower Gemini’s score. API cost increases by 50.7–147.9% even though all three models use fewer decision rounds.
Gemini benefits from a 15-turn window; GPT largely plateaus after 10 turns. Kimi performs best at 10, with a longer window reducing its score and increasing cost.
| Model | GUI | CLI | ||
|---|---|---|---|---|
| Env. | Message | Env. | Message | |
| GPT-5.6 Sol | 15.0 | 31.7 | 63.3 | 68.3 |
| Gemini 3.7 Flash | 3.3 | 10.0 | 38.3 | 40.0 |
| Kimi K3 | 6.7 | 10.0 | 40.0 | 36.7 |
Env.: reference drawing available in the environment. Message: reference drawing included in the initial message.
| Model | With initial image | Without initial image | ||
|---|---|---|---|---|
| Off | On | Off | On | |
| GPT-5.6 Sol | 66.7 | 68.3 | 48.1 | 41.5 |
| Gemini 3.7 Flash | 36.7 | 40.0 | 33.3 | 36.6 |
| Kimi K3 | 41.7 | 36.7 | 31.5 | 26.5 |
Off / On: whether the agent can inspect images through readimg during execution.
Four single-software GUI workflows from the paper. Each passes its task-specific artifact checks and receives a score of 1.0.
Eight distinct data nets establish the required MCU–SRAM mapping. The saved schematic retains the existing circuitry and passes with zero electrical-rule errors.
Two rectangular arrays create 23 tube holes. The exported DXF preserves the required coordinates, radii, entity types, and layer assignments.
The sliding block is positioned against the base. The exported STEP retains both independent solids and their original dimensions.
A 512 × 512 ambient-occlusion image is baked and packed into the saved native scene, preserving the computed pixels in the deliverable.
The recorded failures span import-scale errors, incorrect action targeting, and mismatches between submitted representations and verifier expectations. In the load case, a pressure representation does not satisfy the verifier’s force/displacement check; in the CSV case, an unrecognized diameter-field name stops evaluation before downstream geometry checks.
All four illustrated runs receive zero credit. The examples distinguish execution errors from representation and schema mismatches; they do not establish that every zero-score artifact is physically invalid.
@misc{engiworld2026,
title = {EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?},
author = {Hongcheng Gao and Hailong Qu and Yu Lei and Henghui Sun and Haoyang Li and Yipeng Wei and Naihao Xue and Xiaohan Yu and Zhuo Tao and Yihe Zang and Yajiao Wang and Jingyi Tang and Yi Li and Jingjing Zhou and Jie Luo and Bohan Zeng and Chengyu Shen and Hao Jiang and Chong Chen and Bowen Qu and Olive Huang and Zeqiang Wang},
year = {2026},
eprint = {2609.37686},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2609.37686}
}