EngiWorldWhat Can Frontier Agents Deliver in
Professional Engineering Environments?

1Tsinghua University 2Zhiman Inc. 3Chongqing University 4University of the Chinese Academy of Sciences 5Shandong University 6Beijing University of Posts and Telecommunications 7Fudan University 8Henan Polytechnic University 9Xi'an Jiaotong University 10Peking University 11Zhejiang University

* Equal contribution  ·  † Corresponding author

1,301Expert-curated tasks
26Software platforms
6Engineering domains
6Task types
7Evaluated models
From one task to a world of engineering. 88 successful GUI tasks, edited from real experiment recordings and screenshots.

Overview

Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages.

We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks.

We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.

EngiWorld at a glance. Representative multi-software and software-selection tasks, with software coverage across CAD, CAE, EDA, BIM, CAM, and 3D visualization. View PDF ↗

The benchmark

Six task types, spanning the complete engineering design loop.

Intent grounding120

Image-based modeling

Construct engineering artifacts from visual references and drawings.

Tool selection140

Software-selection

Choose suitable applications from a permitted set of engineering tools.

Precise construction931

Single-software

Complete a specified engineering objective within one designated application.

Specification attainment40

Quantitative design

Meet feasibility constraints and optimize a continuous design-quality objective.

Cross-toolchain delivery60

Multi-software

Transfer intermediate artifacts and preserve dependencies across applications.

Environment bootstrapping10

Open-ended

Choose tools and execution strategies without a predefined software workflow.

691 CLI tasks610 GUI tasksWindows + Ubuntu environments
Software coverage. 1,723 task–software associations, including every participating application and permitted candidate tool. View PDF ↗
Dataset composition. Six disjoint task types totaling 1,301 tasks, divided into 691 CLI and 610 GUI tasks. View PDF ↗
How the dataset is constructed

Experts define representative engineering objectives, initial workspaces, deliverables, and acceptance criteria. AI assists task specification and verifier implementation. The final task packages are executed and reviewed in their designated software environments.

Expert-led, AI-assisted construction. Task specification, task-instance preparation, and evaluation-criteria construction. View PDF ↗

Evaluation

Agents work through GUI or CLI interfaces. Domain-specific verifiers reopen their deliverables and check the engineering properties that matter.

Unified interaction and verification. GUI and CLI agents share the same artifact-based evaluation protocol. View PDF ↗

Interact

Start from a clean, task-specific workspace. Use visual observations and mouse or keyboard actions, or terminal commands and scripts.

Deliver

Submit native models, simulation outputs, circuits, building models, or 3D scenes, including intermediate files when the workflow requires them.

Verify

Recompute geometry, topology, physical quantities, rule compliance, and consistency across workflow stages against task-specific criteria.

EngiScore = 100 × mean task score. Binary tasks receive 1 only when every required criterion passes. Quantitative design tasks receive a continuous score between 0 and 1 after passing feasibility checks; infeasible outputs receive 0.

Results

Seven frontier models evaluated on the same stratified subset of 300 tasks: 152 CLI and 148 GUI tasks across 73 software–category strata.

44.3

Highest overall EngiScore
Claude Opus 5 leads the main evaluation.

3.6%

Multi-software success
6 of 168 attempts succeed across all seven models.

128

Tasks with zero scores for every model
Out of the 300 tasks in the main evaluation.

Download results
EngiWorld main evaluation. Click column headings to sort.
Model
01Claude Opus 544.347.940.570.22.419.02
02GPT-5.6 Sol38.047.128.654.91.09.73
03Qwen3.8 Max25.834.716.6108.12.66.52
04DeepSeek V4.1 Flash25.548.32.0116.29.90.81
05Gemini 3.7 Flash25.336.114.2110.80.72.11
06Kimi K324.131.316.792.50.95.90
07Qwen3.8 Flash16.130.81.0120.51.30.45

EngiScore: higher is better. Steps and API cost: per-task means over all 300 tasks, including unsuccessful runs. Tokens (K): output tokens per decision step, averaged over the 300 tasks and reported in thousands. Lower Steps, Tokens, and Cost are preferred. Click a column heading to sort.

CLI and GUI scores are measured on different task subsets, not a paired comparison of the same tasks. Scores are reported as in the supplied manuscript.

Performance by domain. Tasks spanning multiple domains contribute to each relevant domain. View PDF ↗
Failure analysis. Declared completion with zero credit and decision-turn exhaustion account for 94.1% of analyzed main-evaluation failures. Missing-field and missing-file failures are excluded. View PDF ↗

Analysis

Controlled ablations vary visual inputs, accessibility information, screenshot resolution, and interaction history.

Observation and history ablations. Accessibility and resolution use GUI tasks; history includes both interfaces. Resolution scales are relative to 1920 × 1080. View PDF ↗

Reference images in the message

For GPT-5.6 Sol, GUI EngiScore rises from 15.0 to 31.7 when the reference drawing is shown directly in the task message. Improvements are not uniform across individual tasks.

Richer observations have tradeoffs

Accessibility trees improve Kimi, change GPT little, and lower Gemini’s score. API cost increases by 50.7–147.9% even though all three models use fewer decision rounds.

History length is model dependent

Gemini benefits from a 15-turn window; GPT largely plateaus after 10 turns. Kimi performs best at 10, with a longer window reducing its score and increasing cost.

Initial-image presentation

EngiScore by initial-image presentation
ModelGUICLI
Env.MessageEnv.Message
GPT-5.6 Sol15.031.763.368.3
Gemini 3.7 Flash3.310.038.340.0
Kimi K36.710.040.036.7

Env.: reference drawing available in the environment. Message: reference drawing included in the initial message.

Runtime image access in CLI

EngiScore with runtime image access disabled or enabled
ModelWith initial imageWithout initial image
OffOnOffOn
GPT-5.6 Sol66.768.348.141.5
Gemini 3.7 Flash36.740.033.336.6
Kimi K341.736.731.526.5

Off / On: whether the agent can inspect images through readimg during execution.

Case studies

Four single-software GUI workflows from the paper. Each passes its task-specific artifact checks and receives a score of 1.0.

Successful GUI workflows. Intermediate actions lead to deliverables that pass the artifact verifiers. View PDF ↗
EDA / KiCad

Connect a data bus and preserve the circuit

Eight distinct data nets establish the required MCU–SRAM mapping. The saved schematic retains the existing circuitry and passes with zero electrical-rule errors.

CAD / AutoCAD

Build a precise staggered hole pattern

Two rectangular arrays create 23 tube holes. The exported DXF preserves the required coordinates, radii, entity types, and layer assignments.

CAD / SolidWorks

Align parts while preserving geometry

The sliding block is positioned against the base. The exported STEP retains both independent solids and their original dimensions.

3D visualization / Blender

Deliver a scene with a packed AO bake

A 512 × 512 ambient-occlusion image is baked and packed into the saved native scene, preserving the computed pixels in the deliverable.

Where multi-software workflows break down
Multi-software failure cases. Recorded operations and submitted data compared with independently rendered reference artifacts. View PDF ↗

The recorded failures span import-scale errors, incorrect action targeting, and mismatches between submitted representations and verifier expectations. In the load case, a pressure representation does not satisfy the verifier’s force/displacement check; in the CSV case, an unrecognized diameter-field name stops evaluation before downstream geometry checks.

All four illustrated runs receive zero credit. The examples distinguish execution errors from representation and schema mismatches; they do not establish that every zero-score artifact is physically invalid.

Citation

BibTeX
@misc{engiworld2026,
  title = {EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?},
  author = {Hongcheng Gao and Hailong Qu and Yu Lei and Henghui Sun and Haoyang Li and Yipeng Wei and Naihao Xue and Xiaohan Yu and Zhuo Tao and Yihe Zang and Yajiao Wang and Jingyi Tang and Yi Li and Jingjing Zhou and Jie Luo and Bohan Zeng and Chengyu Shen and Hao Jiang and Chong Chen and Bowen Qu and Olive Huang and Zeqiang Wang},
  year = {2026},
  eprint = {2609.37686},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url = {https://arxiv.org/abs/2609.37686}
}