Skip to content
VectorSpaceLabPublic

About

Can coding agents build the agent you actually want? A benchmark of 25 agent software engineering tasks across creation, editing and optimization.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

AgentSWE

Can coding agents build the agent you actually want?

project page arxiv license

News | Benchmark | Results | Lite | Quick Start | Citation | Chinese

AgentSWE is a benchmark for agent software engineering. A coding agent gets a natural-language commission for an agent, builds or changes it, and the agent it delivers is tested on cases it never saw. The 25 tasks cover the three stages of an agent's life: Creation (10), Editing (10) and Optimization (5).

✨ Highlights

  • The whole lifecycle. Build an agent from an empty directory, add capabilities to ten production agent codebases (Aider, OpenHands, Codex, DeepTutor and others), or improve a working agent on BrowseComp, τ³, Terminal-Bench, PinchBench and OSWorld.
  • A commission, not an exam. The builder sees the requirement and a few example cases. Its frozen delivery is then run on held-out cases that cover the rest of the requirement, behind pass/fail gates and a rubric derived from the requirement. Optimization tasks use the source benchmark's own metric.
  • Today's coding agents fall short. The 50 Creation runs average 11.2 out of 100 and 20 of them score zero. On Optimization, no builder closes more than 5% of the remaining gap on average.
  • Easy to try. Every task runs in a pinned, isolated environment. The 8 Lite tasks need only a DeepSeek API key, so we recommend starting there.

🔥 News

  • 2026-10: Code, all 25 tasks and the Lite set are released.

📦 Benchmark

Stage The builder starts from Tasks
Creation an empty workspace authorized vulnerability validation, database analytics, desktop GUI automation, document-to-editable PPTX, evidence-grounded document QA, formal theorem proving, repository bug repair, schema-guided web extraction, scientific PDF translation, web research report
Editing a pinned production agent AI-Scientist, Aider, Claude Code, Codex, DeepCode, DeepTutor, Dyad, OpenClaw, OpenHands, OpenWiki
Optimization a working starter agent BrowseComp, τ³ (retail), Terminal-Bench, PinchBench, OSWorld

Each Creation and Editing task has 2 development cases and 6 held-out cases. Optimization tasks use disjoint development and held-out subsets of the source benchmark.


The builder iterates on the public cases, then its frozen delivery is evaluated on held-out cases.

📊 Results

Mean held-out score of five builders, all in the Codex harness. Creation and Editing report Result (0 to 100); Optimization reports how much of the gap to the target the delivery closes (negative when it scores below the starter, with no lower bound when the starter is already close to the target). Per-task results are in the paper and on the project page.

Builder Creation Editing Optimization
GPT-5.6 Sol 13.7 23.9 4.7
GPT-5.5 5.8 31.6 1.2
DeepSeek-V4-Pro 19.1 31.6 -4.0
DeepSeek-V4-Flash 15.1 14.0 -7.1
Qwen3.6-35B-A3B 2.3 0.0 0.9

What a builder shows during development is not what it delivers. The trajectories point to three recurring failure modes: overfitting to the visible cases, hollow runtimes behind code that looks complete, and claims of completion that the records contradict.

AgentSWE-Lite

AgentSWE-Lite is 8 of the 25 tasks, built three times per builder in Codex, with DeepSeek-V4.1-Flash (deepseek-flash) as the runtime model and the judge in all three stages (profile lite-v1.1). It needs only a DeepSeek key, no search key and no release assets, so it is the easiest way to evaluate your own coding agent. Scores are the mean of the three builds. Because the runtime model and the judge differ from the main study, the Creation and Optimization levels are not comparable with the table above.

Stage Task GPT-5.6 Sol DeepSeek-V4.1-Flash Qwen3.6-35B-A3B
Creation Scientific PDF translation (pdf) 18.9 31.9 0.0
Database analytics (db) 7.4 38.7 0.0
Repository bug repair (repo) 42.9 30.0 0.0
Editing DeepTutor adaptive remediation (deeptutor) 52.2 27.6 0.0
Aider worktree transaction (aider) 13.0 39.3 6.0
OpenWiki change-impact documentation (openwiki) 22.8 21.0 0.0
Optimization τ³ (tau3) 10.1 13.0 4.3
PinchBench (pinch) 19.4 24.3 2.4
Mean Creation / Editing / Optimization 23.1 / 29.3 / 14.8 33.5 / 29.3 / 18.7 0.0 / 2.0 / 3.4

🚀 Quick Start

You need Linux x86-64 with Docker and Python 3.10 or newer. Some tasks need more (KVM, cgroup v2 with systemd, large disks); doctor tells you what a task needs and whether your host has it. pip install -e . adds an agentswe command, the short form the docs use for python3 -m agentswe.

git clone https://github.com/VectorSpaceLab/AgentSWE.git && cd AgentSWE
cp .env.example .env                  # set AGENTSWE_DEFAULT_API_KEY to your DeepSeek key
python3 -m agentswe list              # the 25 tasks
python3 -m agentswe doctor repo       # host check, no model calls
python3 -m agentswe probe-roles       # one short request per model role with your key
python3 -m agentswe setup repo        # build the task environment (once)
python3 -m agentswe run repo --builder codex --smoke
python3 -m agentswe status
python3 -m agentswe result <run_id>

--smoke runs a small budget to check that everything works: one held-out case, after one development submission for Creation and two for Editing. On a well-connected server the first setup and smoke run of repo take 25 to 45 minutes; an Editing smoke takes 10 to 40 minutes. Drop --smoke for a full run under the paper's protocol. Smoke scores are not comparable with the paper's.

The web research, PPTX and BrowseComp tasks also need a Serper-compatible search key. Behind a firewall or a proxy, see docs/ENV.md.

Run Lite (recommended)

The 8 Lite tasks are repo, db, pdf (Creation), deeptutor, aider, openwiki (Editing), tau3, pinch (Optimization). With the Lite profile every role uses DeepSeek-V4.1-Flash (deepseek-flash), as in the Lite runs above.

echo AGENTSWE_PROFILE=lite-v1.1 >> .env
python3 -m agentswe doctor repo db pdf deeptutor aider openwiki tau3 pinch
python3 -m agentswe setup <task>
python3 -m agentswe run <task> --builder codex

The three Editing tasks also need cgroup v2 with systemd delegation; doctor checks it.

📚 Documentation

🙏 Acknowledgements

AgentSWE runs every task through Harbor and adapts many open-source agents and benchmarks; THIRD_PARTY.md lists them with their licenses.

📝 Citation

@misc{wang2026agentswe,
  title  = {AgentSWE: Can Coding Agents Build the Agent You Actually Want?},
  author = {Jiahao Wang and Hongjin Qian and Yuyang Hu and Jiajun Zhang and Zheng Liu},
  year   = {2026},
  note   = {arXiv preprint coming soon}
}

License

Apache License 2.0; see LICENSE and NOTICE. Task upstreams keep their own licenses and are downloaded during setup, not redistributed.

About

Can coding agents build the agent you actually want? A benchmark of 25 agent software engineering tasks across creation, editing and optimization.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages