arXiv:2609.33382September 2026

WideSWE

Can Coding Agents Coordinate
Changes Across Repositories?

Baoyi Wang1,*Xingliang Wang1,*Jinyang Wu2 Keming Wu2Chen Zhi1,†Jianwei Yin1

1 Zhejiang University2 Tsinghua University

* Equal contribution   † Corresponding author

One shared request. Multiple repositories.
Can an agent complete every part?

Video overview

1 min 11 sec

01 / THE BENCHMARK

One request, across repositories.

A shared feature or bug fix can span multiple repositories. WideSWE brings together 120 real-world tasks across 41 software ecosystems, balanced between 60 bug fixes and 60 features. Each task pairs a shared request with a historical workspace and hidden tests.

Shared requestOne feature or bug fix
Historical workspaceTargets + context repositories
Repository patchesAgent-generated changes
Hidden evaluationRequired behavior + regressions
01

Identify the scope

Determine which repositories need to change, with related repositories available as context.

02

Deliver the changes

Implement every repository-specific obligation of the shared feature or bug fix.

03

Satisfy the request

Pass the required behavior and regression checks in every target repository.

Three observed cross-repository failures: incomplete scope identification, recognized work without delivery, and post-edit failure.
Scope, delivery, and post-edit failures in cross-repository tasks. Figure 1 of the paper.

02 / CONSTRUCTION

From linked changes to executable tasks.

Real issues and pull requests establish the shared requirement. Historical snapshots and fail-to-pass checks make it executable.

  1. 01

    Mine linked changes

    Collect merged PRs across 103 ecosystems and group changes linked across repositories.

  2. 02

    Review the request

    Retain 635 groups addressing one shared feature or bug fix with substantive cross-repository changes.

  3. 03

    Validate behavior

    Executable validation and final review yield 192 eligible cases, with fail-to-pass tests for each target.

  4. 04

    Balance the release

    Select 120 tasks: 60 bug fixes and 60 features, covering 41 software ecosystems.

Shared prompts. Requirements are consolidated from original issues and PRs, without turning reference implementations into step-by-step solutions.

Reviewed hidden tests. Implementation-specific constraints and unrequested checks are removed while required behavior and regression checks are preserved.

Construction and ecosystem coverage WideSWE mining and case selection pipeline, prompt and test construction, and coverage of 41 ecosystems.

03 / MAIN RESULTS

Progress is not completion.

The best-performing configuration fully solves 42.50% of tasks, while solving at least one target repository in 83.33%. Completing the shared request remains a challenge.

Task success: GPT-5.6-sol with Codex CLI 42.50%; Qwen 3.8 Max 37.50%; Claude Opus 5 35.00%; GPT-5.6-sol with Claude Code 32.50%; DeepSeek V4 Pro 26.67%; GLM 5.3 20.00%; Gemini 3.8 Flash 10.83%. All configurations except the first use Claude Code.
Task success on the same 120 tasks. A task is solved only when every target repository passes all required checks.
Results by task type
Task success (%) from Table 1 of the paper
ModelAgentBugfixFeatureOverall
GPT-5.6-solCodex CLI41.6743.3342.50
Qwen 3.8 MaxClaude Code48.3326.6737.50
Claude Opus 5Claude Code41.6728.3335.00
GPT-5.6-solClaude Code33.3331.6732.50
DeepSeek V4 ProClaude Code38.3315.0026.67
GLM 5.3Claude Code30.0010.0020.00
Gemini 3.8 FlashClaude Code21.670.0010.83

04 / TRAJECTORY EVIDENCE

Where does completion break down?

Selected runs expose different gaps between understanding a shared request and delivering working changes.

Incomplete scope identification

Sentry SDKs

Codex CLI · GPT-5.6-sol

Shared requirement

Implement the same strict trace-continuation policy in the Go, Python, and Ruby SDKs.

Observed behavior

The agent treats Go as the target and reads Python code as a reference, but only Go receives a patch.

HIDDEN F2P CHECKS Passed / total

Go SDK14 / 14Required checks pass
Python SDK0 / 22Unmodified
Ruby SDK0 / 13Unmodified

Reading a related repository does not ensure that its required changes are delivered.

Source: Appendix C.1. Selected execution trajectories, not aggregate model rankings.

Read the full case studies in the paper

05 / EXPLORE THE RELEASE

120 tasks. Real development work.

From SDKs and developer tools to cloud infrastructure and machine learning libraries.

RUN WIDESWE

Evaluate your coding agent.

Task definitions, historical revisions, container environments,
and evaluation harness in one open-source release.

Get started

CITATION

Cite WideSWE

@misc{wang2026wideswecodingagentscoordinate,
      title={WideSWE: Can Coding Agents Coordinate Changes Across Repositories?},
      author={Baoyi Wang and Xingliang Wang and Jinyang Wu and Keming Wu and Chen Zhi and Jianwei Yin},
      year={2026},
      eprint={2609.33382},
      archivePrefix={arXiv},
      primaryClass={cs.SE},
      url={https://arxiv.org/abs/2609.33382},
}