SWE-2 benchmarks: results, limits and how to test it

By CodeAgentSwarm · Updated September 11, 2026

Choosing a coding model means checking which tasks it solves, how long it takes and how much work it leaves for the reviewer. This guide brings together published results and a simple way to test whether SWE-2 fits your project.

Results published by Cognition

Devin CLI

Cognition published this comparison on September 10, 2026. These are vendor-reported figures, not tests run by CodeAgentSwarm.

BenchmarkSWE-2Fable 5.1GPT-6 Astra
FrontierCode 1.1 Main50.0%50.9%53.3%
DeepSWE 1.173.0%67.4%74.1%
Terminal-Bench 2.192.8%91.4%89.9%
Terminal-Bench 427.3%55.8%57.9%

SWE-2 stands out on Terminal-Bench 2.1 but trails substantially on Terminal-Bench 4. The choice depends on the task.

How to read this comparison

The methodology appendix combines public results with internal evaluations, uses different agents and selects the best reasoning effort per model. It is not a fixed-configuration comparison.

FrontierCode 1.1 Main contains the 100 hardest tasks from Extended. Cognition develops both this benchmark and SWE-2. Keep that relationship in mind when interpreting the results.

Keep the benchmark name and version attached to every score. An aggregate result is not the probability that an agent will solve your next bug: your project may use different tools, languages and constraints.

Choose a specific reasoning effort

The model documentation lists SWE-2 Medium, High and Max. Record which one you test so you can repeat the comparison.

Devin model picker in CodeAgentSwarm beta
Real beta capture showing the picker. It does not depict a benchmark run.

Start with the effort you would normally use and retry failures that justify more time. If you change effort, keep the failed attempt in your record too: discarding it would make the comparison look better than it was.

A small test on your repository

This is our suggested local evaluation; we have not run this protocol as a SWE-2 benchmark. Pick a reproducible bug, an interface change and a maintenance task that you know well.

  • Use the same starting commit, dependencies and tests for each attempt.
  • Give each model the same request and tools. Keep their changes separate.
  • Record total time, human intervention and actual usage, including retries.
  • Review the diff and run the tests before scoring the solution.
text
Task | Model + effort | Tests passed | Review fixes | Elapsed time | Usage
Bug fix | SWE-2 High | ... | ... | ... | ...

A useful result is a change you would accept in the project. Count your own fixes and review time; a quick response can become expensive if it needs corrections.

Assess cost after checking quality

For Pro pricing, the temporary promotion and its end dates, see the Devin models and quota guide. Record your account terms on the test date so the result remains understandable when the offer changes.

If you have not installed the agent, start with Devin CLI installation and login.

FAQ

No. The table contains results published by Cognition. The suggested local test is a way to evaluate the model on your own project.

The table does not establish that. Test representative tasks and consider correctness, human review, time and usage before choosing.

Yes, but you will be comparing the model together with its execution environment. Record the tools, permissions and settings in each environment.

Devin Chat, installation and history are in CodeAgentSwarm beta testing for an upcoming release. The button downloads the current public app; check its release notes for Devin availability.

Download free betamacOS & Windows