SWE-2 benchmarks: results, limits and how to test it
By CodeAgentSwarm · Updated September 11, 2026
Results published by Cognition
Cognition published this comparison on September 10, 2026. These are vendor-reported figures, not tests run by CodeAgentSwarm.
| Benchmark | SWE-2 | Fable 5.1 | GPT-6 Astra |
|---|---|---|---|
| FrontierCode 1.1 Main | 50.0% | 50.9% | 53.3% |
| DeepSWE 1.1 | 73.0% | 67.4% | 74.1% |
| Terminal-Bench 2.1 | 92.8% | 91.4% | 89.9% |
| Terminal-Bench 4 | 27.3% | 55.8% | 57.9% |
SWE-2 stands out on Terminal-Bench 2.1 but trails substantially on Terminal-Bench 4. The choice depends on the task.
How to read this comparison
The methodology appendix combines public results with internal evaluations, uses different agents and selects the best reasoning effort per model. It is not a fixed-configuration comparison.
FrontierCode 1.1 Main contains the 100 hardest tasks from Extended. Cognition develops both this benchmark and SWE-2. Keep that relationship in mind when interpreting the results.
Keep the benchmark name and version attached to every score. An aggregate result is not the probability that an agent will solve your next bug: your project may use different tools, languages and constraints.
Choose a specific reasoning effort
The model documentation lists SWE-2 Medium, High and Max. Record which one you test so you can repeat the comparison.

Start with the effort you would normally use and retry failures that justify more time. If you change effort, keep the failed attempt in your record too: discarding it would make the comparison look better than it was.
A small test on your repository
This is our suggested local evaluation; we have not run this protocol as a SWE-2 benchmark. Pick a reproducible bug, an interface change and a maintenance task that you know well.
- Use the same starting commit, dependencies and tests for each attempt.
- Give each model the same request and tools. Keep their changes separate.
- Record total time, human intervention and actual usage, including retries.
- Review the diff and run the tests before scoring the solution.
Task | Model + effort | Tests passed | Review fixes | Elapsed time | Usage
Bug fix | SWE-2 High | ... | ... | ... | ...A useful result is a change you would accept in the project. Count your own fixes and review time; a quick response can become expensive if it needs corrections.
Assess cost after checking quality
For Pro pricing, the temporary promotion and its end dates, see the Devin models and quota guide. Record your account terms on the test date so the result remains understandable when the offer changes.
If you have not installed the agent, start with Devin CLI installation and login.
FAQ
No. The table contains results published by Cognition. The suggested local test is a way to evaluate the model on your own project.
The table does not establish that. Test representative tasks and consider correctness, human review, time and usage before choosing.
Yes, but you will be comparing the model together with its execution environment. Record the tools, permissions and settings in each environment.
Devin Chat, installation and history are in CodeAgentSwarm beta testing for an upcoming release. The button downloads the current public app; check its release notes for Devin availability.