13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
Posted by ibragim_bad 3 days ago
Comments
Comment by sathish316 3 days ago
Comment by spullara 3 days ago
Comment by ducktective 3 days ago
[1] : https://martinalderson.com/posts/which-programming-languages...
Comment by revetkn 3 days ago
Comment by dia80 3 days ago
Comment by cbg0 3 days ago
Edit: In DeepSWE Sol High scores the same as Fable High for ~1/3 of the cost.
Comment by stared 2 days ago
Comment by guywithahat 2 days ago
Comment by pseudosavant 2 days ago
Comment by goldenarm 3 days ago
Comment by joshka 2 days ago
Comment by abratabia 3 days ago
Comment by tcdent 3 days ago
Comment by citizenpaul 3 days ago
Comment by stpedgwdgfhgdd 3 days ago
Comment by rsyring 3 days ago
> Potential data contamination: The SWE-bench dataset, comprising a collection of GitHub issues, has been publicly available since the end of 2023. As a result, models released after this date may have seen these exact issues or highly similar data during training. This raises the risk of inflated performance metrics and makes it harder to distinguish genuine generalization from memorization.