@teortaxesTex run some tests before, recent models are more likely to regard itself as Claude in coding context, which I think may result from the massive co-author records in Github. https://t.co/6vYcckQ5Ju
Introducing SWE-1.7, the most capable model we’ve trained yet.
It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s.
RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale