Current benchmarks miss the central challenge in agent optimization.
Most evaluations test one-shot improvement, while real-world use requires recursive optimization as new tasks continue to arrive.
To study this, we built a two-phase continual-learning evaluation on hard Terminal-Bench 2.0 tasks, comparing GEPA, Meta Harness, and RELAI.
The methods behaved very differently:
- GEPA improved initially but transferred negatively to unseen tasks, indicating overfitting.
- Meta Harness transferred well but could not continue improving.
- RELAI was the only method that both generalized to unseen tasks and improved again without regressing prior performance.
Full results are in the article below.
You can access RELAI’s continual-learning engine to run similar optimization studies and apply it to your own agents at https://t.co/xuJxZo1Nqj
@shai_wininger Really interesting and definitely a call to action 😀. At @packmind_app, we're building a tool to solve exactly this problem: capturing your unique context and technical decisions in all their nuances, guiding AI assistants, and deterministically enforcing these rules.
Remember when GitHub said that Copilot helped engineers write code 55% faster?
Well, a first of its kind study has concluded that this will lead to a doubling of the rate of code churn - the percentage of lines that are reverted or updated less than two weeks after being authored.
The report was pretty scathing in its analysis, saying:
Code generated during 2023 more resembles an itinerant contributor, prone to violate the DRY-ness of the repos visited.
Je suis ravi de lancer le podcast "Tech Lead Corner" dédié aux missions des Tech Leaders.
Quel joie d'accueillir @DorraBartaguiz, VP Tech chez @ArollaFr, pour ce premier épisode 😊(dont voici un extrait)
A retrouver sur https://t.co/jJi6M8Y9C6
@libe Cette comparaison n a pas de sens. La fortune de BA vient de son entreprise qui crée de la valeur et dont la valorisation augmente en consequence dans le temps. Conclusion: investissez! Ceci n enleve rien au besoin de taxation mais utilisez des vrais arguments...
Linear in the software engineering project management world feels like what Figma is in the design world.
I don’t know any startup that went back to another tool after using Linear, and have not seen software engineers before raving about *a project management tool*
This is exactly what happened at Uber Eats 3-4 years ago.
A few key leaders who were executing amazing at Uber Eats were denied their L+1 promotion thanks to Uber Prime’s bureaucracy. So they went to L+2 to DoorDash and are still crushing it. What a loss for Uber, win for DD.
If the best public cloud company (Snowflake) trades at 20X EV/NTM revenue with 60% NTM growth at scale and +FCF, mid-stage startups are not going to do rounds in the near term at much bigger multiples
Expectations still meeting the asphalt of reality for private software co’s
Things that don’t exist:
• “Zero to one” product manager
• Viral content marketer
• Full stack engineer
• Designer who can prototype
If you find one, don’t fuck it up because they can start their own company. They don’t actually need you.
It is not even 4PM on the first day back and I have already seen:
$100M pre for pre-product company.
$1BN pre for $1M ARR company.
$10BN pre for company less than 15 months old.
My oh my, welcome 2022.