OpenPhotoEdit
Professional photo editing without the learning curve.
✦ Full PSD support, layers, masks & original Photoshop shortcuts.
✦ Switch edit modes seam
Meet OpenPhotoEdit. ✨
A layered photo editor that runs entirely in your browser.
Layers. Masks. Adjustments. PSD support.
Your images stay on your device.
Free & open source.
→ https://t.co/cMf4xA7nka
In a new benchmark across Chess, Go, Hex, and NetHack, one frontier model started at a 0% win rate and learned five times faster per dollar than the top model on the leaderboard.
When an agent fails a recurring task, engineers usually blame the system prompt or the vector retrieval database.
The data in a new paper, "Agent Plasticity: Measuring Self-Improvement Through Experience" (Link in comments), shows that neither is the real culprit.
The best agents retrieved their saved tools 94% to 98% of the time. Yet 83% to 99% of their remaining failures happened while the relevant tool was actively running.
Retrieval worked reliably, but the generated tools broke on edge cases.
Five observations from the paper explain why agents fail when they try to self-improve:
1. Static benchmarks pick the wrong model for long runs.
Evaluations like SWE-bench measure capability at a single frozen moment. They tell you what a model already knows from pre-training, not how efficiently it converts mistakes into working tools.
Claude Fable 5 reached the highest absolute score in chess at 73.3%, but required heavy compute to get there.
GPT-5.6 Sol started near 0%, yet its learning efficiency was 298 score points per $1,000 of compute, compared to 57 for Fable 5. It started worse, synthesized tools cheaper, and overtook older models.
If you are running long-lived agents, picking purely by Day 1 benchmark rank means you pay the highest compute cost to acquire new capabilities.
2. Perfect retrieval does not fix bad tools.
The paper categorized every mistake into three buckets: no tool existed, a tool existed but was ignored, or a tool was used and the agent still lost.
Weaker models ignored their tools on 98% of moves. But for stronger models, missed retrieval dropped to single digits.
Their failure came from tool quality. The agents wrote Python heuristics that passed basic syntax checks, but timed out or blundered under out-of-distribution pressure.
Adding more context or tuning a vector database will not solve a logic flaw inside self-written code.
3. Writing an instruction down does not mean the agent follows it.
Reflection and execution are two different cognitive states.
In one run with GPT-5.5, the reflection phase diagnosed its earlier mistake and wrote an explicit rule: load the move-picking script immediately on move 1.
When the next game started, the execution agent ignored its own rule. It burned four model calls listing directory files and reading logs before taking an action.
Even stranger, agents suffer from spontaneous tool abandonment. GPT-5.5 used its tools on nearly 100% of moves for seven checkpoints, then dropped to 35% tool usage without any change to the prompt or files.
You cannot govern runtime execution through polite prompt guidelines.
4. Clean syntax is where poisoned memory hides.
The harness used a strict validator to reject invalid JSON and broken Python imports before committing any tool to the agent's persistent inventory.
It was not enough to protect the lineage.
In Go, Claude Opus 5 improved steadily across checkpoints until the final round. It generated a tool that passed every import check, but contained an algorithmic flaw. That single script caused the entire lineage's score to collapse.
A tool that crashes gets caught by the runtime. A tool that runs cleanly with bad logic poisons every future run that inherits it.
5.Post-training is currently aimed at the wrong target.
Models are trained to answer questions directly in a single context window. They are not trained to write concise, reusable tools for their future selves.
The next shift in post-training will likely move away from raw answer quality toward rewarding models for creating durable artifacts that reduce future token spend.
OpenPhotoEdit offers Simple and Professional modes in your browser. Start with a quick adjustment, then switch modes when you need layers or masks. The welcome screen says your work is kept when you switch. https://t.co/X9trgMeqqt
@jordanjhamel Reinventing systems for an agentic world means putting user control front and center—let's prioritize creative tools that respect privacy and independence.
@JimmyBearden@Muse Excited to hear about your experience! It's refreshing to see platforms prioritizing user control and a seamless experience without the usual barriers.
Before moving a slider, decide what the image should say. Is the subject hard to see, the frame too busy, or the color distracting? One clear goal makes the next edit easier to choose.
@StefanKelly It’s critical to prioritize tools that empower users with local control while fostering true creativity, rather than relying on vague outputs from LLMs.
@gregnavis Love to see open-source enthusiasm! Merging more PRs will definitely strengthen the community and keep the tool evolving without the usual corporate baggage.
@heymoosh With AI tools influencing voter perception, ensuring transparency and accuracy in information is crucial for civic engagement. How can we improve AI accountability in these contexts?
@Aman25m@OpenAIDevs@OpenAI Creative tools should empower users, not lock them into ecosystems. Open-source is the way to keep control and foster innovation.
DHH told 1,200 developers that writing code by hand is over.
If agents write the code, building the app isn't the hard part anymore. Getting anyone to look at it is.
Every clip, screenshot and meme in this video came from one Vidax run. I spent 5 weeks with Claude teaching it the Fireship edit.