playing with Jev, i built a small @pidotdev extension that sends my latest Git diff to Jev and asks:
→ how strong is this change according to a defined rubric?
it returns a 0–5 score with confidence:
git diff → Jev score + confidence → review signal
in simple terms, it's a simple code review, not a replacement for tests or human review, but this gives me a very fast signal for deciding which changes deserve a closer look
¡POR FIN! Claude Code añade soporte a AGENTS.md.
No tienes que hacer nada:
→ Por defecto lee el AGENTS.md
→ Si no lo encuentra, leerá el archivo CLAUDE.md
Disponible desde la versión 2.1.277
Sometimes it's easier to show than tell. We're sure this update will help with that. 👀
GitHub CLI now has a repeatable --attach flag that uploads a local image or video. Reference it inline in an issue, pull request, or comment body.
Available now to all users on GitHub across all plans. 🎉
https://t.co/2oRtlSHdPT
my day-to-day Pi setup lived locally and moved faster, so i’m consolidating them, in the same repository.
current extensions, skills, theme, and portable settings.
the recommendation stays same: take it as reference and build setup that fits how you work
https://t.co/clrzPKRjaR
good part: awesome solution, great fan of the tool we use it everyday, love to continuing moving most of the workflows to the cloud
bad part: starting at $299/m 🙃
i killed plan mode in my pi extension and put grilling on the same shortcut
i kept hitting the same loop. turn on plan mode, get a generated document, then bounce on the steps with the agent while the raw idea was still fuzzy. that document was never for me. it was something to pass to the next actor or process. routing factory input through my chair was the slow pat
so planning belongs in a workflow/software factory, on a planner subagent that does not need me in the loop
inspired by @mattpocockuk grilling skill. i made it a mode, not only a skill. constraints + focus:
the grilling skill, and only the tools this hard-test needs. read, search, ask. i sit with the fuzzy thing until we share a goal and the leftover assumptions have names. no obligation to emit a document. keep the contract in the session, or optionally generate an artifact later if another actor needs a handoff
different value, different box. grilling produces shared understanding while i am here. planning designs and produces what the next process consumes
grilling stays on the coding agent, planning goes in the factory box
We're using a new method to hire engineers at @HelloUntangle.
I think this playbook will become the standard for hiring technical talent in the age of agents.
Here's how it works:
1. We ask candidates to submit a screen recording of their whole screen, shipping a new feature on a current app they work on.
Note: This is usually 45 minutes of video - no editing or speeding up - we want to see them thinking, prompting and problem solving with the agent - we also want to see their devops skills.
2. We watch these videos and pick a subset to go through Round 2.
3. Round 2 is giving them a seat on @DevinAI so they have secure access to our repo.
4. We ask them to ship a new real-world feature and get the PR merge-ready.
5. We pay them a fixed fee to do this, and they have 16 hours to complete the work.
6. We watch the videos, review at all the Devin threads, have an agent do the same, and then choose which one to hire.
Note: We don't have any meetings with the folks before any of these things. It's all about the work and the results.
Introducing fx, a tiny, open, native coding agent from Vercel Labs.
Originally an internal tool, fx is a harness and CLI written in Zig, optimized for research and embedding in larger systems. Today, we're open sourcing it.
fx is built on three principles:
1. Fast. A single native binary, no runtime to install. It cold starts in 10µs and does no unnecessary work or I/O before accepting input. fx is the answer to "how fast can a coding agent be?"
2. Light. The 6.3MiB binary uses single-digit megabytes of memory at baseline, made for instant installation and embedding in resource-constrained environments and agent sandboxes.
3. Open. Apache-2.0, model and provider agnostic, suitable for local and cloud inference. Its small core extends through skills, plugins, and MCP.
Minimalism is an obsession throughout the entire harness: system prompt, tools, features, binary. The goal was to keep context usage and time to first token low, and make fx optimal for model benchmarking, sandboxing, evals, and gyms.
You can use fx directly or embed it as infrastructure. The CLI feels more like a Unix shell than an IDE in the terminal: it preserves scroll history, produces minimal output, and uses complex TUI rendering very, very sparingly. Programmatically, 𝚏𝚡 𝚊𝚜𝚔 --𝚓𝚜𝚘𝚗 gives structured output, 𝚏𝚡 𝚊𝚌𝚙 connects to editors and other clients, and WebAssembly can even run the whole thing inside the browser (see: https://t.co/wf2Trg47sC).
Privacy is a design constraint: no product telemetry, sessions and usage stay local, and no source code or prompts are shared with any endpoint other than inference. With local inference and auto-updates off, fx is fully hermetic.
fx is experimental. Use at your own risk and expect frequent changes. Chat with us on X (https://t.co/A2AB2YythC) or file issues (https://t.co/GEjTHSoa1J).
𝚌𝚞𝚛𝚕 -𝚏𝚜𝚂𝙻 𝚏𝚡.𝚜𝚑/𝚜𝚎𝚝𝚞𝚙.𝚜𝚑 | 𝚋𝚊𝚜𝚑
https://t.co/g2uEuXhGnt
should a documenter subagent depend on an external service, or can we build it into the coding agent’s session and let it create documentation directly in the codebase?
i think it should be part of the coding session:
↳ observe decisions, conversations, and implementations
↳ identify what needs documenting
↳ preserve context and constraints
↳ create useful artifacts at the right moment
↳ follow explicit rules
the hard part isn’t generating documents.
it’s deciding when the agent should act, what authority it has, and how its output shapes future work.
i sat on this for weeks testing it out,
@herdrdev it is now an essential part of my workflow, such an awesome tool that fully adapted to me very quickly
did not want to talk about another new tool i would drop after a few days. if i was not going to get real value, or use it daily, i was going to stay quiet
i ignored tmux for years. you can feel the value, but the on-ramp is practice and muscle memory. so you never start.
i tried cmux next. same wall. new app, new configs, time i did not want to spend. that friction kills adoption to me.
herdr skipped that. not another terminal to download. just a package. install it, start from a small config, stay in the terminal you already have.
then it stuck.
not a new category. same old primitive: split sessions, live in the terminal. different job. a workspace for a herd of agents, not one pretty chat window.
one coding agent is fine until you want a reviewer, a researcher, and a runner at the same time. tabs hide that mess. panes do not.
two ways people miss this:
↳ you are scared of multiplexers, so you never start.
↳ you already have a friendly GUI, so you tell yourself you do not need one.
i did both.
so, what actually stuck:
steal other people's plugins and configs. do not steal their whole layout. build your own. that is the real value. custom is the product.
if "terminal multiplexer" still sounds like a senior-sysadmin toy, start dumb. two panes. one agent. one notes buffer. curiosity shows up after days, not day 0.
IMO, this is not another overhyped agent wrapper. it is tmux with the camera pointed at the work we actually do now.
That's right, GPT-5.6 Sol is awesome and can be used pretty much anywhere, including in the CC harness.
To celebrate this, together with the fact that I'm not going anywhere... I have reset usage limits for all paid users of ChatGPT Work and Codex.
Have fun out there!