We are running ai-native org.
Our own agent harness (similar to YC's QM) captures all client and internal meetings, all engineering decisions from Claude Code sessions. Monitors all the logs and fixes the bugs on his own. Allows you to do analytics and does learning loops.
Yesterday our talent person (first ever!) asked us to vibe code simple internal holiday tracker.
At the end as a startup in HR we just decided to build this for our new Employee Experience module, which currently only has exit interviews with agents.
By 12pm today our founding designer already created PR to the frontend. And this afternoon I will couple it to the backend and we release at 5 pm.
It was not in our quarter roadmap, but who cares if it is easy to build? and we can dogfood it before giving it to clients.
yes, for truly local coding, llms must become smaller keeping same perf at software tasks AND local hardware must become better (which is very slow process of upgrading device from M4 to M5, M6 etc). Cloud is iterating faster 100%.
And a lot of software tasks are done from the phone, so you have to run them on cloud 100%.
But I see now more reasons beyond privacy
1. costs motivation (Pro plans not enough anymore!),
2. limited energy and water resources (it is finite, so far energy stocks are going up with ai and why space data center idea makes sense).
@GuliMoreno There are mixed signals.
Meta/Google/HF/China are pushing for local.
But Ollama on the other hand released Cloud.
Grok is cloud 100%, but they need to deliver on their space data centers, so incentives are clear there.
23:59. i am now dumping our hiring agent turn traces into jsonl files (one per client). The plan is to ask CC to analyze each file, the agent code and prompt that produced the trace, create an artifact and then discuss with CC the findings.
Tomorrow (or more likely late tonight), come up with nice skill to put this into our harness Ada for automation. And tomorrow during the day repeat the process for other agents.
Today at Orbio AI we had a big luck to learn from gran @Carles_Reina and his time at @ElevenLabs . Best of luck on your next adventure!
Europa tiene un talento brutal
Product learning today: "Send directly to the agent".
Users should send their bug requests directly to a coding agent, that will debug and try to solve it in the background. Sending it to FDEs or to automatic ticket system is sub-optimal: at the end a human (FDE or on-call) will trigger the same coding agent to debug anyways. Therefore the human is a bottleneck.
So why not let the agent do it automatically? That bug request with all the logs from frontend, screenshot or screen recording, combined with backend logs for that user and time is everything the coding agent needs.
We are running ai-native org.
Our own agent harness (similar to YC's QM) captures all client and internal meetings, all engineering decisions from Claude Code sessions. Monitors all the logs and fixes the bugs on his own. Allows you to do analytics and does learning loops.
Just finished an interview with the candidate I liked and had to "sell" him a bit Orbio AI. It is pity that very few in Spain/Europe (not even mentioning US) know about us, while we are in top quartile in terms of growth/arr according to a16z (b2b ai): https://t.co/SgPwmfhnHd
All of it with a small, talented and ambitious team: <30 people, 15 builders (ai engineers + FDEs), ex-McK/BCG deployment strategists to handle large enterprise deployments.
And we are actively hiring: https://t.co/9XTPZvVpOK
Some notes i had:
1. Having human vs agent on websites (added to my website)
2. Send to agent in the website (will add it to our app)
3. Build own tools
4. Expose some settings to make website more personalized
5. SF art and pins were cool
Saw ads on instagram that Bending Spoons is opening Madrid office, with a lot of graduate level roles.
I realized that this is the same company that recently acquired Airtable, but also owns Evernote, AOL, Vimeo, WeTransfer, Eventbrite, Meetup! 🤯
All operated from Italy.
We at Orbio AI have a similar company wide harness, that
1. checks logs, finds and solves bugs
2. can query any analytics question
3. does preliminary handling of customer requests on Slack channels
4. remembers Claude Code sessions, context on features, etc
We’ve decided to open-source a multi-agent harness we use internally at YC.
We call it “QM” and it’s meant to be easy to customize, like Hermes or OpenClaw, but useful for a whole company. We use it across accounting, legal, events, and engineering (including building QM itself!).
The whole project is under an MIT license. It is cloud-first and has Slack and web UI natively.
LLM models keep getting better every quarter. Building a reliable agent on top of them is a different job entirely. Wrote about why.
https://t.co/Ky18QZnnKA