All of the above makes for confident development for high quality AI application where maintaining accuracy is key regression wise, beyond obviously core agentic + model post training tech stack. That's a different post.
Hope you enjoyed!
Automating your app dev as much as possible. I enjoy ideating to code quickly, therefore i set up for automation upfront, almost manual self learn the env if you will, see below ##ArtificialIntelligence :
5/ big time use of Github workflows to automate CI/CD to fully regress , including evals
6/ api led architecture lead to mcp set up and then CLIs, makes it easy to create all kind of tools to test the app
A big fan of custom models in SML. They work very well, especially when you tie inference prompt/schemas to what it was trained with. Big qualify lift. Thanks @tobi for sharing.
That will get a solid architecture out of the box to iterate on quickly and adding features at a blistering pace. Much of the above can be tooled to enforce it all , then speed develops. #coding
1/ API boundaries
2/ shape your tools for mcp around workflows
3/ Headless from the get go
4/ graph your agents & re-inforce with prompts for more reliability
4/ CLIs for agents to easily consume your functions
5 /Include classical encapsulation pattern (LLM context centralized, services under APIs ..)
6/ Heavily use unstructured document like DB , they work really well with context/conversations
7/ Scope your memory arch early on to retain past knowledge
8/ eval where non deterministic
In a world where AI can code so much, your architectural pattern boundaries is what makes the code maintainable, re-usable. #AI#buildinpublic#AgenticAI
sundays = coding besides night hacking. Incredible how SaaS app can be rewritten nowadays into new native AI stacks. For incumbents, good to go back to Day 1 thinking & founder led execution mode.
@twilio your AI online doc does not work well. I use Claude which scans your website instead to get up to date doc info to code. Hope this feedback is taken well.
A big fan of @stripe's acq. of @OpenRouter. Pricing learnings from card network variables costs transfer to tokens naturally. Most new AI like companies now need to package pricing into agentic tasks and learn how to price their outcome against their token costs.
Brilliant move.
The days of a one person startup are firmly here ( code is AI based mostly, web sites can do sales with co-pilots). Obviously, to scale you need more folks esp if larger customers , yet 2-3 folks can grow fast nowadays, even in b2b. wdyt ? cc @ycombinator
Weโre launching DeepSeek-V4-Pro today! ๐
๐ท Major Agent upgrades with strong production gains!
๐ท Flexible reasoning effort for V4-Pro & V4-Flash: low for simple tasks, high for daily Agent workflows, max for complex tasks.
๐ท Native OpenAI Responses API support, optimized for Codex with one-click setup.
V4 Pro is now available on app/web. Try it via โExpert Modeโ.
V4 Pro is also available via API. Model names remain unchangedโplease refer to the API docs for setup details.
DeepSeek is solid. My app is almost as accurate as in Claude. And DeepSeek is my choice for running test suites to not max out more expensive models. Then take a cheap old model for eval judges. Bless choices.
Designing AI memory architectures is fun. Agent Id, user Id, long/short term, cross sessions, search/recall/decay. So far loving @mem0 , well though out and evolving quickly.๐
Back on X after a hiatus. Doing my little part, Open Models matter, esp. for early stage that goes deep to differentiate in new ai app stacks & need to control costs/destiny. RT :)-
For my first post, Iโm sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb