@ishaansehgal If you have the agent running inside of the sandbox AND record every log/event it does, such that you can reply the session in a new sandbox, is there a reason why this isn’t also viable?
@kunchenguid Have you thought about switching a model mid-task if u detect the agent not being sufficient enough for the task? Ex. A sidecar agent/jev watching the task to see if the agent is not getting enough context or not completing the task correctly?
@scheredev@steipete I’d imagine a lot of the complexity comes from data/api/iOS integrations from instinct
Re: innovation. The biggest innovation of the last year (coding agents) is just a simple LLM in a while loop calling tools, nothing complex about it
Before, the industry turned to DAGs because they made agents more reliable/accurate (Langgraph)
Currently, the industry moved to agents because they can handle more and more dynamic complex inputs and uncertainty with high accuracy (Claude code , codex)
Future/now, the industry is looking back to DAGs as a cost & latency improvement rather than an accuracy improvement. If the task can be broken down into a DAG, it’s 100-300x cheaper/faster now.
Still I agree with you, short term memory, but good to put everything in perspective
This works when the agent flow/user journey is simple.
Most times, need to handle:
1. Conversational queries + large workflows
2. Ambiguity of information, where the agent needs to spend a few minute figuring out the missing details and then creating a plan
Are you saying you can put that into code? And if so, what does it look like?
@miu21590@1yian I’m glad OpenAI started training their models without the reasoning effort/budget in the 3rd line of the system prompt for Astra . Update priors @devagrawal09 .
@mstockton My thoughts too. Industry was forcing agents down everything, which worked because it was easy and prompt tuning got less important. Cost was always a factor but efficiency/latency is now a better indicator to switch some parts of the agent to deterministic JEV functions
@DhravyaShah R u releasing these .md’s? Do they have any structure to them that we can infer based on the md’s? Weekly vs daily structure would be interesting .
@DhravyaShah@mohsen____ Great stuff! I’m not following the git part. Is this local git? Or are user memories actually getting stored in a giant mono repo somewhere …
@CatalinCalistr1@kushbhuwalka@typesafeai RAG usually means there’s some vector/cosine similarity involved . Technically “retrieval” could mean anything, but I was referring to vector cosine similarity. And when I said “rag is dead pt2” I was referring to it being killed by Grep/Bash and agentic coding/search