A large part of interpretability in AI Research is built around an implicit assumption:
If we can understand a general model well enough, we can trust it everywhere.
This has led to years of work trying to make general-purpose LLMs interpretable for any user, safe for any task, explainable across any context.
The problem is not lack of effort. The problem is the target itself.
Universal interpretability seems an impossible abstraction.
@kunchenguid Makes sense, the use cases I measured was related to knowledge work.
For coding It was so problematic to use in a first mate style that I didnt even bottered to measure the cost. 😅
@espindola17848@MiguelNicolelis O problema não foi resolvido pela AI do tipo “Resolver essa questão aqui blz?”. Maematicos liderando o projeto e como comparar eu pilotando um carro de fórmula 1 x Ayrton Senna. Ambos tem o mesmo carro, a diferença é o piloto.
Non tech people just want their problem solved, not to think about how to build something and maintain it working as it goes. A therapist, want to do their job, not to become a builder, people’s passion outside the bubble are completely different.
People are not the same, everyone can build an own harness, yet, mostly are using someone else’s software like Claude code, codex whatever. (Same analogy applying to the bubble in discussion).
Too much tech thinking on the problems, not enough designers thinking about human needs and behavior. (Wish we had more designers on those discussions)
In summary, I agree with you @kunchenguid
@atonse@matthewcanham Things is that AI is exploding the agile cult and jira looks like was built around these dogmas.
Here we r not looking and task size but blast radius and risk to estimate not time, but human proximity.
2 years ago I’d say that it needs better UX, now I’d say that it needs better Agent Experience. It was 10x easier to create an upstream>downstream autonomous workflow on GitHub than Jira…
But I’d also admit that companies some times contribute for the over engineered workflows 😅
Build the right roads for agents ffs, they r the ones moving tickets now and generating reports now.
I did in an enterprise level. It’s more a service design skill than engineering.. thing is that most service designers cannot understand LLMs and most engineers have no idea of what’s is service designs so people are still thinking that Knowledge work is about to organize docs to better retrieval only.
@mattpocockuk it’s more service design skills than engineering…
That’s why most people thinks that “Knowledge work” is about to structure data to better retrieval only
Exactly! We mapped over 900 process and 5000 micro tasks in 40 areas, no one has no idea of the kind of workarounds people had to do to achieve the goals. This is the game changer when we talk about Knowledge work.
it’s more service design than engineering work
That’s why most people thinks that “Knowledge work” is about to structure data to better retrieval only
Worth mentioning the tacit knowledge and conversations in the elevator for instance, things that are not in a doc. I bet 99% of the companies doesn’t know what the team do. We mapped over 900 process, 5000 micro tasks happens every month, mostly are workarounds to handle the own companies entropy to make things work.
@mattpocockuk And surprisingly, it allows you to understand better than code on how instructions, recall and memory works… Because you focus on LLM only and abstract the code complexity.
We built a conversational commerce agent from upstream to downstream with evals on both parts, at first as skills, than we upgraded to DAGs. It allowed us to spend up the process and validate problem-solution to filter hypothesis to downstream which we could run simulations through real and synthetic traces very gong not just success but how it relates to the intent.
What we learned is that it makes things looks magic, so product judgement is the most important thing even on workflows like that to know when it can or cannot run autonomously.