@antirez I should have used the word consistency instead of reproducibility. E,g, different models may require adjustments to prompts/skills to perform their best. Better overall may be worse in particular scenarios. Losing the model one has invested in may require painful adaptation.
@antirez This makes difficult using DS in scenarios where reproducibility is important: agents doing e.g. automated checks. Since you decide what better is, users will need to periodically retest their workflows or be left behind. Not bad, just limits use cases which is a fair UX choice.
@simonw@awnihannun More precisely, it implements an efficient LLM token selector driven by a JSON schema, with application examples on MLX. This is better that GNBF in that mapping to GNBF can produce unwieldy rule sets for things like min and max number of items in an array . @awnihannun@simonw
We evaluated GPT-4-Turbo in our pipeline and found it a very useful addition to the palette of LLMs available, although not a complete game-changer.
✅ Against GPT-4, we saw 66% reduction in latency and cost with an estimated 5% decrease in quality (F1 score).
(1/2)
What’s the point of AI if you can’t rely on its output? We are building a trustworthy assistant for Confluence that tells you exactly how it got its answers. Not just sources, but every reasoning step. Read more in our blog: https://t.co/qUU2lcZXBE
Atlassian has just released the public beta of Atlassian Intelligence. We had a first look and found it still has some way to go before being trustworthy. Link to our report below.
Connie AI can now answer sophisticated questions about tabular data. We translate natural language queries to a pseudo-SQL language and execute it over your Confluence tables. Support for the new Confluence Databases coming soon.
@typesfast Connie AI (https://t.co/LnKt8NFDz5). We’re ex-Atlassians building an AI assistant for Confluence. It’s miles ahead of Confluence search already.
Our internal evaluation tool, Gaucho, allows us to keep tabs on the quality of our system: retrieval, LLM prompts, etc. We wrote about it in the blog: https://t.co/htUB96qUqe