Which LLM agent framework works best for your use case? @aparnadhinak compares the benefits and limitations of LangGraph, LlamaIndex Workflows, and a bespoke code-based agent. https://t.co/dwcdDvXrkp
(1/7) o1-Preview Time Series ⌚Eval Challenge 🥊
Time Series anomaly evaluations trip up most LLMs. We decided to put the new @OpenAI O1 models to the test.
We actually use LLMs for time series analysis in our @arizeai Co-pilot so this test is important to us. This skill is incredibly useful for debugging time series patterns, correlations and problems in data.
TLDR:
✅o1-preview did actually perform a lot better
❌o1-preview is just too slow for our product right now
❌o1-mini was horrible on this Eval
If you take a look at the data - here's a set of traces you can check out on @ArizePhoenix:
https://t.co/ZxiE3PcLNi
If this gets 2-3x faster, we would use it immediately for our time series debugging skill. There is a very real future where we swap in and out stronger models for the “tougher” cognitive tasks.
Results
o1-Preview:
85% of anomalies detected - 110k context window
80% of anomalies detected - 80k context window
95% of anomalies detected - 56k context window
100% of anomalies detected - 30k context window
Claude-Sonnet 3.5:
55% of anomalies detected - 110k context window
75% of anomalies detected - 80k context window
85% of anomalies detected - 56k context window
60% of anomalies detected - 30k context window
o1-mini:
20% of anomalies detected - 110k context window
20% of anomalies detected - 80k context window
25% of anomalies detected - 56k context window
45% of anomalies detected - 30k context window
Tagging relevant LLM Evals folks!
@rown@universeinanegg@ybisk@YejinChoinka@allen_ai@haileysch__@lintangsutawika@hendrycks@markchen90@MillionInt@HenriquePonde@Shahules786@karlcobbe@jerryjliu0@mobav0@lukaszkaiser@gdb@HamelHusain@sh_reya@eugeneyan@DimitrisPapail@_akhaliq@JeffDean@demishassabis@jxnlco@OpenAI@AnthropicAI@GregKamradt@MiqJ
(1/6) Can LLMs Do Time Series Analysis ⏲️? GPT-4 vs Claude 3 Opus 🥊
We have seen a lot of customers trying to apply LLMs to all kinds of data, but have not seen many Evals that show how well LLMs can analyze patterns in data that are not text related - especially timeseries🕰️
Ex: Teams are launching GPT stock pickers 💸without testing how well LLMs are at basic time series pattern analysis!
We set out to answer the following question: if we fed in a large set of time series data into the context window, how well can the LLM detect anomalies or movements in time series data 🤔?
AKA should you trust your money with a stock picking GPT-4 or Claude 3 agent? Cut to the chase - the answer is NO🚫!
We tested both GPT-4 and Claude to find the anomalous time series patterns, mixed in with the normal time series. The goal is to find the anomalous ones ONLY (aka the ones that had spikes).
Tagging relevant LLM Evals folks!
@rown@universeinanegg@ybisk@YejinChoinka@allen_ai@haileysch__@lintangsutawika@hendrycks@markchen90@MillionInt@HenriquePonde@Shahules786@karlcobbe@jerryjliu0@mobav0@lukaszkaiser@gdb@_akhaliq@JeffDean@demishassabis@jxnlco@OpenAI@AnthropicAI@GregKamradt@MiqJ@ArizePhoenix@arizeai
🧵 below shows results:
(1/9) LLM as a Judge: Numeric Score Evals are Broken!!!
LLM Evals are valuable analysis tools. But should you use numeric scores or classes as outputs? 🤔
TLDR: LLM’s suck at continuous ranges ☠️ - use LLM classification evals instead! 🔤
An LLM Score Eval uses an LLM to judge the quality of an LLM output (such as summarization) and outputs a numeric score. Examples of this include OpenAI cookbooks. https://t.co/c7oJwrhWau
In the example here, we ran a spelling eval to evaluate how many words have a spelling error in a document and give a score between 1-10.
✅If every word had a spelling error, we’d expect a score of 10.
✅If no words had a spelling error, we’d expect a score of 0.
✅Everything in the middle would fall into the range with higher percentage of spelling errors landing closer to 10, and lower landing closer to 0.
✅Our expectation: the continuous score value should have some clear connection to the quantity of spelling errors in the paragraph.
‼️Our results however did not show that the score value had a consistent meaning. ‼️
In the example below, we have a paragraph where 80% of words have spelling errors, but has a score of a 10. We also have a paragraph with 10% of words having a spelling error with the same score of 10!
🧵 below is a rigorous study in how well LLMs handle continuous numeric ranges for LLM Evals.
(1/8) Surprising Gemini results for RAG Needle in a Haystack🪡 test ‼️ GPT-4 and Anthropic look better in this eval ‼️
We were expecting better results. Gemini has performed well on all our other Evals, often second to GPT-4, which makes these results more surprising. @demishassabis@JeffDean
Tests run using @ArizePhoenix Evals library.
Code available here: https://t.co/ohB2Sj9bIQ
We wanted to thank @GregKamradt who spearheaded the original idea!
🧵here for more results and analysis:
(1/8) The Needle in the Haystack done by @GregKamradt was an amazing analysis of retrieval performance! Greg has graciously allowed us to build on his work with a repository that is now OSS.
@natfriedman We have a much more rigorous test we’ve put out based on this idea.
@Anthropic we could not reproduce your results. Please reach out to us, we are happy to re-run.
🧵 on how we tested and the results
// Phoenix: Open-source ML Observability //
@arizeai just released an open-source library for LLMs and deep NNs more generally! It's really cool.
Provides, among other features:
- Visual clustering analysis model interpretability in latent space
...
1/2
@aparnadhinak and @kaggle CEO @antgoldbloom will discuss the question "How different is production and the real world compared to build environments?" in their Arize:Observe fireside chat. Register for your spot today to be a part of the conversation!
https://t.co/XNK2YkWCuy
To all cloud founders:
I'm excited to share @BessemerVP's Scaling to $100 Million Report, based on a decade of BVP's cloud investments. In the report we surface the financial and operating benchmarks you should target as you grow.
The 5 Lessons: 🧵
https://t.co/w7if5Ryhh5
We have a major incident on SBG2. The fire declared in the building. Firefighters were immediately on the scene but could not control the fire in SBG2. The whole site has been isolated which impacts all services in SGB1-4. We recommend to activate your Disaster Recovery Plan.