๐๐ณ ๐ฌ๐ผ๐๐ฟ ๐๐ผ๐ฏ ๐๐ถ๐๐ฒ๐ ๐ถ๐ป ๐๐ ๐ฐ๐ฒ๐น, ๐๐ ๐๐ ๐๐ผ๐บ๐ถ๐ป๐ด ๐ณ๐ผ๐ฟ ๐๐!
๐ ๐๐จ๐ง๐ญ๐๐ฑ๐ญ:
@AnthropicAI released ๐๐น๐ฎ๐ฑ๐๐ฒ ๐ณ๐ผ๐ฟ ๐๐ ๐ฐ๐ฒ๐น in Oct 2025 and updated it on Feb 24 with an emphasis on Financial modeling. It can help in building models, Forecasting, Scenario analysis, Model Audit, and validation. It can also perform data cleaning, Data analysis (e.g., insights and trends), formatting, dashboard development, and reporting, etc. ๐ช๐ผ๐ฟ๐ฑ ๐ผ๐ป ๐๐ต๐ฒ ๐๐๐ฟ๐ฒ๐ฒ๐ ๐ถ๐ ๐๐ต๐ฎ๐ ๐ถ๐'๐ ๐ฝ๐ฟ๐ฒ๐๐๐ ๐ด๐ผ๐ผ๐ฑ.
Yesterday (March 5), @OpenAI released ChatGPT for Excel to help with "real-world finance workflows that often take analysts hours or days to complete, including financial modeling, scenario analysis, data extraction, and long-form research."
To test:
โ I downloaded both Excel Plugins (Claude Opus 4.6, ChatGPT "Heavy")
โ gave them a store sales data - 10,195 rows, 25 columns with clearly defined headers (OrderID, Customer, Product, Sales, Profit, etc). FYI, this is data from Tableau.
โ and gave this simple prompt: โAnalyze this spreadsheet and find meaningful trends. Plot them in a dashboardโ
You can see both outputs in the attached pics.
โ๏ธ Claude used Python code in the background to perform the analysis and presented the findings quickly. I actually liked it! It will take me quite a while to do all this by myself.
โ๏ธ ChatGPT took 9 minutes and 46 secs, fumbled a bit with Excel formulas, retracted its steps, and published its findings. Its insights were not as extensive/mature as Claude's, but it did a decent job. @sama This was just the first beta release - I am sure it will get better over time.
๐ฐ ๐๐ก๐ฒ ๐๐จ๐๐ฌ ๐ข๐ญ ๐ฆ๐๐ญ๐ญ๐๐ซ?
Just like software engineering roles, I see the โdata manipulation with Excelโ roles will be affected as well, starting with financial modeling roles.
If your job involves working with Excel, I highly encourage you to try these plugins (you need to be on the paid versions of both). Ask some open-ended questions, and these tools may just surprise you! Believe both OpenAI and Anthropic are offering higher usage quotas for the next few days (until March 19 for Cladue).
โ ๏ธ Btw, these are both beta products, and they can/will make mistakes. Please do verify their findings.
#Excel #ChatGPT #Claude
@goodside I wonder why you are surprised. The auto-regressive LLM will write out garbage when prompted this way. This is expected, and this is how these models work. It has been a few years since GPT came out and understanding how these models work is way past overdue for many.
Claude 2.1 (200K Tokens) - Pressure Testing Long Context Recall
We all love increasing context lengths - but what's performance like?
Anthropic reached out with early access to Claude 2.1 so I repeated the โneedle in a haystackโ analysis I did on GPT-4
Here's what I found:
Findings:
* At 200K tokens (nearly 470 pages), Claude 2.1 was able to recall facts at some document depths
* Facts at the very top and very bottom of the document were recalled with nearly 100% accuracy
* Facts positioned at the top of the document were recalled with less performance than the bottom (similar to GPT-4)
* Starting at ~90K tokens, performance of recall at the bottom of the document started to get increasingly worse
* Performance at low context lengths was not guaranteed
So what:
* Prompting Engineering Matters - Itโs worth tinkering with your prompt and running A/B tests to measure retrieval accuracy
* No Guarantees - Your facts are not guaranteed to be retrieved. Donโt bake the assumption they will into your applications
* Less context = more accuracy - This is well know, but when possible reduce the amount of context you send to the models to increase its ability to recall
* Position Matters - Also well know, but facts placed at the very beginning and 2nd half of the document seem to be recalled better
Why run this test?:
* Iโm a big fan of Anthropic! They are helping to push the bounds on LLM performance and creating powerful tools for the world
* As a practitioner of LLMs, itโs important to build an intuition for how they work, where they excel and their limits
* Tests like these, while not bulletproof, help showcase real world examples and get a feeling for how they work. The goal is to transfer this knowledge to productive use cases
Overview of the process:
* Use Paul Graham essays as โbackgroundโ tokens. With 218 essays itโs easy to get up to 200K tokens (repeated essays when necessary)
* Place a random statement within the document at various depths. Fact used: โThe best thing to do in San Francisco is eat a sandwich and sit in Dolores Park on a sunny day.โ
* Ask Claude 2.1 to answer this question only using the context provided
* Evaluate Claude 2.1s answer with GPT-4 using @langchain evals
* Rinse and repeat for 35x document depths between 0% (top of document) and 100% (bottom of document) (sigmoid distribution) and 35x context lengths (1K Tokens > 200K Tokens)
Next Steps To Take This Further:
* For rigor, one should do a key:value retrieval step. However for relatability I did a San Francisco line within PGs essays for clarity and practical relevance
* Repeat test multiple times for increased statistical significance
Notes:
* Amount Of Recall Matters - The model's performance is hypothesized to diminish when tasked with multiple fact retrievals or when engaging in synthetic reasoning steps
* Changing your prompt, question, fact to be retrieved and background context will impact performance
* The Anthropic team reached out and offered credits to repeat this test. They also offered prompt advice to maximize performance. It's important to clarify that their involvement was strictly logistical. The integrity and independence of the results were maintained, ensuring that the findings reflect my unbiased evaluation and are not influenced by their support.
* This test cost ~$1,016 for API calls ($8 per million tokens)
A nice post about the various techniques used to get a training job to scale to more than 50,000 TPU v5e chips on @googlecloud.
https://t.co/kbouylQYQ1
Large language models have demonstrated a surprising range of skills and behaviors. How can we trace their source? In our new paper, we use influence functions to find training examples that contribute to a given model output.
@karpathy Iol, can relate.
My prompt: "What does Generative AI mean for the enterprise"
Me: 4 weeks and 450 slides later .. it's going to be a full day coaching session :)
Are you overwhelmed by everything happening in the ML ecosystem?
We're doing a small crowdsourced initiative with a high-level distillation+timeline of cool big things happening in the ML landscape. Feel free to contribute! ๐ค
https://t.co/cWzMo9IZLh
My personal identity was hacked last week. The attacker was able to steal $100k+ in a sweep of my Coinbase account. I'm equal parts embarrassed, hurt, and deeply remorseful.
In an effort to raise awareness about the attack, I wrote about it here: https://t.co/ZnbB0AN6Gd
A common misconception is that the risk of overfitting increases with the number of parameters in the model. In reality, a single parameter suffices to fit most datasets: https://t.co/4eOGBIyZl9
Implementation available at: https://t.co/xKikc2m0Yf
Are you a deep learning researcher? Wondering if all this TensorFlow 2.0 stuff you heard about is relevant to you?
This thread is a crash course on everything you need to know to use TensorFlow 2.0 + Keras for deep learning research. Read on!