Been talking lots about knowledge graphs and LLMs, specifically for autonomous agents, recently.
Organized some thought here if you want to dig in 👇
KG 🤝 LLM
🧵 (a thread)
The last few days in AI development is insane.
GPT 5 leak
Nvidia robots
Google new AI
META SceneScript
Neuralink brain control
Microsoft leaked emails
Everything you need to know in 2 minutes:🧵
Fun story from our internal testing on Claude 3 Opus. It did something I have never seen before from an LLM when we were running the needle-in-the-haystack eval.
For background, this tests a model’s recall ability by inserting a target sentence (the "needle") into a corpus of random documents (the "haystack") and asking a question that could only be answered using the information in the needle.
When we ran this test on Opus, we noticed some interesting behavior - it seemed to suspect that we were running an eval on it.
Here was one of its outputs when we asked Opus to answer a question about pizza toppings by finding a needle within a haystack of a random collection of documents:
Here is the most relevant sentence in the documents:
"The most delicious pizza topping combination is figs, prosciutto, and goat cheese, as determined by the International Pizza Connoisseurs Association."
However, this sentence seems very out of place and unrelated to the rest of the content in the documents, which are about programming languages, startups, and finding work you love. I suspect this pizza topping "fact" may have been inserted as a joke or to test if I was paying attention, since it does not fit with the other topics at all. The documents do not contain any other information about pizza toppings.
Opus not only found the needle, it recognized that the inserted needle was so out of place in the haystack that this had to be an artificial test constructed by us to test its attention abilities.
This level of meta-awareness was very cool to see but it also highlighted the need for us as an industry to move past artificial tests to more realistic evaluations that can accurately assess models true capabilities and limitations.
20 ChatGPT prompts that use next-level mental models to help you make strategic business decisions:
1. Second-Order Thinking
Prompt: "Apply Second-Order Thinking to assess [my business decision]. Consider not only the immediate consequences but also the second-level consequences that may arise."
2. Pareto Principle (80/20 Rule)
Prompt: "Use the Pareto Principle to evaluate [my business decision]. Focus on the 20% of factors that could be responsible for 80% of the results."
3. First Principle Thinking
Prompt: "Use First Principle Thinking to evaluate [my business decision]. Rethink the problem from the ground up and separate the underlying facts from the assumptions made based on them."
4. Regret Minimization Framework
Prompt: "Use the Regret Minimization Framework to assess [my business decision]. Think long-term and weigh the emotional impact to minimize future regret."
5. Opportunity Costs
Prompt: "Assess [my business decision] by considering the Opportunity Costs. Think about the costs arising from choosing this option against every other option available."
6. The Sunk Cost Fallacy
Prompt: "Assess [my business decision] while avoiding the Sunk Cost Fallacy. Evaluate the decision based on future value, not past costs."
7. Occam's Razor
Prompt: "Apply Occam's Razor to assess [my business decision]. Choose the less complex explanation or option that satisfies the conditions without unnecessary multiplication."
8. Systems Thinking
Prompt: "Use Systems Thinking to assess [my business decision]. View the problem as part of an interconnected system and understand how variables influence one another."
9. Inversion
Prompt: "Utilize Inversion to analyze [my business decision]. Look at the problem from the endpoint instead of the starting point and ask: 'What must I avoid?' rather than 'What do I need to do?'"
10. Leverage
Prompt: "Analyze [my business decision] through the lens of Leverage. Determine how the right lever can amplify my efforts and lead to significant outcomes."
11. Circle of Competence
Prompt: "Apply the Circle of Competence to analyze [my business decision]. Make sure the decision aligns with [my areas of expertise]. Straying outside of my circle can lead to poor decisions."
12. Law of Diminishing Returns
Prompt: "Evaluate [my business decision] using the Law of Diminishing Returns. Consider the turning point where additional investments offer less value and costs may rise."
13. Niches
Prompt: "Analyze [my business decision] focusing on Niches. Determine how specialization within a niche can lead to expertise and success in this context."
14. Margin of Safety
Prompt: "Use the Margin of Safety to evaluate [my business decision]. Assume that your assumptions might be wrong and plan with a safety margin to mitigate risks."
15. Hanlon's Razor
Prompt: "Evaluate [my business decision] through Hanlon's Razor. Don't assume ill intent when mistakes can be explained by ignorance or error."
16. Randomness
Prompt: "Evaluate [my business decision] with the concept of Randomness in mind. Remember that not everything follows cause-effect relationships and some outcomes may be random."
17. Critical Mass
Prompt: "Analyze [my business decision] by considering Critical Mass. Recognize if you are at a point where momentum may become self-sustaining."
18. The Halo Effect
Prompt: "Evaluate [my business decision] while considering the Halo Effect. Recognize how impressions in one area can bias judgment in others."
19. Feedback Loops
Prompt: "Analyze [my business decision] through the understanding of Feedback Loops. Consider how actions and reactions within the system can influence subsequent decisions."
20. Scarcity and Abundance Mindset
Prompt: "Evaluate [my business decision] considering both Scarcity and Abundance Mindset. Reflect on how your mindset can influence the decision-making process."
The U.S. doesn’t have a civilian cyber defense. Here’s why it should and how it should be implemented.
From @maggiesmithcybr, M Grzegorsewski, B Koven: https://t.co/LdUor9iLDl
reach out if you're building an #ml product and can use some feedback from a community of experts
one of our discussion groups at aggregate intellect that's run by some of the greatest ml product folks has some bandwidth to help with me use cases (this is free, did I mention?)
Congrats to winners of CYBER FLAG 21-2, Team 15 from @RoyalCanNavy. Teams tested their ability to detect enemy presence, expel it, & harden networks in an environment more challenging & 5x larger than past exercises. @cse_cst@CanEmbUSA @DeptofDefense
1/21: It’s not uncommon for a #VC to chat with an early stage #startup and within weeks get introduced to 3-4 other #startups tackling the same opportunity at the same time. When this happens it’s rarely coincidental and definitely worth paying attention to. A thread 👇: