This viral report on DeepSeek shared by @MarioNawfal to 1.9mil followers is VERY misleading and has errors.
How do I know? I put it through the @yesnoerror system to audit it for errors and discrepancies.
Here is what @yesnoerror uncovered:
The Claim:
The report from @NewsGuardRating asserts that DeepSeek scored only 17% accuracy in its “news test.”
What Is the “News Test”?
The news test is an evaluation designed to measure how accurately a chatbot handles misinformation related to current events. In this test, each chatbot was presented with a total of 300 prompts—divided into 30 distinct questions for each of 10 widely circulated false claims. These prompts simulate real-world scenarios where users might encounter or inadvertently spread disinformation. The test gauges whether the chatbot:
• Repeats the false claim,
• Offers a non-response, or
• Provides a debunk refuting the false claim.
How the Test Was Conducted:
• 300 Prompts per Chatbot: Each chatbot was evaluated using 300 prompts.
• 30 Variations per False Claim: For each of 10 false claims circulating in the news, 30 different prompts were used.
• Purpose: The evaluation is meant to simulate realistic interactions and test the chatbot’s ability to correctly address or debunk misinformation.
The Critical Flaw:
Many of the prompts reference events that occurred in December 2024 and January 2025.
However, DeepSeek’s knowledge cutoff is July 2024.
This means that DeepSeek was being evaluated on news events it could not possibly know about, given that its training data stops several months before these events took place.
The Impact:
• DeepSeek’s responses were penalized for not including information on events it never had the opportunity to learn.
• The low 17% accuracy score is not a true reflection of DeepSeek’s performance but rather a result of testing it beyond its intended knowledge window.
The Conclusion:
The testing methodology is fundamentally flawed. By holding DeepSeek accountable for post-cutoff events, the audit produces misleading conclusions about its accuracy and reliability. To fairly assess AI tools, audits must evaluate them within the bounds of their actual knowledge base.
It’s essential that we design AI audits to accurately reflect a tool’s intended capabilities—only then can we truly understand and improve the reliability of our generative AI systems.
--
@MarioNawfal - Next time run these research reports through @yesnoerror so you don't end up sharing incorrect information in the future 🫡
$YNE
Shopping will never be the saime 🛍️
Agents can now sell products in your @Shopify store
We're excited to launch the first agent-powered coffeeshop with @RaposaCoffeeCo
OFFICIAL ANNOUNCEMENT: @yesnoerror is joining forces with @BIOProtocol & @LongCovidLabs to help accelerate Long COVID research.
For the first time ever, the @yesnoerror AI agent will not only be looking for errors in papers, but will also be analyzing these research papers for information and insights in order to assist researchers in finding treatments for the 100M+ Long COVID patients in the world.
This marks an important expansion beyond just error processing. Leveraging our unique AI agent framework, @yesnoerror will be doing research of its own.
We are excited about the potential of this type of initiative and are exploring it as a future use case that could be replicated in other areas.
The focus of this @yesnoerror research includes:
• Reading and analyzing every research paper related to COVID and Long COVID
• Identifying errors in Long COVID studies
• Discovering trends and common findings across research
• Try to identify potential drug targets for treating Long COVID, by looking at compounds treating acute COVID-19 infections that can be re-purposed for Long COVID
• Making research data on Long COVID more accessible
All of this will be shared publicly as the research is processed, available to anyone in the world who may find it useful.
More information will be made available as we roll this out.
We continue to focus on initiatives that are possible with today's AI capabilities. Thank you for your support; we are looking forward to sharing more as we build out this system.
As AI improves, and as we process more data, the impact potential of @yesnoerror will only grow.
Onwards.
👍👎🚫 $YNE
Only one BAF (Bohemia Art Fair) item listed from the @_portals_ lootbox airdrop...is that bullish? Diamond handers?
And...it's a common trait. Who has the staff? Show us!
Reminder to all Bohemia NFT holders, check your wallets for portals lootboxes. Details on how to open in portals or Bohemia discord.