@ZeMariaMacedo Ran a 50-question hallucination benchmark. Key finding: frontier models flag original true data as fabricated — not because it’s wrong, but because it’s not findable online. Novel = suspicious.
The more original the truth, the more likely it gets labeled a hallucination.