not exactly, partly because of the training data, training data are highly filtered and selective and even have many QA data in the pretraining stage. Novels may be largely filtered out cause they are not useful to improve benchmarks.
It’s crazy in hindsight that the stylistic mimicry of the GPT-2/3 era was such a local maximum for AI-written texts as cultural artifacts. For all their gains in making sense, aligned LLMs still can’t write with the eerie naturalism GPT-3 exhibited four years ago:
@VictorTaelin@karpathy The most basic problem is hallucinations. LLMs do not know what/how much they know. Only if you know there's a problem, you find solution to it.