Your palms are sweaty, knees are weak, arms are heavy – it’s not a rap battle but a test for a class which you may or may not have slept through for the entire semester. The test is multiple choice – do you guess or leave the answer blank? There’s no negative marking, so you guess. There’s a one-in-five chance that the answer is correct, and you take it.
I didn’t mean to lose myself in the traumas of my misspent youth 😰, but recent research by OpenAI on why models hallucinate made me take this rather unwelcome trip back to those days.
So, why do models hallucinate?

Let’s take a step back to think about what is happening under the hood.
The first step is called pre-training, where models are taught to predict what word should follow by training on extremely large amounts of text. The problem starts here: some data is very rare or doesn’t exist in the training corpus. Take the birthday of one of the paper’s authors – the model confidently spits out wrong dates because this fact barely appears in training.
The next step usually involves some sort of reinforcement learning (RL). Here, the model is given labelled data and further trained to become more accurate.
OpenAI claims that this training for accuracy is a factor leading to hallucinations. When models are trained to be more accurate, it makes more sense to guess an answer than to say “🤷🏾♂️- I don’t know.” After all, a slight chance of being correct is better than a zero chance of being correct.
So, let’s bring this together: LLMs are first trained to predict plausible answers and then further trained to optimize for accuracy on rare facts where they have limited training data. So we end up with behaviors where a model will confidently BS instead of saying “I don’t know.”
OpenAI suggests that we could fix this by changing how we evaluate models, giving them explicit confidence targets in prompts like “Only answer if you’re more than X% confident” and scoring uncertainty appropriately. The article (and associated paper) is worth a read!