Ask a raw language model a question and it does not know what you actually want. It only knows what tends to come next.

That is the first stage of training, and it produces something strange: a system fluent in language but indifferent to being helpful. It can generate ten plausible answers to the same question and has no preference among them. Nothing tells it that one is better than another. It just predicts.

Reinforcement learning from human feedback is the stage where that changes. Human raters read pairs of outputs and rank which one is better, more accurate, more helpful, less likely to ramble or dodge the question. The model gets nudged toward whatever patterns correlate with the answers humans preferred. Nobody is programming politeness or safety directly. It is statistical reinforcement, repeated at scale, until the probability of a good answer rises and the probability of a bad one falls.

The shift is subtle but total. A raw model treats all ten of its plausible answers as roughly equal. After RLHF, the same underlying model has learned which of those ten a person would actually choose, and it starts generating that one first. The architecture has not changed. The training direction has.

This is worth sitting with, because it means the difference between a raw model and something like Claude is not a different brain. It is a second pass of training, aimed entirely at human judgment rather than statistical likelihood.

But steering toward good judgment only works if the model has good material to judge. Claude can be trained to prefer the accurate, well reasoned answer over the sloppy one, and it will still get a client’s trust structure wrong if the data it is reading is fragmented across five systems that do not agree with each other. RLHF fixes how the model chooses among what it knows. It does not fix what the model knows in the first place.

That second problem is the one LEA solves. Feeding a well trained model reliable, reconciled client data is what turns a model’s good judgment into a correct answer, rather than a confident one built on the wrong facts. The training direction matters. So does what you point it at.