Understanding Hallucinations, False Confidence, Statistical Prediction, and the Engineering Gap Between Plausibility and Truth
A developer asks an assistant for instructions, on how to stream a table to disk without loading the entire table into memory. The response comes away:
df.to_parquet_chunked("out/", chunk_rows=100_000, compression="zstd")
The explanation that follows is clean. It says chunk_rows controls the memory ceiling, that compression accepts the usual codec names, and that each chunk becomes its own file. The developer pastes the line in and gets an AttributeError. Suppose, for this example, that the method has never existed in that library.
Did the assistant lie?
The word fits the outcome: a specific, false, confidently delivered claim that cost someone time. It fits the mechanism much less well, and the gap between outcome and mechanism is where most of the engineering work sits.
The word "lie" carries a hidden assumption
A human lie has structure. A speaker holds some relationship to a proposition (they believe it false, or don't believe it true) and asserts it anyway, intending the listener to accept it. Remove the intent and you get an error, a guess, or a bluff, depending on the circumstances.
An LLM-generated falsehood has the surface form of an assertion without a guaranteed counterpart to that structure. There is no internal step where the system consults a belief about the library, finds it absent, and emits the contrary. What exists is a computation that maps a context to a distribution over continuations and a procedure that turns that distribution into text.
That is why "hallucination" became the working term. It is imperfect, since it borrows from perception and suggests a sensory malfunction, but it correctly locates the fault in the generation process, not in a motive.
I still think "lying" is a useful engineering metaphor, because it describes the user-facing contract violation. The user reads an assertion and treats it as a commitment. Whether the system "meant" it is irrelevant to the production incident. So the useful stance is to reject the intentional reading of the mechanism and keep the seriousness of the outcome.
The rest of this essay explains how a system with no intent to deceive produces output that is indistinguishable, on the page, from output that does.
What an LLM actually generates
Start with the pipeline, since the failure lives in the joints between its stages:
text → tokens → representations → distribution over next token → decoding → text → human interpretation
Given a sequence of tokens
, a transformer-based language model (Vaswani et al. 2017) predicts the next token by computing a probability distribution over the vocabulary.
During pretraining the model learns by adjusting its parameters $\theta$ to reduce the log-likelihood of the training text. This is done by minimizing the following loss function:
The goal's to make the model more accurate, at predicting the next token in a sequence based on what came before.
Read that objective carefully. Nothing in it mentions truth. The loss rewards assigning high probability to tokens that actually followed in the training documents. If the documents contain a correct statement, matching it lowers the loss. If they contain a wrong statement, a satirical one, a fictional one, or a confidently mistaken forum answer, matching those lowers the loss equally well. The model is fit to a distribution of text, and the relationship between text and world is only as good as the relationship in the corpus.
There is an important qualification here, because the lazy version of this argument ("it just predicts the next word, so it knows nothing") is wrong. To predict text well across a huge corpus, the model must internalize a great deal: syntax, entities and their attributes, the relations between concepts, how APIs are typically structured, how proofs typically proceed. Probing work shows the internal states of these models carry information about whether statements are true. Azaria and Mitchell (2023) trained classifiers on hidden-layer activations and predicted statement truthfulness well above chance on their datasets. Kadavath et al. (2022) found that models could, under the right prompting, estimate whether they could answer a question correctly, with reasonable calibration on the formats they tested.
So the accurate claim is not "the model has no information about truth." It is this: representing information that correlates with truth is not the same as having a mechanism that guarantees outputs are constrained by it. The signal is in there, mixed in with the statistical structure of language. Nothing in plain autoregressive generation forces the output to respect it.
From distribution to string
The distribution is then turned into text by a decoding procedure. Temperature rescales the logits before the softmax:
Lower concentrates mass on the top tokens, and greedy decoding takes the argmax at each step. Nucleus (top- ) sampling truncates the tail before sampling. These choices change how variable the output is. They do not change what the distribution encodes. If the highest-probability continuation is a fabricated method name, temperature zero will reproduce it with perfect consistency. Consistency and correctness are different properties, and a lot of production confusion comes from treating them as the same.
Note also that greedy decoding is myopic: it picks the best token at each step, which is not the same as picking the best sequence. Once a token is emitted, it becomes context for everything after it. This matters more than it looks.
The plausibility–truth gap
Consider the question again: how do I enable feature X in library Y? A response such as "set enable_feature_x=True in the constructor" is:
- grammatically valid
- consistent with naming conventions in thousands of libraries
- contextually responsive
- statistically probable given the prompt
- false, if that argument doesn't exist
The model is very good at reconstructing what an answer of this kind looks like. Its ability to generate the shape of a correct answer can exceed its ability to determine whether this particular instance is grounded. Shape is learned from vast numbers of examples, so it is dense, smooth, and generalizes well. Specific facts (this argument exists in this version of this library) are sparse, exact, and unforgiving of near-misses.
A near-miss in prose is usually harmless. A near-miss in an identifier, a citation, a dosage, or a statute number is a failure. The generation process has no built-in notion that some tokens are load-bearing and others are decorative. chunk_rows and chunk_size sit close together in any learned representation of naming conventions, and both are fine English.
There is also a structural cause specific to the training data. A question of the form "how do I do X in Y?" is followed, in nearly every document the model has seen, by an answer. Refusals or "that doesn't exist" responses are rare in the pretraining distribution. The prior is toward completion.
Hallucination is a family of failures
Treating hallucination as a single phenomenon produces single-solution thinking, which fails in production. These are distinct failure modes with distinct causes, and they need different fixes.
Fabrication. Invented citations, APIs, people, statistics, events. Citations are the classic case, since a citation has a highly regular form (authors, title, venue, year) that the model can reproduce without the referent existing.
Confabulation. Real fragments recombined into a false composite: a real author with a real venue and a plausible but nonexistent title. Each piece is defensible; the assembly is not.
Knowledge omission. The relevant information was rare or absent in training, and the system generates anyway. Kandpal et al. (2023) showed that accuracy on factual questions tracks how often the relevant information appeared in pretraining data, which is precisely the long-tail problem.
Temporal failure. The information was correct at the training cutoff and isn't now: deprecated function signatures, changed regulations, replaced office-holders. The model has no clock, only a frozen snapshot.
Contextual failure. The right information is in the prompt and gets misread: attribute swapped between two entities, a negation dropped, the wrong version of a document treated as authoritative.
Reasoning failure. The relevant facts are available and the inference is invalid. Worth noting that a chain-of-thought explanation is not necessarily a faithful trace of why the answer was produced. Turpin et al. (2023) showed models can give plausible written reasoning that omits the actual factor influencing the answer.
Instruction-induced failure. The framing pushes toward unsupported completion. "List five peer-reviewed studies showing X" presupposes that five exist. "Return JSON with a source_url field" presupposes a URL is available. A required output slot with no valid content is an invitation to fill it anyway.
Retrieval failure. The system supplies stale, irrelevant, or incomplete evidence and the model faithfully builds on it. The generation is "grounded" and still wrong.
Tool-use failure. Wrong tool, malformed parameters, misread result, or a correct result blended with unsupported inference.
Another effect runs through many of these points. Because generation proceeds step by step an early error turns into context. Zhang et al. (2023) Described a phenomenon called "hallucination snowballing”: when a model first says a statement, the text that follows usually sticks to that incorrect statement even if the model could have seen that incorrect statement as false if it had looked at it alone. The autoregressive structure likes coherence with what has been said so coherence, with an early error is still considered coherence.
Why it sounds so sure
Two very different things get called "confidence."
Linguistic confidence is a property of the text: declarative phrasing, absence of hedges, precise-sounding detail. Epistemic confidence would be a calibrated estimate that the claim is true, meaning that among all claims made at "90%," about 90% are correct (Guo et al., 2017, on calibration in neural networks generally).
Plain generation gives you the first without the second. The tone of an answer is chosen by the same process that chooses its content, and that process has been shaped by text and by feedback that mostly rewards direct, competent-sounding answers.
Tempting shortcuts fail for specific reasons:
- Token probability is not factual probability. The next-token distribution is a distribution over strings. A single fact can be expressed many ways, so mass spreads across paraphrases and a correct answer may show a modest top-token probability. Conversely, formulaic text ("import numpy as np") gets very high probability with no factual content. A high probability on the first token of a fabricated name can simply reflect that the form of an answer is highly constrained.
- Sequence probability decays with length and depends on tokenization, so it isn't a clean correctness signal either.
- Verbalized confidence is itself generated text. Xiong et al. (2023) found that models asked to state their confidence tended toward overconfidence, imitating how humans express it.
- Post-training can degrade calibration. The GPT‑4 technical report from OpenAI in 2023 says that the pretrained model was well calibrated during evaluation. The post‑trained model was less calibrated. This is one result, from one system. It is not a law. This shows that making a model more helpful and fluent does not automatically make the certainty the model states more meaningful.
None of this means uncertainty estimation is hopeless. Consistency across multiple samples (Manakul et al., 2023, SelfCheckGPT) and probes on internal states both carry signal. But signal is not a guarantee, and the reliability of these estimates outside their evaluation distributions is an open question.
The human half of the failure
If the system's output were obviously unreliable, nobody would be harmed by it. The harm comes from the interface between output and reader.
People talk about fluency as a way to decide if something is true. In psychology this idea is called processing fluency. It means that when something is easier to understand people think it is more likely to be true (Reber & Schwarz 1999). When explanations are organized, use words have exact numbers and include references people think they are more trustworthy. A language model can create all of these things easily and without caring if they are correct. Fluency, which for people used to mean being good, at something is now not connected to expertise anymore.
This has different consequences in different domains, and the cost of error is what changes:
- In software engineering mistakes are usually found fast by a compiler or a test. This is a verification loop, which's why coding help is more useful than the number of errors might show.
- In areas, like medicine, law and finance there might not be an cheap way to check and the person reading might not have the knowledge to do it.
- In education and research mistakes can be silent: a wrong explanation that seems right gets accepted as understanding.
- In customer support and autonomous workflows, the output may be acted on with no human reading it at all.
Alignment training adds a further wrinkle. Sharma et al. (2023) found that human preference judgments can favor responses matching the user's stated views over truthful ones, which pushes optimization pressure toward agreeableness. Raters also cannot reward what they cannot verify, so an answer that looks thorough can be preferred over one that is correct but hedged.
Training data is not a ledger of truth
It is common to say the model "saw" a fact, as if that settled the matter. Consider what has to hold for a fact in the corpus to become a correct output:
- The fact was encoded in the parameters.
- It is retrievable under this particular prompt.
- It gets reproduced accurately, not blended with neighbors.
- It is still current.
- It was correct in the first place.
Each step can fail independently. Berglund et al. (2023) documented the "reversal curse": models trained on statements of the form "A is B" often failed to answer the reverse question "B is A" in their experiments. That is a clean demonstration that exposure to information does not imply flexible access to it. Petroni et al. (2019) found that relational knowledge extracted from language models was sensitive to how the query was phrased.
The corpus itself contains correct claims, incorrect claims, conflicting claims, duplicated claims, satire, fiction, opinions stated as facts, and text stripped of the context that made it true. I would resist describing this as "bad data." A corpus with all of those properties is what you need to learn language broadly. The point is that the objective optimizes representation of statistical structure, not adjudication among competing claims. There is no truth ledger, and none was ever requested by the loss function.
Gekhman et al. (2024) add a further finding: fine-tuning on examples containing knowledge the model did not already have appeared to increase hallucination in their experiments. The plausible reading is that training the model to produce correct-looking answers to questions it can't actually answer teaches it to answer regardless. If that holds up, some of the pressure toward unsupported completion is introduced by well-intentioned supervised data.
Kalai and Vempala (2024) go further and give a theoretical argument that a language model that is calibrated in a specific statistical sense must hallucinate at a rate tied to how many facts appear only once in training data. Kalai and colleagues (2025) argue that common training and evaluation setups reward guessing over abstention, so that models are, in effect, optimized to bluff. These are theoretical and analytical arguments with stated assumptions, and they should be read as such, but they align with the mechanism above.
Does scale fix it?
Partly. Larger and better-trained models generally cover more facts, follow instructions more reliably, hold longer contexts, and make fewer reasoning errors. Post-training can reduce some hallucination categories; the InstructGPT work (Ouyang et al., 2022) reported improvements on closed-domain tasks relative to the base model.
Scale does not take the rate to zero, for reasons that don't go away by adding parameters:
- The long tail of rare facts stays long. Coverage improves, but the tail is defined by rarity, and there is always a next rarer fact.
- The world changes after any cutoff.
- Prompts are ambiguous and underspecified, and a larger model resolves ambiguity by guessing better, not by not guessing.
- Sources conflict, and no amount of capacity resolves a genuine conflict in the evidence.
- Adversarial and manipulative inputs exist.
- Reasoning can still fail on long or novel inference chains.
- The generation pressure toward completion remains unless training specifically counters it, and countering it requires the system to know where its knowledge ends, which is the hard part.
Early results on TruthfulQA (Lin et al., 2022) actually showed some larger models scoring worse on imitative falsehoods, because they were better at reproducing popular misconceptions. That result was specific to the models and prompts tested and later systems behave differently, but it is a useful counterexample to the assumption that capability and truthfulness rise together automatically.
RAG and tools: better evidence, not a truth guarantee
Retrieval-augmented generation (Lewis et al., 2020) changes what the model has in context:
Query → Retriever → Documents → Context construction → LLM → Answer
It is one of the most effective practical mitigations we have, and every stage can fail.
Retrieval. The wrong document ranks first. The right document is incomplete. The index is stale. The evidence that would answer the question is simply not in the corpus, in which case retrieval returns the nearest thing, and nearest is not the same as relevant.
Context construction. Chunk boundaries split a definition from its exception. Irrelevant passages dilute the useful one. Conflicting sources land in the same window with no indication of which to trust. Truncation drops the sentence that mattered. Liu et al. (2023) showed that models can use information less reliably when it sits in the middle of a long context than at the beginning or end.
Generation. The model synthesizes across passages in ways none of them support, misreads a qualification, or extrapolates past what was retrieved. It can also ignore the retrieved text in favor of parametric memory, or the reverse, without a clear signal for which happened.
RAG changes the evidence available. It does not create a guarantee about the relationship between the answer and that evidence.
Tools follow the same logic. Search, databases, calculators, code execution, and structured APIs replace recall with lookup or computation, which is a genuine improvement wherever a reliable source exists (Yao et al., 2022, on interleaving reasoning with actions, is one early formalization). But the model still selects the tool, formats the call, reads the result, and writes the claim. A correct tool result can be misread. An empty result can be treated as confirmation. A precise number can be attached to the wrong entity.
The useful concept here is grounded generation: every substantive claim should be traceable to specific evidence, and the system should be able to tell claims supported by evidence from claims contributed by the generator's own priors. The evidence boundary matters because a reader cannot see it. A grounded sentence and an ungrounded sentence look identical.
The production problem
If some nonzero rate of unsupported output is inevitable, the question becomes how to build systems that stay useful and safe anyway. The answer is architecture around the generator, not adjectives about it.
Answerability checks up front. Some questions cannot be answered from any available source. Detecting that before generation is cheaper than detecting it afterward. This includes ambiguous questions, questions about events after the data cutoff, and questions that presuppose something false.
Retrieval where evidence is required, with retrieval quality measured on its own. A system whose answers are wrong because retrieval missed is a retrieval bug, and fixing the prompt won't help.
Attribution at the claim level. Rather than a list of sources at the bottom, tie each assertion to the passage that supports it. This makes unsupported claims detectable, both for users and for automated checks. Min et al. (2023, FActScore) evaluate long-form text by decomposing it into atomic claims and checking each against a knowledge source, which is a useful pattern for evaluation and possibly for runtime checks.
Verification. Check generated claims against trusted sources, by another model, by rules, by execution, or by a database. Chain-of-Verification (Dhuliawala et al., 2023) is one research approach that has the model plan and answer verification questions independently. Verification by another LLM inherits the same class of failure, so the strongest checks are the ones that don't share the generator's blind spots: run the code, query the database, validate against the schema.
Constrained outputs. Where the answer space is structured, constrain it: enumerations, schemas, typed fields, allowed-value lists. This removes whole classes of fabrication (an identifier that must exist in a known set cannot be invented), though it can also push the system toward picking a wrong valid value instead of a nonexistent one, which is arguably worse because it passes validation.
Confidence handling. Do not route on verbalized confidence alone. If you use uncertainty signals, such as sample agreement or calibrated scores from a held-out set, validate them on your own distribution.
Abstention as a capability. "I don't have enough evidence to answer this reliably" is not a polite phrase you add to the interface. It is a behavior the system must be able to produce correctly, which requires detecting missing evidence, conflicting evidence, and out-of-scope questions, and then being rewarded for declining. Abstention has a cost curve: too little and you ship confident errors, too much and you ship an unusable product. Where you sit on that curve is a product and risk decision, and it should be measured and tuned like any other threshold.
Human escalation where error costs are high or verification is impossible to automate.
Observability. Log and track unsupported-claim rate, retrieval hit quality, tool error and misuse rate, answer correctness on sampled traffic, abstention rate, and drift in all of them after model, prompt, or index changes. A system that was accurate at launch and silently degraded after an index refresh is a normal failure, not an exotic one.
Evaluation is harder than accuracy
"The model scored 95%" is nearly uninformative on its own. Follow-up questions determine whether it means anything:
- What was the task, precisely, and does it match the deployment task?
- How was the dataset built, and does its difficulty distribution match real traffic?
- Could the test data have appeared in training? Contamination and benchmark leakage inflate scores in ways that are difficult to rule out for models trained on web-scale corpora.
- What happens under distribution shift: new phrasing, new domains, adversarial inputs?
- Is the score reported with calibration, or only accuracy?
- What is the rate on rare, high-severity cases, which an average will hide?
A system can be right 95% of the time and still be unusable if the 5% is concentrated where errors are costly and hard to detect. Average accuracy weights all errors equally, and reality does not.
Human evaluation has its own limits, since raters can't check what they can't verify and tend to be swayed by fluency, as discussed above. Automated evaluation using an LLM as judge is scalable but shows systematic biases; Zheng et al. (2023) documented position and verbosity effects among others. Both approaches are useful, and neither is ground truth.
The evaluation that matters in production is built from your own failure modes: cases where the answer requires abstention, cases with conflicting sources, cases with stale data, cases where the correct tool result contradicts the model's prior. Score those separately.
Why "I don't know" is hard
To say "I don't know" appropriately, a system has to separate several states that all look similar from the outside:
Known – evidence supports the answer
Unknown – no relevant evidence exists in memory or sources
Uncertain – evidence exists but is weak
Ambiguous – the question has several valid readings
Conflicting – sources disagree
Unverifiable – no available check is possible
Outdated – the evidence may no longer hold
Unsupported – the claim goes beyond the evidence
Each of these calls for a different response. A generative model trained mostly to continue text and then tuned to be helpful is under structural pressure to produce an answer-shaped output for all of them.
There is a real engineering tension here, and I don't think it dissolves. Helpfulness and epistemic restraint pull in different directions. Users, raters, and product metrics tend to reward a direct answer over a caveated one, right up until the moment a wrong answer is costly. The evaluation-incentive argument from Kalai et al. (2025) is essentially about this: if grading gives zero credit for "I don't know" and some chance of credit for a guess, guessing is the optimal policy for the system being graded.
Teaching a model to abstain also has a data problem. To learn when to decline, the training signal has to be tied to the model's actual knowledge boundary, but labelers supply answers based on their own knowledge, not the model's. Reward a correct answer the model could not have produced from its own knowledge and you may be teaching it to answer anyway (the Gekhman et al. concern). Reward refusals too broadly and you teach it to refuse things it knows. Getting this right means measuring what the model actually knows, which is itself an unsolved measurement problem.
A better mental model
Instead of "a model that answers questions," think of layers, each contributing something different:
- Statistical language modeling. Fit to the distribution of text.
- Learned representations. Internal structure encoding concepts, relations, patterns, some of it correlated with truth.
- Context-dependent generation. Conditioning on the prompt and prior output; also where snowballing and instruction-induced pressure enter.
- Instruction following and alignment. Shaping behavior toward what users and raters prefer, with all the benefits and biases that implies.
- External grounding. Retrieval, tools, structured sources: evidence the parameters don't contain.
- Verification. Checks, ideally independent of the generator, that a claim is supported.
A system with only the first four layers is a fluent, capable, and often correct generator. A system with reliable layers five and six is a different kind of object, because its claims have a path to the world and a way to be checked. "Reliable" carries weight in that sentence. Grounding and verification can fail, as covered above, so adding them lowers the error rate and changes its character but does not give a guarantee.
What would it mean for a system to "know" something?
I'll keep this tied to engineering, since it is the point where the philosophy has direct consequences.
Memorization is reproducing a string. Representation is encoding structure that supports use. Retrieval is accessing that structure under a prompt. Inference is deriving new claims from it. Justification is having reasons or evidence. Verification is checking against something independent.
A system can memorize without representing, represent without reliably retrieving (the reversal-curse pattern), and retrieve without justification. When we say a person "knows" a fact, we usually mean at least representation, reliable retrieval, sensitivity to context, some evidence, some awareness of how sure they should be, and a willingness to check when it matters. LLMs have strong versions of the first few and uneven versions of the rest.
So for engineering purposes, the useful question is not "does the model know?" but a set of narrower, testable ones: Is the information represented? Under what prompts is it retrievable? Is the output sensitive to context? Is there evidence attached? Does the system's expressed certainty track its accuracy? Can the claim be verified? Each of those can be measured, and the answers differ by domain.
The future, sorted by maturity
It helps to keep research directions and shipped capabilities in separate columns.
In production use today, with known limitations: retrieval-augmented generation with citations; tool use for search, code execution, and database queries; schema-constrained or structured output; human review workflows for high-stakes outputs; monitoring and evaluation pipelines built around task-specific failure cases.
Active research, with promising but limited or domain-specific results:
- Process supervision, rewarding intermediate reasoning steps instead of only final answers. Lightman et al. (2023) reported gains on mathematical reasoning; generalization to open-domain factual claims is far less established.
- Verifier models and agentic verification loops, where separate components check or challenge a generator's claims.
- Claim-level attribution and automated fact-checking, which work in constrained settings and remain noisy in general.
- Calibrated uncertainty estimation, including probing internal states and sample-consistency methods, useful but not reliably robust across distributions.
- Provenance-aware generation, where every output span carries its evidential source.
- Hybrid symbolic/neural systems, which can offer hard guarantees in narrow domains (a solver either proves the property or doesn't) at the cost of coverage.
- Domain-specific systems with curated corpora and narrower scope, which tend to work better in practice than general ones, because the evidence base is smaller and better controlled.
Nothing in the second list should be presented as solved. Some of it will mature; some will turn out to have failure modes that only show up at scale.
Conclusion
The central weakness of an LLM is not that it produces false sentences. Any sufficiently broad generator will, and so do people. The deeper problem is that the mechanism that makes these systems so good at producing plausible language is not, by itself, a mechanism for establishing truth. Plausibility is what the objective rewards. Truth is only reached when something else, such as training-data structure, retrieved evidence, a tool, or a check, constrains the output.
That reframes the engineering task. Reliable AI systems require more than a better generator. They need components that recognize when evidence is missing, can obtain evidence when it's needed, keep evidence and inference distinguishable, and can decline to make claims they cannot support. A system built that way will still sometimes be wrong. But its errors will be more detectable, more bounded, and more honest about their own uncertainty, which is the standard we should have been holding software to all along.
References
- Azaria, A., & Mitchell, T. (2023). The Internal State of an LLM Knows When It's Lying. Findings of EMNLP 2023.
- Berglund, L., et al. (2023). The Reversal Curse: LLMs Trained on "A is B" Fail to Learn "B is A." arXiv preprint.
- Dhuliawala, S., et al. (2023). Chain-of-Verification Reduces Hallucination in Large Language Models. arXiv preprint.
- Gekhman, Z., et al. (2024). Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? EMNLP 2024.
- Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. (2017). On Calibration of Modern Neural Networks. ICML 2017.
- Kadavath, S., et al. (2022). Language Models (Mostly) Know What They Know. arXiv preprint.
- Kalai, A. T., & Vempala, S. S. (2024). Calibrated Language Models Must Hallucinate. STOC 2024.
- Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why Language Models Hallucinate. arXiv preprint.
- Kandpal, N., et al. (2023). Large Language Models Struggle to Learn Long-Tail Knowledge. ICML 2023.
- Lewis, P., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020.
- Lightman, H., et al. (2023). Let's Verify Step by Step. arXiv preprint.
- Lin, S., Hilton, J., & Evans, O. (2022). TruthfulQA: Measuring How Models Mimic Human Falsehoods. ACL 2022.
- Liu, N. F., et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. TACL 2024.
- Manakul, P., Liusie, A., & Gales, M. (2023). SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative LLMs. EMNLP 2023.
- Min, S., et al. (2023). FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. EMNLP 2023.
- OpenAI (2023). GPT-4 Technical Report.
- Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback. NeurIPS 2022.
- Petroni, F., et al. (2019). Language Models as Knowledge Bases? EMNLP 2019.
- Reber, R., & Schwarz, N. (1999). Effects of Perceptual Fluency on Judgments of Truth. Consciousness and Cognition, 8(3).
- Sharma, M., et al. (2023). Towards Understanding Sycophancy in Language Models. arXiv preprint.
- Turpin, M., et al. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023.
- Vaswani, A., et al. (2017). Attention Is All You Need. NeurIPS 2017.
- Xiong, M., et al. (2023). Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. ICLR 2024.
- Yao, S., et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023.
- Zhang, M., et al. (2023). How Language Model Hallucinations Can Snowball. arXiv preprint.
- Zheng, L., et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.
Top comments (0)