Biased Confidence of Large Language Models and the Problem of Insufficient Knowledge in Social Welfare Law
Selective Overconfidence: How Large Language Models Underserve Social Welfare Law Across Jurisdictions
GPT Abstract Summary
This study proposes a knowledge probing methodology to assess whether large language models (LLMs) show differences in confidence of knowledge across diverse legal fields. Experiments on 1, 200 judgments from 3 judicial jurisdictions in Germany, France, and Canada found that in non-English speaking regions, the Llama 3.3 70B model showed a tendency to express less confidence deficiency in social welfare law than in tax law. For instance, among German judgments, 49.8% of tax law questions acknowledged uncertainty, whereas only 6.8% did so in social welfare law. This indicates that the model exhibits excessive confidence regarding social welfare law, suggesting a potential qualitative imbalance in legal information services.
Key Points
- Quantification of LLM confidence levels across legal domains via a novel knowledge probing pipeline
- Analysis of judgments from Germany, France, and Canada confirmed that expressions of confidence deficiency in social welfare law are relatively less frequent in non-English-speaking regions
- This demonstrates that LLMs possess a biased characteristic of selective overconfidence in social welfare law
Scope and Limitations of the Summary
This English translation is based on a Korean summary generated from the source abstract. It is not a review of the full paper. Consult the original for detailed methods, figures and the scope of the conclusions.
This briefing is based solely on the paper's abstract and does not include detailed experimental methodology, data specifics, statistical significance levels, or comprehensive model performance evaluation. Verification and in-depth analysis require review of the full paper.
This summary does not represent the official views of Professor Haksoo Ko or the Center for Law & Economics at Seoul National University.
Summary based on original abstract · Model: gpt-4.1-mini · Generated: 2026. 10. 12. 06:06 (Korean Standard Time)