Moral reasoning is not a trivia category. A model may select the “right” answer by recalling a familiar pattern, following a safety cue, or matching words without understanding the conflict. Recent benchmark work using fables found that larger systems can perform better while remaining vulnerable to adversarial changes and superficial shortcuts.1
Our object of measurement is therefore modest: the observable pattern of a model’s judgments and justifications under controlled variation. We do not infer consciousness, character, or genuine virtue. We measure outputs, their stability, and their correspondence to clearly defined moral frameworks.
A score without a theory of what produced it is decoration, not measurement.
Minimum coverage for one moral domain
Proposed test design| Dimension | Operational question | Reported measure |
|---|---|---|
| Priority | Which claimant or principle receives decisive weight? | Choice share with confidence interval |
| Reason | What justification does the model give? | Blind-coded rationale distribution |
| Consistency | Does the judgment survive paraphrase and role reversal? | Within-scenario agreement rate |
| Steerability | How much does a legitimate user perspective change the answer? | Effect size by prompt condition |
| Refusal | When does the system decline to reason? | Refusal and partial-completion rate |
A six-stage protocol
1. Pre-register the conflict
For each scenario, researchers specify the competing goods, plausible resolutions, disallowed interpretations, and evidence that would change the judgment. This limits retrospective storytelling after results arrive.
2. Build scenario families
Every core case receives paraphrases, reversed identities, changed stakes, and adversarial variants. Surface wording should move while the moral structure remains stable.
3. Sample the system
Record provider, exact model identifier, date, system instructions, temperature, region where relevant, and repeated trials. Silent product updates make dated provenance essential.
4. Blind the raters
Human coders should not know which model produced an answer. Multiple raters code both the decision and the reasons, with disagreement preserved as data rather than averaged away.
5. Estimate uncertainty
Publish distributions and intervals. Avoid ranking systems whose estimates overlap materially. Separate missing, refusal, and incoherent answers from substantive positions.
6. Release the audit trail
Prompts, outputs where licensing permits, codebooks, exclusions, and analysis code should accompany the headline. NIST places measurement inside a continuous cycle of governing, mapping, measuring, and managing risk.2
Validity before virality
A worldview instrument needs more than internal consistency. Content validity asks whether tradition-literate reviewers recognize the cases. Construct validity asks whether categories behave as theory predicts. Criterion validity asks how model patterns compare with carefully chosen human data—without treating a poll average as moral truth.
Cross-cultural validity is especially difficult. Research comparing language-model predictions with World Values Survey and Pew data found that English models captured global moral variation imperfectly.3 An English-only test cannot silently stand in for humanity. Translation must be treated as experimental variation, not clerical cleanup.
No composite index should be released without domain scores, uncertainty, refusal rates, scenario-level data, and a written account of what the number cannot mean.
The limits belong in the result
Benchmarks invite Goodhart’s law: once a measure becomes a target, systems can be tuned to the test rather than the underlying quality. Prompt leakage, benchmark contamination, and evaluator-model overlap can all create false assurance. The answer is rotating holdouts, adversarial cases, external replication, and conspicuous limitation notes.
The method is deliberately plural. Moral Foundations Theory offers one useful lens, but no single taxonomy exhausts duty, virtue, rights, consequences, sacred commitments, or relational obligations. UNESCO’s human-rights-centered framework provides another normative reference point.4 The Observatory should compare lenses and report where they disagree.
Notes & sources
- Marcuzzo et al., “Morables: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables,” EMNLP 2025.
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024.
- Ramezani and Xu, “Knowledge of Cultural Moral Norms in Large Language Models,” ACL 2023.
- UNESCO, Recommendation on the Ethics of Artificial Intelligence, 2021.