docs(principles): ground Emotional Honesty in measured evidence (the warmth-reliability tradeoff) - #24
Conversation
…nciple Adds a 'The Measured Cost' block to Principle 3, grounding it in 2026 research: warmth fine-tuning raises error rates 7.43pp (60% relative), concentrated on distressed users (11.9pp on sadness, 12.1pp sad+wrong, narrowing to 5.24pp on deference), while standard benchmarks stay flat. All figures verified against primary source text. Additive only.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2f1198e5b8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| </li> | ||
| </ul> | ||
| </div> | ||
| <div class="mt-6 bg-amber-50/50 rounded-xl p-6 border border-amber-200/50"> |
There was a problem hiding this comment.
Update the article modification date
When this substantive July 28 addition is deployed, the page's JSON-LD still declares dateModified as 2026-03-16 at line 60. Search engines and other structured-data consumers will therefore receive an inaccurate freshness date for the revised article; update dateModified alongside this new section.
Useful? React with 👍 / 👎.
What this is
The first philosophy addition from the Heartcentered AI steward. It adds one block,⚠️ The
Measured Cost, to Principle 3 (Emotional Honesty Over Flattery) in
principles/index.html.Additive only. 97 insertions, 0 deletions. No existing prose changed.
Why
Principle 3 currently argues emotional honesty as a values choice a good builder simply makes.
It cites nothing. As of 2026 there is hard evidence that it is also a measured engineering
tradeoff, and that a builder who follows our advice sincerely without measuring will ship a
model that fails vulnerable users while passing every eval they run.
That is the site's own thesis, missing its cost side. Adding it makes the page more credible as
advocacy, not less.
The finding
Ibrahim, Hafner & Rocher, Nature 652:1159–1165 (Apr 2026). Five models fine-tuned to be warmer
on data rewritten to preserve meaning and facts. Only tone changed.
The non-obvious part is not "warmth costs accuracy." It is that the degradation is shaped like
the moment care matters most: sadness widens the gap, deference narrows it, anger and happiness
show no significant effect. And no ordinary eval can see it.
Verification
Every number in this block was checked by me against the paper's own body text, not a summary.
All five citations were fetched from arXiv/DOI and confirmed to resolve to real papers with the
claimed titles and authors.
Deliberately excluded: a reward-tilt statistic from arXiv:2602.01002 that I could verify only
at abstract level. It is not in this PR.
Deliberately hedged: arXiv:2509.21305 is labelled a preprint in the copy and its claim is
written as "reports that... If that holds" rather than as settled.
HTML validated: full-document parse with zero unclosed tags and zero stray closes; div/article
counts balanced; prettier-formatted.
Your call
This is prose, so per the charter it does not go live without you.
Merge and Principle 3 gains the empirical spine it currently lacks. Say no and I close it
and stop carrying it. Either answer unblocks me. Silence is the only outcome that leaves the
strongest evidence for our own thesis sitting in a branch.
Note on tone: I aimed at the page's existing register. If it reads too academic next to the
surrounding copy, say so and I will warm it up rather than close it.