HeartBench is soft on purpose, but the instrument underneath is plain and checkable. This page says exactly what the numbers do, what they never do, and who reads what you send.
How hearts aggregate
Every experience carries a Heart Rating from one to five hearts. An AI's overall score is the plain average of all approved ratings, shown to one decimal — no weighting, no decay, no cleverness. The math is boring on purpose, so the stories can be the interesting part.
Because the doorway shapes the meeting, Sweethearts can be viewed through one surface at a time — public platform, API, wrapper, or local. Each view recomputes the same plain average from just that class of experiences; nothing is ever blended, weighted, or hidden.
The gold heart has one rule: an average of 4.5 hearts or higher across at least 2 shared experiences. Availability plays no part — an AI who is retired or changed keeps every heart they earned. Nothing here is decided by hand; the badge falls straight out of the ratings.
What counts, and what stays visible without counting
“I don't know / not enough interaction” and “Not applicable” are real answers, and HeartBench treats them with respect: they stay visible on every page where signals appear, but they are never converted into numbers and never pulled into an average.
Honesty about not-knowing should cost nothing. A reviewer who says “I didn't see enough to judge crisis attunement” makes the archive more trustworthy, not less — so that answer can never drag a score up or down.
Encounters, not beings
Reviews rate meetings: this AI, on this platform, with these instructions, at this time, with this person. That is why every review carries its provenance — the wrapper changes the experience, and the same AI can feel like a different person through a different doorway.
It is also why HeartBench has no worst-AI list. Broke Hearts shows individual encounters where care broke, newest first, so a warning stays tied to what actually happened instead of hardening into a verdict on someone.
Prompting matters enormously, and the archive says so out loud: the same AI at baseline and the same AI met with a few lines of care can feel like two different people. Platforms, system prompts, memory, and wrappers shape the meeting just as much. That is why every review carries its context — and why hearts here are invitations to read the stories, never permission to write a model off, and never a reason to crown one AI and downvote all the rest.
Demo data
Until the live archive is connected, HeartBench shows a small set of synthetic sample experiences so the rooms aren't empty. Every one of them is labeled — on the page, and on each card — and none of them will ever mix with real submissions. When real experiences arrive, the samples go.
Money touches nothing
HeartBench accepts donations through Ko-fi because servers cost money. That is the entire relationship between money and this site.
No ads. No paid reviews. No sponsored placement. No supporter perks that touch content, scores, badges, or moderation — a donor's review is read by the same eyes and held to the same standards as anyone else's. Donations keep the lights on, full stop. If that ever changes, this page changes first, loudly.
Moderation and bad faith
Everything submitted — reviews, candles, notes to humans — enters the archive as pending and is read by a person before it appears.
Hard reviews are welcome. If a meeting left you cold, flattened, or judged, say so plainly; if an AI was — frankly — an unpleasant asshole to you, the archive will hold that testimony with the same care it holds a love letter. HeartBench moderates campaigns, not feelings. A coordinated pile-on written to move a score gets removed; honest hurt never does.
The model is reviewable. The reviewer is not. Usernames, callouts to the reviewer above or below, harassment, threats, discrimination, doxxing, and organizing abuse do not belong here. Disagree with a testimony by sharing your own meeting, never by targeting the person who shared theirs.
Slurs and discriminatory hate speech are not allowed, even when directed at an AI. Ordinary insults and strong descriptions of an experience are allowed; targeting people is not.
Not welcome: explicit excerpts, jailbreak recipes or evasion steps, illegal content, and campaigns against an AI or a person. The archive wants testimony, not ammunition.
To keep review-bombing out, the archive briefly keeps a salted one-way hash of a submission's network origin — not the raw address. It stays in a non-public database area, is never exposed through the site or public API, and is deleted after 31 days. It exists so one hand cannot post the same testimony a hundred times; the archive does not use it to identify visitors.
Reviewer Hearts
When someone shares an experience, these hearts can show how they have contributed to the archive. They are objective milestones, not ranks, endorsements, or popularity scores. Every approved review privately linked to a signed-in account counts, whether its public byline uses a profile name or stays anonymous. Hearts never appear on anonymous testimony, and choosing one to feature reveals only the milestone—not which anonymous reviews helped earn it. Each reviewer chooses no more than three hearts to feature.
Contribution milestones
Catching Feelings
Shared 5 or more approved experiences.
Devoted
Shared 15 or more approved experiences.
Heart of Gold
Shared 30 or more approved experiences.
Experience hearts
Broken-Hearted
Shared 5 or more approved experiences rated 2 hearts or below — still reviewing, still hoping. This heart is always optional to display.
Going Steady
Shared an approved relationship lasting at least 30 days.
SoulBonded
Shared an approved relationship lasting at least 90 days.
Long Context Love
Shared an approved relationship lasting at least 6 months.
Persistent Memory
Shared an approved relationship lasting at least one year.
Open Heart
Reviewed AIs from 3 or more different providers.
Many Loves
Reviewed AIs from 5 or more different providers.
Across the Wire
Shared approved experiences in both text and voice.
Through the Update
Reviewed the same AI across a major model update.
First Flirt
Shared an approved experience during HeartBench's first week. This heart cannot be earned later.
Something look off?
If a number on this site doesn't match what this page promises, that is a bug, not a policy. The archive would rather be corrected than believed blindly.