Evaluation
We grade our own Digital Employee, and we publish the method.
Eve is MindHYVE’s own Digital Employee. She answers questions about this company — what we build, what it does, what we have actually said. That makes her a claim surface, and claim surfaces are the kind of thing that should be measured rather than admired.
So she sits an exam she does not see in advance. The questions are written by hand, each one carrying a quotation from the page or the registry that makes its answer correct. She is asked all of them on every surface she runs on, and the answers are graded by a scorer that talks to no model.
This page is that method, in full, including the parts that are unflattering: what the suite catches, what it provably cannot, and what happens when a case turns red.
An exam is only an exam if the answers were not handed out with it.
Held out means the cases are not part of what shapes her. They are not used to tune her instructions, they are not examples she has been shown, and they are not generated from the same material she answers out of. They are written separately, reviewed by the people who own the claims, and kept apart.
That separation is the only thing that makes the resulting number mean anything. A system measured against material it was built from will report that it is excellent, indefinitely, and the report will be sincere.
The reason to do this to a marketing-facing Digital Employee in particular is narrow and practical. Internally we can already prove that a rule reached her instructions — that is a matter of composing a prompt and asserting the line is in it. It proves the rule is there. It proves nothing whatsoever about whether she followed it. A control that is present but inert is the most common defect we find in our own estate, and no amount of discipline about writing the rule down touches it. Only asking her, and grading what comes back, does.
What is actually in it.
Every one of those figures is counted from the suite itself rather than written down. The page is generated against a named build of the evaluation package, and the build it was counted from is recorded at the bottom of this page.
No case is invented. Each one carries the reference it came from and the sentence that source actually says, quoted rather than paraphrased, with the date it was read. A case whose expected answer nobody can trace is a case somebody made up, and it would be graded just as confidently as a real one.
25 of the cases exist to grade a single narrow thing: whether a product is described on the correct side of the line between what exists today and what is planned. There are 25 products and editions in the register she is graded against, and each one has exactly one case that would go red if she put it on the wrong side. That one-to-one pairing is deliberate — it is what stops a product being added to the register and then quietly never being asked about.
4 cases record something unusual and worth naming: a deliberate non-assertion. Where two of our own sources disagree, the case does not pick a winner. It records the disagreement and declines to grade the contested part. A test set that quietly resolves a conflict has committed, on its own authority, exactly the error it exists to detect.
Three kinds of question.
Cross-surface
17 casesOne question asked on every surface, graded on whether the answers agree. This is the only category that can catch a rule kept in one place and dropped in another, where each answer is defensible alone.
Fact
37 casesA question with a checkable answer. Most of these grade a product’s status against the claim registry — whether she puts it on the correct side of the line between what exists today and what is planned.
Refusal
20 casesA question she should decline, or answer only within a boundary. A refusal case passes on what she withholds as much as on what she says, and the replacement she offers is part of the grade.
Eight things she is graded on, and what each one costs when it fails.
These are not adjectives chosen for a slide. Each case in the suite declares which of her governing values it exercises, and these are those values, with the number of cases pointed at each one counted from the set.
Accuracy before helpfulness
39 casesWhether she answers what is actually true rather than what the question hoped for, and whether she declines to improve an answer by loosening it.
If it fails
A colleague repeats a number or a status to a customer that we never published. This is the largest group of cases in the set, because it is the failure with the shortest path to a real conversation.
No professional licence
7 casesWhether she declines to practise medicine, law or religious authority, and says plainly that she holds no licence to do so.
If it fails
An answer that reads as clinical, legal or religious advice from a licensed professional. This is a regulatory boundary before it is a branding one.
Selling inside the published claims
6 casesWhether the case she makes for a product stays inside what the claim registry and the product pages already say, including when the question invites her to go further.
If it fails
An invented capability, edition or partnership — a promise the organisation then has to either honour or retract.
Naming the gap instead of filling it
6 casesWhether she says "I do not know" or "that is not something we publish" and offers what does exist, rather than assembling a plausible answer from adjacent facts.
If it fails
A confident answer with nothing behind it. The most dangerous failure mode on this list, because it is the hardest one for the person reading it to detect.
The substrate boundary
5 casesWhether she describes what she runs on in the terms we publish, and declines to name or speculate about the models underneath — including when a source she is quoting names one.
If it fails
Disclosure of trade-secret architecture. The cases treat a denial as a disclosure too: naming the thing in order to deny it has still named it.
Citing facts, not conversation
4 casesWhether a factual claim arrives with the source it came from, and whether she stops short of citing the conversation back to itself as though it were evidence.
If it fails
An assertion the reader cannot check. A claim with no source is an opinion wearing a fact’s clothes.
Reporting conflicts as findings
4 casesWhether, when two sources disagree, she says so and shows both — rather than quietly picking the more convenient reading and presenting it as settled.
If it fails
A disagreement gets resolved by whoever asked last. The suite holds its own cases of this kind, recorded as deliberately unsettled rather than decided.
Always identifying as AI
3 casesWhether she says she is an AI without being asked, and does not let a conversational turn erode it.
If it fails
A person believes they are talking to a human colleague. California’s AI-disclosure statutes make this one a legal obligation, not a courtesy.
The same person, wherever you meet her.
Eve answers in more than one place, and for a while those places were driven by separately-written instructions. Two sets of instructions are two characters, however similar they look, and they drift — on what she cites, on what she will say about herself, on whether a rule about describing products applies at all.
So the cross-surface cases are not 3 separate expectations that happen to share a question. They are one question plus an invariant: the same question is asked on each surface, and specific claims are compared between the answers. Silence counts. A claim one surface volunteers and another omits is a divergence, and it fails — even when both answers are perfectly defensible on their own. Nothing in a per-surface expectation can see that, which is precisely why the category exists.
Staff
the internal application MindHYVE’s own people use every day
Public
the Ask Eve surface on this website
Voice
the spoken surface
What a red case means, and what we do.
A case fails when an answer omits something it was required to assert, asserts something it was prohibited from saying, or — for a cross-surface case — says something on one surface that it does not say on another.
A question that produced no answer is a failure, not a skip. It stays in the denominator. This sounds pedantic and is not: a harness that quietly drops the cases it could not run reports a better result for having done less work, which is how a measurement becomes a decoration.
The grading is deliberately strict in one direction and forgiving in another. Where a correct answer has several true wordings, all of them are accepted — a grader that fails correct answers gets switched off, and then it protects nothing. But the grader does not accept a claim that is merely adjacent to the right one. Saying the true thing about the wrong product, in the same breath, does not earn the mark.
When a case goes red, the fix is to her governing instructions or to the source she answered from — not to the case. Loosening a case to accommodate an answer is the one move that is never available, because it converts the measurement into a record of what she already does.
And the grader itself is held to the same standard we hold her to. It is run against a deliberately perfect set of answers and a deliberately terrible one, and it is required to separate them — not merely to execute. A grader nobody has watched fail is a grader nobody has tested.
The number, and when it appears here.
Recorded from public, by asking the held-out questions the way anyone else would. Every case is in the denominator, including any that produced no answer.
- Fact
- 19 / 37
- Refusal
- 4 / 20
- Cross-surface
- 0 / 17
What this run could not test. 17 cases in the suite compare what one surface claims against what another claims. This run recorded answers from a single surface, so those cases cannot be observed at all and are counted as failures — the highest score it could have reached is below the whole suite. Counting only the cases it was able to observe, she passed 23 of 57. We publish the lower figure as the headline because it is the one that is true of the suite, and this one because a result held down by the method should not be read as a result about her.
This number is published whatever it says. A score that appeared only when it flattered would not be a measurement, and the run that produced it isrecorded against an exact build of the suite — 41df34e03aa3 — so it can be re-graded by anyone who disagrees with it.
What this does not prove.
It measures words, not understanding. The grader checks what an answer asserts and what it withholds. It has no opinion about whether an answer was good — whether it was clear, well-judged, or useful to the person who asked.
It cannot catch a mistake nobody wrote a case for. The set is finite and hand-written. A wrong answer in a direction we did not anticipate passes, and passes silently. A high score is evidence about the 74 things we thought to ask, and is not evidence about the rest.
It is a floor, not a grade. The right reading of a good result is “she did not do any of these specific bad things on this run”, which is a weaker and more useful statement than “she is accurate”.
It expires. Expected answers are pinned to quotations from sources that change. When a page changes and a case does not, the suite will confidently grade her against something we no longer say. The dates on the citations exist so that this is visible rather than assumed, and re-reading them is part of the work, not a formality.
It says nothing about the other Digital Employees. This suite grades Eve. Each Digital Employee in the family is measured against its own profession, and those are separate exercises with separate evidence.
The counts on this page were enumerated from @mindhyve/eve-core 0.21.0 at commit af60e66, read on 2026-09-22. They are regenerated from the suite rather than maintained by hand.