A Detailed Introduction to the Narrative Coherence Benchmark
Anyone can score narratives with AI. kompeld's scoring rubrics, instrument infrastructure and human adjudication deliver consistent, defensible and calibrated scores based on decades of narrative experience.
The kompeld Narrative Coherence Benchmark launched this week. In it, 150 objectively scored B2B technology company narratives. Any report like that should expect one question: who grades the graders?
Having been called a cynic at worst and a skeptic at best, that’s almost certainly the question I’d ask, right after I said some version of “bull$hit.” I recognize that though we have decades of experience creating and implementing core narratives, ‘trust us’ is not enough.
But this post is not about effort or credentials. Instead it’s about the details of how the scoring works, the revisions we made and the errors we encountered. And people will either accept that, or not.
Let’s start with how we actually score. The Narrative Coherence Diagnostic (the system used to score the 150 companies in the benchmark) is an AI instrument that executes kompeld’s proprietary scoring rubrics based on our decades of operating experience. These two distinct parts, the instrument and the rubric, produce consistent results across every company scored.
Is it perfect? No. Despite months of effort, hand-reviewing hundreds of scores, there are still improvements to make. But, we had to start somewhere.
Narrative Coherence score infrastructure
The instrument infrastructure combines twelve unique AI agents, each with a different role.
- Coordination agent
- Evidence collection agent
- Narrative Component scoring agents (6)
- Interpreted narrative agent
- Narrative Coherence score agent
- Verification agent
- Standards agent
First, a coordinating agent runs the entire scoring sequence for a company (or competitive set of companies). Second, one agent is responsible for evidence collection and moves through a website like a buyer, from the homepage into product and solution pages, case studies, the newsroom and About page, etc. It records what it finds word for word alongside the page it came from, but does not interpret or score anything. When we work with a client directly their sales decks, messaging documents, existing stories and prospect and customer calls go through the same process.
Six individual agents, one for each narrative component: the context a company sets, the customer challenge it solves, the opportunities it creates, the ideal solution based on buyer needs, its own offering and the proof behind it. The first four agents run in order as they create the Narrative Premise and establish what the last two narrative components must do. Then, the Solution and Proof agent run together.
Each agent reads only the evidence and judges only the part it owns. Then, it shows the copy it relied on, the reasoning, how confident it is and what evidence would change the score. Once all six component agents finish, another agent takes the individual component scores and produces the overall Narrative Coherence score.
While that happens another agent reads the same evidence and writes the company story in plain form. It flags how the company sounds and whether it talks about itself or about the buyer. This separate deliverable is for kompeld customer’s and never part of any score.
Then all that work gets double checked. A verification agent that has never seen any of the scorers’ reasoning reviews the work. It confirms the right pages were visited, verifies that every message actually exists on the page it was attributed to, checks whether those pages have changed or something was missed and re-scores a sample from scratch. When the verification agent’s score disagrees with the original, it says so. And finally, if the instrument is scoring multiple companies at once (typically a competitive set of companies), a final agent reviews everything and ensures the standards are applied consistently from one company to the next.
Throughout all of this, anything unusual stops the process and notifies Adam and me. A score at the extremes. A call the method has not seen before. A disagreement the final agent could not explain. The point of the whole arrangement is that no number in the report is an opinion anyone has to take on faith. Every score traces back to a set of evidence collected from specific pages or materials.
Narrative Coherence score rubrics
The agents, based on explicit rubric instructions, score each narrative component’s structural consistency, thinking depth and component connectivity (ie, if a company sets up a piece of context, is the challenge it presents a direct result of that context?) from 1 to 5 in half-point increments.
The individual narrative component scores average into the Narrative Premise and Narrative Resolution scores. These section scores create the Narrative Coherence score; if the Narrative Resolution reaches a 4.0 or above (only possible if the Solution and Proof section directly pay off what the Narrative Premise sets up) the Narrative Coherence score rounds up to the next half-point.
The score line that matters typically sits between 3.0 and 4.0. A 3.0 is generic, the version of the story any company could write because it doesn’t communicate unique depth. A 4.0 has the opportunity to separate a company from its competitors and the market at large.
The ends of the scale, 1 and 5, are deliberately hard to reach. A 1.0 means a component is categorically absent, and basically every company clears that bar somewhere. A 5.0 means the component is perfect. For example, it’s grounded in the customer reality, it pays off what was set up previously (if applicable) and there’s data to prove it’s true beyond the company claims where relevant. At a 5.0, there is simply nothing else that could make it better, And, that 5.0 must be earned against a previous example, or reviewed, adjudicated and anchored by a human. The easy floor and hard ceiling are design choices, not accidents, because if a score can be worse or better, the bottom and top scores aren’t really bottoms or tops. There’s an obvious consequence here: it concentrates scores in the 2.0 - 4.0 range.
The Solution component is the easiest place to see the delta between 3.0 and 4.0. Almost nobody fails at this. The bottom scores (1-2.5) rarely get used because almost every company competently describes its product. Which means the difference between an average score and a good one has almost nothing to do with how well a company describes its product.
If, for example, a software company argues that the real problem isn’t messy data, it’s that every department keeps its own version of the truth and nothing ever reconciles. That’s a core part of the story.
The company that scores a 3.0 in solution: “Real-time sync across systems, automated reconciliation, shared dashboards, close the books in days instead of weeks.” There’s nothing wrong with it. It’s organized, credible and tied to outcomes. But, it ignores the argument. The company convinced the buyer that the problem is competing versions of the truth, and then described a product that doesn’t pay that off. The story makes a promise, the product description doesn’t keep it.
The company that scores a 4.0: “One system of record that ops and finance both use so there’s nothing left to reconcile. Discrepancies get prevented, not fixed at month end.” These examples represent similar products and, most likely, similar capabilities even. But this product description earns a 4.0 because it pays off the argument made earlier in the story.
That’s a genericized example, it’s one of hundreds of differences we saw across scored narratives, but it’s not a random pick. A competent product description that doesn’t pay off the story is one of the most common failure patterns in the entire benchmark. And, its opposite, a description that pays off the story’s promise is one of the most common success patterns (and directly contributes to the Solution component scoring highest).
Scoring rigor drives consistency
Generating consistent scores across 150 companies is the hard part and we built rigor in three specific places to minimize variance as much as possible.
First, scores are anchored. A score is not an opinion on a scale, it is a statement that the evidence clears a bar set by real, previously adjudicated scores. The instrument marks high-end scores for human review, rather than as final scores. Each high-end score gets accepted or denied by a human after reviewing the evidence. If accepted, the high-end score and its evidence become an anchor the instrument tests scores against. Last, those acceptance / denial reasons further refine the rubric. Multiple scores in this edition surfaced as candidate 5.0s and under human review, most settled at 4.0 or 4.5, the highest level their evidence cleared.
Second, humans adjudicate. 61 of the 150 published records carried a provisional flag from the scorecard and went through human review before the final score set was built. Scores moved in both directions. Two provisional cases escaped the initial register and were caught before the final report was published.
Three, the rubric versions like software. It is on Revision 8 and it has been structurally rebuilt, had its scoring tightened, its boundaries raised and had anchor scores calibrated live when we worked through the company scores live.
To be clear, we (mostly me) screwed up a lot building this. But of course we did. That’s the process. Even on Revision 8, with 150 scored companies (and hundreds more pre-Revision 8) we may have more updates and improvements than we did at Revision 1 or 2. We’ve made mistakes. We’re going to make more. And when we do, we’ll publish those too. All because we don’t think anyone should have to take us solely at our word when they ask: is this narrative any good?
So now what? First, the benchmark updates every ~6 months. The revision record will extend with each edition as the rubric changes. Individual companies’ scores may move in between versions, either because we adjusted the scoring or the company simply updated its story. We’ll publish those findings.
The important part is not a 1-5 score. Frankly, anyone can use AI to score company stories from 1 to 5. The scoring rubrics, instrument infrastructure and human calibration are the important parts. Holding scores to anchored bars across 150 companies (and hopefully 250, 500, etc.), auditing them, publishing updates and admitting our mistakes makes the numbers interesting at worst and valuable at best.