Management consulting and artificial intelligence have the same challenge: consistency. Ask the same question twice and you get two different answers.
Two consultants assess the same company and produce different priorities. Run a commercial question through an AI model twice and you get two confident, well-written, materially different responses.
In both cases the output looks authoritative, and the variance stays invisible, because you only ever see one sample.
That matters more than it sounds. If the instrument lacks reproducibility, you cannot measure anything with it. You cannot assess before a raise, do twelve months of work, assess again, and claim the difference means something, because you have no idea how much of the difference is you and how much is the instrument. An unreliable measure is not just a weak measure. It is not a measure at all.
What follows is the over six month journey to solving the single biggest challenge facing both AI and management consulting: consistency.
Phase One: Two Dozen Rules and a Lot of Faith
When I first began feeding the PRIYA platform with product-market fit and commercial readiness inputs, I noticed a glaring issue. The consistency just wasn't there to provide the solution I was hoping to deliver as a management consultant. I started trying to solve for the consistency challenge by writing about two dozen rules from my own experience in the life sciences covering the things that are objectively true or not true about a submission: incomplete information, missing information, basic contradictions between GTM dimensions. Then I assumed AI would handle everything else. The edge cases. The permutations. The combinations I had not thought to write down.
It is a reasonable assumption. Modern LLMs are genuinely good at reading a document and telling you what is wrong with it. They are also good at telling you something slightly different every time you ask. With only two dozen rules underneath, almost the entire assessment was riding on the AI layer, and the AI layer is the part that was changing from one run to the next.
Phase Two: Run Multiple AI Agents in Parallel and Take a Consensus
The next phase was to stop trusting a single agent and throw raw AI processing power at the problem, incorporating multiple agents running independently in parallel. Five agents assess the same submission, and only findings that enough of them agree on survive into the deliverable.
This helped. It was a real improvement, and it remains part of the architecture today. But it didn't solve the problem, it averaged it. Five agents drawing on their own reading still disagreed about what counted as a finding, described the same issue in language different enough that the merge treated it as two, and rated the same condition differently.
Consensus among unanchored opinions is still an opinion. Run the whole thing again tomorrow and the deliverable was measurably different, with different conclusions, a different narrative, and different action items.
Phase Three: Say It Louder
The third phase was whack-a-mole. I ran the pipeline more than a hundred times, watched where the discrepancies surfaced, and wrote a rule into the agent instructions for each one as it appeared.
Then I ran it again, found the next set, and wrote those in too. Every fix was a bandaid on one specific symptom, and every round surfaced new symptoms somewhere else.
This was the stretch where I had real doubts about whether the thing was possible at all.
An example of one of the rules was simple and absolute: if the market sizing does not carry a complete derivation, coherence cannot be rated Strong. No exceptions, no judgment. I stated that in explicit, capitalized, unmissable language, and then I measured it across twenty-five agent executions.
Three applied it. Compliance was twelve percent.
A later revision took the same condition to full compliance across a test set and dropped a related misclassification rate from twenty-four percent to zero. But by then the more useful lesson had landed, and it had nothing to do with wording, or using ALL CAPS to try to get the AI to do what I wanted it to.
Success: Flipping Everything Upside Down
The breakthrough was not a better prompt. It was realizing I had the sequence backwards.
I had been running AI first and using rules as a backstop. The fix was to build the rules engine out properly, all the way to 225 rules covering the conditions, contradictions, thresholds and dependencies across all product-market fit and commercial readiness dimensions, with highly specific call outs that are unique to life sciences and diagnostics (including evidence-based value propositions, KOL strategy, market access and regulatory strategy) and run it first. Its output becomes the fixed, rich context every AI agent receives before it forms a single opinion.
Agents no longer begin from relatively fewer inputs and their own reading of the document. They begin from the same set of counted facts. Their job narrows to what models are genuinely superb at and rules are hopeless at: reading a submission as a narrative, noticing when two sections tell stories that cannot both be true, and telling the difference between evidence and assertion.
Everything requiring interpretation goes to five independent agents whose consensus is then bounded by constraints enforced after the vote rather than requested before it. Stated in a prompt, that constraint held twelve percent of the time. Enforced in code after the vote, it holds every time, by construction.
What That Produced
We developed a validation protocol that runs the entire pipeline several times against one unchanged submission and compares the deliverables. We have now run more than ten validation studies across products spanning multiple regulatory contexts, including a European IVDR assay, a US laboratory-developed test, and non-clinical manufacturing software.
The result has been the same in every one. Every finding appeared in every run. Every finding held the same level of agreement. No action item was dropped.
In the largest study, forty findings were identified by every agent in every run, with identical agreement counts throughout. The only thing that differed was the prose.
Where the Variability Still Lives
It has not gone away entirely, and I would be suspicious of anyone claiming it had. Five agents that agree on which findings exist, at what level of agreement, in which sections, at which severity, will still write five different paragraphs explaining each one.
So we measure that too. An independent model reads every version of every explanation, blind to which run produced which, and scores whether they make substantively the same argument.
In the study shown above that came back at eighty-six percent average prose similarity. In the larger study it averaged eighty-nine, with the weakest finding at seventy-two, where the runs agreed completely on the conflict but emphasized different downstream consequences of it.
Here is the important part: none of that changes the output. The findings are the same, the agreement counts are the same, the severities are the same, the action plan is the same, and the readiness classification is the same. The prose varies. The assessment does not. That is the correct place for the remaining variance to sit, and it is the only metric in the suite that moves at all, which is exactly what makes it worth reporting. A metric that always returns a perfect score cannot distinguish a reproducible system from an insensitive instrument.
Why This Is the Part That Deserves Attention
Reproducibility is not a nice property of this system. It is the foundation the entire approach rests on, and without it nothing else we claim would mean anything.
It is what converts an opinion into an instrument. A traditional engagement hands you one sample from a distribution you never get to see. It might be an excellent sample. You have no way to know, and no way to re-run it in nine months and claim the delta reflects your progress rather than a different consultant on a different Tuesday. When the same submission produces the same assessment every time, the assessment becomes a baseline. Act on it, do the work, run it again, and show an investor a difference that is genuinely yours.
Why This Is Not “Just AI Doing Go-To-Market”
If you have gotten this far, this is the part I most want understood, because it is the part that gets flattened into a false headline.
You cannot prompt a 225-rule engine into existence. Every one of those rules is a judgment about what actually matters in life science commercialization, encoded from more than twenty years of doing the work: which contradictions between sections are fatal and which are cosmetic, where a reimbursement assumption quietly caps a market, which evidence tier a clinical claim requires, what a value proposition is missing when it reads well and says nothing.
No model generates that from a blank page, because it is not retrievable information. It is accumulated operating experience, written down as testable conditions.
And the reverse is just as true. Those 225 rules, on their own, cannot read a submission as a story. They cannot notice that the regulatory intent and the clinical language in the value proposition are describing two different companies. Rules do not interpret. That is precisely what the AI layer is for.
Take either one away and you land back where we started. Domain expertise without the deterministic layer is a very experienced person producing a different answer each time they look, which is the honest description of most consulting. AI without the domain layer is fluent, fast, and unanchored, which is the honest description of most AI-generated strategy. Both fail the same way, into subjectivity you cannot reproduce.
Together they do something neither can do alone. The expertise supplies the standard, the code enforces it, and the model does the interpretive reading on top of a foundation that does not move. That pairing is the product. The AI is not the clever part. The pairing is.
The Bottom Line
If you are building anything on top of an AI model and relying on the prompt to enforce your rules, measure it before you trust it. My instruction was as clear as I knew how to write it, in capitals, marked mandatory. It was followed twelve percent of the time.
What worked was giving up on asking. Establish in code everything the data can already answer, hand that to the model as fixed ground, and reserve the model for the reading and interpretation that genuinely require it. Then measure what is left over and publish the number.
That is a harder thing to build than a good prompt. It is also the only version of this I would be willing to put in front of a board.
See what a reproducible commercial readiness assessment looks like.
Back to Blog