By Aleksander Niebylski · OloLand
Research checked September 12, 2026
Two experienced investors review the same memo. One challenges the reliability of the evidence. The other accepts the evidence but rejects the assumption connecting it to the recommendation. An AI product that records both responses as a thumbs-down has lost the most useful part of the review.
My view is that a team can make expert judgment more reusable by stating the decision, preserving the evidence and disagreement, and testing the result. A structured decision record gives another person something they can inspect, challenge, and improve.
Recent research gives product leaders a reason to take this work seriously. A Bridgewater AIA Labs and Thinking Machines collaboration describes training a model on expert-labelled financial information-triage tasks across several related activities. The authors report improved results on their tested tasks after incorporating expert data. The authors’ account of learning to replicate expert judgment is evidence that proprietary labels can teach a bounded workflow. It is not independent proof of investment alpha, general investor judgment, or performance outside that organization’s data and tests.
PRBench makes a complementary contribution: it frames professional finance and legal work as open-ended tasks evaluated with expert-authored, weighted criteria. That is a better starting point than asking a general model to sound experienced, because it exposes what a reviewer believes matters in the answer. The PRBench paper is evidence for rubric-based evaluation of professional reasoning; it does not make an automatic reviewer a substitute for the accountable professional.
The product lesson is to turn judgment into hypotheses. “Flag a revenue-quality concern when the filing’s explanation conflicts with the supporting schedule” is testable. “Be commercially savvy” is not. For each hypothesis, define the evidence that would support it, the evidence that would weaken it, the allowed uncertainty, and a kill criterion. At the case level, a kill criterion identifies evidence that defeats the hypothesis. At the product level, a separate stop rule defines a failure that makes the pilot unacceptable.
An illustrative example: a diligence assistant reviews customer contracts and proposes that a renewal risk deserves attention. Its record includes the clause, the contract date, the reason for the proposal, a counterargument, the reviewer’s decision, and the later outcome when the renewal is known. The example is hypothetical. Its purpose is to show how judgment can become a sequence of inspectable claims rather than an impressive paragraph.
The evaluation design should anticipate disagreement. Seeded tests can place known conflicts, omissions, and misleading context into a case. Held-out cases can test whether the behavior survives new documents rather than memorized phrasing. Human reviewers can score the same case independently, then adjudicate disagreements with the rationale preserved. Agreement is useful evidence, but forced consensus can hide a real uncertainty. Record the minority view when it changes the decision.
Delayed outcomes require care. A recommendation may be sensible even when a later event turns out badly, and a lucky outcome can conceal weak reasoning. Separate process quality from outcome quality. Did the record cite the relevant evidence? Did it identify what was unknown? Did the proposed action match the stated risk? When the outcome arrives, update the record without rewriting what was known at the time. Otherwise the training set quietly learns hindsight.
This is where vendor evidence and expert judgment must remain separate. A vendor can show a benchmark score, a case study, or a training recipe. The firm still has to establish whether the task, rubric, data rights, and failure costs match its own work. A result on a private triage dataset may justify a pilot. It does not justify saying that a model understands the firm’s investment process.
My recommendation is to build a judgment ledger before fine-tuning. Each entry should contain the decision question, evidence references, hypothesis, competing interpretation, reviewer labels, disagreement notes, confidence, and kill criterion. Add a versioned rubric and a held-out test set. When the rubric changes, preserve the prior version so the trend remains interpretable. When the model changes, rerun the same cases and inspect the failures by type.
That ledger also changes the role of experts. Their scarce time goes to defining material distinctions, adjudicating hard cases, and deciding which errors are unacceptable. They do not need to narrate every private thought for the system to learn. A concise rationale tied to evidence is more useful than an unstructured transcript, because it can be reviewed and compared.
For OloLand, the near-term opportunity is to make expert corrections and disagreements useful before claiming that we have captured expert judgment. The proposed path is specific: pick one recurring diligence decision, measure the current baseline, create a rubric, run seeded and held-out cases, and decide in advance what would cause us to stop.
For a first exercise, select a small set of disputed recommendations. Have two reviewers identify the contested claim, cite the relevant evidence, and explain what would change their conclusion. Use the disagreements to write the initial rubric, then test it on new cases. That gives the next model experiment a concrete standard to meet.
To turn a recurring judgment into a tested workflow with evidence and review, bring the decision to OloLand Studio.
Read the four-part AI Product Strategy series:
