Loading research article...
A useful AI assessment connects the claim it makes to a task that can elicit the right behaviour, observable evidence, conservative interpretation and visible limits.

Put the framework to work
Two Prune uses realistic work simulations to produce structured assessment evidence—not automated employment decisions.
Request an assessment
3 August 2026 · 3 min read
Her assessment revealed a workflow shaped by early corporate AI training: use AI to write, then manually check everything.

28 July 2026 · 5 min read
Two Prune is beginning a research programme—and inviting academic collaboration—to develop task-bounded AI-fluency assessment into a responsible measurement method.
Neither is sufficient alone. A strong memo may be paired with a fragile recorded process. A disordered sequence may still include a valuable correction after an error appears. Recorded sequence does not establish intent, causation or a stable trait. In one anonymised participant case, a coherent and strongly authored report still carried a material contradiction because the recorded work used AI to synthesise the evidence but did not establish verification of the final claims. Read the full verification case note.
Two Prune uses a versioned AI evaluator to interpret bounded evidence. Separate validation checks that every rating links back to the evidence, calculates the six pillar ratings and overall score, and controls whether the report can be published automatically. Publication occurs only after checks of the required report format, PDF, stored file and other gates pass. If evidence remains materially insufficient after bounded recovery, the overall score is withheld and a score-free limited-evidence report is produced. Missing capture is never turned into poor performance or zero.
Population benchmarks are currently unavailable. Any future benchmark should remain unavailable or be explicitly directional until its comparison group is sufficiently mature. This is not a claim of completed psychometric validation. Calibration is an ongoing empirical programme: compare scoring across cases and reviewers, examine whether tasks elicit the intended behaviours, monitor evidence coverage and revise the evaluator contract or scoring rules when they produce unstable interpretations.
NIST describes AI evaluation as a continuing practice involving quantitative, qualitative or mixed methods, formal reporting and measures of uncertainty.[1] Its ARIA programme also focuses on testing AI systems in realistic societal contexts rather than treating model performance alone as the whole system.[2] The same discipline is useful here: methodology is not a declaration on a webpage. It is a chain that must survive contact with real evidence.
What exact claim is the assessment designed to support?
Does the task create a credible opportunity for the relevant behaviour to appear?
Can a customer trace each material interpretation through readable evidence references and source-location cues?
Does the report explain unavailable evidence, missing ratings and the resulting limitations?
Does the report make clear that population benchmarks are currently unavailable?
A trustworthy five-page report should let a customer understand what was measured, what the recorded work supports, which readable evidence references and source-location cues support material claims, why a rating is missing, and the limits on interpretation. Report hashes support publication integrity. Raw AI transcripts, private notes, full event logs, full source text and the complete internal evidence package are not exposed by default. The score is useful only when the visible evidence chain supports it.
24 July 2026 · 3 min read
A stray AI editing note in a legislative speech shows why exploration, production and delivery need clear boundaries.