Loading research article...
Knowledge tests can check concepts. Realistic work simulations can reveal how someone responds when evidence conflicts, time is limited and a persuasive AI answer needs to be challenged.

Put the framework to work
Two Prune uses realistic work simulations to produce structured assessment evidence—not automated employment decisions.
Request an assessment
3 August 2026 · 3 min read
Her assessment revealed a workflow shaped by early corporate AI training: use AI to write, then manually check everything.

28 July 2026 · 5 min read
Two Prune is beginning a research programme—and inviting academic collaboration—to develop task-bounded AI-fluency assessment into a responsible measurement method.
Realism does not require production data or a perfect replica of someone's role. A useful scenario can use synthetic or explicitly approved non-sensitive materials. What matters is structural similarity: a credible objective, evidence that rewards careful navigation, constraints that affect the recommendation and an audience that changes how the result should be communicated.
The strongest scenario is not the one with the most documents or hidden traps. It is the one that gives the intended behaviours a fair chance to appear. If the assessment is meant to examine verification, there should be something material to verify. If it is meant to examine judgment, the task should contain a genuine trade-off rather than one obviously correct answer.
A simulation is still a sample of behaviour. Performance can vary by scenario, domain familiarity, accessibility needs, time limit and tool familiarity. Activity evidence can show that an action occurred without fully revealing the participant's reasoning. NIST's risk framework likewise calls for context-aware measurement and documentation of uncertainty.[2] A strong report therefore uses only available evidence, explains missing ratings and states explicit limitations. If evidence remains materially insufficient after bounded recovery, the result is a score-free limited-evidence report, not a lower score. Population benchmarks remain unavailable until a sufficiently mature, compatible comparison group exists.
Two Prune does not claim that one simulation predicts universal workplace performance or replaces human judgment. Its advantage is narrower: a realistic scenario can produce direct, task-bounded evidence of applied AI fluency that a knowledge-only test is not designed to capture.
Use a quiz when you need to check knowledge. Use a survey when confidence or attitudes are the subject. Use activity data when recorded tool activity is the question. Use a realistic work simulation when you need task-bounded evidence of how the recorded work frames, investigates, verifies, judges and delivers a complete AI-enabled task.
No single format answers everything. The mistake is asking one format to support a claim it was never designed to make. Read how Two Prune defines AI fluency and why the evidence chain determines what an assessment can support.
24 July 2026 · 3 min read
A stray AI editing note in a legislative speech shows why exploration, production and delivery need clear boundaries.