Loading research article...
twoprune is beginning a research programme—and inviting academic collaboration—to develop task-bounded AI-fluency assessment into a responsible measurement method.

Put the framework to work
twoprune uses realistic work simulations to produce structured assessment evidence—not automated employment decisions.
Request an assessment24 July 2026 · 3 min read
A stray AI editing note in a legislative speech shows why exploration, production and delivery need clear boundaries.

21 July 2026 · 3 min read
A participant case note showing why a strong AI-assisted synthesis still needs a separate, deliberately sceptical verification pass against its sources.

13 July 2026 · 4 min read
Human–computer interaction research contributes a further distinction: using AI is not the same as relying on it appropriately. In bounded advice-taking studies, appropriate reliance involves accepting correct advice while rejecting or correcting incorrect advice.[4]
These traditions capture useful but different pieces—self-perception, knowledge, task-centered skills, process traces and reliance decisions. We see an underdeveloped opportunity to connect them across a complete, bounded AI-assisted work episode.
We use task-bounded AI fluency as a working construct: the observable capacity to frame, conduct, verify, and communicate a realistic task with AI support while retaining human judgment, evidence discipline and accountability under specified conditions.
Every part of that definition sets a boundary. Observable means conclusions should connect to evidence rather than confidence alone. Realistic task means the person has a credible objective, audience, constraints and source material. With AI support means the assessment includes the tool as part of the work rather than treating it as outside assistance. Under specified conditions means the result belongs to that task and environment; it is not a universal label.
Our current product framework uses six connected lenses: Context Framing; Evidence Navigation; AI Orchestration; Verification & Risk Control; Judgment & Synthesis; and Communication & Delivery. They are a practical way to organize observations and guide research. We are not presenting them as empirically established factors.
This is why we prefer task-bounded AI-fluency assessment to testing. A single score can create an illusion of completeness. A more responsible account makes the task, evidence, inference and uncertainty visible.
twoprune provides structured assessment evidence from timed, realistic work simulations. Because the product sells evidence, the meaning and boundaries of that evidence must be defensible.
That is why we are committing to advance this area through internal research and external academic collaboration. The aim is not borrowed prestige. It is to test whether a promising product architecture can develop into a responsible measurement method—and to be explicit about what the evidence cannot yet support.
Our research programme is focused on:
construct boundaries: defining which claims belong inside task-bounded AI fluency and which do not;
evidence-centred scenario design: beginning with the claim, the evidence that could support it and the task conditions needed to elicit that evidence;[5]
integrated process-and-outcome evidence: interpreting the submitted work alongside deliberately structured traces of the work episode, rather than treating generic activity logs as proof;[6]
opportunity-to-observe: checking whether a scenario genuinely gave a participant the chance to demonstrate the behaviour being interpreted;
explicit missingness and confidence handling: Missing evidence reduces confidence; it is not evidence of poor performance.
a responsible validation programme that tests reliability, validity, fairness and appropriate use before stronger claims are made.
Large-scale assessment provides useful design precedents without settling the AI-fluency question. OECD's adaptive-problem-solving framework, for example, treats problem solving as dynamic and uses both item scores and selected log-file indicators to illuminate parts of the process.[7] The lesson is not that every click is meaningful. It is that process evidence must be designed and interpreted against a clear evidentiary argument.
External scrutiny matters because the hard questions are not twoprune's alone to answer. What belongs inside the construct? Which observations can support which claims? When is evidence too sparse? Where might a scenario or interpretation create unfairness? Those questions require methods and perspectives from beyond one company.
The long-term opportunity is to connect AI-literacy theory, HCI research on human–AI reliance, performance assessment and Evidence-Centred Design into an end-to-end method for realistic AI-assisted work.
That method will need challenge, replication and limits. It should preserve human interpretation rather than automate consequential decisions. twoprune is a desktop AI-fluency assessment platform; it does not replace human judgment. Its reports are intended for human interpretation, and task evidence remains bounded to the assessed conditions.
If you work in HCI, performance assessment, learning science, psychometrics or workplace AI, we would like to hear from you. Challenge the construct. Help design studies. Test what the evidence can responsibly support. Contact us at hello@twoprune.com.
A useful AI assessment connects the claim it makes to a task that can elicit the right behaviour, observable evidence, conservative interpretation and visible limits.