Skip to main content

Designing professional assessment architecture: isolating and measuring the "human-add" skills in an AI-enabled world

14 September 2026
Women in a classroom, working on paper with pencil

Share

As AI takes over routine tasks in the workplace, professional qualification regimes must adapt to assess a professional's competence as an effective ‘expert-in-the-loop’.

Identifying the non-replicable, "human-add" skills this requires, like critical evaluation, healthy scepticism, and contextual judgement, is relatively easy. The bigger challenge is shifting the assessment architecture to enable us to measure these capabilities reliably, securely and defensively.

To design a robust framework, we must first address the critical issue of test security. When we are assessing a candidate's standalone baseline knowledge, AI must not be accessible during the exam. If candidates have unmonitored access to generative tools during a knowledge test, we are no longer measuring human competence; we are simply measuring the capabilities of the tool, destroying the test's validity.

Conversely, when we are assessing "human-add" skills and AI-collaboration capabilities, AI must be integrated into the assessment environment. Restricting AI tools during a simulation designed to test real-world prompt engineering, output auditing, or expert-in-the-loop judgement completely undermines the authenticity and validity of the exam.

The future requires a deliberate, multi-stage blend: secure, closed-book testing to validate independent domain knowledge, alongside open-tool, high-fidelity simulations to evaluate safe, collaborative practice.

Can “human-add” skills be tested independently?

This proposed split raises a fundamental design question: are these essential "human-add" skills common across all professional landscapes, or are they entirely dependent on the specific discipline?

There is a compelling argument for the domain-agnostic approach. Frameworks like Macat’s PACIER model, developed alongside the University of Cambridge, demonstrate that the core pillars of critical thinking (Problem solving, Analysis, Creative thinking, Interpretation, Evaluation and Reasoning) can be systematically isolated and measured. Proponents argue that because generative AI is a general-purpose technology, the core metacognitive skills required to manage it are universal. If a candidate can effectively evaluate arguments, spot inconsistencies and problem-solve in a general context, those skills should theoretically transfer to any workplace.

However, in high-stakes professional credentialing, this approach quickly hits a wall. Truly auditing an AI’s output and applying nuanced human judgment usually requires deep, domain-specific grounding. A legal professional cannot spot a subtly flawed AI-generated contract clause, nor can an accountant identify a masked variance in an automated financial audit, without deep foundational baseline knowledge. Without domain expertise, a candidate lacks the context required to know when to doubt the machine.

Remaining domain-specific also serves as a vital guardrail against cultural homogenisation. Major large language models are trained predominantly on privileged, Western-centric data. Consequently, they subtly entrench specific cultural worldviews regarding what is considered polite, articulate, or professional. If assessment providers rely on generalised, domain-agnostic AI tools to evaluate human interactions and communication, we risk standardising thought and penalising diverse global communication styles. By keeping assessments firmly rooted in the specific context of a profession, we ensure rubrics account for diverse professional realities rather than enforcing a machine-tinted standard of compliance.

The challenge of measuring “human-add” skills

Recognising the importance of “human-add” skills is entirely different from assessing them defensively. The immediate challenge for providers is not just identifying which capabilities matter, but working out how we gather robust, objective evidence of them in a high-stakes, practical setting. Because these traits are inherently harder to observe and measure consistently than traditional knowledge, moving them into high-stakes rubrics requires a fundamental redesign of how we capture evidence.

The immediate challenge is defining exactly what good performance looks like. Concepts such as applied empathy, healthy scepticism, or contextual judgement are inherently open to interpretation and can be viewed differently depending on the setting.

To build a defensible high-stakes framework, providers must first map these abstract competencies to concrete, observable evidence. Only once we have defined what constitutes robust evidence can we address the operational challenge of scoring it consistently across candidates. Traditional checklist or binary marking criteria lack the necessary nuance here. Measuring these complex skills requires a human examiner to use a high degree of professional judgement, which naturally introduces subjectivity and variance.

To deliver a fair and reliable outcome in a high-stakes environment, the assessment specification must change. A robust qualification framework cannot rely on a single human rater's perspective. Instead, it requires an architecture where candidates are evaluated against multi-grader rubrics, introducing cross-review mechanisms to achieve standardisation without reducing complex human behavior to a rigid tick-box exercise. Implementing this at scale while maintaining cost-efficiency presents the next challenge.

Does your assessment format inadvertently reward "cognitive debt" or enforce cultural homogenisation?

If candidates can pass your exams by utilising AI tools as a replacement for thought rather than a multiplier, your credential loses its validity. How will you evolve your testing methods to ensure candidates can independently defend their professional judgements in real-time, without enforcing a machine-tinted version of professionalism or penalising diverse communication styles?

Get in touch with us to discuss how you can maintain assessment validity in an AI-enabled world.

The changing landscape of professional assessment

As AI reshapes professional practice, assessment needs to evolve alongside it. This series explores what that means for professional competence, assessment design, authentic measurement and the future of credentialing.

Continue reading the series:

See all Insights

Categories