Skip to main content

Maintaining assessment validity: the shift in what it means to be ‘professionally capable’

14 September 2026
Man at desk writing with a pencil

Share

The rapid integration of AI into the workplace fundamentally shifts what it means to be a competent professional. It also significantly heightens the risks if a practitioner lacks the skills to navigate this new landscape safely. Consequently, this challenges the very validity of professional qualification. If we are not actively measuring these critical AI-collaboration skills in an environment that closely mirrors modern professional practice, assessments risk failing to provide the public assurance they were designed to guarantee.

To understand how qualifications must adapt, we must first look at how we have historically measured professional capability. Traditionally, professional assessments have focused on two core areas: validating a candidate's knowledge and testing their ability to apply that knowledge to realistic scenarios.

To test application, there have generally been two different paths. One uses high-fidelity, authentic simulation assessments, often considered the "Rolls-Royce" of testing. While these simulations offer exceptionally strong validity, they are also complex, logistically demanding and expensive to deliver. Alternatively, we use pragmatic assessment formats where the delivery method differs from real-world practice, such as multiple-choice questions or written essays, but where the underlying capability to apply knowledge is still successfully tested. While the direct validity of these latter formats may not be as robust as a full simulation, they have been good enough to safely certify competence.

In an AI-driven world, however, this is shifting. The way competent professionals have to apply their knowledge is no longer by drafting documents or writing reports. It’s by being an effective, critical human-in-the-loop, or rather “expert-in-the-loop”, crafting prompts, refining outputs and exercising judgement to drive the right outcome.

Because assessments play such a profound role in driving learning and skill adoption, we must adapt what we measure. The old adage holds true: what gets measured gets done, and in the world of education, that means what gets assessed gets taught. If our qualification regimes continue to test elements that AI can handle, educators will continue to prepare candidates for an obsolete workplace. By bringing non-replicable “human” elements into high-stakes rubrics, we force the entire training ecosystem to prioritise genuine behavioural and interpersonal development.

Does knowledge still matter?

Even in a world where AI handles all knowledge retrieval, foundational knowledge is not dead. In fact, it has never been more critical and is a prerequisite to success in an AI-driven workplace.

Candidates with strong conceptual foundations use AI as a multiplier to elevate their thinking, whereas those lacking foundational knowledge use it as a complete replacement, automatically accepting answers they are unequipped to evaluate.

In high-stakes, regulated professions, this passive reliance introduces the risk of hallucinated negligence. If a professional lacks the baseline domain expertise to spot an AI error or fabrication, the real-world consequences can be legally and financially catastrophic. High-stakes assessment must therefore certify that a candidate possesses the deep foundational knowledge required to act as an effective, responsible expert-in-the-loop, capable of supervising, auditing and challenging AI outputs rather than just operating the software.

Cognitive laziness and the junior deskilling crisis

The use of AI tools can result in two fundamentally different cognitive behaviours:

  • Cognitive off-load: This is the legitimate, efficient outsourcing of routine tasks. Much like a calculator accelerating arithmetic, off-loading allows a professional to free up mental capacity for higher-order strategy, nuance and critical thinking. This is a valuable, everyday skill that should be encouraged.
  • Cognitive laziness: This occurs when a user entirely disengages their critical faculties and unquestioningly follows technology. It is the modern "Sat-Nav" effect - drivers following GPS instructions without questioning whether they are being steered down a dead-end or into a river. Instead of using the tool to multiply their capability, decision-making is abdicated to the machine.

Cognitive laziness creates severe cognitive debt. A striking four-month study from the MIT Media Lab utilising EEG scans found that when individuals relied heavily on generative AI, the brain rhythms associated with deep memory formation actively declined. By the final session, 78% of participants who relied on ChatGPT could not quote anything from material they had produced just minutes earlier.

When cognitive laziness abounds in the workplace, it triggers a broader crisis: the deskilling of junior professionals. In many industries, entry-level tasks, such as drafting basic legal documents, compiling initial financial reports, or writing routine code, are the traditional training grounds where junior workers build foundational intuition and professional judgment. As AI automates these entry-level tasks, junior professionals are more likely to lack the foundational expertise to challenge the machine and are therefore more likely to succumb to cognitive laziness because they are denied the opportunity to develop the practical, “human-add” skills required to become tomorrow’s experts.

Assessing AI capabilities

To defend against cognitive laziness and ensure professionals can safely navigate an AI-enabled workplace, qualification regimes must adapt to measure how candidates harness the power of AI to elevate an outcome, rather than using it to replace thought. The focus of skills-based testing must pivot from measuring whether a candidate can generate a baseline answer, to measuring how they manage the tool and critically evaluate its output.

To act as a safe, effective expert-in-the-loop, a candidate must demonstrate a specific set of AI capabilities. These capabilities represent the core "human-add" skills of the modern workplace - the non-replicable human traits that ensure we are multiplying the power of AI rather than passively accepting its output.

Assessments need to target the practical, cognitive skills required to supervise a machine. These capabilities include:

  • Critical evaluation: The ability to actively audit AI-generated outputs, spot subtle fabrications or hallucinations and identify factual or contextual errors.
  • Justified deviation: The confidence and professional judgement to challenge an automated recommendation, deviate from the machine’s path and independently defend that decision.
  • Prompt calibration: The capacity to direct, refine and iterate with generative tools to elevate the quality of the final outcome, rather than passively accepting the first draft.
  • Metacognitive awareness: The ability to plan and monitor one's own thinking, recognising when to rely on automated efficiency and when to pause, slow down and apply deep human intervention.

By explicitly defining these capabilities, professional bodies can establish what it actually means to be competent in an AI-driven environment. Once we define what we need to measure, the challenge then shifts to the design of the assessment itself and how we gather robust, defensible evidence of these skills in a high-stakes setting.

Are you measuring obsolete constructs or true AI-ready competence?

If your assessments primarily test rote knowledge retrieval, you are measuring capabilities that machines can now execute flawlessly. How will you redesign your assessment regime to explicitly certify that a candidate has the deep foundational knowledge required to act as an effective, expert-in-the-loop, and the AI Capabilities focusing on concrete, observable behaviours, such as a candidate's ability to audit AI outputs and justify professional deviations?

Talk to our assessment experts about designing assessments that measure foundational knowledge and the AI Capabilities required to act as an effective expert-in-the-loop.

Explore other articles:

See all Insights

Categories