Exam Security Improves When Monitoring Matches Risk

Exam Security Improves When Monitoring Matches Risk

Exam security becomes more defensible when the level of monitoring begins with the risk an assessment actually presents. A diagnostic quiz, a final subject examination and a professional certification may all use the same delivery technology, but they do not require the same evidence or oversight.

That distinction gives assessment directors and academic leaders a more useful starting point than simply deciding whether an exam should be supervised. The better questions are what must be protected, how the assessment is structured, what could realistically compromise the result and how much assurance is appropriate given the consequences.

An exam can still be watched closely without answering those questions. A candidate may remain visible on camera while receiving help through another device, while a monitoring system may generate numerous alerts without establishing whether a rule was broken. Stronger security therefore depends less on the volume of observation than on whether each control addresses a defined risk and produces evidence that can support a fair decision.

Assessment Consequence Should Set the Security Level

An online exam is not inherently low risk or high risk. Risk should be judged through the purpose, format and consequences of the assessment rather than inferred from whether it is delivered online or in person.

TEQSA’s latest security framework makes that distinction especially clear in its Assessment Security: Understanding the Risks guidance. Its framework places online proctored examinations in a medium security category rather than assuming that supervision automatically creates the highest level of protection. It recognises that credentials can still be shared and that some forms of collusion may remain difficult to detect.

For assessment teams, this creates a practical sequence for decision making. First consider the purpose of the result and what happens if it is unreliable. A low consequence assessment designed to guide subsequent teaching may tolerate greater uncertainty because no qualification or major progression decision depends on one result. An examination that confirms professional competence may require much stronger assurance because the consequences extend beyond the individual candidate.

Format should then shape the risk assessment. A timed calculation test, an open book analysis and an oral examination create different opportunities for unauthorised assistance even when they contribute to the same programme. Looking at purpose, format and consequence together gives educators a clearer basis for deciding which risks are material.

The next question is how the result could realistically be compromised. Identity substitution, unauthorised materials, communication with another person and prior exposure to questions are different risks. Once those risks are identified, educators can select controls that address them directly rather than increasing monitoring simply because an assessment carries importance.

This creates a proportionate relationship between consequence, threat and control. It also gives academic leaders a clearer basis for explaining why different assessments require different forms of oversight.

Reviewable Evidence Matters More Than Alert Volume

Monitoring becomes useful when the information it produces can be interpreted consistently. An alert is therefore best treated as a prompt for professional review rather than evidence of misconduct in itself.

A different concern emerges from a 2025 systematic narrative review of online proctoring systems. The research examined publicly available descriptions of 33 systems and identified limited transparency around the use of artificial intelligence in decisions about possible rule violations. It also highlights the difficulty of interpreting behavioural indicators without considering their context.

Looking away from a screen, for example, may indicate access to outside assistance, but it may also reflect concentration or thought. The value of monitoring therefore depends on what happens after an indicator is generated.

Assessment teams can strengthen that process by defining the review pathway before the exam begins. Reviewers need to know which indicators are material, what supporting evidence is required, how uncertainty is recorded and when a case should move to an academic decision maker. These arrangements turn monitoring data into something that can be assessed fairly rather than leaving individual reviewers to interpret alerts differently.

Capacity matters as well. A control that generates a substantial volume of evidence requires enough trained reviewers to assess that material within an appropriate timeframe. Selecting monitoring intensity should therefore include an operational question: can the organisation examine what this control produces with the consistency the assessment requires?

This makes human review part of security design rather than an activity added after the technology has done its work.

Different Risks Call for Different Controls

A stronger security model separates the risks that arise from assessment design, candidate identity, resource access and behaviour during the examination. Monitoring can address some of these areas, but each requires a control suited to the weakness being protected.

Assessment design, for instance, influences how useful outside assistance would be. A heavily reused question bank may be vulnerable before the exam session starts. A task based almost entirely on recall may be easier to answer using unauthorised resources than one requiring candidates to apply knowledge to a new scenario. Clear rules around permitted materials also reduce uncertainty by making the boundary between legitimate and prohibited behaviour easier to interpret.

The design side of the argument is reinforced by the Jisc report on trends in higher education assessment. Its findings describe movement towards programme-level assessment design and a reconsideration of how different assessment formats should be used as institutions respond to digital transformation and generative AI.

For educators, the implication is that integrity becomes stronger when assessment design and delivery controls support one another. Monitoring does not need to carry the entire security burden if the task itself reduces opportunities for unauthorised assistance and makes genuine competence easier to distinguish.

Identity assurance should also be treated separately from behavioural monitoring. Confirming who enters an examination addresses a different question from establishing what happens during the session. A lower risk assessment may require a straightforward identity check. A more consequential examination may justify continuing assurance, restricted access to resources and a session record capable of supporting later review.

Where candidates must sit remotely and the consequences are material, remote proctoring technology can form one part of that control set. Its relevance depends on the specific risk it is intended to address, whether that is identity substitution, access to prohibited resources, communication with another person or the need for a reviewable record. It does not need to become the default response simply because an assessment is online.

A Tiered Framework Makes Decisions More Consistent

A practical risk framework allows assessment teams to apply common principles without requiring identical supervision for every examination.

At a lower level, the result may be formative or carry limited weight. Clear candidate instructions, basic identity assurance, question variation and sensible time controls may provide sufficient protection. The objective is to preserve confidence in the result without introducing controls that are disproportionate to its consequences.

At a middle level, an assessment may contribute materially to progression or a final grade. Stronger identity checks, controlled resources, selected monitoring and a defined review process may be appropriate. The important step is to identify which behaviours would genuinely compromise the result and what evidence would be needed before an academic judgement is made.

At the highest level, the result may confirm safety critical competence, grant entry to a profession or determine a significant progression decision. Assessment teams may then combine several controls, including stronger identity assurance, secure task design, closer oversight, reliable records, human review and a formal escalation process.

The value of the tiers is not in creating rigid categories. It is in helping educators make the reasoning behind each security decision explicit. Two assessments may use similar technology but sit at different levels because their purpose, format and consequences create different integrity risks.

Administrative capacity can be considered at each tier as well. More intensive monitoring may be appropriate only when the institution can review the resulting evidence thoroughly and consistently. This prevents security decisions from being driven by what technology can collect rather than by what the assessment actually needs.

Proportion Makes Security More Defensible

Consistency in assessment security does not require every candidate to be monitored in the same way. It requires shared principles for deciding which controls are appropriate and how resulting evidence will be judged.

An open book essay, a timed calculation test and a professional certification exam create different opportunities for unauthorised assistance. Applying identical monitoring to all three may look consistent, but it can remove the contextual judgement needed to interpret behaviour fairly.

A more useful standard is to align the purpose and consequence of the result, the assessment format, the realistic threat, the quality of the evidence and the institution’s capacity to review it. That gives educators a framework for explaining not only what security measures were selected, but why they were appropriate for that particular assessment.

The objective is not to maximise observation. It is to create enough relevant, reviewable evidence to protect the validity of the result while keeping oversight proportionate to the risk and manageable for the teams responsible for applying it.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *