When AI Can Produce the Assignment, What Are We Actually Assessing?

When AI Can Produce the Assignment, What Are We Actually Assessing?

Riveral Advisory · Research Insight

When AI Can Produce the Assignment, What Are We Actually Assessing?

Generative AI is forcing higher education to reconsider not only academic integrity, but what counts as credible evidence of learning.

AI & Digital Transformation September 2026 Research Insight
Author

Dr. Ahmad Alhalak

Founder & Principal, Riveral Advisory

Dr. Alhalak brings more than 20 years of experience in higher education, with work spanning academic development, institutional strategy, leadership, organizational development, and emerging technologies.

Research & analysis by Riveral Advisory

For generations, higher education has relied heavily on a relatively simple assumption: when a student submits an essay, report, analysis, presentation, piece of code, or other academic work, the quality of that artifact tells us something meaningful about what the student knows and can do.

Generative artificial intelligence is making that assumption increasingly difficult to sustain.

This is no longer a hypothetical issue. Research published by Lumina Foundation and Gallup in 2026, based on 3,801 students pursuing associate or bachelor's degrees in the United States, found that 57% use AI for schoolwork daily or weekly, while only 13% said they never use it. At the same time, 52% reported that at least some of their courses lacked clear guidance about AI use. [1]

Meanwhile, EDUCAUSE surveyed 438 faculty and staff directly involved in designing or delivering learning assessments in March 2026, documenting a sector in which assessment practices, policies, expectations, and attitudes are already changing in response to AI. [2]

The central question

If AI can produce the artifact we traditionally grade, what evidence actually demonstrates that the student has learned?

That question shifts the conversation away from simply asking whether a student used AI and toward something more fundamental: whether the assessment still produces sufficiently trustworthy evidence of the capability it claims to measure.

The issue in context
57%
of surveyed U.S. college students reported using AI for schoolwork daily or weekly.
52%
reported that at least some courses lacked clear guidance on specific AI-use policies.
438
higher education assessment practitioners participated in EDUCAUSE's 2026 assessment study.
1,066
students were included in a 2026 natural experiment examining assessment conditions and grade outcomes.

The real issue is assessment validity

Assessment has never provided direct access to learning.

Learning itself is largely internal. Educators cannot directly observe a student's understanding, disciplinary reasoning, judgment, creativity, problem-solving capability, or conceptual development. Instead, assessment asks students to produce evidence from which educators make inferences about those capabilities.

An essay may be treated as evidence of analytical reasoning. A laboratory report may be used as evidence of scientific understanding. A coding project may support a judgment about programming capability. A business case may be used to infer strategic thinking.

Fawns, Boud, and Dawson describe this challenge through a framework of assessment evidence based on four types of proxy: product, process, performance, and practice. Their argument is that the critical question is not simply whether an assessment appears appropriate, but whether the evidence generated actually warrants the claim being made about what a student knows or can do. [4]

Generative AI complicates this relationship because the final product may increasingly reflect a combination of human and machine capability.

A polished essay could demonstrate strong student reasoning. It could also demonstrate strong prompting, effective editing of machine-generated material, extensive external assistance, or sophisticated collaboration between a learner and an AI system.

Those possibilities are not equivalent.

The assessment challenge is no longer only whether the work is original. It is whether the work is valid evidence of learning.

Better performance is not necessarily better learning

One of the most important distinctions emerging from current research is the difference between successfully performing a task and learning from the task.

The OECD's Digital Education Outlook 2026 synthesizes emerging research suggesting that general-purpose generative AI can improve the quality of student outputs without necessarily creating corresponding learning gains. The OECD notes that advantages visible while students use general-purpose AI can disappear, and sometimes reverse, when access to that assistance is removed during later assessment. [3]

Importantly, the OECD does not conclude that AI inherently undermines learning. It finds that AI can support meaningful learning when it is used with clear pedagogical intent. The distinction is therefore not simply between AI and no AI. It is between AI that supports learning and AI that substitutes for cognitive activity that the learner was expected to develop.

A systematic review and meta-analysis by Deng and colleagues examined 69 experimental studies and reported positive effects of ChatGPT interventions on academic performance, affective and motivational outcomes, and higher-order thinking tendencies. At the same time, the authors specifically cautioned researchers and educators to distinguish the quality of AI-supported outputs from evidence that learners had developed the relevant capability. [7]

Important distinction

AI can improve what a student produces without necessarily improving what the student can subsequently do.

A 2026 natural experiment raises an uncomfortable question

Recent empirical evidence illustrates why assessment conditions matter.

Brattli, Utne, and Lynch examined grade data from a compulsory undergraduate course delivered over five years, involving 1,066 students. [5]

From 2021 through 2024, the course used an AI-accessible take-home examination. In 2025, assessment moved to an AI-restricted, supervised in-person format. According to the researchers, course content, intended learning outcomes, grading criteria, examiner continuity, and the structural design of the examination tasks remained broadly stable.

The grade distribution changed substantially.

Failure rates had ranged from approximately 2% to 6% during the four earlier cohorts. Under the 2025 supervised assessment condition, the failure rate increased to 18.4%. The researchers found a statistically significant association between examination period and grade distribution. [5]

The study cannot establish that generative AI caused the change. It was a natural experiment rather than a randomized controlled trial, and factors such as cohort variation, assessment environment, or examination anxiety may also have contributed.

But it raises an important measurement question:

Assessment validity

Were the two assessment conditions producing evidence of the same underlying capability?

If performance changes substantially when external cognitive assistance is removed, grades under AI-permissive conditions may partly represent the capability of a student-plus-AI system rather than the student's independent competence.

Neither capability is necessarily unimportant. In modern professional environments, knowing how to work effectively with AI may itself be valuable.

But independent capability and AI-augmented capability are not identical—and institutions need to know which one they intend to certify.

Policy alone cannot solve a measurement problem

Much of higher education's early response to generative AI focused on policy: whether AI was prohibited, restricted, permitted, or encouraged.

Clear expectations remain important. The Lumina Foundation–Gallup findings illustrate the problem of inconsistency: 53% of students surveyed said their institution discouraged or prohibited AI use, while 52% reported unclear guidance in at least some individual courses. [1]

But policy clarity and assessment validity are two different issues.

Corbin, Dawson, and Liu distinguish between changes that primarily communicate permissible behavior and structural changes to assessment itself. They argue that instructions alone depend heavily on student compliance and cannot guarantee that submitted evidence genuinely represents the capability being assessed. [6]

The implication is significant:

An AI policy can clarify expectations. It cannot do the work of assessment design.

Detection is unlikely to be the long-term answer

Another response has been to focus on detecting whether AI was used.

Research by Kofinas, Tsay, and Pike examined human-authored, AI-influenced, and AI-generated assessment work across two UK universities. Their findings indicated that markers generally could not reliably distinguish work containing generative-AI input from work that did not, and that suspicion of AI could itself influence the marking process. [8]

This matters because detection is addressing a different question from learning assurance.

Even a perfect answer to “Was AI used?” would not necessarily answer “What did this student learn?”

TEQSA's 2025 work on assessment reform therefore emphasizes forming trustworthy judgments about learning through multiple, inclusive, contextualized forms of assessment evidence, alongside opportunities for students to engage appropriately with AI and secure points at which learning can be assured. [9]

The more sustainable objective is not to design an assignment that AI can never complete.

It is to design an assessment system from which the institution can make a sufficiently credible judgment about what the student knows and can do.

Authentic assessment is important—but not automatically AI-resilient

One common response to generative AI has been to recommend more authentic assessment: professional projects, applied problems, real-world cases, workplace simulations, and tasks connected to disciplinary practice.

There are strong educational reasons for doing so.

However, authenticity should not be confused with proof of independent learning.

Kofinas and colleagues found that increasing the authenticity of an assessment did not, by itself, allow markers to reliably determine whether generative AI had influenced the work. [8]

A highly realistic consultancy report, engineering proposal, policy paper, strategic plan, marketing project, or case analysis may be professionally authentic while still allowing a large portion of the underlying intellectual work to be generated externally.

The future of authentic assessment may therefore depend not only on what students produce, but also on the evidence surrounding that production.

Higher education may need to assess two kinds of capability

The response to AI does not have to be a choice between “AI everywhere” and “AI nowhere.”

Instead, programs may need to become much more explicit about the distinction between independent capability and AI-augmented capability.

01 — Independent

Independent Capability

Knowledge, reasoning, judgment, communication, disciplinary understanding, or professional performance that the graduate must be able to demonstrate personally.

02 — Augmented

AI-Augmented Capability

The ability to use AI appropriately while framing problems, interrogating outputs, verifying information, identifying limitations, exercising judgment, and remaining accountable for the result.

This distinction is increasingly important in professional education.

A nurse needs clinical judgment even when technology provides support. An engineer requires foundational technical understanding. A teacher needs pedagogical judgment. A researcher needs to evaluate evidence and methodology. A business graduate should understand the principles underlying the decisions being made.

At the same time, graduates are entering environments in which AI-supported work is increasingly normal.

UNESCO's AI Competency Framework for Students emphasizes a combination of human-centered thinking, ethical judgment, understanding of AI techniques and applications, and responsible creation. [11]

The OECD similarly argues that education should preserve valued human knowledge and skills while enabling purposeful learning without AI, with educational AI, and with general-purpose AI where appropriate. [3]

A question for every program

What must this graduate demonstrate independently—and what should this graduate demonstrate with AI?

Riveral Advisory Synthesis

The Evidence-of-Learning Framework

Instead of asking whether an assessment is “AI-proof,” institutions can examine whether the overall evidence produced by assessment supports a credible judgment about student capability.

This framework is a Riveral Advisory synthesis of the research discussed in this article. It is intended as a practical decision-making framework and is not presented as a validated psychometric instrument.

01

Capability

What must the student personally know, understand, judge, explain, or perform?

02

Augmentation

Where should AI appropriately enhance professional or academic performance?

03

Process

What evidence makes the student's reasoning, decisions, development, and iteration visible?

04

Verification

Where is authenticated evidence of individual capability necessary?

05

Progression

How does credible evidence accumulate across the student's entire program?

The design question becomes: What combination of evidence would allow us to make a credible judgment about this student's capability?

Assessment may need to become an evidence system, not a single artifact

One of the most important implications of the emerging research is that institutions may have relied too heavily on individual submitted products as sufficient evidence of learning.

AI creates an opportunity to reconsider that assumption.

Instead of relying on a single artifact, assessments can deliberately combine different forms of evidence.

Product + Defense

A substantial written analysis followed by focused questioning about assumptions, evidence, and decisions.

AI Use + Critique

Students use AI openly, then identify weaknesses, verify claims, revise outputs, and justify their decisions.

Portfolio + Process

Finished work is considered alongside drafts, decisions, reflection, feedback, and evidence of development.

Applied Performance

Demonstrations, simulations, presentations, practical tasks, or professional scenarios reveal capability directly.

The purpose is not to make every assessment more complicated.

It is to avoid making consequential claims about student capability from evidence that has become increasingly ambiguous.

Oral and performance assessment offer possibilities—but not a universal solution

Oral assessments, demonstrations, simulations, and other forms of synchronous performance have received renewed attention because they can make student reasoning more visible.

A 2026 study by Ogle and Jarrett examined the replacement of a written nursing case study with a structured oral assessment in a cohort of 740 students. Overall unit performance remained stable, assessment outcomes were comparable, end-of-semester examination performance improved, and educators reported greater confidence in authorship and alignment with clinical judgment. [10]

The findings are promising—but they should not be interpreted as evidence that every essay should become an oral examination.

Oral assessment can introduce other challenges: scheduling, accessibility, student anxiety, consistency between assessors, staff training, scalability, and resource requirements.

The principle is therefore broader:

Choose the form of evidence that makes the capability you actually care about sufficiently visible.

Nor should higher education simply retreat to traditional examinations

If unsupervised coursework becomes difficult to authenticate, returning all assessment to examination halls can appear to offer a straightforward answer.

In some circumstances, supervised assessment will remain important. Institutions need credible points at which students demonstrate independent capability.

But making all assessment secure by default could sacrifice authenticity, inclusion, flexibility, collaboration, creativity, and opportunities to demonstrate complex professional capability.

TEQSA's assessment-reform work instead identifies multiple institutional pathways, including program-level approaches, unit-level approaches, and combinations of both. It emphasizes meaningful secure points alongside authentic engagement with AI, the process of learning, and program-level assessment design. [9]

Jisc's review of assessment trends similarly identifies growing interest in program-focused assessment, portfolios, oral and investigative vivas, capstones, and process-oriented assessment approaches. [12]

The strongest future assessment systems may therefore deliberately combine different conditions.

Some learning may need to be demonstrated independently. Some collaboratively. Some with AI. Some without it. Some through process. Some through product. Some through live performance.

What matters is that those choices are intentional.

The bigger shift may be from assignment redesign to program-level assurance

Perhaps the most significant implication of generative AI is that institutions may eventually need to stop asking every individual faculty member to solve the assessment problem alone.

A degree is ultimately a claim about what a graduate knows and can do.

Yet assessment is frequently designed course by course and assignment by assignment.

That fragmentation can produce repetition, gaps, inconsistent expectations about AI, unnecessary assessment burden, and uncertainty about where important graduate capabilities are actually demonstrated.

TEQSA's recent work explicitly examines program-wide assessment approaches in which learning assurance is considered across the structure of an entire qualification rather than being solved independently inside every unit. [9]

The 2026 EDUCAUSE Horizon Report also identifies AI as a force reshaping assessment and notes movement toward more authentic and process-based demonstrations of learning. [13]

This changes the institutional question from:

From assignment to program

Across this entire program, where and how do students convincingly demonstrate the capabilities we say our graduates possess?

That is no longer simply an assessment question.

It becomes a question of curriculum design, faculty capability, academic leadership, governance, learning outcomes, technology, student communication, quality improvement, and institutional strategy.

What institutional leaders should be asking now

Institutions do not need to redesign every assessment immediately. Nor should they respond to a rapidly changing technology with hurried, institution-wide solutions that have not been evaluated.

A better starting point is to identify where assessment evidence is most consequential, most ambiguous, or most vulnerable to changes in how students complete intellectual work.

Questions for institutional and academic leaders

  • Which graduate capabilities must students be able to demonstrate independently?
  • Which capabilities should now explicitly include effective and responsible AI use?
  • Where are important judgments about student capability based almost entirely on an unverified final artifact?
  • Where does existing assessment make student reasoning and decision-making visible?
  • Where are authenticated demonstrations of capability most important?
  • Are AI expectations coherent across courses within the same academic program?
  • Do students understand why different assessments have different expectations around AI?
  • Do faculty have the time, capability, guidance, and support required to redesign assessment thoughtfully?
  • Where across the program does evidence accumulate that the institution's stated graduate outcomes have actually been achieved?

There is still much we do not know

A rigorous response to generative AI also requires intellectual humility.

The technology is developing faster than conventional academic research cycles. Much of the current evidence comes from relatively recent experiments, natural experiments, individual disciplines, conceptual work, institutional case studies, and emerging models of assessment reform.

Findings from nursing cannot automatically be generalized to business, engineering, history, or the arts. Findings from one country or institutional system cannot simply be transferred everywhere. New AI models may also alter what is technically possible before longitudinal research has had time to evaluate the consequences.

The evidence is also not uniformly negative.

Experimental research indicates that well-designed AI-supported learning can improve academic performance and other learning outcomes. [7] OECD's synthesis likewise identifies meaningful potential when generative AI is used intentionally in support of pedagogy. [3]

The conclusion should therefore not be that AI prevents learning.

It is that institutions can no longer assume that a high-quality AI-assisted output is sufficient evidence of the learner's underlying capability.

That uncertainty is not a reason to wait.

It is a reason to experiment carefully, evaluate outcomes, develop faculty capability, examine assessment at program level, and avoid simplistic responses.

From detecting AI to demonstrating learning

Generative AI has exposed something that may always have been true about higher education assessment.

A completed assignment is not learning itself.

It is evidence from which educators infer learning.

For many years, the relationship between the two was sufficiently stable that institutions could often treat the submitted product as a reasonable representation of the student's intellectual work.

Generative AI has weakened that convenience.

But that disruption may ultimately be constructive.

It creates an opportunity for higher education to become more precise about the capabilities that matter, distinguish between independent and AI-augmented competence, make student reasoning more visible, strengthen assessment across whole programs, and build more credible connections between learning outcomes and the evidence used to judge them.

The future of assessment therefore does not need to revolve around proving that artificial intelligence was absent.

It should revolve around something more important:

The central principle

Designing credible evidence that learning is present.

AI may not make assessment impossible. It may simply make weak assessment harder to ignore.

Research Base

References

This Insight draws on recent peer-reviewed research and sector reports available at the time of publication.

  1. Lumina Foundation & Gallup. (2026). AI in Higher Education: Widespread Use, Unclear Rules. State of Higher Education research. View source
  2. Robert, J. (2026). The Impact of AI on Learning Assessment. EDUCAUSE. View source
  3. OECD. (2026). OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education. OECD Publishing. View source
  4. Fawns, T., Boud, D., & Dawson, P. (2026). Identifying what our students have learned: A framework for practical assessment validation. Assessment & Evaluation in Higher Education. View study
  5. Brattli, H., Utne, A., & Lynch, M. (2026). Assessment validity in the age of generative AI: A natural experiment. Informatics, 13(4), 56. View study
  6. Corbin, T., Dawson, P., & Liu, D. (2025). Talk is cheap: Why structural assessment changes are needed for a time of GenAI. Assessment & Evaluation in Higher Education, 50(7), 1087–1097. View study
  7. Deng, R., Jiang, M., Yu, X., Lu, Y., & Liu, S. (2025). Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies. Computers & Education, 227, 105224. View study
  8. Kofinas, A. K., Tsay, C. H.-H., & Pike, D. (2025). The impact of generative AI on academic integrity of authentic assessments within a higher education context. British Journal of Educational Technology, 56, 2522–2549. View study
  9. Lodge, J. M., Bearman, M., Dawson, P., Gniel, H., Harper, R., Liu, D., McLean, J., & Ucnik, L. (2025). Enacting Assessment Reform in a Time of Artificial Intelligence. Tertiary Education Quality and Standards Agency. View report
  10. Ogle, T., & Jarrett, P. (2026). Replacing written assessment with an oral ISBAR assessment in a 700-student cohort: Cohort-level outcomes for performance, integrity, and staff experience. Assessment & Evaluation in Higher Education. View study
  11. Miao, F., Shiohira, K., & Lao, N. (2024). AI Competency Framework for Students. UNESCO. View framework
  12. Walker, S. (2025). Trends in Assessment in Higher Education: Considerations for Policy and Practice. Jisc. View report
  13. Robert, J., Muscanell, N., McCormack, M., & Arnold, K. (2026). 2026 EDUCAUSE Horizon Report: Teaching and Learning Edition. EDUCAUSE. View report
For Higher Education Institutions

What Does Credible Evidence of Learning Look Like in Your Programs?

Riveral Advisory works with higher education institutions to examine academic programs, assessment approaches, faculty capability, and the implications of AI for teaching, learning, and institutional development.

Schedule a Conversation
Scroll