Speaker
Description
Designing an AI-Aware multimodal assessment rubric for technical presentations in an ESP course for engineering students
Nguyen Thi Thu Thuy
University of Languages and International Studies, Vietnam National University, Hanoi
Hanoi University of Civil Engineering
nguyenworkaholic@gmail.com
Abstract
The rapid uptake of generative AI is transforming how engineering students produce multimodal texts slides, diagrams, and AI-generated images for technical presentations yet existing assessment rubrics rarely account for work that is partly machine-created. This raises pressing questions of fairness, validity, and academic integrity: how can teachers assess multimodal communicative competence when authorship is shared between student and AI? This study designs and pilot-tests a multimodal assessment rubric for AI-assisted technical presentations in an English for Specific Purposes (ESP) course for engineering students. Drawing on multiliteracies pedagogy and validity theory, the rubric distinguishes students' communicative and design decisions from AI-generated content. It was developed and trialled with a cohort of undergraduate presentations and refined through teacher–student feedback. The study is expected to yield practical, transparent criteria for evaluating AI-mediated multimodal work while safeguarding academic integrity. For TESOL and ESP practitioners, it offers a classroom-ready instrument and design principles for AI-aware assessment, supporting fairer evaluation and stronger learner agency in technical English contexts.
Keywords: multimodal assessment, rubric design, AI-assisted learning, English for Specific Purposes, academic integrity
1. Introduction
The technical presentation is one of the most consequential genres an engineering graduate has to master in English: explaining a structural solution to a client, walking a contractor through a revised detail, defending a design choice to a panel. In every case the spoken word is only part of the message; the argument lives in the drawing on the screen, in the section enlarged at the right moment, in the sequence in which the slides unfold. Research on technical talk has shown for two decades that visuals in such presentations do not decorate the speech but carry propositional and argumentative work of their own, and that effective speakers orchestrate verbal, written and visual resources together rather than in sequence (Morell, 2015; Rowley-Jolivet, 2002). The implication for English for Specific Purposes (ESP) teaching is direct: a presentation assessed only for pronunciation, fluency and grammatical accuracy under-represents the competence that actually matters.
Within a very short period, generative artificial intelligence (GenAI) has entered the production pipeline of exactly this genre. Students now draft slide text with a chatbot, generate an illustrative image from a prompt, translate a Vietnamese explanation into English, or upload a drawing and ask for a summary of what it shows. None of these acts is exotic; several mirror what practising engineers already do. What has changed is the evidentiary status of the artefact the teacher receives. A polished English sentence on a slide no longer licenses an inference about the student's writing; a competent technical diagram no longer licenses an inference about the student's ability to select and adapt a visual for an audience.
The instruments used to judge these performances, however, have not moved. Most presentation rubrics in ESP courses including the one previously used in the course reported here describe a single human author. They contain criteria such as "slides are clear, well organised and visually appropriate" or "language is accurate and appropriate to the field". Read literally, such criteria now describe properties of an artefact whose origin is unknown. Applied honestly, they force the marker into one of two unsatisfactory positions: either to assume the work is entirely the student's, which rewards undisclosed delegation, or to suspect it is not, which turns assessment into an adversarial inquiry that the evidence cannot support.
The temptation to resolve this by detection should be resisted: independent testing of tools that claim to identify machine-generated text has found them neither accurate nor robust, with performance degrading sharply once text is lightly paraphrased (Weber-Wulff et al., 2023). Prohibition is equally unpromising. A ban is unenforceable at the artefact level, pedagogically perverse in a discipline where AI-assisted drafting is becoming ordinary professional practice, and it pushes use underground precisely the condition under which nothing can be assessed. International guidance has consequently moved towards a human-centred position in which institutions specify permitted use, require disclosure, and redesign tasks so that learning remains visible (Miao & Holmes, 2023).
The problem facing the ESP teacher, then, is not primarily a policing problem. It is a judgement problem, and it can be stated in one question: when I put a mark next to this presentation, what exactly am I giving credit for? This paper takes that question as its design brief. It reports the development of a multimodal assessment rubric for AI-assisted technical presentations in an ESP course for undergraduate engineering students, and it argues for a particular answer: credit belongs to the student's communicative and design decisions, including decisions about what to take from a machine, what to reject, what to verify and what to change, rather than to the surface properties of the artefact those decisions produced.
Three questions guided the work:
• RQ1. What should be assessed in an AI-assisted multimodal technical presentation in an ESP engineering course?
• RQ2. How can a rubric make the boundary between a student's own design decisions and AI-generated content visible, and how can that boundary be turned into assessable criteria?
• RQ3. What transferable design principles follow for AI-aware assessment in other ESP and TESOL settings?
Section 2 reviews the three literatures the design draws on. Section 3 sets out the conceptual framework and Section 4 the design procedure. Section 5 presents the instrument, Section 6 what the classroom trial showed about judging with it, and Sections 7 to 10 discuss validity, fairness and agency, distil design principles, and state limitations.
One clarification about scope is owed to the reader at the outset. This is a design study. It reports an instrument, the reasoning that produced it, and the revisions that classroom use forced upon it. It does not report score data, rater agreement indices or measured gains, and it makes no claim that presentations marked with this rubric are better or more reliably marked than presentations marked with another. Section 4.5 explains why that restraint is deliberate rather than merely a limitation.
2. Literature review
2.1 The technical presentation as a multimodal ESP genre
Social semiotic accounts treat meaning as distributed across modes, each with its own affordances and constraints (Kress, 2010; Kress & van Leeuwen, 2021). A plan view affords spatial relation in a way a sentence does not; a sentence affords causal qualification in a way a plan view does not. Competence in a multimodal genre is therefore not the sum of competences in separate channels but the capacity to combine them so that each does the work it is best suited to.
Genre studies of technical talk bear this out. Rowley-Jolivet (2002) analysed the visuals projected during conference papers in geology, medicine and physics and showed that they form a genre-specific repertoire performing recurrent rhetorical functions: structuring the discourse, expressing logical relations, and grounding claims in evidence. Morell (2015) found that effective presenters deploy verbal, written, non-verbal material and body language in overlapping combinations, that speakers in the technical sciences rely comparatively more on non-verbal materials than those in the social sciences, and that strong use of the visual modes can compensate for limitations in the verbal mode.
That last finding has a double edge. It is encouraging, because a student whose spoken English is still developing is not debarred from communicating a technical point well. It is also a warning, because compensation shades into concealment once the visual material is generated on request. Any rubric treating visual quality as an unanalysed good will, under AI conditions, reward the tool rather than the learner.
For civil engineering students in particular, the visual channel is not an accessory to the discourse but its substance. Construction drawings plans, sections, elevations, details, schedules are the field's primary technical texts, and the spoken presentation typically indexes them: it points, enlarges, compares, and narrates a path through them. An assessment that captures only what was said, and not how the speaker moved between saying and showing, misses the genre almost entirely.
2.2 Assessing digital multimodal composing
The pedagogy of multiliteracies supplies the vocabulary that makes such composing assessable. The New London Group (1996) framed meaning-making as Design, distinguishing Available Designs the semiotic resources a designer can draw on from Designing, the act of selecting and recombining them, and from the Redesigned that results. Cope and Kalantzis (2009) elaborated this into a pedagogy moving between experiencing, conceptualising, analysing and applying. The distinction tells an assessor where to look: the learning is in Designing, and the artefact is evidence of it rather than the object of interest in itself.
Hafner and Ho (2020) pursued this insight into practice. Studying an English for Science course in which students produced digital video documentaries, they proposed a process-based model in which assessment is planned across four stages pre-design, design, sharing and reflecting and combines formative and summative moments. Their teacher interviews are instructive here: teachers reported real difficulty in attending simultaneously to sound, image, fluency and accuracy while a performance unfolded, and found their rubric lacked exemplars for distinguishing strong from merely adequate work and omitted delivery features altogether. Belcher (2017) added that as multimodal composing becomes normal, teachers are pushed into the role of facilitators of digital design a role for which assessment practice has lagged behind pedagogy.
The rubric literature itself offers design discipline. Jonsson and Svingby (2007) concluded that rubrics improve scoring consistency principally when criteria are explicit and raters are supported by exemplars and training, not merely by the existence of a grid, while Panadero and Jonsson (2013) documented the formative effects of transparent criteria on self-regulation. Dawson (2017) set out fourteen design elements among them specificity, secrecy, exemplars, scoring strategy, judgement complexity, users and uses, and explanation that distinguish one rubric from another. Any serious claim to have designed a rubric should state where it sits on those dimensions, and Section 5.6 does so.
Rubrics are also not only measuring devices. Sadler (1989) argued that students cannot improve work whose quality they cannot appraise, and Tai et al. (2018) developed this into evaluative judgement: the capability to decide about the quality of work of oneself and others, proposed as a goal of higher education in its own right. A rubric that students see, discuss and apply builds that capability rather than merely recording its absence.
2.3 Generative AI, authorship and academic integrity
The first wave of responses to GenAI was dominated by misconduct framing. Perkins (2023) surveyed the academic integrity implications of large language models, and Cotton et al. (2024) articulated the challenge of maintaining integrity when text of acceptable quality can be produced on demand. Empirical work on detection has since narrowed the options: Weber-Wulff et al. (2023) tested many public and commercial detectors and found them unreliable, easily defeated by paraphrase or machine translation, and biased towards classifying text as human-written, concluding that they cannot bear the evidentiary weight disciplinary processes would place on them.
Attention has consequently shifted from detection to design. The AI Assessment Scale (AIAS) asks the teacher to specify, for each assessment, the level of AI involvement permitted, from tasks in which no AI is used through to tasks in which its use is expected and explored (Perkins, Furze, et al., 2024); a pilot implementation followed (Furze et al., 2024), and a refined version has since been published (Perkins, Roe, & Furze, 2025). Its contribution is to replace an unenforceable general prohibition with an explicit, task-level permission that students and teachers share.
The AIAS is a permission instrument, however, not a credit instrument. Knowing that a task sits at a collaborative level tells a student what is allowed; it does not tell a marker how much a competently selected and corrected AI image is worth beside a student's own photograph of a site defect. UNESCO's guidance points the same way without closing the gap, urging human-centred governance and transparency of use while leaving the criterion-level question to practitioners (Miao & Holmes, 2023). That gap between permission and credit is where this study is located.
2.4 Validity and fairness when authorship is shared
Validity theory provides the terms in which the gap can be described precisely. Messick (1989) established that validity is a property of the interpretations and uses of scores rather than of instruments, and identified construct-irrelevant variance and construct under-representation as the two principal threats. Kane (2013) developed the argument-based approach, in which the claims implicit in a score interpretation are laid out as a chain of inferences scoring, generalisation, extrapolation and implications each resting on examinable assumptions. The Standards (AERA et al., 2014) situate fairness as a validity concern rather than a separate virtue.
Undisclosed AI assistance is, in these terms, a textbook source of construct-irrelevant variance: it inflates observed performance for reasons unrelated to the construct the score should represent. The converse error is just as damaging. A rubric that refuses to credit anything a machine touched under-represents the construct, because selecting, evaluating and correcting machine output is now part of competent technical communication. Bearman et al. (2024) make this case directly: under generative AI, evaluative judgement becomes more central rather than less. Dawson (2021) adds the institutional argument that assessment security is best pursued through design rather than surveillance.
2.5 The gap addressed by this study
These three literatures do not yet meet. Work on assessing digital multimodal composing was largely developed before generative tools were widely available, and its process models, though highly compatible with the present problem, do not address machine authorship. Work on AI and assessment operates at task and policy level, specifying what is permitted rather than what is credited. ESP presentation assessment remains, in most instruments, oriented to language and delivery. What is missing is an instrument at criterion level that is simultaneously multimodal, AI-aware and defensible as a basis for judgement.
3. Conceptual framework
The design rests on two pillars and three commitments, shown in Figure 1.
Figure 1. Conceptual framework of the study.
From multiliteracies pedagogy the design takes its object: the assessable thing is Designing, the student's act of selecting and orchestrating available resources for a communicative purpose (Cope & Kalantzis, 2009; New London Group, 1996). Under AI conditions this framing does unexpectedly useful work, because a generative model is best understood not as a co-author but as an unusually productive supplier of Available Designs. A generated image, a drafted sentence, a suggested structure — each is a resource offered to the designer, and none is a design. The question the marker must answer is therefore unchanged in form: what did this student do with what was available?
From validity theory the design takes its discipline: every criterion must be answerable to the inferences it will be asked to support (Kane, 2013; Messick, 1989). This forces a concreteness that rubric design often lacks. A criterion is not justified because it names something desirable; it is justified because a judgement made against it supports a claim someone will act on.
Three commitments follow, and they govern every subsequent decision in the instrument:
• Transparency of provenance. The origin of each element of the presentation is declared before it is judged. Provenance is treated as a condition of entry rather than as something to be inferred, suspected or scored.
• Credit for decisions, not artefacts. Descriptors are written so that they interrogate what the student chose, checked, rejected and changed. Surface quality alone cannot lift a performance into a higher band.
• Evidence across the process. Judgement draws on traces from several points in the task cycle rather than on the performance alone, following the process-based logic of Hafner and Ho (2020).
The construct these commitments serve can be stated as follows: AI-aware multimodal communicative competence in technical English is the capacity to construct, deliver and defend a technical message for a specified professional audience by orchestrating spoken language, written text and graphic-symbolic resources, while selecting, evaluating, correcting and honestly accounting for any machine-generated material used in the process. Table 1 sets out its dimensions and the evidence that bears on each.
Table 1. Dimensions of the construct and their evidentiary basis
Dimension What it covers Where the evidence comes from
Technical message Accuracy, relevance and depth of the engineering content; appropriate scoping for the stated audience The talk itself; responses to unscripted questions; planning notes
Spoken technical English Intelligibility, control of field terminology, register, signposting, interaction with the audience The talk; the question-and-answer exchange; rehearsal recordings
Multimodal orchestration Fit between what is shown and what is said; use of drawings and images to carry argument rather than to decorate Slides and drawings; the delivery; successive draft versions
Critical use of AI Purposeful selection of tools; verification of generated content against technical sources; correction and adaptation Prompt and output log; comparison of drafts; the reflection; questioning
Integrity and attribution Completeness and specificity of the account of authorship, including third-party technical sources The AI-use plan and the disclosure statement, read against the artefacts
Note. The dimensions correspond to domains D1 to D5 of the rubric presented in Section 5.
3.1 What the five domains do not cover
The correspondence between construct and domains is deliberate but not complete. The domains cover what a short, individually delivered, slide-based presentation can evidence; several capabilities belonging to the construct as stated fall outside them. Collaborative design the negotiation through which a project team arrives at what to show is invisible in an individual task. Sustained interaction of the kind a site meeting demands is only briefly sampled in the question stage. Listening, in the sense of reconstructing a technical argument someone else is making, is not assessed at all, nor is the written documentation that normally accompanies a technical talk.
This is construct under-representation in Messick's (1989) sense, and naming it is not a formality. A user of the rubric is entitled to know that a high mark supports a claim about this student's capacity to design and defend a technical message alone, in English, with preparation time not a claim about their communicative competence in an engineering workplace as a whole. Where a programme needs the wider claim, the instrument belongs in a sequence of tasks rather than standing alone.
4. Design Approach
4.1 An educational design study
The work follows the logic of educational design research: a practical problem is analysed, a prototype intervention is built from theory, it is put into use, and the difficulties encountered in use are treated as data about the design rather than as noise. The output of such a study is an artefact together with the reasoning that justifies it and a set of design principles that may transfer. Figure 2 shows the four phases followed here.
Figure 2. The four-phase design cycle followed in developing the rubric.
4.2 Setting and task
The setting is an ESP course for undergraduate students at a university of civil engineering in Hanoi. Students are preparing to read, discuss and present technical material in English within their discipline; many will work with international consultants or suppliers, and most meet English primarily through drawings, standards and manufacturers' documentation rather than through general texts.
The trial drew on a single intact class group of 25 second-year students, all of whom had completed the programme's general English requirement. They worked in five groups, producing five presentations marked with the draft instrument. No selection was applied: this was the class the researcher taught in the semester concerned, and every student in it completed the task.
The trial drew on a single intact class group of 25 second-year students, all of whom had completed the programme's general English requirement. They worked in five groups, producing five presentations marked with the draft instrument. No selection was applied: this was the class the researcher taught in the semester concerned, and every student in it completed the task.
Five presentations is a small set, and it is worth saying plainly what such a set can and cannot support. It cannot show how the descriptors behave across a range of performances, and the difficulties reported in Section 6 are difficulties that arose rather than difficulties whose frequency is known. What it does support is a claim about the instrument rather than about students: whether the wording decided cases, and where it did not. The class size also bears on Section 7.6 the full task cycle described here was designed for, and run in, a class of this size, so the reductions proposed there for larger cohorts are reasoned extrapolation rather than tested practice.
4.3 Sources of design evidence
Four sources informed the successive versions of the instrument. First, an audit of the course's previous presentation rubric and of published rubrics and frameworks for multimodal composing and AI-aware assessment. Second, consultation with colleagues teaching on the course about the criteria they found hardest to apply. Third, the trial itself: the rubric was released to students in advance and used to mark the presentations. Fourth, the marker's own record of the points at which judgement stalled. That last source proved the most productive, and Section 6 is largely built from it, so the way it was produced deserves description.
Marking was carried out with the draft rubric open in a document that carried, beside each domain, an empty field headed “where the descriptor did not decide the case”. Whenever the marker hesitated that is, whenever the wording in front of her admitted two defensible placements she stopped, wrote one or two sentences naming the difficulty and the presentation it arose from, placed the work provisionally, and continued. This produced a running log of entries spanning the whole set.
After marking, the entries were sorted by the kind of difficulty they named rather than by the domain in which they surfaced: a mismatch between artefact and speech, a term the speaker could not gloss, an ambiguity in the declaration, a case the descriptors had not anticipated. Categories were derived from the entries rather than fixed in advance, and revised twice as the sorting proceeded; six accounted for the log and correspond to the six subsections of Section 6. A category was treated as calling for revision when more than one entry fell into it and when the difficulty was attributable to the wording of the instrument rather than to the quality of the work. That second test excluded a small number of entries recording ordinary borderline judgements of the kind any rubric produces.
Feedback from students and colleagues was gathered less formally and is reported as such. One debriefing session of about 30 minutes was held with the class after marks were returned, in which students were asked what writing the declaration had felt like and whether the criteria matched what they had actually spent their effort on; the researcher took written notes during and immediately after. Two colleagues teaching the same course read the draft rubric together with the log of difficulties and commented in writing. Where a revision described in Section 6 followed from these conversations rather than from the marking log, the text says so.
4.4 Ethical considerations
Three provisions were adopted. Disclosure was made non-punitive: declaring the use of a tool could not by itself reduce a mark, and students were told so in writing before the task. The AI-use plan was agreed in advance so that the declaration was prospective planning rather than retrospective confession. And a fully viable no-AI route was preserved: a presentation produced without any generative tool could reach the highest band in every domain, including D4, which is then judged on the reasoning behind the decision not to use one. This also addresses equity, since access to paid tools is not evenly distributed.
4.5 Reporting stance
This paper reports no score data, for methodological rather than presentational reasons. After one classroom cycle, the validity argument concerns the plausibility of its assumptions, not the accumulation of evidence for claims the design is not yet entitled to make (Kane, 2013). Score distributions, agreement indices or gain comparisons drawn from a single cohort marked by the instrument's designer would convey a precision the design does not possess, and would invite readers to treat as findings what are properly hypotheses. What the paper offers instead is an instrument specified in enough detail to be criticised, adopted or tested by others, and an account of where its descriptors failed under use.
5. The instrument
5.1 Architecture
The rubric has three layers, shown in Figure 3. Layer A is a disclosure gateway through which the work passes before it is judged. Layer B contains the five assessed domains, each with four quality levels. Layer C is the base of process evidence on which judgements in Layer B draw. The layering is itself a design claim: provenance is a precondition of interpretation, not a criterion competing with others for weight.
Figure 3. The three layers of the rubric and their relationship.
5.2 Layer A: the disclosure gateway
Before marking, the student submits three short items: the AI-use plan agreed at the start; a declaration of the tools used, their purposes and the material they produced; and a statement of what the student changed in that material and why. The third does most of the work. A declaration that a model generated an image tells the marker very little; a statement that the generated image showed a reinforcement arrangement not matching the detail under discussion, and was therefore replaced with a redrawn figure, tells the marker a great deal — and is itself evidence for D4.
Nothing in Layer A is scored. This is the design's most counter-intuitive feature and its most important one. If disclosure carries marks, students optimise the disclosure; if it is penalised, they suppress it; either way the marker loses the information the judgement depends on. Making it a condition of submission aligns the student's interest with the marker's need. Table 2 gives the provenance categories used.
Table 2. Provenance categories used in the disclosure statement
Category Description What the student must show
Student-authored Text, drawing, photograph or slide produced by the student without generative assistance Nothing beyond the artefact itself
Machine-generated, used as produced Content taken from a generative tool and included essentially unchanged The tool, the request made of it, and the reason the output was judged fit for purpose
Machine-generated, substantially reworked Content originating with a tool but altered in substance by the student The original output alongside the final version, and an account of the change
Machine-assisted student work Student content on which a tool gave feedback, correction, translation or language support The nature of the support and which suggestions were accepted or declined
Third-party technical source Standards, manufacturers' drawings, published details, site photographs by others A conventional citation and permission status where relevant
Note. Categories are applied element by element, not to the presentation as a whole; a single talk will normally involve several of them.
5.3 Layer B: the five assessed domains
The five domains are set out in Table 3. Three points about their wording deserve emphasis, because they are where the instrument departs from a conventional presentation rubric.
First, domain D3 is written around fit rather than quality. The question is not whether a slide is attractive but whether what is shown and what is said do different and complementary work. This reorientation was forced by the trial (Section 6), and it is the change that most affects how the rubric behaves in the presence of machine-generated visuals: a generated image that is beautiful and irrelevant scores lower than a rough hand-redrawn section that carries the argument.
Second, domain D4 credits scepticism. The highest band is reached not by producing good output but by testing output against technical sources, catching where it is wrong, and acting on what is found. Engineering is an area where generative models fail plausibly: a generated detail may be geometrically coherent and constructionally impossible. Rewarding verification builds the disposition the discipline requires.
Third, domain D5 judges honesty and specificity, not abstinence. A student who declares extensive, well-reasoned use in precise terms sits in a higher band than one who declares vaguely, and far higher than one whose declaration is contradicted by the artefacts.
The four levels are labelled Accomplished, Competent, Developing and Emerging, avoiding the language of pass and failure because the rubric is used formatively at rehearsal as well as summatively after delivery.
5.4 Layer C: evidence across the task cycle
A single short performance is a thin evidentiary base at the best of times; under AI conditions it is thinner still, because its most polished parts are the least diagnostic. Following Hafner and Ho (2020), judgement is distributed across the cycle. Figure 4 shows which evidence is gathered when, and which domains it informs.
Figure 4. Assessment points distributed across the task cycle.
The unscripted question stage carries a disproportionate share of the load. Asking a student to explain, without slides, why a particular section was chosen for enlargement, or to gloss a term from a slide, produces evidence no artefact can supply and no tool can pre-empt. It is where the extrapolation inference is most directly supported, and it is cheap: a few minutes per student, no technology.
5.5 Relationship to the AI Assessment Scale
The rubric presupposes that the task's permitted level of AI use has been settled, and the AIAS is a convenient way to settle it (Perkins, Furze, et al., 2024; Perkins, Roe, & Furze, 2025). The task here sits at a collaborative level: tools may be used throughout preparation provided use is declared and the student can account for it. The two instruments are complementary, as Table 4 sets out.
Table 4. Division of labour between the permission framework and the rubric
Question Answered by the AIAS level Answered by the rubric
May a student use a generative tool for this task? Yes, throughout preparation, at the collaborative level set for the task Not addressed
What must the student tell the teacher? That use must be declared The categories, level of detail and format of the declaration (Layer A)
What earns credit in the finished work? Not addressed The five domains and their level descriptors (Layer B)
On what evidence is the judgement made? Not addressed Artefacts and process traces across the cycle (Layer C)
5.6 Design decisions in Dawson's terms
Stating design parameters explicitly makes a rubric criticisable, which is the point of Dawson's (2017) framework. The scoring strategy is analytic, with domains reported separately and no compensatory total. Specificity is mixed: D1 to D3 travel to other technical presentation tasks, while the accompanying exemplars are task-specific. There is no secrecy: the rubric is released with the brief and applied in class to two contrasting samples, which also serves the evaluative-judgement purpose identified by Tai et al. (2018). Judgement complexity is deliberately high in D4 and D5, which require reading artefacts against declarations rather than observing and ticking; this is a cost, and Sections 6.6 and 7.6 treat it as one.
6. What the trial showed
The rubric was released to a class group, used to mark their presentations, and discussed with students and colleagues afterwards. What follows is an account of where judgement proved difficult and what was changed in response. It is offered as design evidence, not as findings about students.
6.1 Polish without fit
The largest category in the log was the presentation whose slides were visibly beyond the standard the course had previously seen clean typography, coherent colour, well-composed generated illustrations accompanied by a spoken account that did not correspond to them. The speaker described one thing while the screen showed another, or referred to a figure without indicating what in it mattered. Under the previous rubric such work scored well on slide quality and poorly on delivery, producing a middling mark that described nothing real. D3 was rewritten so that its object is the relation between showing and saying: it is now impossible to score highly on multimodal design while the two channels pull apart, however accomplished either is alone.
6.2 Terminology that could not be glossed
A second category concerned technical vocabulary that appeared on slides or in scripted passages and that students could not explain when asked. This is not new in ESP, but generated text raises the sophistication of the borrowed material and lowers the effort of borrowing it. The question stage identified these cases reliably and without accusatory framing: a request to say the same thing in other words settles the matter in seconds. D1 and D2 were amended to refer explicitly to unscripted explanation, so that what the questioning revealed had a place in the record rather than merely colouring the marker's impression.
6.3 Disclosure read as confession
Students' first reaction to the declaration requirement was anxiety, despite the written assurance that disclosure could not reduce a mark; several asked whether it would be better to use nothing. The remedy was not a better form but a change of sequence: moving the declaration forward into a plan agreed before preparation began. Once the document was something to be planned rather than confessed, the tone of the conversation changed. This revision came from the debriefing session rather than from the marking log. It is a point about framing and teacher talk rather than instrument design, and it may be the most transferable observation in this section.
6.4 Translation as scaffold and translation as substitute
Machine translation from Vietnamese formed a category of its own. For some students it was a scaffold: the technical thinking was done in Vietnamese, translated, then worked over terms checked against field usage, sentences shortened for speaking. For others it was a substitute: the English arrived fully formed and was read aloud. The artefacts alone do not distinguish these; the process evidence and the question stage do. Table 2 was extended to make language support a separate declarable category, and D4 written so that reworking a translation is creditable in the same way as reworking a generated image.
6.5 The no-AI route
Some students chose not to use generative tools at all, by preference or because of cost. An early draft of D4 had no home for this choice: the descriptors assumed AI use and could only record its absence as a gap. That was unfair and pedagogically wrong, since deciding that a tool is not worth using is itself a judgement of the kind the domain rewards. D4 was rewritten so that its highest band is reachable either by well-evidenced critical use or by a reasoned, demonstrated decision not to use, evidenced in the plan and the reflection. Both colleagues who read the draft independently raised this gap, and the students who had chosen not to use tools raised it in the debriefing.
6.6 Marker workload
Reading artefacts against declarations takes longer than ticking observable features, and the additional time was concentrated in D4 and D5. Two mitigations are now in the guidance notes: capping the declaration at one page, which forces students to be specific, and marking D1 to D3 live during delivery while reserving D4 and D5 for a short reading afterwards. Whether the remaining overhead is acceptable is a judgement each teacher must make; the trial did not establish that it is.
7. Discussion
7.1 The validity argument
Figure 5 lays out the four inferences that a mark from this rubric is asked to support, the threat each faces when authorship is shared, and the feature of the design that responds to it.
Figure 5. Inferences, threats and the design responses built into the rubric.
Two things should be said about it. It is an argument, not a result: each response is a reasoned attempt to make an assumption plausible, and none has been tested. And the responses interact the disclosure gateway is what makes the decision-focused descriptors usable, since a descriptor about what the student changed is unanswerable without a record of what was there before. Remove Layer A and Layer B degrades into a conventional rubric applied hopefully.
7.2 Fairness
Fairness runs in more than one direction here, which is why the Standards treat it as an aspect of validity rather than a separate box to tick (AERA et al., 2014). Undisclosed use advantages some students over others; that is the familiar concern. But a disclosure regime creates its own risks. Students differ in access to capable tools, which is why the no-AI route must be genuinely equivalent rather than nominally permitted. And disclosure asks students to reveal their working in a way that could feel intrusive, which is why nothing in Layer A is scored.
A second source of construct-irrelevant variance deserves separate treatment, because the design itself creates it. The declaration is written, and Appendix A asks for most of it in English. A student who writes English fluently can present a modest amount of critical work in convincing detail; a student whose writing is weaker may have done more and shown less. Since D4 and D5 are judged by reading the artefacts against the declaration, the quality of the declaration's prose can contaminate two domains that are not measuring writing at all. Nothing in the design removes this risk. Three partial responses are built in: the plan may be written in Vietnamese; the declaration is read for content, with its language accuracy explicitly excluded from the mark; and the question stage offers a spoken route to the same evidence, so a student whose written account is thin can still demonstrate in speech what was checked and changed. Whether these suffice is an empirical question this trial cannot answer, and a marker adopting the instrument should watch for the tell-tale pattern the strongest writers scoring highest on D4 as a sign that they do not.
7.3 Agency
The design's most substantial pedagogical claim is that it puts the right thing in front of students. If evaluative judgement is a graduate capability (Tai et al., 2018), and if generative tools make it more rather than less central (Bearman et al., 2024), then a rubric requiring students to say what they took from a machine, what they rejected, and on what grounds teaches the capability while assessing it. The reflection is not an administrative appendix; it is one of the places where the learning happens.
7.4 Integrity reframed
The instrument treats integrity as a positive competence to be demonstrated rather than an offence to be detected consistent with international guidance (Miao & Holmes, 2023), with the case for pursuing assessment security through design (Dawson, 2021), and with the evidence on detection tools (Weber-Wulff et al., 2023). It does not make dishonesty impossible; a student may still declare falsely. What it changes is what dishonesty requires: a false declaration must be sustained against the artefacts, the drafts and the questions, and must be made explicitly rather than by silence.
7.5 Washback on task design
Adopting the rubric is not a drop-in change. It obliges the teacher to collect process artefacts, to run an unscripted question stage, and to release and teach the criteria building checkpoints into the schedule and spending at least one session on assessment literacy rather than content. Each is defensible on independent pedagogical grounds, but together they amount to a redesign of the task cycle rather than a new marking sheet.
7.6 Feasibility in constrained settings
Section 6.6 reported that the instrument costs marking time. The wider question is whether the model survives conditions common in ESP teaching in Vietnam and the region: a large class, one teacher without an assistant, no technical support, and a marking window measured in days. Under those conditions the task cycle as described here is not realistic, and it would be misleading to present it as though every reader's context resembled the one it was built in.
Three reductions preserve most of the design's logic at lower cost. The question stage can be shortened to a single question per speaker, asked of everyone and drawn from a short standing list; this retains the extrapolation evidence, the element least replaceable by anything else, while bounding the time it takes. The process-evidence base can be reduced to two artefacts the AI-use plan and one intermediate draft rather than the full set shown in Figure 4, since the marginal value of further traces falls quickly once a before-and-after comparison is possible. And peer assessment at the rehearsal stage can carry part of the formative load, which the design wants in any case on evaluative-judgement grounds (Tai et al., 2018).
What cannot be reduced without losing the design is the disclosure gateway: it costs almost nothing to collect and is what makes every decision-focused descriptor interpretable. A teacher with time for one element only should keep Layer A and simplify Layer C.
7.7 What the instrument asks of the teacher
The rubric assesses students' critical use of generative tools, and in doing so it quietly assumes a corresponding capability in the person marking it. To judge D4 the marker must recognise the kinds of error these tools characteristically make in this subject matter: a reinforcement arrangement drawn coherently that cannot be built, a standard cited with a plausible but non-existent clause number, a confident gloss of a term that the local code uses differently. A marker who cannot see such errors will credit a student who did not catch them, which inverts the purpose of the domain.
This is a prerequisite rather than a peripheral concern, and it shapes how the instrument should be introduced. A department adopting it would do well to spend part of a staff meeting generating output in the relevant topic areas and examining where it fails, and to keep a shared file of the examples found, which serves both as marker preparation and as classroom exemplars. Teachers are not obliged to be expert users of these tools; they are obliged to be informed sceptics, which is precisely what the rubric asks of students.
8. Design principles
Seven principles generalise beyond the particular instrument. They are offered as hypotheses for other designers to test rather than as established rules.
Table 5. Design principles for AI-aware multimodal assessment
Principle What it means in practice
1. Make provenance a condition, not a criterion Require declaration before marking and give it no weight in the mark. Scoring disclosure corrupts it; penalising it suppresses it.
2. Assess decisions, not artefacts Write every descriptor so that it asks what the student selected, checked, rejected or changed. Surface quality alone must not lift a band.
3. Distribute the evidence Gather traces from planning, drafting, rehearsal, delivery and reflection. A single performance is the least diagnostic evidence available.
4. Keep an unscripted moment Reserve time for live questioning that no artefact can anticipate. It is the cheapest and strongest support for extrapolation.
5. Keep the no-AI route fully viable Every band must be reachable without generative tools, with the decision not to use them creditable as a judgement in its own right.
6. Publish the rubric and teach it Release criteria with the brief and spend class time applying them to samples, so that the instrument builds evaluative judgement as well as recording it.
7. Prepare the marker before adopting the instrument A marker who cannot recognise how generative tools fail in the subject matter cannot judge critical use. Treat marker preparation, and a shared file of failure examples, as part of adoption.
For practitioners outside engineering, the principles transfer more readily than the domains. A business English course assessing client pitches, a nursing programme assessing handover simulations or a tourism course assessing site commentaries would each need different content descriptors in D1 and different visual conventions in D3, but the three-layer architecture, the provenance categories and the principles above apply unchanged.
9. Limitations and directions for further work
The limitations are substantial and should be read as a research agenda rather than an apology. The instrument has been developed and used in one course at one institution, with a single teacher who is also its designer; the risk that its descriptors work because their author knows what she meant by them is real and unaddressed. No reliability evidence exists, and the domains most likely to strain agreement D4 and D5, which require interpretation rather than observation are exactly the novel ones. No comparative claim can be made about the quality of judgements produced with it, and student perceptions were gathered informally.
One cheap step towards usability evidence was available and was not taken: a colleague applying the rubric independently to a small number of recorded presentations, keeping the same log of difficulties, would show whether the descriptors work for someone who did not write them, even without yielding any index of agreement. That is the immediate next move rather than a distant one. Four fuller studies would address the remaining gaps. A multi-rater trial, including markers who did not write the descriptors, would establish whether the domains can be applied consistently and where they diverge. A think-aloud study of marking would show what raters actually attend to when reading a disclosure against an artefact. A study of student perceptions, conducted by someone other than the teacher, would test whether the declaration is experienced as planning or as surveillance. And a transfer study in another discipline would separate what is general in the architecture from what is particular to construction drawings. Only after such work would claims about validity in Kane's (2013) fuller sense be reasonable; until then, Section 7 stands as a statement of what would need to be true.
10. Conclusion
Generative AI has not created a new assessment problem so much as exposed an old evasion. Rubrics for multimodal work have often credited the artefact and hoped it stood for the student. That hope is no longer tenable, and its collapse is an opportunity to say more precisely what an ESP course develops: not the capacity to produce a clean slide, but the capacity to decide what a technical audience needs to see, to find or make it, to judge whether it is right, and to stand behind it under questioning.
The rubric presented here is one attempt to write that down in a form a teacher can use. It makes provenance a condition of judgement rather than an object of suspicion, credits decisions rather than surfaces, and spreads its evidence across a task cycle. It has been used, it failed in specific and instructive ways, and it was changed in response. What it has not been is tested by anyone other than its author, and this paper has tried to be exact about that boundary. The instrument is offered in that spirit: detailed enough to be argued with, and specific enough to be improved.
References
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Bearman, M., Tai, J., Dawson, P., Boud, D., & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. Assessment & Evaluation in Higher Education, 49(6), 893–905. https://doi.org/10.1080/02602938.2024.2335321
Belcher, D. D. (2017). On becoming facilitators of multimodal composing and digital design. Journal of Second Language Writing, 38, 80–85. https://doi.org/10.1016/j.jslw.2017.10.004
Cope, B., & Kalantzis, M. (2009). “Multiliteracies”: New literacies, new learning. Pedagogies: An International Journal, 4(3), 164–195. https://doi.org/10.1080/15544800903076044
Cotton, D. R. E., Cotton, P. A., & Shipway, J. R. (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 61(2), 228–239. https://doi.org/10.1080/14703297.2023.2190148
Dawson, P. (2017). Assessment rubrics: Towards clearer and more replicable design, research and practice. Assessment & Evaluation in Higher Education, 42(3), 347–360. https://doi.org/10.1080/02602938.2015.1111294
Dawson, P. (2021). Defending assessment security in a digital world: Preventing e-cheating and supporting academic integrity in higher education. Routledge.
Furze, L., Perkins, M., Roe, J., & MacVaugh, J. (2024). The AI Assessment Scale (AIAS) in action: A pilot implementation of GenAI-supported assessment. Australasian Journal of Educational Technology, 40(4).
Hafner, C. A., & Ho, W. Y. J. (2020). Assessing digital multimodal composing in second language writing: Towards a process-based model. Journal of Second Language Writing, 47, Article 100710. https://doi.org/10.1016/j.jslw.2020.100710
Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational Research Review, 2(2), 130–144. https://doi.org/10.1016/j.edurev.2007.05.002
Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
Kress, G. (2010). Multimodality: A social semiotic approach to contemporary communication. Routledge.
Kress, G., & van Leeuwen, T. (2021). Reading images: The grammar of visual design (3rd ed.). Routledge.
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). Macmillan.
Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
Morell, T. (2015). International conference paper presentations: A multimodal analysis to determine effectiveness. English for Specific Purposes, 37, 137–150. https://doi.org/10.1016/j.esp.2014.10.002
New London Group. (1996). A pedagogy of multiliteracies: Designing social futures. Harvard Educational Review, 66(1), 60–92.
Panadero, E., & Jonsson, A. (2013). The use of scoring rubrics for formative assessment purposes revisited: A review. Educational Research Review, 9, 129–144. https://doi.org/10.1016/j.edurev.2013.01.002
Perkins, M. (2023). Academic integrity considerations of AI large language models in the post-pandemic era: ChatGPT and beyond. Journal of University Teaching and Learning Practice, 20(2). https://doi.org/10.53761/1.20.02.07
Perkins, M., Furze, L., Roe, J., & MacVaugh, J. (2024). The Artificial Intelligence Assessment Scale (AIAS): A framework for ethical integration of generative AI in educational assessment. Journal of University Teaching and Learning Practice, 21(6), 49–66. https://doi.org/10.53761/q3azde36
Perkins, M., Roe, J., & Furze, L. (2025). Reimagining the Artificial Intelligence Assessment Scale (AIAS): A refined framework for educational assessment. Journal of University Teaching and Learning Practice, 22(7). https://doi.org/10.53761/rrm4y757
Rowley-Jolivet, E. (2002). Visual discourse in scientific conference papers: A genre-based study. English for Specific Purposes, 21(1), 19–40. https://doi.org/10.1016/S0889-4906(00)00024-7
Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18(2), 119–144. https://doi.org/10.1007/BF00117714
Tai, J., Ajjawi, R., Boud, D., Dawson, P., & Panadero, E. (2018). Developing evaluative judgement: Enabling students to make decisions about the quality of work. Higher Education, 76(3), 467–481. https://doi.org/10.1007/s10734-017-0220-3
Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, Article 26. https://doi.org/10.1007/s40979-023-00146-z
Appendix A: Student-facing AI-use plan and disclosure statement
Submitted on one page. Nothing on this page is scored; a complete page is a condition of marking. Part 1 may be written in Vietnamese; Parts 2 and 3 are read for content, not for language accuracy.
Part When What to write Guidance
1. Plan Agreed before preparation begins Which tools you intend to use, and for what. What you will do yourself. If you intend to use none, why. This is planning, not confession. Nothing here can reduce your mark, and choosing to use no tools costs you nothing.
2. Declaration Submitted before you present For each element — each slide, image, drawing, section of script — its provenance category: student-authored; machine-generated, used as produced; machine-generated, substantially reworked; machine-assisted student work; third-party technical source. Work element by element, not presentation by presentation. A list is enough; sentences are not required.
3. Changes Submitted before you present Two things a tool produced that you changed or rejected: what was wrong with them, how you checked, and what you did instead. This is the part that earns credit in D4. Name the technical source you checked against. Keep the original output so it can be compared.
Note. Total length: one page. Bring the original tool outputs referred to in Part 3 to the presentation.
Table 3. The AI-aware multimodal assessment rubric for ESP technical presentations
Domain Accomplished Competent Developing Emerging
D1 Technical content and accuracy Accurate, well scoped for the stated audience and developed in depth. Claims grounded in named technical sources. Unscripted questions answered with confident, correct detail. Accurate and appropriately scoped, with reasonable depth. Sources mentioned. Questions answered correctly, though sometimes only at the level given on the slides. Content is broadly correct but thin, over-general or loosely scoped. Sources are vague. Some questions cannot be answered beyond what was shown on screen. Content contains technical errors, or is copied at a level the speaker cannot explain. Questions about substance are not answered.
D2 Spoken technical English Intelligible and well paced throughout. Field terminology accurate and reworded on request. Register suits the audience. The speaker signposts, responds and repairs comfortably. Generally intelligible with occasional strain. Terminology is mostly accurate and can usually be glossed. Signposting is present. Interaction is managed with some hesitation. Intelligibility varies and depends on the slides. Terminology is used but not always understood when questioned. Delivery is largely read. Interaction is minimal. Delivery is read or recited with little intelligible spontaneous speech. Terms cannot be explained in other words. Questions are not engaged with.
D3 Multimodal design and orchestration What is shown and what is said do different, complementary work. Drawings and images carry the argument; the speaker directs attention within them. Sequence and emphasis are purposeful. Visual and verbal channels are aligned and generally support each other. Key figures are pointed to and explained. Some slides carry more decoration than argument. Alignment is inconsistent: some figures are displayed without being used, or the talk restates what is already written. Attention is not directed within complex graphics. Visual and verbal channels diverge. Figures are unexplained, irrelevant or merely decorative, however polished. The talk could be given without them.
D4 Critical and purposeful use of AI Tools chosen for stated purposes. Generated content checked against technical sources; errors found, corrected or rejected with reasons. Alternatively, a reasoned and evidenced decision not to use tools. Tools are used purposefully and some checking is evident. At least one substantive correction or rejection is described, though the grounds are stated briefly. Tools are used without a clear purpose, or output is accepted with only cosmetic change. Checking is asserted rather than shown. Generated content is used as produced, with no evidence of checking; technical errors introduced by a tool remain in the work.
D5 Integrity and attribution Every element assigned a provenance category. The account of changes is specific and confirmed by the drafts and artefacts. Third-party technical sources cited conventionally. All substantial elements are assigned a category. The account of changes is accurate but general. Most sources are cited. The declaration is partial or imprecise, leaving the origin of some elements unclear. Sources are named inconsistently. The declaration is absent, or is contradicted by the artefacts or by what the speaker says under questioning.
Note. Domains are reported separately; no compensatory total is calculated. The rubric is released to students with the task brief and is applied by students to sample presentations before their own preparation begins.