Nov 27 – 28, 2026
Asia/Ho_Chi_Minh timezone

Designing an AI-Aware multimodal assessment rubric for technical presentations in an ESP course for engineering students

Not scheduled
20m
Agency, Literacy & Assessment

Speaker

Thuỷ Nguyễn Thị Thu

Description

Designing an AI-Aware multimodal assessment rubric for technical presentations in an ESP course for engineering students
Nguyen Thi Thu Thuy
University of Languages and International Studies, Vietnam National University, Hanoi
Hanoi University of Civil Engineering
nguyenworkaholic@gmail.com
Abstract
The rapid uptake of generative AI is transforming how engineering students produce multimodal texts slides, diagrams, and AI-generated images for technical presentations yet existing assessment rubrics rarely account for work that is partly machine-created. This raises pressing questions of fairness, validity, and academic integrity: how can teachers assess multimodal communicative competence when authorship is shared between student and AI? This study designs and pilot-tests a multimodal assessment rubric for AI-assisted technical presentations in an English for Specific Purposes (ESP) course for engineering students. Drawing on multiliteracies pedagogy and validity theory, the rubric distinguishes students' communicative and design decisions from AI-generated content. It was developed and trialled with a cohort of undergraduate presentations and refined through teacher–student feedback. The study is expected to yield practical, transparent criteria for evaluating AI-mediated multimodal work while safeguarding academic integrity. For TESOL and ESP practitioners, it offers a classroom-ready instrument and design principles for AI-aware assessment, supporting fairer evaluation and stronger learner agency in technical English contexts.
Keywords: multimodal assessment, rubric design, AI-assisted learning, English for Specific Purposes, academic integrity
1. Introduction
The technical presentation is one of the most consequential genres an engineering graduate must master in English, and the spoken word is only part of the message: the argument lives in the drawing on the screen, in the section enlarged at the right moment, in the sequence in which the slides unfold. Research on technical talk has shown that such visuals carry argumentative work of their own, and that effective speakers orchestrate verbal, written and visual resources together rather than in sequence (Morell, 2015; Rowley-Jolivet, 2002). A presentation assessed only for pronunciation, fluency and grammatical accuracy under-represents the competence that matters.
Generative artificial intelligence (GenAI) has now entered the production pipeline of this genre, and what has changed is the evidentiary status of the artefact the teacher receives: a polished English sentence no longer licenses an inference about the student's writing, and a competent diagram no longer licenses one about their ability to select and adapt a visual. Yet most ESP presentation rubrics still describe a single human author, with criteria such as “slides are clear, well organised and visually appropriate”. Applied honestly, they force the marker either to assume the work is entirely the student's, which rewards undisclosed delegation, or to suspect it is not, which turns assessment into an inquiry the evidence cannot support. Detection does not help, since tools claiming to identify machine-generated text are neither accurate nor robust (Weber-Wulff et al., 2023), and prohibition pushes use underground precisely the condition under which nothing can be assessed. International guidance has moved instead towards specifying permitted use, requiring disclosure, and redesigning tasks so that learning remains visible (Miao & Holmes, 2023).
The teacher's problem is therefore not policing but judgement: when I put a mark next to this presentation, what exactly am I giving credit for? This paper argues that credit belongs to the student's communicative and design decisions what to take from a machine, what to reject, what to verify and what to change rather than to the surface properties of the artefact. Three questions guided the work:
• RQ1. What should be assessed in an AI-assisted multimodal technical presentation in an ESP engineering course?
• RQ2. How can a rubric make the boundary between a student's design decisions and AI-generated content visible and assessable?
• RQ3. What transferable design principles follow for AI-aware assessment in other ESP and TESOL settings?
This is a design study: it reports an instrument, the reasoning that produced it and the revisions classroom use forced upon it, but no score data or measured gains.
2. Literature review
2.1 Multimodal competence and its assessment
Social semiotic accounts treat meaning as distributed across modes, each with its own affordances (Kress, 2010; Kress & van Leeuwen, 2021), so competence in a multimodal genre is the capacity to combine channels such that each does the work it is best suited to. Rowley-Jolivet (2002) showed that visuals in scientific conference papers form a genre-specific repertoire performing recurrent rhetorical functions, and Morell (2015) found that effective presenters deploy verbal, written and non-verbal resources in overlapping combinations, and that strong use of the visual modes can compensate for limitations in the verbal mode a finding with a double edge, since compensation shades into concealment once the visual material is generated on request. For civil engineering students the visual channel is the substance of the discourse: construction drawings are the field's primary technical texts, and the talk points, enlarges and narrates a path through them.
The pedagogy of multiliteracies makes such composing assessable. The New London Group (1996) framed meaning-making as Design, distinguishing Available Designs from Designing the act of selecting and recombining them and from the Redesigned that results (see also Cope & Kalantzis, 2009). The distinction tells an assessor where to look: the learning is in Designing, and the artefact is evidence of it rather than the object of interest. Hafner and Ho (2020) pursued this into practice with a process-based model spanning pre-design, design, sharing and reflecting; their teacher interviews report difficulty in attending simultaneously to sound, image, fluency and accuracy while a performance unfolds, and rubrics lacking exemplars for distinguishing strong from merely adequate work (see also Belcher, 2017). The rubric literature adds that consistency improves principally when criteria are explicit and raters are supported by exemplars and training (Jonsson & Svingby, 2007; Panadero & Jonsson, 2013), that any rubric should be locatable on Dawson's (2017) fourteen design elements, and that a rubric students see and apply builds evaluative judgement rather than merely recording its absence (Sadler, 1989; Tai et al., 2018).
2.2 Generative AI, authorship and validity
The first wave of responses to GenAI was dominated by misconduct framing (Cotton et al., 2024; Perkins, 2023), but detectors have proved unreliable and easily defeated by paraphrase or machine translation (Weber-Wulff et al., 2023), so attention has shifted from detection to design. The AI Assessment Scale (AIAS) asks the teacher to specify, for each assessment, the level of AI involvement permitted (Furze et al., 2024; Perkins, Furze, et al., 2024; Perkins, Roe, & Furze, 2025). It is, however, a permission instrument rather than a credit instrument: it tells a student what is allowed, not how much a competently selected and corrected AI image is worth beside a student's own photograph of a site defect. UNESCO's guidance points the same way while leaving the criterion-level question to practitioners (Miao & Holmes, 2023).
Validity theory names the stakes. Messick (1989) identified construct-irrelevant variance and construct under-representation as the principal threats; Kane (2013) laid the claims implicit in a score interpretation out as a chain of inferences scoring, generalisation, extrapolation, implications; the Standards (AERA et al., 2014) treat fairness as a validity concern rather than a separate virtue. Undisclosed AI assistance is a textbook source of construct-irrelevant variance, but the converse error is as damaging: a rubric refusing to credit anything a machine touched under-represents the construct, because selecting, evaluating and correcting machine output is now part of competent technical communication (Bearman et al., 2024; Dawson, 2021). What is missing is an instrument at criterion level that is simultaneously multimodal, AI-aware and defensible as a basis for judgement.
3. Conceptual framework
The design rests on two pillars and three commitments, shown in Figure 1.

Figure 1. Conceptual framework of the study.
From multiliteracies pedagogy the design takes its object: the assessable thing is Designing. Under AI conditions a generative model is best understood not as a co-author but as an unusually productive supplier of Available Designs a generated image or a drafted sentence is a resource offered to the designer, not a design so the marker's question is unchanged in form: what did this student do with what was available? From validity theory the design takes its discipline: a criterion is justified not because it names something desirable but because a judgement made against it supports a claim someone will act on (Kane, 2013; Messick, 1989). Three commitments follow:
• Transparency of provenance. The origin of each element is declared before it is judged a condition of entry, not something to be inferred or scored.
• Credit for decisions, not artefacts. Descriptors interrogate what the student chose, checked, rejected and changed; surface quality alone cannot lift a band.
• Evidence across the process. Judgement draws on traces from several points in the cycle (Hafner & Ho, 2020).
The construct they serve is the capacity to construct, deliver and defend a technical message for a specified professional audience by orchestrating spoken language, written text and graphic-symbolic resources, while selecting, evaluating, correcting and honestly accounting for any machine-generated material used. Five domains operationalise it: technical content (D1), spoken technical English (D2), multimodal orchestration (D3), critical use of AI (D4), integrity and attribution (D5).
3.1 What the five domains do not cover
The correspondence between construct and domains is deliberate but not complete. Collaborative design is only partly visible: the task is done in groups, so the rubric records what a group arrived at, not how the negotiation ran, and individual contribution is sampled only in the question stage. Listening is not assessed at all, nor is the written documentation that accompanies a technical talk. This is construct under-representation in Messick's (1989) sense: a high mark supports a claim about this student's capacity to design and defend a technical message in English with preparation time, not about their communicative competence in an engineering workplace as a whole.
4. Design Approach
4.1 An educational design study
The work follows the logic of educational design research: a practical problem is analysed, a prototype is built from theory, it is put into use, and the difficulties met in use are treated as data about the design rather than as noise. The output is an artefact, the reasoning that justifies it, and design principles that may transfer (Figure 2).

Figure 2. The four-phase design cycle followed in developing the rubric.
4.2 Setting, task and scale
The setting is an ESP course at a university of civil engineering in Hanoi, where students meet English primarily through drawings, standards and manufacturers' documentation. The assessed task is a short technical presentation prepared and delivered in groups: each group chooses a topic within a specified area of construction practice, prepares slides including at least one construction drawing, delivers the talk with every member speaking, then answers unscripted questions addressed to named individuals. The task predates this study; what changed was the framing an AI-use plan agreed in advance, the retention of process artefacts, and the rubric.
The trial drew on a single intact class group of 25 second-year students, all of whom had completed the programme's general English requirement. They worked in five groups, producing five presentations marked with the draft instrument; no selection was applied. Five presentations cannot show how the descriptors behave across a range of performances, but they can show whether the wording decided cases, and where it did not. The class size also bears on Section 7.4: the full cycle was designed for and run in a class of this size, so the reductions proposed there for larger cohorts are reasoned extrapolation rather than tested practice.
4.3 Design evidence and its analysis
Four sources informed the successive versions: an audit of the course's previous rubric and of published frameworks; consultation with colleagues about the criteria they found hardest to apply; the trial itself; and the marker's record of the points at which judgement stalled, which proved most productive. Marking was carried out with the draft rubric open in a document carrying, beside each domain, an empty field headed “where the descriptor did not decide the case”; whenever the wording admitted two defensible placements, the marker stopped, named the difficulty and the presentation it arose from, placed the work provisionally and continued, producing a running log across the whole set.
The entries were then sorted by the kind of difficulty they named rather than by the domain in which they surfaced. Categories were derived from the entries rather than fixed in advance and revised twice as sorting proceeded; six accounted for the log and correspond to the subsections of Section 6. A category called for revision when more than one entry fell into it and when the difficulty was attributable to the wording of the instrument rather than to the quality of the work. Feedback was gathered less formally: a 30-minute debriefing with the class after marks were returned, and written comments from two colleagues who read the draft rubric alongside the log.
4.4 Ethics and reporting stance
Disclosure was made non-punitive: declaring the use of a tool could not by itself reduce a mark, and students were told so in writing beforehand. The plan was agreed in advance so that the declaration was planning rather than confession, and a fully viable no-AI route was preserved, with D4 then judged on the reasoning behind the decision not to use a tool which also addresses equity. The paper reports no score data: after one cycle the validity argument concerns the plausibility of its assumptions, and agreement indices drawn from five presentations marked by the instrument's designer would convey a precision the design does not possess (Kane, 2013).

  1. The Instrument
    5.1 Architecture
    The rubric has three layers (Figure 3). Layer A is a disclosure gateway through which the work passes before it is judged; Layer B contains the five assessed domains, each with four quality levels; Layer C is the base of process evidence on which Layer B judgements draw. The layering is itself a design claim: provenance is a precondition of interpretation, not a criterion competing for weight.

Figure 3. The three layers of the rubric and their relationship.
5.2 Layer A: the disclosure gateway
Before marking, students submit the AI-use plan agreed at the start, a declaration of the tools used and the material they produced, and a statement of what was changed and why. The third does most of the work: a declaration that a model generated an image tells the marker little, whereas a statement that the image showed a reinforcement arrangement not matching the detail under discussion, and was therefore replaced with a redrawn figure, tells the marker a great deal and is itself evidence for D4. Nothing in Layer A is scored, which is the design's most counter-intuitive feature and its most important one: if disclosure carries marks students optimise it, if it is penalised they suppress it, and either way the marker loses the information the judgement depends on.
Table 2. Provenance categories used in the disclosure statement
Category Description What the student must show
Student-authored Produced without generative assistance Nothing beyond the artefact itself
Machine-generated, used as produced From a generative tool, included essentially unchanged The tool, the request made of it, and why the output was judged fit for purpose
Machine-generated, substantially reworked From a tool but altered in substance by the student The original output beside the final version, and an account of the change
Machine-assisted student work Student content on which a tool gave feedback, correction or translation The nature of the support and which suggestions were accepted or declined
Third-party technical source Standards, manufacturers' drawings, published details, others' photographs A conventional citation, and permission status where relevant
Note. Categories are applied element by element, not to the presentation as a whole.
5.3 Layer B: the five assessed domains
The domains are set out in Table 3. Three points mark where the instrument departs from a conventional rubric. D3 is written around fit rather than quality: a generated image that is beautiful and irrelevant scores lower than a rough hand-redrawn section that carries the argument. D4 credits scepticism the highest band is reached not by producing good output but by testing it against technical sources and acting on what is found, which matters because generative models fail plausibly here, producing details that are geometrically coherent and constructionally impossible. D5 judges honesty and specificity, not abstinence.
5.4 Layer C: evidence across the task cycle
A single short performance is a thin evidentiary base, and under AI conditions its most polished parts are the least diagnostic, so judgement is distributed across the cycle (Figure 4). The unscripted question stage carries a disproportionate share of the load: asking a student to explain, without slides, why a section was chosen for enlargement, or to gloss a term, produces evidence no artefact can supply and no tool can pre-empt. It is where the extrapolation inference is most directly supported, and it is cheap.

Figure 4. Assessment points distributed across the task cycle.
5.5 Permission, credit and design parameters
The rubric presupposes that the task's permitted level of AI use has been settled, and the AIAS is a convenient way to settle it (Perkins, Furze, et al., 2024). The task sits at a collaborative level: tools may be used throughout preparation provided use is declared and can be accounted for. The AIAS answers what is permitted; the rubric answers what is credited and on what evidence. In Dawson's (2017) terms the scoring strategy is analytic with no compensatory total, there is no secrecy the rubric is released with the brief and applied in class to two contrasting samples and judgement complexity is deliberately high in D4 and D5.
6. What the trial showed
What follows is an account of where judgement proved difficult and what was changed in response. It is design evidence, not findings about students.
6.1 Polish without fit, and terminology that could not be glossed
The largest category in the log was the presentation whose slides were visibly beyond the standard the course had previously seen, accompanied by a spoken account that did not correspond to them: the speaker described one thing while the screen showed another, or referred to a figure without indicating what in it mattered. Under the previous rubric such work scored well on slide quality and poorly on delivery, producing a middling mark that described nothing real; D3 was rewritten so that its object is the relation between showing and saying. A second category concerned technical vocabulary that students could not explain when asked. Borrowing is not new in ESP, but generated text raises the sophistication of the borrowed material and lowers the effort of borrowing it. The question stage identified these cases reliably and without accusatory framing, and D1 and D2 were amended to refer explicitly to unscripted explanation.
6.2 Disclosure, translation and the no-AI route
Students' first reaction to the declaration was anxiety, despite the written assurance that it could not reduce a mark. The remedy was not a better form but a change of sequence moving the declaration forward into a plan agreed before preparation began and it came from the debriefing rather than from the marking log. Machine translation from Vietnamese formed a category of its own: for some a scaffold, the thinking done in Vietnamese and then worked over with terms checked against field usage; for others a substitute, the English arriving fully formed and read aloud. The artefacts alone do not distinguish these; the process evidence and the question stage do, and Table 2 was extended to make language support a separate declarable category. Some students also chose not to use generative tools at all, and an early draft of D4 could only record that choice as a gap; D4 was rewritten so that its highest band is reachable either by well-evidenced critical use or by a reasoned, demonstrated decision not to use.
6.3 Marker workload
Reading artefacts against declarations takes longer than ticking observable features, and the extra time was concentrated in D4 and D5. Two mitigations are now in the guidance notes: capping the declaration at one page, which forces students to be specific, and marking D1 to D3 live during delivery while reserving D4 and D5 for a short reading afterwards. Whether the remaining overhead is acceptable is a judgement each teacher must make.
7. Discussion
7.1 The validity argument
Figure 5 lays out the four inferences a mark from this rubric is asked to support, the threat each faces when authorship is shared, and the design feature that responds to it. It is an argument, not a result: each response is a reasoned attempt to make an assumption plausible, and none has been tested. The responses also interact the disclosure gateway is what makes the decision-focused descriptors usable, since a descriptor about what the student changed is unanswerable without a record of what was there before.

Figure 5. Inferences, threats and the design responses built into the rubric.
7.2 Fairness
Fairness runs in more than one direction, which is why the Standards treat it as an aspect of validity rather than a separate box to tick (AERA et al., 2014). Undisclosed use advantages some students over others, but a disclosure regime creates its own risks: students differ in access to capable tools, which is why the no-AI route must be genuinely equivalent, and disclosure asks students to reveal their working in a way that could feel intrusive, which is why nothing in Layer A is scored.
A second source of construct-irrelevant variance deserves separate treatment, because the design creates it. The declaration is written, and Appendix A asks for most of it in English, so a fluent writer can present a modest amount of critical work convincingly while a weaker one may have done more and shown less. Since D4 and D5 are judged by reading artefacts against the declaration, the quality of its prose can contaminate two domains that are not measuring writing at all. Three partial responses are built in: the plan may be written in Vietnamese; the declaration is read for content, with language accuracy excluded from the mark; and the question stage offers a spoken route to the same evidence. Whether these suffice this trial cannot answer, and a marker should watch for the tell-tale pattern the strongest writers scoring highest on D4.
7.3 Agency and integrity
If evaluative judgement is a graduate capability (Tai et al., 2018) that generative tools make more rather than less central (Bearman et al., 2024), then a rubric requiring students to say what they took from a machine, what they rejected and on what grounds teaches the capability while assessing it. Correspondingly, the instrument treats integrity as a positive competence to be demonstrated rather than an offence to be detected (Dawson, 2021; Miao & Holmes, 2023). It does not make dishonesty impossible, but it changes what dishonesty requires: a false declaration must be sustained against the artefacts, the drafts and the questions, and must be made explicitly rather than by silence.
7.4 Feasibility in constrained settings
Adopting the rubric obliges the teacher to collect process artefacts, run an unscripted question stage, and release and teach the criteria a redesign of the task cycle rather than a new marking sheet. Under conditions common in the region a large class, one teacher without an assistant, a marking window measured in days that cycle is not realistic. Three reductions preserve most of the logic: shortening the question stage to one question per speaker drawn from a standing list, which retains the extrapolation evidence while bounding the time; reducing the process-evidence base to the AI-use plan and one intermediate draft; and letting peer assessment at rehearsal carry part of the formative load (Tai et al., 2018). What cannot be reduced is the disclosure gateway, which costs almost nothing to collect and is what makes every decision-focused descriptor interpretable.
7.5 What the instrument asks of the teacher
The rubric assesses students' critical use of generative tools and so assumes a corresponding capability in the marker. To judge D4 the marker must recognise the errors these tools characteristically make here: a reinforcement arrangement drawn coherently that cannot be built, a standard cited with a plausible but non-existent clause number, a confident gloss of a term the local code uses differently. A marker who cannot see such errors will credit a student who did not catch them, inverting the purpose of the domain. A department adopting the instrument should spend part of a staff meeting generating output in the relevant topic areas and examining where it fails, keeping the examples as marker preparation and classroom exemplars.
8. Design Principles
Seven principles generalise beyond the particular instrument. They are offered as hypotheses for other designers to test rather than as established rules.
Table 5. Design principles for AI-aware multimodal assessment
Principle What it means in practice
1. Make provenance a condition, not a criterion Require declaration before marking and give it no weight in the mark: scoring disclosure corrupts it, penalising it suppresses it.
2. Assess decisions, not artefacts Write every descriptor so that it asks what the student selected, checked, rejected or changed.
3. Distribute the evidence Gather traces from planning, drafting, rehearsal, delivery and reflection; a single performance is the least diagnostic evidence available.
4. Keep an unscripted moment Reserve time for live questioning no artefact can anticipate — the strongest support for extrapolation.
5. Keep the no-AI route fully viable Every band must be reachable without generative tools, the decision not to use them creditable in its own right.
6. Publish the rubric and teach it Release criteria with the brief and apply them in class to samples, so the instrument builds evaluative judgement.
7. Prepare the marker before adopting the instrument A marker who cannot recognise how generative tools fail in the subject matter cannot judge critical use; treat marker preparation as part of adoption.
Outside engineering the principles transfer more readily than the domains: a business English course assessing client pitches or a nursing programme assessing handover simulations would need different descriptors in D1 and different visual conventions in D3, but the three-layer architecture, the provenance categories and the principles apply unchanged.
9. Limitations and Directions for Further Work
The instrument has been developed and used in one course at one institution, with a single teacher who is also its designer; the risk that its descriptors work because their author knows what she meant by them is real and unaddressed. Five presentations cannot show how the wording behaves across a range of performances, no reliability evidence exists, and the domains most likely to strain agreement D4 and D5 are exactly the novel ones. One cheap step was available and not taken: a colleague applying the rubric independently to a few recorded presentations, keeping the same log, would show whether the descriptors work for someone who did not write them. Beyond that, a multi-rater trial would establish whether the domains can be applied consistently; a think-aloud study would show what raters attend to when reading a declaration against an artefact; a study of student perceptions conducted by someone other than the teacher would test whether the declaration is experienced as planning or as surveillance; and a transfer study would separate what is general in the architecture from what is particular to construction drawings.
10. Conclusion
Generative AI has not created a new assessment problem so much as exposed an old evasion: rubrics for multimodal work have often credited the artefact and hoped it stood for the student. That hope is no longer tenable, and its collapse is an opportunity to say more precisely what an ESP course develops not the capacity to produce a clean slide, but the capacity to decide what a technical audience needs to see, to find or make it, to judge whether it is right, and to stand behind it under questioning. The rubric presented here makes provenance a condition of judgement rather than an object of suspicion, credits decisions rather than surfaces, and spreads its evidence across a task cycle. It has been used, it failed in instructive ways, and it was changed in response; what it has not been is tested by anyone other than its author.

References
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Bearman, M., Tai, J., Dawson, P., Boud, D., & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. Assessment & Evaluation in Higher Education, 49(6), 893–905. https://doi.org/10.1080/02602938.2024.2335321
Belcher, D. D. (2017). On becoming facilitators of multimodal composing and digital design. Journal of Second Language Writing, 38, 80–85. https://doi.org/10.1016/j.jslw.2017.10.004
Cope, B., & Kalantzis, M. (2009). “Multiliteracies”: New literacies, new learning. Pedagogies: An International Journal, 4(3), 164–195. https://doi.org/10.1080/15544800903076044
Cotton, D. R. E., Cotton, P. A., & Shipway, J. R. (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 61(2), 228–239. https://doi.org/10.1080/14703297.2023.2190148
Dawson, P. (2017). Assessment rubrics: Towards clearer and more replicable design, research and practice. Assessment & Evaluation in Higher Education, 42(3), 347–360. https://doi.org/10.1080/02602938.2015.1111294
Dawson, P. (2021). Defending assessment security in a digital world: Preventing e-cheating and supporting academic integrity in higher education. Routledge.
Furze, L., Perkins, M., Roe, J., & MacVaugh, J. (2024). The AI Assessment Scale (AIAS) in action: A pilot implementation of GenAI-supported assessment. Australasian Journal of Educational Technology, 40(4).
Hafner, C. A., & Ho, W. Y. J. (2020). Assessing digital multimodal composing in second language writing: Towards a process-based model. Journal of Second Language Writing, 47, Article 100710. https://doi.org/10.1016/j.jslw.2020.100710
Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational Research Review, 2(2), 130–144. https://doi.org/10.1016/j.edurev.2007.05.002
Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
Kress, G. (2010). Multimodality: A social semiotic approach to contemporary communication. Routledge.
Kress, G., & van Leeuwen, T. (2021). Reading images: The grammar of visual design (3rd ed.). Routledge.
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). Macmillan.
Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
Morell, T. (2015). International conference paper presentations: A multimodal analysis to determine effectiveness. English for Specific Purposes, 37, 137–150. https://doi.org/10.1016/j.esp.2014.10.002
New London Group. (1996). A pedagogy of multiliteracies: Designing social futures. Harvard Educational Review, 66(1), 60–92.
Panadero, E., & Jonsson, A. (2013). The use of scoring rubrics for formative assessment purposes revisited: A review. Educational Research Review, 9, 129–144. https://doi.org/10.1016/j.edurev.2013.01.002
Perkins, M. (2023). Academic integrity considerations of AI large language models in the post-pandemic era: ChatGPT and beyond. Journal of University Teaching and Learning Practice, 20(2). https://doi.org/10.53761/1.20.02.07
Perkins, M., Furze, L., Roe, J., & MacVaugh, J. (2024). The Artificial Intelligence Assessment Scale (AIAS): A framework for ethical integration of generative AI in educational assessment. Journal of University Teaching and Learning Practice, 21(6), 49–66. https://doi.org/10.53761/q3azde36
Perkins, M., Roe, J., & Furze, L. (2025). Reimagining the Artificial Intelligence Assessment Scale (AIAS): A refined framework for educational assessment. Journal of University Teaching and Learning Practice, 22(7). https://doi.org/10.53761/rrm4y757
Rowley-Jolivet, E. (2002). Visual discourse in scientific conference papers: A genre-based study. English for Specific Purposes, 21(1), 19–40. https://doi.org/10.1016/S0889-4906(00)00024-7
Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18(2), 119–144. https://doi.org/10.1007/BF00117714
Tai, J., Ajjawi, R., Boud, D., Dawson, P., & Panadero, E. (2018). Developing evaluative judgement: Enabling students to make decisions about the quality of work. Higher Education, 76(3), 467–481. https://doi.org/10.1007/s10734-017-0220-3
Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, Article 26. https://doi.org/10.1007/s40979-023-00146-z
Appendix A: Student-Facing AI-Use Plan and Disclosure Statement
Nothing on this page is scored; a complete page is a condition of marking. Parts 1 and 2 are completed by the group on a single page; Part 3 is written individually. Part 1 may be written in Vietnamese; Parts 2 and 3 are read for content, not for language accuracy.
Part When What to write Guidance
1. Plan Agreed before preparation begins Which tools you intend to use, and for what. What you will do yourselves. If none, why. Planning, not confession. Nothing here can reduce your mark.
2. Declaration Submitted before you present For each element slide, image, drawing, section of script its provenance category from Table 2. Work element by element. A list is enough; sentences are not required.
3. Changes Submitted before you present, individually Two things a tool produced that you changed or rejected: what was wrong, how you checked, what you did instead. This part earns credit in D4. Name the source you checked against; keep the original output for comparison.
Note. Bring the original tool outputs referred to in Part 3 to the presentation.

Table 3. The AI-aware multimodal assessment rubric for ESP technical presentations
Domain Accomplished Competent Developing Emerging
D1 Technical content and accuracy Accurate, well scoped, developed in depth. Claims grounded in named sources. Unscripted questions answered with correct detail. Accurate and appropriately scoped. Sources mentioned. Questions answered correctly, sometimes only at slide level. Broadly correct but thin. Sources vague. Some questions unanswerable beyond what was on screen. Technical errors, or content copied at a level the speaker cannot explain.
D2 Spoken technical English Intelligible and well paced. Terminology accurate and reworded on request. Signposts, responds and repairs comfortably. Generally intelligible. Terminology mostly accurate and usually glossable. Interaction managed with some hesitation. Intelligibility varies and depends on the slides. Terminology not always understood when questioned; largely read. Read or recited, with little spontaneous speech. Terms cannot be explained in other words.
D3 Multimodal design and orchestration Showing and saying do different, complementary work. Drawings carry the argument; attention is directed within them. Channels aligned and generally supportive. Key figures pointed to and explained; some slides carry more decoration than argument. Alignment inconsistent: figures displayed without being used, or the talk restates what is written. Channels diverge. Figures unexplained or merely decorative, however polished; the talk could be given without them.
D4 Critical and purposeful use of AI Tools chosen for stated purposes; output checked against technical sources and errors corrected or rejected with reasons. Or: a reasoned, evidenced decision not to use tools. Tools used purposefully with some checking evident; at least one substantive correction or rejection described, grounds stated briefly. Tools used without clear purpose, or output accepted with only cosmetic change; checking asserted rather than shown. Generated content used as produced; technical errors introduced by a tool remain in the work.
D5 Integrity and attribution Every element assigned a provenance category. Account of changes specific and confirmed by drafts and artefacts. Sources cited conventionally. All substantial elements assigned a category. Account of changes accurate but general. Declaration partial or imprecise, leaving the origin of some elements unclear. Declaration absent, or contradicted by the artefacts or by the speaker under questioning.
Note. Domains are reported separately; no compensatory total is calculated. The rubric is released with the task brief and applied by students to sample presentations before their own preparation begins.

Author

Presentation materials

There are no materials yet.