SRHE Blog

The Society for Research into Higher Education


Leave a comment

Vibe-coding and co-creation: human-AI collaboration in higher education

by Andrew Williams

The debate around artificial intelligence (AI) in higher education has largely focused on students. Are they using ChatGPT in assessments? How do we maintain academic integrity? How should teachers and curricula adapt?

Much less attention has been given to academics working with generative AI (GenAI) as a creative collaborator or co-pilot, rather than merely attempting to mitigate its impact. The potential for productive and effective AI collaboration deserves greater consideration because it challenges longstanding assumptions about expertise, creativity and educational design.

Ever since the release of ChatGPT in November 2022, I have been fascinated by how intuitive the interaction with GenAI tools can be and the ease with which ideas can be explored through ordinary conversation. GenAI is a sophisticated technology that allows users to test possibilities, refine outputs, and rapidly turn preliminary concepts into usable material.

My background is in biological sciences, with no grounding or expertise in computer science, software engineering or coding. For me, this technology has helped me overcome technical barriers and allowed me to explore new ways to teach and to develop bespoke educational resources. I envisioned a learning experience for my undergraduate medical science students (studying human genetics and evolution) that moved beyond static textbook diagrams and dry explanations to a more interactive and visually stimulating format. Realising that vision required a simulation that let students tweak parameters, observe real‑time population dynamics, and grasp complex evolutionary ideas.

Traditionally, such a project would rely on grants, external software developers, or months of self‑taught coding. For most of us, the gap between an innovative pedagogical idea and a digital reality was confined by technical expertise. Today, through iterative collaboration with GenAI, natural language can function as a design interface, allowing educators to transform ideas into working educational tools.

I was also intrigued by the ever improving, and often publicised, capabilities of GenAI tools as coding agents. With no coding experience myself, I decided to try my hand at ‘vibe-coding’.

Natural language as a design interface

‘Vibe-coding’ is an emerging term within technology communities that describes a process where functional intent is communicated via natural language prompts, allowing AI to generate and refine code. This allows the user to focus on what the tool should do rather than how it must be coded. Essentially, ‘vibe-coding’ involves directing the AI to write and co-develop code aligned with your vision and pedagogical objectives. In other words – no code, no problem!

My recent research describes an iterative prompting strategy and the outcomes of a human-AI collaboration to create an interactive simulation that models certain aspects of evolution (Williams, 2026a). Instead of learning JavaScript or working with software developers, I co-created an application with ChatGPT using only natural language prompting. The result was a browser-based simulation capable of supporting inquiry-based learning and formative assessment in undergraduate education.

Rather than blindly generating content based on loosely defined instructions, the human-AI collaboration involved a structured iterative prompting strategy, careful testing of AI-generated output and re-prompting until the desired aims of the project were met. This enabled the creation of an interactive simulation in which students could investigate biological systems, manipulate variables, export datasets and construct evidence-based explanations (figure 1). This allowed students to move beyond consuming information, to generating and interpreting data themselves in real-time.

Figure 1. Snapshot of ‘Survival Island’ – an interactive simulation.

The Illusion of automated expertise

There is a prevailing narrative that AI democratises expertise, in that it allows anyone to do anything more quickly and with less effort. My experience of ‘vibe-coding’ suggested something different. The  credibility of the final simulation depended on actively embedding disciplinary expertise directly into the co-design process.

While the AI had removed the technical barrier (the coding), it had increased the importance of my academic expertise and oversight. The AI could generate a thousand lines of code in seconds, but it couldn’t tell me if the code represented a plausible scientific simulation. My epistemic judgement was required to verify the AI-generated output. The final student-ready simulation required multiple rounds of iteration, testing and re-prompting, until the collaboration achieved a credible and biologically grounded application suitable for teaching.

Rather than replacing my expertise, ChatGPT provided capabilities I do not possess: writing HTML; JavaScript; debugging code. I contributed disciplinary knowledge, pedagogical intent and continuous evaluation. Every feature reflected educational decisions rather than technical ones. The simulation succeeded because the educational need was clearly defined before AI generated the code.

Higher education has often treated GenAI as an automated assistant or as a problem to be mitigated. Increasingly, it may be more productive to think of AI as a collaborator whose contributions depend upon human judgement. Expertise remains indispensable, as humans ultimately remain responsible for judging whether outputs are educationally, scientifically and ethically sound.

Beyond the prompt: the AI literacy gap

Universities have understandably invested considerable effort in developing student AI literacy. This project led me to a broader appreciation of the need to promote teacher AI literacy across higher education, a critical area often overlooked.

AI literacy for teachers is not just about knowing how to write a good prompt or choosing the best online GenAI tool. Rather, it is a multi-dimensional competency that considers foundational knowledge (prompting, tool use), ethical, emotional and technical proficiencies. For educators to co-create teaching resources in collaboration with AI, for example through vibe-coding, more than technical curiosity is required.

The iterative ‘vibe-coding’ process I followed mirrors the ‘critical evaluation’, ‘epistemic’ and ‘innovation’ dimensions of an integrated AI literacy framework for higher education teachers (Williams, 2026b). This recently published conceptual framework comprises eight core AI literacy dimensions, which you can access in full using the link.

In the current context, the important  AI literacy dimensions include:

  • Critical evaluation: the ability to treat the AI model as a “black box” and rigorously verify outputs against authoritative sources.
  • Epistemic authority: a clear understanding of the probabilistic nature of GenAI and the domain-specific expertise to judge AI output, thereby retaining human agency during interactions with AI.
  • Innovation mindset: the willingness to move from being a consumer of tools to a co-creator of bespoke learning environments, including an understanding of how to collaborate effectively with GenAI.

These are becoming essential academic capabilities rather than optional technical skills. This transition presents considerable challenges, as academic practice has historically been predicated on either individual resource creation or the outsourcing of technical development to third parties. GenAI now enables academics to become directors of increasingly sophisticated creative processes. They define the educational challenge, AI performs the technical implementation, and academics evaluate and verify the final outcome.

As technical production becomes easier with GenAI, educational judgement becomes increasingly important. Our value is no longer found in the technical labour of production, but in our ability to define the educational challenge, critique the iterations, and curate the final outcome. The “vibe” we are coding is our pedagogical intent.

Key takeaways

  • AI is a tool, not a replacement.
  • Human-AI collaboration enables pedagogical creativity.
  • ‘Vibe-coding’ uses AI’s coding capabilities but demands epistemic oversight.
  • Teacher AI literacy is multi‑dimensional.
  • Institutional support is required to build that AI literacy.

GenAI is redefining the landscape of higher education, not through incremental improvements to traditional teaching materials, but by enabling entirely new educational content to be created. However, despite the promises of this new technological revolution, we still need robust evidence about student learning gains, engagement, accessibility and the sustainability of AI-generated educational resources.

It is easy to provide AI tools for HE educators to use. The challenge is to give educators the technical, critical and ethical literacy required to use them effectively.

Andrew Williams is a Professor (Teaching) in the Faculty of Medical Sciences at UCL. He teaches immunology and cell and molecular biology to undergraduate and postgraduate students. His research aims to explore the impact of artificial intelligence (AI) on assessment and learning in higher education and the ways in which we can promote student and staff AI literacy.


1 Comment

A narrator Is not a witness: what AI got right — and wrong — across three Ghanaian MBA classrooms

by Carlene Kyeremeh

It was late, and I was building slides on fiscal policy for a class of twenty, most of whom had never taken economics before. I asked an AI tool for a worked example of expansionary policy — something concrete enough to anchor the mechanism. It answered in seconds: a government cuts taxes and raises infrastructure spending; roads get built; contractors hire; demand rises. Clean. Usable. American. The mechanism travelled, but the roads were interstate highways, and the government spending the money was not one my students would ever petition. The example had arrived so quickly, and so confidently, that I nearly pasted it in. For a moment, convenience almost became curriculum.

Then, in the same session, the same tool did something I could not have done as well alone. I asked it to scaffold the vocabulary — a glossary that assumed nothing and met students who did not yet have the words. It was patient in a way I am not always patient at that hour. It found analogies; it sequenced the terms so that each rested on the one before. That slide was better for the machine’s help.

Within a single sitting, the tool had lowered the barrier for my students and quietly imported the exact default my work exists to help people notice. Across three MBA courses this term, I began to see the same pattern: the tool improved many aspects of how I prepared to teach, and became least dependable where the teaching required local knowledge, current evidence or an institution named correctly. The help landed on form. The failures landed on specificity.

In my international trade finance course, I extend a running case — Adom Naturals Ltd, a Kumasi agribusiness I built so that trade finance stops being abstract and acquires a location, a product and a currency exposure. I asked the tool to situate the firm against current conditions: what the African Continental Free Trade Area (AfCFTA) changes for a business like this, and where COCOBOD now sits in the picture. It answered in assured paragraphs and told the familiar story: Ghana grows some of the world’s finest cocoa, ships much of it out with limited processing, and watches the greater share of value accrue downstream. Fluent, orderly, and a season out of date.

The confident narrator

It missed what I happened to be holding in a government source that week: in February 2026, Cabinet directed that, from the 2026/27 crop season, a minimum of 50 per cent of Ghana’s cocoa beans should be processed locally. For Adom Naturals, that reform changes the opportunity set. Cocoa liquor, butter, cake and other processed products can retain more value within Ghana and may qualify for preferential treatment in African markets where the relevant AfCFTA rules of origin and tariff requirements are satisfied. The tool had narrated the extractive arrangement in the present tense and missed the policy intended to change it.

Nothing marked the claim as stale. The tool warns that it can make mistakes, in a line printed beneath every answer, but a caveat attached equally to everything is not calibration; it is the absence of it, dressed as candour. A colleague says: I am sure of this; check me on that. The tool says it might be wrong about anything, then says everything in the same even voice. The danger was never that it made mistakes. Every source makes mistakes. The danger was that it sounded exactly as certain when it was wrong as when it was right.

The wrong institution

Financial regulation showed me the problem from another angle. Ask a general-purpose tool about capital adequacy, disclosure or market conduct and it is fluent, because the published record is thick with Basel standards, US and UK regimes, and decades of commentary. Ask it to route the same questions through Ghana’s regulatory architecture and the fluency thins.

When I asked which body supervises an insurer in Ghana, and then which oversees a securities offering, it reached both times, confidently, for the Bank of Ghana. The answer was plausible because the central bank is prominent in Ghana’s financial system. It was nevertheless wrong. Insurance supervision belongs to the National Insurance Commission under the Insurance Act, 2021; securities-market regulation belongs to the Securities and Exchange Commission under the Securities Industry Act, 2016, as amended. When I named the specific commissions, the tool corrected itself at once. The information was retrievable; it was not the default.

Three defaults, one voice

Set the three moments side by side and a more complicated pattern emerges. The fiscal-policy example exposed a geographical default; the cocoa case, a temporal one; and the regulation case, an institutional one. These were different failures, but they arrived in the same confident voice.

I cannot inspect the tool’s training archive, so I cannot attribute every error to missing data alone. A stale policy claim may reflect a knowledge cut-off or the absence of live search. A regulatory error may reflect weak retrieval, poor weighting or the greater prominence of a general institution over a specialised one. What I can observe is an asymmetry of retrieval: general and North Atlantic formulations arrived unprompted, while Ghanaian specificity had to be named, sourced and verified into view.

That asymmetry belongs in the larger conversation about AI and epistemic justice — about whose knowledge is dense enough, accessible enough and prominent enough to be retrieved fluently, and whose is thin enough to be flattened, displaced or missed. The tool did not invent the hierarchy of whose knowledge counts. It inherited a record shaped by that hierarchy and can reproduce it at scale, in fluent prose. Better models may reduce some errors, but model improvement alone cannot repair knowledge that remains absent, inaccessible or systematically underrepresented.

I develop that argument more formally elsewhere, in work currently under review. Here, I want only to report what it looks like from inside three classrooms, at the point where defaults become examples and examples become curriculum.

Verification is the work

I use these tools daily and they earn their place, so let me be honest about the difficulty. The answer is not refusal; refusing the help is not a decolonial act, only less help. The answer is the discipline I have argued for all along, now turned on the machine: no sentence enters the curriculum until it points to a source I can hold. Verification is not the friction that slows the tool down. With a tool like this, verification is the work.

I have also begun turning that work into a learning activity. I place selected AI outputs beside the relevant primary or institutional sources and ask students to identify what the model has generalised, dated or assigned to the wrong body. Verification becomes not only my quality-control procedure but part of the curriculum itself.

Perhaps that is the graduate skill this moment now asks for. Not simply how to find information, the tool is generous with information, but how to test information whose presentation gives no sign whether it has earned our trust. The scarce skill is no longer retrieval. It is discernment. Teaching has always required two kinds of expertise: explaining ideas well, and knowing where they belong. The tool is becoming remarkably good at the first; the second is still ours. It narrates beautifully — but a narrator is not a witness, and decolonising the curriculum now includes learning to interrogate the archive that speaks back.

If your tool has ever been confidently wrong about your own institution, your own regulator or your own country’s data, I would like to know what it got wrong — and whether a student would have caught it.

Dr Carlene Kyeremeh is an Associate Professor and Vice President, University Advancement, Recruitment & Research, at All Nations University, Ghana, where she teaches managerial economics and international trade and finance on the MBA programme. Her research examines decolonial curriculum reform, gender equity and academic mobility in African higher education, with the African Continental Free Trade Area as a recurring empirical anchor. She is currently researching the reintegration of diaspora-return faculty in Ghanaian universities. She writes The Decolonized Curriculum, a newsletter on curriculum decolonisation in African higher education. 

LinkedIn [https://www.linkedin.com/build-relation/newsletter-follow?entityUrn=7412983304175534080]

Author’s note: This article is adapted and substantially expanded from Issue 13 of The Decolonized Curriculum.


2 Comments

Overcoming Built-In Prejudices in Proofreading Apps

by Ann Gillian Chu

As I am typing away in Microsoft Word, the glaring, red squiggly underline inevitably pops up, bringing up all the insecurities I have with academic English writing, as an ethnically Chinese, bilingual Chinese-English speaker. So what if I speak English with a North American accent? So what if English has been my medium of instruction for my entire life? So what if I graduated with a Master of Arts with honours in English Language from the University of Edinburgh? My fluency in Chinese somehow discredits my English fluency, as if I cannot be equally competent in both. Because I am not white, my English will always somehow be inadequate.

The way others, and I, perceive my English ability reflects how ‘standard English’ as an idea is toxic to the identity-building of those who are not middle-class, cisgender, heterosexual, white men from the Anglophone world. April Baker-Bell talks about the concept of linguistic justice, arguing that promoting a type of ‘correct English’ has inherent white linguistic supremacy. Traditional approaches to language education do not account for the emotional harm, internalised linguistic racism, or the consequences these approaches have on the sense of self and identity of non-white students. Extending Baker-Bell’s theory, how would this apply to the use of proofreading apps?

Apps are created by people who have their own underlying assumptions and worldviews, even if these assumptions are not explicitly written in any of the apps’ documentation. When using these tools, users need to have a sense of the kind of assumptions these apps carry into their corrections. More importantly, as programmes are written by people and applied in a formulaic way, they should not have the power to define their users’ sense of identity, or even their ability to communicate in English. Algorithm-based tools will always fall short in understanding the nuances and eccentricities that make human writing exciting and intriguing. The app should not have authority over its users, and its feedback should never be taken uncritically.

However, proofreading apps could be used as a pedagogical tool when thoughtfully and critically engaged. Evija Trofimova created a resource titled ‘Digital Writing Tools: Spelling, Grammar and Style Checkers,’ which investigates how different proofreading apps can or should be used. Trofimova’s project assesses how each app can be used for best didactic experiences, with exercise suggestions and classroom activities available for users to begin to see these proofreading apps as a possible pedagogical tool, rather than law enforcement of sorts. Users of proofreading apps should always treat each ‘error’ as a learning opportunity, investigate the rule behind the correction, and actively consider whether it is indeed a correction they want to take up in their writing. If a correction is unexpected, users should be encouraged to investigate why the app suggested it and what is the underlying principle. Crucially, app users need to have a sense of where to draw the fine line between what is conventional (rather than right) and what makes writing comprehensible to readers, and what expresses the unique voice and identity of a writer. Writers should explore more ways to communicate within and outside of conventions, in a way that best represents them. This sort of creativity will go beyond simply relying on the algorithm of a proofreading app.

It needs to be said more often that English as an academic lingua franca is no one’s first language (see Marion Heron and Doris Dippold, Meaningful Teaching Interaction at the Internationalised University). Just because someone is a monolingual English speaker does not necessarily mean that they are good at academic English writing. Just because someone writes in an unexpected or unconventional way does not necessarily mean that they are wrong. An essential purpose of writing is to communicate. As academics, we need to ask ourselves, how much can someone deviate from a standard and still be comprehensible? How much room can we leave to allow students to be themselves and express themselves fully in their writing? How much are proofreading apps stifling their ability to flourish as writers? We are not teaching students to become machines. There is no point in having different students write the same essay in the exact same way. Rather, it is precisely their unique and different voices that should be celebrated.

In ‘The Danger of a Single Story,’ Chimamanda Ngozi Adichie talks about her childhood in Nigeria, reading books about white, blue-eyed characters who played in the snow, ate apples, and talked a lot about the weather, which does not reflect her experience of the world at all. Growing up, Adichie struggled with characters in novels being made up of white foreigners from the West alone, as if the Western world is a cultural ground zero. She and other Nigerians are not represented in the literary works she read. Non-white proofreading app users may easily fall into the same impressionable and vulnerable position as Adichie did. These prescriptive ‘corrections’ made by proofreading apps, just like the children stories that Adichie read, implies an ideological position that a specific language standard, such as standard British or American English, is somehow superior and ‘correct.’ However, the West is not a cultural ground zero, nor is English a neutral medium of communication. Other varieties of English used in non-Western worlds, far from being inappropriate or incorrect, should be celebrated for their ability to reflect the culture and experiences of non-Western writers. The attempt to make a piece of written work meet certain linguistic standards should not be above rhetoric, creativity, and cultural expression.

What makes proofreading apps dangerous is that their underlying assumptions remain invisible to their users. A ‘correct’ grammar may reinforce existing biases in our society, creating linguistic violence, persecution, dehumanisation, and marginalisation, which non-standard English speakers endure when using their own language in schools and everyday life. One thing that stuck with me the most from my undergraduate degree is that proper English changes throughout the ages. What is now deemed suitable was once upon a time a deviant use of the language. And what is deviant now may become mainstream in the future. People have always reacted badly to these linguistic changes, but the changes in usage stick nonetheless. As users of proofreading apps and teachers of students who use these apps, it is important to encourage everyone to think about who the apps were built for and what purposes they were meant to serve. What spelling and grammatical rules do they enforce, and why? In After Whiteness: An Education in Belonging, Willie James Jennings pushes against (theological) education that is ultimately set up for white self-sufficient masculinity, a way of organising life around a persona that distorts authentic identity. This way of being in the world forms cognitive and affective structures that seduce people into its habitation and its meaning-making. When a white Anglophone world is presumed as a norm, and others somehow have to cater to its expectations, it strangles intellectual pursuits from the perspectives of the other. It is the freedom of expression between interlocutors that will create a space for students from all backgrounds to flourish.

Ann Gillian Chu (FHEA) is a PhD (Divinity) candidate at the University of St Andrews. She has taught in higher education contexts in Britain, Canada, and China using a variety of platforms and education tools. As an ethnically Chinese woman who grew up in Hong Kong, Gillian is interested in efforts to decolonise academia, such as exploring ways to make academic conferences more inclusive.