Skip to main content Accessibility Statement
23 June 2026

Beyond the digital divide: how institutional strategy and organisational knowledge shape AI-powered assessment design

Authors

 

Dr Yang Yang

Senior Academic Developer, London School of Economics and Political Science

 

Prof. Isabel Fischer

Professor of Digital Innovation, Warwick Business School (WBS)

Tom Hinks

Senior Learning Designer Digital Education Studio, Faculty of Medicine and Dentistry, Queen Mary University of London


Adrian Chong

Student intern, University of Southampton

 

Prof. Simon Walker

Professor of Practice, Institute of Education, University of Southampton

 

What happens when several universities build the same AI agent, but implement it through different AI platforms, procurement decisions and institutional strategies?


The question emerged from a QAA-supported collaborative project, led by the London School of Economics and Political Science (LSE), in which colleagues from the universities of Warwick, Queen Mary and Southampton developed and evaluated ADA, an AI agent designed to help educators reflect on and co-design assessment. Although ADA’s intended workflow remained broadly consistent, the ways in which it was implemented within each institution differed considerably. Colleagues experimented with Claude, Microsoft Copilot, Gemini and smaller open-source language models, alongside different licensing arrangements, retrieval architectures (how AI finds its answers) and institutional infrastructures.


Organised around a Collaborative Assessment Design Decision-Making Map, ADA was designed to scaffold reflection, ask meaningful questions, present alternative design options and support educators in making context-sensitive decisions. Yet our experiences varied substantially: an agent that worked well on one AI platform could become verbose, generic or lose conversational context in another.

 

Our initial interpretation focused on unequal AI access, however as the project progressed, a more complex picture emerged. Platform capabilities change rapidly, procurement decisions evolve and universities continually adapt their AI strategies. The more durable differences may not lie in which AI platform institutions can access, but in how effectively institutions evaluate, govern, integrate and adapt AI within educational practice.

 

This raises broader questions for the sector. How transferable are AI-powered educational innovations across institutions? How much does the performance of an AI agent depend on organisational knowledge rather than generic LLM model capability? And what organisational capabilities will universities need as AI platforms continue to evolve?


Different AI platforms, different performances of ADA

What follows should be read as a set of implementation observations, not a definitive ranking of AI platforms. Models, interfaces and licence arrangements change quickly, so these comparisons are snapshots of the configurations we tested.

 

The Claude implementation generally followed ADA’s intended decision-making process and made effective use of the project knowledge base, although its behaviour varied across model versions (e.g. Haiku, sonnet and Opus). Microsoft Copilot could generate useful alternative assessment design ideas. But colleagues using free or restricted versions found it difficult to verify whether Copilot was drawing on the supplied knowledge base. Its behaviour also appeared to shift with the licence tier and with whether it was accessed through a browser or Microsoft Teams. Gemini referred more explicitly to supplied documents, but this could fragment the conversation, and it sometimes moved towards assessment design outputs before the underlying pedagogic problem had been sufficiently explored.

 

These differences made it difficult to separate the effects of ADA’s blueprint prompt from those of the model, retrieval architecture, interface, licence or account configuration. The same educational design concept can become a substantially different tool depending on the platform and configuration through which it is implemented. Platform choice, configuration and ongoing evaluation are part of educational design, not just technical implementation.



Image showing Claude chat about AI and assessment



The knowledge base is more than a collection of documents

ADA combines its blueprint prompt with a knowledge base containing assessment frameworks, pedagogical principles, institutional policies and guidance on matters such as feedback and the use of AI in assessment. It is tempting to treat that knowledge base as a collection of documents that helps the LLM model produce more accurate answers, but this description understates its significance.

 

An institutional knowledge base embodies organisational knowledge. It contains accumulated expertise, policy decisions, educational values and assumptions about what good assessment looks like in a particular context, and reflects the institution's programmes, disciplines, student population, regulatory environment and appetite for educational innovation. If the prompt tells ADA how it should conduct a conversation, the knowledge base helps determine whose knowledge informs that conversation.

 

This matters because a knowledge base only counts if the AI agent can actually use it. If the AI platform cannot reliably retrieve and prioritise those documents, they may be technically present but absent from the resulting advice. The practical challenge is to establish whether the AI platform can use the knowledge base at all. This is surprisingly difficult to verify. A fluent answer can create the impression that a model has consulted institutional guidance when it may instead be drawing on general knowledge from the wider web.

 

The distinction becomes clearer when the knowledge base contains institution-specific terminology, rules or constraints: can the agent reproduce these accurately, explain which document informed its recommendation, distinguish institutional requirements from optional external practices, and acknowledge when the available documents do not contain an answer? Without that kind of evidence, a knowledge base risks becoming symbolic. Its presence reassures users that the agent is institutionally grounded, while its actual recommendations remain generic. This has real implications for trust. 



Image showing ADA's artefact for decision tree


Institutional context is more than a knowledge base

It’s important to acknowledge that the knowledge base is just one kind of knowledge needed for the human-centred decision-making process. Academics and students involved in using ADA will bring their own needs, situated practices and disciplinary knowledge. However, these may not align with the given knowledge base. Facilitating assessment redesign is a challenging process with many competing interests that need to be negotiated. For example, it’s not enough to say assessments must be valid and reliable when in many contexts increasing one can detrimentally affect the other. ADA needs not only to be able to retrieve the knowledge base but also to interpret it in the context of the current conversation it is having.

 

This kind of negotiation and interpretation requires a firm understanding of which parts of the knowledge base represent firm boundaries and which parts can be treated more as suggestions, good practice or starting points to be augmented by the human inputs. For example, when ADA was prompted about practices not present in the knowledge base but with some support in the literature, some versions would be willing to provide suggestions. A human facilitator with an understanding of the broader context  would have recognised the omission of these practices as an indicator of insufficient institutional support and reframed the ideas to make them more institutionally acceptable.

 

This highlights how policies and guidance are enacted within a context which affects how they are interpreted. What a document means is not just the words on the page. As well as issues of whether the knowledge base is being used, we must also contend with questions of how it’s being interpreted and how we illuminate more of the surrounding institutional context.



The risk of homogenisation

AI platforms are very good at producing plausible versions of familiar practice. This can be helpful because much educational guidance is shared across institutions. The risk arises when generic practice is presented as context-sensitive advice.

 

If an AI agent does not reliably use institutional knowledge, it must fill the gaps from elsewhere, typically from dominant patterns in its training data, including conventional assumptions about modules, classrooms, students and assessment formats. We noticed, for example, ADA’s tendency to assume that teaching would take place in person: recommendations involving classroom activities or live participation were offered without first checking the delivery mode. The advice could be revised once the user explained that the programme was online, but the initial assumption revealed the model's default educational setting. Similar defaults may favour undergraduate provision, module structures and familiar approaches to grading.

 

The mechanism of homogenisation is subtle. Institutions supply their own policies and frameworks, but platforms may use these materials inconsistently. Generic model knowledge fills the resulting gaps, and the assistant recommends familiar and statistically dominant practices. Educators at different institutions may consequently receive increasingly similar advice. Over time, an AI agent can appear to personalise assessment design while quietly narrowing the range of possibilities considered.

 

This does not mean that every AI-generated similarity is undesirable. Shared principles such as fairness, validity and constructive alignment remain important. The concern is convergence: recommendations shaped by what is most prominent in training data rather than what is educationally appropriate in a specific setting. Preventing this kind of convergence is not simply a matter of choosing a more capable model. It increasingly depends on how effectively institutions govern, evaluate and maintain their AI-powered educational environments over time.



Institutional strategy is more than AI access

At sector level, the implementation differences we observed have consequences for sustainability, licensing, procurement and infrastructure. Some institutions host open-source models or build their own retrieval infrastructure, while others do not want to or cannot prioritise the development of these capabilities across their institutions.

 

These institutional arrangements are dynamic rather than fixed. Universities regularly review procurement decisions, governance arrangements and AI platforms as technologies evolve. Likewise, providers continuously update foundation models, retrieval architectures and interfaces, meaning that the performance of the same knowledge base may change over time. Institutional agility—the ability to evaluate and respond to these changes—may prove as important as access to any particular model.

 

Even when an institution can create a sophisticated AI agent, sharing it may generate additional costs. Users without equivalent licences may need to be covered through pay-as-you-go charging or pre-purchased capacity, so a successful pilot can become difficult to distribute sustainably.

 

Taken together, licensing, infrastructure, governance, leadership, procurement decisions and institutional agility determine whether a pilot can be maintained, distributed and adapted. A well-resourced institution may prioritise governed, context-sensitive AI support grounded in local policy, while another may prioritise generic public tools that do not reliably draw on its organisational knowledge. The resulting inequality is therefore practical as well as technical: it concerns whether institutions can provide AI that is reliable, contextualised and governable for educational use.



What institutions can do

Before deploying AI for assessment design, institutions should test whether the system genuinely uses their organisational knowledge. Much of this can be done directly. An agent can be asked to retrieve distinctive information from the knowledge base, its quotations can be checked against their original sources, and its responses can be compared with and without the institutional documents in place to see what difference they make. Running identical scenarios across different platforms and licence tiers helps reveal how far behaviour depends on configuration rather than design, and deliberately testing online, postgraduate, interdisciplinary and other non-conventional contexts shows whether the assistant can move beyond its default assumptions. It is also worth examining whether it distinguishes policy requirements from general suggestions, and recording the model version, interface and configuration used for each test so that results remain comparable over time.

 

Institutions should therefore evaluate not only the technical performance of AI platforms but also examine their own organisational capability to embed and govern them effectively. Procurement decisions, staff expertise, evaluation processes, leadership and the ability to respond quickly to technological change should all be considered part of an institution's AI capability, rather than simply technical implementation issues.



A sector-wide challenge

Our experience leaves several questions unresolved. What minimum retrieval standard should an AI agent meet before it can credibly claim to use institutional knowledge? Can knowledge bases, divorced from the contexts in which they’re enacted, be reliably interpreted by AI? How can institutions preserve their educational identity while building AI agents that remain transferable across different technical environments? How should universities evaluate AI agents whose behaviour changes as models, interfaces and retrieval architectures evolve? And how can institutions collaborate on educational innovation without becoming dependent on a single provider or technological ecosystem?

 

The future of AI-powered educational innovations will not be determined simply by which institutions stipulate access to particular AI models. It will increasingly depend on how effectively institutions evaluate, govern, integrate and continuously adapt AI within their educational processes. As foundation models continue to evolve, today's technological advantages may prove temporary. More enduring advantages are likely to come from informed leadership, thoughtful procurement, robust governance, effective evaluation of situated practices and the ability to respond strategically to rapid technological change.

 

The challenge for higher education is therefore broader than technological inequality. It requires an understanding of how technology and values are mutually shaped. When we consider how we could or should teach and assess, technology is already built into our thinking. From whiteboards to PowerPoints to the buildings and chairs themselves, it is difficult to make these value-based judgments divorced from technology. Universities that develop the organisational capabilities to evaluate the shifting relationships between students, teachers, systems and technologies in the contexts where they are enacted will be best placed to shape and reshape these relationships in the future.

 

Viewed in this way, the central question is no longer simply who has access to AI. It is how staff and students are supported to evaluate, challenge and transform their relationships with AI as technology, universities and industry continue to evolve.