SRHE Blog

The Society for Research into Higher Education


Leave a comment

From data to policy: building an evidence-based future for skills, work, and learning in Latin America

by Sabur Butt, Hector G. Ceballos, and Michael Fung

Latin America stands at a familiar crossroads. Once again, a technological revolution is reshaping how work gets done, what skills employers value, and how quickly the workforce must adapt. And once again, the region risks being a generation behind.

During the Third Industrial Revolution of the 1970s–1990s, East Asia invested decisively in microelectronics, computing, and technical education. Latin America, consumed by debt crises and macroeconomic instability, missed the wave. The cost was measured not only in lost output but in institutional habits of looking backward while others prepared to look forward. Today, the Fourth Industrial Revolution is unfolding faster than its predecessor, and the region’s central deficit is not money or talent. It is foresight leading to informed interventions.

A structural mismatch between work and learning

Industries now change in cycles of months. University curricula change in cycles of years. That gap is no longer a minor inefficiency; it is the bottleneck that determines whether a country prepares its workforce in time or trails behind it.

The numbers are sobering. Across the six largest Latin American economies, between 19% and 21% of workers face a high probability of automation-related displacement. Almost half work in informal jobs. University dropout rates exceed 50% in several countries, climbing as high as 76% in the Dominican Republic. Students perform consistently below OECD averages in foundational literacy and numeracy. By the time graduates from the educational system enter the labour market, there is already a structural gap between what they can do and what employers want.

What the region needs is not another labour-market report. It needs skills intelligence, a continuous, forward-looking system that informs policymakers what is rising, what is fading, and where the workforce can realistically move next.

Why traditional forecasting falls short

Most forecasting in the region still relies on expert panels, occupational surveys, and static taxonomies that take years to update. These tools were designed for a slower-paced world.

To illustrate the issue, consider how such static taxonomies treat “database management” as a single skill, when the real labour market is moving between MySQL, PostgreSQL, MongoDB, Redis, Cassandra, and now vector databases, each with a different demand trajectory. Or consider the noise in a single automotive dataset, where the same assembly-line job appears as “ensamblador de carrocerías”, “operador de ensamble”, “ensamblador automotriz”, and “técnico en armado de vehículos”. Human annotators cannot keep up; even expert raters disagree on whether two skill descriptions refer to the same thing.

By the time a traditional taxonomy recognizes “prompt engineering” as a skill category, the labor market has already moved on to whatever comes next.

A data-driven alternative, already running

In our paper, we describe an operational deployment at the Institute for the Future of Education at Tecnológico de Monterrey. It is not a proposal; it is a working system (See Figure 1).

Figure 1. The framework links real-time labour market signals to an AI-assisted, expert-validated skills intelligence layer, which in turn informs curriculum design,reskilling pathways, workforce policy, and governance.

The approach combines large language models with retrieval-augmented generation to build dynamic, hierarchical skill taxonomies that update themselves as new job postings flow in. Each new skill is matched semantically against the existing taxonomy. Known skills are normalized to canonical terms. Genuinely new ones are flagged, classified, and added.

In Mexico’s automotive sector alone, the system has mapped more than 11,000 skill variations across 220 hypernym categories, identified 847 unique skills clustered into 12 occupational groups, and tracked the rise of electric-vehicle competencies in real time. Generative AI skills surfaced in the data months before they appeared in any official classification.

The infrastructure required is modest. A taxonomy covering 10,000 skills across major sectors can be maintained on standard cloud infrastructure for roughly $500–1,000 USD per month, within reach of education ministries in developing economies.

From data to policy

Technology alone does not change a system. The harder work is institutional. To shorten the lag between detecting a skill shift and updating a training program (currently 18 to 24 months) in most regions, curriculum committees must meet more often, must accept real-time data alongside surveys, and procurement procedures must allow timely equipment purchases according to the emerging skills.

Governance matters equally. Sustainable implementations bring together labour ministries, education ministries, economic development ministries, national statistics offices, industry associations, and universities. Each contributes something the others cannot: data access, curriculum authority, methodological rigor, domain expertise, and research capacity. No single actor owns skills or intelligence; the legitimacy of the system depends on shared ownership.

Localization also matters. Global taxonomies like ESCO and O*NET are useful starting points, but they need to incorporate regional terminology, indigenous skill categories, and sector-specific competencies. A skill system that does not speak the local language of work will not be trusted by the people meant to use it.

We also acknowledge real-world limitations. Many countries still hire through newspaper classifieds, physical noticeboards, and informal networks. Supplier tiers and small enterprises seldom advertise online. A system built solely on digital postings produces a geographically and structurally biased picture. Integrating offline data sources, replicating the approach across countries, and validating its predictions over time are the necessary next steps.

What success would look like

Executed well, skills intelligence reshapes reskilling itself. Instead of generic, fixed-duration programs, workers receive personalized pathways: a manufacturing technician’s quality-control experience mapped onto automated-systems monitoring; a mid-career professional alerted 12–24 months ahead of emerging demand; a young job-seeker pointed toward a stackable micro-credential that the market will actually reward.

For Latin America, this is more than a technical upgrade. It is a chance to break a historical pattern of arriving late to every industrial revolution and start arriving prepared. The region has the data, the talent, and increasingly the tools. What it has lacked is the institutional capacity to anticipate. That capacity is now within reach.

The Fourth Industrial Revolution will not wait. But for the first time, neither does the evidence.

Reference: Butt, S, Ceballos, HG, and Fung, M (2026) ‘From Data to Policy: Building an Evidence-Based Future for Skills, Work, and Learning in Latin AmericaPolicy Reviews in Higher Education 

Sabur Butt is a research professor at the Institute for the Future of Education, Tecnológico de Monterrey. His work focuses on artificial intelligence, natural language processing, and dynamic skills taxonomies, with a particular interest in how AI-driven labour-market intelligence can inform education and workforce policy across Latin America.

Hector G Ceballos is Director of the Living Lab & Data Hub of the Institute for the Future of Education (IFE) at Tecnológico de Monterrey. His research spans data science, knowledge engineering, and educational analytics, with a focus on building evidence-based systems that connect higher education to evolving industry and labour-market needs.

Michael Fung is Executive Director of the Institute for the Future of Education at Tecnológico de Monterrey. He was formerly the Deputy Chief Executive at SkillsFuture Singapore (SSG). He led the development of a comprehensive education and training ecosystem under the national SkillsFuture movement, which has become a global benchmark and reference for workforce skills development and lifelong learning across society.


Leave a comment

Institutional constraints to higher education datafication: an English case study

by Rachel Brooks

‘Intractable’ datafication?

Over recent years, both policymakers and university leaders have extolled the virtues of moving to a more metricised higher education sector: statistics about student satisfaction with their degree programme are held to improve the decision-making processes of prospective students, while data analytics are purported to help the shift to more personalised learning, for example. Moreover, academic studies have contended that datafication has become an ‘intractable’ part of higher education institutions (HEIs) across the world.

Nevertheless, our research (conducted in ten English HEIs, funded by TASO) – of data use with respect to widening participation to undergraduate ‘sandwich’ courses (where students spend a year on a work placement, typically during the third year of a four-year degree programme) – indicates that, despite the strong claims about the advantages of making more and better use of data, in this particular area of activity at least, significant constraints operate, limiting the advantages that can accrue through datafication.

Little evidence of widespread data use

Our interviewees were those responsible for sandwich course provision in their HEI. While most thought that data could offer useful insights into the effectiveness of their area of activity, there was little evidence of ‘intractable’ data use. This was for three main reasons. First, in some cases, interviewees explained that no relevant data were collected – in relation to access to sandwich courses and/or the outcomes of such courses. Second, in some HEIs, relevant data were collected but not analysed. Such evidence tends to support the contention that ‘data lakes’ are emerging, as HEIs collect more and more data that often remain untapped. Third, in other cases, appropriate data were collected and analysed, but in a very limited manner. For example, one interviewee explained how data were collected and analysed in relation to the participation of students from under-represented ethnic groups, but not with respect to any other widening participation categories. This limited form of datafication, in which only some social characteristics were datafied, was not, therefore, able to inform any action with respect to the participation of widening participation students generally. Indeed, across all ten HEIs, there was only one example of where data were used in a systematic fashion to help analyse who was accessing sandwich courses within the institution, and the extent to which they were representative of the wider student population.

Constraints on data use

Lack of institutional capacity

In explaining this absence of data use, the most commonly identified constraint was the lack of institutional capacity to collect and/or analyse appropriate data. For example, one interviewee commented that they did not have a very good data system for placements – ‘we are still quite Excel- based’. Excel spreadsheets were viewed as limited as they could not be easily shared or updated, and data were relatively hard to manipulate. This, according to the interviewee, made collection of appropriate data laborious, and systematic analysis of the data difficult. Interviewees also pointed to the limited time staff had available to analyse data that the institution had collected.

Prioritisation of ‘externally-facing’ data

Several interviewees described how ‘externally-facing data’ – i.e. that required by regulatory bodies and/or that fed into national and international league tables – was commonly prioritised, leaving little time for information officers to devote to generating and/or analysing data for internal purposes. One interviewee, for example, was unclear about what data, if any, were collected about equity gaps but believed that they were generally only pulled together for high-level reports ‘such as for the TEF’.

Institutional cultures

A further barrier to using data to analyse access to and outcomes of sandwich courses was perceived to be the wider culture of the institution, including its attitude to risk. An interviewee explained that the data collected in their institution was limited to two main variables – subject of study and fee status (home or international) – because of ‘ongoing cautiousness at the university about how some of that data is used and how it’s shared with different teams’.

In addition, many participants outlined the struggles they had faced in gaining access to relevant data, and in influencing decisions about what should be collected and what analyses should be run. Several spoke of having to ‘request’ particular analyses to be run (which could be turned down), leading to a fairly ad hoc and inefficient way of proceeding, and illustrating the relative lack of agency accorded to staff – typically occupying mid-level organisational roles – in accessing and manipulating data.

Reflections

Examining a discrete set of activities within the UK higher education sector – those relating to sandwich courses – provides a useful lens to examine quotidian practices with respect to the availability and use of data. Despite the strong emphasis on data by government bodies and HEI senior management teams, as well as the claims made about the ‘intractability’ of HEI data use in the academic literature, our research suggests that datafication is perhaps not as widespread as some have claimed. Indeed, it indicates that some areas of activity – even those linked to high profile political and institutional priorities (in this case, employability and widening participation) – have remained largely untouched by ‘intractable’ datafication, with relevant data either not being collected or, where it is collected, not being made available to staff working in pertinent areas.

As a consequence, the extent to which students from widening participation backgrounds were accessing sandwich courses – and then succeeding on them – relative to their peers typically remained invisible. While the majority of our interviewees were able to speculate on the extent of any under-representation and/or poor experience, this was typically on the basis of anecdotal evidence and their own ‘sense’ of how inequalities were played out in this area. Although reflecting on professional experience is obviously important, many inequalities may not be visible to staff (for example, if a student chooses not to talk about their neurodiversity or first-in-family status), even if they have regular contact with those eligible to take a sandwich course. Moreover, given the status often accorded to quantitative data within the senior management teams of universities, the lack of any statistical reporting about inequalities by social characteristic, as they pertain to sandwich courses, makes it highly likely that such issues will struggle to gain the attention of senior leaders. The barriers to the effective use of metrics highlighted above may thus have a direct impact on HEIs’ capacity to recognise and address inequalities.  

The research on which this blog is based was carried out with Jill Timms (University of Surrey) and is discussed in more detail in this article Institutional constraints to higher education datafication: an English case study | Higher Education

Rachel Brooks is Professor of Higher Education at the University of Oxford and current President of the British Sociological Association. She has conducted a wide range of research on the sociology of higher education; her most recent book is Constructing the Higher Education Student: perspectives from across Europe, published (open access) with Policy Press.

Vicky Gunn


Leave a comment

Learning Analytics, surveillance, and the future of understanding our students

By Vicky Gunn

There has been a flurry of activity around Learning Analytics in Scotland’s higher education sector this past year. Responding no doubt to the seemingly unlimited promises of being able to study our students, we are excitedly wondering just how best to use what the technology has to offer. At Edinburgh University, a professorial level post has been advertised; at my own institution we are pulling together the various people who run our student experience surveys (who have hitherto been distributed across the institution) into a central unit in Planning so that we can triangulate surveys, evaluations and other contextual data-sets; elsewhere systems which enable ‘early warning signals’ with regards to student drop-out have been implemented with gusto.

I am one of the worst of the learning analytics’ offenders.  My curiosity to observe and understand the patterns in activity, behaviour, and perception of the students is just too intellectually compelling. The possibility that we could crunch all of the data about our students into one big stew-pot and then extract answers to meaning-of-student-life questions is a temptation I find too hard to resist (especially when someone puts what is called a ‘dashboard’ in front of me and says, ‘look what happens if we interrogate the data this way’). Continue reading