When the 2026 FIFA World Cup kicked off, it came with AI-assisted offside calls, a sensor-equipped ball, 3D scans of all 1,248 players, and an AI assistant for every national team. Real-time tracking, recruitment models and tactical dashboards are now ordinary tools of elite football. There are workers behind every one of those data points are.

World Cup conversations about technology usually stop at the screen: the offside line or the live stats. Few people ask who produces the data underneath, where, and under what conditions. In my new research project, Tech Workers in Football, based at the University of Toronto’s Creative Labour and Critical Futures cluster, I am mapping the workforce behind football’s data value chains.

Artificial intelligence, anywhere, runs on data, on the human labour that annotates and validates that data, and on physical infrastructure. Football has been relying on this kind of work far longer than the current AI excitement suggests. More than a decade ago, one of UK’s biggest clubs, Arsenal, bought a small data analytics company it had been working with and folded it into the club as an in-house data science department, a company whose match footage was coded, in turn, by a data workforce in Cambodia and Laos. Football’s data workforce has been taking shape for more than ten years, and it has been overlooked in scholarship and media.

There are three main layers in football’s data value chains. Closest to the game are the in-house tech workers: the analysts and data scientists clubs employ directly, working alongside coaches. Even among clubs in the same league, there is no single way to organize this work. The teams go by different names and sit in different corners of the club, the working arrangements differ, and the staff are drawn from sharply different backgrounds, such as doctorates in physics or mathematics or people recruited from the big tech companies, and, in places like Brazil, a physical-education degree. These setups vary from one club to the next, and clubs tend to keep them secret. To date, no research documents the profile of these tech workers in football.

Beyond the clubs are the data vendors, and they are not only one kind of company. Some collect the official event data, the structured log of on-ball actions, and a handful of them also hold the rights to gather and distribute a league’s official feed for media and betting. Others specialize in tracking, using stadium cameras to fix the position of every player on the pitch, and their engineers are the ones who turn raw video into data. One Paris startup, founded in 2016, uses computer vision and machine learning to extract player-tracking data from broadcast camera feeds. The core problem it solved is that broadcast cameras follow the ball, leaving players away from the action off-screen. The company built models to extrapolate the positions of all 22 players throughout the full 90 minutes, even when they are not visible on screen.

Around them there is a wider ecosystem: makers of GPS and inertial wearables that measure how far and how hard players run, video platforms that record and tag matches, scouting databases clubs mine to find their next signing, data analytics consultancies that model performance on their own or others’ data, companies born in the betting industry that now sell predictions, and athlete-management systems that try to forecast injury risk. Across these layers the people doing the work are often bound by non-disclosure agreements and clustered in a few business hubs. And the field is narrowing: a small number of companies now control the data most football clubs rely on, several assembled through waves of acquisitions, and the sector has drawn private-equity and stock-market money as it consolidates.

Furthest from view, at the base, are the data workers who annotate what happens during the game. They watch matches and turn each pass, tackle and shot into structured data, racing the broadcast as they go. The work is concentrated in lower-wage cities: more than 100 data workers tag matches from a single office in Ternopil, Ukraine, and a similar workforce does it in Cairo. At the lowest tier, much of this live data is gathered by people hired match by match as independent contractors. One German company, now part of an Australian sports-tech company, has its matches annotated by a data work team in the Philippines, where data workers can spend three to four hours on a single game.

In the book Expected Goals, the journalist Rory Smith recounts that new data workers in Manila learn the job on one match: Germany’s 7–1 win over Brazil in the 2014 World Cup semi-final. Despite Brazil shooting more often and registering more touches, they were crushed. It teaches annotators which other factors to weigh when they watch and tag a game. This work is behind the stands in almost all public debate about football.

These layers sustain how football is now watched and managed in its different dimensions, such as the graphics on the broadcast and the win-probability figure. These different workers in data value chains are now essential to football. However, they are little known to the wider public. Even the data scientists, the most prominent of these workers, are rarely known by name outside their clubs.

The data value chain in football also has a geographical component. The data science work is located in a handful of wealthy centres, while the data annotation work concentrates in cities across Eastern Europe, Africa, South Asia and Southeast Asia. Clubs outside the dominant leagues, in Brazil, for example, often pay those vendors in foreign currency for the data their own tech workers depend on.

But it would be a mistake to understand such leagues outside Europe as simply behind. Brazilian football has built an arrangement of its own. It has data companies that field their own teams to log every match off the television feed, and data consultancies that train the analysts clubs now scramble to hire. And it has a data science labour market that runs within South America, built around in-house data analytics departments, circulating between Brazilian clubs.

Argentina has  a similar story. ATENEA ID supplies official event data to all 30 clubs in the Liga Profesional. LibroDePases, an AI-assisted scouting platform, is the league’s official AI partner. Genius Sports posts job listings for football data collectors in Argentina, and the workers are contracted game by game via a proprietary app.

Understanding the complex politico-economic and geographic relations in football’s data value chains is a task for future research. For instance, a growing number of investors own several clubs in different countries and run them as a single portfolio. Methods, data and staff move between those clubs as internal transfers: one Brazilian club’s leadership says it uses, by contract, the same performance analysts, data and software as the English club at the top of its ownership group. RB Leipzig and Red Bull Bragantino in Brazil share scouting tools and players within the Red Bull network.

The World Cup will put football’s data and AI on the spotlight. None of it would exist without the workforce behind it. The football we watch runs on their work as much as on the players’. Policy makers, researchers and the media need to take that workforce more seriously: who they are, where they work, what they earn, and what say they have over the technologies they rely on. And unlike baseball, the game data conquered first, football never fully yields to the numbers. It stays magical, unpredictable. As a Brazilian who loves the game, I know that better than anyone.

Rafael Grohmann

CLCF Co-Director & Assistant Professor

Rafael Grohmann is a Co-lead and Co-Director of the Creative Labour and Critical Futures (CLCF) cluster and an Assistant Professor of Media Studies (Critical Platform Studies) at the University of Toronto. Rafael is the leader of the DigiLabour initiative and founding editor of the Platforms & Society journal.