Prompt and Model

Labour Behind The Models

Deployment statusActive, deprecated, or restricted
Model typeClassification, regression, generative, etc.
Input data typeText, image, tabular, etc.

Origin and history

Labour Behind The Models is an investigative project and digital platform originating in the United Kingdom. It was first launched and documented in the late 2010s, emerging from a growing public and academic focus on ethics within the technology sector. The project was founded by researchers and journalists concerned with the social and economic conditions of data work that fuels artificial intelligence. Its creation aligns with a period of increased scrutiny of the global supply chains behind AI systems, moving beyond a sole focus on algorithms to the human labor involved. The platform's foundational research involved tracing the networks of data annotation and content moderation work outsourced to various countries. Its historical context is firmly rooted in the critical study of the political economy of digital technology and its workforce.

What it is designed for

Labour Behind The Models is designed to audit, document, and publicize the human labor involved in creating and maintaining large-scale machine learning models. Its primary purpose is to make visible the often-invisible workforce that performs tasks like data labeling, content moderation, and algorithmic training. The project specifically investigates the working conditions, pay, management practices, and geographic locations of these workers. It is engineered to serve as a resource for journalists, policymakers, academics, and ethicists seeking to understand the full lifecycle of an AI model. The design is inherently advocacy-oriented, aiming to foster accountability among technology companies that utilize such labor. Ultimately, it functions as a corrective to narratives that frame AI development as purely automated or confined to elite engineering hubs.

Development and versions

The project has developed through successive phases of investigative reporting and research publication rather than discrete software versions. Its initial development focused on building a public repository of case studies and reports detailing specific instances of data labor across different industries. Subsequent development expanded its scope to include interactive maps and databases that visualize the global flow of data work, linking client companies in the Global North to subcontractors and workers in the Global South. The platform has also developed methodological guides for others to conduct similar audits of AI supply chains. Its evolution includes incorporating worker testimonials and collaborating with labor rights organizations to amplify findings. The development trajectory shows a shift from isolated reports toward a more systemic, tool-oriented resource for ongoing investigation.

Overview

Labour Behind The Models operates as a cross between a research institute, a journalism outlet, and a public archive focused on the AI data supply chain. The core of the project is its published investigations, which dissect the labor processes behind specific technologies like facial recognition datasets, large language models, and content moderation systems. It provides a structured framework for analyzing the stages of data work, from collection and cleaning to annotation and validation. The platform categorizes labor types, such as clickwork, microwork, and cognitive piecework, to clarify the nature of the tasks performed. It also tracks the corporate and contractual relationships that distance large tech firms from the individuals performing the foundational work for their AI products. This overview positions it as a key entity in the field of AI ethics and governance.

What to know

A key thing to know is that Labour Behind The Models explicitly challenges the term "artificial" intelligence by highlighting the extensive human effort required. The platform reveals that many celebrated AI breakthroughs are dependent on vast, low-wage, and precarious workforces often located in countries like Kenya, Venezuela, and the Philippines. It is important to understand that the project defines "the model" not just as the algorithmic architecture but as the entire socio-technical system, including its human components. Readers should know that its findings are used to argue for stronger labor protections and ethical sourcing guidelines within the tech industry. The platform also demonstrates how data work replicates historical patterns of colonial extraction, where raw material (data) is processed abroad for value creation elsewhere. Knowing this context is crucial for anyone involved in deploying or governing machine learning models responsibly.

Common questions

A common question is whether the labor documented is directly employed by major AI companies or through complex subcontracting chains, with the answer almost always being the latter. People often ask what specific job platforms or vendors are involved, with the project naming firms like Sama, Appen, and Scale AI as frequent intermediaries. Many inquire about the typical pay for such tasks, and reports consistently cite rates far below minimum wage standards, often calculated per task rather than per hour. A recurring question is about the psychological impact of this work, particularly for content moderators, with documented cases of trauma and insufficient mental health support. Readers frequently seek recommendations for ethical alternatives, a question the project addresses by discussing the principles of fair work and data dignity rather than endorsing specific vendors. Another common query is how developers can audit their own supply chains, prompting the project's methodological publications.

Pros and cons

A significant pro is that the project provides an essential, evidence-based counter-narrative to the myth of fully automated AI, grounding ethical discussions in concrete labor practices. It offers a valuable due diligence resource for organizations seeking to understand and mitigate the social risks of their AI deployments. A con is that its focus on exposing problems can sometimes outpace the presentation of immediately actionable, scalable solutions for developers under commercial pressure. Organizations may regret consulting it if they are seeking a simple, compliant vendor list, as its work often complicates procurement decisions by revealing flaws in major subcontractors. A common mistake is to treat its findings as isolated ethical concerns rather than as systemic features of the current AI economy. The project's advocacy stance, while a strength for accountability, can be perceived as overly adversarial by entities looking for collaborative, incremental reform.

Who it suits

This resource best suits AI ethicists, academic researchers, and investigative journalists who require deep, sourced analysis of the AI labor supply chain. It is highly suitable for policymakers and regulatory bodies drafting legislation or guidelines around ethical AI and digital labor rights. Corporate social responsibility officers and procurement teams at technology companies can use it for risk assessment, though they must be prepared for challenging findings. The platform suits educators and students in fields like science and technology studies, sociology, and computer science who need to understand the human dimensions of technical systems. It is less suited for software engineers or product managers seeking quick, technical implementation guides without this contextual depth. Ultimately, it is a critical tool for anyone whose role requires a comprehensive understanding of the true cost and composition of contemporary machine learning models.

Latest Labour Behind The Models news

Latest reporting