# Generative AI is a Crisis for Copyright Law

**Authors:** Kate Crawford, Jason M. Schultz
**Citation:** "Generative AI is a Crisis for Copyright Law," *Issues Sci. & Tech.*, Winter 2024, at 79 (with Kate Crawford)
**Source:** https://issues.org/wp-content/uploads/2023/12/76-88-An-AI-Society-Winter-2024.pdf

## artificial intelligence

*p. 3*
stories it can generate-and what stories it will not generate. It's no secret that AI is biased. Researchers recently asked the image generator Midjourney to create images of Black physicians treating impoverished white children, but the system would only return images depicting the children as Black. Even a er several iterations, Midjourney failed to produce the speci ed results. e closest it got to the prompt was a shirtless medicine man with feathers, leather bands, and beads, gazing at a similarly garbed blond child.

*p. 3*
Here's something that hits close to home: the potential loss of Cardamom Town. orbjørn Egner's Folk og røvere i Kardemomme by (When the Robbers Came to Cardamom Town) is a children's book and musical well known to anyone who grew up in Norway or Denmark a er 1955. e songs and stories have been played, read, and sung in homes and preschools for decades; there's even a theme park inspired by the book in the city of Kristiansand. e story features three comical thieves who steal food because they are hungry and Amy Karle is a contemporary artist who uses artificial intelligence as both a medium and a subject in her work. Karle has been deeply engaged with AI, artificial neural networking, machine learning, and generative design since 2015. She poses critical questions about AI, illuminates future visions, and encourages us to actively shape the future we desire.

*p. 3*
One of Karle's projects focuses on how AI can help design and grow biomaterials and biosubstrates, including guiding the growth of mycelium-based materials. Her approach uses AI to identify, design, and develop diverse bioengineered and bioinspired structures and forms and to refine and improve the structure of biomaterials for greater functionality and sustainability. Another project is inspired by the seductive form of corals. Karle's speculative biomimetic corals leverage AI-assisted biodesign in conjunction with what she terms "computational ecology" to capture, transport, store, and use carbon dioxide. Her goal with this series is to help mitigate carbon dioxide emissions from industrial sources such as power plants and refineries and to clean up highly polluted areas.

*p. 3*
Amy Karle: AI-Assisted Biodesign "The future with AI does not have to be something that happens to us, it is something that we can cocreate." -Amy Karle don't understand that work is necessary. A er being caught stealing sausages and chocolate, they are rehabilitated by the kind police o cer and townsfolk, then end up saving the town from a re.

*p. 3*
is story is more than a shared cultural reference-it supports the Norwegian criminal justice system's priority of rehabilitation over punishment. It is distinct from Disney movies, with their unambiguous villains who are punished at the end, and from Hollywood bank heists and gangster movies that glorify criminals. Generative AI might well bury stories like Cardamom Town by stu ng chatbot responses and search results worldwide with homogenized American narratives.

*p. 3*
Narrative archetypes give us templates to live by. Depending on the stories we hear, share, and create, we shape possibilities for action and for understanding. We learn that criminals can be rehabilitated, or that they deserve to come to a bad end.

*p. 4*
Generative arti cial intelligence is driving copyright into a crisis. More than a dozen copyright cases about AI were led in the United States last year, up severalfold from all lings from 2020 to 2022. In early 2023, the US Copyright O ce launched the most comprehensive review of the entire copyright system in 50 years, with a focus on generative AI. Simply put, the widespread use of AI is poised to force a substantial reworking of how, where, and to whom copyright should apply.

*p. 4*
Starting with the 1710 British statute, "An Act for the Encouragement of Learning," Anglo-American copyright law has provided a framework around creative production and ownership. Copyright is even embedded in the US Constitution as a tool "to promote the Progress of Science and useful Arts." Now generative AI is destabilizing the foundational concepts of copyright law as it was originally conceived.

*p. 4*
Typical copyright lawsuits focus on a single work and a single unauthorized copy, or "output," to determine if infringement has occurred. When it comes to the capture of online data to train AI systems, the sheer scale and scope of these datasets overwhelms traditional analysis. e LAION 5-B dataset, used to train the AI image generator Stable Di usion, contains 5 billion images and text captions harvested from the internet, while CommonPool (a collection of datasets released by nonpro t LAION in April to democratize machine learning), o ers 12.8 billion images and captions.

*p. 4*
Generative AI systems have used datasets like these to produce billions of outputs.

*p. 4*
For many artists and designers, this feels like an existential threat. eir work is being used to train AI systems, which can then create images and texts that replicate their artistic style. But to date, no court has considered AI training to be copyright infringement: following the Google Books case in 2015, which assessed scanning books to create a searchable index, US courts are likely to nd that training AI systems on copyrighted works is acceptable under the fair use exemption, which allows for limited use of copyrighted works without permission in some cases when the use serves the public interest. It is also permitted in the European Union under the text and data mining exception of EU digital copyright law.

*p. 4*
Copyright law has also struggled with authorship by AI systems. Anglo-American law presumes that work has an "author" somewhere. To encourage human creativity, some authors need the economic incentive of a time-limited monopoly on making, selling, and showing their work. But algorithms don't need incentives. So according to the US Copyright O ce they aren't entitled to copyright. e same reasoning applied to other cases involving nonhuman authors, including the case where a macaque took sel es using a nature photographer's camera. Generative AI is the latest in a line of nonhumans deemed un t to hold copyright.

*p. 4*
Nor are human prompters likely to have copyrights in AIgenerated work. e algorithms and neural net architectures behind generative AI algorithms produce outputs that are inherently unpredictable, and any human prompter has less control over a creation than the model does.

*p. 5*
Where does this leave us? For the moment, in limbo. e billions of works produced by generative AI are unowned and can be used anywhere, by anyone, for any purpose. Whether a ChatGPT novella or a Stable Di usion artwork, output now exists as unclaimable content in the commercial workings of copyright itself. is is a radical moment in creative production: a stream of works without any legally recognizable author.

*p. 5*
ere is an equivalent crisis in proving copyright infringement. Historically, this has been easy, but when a generative AI system produces infringing content, be it an image of Mickey Mouse or Pikachu, courts will struggle with the question of who is initiating the copying. e AI researchers who gathered the training dataset? e company that trained the model? e user who prompted the model? It's unclear where agency and accountability lie, so how can courts order an appropriate remedy?

*p. 5*
Copyright law was developed by eighteenth-century capitalists to intertwine art with commerce. In the twentyrst century, it is being used by technology companies to allow them to exploit all the works of human creativity that are digitized and online. But the destabilization around generative AI is also an opportunity for a more radical reassessment of the social, legal, and cultural frameworks underpinning creative production.

*p. 5*
What expectations of consent, credit, or compensation should human creators have going forward, when their online work is routinely incorporated into training sets? What happens when humans make works using generative AI that cannot have copyright protection? And how does our understanding of the value of human creativity change when it is increasingly mediated by technology, be it the pen, paintbrush, Photoshop, or DALL-E?

*p. 5*
It may be time to develop concepts of intellectual property with a stronger focus on equity and creativity as opposed to economic incentives for media corporations. We are seeing early prototypes emerge from the recent collective bargaining agreements for writers, actors, and directors, many of whom lack copyrights but are nonetheless at the creative core of lmmaking. e lessons we learn from them could set a powerful precedent for how to pluralize intellectual property. Making a better world will require a deeper philosophical engagement with what it is to create, who has a say in how creations can be used, and who should pro t.

*p. 5*
Kate Crawford is a research professor at the University of Southern California Annenberg, a senior principal researcher at Microso Research, the inaugural chair of AI & Justice at the École Normale Supérieure, and the author of Atlas of AI (Yale University Press, 2021). Jason Schultz is a clinical professor of law at New York University, the codirector of the Engleberg Center on Innovation, Law, and Policy, and the author of e End of Ownership ( e MIT Press, 2016).

## LINNET TAYLOR

*p. 5*
Following Nazi medical experiments in World War II and outrage over the US Public Health Service's fourdecade-long Tuskegee syphilis study, bioethicists laid out frameworks, such as the 1947 Nuremberg Code and the 1979 Belmont Report, to regulate medical experimentation on human subjects. Today social media-and, increasingly, generative arti cial intelligence-are constantly experimenting on human subjects, but without institutional checks to prevent harm.

*p. 5*
In fact, over the last two decades, individuals have become so used to being part of large-scale testing that society has essentially been con gured to produce human laboratories for AI. Examples include experiments with biometric and payment systems in refugee camps (designed to investigate use cases for blockchain applications), urban living labs where families are o ered rent-free housing in exchange for serving as human subjects in a permanent marketing and branding experiment, and a mobile money research and development program where mobile providers o er their African consumers to rms looking to test new biometric and ntech applications. Originally put forward as a simpler way to test applications, the convention of so ware as "continual beta" rather than more discrete releases has enabled business models that depend on the given that it is already being incorporated into the base layers of applications we would be hard-pressed to avoid, the boundaries between the sandbox and the world are unclear.

*p. 6*
Generative AI is an extreme case of unregulated experimentation-as-innovation, with no formal mechanism for considering potential harms. ese experiments are already producing unforeseen ruptures in professional practice and knowledge: students are using ChatGPT to cheat on exams, and lawyers are ling AI-dra ed briefs with fabricated case citations. Generative AI also undermines the public's grip on the notion of "ground truth" by hallucinating false information in subtle and unpredictable ways.

*p. 6*
ese two breakdowns constitute an abrupt removal of what philosopher Regina Rini has termed "the epistemic backstop,"-that is, the benchmark for considering something real. Generative AI subverts information-seeking practices that professional domains such as law, policy, and medicine rely on; it also corrupts the ability to draw on common truth in public debates. Ironically, that disruption is being classed as success by the developers of such systems, emphasizing that this is not an experiment we are conducting but one that is being conducted upon us.

*p. 6*
is is problematic from a governance point of view because much of current regulation places the responsibility for AI safety on individuals, whereas in reality they are the subjects of an experiment being conducted across society.

*p. 6*
e challenge this creates for researchers is to identify the kinds of rupture generative AI can cause and at what scales, and then translate the problem into a regulatory one. en authorities can formalize and impose accountability, rather than creating di use and ill-de ned forms of responsibility for individuals. Getting this right will guide how the technology develops and set the risks AI will pose in the medium and longer term.

*p. 6*
Much like what happened with biomedical experimentation in the twentieth century, the work of de ning boundaries for AI experimentation goes beyond "AI safety" to AI legitimacy, and this is the next frontier of conceptual social scienti c work. Sectors, disciplines, and regulatory authorities must work to update the de nition of experimentation so that it includes digitally enabled and data-driven forms of testing. It can no longer be assumed that experimentation is a bounded activity with impacts only on a single, visible group of people. Experimentation at scale is frequently invisible to its subjects, but this does not render it any less problematic or absolve regulators from creating ways of scrutinizing and controlling it.

*p. 6*
Linnet Taylor is a professor of international data governance at Tilburg University, Netherlands, and leads the European Research Council-funded Global Data Justice project.

## AI Aids the Pretense of Military "Precision"

*p. 6*
LUCY SUCHMAN cial intelligence is the latest promise of a technological solution to the intractable "fog of war." In Ukraine and Gaza, enthusiasts have proclaimed the advent of AI-driven war ghting. In October 2023, Ukrainian technologists con rmed that AI-enabled drones identify and target 64 types of Russian "military objects" without a human operator; meanwhile the Israeli Defense Forces website states that an AI system generates recommended targets, reportedly at an unprecedented rate. Enormous questions arise regarding the validity of the assumptions built into these systems about who comprises an imminent threat and about the legitimacy of their targeting functions under the Geneva Conventions and the laws of war.

*p. 7*
Considering military investments in AI as part of a sociotechnical imaginary is helpful here. Developed within the eld of science and technology studies, the concept of sociotechnical imaginaries describes collectively imagined forms of social order as materialized by scienti c and technological projects. ese include aspirational futures that sustain investments in the military-industrial-academic complex. Iconic examples of AI-enabled war ghting in the present moment include battle management interfaces like Palantir's AI platform.

*p. 7*
To function in the real world, these platforms require very large, up-to-date datasets (of labeled "military objects" or biometric pro les of "persons of interest," for example), from which models can be developed. In the case of threat prediction and targeting, neither the US Department of Defense nor allied militaries make public the details necessary to assess validity. But in the case of predictive policing, an investigation by e Markup found that fewer than 1% of data-based predictions actually lined up with reported crimes. And generative AI introduces new uncertainties: both the provenance of the data and reliability of information are hard to check. at is particularly dangerous for "actionable military intelligence," which is used for targeting and to designate imminent threats.

*p. 7*
We should be deeply skeptical of the promotion of AI as a solution to the fog of war, which imagines that the right technology will nd the important signals amid the noise. is faith in technology constitutes a kind of willful ignorance, as if AI is a talisman that sustains the wider magical thinking of militarism as a path to security. In the words of performance artist Laurie Anderson (quoting her meditation teacher), "If you think technology will solve your problems, then you don't understand technology-and you don't understand your problems."

*p. 7*
Critical inquiry into the realities of war can help challenge the logics through which militarism perpetuates its imaginary of rational and controllable state violence while obscuring war's ungovernable chaos and unjusti able injuries. Although there are valid reasons that military forces exist in today's world, we should question the narratives that underwrite the billions of dollars funnelled into algorithmically based war ghting. We need to redirect resources to creative projects in de-escalation, negotiated settlements that o er true security for all, and eventual demilitarization. While the techno-solutionist imaginaries of militarism are longstanding, so are their limits as a basis for sustainable peace.

*p. 7*
Lucy Suchman is professor emerita of the anthropology of science and technology at Lancaster University in the United Kingdom.

## MARK ANDREJEVIC

*p. 7*
Already, content generated by arti cial intelligence populates the advertisements, news, and entertainment people see every day. According to OpenAI's cofounder Greg Brockman, the technology could fundamentally transform mass culture, making it possible, for example, to customize TV shows for individual viewers: "Imagine if you could ask your AI to make a new ending … maybe even put yourself in there as a main character."

*p. 7*
Brockman meant this as a sort of paradise of customization, but it's not hard to see how such tools could also spew misinformation and other content that would disrupt civic life and undermine democracy. Bad content would drive out good, enacting "Gresham's Law"-the principle that "bad money drives out good"-on steroids. Even top AI executives are begging for regulation, albeit at the level of individual products and their potential dangers. I think a more productive way to frame regulation is as a means of protecting the shared information environment.

*p. 7*
In decades past, the rationale for regulating the information space pivoted on the limited availability of broadcast channels, or "channel scarcity." Public attention can also be considered a nite resource, rationed by what information theorist Tiziana Terranova describes as "the limits inherent to the neurophysiology of perception and artificial intelligence the social limitations to time available for consumption." For democracy to function, people need to pay attention to matters of public import. In an information environment swamped with automatically generated content, attention becomes the scarce resource.

*p. 8*
A world in which attention is monopolized by an endless ow of personalized entertainment might be a consumers' paradise-but it would be a citizen's nightmare. e tech sector has already proposed a model for dispensing with public attention, one that is far from democratic. In 2016, a team at Google envisioned a "Sel sh Ledger"-a data pro le that would infer individuals' goals and then prompt aligned behavior, such as buying healthier food or locally grown produce, and seek more data to tweak the customized model. Similarly, physicist César Hidalgo suggested providing every citizen with a so ware agent that could infer political preferences and act on their behalf. In such a world, the algorithm would pay attention for us: no need for people to learn about the issues or even directly express their opinions.

*p. 8*
Such proposals show how important it is for citizens to actively regulate the information commons. Preserving scarce attention is essential to recapturing an increasingly elusive sense of shared, overlapping, and common interests.

*p. 8*
e world is moving toward a state where the data we generate can be used to further capture and channel our attention according to priorities that are neither our own, nor those of civic life. So ware, and whoever it serves, cannot be allowed to substitute for citizenship, and the economic might of tech giants must be balanced by citizens' ability to access the information they need to exercise their political power.

*p. 8*
Mark Andrejevic is a professor of communication and media studies in the School of Media, Film, and Journalism at Monash University, Australia. He is the author of Facial Recognition (Polity Press, 2022).

## KARINE GENTELET

*p. 8*
As part of my job, I give talks about how arti cial intelligence a ects human rights: to criminology experts, schoolteachers, retirees, union members, First Peoples, and more. Across these diverse groups, I hear common themes. One is that although AI programs could impact how they do their jobs and live their lives, people feel their experience and expertise are completely le out before programs are deployed. Some worry, legitimately, about facing legal action if they protest.

*p. 8*
Plans and policies to regulate AI systems in Europe, Canada, and the United States are not likely to improve the situation. Europe plans to assign regulatory requirements based on application. For example, the high-risk category includes technology used in hiring decisions, police checks, banking, and education. Canadian legislation, still under review by the House of Commons, is based on the same risk assessment. e US president has outlined demands for rigorous safety testing, with results reported to the government. e problem is that these plans focus on laying out guardrails for anticipated threats without establishing an early warning system for citizens' actual experiences or concerns.

*p. 8*
Regulatory schemes based on a rigid set of anticipated outcomes might be a good rst step, but they are not enough. For one thing, some harms are only now emerging. And they could become most entrenched for marginalized, underserved groups because generative AI is trained on biased datasets that then generate new datasets that perpetuate the vicious cycle. A 2021 paper shows how prediction tools in education systems incorporate not just statistical biases (by gender, race, ethnicity, or language) but also understudied sociological ones such as urbanity. For instance, rural learners in Brazil are likely to di er from their urban counterparts with regards to uency in the o cial state language and their access to relevant educational materials, up-to-date facilities, and teaching sta . But because there aren't enough data on speci c groups' learning and schooling issues, their needs would be aggregated into a larger dataset and made invisible. Given the lack of knowledge, it would be di cult to even predict any kind of bias.

*p. 8*
What's needed are mechanisms that support citizens' direct engagement with AI deployments to document, from the ground, potentially high-risk impacts on collective equity. ere are democratic formats already in place to support citizens' perspectives. In Canada, for example, the mandate of the general solicitor or privacy commissioner could be strengthened to review AI deployments in the public sector (audits of datasets, mandatory impact assessments, etc.). ese mechanisms would provide transparent and accountable standards to keep citizens adequately informed about AI deployments, help balance the civic power dynamic, and strengthen social justice.

*p. 8*
Citizens' direct engagement could also be supported through access to courts. ere are few (if any) direct legal recourses available for ordinary people to challenge algorithmic harms in current AI regulatory schemes. Access to courts-and implicitly to justice-could send a clear message about citizens' power to corporations, governments, and, most importantly, to citizens artificial intelligence themselves. In combination with other mechanisms to increase citizen oversight, legal suits would o er not only access to rightful reparations, but also give societal recognition of citizens' rights.

*p. 9*
Sometimes at my talks people tell me they feel illegitimate asking questions about AI's impacts, given their lack of expertise. What I tell them is that they don't need to be a mechanic to know how bad it would be to be hit by a car. Harms from AI are bound to be more subtle, but the point stands. Citizens are the ones primarily a ected, so they must have an active role within AI governance. Emerging regulatory systems should highlight the role of citizens as social actors who contribute-as they should-to the collective good.

*p. 9*
Karine Gentelet is an associate professor of social sciences at the Université du Québec en Outaouais, Gatineau, Canada.

*p. 9*
including tech luminaries such as Elon Musk and Steve Wozniak, signed a call to "pause giant AI experiments" to deal with "profound risks to society." But the question is more complex than restraint versus unfettered technological development. It is about di erent ways to articulate ethical values and, above all, di erent visions of what society should be.

*p. 9*
A double interview in the French journal Le Monde illustrates the distinction. e interviewees, Yoshua Bengio and Yann Le Cun, are friends and collaborators who both received the 2018 Turing Award for their contributions to computer science. But they have radically di erent views on the future of generative AI.

*p. 9*
Bengio, who works at a nonpro t AI think tank in Montreal, believes ChatGPT is revolutionary. at's why he sees it as dangerous. ChatGPT and other generative AI systems work in ways that cannot be fully understood and o en produce results that are simultaneously wrong and credible, which threatens news and information sources and democracy at large. His argument mirrors philosopher Hans Jonas's precautionary principle: since humanity is better at producing new technological tools than foreseeing their future consequences, extreme caution about what AI can do to humanity is warranted. e solution is to establish ethical guidelines for generative AI, a task that the European Group on Ethics, the Organisation for Economic Co-operation and Development, UNESCO, and other global entities have already embraced.

*p. 9*
Le Cun, who works for Meta, does not consider ChatGPT revolutionary. It depends on neural networks trained on very large databases-all technologies that are several years old. Yes, it can produce fake news, but dissemination-not production-is the real risk. Techniques can be developed to ag AI-generated outputs and reveal what text and images have been manipulated, creating something akin to antispam so ware today. For Le Cun, the way to quash the dangers of generative AI will rely on AI. It is not the problem but the solution-a tool humanity can use to make better decisions. But who de nes what is a "better decision"? Which set of values will prevail? Here I see in Le Cun's arguments parallels to the economist and innovation scholar Joseph Schumpeter, who argued that within a democracy, the tools humans use to institutionalize values are the law and government. In other words, regulation of AI is essential.

*p. 9*
ese radically disparate views land on solutions that are similar in at least one aspect: whether generative AI is seen as a technological revolution or not, it is always embedded within a wider set of values. When seen as a for humanity, ethics are mobilized. When social values are threatened, the law is brought in. Either way, the solution is oversight of the corporations building AI.

*p. 9*
The Question Isn't Asset or Threat; It's Oversight

## EMMANUEL DIDIER

*p. 9*
As part of a research group studying generative AI with France's Académie Nationale de Médecine, I was surprised by some clinicians' technological determinism-their immediate assumption that this technology would, on its own, act against humans' wishes. e anxiety is not limited to physicians. In spring 2023, thousands of individuals,

## FLORIAN JATON

*p. 10*
Arti cial intelligence algorithms are human-made, cultural constructs, something I saw rst-hand as a scholar and technician embedded with AI teams for 30 months. Among the many concrete practices and materials these algorithms need in order to come into existence are sets of numerical values that enable machine learning. ese referential repositories are o en called "ground truths," and when computer scientists construct or use these datasets to design new algorithms and attest to their e ciency, the process is called "ground-truthing."

*p. 10*
Understanding how ground-truthing works can reveal inherent limitations of algorithms-how they enable the spread of false information, pass biased judgments, or otherwise erode society's agency-and this could also catalyze more thoughtful regulation. As long as groundtruthing remains clouded and abstract, society will struggle to prevent algorithms from causing harm and to optimize algorithms for the greater good.

*p. 10*
Ground-truth datasets de ne AI algorithms' fundamental goal of reliably predicting and generating a speci c output-say, an image with requested speci cations that resembles other input, such as web-crawled images. In other words, ground-truth datasets are deliberately constructed. As such, they, along with their resultant algorithms, are limited and arbitrary and bear the sociocultural ngerprints of the teams that made them.

*p. 10*
Ground-truth datasets fall into at least two subsets: input data (what the algorithm should process) and output targets (what the algorithm should produce). In supervised machine learning, computer scientists start by building new algorithms using one part of the output targets annotated by human labelers, before evaluating their built algorithms on the remaining part. In the unsupervised (or "selfsupervised") machine learning that underpins most generative AI, output targets are used only to evaluate new algorithms.

*p. 10*
Most production-grade generative AI systems are assemblages of algorithms built from both supervised and self-supervised machine learning. For example, an AI image generator depends on self-supervised di usion algorithms (which create a new set of data based on a given set) and supervised noise reduction algorithms. In other words, generative AI is thoroughly dependent on ground truths and their socioculturally oriented nature, even if it is o en presented-and rightly so-as a signi cant application of self-supervised learning.

*p. 10*
Why does that matter? Much of AI punditry asserts that we live in a post-classi cation, post-socially constructed world in which computers have free access to "raw data," which they re ne into actionable truth. Yet data are never raw, and consequently actionable truth is never totally objective.

*p. 10*
Algorithms do not create so much as retrieve what has been supplied and de ned-albeit repurposed and with varying levels of human intervention. is observation rebuts certain promises around AI and may sound like a disadvantage, but I believe that it could instead be an opportunity for social scientists to begin new collaborations with computer scientists. is could take the form of a professional social activity, people working together to describe the ground-truthing processes that underpin new algorithms, and so help make them more accountable and worthy.

*p. 10*
Florian Jaton is a senior researcher and lecturer in sociology of science and technology at the Geneva Graduate Institute of International and Development Studies, Switzerland. He is the author of e Constitution of Algorithms: Ground-Truthing, Programming, Formulating ( e MIT Press, 2021).

## XIAOCHANG LI

*p. 10*
Current technical approaches to preventing harm from arti cial intelligence and machine learning largely focus on bias in training data and careless (even malicious) misuse. To be sure, these are crucial steps, but they are not su cient solutions. Many risks from AI are not simply due to awed executions of an otherwise sound strategy: AI's artificial intelligence penchant for enabling bias and misinformation is built into its "data-driven" modeling paradigm.

*p. 11*
is paradigm forms the foundation of present-day machine learning. It relies on data-intensive pattern recognition techniques that generalize from past examples without direct reference to, or even knowledge about, what is being modeled. In other words, data-driven methods are designed to predict the probable output of processes that they can't describe or explain. at deliberate omission of explanatory models leaves these methods particularly receptive to misdirection.

*p. 11*
Today, this data-intensive, brute-force approach to machine learning has become largely synonymous with arti cial intelligence and computational modeling as a whole. Yet history shows that the rise of data-driven machine learning was neither natural nor inevitable. Even machine learning itself was not always so data-centric. Today's dominant paradigm of data-driven machine learning in key areas such as natural language processing represents what Alfred Spector, then Google's vice president for research, lauded in 2010 as "almost a 180-degree turn in the established approaches to speech recognition." rough its early decades, AI research in the United States xated on replicating human cognitive faculties, based on an assumption that, as historian Stephanie Dick puts it, "computers and minds were the same kind of thing." e devotion to this human analogy began to change in the 1970s with a highly unorthodox "statistical approach" to speech recognition at IBM. In a stark departure from the established "knowledge-based" approaches of the period, IBM researchers abandoned elaborate formal representations of linguistic knowledge and used statistical pattern recognition techniques to predict the most likely sequence of words, based on large quantities of sample data. ose very researchers described to me how this work owed much of its success to the unique computing resources available at IBM, where they had access to more computing power than anyone else. Even more importantly, they had access to more training data in a period where digitized text was vanishingly scarce by today's standards. During a federal antitrust case against the company from 1969 to 1982, IBM had manually digitized over 100,000 pages of witness testimony using a warehouse facility full of keypunch operators to manually encode text onto Hollerith punched cards. is material was repurposed into a training corpus of unprecedented size for the period, at around 100 million words.

*p. 11*
What resulted was an abandonment of knowledge-based approaches aimed at simulating human decision processes in favor of data-driven approaches aimed solely at predicting their output. is signaled a fundamental reimagining of the relation between human and machine intelligence. Director of IBM's Continuous Speech Recognition group Fred Jelinek described their approach in 1987 as "the natural way for the machine," quipping that "if a machine has to y, it does so as an airplane does-not by apping its wings." e success of this approach directly triggered a shi to data-driven approaches across natural language processing as well as machine vision, bioinformatics, and other domains. In 2009, top Google researchers pointed the earlier success of the statistical approach to speech recognition as proof that "invariably, simple models and a lot of data trump more elaborate models based on less data." machine intelligence as something fundamentally distinct from, if not antithetical to, human understanding set a powerful precedent for replacing expert knowledge with data-driven approximation in computational modeling. Generative AI takes this logic a crucial step further, using data not only to model the world, but to actively remake it.

*p. 11*
Large language models are both ignorant of and indi erent toward the substance of the statements they generate; they gauge only how likely it is for a sequence of text to appear. Which is to say, if the results pushed to our social media feeds are decided by algorithms that are intentionally designed to only predict patterns, but not to understand them, can the ourishing of misinformation really come as such a surprise? A failure to recognize how such problems may be intrinsic to the very logic of data-driven machine learning inspires omisguided technical xes, such as increased data collection and tracking, which can lead to harms such as predatory inclusion (in which outwardly democratizing schemes further exploit already marginalized groups). Such approaches are limited because they presume more machine learning to be the best recourse.

*p. 11*
But the lens of history helps us break out of this circular thinking. e perpetual expansion of data-driven machine learning should not be seen as a foregone conclusion. Its rise to prominence was embedded in certain assumptions and Many risks from AI are not simply due to awed executions of an otherwise sound strategy: AI's penchant for enabling bias and misinformation is built into its "data-driven" modeling paradigm.

*p. 12*
priorities that became entrenched in its technical framework and normalized over time. Instead of defaulting to tactics that augment machine learning, we need to consider that in some circumstances the very logic of machine learning might be fundamentally unsuitable to our aims.

## STEPHANIE DICK, WENDY HUI KYONG CHUN, AND MATT CANUTE

*p. 12*
We are watching "intelligence" being rede ned as the tasks that an arti cial intelligence can do. Time and again, generative AI is pitted against human counterparts, with textual and visual outputs measured against human abilities, standards, and exemplars. AI is asked to mimic, and then to better, human performance on law and graduate school admission tests, advanced placement exams and more-even as those tests are being abandoned because they perpetuate inequality and are inadequate to the task of truly measuring human capacity. e narratives trumpeting AI's progress obscure an underlying logic requiring that everything be translated into the technology's terms. If it is not addressed, that hegemonic logic will continue to narrow viewpoints, hamper human aspirations, and foreclose possible futures by condemning us to repeat-rather than learn from-past mistakes.

*p. 12*
e problem has deep roots. As AI evolved in the 1950s and '60s, researchers o en made human comparisons. Some suggested that computers would become "mentors" and "colleagues," others "assistants," "servants," or "slaves." As science and technology scholars Neda Atanasoski, Kalindi Vora, and Ron Eglash have shown, these comparisons shaped the perceived value not only of AI, but also of human labor. ose relegating AI to the latter categories usually did so because they believed computers would be limited to menial, repetitive, and mindless labor. ey were also reproducing the ction that human assistants are merely mechanical, menial, and mindless. On the other hand, those celebrating potential mentors and colleagues were tacitly assuming that human counterparts could be stripped of everything beyond e cient reasoning.

*p. 12*
Comparisons between AI and human performance o en correlate with social hierarchy. As society and technology scholars Janet Abbate, Mar Hicks, and Alison Adam have shown, in the 1960s and 1970s, women and minorities were encouraged to advance in society by learning to code-but those skills were then devalued, while domains dominated by white men were seen as the realm of the truly technically skilled. More recent OpenAI measures of AI against standardized exams endorse a positivist, adversarial, and bureaucratic understanding of human intelligence and potential. Similarly, AI-generated "case interviews" and artworks encode mimicry as the de nition of intelligence. For a result from generative AI to be validated as true-or to shock others as "true"-it has to be plausible, that is, recognizable in terms of past values or experiences. But looking backward and smoothing out outliers forecloses the rich wellsprings of humanity's imagination for the future.

*p. 12*
Such practices will ultimately a ect who and what is perceived as intelligent, and that will profoundly change society, discourse, politics, and power. For example, in "AI ethics," complex concepts such as "fairness" and "equality" are recon gured as mathematical constraints on predictions, collapsed onto the underlying logic of machine learning. In another example, the development of machine learning systems for game-playing has led to a reductive rede nition of "play" as simply making permissible moves in search of victory. Anyone who has played Go or chess or poker against another person knowns that, for humans, "play" includes so much more.

*p. 12*
e portrayal of AI's history is usually one of progress, where constellations of algorithms attain humanlike general intelligence and creativity. But that narrative might be more accurately inverted with a shrinking de nition of intelligence artificial intelligence that excludes many human capabilities. is narrows the horizon of intelligence to tasks that can be accomplished with pattern recognition, prediction from data, and the like. We fear this could set limits for human aspirations and for core ideals like knowledge, creativity, imagination, and democracymaking for a poorer, more constrained human future.

## MIKE ANANNY

*p. 13*
O en, problems that seem narrow and purely technical are best tackled if they're recast as "public problems," a concept put forth almost a century ago by philosopher and educator John Dewey. Examples of public problems include dirty air, polluted water, global warming, and childhood education. Public problems bring harms that are not always felt individually but that nonetheless shape what it means to be a thriving person in a thriving society. ese problems need to be noticed, discussed, and collectively managed. In contrast to problems that are personal, private, or technical, Dewey wrote, public problems happen when people experience "indirect consequences" that need to be collectively and "systematically cared for," regardless of an individual's circumstance, wealth, privilege, or interests. Public problems de ne our shared realities.

*p. 13*
Although generative AI has been framed as a technical problem, recasting it as a public problem o ers new avenues for action. Generative AI is quickly becoming a language for telling society's collective stories and teaching us about each other. If you ask generative AI to make a story or video that explains climate change, you are actually asking a probabilistic machine learning model to create a statistically acceptable account of a public problem. Tools such as ChatGPT and Midjourney are fast becoming languages for understanding public problems, but with little analysis of their power to shape the stories that humans use to understand the shared consequences that Dewey told us create public life.

*p. 13*
To grapple with generative AI e ectively, consumers and developers alike need to see it not only as biased datasets and machine learning run amok-we need to see it as a language that people are using to learn, make sense of their worlds, and communicate with others. In other words, it needs to be seen as a public problem.

*p. 13*
First, researchers need to see generative AI as a powerful language-as "boundaries," "infrastructures," and "hinges" that scholars of science and technology tell us create technologies. is means tracing the connections among the people and machines that make synthetic language: engineers who build machine learning systems, for example, entrepreneurs who pitch business models, journalists who make synthetic news stories, audiences who struggle to know what to believe. ese are the complex and largely invisible assumptions that make generative AI a language for representing knowledge, fueling innovation, telling stories, and creating shared realities.

*p. 13*
Second, as a society, we need to analyze the harms created by generative AI. When statistical hallucinations invent facts, chatbots misattribute authorship, or computational summaries bungle analyses, they produce dangerously wrong language that has all the con dence of a seemingly neutral, computational certainty. ese errors are not just rare and idiosyncratic curiosities of misinformation; their real and imagined existence makes people see media as unstable, unreliable, and untrusted. Society's information sources-and ability to gauge reality-are destabilized.

*p. 13*
Finally, all members of society should reject the assertions of technology companies and AI "godfathers" who claim that generative AI is both an existential threat and a problem that only technologists can manage. Public problems are collectively debated, accounted for, and managed; they are not the purview of private companies or self-identi ed caretakers who work on their own timelines with proprietary knowledge. Truly public problems are never outsourced to private interests or charismatic authorities.

*p. 13*
A public problem is not merely a technical curiosity, a moral panic, or an inevitable future. It is a system of relationships between people and machines that creates language, makes mistakes, and needs to be systematically cared for. Once we understand generative AI as a vital language for creating shared realities and tackling collective challenges, we can start to see it as a public problem, and then we will be in a better place to solve it. These articles arose from a working group on arti cial intelligence and justice convened by Kate Crawford at the École Normale Supérieure.

## Footnotes

> AMY KARLE, BioAI-Formed Mycelium, 2023
