# Researcher Access to Social Media Data: Lessons From Clinical Trial Data Sharing

**Authors:** Christopher J. Morten, Gabriel Nicholas, Salomé Viljoen
**Citation:** "Researcher Access to Social Media Data: Lessons From Clinical Trial Data Sharing," 39 *Berkeley Technology Law Journal* 109 (2024) (with Gabriel Nicholas and Salomé Viljoen)
**Source:** https://btlj.org/wp-content/uploads/2024/04/0003_39-1_Morten.pdf

## INTRODUCTION

*p. 4*
In 2018, researchers at Harvard University announced that they had entered into a landmark voluntary partnership with Facebook called Social Science One (SS1) to gather and share data on the inner workings of the social media goliath. The announcement was met with great fanfare. Researchers had been clamoring for data access in order to better understand the dynamics of social media and its effects on everything from elections to teenage mental health to free speech online. Today, however, this grand experiment in voluntary social media data sharing is remembered as a fiasco. Facebook delivered only a fraction of the data it had promised; technical "fixes" made by the company to protect user privacy rendered certain data useless for research; and funders, academics, and civil society partners all eventually withdrew from the project. 1 Two years after SS1, researchers at New York University's (NYU) Ad Observatory announced that they were taking a different approach to studying Facebook: conducting large-scale research, with or without the company's consent. The Ad Observatory focused on understanding political advertising on Facebook and tracked electoral races across the country. Ad Observatory researchers developed a browser extension, externally audited for security and privacy, that scraped ad data from Facebook and contributed it to an NYUrun database. Months later, Facebook suspended the Ad Observatory researchers' access to Facebook. Facebook's stated justification was to "protect people's privacy." 2 These two abbreviated anecdotes illuminate a few things about the current state of researchers' access to social media. First, they highlight that significant numbers of researchers in academia and civil society actively want to research social media and will go to great lengths to do so. Second, they show that independent researchers lack sufficient access to various forms of social media data, including content data about what users see, moderation data about how platforms such as Facebook promote and censor content, and distribution data about what kinds of users see what kinds of content. Third, they show that when platforms themselves wield absolute control over which researchers get access to data (and how much, and on what terms), platforms can thwart critical research and shape the literature that emerges by selectively providing access to data.

*p. 4*
As we explain in this Article, we need research on social media to flourish if we, as a social-media-obsessed world, are to flourish. For example, Yet researchers' access to data remains controversial. Independent privacy advocates raise concerns over the sensitivity of social media data held by companies and the potential threats of researchers using such data irresponsibly. 7 Social media companies themselves increasingly deploy (or perhaps "weaponize") arguments about individual privacy to justify intense secrecy. 8 These companies wield privacy arguments at both the doctrinal and theoretical levels, arguing that researcher access (1) would violate various extant laws, such as the European Union's General Data Protection Regulation (GDPR), and (2) is normatively undesirable because it would expose individuals who use social media to a raft of harms that outweigh the research's foreseeable benefits. 9 In addition, the same social media companies raise separate but equally serious concerns over intellectual property. Again, these companies raise commercial secrecy objections at both the doctrinal level and the theoretical level, asserting that researcher access would (1) violate state and federal trade secrecy law, and (2) be normatively undesirable because it would encourage "free riding" by competitors and thereby erode crucial "incentives to innovate." 10 In industry's telling, and in much popular discourse, privacy and incentives to innovate have become a kind of "Scylla and Charybdis" of sharing social media data-two obstacles that any data-sharing effort must navigate to MISINFORMATION REV. 1, 8 (2020) (on social media platforms' role in propagation of misinformation).

*p. 6*
7. SHAPIRO ET AL., supra note 3, at 45 ("The rules and regulations around user privacy, combined with the political force of privacy advocates, are by far the biggest barrier to platform companies' ability and willingness to share data with researchers."); VOGUS, supra note 3, at 33 ("Properly balancing competing interests, such as the risks to user privacy, may require policymakers to take incremental steps to improve researchers' access to data, and to carefully assess whether those steps are serving the public interest.").

*p. 6*
8. For a broad analysis, see Rory Van Loo, Privacy Pretexts, 108 CORNELL L. REV. 1 (2022). For a specific example, see generally AMY O'HARA & JODI NELSON, EVALUATION OF THE SOCIAL SCIENCE ONE-SOCIAL SCIENCE RESEARCH COUNCIL-FACEBOOK PARTNERSHIP (2020) (explaining how Facebook concluded it could not provide previouslypromised data access to researchers because of concerns over user privacy). 9. Id.; see also Matias Vermeulen, The Keys to the Kingdom (July 27, 2021), https:// knightcolumbia.org/content/the-keys-to-the-kingdom (analyzing whether GDPR creates barriers to researcher access).

*p. 6*
10. FACEBOOK, COMMENTS TO THE FEDERAL TRADE COMMISSION ON DATA PORTABILITY 13 (2020) (arguing that portability of and access to "all observed and inferred data could also result in a different sort of burden: the disclosure of trade secret or other proprietary information developed by a business to enhance or differentiate its services. Enabling people to port that kind of information could reduce incentives for businesses to develop it in the first place"); VOGUS, supra note 3, at 25 ("[A]ccess to non-public data raises greater risks of invading users' privacy and revealing trade secrets or security measures used by hosts.").

*p. 7*
succeed. 11 Social media companies cast this two-headed trap as so fearsome that it may ultimately doom even the cleverest efforts. Some regulators and legislators have nonetheless persisted in proposing and enacting new laws to expand researcher access to social media data, 12 but they face stiff headwinds. Concerns over privacy and incentives to innovate have chilled nascent efforts toward real transparency and accountability of social media. 13 The key question that this Article addresses is this: Does a regulatory pathway exist to achieve meaningful researcher access to social media data while protecting privacy and incentives to innovate?

*p. 7*
This is an urgent question, and we are far from the first to write on it. Daphne Keller; 14 Aline Iramina, Maayan Perel & Niva Elkin-Koren; 15 Rebekah Tromble; 16 and the Working Group established by the European Digital Media Observatory 17 are among those who have offered important views on this question. The European Union is already moving to mandate researcher access to social media platform data. 18 Its Digital Services Act, among other initiatives, requires qualifying platforms to grant access to certain 11. E.g., Paddy Leerssen, Platform Research Access in Article 31 of the Digital Services Act, VERFASSUNGSBLOG (Sept. 7, 2021), https://verfassungsblog.de/power-dsa-dma-14/. The twin obstacles of privacy and incentives to innovate are discussed in greater detail in infra Part II.

*p. 7*
12. See generally VOGUS, supra note 3 (discussing U.S. legislative proposals to guarantee researcher access to social media data); Alex Engler, Platform Data Access Is a Lynchpin of the EU's Digital Services Act, BROOKINGS INST. (Jan. 15, 2021), https://www.brookings.edu/blog/ techtank/2021/01/15/platform-data-access-is-a-lynchpin-of-the-eus-digital-services-act/ (presenting researcher access provisions of EU's Digital Services Act).

*p. 7*
13. See SHAPIRO ET AL., supra note 3; VOGUS, supra note 3. 14. Daphne Keller, Delegated Regulation on data access provided for in the Digital Services Act-requested information to vetted researchers, although the processes for doing so have not yet been finalized. 19 We think the answer is yes-a regulatory pathway does exist to achieve meaningful researcher access to social media data while protecting privacy and incentives to innovate. While the Digital Services Act's vetted researcher access mandate is a valuable source of insight and inspiration, we choose to make a complementary case focusing and drawing on U.S. law to argue that researcher access can be achieved here in the United States-because, indeed, in other technology industries, it has already. We do not have to look only to "proregulatory" Europe for comparative lessons on the potential virtues of regulation: our own regulatory history and landscape offers such lessons, too. 20 The main contribution of this Article is comparative. It imports hard-won lessons from other fields of technology-pharmaceuticals 21 and medical devices-to enrich the current debate over researcher access to social media data. 22 The complexity of these technologies rivals that of social media-as does the power of their industries and lobbies, especially in the United States. And yet in pharma and medical devices, we have successfully established mechanisms for broad sharing of what would otherwise be secret industry data. 23 Along the way, these fields successfully navigated a similarly narrow strait between potential harms to individual privacy and harms to incentives to innovate.

*p. 8*
19. Regulation on a Single Market for Digital Services (Digital Services Act), 2022 O.J. (L 277) 1, 27 ("This Regulation therefore provides a framework for compelling access to data from very large online platforms and very large online search engines to vetted researchers affiliated to a research organisation within the meaning of Article 2 of Directive (EU) 2019/ 790, which may include, for the purpose of this Regulation, civil society organisations that are conducting scientific research with the primary goal of supporting their public interest mission."). For an explainer of researcher access and the processes ahead, see John Albert, A Guide to the EU's New Rules for Researcher Access to Platform Data, ALGORITHM WATCH (Dec. 7, 2022), https://algorithmwatch.org/en/dsa-data-access-explained/.

*p. 8*
20. This point is not meant to undercut the significance of the Digital Services Act for non-EU researchers who will likely, under the delegated acts, gain access to hitherto unavailable social media platform data.

*p. 8*
21. Throughout this Article, for concision, we generally use the terms "pharmaceutical" and "drug" broadly to describe both small-molecule drug products and biologic drug products. This broad usage is admittedly inexact but consistent with the common practice of the Food & Drug Administration (FDA) and others. See, e.g., Drugs@FDA Glossary, FDA, https:// www.accessdata.fda.gov/scripts/cder/daf/index.cfm?event=glossary.page (last visited Dec. 27, 2023) (defining "Drug" to include biological products).

*p. 8*
22. Small portions of a preliminary version of the ideas in this Article were published in a 2022 white paper, GABRIEL NICHOLAS & DHANARAJ THAKUR, LEARNING TO SHARE: LESSONS ON DATA-SHARING FROM BEYOND SOCIAL MEDIA (2022).

*p. 9*
In this Article, we focus on one specific kind of data generated by pharmaceutical and medical device companies: clinical trial data. Clinical trials are research studies on human volunteers that answer questions about the safety and efficacy of different health interventions, such as drugs, vaccines, and devices. They are the "gold standard" of evidence-based medicine. They are expensive to conduct, and their data is enormously valuable to doctors' care for patients, regulatory approval, businesses' decision-making and marketing, and scientific research.

*p. 9*
Until the 1990s and 2000s, the pharmaceutical and medical device industries could and did keep clinical trial results proprietary. The result was a comparative dark age of information, with drug companies "cherry-picking" only their most favorable data for publication in the medical literature, and falsely marketing unsafe and ineffective products as wonder drugs. A series of high-profile scandals ensued, which involved companies that hid unfavorable data from independent researchers and the broader public, leading to widespread patient harm. These scandals ultimately provoked landmark federal legislation in 2007 that, for the first time, mandated that industry share certain clinical trial data at an across-the-board baseline level. Today, independent researchers around the world use this data to double check the industries' claims and the work of the industries' central regulator, the Food & Drug Administration (FDA), identify unsafe and ineffective products, and advance science.

*p. 9*
Before the 2007 clinical trial data-sharing mandate, the pharmaceutical and medical device industries fought it by advancing privacy and incentives-toinnovate arguments similar to those that social media companies deploy today. 24 For example, the largest pharmaceutical lobby warned that mandatory clinical trial data sharing would "fail to protect adequately trade secrets and confidential commercial information," and therefore "harm the public health by discouraging the very innovation necessary to bring new medical advances to the market." 25 And like most social media data, much clinical trial data implicates acute privacy concerns, as individuals' detailed medical statuses are encoded in the data, including many statuses that expose people to discrimination and exploitation. 26 24. Supra Section III.C. 25. Letter from William W. Chin, Executive Vice President, and Jeffrey K. Francer, Vice President & Senior Counsel Scientific & Regulatory Affairs, PhRMA, to Jerry Moore, NIH Regulations Officer, National Institutes of Health (Mar. 25, 2015) (on file with the National Institutes of Health).

*p. 10*
Yet in the years since Congress legislated the clinical trial data-sharing mandate, no real harm to privacy or to incentives has occurred, even as independent research on that data has unlocked new uses and social benefits. If anything, the trend in clinical trial data sharing today is to push further, expanding researcher access to the most sensitive kinds of data, especially individual patient-level data (IPD) and methodological protocols that reveal exactly how companies conduct their trials and generate and interpret their own data. 27 As we show below, there are important proof-of-concept datasharing initiatives led by academic centers and by administrative agencies in the United States and Canada that demonstrate even the most highly sensitive data can, under the right conditions, be shared responsibly with researchers.

*p. 10*
We recognize that the parallels between social media data and clinical trial data are inexact. Clinical trial data sets are more standardized and far smaller than that of social media platforms. The data subjects in clinical trials are volunteers, enrolled pursuant to elaborate and independently vetted processes of informed consent, while the quality of informed consent for data collection from users of social media is widely perceived as laughable. 28 Some individuals' social media data is intensively sensitive in ways that even the most detailed medical data is not; social media data may reveal, for example, users' political affiliations and organizing activities, romantic preferences, travel histories, and more. The variety and profundity of harms that flow from discriminatory and other unwanted uses of social media data can therefore be even greater than the harms that flow from unwanted uses of medical data. Furthermore, social media and medical products implicate very different tradeoffs. Medical products are generally seen as innovations vital for society; social media innovations, such as algorithms targeting ads or recommending content, for example, are increasingly seen as socially deleterious. 29 Clinical trial and social media data access systems both need to manage tradeoffs between protection of trade secrecy and utility to researchers, but where they draw those lines will be very different.

*p. 10*
Yet as we endeavor to show in this Article, the benefits of sharing are likely to be broadly similar. Indeed, we argue that important parallels do exist and that the history of clinical trial data sharing therefore holds important lessons for social media data sharing. 30 We focus on clinical trial data not because this 27. Supra Section III.D. 28. See Ari Ezra Waldman, Privacy, Notice, and Design, 21 STAN. TECH. L. REV. 74 (2018). 29. See generally MARIANA MAZZUCATO, THE VALUE OF EVERYTHING: MAKING AND TAKING IN THE GLOBAL ECONOMY (2018).

*p. 10*
30. Social media companies sometimes insist that their technologies are unprecedented and sui generis, and thus cannot be regulated like technologies past; a rich literature shows that's false. See, e.g., MARIANA MAZZUCATO, THE ENTREPRENEURIAL STATE (1st ed. 2013) data is, as a technical matter, most similar to social media data, but because the technical, institutional, and legal structures that govern clinical trial data sharing are particularly mature, tested, and successful, as we show below. In future work, we and other scholars may draw other instructive lessons from efforts to share other kinds of medical data, such as electronic medical record data. 31 In this Article, we offer three primary lessons for those studying, advocating, and legislating social media data sharing: first, the benefits of research on otherwise secret data are cascading and unpredictable; second, law without institutions to implement the law is insufficient; and third, different kinds of data must be treated differently. 32 The history of clinical trial data sharing shows that effective researcher access and use of industry data is impossible without powerful independent institutions that can serve as counterweights to extraordinarily powerful industries. Such counterweight institutions, whether public agencies, private independent institutions, or both, could serve as "regulators" of the social media industry. To support research, these regulators may serve many roles:

*p. 11*
(technology and pharmaceutical companies arguing they deserve regulatory exceptions); Rebecca Haw Allensworth, Antitrust's High-Tech Exceptionalism, 130 YALE L.J. F. 588 (2021) (detailing how courts granted tech companies special exceptions to antitrust rules due to "views about digital markets in the early 2000s-that they were uniquely dynamic, innovative, and competitive" that are not only false, but have also prevented competition in the tech sector); Yaël Eisenstat & Nils Gilman, The Myth of Tech Exceptionalism, NOEMA MAGAZINE (Feb. 10, 2022), https://www.noemamag.com/the-myth-of-tech-exceptionalism/ (detailing how big tech companies use the narrative of innovation to ward off regulation); Richard Waters, Tech's Self-Declared Exceptionalism is Coming to an End, FIN. they monitor and enforce industries' compliance with data sharing laws; collect, standardize, curate, steward, and share data; govern researchers' access and use of data; explain to researchers and the broader public how to use data; and sometimes fund worthy research. These institutions need not be public (though most are in the world of clinical trial data sharing); they can be academic or non-governmental organizations. But they do need to be functionally independent from industry; pharmaceutical industry-funded clinical trial data sharing initiatives failed to spark useful research and to check the industry's worst excesses.

*p. 12*
The history of clinical trial data sharing also shows that different kinds of data should be treated differently. Perhaps the point is self-evident, but it is also vital. Today federal legislation mandates sharing of certain clinical trial data-so-called "summary data" characterizing broad trends, as well as certain "metadata" on how data is generated-on a public website accessible from anywhere in the world. This kind of blunt mandatory disclosure works well for data of high value to researchers and for which sharing poses low risk. For more sensitive data-individual participant data (IPD), which can easily be reidentified, or certain trial protocols that reveal industries' innovative and confidential scientific methods-blunt disclosure to the general public is inappropriate. Instead, more sensitive data tends to be shared only with trusted researchers subject to a raft of constraints on access and use.

*p. 12*
Before we turn to the body of the Article, a word on the Article's limitations-on what this Article is and is not. First, we intend this Article as a primarily descriptive, positivist account of how law and technology currently work. Much of the description and analysis of clinical trial data sharing (and sharing of other kinds of medical data) is in the medical and scientific literature rather than the law review literature, and thus has received comparatively little attention from legal scholars, activists, and other researchers focused on social media. 33 We see value in building a bridge between distinct literatures and distinct readerships.

*p. 12*
Second, we recognize and decline to address, in this Article, a large set of important theoretical and doctrinal questions attached to the value of social media data sharing. For example, what is the fundamental value of social media? Is the collection of social data ethical and desirable in the first place? What theory (or theories) of privacy should inform laws governing social media? Under existing doctrine, does any form of social media data qualify for trade secrecy protection, or other forms of intellectual property protection? Should it, from a public policy perspective? The three of us have grappled with some of these questions in other work, 34 and will continue to, but we put these questions aside for this Article.

*p. 13*
Third, this Article largely accepts the social media industry's professed concerns over privacy and incentives to innovate. There are, of course, compelling reasons to be skeptical. 35 But here we endeavor to show that it is possible to take the social media industry's concerns seriously and overcome them. This Article argues that legislators and regulators concerned with protecting privacy and intellectual property rights in sensitive privately held data can nonetheless devise rules and institutions to share that data with independent researchers responsibly. This is, at the very least, precisely what has happened with the pharmaceutical and medical device industries.

*p. 13*
The Article proceeds as follows. Part II provides a legal and technical description of the current state of researchers' access to social media data and presents a novel taxonomy of its problems. It also describes the law and normative arguments that created and perpetuate today's status quo, with focus on trade secrecy and privacy in the United States. Part III lays out relevant lessons from clinical trial data, explaining what clinical trial data is, how it compares to social media data, and how regulatory and voluntary efforts managed to responsibly share even the most sensitive personal and trade secret data with independent researchers. Part III also gives the history of these efforts, describing first the "dark ages" of clinical trial data secrecy, when the pharmaceutical companies that created and exploited this data wielded neartotal control over access to it, and then how the industry emerged from these dark ages after Congress passed data sharing requirements and invested in countervailing public and nonprofit institutions. Part IV applies the clinical trial data sharing framework's legal and institutional lessons to social media data and charts a strategic course forward toward responsible and effective social media data sharing. As noted above, one key lesson is the need to empower public or nonprofit institutions capable of confronting the powerful social media industry. Another is the value of treating different kinds of data differently. In particular, clinical trial data's tripartite distinction of individual data, summary data, and metadata promotes distinct governance structures 34. that maximize researcher utility while minimizing risks to data subjects and incentives to innovate. Part V briefly concludes with a discussion of proposed legislation.

## II. THE STATE OF SOCIAL MEDIA DATA SHARING

*p. 14*
Social media companies have a wide range of approaches they can take to sharing data with researchers. This Part offers a snapshot of the status quo of how sharing occurs currently and the legal and technical arrangements that support that sharing. It also lays out the primary legal challenges to addressing the problem of researcher access to data that animate the rest of the Article.

## A.

*p. 14*
HOW RESEARCHERS USE SOCIAL MEDIA DATA Researchers are interested in all sorts of social media data for all sorts of reasons. Many seek to better understand the dynamics and external effects of social media ecosystems. Social and computer science researchers use platform data to better understand widespread popular problems such as the spread of mis-and dis-information, 36 the effects of algorithmic speech systems, 37 online extremism, 38 child welfare, 39 free speech online, 40 and online discourse around elections and other democratic processes. 41 Some smaller scale work may not require researchers to have access to more or different data than is available to ordinary users. For instance, sociological research that focuses on small online communities can be done without special access to data, so long as researchers can embed themselves within those communities. 42 Larger scale and more macro-level research, however, requires access to more data than any one regular user has access to through non-automated means. For example, researchers looking to understand public views of gender-based violence on X, née Twitter (referred to from here as "Twitter"), need access to hundreds of thousands or millions of posts to be able to discern recurring behaviors and rhetorical patterns. 43 Researchers who attempt to reverse engineer or uncover patterns in recommendation algorithms require particularly large volumes of detailed data to produce significant results, since any one user's recommendations only reflects their own tastes, not the system as a whole. 44 Researchers are also interested in accessing social media data in order to confirm or refute otherwise unverifiable claims made by companies, particularly about changes in their practices. The Markup used data collected from its Citizen Browser to reveal that Facebook had not stopped recommending anti-vaccine groups as it claimed it had. 45 In April 2022, researchers used data collected from Russian TikTok to show that TikTok had not had as complete of a ban of Russian pro-war propaganda as it had claimed. 46 Researchers have also used data to show when social media services have made good on their promises to improve. For example, researchers used data scraped from YouTube to confirm that it had reduced the prevalence of conspiratorial content in its recommendation algorithms. 47 Giving researchers access to social media data can confirm theoretical problems on social media or uncover new problems not previously known to exist. The now-famous "filter bubble" phenomenon, for example, was able to be confirmed by researchers with access to data donated by social media users. 48 Work from Jonas Kaiser and Adrian Rauchfleisch studying YouTube's recommendation algorithm in Brazil found that users could go down rabbit holes of videos of sexually suggestive videos of children. 49 The Stanford Internet Observatory used data from Mastodon to discover a large decentralized distribution network of human-and computer-generated child sexual abuse material. 50 Some areas of social media research require access beyond what is available on the internet publicly. For example, most research related to personalization requires information on real people's profiles, activities, and recommendations, which, if not public, can only be obtained through donation by the users or the platform itself. Though more challenging from a privacy perspective, this research is still critically important. For instance, research using data donated from Facebook users found that the platform drastically overcounted some and undercounted other political ads, including tens of thousands of ads that ran during its "moratorium" on political ads around the U.S. 2020 elections, raising questions about the company's ability to effectively enforce its own policies. 51 Researchers that study topics beyond social media may also be interested in data from platforms. Linguists, for example, use social media to understand emerging subject areas such as how emojis are used and how people from different generations speak online. 52 Machine learning researchers use labeled image data and unlabeled text data from social media to train generative AI models. 53 Hundreds of scientific articles have sought to use social media posts to detect mental illness. 54 And at least while it was publicly available, the U.S. Geological Survey used Twitter data to track earthquakes, which in some cases has been shown to work even better than a Richter scale. 55 It is easy to imagine many other use cases of social data: ornithologists accessing photos of birds on Instagram, social scientists accessing relational information to predict gun violence, and so on.

*p. 18*
Social media companies themselves of course stand to gain a lot of value from the data generated by their own services, and many have business models that entirely depend on such data. 56 Companies can use their data to target advertisements, increase the amount of time users spend on a service, or sell it to data brokers and other actors that can monetize the data. For instance, when Reddit began charging for its API in 2023, the company claimed it was because Google and OpenAI were using their data to train large language models, although critics argued it was also for them to wrestle control over their advertising revenue from third-party apps. 57 Companies can also use data from their platforms to better understand how users use their services, and use that information to improve the user experience or the safety and integrity of their communities. Many legal scholars have written about the market benefits of requiring social media companies to make certain data available to competitors, 58 but those efforts have different normative values from providing researchers with data-facilitation of markets as opposed to the generation of knowledge-and entail very different governance decisions outside the scope of this Article.

## B.

*p. 18*
CURRENT RESEARCHER ACCESS TO SOCIAL MEDIA DATA Social media companies vary widely in what data they share with researchers and how they make it available. Many platforms make little to no data available to researchers, including private messaging apps such as WhatsApp, Telegram, and iMessage; team chat apps such as Slack and Discord; semi-private social networks such as Snapchat; and public social networks such as LinkedIn and Pinterest. There are also large public social networks such as YouTube and TikTok that, as of this writing, make some data available to researchers-and recently increased that amount due to recent regulatory efforts-but still not enough or under too restrictive agreements to be adopted by researchers en masse. 59 Other large public-facing platforms offer data access but only under certain conditions. Most platforms at least have their data protected under terms of service, but some have additional restrictions they impose upon researchers in exchange for access to more data. Meta, for instance, allows approved academics and independent researchers to access data sets about election ads and URL shares on Facebook. 60 However, those researchers are required to sign a data agreement that, among other things, limits their ability to share data with third party reviewers, prevents them from using Facebook data in conjunction with other data, and allows Meta to review any published material ahead of time for "any Confidential Information or any Personal Data that may be included or revealed in those materials and which need to be removed prior to publication or disclosure." 61 There are two primary ways companies make data available to researchers: static public datasets and application programming interfaces (APIs). Static data datasets allow companies to share a snapshot of the data on their platform, but since they are not dynamic, they can go out of date. APIs, on the other hand, allow live, up-to-date access to data hosted on a platform. They are more expensive for companies to build, maintain, and operate, but unlike static data sets, they allow companies to retain extensive control over who can access what data and how much.

*p. 19*
Social media companies have not shied away from severely limiting data access through APIs, even to researchers. Before Twitter raised the cost of its API from free to $42,000 per month (a move many see as Elon Musk thumbing his nose at researchers), 62 Twitter offered researchers with university affiliations an Academic API, which allowed them to access Twitter's full archive of historical tweets and perform more refined searches. 63 However, it limited researchers to accessing ten million tweets per month, or the equivalent of about one fiftieth of all tweets sent per day. 64 YouTube's API is far more limited: by default, it allows researchers to make 100 search requests or 10,000 video information requests per day. 65 While some of these numbers sound large, they constitute a very small fraction of the activity that happens on these platforms. 66 Researchers complain that these limitations significantly stifle or prevent research. 67 With APIs, platforms can also change the data they make available or revoke data access to individuals as they see fit. Facebook and Twitter, for example, both drastically reduced what and how much data users, including researchers, could access shortly after news of the Cambridge Analytica scandal broke. 68 specific companies to access data may be particularly vulnerable to losing data access without warning. 69 Social media companies can also censure specific researchers for using data in ways they deem improper, as will be discussed further in Section II.B.2 with the case of NYU Ad Observatory.

*p. 21*
Finally, social media companies can withdraw support for their data sharing tools or remove them entirely. Twitter and Reddit have both recently been in the news for starting to charge extremely high prices for their oncefree APIs. 70 More quietly, Meta appears to be slowly sunsetting CrowdTangle, a popular social media monitoring tool acquired by Facebook in 2016. 71 CrowdTangle is a particularly popular tool with researchers for studying COVID misinformation, 72 election misinformation, 73 and online hate 74 in a wide range of languages. Recently however, Meta has reduced support for the product, allowing it to become buggy and less usable, and has plans to shut it down entirely. 75 Critics argue that Meta is deprecating CrowdTangle because it has contributed to negative press about the company. 76 Platformed-sanctioned methods, however, are not the only ways for researchers to be able to access social media data. Researchers can appeal directly to users themselves to give permission to read their data, usually either through authenticating a third-party application (aka a "Sign in with __" button) or through installing a browser extension that scrapes websites on their behalf. These platform-unsanctioned methods can pose additional risks for users because bad actors can use elevated permissions to exfiltrate data. Researchers who build these tools are also at risk of violating a platform's Terms of Service, if not the Computer Fraud and Abuse Act. 77 However, unsanctioned methods allow for research that could not be otherwise possible under platform sanctioned methods, including research a platform may try to preclude since it could reflect unfavorably on the platform. 78 C.

*p. 22*
WHAT HAPPENS WHEN RESEARCHERS TRY TO USE THIS ARCHITECTURE?

*p. 22*
Two public controversies showcase the deficiencies and barriers of the current state of social media data access: Social Science One and the New York University Ad Observatory. 79 In the first case, researchers tried to work within the platform's data sharing architecture but ran into shortcomings and had no way to negotiate the additional access they needed, despite being well connected and resourced. In the second, researchers tried to work outside the platform's data sharing architecture, but the platform rejected them, despite their research being safe, secure, socially beneficial, and impossible to do within the company's platform-sanctioned methods.

## Social Science One (SS1)

*p. 22*
On March 17, 2018, The New York Times and The Observer revealed that the conservative political consulting firm Cambridge Analytica had harvested private information from more than fifty million Facebook profiles and used that data to influence elections around the world. 80 the center of controversy for its role in the 2016 United States presidential election, Brexit, and the spreading of Russian-influenced propaganda, but Cambridge Analytica turned a gradual public relations crisis into an acute one.

*p. 23*
Facebook higher ups soon after began to look for new ways to support independent research to help avoid future election interference, and honed in on one method proposed by Harvard social scientists Gary King and Nate Persily. 81 King and Persily argued that researchers inside social media companies had access to data but no credibility or independence, while researchers outside the companies had the inverse. To resolve this, they proposed giving some academics access to a company's data but having them sign NDAs and preventing them from publishing. Those academics on the inside could then help decide what data is important and how to share it with third-party researchers in a privacy-preserving way. 82 Facebook quickly put the proposal into practice. About three weeks after the Cambridge Analytica leak (and one day before Zuckerberg was slated to testify before the Senate), Facebook announced a new initiative to allow academics independent access to Facebook data. 83 King and Persily established SS1 as the organization that would operate within Facebook, and they brought on the Social Science Research Council (SSRC) to manage external researchers, who would apply for access to the data they made available. King and Persily raised ten million dollars for the initiative from an ideologically diverse group of seven foundations. 84 In July 2018, SS1 announced the data set Facebook would release: every URL that had ever been shared publicly on Facebook between January 1, 2017 and June 11, 2018, along with information about who shared it, how often it was shared, and how many people saw it. 85 SSRC and SS1 put out a request for proposals for research projects and granted $50,000 to each project along with access to the URL share dataset. 86 However, the endeavor faced legal and political headwinds. Europe's General Data Protection Regulation (GDPR) took effect in May 2018 and California passed the California Consumer Privacy Act a month later, introducing new legal complexities. The Electronic Privacy Information Center (EPIC) also sent an open letter to SS1 claiming that the project complied with neither GDPR's personal data protection requirements nor Facebook's 2011 consent decree from the Federal Trade Commission to obtain user consent before sharing data. 87 The project also faced technical headwinds. Facebook needed to comb through a huge volume of data to create the URL shares dataset. Facebook had over two billion active users, the URL shares dataset was initially calculated to include sixty billion public posts, and preparing just the shares and interaction metrics required processing more than fifty terabytes per day. 88 Many researchers would likely not have the computing resources to ingest this much data. Simultaneously, Facebook tried to respond to privacy concerns by implementing differential privacy, a statistical method that adds noise to a dataset to make individuals less identifiable, while still maintaining certain core patterns in the data. 89 Facebook engineers underwent lots of trial and error to make it work at such a scale. 90 The technical and legal challenges plagued the project with delays and eventually led to its collapse. SS1 and SSRC believed that Facebook would be able to provide the URL shares data by fall 2018, but they gave no information until January 2019, when they admitted to further delay. SSRC announced the first research grant winners in April 2019, which included more than sixty researchers from thirty academic institutions in eleven countries, but Facebook still had no URL shares data. 91 When Facebook did finally share data, it was a "light" version of the dataset, which excluded demographic and exposure data. This meant researchers could not study who and how many people different posts reached, likely hampering research on such topics as mis-and disinformation. At the end of SS1's year-long funding period, all seven funders sent a joint letter to SSRC announcing that they would discontinue funding. As they explained:

*p. 25*
It now seems clear that the technical and legal complexities associated with making proprietary data available to independent scholars are greater than any of the parties originally understood, and Facebook has as a result been unable to deliver all the data initially anticipated. 92 Facebook continued the project on its own, and the full URL shares dataset was finally made available to researchers in February 2020. However, statistical analysis from King and others suggest that the differential privacy methods Facebook used added significant statistical bias. 93 In 2021, Facebook also revealed that the data accidentally excluded URLs shared by any U.S. user without detectable political leanings, about half of all US Facebook users. 94 A 2019 post-mortem released by the Hewlett Foundation offered multiple interpretations of the events of SS1. One is that funders, SS1, and SSRC put the cart before the horse: "investing in research was premature given the uncertainty of data access." 95 Another is that SS1 was unable to motivate Facebook to share data. 96 Both may be right.

*p. 26*
All in all, SS1 is widely seen as a failure, or as Persily put it to the press, "I'm happy to be quoted saying this: This was the most frustrating thing I've been involved in, in my life." 97 Persily later stated that the demise of SS1 "demonstrates why we need government regulation to force social media companies to develop secure data sharing programs with outside independent researchers." 98

## NYU Ad Observatory

*p. 26*
Laura Edelson and Damon McCoy of the NYU Cybersecurity for Democracy group started the NYU Ad Observatory on September 15, 2020. 99 The Observatory was meant to increase political ad transparency on social media ahead of the 2020 elections, and let researchers independently search for and analyze political ads by state, races, targeting criteria, funding sources, money spent, and messaging. The Observatory quickly saw adoption, particularly from journalists reporting on federal and local elections, including in Florida, Kentucky, Missouri, and Utah. 100 Data for the NYU Ad Observatory came from a mix of platform sanctioned and unsanctioned sources. It used reports provided by Facebook such as the Facebook API, CrowdTangle, and Ad Library reports, as well as an unsanctioned browser extension called the Ad Observer that users could install to scrape ad data from the Facebook website to donate to the Observatory. The Ad Observer is an open-source tool that underwent independent reviews of its code and privacy practices to ensure it adequately obtained user consent and collected only the data it needed. 101 Edelson claimed that they could not depend solely on data Facebook made availableparticularly Facebook's Ad Library-because it had many reporting inconsistencies and thousands of missing ads. 102 In late October 2020, Facebook sent a cease and desist letter to Edelson and McCoy, demanding NYU Cybersecurity for Democracy shut down its Ad Observer plug-in and delete any data collected from it. Civil society groups lashed back: more than fifty signed onto a letter from Mozilla demanding Facebook withdraw the cease and desist. 103 The Knight First Amendment Institute provided legal representation for Edelson and McCoy. 104 Little was heard from the case for the next several months while negotiations between Facebook and NYU Cybersecurity for Democracy continued behind closed doors.

*p. 27*
On August 3, 2021, negotiations broke down and Facebook suspended Edelson, McCoy, and others' Facebook accounts, thereby cutting off their access to Facebook's sanctioned tools, the API, Ad Library, and CrowdTangle. Facebook had cut off other ad transparency tools in the past, including ones from ProPublica, Mozilla, and Who Targets Me, but they largely did this by updating their own website in a way that broke those tools, not by suspending 101. JASON CHUANG, AD OBSERVER PRIVACY PROPERTIES & DATA COLLECTION 1, MOZILLA BUGZILLA, https://bug1676407.bmoattachments.org/attachment.cgi?id=9187255 (last visited Jan. 30, 2024). Specifically, that info was from the "Why am I seeing this ad?" box of each ad a user saw. Id. at 2. 102. Jeremy B. Merrill, How Facebook's Ad System Lets Companies Talk Out of Both Sides of Their Mouths, MARKUP (Apr. 13, 2021), https://themarkup.org/citizen-browser/2021/04/13/ how-facebooks-ad-system-lets-companies-talk-out-of-both-sides-of-their-mouths; Laura Edelson, Audit of Facebook Ad Transparency Finds Missed Political Ads, MEDIUM (Oct. 22, 2020), https://medium.com/online-political-transparency-project/audit-of-facebook-ad-transparency-finds-missed-political-ads-603f95027cc6. researchers' accounts. 105 In a blog post titled, "Research Cannot Be the Justification for Compromising People's Privacy," Facebook claimed that they "took these actions to stop unauthorized scraping and protect people's privacy in line with our privacy program under the FTC Order," and offered the Ad Library as an alternative. 106 There was immediate public outrage from academics, civil society, journalists, and lawmakers. 107 Edelson published an opinion piece in The New York Times a week after the incident arguing against Facebook's justifications blocking their work. 108 Edelson testified before Congress at the end of September, where she argued that to use the Ad Library, researchers were required to "sign an agreement that limits how they use and share the data, which significantly hampers meaningful publication of any research findings, as the dataset that would be necessary for other researchers to reproduce any findings cannot be publicly shared." 109 Edelson also argued that many ads were missing from the Ad Library and that others were intentionally mislabeled as non-political by bad actors. 110 FTC Acting Director of the Bureau of Consumer Protection Samuel Levine soon sent a letter clarifying that the NYU Ad Observer did not break Facebook's consent decree:

*p. 28*
Had you honored your commitment to contact us in advance, we would have pointed out that the consent decree does not bar Facebook from creating exceptions for good-faith research in the public interest. Indeed, the FTC supports efforts to shed light on opaque business practices, especially around surveillance-based advertising. While it is not our role to resolve individual disputes between Facebook and third parties, we hope that the company is not invoking privacy-much less the FTC consent order-as a pretext to advance other aims. 111 Despite being cut off from some data, NYU Cybersecurity for Democracy was able to release a new version of the Ad Observatory ahead of the 2022 elections. 112 Facebook (now Meta) has not shared whether or not they have reinstated any of the researchers' accounts as of this writing, but the company has expanded their own Ad Library to include more in-depth targeting information about political ads. However, researchers continue to argue that Ad Library misses several political ads since those running the ads do not identify them as political.

## D.

*p. 29*
A TAXONOMY OF PROBLEMS WITH RESEARCHER ACCESS TO SOCIAL MEDIA DATA

*p. 29*
We identify two broad categories of problems that currently afflict social media data sharing. The first is poor research quality; existing approaches to giving researchers access to data negatively impact the quality and utility of research that gets produced. The second is unrealized research; some socially beneficial types of research cannot be done at all with the data currently made available.

## Poor Research Quality a) Limited by Data Access Arrangements

*p. 29*
Platforms sometimes require researchers to sign burdensome contracts in order to gain access to data, as the NYU Ad Observatory argued Facebook has done. 113 Platforms can also impose large technical burdens, like how TikTok requires researchers using its research API to refresh results "at least every fifteen (15) days, and delete data that is not available from the TikTok Research API at the time of each refresh." 114 Even without requiring prepublication approval, a platform has unilateral power over the data it makes available and may be able to pressure researchers to suppress results that reflect on it negatively. This is particularly acute when companies provide ad hoc access to individual researchers, or when researchers receive direct funding from companies. The inability to share data further makes research results less robust and more difficult to publish since it is unreproducible and unverifiable.

## b) Unstable Data Access

*p. 30*
Platforms regularly change which data they make available to researchers and under what terms, often with little warning. Shortly after Musk acquired Twitter, for instance, the service very suddenly raised the cost of its API from free to $42,000 a month, making it inaccessible to nearly all academic researchers and jeopardizing hundreds of in-progress research projects. 115 Data access can change also because new threats to privacy and security are uncovered, as happened with SS1 and the Facebook Graph API in the wake of Cambridge Analytica. 116 The possibility of data access changing precludes entire research methodologies, such as longitudinal research, and threatens inprogress projects.

## c) Decontextualized Data Production

*p. 30*
Platforms often share only limited information about how they generate the data they share and how it has been filtered. Without understanding the provenance of data from platforms' tools, researchers often cannot know or predict how their data is skewed. This problem is not just theoretical. Studies show that tweets from Twitter's livestream API, which shares 1% of all live traffic, are not randomly sampled. 117 Often, platform-permissioned tools are not designed with research in mind, so they can be missing basic information. 118 Even when these tools are designed for researchers, opacity around the processes in which they are built can lead to huge oversights that even the platforms themselves miss, as occurred with SS1. 119

## d) Streetlight Effect

*p. 31*
The streetlight effect is a type of bias wherein people only search for something where it is easiest to look, just as someone who lost their keys outside at night might only look where there are streetlights. 120 A similar effect plays out in social media research: researchers often study the platforms where they can access the most data, not necessarily the ones most relevant to the effect they are trying to study. 121 Entire domains of research can end up centralizing around non-representative data sources, as some argue occurred with Twitter. 122 The streetlight effect also creates perverse incentives for companies not to share data. Companies that provide data may end up receiving more scrutiny and criticism from researchers. They may not even experience the public relations benefits of openness because they may be publicly criticized, as frequently and harshly as companies that share no data at all, for sharing insufficient data or in ways that make it difficult to use. 123

## e) Denominator Problem

*p. 31*
The denominator problem is when researchers are unable to use the volume of overall activity on a platform to contextualize their findings. 124 For instance, imagine that a researcher found five thousand tweets in Hindi over a week-long period of time that promote ethnic violence against Muslims. Without certain baseline information, such as the total number of tweets per week, tweets in Hindi per week, or total active users versus active Hindispeaking users, that researcher will not know whether their five thousand tweets should be considered a lot or a little. Social media companies frequently roll out changes to their systems. Sometimes, these changes are publicly announced and are meant to address controversies or harms uncovered by research. 125 However, without access to adequate data, researchers are unable to evaluate the effectiveness of these interventions, or whether they have been rolled out at all. Claims related to opaque technical systems, such as recommendation algorithms and content moderation practices, are nearly impossible to evaluate, making it difficult for the public to distinguish between public relations puffery and meaningful changes.

## b) Unequal Access Leads to Less Diverse Research

*p. 32*
Researchers with personal connections to large social media companies are more easily able to gain access to data through both informal and formal means. Well-connected researchers are more likely to convince companies to share data in ad hoc ways for one-off projects. 126 They are also better able to defend their unsanctioned access since they may have powerful allies, such as when the Knight First Amendment Institute offered legal defense to NYU Cybersecurity for Democracy for its Ad Observatory. 127 Less resourced and connected researchers may not even have the budget to purchase the computing power necessary to do certain research.

*p. 32*
This unmeritocratic approach to doling out access to data may lead to worse outcomes. The best-connected researchers are not necessarily the ones who come up with the best research questions or plans of execution. Underrepresented researchers may bring unique insights and approaches that more well-connected researchers do not.

## c) Inability to Discover Unexpected Effects

*p. 32*
Social media companies share non-public data with researchers in some areas more than others. Meta, for example, offers more information about political advertisements than it does non-political advertisements, in part because researchers have appealed to democratic values to gain such access. 128 By limiting access to other data not deemed as important, however, platforms may prevent researchers from discovering new, unexpected effects of different technological architectures, user interfaces, and policy designs. A change in the way a social network displays advertisements, for instance, could drastically increase how often users fall for cryptocurrency fraud. This effect would be unexpected and important, but impossible for researchers to discover for a number of reasons: researchers do not have access to data regarding how the company rolled out the change to advertisements (e.g., A/B test data), which content gets flagged as cryptocurrency fraud, which ads can be categorized as cryptocurrency ads, or how much engagement those ads receive. Companies are disincentivized from finding or sharing with the public new negative social impacts of their services.

## d) Slow Responses to Sudden Problems

*p. 33*
Sudden social, economic, and political upheavals often play out on social media. Fast evolving and paradigm shifting events such as COVID-19, the January 6th attacks, and the Russian attack on Ukraine are both reflected on and affected by the online information ecosystem. 129 Researcher organizations that use platform data access mechanisms to run social media monitoring programs, including the Stanford Internet Observatory and the Global Disinformation Lab at UT Austin, may be uniquely poised to give platforms the information they need to act quickly. Sharing timely data with external researchers, such as watchdog organizations and journalists, could help companies and the public better understand what is happening on platforms, and in turn, improve responses to such upheavals. Platforms, however, do not have policies to allow emergency access to data, even if it may be useful for all parties.

## E. THE LEGAL LANDSCAPE OF DATA SHARING

*p. 33*
This Section takes a step back to consider the state of social media data sharing from a legal point of view.

## What Made Things This Way?

*p. 34*
As the case studies above highlight, the barriers to data access are not only technical, but also legal. Subject to a few narrow exceptions outlined in state and federal privacy laws, social media data is subject to private ordering: once data subjects have consented to their data being collected, companies enjoy broad discretion to determine who gains access to social media data and on what terms such access is granted.

*p. 34*
Companies assert both legal rights and legal duties to control and manage access to proprietary data. Technically, there is no recognized legal property right in data per se, despite enduring debate over recognizing one. 130 Instead, companies rely on two kinds of legal claims to approximate full-throated entitlement rights over data access and control: rights to limit access to data to protect commercial secrets and competitive advantage, and obligations companies owe data subjects to limit access to data, which may arise under companies' terms of service or privacy laws. Together, these two kinds of legal claims allow companies to justify broad, contractually governed discretion over how researchers gain access to data.

*p. 34*
Both trade secrecy and privacy claims generally arise out of underlying contractual legal relationships that structure companies' claims to and obligations regarding social media data. Two kinds of contractual relationships govern, to a large degree, how social media data is collected, processed, and used. First, terms of service govern collection and the relationship between companies and data subjects, and second, data use agreements govern data access and the relationship between companies and researchers.

*p. 34*
Companies have been able to constrain access to data in the contractual realm because of their success at invoking underlying privacy and trade secrecy rationales-rationales that companies use as obstacles to increased public oversight and control over researcher access. Thus, we focus on privacy and trade secrecy because these are the doctrinal obstacles and normative justifications that platforms invoke in public statements against researcher access. To retrieve affirmative public rights of researcher access from the realm of private contractual ordering requires us to address these privacy and trade secrecy claims. First, companies make trade secrecy claims to protect their commercial interests in data acquired from users and used to develop their products. 131 As Tait Graves and Sonia Katyal have written (in a broad survey of recent trends in trade secrecy law), "companies are increasingly exploiting [gaps in trade secrecy doctrine] to assert trade secret rights in a growing range of nontraditional contexts." 132 Under now-dominant definitions of a trade secret, information qualifies for trade secret protection if it (1) is generally not known to others in the same industry; (2) is not readily ascertainable from the use of limited time and effort; (3) has actual or potential independent economic value to competitors; and (4) is reasonably guarded as secret. 133 This broad definition permits companies to claim-often without substantiation-proprietary rights over a sweeping range of information. 134 Once a claim of trade secrecy is made, companies wield the claim to withhold the information from researchers and even regulators. 135 These companies argue that disclosure of the secret information-even to these noncommercial audiences-will inevitably lead to some leaks to competitors, encouraging free riding and thereby eroding crucial incentives to innovate. 136 Tech platforms have a track record of making such trade secrecy claims. For instance, in its 2020 comments to the FTC on data portability, Facebook alleged that data such as granular use logs, non-human understandable data, and data stored in formats that rely on proprietary technology "make clear that including all observed and inferred data could also result in a different sort of burden: the disclosure of trade secret or other proprietary information developed by a business to enhance or differentiate its services. Enabling people to port that kind of information could reduce incentives for businesses to develop it in the first place." 137 In 2021, Facebook withheld internal research on the impact of its platforms on youth mental health from senators, stating that "its internal research is proprietary and 'kept confidential to promote frank and open dialogue and brainstorming internally.'" 138 In 2023, the Information Technology Industry Council (ITI), issued a statement expressing concern over the European Union Data Act's data sharing provisions. ITI, which includes Google, Meta, Microsoft, and Snap as members, argued the law should be amended to permit companies to "refus[e] to share data in specific circumstances where disclosure of trade secrets would be likely to cause serious damage to the data holder." 139 A bit further afield, Uber Eats and two other food delivery platforms challenged a New York City municipal ordinance requiring platforms share customer data with the underlying restaurant fulfilling an order. All three platforms asserted that the law constitutes a violation of their trade secrecy rights under the Second Circuit standard. 140 Companies also use other entitlement-like claims to limit extra-contractual researcher access. For instance, despite recent cases limiting the application of such laws to certain forms of research, many companies still include language in their terms of service indicating that activity that violates their terms will be referred to law enforcement for prosecution under the Computer Fraud and Abuse Act (CFAA). 141 Copyright enforcement has similarly endowed platforms with legal rights to control and manage access. For instance, the Digital Millennium Copyright Act (DMCA) not only establishes a takedown regime for unauthorized content, but also includes prohibitions against circumventing technical access protections, knowingly and improperly obtaining valuable trade secrets, and distributing technologies that facilitate circumvention. 142 The practical upshot of the DMCA, particularly the provision against trafficking in circumvention technologies themselves, is that platforms enjoy strong rights over access control protocols. 143 b) Privacy Second, companies assert that the privacy obligations they owe consumers (either via the contractual promises they make to data subjects or due to privacy regulations with which they must comply) are reasons to deny researcher access. 144 These concerns, while sometimes used as pretexts by companies to protect the value of walled-off data assets, are not always levied in bad faith or without merit. Users have legitimate privacy interests in the data at issue in researcher access; protecting this legitimate interest makes researcher access a legally and ethically tricky problem. 145 Indeed, researchers therefore that the Plaintiffs' proposed research plans were not criminal activity under the CFAA); Facebook, Inc. v. Power Ventures, Inc., 844 F.3d 1058, 1065-69 (9th Cir. 2016) (holding a third-party platform civilly liable under the CFAA for accessing Facebook users' data). 142. To be clear, the DMCA does not directly apply to social media data (which is not as a general matter copyrightable), but it has featured significantly as a background law governing the relationship between online platforms and external researchers of those platforms, and depending on the research in question, may be implicated in a given form of social media research.

*p. 37*
143 145. The argument that researcher access is normatively good for user privacy is orthogonal to the argument of this Article. That said, there are compelling reasons to think that well-designed researcher access mechanisms for social media data may have salutary effects on the overall privacy of social media users. This view is suggested by the FTC's favorable response to NYU's Ad Observatory and other research that seeks to "shed light on themselves have recognized that proposals to increase access to social data pose privacy risks to platform users. 146 Some information privacy laws affirmatively grant data subjects additional rights and impose additional duties on platforms. For example, the California Consumer Privacy Act (CCPA) grants data subjects rights to request information about what data is being collected about them and whether any of their personal data is being sold or disclosed to third parties. 147 It also grants data subjects the right to opt out of the sale of their personal information. 148 The Children's Online Privacy Protection Act (COPPA) imposes additional obligations on platforms regarding data collected from children under thirteen years of age. 149 To comply, platforms must post comprehensive policies regarding their practices for such data and obtain verified parental consent prior to any data collection, among other requirements. Although COPPA does not prohibit children under the age of thirteen from sharing their data with platforms, many social media platforms prohibit children under age thirteen from using their services due to the costs and risks associated with violating COPPA. 150 These contractual terms-and several federal privacy laws, including COPPA-are in turn regulated by the Federal Trade Commission Act's § 5 authority and state consumer protection laws. 151 opaque business practices," in the wake of Meta's efforts to use obligations under its 2012 FTC consent decree as a justification to shut down that research. See supra Section II.C. Such cases, where two sides of a dispute both marshal privacy arguments to advance their claims (in this case, companies and social media researchers), present an instance of what David Pozen calls a 'privacy-privacy tradeoff. ' As the case studies above highlight, the legal barriers erected by privacy obligations to researcher access (as well as the perceived legal risks accompanying these barriers) are significant. In the case of SS1, the growing legal complexities around compliance with the GDPR and CCPA were key contributors to the consortium's failure. In the case of the NYU Ad Observatory, Facebook invoked privacy duties-its supposed obligations under its FTC consent decree, and its obligations to users under their terms of service-to cut off researcher access.

*p. 39*
These cases also demonstrate additional complexities when it comes to assessing the merit of privacy claims. On the one hand, social media companies may invoke privacy obligations in bad faith to withhold data that makes them look bad. 152 In the case of the NYU Ad Observatory, for example, Facebook's attempt to use its FTC consent decree to block access to data was undermined by the FTC itself. 153 The agency clarified that it welcomed and encouraged greater researcher access to platform data.

*p. 39*
On the other hand, companies also underinvest in privacy, and sharing data with researchers can raise legitimate privacy risks. Perhaps the most infamous example here is the Cambridge Analytica scandal, which nominally involved data harvested for a research project. SS1 sits somewhere in between this example and the NYU Ad Observatory example. Researchers and Facebook became mired in concerns over what SS1 would mean for Facebook's obligations under significant, new data protection laws. Some viewed Facebook's privacy concerns as pretextual; the company used exaggerated estimates of the perceived legal risk of new laws to wriggle out of obligations it no longer wanted to fulfill. However, Facebook was not alone in its assessment of risk. Credible third-party groups, including EPIC, clearly thought that SS1 raised genuine privacy concerns. Regardless of whether companies raise privacy concerns in good or bad faith, courts and would-be legislators must consider the merit of such claims. 155 On this count, the privacy concerns of data sharing clearly present a challenge to unfettered researcher access, and they require good faith engagement.

## Navigating a Path Forward Between Privacy and Trade Secrecy

*p. 40*
Alongside the strong legal claims of companies over social media data is the conspicuous absence of rights to access for other entities. Users themselves have some individual rights over their data, but researchers and even government agencies have limited countervailing legal rights over data to supersede those of companies. 156 This is notable, given that absolute rights of any kind are rare in law, particularly with respect to intangible goods, and that government claims that limit or supersede private (commercial) claims of right in the course of ordinary socioeconomic legislation were once more common. 157 The lack of public rights in social media is also extraordinary given the magnitude of the public interests at stake. Social media companies are some of the largest companies in the world. They exert significant influence on the public sphere, affecting how billions of people around the world interact with one another and with the news of the day. These spaces are key to self, social, and political formation. They generate billions, if not trillions, of dollars of revenue. And yet very little is known about how they actually work.

*p. 40*
As this Article endeavors to show, we must overcome the legal barriers to researcher access imposed by trade secrecy and privacy claims to examine how these platforms work. Or, more accurately (and more humbly), we must find ways to navigate safely past these barriers. For this, we now turn to the lessons of another powerful industry where researchers have been granted access to valuable and sensitive commercial data: pharmaceutical and medical device companies' clinical trials.

## III. CLINICAL TRIAL DATA SHARING: MANDATE AND EXPERIMENTS

*p. 41*
What is clinical trial data? What is the clinical trial data sharing mandate, and why might it matter for governance of social media data? What mechanisms have emerged for responsible sharing of even the most sensitive components of clinical trial data? This Part answers these questions.

*p. 41*
In this Part, Section III.A introduces clinical trial data. Section III.B provides historical context for the clinical trial data sharing mandate that emerged in the United States in the 21st century. Section III.C then describes the 2007 legislation-the Food & Drug Administration Amendments Act (FDAAA)-that forms the foundation of that mandate. The law works, albeit imperfectly, and it has unlocked benefits for researchers, patients, and the broader public. Section III.D then describes the institutions that implement FDAAA and other laws that govern researcher access to clinical trial data. Section III.D also shows that some institutions that share clinical trial data have been able to achieve deeper sorts of data sharing with researchers. These relationships have made the most sensitive components of trial dataindividual patient data (IPD) and detailed trial methodologies that may implicate companies' trade secrets-accessible to researchers. Section III.E distills key features.

*p. 41*
Today researchers have meaningful access to much of the very same data that companies rely on for their research and development (R&D), regulatory approvals, and profits. So far, at least, clinical trial data sharing also capably protects the interests of the people who create this data by volunteering for clinical trials.

*p. 41*
WHAT IS CLINICAL TRIAL DATA, AND WHY DOES IT MATTER?

## Clinical Trial Data Defined

*p. 41*
Clinical trials are research studies on human volunteers. Clinical trials answer questions about different health interventions, such as surgeries, drugs, vaccines, knee replacements, and changes in exercise or diet.

*p. 41*
The highest quality clinical trials are randomized and controlled. Human subjects are assigned at random to different "groups" within the trial; one of the groups is a "control group" that receives a standard intervention, a placebo, or no intervention at all. By comparing outcomes in the treatment and control group, the safety, efficacy, and other properties of the intervention under study can be measured. Randomized controlled trials are the most important means of testing whether a particular intervention is safe and effective-the "gold standard" of evidence-based medicine. 158 Clinical trials are traditionally categorized into one of four "phases." "Phase 1" trials are the first trials conducted on a new intervention. Small and cautious, they are primarily used to evaluate safety. "Phase 2" trials are larger and longer; they gather more safety information and begin to explore the intervention's efficacy. "Phase 3" trials are still larger; they weigh benefits and harms and examine rare adverse events in a larger population. "Phase 4" trials are done after an intervention is already on the market and in wide use, to study longer-term safety and effectiveness, new uses in new patient populations, and other outstanding questions. 159 Clinical trials generate lots of data, especially large Phase 3 and Phase 4 trials. One 1999 estimate concluded that a typical Phase 3 clinical trial design with 2,000 patients studied for twelve months could "generate up to 3 million data points." 160 Thousands of clinical trials are conducted every year, making the total quantity of trial data enormous.

*p. 42*
There are numerous components of clinical trial data, each with its own properties, utility, and sensitivities. Before proceeding further, we provide a brief taxonomy of clinical trial data. Clinical trial data contains three distinct components: (1) individual patient-level data (IPD), (2) summary data, and (3) metadata. 161 Together, these three components constitute the body of information collectively referred to as clinical trial data. The first and perhaps most obvious component of clinical trial data is individual patient-level data (IPD). 162 IPD is the "raw" data collected on individual patients. Among other things, it reveals the precise health statuses of different patients-the testing, care, and diagnoses they receive; the side effects and other "adverse events" they experience; and so on. 163 Expert users of data, such as academics and the FDA's regulatory scientists, may be most interested in IPD, but other users, such as journalists and patient groups, may find it difficult to use and understand.

*p. 43*
IPD is the most sensitive data component from a patient privacy perspective. It is the rich, detailed personally identifying information (PII) of the clinical trial world, as it links specific health status information with specific individuals. 164 IPD can be de-identified by redacting obvious identifiers such as name, birth year, and zip code, 165 but it remains IPD after de-identification as it continues to characterize the health status of individual people rather than larger groups. Thus, even after de-identification, IPD remains at risk of reidentification and subsequent effects on individual patients.

## b) Summary Data

*p. 43*
The second data component of clinical trial data is summary data, also known as aggregate data. As the name suggests, this data does not reveal the health status of individual people but instead reveals something about groups of people-e.g., the treatment and control arms of a clinical trial, or demographic subgroups of patients in the trial (such as patients over age sixty-five). Some summary data includes explanations and simple "takeaways" digestible to nonexpert readers, such as high-level conclusions about a drug's safety and efficacy (or lack thereof) in a group of people.

*p. 44*
Summary data may span multiple trials. The FDA, for example, synthesizes IPD from multiple trials to produce summary data useful to patients and doctors. 166 The term "summary" clinical trial data suggests brevity, but some important summary data runs long. Standard summary clinical study reports (CSRs) can run many thousands of pages and provide expert readers with a wealth of information. 167 c) Metadata

*p. 44*
The third component of clinical trial data is metadata. Metadata is data about the other data components. It describes how, exactly, IPD and/or summary data is generated, recorded, analyzed, and presented. Analysis of metadata alongside IPD and summary data can confirm that IPD and summary data are trustworthy-and reveal and discourage manipulation and mistakes. 168 In the context of clinical trials, the term "metadata" commonly refers to specific standardized documents and data elements: the clinical trial protocol, the statistical analysis plan (SAP), and any analytic code used in connection with the SAP. Together, these resources provide a trial's precise methodology: what questions the trial was intended to answer; what patients were included in and excluded from the trial; what patient "outcomes" it measured (such as tumor size or cholesterol levels); how those measurements were taken and processed; and more.

## The Value of Clinical Trial Data and Clinical Trial Data Sharing

*p. 44*
Clinical trial data is terrifically valuable and expensive to generate. Even a simple trial costs millions of dollars to run; larger, longer Phase 3 trials typically cost tens of millions of dollars. 169 The costs are worth incurring because the 166. For example, the FDA publishes simple "Medication Guides" that explain to patients how to make safe use of certain relatively risky drugs. 21 C.F.R. § 208.24 (2022). The FDA also publishes more detailed "approval" packages, described infra Section III.C.1, that likewise synthesize findings from multiple trials.

*p. 44*
167 data generated is scientifically and commercially valuable. Drug, vaccine, device, and other for-profit companies around the world spend tens of billions of dollars on clinical trials to guide their research, to support marketing efforts, and to generate sufficient data to earn approval from the FDA and other regulators around the world. 170 As with social media data, the stakeholders in clinical trial data are numerous. Key stakeholders include patients themselves; doctors, nurses, and other providers whose care is shaped by trial results; hospitals, clinics, and other organizations that employ the providers (and are liable for many of their actions); innovative companies that develop and sell new drugs, devices, and vaccines; generic and biosimilar companies that seek to sell similar products at lower prices; scientific researchers in academia, government, and nonprofit nongovernmental organizations who do basic research; government regulators who conduct, referee, and pay for research; journalists, academics, and civil society researchers who watchdog those regulators and the healthcare system as a whole; and the public at large, who pay for the regulators and pay a fortune for healthcare. 171 All these stakeholders are important, but for purposes of this Article, we focus on researchers and research uses of clinical trial data that provide benefits to the broader public. In 2015, a landmark report from the Institute of Medicine (now known as the National Academy of Medicine) characterized the benefits of IPD sharing as follows: 172 Clinical Benefit Trials Supporting the US Approval of New Therapeutic Agents, 2015-2017: A Cross-Sectional Study, 10 BMJ OPEN 1 (2020). Some industry-funded estimates based on industryprovided data put the average cost of a Phase 3 trial over $200 million. From the perspective of society as a whole, sharing of data from clinical trials could provide a more comprehensive picture of the benefits and risks of an intervention and allow health care professionals and patients to make more informed decisions about clinical care. Moreover, sharing clinical trial data could potentially lead to enhanced efficiency and safety of the clinical research process by, for example, reducing unnecessary duplication of effort and the costs of future studies, reducing exposure of participants in future trials to avoidable harms identified through the data sharing, and providing a deeper knowledge base for regulatory decisions.

*p. 46*
In the long run, sharing clinical trial data could potentially improve public health and patient outcomes, reduce the incidence of adverse effects from therapies, and decrease expenditures for medical interventions that are ineffective or less effective than alternatives. In addition, data sharing could open up opportunities for exploratory research that might lead to new hypotheses about the mechanisms of disease, more effective therapies, or alternative uses of existing or abandoned therapies that could then be tested in additional research.

*p. 46*
In the following Sections, we show in more detail how independent researchers have used access to IPD and other clinical trial data to interrogate manufacturers' claims about their products and help protect the public from unsafe, ineffective, or exaggerated products (think Ad Observatory, but for drugs). For now, one vivid example of the value of clinical trial data sharing: the antidepressant paroxetine ("Paxil").

*p. 46*
Paroxetine was never approved for use in children but became popular with providers, who wrote over two million prescriptions for children per year in the early 2000s on the basis of a 2001 medical journal article. The drug's manufacturer, GlaxoSmithKline, funded and disseminated the article, 173 which claimed that the medicine was "generally well tolerated and effective" in young patients. 174 In fact, paroxetine caused suicidal thinking and suicide in many children. 175 In 2003 and 2004, after widespread anecdotal reports of teen suicides caused by paroxetine, FDA scientists reanalyzed earlier-submitted clinical trial data and concluded that the drug causes increased risk of suicide and suicidal ideation. 176 This led to stricter prescribing rules and a wave of litigation against GlaxoSmithKline. GlaxoSmithKline ultimately pled guilty to fraud, 177 and paroxetine is no longer widely prescribed to children.

*p. 47*
In the 2010s, independent academic researchers eventually convinced GlaxoSmithKline to share more comprehensive data from the trial described in the 2001 article. They found that the trial data had shown the risks all along and that GlaxoSmithKline had misrepresented the data. 178 The researchers concluded that the affair "illustrates the necessity of making primary trial data and protocols available to increase the rigor of the evidence base." 179 Had GlaxoSmithKline's data been shared with independent researchers in 2001, they might have raised the alarm then, and years of harm might have been averted.

*p. 47*
Independent research conducted with clinical trial data is not limited to investigation of questions of safety and efficacy, vital as those questions obviously are. Independent research also helps private and public payers allocate resources better. For example, the nonprofit organization Institute for Clinical and Economic Review (ICER) uses trial data and other medical data to undertake detailed analyses of the cost-effectiveness of various medical interventions, including everything from comparison of all FDA-approved multiple sclerosis drugs 180 to service dogs as treatment for post-traumatic stress disorder. 181 Meta-analysis of pooled clinical trial data established that the blockbuster influenza drug oseltamivir ("Tamiflu") is only modestly effective and that massive stockpiling was a poor use of billions of dollars of public money. 182 [Vol. 39:109

## The Dangers of Clinical Trial Data Sharing

*p. 48*
Of course, sharing clinical trial data with researchers has risks, too. There are legitimate and strong countervailing interests that often militate against sharing. The two predominant interests here are patients' privacy (especially as to IPD) and innovative companies' competitive interests. 183 The latter are often articulated in terms of "incentives to innovate" and "protection from free-riders," or framed in terms of specific intellectual property doctrines, such as trade secrecy.

*p. 48*
Others' work has thoroughly analyzed both these important interests, in the context of clinical trial data, in the context of healthcare more broadly, and in the context of valuable data writ large. 184 In the Sections that follow, we will show specific instances of such arguments being raised by the pharmaceutical or medical device industry and then accommodated or rebutted by the legislators and governors of the clinical trial data-sharing mandate.

*p. 49*
Note here that the parallels with social media data are strong. Just as platform companies have invoked patient privacy and innovation to limit sharing their data with researchers, so too have large, incumbent companies that hold and profit from clinical trial data. 185 For example, in 2015, shortly after the National Institutes of Health (NIH) proposed a new rule mandating expanded sharing of certain summary and metadata from clinical trials with researchers and the broader public, 186 the Pharmaceutical Research and Manufacturers of America (PhRMA) association warned, ominously, that "the rule does not adequately protect the process of medical research innovation. Failure to protect adequately trade secrets and confidential commercial information would harm public health by discouraging the very innovation necessary to bring new medical advances to the market." 187 NIH responded that PhRMA's concerns were overblown and that NIH's rule struck an appropriate balance. 188 Since NIH's rule went into effect in 2017, NIH has proven correct-as the next two Sections show.

## B. "DARK AGES" OF CLINICAL TRIAL SECRECY: LITTLE RESEARCHER ACCESS, UNREALIZED BENEFITS, AND HARM TO PATIENTS

*p. 49*
This Section explains how today's clinical trial data sharing mandate emerged out of comparative "dark ages" of data secrecy, contestation, and unnecessary human suffering.

*p. 49*
Consider the United States in 1960. There was then no explicit law governing researcher access to clinical trial data and other kinds of medical research data. In addition, more rudimentary information technology meant that data was more difficult to share and use.

*p. 49*
Because no law mandated researcher access, drug companies, medical device manufacturers, universities, and other entities that conducted clinical trials were free to disseminate or withhold data as they saw fit. They massaged data, such as by publishing selective data in medical journals that painted their 185. For a broad, independent view of the pharmaceutical industry's claims of trade secrecy in clinical trial data, see W. products in the best possible light. 189 The medical literature was thus incomplete and manipulated.

*p. 50*
In fact, as of 1960, drug companies sometimes withheld clinical trial data not just from researchers but from the FDA itself. An infamous example: In 1960 and 1961, one FDA scientist, Frances Kelsey, grew concerned over a lack of safety data to the FDA on the drug thalidomide, even as the drug had been approved and entered widespread use in Europe and Australia. Kelsey came to suspect that the drug's manufacturer, the William S. Merrell Company, was withholding safety data from the FDA 190 and requested this missing data from the company. 191 Kelsey's insistence on receiving the data parallels the FTC's recent insistence that social media companies turn over certain data pursuant to a past consent decree. 192 Kelsey's lengthy review of thalidomide prevented widespread use in the United States. By late 1961, reports of thousands of horrifying birth defects and fetal deaths caused by the drug in other countries led to its withdrawal from pharmacies worldwide. Kelsey was justifiably hailed as a hero for protecting Americans from its harms.

*p. 50*
The thalidomide catastrophe, and a broader "full disclosure movement" that coalesced in drug regulation in the wake of other, smaller drug scandals, 193 prompted Congress to enact the first important federal clinical trial data sharing legislation: the 1962 Kefauver-Harris Amendments to the Food, Drug & Cosmetics Act. This legislation mandated, for the first time, that drug companies submit clinical trial data to the FDA as a condition of market approval, and it gave the FDA legal authority to dictate exactly how that data was packaged and presented to the agency. If companies didn't comply, the FDA could keep products off the U.S. market. The FDA became the world's largest reservoir of clinical trial data, which it remains today. 194 But the Kefauver-Harris Amendments did not guarantee researcher or public access to that data. The "full disclosure" in "full disclosure movement" meant full disclosure to the FDA, not to independent researchers. Against a statutory blank canvas, the FDA had no legal obligation to disclose any of the clinical trial data in its possession to the broader public. 195 Through the 1960s, the FDA's choice was to keep most of this data confidential; its expert reviewers worked mostly in secret. Independent researchers outside the FDA typically learned the results of clinical trials from the medical literature, where industry continued to cherry-pick the data it wanted to share.

*p. 51*
At least as early as 1969, some FDA officials expressed a desire to change this state of affairs and make all clinical trial data held by the agency public once the product in question had been approved for sale. 196 Tentative, inconsistent efforts to do so through discretionary agency action proved unsuccessful, in part because they were undone by a rotating cast of more industry-friendly, pro-secrecy FDA commissioners, and in part because for years, the FDA was threatened with legal challenge by the powerful pharmaceutical industry. From the 1970s to the 1990s, there remained no coherent statutory regime guaranteeing researcher access to clinical trial data, even as researchers clamored for access. In 1978, a bill that would have mandated disclosure of summary data, metadata, and IPD, called the Drug Regulation Reform Act (DRRA), was defeated in Congress. 198 The bill failed to pass despite support from the Center for Law and Social Policy, the Environmental Defense Fund, and Public Citizen. 199 In 1980, McGarity and Shapiro published an article in the Harvard Law Review criticizing the FDA's then-skimpy disclosure of industry-generated clinical trial data in the agency's possession; 200 this practice contrasted with the EPA's much richer data disclosure of testing data on pesticides 201 and the FDA's own richer data disclosure on food additives. 202 The 1984 Hatch-Waxman Act was, in early drafts of the legislation, to have included a DRRA-like provision that would have required the FDA to publish volumes of clinical trial data when product applications were approved or denied. 203 The pharmaceutical industry's lobby watered down the statutory language, arguing that mandatory disclosure would undermine patient privacy and its trade secrecy interests. 204 At the same time, the FDA Commissioner testified in Congress on the alleged benefits of data secrecy and urged construction of the watered-down statutory language in ways that perpetuated the secretive status quo. 205 In the late 1990s, the FDA began voluntary, 205. Then-FDA Commissioner Frank Young intervened during the negotiation and passage of the Hatch-Waxman Act in 1984 to express the view that the statutory text of Act did and should not expand the agency's obligation to disclose safety and efficacy data, despite statutory language mandating that "[s]afety and efficacy data" "be made available to the public, upon request," under various circumstances. O'Reilly, supra note 203, at 20-21; Fisher, supra discretionary disclosure of some summary data and metadata from clinical trials, but shared this data only after product approval for a subset of approved products 206 and on a leisurely timeline. 207 The pharmaceutical industry largely thwarted researcher access into the 2000s. 208 For example, in 2000, David Willman of The Los Angeles Times reported a meticulous, Pulitzer-Prize-winning series of articles 209 on seven drugs that had been withdrawn between 1993 and 2000 for causing death and other serious side effects, revealing weaknesses in the FDA's drug approval process and in the pharmaceutical industry's ethics. 210 Willman remarked on the difficulty of his investigation and the FDA's then-still-prevalent culture of data secrecy. For example, data from one important clinical trial showing deaths in kidney transplant patients taking the immunosuppressive drug tacrolimus ("Prograf") had been disclosed to the FDA but not made readily available to outside researchers; per Willman, "the only way for doctors or patients to find that data is to search the medical literature or seek the FDA's review documents" through FOIA. 211 Similarly, in 2004, Barry Meier of The New York Times reported that medical researchers seeking to investigate the safety of antidepressants "could get only pieces of" relevant trial data, as "drug companies refused to turn over data . . . even though these researchers had helped come up with it." 212 Meier added that companies blocked researchers from "shar[ing] their own data with colleagues who had not worked" on a particular trial, siloing researchers from one another. 213 In 2006, two representatives of the prominent nonprofit Public Citizen, Peter Lurie and Allison Zieve, summarized the lamentable state of affairs: "Those committed to the free exchange of scientific information have long complained about various restrictions on access to [the FDA's] pharmaceutical data and the resultant restrictions on open discourse." 214 During this time, there was some voluntary sharing of data by drug and device manufacturers. As noted above, these companies selectively published data in medical literature. Some companies went further and made databases of certain clinical trial data and other data (e.g., genetic data) available to academic and other researchers. Companies that shared more were praised for "transparency," but this transparency was selective and subject to some of the same "pathologies" of voluntary sharing of social media data identified in Part II-decontextualization and streetlight effects especially. (For example, Merck, a company that received praise in the 1990s for voluntary sharing of some kinds of data, 215 was later shown to have hidden other data on the safety of rofecoxib ("Vioxx") that contributed to the deaths of tens of thousands of people. 216 ) As Deborah Zarin and Tony Tse stated in 2007, there were twelve "pharmaceutical industry-sponsored clinical trial databases," but they were "generally not reviewed by experts external to the company." An independent investigation "found that when conclusions were listed in these databases, they tended to be more favorable for the company's product than those found in published articles or FDA reviews of the same trials." 217 Perhaps not coincidentally, the 1990s and 2000s were marked by a series of increasingly high-profile scandals involving drug companies that hid unfavorable clinical trial data from independent researchers and the broader public, leading to widespread harm to patients. Two of the highest-profile scandals involved the drugs paroxetine ("Paxil") and rofecoxib ("Vioxx"). The basic details of the paroxetine scandal are summarized above. GlaxoSmithKline gathered evidence that its drug fueled tens of thousands of teen suicides, then intentionally hid that evidence from the public. 218 The paroxetine scandal gripped the public consciousness and helped spur Congress to action. When the FDA decided to warn doctors and parents to stop giving paroxetine to children in 2003, the story made headline news. 219 Media not only covered paroxetine's contributions to a spike in teen suicides but also big pharma's culture of data secrecy. A 2004 New York Times Magazine story observed "public outrage at revelations that a number of pharmaceutical companies had deliberately withheld damning information about [antidepressants including Paxil]-specifically, data from clinical trials that suggested that these drugs were both more dangerous and less effective for adolescents than millions of consumers had been led to believe." 220 The rofecoxib ("Vioxx") scandal was perhaps even more shocking. Rofecoxib, a painkiller, was approved by the FDA in 1999 and quickly became a blockbuster, earning Merck billions of dollars. 221 Then, in 2004, Merck abruptly removed the drug from the market, with encouragement from the FDA and other drug regulators because-Merck admitted-it caused heart attacks, strokes, and heart failures. 222 Merck held, internally, clinical trial data establishing these deadly side effects but did not disclose the data to independent researchers or the broader public. Merck moved to withdraw the drug only because a courageous FDA scientist with access to the data, David Graham, double-checked the agency's analysis and raised concerns, first with the agency and then with the U.S. Senate and the broader public. 223 The relevant trial data was first made available to independent researchers only years later, through litigation. 224 These researchers quickly proved that signals of these risks were present in data held by Merck and the FDA nearly 3.5 years before the drug was withdrawn from the market. 225 Had independent researchers gotten access to the data sooner, they could have caught the problem and averted at least 39,000 deaths. 226 Prominent scientists pointed to Vioxx as evidence that clinical trial data should be "stored on an academic site, analysed by non-company investigators, and eventually made accessible to the public for scrutiny." 227 The New York Times covered the Vioxx scandal at length, publishing stories on the FDA's promises of greater clinical trial data sharing 228 and the pharmaceutical industry's unreliable commitments to transparency. 229 In short, Vioxx and Paxil were "Cambridge Analytica moments" for the pharmaceutical industry. GlaxoSmithKline's efforts to downplay safety problems with a different drug, rosiglitazone ("Avandia"), constituted a third such moment, prompting more Congressional hearings and calls for reform. 230 Pharmacia's manipulation of data on another blockbuster painkiller drug, celecoxib ("Celebrex"), arguably created yet a fourth. 231 Clinical trial data secrecy had become a matter of national attention. Resulting public outrage 232 prompted Congress to revisit the possibility of legislation mandating data sharing by pharmaceutical and medical device companies and resulted in breakthrough federal legislation that forms the foundation of today's data sharing mandate.

*p. 57*
The pharmaceutical and medical device industries fought data-sharing legislation from the start. As Galbraith details, 233 Not surprisingly, the pharmaceutical industry's trade group did not support the FACT Act [proposed federal legislation that would mandate sharing of clinical trial data]. Originally, the Pharmaceutical Research and Manufacturers of America (PhRMA) asserted that a results reporting requirement was unnecessary. 202 However, in January of 2005, faced with pressure from lawmakers, the medical community, and the public, the four largest pharmaceutical trade groups in the world, including PhRMA, released a joint statement on the disclosure of clinical trial information. 203 While the group members pledged to release a nominal amount of information regarding ongoing trials, they did not commit to submitting the data to a comprehensive, government-sponsored registry. 204 Instead, the provisions left open the possibility of publishing the information on individual, company-sponsored websites that could contain internal rules that might not be publicly disclosed and consequently may differ from one site to the next . . . . Furthermore, with regard to completed trials, the pharmaceutical manufacturers agreed only to make public "summary results" of the studies and, additionally, asserted such disclosure "must maintain protections for . . . intellectual property and contract rights." Just as Facebook and other social media platform companies claim today, pharmaceutical companies in the 2000s argued that laws mandating data sharing would compromise their trade secrets and the privacy of individual data subjects. 234 Congress enacted mandate legislation anyway. 235 When NIH then proposed the rule implementing the legislation, the pharmaceutical lobby again sang the same tune, warning that "the rule does not adequately protect the process of medical research innovation. Failure to protect adequately trade secrets and confidential commercial information would harm public health by discouraging the very innovation necessary to bring new medical advances to the market." 236 As we describe in the next Section, the pharmaceutical lobby's concerns proved unfounded. NIH and other stewards of sensitive and previously secret clinical trial data have proven capable of collecting it from industry and sharing it with researchers without compromising patient privacy or incentives to innovate. The legislation that the pharmaceutical lobby resisted now forms the cornerstone of today's clinical trial data sharing mandate, pushing the industry out of the dark ages.

## C. LEGISLATING TODAY'S CLINICAL TRIAL DATA SHARING MANDATE

*p. 58*
The story of today's clinical trial data sharing mandate begins with the legal system: first legislation, and then regulation to implement and extend legislation. As in many other contexts, public law provided a necessary counterweight to private power. Public law mandated that drug and device companies make clinical trial data available to researchers and empowered federal regulators, such as the FDA and the NIH, to enforce compliance and govern that data.

*p. 58*
The single most important piece of American law in the clinical trial data sharing mandate is the Food and Drug Administration Amendments Act (FDAAA), enacted in 2007. FDAAA was described by the then-FDA commissioner as "massive legislation" informed by a "spirit of transparency." 237 A key achievement of FDAAA was to mandate universal disclosure of summary and metadata from clinical trials (though not IPD).

*p. 58*
FDAAA achieved much broader researcher access to clinical trial data in two ways: (1) mandatory publication by the FDA of "approval packages" that contain clinical trial data (and more); and (2) mandatory submission of clinical trial data to NIH, for validation and posting by NIH on a public website, ClinicalTrials.gov. We discuss each in turn.

## Mandatory Publication of Approval Packages

*p. 58*
FDAAA mandates that every time the FDA approves a new drug or vaccine, the agency must publish an "approval package" 238 data and metadata from all the clinical trials on which it relied for approval. 239 The approval package provides a summary of both the drug manufacturer's data and the FDA's independent analysis. 240 The FDA must publish the approval package within thirty days of approval. 241 In effect, this provision of FDAAA obliges the FDA to share some of its vast reservoir of data with the public.

*p. 59*
Today the FDA publishes approval packages as a matter of routine practice, on a website it calls "Drugs@FDA." 242 These packages fuel important research. 243 For example, a 2013 review article observed that "FDA documents contain unpublished evidence that can be highly useful in resolving publication bias and selective outcome and analysis reporting, identifying important harms, and filling gaps in knowledge about understudied subpopulations, outcomes, and comparisons." 244 In effect, approval packages equip independent researchers to overcome structural problems that afflict independent research, including the problem of decontextualized data production (by giving researchers more objective context, including the FDA's own analysis) and the streetlight effect (by giving researchers access to the FDA's data, rather than simply to a cherry-picked subset that drug manufacturers choose to publish in the medical literature).

*p. 60*
To show the value of the FDA's approval packages to independent researchers and the broader public, a few concrete examples: In 2014, independent researchers used an approval package to detect and publicize errors in clinical trial data reporting by the drug company Roche on its antiinfluenza drug oseltamivir ("Tamiflu"). 245 In the same year, different researchers used an approval package to establish that the anti-inflammatory drug roflumilast ("Daxas") provides net benefits to patients with severe chronic obstructive pulmonary disease (COPD), but not patients with milder disease, reshaping prescribing habits. 246 In similar ways, independent academic and nonprofit researchers have used approval package data in combination with other data (from the medical literature and other sources) to conduct research on the diabetes drug rosiglitazone ("Avandia"), 247 the painkiller valdecoxib ("Bextra"), 248 and cosmetic injections of botulinum toxin (better known under the brand name Botox). 249 The FDA's data transparency has benefits for the agency's public credibility, as well. In November 2020, at a moment when the American public's trust in the FDA had been damaged by interference in its COVID-19 vaccine review process from then-President Trump and his political appointees, 250 the agency was able to restore some trust in the agency and in the vaccines themselves by committing to publish complete approval packages even as the agency was short-cutting other steps of the standard vaccine approval process in the emergency setting of a global pandemic. 251 Independent researchers dissected these approval packages once published and, by and large, confirmed COVID vaccines' safety and efficacy, and the wisdom of the FDA's decision to hurry them into patients' arms. 252

## Mandatory Submission and Publication of Clinical Trial Data to ClinicalTrials.gov

*p. 61*
A separate provision of FDAAA mandates that an even broader set of summary data and metadata must be shared with researchers via an independent means: ClinicalTrials.gov, a free and publicly accessible website administered by the NIH. 253 Regardless of whether a particular drug, vaccine, or device is approved or unapproved by the FDA, the results of Phase 2, 3, or 4 trials studying the drug or device in the United States must, by law, be published on ClinicalTrials.gov. 254 FDAAA's ClinicalTrials.gov mandate requires that the results of clinical trials be individually submitted to NIH by the companies, universities, and other entities ("responsible parties" per the statute) that run them.

*p. 62*
FDAAA is detailed and exacting. It specifies the precise summary data and metadata that responsible parties must submit to ClinicalTrials.gov and thereby disclose, data element by data element. 255 When FDAAA was being debated and implemented, many entities that conduct clinical trials protested that the statute's and subsequent rule's data elements were overly detailed, overly rigid, or unreasonably different from the idiosyncratic ways in which they formatted their own data. 256 However, the consistent, predictable format of summary data provided on ClinicalTrials.gov has helped independent researchers understand and use its data.

*p. 62*
The mandatory metadata-sharing provisions of FDAAA merit attention too, as they likewise help independent researchers contextualize trial results and perform useful research. FDAAA requires responsible parties to share detailed metadata: "[t]he full protocol or such information on the protocol for the trial as may be necessary to help to evaluate the results of the trial." 257 NIH has elaborated on this statutory provision with a rule specifying that responsible parties must also share their statistical analysis plans. 258 This mandatory sharing of metadata makes the summary data richer for researchers, and permits researchers to root out errors and manipulation.

*p. 62*
FDAAA's mandatory metadata-sharing requirement was fought by the pharmaceutical and medical device industries. As NIH observed when it promulgated the rule that implemented this provision of FDAAA, multiple commentators from relevant industries alleged that requiring disclosure of trial protocols would violate privacy and intellectual property interests: "Some asserted that protocols contain personally identifiable information, proprietary information, or other information that, if publicly disclosed, could be damaging to business interests." 259 The largest biotech industry lobbying group, the Biotechnology Innovation Organization (BIO), argued that NIH's commitment to sharing protocols (and summary data, too) "may undermine Reg., supra note 188, at 64,982, 65,006 ("While the Agency appreciates that accepting a variety of submission formats . . . may be less burdensome for responsible parties, [FDAAA] requires the final rule to establish a standard format for the submission of clinical trial information. This standard format will, in turn, facilitate search and comparison of entries in the registry data bank, as is also required under the statute.").

*p. 62*
257 incentives to innovate by forcing premature disclosure of proprietary information." 260 The largest medical device industry lobbying group, AdvaMed, echoed BIO and went further, threatening litigation over NIH's interference with its alleged trade secrets: [NIH's] disclosure of "trade secret and confidential commercial information" would constitute a taking in violation of the Fifth Amendment, AdvaMed stated. The device lobby group also asserted the disclosure of proprietary, confidential clinical trial data for products not approved would chill interest in developing new and innovative devices. 261 NIH proceeded anyway. However, in a concession to industry, NIH allows companies to redact portions of their trial protocols that they consider trade secrets before posting them to ClinicalTrials.gov, 262 "so long as the redaction does not include any specific information that is otherwise required to be submitted under" the law. 263 NIH held the line on summary data and, through rulemaking, extended FDAAA's disclosure mandate to reach experimental products not yet approved by the FDA. 264 NIH does not permit companies to redact any portion of their summary data, even if they fear competitors' use of the information. 265 Industry's threats of litigation proved hollow. NIH has never been sued by industry over its implementation of FDAAA. Nor have the FDA or the U.S. Department of Health and Human Services (HHS). The pharmaceutical and 260. Erin Durkin, Califf, Biden Task Force Tout NIH Rule Requiring Failed Trial Data be Posted, 22 INSIDEHEALTHPOLICY.COM'S FDA WEEK 11 (2016).

## Clinical Trials Registration and Results

*p. 63*
Information Submission, 81 Fed. Reg., supra note 188, at 64,982, 65,000 ("[I]f there is a case in which a responsible party believes that a protocol does contain trade secret and/or confidential commercial information, the responsible party may redact that information, so long as the redaction does not include any specific information that is otherwise required to be submitted under this rule.").

*p. 63*
263 265. Id. at 64,982, 64,996 ("A few commenters suggested that if the proposal is adopted, only a limited number of primary or key secondary outcomes prior to regulatory approval should be required to be submitted, or the final rule should allow the submission of redacted results information, especially when the product has not been approved, licensed, or cleared by FDA. The Agency disagrees; we believe that results information submission for all prespecified primary and secondary outcomes, as required in the statute, is necessary to serve the public interest in having access to full and complete information."). medical device industries have stopped criticizing FDAAA and quietly begun complying with its mandates.

*p. 64*
To be sure, compliance with FDAAA's ClinicalTrials.gov reporting rules is less than perfect: Independent analysis by "FDAAA Trials Tracker," a project of the Bennett Institute for Applied Data Science at Oxford University, suggests that only about 78% of trials with a legal obligation to comply with reporting rules have done so. 266 In addition, many trials that do report are late; in 2021, independent experts estimated that fewer than 50% of covered trials report results on time. 267 But this data sharing is meaningful, as much of this data is unavailable elsewhere. NIH's ClinicalTrials.gov has become the world's largest publicly accessible database of clinical trial data. 268 And ClinicalTrials.gov has proven the value of the clinical trial data sharing mandate. Since assuming its modern form in 2017, 269 ClinicalTrials.gov's vault of data has been used in a wide range of socially beneficial research. For example, a 2014 study compared data reported on ClinicalTrials.gov with data reported in medical literature and found that "nearly all had at least 1 discrepancy in the cohort, intervention, or results reported between the two sources." 270 This study underscored ongoing errors in and manipulation of medical literature (where data reporting is less standardized and, in some journals, less scrutinized than ClinicalTrials.gov). Researchers used ClinicalTrials.gov-primarily the metadata reported pursuant to FDAAA-to critique the proliferation of many small, relatively low-quality trials of COVID therapeutics in 2020 and early 2021. 271 Such critique helped to prompt the U.S. government to promise better coordination of government-funded trials. 272 Deborah Zarin, Director of ClinicalTrials.gov from 2005-2018, wrote in 2022, 273 [The ClinicalTrials.gov] database has been in existence since 2008, and has been continually updated and improved during that time. Thousands of responsible parties have used it to submit over 51,000 sets of results. Research has shown that about half of these-results for about 25,000 trials-are not available in the published literature, making ClinicalTrials.gov the unique public source of this information.

*p. 65*
Research into safety, efficacy, and the accuracy of companies' claims often complements the work of government regulators. Independent research critiques and ultimately reinforces the credibility and reliability of government regulators such as the FDA. This sort of research not only informs the public, but also actively checks and reshapes the regulatory process. For example, independent analysis of drug safety by the nonprofit organization Public Citizen, using data from ClinicalTrials.gov, Drugs@FDA, and other sources, helped convince the FDA to remove at least twenty-three dangerous drugs from the U.S. market, as of 2019. 274 Independent analysis of the clinical trial data that supported approval of Purdue Pharma's addictive oxycodone product, Oxycontin, and other opioid painkillers by drug regulators worldwide has underscored the paucity of evidence on addiction that regulators initially demanded, and has helped shape a present-day consensus that regulators must more carefully scrutinize new drugs for addictive potential. 275 In the past two years, independent analysis of the results of COVID-19 vaccines clinical trials has consistently corroborated the FDA's conclusion that the vaccines are safe, and helped to counter some of the hesitance and misinformation that have surrounded the vaccines. 276 The ClinicalTrials.gov database is free and accessible all over the world. 277 As such, it reduces longstanding inequities in access to trial data 278 and has catalyzed research not just in the United States but around the world. Some of the research conducted with ClinicalTrials.gov is conducted by researchers outside the United States. 279 Data from ClinicalTrials.gov has also been used to study the extent of research conducted in Global North-South collaboration. 280

## D. IMPLEMENTATION OF THE CLINICAL TRIAL DATA SHARING MANDATE AND EXPERIMENTATION WITH RESEARCHER ACCESS TO MORE SENSITIVE DATA

*p. 66*
This Section elaborates on FDAAA's data-sharing mandate in two important regards.

*p. 66*
First, this Section elaborates on implementation: How, exactly, does clinical trial data sharing work? For example, who enforces compliance with data-sharing mandates, and how? Because this Section focuses on implementation, it necessarily focuses on institutions. These institutions perform a number of important roles in the clinical trial data sharing ecosystem: they request or mandate submission of clinical trial data by industry, academia, and other sectors that perform clinical trial research; verify clinical trial data and hold it securely; mediate access to it; oversee uses by researchers; and monitor and enforce compliance with the laws that govern each of these steps.

*p. 67*
Second, this Section describes how some institutions have begun pioneering giving researchers access to more sensitive data. As traced in Section III.C, FDAAA's clinical trial data sharing mandate is limited to highreward, low-risk data: summary data and some metadata. The mandate does not reach IPD-the most sensitive data, from a privacy perspective-nor does it reach all information industry describes as its trade secrets. Yet, as we show, some institutions have pioneered mechanisms for sharing this data responsibly.

*p. 67*
A key theme is that institutional governance of medical data sharing is vital to the success of legal governance of the same. Law on paper is only modestly effective without associated institutions to implement, elaborate, and enforce that law. It is institutions-people-that get things done.

## Key Institutional Governors of the Clinical Trial Data Sharing Mandate:

*p. 67*
FDA and NIH FDAAA's results-sharing mandate did not effectuate itself; FDAAA requires two federal agencies, the FDA and NIH, to implement the legislation's data-sharing mandate, and govern access to and use of clinical trial data.

*p. 67*
The FDA, NIH, and other federal scientific agencies play a variety of important roles in managing not just clinical trial data but a wealth of other scientific and technical data. As Contreras observed, "the state's role in fostering innovation and scientific advancement is often analyzed in terms of incentives that the state may offer to private actors" such as tax credits, IP protections, direct grants, and provision of infrastructure. 281 Yet Contreras convincingly argues that this view is incomplete, at least in the fields of medicine and biotechnology. In the United States, the medical "innovation system" depends on the U.S. government not just as incentive-setter but as a central actor in the "information economy," managing data flows:

*p. 67*
The state plays a number of well-understood roles with respect to the planning, provisioning, and maintenance of publicly owned infrastructure resources such as highways, prisons, and public utilities. Likewise, the state is often involved in the oversight, regulation, and operation of private and public-private infrastructural resources such as airports and telecommunications networks. Why then should the same types of complementary and overlapping relationships not arise with respect to data resources that form an integral part of the research infrastructure? 282 Contreras maps nine distinct roles that U.S. government agencies play in the governance of medical data, writ large: (1) creator, (2) funder, (3) convenor, (4) collaborator, (5) endorser, (6) curator, (7) regulator, (8) enforcer, and (9) consumer. 283 In the world of clinical trial data sharing, the FDA and NIH play all nine roles, but in this Section, we focus on four overlapping roles we consider particularly important to the success of clinical trial data sharing: curator, funder, regulator, and enforcer. a) FDA and NIH Curate Data Institutions curate data by aggregating, hosting, and explaining data for other stakeholders to access and use. FDAAA mandates that NIH and the FDA play these curatorial roles: NIH with ClinicalTrials.gov, and FDA with Drugs@FDA. 284 NIH's National Library of Medicine (NLM) aggregates and hosts the massive ClinicalTrials.gov database. NLM also actively safeguards the quality, accuracy, and usability of each submission of clinical trial data. 285 NLM conducts an extensive quality control process to ensure that data is submitted to ClinicalTrials.gov completely and in the correct format. 286 NLM also maintains an elaborate "customer support" site and helpline for staff at universities, drug companies, and other institutions who encounter problems when preparing and submitting data to the database. 287 In this way, NLM protects the credibility and usability of the database.

*p. 68*
NLM has curated not just data submission but data use by researchers and the general public; it maintains an extensive Glossary and FAQ page to guide researchers through searching and interpreting the database. 288 NLM has also published research guides in the medical literature, detailing how to make effective use of ClinicalTrials.gov. 289 The FDA similarly aggregates, hosts, and explains the data it publishes on its own Drugs@FDA website in the form of the approval packages required by FDAAA. The FDA does not simply republish industry-submitted trial data, but also independently reviews the data and provides its own written critique and summary. 290 Like NLM, the FDA maintains a glossary 291 and FAQ 292 to help researchers use Drugs@FDA.

## b) FDA and NIH Fund Data-Sharing Initiatives and Research Itself

*p. 69*
The FDA and NIH serve separate roles as funders. They fund private initiatives to steward and share data, and they fund academic researchers who make socially beneficial uses of data. This role flows from law; Congress's appropriations bills earmark public money to the agencies for these very purposes. This role, too, explains the success of the clinical trial data sharing mandate.

*p. 69*
NIH is the world's largest medical research grant-maker, 293 and some of the billions disbursed go to researchers who use ClinicalTrials.gov in their research. 294 The FDA has formed multi-year partnerships with Johns Hopkins, Stanford, the University of Maryland, the Mayo Clinic, and Yale to study pharmaceutical and medical device regulation to scrutinize and improve the FDA's regulatory work. These government-academic initiatives are called Centers of Excellence in Regulatory Science and Innovation (CERSI). 295 Researchers funded by the FDA in this way have critiqued and improved the FDA's own work, e.g., by using FDA-published trial data and other data to question the use of "real-world evidence" in lieu of traditional clinical trials 296 and asking whether the FDA is sufficiently attentive to evidence of side effects gathered after drug approval. 297 The FDA has also experimented with funding academic institutions, nonprofits, and patient groups to become data-sharing platforms themselves. That is, the FDA has sponsored private institutions to aggregate and share certain clinical trial data. These initiatives include the Rare Disease Cures Accelerator-Data and Analytics Platform (RDCA-DAP). 298 Indeed, some other emerging "private" medical data-sharing initiatives led by patients, academia, and/or industry are funded partly with public resources; they do not always emerge entirely "organically" without the hand of the state. One such example is the Yale Open Data Access (YODA) Project, discussed more below. c) FDA and NIH Regulate and Enforce the Data Sharing Mandate Finally, we consider the roles of NIH and the FDA as regulators and enforcers of FDAAA's clinical trial data sharing mandate. NIH and the FDA force the pharmaceutical and medical device industries to share otherwise proprietary clinical trial data, consistent with FDAAA's mandate. Congress gave FDAAA "teeth" by specifying draconian potential consequences for failing to submit clinical trial results to ClinicalTrials.gov, including fines of over $10,000 per day per missing trial and a "freeze" on any grant money disbursed by NIH, the FDA, and other constituent agencies of HHS. 299 FDAAA also requires the FDA to name and shame responsible parties out of compliance with FDAAA's reporting rules, via public "Notices of Noncompliance" on a FDA-managed website crosslinked to ClinicalTrials.gov. 300 The FDA and NIH have performed poorly in their role as enforcers. Since FDAAA's enactment, the FDA's enforcement efforts have been almost laughably minimal: just five Notices of Noncompliance issued and zero fines imposed, despite thousands of trials out of compliance (among tens of thousands of trials with results required under FDAAA). 301 It was only in 2022 that NIH began sending letters threatening to withhold grant money from grantees out of compliance with FDAAA's data sharing mandate. 302 NIH and the FDA have been criticized from many sides for not doing more enforcement, including by researchers seeking access to missing data, 303 civil society groups, 304 journalists, 305 a former director of ClinicalTrials.gov, 306 HHS's Office of Inspector General, 307 and one of us. 308 Yet even the FDA and NIH's meager enforcement has contributed to a significant increase in data-sharing compliance rates. Since 2020, when the FDA first promised to begin issuing Notices of Noncompliance and threatened fines, 309 the percentage of applicable clinical trial results reported to the database rose from approximately 60-65% 310 to about 75-80%. 311 Even light-touch enforcement prompts compliance. A 2021 analysis showed that when the FDA simply sent a few dozen short letters to responsible parties, stating that the agency had reason to believe their trials might be out of compliance with FDAAA's data reporting rules, more than 90% of recipients provided the missing data with a median response time of just a few weeks. 312 And the present, C-grade state of enforcement and compliance with ClinicalTrials.gov's reporting mandate is nonetheless sufficient to unlock enormous benefits. 313 As former ClinicalTrials.gov Director Zarin wrote in 2022, there are approximately 25,000 trial results reported on ClinicalTrials.gov that are unreported in the medical literature, and thus presumably accessible to researchers nowhere but ClinicalTrials.gov. 314 Why such meager enforcement from the FDA and NIH? One major reason is that FDAAA imposed new regulatory obligations on both agencies without allocating new funding. 315 Both the FDA and NIH have many other obligations, and neither agency had strong incentives to dedicate personnel and attention to ClinicalTrials.gov. In addition, HHS's choice to divide enforcement responsibilities between the two agencies 316 rather than vesting responsibility entirely with one has made it easier for each agency to point to the other as the laggard.

## Pioneering Researcher Access to More Sensitive Data

*p. 73*
The entire clinical trial data sharing mandate described above requires sharing of just two components of clinical trial data: summary data and metadata. To recap, FDAAA mandates that summary data be disclosed without redaction. 317 It mandates that metadata be disclosed as well, 318 though NIH rules permits companies (and other trial sponsors) to redact information in trial protocols deemed a trade secret or confidential commercial information. 319 This means that FDAAA's clinical trial data sharing mandate does not reach IPD, the most detailed and most sensitive trial data. 320 The mandate also does not reach some metadata in trial protocols that companies deem trade secrets.

*p. 73*
Yet some institutions that share clinical trial data have pioneered ways to share sensitive information with independent researchers. These efforts show it is possible to navigate treacherous hazards to privacy and trade secrecy with careful institutional and legal design. a) Sharing IPD Sharing raw clinical trial data that describes, in detail, the health statuses of individual patients-IPD-poses profound risks to patient privacy. 321 As the Institute of Medicine put it in 2015, "privacy concerns have been stated as a key obstacle to making these data available." 322 Yet some kinds of research depend on IPD and cannot be done without it. For example, only researchers with access to IPD and the trial's complete methodology can conduct reanalysis to confirm the correctness of the trial sponsor's conclusions.

*p. 74*
Numerous institutions now share IPD with researchers, and do so responsibly. 323 Some of these databases are public-e.g., NIH's Biologic Specimen and Data Repositories Information Coordinating Center (BioLINCC). Other databases are nonprofit and academic-e.g., the Yale Open Data Access Project (YODA). Others are industry-run. 324 We describe these two IPD-sharing databases here. We do not attempt a comprehensive survey of IPD-sharing initiatives but instead present these as proofs-of-concept. Key features permit them to share sensitive data with researchers while protecting the data's integrity and the interests of the data subjects.

*p. 74*
As we trace below, a constant of these databases is that they are not universally accessible; they do not publish data for use by any and all comers. Instead, they discriminate among prospective users and provide access only to researchers deemed sufficiently responsible.

*p. 74*
Further, the institutions that manage these databases use legal and/or technological constraints to limit researchers' access to and use of the data, reducing the risk of harmful uses. Researchers components or kinds of data. All this underscores the vital role of institutions in clinical trial data sharing; these databases require active stewardship.

## i) NIH BioLINCC

*p. 75*
In addition to the enormous ClinicalTrials.gov database, NIH curates and controls smaller databases of clinical trial data. A notable one is BioLINCC, a database that contains sensitive IPD from clinical trials in cardiovascular, pulmonary, and hematological diseases. 325 BioLINCC has been in operation since the 2000s. 326 NIH created and administers the center, but much of the information contained in BioLINCC's databases is contributed not by NIH itself but by nongovernmental entities, including drug and device companies. 327 NIH requires these entities to submit data to BioLINCC as a condition of accepting NIH funding for their research. This straightforward quid pro quo leverages NIH's separate role as funder.

*p. 75*
Because BioLINCC data typically contains IPD, NIH shares data conditionally, limiting access and use. BioLINCC requires would-be researchers to submit data use applications, which document the intended uses of specific data sets (prospective researchers' "Research Plan"), data security practices, and commitments. NIH discriminates among users; NIH provides commercial users access only to a subset of BioLINCC's data and provides no access at all to would-be researchers that do not submit a credible Research Plan. 328 NIH then enforces researchers' compliance with their Research Plans through contract. NIH imposes a data use agreement on every researcher who obtains access to IPD from BioLINCC. The data use agreement governs transfer, maintenance, and use of protected data. The agreement imposes constraints on researchers, both positive (incentivizing users to do beneficial things) and negative (disincentivizing users from doing harmful things). BioLINCC's current standard agreement includes all the following: 329 Provisions that prohibit . . .

*p. 76*
• reidentification of or contact with any patient whose IPD is in the data set.

*p. 76*
• regular updates to NIH on the status of research;

*p. 76*
• notification to NIH in the event of data breach;

*p. 76*
• notification to NIH and the FDA in the event the data user identifies in the data an ongoing risk to public health and safety; • dissemination of any findings to the public, e.g., by publication in the peer-reviewed medical or scientific literature; and • destruction of data when research is complete.

*p. 76*
Data use agreements can specify penalties in the event a researcher breaches the agreement. These penalties can be financial or non-financial. BioLINCC's data use agreement does not contemplate financial penalties but does promise to ban breachers from any future access to data. 330 BioLINCC's information-sharing program has succeeded. Hundreds of requesters have sought and received access to thousands of data sets, leading to dozens of high-profile scientific and medical publications in cardiology, infectious disease, and other fields of medical research. 331 Over 250 articles were published based on BioLINCC data accessed between January 2000 and May 2016. 332 In practice, NIH's scrutiny and data use agreements seem to work. No researcher misuse of BioLINCC data covered by a data use agreement has been reported in the years of BioLINCC's existence.

*p. 77*
ii) Yale Open Data Access Project (YODA) Another prominent institution with a track record of successfully sharing IPD is YODA, a nonprofit academic data center that holds complete data sets (including IPD) on over 400 trials. 333 YODA is not the only non-governmental, not-for-profit institution that shares IPD. Two additional examples are Vivli and Project Data Sphere. 334 YODA operates similarly to NIH's BioLINCC. Like BioLINCC, YODA holds data on its own servers, gatekeeps requests for access to data, and enforces compliance with its own rules for data sharing and use. To get YODA data, researchers must establish that they have a credible research plan and proper security measures in place. 335 YODA refuses some applicants, especially when those applicants seek access to sensitive IPD. In difficult cases, YODA uses a peer-review-like process: it solicits reviews from two independent scientists to help decide whether to approve or deny applications. 336 Like BioLINCC, YODA imposes data use agreements on all researchers who get access to the data.

*p. 77*
YODA has convinced major medical technology companies-including Medtronic and Johnson & Johnson-to share, voluntarily, complete clinical trial data sets that would otherwise be proprietary. These companies benefit in various ways from contributing data to YODA, including a "halo effect" of good publicity and early access to scientific insights contributed by the researchers who use their data. 337 The companies that contribute data to YODA reserve their own rights to bring breach-of-contract claims against researchers who breach YODA's data use agreements.

*p. 78*
Though rather small, YODA has been a success thus far: between 2014 and 2018, Johnson & Johnson voluntarily shared data from 200 clinical trials through YODA, generating at least a dozen new scientific publications, 338 including analyses of the safety of ulcerative colitis treatments 339 and the efficacy of schizophrenia drugs (which critiqued exaggerated claims made in the medical literature). 340 All this occurred without evidence of privacy violations, breaches of the data use agreements, or harmful use of data by Johnson & Johnson's competitors. 341 YODA operates on a mixture of grants provided by industry (Medtronic and Johnson & Johnson), philanthropy, and government. The FDA and the Centers for Medicare & Medicaid Services (CMS) have both funded YODA, showing the role public money and institutions can play in nurturing private governors of data. 342

## b) Sharing Metadata That Contains Alleged Trade Secrets

*p. 78*
In this Section III.D.2.b, we turn to an institution that has pioneered responsible sharing of (purported) trade secret data with researchers: Health Canada, Canada's central drug regulator.

*p. 78*
Since 2019, Health Canada has shared rich data sets from clinical trials of agency-approved products, under a program called Public Release of Clinical accessible to routine users of PRCI. 347 Users who wish to access and use these redacted data sets may do so with few restrictions, much like Drugs@FDA and ClinicalTrials.gov.

*p. 80*
Yet Health Canada shares even more information with select researchers, including unredacted trade secrets. According to Paragraph 21.1(3)(c) of the Canadian Food and Drugs Act, 348 Health Canada will share trade secrets (CBI) on certain conditions. First, researchers must submit a data use application that proves their proposed use is noncommercial and relates to "protection or promotion of human health or the safety of the public." 349 Second, the application must also explain "[h]ow the results of the proposed project will be disseminated to the Canadian public." 350 Any researchers granted access must then sign data use agreements insisting "the specified CBI can be used only for the purposes of the proposed project and must be kept confidential using appropriate safeguards." 351 In the event a researcher detects a safety, efficacy, or quality problem in the data, Health Canada requests the researcher notify Health Canada as well as the public at large. 352 In 2016, a medical researcher, Peter Doshi, used Paragraph 21.1(3)(c) to obtain detailed, previously secret data on the safety and efficacy of several medical products, including oseltamivir ("Tamiflu") and vaccines for human papillomavirus (HPV). 353 Doshi's access to this CBI-and his legal authority to disseminate analysis of it-was upheld by the Canadian Federal Court. 354 Doshi has not made inappropriate use of the data, and industry has not subsequently sued Health Canada to block similar disclosures.

## E. CLINICAL TRIAL DATA IN ACTION: A RECAP

*p. 81*
Perhaps the single most important lesson of Part III is that clinical trial data sharing works. Today's clinical trial data sharing mandate guarantees researchers meaningful access to components of clinical trial data that the R&D-driven pharmaceutical industry kept proprietary for decades. The mandate has fostered beneficial research that could not have occurred otherwise, some of which has challenged industries' overblown claims and improved the FDA's regulation. Indeed, the mandate seems to have contributed to a "new normal" of improved drug safety; in the years since FDAAA was enacted, we have not had scandals of unsafe products and manufacturer cover-ups on the level of Paxil or Vioxx. 355 The pharmaceutical and medical device industries resisted clinical trial data sharing on the argument that sharing would harm privacy and incentives to innovate. But so far, clinical trial data sharing has capably protected those interests.

*p. 81*
The clinical trial data sharing mandate emerged over years, not overnight, and remains a work in progress. Key to the mandate's qualified success are the institutions that give ongoing effect to its underlying law, especially FDAAA. Law cannot simply proscribe or prescribe behavior, nor can it reallocate power with the stroke of a pen. In our view, law must also create and nurture institutions to give law meaning and teeth. For the clinical trial data sharing mandate, the key institutions are the FDA and NIH, but they are surrounded by an array of other institutions, some private and some independent but government-funded.

*p. 81*
Another key to the success of clinical trial data sharing, in our view, has been the recognition that different components of clinical trial data deserve different treatment. Clinical trial summary data and most metadata are low risk and high reward; they can be shared freely with users without restrictions on access and use. A small fraction of metadata may implicate trade secrecy, but such data can be shared carefully; data use agreements and other constraints preventing competitive use can protect innovative companies' first-mover advantages. Sharing IPD poses profound privacy risks, but IPD too can be 355. That is not to say that the pharmaceutical and medical device industries, or the FDA, have had a perfect track record since 2007. shared responsibly with some users, subject to appropriate institutional and technical constraints.

## IV. TOWARD A SOCIAL MEDIA DATA SHARING MANDATE

*p. 82*
Part IV applies some of the primary lessons learned from clinical trial data sharing and charts a course toward responsible and effective social media data sharing. Section IV.A focuses on how the benefits of independent research cascade, emerge, and are unpredictable at the outset. Section IV.B focuses on the need for regulators. Here we use the term "regulators" to refer to both public and private entities that can impose accountability and exert countervailing power over social media companies by providing alternative forms of expertise, employment, and perspectives. Section IV.C, drawing from the concept of contextual integrity, transposes many of clinical trial data sharing's solutions for navigating the Scylla and Charybdis of trade secrecy and privacy. These solutions apply context-specific controls over social media data to treat contextually and normatively distinct kinds of data differently, using tiered access and a variety of constraints on data access and use tailored to the goals and needs of particular applications.

*p. 82*
In our view, clinical trial data sharing's hybrid, "both and" approaches are successful. Various clinical trial data sharing initiatives deploy a mix of mandated sharing and voluntary arrangements, across data types of varying sensitivity, in order to balance the interests of commercial secrecy, individual privacy, and public benefits of research.

*p. 82*
Clinical trial data sharing also shows that meaningful independent researcher access cannot be achieved without laws mandating that industry share more data. Clinical trial data's journey from the dark ages to today's robust ecosystem was made possible by the legal transformation of the rights in such data. What began as data governed almost exclusively by private ordering eventually incorporated public demands to constrain those interests and indexed a public right to quality research to provide accountability to a high-stakes sphere of life. Clinical trial data's iterative process of legislation and regulation to enact and build on that legislation was the legal foundation needed to build a robust data sharing ecosystem.

*p. 82*
Finally, the example of clinical trial data shows the importance of ensuring that data access mandates do not operate as mere transparency requirements. Laws to grant researcher access must materially and legally empower regulators to avoid this pitfall.

*p. 82*
Data access mandates that allow companies to retain either discretionary control over who is granted access or financial control over how the work of access is funded do more harm than good. At best, such proposals will empower a subset of well-connected and resourced researchers through narrow interpretations of such rules. At worst, such proposals may weaken pressure to impose more substantive regulation over the digital economy.

*p. 83*
We do not believe transparency alone can provide a sufficient solution to the larger issues surveyed above in the social media research ecosystem, as the case of SS1 amply demonstrates. As AI Now noted in its 2023 annual report, data access regulation alone is not enough to promote a stronger and more robust independent researcher ecosystem. 356 A.

*p. 83*
CASCADING (AND UNPREDICTABLE) BENEFITS OF BASIC RESEARCH One lesson of clinical trials is that the benefits of basic research are not always obvious before research begins. Benefits are instead cascading and unpredictable. Just because these benefits are not readily apparent at the time access is granted does not mean such benefits will not be significant. (And to be clear, in the case of social media data, many pressing societal benefits for researcher access are already readily apparent, as we have argued in Part II).

*p. 83*
Basic research is infrastructural. It is the first step in the process of refining unknown unknowns into known unknowns or known knowns. 357 Basic research provides the scientific building blocks upon which many other forms of research and productive innovation rely. At the outset, the cascading, indirect benefits of basic research are near-impossible to predict because the stuff of value being built on or adapted for commercial use-a useful material or a surprising scientific breakthrough-is not even known to exist at the time. 358 It seems obvious to say, but discovery of the previously unknown is the point of basic research.

*p. 83*
The value and unpredictability of discovery are important to emphasize when weighing the potential benefits of researcher access against claims of the risks to secrecy and privacy. Addressing direct, currently known needs are just one of the emergent beneficial properties of the new institutions that will be created to facilitate social media access.

*p. 83*
As Part III showed, researchers' access to clinical trial data has led to many cascading benefits: illumination of harms that regulators missed, improved patient care and public health, higher quality trials, combating misinformation, and more. Nonprofit and broadly accessible clinical trial databases, including ClinicalTrials.gov, Drugs@FDA, and BioLINCC expand and democratize access to scientific data.

*p. 84*
Reliable and growing access to clinical trial data has also helped to create a cadre of independent researchers able to use that data. Grants from NIH and FDA have contributed to a corps of independent experts able to manage and use this data for public benefit. This material independence in turn has fostered a larger ecosystem of expertise and knowledge production that exists outside of-and largely independent of-the pharmaceutical and medical device industries.

*p. 84*
Independent access to social media data, done right, can also empower a greater diversity of researchers with the tools to access this data, and thus conduct scientific research with this resource. Because researchers will no longer need to rely on individual, bespoke relationships with companies, or be willing to assume the legal risk of proceeding without such relationships in place, it is reasonable to assume that greater numbers of researchers from less well-resourced institutions will be able to gain access to social media data. The same goes for researchers that may be interested in U.S. social media data but reside outside of the United States-making this data available to qualified researchers opens up access to a global research community. Indeed, we have already seen a similar benefit of the European Union's recent efforts to grant researchers access to E.U. data; many U.S. researchers are extremely enthusiastic about the research potential of accessing E.U. data. 359 Robust ecosystems of researcher data access take time to develop. They cannot be achieved in a day. Nevertheless, achieving a successful state of social media data access depends in part on the steps taken now. The cascading benefits of clinical trial data have taken years to realize and are still emerging. We are only at the very beginning of the process of implementing researcher access to social media data, and whether the process realizes its potential depends on the steps taken today.

## B. EMPOWERING REGULATORS

*p. 84*
To be successful, researcher access laws and policies must create and empower institutions, inside and outside government, with the funding, mandate, and expertise to manage the technical governance mechanisms of research data and to keep social media companies in compliance with existing law and accountable if they are not.

*p. 85*
Legislation to require access and prescribe certain data practices is an important first step. But to produce real results, the experience of clinical trial data sharing suggests that laws also need to empower regulators to engage in the day-to-day work of both keeping social media companies compliant with data sharing requirements and managing the technical governance mechanisms of access.

*p. 85*
Empowerment of such regulators means a few different things, and it can take a range of forms. Below we offer a menu of options, a mix of which have been successfully deployed in the clinical trial data setting. Given the early days of social media data sharing, we endorse experimentation, hybridization, and pluralism in approach among the options surveyed below. But the key lesson behind all these options is that social media platforms should not retain gatekeeping (or funding) authority over who is granted access to data, what studies are deemed fundable or feasible, or which results may be published.

## Independent, Preferably Public, Funding

*p. 85*
First, empowered regulators must have access to secure, reliable public funding. Currently, much of the funding (directly or indirectly) for researcher access to social media data is provided by companies themselves. This leaves researchers vulnerable to changes in market forces or company priorities. 360 It also produces a chilling effect on research considered overly critical. It is neither a sustainable model on which to build long-term access nor conducive to robust independent research.

*p. 85*
As seen in Section III.D.2.a, public funding does not have to mean servers running under direct government control. Government agencies can and do fund several different institutional models of data curation and sharing. NIH directly funds, manages, and hosts its own databases, including ClinicalTrials.gov and BioLINCC. But the FDA and NIH also provide funding to private data stewards, including YODA. Recipients of public funding can be other public institutions (like public universities or research consortia), private academic or non-profit research institutions, or clusters of all the above (similar to CERSI).

*p. 85*
Access mandates that both empower public and civil society institutions with independent funding and foster non-industry expertise in managing and providing access to such data can build these communities' material and intellectual capacity to do their work. Researcher access done right can thus play a key role in fostering the growth of meaningful regulators in the digital economy. As Part III shows, such institutions can play key roles in movement and coalition building. Free from material dependency on the companies, independent technology research ecosystems can provide the intellectual and civic seeds of the broad political mobilization needed to transform how we develop and manage the digital infrastructures of social and public life.

## Control Over Standards and Terms of Access and Use

*p. 86*
Second, empowered regulators are those that have meaningful control over (1) standards and processes of data sharing and (2) researchers' data access and use. Control over the standards and processes of data sharing means regulators must curate and safeguard data by protecting its quality, accuracy, and useability. Control over researchers' access and use means just that. Control can be effectuated through technical means, contracts (data use agreements), guides and protocols for use, and more.

*p. 86*
Regulators can exert control via a range of options that empower them in their relationships with both companies and researchers. At its most simple and direct, institutional control begins with laws that require companies to share certain data with regulators, as seen with ClinicalTrials.gov and in the Canadian example of trusted researcher access in Sections III.C.2 and III.D.2.b. We believe some degree of compulsory data sharing is required to foster successful, independent research. However, as Section III.D more broadly shows, voluntary forms of sharing can supplement mandatory forms, expand the universe of data made available to researchers, and build on their success. As Part III also shows (particularly in Section III.C.2) and as will be discussed below, when companies do not provide the data they are required to share, regulators should also be empowered to enforce sharing requirements.

*p. 86*
Importantly, institutional control also means data stewards should be tasked with administering researcher access and use of data to ensure researchers comply with necessary controls and safeguards.

*p. 86*
The destination of compelled data can be a government curator, as is the case with ClinicalTrials.gov. This approach is particularly promising for managing datasets on features shared across social media companies, like active users, volume of activity, distribution of that activity, language, and country of origin.

*p. 86*
However, curators need not be government entities. In the United States, the FDA funded RDCA-DAP and YODA, two exemplary non-governmental data sharing platforms. Non-governmental options may be particularly attractive for data that is more sensitive to privacy concerns that militate against permitting government agencies the capacity to hold, see, or use such data. Regardless of whether institutions are public or private, they should be given the means to manage data responsibly. This means funding to keep servers running and curatorial experts employed. This also means: legal rights to determine how data is to be shared from companies; rights to curate and assess data for quality; and rights to set the terms (and/or manage the process) of screening applicants for access via their own data use agreements. Curatorial institutions ought to have the rights to hold data on their own servers, serve as gatekeepers for access to data, and develop internal protocols for screening and evaluating researcher access proposals, including peer-review mechanisms for access to particularly sensitive data.

## Meaningful Regulatory Enforcement

*p. 87*
Part III also highlights the importance of meaningful enforcement of data sharing mandates to ensure compliance. The experience of ClinicalTrials.gov presented in Section III.C.1 suggests both that some enforcement is necessary and that even minimal enforcement through "naming and shaming" a handful of noncompliant entities can spur significant compliance. 361 One condition of granting private entities data curation roles might be a requirement to regularly report noncompliance to the relevant public regulator. Public data stewards and regulators can be given the capacity to enforce compliance directly via mechanisms like naming and shaming, imposing fines, or a court-enforceable right of action to compel access, to name a few. If public stewards lack authority to enforce the law themselves, then they should at least be able to highlight non-compliance to the public and the relevant regulator.

*p. 87*
The experience of clinical trial data sharing shows the modest but meaningful effectiveness of simple "naming and shaming" companies and other entities that withhold data from researchers despite a mandate to share. For instance, the FDAAA Trials Tracker, built by the Bennett Institute for Applied Data Science at Oxford, keeps track of which companies and clinical trials have shared their results as required under FDAAA. 362 For social media, regulation can help remove barriers to third-party development of similar accountability mechanisms.

## C. TREATING DIFFERENT DATA DIFFERENTLY

*p. 87*
Existing models of clinical trial data sharing show that it is possible to share data with researchers while also protecting data subjects from harm and preserving incentives to innovate. Clinical trial data sharing offers lessons 361. See supra Section III.D.1.c, on public institutional governors as regulators and enforcers.

*p. 87*
362. As they say on their website, "The FDA are not publicly tracking compliance. So we are, here." FDAAA Trials Tracker, BENNET INST. FOR APPLIED DATA SCI., https:// fdaaa.trialstracker.net/rankings/ (last visited Jan. 28, 2023).

*p. 88*
about the design of both the technology and the law. In both domains, the clinical trial sector has developed data-sharing mechanisms that are specific, contextual, and allow researchers to access useful data while remaining independent.

*p. 88*
Valid privacy and trade secrecy concerns should be treated with a scalpel, not a broadsword. In order to do this, data sharing mechanisms need to be tailored to the affordances of the data they offer and the risks posed by that data to data subjects, researchers, and platforms. This basic insight is not new. Scholars including Helen Nissenbaum, Dan Solove, and Neil Richards have argued for some time that theories and applications of information privacy need to be attentive to the contextually specific purposes and norms that both motivate and constrain information sharing. 363 However, Section III.D.2 shows how the legal, institutional, and technological responses that structured the still-evolving clinical trial data governance regime paralleled-perhaps even prefigured-these theoretical developments in information privacy law. Different tiers and mechanisms of access for different kinds of clinical trial data, users, and uses gradually emerged in response to live policy considerations of how to balance the risks to commercial secrecy and privacy with the social benefits of access. In other words, the solutions that emerged in clinical trial sharing look quite similar to what information privacy theorists have long observed and recommended for digital personal information subject to privacy and other concerns. This Section, IV.C, transposes many of clinical trial data sharing's solutions for navigating the twin barriers of trade secrecy and privacy. In line with existing theories of privacy law, these apply context-specific controls over social media data to treat contextually and normatively distinct kinds of data differently.

*p. 88*
To this end, we argue that, as an initial matter, social media should adopt clinical trial data's useful tripartite distinction of data types: individual data, summary data, and metadata. Social media companies tend to lump all these types of data together, raising the lowest common denominator of necessary protection. In other words, all data gets treated with the privacy and security sensitivity of individual data and the trade secrecy sensitivity of metadata, even though certain data-especially summary data-could easily be shared that does not raise those concerns.

*p. 88*
When resisting sharing data with researchers, social media companies by and large focus on the promises and pitfalls associated with sharing data about individuals' social media activity. This is evident in their most common methods, such as APIs and static data sets, and large data sharing initiatives such as SS1. Yet the same companies provide little information on how these data are generated (metadata) or aggregate data on their users and their activity (summary data).

*p. 89*
Without metadata, there are looming questions about the provenance and representativeness of data available to researchers. Without metadata, researchers must trust companies to have answered these questions in their own undocumented methodologies, despite evidence that some of these companies unreliable and unrepresentative data before. 364 For instance, Facebook's Ad Library comes in part from ads that the company's automatic detection algorithm flags as political. 365 However, Facebook does not offer any metadata on what classifiers it uses. Therefore, some entire topics may not be included in the library, and researchers would have no idea.

*p. 89*
Without summary data, researchers face difficulty contextualizing their results (e.g., understanding relative effect size) and verifying the numbers they receive from companies. For instance, researchers did not know that nearly half of all data was missing from SS1, or that so many advertisements were mislabeled on Facebook's Ad Library (before the NYU Ad Observatory uncovered it) because it was not possible to see if the numbers made sense.

*p. 89*
The minimal metadata and summary data that social media companies do currently provide to researchers lacks the requisite methodological clarity and specificity to be useful. Instagram, for instance, shares some information about how it ranks posts for users' feeds or explore pages, but the information provided is too general to be used in academic research. 366 The primary method companies use to share summary data is content moderation transparency reports, but these contain little information beyond how much content governments have requested be taken down and how often the platform complied. 367 Social media companies keep secret even basic platform usage information such as monthly active users and volume of uploads. For instance, the public learned that Instagram passed two billion monthly active users only when journalists leaked the information. 368 Existing legal frameworks do little better. Proposed laws in the United States and passed laws in the European Union almost always focus on access to individual data, rather than summary data and metadata, and in turn, impose severe limitations to maintain privacy and trade secrecy. The Platform Accountability and Transparency Act, for instance, mostly focuses on sharing individual data with researchers, particularly high-profile users and content moderation actions taken against them. The Ad Transparency Act also focuses on individual ads instead of requiring companies to describe underlying ad targeting systems. And while the Digital Services Act in theory allows researchers to access all three types of data, this data is only available to certain vetted researchers. 369 Below, we elaborate how not just individual data but summary and metadata on social media could be made available to researchers, and how access could be tailored to accommodate the privacy and trade secrecy considerations of each.

## Summary Data

*p. 90*
Summary data can be used by researchers to better understand who, how, and how many people use social media, while posing little trade secrecy or privacy risk. High level metrics (e.g., number of users, frequency of posts, or time spent on platform) broken down into certain categories (e.g., language or country of origin) can contextualize research and guide directions of future research. And if those categories are standardized, researchers can make comparisons across platforms. Summary data can also reveal self-sorted categories based on individual data, such as how many people use a given hashtag or remix a certain sound clip. For clinical trials, it took years of regulatory battles and clarification to get pharmaceutical and medical device companies to share summary data, but the resulting data sharing paradigm directly benefited the public, including by revealing discrepancies between Metadata does pose some real privacy risks. Metadata for social media encompasses a broader range of forms than metadata for clinical trials, and some social media metadata may reveal things about individual users, such as information on the users that initially posted content banned or restricted by a social media platform.

*p. 92*
The trade secrecy risks posed to social media companies by sharing metadata likewise vary along on a sliding scale. Divulging methods of how summary data-i.e., statistics on hashtags-get generated is on the low-risk end of the spectrum, as is divulging the methods by which individual data gets produced and organized. Moderation and recommender systems pose greater risk to trade secrecy interests, as does information on systems for evaluating whether features should be rolled out. Metadata on how ad targeting systems work is perhaps still higher risk, as these ad targeting systems are currently social media platforms' main drivers of revenue. This sliding scale moves slowly from what is clearly data about data to what is data about how larger systems work. As such, it becomes harder to fit clearly into the category metadata and moves further from the factual parallelism of medical data.

*p. 92*
We expect that controlled sharing of metadata from social media companies will yield real public benefits, broadly similar to those achieved by sharing metadata from clinical trials. With clinical trials, for instance, data sharing revealed limitations-even profound problems-with Tamiflu, Paxil, and Vioxx, but improved trust in certain COVID-19 vaccines. Similarly, social media metadata could be used to reveal the harms of some systems, but also to bolster public trust of others.

## Individual Data

*p. 92*
The concerns with individual data are a mirror of the concerns of those with summary data: they are not very likely to implicate trade secrecy concerns but can raise privacy concerns on a sliding scale from moderate to severe. And again, the tactic to manage this variance is to treat different data differently. Clinical trial data sharing initiatives do this very effectively. Clinical trial IPD is made available through tiered, tightly controlled access systems such as BioLINCC and YODA. The level of access provided to researchers and the sorts of research permitted depends on the data, the researchers, the intended research, and the associated privacy risks. More than two tiers of researcher access can exist, beyond one tier for "trusted researchers" and another for the broad public. The tailored access that YODA and BioLINCC provide useful models here. 374 374. Infra Section II.D.

## 2024] RESEARCHER ACCESS TO SOCIAL MEDIA DATA 201

*p. 93*
Some social media companies already tier data access, corroborating the notion that it can be done. For example, when Twitter offered its public facing API it had regular, enterprise, and academic versions. 375 Facebook has some data it shares publicly and other data it shares with those who sign an agreement, including now the data from SS1. 376 The experience of clinical trial data sharing shows that platforms can share more individual data than they already do, and that the stewards of that data can be trusted actors outside of social media companies themselves. Social media companies could, through tiered access data sharing programs, share some of the most sensitive social media data with trusted researchers who commit to avoid harmful uses. This sensitive data includes complete lists of removed posts, individual ad targeting information, and inferred data. Some of the most sensitive social media data that poses the greatest privacy risks, such as personally identifiable information and direct messages, may remain off-limits to even the most trusted researchers.

## V. CONCLUSION

*p. 93*
Social media is in its data secrecy dark age, just as pharmaceuticals were in previous decades. 377 This Article has traced parallels between clinical trials' past and social media's present. For instance, both have witnessed high-profile crises caused by a lack of accountability and transparency: for clinical trials, Paxil's teen suicides and Vioxx's heart failures; for social media, Cambridge Analytica, the rise of online populism, and the degradation of truth in media and democracy. Just as intrepid health journalists in the 1990s and 2000s used the limited tools they had to shine a light on the shadowy pharmaceutical industry, so too have tech journalists and social media company whistleblowers bravely revealed some of the public consequences of surveillance capitalism and the attention economy. Pharmaceutical, medical device, and social media companies have all adopted similar tactics to appease or deflect popular demand for more information, including limited, cherry-picked "transparency" efforts.

*p. 93*
In the past few years, a rash of new federal laws have been proposed that would mandate social media companies to share data with researchers-and, perhaps, bring in the light sufficient to end these dark ages. The Platform Accountability and Transparency Act, for instance, would empower the FTC 375. Adam Torres, Enabling The Future of Academic Research with the Twitter API, X DEVELOPER PLATFORM (Jan. 26, 2021), https://developer.twitter.com/en/blog/product-news/2021/enabling-the-future-of-academic-research-with-the-twitter-api.

*p. 94*
to compel social media companies to share data with qualified researchers approved by the National Science Foundation. 378 The Social Media Data Act proposes requiring platforms to create in-depth ad libraries for academic researchers. 379 Other proposed U.S. laws such as the Kids Online Safety Act, the Digital Services Oversight and Safety Act, and the ACCESS Act could also allow researchers to access social media data in other ways. 380 As of this writing, none of these proposals have become law. They remain the subject of intense debate, even controversy. Social media companies have fought them, just as pharmaceutical and medical device companies fought the legislation that mandates transparency of their clinical trial data. As if on cue, social media companies have invoked privacy and trade secrecy-this Article's "Scylla and Charybdis"-as doctrinal and normative reasons to oppose these proposals. 381 This Article has argued it is possible for legislation and regulation to protect privacy and trade secrecy while simultaneously mandating and mediating researcher access to sensitive data. The precedent of clinical trial data sharing reveals both some pitfalls that await lawmakers seeking to create an effective social media data sharing mandate and some paths to avoid them. Even when clinical data sharing rules were enacted into law, it took years of rulemaking, enforcement, and public pressure to get pharmaceutical companies to actually share their data. And though those battles continue today, the fight has produced safer medical products. For those regulating social media in the United States, the history of sharing clinical trial data shows that merely requiring data access, as legislative proposals do now, is necessary but not sufficient: law also needs to empower regulatory institutions that can enforce those laws and tailor data sharing systems to narrowly manage the privacy and trade secrecy risks that accompany each data type.

*p. 94*
In Part IV, we have done our best to distill useful lessons for governance of social media. Undoubtedly many readers will disagree that these are the right lessons. We hope, at very least, that the "thick" accounts of the need for researcher access to social media data and the history of clinical trial data sharing offered in Parts II and III inspire readers to make their own comparisons and derive their own lessons.

## Footnotes

> See discussion of the Digital Services Act's mandated access for vetted researchers in the Introduction, supra.

> See, e.g., Calma, supra note 115.

> TECH POL'Y PRESS (Apr. 26, 2023), https://techpolicy.press/tools-for-platform-researchlessons-from-the-medical-research-industry/. 32. See infra Part II, especially Section II.C through Section II.E.

> 105. Jeremy B. Merrill & Ariana Tobin, Facebook Moves to Block Ad Transparency Tools--Including Ours, PROPUBLICA (Jan. 28, 2019), https://www.propublica.org/article/facebookblocks-ad-transparency-tools. 106. Mike Clark, Research Cannot be the Justification for Compromising People's Privacy, 109. Testimony of Laura Edelson, NYU Cybersecurity for Democracy, Before the Subcomm. on Investigations & Oversight of the H. Comm. on Sci., Space, and Tech., 117th Cong. 2 (2021), https://docs.house.gov/meetings/SY/SY21/20210928/114064/HHRG-117-SY21-Wstate-EdelsonL-20210928.pdf. 110. Laura Edelson, Tobias Lauinger & Damon McCoy, A Security Analysis of the Facebook Ad Library, 2020 IEEE SYMP. ON SEC. & PRIV. (SP) 661, 667 (2020).

> developer-issues-apps. 116. See Graph API Reference, META FOR DEVELOPERS, https://developers. https:// developers.facebook.com/docs/graph-api/changelog/version3.0#gapi-90 (last visited Aug. 8. 2023). 117. Fred Morstatter, Jürgen Pfeffer & Huan Liu, When is it Biased? Assessing the Representativeness of Twitter's Streaming API, WWW '14 COMPANION: PROCS. 23 RD INT'L CONF. ON WORLD WIDE WEB 555 (2014).

> 119. See supra Section II.B.1. 120. DAVID H. FREEDMAN, WRONG: WHY EXPERTS* KEEP FAILING US-AND HOW TO KNOW WHEN NOT TO TRUST THEM (2010). 121. E.g., Tromble, supra note 16; Michael Zimmer & Nicholas Proferes, A Topology of Twitter Research: Disciplines, Methods, and Ethics, 66 ASLIB J. INFO. MGMT. 250 (2014). 122. Nicolas Kayser-Bril, Under the Twitter Streetlight: How Data Scarcity Distorts Research, ALGORITHM WATCH, https://algorithmwatch.org/en/data-access-researchers-left-on-read/. 123. SHAPIRO ET AL., supra note 3, at 24-26. 124. Id. at 46. [Vol. 39:109 2. Unrealized Research a) Inability to Evaluate Social Media Claims

> 150. Id. COPPA applies both to services that are "directed to children" under 13, such as children's online games, and those that knowingly collect personal information from people under 13. Platforms look to avoid charges of "actual knowledge" under COPPA by requiring users to input a birthdate on their registration page, and disallowing any user that responds with a year that suggests they are under 13.151. See 15 U.S.C. § 45(a)(1) (2018) (prohibiting "unfair or deceptive acts or practices in or affecting commerce"). All states have incorporated similar consumer protection clauses into

> 152. Van Loo, supra note 8. 153. See supra Section II.C. 154. Letter from Marc Rotenberg, EPIC President, Christine Bannan, EPIC Administrative Law and Policy Fellow, Sunny Kang, EPIC International Consumer Council, and Sam Lester, EPIC Consumer Privacy Fellow, to Gary King and Nathaniel Persily, ELEC.

> 168. See John P.A. Ioannidis, Arthur L Caplan & Rafael Dal-Ré, Outcome Reporting Bias in Clinical Trials: Why Monitoring Matters, 2017 BMJ 356 (2017) (describing value of comparing published trial results against trial protocols). 169. Linda Martin, Melissa Hutchens, Conrad Hawkins & Alaina Radnov, How Much Do Clinical Trials Cost?, 16 NATURE REVS. DRUG DISCOVERY 381 (2017); Thomas J. Moore, James Heyward, Gerard Anderson & G. Caleb Alexander, Variation in the Estimated Costs of Pivotal

> 172. INSTITUTE OF MEDICINE, supra note 161, at 32 (citations omitted); see also NATIONAL ACADEMIES OF SCIENCES., ENG'RS & MED., REFLECTIONS ON SHARING CLINICAL TRIAL DATA: CHALLENGES AND A WAY FORWARD (2020).

> documents/regulatory-procedural-guideline/external-guidance-implementation-europeanmedicines-agency-policy-publication-clinical-data_en-3.pdf.

> 186. See infra Sections III.B & III.C.1. NIH did not propose, and has not proposed, mandatory sharing of IPD from the same broad swath of clinical trials. 187. Letter to Jerry Moore, supra note 25, at 2. 188. See, e.g., 81 Fed. Reg. 64982, 64968, 64995.

> 196. See Silverman & Lee, supra note 189, at 241 (recounting that the then-FDA Commissioner, appointed in 1969, "urged . . . that the results of all animal and human trials and similar clinical data should be made public"). 197. Robert M. Halperin, FDA Disclosure of Safety and Effectiveness Data: A Legal and Policy Analysis, 1979 DUKE L.J. 286, 294 (1979); Thomas O. McGarity & Sidney A. Shapiro, The

> 204. O'Reilly, supra note 203; Fisher, supra note 203.

> FOOD & DRUGS L.J. 579, 581 (2008) 238. Also known as an "action package."

> . 42 U.S.C. § 282 (j)(3)(D)(iii)(III). A trial's protocol is "[t]he written description of a clinical study. 258. 42 C.F.R. § 11.48(a)(5). 259. Clinical Trials Registration and Results Information Submission, 81 Fed. Reg., supra note 188, at 64,982, 65,000.
