Engelberg Center mark Engelberg Center on Innovation Law & Policy Corpus

Law and the Science of Networks: An Overview and an Application to the 'Patent Explosion,'

Katherine J. Strandburg, Gábor Csárdi, Jan Tobochnik, Péter Érdi, László Zalányi
Articles
"Law and the Science of Networks: An Overview and an Application to the 'Patent Explosion,'" 21 Berkeley Tech. L.J. 1293 (2006) (with Gábor Csárdi, László Zalányi, Jan Tobochnik and Péter Érdi)
Abstract: The article presents an overview on the evolution and structure of technical relationships between patents in the U.S. patent system. It argues that the legal scholars should engage on network science trend because it provides significant conceptual advances and analytical tools. It illustrates the application of a network approach to empirical data by identifying the result of a network science study of patent acknowledgement.
This is an author copy made available for research purposes. Publisher version →

INTRODUCTION

p. 2

Networks are powerful tools for understanding the interconnectedness of the modern era. The internet no doubt motivates our fascination with networks, as it not only ties us together technologically in ways that were not previously possible, but also leads us to conceptualize the world in a This Article seeks to accomplish two tasks with respect to the application of network science to the study of law. First, in Part II, it seeks to persuade readers that legal scholars should jump on the network bandwagon in greater numbers because of the important conceptual advances and analytical tools that network science provides. Specifically, network science highlights the potential importance of heterogeneity and local relationship patterns in determining collective behavior. Part II also underscores and begins to address the difficulty of predicting collective behavior from individual interactions or transactions. As the science of networks continues to develop, it promises to provide new approaches to the analysis of empirical data and new tools for modeling the expected results of legal For example, a LEXIS search for law review articles citing either of the authors of the most well-known popular works on network science (BARABASI, supra note 2 or WATTS, supra note 2), uncovers only about ten articles making more than a passing reference to network science. Yet even this small group demonstrates the breadth of potential applications of network science to legal scholarship. See change. These analytical approaches and modeling tools complement the existing empirical methods and theoretical models of law and economics and other interdisciplinary approaches.

p. 5

Second, in Parts III, IV, and V, the Article illustrates the application of a network approach to empirical data by describing the results of a network science study of patent citations. The overall goal of the study is to gain a better understanding of the evolution and structure of technical relationships between patents in the United States patent system, and to investigate current patenting behavior so as to gain insight into innovation policy and law. The United States patent system has been the subject of increasing criticism in recent years. 7 There is a perception that recently issued patents sweep more broadly than their inventors' contributions justify, and that some are being issued for inventions that should be deemed obvious. Critics also contend that patents increasingly impose transaction costs in areas, such as scientific research, software, and business methods, in which patents may not be needed to promote technological advances. Recently proposed legislation 8 and renewed Supreme Court attention to patent law reflect these criticisms. 9 Responses to the criticism highlight the contributions of intellectual property to the nation's economy and ar-gue that increasing numbers of patents do not necessarily signify an erosion of patent quality. 10 Part III of this Article employs a network science approach to interpreting the burgeoning patenting of the past few decades. 11 In our study, the "nodes" of the network are United States patents and the "links" are citations of one patent by another. 12 Patent citations indicate technological relationships between the patented technologies. For the most part, a citation from one patent to another indicates either that the later patent builds upon the technology of the earlier patent or that the claimed inventions are closely enough related that the earlier technology was material to determining whether the later patent should issue. The citation network thus provides a "map" of the relationships between patented technologies which can be explored using network science techniques.

p. 6

We have performed a detailed study of the evolution of the patent citation network since 1975. In looking at how the network structure has evolved, we distinguish between three possible explanations for the recent increase in patenting: faster technological progress; increasing breadth of patented technologies due either to expansion of patentable subject matter or to new fields of invention; and a lowering of the legal standard for patentability leading to the issuance of patents for more trivial innovative steps. These possibilities are not mutually exclusive, of course, nor are they the only factors that might affect citation patterns. The question we ask here is whether one can explain the observed evolution of the patent network solely in terms of an increasing pace and breadth of technological advance or whether some additional factor is needed to understand the observed network growth and structure.

p. 6

Our analysis leads us to conclude that increasing pace and breadth of innovation alone are unlikely to explain the way in which the patent citation network has evolved in recent years. This is primarily because, as we explain in detail in Part III, the citation network has undergone a measurable increase in stratification since the late 1980s, following a modest decrease earlier in the decade. By "increasing stratification" we mean that 10. See, e.g., Perception Gap Hindering Efforts to Improve Patent System, Dudas says, 71 PAT. TRADEMARK & COPYRIGHT J. (BNA) 374 (2006).

p. 6

11. Our patent citation study is similar in flavor to recent studies of citations in scientific journal articles, see Sidney Redner, Citations Statistics from 110 Years of Physical Review, 58 PHYSICS TODAY 49 (2005), and of the "web of law" consisting of cases and other legal authorities and the citations linking them. See Smith, supra note 6; Chandler, supra note 6.

p. 6

12. United States patents also cite scientific literature and foreign patents. Our data do not include these additional links.

p. 7

there is an increasing disparity in likelihood of citation between the patents most and least likely to be cited again. In the empirical analysis described below, the probability that a patent will be cited again depends upon the patent's age and number of previous citations received. Because econometric studies have indicated that patent citations are linked to the technological importance of the patented technology, 13 a reasonable interpretation of this observation is that patents are being issued for increasingly trivial advances. 14 The timing of the increasing stratification roughly correlates with increasing reliance by the U.S. Court of Appeals for the Federal Circuit on the widely criticized "motivation or suggestion to combine" test for nonobviousness, 15 and with other indications of an increasingly lax legal standard of patentability. 16 In Part IV of this Article we describe how network science illuminates the question of how innovation proceeds by combining prior technologies. We examine the dependence of patent "citability," or probability that a patent will be cited again, on the age of the patent. Patent citability dies out surprisingly slowly with age even for patents that have rarely or never been previously cited. This means that there are "sleeper" patents which are cited after long periods of neglect, perhaps signifying how innovation proceeds not only by building incrementally on previous advances but also by "doubling back" to make new use of previously neglected technology. These results cast doubt on the possibility that the value of an innovation can be identified confidently at the time of patenting. 17 We briefly review some earlier studies of the innovative process using a network approach and discuss how these approaches might be combined with ours to investigate the innovative process.

p. 8

Part V of the Article describes several other possibilities for applying network science to the patent citation network to explore questions of legal significance. We first describe how the network concept of path length may be useful in quantifying the "closeness" of patented technologies. Numerous patent law doctrines, including patent validity and patent infringement under the doctrine of equivalents, are based on the perspective of the "person having ordinary skill in the art" or PHOSITA. 18 In addition, the obviousness of a claimed invention is judged in light of prior inventions in analogous arts. 19 Assessing the "closeness" of different technologies is essential to interpreting the concept of an "art" in applying these doctrines and is also of practical importance to patent classification.

p. 8

We then describe preliminary studies of the "small world" property of the patent citation network and discuss prospects for further study of the implications of the increasing "connectedness" of the network. Finally, we describe network metrics that may be of use in exploring the extent to which closely related patented technologies are complements or substitutes, which is relevant to the problem of patent thickets and to issues of anti-competitiveness in patent licensing.

p. 8

Part VI offers conclusions and returns to the broader point of this Article that network theory and modeling, as they are applied in social network analysis, computer science, statistical physics, and elsewhere, have great potential to complement other interdisciplinary approaches to the analysis of present and proposed legal doctrine.

A. What Network Science Is

p. 8

Network science, whether it is rooted in the social sciences, computer science, or the natural sciences such as physics or biology, has three general, interrelated, and ongoing goals: (1) to measure, describe, and categorize network structure and the patterns of relationships between network nodes; (2) to understand network evolution and growth and its relationship 18. See, e.g., Dan L. Burk & Mark A. Lemley, Is Patent Law Technology-Specific?, 17 BERKELEY TECH. L.J. 1155, 1185-91 (2002) (discussing the importance of the PHO-SITA in patent law).

p. 8

19. See, e.g., In re Clay, 966 F.2d 656, 658 (Fed. Cir. 1992).

p. 9

to network structure; and (3) to understand how the collective behavior of entities connected in a network depends on and derives from the network's structure. Many open questions exist at all three levels. The overall intuition behind the interdisciplinary conversations that characterize "network science" is that common structures, growth patterns, and collective behaviors will arise in networks composed of very different kinds of elements and linkages. If this is the case, common concepts and methods will be useful in understanding widely varying networks and in answering the very different substantive questions posed by physicists, biologists, computer scientists, sociologists, and, most recently, by legal scholars.

How to Describe a Network: Common Structural Signatures of Diverse Networks

p. 9

Network science seeks at the outset to measure and describe various relational structures found in the social or physical world. In a network in which the node is an individual, the links might be relationships of friendship, kinship, sexual involvement, business association, and so forth. When the node is an airport, the links might be flight paths. When the node is a business entity, the links might be supply and distribution relationships. And when the node is a patent, scientific journal, or legal opinion, the links can be citations. Links can be "directed," meaning that the nodes on either end of it relate to each other in an asymmetric way. For example, a link between a patent and a reference that it cites has a direction to it: if one reversed the order of the nodes in the link, it would change what the link meant. Links can also be "undirected." For example, a link indicating friendship is normally symmetric: "A is friends with B" normally means the same as "B is friends with A."

p. 9

The descriptive task of network science necessitates thinking about what quantities characterize networks-what metrics provide concise yet illuminating means of describing and comparing relational structures. Narrowly, it aims at providing a means to parameterize individual network structures. The broader descriptive goal is to delineate the similarities and differences between networks, with the hope of categorizing these structures in a logical way that can provide substantive insights into the underlying systems. Part of the fascination of network science is that similar metrics-degree distribution, "transitivity," network "path lengths," and others (some of which are discussed in this Article)-may be used to describe networks of very different "nodes" and "links." 20 Strikingly similar results are obtained for some of these metrics when measuring seemingly disparate networks. a) Not Your Average Node: "Scale-Free" Degree Distributions One metric by which one may distinguish networks is the frequency distribution of the average number of connections between network nodes. The "degree" of a node is the number of connections or "links" it has to other nodes. Node degree is an intuitive measure of "importance." A person in a social network who has many connections to others likely plays an important role in that network, for example. Of course, the meaning of such importance will be very different in a network of friends than it is in a network of legal citations, but the network approach brings to light underlying similarities between very different networks. One way to characterize a network is to study the "node degree distribution." The node degree distribution charts the fraction of nodes in the given network that have a particular number of connections. 21 If a network is produced by randomly connecting nodes to one another with a uniform probability of connection, the degree distribution will be symmetrically peaked around a "typical" number of edges determined by that probability of connection. 22 For large networks, the peak will be very narrow, so that most nodes will have approximately the "typical" number of edges. In this sense, the nodes of a random network can be said to be relatively homogeneous. In the random network analyzed in Figure 2, for example, the most common number of connections for a node is twenty, which is also the average and median number. Though there are plenty of nodes with ten to thirty connections, the numbers drop drastically outside the peak width. There are virtually no nodes with 100 connections.

p. 10

The symmetrical, peaked distribution of Figure 2 is similar to the "normal" distribution or "Bell curve" that characterizes the probability distribution of many things in our everyday experience. 23 (See Figure 3 for an example of a normal distribution.) The fact that the height of human beings, for example, has this kind of normal probability distribution makes it possible to design furniture, showers, cars, and the like for the average 21. See, e.g., infra Appendix, Figure 1 (plotting the node degree distribution for the patent citation network).

p. 10

22. See infra Appendix, Figure 2. 23. The probability distribution for the node degree of a random network is not precisely a normal or "Gaussian" distribution. Mathematically, the random network degree distribution takes another form typical for probability distributions-the Poisson distribution. Newman, supra note 2, at 197-99. For present purposes, the difference is unimportant. What is important is the peaked, symmetrical shape.

p. 11

person and to expect that nearly all people will be able to use them, perhaps with minor adjustments.

p. 11

One of the first and most interesting results of network science was the discovery that many networks encountered in the real world are not homogeneous like the random networks with degree distributions as illustrated in Figure 2. Instead, many real world degree distributions are highly skewed, and very broad, with what are called "fat tails." 24 In these networks, rather than all nodes having roughly the same "typical" number of links, there are a large number of nodes with very few links and a small, but significant, number of nodes with very many links. These highly connected nodes may have hundreds or thousands of times as many connections as the "average" or "typical" node. Depending upon the network, the few highly connected nodes may play a much more important role than the large number of minimally connected nodes. In this kind of situation, it is not enough to make decisions based on the likely behavior of "typical" nodes. Because a significant fraction of nodes are so far from typical, the effects of non-typical nodes must be taken into account in understanding and predicting the behavior of the network. For example, in a communications network, the system might continue to operate despite failures of many "typical" nodes but may be completely defeated by the failure of one highly connected node. Any attempt to understand the collective behavior of such a network by focusing on the "typical" node is doomed to fail.

p. 11

When the degree distribution is sufficiently broad, it will have a "power law" decaying tail. In some cases, network degree distributions may be so broad that they do not even have well-defined averages or standard deviations. Figure 4 illustrates a power law degree distribution, meaning that the fraction of nodes with a particular number of connections decays as d -γ once d is large (where d is the degree of a node). This shows how it becomes more difficult to identify a "typical" node as the decay power γ becomes smaller. When γ is 3, the most likely number of connections is 0, while the median number is about 15 and the average number is about 20. When γ is even smaller, the most likely number of connections is still 0, but the median number is 40 and the average is essentially infinite because there are so many nodes with very large numbers of connections. The likelihood of having a particular number of connections for γ = 1.5 is fairly flat and it would be a mistake to say that a node with 10, 40, or 100 connections is "typical."

p. 12

Because there is no "typical" number of links for a node in a power law distribution, such networks are often termed "scale-free," meaning roughly that there is no typical number of connections or "scale" that can be used to describe the network. In such a scale-free situation, knowing, for example, that the average number of links per node is 12 is not useful for predicting the collective behavior of the network because there are significant numbers of nodes with no links and significant numbers of nodes with hundreds or thousands of links. 25 Decisions, plans, or predictions based on the average node degree can go badly awry. An example of this kind of situation is the attempt to estimate the spread of a virus, such as HIV. While the average number of sexual partners per person may be quite small, if, as appears to be the case, the distribution is sufficiently broad (effectively scale-free), there are a small but significant number of individuals who have hundreds of sexual partners. 26 Any predictions of how fast the virus will spread based on the average number of sexual partners will be completely wrong. 27 Moreover, a program of random condom distribution or HIV testing may be much less effective than anticipated if it happens to miss the few most highly active individuals.

p. 12

Highly skewed and broad degree distributions are common in networks of widely varying types, including the physical connections and hardware of the internet, the hyperlinks of the World Wide Web, the networks of co-authorship of high energy physics papers and of coappearances of movie actors, networks of sexual contacts, metabolic networks, phone call networks, networks of word uses, protein interaction networks, and many others. 28 Other observed networks, such as the power grid, do not have a scale-free or nearly scale-free form, however. Thus, the 25. More precisely, a "scale-free" network is one in which the large degree tail of the node degree distribution decays as a power of the degree, P(k) ~ k -a , rather than exponentially with the degree, as more well known probability distributions do. Where the degree distribution decays exponentially, it has a well-defined "scale" given by the average number of links. The term "scale-free" has its roots in the physics of critical phenomena, in which scale referred to the typical length scale of correlations between atoms or magnetic spins. At a critical point, the correlations become "scale-free", meaning that they occur on all length scales and hence no characteristic length scale can be identified. Just as the analysis of social issues becomes more difficult when no typical behavior can be identified, the physics of critical phenomena is difficult because there is no typical length scale. Technically, the term "scale-free" should be reserved for networks in which the node degree distribution has a power law decay. In reality, it is sometimes difficult to tell a true "power law" decay from a very broad, skewed distribution with an exponentially decaying tail.

p. 12

26. Ayres & Baker, supra note 6, at 607-11. 27. Id. 28. See Albert & Barabasi, supra note 2; Newman, supra note 2. degree distribution is a metric which can be used to categorize networks and to differentiate between them with respect to how they may behave. In the legal context, knowing how networks-such as drug distribution networks (legal and illegal), communications networks, infectious networks, terrorist networks, and the social networks of ordinary citizens-tend to be organized may be very important in devising effective policy and law.

p. 13

b) Building Bridges: The Betweenness Metric Node degree is only one possible way to measure the importance of a network node. Depending on the circumstances, the number of connections a node has may be less important than its specific place in the network structure. A node with few links may play a critical part in connecting different parts of a network. The friend who moved to New York may not have been the most popular person in the Chicago social network, but she may be the only connection to New York which the social group has. In that case, she will have a relatively low node degree but a relatively high "betweenness centrality." 29 The betweenness centrality of a particular node can be measured by counting the number of paths between two points in the network which pass through that node. Chicago's O'Hare airport, for example, has a high betweenness centrality because many flight paths between two United States cities pass through that airport. For some purposes, such as communicating between disparate parts of a network, for example, a network may be more dependent on high "betweenness" nodes than on those with high node degree. Of course, often, as in the case of O'Hare airport, nodes with high degree centrality also have high betweenness centrality. c) Six Degrees of Separation: The "Small World" Property Another commonly observed attribute of real world networks is the "small world" property, in which a relatively small number of "hops" between nodes is needed to connect any two nodes in the network. 30 The small world property can be important in determining the rate at which something-which might be a fad, a piece of information, or a deadly virus-spreads across the network. The significance of the small world property depends on the particular network. In general, it determines how tightly connected a network is. The small world property is often related to the presence of highly connected "hubs" in the network, which connect 29. See, e.g., Carrington, Scott & Wasserman, supra note 2, at 65-66 (defining this and other network metrics). 30. Id.

p. 14

otherwise widely separated nodes (as in the sexual contact example already mentioned). This is not always the case, however. A random network with a degree distribution like that shown in Figure 2 can have the small world property, as can a regular network (think of a fishnet, for example), after a small amount of re-wiring to connect distant parts of the network. 31 Re-wiring occurs in a social network, for example, whenever one person moves to a distant location and maintains social contacts in both locations. The person who moves from Chicago to New York broadens not only her own choice of friends, but also the choices of all her friends and relations who may meet New Yorkers through her.

p. 14

The presence or absence of the small world property is thus an additional metric which one may use to categorize observed networks and to understand their behavior. There are various ways to measure the small world property, which have somewhat different significance. The most common measure is to compute, for each pair of nodes in the network, the minimum path length (number of hops needed to get from one node to the other), and then to identify the largest of this set of lengths as the "diameter" of the network. 32 The comparison between this network diameter and the size of the network determines whether the network has the small world property.

d) Transitivity and Local Network Structure

p. 14

The small world property is one example of the ways in which network science probes beyond both average properties of individual nodes (such as node degree) and overall properties of the network (such as its 31. See Duncan J. Watts & Steven H. Strogatz, Collective Dynamics of "Small World" Networks, 393 NATURE 440 (1998) (discussing the basic properties of small world networks).

p. 14

32. Some technical issues arise here. The term "small world" is used in two slightly different senses: to denote the property of short "distance" between nodes and to denote that property joined with the additional element of local clustering (more technically, high "transitivity"). By the "small world property" here we will mean short "distances" between nodes. By short, we mean, technically, that the longest distance between two nodes grows no more than logarithmically with the number of nodes. Intuitively, this means that it is impossible to make a spatial "map" of the network in which nodes that are connected to one another are near to one another. You will always end up with some connections between "far apart" nodes. An additional technical issue arises because many networks are not completely connected-there are isolated nodes or small clusters of nodes. Path lengths between these isolates and the bulk of the network are technically infinite. A network is still considered to have the small world property if it contains a socalled "giant cluster" which contains the vast majority of nodes and if the giant cluster has the small world property. size) to investigate network structure. Another important measure of network structure is "transitivity" (often called "clustering coefficient" in the physics literature). 33 The transitivity of a node can be measured by determining the fraction of its network neighbors which are connected to one another. In the social context, a node with high transitivity is an individual whose friends are friends. Such individuals tend to be part of a densely connected social network. Often these multiply connected groups are connected by relatively strong ties of friendship or kinship. A node with low transitivity is an individual with friends who do not know one another. These individuals generally have more acquaintances (as opposed to close friends) and often their social contacts span a wider social context. 34 In other contexts, transitivity has different significance. In a citation network, for example, the transitivity of a particular scientific article may indicate its breadth or degree of interdisciplinarity. An article that cites or is cited by articles that cite each other will tend to have a narrower focus than an article that cites or is cited by articles that do not cite one another (and thus tend to be less closely related). Transitivity distinguishes nodes of the same degree. A highly cited article may be important in a single "hot" field or have more general significance. A person may be very popular within a close-knit group or may be more widely "connected." Transitivity is thus an important means of probing the local structure of a network.

p. 15

The above examples give a brief glimpse of the ways networks can be characterized by their patterns of node relationships and of how the network structure can affect the collective behavior of the network. The next Section discusses how the structure of a network is related to the way in which it evolves.

How Does Your Network Grow? "Preferential Attachment" and Other Factors in Network Evolution

p. 15

Degree distributions, small world properties and other measures are useful in describing and categorizing real world networks. The aim of this descriptive task is not simply the "gee-whiz" satisfaction of seeing similar graphs show up for widely differing networks (though physicists, at least, do derive an inordinate amount of pleasure from such observations of universality). The underlying intuition driving the descriptive enterprise is 33. See Carrington, Scott & Wasserman, supra note 2, at 164-67; Albert & Barabasi, supra note 2, at 49; Newman, supra note 2, at 183 (giving definitions and discussing transitivity (also known as "clustering coefficient")).

p. 15

34. See Mark Granovetter, The Strength of Weak Ties: A Network Theory Revisited, 1 SOC. THEORY 201, 218 (1983) (discussing the relationship between transitivity and "strong" and "weak" social ties).

p. 16

that similar structures probably result from similar mechanisms of network growth and evolution, and that networks with similar structures will behave in similar ways. Making the connection between the mechanisms by which a network evolved and the network's eventual structure is the second task of network science.

p. 16

In many cases, for example, a scale-free degree distribution results from what is called the "preferential attachment" mechanism. 35 Preferential attachment (also called the "rich get richer" effect) describes a process for creating a network in which "popular" nodes, which already have many links, are more likely than others to gain additional links as nodes are added to the network. It can lead to a highly heterogeneous, scale-free degree distribution in which some nodes eventually acquire a very large number of links while others remain relatively unconnected.

p. 16

Why does this happen? In some cases, the nodes of the network may have a highly skewed distribution of "quality" which becomes reflected in the network connectivity. In a social network sociable people attract friends; in scientific journal citation networks the most important seminal papers are increasingly cited; in transportation networks "all roads lead to Rome" and so forth. Probably more often, however, the very fact that nodes are highly linked may make them more "linkable"-it is useful to fly through Chicago's O'Hare airport because so many other connecting flights already do; there may be reputational benefits to associating with popular people; popular sites garner more links on the World Wide Web. 36 Once one identifies the concept of preferential attachment, it can provide insights into the growth of many different kinds of networks. While the particular reasons that well-connected nodes acquire additional links cer-35. See Albert & Barabasi, supra note 2, at 71; Newman, supra note 2, at 212-18. 36. The preferential attachment mechanism is related to the economic concept of "network effects" which explains, for example, why a lesser quality technology (such as VHS rather than Betamax for video recorders) can become entrenched. See, for a definition, S. J. Liebowitz & Stephen E. Margolis, Network Externalities (Effects), in THE NEW PALGRAVE DICTIONARY OF ECONOMICS AND THE LAW 671-75 (Peter Newman ed., 1998). Network effects can be a source of preferential attachment. People choose a particular technology because others have chosen it. "Network effects" are important causes of market failure and sub-optimal social outcomes. More generally, preferential attachment can lead to highly skewed outcomes even when attachments are selected entirely at random in the first instance. Disentangling pure preferential attachment ("the rich get richer") from the effects of underlying node quality ("the talented or hard-working get richer") is a difficult and open question in network science. Understanding the reasons for preferential attachment in a particular context may be very important in predicting or understanding the network's collective behavior.

p. 17

tainly vary from network to network, the general phenomenon of preferential attachment is widespread.

p. 17

Though preferential attachment is common, in many real situations it is curtailed by cost or congestion. The capacity of a power hub limits the number of substations to which it can connect; the capacity of an airport limits the number of flights it can handle; popular people may run out of time or otherwise wish to limit their numbers of friends. Measuring the degree distribution of a network can thus lead to insights into how the network evolved. If the network is scale-free or very broad, one may search for some mechanism for preferential attachment. Similarly, if the number of connections cuts off at some point, a hunt for a congestion or cost mechanism may be warranted.

p. 17

Citation networks, such as the patent network discussed in Part III of this Article or a network of scientific journal citations, tend to have very similar degree distributions, which are broad and skewed, but differ from the pure preferential attachment scale-free form. Figure 1 shows the distribution of number of citations received for the patent citation network at different times. It is quite broad-there are many patents that have never been cited, but others that have received nearly 1000 citations. To look for possible scale-free behavior, we plot the distribution function on a "loglog" graph. On such a graph, a power-law decay would show up as a downward sloping straight line. (This is illustrated in Figure 4 at the bottom, which shows the power law distributions on a log-log graph.) The patent citation distribution is clearly quite broad and skewed, but the tail of the distribution is not quite a straight line on the graph. Instead, it curves downward, cutting off the probability of finding a patent that has received an extremely high number of citations.

p. 17

While the precise mathematical form of these citation network distributions is not completely understood, our results, discussed in more detail in Part III, and those of others 37 suggest that the shape reflects a competition between preferential attachment-which leads to some patents acquiring very large numbers of citations-and the rate at which patents age or go out of date. The preferential attachment leads to a broad, skewed distribution, while the aging of nodes slows their accumulation of citations and cuts off the tail of the degree distribution. If a similar distribution is observed in some new network, it will be reasonable to predict that the growth of that network involved some mechanism for preferential attachment and for aging.

p. 18

As more networks are characterized using a variety of metrics, it becomes possible to infer much about the means by which a particular network evolved (or is evolving) by measuring aspects of the network's current structure. It may conversely be possible to predict many aspects of the network structure that will result from particular means of network growth.

Network Structure and Collective Behavior

p. 18

The third task of network science is to understand the kinds of collective behavior that may emerge when elements of a network interact with one another and how the resulting behavior may depend upon the structure of the network. For example, the flow of information (or the progress of an epidemic) within a social group, the flow of oil through the ground, the susceptibility of an electrical or communication grid to the failure of individual elements, the cost of licensing the necessary patents to commercialize a particular technology, and the propagation and acceptance of a changing behavioral norm may all depend on the relational structure of an underlying network. 38 Network science will attempt to delineate and interpret the common features of network structures and interactions between nodes that result in particular varieties of collective responses. Such understandings may be used to predict the collective effects of changes in network structure, changes in the interactions between neighboring elements, and changing global influences-including changes in legal rules.

B. Why Legal Scholars Should Care About Network Science

p. 18

In some sense, legal scholarship, in its descriptive form, is network science-the study of how particular social, commercial, infrastructural, and other kinds of networks adapt and react to particular rules of interaction between nodes that might be individuals, corporations, communities, "things," or institutions. Moreover, the common law itself is a growing network of precedent, as reflected in citations in judicial opinions. 39 More than this, though, normative legal scholarship, legal policymaking, statutory legislation, and common law jurisprudence are in some respects network engineering-attempts to produce particular collective social re-38. See BARABASI, supra note 2, WATTS, supra note 2, and MALCOLM GLADWELL, THE TIPPING POINT (2000) for examples of situations in which the local structure of a network can determine the global behavior.

p. 18

39. See Post & Eisen, supra note 6, Chandler, supra note 6, and Smith, supra note 6 for studies of the network of legal citations.

sponses by adjusting the interactions and incentives experienced by individual network elements. Of course, this observation is important (and not merely semantic) only if thinking about society as a network mattersonly if it tells us something different from what we have already gained from our existing conceptual frameworks.

p. 19

There are good reasons to think that network analysis will matter in this way-and that it may matter a lot. Here we discuss three inter-related concepts highlighted by network science which may have broad impact on legal and policy analysis: (1) node heterogeneity, (2) the importance of network relational structure, and (3) the complicated relationship between local structure and interactions and global, collective behavior.

The Implications of Heterogeneity

p. 19

Much of our legal doctrine and scholarship-including much law and economics analysis-depends implicitly, but heavily, on assumptions and predictions about average or typical behavior. Thus, in legal decisionmaking we rely on the "reasonable person" and the "person having ordinary skill in the art;" we seek to deter the "typical" criminal; and we attempt to avoid confusing the "ordinary consumer." We predict the results of policy changes by considering the rational economic actor (or, perhaps, the "boundedly rational" economic actor). 40 All of these modes of analysis rely on an intuitive assumption that various traits, propensities, preferences, and so forth are distributed among the population according to a more or less normal distribution. In other words, we assume that the average behavior is the most typical behavior and that extreme deviations from the typical are extremely rare-and thus generally unimportant.

p. 19

This assumption is reasonable for many purposes since many properties are distributed according to a normal distribution-there are no hundred foot tall or one inch tall human beings, for example. However, network science provides a warning about unquestioning reliance on the typicality assumption. It points out that there are common examples of quantities-such as the number of links to a node-that are distributed in a highly skewed and very broad manner-indeed, scale-free or close to it. While we have always known that such broad, skewed distributions were possible, the study of networks and their societal ubiquity brings us faceto-face with the fact that highly skewed, scale-free or nearly scale-free, 40. See, e.g., Christine Jolls, Cass R. Sunstein & Richard H. Thaler, A Behavioral Approach to Law and Economics, in BEHAVIORAL LAW AND ECONOMICS 13, 14-15 (Cass Sunstein ed., 2000) (providing an overview of the significance of bounded rationality).

p. 20

distributions are not only possible, but are common in some contexts and that taking account of the non-typical may be crucial. 41 As discussed above, if a network is scale-free or nearly so, it is meaningless to talk about a "typical" node because the average node is not the most typical node-and the most important nodes are likely neither average nor typical. As a result, social interventions, including legal rules, which are aimed at the "typical" member of a social group may be ineffective-it may be necessary to consider a broad spectrum of characteristics to devise a workable rule. Ayres and Baker, for example, have highlighted the importance of the extreme heterogeneity of sexual behavior to designing legal means to deal with the threat of HIV-AIDS. 42 Thus, the first lesson from network science is that legal analysis should consider the possibility of heterogeneity-in some cases radical heterogeneity-among network elements and environments. Indeed, the realization that a simple and common mechanism like preferential attachment in a network can result in radical heterogeneity should motivate us to consider what other commonplace deviations from "typicality" may arise. 43 Perhaps deviations from the normal distribution are not so abnormal after all.

p. 20

Currently, there is no complete fundamental understanding of when scale-free distributions occur. Intuitively, however, it seems likely that these distributions are possible when scale is not strongly constrained by some form of cost. 44 For example, the network of World Wide Web hyper-

See CHRIS ANDERSON, THE LONG TAIL: WHY THE FUTURE OF BUSINESS IS SELL-ING LESS OF MORE (2006) (discussing the prospective effects of consumer heterogeneity on business).

p. 20

42. See Ayres & Baker, supra note 6, at 630-58 (arguing in favor of a crime of "reckless sex" based on the importance of targeting one-time sexual encounters to prevent the spread of sexually transmitted diseases, in part because of the highly skewed distribution of numbers of sexual partners).

p. 20

43. The statistical physics interest in scale-free distributions, for example, pre-dates network theory and stems from the observation of scale-free behavior at phase transitions and in observations of self-organized criticality. 44. In statistical physics, for example, interest in scale-free phenomena arose in the study of critical phenomena. Critical points occur when energy and entropy balance in such a way that fluctuations on all length scales have equal free energy and thus all occur links (which is relatively unconstrained by linkage costs) is approximately scale-free, whereas the railway system (which is constrained by the costs of building additional track "links") and the power grid (which is constrained by the cost of wiring and the loss of energy in transportation) are not. 45 An observation of extreme heterogeneity should trigger an inquiry not only into its effects but also into its causes. It is conceivable, for example, that there is a relation between the well-known highly skewed distribution of values of technological innovations and a fundamental lack of proportionality between research investment and research result. If this (highly speculative) premise were correct, it might have important implications for the use of patent law to provide incentives to invest. Similarly, a very skewed distribution of network connections which results from some particular combination of legal rules and other norms and constraints might reflect an underlying skewed distribution of talent or effort, but we should recognize that preferential attachment alone is enough to produce such a distribution even if all nodes are identical. Depending on whether the skewed distribution of connections is socially desirable, this recognition may affect the choices of policies to pursue.

The Importance of Network Structure in Determining Individual Responses to Legal and Social Change

p. 21

A second lesson from network science is the importance of local network context in determining individual responses to legal and social change. Legal analysis often conceptualizes individual legal actors as responding independently to legal rules in the context of global average social forces. 46 Network science demonstrates that individual responses to legal intervention or changing social norms may be determined not only by the average impact of global social forces but also by the specific network structures by which these social forces are mediated in a particular case. Social network studies demonstrate that access to information deat once. Away from critical points, fluctuations in physical systems have a characteristic scale. pends on one's position in the network. 47 An individual's susceptibility to legal and social sanctions for socially undesirable behavior will be similarly dependent on his or her local social network.

p. 22

Rather than thinking of an individual person or firm as being immersed in an average social ether of information and influences, it may sometimes be important to take into account the specific local relationships in which particular individuals are embedded. Dickerson, for example, discusses how the network structure within corporations and the distinctions between the connection environments of particular nodes has been implicitly, and might be explicitly, used to effect corporate change by targeting particular corporate actors. 48 The realization that a model in which an individual particle is subjected to the average influence of its fellows cannot always predict behavior was important in statistical physicsit led to the rejection of mean field theory and inaugurated the modern study of collective physical phenomena. 49 The statistical physics approach to network science stems from this background understanding, which may be quite important for legal analysis as well. 50

The Complicated Influence of Local Network Structure on Collective Behavior

p. 22

A corollary of the premise that individual behavior may be significantly affected by local network structure rather than simply by average social forces is that one cannot simply predict collective behaviors, such as social norms and behavioral regularities, the collective impact of legal 47. Similarly, statistical physics studies show that a randomly constructed network may make a transition from a set of disconnected clusters to a single connected network at a particular critical density of links. A scale-free network may or may not experience such a transition depending upon the precise form of the degree distribution and other factors. See Newman, supra note 2, at 225-28.

p. 22

48. See Dickerson, supra note 6, at 556-67. 49. In statistical physics, the approximation that an individual atom reacts to the average influence of the other atoms is called "mean field theory." While mean field theory is very useful in some situations, it is particularly bad at explaining "phase transitions" (such as melting and freezing, evaporation, and magnetization), in which the collective behavior of a system changes from one macroscopic state to another. Part of the reason for the failings of mean field theory is its inability to handle the "scale-free" correlations between blocks of atoms of widely varying sizes that appear near many phase transitions. See, e.g., MA, supra note 43, at 34-39 (discussing mean field theory and its shortcomings in the phase transition context). 50. See Albert & Barabasi, supra note 2, at 91-92 (discussing the implications of local correlations for network structure); Newman, supra note 2, at 224-40; see also PHILIP BALL, CRITICAL MASS: HOW ONE THING LEADS TO ANOTHER (2004) (discussing, for a general audience, the application of statistical physics methods to social problems).

p. 23

change, and flows of information, influence, and other goods by looking at local relationships and individual transactions.

p. 23

Optimal collective behavior, for example, may depend not just on individual relationships but on the particular network structure in which those relationships are embedded. A simple example shown in Figure 5 illustrates the point. Assume that each dot represents an individual and that the associated + or -represents the individual's yes or no decision about some question. Also assume that each individual must make a decision and prefers, for some reason, to make the opposite decision from his or her immediate neighbors. (In physics this is known as the anti-ferromagnetic Ising model. 51 ) If the network of relationships between individuals forms a square pattern, as shown in Figure 5 at the top, every pair can be made happy by making opposing decisions, and the global arrangement will reflect the optimization of the pairwise interactions. However, if the network of relationships between individuals forms a triangular pattern, as shown in the bottom panel of Figure 5, it is simply not possible for everyone to disagree with all of his or her neighbors. Even though there is a preferred arrangement for every pair of neighbors (disagreement), the network as a whole will be unable to get to a completely "disagreeable" state. The network will be, in the terminology used by statistical physicists, frustrated. 52 The best global arrangement (the one with the maximum number of local disagreements) is not predictable by scrutinizing the preferences of each pair in isolation. The collective behavior that results from the relationship pattern shown at the top of Figure 5 differs entirely from the collective behavior that results from the pattern shown at the bottom of Figure 5, even though the local interactions and preferences are identical.

p. 23

Figure 5 is just an illustration of the more general point that one cannot always predict global behavior from interactions between pairs or in small groups. In the legal context, this insight suggests that focusing only on the efficiency of individual transactions between individual legal actors may, in some cases, lead to suboptimal legal policies because these actors are embedded in a network of relationships. Yet legal analysis and implementation frequently does focus on such individual transactions, either because legal rules are forged in the context of specific disputes or because, as a theoretical matter, generalizing beyond these individual transactions is difficult. 53 51. SORNETTE, supra note 44, at 441-43. 52. See, for example, id. at 441-43, for a discussion of this example of frustration. 53. For example, applications of game theory to legal theory very frequently focus on analyzing a two-person or, at most, three-person game and then extrapolating to form predictions about collective behavior. See generally, DOUGLAS G. BAIRD ET AL., GAME As another example of the importance of local network structure for collective behavior, diffusion of anything on a network-whether it be information, opinion, resources, or infection-may be a complicated function of the number and placement of network links and not simply of the average likelihood of transmittal from one node to its neighbors. Depending on the specifics of relationships, diffusion of information from one social group to another may be rapid or extremely slow. 54 Lior Jacob Strahilevitz, applying these ideas, argues that analysis of the privacy tort of "public disclosure of private facts" should rely on the network concepts of weak and strong ties to distinguish between information that likely would have spread without the challenged disclosure and information that likely would not. 55 Merely recognizing these conceptual points may sometimes be enough to provide insight into an important legal question. 56 However, network science's potential usefulness to legal analysis extends beyond highlighting these general concepts. Indeed, the observations that important social and economic quantities may be distributed in a highly skewed, scale-free manner and that collective behavior may not be accurately predictable by focusing either on average global influences or on individual transactions are not entirely new. Legal scholars and decision makers have employed the approximations of typicality, of average social forces, or of the generalizability of pairwise interactions in the past not because they were unaware that reality is more complicated but to reduce the complexity of the analysis.

p. 24

Network science can contribute in that it not only highlights these conceptual issues, but also promises to provide ways to determine when these issues are likely to be important and tools to perform more accurate analyses in various contexts. These tools may be analytical, such as the introduction of explanatory concepts like "preferential attachment" and "small worlds." They may be mathematical, as illustrated by the analysis of patent citations in Part III of this Article. Or they may be computational, as exemplified by many social network theory software tools or by the computer simulation methods of statistical physics. 57 Network science concepts may provide a basis for critiquing extant or proposed legal rules. 58 Network science methods may also provide, as we begin to demonstrate in Parts III and IV of this Article, means of empirical analysis that are applicable to legally significant systems in which heterogeneity and network structure are important. Network science techniques may also help to model the collective behavior that results from particular local interactions and preferences and to predict whether legal changes will lead to global behavioral change. For example, one may begin by postulating particular pairwise interactions, such as those embodied in many law and economics game theory models. Network science methods may then be employed to determine the collective behavior that results from such pairwise interactions. 59 Computer simulations will often be the most feasible approach to such questions and can in principle be devised to account for heterogeneous preferences and for various network structures.

p. 26

Just as has been done with game theory models, it may well be possible to map some problems of legal significance onto models that physicists, sociologists, and other network scientists have already studied. Additionally, legal problems may eventually motivate network scientists to perform computer simulations or mathematical studies of models derived from those problems.

p. 26

In sum, network science is an emerging discipline that holds great promise to provide insights, tools, and models that will make important contributions to legal analysis. In the next Part we provide a sample of the application of network science to law by investigating what the evolution of the patent citation network can tell us about how the standard of patentability is changing.

III. NETWORK CLUES TO A DECREASING PATENTABILITY STANDARD

p. 26

We now move from the lofty realms of possibility to a specific application of network science in the legal arena. We have performed a quantitative analysis of United States patents and their citations, treating the patents as nodes and citations from one patent to another as network links. Our analysis suggests that the "patent explosion" of recent years is not completely explained by either a rapidly increasing pace of technological advance or a broadened scope of patented technology. Something more has happened in the network-an increasing stratification of patent "citability," 60 by which we mean the probability that a patent with particular characteristics will be cited. In this Part, we describe how our network science approach reveals this dynamic change and hypothesize that it may result from a decreasing patentability standard. We distinguish carefully between our empirical results-which any theory of the way in which patents and the citations between them have evolved will need to explainand our proposed interpretation. A brief technical report focusing on the v1; Sitabhra Sinha & Sudeshna Sinha, Robust Emergent Activity in Dynamical Networks (Oct. 22, 2005), http://arxiv.org/PS_cache/condmat/pdf/0510/-0510603.pdf. 60. Citability is defined above as the probability that a patent will be cited again depending on its current age and number of previous citations. methodological contributions of this research to network science has been accepted for publication in the physics journal, Physica A. 61

The Patent System and Its Discontents 62

p. 27

Recent years have seen a major upsurge in patenting, an expansion of the range of innovations which are eligible for patent protection, and a perception that the United States economy relies more and more heavily on knowledge and innovation for its success. 63 At the same time, developments in the law, including the establishment of a single appellate court-the U.S. Court of Appeals for the Federal Circuit-to hear the vast majority of patent appeals in the United States, have led to debate as to whether the legal system is becoming increasingly patent-friendly; whether patents are being issued for lower quality innovations; and whether the legal rights awarded to patentees are becoming stronger. Empirical evidence increasingly raises questions as to the extent to which patents are needed to provide incentives for research and development. 65 These trends have converged to raise growing concerns among academics and policymakers about whether patent law and policy are adequately designed to "promote the Progress of . . . useful Arts." 66 This discontent has gained the attention of members of Congress, who have proposed various patent reform bills, 67 of the Federal Trade Commission, 68 of the National Academies of Sciences, 69 and of the Supreme Court, which has granted review in six patent cases since 2005 70 -a level of Supreme Court interest unheard of in at least 25 years. It has also provoked responses from defenders of the present system, who argue that criticisms are overblown. 71 To accomplish its constitutionally mandated objective, patent protection must be carefully tailored to balance its benefits against its costs. The benefits may include providing incentives to invent, functioning as a signal of technical competence, and facilitating a market for intangible knowledge. However, because a patent provides exclusive rights to practice the patented technology, patents impose costs on society that may include not only supra-competitive pricing of patented products but also increased barriers to building upon existing technology. These barriers arise because improving upon a patented technology may require either using the patented technology during development or incorporating it into the improved result. In either case, United States law generally requires that the improver obtain a license, usually requiring the payment of royalties, from the holder of a patent on the foundational technology or technologies. Negotiating authorization to build upon a patented technology can be expensive, especially when there is disagreement as to the relative value of various contributions to the improved technology. In the extreme case, when the holders of the original patent and an improvement patent cannot agree on a licensing arrangement, this "blocking patent" situation can deprive the public of access to the improved technology altogether.

p. 29

Part of the patent cost-benefit analysis is a requirement that the USPTO issue patents only for inventions that are novel and nonobvious 72 -that is, the inventions differ sufficiently from presently available technology to justify an award of legal exclusivity to the inventor. The legal standard of nonobviousness sets the height of the bar for "sufficient" difference. Along with other patentability requirements, it determines the tradeoff between social costs and benefits that results from the issuance of a patent. Because patents can impose substantial costs on consumers and subsequent innovators, one major objective of studies of the patent system is to determine whether the procedures and substantive standards that guide the issuance of patents are appropriately tuned.

p. 29

Recent developments in the law, along with the issuance of a burgeoning number of patents, have led to widespread perceptions among patent law scholars and policymakers that the system is off balance. 73 Many argue, for example, that the legal standard of nonobviousness is insufficiently rigorous; 74 that patent examiners have insufficient access to potentially relevant prior art in some fields such as software or business methods; 75 and that patents issue on increasingly upstream technologies that form the basis for further advances. 76 There are dire predictions of a patent thicket, 77 in which technological progress is made increasingly difficult by the need to negotiate multiple levels of blocking patent rights on each of the many patented components which may be needed to produce a new commercial product. 78 One way to avoid a potential thicket is for competing patent holders to negotiate cross-licenses or patent pools. Such agreements between competitors raise concerns about collusion, however, and the societal ramifications depend upon the extent to which cross-licensing lowers barriers to the use of complementary technologies, as opposed to allowing competitors to avoid competition from substitute technologies. 79 Understanding the interaction between innovation and the patent system is difficult for many reasons. Increased patenting, for example, can stem from various causes, including an increased pace of technological change, an increased range of patented technology due either to expansion of the scope of legally patentable subject matter or to the birth of new fields of technology, a growing perception of the usefulness of patents as business tools, or the issuance of lower quality patents. Empirical investigation of the patent system can play an important role in understanding how to maintain the appropriate balance. In what follows, we report the results of a detailed study of the evolution of the patent citation network and argue that the evolution suggests that the patentability standard has been decreasing. Before presenting our results, we detour to provide an overview of patent issuance and the meaning of patent citations and discuss how patent data has been used in previous studies. With that context in place, we then describe our study.

Patent Issuance and Prior Art Citations

p. 31

The United States Patent and Trademark Office (USPTO) issues patents after applications are examined to determine, among other things, whether the patent claims meet the legal requirements of novelty and nonobviousness. 80 Patent claims are specific statements of the scope of the legal coverage of a patent. As noted above, the legal effect of a patent is to provide the patentee a right to exclude others from using the claimed technology without a license, as detailed in the infringement provisions of the patent statute. 81 In the course of the examination of a patent application for novelty and nonobviousness, patent claims are compared against potential prior art, consisting in large part of prior patents and other publications in relevant technical fields. Applicants, their patent attorneys, and the official patent examiners all identify potential prior art. To qualify for a patent, the claimed invention must be novel, meaning there is no prior patent or other prior art that is identical to what is claimed. More importantly, the claimed invention must be nonobvious, meaning that at the time it was invented, the invention would not have been obvious to a person having ordinary skill in the art in the field of the invention. When an invention is deemed obvious, it is usually because it is an obvious combination of prior art technology. Determining whether a combination of prior technology would have been obvious is a tricky matter, since it requires both an exercise in hindsight and putting oneself in the shoes of the skilled practitioner of the patented technology. 82 The way in which obviousness is determined essentially sets the threshold of patentability. To guide this endeavor, and to try to avoid hindsight bias, the Federal Circuit employs a controversial test. Under this "suggestion test," a patent examiner may not reject a patent application as an obvious combination of prior art elements unless a "suggestion, teaching, or motivation to combine" the prior art elements appears "1) in the prior art references themselves; 2) in the knowledge of those of ordinary skill in the art that certain references . . . are of special interest or importance in the field; or 3) from the nature of the problem to be solved, 'leading inventors to look to references relating to possible solutions to that problem.'" 83 The test has been widely, though certainly not universally, criticized as lowering the barrier to patenting of obvious combinations of or improvements on old technology 84 and is pending review by the Supreme Court. 85 Seeking out prior art patents (and other sources of prior art) is key to determining both novelty and nonobviousness. The search for related prior patents is guided to a significant extent by an ad hoc classification scheme that has been developed by the USPTO over the years. 86 References are cited in the issued patent document if their technical relationship to the claimed technology is close enough that they are relevant to determining whether the claimed technology is new and nonobvious. 87 Recent studies show that patent examiners provide a large fraction of the cited references. During the 2001-2003 period, for example, examiners provided 67 per cent of all citations. Indeed, in 40 per cent of the patents granted, all citations were provided by examiners. 88 Because such a large fraction of ref-erences are provided by patent examiners, and another large group by patent attorneys, citations do not necessarily indicate a direct flow of knowledge and thus we treat them only as indications of technological relationships.

p. 33

Patents and their citations form a directed network, meaning that citations go from later patents to earlier patents and not in the opposite direction, in which patents are the network nodes and citations are directed links. Citations convey valuable information about the relationships between the technologies covered by the citing and cited patents. One can thus view the patent citation network as a kind of map of the space of patented technology, indicating the relationships between various pieces of "property" in that space. 89 As discussed in Section III.B, the evolution of the network may help to illuminate whether the PTO is awarding patents for more trivial technological steps.

p. 33

While the precise significance of a patent citation varies, a citation sometimes indicates that the claims of the cited patent encompass the claims of the citing patent and that a blocking patent situation exists. Consequently, one must obtain permission from both patent owners in order to use the invention claimed in the citing patent. As will be discussed in Part V, we believe it is likely that one can mine the structure of the patent citation "map" for signatures of patent thickets, in which there is a high density of overlapping patent claims, so as to test, for example, whether such thickets are increasingly prevalent in the patent system.

Relationship of this Study to Previous Uses of Patent Citation Data

p. 33

The United States patent system provides a historical record that encompasses much of the history of innovation in this country and, increasingly, abroad. Records of patent prosecution, patent citations, and patent litigation over many years are publicly available. The history provided by patent records is incomplete, of course. Many technical advances are either unpatented trade secrets or unpatentable technical know-how. Moreover, the patent record was historically limited to traditionally industrial innovations, though the scope of patentable subject matter has expanded to cover an increasingly broad range of innovations at an increasingly early stage of development. Whatever its limitations, the patent record has long been recognized as a rich source of data about innovation and innovation 89. Of course, this map is neither perfect nor complete. Examiners and applicants may miss relevant connections between patents, cite particular patents because they are familiar, and so forth. The analysis here assumes only that citations generally indicate technological relationships between citing and cited patents. policy. 90 In the past, the form in which the data were available made quantitative analysis a difficult and painstaking process and limited the kinds of analyses that could be performed. Recent advances in computer technology, however, along with the efforts of empirical economists, have rendered a wealth of patent data suitable for large-scale quantitative analysis. 91 A surge of attempts by economists, social scientists, and legal academics to capitalize on this new availability has resulted. 92 For the most part, that work has applied statistical regression techniques to connect patent characteristics-such as the number of citations made and received, the number of patent claims and whether the patent was renewed; inventor characteristics-such as geographical location and employment context; and, most recently, patent examiner characteristics-such as length of experience-to financial, economic, and technological indicators. In this way, commentators have used patent data to investigate knowledge flows and spillovers; 93 to attempt to determine the characteristics of valuable patents; 94 and to try to understand the role innovation plays in the behavior of various types of institutions, from universities to small and large firms. 95 Because of its heterogeneity and local structure, the network of patents and citations is a much richer source of information about the patent system and the associated technological development than is generally captured by these techniques. Network science analysis of the patent citation network has great potential both to complement existing econometric studies of patent citations and to advance the understanding of networks in general because of the fact that the patent citation network is one of the largest and most completely characterized networks available for study. Though legal scholars and economists have generally not employed a network approach, 96 a few studies of innovation by social scientists have applied social network analysis techniques to patent data. These studies, some of which we describe in a bit more detail in Part IV, view patents as footprints of innovation. They interpret the pattern of patent citations to indicate the combination of past technologies into new innovations and use the patent citation network to investigate theories of innovation as recombinant search. 97 These theories of innovation should be of interest to legal scholars since they move beyond a linear model of sequential or cumulative innovation and attempt to use network concepts and measures to incorporate some of the complexities of the innovative process. Dealing with inventive combinations is a critical and ill-understood aspect of patent law, as demonstrated by the difficulty in finding an effective approach to nonobviousness.

p. 36

Our treatment of patent citations differs from that of a number of previous studies because we attempt to make minimal assumptions about the significance of individual citations. Many previous studies have assumed that when a patent cites another patent it is an indication of knowledge flow from the cited patent to the inventors of the citing patent. 98 This assumption is questionable in light of data, only recently available in electronic form, showing the extent to which citations are inserted, not by the patent applicant, or even by the applicant's attorneys, but by patent examiners long after the time of invention. 99 Our approach to patent citations is parsimonious and limited to an assumption that citations generally indicate significant technological relationships between patented technologies. We do not attempt to probe knowledge flows.

p. 36

Finally, this study differs from many econometric studies because we make no assumptions about the distribution or functional form that describes the citation data. Econometric studies using patent data have been criticized due to their reliance on such assumptions because citation data are highly skewed and far from the "normal distributions" which are typi- cally used in statistical analysis. 100 Thus, one must interpret those econometric analyses with particular care. 101 Network science approaches avoid making such assumptions and are tailored to highly heterogeneous systems, in which broad, skewed distributions are typical.

B.

p. 37

The Evolving Patent Citation Network: Is Patenting Getting Out of Hand?

p. 37

Commentators have widely remarked that patenting has burgeoned since the 1980s and almost equally widely alleged that patenting is getting out of hand-that patents are increasingly of low quality, providing the transaction costs of a divided and proprietary knowledge base without the benefit of spurred innovative progress. In fact, the number of patents issued by the USPTO has increased more or less exponentially since the patent system was inaugurated in 1790, and it is true that the rate of increase, which had been more or less constant since 1870, sharpened noticeably in the early 1980s. 102 An increase in patenting is an ambiguous signal, however. The early 1980s was a time of both technological ferment and changes in patent law. Both the computer revolution and the biotechnology industry got their starts at around this time. Moreover, the Supreme Court in 1980 and 1981 issued its opinions in Diamond v. Chakrabarty 103 and Diamond v. Diehr, 104 putting its stamp of approval on the patentability of biological materials and computer software, respectively. In 1983, the Federal Circuit published its first patent opinion, 105 inaugurating what many have argued has been an era of patent-friendly legal review after a period of purported judicial hostility to patentees. 106 Any evaluation of the significance of the patenting boom must thus attempt to distinguish among potential causes, which encompass at least three possibilities. Increased patenting might stem from a faster pace of technological change, from a broader range of patented technology (which may have resulted both from the extension of the scope of legally patentable subject matter and from technological advances that have broken new ground), or from a weakened legal patentability standard. In this Section we use a network approach to separate these effects. Our results indicate that not only have the pace of innovation and breadth of patented technology changed in recent years, but the pattern of patent citations has also changed. The change in citation patterns indicates an increasingly skewed distribution of citability for patents issued beginning in the late 1980s. If citability reflects technical value, then the increasing gap between the most and least citable patents reflects an increasing gap between most and least valuable patents. While a definitive interpretation of this increasing gap is not possible without further investigation, we argue that this change is not likely to be due entirely to an increased pace and scope of patentable innovations. Instead, the increasing gap between most and least citable patents suggests that the USPTO may be issuing patents on comparatively more trivial advances. The timing of the increase in citability stratification is suggestive of an association with the Federal Circuit's increasing reliance on the "motivation or suggestion to combine" test for nonobviousness, which arguably has lowered the legal standard of patentability. It is also possible, of course, that patents are being issued on more trivial advances because inventors are applying for patents more readily. If the perceived business value of patents increases generally, the cost-benefit balance might shift toward patenting of less significant inventions. In other words, an effectively lower patenting threshold might reflect increased propensity to apply for patents rather than weakened standards for issuing patents. The legal patentability standard places a lower limit on the extent to which businesses can simply "choose" to patent more trivial inventions, however.

p. 38

In this Section we describe the results of our study of the evolution of the pattern of citations in the patent citation network. In analyzing the patent citation network, we have used the Hall, Jaffe, and Trajtenberg dataset, 107 which includes the approximately 16 million citations made by the more than two million patents issued by the USPTO between 1975 and 1999, as well as updated citation data from the USPTO that extends to 2006. 108 We begin by discussing the effects of innovative pace and breadth on the network's evolution and explain why these effects are not sufficient to explain what is happening in the network. We then turn to the heart of our analysis, in which we demonstrate that since the late 1980s patents have become increasingly stratified in their citability. We interpret this increasing stratification as evidence of an increasing gap in technical value between the most and least valuable patents. We then argue that this in-107. Hall, Jaffe & Trajtenberg et al., supra note 91. 108. Our analysis includes all of the patents and citations in the database. We do not randomly sample the data. creasing value gap is likely to be due to a lowering patentability threshold and that changes in the nonobviousness standard may be responsible.

Concerning the Pace and Breadth of Innovative Progress

p. 39

Before describing our analysis of the changing pattern of citations, which we argue suggests a lowering of the threshold of patentability, we consider how to separate out the effects of either a faster pace of technological advance or a broader landscape of patented technologies. We certainly do not deny that the pace of technological progress may be increasing and that the breadth of patented technology has expanded. However, we do not aim to investigate those effects here. Rather, we want to find out whether additional important changes in patenting may be occurring.

p. 39

For that reason, in our analysis we measure time and patent age in units of patent numbers, rather than in units of months and years. We do this to mask out any effects of a mere speed-up of technological progress. If innovation is simply occurring at a faster pace without any change in the innovative step between patents or the character of the relationships between patents, then measuring age and time in terms of patent numbers should cancel out the effects of the increasing pace. After canceling out this source of time-dependence, remaining changes that we observe must be due to some more fundamental change in the way in which the citation network is evolving.

p. 39

Another way in which the patent citation network has undoubtedly changed in recent years is through a broadening of the technologies covered by patents. To understand this issue, imagine a "landscape" of patented technology. As more patents are issued, they might overlap the existing "landscape" by claiming a somewhat improved or different version of a well-established technology-such as coffee filters or disposable diapers, for example. Alternatively, they might strike out into new territory by, for example, being the first patent to claim a practical application of a scientific advance, such as stem cell research. While this is surely an oversimplified view of the complexities of technological advancement, it is perhaps sufficient to illustrate that there is a distinction between the size of the innovative step represented by a new patent (how much better is the claimed disposable diaper, for example) and the extent to which it represents a broadening of the technological landscape. We need to consider, therefore, whether the changes in the citation network that we observe can be explained by the broadening range of patented technology rather than by a lower patentability threshold.

p. 39

We have measured two quantities that are relevant to the question of to what degree the "space" of patented technologies is "spreading out" as opposed to "filling in." The first is the average number of citations made per patent. As shown in the top graph in Figure 7, the average number of citations made per patent has been steadily growing (in fact, nearly proportional to the number of patents, which is our measure of time) throughout the period of our data. This increasing number of citations made per patent seems unlikely to have been caused by mere expansion of the technological frontier. Patents that are breaking new ground by heading off in uncharted technological directions would, if anything, seem likely to encounter fewer material prior art patents to cite, rather than more. And the addition of patents in a new technical field seems unlikely to mean that patents in older fields will have to cite more patents. It seems unlikely, for example, that the addition of the field of biotechnology patents would require that patents on mechanical devices cite more prior art.

p. 40

On the other hand, the increase in citations made per patent is consistent with an increasingly dense space of narrow patents, with new patents having an increasing number of patents to cite because more and more patents are material. It is similarly consistent with patenting of smaller and smaller incremental advances, such that inventors need to cite more patents to reach back to all those which are relevant to the claims of the new patent. It would also be consistent with an increased issuance of patents that are based on combinations of older technology. Thus, the increasing number of citations made suggests that something is going on in addition to the broadening of the landscape of patented technology. Some criticisms of the recent evolution of the patent system, particularly those in the popular press, have highlighted another way in which the breadth of patented technologies might be expanding by providing examples of "outlier" patents, often of a somewhat humorous character, such as the much-discussed peanut-butter-and-jelly sandwich patents. 109 If the addition of such "outlier" patents were significantly expanding the landscape of patented technology, one would expect that the average citability of brand new patents would decline since such patents would be less likely to spawn improvements and follow-on innovation than patents on more mainstream technology. In fact, our analysis of the citation network's evolution finds that the average likelihood that a new patent will be cited by the next patent that issues has increased somewhat in recent years due to the increasing number of citations made by each patent, 110 suggesting that true technological "outlier" patents, while making for entertaining rhetoric, are not a major part of the patenting explosion. On the whole, patent examiners and applicants have deemed it necessary to cite more patents, suggesting that the density of "mainstream" patents is increasing despite the undeniable increase in the breadth of patented technology.

p. 41

Having accounted for a possible speed-up in innovative progress through our choice to measure time and age in units of patent numbers and having concluded that recent citation patterns are not likely to be explained solely in terms of an expansion of the breadth of patented technology, we now turn to the question of whether something else is going on in the patent system beyond an overall "speeding up" or "spreading out."

The Evolving Citability Distribution

p. 41

Our approach is motivated by statistical physics studies of a diverse range of other growing networks. We describe the evolution of the patent citation network in terms of the probability that a patent with given characteristics will be cited, which we call "citability." To get a result that will provide insight into the general features of the network evolution, we attempt to describe the probability of citation using only those characteristics that strongly affect the likelihood that a given patent will be cited. Because a case-by-case evaluation of the underlying reasons that one patent might cite another is impossible for a large network of citations, we look for objective characteristics that we expect, on average, to be good proxies for the underlying citation process. In our study, we hypothesize that we will be able to describe the evolution of the patent network to a very good approximation by assuming that the probability that a particular patent will be cited at a given time depends primarily on its age, which we will call l, and on the number of times it has already been cited, which we will call k. 111 Our results bear out the hypothesis that, on average, these two characteristics are highly determinative of the likelihood that a patent will be cited.

p. 41

Our assumption that the probability that a patent will be cited depends on its age requires little explanation-technology tends to become obsolete. Our expectation that the probability that a patent will be cited de-110. See infra Appendix, Figure 7, bottom graph. 111. A more complete description of our data analysis procedure is available in our technical publication, Gábor Csárdi, Katherine Strandburg, László Zalányi, Jan Tobochnik & Péter Érdi, Modeling Innovation by a Kinetic Description of the Patent Citation System, PHYSICA A (forthcoming 2006), available at http://www.arxiv.org/abs/-physics/0508132 (last visited Nov. 13, 2006). pends on how many citations it has already received derives both from our intuitions about patents and from experience with other evolving networks. As we discussed in Part II, one of the most interesting discoveries of network science has been that it is very common for nodes that already have many links to acquire additional links more quickly than nodes with fewer links. One might reasonably expect this preferential attachment or, colloquially, "rich get richer" phenomenon to occur in the patent citation network.

p. 42

There is considerable statistical evidence that highly cited patents are more valuable and may be of greater technological merit. 112 Moreover, technology has its own "popular crowd," depending on what field is "hot" at a particular time. In our analysis, we thus assume that the number of previous citations to a patent, which we call k, will be relevant to its likelihood of being cited again. Our computations confirm that the likelihood of being cited increases as the number of previous citations increases, as we would have guessed, but also permit us to extract empirically the detailed structure of that increase, as we explain below.

p. 42

Of course, the fact that we extract a citation "probability" is not meant to suggest that the particular citation choices made by patent examiners or applicants are actually random. Our results mean only that, cumulatively, those individual citation decisions result in a likelihood of citation which depends in interesting ways on k and l. 113 Patents presumably have inherent "quality" that affects whether they are cited, just as individuals have personalities that make them more or less popular and websites have content that make them more or less useful. The higher citability of highly cited patents presumably depends on some combination of inherent quality and the extent to which a particular patent is already known to the relevant examiner. Our present analysis says nothing directly about the inherent patent quality of any particular patent, but does allow us to make reasonable inferences about the patent system as a whole based on the observed citation distribution and the way in which it has changed in recent years.

See, e.g., Allison et al., supra note 13 and references therein (discussing studies of the relationship between patent citations and patent value).

p. 42

113. There is also an overall time-dependent scale factor, S(t) (where t is measured by patent numbers). More specifically, we find that the probability function P(l,k,t) can be written to a good first approximation as the ratio of a time-independent function, A(k,l), and a time-dependent scale factor, S(t). The scale factor S(t) is just the sum of A(k,l) over all existing patents at time t. Thus, S(t) changes over time only because the number of patents of age l and connectedness k changes. See our technical paper, Csárdi et al., supra note 111, for a detailed explanation of this factor.

p. 43

To find the likelihood of citation for given k and l we developed a novel iterative technique, which we used to extract from the patent citation data the average likelihood that a patent with k previous citations of age l will be cited again. We do not assume a particular functional form for this probability-we derive the functional dependence on citations previously received and patent age directly from the data. The likelihood that a patent will be cited depends on both its age and on the number of citations it has previously received. Figure 9 shows the dependence of citability on age. Very young patents have low citability (A l (l) is small); the likelihood of being cited grows rapidly in the first couple of years after a patent issues, peaks, and then it falls slowly thereafter. As it turns out, however, even though the probability of being cited depends on age, the way in which the probability of being cited depends on the number of citations previously received is more or less the same for patents of any age. 114 No matter the age of the patent, the ratio of the likelihood that a patent with ten previous citations will be cited to the likelihood that one with three previous citations will be cited is nearly the same. So if, for example, patents that are 10,000 patent numbers "old" are five times more likely to be cited if they have ten previous citations than if they have three, roughly the same will be true for patents that are 100,000 patent numbers "old."

p. 43

This means that we can look separately at the age dependence of the citability and the dependence on number of citations already received. 115 In this Section, we focus on the dependence on citations received, k. In Part IV, we discuss some of our results for the age dependence of citability.

p. 43

Figure 8 shows the way in which the likelihood of being cited depends on the number of citations already received, which we call A k (k). Figure 8 also displays the results extracted from the data for patents of several different ages. The more often a patent has already been cited (the higher its value of k) the more likely it is to be cited again (the higher its value of A k (k))-the signature of preferential attachment.

p. 43

This demonstration of preferential attachment in the patent citation network is not especially surprising-preferential attachment is a common property of growing networks and is intuitively sensible in the patent system. Preferential attachment is cumulative-highly cited patents are more likely to be cited, hence becoming even more highly cited and even more 114. See infra Appendix, Figure 8. 115. Mathematically, this means that the likelihood of being cited can be written approximately as a product of two functions, i.e., A(k,l) = A l (l)A k (k). A l (l) depends only on age while A k (k) depends only on the number of citations previously received. likely to be cited, and so forth. The pattern of citability that Figure 8 shows thus eventually leads to an extremely skewed distribution of citations eventually received by patents. Most patents are hardly cited at all, while a few patents become citation "billionaires" (well, "hundredaires," really). 116 If we make the reasonable (and empirically supported) assumption that the number of times a patent is cited signals its technological value, 117 we can infer that the very skewed and broad citation distribution in Figure 1 reflects a very stratified distribution of patent values, with a few superstars and a vastly larger number of patents that go nowhere.

p. 44

This general picture of a highly skewed patent value is well known by now. 118 Our network analysis allows us to get behind this general observation, however, to ask just how stratified patent value is and how the extent of stratification has evolved over time. If more patents are issuing simply as a result of faster or broader technological progress, we would expect the degree of stratification to remain about the same over time. On the other hand, if the patentability standard decreases, there would be not only more patents issued, but a higher proportion of them would be less important; the degree to which highly citable patents dominate trivial patents should increase.

p. 44

The citability function, A k (k), gives us a quantitative handle on the degree of stratification among issued patents. We arrive at this function after factoring out the effects of patent aging due to obsolescence (which, as noted, turn out to be roughly the same no matter how may citations a patent has previously received). The function thus gives a relatively simple and unobscured view of the underlying citability of patents. Looking at Figure 8 we see that the citability of a patent on average increases somewhat more than proportionally to the number of citations already received. We can quantify this "somewhat more" by noting that A k (k) is closely fit by the form A k (k) ~ k α . The parameter α is a measure of the extent to which highly cited patents are preferred. When averaged over all of the data from 1975 on, α ≅ 1.19 (with an estimated error of less than .01).

p. 44

To give us a feel for the interpretation of the values of α, we can compare the patent system's α value with known α values of other network studies. If there were no aging or obsolescence, nothing would inhibit old patents from acquiring extremely large numbers of citations. In that situa-tion, theoretical network models have shown that when α is larger than 1 the network "condenses" to a highly unequal situation in which nearly all nodes in the system have very low connectivity. These "unimportant" nodes are all connected to a very small (finite even in an infinite network) number of very highly connected nodes 119 -almost all nodes are "peons" connected to a very few "royal" nodes and there is no "middle class" of moderately connected nodes. In the patent system, such an extreme stratification does not occur because highly cited patents eventually become obsolete and generally cease to be cited. However, the fact that α > 1 in the patent network still suggests that citations are highly concentrated. There is a highly unequal distribution of patent citability-and hence, most likely, of patent value.

The Increasing Stratification of Patent Citation Patterns

p. 45

The network analysis not only shows us that patent citability is stratified, but also gives us a means to investigate whether the degree of stratification has varied over time. We find that patent citability has been becoming increasingly stratified since the late 1980s. We investigate this phenomenon by calculating α using only the patents within a 500,000patent sliding time window and calculating a value for α after every 100,000 patents. The value of α-and hence the degree of stratification of patent citability-has varied in an interesting way. As Figure 10 shows, the stratification of citability, as reflected in the value of α, began to rise in the late 1980s and has continued to rise throughout the period of our study. The rise followed a period of slightly decreasing stratification. (We do not know what was happening before 1982 because we do not have sufficient earlier data.) Figure 6 shows that the number of patents issued annually has been rising essentially since the inauguration of the patent system (and rising very rapidly since the early 1980s), Thus, during a period throughout which the absolute number of patents issued was rising rapidly, the relationships between those patents were also changing-but not in a way that simply reflects increasing numbers. Patent citability became first less and then increasingly stratified. The increase in αcorresponding to increasingly stratified citation patterns-began nearly 10 years later than the beginning of the recent rise in patent issuance and after the major change in the patent system represented by the inauguration of the Federal Circuit. How should we interpret the increasing stratification of patent citability since the late 1980s? There are several possibilities. One possibility is that the legal patentability standard has decreased, resulting in the issuance of a larger fraction of more trivial-and hence less citable-patents. This explanation is consistent with growing societal concern with "low quality" patents. 120 A second possibility is that the proportion of highly citable pioneer patents has increased. A third possibility is that the change in stratification parameter reflects changes in the subject matter mix of patented technologies. Finally, it is possible that the change in citability stratification reflects some kind of change in citation practices, rather than a change in inherent patent characteristics.

p. 46

The increasing stratification of patent citability is correlated with increasing concerns with patent quality. Anecdotal and survey evidence suggests that patent quality has been decreasing in recent years, resulting in the issuance of a larger fraction of more trivial-and hence less citable-patents. One possible reason for the issuance of lower quality patents is a lower legal threshold of nonobviousness. At around the time that patent citability stratification began to increase, the Federal Circuit (followed of course by the USPTO and, presumably, by patent applicants themselves in making their filing decisions) increasingly adopted the "teaching, suggestion, or motivation to combine" test for nonobviousness 121 While the effect of the "suggestion test" on patent issuance is a subject of debate 122 (and the test is, as already mentioned, the subject of a pending Supreme Court case 123 ), there are doctrinal reasons to believe that it is a systematically lower standard of patentability than the previous legal standard. It is thus plausible that the increasing stratification reflects a lower threshold of nonobviousness resulting from increasing application of the "suggestion test."

p. 46

While we believe that a weakening patentability standard is the most likely explanation for the increasing stratification of patent citabilityand, inferentially, the increased disparity in patents' technical worththere are other possible explanations. Perhaps the increasing stratification is due not to an increasing issuance of trivial patents, but to an increasing issuance of more highly citable pioneer patents. An increase in pioneer patents could result from increased patenting of upstream research results, such as has been occurring, for example, in the field of biotechnology. A third possibility, which we plan to explore in future studies, is that the average stratification parameter α is changing because of changes in the subject matter mix of patented technologies. If patent "importance" is inherently more stratified in one field of technology than in another (because of a difference in the importance of "pioneer" patents, for example), then an increasing prevalence of patents in that field could change the average degree of stratification that we observe. Preliminary studies of a rough division of patents into six technological categories did not turn up any significant variations in α, but a more in-depth study of differences between technological fields is under way.

p. 47

A final possibility is that the change in citability over the years reflects a change in citation practice, rather than a change in inherent patent characteristics. This possibility is especially interesting because the timing of the increased stratification in the late 1980s corresponds to the time at which computerized searching became increasingly prevalent. We are inclined to reject this possibility at present because most of the trends in patent citation practice that we can think of-most notably the increased ease of computerized searching for prior art-seem unlikely to have changed direction from decreasing α to increasing α during the 1980s. Computerized searching seems likely to have had a one-way influence. One way to check for the influence of search technology would be to compare the behavior of the United States patent citation network with other citation networks, such as the European patent citation network or the network of citations in scientific journals.

p. 47

To summarize, in this Part we have shown empirically that the distribution of patent citability has been changing in recent years. Prior to the late 1980s, citability was becoming slightly more egalitarian-the difference between the citability of the most highly cited patents and that of less cited patents was decreasing. From the late 1980s on, that trend has been reversed. Citability has become more stratified, with highly cited patents becoming more and more citable compared to less cited patents. In line with our intuition and with earlier studies of patent value, we interpret citability as a reflection (on average) of technological importance. We thus conclude that the distribution of patent importance has also become more stratified since the late 1980s. There are several possible explanations of this trend. We hypothesize, consistent with more general evidence of decreasing patent quality, that the increased stratification may result from a lowered patentability threshold, possibly reflected in the "teaching, suggestion, or motivation to combine" test of nonobviousness, and results in the issuance of more trivial patents.

IV. OF SLEEPER PATENTS AND INNOVATIVE PROGRESS: HOW INNOVATION BUILDS ON PRIOR TECHNOLOGY

p. 48

We now turn to our results for the age dependence of the likelihood that a patent will be cited. Our study of how patent citability depends on patent age contributes to understanding the way in which innovation proceeds by building on prior technology. Most studies of the process of innovation reflected in the legal and economic literature have, to render the analysis tractable, taken a simplified, linear view in which innovation is pictured as sequential or cumulative, each new advance building on a previous innovative step. In fact, innovation is a much more complicated process involving new advances, combinations of old technologies, and potentially reaching back into the past to make new uses of technologies that acquire new relevance in light of some new advance. 124 Our results for how citability depends on age provide a window into this process.

p. 48

As already mentioned, in the course of our analysis of patent citability, we determined how patent citability depends on a patent's age (in patent numbers). We noted that the way in which a patent's citability varies with the number of citations it has already received is roughly the same for patents of all ages. Correspondingly, the way in which patent age affects the likelihood of being cited is nearly the same no matter how many times a patent has been cited before. Patents that have rarely been cited have a lower overall chance of being cited than highly cited patents, but the "age profile" of their chance of being cited is about the same. Mathematically, the dependence of citability on age, which we call A l (l), is relatively independent of k, especially at large l and k, as shown in Figure 9 for several values of k.

p. 48

The observed dependence of citation probability on patent age provides insight into the complicated dynamics of technical innovation. Not surprisingly, citability peaks at a relatively young age (small l). More interestingly, citability decays unexpectedly slowly for older patents. 125 For large l, A l (l) ~ l -β , where β ≅ 1.6. This power law form signifies a long, slow decay-very old patents are still being cited. The peak in citations relatively soon after issuance likely corresponds to "typical," incremental improvements, such as those reflected in linear models of cumulative in- 125. Previous studies of citation lags have also noted that old patents continue to be cited. See, e.g., Hall, Jaffe & Trajtenberg, supra note 91, at 421-24, 448-51. novation. Such incremental advances are likely to occur relatively soon after a patent issues. The tail of citations that occur long after issuance no doubt includes "pioneer"-type patents that continue to influence innovation (and hence continue to be cited) over long periods of time. However, not all of these late citations are to broad, pioneer patents. The form of A l (l)-including the tail of citability long after issuance-is the same even for patents that have rarely been cited in the past. This means that even patents that have been cited very rarely or never in the past are sometimes cited long after issuance.

p. 49

Thus, there is no age at which patents can be pronounced "dead." While most patents that have not been cited shortly after they are issued will never be cited, there are "sleeper" patents that go without citation for long periods of time, only to reawaken at a later time. If one makes the reasonable assumption that receiving a citation is some indication of social value, this observation suggests that inventive progress is not simply a matter of steady accumulation of incremental progress, but a complicated process in which old patents may gain new significance in light of later advances. The long tail of inventive relevance also underscores the difficulty of predicting the social value of innovations in advance.

p. 49

A network approach permits us to take into account the multidimensional and combinatorial nature of technological progress. Our age profile is just one simple way to use citation network analysis to probe the innovative process. Social scientists have also begun to use network measures of patent citations to explore the way in which technological change occurs. 126 From this perspective, one may view citations as indicators of the ways in which prior technologies have been combined to produce a new invention. In this vein, Fleming, Sorenson, and their collaborators 127 used patent citations in combination with the USPTO classification system to explore the way in which innovation proceeds as a process of search and recombination of prior technologies. They argue that successful innovation is a balance between re-using familiar components-an approach that is more certain to succeed-and combining elements that have rarely been used together-an approach that is more likely to fail entirely, but also more likely to result in radical improvements.

p. 50

Podolny and Stuart and collaborators also made explicit use of social network analysis by using a basic network measure of local network structure, the transitivity, to study the innovative process. 128 Transitivity measures the likelihood that two nodes that are connected to a specific third node are also connected to one another. Figure 12 illustrates this concept. In the patent context, transitivity measures the likelihood that two patents that cite or are cited by the same patent also cite one another.

p. 50

Podolny and collaborators used transitivity-type quantities to define local measures of competitive intensity and competitive crowding based on indirect patent ties. 129 A basic insight of their work is that if many innovations are building on the same technological antecedents, and thus contributing very similar technological outputs, the likelihood of a new entrant into the associated niche is lowered. They used network transitivity measures to observe how a technological niche can become popular, but then overly crowded and "exhausted" as a result of a flurry of inventive activity in the niche. They point out that counting the numbers of citations made or received is not sufficient to identify crowded niches-an innovation that provides a technological foundation for a variety of unrelated advances will be highly cited without crowding. Transitivity measures thus show promise as means to distinguish patents in hot fields from broad, important patents.

p. 50

Studying network metrics and their evolution over time thus allows us to delve into the way in which inventions are combined, recombined, and re-used to produce new innovations. In the future, we can build on our work and the social science studies discussed above by combining the approaches. For example, we might investigate the extent to which delayed citations in the power law tail of the A l (l) function target patents with high "originality" (low transitivity) which connect disparate technologies.

V. FUTURE DIRECTIONS AND SOME PRELIMINARY RESULTS

p. 50

The network approach has the potential to elucidate other questions of patent and innovation policy, some of which we discuss briefly here.

A. Patent Classification and the Meaning of Analogous Arts

p. 51

The categorization of patents by technological field of endeavor is of both practical and analytical significance. Patent classification plays an important role in prior art searching. Moreover, under the patent doctrine of analogous arts, the technological field of endeavor determines which prior art patents one must consider in determining whether a claimed invention is nonobvious. 130 Though the law requires that novelty be absolute, courts judge nonobviousness only with respect to prior technology in analogous arts of which a person having ordinary skill in the art would reasonably be apprised. To apply this standard, it is necessary, of course, to determine which fields of technology are analogous to the field of the patented invention and which patents are in a relevant art. Moreover, any study of technological innovation, or of the functioning of the patent system, which relies on patent data and seeks to inquire into potential differences between fields of technology must find some means of categorizing patents.

p. 51

At present, a USPTO examiner categorizes each patent application according to a scheme of classes and subclasses that has developed in an ad hoc manner over the years. The classification assists examiners and patent applicants in searching for relevant prior art. Researchers have also used USPTO patent classifications to assess such things as the generality of a patent (evaluated by the extent to which it is cited by patents from different classes), the originality of a patent (evaluated by the extent to which it cites patents from different classes), the technological closeness of firms involved in patent litigation (evaluated by the extent to which the patents assigned to those firms are in the same classes), and the way in which innovations recombine components from the prior art. 131 The usefulness of these measures is limited, however, by the fact that the classification scheme is ad hoc and does not provide a well-defined way to measure the degree of distinctiveness between inventions in different subclasses or classes. A more objective and quantitative measure of patent "closeness" is desirable.

p. 51

Network analysis provides an alternative way to measure technological distance that is more fine-grained and quantitative than comparing USPTO classifications and less ad hoc than the classifications themselves. The number of citation "hops" that it takes to get from one patent to another is one way of measuring how "far apart" two patented technologies are. Essentially, patented technology is likely to be related most closely to the technology of the patents it cites or is cited by directly, less closely to the technology of the patent that is two steps away, and so on. Path length measures may provide a more fine-grained (and complementary) measure of technological closeness than metrics that use the number of overlapping subclasses in the USPTO classification system.

p. 52

One may also use the network path length to evaluate and explore existing patent classification schemes. One may compare the distances between patents in the same USPTO category with distances between patents in different categories to evaluate the accuracy of the classification system. Moreover, one can compute the average distances between different categories to get a more qualitative measure of the relationships between different categories.

B. The Shrinking Patent Citation Network

p. 52

In a preliminary study, we have measured the overall "size" of the patent citation network by looking at the shortest path lengths between patent nodes. As discussed in Part II, one of the most interesting results of network science studies is that many networks have what has come to be called the "small world" property. 132 This property, made famous by the phrase "six degrees of separation" 133 and by the "Kevin Bacon game," 134 measures how closely knit a network is. To determine whether a network has the small world property, one counts how many "hops" from node to node along the network one must make to travel between any two nodes. The number of hops along the shortest path between two nodes is sometimes called the "geodesic distance." The maximum geodesic distance between any two nodes in the network is the "diameter" of the network. If the maximum geodesic distance between any two nodes in a network re- 134. The Kevin Bacon Game, which may be played at http://www.cs.virginia.edu/--oracle/, counts any two actors as linked if they have had roles in the same movie. The "Oracle of Kevin Bacon" shows that any two of the 800,000 or so actors in the Internet Movie Database are linked to Kevin Bacon in an average of fewer than 3 steps. mains relatively small, 135 the network has the "small world property." The diameter of many real world networks is surprisingly small. 136 We have determined so far that the patent citation network displays the small world property discussed in Section II.A.1.b. The patent citation network is not completely connected. There are a few isolated patents and clusters. However, the network contains a "giant" component which contains the vast majority of patents. Because citations are directed links-a particular citation goes from one patent to another and not vice versa-one can measure the small world property in two different ways. We have measured the small world property using directed paths, meaning that we permit only hops from citing to cited patents. The longest directed geodesic distance between nodes in the "giant" component (of nearly 4 million patents) is only 24 steps, 137 while the average is 6.9 steps. It is also possible to investigate the small world property of the undirected citation network, in which hops are permitted either from citing to cited patents or from cited to citing patents. An exact calculation of the diameter of the undirected network is computationally difficult because of the large size of the patent citation network. An approximate calculation on the giant component of the undirected patent citation network has been performed by Leskovec et al. 138 They determined, using a sampling method, that the average undirected geodesic distance is about 10 steps. These two calculations measure somewhat different quantities, but the qualitative conclusion is the same-the patent citation network is very closely connected. 139 Lescovec et al. also determined that the "size" of the patent citation network has been shrinking over time, even though the number of patents has increased dramatically. This shrinking diameter is not typical of small 139. The fact that the average undirected path length is longer than the average directed path length is somewhat counterintuitive. This occurs because there are pairs of patents which are not connected at all by a directed path. The directed path length between such patents is effectively infinite, but these infinite path lengths are not included in the average of 6.9 steps which we report. Once undirected hops are permitted, some of these nodes are connected, but the path lengths required to connect them may be long enough to raise the average. There are also difficult questions about how best to sample a network to determine average properties. world networks, which generally have very slowly increasing diameters, and may indicate increasing interdisciplinarity (interdisciplinary patents will cite patents in disparate fields, thus providing a network "shortcut" between disciplines) or the increasing importance of "bridge" technologies which are used in a number of fields. We plan to study the distribution of directed and undirected path lengths in more detail and, if the shrinking "size" is confirmed, will combine path length measures with the USPTO classification scheme to understand the source of the closer connectedness of the network.

p. 54

The small world property of the network also suggests the intriguing possibility that a relatively small number of "hub" or "bridge" patents are integral to many apparently separate fields (and thus provide "shortcuts" between patents in disparate fields). In future work, we hope to investigate that possibility and to study qualitatively any "hub" or "bridge" patents that we can identify.

C. Patent Thickets and Potential Anti-Competitive Licensing Practices

p. 54

The proliferation of patents results in a situation in which many commercial products require the use of more than one-and in some cases hundreds of-patented technologies. This situation raises a number of sometimes conflicting concerns. On the one hand, there is the fear of a patent thicket in which the transaction costs associated with obtaining the necessary patent licenses to do something of practical usefulness become so high as to undermine the social value of the patents. 140 There can even be a hold-up problem in which the owners of contributing technologies cannot come to an agreement at all, as a result of differing assessments of the relative value of various contributions to a commercial product. 141 There are two distinct ways in which such a transaction costs nightmare could arise. It might be that a commercial application requires many distinct and complementary patented technologies. This could happen either as a result of increasing complexity in the technology itself or as a result of patenting more and "smaller" pieces of technology space. A need for many patent licenses might also be due to patent blocking. Patent blocking occurs when a patent improves on the invention claimed in another patent in such a way that the claims of the earlier patent still cover the improved invention. Where patent blocking occurs, improvers may require licenses to a string of "nested" improvement patents to use one 140. See, e.g., Shapiro, supra note 64; Bessen, supra note 77. 141. See, e.g., Merges, supra note 78, at 84-91. particular technology. In either case, one way to handle the necessary cross-licensing and cut down on transaction costs is for the owners of the relevant patents (who are often commercial entities in the same industry) to form patent pools. 142 A high density of patents in a particular technological "niche" need not always indicate a patent thicket, however. Closely related patented technologies may be potential substitutes for one another. This situation creates something more like patent supermarkets, offering many nearly interchangeable options, than patent thickets. If these patents are separately owned, competition between patent holders will reduce licensing fees and the issue of hold-up will not arise. In situations like this, in contrast to the patent thicket situation, patent pools and cross-licensing agreements between industry actors may be anti-competitive. Rather than serving to reduce transaction costs and avoid hold-up, such agreements may permit firms to avoid competing for licensing revenues and drive up royalties.

p. 55

Distinguishing between these two kinds of relationships between patents is of great theoretical and practical interest. Citation network structural measures may be useful in addressing questions related to the patent thicket issue. Prof. Gavin Clarkson has suggested that a high density of citations (meaning the fraction of possible citations that are actually made) in a particular patent neighborhood may suggest where patent pools might legitimately form. Setting aside some technical issues about how to compare citation densities from different sized network samples, citation density is a useful way to assess the level of technological interrelatedness in a particular field. 143 Density alone cannot distinguish between patent thickets and patent supermarkets, in which closely related technologies compete for licensing customers, however. The mere existence of a citation from one patent to another cannot tell us which scenario is most likely. Determining the meaning of a particular citation is an extremely labor-intensive process requiring understanding of the legal and technical relationship between the citing and cited patents. To investigate the existence of patent thickets and the potential for related antitrust problems, some structural metric that is sensitive to the character of a citation is highly desirable. It may be possible to design such a metric based on transitivity concepts.

p. 55

Podolny and Stuart note that a technological tie (and hence a likelihood of a citation) exists between two patents if "the contribution of [the second] incorporates, builds on, or is bounded by a technological contribu-142. See, e.g., Shapiro, supra note 64, at 123; Goller, supra note 79, at 727. 143. Clarkson, supra note 6.

p. 56

tion of [the first.]" 144 The distinction between these relationship types is critical to understanding the extent to which increased patenting in a particular area signifies patent blocking and high-transaction-cost thickets or desirable designing around and competition.

p. 56

As we alluded to above, patents may cite an earlier patent because they build on its technology (what might be termed "lovely" citations) or because they replace or distinguish its technology ("dangerous" citations). 145 Depending on the scope of a patent's claims, a lovely citation may mean the owners of both patents must authorize that use of the later technology. A dangerous citation, on the other hand, may mean that a prospective user of the related technologies can play them off against one another in bargaining for a license. The existence of lovely and dangerous citations is thus related to the potential for patent thickets and for antitrust problems when industry competitors sign cross-licenses or form patent pools.

p. 56

We suggest that a patent C that cites a patent B which makes a lovely citation to an earlier patent A is relatively likely to cite the earlier patent as well. 146 This is because for a lovely citation, the claims of A are likely to be broader in scope than the claims of B, and patent C must therefore be distinguished from both A and B to be patentable. If the citation is dangerous, on the other hand, one may not need to distinguish patent C from patent A. Thus, on average, we would expect a group of patents in a hightransaction cost thicket to have a higher value of the variant of transitivity illustrated in Figure 12c. We plan to investigate whether the Figure 12c transitivity will provide insight into the extent to which patents in a given technical field tend to be competing substitutes or blocking patent complements. If we are successful in identifying a structural measure of this kind, it may be both of analytical use in understanding the prevalence of blocking patents in particular technical fields and of practical use in evaluating the potential anti-competitive effects of patent pools and crosslicensing agreements.

VI. CONCLUSIONS

p. 56

This Article began by arguing that network science is poised to begin making important contributions to legal scholarship by offering relevant concepts, empirical methods, and modeling techniques that will be better able than current approaches to account for the importance of heterogeneity and local network structure in determining collective behavior. Be-cause networks are ubiquitous in the social problems to which the law addresses itself, network science, we argued, may make a significant contribution to legal analysis. Network science is particularly promising for dealing with situations in which local relationships are important and heterogeneous. It provides means to illuminate the complicated relationship between local relationship patterns and structure and interactions and global, collective behavior.

p. 57

In the second half of this Article, we illustrated how one can apply the network approach to study the patent citation network. In Part III we showed empirically that, not only has the number of patents been increasing rapidly in recent years, but citation patterns have been changing as well. The average citability of a patent increases very rapidly with the number of times it has already been cited, demonstrating the preferential attachment, or "rich get richer," phenomenon observed in many complex networks. Moreover, the extent to which highly cited patents are more "citable" than less cited patents has changed over time. Citability has become more stratified since the late 1980s. Since citability is likely related to a patent's technical importance, the increasing stratification suggests that the patentability standard may have decreased, resulting in the issuance of a larger fraction of more trivial patents. Neither a general increase in the pace of technological progress, nor a general broadening of patented technology seems to explain the increasing stratification of patent citability. In fact, the average number of citations that a patent makes has increased over time, suggesting that patents inhabit a more and more densely crowded technological space on average.

p. 57

The increasing stratification began in the late 1980s, a period which has no obvious connection with legal changes related to biotechnology, software, or business methods patenting, but is around the time that the "suggestion or motivation to combine" test for determining whether a claimed invention is obvious became increasingly established as the definitive Federal Circuit approach. 147 One possible hypothesis, based on our study so far, is that a weakening of the nonobviousness requirement, per-haps associated with the suggestion test, has given rise to a proliferation of patents on minor improvements over the prior art. These improvements typically have little impact on the development of the technology and are thus cited only for a brief period after issuance. We cannot rule out other explanations for the change in the pattern of citation network evolution, however.

p. 58

In Part IV we discussed how network-based analysis provides insights into the innovative process. For example, consistent with earlier statistical studies of citation lags, we found that the probability that a patent will be cited peaks at a relatively young patent age. But the citation probability also has a long, slow power law decay at older ages, suggesting that some patents retain their influence over very long times. 148 Surprisingly, we see this persistent vitality of some older patents even for patents which have received few or no citations-thus even "unpopular" patents have a significant probability of being revived after a long period of dormancy. This complicated dependence on patent age is evidence of the highly nonlinear nature of innovation and suggests that there may be two different innovation types. First, there is incremental innovation, for which the standard models of patents as incentives to investment may be relevant. Second, there is unpredictable innovation, which may be less susceptible to a simple incentive theory and more in line with "percolation" models of invention, in which incremental progress is coupled with complicated linkages back to previous technology. 149 Social scientists have begun to apply social-network-based approaches, looking at local network structure, to validate theories of innovation based on concepts of search and recombination. It should be possible to incorporate and build upon these studies in the analysis of innovation policy from a legal perspective.

p. 58

In Part V we reported some preliminary results of other network-based approaches to understanding the patent system. We suggested that the concept of network path length or "distance" may be a fruitful means to explore the connections between different technical fields. One could use a network measure of "technological distance" to evaluate the USPTO classification scheme and as an alternative and more quantitative way to classify patents and to evaluate concepts such as the patent law concept of analogous arts that identifies the prior art that courts and examiners must consider in assessing nonobviousness.

p. 58

148. This long tail in the probability of citation as a function of patent age is masked to some extent by the rapidly increasing number of younger patents in earlier studies which simply compute average distributions of citation lags. Our dynamic analysis disentangles these two effects.

p. 59

Another area in which a network approach shows promise is the evaluation of the extent to which increased patenting relates to the emergence of patent thickets and blocking patents. Network-based measures of local structure may be able to distinguish on average between patent thickets and areas of patenting of competing technologies. This approach may even help to distinguish between socially valuable patent pools and anti-competitive cross-licensing arrangements.

p. 59

In this Article we have provided only a small sample of the application of network science to legal problems. Much remains to be done in applying network science to the patent system and to other legal issues. Figure 4: Top: This figure illustrates power law or "scale-free" probability distributions of node degree for two values of the decay exponent. Unlike the normal distribution, these distributions are highly skewed and cannot be meaningfully characterized by a "typical" value. Depending upon the decay exponent, the mean and median values can be quite far from one another. Indeed, if the decay exponent is small enough, the mean (or average) value is infinite even though the most likely value is zero. The standard deviation is infinite for both of the exponent values shown. Bottom: The same distributions are shown on a "log-log" plot, in which power law distributions are straight lines. Fraction of Federal Circuit cases involving obviousness that referred to the "suggestion, teaching, or motivation to combine" test as a function of time. These numbers were obtained using LEXIS searches for "(suggestion or motivation or teaching) w/s combine" (to count references to the suggestion test) and either a reference to "obvious" in the headnotes (dashes) or at least 5 uses of the word "obvious" in the case (diamonds). These two methods of counting yield the same qualitative results showing an increase in use of the suggestion test throughout the 1990s.

Footnotes

See, e.g., sources cited supra note 2.
See infra Appendix, Figure 1.
BERKELEY TECHNOLOGY LAW JOURNAL[Vol. 21:4
58. See, e.g., Ayres & Baker, supra note 6, at 607-18; Matwyshyn,
75. See, e.g., JAFFE & LERNER, supra note 7, at 198-202. 76. See, e.g., FEDERAL TRADE COMMISSION, supra note 7 and sources cited therein. 77. See, e.g., Shapiro, supra note 64; James Bessen,