The overriding question in most federal trademark infringement litigation is a simple one: is the defendant's trademark, because of its similarity to the plaintiffs trademark, causing or likely to cause consumer confusion as to the true source of the defendant's goods? In answering this question, each circuit requires that the district court conduct a multifactor analysis of the likelihood of consumer confusion according to the factors set out by that circuit. As the Seventh Circuit has recently explained, the multifactor test operates "as a heuristic device to assist in determining whether confusion exists."' This heuristic device is the fulcrum of American trademark law, and yet, for all of its importance, the test is in a severe state of disrepair. Its current condition is Babelian. Each circuit has developed its own formulation of the test.' In the Eighth and Tenth Circuits, the test consists of six factors. 3 Other circuits, such as the Seventh, 4 typically use seven factors, 5 while still others, such as the Second 6 and the Ninth, 7 typically use eight. 8 The Third Circuit uses ten 9 and the Federal Circuit thirteen,'" of which the last is "any other established fact probative of effect of use."" While there is overlap among some of the factors used, there is also great diversitynot just in which factors are employed, but in how they are employed. Some circuits claim to weigh heavily under certain factors what other circuits claim to ignore, 2 and nearly every factor or combination of factors has been called the "most important" by one court or another. To make matters worse, scattered among the circuits are factors that are clearly obsolete, redundant, or irrelevant, or, in the hands of an experienced judge or [Vol. 94:1581 litigator, notoriously pliable. 3 This is a circuit split of considerable proportions, one which produces, as this Article will show, excessive intercircuit variation in the application and outcome of the multifactor tests.
Despite their critical importance, the multifactor tests have received little academic analysis beyond that of the treatise writers 4 and no empirical analysis.' 5 For their part, the circuits have not sought to harmonize their tests and the Supreme Court, despite having issued seven trademark opinions in the last twelve years, " has so far declined to intervene. Courts, commentators, and practitioners have all the while speculated about which factors, if any, actually drive the outcome of the test, how the factors interact, and most importantly, whether the different tests, given the same facts, would yield different outcomes. No convincing answers to any of these questions have yet emerged.
We lack knowledge of the multifactor tests because these are empirical questions that can only be answered through the use of empirical methods. This Article sets forth the results of an empirical study of all reported district court opinions for the five-year period from 2000 to 2004 that made substantial use of a multifactor test for the likelihood of consumer confusion. 7 In working from an original data set of 331 opinions, it pursues two related goals, the first specific to trademark law, the second broader and more theoretical in nature. First, this Article seeks to show how the tests actually work in practice-and to do so by using easily understood statistical methods. My hope is that the study's findings will enhance predictability and proficiency in the application of the circuits' current tests and serve as evidence in support of a national standard multifactor test.
Second, this Article draws upon recent social science learning on cognition and decision making to reflect more generally on the nature of legal multifactor decision making. This is an important inquiry not simply for trademark law, where so much of adjudication has taken the form of either statutorily-prescribed or judge-made multifactor tests, 18 but for a wide and diverse assortment of other areas of law as well, including copyright, 19 takings, 2 " evidence, 2 conflict of laws, 22 and criminal law, 23 to name only a 18. See, e.g., 15 U.S.C. § 1 125(c)(1)(A)-(H) (2001) (prescribing a nonexhaustive eight-factor test to determine whether a mark is "distinctive and famous"); 15 U.S.C. § 1125(d)(1)(B)(i) (2001) (prescribing a nonexhaustive nine-factor test to determine whether a defendant has a "bad faith intent to profit" from the use of a domain name "identical or confusingly similar" to plaintiffs trademark); Heartland Bank v. Heartland Home Fin., Inc., 335 F.3d 810, 819 (8th Cir. 2003) (setting forth a nonexhaustive seven-factor test to determine secondary meaning); Echo Travel, Inc. v. Travel Assocs., Inc., 870 F.2d 1264, 1267 (7th Cir. 1989) (same); E-Systems, Inc. v. Monitek, Inc., 720 F.2d 604, 607 (9th Cir. 1983) 19. See, e.g., 17 U.S.C. § 107 (2001) (setting forth a four-factor test for the fair use of copyrighted material); Cmty. for Creative Non-Violence v. Reid, 490 U.S. 730, 751 (1989) (setting forth a nonexhaustive multifactor test to determine whether a hired party is an employee under the general common law of agency). See also Barton Beebe, An Empirical Study of U.S. Copyright Fair Use Cases, 1978-2005 (working paper on file with author and the California Law Review). 20. See Penn Cent. Transp. Co. v. New York City, 438 U.S. 104, 124 (1978) (setting forth factors to be considered in resolving certain regulatory takings claims); See also Eduardo Mois~s Pefialver, Regulatory Taxings, 104 Colum. L. Rev. 2182, 2193-98 (2004) (discussing Penn Central's progeny and the establishment of various per se rules in takings jurisprudence); Cf Tahoe-Sierra Pres. Council, Inc. v. Tahoe Reg'l Planning Agency, 535 U.S. 302, 342 (2002) (holding that "the interest in 'fairness and justice' will be best served by relying on the familiar Penn Central approach when deciding [regulatory takings cases involving temporary regulations], rather than by attempting to craft a new categorical rule").
21. Cf Crawford v. Washington 541 U.S. 36, 63 (2004) (discussing "unpredictability" of multifactor tests for reliability of hearsay evidence).
22. See, e.g., Reinsurance Co. of Am. v. Administratia Asigurarilor de Stat, 902 F.2d 1275, 1279-83 (7th Cir. 1990) (applying the multifactor balancing test set forth in Restatement (Second) of Foreign Relations Law of the United States § 40); see id. at 1283 (Easterbrook, J., concurring) (expressing a "reluctan [ce] to accept an approach that calls on the district judge to throw a heap of factors on a table and then slice and dice to taste. Although it is easy to identify many relevant considerations, as the ALI's Restatement does, a court's job is to reach judgments on the basis of rules of law rather than to use a different recipe for each meal."). few. 24 For reasons described below, the multifactor test for the likelihood of consumer confusion provides an ideal case study in legal multifactor decision making, one that conduces to the development of a methodology and theoretical toolkit for the study of this form of legal analysis in the many other areas of law in which it is used.
Part I provides background. It briefly reviews the common origin of the multifactor tests for trademark infringement and their current diversity, and explains why judicial application of the tests is particularly amenable to empirical analysis. Part II sets out summary statistics relating to the venue, posture, and multifactor test win rates of the opinions sampled. It reveals significant intercircuit and interdistrict variation in plaintiff multifactor test win rates (the percentage of cases in which the plaintiff prevailed in the multifactor test).
Part III shows that judges employ fast andfrugal 5 heuristics to shortcircuit the multifactor test. Perhaps as an expression of their cognitive limitations, but more likely as an expression of their cognitive ingenuity, judges rely upon a few factors or combinations of factors to make their decisions. The rest of the factors are at best redundant and at worst irrelevant. Part III.A uses a series of classification trees to demonstrate that a limited number of core factors drive the test outcomes across the circuits.
See, e.g., Barker v. Wingo, 407 U.S. 514, 530 (1972) (identifying four factors that courts should weigh in determining whether a defendant has been deprived of his right to a speedy trial); Doggett v. United States, 505 U.S. 647, 657 (1992) (applying the Barker factors); see also id. at 670 (O'Connor, J., dissenting) (lamenting that "Barker's factors now appear to have taken on a life of their own").
24. See also Roth Steel Tube Co. v. Comm'r of Internal Revenue, 800 F.2d 625, 630 (6th Cir. 1986) (setting forth factors to be considered in recharacterizing debt as equity); In re Bendectin Prods. Liab. Litig., 749 F.2d 300, 304 (6th Cir. 1984) (setting forth factors to be considered in issuing writ of mandamus in class action context).
25. satisficing algorithms [which] operate with simple psychological principles that satisfy the constraints of limited time, knowledge, and computational might, rather than those of classical rationality. At the same time, they are designed to be fast and frugal without a significant loss of inferential accuracy, because the algorithms can exploit the structure of environments. Gigerenzer & Goldstein, supra at 651. In contrast to the heuristics and biases tradition, the fast and frugal view insists that fast and frugal heuristics are not a source of flawed decision making. On the contrary, "whereas the heuristics-and-biases program portrays heuristics as a frequent hindrance to sound reasoning, rendering Homo sapiens not so sapient, we see fast and frugal heuristics as enabling us to make reasonable decisions and behave adaptively in our environment-Homo sapiens would be lost without them." Gerd Gigerenzer & Peter M. Todd, Fast and Frugal Heuristics: The Adaptive Toolbox, in Simple Heuristics That Make Us Smart 3, 29 (1999). Indeed, Gigerenzer and his collaborators assert that fast and frugal heuristics in many cases "outperform" rational inference in that they achieve satisfactory results while requiring less time, information, and computation than more rational, integrative decision strategies. See, e.g., Gigerenzer & Goldstein, supra at 660.
Courts have long asserted that no single factor in the multifactor test is dispositive. 26 The data contradict this, and do so in dramatic fashion. Part III.B reveals the degree to which judges stampede specific factor outcomes to conform to or support the overall test outcome. The data suggest that judges determine the test outcome based on a limited number of core factors and then adjust the rest of the factor outcomes to accord with that result. This represents strong evidence of coherence-based reasoning 27 in the courts.
Part IV addresses specific characteristics of each of the core factors and pays particular attention to intercircuit variations in the application of these factors. The data on the intent, strength, and actual confusion factors are of special interest and run counter to conventional wisdom. In light of this study's findings, Part V proposes principles for the formulation of a national standard multifactor test for trademark infringement and suggests specific language. Part VI concludes.
Of the multifactor tests for trademark infringement, one commentator has observed that "[a]fter a [brief] period of disparity, the lists developed by the various federal circuits have converged; differences from one list to another have become fairly minimal." 2 " This is not accurate. The "period of disparity" was not brief and we are still in it.
26. See, e.g., Team Tires Plus, Ltd. v. Tires Plus, Inc., 394 F.3d 831, 833 (10th Cir. 2005) ("[N]o single factor is dispositive."); Gateway, Inc. v. Companion Prods., Inc., 384 F.3d 503, 509 (8th Cir. 2004) ("No single factor is dispositive."); AHP Subsidiary Holding Co. v. Stuart Hale Co., I F.3d 611, 616 (7th Cir. 1993) ("None of the seven confusion factors alone is dispositive in a likelihood of confusion analysis."); Plus Prods. v. Plus Disc. Foods, Inc., 722 F.2d 999, 1004 (2d Cir. 1983) ("No single Polaroid factor is determinative."); Lever Bros. v. Am. Bakeries Co., 693 F.2d 251, 253 (2d Cir. 1982) ("No single ... factor is preeminent, nor can the presence or absence of one determine, without analysis of the others, the outcome of an infringement suit.").
27. Coherence-based reasoning is a model of human decision making recently proposed and tested empirically by Professor Dan Simon and others. This model hypothesizes that decision making proceeds "bidirectionally:" premises and facts determine conclusions but those conclusions also inform and alter the premises and facts on which they are based. The decision maker's mental model of the decision task cycles towards a state of maximal possible coherence among premises, facts, and conclusion, each modifying the others. This results in a skewing of facts and premises towards inflated support for the ultimate decision. ("Although the factors of this test vary from circuit to circuit, there is little substantive variation among the tests."); Cf Bierman & Wexler, supra note 14, at 2 ("Although the circuits differ as to the precise formulation of the test for the likelihood of confusion, there is broad general agreement as to the factors which should be considered.").
[Vol. 94:1581 The idiosyncrasies of tradition rather than of reason governed the development of the multifactor tests across the circuits. Each of the circuits' current multifactor tests originated either directly or indirectly from the 1938 Restatement (First) of Torts, 29 and this may largely account for the current muddle. Due to a controversy of the time concerning the proper scope of trademark rights, 3 " the Restatement (First) failed to set forth a single, unified multifactor test for trademark infringement. Instead, it proposed a four-factor test to be considered in cases in which the parties' goods were competitive, that is, substitutable, 3 1 and an additional ninefactor test to be considered in cases where the parties' goods were not competitive.32 Initially, most circuits followed the example of the Restatement (First) and applied one multifactor test to competitive goods cases and another to non-competitive goods cases. This distinction eventually broke down, however, and the circuits each began to use a single, unified multifactor test regardless of whether the parties' goods were competitive or not. 33 For the most part, the peculiarities of the particular cases in which the circuit's multifactor test first coalesced determined which factors the circuit still considers today. This is certainly true of the influential Polaroid factors in the Second Circuit, 34 the Roto-Rooter factors 29.
See, e.g., AMF Inc. v. Sleekcraft Boats, 599 F.2d 341, 348 n. 1I (9th Cir. 1979) (citing to the Restatement (First) of Torts § 729); Scott Paper Co. v. Scott's Liquid Gold, Inc., 589 F.2d 1225, 1229 (3d Cir. 1978) (citing to Scott Paper Co. v. Scott's Liquid Gold, Inc., 439 F. Supp. 1022, 1036-37 (D. Del. 1977) (citing to the Restatement (First) of Torts § § 729, 731)); Roto-Rooter Corp. v. O'Neal, 513 F.2d 44, 45 (5th Cir. 1975) (citing Am. Foods, Inc. v. Golden Flake, Inc., 312 F.2d 619 (5th Cir. 1963) (citing to the Restatement (First) The controversy concerned whether a trademark owner should have the right to exclude the use of a confusingly similar mark only on "directly competitive" (directly substitutable) goods or services, or should have the right to exclude the use of such a mark at least on "goods of the same descriptive properties," if not more broadly on any goods or services where consumer confusion was likely. Cir. 1961). The plaintiff in Polaroid offered evidence that it might expand into the market for defendant's goods; consequently, the "bridge the gap" factor became one of the Polaroid factors, and was then taken up in various forms by the Third, Sixth, Ninth, and D.C. Circuits. Id. at 495-96. Similarly, the defendant in Polaroid argued in the Fifth, 35 and the Lapp factors in the Third. 3 6 In some circuits, however, such as the Seventh, the adoption of specific factors appears to have occurred more or less randomly. 37 The result is the great diversity of factors considered by the circuits as represented in Table 1.
Common to all of the circuits' tests are four factors: the similarity of the marks, the proximity of the goods, evidence of actual confusion, and the strength of the plaintiffs mark. A fifth factor, the intent of the defenthat "that there is no evidence that plaintiff has suffered either through loss of customers or injury to reputation, since defendant has conducted its business with high standards," and thus Judge Friendly considered the "quality of defendant's product" as well. Id. at 495. Courts ever since have wondered how the quality of defendants' goods speaks to the question of whether consumers are likely to be confused. Cir. 1975) ("In this circuit likelihood of confusion is determined by evaluating a variety of factors including the type of trademark at issue; similarity of design; similarity of product; identity of retail outlets and purchasers; identity of advertising media utilized; defendant's intent; and actual confusion."). The Roto-Rooter test, which strongly influenced the First, Fourth, and Eleventh Circuits' tests, failed to include a factor that was at the heart of the appeal in that case: the sophistication of the relevant consumer population. The district court in Roto-Rooter discounted the plaintiffs anecdotal evidence of actual confusion, finding that the four consumers who testified to having been confused had simply made an "'error' which 'resulted from carelessness or inadvertence rather than from any confusing similarity."' Id. at 46. The Fifth Circuit reversed. Apparently, since the very issue on appeal was whether the four consumers who were actually confused were or were not representative of the general sophistication of the relevant consumer population, the consumer sophistication factor did not make it into the opinion's recitation of the (other) factors to be considered. As a result, the Fifth Circuit, and the Fourth and Eleventh Circuits along with it, still do not explicitly consider the consumer sophistication factor. 36. 1977). Judge Stapleton observed that in a case involving non-competing goods, the court should consider "many of the same factors which are relevant in a competing goods case," of which he listed six. Id. at 1036. He then added that there "are additional factors, however, which are particularly relevant in a case involving non-competing goods," of which he listed four more. Seven years later, in its Helene Curtis opinion, the Seventh Circuit quoted from Carl Zeiss Stiftung as the sole authority for its own seven-factor test, which is still the test used in that circuit. Helene Curtis, 560 F.2d at 1330. It is not clear from Helene Curtis whether the Seventh Circuit chose the alternative authority on principle or simply because it was the case most readily at hand. dant, is found in all but the Federal Circuit's test. 38 For purposes of com- (First) of Torts and the current Restatement (Third) of Unfair Competition. Despite the great weight that many courts and commentators have long purported to place on the evidence of actual confusion factor, 39 it was never listed in the Restatement (First). Similarly, despite the centrality of the proximity of the goods factor to the current likelihood of confusion analysis, it does not explicitly appear among the factors proposed by the Restatement (Third). 40 Beyond the five core factors, the circuits currently consider a wide variety of additional factors. Also, notwithstanding the apparent intersections and equalities among certain circuits' sets of factors, the precise wording of their individual factors varies, sometimes dramatically. This often leads to strikingly different forms of analysis. Part IV below will address this phenomenon in detail.
Finally, with respect to each circuit's ordering of its factors, the similarity factor tends to come early on in the various tests as does the strength factor. Otherwise, the ordering of the factors is altogether haphazard. If the circuits gave any attention to the issue, it is not apparent from the opinions in which each first established its test. This is particularly surprising (or disturbing) in light of the well-established social science finding, if not the simple intuition, that the ordering of cues is of decisive importance to the success of a decision making strategy. 4 '
The story of American trademark doctrine over the past century is at least in part a story of flexible and intensely pragmatic practitionerand judge-made rules of thumb slowly degenerating into inflexible and formal doctrine. Lawyers and judges undoubtedly remain as flexible and 38.
Intent is not explicitly listed among the DuPont factors. However, the Federal Circuit will consider it when it is relevant.
pragmatic in their decision making as they ever were, but the doctrine has tended to take on a life of its own and to demand that its forms be followed, however perfunctorily. This certainly appears to have been the case with the multifactor test for the likelihood of consumer confusion. The Restatement (First) merely stated that "the following factors are important," 4 and the cases in which the various circuits first promulgated their respective tests contain similarly precatory language. 43 Nevertheless, the multifactor analysis has since become an essentially compulsory and formal exercise. While such formalism is regrettable, it also makes the empirical study of the multifactor test possible. The empirical study of judicial reasoning is notoriously problematic, 4 4 perhaps even more so than the empirical study of judicial outcomes (win rates). 45 Nevertheless, the highly routinized and explicit manner in which most district courts employ the multifactor test makes it especially susceptible to empirical methods. Typically, the district court begins the multifactor analysis by citing to or quoting from circuit authority and listing in numerical order the precise wording of the factors considered in its circuit. It may also state various general principles: the test should not be applied mechanically; 46 the list of factors is not exhaustive; 47 some factors may be more important than others, while some may be irrelevant; 48 the outcome of the test should not be driven simply by determining which party has won the most factors. 49 The court then proceeds methodically through each factor, often giving it its own heading, and typically states whether-or the degree to which-the factor favors or disfavors a likelihood of confusion, is neutral, irrelevant, or, in a summary judgment analysis, presents an issue of fact. After considering each of the factors, courts often explicitly balance the factors and summarize the results of their analysis. 5 0 In the process, district courts give every appearance of scrupulously following a basic weighted additive decision strategy, 5 and they do so for good reason. In most circuits, district courts that fail to address each factor of the multifactor analysis risk remand or reversal on that basis. 52 This is especially true in the Second Circuit where the multifactor test is most often applied and where appellate panels have repeatedly emphasized that the multifactor analysis must be exhaustive and explicit. 53 The Ninth Circuit appears to be alone in declining to establish a "rigid test for analyzing likelihood of confusion in trademark cases" 54 and allowing district courts to consider a "subset" 55 of factors. Particularly in the Internet context, the Ninth Circuit has counseled against "excessive rigidity." 56 Overall, the circuit-wide mean of the proportion of factors not explicitly addressed per opinion is very low (.096)," 7 and notwithstanding the Ninth Circuit's stated liberality, the mean in the Ninth Circuit (.113)58 is not significantly 50. At least one court has gone so far as to conclude its analysis with a helpful table showing which way each factor tilts. See Therma-Scan, Inc. v. Thermoscan, Inc., 118 F. Supp. 2d 792, 804 (E.D. Mich. 2000).
51. See John W. Payne et al., The Adaptive Decision Maker 24 (1993) ("The weighted additive rule considers the values of each alternative on all the relevant attributes and considers all the relative importances or weights of the attributes to the decision maker. Further, the conflict among values is assumed to be confronted and resolved by explicitly considering the extent to which one is willing to trade off attribute values, as reflected by the relative importances or weights.") (internal emphasis omitted).
52. different from the mean of all other circuits (.093)," 9 or, for that matter, from the Second Circuit's (.072).6o
District courts also tend to adhere very closely to their circuits' specific tests. Of the 331 opinions sampled, only six (2%) explicitly considered factors beyond those included in their circuit's test, 6 ' and in doing so, only one of these looked to precedent from other circuits. 6 2 What we are left with is a collection of multifactor tests each explicitly and uniformly applied in their respective circuits. This creates a stable platform for empirical analysis. The irony is that this uniformity of application within each circuit may very well generate disparities in outcomes across the circuits. This is the theme of Part 1I.
Most trademark lawyers have been content to assume-or, at least, hope-that given the same facts, the circuits' various tests would yield the same outcome. This Part suggests that this long-held assumption is incorrect. Different tests appear to yield different outcomes. First, however, to build a foundation for this claim, the Part briefly considers the venue and posture of the opinions sampled. 63
For the 331 opinions sampled, Table 2 sets out the number and percentage of opinions sampled per circuit, as well as their procedural posture. As expected, the district courts of the Second Circuit contributed a large plurality of opinions to the sample, producing nearly one-third of the total 9) the extent and nature of the changes made to the product, (10) the clarity and distinctiveness of the labeling on the rebuilt product, and (11) the degree to which any inferior qualities associated with the reconditioned product would likely be identified by the typical purchaser with the manufacturer" (citation omitted)); Savin Corp. v. opinions and nearly one-half of the bench trial opinions.' Also as expected, preliminary injunction opinions formed by far the most common posture. The court engaged in no further application of the multifactor test in all but four of the 145 cases that produced a preliminary injunction opinion. The data are thus consistent with the conventional view that New York is the primary venue and the preliminary injunction hearing the primary forum for the adjudication of trademark infringement claims. There was no significant variation across the years sampled in the number of opinions by circuit or posture. 65
Table 2 also reports the rate at which plaintiffs or, in the case of summary judgment motions, movants won the multifactor test in the opinions 64. The Southern District of New York in particular accounted for 78% of the opinions in the Second Circuit and 25% of the opinions nationally. The two other leading districts were the Northern District of Illinois and the Central District of California, accounting for 11% and 8% of the opinions in the sample, respectively (and accounting for 85% of the opinions in the Seventh Circuit and 47% of the opinions in the Ninth Circuit, respectively).
65. The data reveal certain other aspects of the opinions sampled that will be of particular interest to trademark specialists. With respect to the nature of the trademark property at issue, 81 % (267) of the opinions addressed a claim for the infringement of a trademark only, 16% (53) addressed a claim for the infringement of trade dress only, and 3% (11) addressed a claim for the infringement of both forms of trademark property. Of the seven trademark cases the Supreme Court has considered since 1992, four have been trade dress cases. (2005), claims of initial-interest and post-sale confusion also appear to be outside of the mainstream of trademark litigation. Only 12% (forty) of the opinions sampled considered a claim of initial-interest confusion and only 6% (twenty-one) of the opinions sampled found initial-interest confusion. The Second Circuit appears to be particularly hostile to claims of initial-interest confusion: it found initial-interest confusion in only two of the twelve opinions which considered it. In contrast, the Seventh Circuit found initial-interest confusion in each of the six opinions which considered it. As for post-sale confusion, only 6% (nineteen) of the opinions sampled considered a claim of post-sale confusion, and thirteen of these came from the Second Circuit, which found post-sale confusion in only three of these opinions. Overall, post-sale confusion was found in 1.5% (five) of the opinions sampled.
Also of interest is the exceptionally low proportion of cases in which the defendant made out a defense of parody, another area of the law to which trademark scholarship has long been especially attentive. See, e.g., Julie Zando-Dennis, Note, Not Playing Around: The Chilling Power of the Federal Trademark Dilution Act of 1995, 11 Cardozo Women's L.J. 599 (2005). Of the seven opinions which considered a defense of parody, five were from the Second Circuit, and only one of the seven found in favor of the plaintiff. This is consistent with the conventional wisdom that those defendants who have the wherewithal to litigate the issue are generally successful in doing so.
Finally, nineteen of the opinions sampled addressed claims of reverse confusion. Remarkably, the plaintiff prevailed in only two of these opinions for a win rate of .105 as against a win rate of .452 in opinions not addressing a claim of reverse confusion.
sampled. It must be emphasized that the data set consisted only of litigated federal trademark infringement cases that produced written opinions available from the Westlaw and Lexis databases. Furthermore, as explained in Appendix A, the data set excluded counterfeiting and licensing factpatterns. 66 For these reasons, the inferences we can draw from the multifactor test win rate results are limited. 6 7 Nevertheless, two general observations may be made.
First, there is substantial intercircuit variation in plaintiff multifactor test win rates. Across the opinions sampled, the Second Circuit yielded a significantly lower overall plaintiff multifactor test win rate (37%) than all other circuits (51 %).68 Meanwhile, the Ninth Circuit yielded a significantly higher overall plaintiff win rate (64%) than the rest of the circuits (43%).69 This result may largely stem from the differences in the plaintiff multifactor test win rates in preliminary injunction opinions: in the Second Circuit, the plaintiff win rate in such opinions (41%) was significantly lower than that in all other circuits (59%),70 while, in the Ninth Circuit, the plaintiff win rate (69%) was higher than all other circuits (50%), though the difference was marginally significant. 7 Bench trial plaintiffs also faired relatively poorly in the Second Circuit. Their multifactor test win rate was the lowest among the circuits."7
More specifically, the Southern and Eastern Districts of New York appeared to be particularly hazardous venues for the litigation of trademark infringement claims-or appeared to attract particularly foolhardy plaintiffs. The Southern District's overall plaintiff multifactor test win rate in the opinions sampled (36%) was well below the overall plaintiff win rate of 66.
Furthermore, the win rates reported in all other districts (51%).73 The Eastern District's overall plaintiff win rate was even worse (18%), 74 and was the lowest in the nation among the twenty-seven districts that contributed more than two opinions to the sample. No other districts emerged from the sample as especially proor antiplaintiff. 75 Second, in the opinions sampled, motions for summary judgment that were not met with cross-motions for summary judgment were significantly more successful than those that were met with cross-motions for summary judgment. A simple selection bias may account for this result: a party is more likely to bring a summary judgment motion when it has strong grounds for doing so, and a judge is more likely to produce a written opinion when she grants a summary judgment motion rather than denies it. Yet the striking disparity in win rates leaves open the possibility that crossmotions for summary judgment tend to some degree to cancel each other out. 76 III Interfactor Analysis It is something of a pastime in trademark law to speculate on which factors, if any, drive the outcome of the multifactor test and how the factors interact. The circuits themselves may largely be responsible for this 73. Southern District of New York: .361, n=83; all other districts: .508, n=248; p=.021. 74. Eastern District of New York: .182, n= 11.
The nine opinions sampled from the Eastern District of Virginia yielded an overall plaintiff win rate of .778, but five of these opinions involved a claim of trademark infringement by means of a domain name.
The overall plaintiff win rates in preliminary injunction and bench trial opinions are consistent with the "fifty percent hypothesis." (1990). William Landes hypothesizes that intellectual property plaintiffs should enjoy higher win rates at trial than other civil plaintiffs because intellectual property plaintiffs face the risk that their intellectual property will be declared invalid. Thus, they are more likely to settle or withdraw from even the near-close cases. See William M. Landes, An Empirical Analysis of Intellectual Property Litigation: Some Preliminary Results, 41 Hous. L. Rev. 749 (2004). Landes finds confirmation for this hypothesis in data provided by the Administrative Office of the United States Courts. Id. at 771-72. Appendix B explains why this data should not be trusted. Even so, if we except Landes's theoretical point, then how can we explain trademark plaintiffs' near-fifty percent win rates in the sampled preliminary injunction and bench trial opinions? Perhaps for trademark plaintiffs, the risk that their trademark will be declared invalid is offset by the well-established doctrinal principle that a trademark owner must aggressively protect its mark from use by third-parties lest the mark lose its source-distinctiveness and thus its validity. See, e.g., Herman Miller, Inc. v. Palazzetti Imports and Exports, Inc., 270 F.3d 298, 317 (6th Cir. 2001) (discussing how failure to prosecute may lead to the loss or weakening of trademark rights). For this reason, trademark plaintiffs are prone to proceed to trial unless the defendant agrees to discontinue use of the mark at issue. custom, as most have identified at one time or another certain individual factors or combinations of factors as being the most important in their particular circuit, if not in general. The Second Circuit, for example, has pointed to the similarity of the marks and the proximity of the goods, 7 7 or to the strength of the plaintiff s mark, similarity, and proximity; 8 the Third to similarity; 79 the Fourth to evidence of actual confusion; 80 the Sixth to proximity; 8 " the Seventh to similarity, the defendant's intent, and actual confusion; 82 the Ninth to similarity, proximity, and the commonality of the parties' marketing channels; 83 and the Eleventh to the "type of mark" and actual confusion. 84 In identifying these factors as the leading factors, the circuits rarely specify whether they mean that these factors are in fact the most important, or simply should in principle be the most important, or both. [Vol. 94:1581 This Part seeks to settle the debate, at least with respect to the question of is rather than ought. It shows that, in practice, a limited number of core factors determine the outcome of the test, and tend in the process to stampede the rest of the factors. The similarity of the marks factor is by far the most influential. Two other factors are decisive: the defendant's intent factor, but only when it favors a likelihood of confusion, and the proximity (or relatedness) of the parties' goods factor, but only when it disfavors a likelihood of confusion. The intent and actual confusion factors also appear to exert an inordinate degree of influence on the outcomes of the rest of the factors.
To demonstrate the importance of the core factors, I present the data in Subpart A from a variety of perspectives. Regression analysis is inappropriate--or at least unhelpful-here, not only because the data are so skewed, 85 but also because regression analysis tends to assume a fully integrative model of decision making, a model that this Article rejects. 86 Instead, I resort to a series of simple classification trees of the multifactor test outcome that classify that outcome according to certain factor outcomes. I then present crosstabulations of the test outcome by factor outcomes and factor outcomes by test outcome. This approach facilitates the 85.
As Tables 3 and4 show, the overall test outcome was invariant for certain factor outcomes. This raises the problem of "zero cell count" in which the dependent variable, here, the outcome of the multifactor test, is invariant for one or more values of an independent variable, for example, the similarity factor. See Scott Menard, Applied Logistic Regression Analysis 78-81 (2d ed. 2002). This invariance produces enormous regression coefficients and standard errors that severely limit the utility of the regression results. Furthermore, the factor outcomes tend very strongly to correlate with each other, as is shown in Subpart III.A.4, which raises the problem of multicollinearity. See Jacob Cohen et al., Applied Multiple Regression/Correlation Analysis for the Behavioral Sciences 390 (3d ed. 2003) (describing multicollinearity as occurring "in data sets in which one (or more) of the [independent variables] is highly correlated with the other [independent variables] in the regression equation. The estimate of the regression coefficient Bi for this correlated predictor will be very unreliable because little unique information is available from which to estimate its value.").
86. This problem is not generally recognized in the legal literature and has only recently been addressed elsewhere. See generally Simple Heuristics, supra note 25; Dhami & Harries, infra note 99; Gigerenzer & Goldstein, supra note 25. Cf Kenneth Hammond, Upon Reflection, 2 Thinking & Reasoning 239, 244 (1996) ("[A] sin of commission on my part was to overemphasize the role of the multiple regression (MR) technique as a model for organising information from multiple fallible indicators into a judgment."). If we accept the proposition that the boundedly rational decision maker does not integrate all relevant information and frequently engages in noncompensatory decision strategies, flexibly choosing primary cues based on the decision task, then a fully integrative model of decision making may lead us astray. Cf Societe Anonyme de la Grande Distillerie v. Julius Wile Sons & Co., 161 F. Supp. 545, 547 (S.D.N.Y. 1958) (Likelihood of confusion "does not readily lend itself to resolution by scientific appraisement or comparison. Rather, a finding on infringement is by necessity a subjective determination by the trial judge based on his visceral reactions .... ). Furthermore, if we accept the proposition that judges make a decision based upon a limited number of core factors and then stampede the rest of the factors to support that decision, then regression analysis may not adequately account for the degree to which the factors influence each other. At the very least, such an analysis may produce a suboptimal fit with the data. See Dhami & Harries, infra note 99. Indeed, no regression analysis of the factors considered in this study was able to achieve the degree of fit that a simple five-factor--or two-factor, for that matter-classification tree could achieve. assessment of which factor outcomes necessitate a particular test outcome and which test outcomes necessitate particular factor outcomes. Finally, I present the results of correlation analysis of the factor outcomes and test outcomes. This analysis will prepare the ground for our consideration in Subpart B of the phenomenon of stampeding.
A diverse body of empirical work supports the intuition that when confronted with complex decision tasks we seldom seek to consider all relevant information or reduce uncertainty to the maximum extent conceivable, even if we were capable of doing so. Instead, we use various strategies to decide when to stop acquiring and analyzing information and commit to a course of action. Empirical studies of decision making generally, 87 and of judicial decision making in particular, 8 " consistently show that decision makers, even when making complex decisions, reach their stopping threshold and make a decision after considering a remarkably low number of decision-relevant factors. Social science researchers have demonstrated that in regression-based modeling of human decision making, only a small number of cues-three on average, one author has 61, 93 (2000) (observing that "courts tend-to identify and adapt to the influence of cognitive illusion on the determination of issues that juries are likely to resolve and ignore or fall prey to the influence of cognitive illusions on the determination of issues that judges are likely to resolve"); Jeffrey J. Rachlinksi, A Positive Psychological Theory of Judging in Hindsight, 65 U. Chi. L. Rev. 571 (1998) (discussing the legal systems adaptation to and accommodation of hindsight bias).
[Vol. 94:1581 speculated "--emerge as statistically significant. 9° Perhaps more interestingly for our purposes, empirical work suggests that decision makers tend to use a core attributes heuristic by which they stop acquiring and analyzing information once the last in their set of most important, determinant attributes has been acquired and analyzed. 9 ' Empirical work also suggests that such decision making strategies are rational. Social science research dating back to the 1950s has demonstrated that the consideration of too much information impairs decision-making accuracy. 92 More recently, a strong body of research in the fast and frugal tradition shows that in certain contexts, we make complex decisions based on a single cue, and that this simple strategy, known as Take The Best, is as good as or even outperforms computation-intensive decision making. 93 The data collected for this study support both the general hypothesis that decision makers, even when making complex decisions, consider only a small number of factors and the more specific hypothesis that, in doing so, decision makers use a core attributes heuristic. 94 Furthermore, the data suggest that this decision making strategy need not be understood as a regrettable expression of human cognitive limitations. On the contrary, judges take advantage of the ecology 95 of the multifactor test and, in particular, the redundancy 96 of many of its factors to reach a conclusion efficiently. The heuristic they employ is evidence, I suggest, of human ingenuity rather than human fallibility.
The circuits' various multifactor tests studied in this paper use from six to ten factors-the Federal Circuit's majestic thirteen factor test is not considered. 9 7 The average number of factors across the twelve circuits studied is 7.5, but the data suggest that knowledge of far fewer factor outcomes is sufficient to predict the overall outcome of any circuit's particular multifactor test. Figures 1 and2 set forth simple classification trees of the multifactor test outcome in the 192 preliminary injunction and bench trial opinions sampled. 9 ' Figure 1 shows that a one-factor classification scheme based on the outcome of the similarity factor alone (if the similarity factor favors a likelihood of confusion, predict that plaintiff wins; if it does not, predict that defendant wins) would yield an accurate classification of 173 or 90% of the outcomes of the 192 preliminary injunction and bench trial opinions sampled. In Figure 2, the addition of a second node, going to the proximity factor, would increase the accuracy of the classification scheme to 96%, an accuracy rate far above what any regression analysis of the test outcome on all the core factor outcomes proved capable of yielding. 99 95. See Gigerenzer & Todd, supra note 25, at 5 (describing "ecological rationality" as "rationality that is defined by its fit with reality"); id. at 13 ("A heuristic is ecologically rational to the degree that it is adapted to the structure of an environment. . . . Thus, simple heuristics and environmental structure can both work hand in hand to provide a realistic alternative to the ideal of optimization, whether unbounded or constrained.").
96. See Entrepreneur Media v. Smith, 279 F.3d 1135, 1141 (9th Cir. 2002) ("Since each factor represents only a facet of the single dispositive issue of likely confusion, the factors, not surprisingly, tend to overlap and interact, and the resolution of one factor will likely influence the outcome and relative importance of other factors. . . .[T]he determination of one factor is often, in essence, only another way of viewing the same considerations already taken into account in finding the presence or absence of another one."). See also Gigerenzer & Goldstein, supra note 25, at 654-55 (discussing pairwise correlation as a means of measuring cue-redundancy); id. at 665 (concluding that "[h]igh cue redundancy . . . does seem sufficient but is not necessary for the successful performance of the satisficing algorithms").
97. None of the district court opinions sampled used the Federal Circuit's Du Pont factors. 98. For much of the analysis in this Subpart, I grouped together the 146 preliminary injunction opinions with the forty-six bench trial opinions in order to achieve a sufficiently large sample size to yield meaningful results. Quite obviously, preliminary injunction and bench trial opinions arise out of very different stages of the litigation process. However, as is reflected in the remarkably similar win rates of the two groups of opinions, courts' application of the multifactor test does not appear appreciably to differ between the two groups. A closer comparison of the data for the two groups would confirm this.
99. See Mandeep K. Dhami & Clare Harries, Fast and Frugal Versus Regression Models of Human Judgement, 7 Thinking & Reasoning 5, 6 (2001) ("Fast and frugal models are easier to understand, and are psychologically more plausible than regression models because they are more Though it is hard to believe that a federal district court judge would employ, even under exceptional time pressure, anything like a "take the best, ignore the rest"' ° heuristic, both figures offer tantalizing evidence of such a heuristic in action.
Figure 3 sets forth a more elaborate five-node classification tree of the outcomes of the 192 preliminary injunction and bench trial opinions sampled. Here, the five nodes are ordered at each stage according to the accuracy of their classification of the remaining opinions, and if their accuracy is equal, according to the number of remaining opinion outcomes they accurately classify. 10 ' This node order produces, for this data, the most accurate classification scheme (97%) among all possible node orders.
The distribution of the classification tree in Figure 3 is quite stunning. Of the seventy-one opinions in which the similarity factor did not favor a likelihood of confusion, the defendant won the multifactor test in each of these opinions. Of the 121 remaining opinions, the plaintiff won the multifactor test in each of the sixty-five opinions in which the court found (in addition to similarity) that the defendant intended to confuse consumers as to source. Of the remaining fifty-six opinions, twenty-one found evidence of actual confusion, and in each of these, plaintiff won the multifactor test. Thus, by the third node, we have classified 82% of the 192 opinion outcomes with none misclassified. Of the thirty-five remaining opinions, in thirteen the court found that the proximity of the goods factor disfavored a likelihood of confusion, and in twelve of these, the defendant won the multifactor test. This leaves twenty-two opinions. In twelve of these, the court found that the strength factor favored a likelihood of confusion, and in eleven of these twelve, the court found in favor of the plaintiff. Of the ten opinions in which the court found that the strength factor did not favor a likelihood of confusion, the court found in favor of the defendant in six of them. Thus, the application of five nodes accurately classifies 186 of 192 opinion outcomes for an overall accuracy of 97%. The average number of nodes (factor outcomes) required to classify an opinion outcome under this scheme was 2.22.
A word of caution: though the results of the classification trees set forth in Figures 1,2, and 3 are compelling, there is little in the opinions sampled to suggest that this rudimentary outcome analysis accurately reflects the process of reasoning the judges actually employed. That is, there is little to suggest that the judges conducted a separate, independent analysis of each successive factor in the order and with the stopping points represented. On the contrary, there is much to suggest that the judges engaged in what Professor Dan Simon has called coherence-based reasoning," 0 2 that certain factor outcomes affected the outcomes of certain other factors, and that the test outcome may have influenced various factor outcomes as much as various factor outcomes influenced the test outcome. Of this, I will have more to say below. Nevertheless, the distributions of the classification trees suggest that at least some aspects of the multifactor analysis of the likelihood of confusion are noncompensatory in nature. Put differently, certain factor outcomes are weighted so strongly as to outweigh the combined weights of all other factor outcomes. 1 03 For example, a finding that the similarity of the marks factor does not favor a likelihood of confusion is sufficient to trigger an overall finding of no likelihood of confusion, regardless, it appears, of the outcomes of any other factors. Similarly, two findings-that the similarity of the marks factor favors a likelihood of confusion and the defendant's intent factor also favors a likelihood of confusion-are together sufficient to trigger an overall finding of a likelihood of confusion, again, regardless of the outcomes of any other factors. "4
The multifactor test outcome classification trees do not do justice to the association between the test outcome and certain individual factors listed lower-or not at all-in Figure 3 (1995). An example of a noncompensatory decision-making strategy is the lexicographic rule (consider the process of ordering words in a dictionary), in which the decision maker compares decision alternatives according to attributes ordered by importance, and decides based on the first attribute that discriminates between the alternatives. take a closer look at the data. For the 192 preliminary injunction and bench trial opinions sampled, Table 3 sets out by factor the number and proportion of opinions that held that the factor favored or disfavored a likelihood of confusion. The table also sets out the plaintiff multifactor test win rates in those opinions. Thus, taking the similarity factor as an example, 121 opinions (63%) found that the similarity factor favored a likelihood of confusion, and plaintiff won the multifactor test in 84% of these opinions; sixty-five opinions (34%) found that the similarity factor disfavored confusion, and plaintiff won the multifactor test in none of these opinions.
Compare these plaintiff win rates under the similarity factor to plaintiff win rates under the actual confusion and intent factors. The court found an intent to confuse consumers in sixty-seven opinions. In sixty-five (97%) of these opinions, the court found an overall likelihood of confusion.' 05 A finding of bad faith appears to exert a substantial, if not dispositive, influence on the outcome of the test. With respect to the actual confusion factor, the court found that the factor favored a likelihood of confusion in sixty-six opinions. In sixty-one (92%) of these opinions, the court found an overall likelihood of confusion. Finally, compare the consumer sophistication factor win rates to the similarity factor win rates. The plaintiff win rate in the fifty-five cases in which it won the consumer sophistication factor is comparable to the 121 cases in which it won the similarity factor.
We find dramatically skewed results as well in opinions in which the plaintiff lost certain factors. In forty-one of the 192 opinions, the court found that the parties' goods were not proximate. The plaintiff lost the multifactor test in all but one of these opinions." °6 As a practical matter, in order to win the multifactor test, the plaintiff must not lose this factor--or alternatively, when the judge finds an overall likelihood of confusion, the judge almost invariably finds that the proximity factor favors this result. The same may be said of the strength factor. The plaintiff lost this factor in fifty-three of the 192 opinions and lost the overall test in fifty of these fiftythree opinions.
For a discussion of the two outlying opinions, see infra notes 117-126 and accompanying text.
106. In the one outlying case, Welch Allyn Inc. v. Tyco Int'l Servs. AG, 200 F. Supp. 2d 130 (N.D.N.Y. 2002), the plaintiff sold stethoscopes and sphygmomanometers under the mark TYCOS while the defendant sold medical "disposable items or small items purchased in bulk," id. at 140-41, under the mark TYCO/HEALTHCARE or TYCO with a corporate tag line "A Tyco International Ltd. Company." Id. at 139-41. Though the court found that "[p]laintiff has not presented sufficient evidence to suggest that this [proximity] factor weighs in its favor." Id. at 141. The court found a likelihood of confusion and granted a preliminary injunction against the defendant with respect to the sale of any "non-disposable medical products or medical instruments..." Id. at 151.
While Table 3 tabulates the overall multifactor test outcome by the outcome of each individual factor, Table 4 does the inverse. It tabulates the outcomes of the individual factors by outcome of the overall multifactor test. In doing so, it provides some insight, however limited, on the degree to which test outcomes may drive specific factor outcomes.
Consider, for example, the 102 opinions that found a likelihood of confusion. Of these, only sixty-five (64%) found that the intent factor favored a likelihood of confusion, which suggests that, while probably nearly sufficient, a finding of bad faith intent is by no means necessary to trigger an overall finding of a likelihood of confusion. By comparison, 94% of the 102 opinions that found a likelihood of confusion found that the proximity of the goods factor favored this result and 90% found that the strength factor favored the result-and, of course, all of them found that the similarity factor favored the result. From this, we can infer, albeit weakly, that judges tend to rely on these three factor outcomes to form the foundation for a finding of a likelihood of confusion. Indeed, ninety (88%) of the 102 opinions that found a likelihood of confusion found that each of these factors favored that result.
Consider next the ninety opinions that found no likelihood of confusion. Nineteen of these (21%) nevertheless found that the similarity of the marks disfavored that result. This confirms that a finding that the similarity factor favors a likelihood of confusion is necessary but not sufficient to trigger an overall finding of a likelihood of confusion. Similarly, thirtyeight (42%) found that the proximity of the goods factor favored a likelihood of confusion. Thus, while the plaintiff must not lose the proximity factor in order to win the multifactor test, winning the factor does not guarantee success.
Table 5 shows, with respect to the 192 preliminary injunction and bench trial opinions sampled, the pairwise correlation coefficients 1 1 7 between, on the one hand, each of the two outcomes of the multifactor test and on the other, each of the two most common outcomes (favors or 107.
Generally speaking, correlation analysis produces a correlation coefficient from -1, which, if the coefficient is statistically significant, demonstrates a perfect inverse relation between two variables, to +1, which, again if the coefficient is statistically significant, demonstrates a perfect positive relation between two variables. disfavors a likelihood of confusion) of the nine factors most commonly used among the circuits. It is prudent not to place too much weight on these results, first, because a correlation coefficient can underreport the strength of correlations between dichotomous variables," 8 and second, because the correlations reported are weakened by their failure to take into account outcomes other than favors or disfavors a likelihood of confusion, such as findings that a factor was neutral, irrelevant, or not argued. 109 Even so, the coefficients given in the first two columns of the table are generally consistent with the inferences made above. The outcomes of the similarity factor enjoy the strongest correlation with the overall test outcomes. Additionally, the outcomes of the other four core factors also each correlate fairly strongly with the larger test outcomes, with strength and intent correlating slightly more strongly than actual confusion and proximity." 0 Outcomes under the remaining four factors listed in the tablepurchaser sophistication, similarity of advertising/marketing, similarity of sales facilities, and likelihood of bridging the gap-show weak correlations with the test outcomes, with one exception. A finding that the consumer sophistication factor disfavors a likelihood of confusion correlates fairly strongly with an overall finding of no likelihood of confusion. This is consistent with the plaintiff's relatively low multifactor test win rate (19%) in opinions in which the court made this specific finding. As for the three other factors used by the circuits but not listed in Table 5, both the comparative quality of the marks factor and the similarity of the targets of the parties' sales efforts factor showed no strong correlation with the outcome of the test."' Finally, as expected, the factor analyzing the length of time of 108.
See, e.g., Cohen et al., supra note 85, at 53-55. A dichotomous variable is a binary variable typically coded either as one or zero.
109. For this correlation analysis, each factor outcome is represented with two binary variables: favors a likelihood of confusion (I=yes, 0=no) and disfavors a likelihood of confusion (l=yes, 0=no). Thus, if the first variable is coded as one, then the second variable will be coded as zero, and vice-versa. But if the court found the factor to be neutral, irrelevant, or not argued, then both variables were coded as zero. This explains why, in Table 5, the sum of the absolute values of the correlation coefficients for any particular factor variable as against another factor (or test) outcome do not equal one.
The correlation coefficient of the similarity factor with the test outcome is relatively high. Its confidence interval at the .05 level does not overlap with the confidence intervals of any other factors' correlation coefficients.
Ill. For preliminary injunction and bench trial opinions, the correlation between an overall finding of a likelihood of confusion and a finding that the comparative quality factor favored that result was .289 (p=.020, n=66); the correlation for preliminary injunction and bench trial opinions between an overall finding of no likelihood of confusion and a finding that the comparative quality factor favored that result was .324 (p=.008, n=66). The correlations between the outcomes of the similarity of the targets of the parties' sales efforts and the outcomes of the multifactor test were not significant at the .05 level for preliminary injunction and bench trial opinions.
In general, the data suggest, first, that when district judges use the multifactor test, very few factor outcomes, often merely two or three, are sufficient to trigger one of the two test outcomes, but second, that both of these test outcomes require or nearly require certain factor outcomes. Can we conclude, then, that judges tend to short-circuit the multifactor test?
The answer is yes. In theory, the multifactor test is a fuli-fledged balancing test. In practice, it is a complex of per se rules. But in relying only on certain leading factors, are judges making flawed decisions? If recent research in human decision making is any guide, then the answer is very likely no. Like any human decision makers, district judges attempt to decide both efficiently and accurately. " 3 In pursuit of efficiency, they consider only a few factors. In pursuit of accuracy, they consider the most decisive factors. In essence, as consummate pragmatists, they "take the best," a strategy which empirical work suggests is an altogether successful-and rationalapproach to decision making." 4 The crucial problem, however, is that in explaining the reasoning that led to their decision, district judges are not permitted to "ignore the rest." Each factor must be addressed, not simply those which actually formed the basis for their ruling. This creates the conditions for the stampeding of factor outcomes, a phenomenon to which I now turn.
The correlation coefficients reported in Table 5 among the factor outcomes themselves, rather than with the test outcomes, show that all of the core factor outcomes bear a statistically significant association with each other. Additionally, these coefficients reveal that many of the non-core factor outcomes also produce a statistically significant association with the core factor outcomes, if not also to some extent with each other's outcomes. This hints at an interesting phenomenon: judges tended to stampede the factor outcomes to favor the test outcome, especially when they found a 112. For preliminary injunction and bench trial opinions, the correlation between an overall finding of a likelihood of confusion and a finding that the length of time of concurrent use factor favored that result was .535 (p=.022, n= 18); the correlation between an overall finding of no likelihood of confusion and a finding that the length of time of concurrent use factor favored that result was .570 (p=.014, n=18 likelihood of confusion. What emerges is a picture of legal multifactor decision making in which certain factors drive the outcome and the rest of the factors subsequently fall in line to support that outcome. The data are all the more compelling in light of the fact that the sampled opinions arise out of cases that we would generally assume to be close cases, that is, cases that failed to settle before the issuance of the opinion. "' Furthermore, the data set excluded counterfeiting and licensing opinions, which we would expect to stampede." 6 Before turning to the data, however, I need to address more generally the practice of coherence-based reasoning.
In a remarkable series of law review and experimental psychology journal articles, researchers have developed and tested a model of what they have termed coherence-based reasoning. 1 7 This model hypothesizes that the "decision-making process progresses bi-directionally: premises and facts both determine conclusions and are affected by them in return. A natural result of this cognitive process is a skewing of the premises and facts toward inflated support for the chosen decision."" 8 As Professor Simon explains, the decision-making process begins with a "mental model of the decision task," which: contains a myriad of variables that point in more than one direction and thus do not all fit into a coherent mental model. One subset of variables [a,, a 2 ,... a 3 ] supports conclusion A, and the other subset easy cases, neither subset of variables dominates the other. Since each variable has some bearing on the task, it can be said to impose a constraint on the network.... Each and every constraint influences, and is influenced by, the entire network, so that every processing cycle results in a slightly modified mental model. " 9 The decision maker cycles through the mental model in an effort to satisfy, and in the process conform, the model's myriad constraints until "the 115.
See constraints settle at a point of maximal coherence.""' 2 By means of a kind of "reversed induction," "[c]onstraint satisfaction processes force the task variables to change toward a better fit with the gradually emerging state of coherence."' 2 In time, the mental model shifts toward this state of coherence, in which the considerations favoring an alternative are seen strongly to do so, and those disfavoring [that] alternative are seen to do so weakly, if at all. 122 "In sum," Professor Simon writes, the ultimate state of coherence is essentially a byproduct of the cognitive system's drifting toward either one of two skewed mental models. Within each of these models, the initially complex and incoherent mental model has been spread into two subsets, one of which dominates the other, thereby enabling a relatively easy and confident choice. This skewed representation reflects an artificial polarization between the inflated representation of the variables that support the chosen conclusion and the deflated ones that support the rejected conclusion; it differs considerably from the way the task variables were perceived before the decision-making process got underway, and it differs also from the way they will be perceived some time after the completion of the task. 3 Two aspects of the coherence-based reasoning model must be emphasized from the outset. First, while the model shares in the general spirit of the biases and heuristics research that underlies behavioral law and economics, it addresses different cognitive phenomena. It is more concerned with the underlying cognitive processes that drive complex decision making than with the biases and heuristics that inform various discrete judgments and decisions. Second, coherence-based reasoning theory is distinct from cognitive dissonance theory. 24 The former does not conceive, as the latter does, of coherence shifts as simply a matter of post hoc rationalization, as something triggered in an effort to reduce post-decision regret. On the contrary, empirical work suggests that coherence shifts precede the decision and form the basis for it. In fact, they can occur early on in the decision-making process and can even be triggered by a single attribute. 2 5 "Much as forcing one interpretation on a vertex of a Necker cube can cause 120. Id. at 522. 121.
Id. 122. Id. at 516. ("A mental model of a decision task is deemed 'coherent' when the decisionmaker perceives the chosen alternative to be supported by strong considerations while the considerations that support the rejected alternative are weak .... A mental model is considered 'incoherent' when the decision-maker perceives the considerations as providing equivocal support for both alternatives. As defined, coherence is an empirical phenomenon, not a jurisprudential ideal." (internal citations omitted)).
123 the entire set of vertices to be perceived as a particular three-dimensional form, biasing the assessment of one particular point of dispute should initiate spreading coherence, causing systematic changes in the entire set of assessments related...,1 to the decision task.
The data collected for this study support the coherence-based reasoning model, and specifically, the hypothesis that coherence shifts trigger an artificial polarization between the set of considerations favoring a decision and the set disfavoring a decision. Table 6 reports the mean multifactor stampede scores for the 331 opinions sampled by outcome and posture. An opinion's multifactor stampede score is the difference between the proportion of factors considered that favored a finding of a likelihood of confusion and the proportion of factors considered that did not favor a finding of a likelihood of confusion. Thus, in an opinion where all factors in the multifactor test were held to favor a likelihood of confusion, the opinion yielded a stampede score of 1.000, and in an opinion where all factors were held not to favor a likelihood of confusion, the opinion yielded a stampede score of -1.000. In opinions where the factors were tied (for example, four in favor of a likelihood of confusion and four against), the stampede score was 0.000. Figure 4 sets forth the distribution of stampede scores by disposition for the 287 dispositive opinions sampled.
The remarkably high mean stampede scores for opinions that found a likelihood of confusion provide strong evidence of polarization, as does the strong skew in the distribution of the stampede scores of such opinions. The twenty-five bench trial opinions that found a likelihood of confusion yielded a mean stampede score of .788,127 with nine (36%) yielding a stampede score of 1.000.128 Similarly, the seventy-eight preliminary injunction opinions that found a likelihood of confusion yielded a mean stampede score of .706, with twenty-eight (36%) yielding a stampede score of 1.000. Opinions granting plaintiff summary judgment provided similarly high 126.
Holyoak & Simon, supra note 117, at 12. The Swiss geologist Louis Albert Necker is often given credit for developing the following optical illusion:
See, e.g., Frederick Burwick, The Grotesque: Illusion vs. Delusion, in Aesthetic Illusion: Theoretical and Historical Approaches 122, 132 (Frederick Burwick & Walter Pape eds., 1990) (discussing Necker's first encounter with the optical illusion that bears his name).
127. In an eight factor test, this is roughly the equivalent of, among other combinations, seven factors in favor of confusion and one against.
128. Again, I emphasize that counterfeiting, licensing, and similar cases where we would expect such scores were excluded from the sample. scores, even when the defendant submitted its own cross-motion for summary judgment. This suggests that the tendency of courts to stampede the factors when they find an overall likelihood of confusion is robust against the various presumptions and standards of the opinion postures.
When the multifactor test led to a finding of no likelihood of confusion, however, there was little, if any stampeding of the factors. This is apparent in each disposition's relatively modest stampede score and in the low proportion of opinions that yielded a score of -1.000. The relatively flat distribution of the stampede scores of such opinions also supports this characterization. Once again, this characterization is robust against the posture of the opinion.
Why do the factors tend to stampede in one direction but not in the other, regardless of posture? There may be a simple explanation, one that goes to the underlying merits of the cases brought. Generally, a plaintiff will not bring an action for trademark infringement unless the facts of its case are such that it will win at least a few of the multifactor test factors. The overall mean stampede score, regardless of outcome, for the 331 opinions sampled is .120,129 whereas the plaintiff multifactor test win rate for these opinions is about even at .471. Similarly, the overall mean stampede score for the 192 preliminary and bench trial opinions sampled is .170,13°w hereas the plaintiff multifactor test win rate for these opinions is also about even at .531. These results are consistent with an overall lean in the test by one or two factors toward a likelihood of confusion.
Furthermore, a finding of no likelihood of confusion generally represents an endorsement of the status quo, while a finding of a likelihood of confusion generally leads to an injunctive intervention in the status quo. Regardless of the presumptions and standards of the preliminary injunction, summary judgment, or bench trial postures, courts may feel, perhaps as a function of "status quo bias' ' 31 or "omission bias,'' 132 multifactor test must tilt strongly toward a likelihood of confusion to justify such an intervention, whereas an endorsement of the status quo does not require as strong of a foundation. As their mental models cycle toward a finding for the plaintiff, judges may thus feel the need to bolster their findings under various factors, if only to insulate their more activist judgment from appeal.
Table 6 also shows the highest and lowest stampede scores for each disposition. Note that in eighteen of the 331 cases sampled (5%), the loser of the multifactor test won more factors than the winner of the test. 3 3 For example, Table 6 shows that in two preliminary injunction opinions, the court found a likelihood of confusion even though the defendant won more factors in the multifactor test than the plaintiff, and that in seven preliminary injunction opinions, the court found no likelihood of confusion even though the plaintiff won more factors than the defendant. This confirms courts' frequent observation that the outcome of the test is not necessarily determined by which party wins the most factors., 34
In light of the distribution of the outcome classification tree shown in Figure 3 and the prevalence of certain factor outcomes in opinions that found a likelihood of confusion as shown in Table 4, it is tempting to suggest that the multifactor test is essentially a two-stage test, the first stage of which is not a balancing test. Instead, the first stage could be conceived of as consisting of certain requirements, each of which the plaintiff must meet: The marks must be similar, the plaintiffs mark must be strong, and the goods must be proximate. If these requirements are met, then the test proceeds to the second stage, which balances other factors, the most important of which are intent, actual confusion, and perhaps also consumer sophistication.
The problem with this characterization is that it fails properly to account for the strong influence of two core factors, from stage two of the above hypothesized test, on their stage one peers. These two factors are intent and actual confusion, whose outcomes appear to stampede the rest of the test factors. There is thus yet another possible explanation for the divergence by outcome in the degree of stampeding in the opinions sampled. As we saw above, Simon suggests that a single variable can initiate spreading coherence in the sense that featuring one vertex of a Necker cube can shift the viewer's perception of all the other vertices. Intent and, to a lesser 133. See, e.g., Butcher Co., Inc. v. Bouthot, 124 F. Supp. 2d 750, 760 (D. Me. 2001) ("It is true, of course, that in terms of simple numbers there are only three factors weighing against finding a likelihood of confusion but five factors in favor of finding a likelihood of confusion. Nonetheless, the former are determinative in this case.").
degree, actual confusion appear to exert such a coherence-shifting influence when they favor a likelihood of confusion. Indeed, in the forty-nine opinions in which both findings were made, thirty-four (69%) of them found that all the factors favored a likelihood of confusion.
In addition to setting out the test outcomes by factor outcome, Table 3 reports the stampede scores for the preliminary injunction and bench trial opinions in which a factor was found to favor or disfavor a likelihood of confusion (or some other outcome). When the intent factor is found to favor a likelihood of confusion, it reports a remarkably high mean stampede score of .798 with severe negative skewing"' and peakedness. ' 36 Similarly, when the actual confusion factor is found to favor a likelihood of confusion, it reports a mean stampede score of .765, also with severe negative skewing and peakedness. The other factors do not show such characteristics.
The data show some variation among the circuits in their district courts' propensity to stampede the factors. Due perhaps in part to the limited number of opinions sampled from certain circuits, there are few statistically significant differences between the mean stampede score of an individual circuit as compared to that of all other circuits. The main exception is the Second Circuit. Its districts' mean stampede score for preliminary injunction and bench trial opinions that found a likelihood of confusion (.593)137 is the lowest among the circuits reporting more than two opinions, and is significantly different from the mean stampede score of such opinions from all other circuits (.768).' If, as the data suggest, Second Circuit district courts appear far less prone to stampede the factors when they find a likelihood of confusion, then this may be the result of these courts' and their bars' more frequent exposure to and greater sophistication in the use of the multifactor test. However, the Ninth Circuit, which also produced a relatively large number of multifactor opinions, re- ports a relatively high stampede score (.801) 139 in preliminary injunction and bench trial opinions that found a likelihood of confusion, one that is significantly different from the Second Circuit's, but not from that of all other circuits. 1 40 More speculatively, the Second Circuit's relatively low mean stampede score may be further evidence of what Part II.B suggested was the circuit's slight bias as compared to other circuits against finding a likelihood of confusion.
For preliminary injunction and bench trial opinions in which no likelihood of confusion was found, intercircuit variation in mean stampede scores is muted. Interestingly, however, the Second Circuit reports a relatively large negative mean stampede score (-.502)1 4 1 -though we cannot say that this difference is statistically significant from the mean stampede score in such cases of all other circuits (-.426). 142
We have seen that the core factors drive the outcome of the test, and that some factors are far more influential than others. This Part looks more closely at how each of the five core factors operates, beginning with the similarity of the marks factor, and then considering, in rough order of the importance of the factors, the defendant's intent, the proximity of the goods, the strength of the plaintiff's mark, and the evidence of actual confusion. The Part reports a number of findings that contravene conventional wisdom in trademark law. Most notably, it demonstrates that the intent factor, thought by some to be irrelevant, is of decisive importance, that survey evidence, thought by many to be highly influential, is in practice of little importance, and that the doctrine of trademark strength, particularly as it relates to the concept of inherent distinctiveness, has broken down. In the process of reviewing the factor-specific data, this Part also briefly speculates on the implications of the data for the reform of the multifactor analysis. As originally conceived, the multifactor analysis was simply a heuristic to aid the judge in making an accurate finding of fact as to the likelihood of consumer confusion. Much that is external or contrary to that purpose has since insinuated itself into the multifactor analysis. The goal of any reform should be to restore to the multifactor analysis its empirical, fact-finding purpose.
The data clearly show that the similarity of the marks factor is by far the most important factor in the multifactor test. Of course, courts and commentators have long said as much, if not more. Courts have suggested that the similarity factor can be "dispositive, ' 143 and an authoritative treatise has asserted that this factor "is usually controlling." ' 144 However, these characterizations are accurate only in the limited sense that the plaintiff must win the similarity factor in order to have any chance of winning the multifactor test. 4 ' As some courts have recognized, the similarity inquiry is a threshold inquiry. 146 This makes intuitive sense. It is hard to imagine a judge finding that the marks are not similar, and yet that consumers are likely to confuse them. 47 Remarkably, we might expect cases which produce reported preliminary injunction or bench trial opinions to be close cases that would turn on factors other than this threshold requirement. Yet, in sixty-five out of the 192 opinions sampled, the defendant failed to 143.
See, e.g., Nabisco, Inc. v. Warner-Lambert Co., 220 F.3d 43, 48 (2d Cir. 2000) ("Having determined that the parties' use of their DENTYNE ICE and ICE BREAKERS marks is so dissimilar as to require judgment for Warner-Lambert, we need not examine the remaining Polaroid factors and express no view of the district court's analysis of them."). See also Kaufman & Fisher Wish Co. v. F.A.O. Schwarz, 184 F. Supp. 2d 311,323 (S.D.N.Y. 2001) ("Here, the 'similarity of the marks' factor is dispositive, for the design and packaging of Alluwishes bears so little resemblance to that of Amanda Love that no reasonable factfinder could conclude that there is any likelihood of confusion between them."); E. Am. Trio Prods., Inc. v. Tang Elec. Corp., 97 F. Supp. 2d 395, 414 (S.D.N.Y. 2000) ("The strength of plaintiff's trade dress, the proximity of the products, and even plaintiff's evidence of actual confusion are, in the eyes of the Court, far outweighed by the fact that the overall appearance of the packaging at issue is so dissimilar. The remainder of the Polaroid factors carry little weight in this analysis for the reasons set forth above.").
144. Cir. Aug. 18, 1999). In affirming the district court's granting of summary judgment to plaintiff, the appellate court stated:
[W]e have neither required district courts to make specific findings on this factor nor mandated reversal when a district court failed to do so. Thus, in the present case, so long as the other factors the district court considered support its ultimate determination that a likelihood of confusion exists, reversal is not mandated simply because the district court failed to make a determination of the similarity, or lack thereof, between the competing marks. Id. at *4. The Ninth Circuit then found that based on the record before it, the marks were indeed similar. Id. at *5.
146. See, e.g., Sun-Fun Prods., Inc. v. Suntan Research & Dev., Inc., 656 F.2d 186, 189 (5th Cir. 1981) ("The two marks must bear some threshold resemblance in order to trigger inquiry into extrinsic factors.
... ). See also Kirkpatrick, supra note 14, at § 4:1 ("Without a threshold similarity of the marks that might result in confusion, it may even be unnecessary to weigh the other factors."). Cf Fisons Horticulture, Inc. v. Vigoro Indus., Inc., 30 F.3d 466, 476 n. 11 (3d Cir. 1994) ("We have emphasized the importance of the similarity of the marks in likelihood of confusion, but we have not ranked the factors otherwise." (citation omitted)).
147. See Playmakers, LLC v. ESPN, Inc., 297 F. Supp. 2d 1277, 1282 (W.D. Wash. 2003) ("Without similarity, there can be no confusion."). overcome even this initial hurdle. Perhaps in these opinions, courts decided the outcome of the close case and then conformed their finding under the similarity factor to that outcome. However, in the sixty-five opinions in which courts found that the similarity factor disfavored a likelihood of confusion, courts nevertheless found that only about half of the other factors considered also disfavored a likelihood of confusion. 148 The courts declined, in other words, to conform the outcomes of other, less important factors to the outcome of the overall test.
As for the degree of similarity between the marks, the data call into question the conventional wisdom that plaintiffs stand a better chance of establishing a likelihood of confusion when the parties' marks are identical rather than merely similar. 1 49 In seventy (21%) of the 331 opinions sampled, the court explicitly found that the marks were "identical," the "same," or "nearly" so. 5 ° For opinions in which the parties' goods were found to be competitive, the plaintiffs' multifactor test win rate when the parties' marks were found to be identical (.881)"'1 was higher than, but not significantly different from, the plaintiffs' win rate when the parties' marks were found merely to be similar, but not identical (.806). 152 The same is true for opinions in which the parties' goods were found not to be competitive. In these opinions, the plaintiffs' multifactor test win rate when the parties' marks were found to be identical (.363) 1 3 was again higher than, but not significantly different from, the plaintiffs' win rate when the parties' marks were found to be merely similar (.214).1 4 The implications of the similarity factor data for the reform of the multifactor test are not completely clear. What is certain, though not easily shown empirically, is that the similarity inquiry in trademark law, as in 148. As Table 3 shows, the multifactor stampede score in such opinions was -.502. 149. See, e.g., 3 McCarthy, supra note 2, at § 23:20 ("Cases where a defendant uses an identical mark on competitive goods hardly ever find their way into the appellate reports. Such cases are 'open and shut' and do not involve protracted litigation to determine liability for trademark infringement."), quoted in Wynn Oil Co. v 2004) (arguing that "a rule conclusively presuming confusion when marks are identical and the defendant competes directly for the same consumers will reduce administrative costs and eliminate erroneous acquittals and their associated costs"). copyright law, 1 5 5 is a frustratingly nebulous and unsystematic inquiry, one that is typically little more than an exercise in abstract formal comparison. 1 5 6 Cognitive science has struggled to define and discipline the notion of similarity, 5 7 and trademark doctrine has done no better. The doctrine urges judges to consider similarities of "sound, sight, and meaning,"' ' 58 to view each mark as a whole,' 59 and to emphasize similarities over differences. '16 But beyond that, the inquiry is wide-open and not well-suited to the fact-finding purpose of the multifactor test. It may be reassuring to discover that, in practice, a court must find the marks to be similar if it is also to find an overall likelihood of confusion, and furthermore, that the degree of similarity of the marks does not appear to significantly affect the outcome of the test; instead, having crossed the similarity threshold, the plaintiff must rely on other factors, some of them intensely empirical in orientation, to prevail in the ultimate analysis. Nevertheless, it is undeniable that those plaintiffs who do win the similarity factor tend also to win the multifactor test. In preliminary injunction and bench trial opinions, 83% of plaintiffs who won the similarity factor won the test. The corresponding figure is 90% in opinions addressing summary judgment motions brought by the plaintiff. These kinds of win rates are tantalizing, though far from conclusive, evidence that the formal inquiry as to similarity exerts an inordinate degree of influence, amounting to a de facto strong presumption, on the outcome of what should otherwise be an essentially empirical rather than formal fact-finding analysis.
See, e.g., William Patry, Does the Substantial Similarity Analysis Make Sense?, The Patry Copyright Blog, Sept. 8, 2005, http://williampatry.blogspot.com/2005/09/does-substantial-similarity-analysis.html (last visited Sept. 16, 2006).
156. See, e.g., Thane Int'l, Inc. v. Trek Bicycle Corp., 305 F.3d 894, 903 (9th Cir. 2002) ("OrbiTrek contains the two syllable prefix 'Orbi,' while TREK does not. So OrbiTrek has three times as many syllables as TREK and twice as many letters. In Entrepreneur Media we held that a reasonable fact finder could find 'Entrepreneur' dissimilar from both 'Entrepreneur Illustrated' and 'EntrepreneurPR.' In particular, we reasoned that 'Entrepreneur Illustrated' is almost twice as long-to both the eye and ear-as 'Entrepreneur,' and 'EntrepreneurPR' contains two more syllables than 'Entrepreneur.' Moreover, the TREK trademark appears with all four letters capitalized, distinguishing it visibly from OrbiTrek." (citation omitted)). Cir. 1994) (stating that, in evaluating a trademark infringement claim, "similarities between marks should be given more weight than differences").
In their treatment of the similarity factor, then, judges may often quite successfully employ a kind of Take The Best heuristic to come quickly to a conclusion, especially when that conclusion is that there is no likelihood of confusion in light of the clear formal dissimilarity of the marks.' 6 ' But judges should also be wary of allowing their initial intuitions, formed in considering the formal similarity of the marks, to predispose the remainder of their analysis. Here, the data with respect to a finding that the marks are identical is indeed reassuring. Such a finding does not appear to bias the remainder of the analysis any more than does a finding that the marks are merely similar.' 62
Courts have long expressed conflicting views on what role the intent factor should and does play in the multifactor test. Some circuits have held, as the Second has, that a finding of bad faith intent creates a "rebuttable legal presumption that the actor's intent to confuse will be successful."' 6 3 Others have held that a finding of bad faith intent may "justify the inference" 164 of confusing similarity or is entitled to "great weight,"' 16 5 but have declined to establish a presumption.1 66 161.
Cf Nat'l Collegiate Athletic Ass'n. v. Bd. of Regents of Univ. of Okla., 468 U.S. 85, 110 n.39 (1984) ("The essential point is that the rule of reason can sometimes be applied in the twinkling of an eye." (citation omitted)).
162. There is no substantial variation among the circuits in the wording of the similarity factor itself or in the doctrine underlying the factor. The circuits tend to phrase the factor as simply the "similarity of the marks" or the "degree of similarity between the marks." ("Evidence of conscious imitation is pertinent because the law presumes that an intended similarity is likely to cause confusion."); My-T Fine Corp. v. Samuels, 69 F.2d 76, 77 (2d Cir. 1934) (Hand, J.) ("[S]uch an intent raises a presumption that customers will be deceived."). Cf Kirkpatrick, supra note 14, at § 8:3.2 ("Some panels of the Second Circuit Court of Appeals still honor the presumption, sometimes in the breach."). See also Osem Food Indus. Ltd. v. Sherwood Foods, Inc., 917 F.2d 161, 16 U.S.P.Q.2d 1646, 1649 (4th Cir. 1990) ("Logic requires.., that from such intentional copying arises a presumption that the newcomer is successful and that there is a likelihood of confusion. It would be inconsistent not to require one who tries to deceive customers to prove they have not been deceived.").
164. P]roof of intent to cause confusion is entitled to great weight, not that it creates a presumption of confusion that shifts the burden of proof to the other party") (emphasis in original). Cf Nautilus Group, Ind. v. Icon Health & Fitness, Ind., 372 F.3d 1330, 1337 (Fed. Cir. 2004) ("Recently... the At the same time, certain circuits have declared that "intent is largely irrelevant in determining if consumers likely will be confused as to source." 1 6 7 As one district court put it, "the presence or absence of intent does not impact the perception of consumers whose potential confusion is at issue."' 68 Notwithstanding the presumptive force it appears to accord to the intent factor, the Second Circuit has been particularly critical of it. Judge Leval, the circuit's most influential intellectual property jurist, has addressed intent at length: Bad faith on the part of a party can influence the court in at least two ways. First, where a substantive issue such as irreparable harm or likelihood of confusion is a close question that could reasonably be called either way, a party's bad faith could cause it to lose the benefit of the doubt. Second, if prospective entitlement to relief has been established, the good or bad faith with which the parties had conducted themselves could influence the court in the fashioning of appropriate equitable relief, or even cause it to deny equitable relief to a party that had conducted itself without clean hands. A preliminary injunction can have drastic consequences-potentially putting a party out of business prior to trial on the merits. A court may be less concerned about imposing such drastic consequences on a party that had conducted itself in bad faith. 1 6 9 Other circuits, such as the Fifth, have declared the intent factor to be "critical"' 17 1 or "important"'' in the limited sense that while the "presence of intent constitutes strong evidence of confusion, the absence of intent is Ninth Circuit may have de-emphasized somewhat the role of the intent factor in the likelihood of confusion analysis.").
166. Cir. 1992) (referring to the intent factor as an "important factor bearing on the likelihood of confusion" (internal quotations omitted)).
[Vol. 94:1581 irrelevant in determining likelihood of consumer confusion."' 1 7 2 Fifth Circuit courts frequently state that "if the mark was adopted with the intent of deriving benefit from the reputation of [the plaintiff,] that fact alone may be sufficient to justify the inference that there is confusing similarity."' 73 However, they add, "the lack of guilt is immaterial to the evaluation."' 1 7 4 Courts that follow this approach typically reason that "one who intends to confuse is more likely to succeed in doing so"' 1 7 5 or that the defendant's intent "is relevant because it demonstrates the junior user's true opinion as to the dispositive issue, namely, whether confusion is likely."' 7 6 Some hold more precisely that "defendant's intent will indicate a likelihood of confusion only if an intent to confuse consumers is demonstrated via purposeful manipulation of the junior mark to resemble the senior's." 177 The data strongly reject the hypothesis that the intent factor is irrelevant to the outcome of the multifactor test. In fact, they suggest that a finding of bad faith intent creates, if not in doctrine, then at least in practice, a nearly un-rebuttable presumption of a likelihood of confusion. 17 ' All but one of the fifty preliminary injunction opinions in which the court found bad faith intent resulted in a finding of a likelihood of confusion, 179 and all but one of the seventeen bench trials in which the court found bad faith ("[A] defendant who purposely chooses a particular mark because it is similar to that of a senior user is saying, in effect, that he thinks that there is at least a possibility that he can divert some business from the senior user-and the defendant ought to know at least as much about the likelihood of confusion as the trier of fact.").
177. A & H Sportswear, Inc. v. Victoria's Secret Stores, Inc., 237 F.3d 198, 226 (3d Cir. 2000). 178. This result is all the more remarkable in light of the fact that the data set excluded counterfeiting cases, in which the defendant's bad faith intent is typically quite clear.
179. For the one preliminary injunction opinion sampled in which bad faith intent was found but a likelihood of confusion was not found, see Do the Hustle, LLC v. Rogovich, 03 Civ. intent produced the same result. "' In both of the outlying opinions, a finding of dissimilarity between the marks trumped the finding of bad faith intent,' and in both, the finding of intent was weak. 1 8 2 The force of a finding of bad faith intent is no less apparent in the district courts of the Second Circuit, which routinely cite to Judge Leval's reasoning on the intent factor, 83 but are otherwise no different from other district courts in appearing to treat a finding of bad faith intent as nearly dispositive. ' 8 4
As with the similarity factor, the precise wording and application of the intent factor does not vary substantially by circuit with one exception: the Seventh Circuit. Logistic regression analysis suggests that, all else being equal, district courts in the Seventh Circuit are less likely to find that the intent factor favors a likelihood of confusion than are the district courts of the other circuits-the coefficient was only marginally significant, however.' 85 It is not clear why the Seventh Circuit should stand out in this way. One possible explanation goes to the wording of the intent factor in that circuit. While most circuits phrase their intent factor as simply "the defendant's intent"' 86 or "the defendant's intent in selecting the mark,"' 87 the Seventh Circuit is unique in inquiring into the defendant's "intent to palm 180. For the one bench trial opinion sampled in which bad faith intent was found but a likelihood of confusion was not found, see Choice Hotels Int'l, Inc. v. Kaushik, 147 F. Supp. The bad faith element may weigh in favor of Plaintiffs, since Rogovich was admittedly involved from near the inception of Breakfast Club and was likely responsible for many of the similarities between the two establishments, in violation of the Purchase Agreement. Despite Rogovich's assertion in his submissions and at the Hearing that he mistakenly understood the Purchase Agreement to mean 30 miles from Manhattan as opposed to the plain language, 'New York City,' such a lack of understanding of the plain words in the contract does not excuse his breach."); Choice Hotels, 147 F. Supp. 2d. at 1253-54 ("The court notes that it makes this finding in the absence of any evidence of actual intent to infringe ... [Kaushik's] acts were intentional under the law, while not intentional in the literal sense of the word.").
183. [Vol. 94:1581 off its goods" as those of the plaintiff. 1 88 This phrasing, which comes from the original Helene Curtis opinion that established the Seventh Circuit's factors, refers to the old doctrine of "palming off'--or "passing off'--as it is otherwise known.' 89 Arguably, the Seventh Circuit's phrasing of the intent factor calls for something more than a mere general finding of bad faith, which may be made in situations where the defendant knowingly adopted a mark to which the plaintiff had a prior claim, but where the defendant did not, strictly speaking, intend to cause consumer confusion. The Seventh Circuit's phrasing calls instead for a more specific finding that the defendant knowingly adopted a mark similar to the plaintiffs with the intent to confuse consumers as to source. 90 It is black-letter doctrine across the circuits, including the Seventh Circuit, that bad faith intent may be inferred solely from the fact that the parties' marks are similar and the fact that the defendant had knowledge of the plaintiff's mark when it adopted its own, similar mark. 191 The data suggest that this circumstantial inference is the leading basis for a finding of bad faith intent. District courts found bad faith in 102 of the 331 opinions sampled. In fifty-eight of these 102 opinions, the court based its finding of bad faith at least in part on the combination of similarity and defendant's knowledge. 92 In thirty of these fifty-eight opinions, the court also based its finding on direct evidence of bad faith, such as documents produced by the defendant or actions of the defendant after receiving a cease and desist T]he relevant intent is not just the intent to copy, but to 'pass off one's goods as those of another. Given that Bridgewater prominently displayed its trade name on its candles, we do not think that the evidence of copying was sufficiently probative of secondary meaning." (citation omitted)). Cf Blanchard v. Hill, 2 Atk. 484, 26 Eng. Rep. 692 (Dec. 18, 1742) (holding that the defendant's use of a mark identical to plaintiffs on identical goods was not actionable in the absence of fraudulent intent to pass off defendant's goods as those of plaintiff). Cf Kirkpatrick, supra note 14, at § 1: 1.3 ("In sum, the terms 'passing off and 'palming off principally serve plaintiffs as colorful, pejorative accusations of literally underhanded misconduct by defendants.").
191. See, e.g., Bliss Clearing Niagara, Inc. v. Midwest Brake Bond Co., 339 F. Supp. 2d 944, 967 (W.D. Mich. 2004) ("Direct evidence of intentional copying is not necessary to prove intent. Rather, the use of a contested mark with knowledge of the protected mark at issue can support a finding of intentional copying." (citation omitted)); Pfizer, Inc. v. Y2K Shipping & Trading, Inc., No. 00 CV 5304(SJ), 2004 WL 896952, at *5 (E.D.N.Y. Mar. 26, 2004) ("Bad faith... is established where there is evidence of actual knowledge of the senior user's mark and the marks are so similar that it seems clear that deliberate copying has occurred." (citation omitted)). demand from the plaintiff. 9 3 In only thirty-seven of the 102 opinions did the court base its finding solely on direct evidence without explicitly mentioning the combination of similarity and defendant's knowledge. ' 94 Finally, what light do the data shed on the conflict among the courts concerning the proper role of intent in the multifactor analysis? Even more so than the similarity factor data, the intent factor data suggest that a finding of bad faith intent exerts excessive influence on the outcome of the multifactor test. The facile assumption, evidently quite pervasive among the courts, that if the defendant intended to confuse, then it succeeded in doing so does not do justice to the great diversity of trademark infringement fact patterns before the courts. Further, it loosens the focus of the multifactor analysis on what should be the overriding empirical question of whether consumers are likely to be confused."' 9 To be sure, in light of the defendant's bad faith, courts employ the multifactor test to reach what they deem to be the right result. But if trademark law seeks to prevent commercial immorality, then it should do so explicitly. An injunction should issue and damages be granted on that basis alone, and not on the basis of possibly distorted findings of fact as to the likelihood of consumer confusion.
As they have with the intent factor, courts have expressed conflicting views about the importance of the proximity of the goods factor. The purpose of the proximity factor is to consider whether "the [parties'] goods are similar enough that a customer would assume they were offered by the same source,"' 9 6 and also to consider whether "buyers and users of each parties' goods are likely to encounter the goods of the other, creating an assumption of common source affiliation or sponsorship."' ' 9 7 We saw above that the Sixth Circuit has called the proximity factor "the most important inquiry" in the multifactor analysis. 9 8 Another court has referred to it as "extremely important." ' 199 Yet the Second Circuit has speculated that "since modem marketing methods tend to unify widely different types of products in the same retail outlets or distribution networks, this factor is not of overriding importance. 2 0 The Sixth Circuit has also more precisely characterized the proximity factor as important at the extremes, when it weighs strongly in favor of one party or the other, but not important otherwise:
This court has identified three categories regarding the relatedness of the goods or services with which trademarks are associated: First, if the parties compete directly by offering their goods or services, confusion is likely if the marks are sufficiently similar; second, if the goods or services are somewhat related but not competitive, the likelihood of confusion will turn on other factors; third, if the goods or services are totally unrelated, confusion is unlikely. 20 '
The data support this account. In preliminary injunction and bench trial opinions in which the parties' goods were found to be identical and their marks were found to be similar, plaintiffs won the multifactor test 100% of the time (n=28). In summary judgment opinions of this nature, plaintiffs' multifactor test win rate was also very high (.813).22 On the other extreme, in all opinions, regardless of posture, in which the proximity factor was found to disfavor a likelihood of confusion, the plaintiffs' multifactor test win rate was exceedingly low (.027).23 Finally, in those opinions in which the proximity factor was found to favor a likelihood of confusion but the goods were not identical, the plaintiffs win rate was unexceptional (.624),04 suggesting that the outcome of the test did indeed turn on other factors.
Overall, the degree of proximity of the goods appears significantly to affect the outcome of the test. We saw above, under the similarity factor, that the plaintiffs' win rate in opinions in which the parties' marks were found to be identical was not significantly different from their win rate in opinions in which the parties' marks were found to be similar but not identical. Under the proximity factor, however, the story is different. The overall plaintiff win rate in identical goods cases (.808)205 was significantly higher than the plaintiff win rate (.624)206 in opinions that found the parties' goods to be similar, yet not identical. 2 0 7 If the ideal multifactor analysis seeks only to discover the facts on the ground, then the proximity data should be seen as encouraging. Courts appear still to pay attention to the specific circumstances of the marketplace rather than rely on such general principles as those articulated above by the Second Circuit. Relatedly, the proximity data are also an interesting corrective to the general tendency of trademark policymaking 2 8 and scholarship 209 to focus on the big marks and big marketplaces. The proximity data suggest that the daily workings of trademark litigation are still characterized by the clash of small, or at least non-famous, marks, for which the proximity of the goods consideration is still important.
Courts assess the degree of strength or distinctiveness of the plaintiffs mark on the assumption that the stronger the mark, the more protection it should receive. In conducting their analysis, courts are generally instructed to consider two forms of trademark distinctiveness:
[I]nherent distinctiveness[] examines a mark's theoretical potential to identify plaintiffs goods or services without regard to whether it has actually done so. . . . [A]cquired distinctiveness[] refers to something entirely different. This measure looks solely to that recognition plaintiffs mark has earned in the marketplace as a designator of plaintiff's goods or services. 21 0 The data on the strength factor yield what are probably the most interesting factor-specific results in the study, if also the most ambiguous. The data suggest that, at least in the context of the multifactor test, the doctrine of trademark strength has broken down. Basic concepts are no longer consistently applied and mistakes of doctrine are common." 1 ' Nevertheless, in [Vol. 94:1581 opinions that do address the issue of trademark strength, and inherent strength in particular, there is a surprisingly good correlation between inherent strength and success in the multifactor test.
Established by Judge Friendly in 1976, the Abercrombie spectrum of trademarks classifies marks according to their degree of inherent distinctiveness."' Fanciful marks are coined terms and are thought to have the highest degree of inherent distinctiveness (i.e., xerox). Second in the hierarchy are arbitrary marks, which have no semantic connection to the products to which they are affixed (i.e., apple computers). Third are suggestive marks that are suggestive of or metaphorically related to their products' characteristics (i.e., ivory soap). Descriptive marks (i.e., coca-cola) are thought to lack inherent distinctiveness, as are generic marks (i.e., aspirin 2 3 ). It remains black-letter doctrine that the more inherently distinctive a mark is, the greater the scope of protection it should receive. Nevertheless, courts have occasionally expressed their dissatisfaction with the Abercrombie hierarchy. Judge Easterbrook in particular has criticized its formalism:
We have said before that "arbitrary," "suggestive" and the other words in the vocabulary of trademark law may confuse more Cir. 1976) ("The cases, and in some instances the Lanham Act, identify four different categories of terms with respect to trademark protection. Arrayed in an ascending order which roughly reflects their eligibility to trademark status and the degree of protection accorded, these classes are (1) generic, (2) descriptive, (3) suggestive, and (4) arbitrary or fanciful."). See also Wal-Mart Stores, Inc., 529 U.S. at 210-11 ("[W]ord marks that are 'arbitrary' ('Camel' cigarettes), 'fanciful' ('Kodak' film), or 'suggestive' ('Tide' laundry detergent) are held to be inherently distinctive.").
213. The Bayer Company originally coined the term "aspirin" as a trademark for acetyl salicylic acid. In 1921, Judge Learned Hand found that aspirin had lost its significance as the designation of a particular source of acetyl salicylic acid and had become a generic term for the substance itself. seriously before arguing cases so that everything turns on which word we pick. it is better to analyze trademark cases in terms of the functions of trademarks. That frees the arguments from the clutches of Webster's Third and the conflicting advice of text writers. 2 , 4 As for the bar, one respected trademark practitioner and commentator, the late Beverly W. Pattishall, went so far as to refer to the "artificial and regrettable 'four pigeon hole' rule" established in Abercrombie as "[o]ne of the worst blights [on the law] ... which has spread from the Second Circuit and now appears to be settling in generally. 2 1 5 The data suggest that Mr. Pattishall need not have worried, and that Judge Easterbrook's altogether sensible call has been answered. First, in the context of the likelihood of confusion inquiry, district courts appear to make little use of the Abercrombie spectrum and the concept of inherent distinctiveness that underlies it. Courts failed to specify whether or not the mark at issue was inherently distinctive in 40% of the 192 preliminary injunction and bench trial opinions sampled and in 50% of the 139 summary judgment opinions sampled, for an overall failure rate of 44% in the 331 opinions examined. Overall, only 193 or 58% of the 331 opinions sampled made some use of the Abercrombie spectrum, and twenty-nine of these opinions neglected to place the plaintiff's mark in one specific Abercrombie category. Instead, they opted to make such estimations as that the mark was "suggestive... even though it contains descriptive elements," ' 216 "fanciful and arbitrary," 2 7 "arbitrary or suggestive," 218 or is "at least a suggestive mark and is arguably an arbitrary or fanciful mark. 21 9 The breakdown of the Abercrombie analysis is even more apparent in the context of claims for trade dress infringement, in other words, for the infringement of a product's packaging or configuration. 220 The court failed to specify where in the Abercrombie spectrum the trade dress fell in twentyseven of the fifty-three opinions (51%) that adjudicated a claim only for trade dress infringement and in nine of the eleven opinions (82%) that adjudicated a claim for both trademark and trade dress infringement. 2 Second, and more strikingly, in those opinions in which the court's assessment of the mark's inherent strength was at odds with its assessment of the mark's acquired strength, a finding of acquired strength (or weakness) almost invariably trumped a finding of inherent weakness (or strength). For example, courts found marks to be inherently weak but commercially strong in twenty-three of the opinions sampled. In twentytwo of these opinions, the court found that the strength factor favored confusion. The one outlying opinion found, somewhat ambiguously, "a mark strength somewhere between strong and weak." ' 222 As for the inverse situation, courts found marks to be inherently strong but commercially weak in twenty-seven of the opinions sampled. In twenty-four of these opinions, the court found that the strength factor did not favor confusion. Of the three outliers, one found that the strength factor "favors neither party, ' 2 23 another explicitly narrowed its finding of strength, 2 4 and a third found as it did because the plaintiff failed to present any evidence of commercial strength. 25 Ultimately, these results should not be surprising. Though most appellate courts have not yet come around to acknowledging it, district courts appear in practice to have recognized that the mark's acquired or "actual strength" in the marketplace logically incorporates the effects of the mark's inherent strength. Indeed, the results suggest that courts need not even consider inherent strength in their assessment of the strength factor, or that if they do, inherent strength should properly be understood as merely one factor-among others such as advertising expenditure, length of time of use of the mark, revenues associated with the mark, and third-party uses 22 -that a court should consider in assessing a mark's actual strength. 27
Despite the apparent infirmity of the concept of inherent strength, we cannot close the book on it entirely. When courts did address the issue of the inherent strength of the plaintiffs mark, their findings correlated quite cleanly with the plaintiff multifactor test win rate, as Table 7 reveals. The multifactor test win rate in dispositive opinions for inherently distinctive marks (.628)22 was significantly higher than the win rate in such opinions for non-inherently distinctive marks (.242).229 More specifically, in the ninety dispositive opinions in which the court placed the plaintiffs mark in one of the five Abercrombie categories, the plaintiff multifactor test win rate steadily declined with the inherent strength of its mark: fanciful marks enjoyed the highest win rate, followed by arbitrary marks, suggestive marks, descriptive marks, and then generic marks. More specifically still, and underlying these win rate results, inherently distinctive marks did better on each of the core factors, and the degree of their inherent strength often closely tracked the proportion of opinions in which each of the core factors favored a likelihood of confusion.
These data support two important propositions, one championed by the trademark bar and the other by commentators on trademark doctrine. First, the data support the trademark lawyer's common advice to her clients that, at least as a matter of trademark law, if not of marketing, firms should choose inherently distinctive marks and, ideally, marks that are fanciful. As for the proposition of trademark commentators, courts frequently group into the same category trademarks that are fanciful and those that are arbitrary. Judge Friendly made this mistake when he first formulated the Abercrombie spectrum,"' and the Supreme Court has since failed to correct the likelihood of confusion."' (quoting Restatement of Torts § 729 (1938)). Of the sixty-three opinions that considered the effect of third-party uses on the strength of the plaintiff's mark, thirty-nine found that these uses mitigated strength (and thirty-three of the thirty-nine found no overall likelihood of confusion), while twenty-four found that third-party uses did not mitigate strength.
Another formalism has had a deleterious effect on the strength inquiry. This is the principle that incontestable marks are presumptively strong. See, e.g., Data Concepts, Inc. v. Digital Consulting, Inc., 150 F.3d 620, 625 (6th Cir. 1998) ("A mark that has been registered and uncontested for five years... is entitled to a presumption that it is a strong mark."). The data suggest that courts make only limited use of this principle. Of the thirty out of 331 opinions that addressed the incontestable status of the plaintiff's mark as part of the strength inquiry, twenty-two found that this status supported a finding of strength (and 21 eventually found that the mark was strong), while eight found that this status did not support a finding of strength. Most of these opinions came from the Second, Sixth, and Eleventh Circuits. error, if not compounded it. 231 Trademark commentators have argued that the two kinds of trademarks are in fact quite different, and that fanciful marks deserve a heightened degree of protection over arbitrary marks.1 3 2 This argument is typically made on the basis of functionality concerns. 2 3 3 The data bolster this proposition from a different angle. Perhaps most fanciful marks manage to achieve, due to their fanciful nature, greater actual strength in the marketplace than arbitrary marks, and thus they do better than arbitrary marks in trademark infringement litigation. But we should be clear that the relative success of fanciful marks in trademark infringement litigation is not due simply to their ability to satisfy some category of trademark doctrine-the data show that courts place little weight on the doctrine of inherent strength. Rather, their relative success appears to be due to the degree to which their inherent strength manifests itself in the form of actual marketplace strength.
Ultimately, as with the intent factor, the doctrine of inherent strength threatens to distort the fact-finding inquiry as to the likelihood of consumer confusion by insinuating into that inquiry policy-oriented goals that are better served elsewhere. With intent, the goal is to discourage commercial immorality. Here, with the doctrine of inherent strength, the goal is encourage the use of inherently distinctive rather than descriptive marks. These are both worthy objectives, but courts should not pursue them when they are making findings of fact about the likelihood of consumer confusion. Perhaps the standard for a finding of a likelihood of confusion should be lowered when the defendant has acted in bad faith or when the plaintiff is using an inherently distinctive mark, but the court's estimate of the likelihood of consumer confusion in the marketplace should not simply be raised in light of the presence, without more, of bad faith intent or an inherently distinctive mark. It is at least reassuring that courts appear already to have recognized this principle when they have encountered an inherently distinctive mark.
The doctrine underlying the strength factor varies considerably among the circuits. District courts in the First, Sixth, Eighth, and Ninth Circuits consider somewhat eccentric sub-factors when evaluating trademark strength. 234 The Ninth Circuit's are arguably the most peculiar: Two tests are commonly used to measure the strength of a mark, the "imagination test" and the "need test." The "imagination test" focuses on the amount of imagination required in order for a consumer to associate a given mark with the goods or services it identifies. If a consumer must use more than a small amount of imagination to make the association, the mark is suggestive and not descriptive. The "need test" focuses on the extent to which a mark is actually needed by competitors to identify their goods or services.... The two tests are related, because the more imagination that is required to associate a mark with a product or service, the less likely the words used will be needed by competitors to describe their products or services. 3 5 It is unclear how either of these tests helps to determine the actual strength of the mark in the marketplace. For example, "United Airlines" is hardly imaginative, and the term "united" is certainly needed by competitors, yet it is generally thought to be a very strong mark.236
Despite the wide diversity of circuit-specific doctrine underlying the strength factor, regression analysis demonstrates no significant intercircuit variation in the application of the factor, not even in the Ninth Circuit.
If the factor-specific results relating to the strength factor are the most interesting in the study, the results relating to the actual confusion factor are the most disturbing. Courts consider two forms of evidence of actual confusion: (1) survey evidence and (2) direct evidence, such as testimony of time a mark has been used and the plaintiffs relative renown in its field; the strength of the mark in plaintiff's field of business, especially by looking at the number of similar registered marks; and the plaintiff's actions in promoting its mark." (citations omitted)). For the Sixth Circuit, see Therma-Scan, Inc. v. Thermoscan, Inc., 295 F.3d 623, 631 (6th Cir. 2002) ("Generally, the strength of a mark is the result of its unique nature, its owner's intensive advertising efforts, or both."). See also Midwest Guar, Bank v. Guar. Bank, 270 F. Supp. 2d 900, 910 (E.D. Mich. 2003) ("The Sixth Circuit has commented that a mark's strength is generally a result of(l) its unique nature; (2) its owner's intensive advertising efforts; and (3) which of the four categories the mark occupies-generic, descriptive, suggestive or arbitrary/fanciful." (citing Therma-Scan, 295 F.3d at 631)). For the Eighth Circuit, see Gateway, Inc. v ("Placement of the mark on the continuum, however, is only the first step; the Court must also determine the strength of the mark in the marketplace. To this end, the Court applies both the 'imagination test' and the 'need test."' (citations omitted)); Id. ("As opposed to 'bold' or 'italics,' the word 'lollipop' does not describe the font; rather, it takes imagination to associate the two. Moreover, the need aspect is low because competitors such as CommCut do not need to use the word 'lollipop' to describe their products. Thus, this factor weighs in favor of Ellison."). by consumers who were confused by the defendant's use of its mark or documents indicating such confusion. I focus here on survey evidence.
It is generally thought that survey evidence is the best evidence of actual confusion, and indeed, that a good survey has the potential to supersede the rest of the multifactor analysis. 237 In a recent statement before Congress, the American Bar Association set forth the conventional view: "survey evidence is traditionally one of the most classic and most persuasive and most informative forms of trial evidence that trademark lawyers utilize in both prosecuting and defending against trademark claims of various sorts. '238 Some circuits even apply an adverse inference of no likelihood of confusion if the plaintiff has the resources and time to produce survey evidence but fails to do so 2 39 -though only six opinions out of the 331 sampled drew such an inference. 240 The data suggest that the conventional view of the utility of survey evidence may be incorrect and that the application of an adverse inference may be inappropriate. Of the 331 opinions sampled, only sixty-five (20%) addressed survey evidence, only thirty-four (10%) credited the survey evidence, and only twenty-four (7%) ultimately ruled in favor of the outcome that the credited survey evidence itself favored. More specifically, of the fifty-three opinions that addressed survey evidence presented by the plaintiff, twenty-two (42%) credited that evidence, with one opinion using the evidence against the plaintiff. Of the nineteen opinions that addressed survey evidence presented by the defendant, twelve (63%) credited that evidence, but three used the evidence against the defendant. Finally, of the seven opinions that addressed survey evidence presented by both the plaintiff and the defendant, three credited the defendant's evidence and none credited the plaintiffs. 240. The data set included two relevant variables: whether the court drew a general adverse inference from the plaintiff's failure to present any evidence of actual confusion and whether the court drew a specific adverse inference from the plaintiff's failure to present survey evidence. Of the 331 opinions sampled, nine drew only a general adverse inference, two drew only a specific adverse inference, and four drew both kinds of adverse inference.
It may be objected that trademark litigation is typically resolved at the preliminary injunction stage before either party has had the time or can be expected to conduct a creditable survey. It is true that the percentage of bench trial opinions that addressed survey evidence (24%) was higher than the percentage of preliminary injunction opinions that did so (16%), and that the percentage of preliminary injunction opinions that explicitly held that it was too soon to expect any evidence of actual confusion (11%) was somewhat higher than the percentage of bench trial opinions that espoused this view (9%). Yet it is still striking that survey evidence played a relatively minor role even in the bench trial context. In only thirteen out of forty-six bench trial opinions did the court address survey evidence, which it credited in eight of those opinions. Further, as stated above, only six opinions out of the 331 sampled explicitly drew an adverse inference from plaintiffs failure to present survey evidence. Four of these were preliminary injunction opinions.
There was no significant intercircuit variation in the courts' consideration of survey evidence. It may be of interest, however, that the district courts of the Second Circuit credited the plaintiffs survey evidence in only seven of the twenty-five opinions in which they considered it, and credited the defendant's survey evidence in only six of the ten opinions in which they considered it. 24 '
We saw above that the non-core factors generally have little, if any, effect on the outcome of the multifactor test. This should not be surprising. Several of these factors have no business being in the multifactor test in the first place, and the courts appear to have recognized this in practice, if not yet in doctrine. Happily, one sees here how district courts, in applying a multifactor test, may resist the dead-hand influence of various idiosyncratic and rarely relevant factors that tend to accumulate over time.
Probably the only non-core factor that deserves to be in the multifactor test is the consumer sophistication factor. It makes sense, and has been confirmed empirically, that the more sophisticated the consumers, the more care with which they will treat their search and purchasing decisions. 242 Nevertheless, a fairly high proportion (16%) of the opinions sampled from circuits that include the factor in their multifactor test either failed to address the factor or stated that it was not argued by the parties. The majority of these opinions were preliminary injunction or bench trial opinions. On the whole, across the 292 opinions sampled from circuits that explicitly consider the factor, the factor was found to disfavor a likelihood of confusion (that is, consumers were seen as sufficiently sophisticated not to be confused) 39% of the time and to favor a likelihood of confusion 28% of the time. Interestingly, although no variation exists among the circuits in the doctrine underlying the factor, the data reveal some variation among the circuits in terms of test outcomes: Logistic regression analysis suggests that the Second Circuit is significantly less likely than other circuits to find that the consumer sophistication factor disfavors a likelihood of confusion, which, if win rates are any guide, runs counter to the Second Circuit's apparent bias against finding a likelihood of confusion. 2 43 It is not clear, and probably doubtful, however, that this had any effect on the overall outcome of the multifactor test.
The factors relating to the similarity of the parties' advertising, marketing, and sales facilities all tended to be redundant of the proximity of the goods factor in the circuits that consider these issues separately from the proximity factor. For example, among the seven possible outcomes coded for each factor, the similarity of the sales facilities factor produced exactly the same outcome as the proximity factor in 74% of the opinions sampled (n=90) that considered both factors. Of the twenty opinions in which the two factors produced divergent outcomes, six did so because they found that the proximity factor favored a likelihood of confusion and then simply did not address the sales facilities factor. The similarity of advertising or marketing methods factor produced exactly the same outcome as the proximity factor in 66% of the relevant cases sampled (n=156), and of the fiftythree opinions in which these two factors produced divergent outcomes, twelve did so because, again, they found that the proximity factor favored a likelihood of confusion and then neglected to address the advertising or marketing methods factor. The Second, Eighth, Tenth, and D.C. Circuits already consider proximity, advertising, marketing, and sales facilities together under the proximity factor. 2 ' The data suggest that, at least in practice, the district courts of the other circuits do so as well.
Also redundant of the core factors are two factors used by the Third Circuit. Of the twenty-eight opinions sampled from the Third Circuit,
Logistic regression of the consumer sophistication factor was performed on a dummy variable for the Second Circuit and dummy variables for the two most common outcomes (favors or disfavors confusion) of each of the five core factors. The same regression was performed for each of the circuits with dummy variables for the circuit. Only the Second Circuit yielded a marginally significant coefficient (z--1.89, p>lzl=.058, N=292, x=95.96, p>x2=.000, pseudo R 2 =.245).
244. See, e.g., Best Cellars, Inc. v. Grape Finds at Dupont, Inc. 90 F. Supp. 2d 431, 456 (S.D.N.Y. 2000) ("The 'proximity-of-the-products' inquiry concerns whether and to what extent the two products compete with each other. The court must consider 'the nature of the products themselves and the structure of the relevant market,' including 'the class of customers to whom the goods are sold, the manner in which the products are advertised, and the channels through which the goods are sold." (citations and internal quotations omitted)). twenty (71%) explicitly grouped their analysis of the factor relating to the length of time of concurrent use without evidence of actual confusion with their analysis of the actual confusion factor. Of the eight opinions that did not group their analysis, none reached substantively different outcomes under the two factors. The Third Circuit also considers the extent to which the targets of the parties' sales efforts are the same. Of the twenty-eight opinions from the circuit, seventeen (61%) produced exactly the same outcome under this factor as under the proximity factor, and only three of the remaining eleven opinions produced substantively different outcomes under the two factors.
Finally, the likelihood of bridging the gap factor, and the comparative quality of the parties' goods factor, are remarkable for the degree to which courts either ignore them or bend them to conform to the outcome of the test. Of the 217 opinions sampled from the five circuits that consider the bridge the gap factor, fifty-seven (26%) did not address the factor and another forty-four (20%) explicitly found it to be irrelevant. Strangely, twenty (nearly half) of the opinions that found the factor to be irrelevant nevertheless held that it favored a likelihood of confusion on the grounds that "there is no gap to bridge. As for the quality factor, which is considered only by the Second and D.C. Circuits, twenty-one of the 109 opinions from these circuits did not address the factor. Of the eighty-eight opinions that did, forty-five subscribed solely to the tarnishment theory of the factor, 246 three subscribed solely to the similarity theory of the factor, 2 47 and twenty-one opinions subscribed to both theories. 24t Perhaps more than any other, the quality factor 245.
Id. at 456 ("Here, there is no gap to bridge: Best Cellars and Grape Finds sell the same products in the same field. This factor, therefore, also favors Best Cellars."). See also Blue & White Food Prods. Corp. v. Shamir Food Indus., Ltd., 350 F. Supp. 2d 514, 521 (S.D.N.Y. 2004) ("The fourth factor favors Blue & White because if Shamir Food were permitted to bring its own 'Shamir' products into the market, there would be no gap to bridge."); Macia v. Microsoft Corp., 335 F. Supp. 2d 507, 518 (D. Vt. 2004) ("As the products compete directly there is no 'gap' to be bridged.").
246. The tamishment approach to the comparative quality factor posits that if the defendant's goods are of lower quality, then this will tarnish plaintiff's reputation and the factor should thus somehow favor a likelihood of confusion. See, e.g., Prof I Sound Servs., Inc. v. Guzzi, 349 F. Supp. 2d 722, 735 (S.D.N.Y. 2004) ("The seventh factor, quality of defendants' products, asks 'whether the senior user's reputation could be jeopardized by virtue of the fact that the junior user's product is of inferior quality."' (citation omitted)).
247. The similarity approach to the comparative quality factor posits that if the defendant's goods are of similar quality to or even higher quality than the plaintiffs, then consumers are more likely to confuse the two parties' marks. See, e.g., Landscape Forms, Inc. v. Columbia Cascade Co., 117 F. Supp. 2d 360, 367 (S.D.N.Y. 2000) ("There is also little dispute that the two product lines are of similar quality-a factor also weighing in plaintiffs favor.").
248. See Hasbro, Inc. v. Lanard Toys, Ltd., 858 F.2d 70, 78 (2d Cir. 1988). ("The next factor, quality of the junior user's product, is the subject of some confusion. One view is that an inferior quality product produced by the junior user injures the senior user's reputation insofar as consumers might think that the source of the inferior product is the senior user. Another view is that a junior user's product of equal quality to a senior user's product injures the senior user by the increased tendency of similar quality products to promote consumer confusion." (citations omitted)).
is an embarrassment to the multifactor test, and not simply because tarnishment should have no relevance to a finding of fact as to the likelihood of consumer confusion, nor because similarity in quality should already be addressed under the proximity factor, but because the factor is so utterly pliable.
There is a special irony in an empirical study of the multifactor test for the likelihood of consumer confusion. The test itself is essentially a substitute for empirical work. Ideally, a court would determine the likelihood of consumer confusion by taking testimony from every consumer who has been or will be exposed to the plaintiffs and defendant's marks. The court would then establish what proportion of this population of consumers is or will be confused, and decide whether that proportion is sufficiently high to justify a ruling in favor of the plaintiff. The court, in other words, would conduct a survey. But because a court lacks the time, resources, and capacity to do so, 249 it must instead consider a variety of factors designed to help it estimate the results of that ideal survey. These are the factors of the multifactor test.
When the multifactor tests of the various circuits are held up against the standard of the ideal survey,... most, if not all of them, are found wanting. Certain of their factors as well as their overall design often distract from their ultimate purpose: to estimate what is actually occurring or will occur in the marketplace. Clearly, considerations such as the comparative quality of the parties' goods or the inherent distinctiveness of the plaintiffs mark rarely aid in this inquiry. In the case of intent, the precise wording of the factor appears to affect courts' analyses. More generally, multifactor tests of ten or even eight factors appear to ask too much of the judge's ability simultaneously to weigh competing concerns and may simply result in 249. But see Triangle Publ'ns, Inc. v. Rohrlich, 167 F.2d 969, 974 (2d Cir. 1948) (Frank, J., dissenting). In Triangle Publications, involving the trademark SEVENTEEN for the plaintiffs magazine and for the defendant's girdles, Judge Frank took it upon himself to conduct his own survey:
Like the trial judge's, our surmise must here rest on "judicial notice." As neither the trial judge nor any member of this court is (or resembles) a teen-age girl or the mother or sister of such a girl, ourjudicial notice apparatus will not work well unless we feed it with information directly obtained from "teen-agers" or from their female relatives accustomed to shop for them. Competently to inform ourselves, we should have a staff of investigators like those supplied to administrative agencies. As we have no such staff, I have questioned some adolescent girls and their mothers and sisters, persons I have chosen at random. I have been told uniformly by my questionees that no one could reasonably believe that any relation existed between plaintiff's magazine and defendants' girdles. the stampeding of less significant factors."' Finally, as this Article has sought to demonstrate, the diversity of tests has made judicial analysis under them less uniform and less predictable.
This Article thus recommends adopting a new national multifactor test, one whose sole purpose should be to aid the judge in estimating the results of an ideal survey of the relevant consumer population. Of course, knowledge of the empirical data will only take us so far in designing such a revised test. Nevertheless, the data do recommend a few general principles. First and most importantly, the basic test should not seek to be exhaustive in its list of possible considerations. Rather, as social science work recommends, the list should consist of a limited number of core factors, ideally no more than three or four. This list may be set forth on an illustrative rather than limitative basis. However, if the history of the multifactor test for trademark infringement teaches us anything, it is that judges, especially circuit court judges, should be wary of adding factors to the list that are only occasionally relevant lest those factors end up being institutionalized. If unlisted factors must be considered in light of the facts of a given case, judges should emphasize the narrow context in which these new factors apply. Second, as the data on the intent factor demonstrated, the precise wording of the factors can be important. The adopted wording should emphasize the empirical rather than formal nature of the inquiry. In trademark law, the question is always of consumer perception in the marketplace rather than judicial perception in the courtroom. Third, the order in which the factors are listed should reflect as much as possible the weight that should be given to them. As such, threshold factors should be listed first. Fourth and finally, if factors are introduced into the test that do not address the empirical question that the test seeks to answer, then that should be stated explicitly. For example, if judges believe that factors such as the defendant's intent or the comparative quality of the goods are worth considering even though they do not go directly to the question of the actual likelihood of consumer confusion, then judges should make that clear.
In light of these principles, a starting point for reform might consist of the following statutory language:
In determining whether a mark is likely to cause confusion, or to cause mistake, or to deceive, the court may consider all relevant factors, including the following: (i) the degree of similarity of the marks as perceived by the relevant consumer population; (ii) the degree of proximity of the goods as perceived by the relevant consumer population, including the degree of proximity of marketing methods and channels of distribution and sale;
(iii) evidence of actual confusion, mistake, or deception, including survey evidence; (iv) the marketplace strength of the mark allegedly infringed; and (v) the purpose of the alleged infringer in adopting and using its mark and if the purpose is to cause confusion, or to cause mistake, or to deceive, then the likelihood that the alleged infringer will accomplish that purpose. Admittedly, this list is far from ideal, and violates some of the principles set forth above. First, it proposes more than three or four factors. Nevertheless, the data suggest that, in practice, judges consider each of the five factors proposed to be important to the outcome of the multifactor test. Additionally, the list conflates factors in a way that may complicate the multifactor analysis. Specifically, it forces consumer sophistication considerations into both the similarity and proximity inquiries. The goal here is to emphasize that the multifactor inquiry is an empirical-rather than formal-inquiry that seeks to determine the likely perception of consumers in the marketplace. Finally, though the data suggest that courts place great weight on a finding of bad faith intent, the list proposes the intent factor as the fifth and final factor to be considered. It does so in an effort to limit the impact of this factor on what should remain a tightly-focused fact-finding inquiry into the likelihood of consumer confusion rather than into the commercial morality of the defendant.
Unlike American copyright law and patent law, American trademark has not gone through a postwar phase of reform and remains largely unrationalized. 252 This is nowhere more evident than in the basic test for trademark infringement, and it is here, at the doctrinal fulcrum of trademark law, that the reform process should begin. Indeed, it is remarkable that other arguably less centrally important, but more up-to-date areas of U.S. trademark law, such as fame. 53 or cybersquatting doctrine, 54 enjoy statutorily-prescribed multifactor tests, but the likelihood of confusion test itself continues to be left to the circuit courts.
This Article has sought to present evidence in support of such reform. In the process, it has sought more generally to develop an array of theoretical approaches to the legal multifactor test. Obviously, much more work remains to be done in this regard, not only in developing our theoretical understanding of the legal multifactor test as applied, but in testing those theories against data on other legal multifactor tests in U.S. and foreign law. Certain questions-are of particular importance going forward. First, to what extent do core factors and a core factor heuristic guide other multifactor tests? Specifically, can we predict the outcomes of other multifactor tests based on the outcomes of only one or two factors within the tests, and is a simple classification-tree approach viable? Are threshold factors prevalent among the wide variety of legal multifactor tests? Second, to what extent do other legal multifactor tests tend to stampede, and is stampeding more common or pronounced when judges make a strong decision that intervenes in the status quo? Relatedly, is it appropriate to model judges' application of the legal multifactor tests by means of regression analysis, particularly if judges are employing coherence-based reasoning? Third, as a historical matter, have other multifactor tests as applied decayed into mechanical formalism? Do they tend to accumulate irrelevant factors? Fourth and finally, does a judge's political ideology affect her use of the multifactor test? 255 These are not merely academic questions. Good answers will enable us to improve our design and use of multifactor tests.
Though this Article has been critical of many aspects of the multifactor test for the likelihood of confusion, it has also sought to be ameliorist in orientation. One thing it has not done is question the viability or utility of the multifactor form of judicial analysis itself. Elsewhere in the world of legal multifactor tests, judges have expressed dissatisfaction with the multifactor heuristic. Judge Easterbrook has made clear his "reluctan[ce] to accept an approach that calls on the district judge to throw a heap of factors on a table and then slice and dice to taste. 256 In criticizing the Barker v. Wingo 257 criteria for evaluating speedy trial claims, Justice Thomas
With a view to shedding light on the degree to which political ideology might affect trademark infringement adjudication, the data set for this study included a number of judge-specific variables, such as the judge's gender, age at the time of the opinion, length of tenure at the time of the opinion, and a rough quantification of the judge's political ideology based on the Poole common space score of the judge's appointing president. Poole common space scores place presidents, senators, and representatives on a scale ranging from -1.000 (most liberal) to 1.000 (most conservative) that seeks to be consistent across time and institutions. Senators' and representatives' scores are based on their voting records. lamented that the Barker "factors now appear to have taken on a life of their own." ' 258 Another justice has asserted that "multifactor balancing tests generally tend to produce negative answers." ' 259 Still, multifactor tests appear to be the least worst alternative, if not the only alternative, to a wideopen "totality of the circumstances ' or "rule of reason" ' 61 type of analysis. This is not simply the case in fact-intensive inquiries such as the likelihood of confusion analysis, but also in broadly political inquiries such as those found in constitutional law. Our goal, then, should be to make the best of this situation. This means designing multifactor heuristics not so much in light of judges' cognitive limitations, but, as the fast and frugal tradition might say, in light of their cognitive ingenuity, their ability, in short, to bring heuristics to heuristics. initial sample of 1,252 opinions from the five-year period. I then reviewed each of these opinions to determine whether it made substantial use of the multifactor test. I defined substantial use liberally as any use beyond the mere citation without analysis of the test. Out of the initial sample of 1,252 opinions, 364 met this criterion.
From this population of 364 opinions, I excluded a small minority of fact patterns that led courts to apply the multifactor test in ways that could skew the results of the study. In most counterfeiting opinions, for example, the likelihood of confusion is very clear and the factors tend to weigh overwhelmingly in favor of the plaintiff. 2 63 The same is true of opinions involving an alleged breach of a franchising, 2 " licensing, 26 or distribution 2 66 agreement. These opinions were thus excluded from the sample. For similar reasons, I also excluded opinions on motions to dismiss or on motions where the non-moving party failed to appear. 2 67 I retained and noted opinions involving claims of reverse confusion, and fact patterns in which the defendant repackaged plaintiff s goods.
This resulted in a sample of 337 opinions. I excluded the six opinions in which the outcome of the multifactor test was reversed, which yielded a final sample of 331 opinions.
The data set was coded entirely by the author. 2 68 I conducted two rounds of coding. In the first round, I began with the Second, Seventh, and Ninth Circuits, and then proceeded through the remainder of the circuits in numerical order. In the second round, I worked through the circuits in numerical order, both checking my initial round of coding and adding various information that I had neglected to record in the first round. The coding was done directly into an author-designed Microsoft Access form. It was then converted into an Excel spreadsheet and finally into Stata format. All statistical analysis was done using Intercooled Stata 9.2.
I recorded the caption of the opinion, its citation, its district, the judge who authored the opinion, the date the opinion was filed, and its posture. The posture of the opinion was coded as: (1) preliminary injunction; (2) summary judgment motion by plaintiff; (3) summary judgment motion by defendant; (4) cross-motions for summary judgment; or (5) bench trial. 269 I noted those cases in which the plaintiff sought a declaratory judgment of no likelihood of confusion and coded the plaintiff as the defendant and vice-versa. I further recorded whether the opinion involved a word/image mark and/or trade dress, and whether the defendant formulated a defense of parody.
For each factor of the multifactor test, I used dummy variables (a series of binary 1/0 variables) to code the factor as being: (1) found to favor a likelihood of confusion or otherwise not to favor no likelihood of confusion; (2) found to favor no likelihood of confusion or otherwise not to favor a likelihood of confusion; (3) found to be neutral; (4) found to be irrelevant; (5) found to be a fact issue or otherwise premature; (6) not addressed by the court; or (7) unclear.
For certain factors, I coded additional information relevant to that factor. Under the similarity of the marks factor, I noted those cases in the marks were found to be "identical," the "same," or "substantially" the same. I followed the same process under the similarity of the goods factor. Under the actual confusion factor, I coded whether the plaintiff presented survey evidence and whether it was credited by the court, and coded the same for defendant. I further coded whether the opinion explicitly stated that it was too early to expect the plaintiff to produce survey evidence of actual confusion and whether the court explicitly drew an adverse inference from plaintiffs lack of anecdotal and/or survey evidence of actual confusion. Under the intent factor, in cases in which the court found the factor to favor confusion, I coded whether the court made its determination based on direct evidence of defendant's intent, inferred intent from the identity of the marks and defendant's knowledge of plaintiff's mark, or was unclear in its determination. Under the quality factor, I coded whether the theory underlying the court's determination was that similar quality facilitated confusion or that the defendant's lower quality harmed plaintiff or both.
Under the strength factor, I noted where in the Abercrombie spectrum the mark was placed (and whether the spectrum was even employed at all), if a "dual test" was explicitly used to determine strength, or if the court 269. I encountered one opinion that addressed a motion for a temporary restraining order that had not been converted into a motion for a preliminary injunction. I coded this opinion as a preliminary injunction opinion. provided no analysis of trademark strength. Further, with respect to inherent distinctiveness, I noted whether the court found the mark to be inherently strong or weak, found inherent strength to be a fact issue, or did not address inherent strength. I coded the same categories with respect to acquired distinctiveness. Finally, I noted whether the court found third party uses of the mark to disfavor strength, and whether the court found the incontestability of the mark to favor strength.
For purposes of assessing any effect of certain judge-specific factors on the uses and outcome of the test, the data set also included the birth date and gender of the judge, the date the judge received his or her commission, the appointing president, and the Poole common space score 270 of the appointing president.
The Unreliability of the Federal Court Cases: Integrated Data Base of the Administrative Office of the U.S. Courts
Researchers have made extensive use of federal court data assembled by the Administrative Office of the United States Courts (AO) and the Federal Judicial Center (FJC). 27 1 The AO data set currently contains information about every case filed in federal district court from 1970 to 2004, as well as every appeal filed in the twelve non-specialized federal appellate courts. Legal scholars have long been suspicious of the accuracy of the AO data. AO bankruptcy data in particular have been judged "error ridden" 2 7 2 and "utterly inadequate for policy purposes. '27 3 Legal scholars working in other fields have defended the AO data as at least serviceable depending on the research question and the subtlety of the statistical techniques used. For example, Theodore Eisenberg and Margo Schlanger compared AO data on tort and inmate cases to district court docket sheets. 2 74 They explain that the "implications of our findings depend in part on whether researchers are interested in assessing win rates or award levels. 2 75 AO data on the for-mer, they conclude, are "overwhelmingly accurate." ' 276 On the latter, AO data "substantially overstates mean awards." ' 277 The circuit-wide bench trial win rate reported in Table 2 diverges substantially from that reported by the AO data. The AO data set, which includes a far more comprehensive sample of trademark litigation than the one used in this study, yields a bench trial win rate in trademark cases for the years 2000 through 2003 of .591 (n=76), while this study's data set yields a win rate of .487 (n=37) for bench trial opinions sampled from the same four-year period. 278 To understand this discrepancy, I sought to locate the AO observations for each of the thirty-seven bench trial opinions I sampled from the four-year period from 2000 to 2003. By matching docket numbers-the publicly available AO data set drops the names of the parties, the judges, and certain other information-I found AO observations for thirty-two of these opinions. Comparing these cases' opinions and dockets to the AO observations yielded quite troubling results, particularly since important empirical work in intellectual property relies so heavily on the AO data. 79 Of the thirty-two cases that ended with a bench trial, only eleven of these cases were classified in the FJC database as being "disposed of' by "6court trial." The following table sets forth, for the thirty-two cases, the AO classification of "[t]he manner in which the case was disposed of' under the AO variable "Disposition": FJC "Disposition" N % "Dismissalswant of prosecution" 1 3.1 "Judgment onmotion before trial" 4 12.5 Id. 278. The difference between the two samples in the proportions of bench trials won by plaintiffs is not statistically significant (z=--l.044, p=.
2 9 6 ), but is nevertheless troubling. The difference may partly be explained by the exclusion of counterfeiting, licensing, and similar fact-patterns from my data set. However, nearly all of the opinions excluded on that basis were preliminary injunction opinions. The difference may also be explained if certain bench trials that resulted in judgments for the plaintiff tended not to produce written opinions reported on Westlaw or Lexis, but were nevertheless sampled by the AO data set. 279. See, e.g., Landes, supra note 76.
which the case had progressed when it was disposed of' variable "Procedural Progress":
[Vol. 94:1581 under the AO FJC "Procedural Progress" N % "Before issue joined-order entered" 2 6.3 "After issued joinedno court action" 3 9.4 "After issue joined -judgment on motion" 3 9.4 "After issue joinedpretrial conference held" 6 18.8 "After issue joinedafter court trial" 11 34.4 "After issue joinedafter jury trial" 3 9.4 "After issue joinedother" 3 9.4 "Unknown" 1 3.1 Total 32 100.2
The AO data on the outcomes of the cases were also unreliable. The following table cross-tabulates, for the thirty-two cases, the AO variable "Nature of Judgment" with this study's coding of the outcome of the trademark infringement claim in the reported opinion.