In 2023, nearly 740,000 trademark applications were filed with the US Patent and Trademark Office (USPTO) (U.S. Patent and Trademark Office, 2023). For each of these trademark applications, an examining attorney at the USPTO must decide whether the criteria for trademark protection are met. Companies developing a new brand that would like to register their trademarks must predict the outcome of this examination when devising their brand strategy. And courts must evaluate the outcome of the USPTO's examination when deciding trademark registrability or infringement disputes.
Examining a single trademark application can be a time-intensive task. Examining hundreds of thousands of trademark applications can easily clutter an administrative agency. As of September 2024, the average processing time for a new trademark application at the USPTO was 14.3 months, and it took the USPTO about 7.8 months to issue a first office action after a trademark application was filed (U.S. Patent and Trademark Office, 2024a). Given the costs of these delays for companies and the economy, there is substantial social value in making this process more efficient.
A key component of the trademark application process is determining whether a mark is "inherently distinctive," meaning that consumers immediately recognize that the mark, because of its semantic or other qualities, is a designation of the source of the product or service. US trademark law traditionally decides distinctiveness following Abercrombie & Fitch Co. v. Hunting World, Inc., a 1976 decision by the US Court of Appeals for the Second Circuit (Abercrombie & Fitch, 1976). The Abercrombie spectrum includes five categories. The first three categories indicate inherent distinctiveness: (1) Fanciful marks, which are completely invented terms like "Xerox," have the highest level of distinctiveness; (2) Arbitrary marks, such as "Apple" for computers, have no direct link to the product; and (3) Suggestive marks, like "Ivory" for soap, imply something about the product. The latter two categories are not inherently distinctive: (4) Descriptive marks, for instance "Pizza Hut" or "iPhone," directly describe the product; (5) Generic terms, like "Escalator" or "Aspirin," are common names for the products themselves. The Abercrombie determination is important in trademark applications because if a mark is deemed not to be inherently distinctive, the mark will register only if the applicant additionally shows that the mark is descriptive (rather than generic) and has acquired distinctiveness, typically through advertising and use of the mark in the marketplace. Showing acquired distinctiveness is often difficult and costly.
This article explores the extent to which natural language processing techniques can help to automate Abercrombie determinations. We present a machine-learning pipeline that is trained on the text of trademark applications to predict whether a trademark is inherently distinctive. As the training dataset, we use the 1 million trademark applications filed with the USPTO between 2012 and 2017. We then validate and evaluate the out-of-sample performance using the 234,000 trademark applications filed in 2018 and the 264,000 trademark applications filed in 2019. The model takes as input the mark and associated application data, and outputs a predicted probability that the mark is inherently distinctive for the associated product or service.
We adapt deep-learning-based machine learning methods for text (Devlin et al., 2018). Our method is context-sensitive, in the sense that it works not just by counting individual words but by learning how words are connected, thereby aiming at a semantic understanding of trademark applications. In our approach, the model input includes the mark, the product description, and the product class. From the mark text, the model can directly learn fancifulness-that is, made-up words such as "Xerox." From the contextual connections between the mark and the product description and Nice class, the model can learn the more subtle category of arbitrariness-that is, mark terms that are in the dictionary yet semantically orthogonal to the product attributes (e.g., "Apple computer").
We find that, in the aggregate, our model's binary predictions agree with the trademark office's decisions about 86% of the time (AUC = 0.77). The relatively high error rate of 14% reflects, in part, that Abercrombie is a subjective determination where humans often disagree, so a significant error rate is unavoidable. However, we also find that the model produces a reliable confidence score for distinctiveness-that is, a predicted probability for various USPTO outcomes. For trademarks where the model has high confidence-that is, the predicted probability is close to 0% or close to 100%-the precision of the model decision is very high (above 90%). We provide a number of robustness checks to show that our approach is fit for purpose and preferable to alternatives.
We argue that the varying confidence in predicting Abercrombie is an important feature of our model for trademark theory and practice. On a practical level, we propose a "robot trademark clerk" which would provide a decisionsupport system for trademark examiners, but also judges, attorneys, and firms who are interested in whether a particular mark is inherently distinctive. Such a system could also help trademark experts understand which features of a trademark application contribute the most toward a trademark's distinctiveness. On a normative level, we discuss whether the varying confidence in predicting Abercrombie should lead us to look for new normative foundations on whether to grant protection to middle-ground trademarks, whose inherent distinctiveness is unclear. These points suggest that applying machine-learning methods to trademark law may not only inform us about the extent to which trademark procedure can be automated. It may also shed light on normative tradeoffs that trademark theory must engage in that were not visible without the assistance of machine learning.
The article proceeds as follows. Section 2 provides the relevant legal background on US trademark law and registration procedure. It also gives an overview of the existing literature that applies machine-learning methods to trademark law or analyzes about such applications in theoretical terms. Section 3 presents our approach toward predicting whether a mark is inherently distinctive on the Abercrombie spectrum, the data we use for our approach, and the concrete implementation of our approach in a natural-language processing setting. Section 4 presents the results of our model predictions and explores which factors in the trademark application drive our model's predictions. Section 5 discusses how our framework could be used to develop a robot trademark clerk, and how it sheds light on normative limitations of the Abercrombie spectrum. Section 6 discusses limitations of our framework and next steps to overcome them. Section 7 concludes the article.
When a trademark registration application is filed with the USPTO, an examining attorney reviews it to determine whether it meets all the criteria for registration. Among other criteria, the examining attorney determines whether the mark is descriptive or deceptive, and whether conflicting marks are already registered. If the trademark application does not fulfill all registration criteria, the USPTO issues an office action to the applicant, listing the reasons for refusal. The applicant can then respond, and the USPTO will then either proceed with the application or eventually issue a final office action refusing registration. If all the registration criteria are met, the examining attorney approves the application for publication in the official journal of the USPTO ("Official Gazette"). Competitors or other parties who believe they could be harmed by the mark's registration can then file an opposition against it, which will be decided by the Trademark Trial and Appeal Board. If there is no opposition or the opposition fails, the USPTO will register the trademark (Gilson & Gilson, 2024, §16.01;U. S. Patent & Trademark Office, 2024b, §704.1). Once a trademark is registered in the USPTO's Principal Register, the registration is prima facie evidence of the mark's validity, including its distinctiveness (Gilson & Gilson, 2024, §2.05[2]).
As part of the substantive examination, the examining officer must determine whether a mark is able to "identify and distinguish […] goods […] from those manufactured or sold by others and to indicate the source of the goods" (15 U.S.C. §1127). Conceptually, trademark law follows a classification developed by Judge Friendly in the Abercrombie decision (Abercrombie & Fitch, 1976). The Abercrombie spectrum classifies trademarks according to their degree of inherent distinctiveness: (1) fanciful marks, which have the highest degree of inherent distinctiveness, are coined terms (e.g., Xerox for copiers);
(2) arbitrary marks, which have no semantic connection to the relevant product (e.g., Apple for computers); (3) suggestive marks, which suggest or are metaphorically related to product characteristics (e.g., Ivory for soap); (4) descriptive marks, which describe product characteristics (e.g., Coca-Cola a cola drink or iPhone for a mobile phone); and (5) generic marks, which refer to the type of product (e.g., Escalator or Aspirin, Beebe, 2006Beebe, , p. 1634;;Beebe & Fromer, 2018, p. 957; see Figure 1). Arbitrary and fanciful marks are understood to be the most highly inherently distinctive with suggestive marks understood to be less inherently distinctive. Courts rely on a variety of evidence to determine whether a mark is inherently distinctive, including consumer surveys, dictionary definitions, expert testimony, and evidence of use over time (Gilson & Gilson, 2024, §2.05[1]).
By virtue of their inherent distinctiveness, fanciful, arbitrary, or suggestive marks can be registered on the Principal Register, provided they meet all other requirements for registration. For descriptive marks, trademark law treats them as not inherently distinctive, as consumers may interpret them as mere product descriptions. Therefore, the trademark applicant must produce evidence that the mark is distinctive of source to consumers (so-called "secondary meaning" or "acquired distinctiveness"). A descriptive mark may acquire distinctiveness over time through use, in particular if consumers learn to recognize the mark as an indication of source (Beebe & Fromer, 2018, pp. 957-958). As a general matter, generic marks cannot be registered as trademarks at the USPTO (on the distinction between descriptive and generic marks after the Supreme Court's Booking. com decision, see Fromer, 2022).
If a mark does not achieve distinctiveness-either inherently (for fanciful, arbitrary, or suggestive marks) or through "secondary meaning" (for descriptive marks)-it cannot be registered or otherwise protected as a trademark. While inherently distinctive trademarks that register are placed on the Principal Register (generating prima facie evidence of the mark's validity, as mentioned before),
Inherently distinctive mark F I G U R E 1 Abercrombie spectrum: from generic to fanciful marks.
descriptive trademarks with no secondary meaning can only be registered on the Supplemental Register (providing only limited notice and standing functions; 15 U.S.C. §1091; U.S. Patent and Trademark Office, 2024b, §1209.1; Gilson & Gilson, 2024, §3.05[3], §4.07[5]). Once a mark published only in the Supplemental Register has acquired secondary meaning, it can be published in the Principal Register.
Determining the distinctiveness of a mark is not only important for determining whether a mark is protectable. It is also important because the scope of trademark protection is proportional to the inherent distinctiveness of a mark, which can be an important factor in determining the likelihood of confusion in trademark infringement litigation, along with factors such as the extent of marketplace use and advertising of the mark (Beebe, 2006(Beebe, , pp. 1634(Beebe, , 1637;;Gilson & Gilson, 2024, §2.01[6]). For the purposes of trademark registration, however, the five-tier Abercrombie spectrum boils down to a simple dichotomy: those marks which are distinctive of source and those marks which are not. Among the former group of marks, some marks are inherently distinctive (fanciful, arbitrary, and suggestive), while descriptive marks may acquire distinctiveness through secondary meaning (descriptive; Beebe, 2004, p. 671;Gilson & Gilson, 2024, §2.01[6]).
This article builds on various strands of literature. Tushnet (2017, pp. 870-871) points out that the procedural and substantive aspects of the trademark registration process are typically neglected in trademark law scholarship. With our study, we hope to narrow this research gap. Various scholars have explored the relationship between trademark registrations and language. Beebe (2006) analyzes federal district court opinions from 2000 to 2004 to determine how courts use multifactor tests when determining likelihood of consumer confusion. Beebe and Fromer (2018) measure the extent to which the most common English words and syllables are registered as trademarks, thereby potentially depleting the supply of possible marks. On a conceptual level, Hemel and Ouellette (2021) argue that such depletion can create two distinct types of costs for the trademark system: proximity costs (where different firms use similar marks) and distance costs (where firms eschew these proximity costs by using marks that are difficult for consumers to remember).
Scholars, trademark offices and the consulting industry have started to explore the implications of natural-language processing tools for the trademark system (for related attempts in the patent system, see Ashtor, 2022;Hain et al., 2022). Legal scholars have often debated the impact of artificial intelligence on the trademark system on a conceptual level (Gangjee, 2021;Katyal & Kesari, 2020;Lim, 2022aLim, , 2022b;;Moerland & Freitas, 2021). Trademark offices have started to explore systems classifying trademark applications into the respective Nice classes of goods and services (such as apparel, computer software, and beverages), to identify similar marks and to determine whether a sign is descriptive (Gangjee, 2021, pp. 178-184;Moerland & Freitas, 2021, p. 226).
The consulting industry has started to offer related trademark-search and decision support systems (see Katyal & Kesari, 2020). Detailed information about the inner workings of such systems is often scarce, however.
Data scientists have applied machine-learning tools to develop a recommendation system identifying semantically similar trademark litigation judgments for a trademark of interest (Trappey et al., 2020). Shackell and De Vine (2022) use word-embedding metrics to analyze whether trademarks have become generic. Showkatramani et al. (2019) use natural language processing methods to automate manual efforts in classifying trademarks into the 45 classes of the Nice classification system of goods and services.
Most relevant to our article are projects that attempt to determine trademark distinctiveness with existing data. Even before the big data revolution, Ouellette (2014) proposed to determine trademark strength through conducting Google searches. She argued and demonstrated through case studies that the stronger a trademark is, the more top search results for the trademark will appear on Google. More recently, Goodhue and Wei (2023) explored whether a large-language model such as GPT 3.5 could be used to classify trademarks along the Abercrombie spectrum. In a case study of 24 trademarks with some prompt engineering and limited fine-tuning, they report that GPT 3.5 did a mediocre job in distinguishing descriptive from suggestive, fanciful and arbitrary marks. To determine whether their 24 benchmark trademarks are at least suggestive or merely descriptive, they use the fact of whether the trademarks are registered on the Principal Register or the Supplemental Register as a proxy. In their best specification, they find that GPT 3.5 determines 6 out of 15 on the Principal Register as at least suggestive and 7 of 9 marks on the Supplemental Register as descriptive. In a follow-on study, Goodhue and Xing (2023) use GPT to identify special issues in trademark applications, such as personal names or geographically descriptive marks. They then use a logistic regression model applied to TF-IDF features to classify trademarks on the Abercrombie spectrum. Guha et al. (2022) take a broader approach and build a legal benchmark suite for language foundation models-such as GPT-that would enable comparisons on how various foundation models perform on these tasks. One of the over 170 available tasks is to locate a trademark on the Abercrombie spectrum (LegalBench, 2024). Their initial results provide an F1 score of 0.42 for the bestperforming foundation model (GPT-3 davinci, Guha et al., 2022, p. 8; see also Guha et al., 2023).
Overall, while data scientists and legal scholars have begun to apply machine-learning methods to select issues in trademark law, to the best of our knowledge no scientific work has developed natural-language processing tools to determine trademark distinctiveness on the Abercrombie spectrum on a large scale. 1 More specifically, compared with the existing empirical literature on the Abercrombie spectrum, this article is the first to train a machine-learning model on actual trademark decisions at a large scale, using a dataset with over 1.5 million trademark registrations, while earlier contributions are typically based on querying a standard language model or training a language model with about 3000 trademark registrations only. In addition, by combining trademark registrations with information on USPTO office actions, we present a clean dataset that enables us to accurately determine whether the USPTO has classified a trademark as distinctive or not (for details, see Section 3.2). Finally, our research design enables us to provide information on how confident our model's predictions are, and we can show which components of trademark applications are driving factors in the model predictions (for details, see Sections 4.2 and 4.3). All of these features, which we add to the existing body of literature, have important implications for the trademark discourse (see Section 5).
In this article, we want to use natural-language processing techniques to predict whether a USPTO examining attorney would determine that a particular mark is inherently distinctive or not. We want to use existing trademark registrations (including, e.g., the name of the trademark, its Nice class, and its product description) to train a classifier to predict distinctiveness. As we are interested in exploring the semantic features of registered trademarks, we treat only inherently distinctive marks as "distinctive" and do not include in this category descriptive marks that may acquire or have already acquired secondary meaning among consumers (on this distinction, see supra Section 2.1). While the Abercrombie spectrum can be used to determine the strength of a trademark on a five-tier scale, for the purposes of trademark registration, our task gets reduced to a simple dichotomy (see supra Section 2.1): distinguishing inherently distinctive marks (either fanciful, arbitrary, or suggestive) from non-inherently distinctive marks (either descriptive or generic, regardless of acquired secondary meaning).
To train our classifier whether a mark is inherently distinctive, we ideally would need a comprehensive dataset of all active trademarks with information about their validity, location on the Abercrombie spectrum, product description, 1
The trademark offices of Australia, the European Union and Singapore have reportedly worked on systems to assist distinctiveness determination, see Gangjee (2021, p. 184). However, information on these systems is scant. Also, the Abercrombie test, which is the focus of this study, does not apply in these jurisdictions.
and Nice class. Two types of potential datasets come to mind. First, one could explore litigated trademark cases. While trademark scholars have coded and analyzed trademark court decisions empirically (see Beebe, 2006), such datasets are unfortunately not well-suited for our purposes. The most important limitations are that courts very often do not reveal in their decision what the trademark's registration number is, in which Nice classes the trademark is registered, and where exactly on the Abercrombie spectrum the trademark is located (Beebe, 2006(Beebe, , pp. 1635(Beebe, -1636)). Even if one matched litigated trademarks to trademark registration data, where such information is available, this would not be sufficient. The number of court decisions dealing with trademark distinctiveness is limited and would not suffice to train machine-learning models which typically require very large training datasets. 2A second potential dataset is trademark registration data. The advantage of trademark registration data is that all of the information about trademarks mentioned above is readily available from the USPTO (Graham et al., 2013). When a USPTO examining attorney decides to register a trademark, the registration record does not directly reveal whether the trademark is inherently distinctive (as also descriptive trademarks with acquired secondary meaning can be registered). Rather, one needs to use a proxy to determine the inherent distinctiveness of a trademark. In their case study, Goodhue and Wei (2023) employ as a proxy whether a trademark has been registered on the Principal Register or the Supplemental Register. Mapping inherently distinctive trademarks to the Principal Register and descriptive trademarks to the Supplemental Register is imprecise, however. As mentioned in Section 2.1, a descriptive trademark that has acquired secondary meaning can and will be registered on the Principal Register. Therefore, in our view, using the Principal versus Supplemental Register as a proxy to determine the inherent distinctiveness of a mark is problematic.
In this article, we develop a distinctiveness indicator using a combination of publications of trademarks in the Primary and the Supplemental Register, USPTO office actions, and data available through USPTO case files. We start by considering the publicationfoot_2 of a trademark application in the USPTO Official Gazette as a proxy for inherent distinctiveness. As explained in Section 2.1, when a trademark application is filed with the USPTO, an examining attorney conducts a complete examination, including whether a trademark is descriptive. If the examining attorney decides that the trademark is descriptive and lacks secondary meaning, the application will be refused through an office action. If, however, the examining attorney decides that all the registration criteria are met, the trademark application is published in the USPTO Official Gazette.
Thereafter, a competitor can oppose the mark's registration if, for example, the competitor holds rights in a confusingly similar mark. If there is no opposition or the opposition fails, the USPTO will register the trademark. The trademark's publication provides information about the examining attorney's assessment of the trademark registrability, including the trademark's distinctiveness: if the examining attorney finds a mark merely descriptive and lacking in acquired distinctiveness, the application will not be published in the Official Gazette.
Unfortunately, whether a trademark gets published in the Official Gazette is an imperfect proxy for whether the USPTO examining officer treats a mark as inherently distinctive (publication: yes; no publication: no). First, as mentioned in Section 2.1, descriptive marks may also be published in the Official Gazette if the applicant has provided evidence of secondary meaning. Second, marks that are inherently distinctive may not get published in the Official Gazette for other reasons (e.g., because they are confusingly similar to an already-registered mark). Therefore, if we took publication in the Official Gazette as an indicator of inherent distinctiveness, our analysis might suffer from substantial false positive and false negative errors.
To remedy such errors, we use the Principal and Supplemental Register, but combine them with information about USPTO office actions and case files. More specifically, we treat trademarks that get published for registration on the Principal Register as inherently distinctive, but exclude trademark applications for which the applicant did not claim inherent distinctiveness, but provided evidence of secondary meaning (so-called "2(f) applications," see 15 U.S.C. §1052 (f)). We also treat trademarks as inherently distinctive if they did not get published but, according to their corresponding office actions, did not get refused due to a lack of distinctiveness (for example, they are inherently distinctive, but get rejected on other grounds, such as 15 U.S.C. §1502(d), for being confusingly similar with an already-registered mark). By contrast, we treat trademarks published for registration on the Principal Register that resulted from a "2(f) application" as not inherently distinctive marks. Trademark applications that do not get published and, according to their corresponding office actions, get refused due to a lack of distinctiveness (even if they also get refused on other grounds as well) are treated as not inherently distinctive marks as well. Finally, trademarks published in the Supplemental Register are also not inherently distinctive.
In our view, combining information about which trademarks get published in the Principal and the Supplemental Register with information from USPTO office actions and case files is the most accurate way of creating a training data set for the distinctiveness threshold on the Abercrombie spectrum. Table 1 provides an overview of our categorization. We use this categorization, coupled with all of the other features that can be retrieved from the application about the trademark, to train a classifier. If we provide the system with a trademark with certain features, the system should then be able to predict whether the trademark is inherently distinctive.
As outlined in the preceding section, we use the USPTO Trademark Case Files Dataset (Graham et al., 2013) as our central data source. This dataset includes records on 12.4 million trademark applications and registrations associated with the USPTO. The earliest entries in the dataset go back to 1870. The dataset originates from the primary USPTO database responsible for managing trademarks. It encompasses a wide range of information, including details about the characteristics of marks, prosecution events, ownership, classification, thirdparty oppositions, and the history of renewals. As a result, the dataset offers an extensive labeled data source for analyzing trademark publication, registration, and acquired distinctiveness.
We also use a dataset developed by Beebe and Fromer (used in Beebe & Fromer, 2018) of all trademark office actions that the USPTO issued from 2003, when it began posting office actions online, through 2019. For applications filed from 2012 through 2019, the dataset includes the full text of 2.2 million office actions issued by USPTO trademark examining attorneys concerning 1.5 million trademark applications. We autocoded each office action according to certain keywords and phrases that trademark examining attorneys consistently use when refusing registration based on descriptiveness or genericism.
To streamline the analysis and prioritize recent data, we use data from 2012 to 2019 in our study. To simplify the analysis and focus on text features of trademarks, we drop marks with images (about 20% of the sample). For similar reasons, we also drop marks with five or more words (about 6% of the sample, Figure B1 shows a histogram of word lengths). After preprocessing, we divide the data into training, validation, and test sets. For training, we use trademarks that have been filed between 2012 and 2017. We use trademarks filed in 2018 as our validation dataset, and trademarks filed in 2019 as our test dataset. Overall, we have about 1 million data points for training, about 234,000 data points for validation and about 264,000 data points for testing purposes.
Our primary outcome to be predicted is an indicator variable for "ïnherent distinctiveness," indicating whether the trademark application sought to register an inherently distinctive mark, as described in Table 1. We assign 1 to trademark entries which are inherently distinctive, and 0 otherwise. Approximately 84% of all trademark applications which we use as our test dataset are classified as inherently distinctive trademarks. The incidence rates for all outcome labels are provided in Table B1.
For each trademark, we use the following variables in our predictions (where features 1-4 come directly from the USPTO Trademark Case Files Dataset):
1. Statement text: This provides a description of the proposed mark phrase and the product to which it is supposed to be affixed. As recently developed language models can use all the information from input text, including punctuations and stopwords, we do not preprocess the statement text. Since there are different statement type codes for each trademark (each trademark is represented by a serial number), we aggregate the data on a trademark level by concatenating different statement texts available for a given trademark. We do this for all different statement types, except for pseudo marks. As pseudo marks can be thought of as "wordplay" on trademarks, we keep them as a separate feature. By including pseudo marks, we provide our model with some information on how humans understand a particular word from a linguistic perspective, thereby enriching our model's understanding of context. A summary of the features used is presented in Table 2.
As described in Section 3.1, we want to be able to provide a machine-learning model with a trademark with certain features so that the model predicts whether the trademark is inherently distinctive (according to our indicator presented in Table 1). In choosing our machine classification method, we are guided by the importance trademark practice puts on context. In particular, we are looking for methods that enable the classifier to get a semantic understanding of trademark applications.
We are interested in using a context-sensitive neural net that learns semantic features of trademark applications, along with text representations of covariates. We follow a text classification approach using a BERT-based model. BERT (Devlin et al., 2018) is a large pre-trained transformer model that learns to understand language by guessing left-out words in short documents. Through this process, BERT obtains a semantic understanding of language, where words are understood in the context of other surrounding words. That is, BERT is preferred to other approaches because it not only counts words and phrases, but also learns first-order and second-order connections between words, even at a distance. BERT models can then be applied to many downstream tasks such as classification, question-answering, or summarization. In our case, BERT can learn subtle connections between the different words in the text input, such as how the words in a mark are related to the words in the product description or class description.
BERT takes only text as inputs. We concatenate all the features of a trademark application to form a paragraph, which acts as an input to the BERT model. Figure 2 shows the input format using an example ("ZIPSCENE"). 4 We also compare our main results with Xgboost (Chen & Guestrin, 2016), a more classical machine-learning model. Our fine-tuned DistilBERT outperforms XgBoost model for all the standard evaluation metrics used for classification tasks. Additional details for Xgboost can be found in Appendix A.6.
We match these text inputs with the associated label to form a dataset. We then take a pre-trained BERT model and feed in the new pairs to fine-tune it on the new classification task: whether the associated trademark is distinctive (according to our indicator presented in Table 1). 5 Choosing a threshold for a classification objective is subjective, and depends on the incidence rate of the dependent variable, as well as on the extent to which output probabilities truly estimate the data distribution. To support a decisionmaking framework using our fine-tuned BERT model, it is important that the output probabilities are reliable, and correctly estimate the true correctness likelihood. Hence, we calibrate the output probabilities to reflect an estimate of the data distribution (Guo et al., 2017). The resulting adjusted model will give predicted probabilities that reflect the actual outcome rates in the data.
In our study, the specific model variant used is DistilBERT, an efficient and fast BERT implementation (Sanh et al., 2020). As a comparison to DistilBERT, we report results for two other models. As a more capable model, we use RoBERTa (Liu et al., 2019), a larger BERT-style model that has been pretrained on more data than BERT. As a classical-machine-learning baseline (that is, not using deep learning), we implement XGBoost, which is an ensemble of decision trees that vote on the outcome based on the text inputs. This baseline is explained in detail in Appendix A.6.
This section reports the results from predicting whether a trademark is inherently distinctive.
Table 3 presents the results obtained from our models on the test dataset. We report the following classical machine-learning evaluation metrics: Accuracy means the proportion of correct predictions. Balanced accuracy means the average recall (accuracy) for the two classes. Precision is the accuracy conditional on prediction; F1 is the geometric mean of precision and recall. Finally, AUC (Area Under the ROC Curve) gives the accuracy with which the model can rank two randomly selected observations by their predicted probabilities across classes. The first column provides the weakest baseline of guessing the modal class (distinctive). Accuracy is the proportion of the modal class, and balanced accuracy and AUC are both 0.5. This shows the minimum performance one could expect. The other models should be compared based on their improvement from this baseline. Note: Main classifier evaluation metrics for the distinctive outcome using a decision threshold of 0.5. Guess Distinctive has been calculated by guessing the modal class for every data point (every trademark is distinctive). We use weighted metrics to take the uneven distribution of mark distinctiveness in the dataset into account (84% of trademarks are distinctive, and 16% not distinctive). Balanced accuracy is calculated by taking the unweighted average recall across the two classes. The numbers in the bracket below represent a 95% confidence interval on the test set by bootstrapping 30 times with replacement. As the dataset is imbalanced, we also provide metrics for the minority class (mark not distinctive).
Next, we find that DistilBERT and RoBERTa models are 86% accurate, with similar numbers for recall, precision, and F1. AUC is 0.77 for RoBERTa and 0.76 for DistilBERT. The classical machine-learning baseline XgBoost is significantly worse in terms of AUC. Additional metrics and evaluation are reported in Appendix A.1.
Next, we want to check whether the models are well calibrated. In Figure 3, we divide the test dataset into 10 bins by predicted outcome probability (probability that mark is distinctive) and compare them with the true probability in each bin. Ideally, the calibrated probability should be equal to the true probability for each bin. As can be seen in the figure, as the calibrated probabilities and true probabilities are very close for all the bins, we can say that the model is well calibrated, and can predict probabilities correctly. A calibrated probability of 0.9 implies that we can expect the mark to be distinctive 9 out of 10 times. While this is useful for decision-making, the graph also shows that the model is not that confident most of the time, with most of the probability mass being centered around the outcome base rate (84%). We provide the calibration plot for F I G U R E 3 Calibration Plot for DistilBERT. The dashed black line shows the reference/ideal calibration. The x-axis gives bins of the predicted probabilities, and the y-axis (blue line) shows the true probability in each bin. The histogram depicts the distribution of predicted probabilities using 10 bins in percentage.
RoBERTa in Figure A3. As DistilBERT provides better probability calibration for our data, we continue our main analysis with this model.
As our model delivers 86% accuracy, with similar numbers for recall and precision, one may wonder how useful such a model may be for trademark law and practice. At first sight, 86% accuracy might seem nowhere close to approximating decisions by humans, and pretty much the same accuracy to guessing the modal category "distinctive" (84%). However, the performance metrics are comparable to other studies making predictions to support judicial decisions (such as Ash et al., 2024;Kleinberg et al., 2018). Human decision-makers make errors, they render inconsistent decisions, and legal decisions get appealed and overturned on a frequent basis. Therefore, the accuracy of human decisions has important limitations as well.
More importantly, anecdotal evidence from trademark practice raises the question what the relevant benchmark of human decision-making is in such cases. As the Court of Appeals for the Second Circuit has acknowledged, placing a trademark on the Abercrombie spectrum "is far from an exact science, and […] the differences between the classes, which is not always readily apparent, makes placing a mark in its proper context and attaching to it one of the [Abercrombie] labels a tricky business at best" (Banff, Ltd. v. Federated Dep't Stores, Inc., 1988, p. 489). The Court of Appeals for the Seventh Circuit noted that distinguishing between suggestive and descriptive marks is difficult and is "often made on an intuitive basis rather than as the result of a logical analysis susceptible of articulation" (Union Carbide Corp. v. Every-Ready Inc., 1976, p. 379). And, in the words of the Court of Appeals for the Fifth Circuit: "The labels [of the Abercrombie spectrum] are more advisory than definitional, more like guidelines than pigeonholes. Not surprisingly, they are somewhat difficult to articulate and to apply" (Zatarain's, Inc. v. Oak Grove Smokehouse, Inc., 1983, p. 790). A leading treatise notes that delineating between descriptive and suggestive marks is "subjective" and "intuitive" (McCarthy, 2024, §11.70). Perhaps as a result of these difficulties, Beebe (2006Beebe ( , p. 1635) finds that many district courts make little use of the Abercrombie spectrum in the context of likelihood of confusion inquiries (see also Fromer, 2011Fromer, , pp. 1912Fromer, -1913)).
As a result, it is not clear how well humans will perform when deciding whether a trademark is inherently distinctive. It is unrealistic to expect 100% accuracy for humans, as they may make mistakes, may lack clear guidance from the vague Abercrombie spectrum test, or may just disagree with other humans in particular cases. An advantage of our approach is that our model provides predicted probabilities, which give a score of the difficulty of a case. This confidence score indicates how difficult it is for the machine to determine whether a trademark is distinctive. Note that this might not perfectly reflect the difficulty for humans. Indeed, some cases that our model predicts easily might still be difficult for humans.
To explore the varying confidence in predicting a trademark's location on the Abercrombie spectrum, we deviate from a uniform decision threshold of 0.5 for all trademarks and analyze model precision of the calibrated probabilities at various decision thresholds. We simulate decisions based on nine decision thresholds, corresponding to predicted outcome probabilities. For a particular decision threshold (10%, 20%, …, 90%), if the calibrated probability is above the decision threshold, we classify it as distinctive, otherwise as non-distinctive. Evaluating separate precision values for distinct and non-distinct marks at different decision thresholds provides us with a metric for how often the model is predicting correctly, conditional on the decision specified by that threshold.
Table 4 illustrates the precision for distinctive and non-distinctive marks at different decision thresholds for DistilBERT (see Figure A2 for a graphical representation). In this table, rows correspond to decision thresholds. That is, for a given threshold X (between 0.0 and 1.0), we assign trademarks as non-distinctive if the predicted probability Y-hat
As Table 4 shows, the model becomes more precise for each class as we become more stringent in deciding for that class. For example, for a decision threshold of 0.8, a decision for the distinctive class (class 1) is accurate 90.39% of the time. Analogously, for a decision threshold of 0.2, the mark has an 88.29% chance of not being inherently distinctive. Overall, the model is particularly good at the tails of the distribution: for trademarks where the model predicts that the mark can be published with a very high (or very low) probability, the precision of this prediction is actually very high (>90%). In Appendix A1, we present a similar analysis for a fine-tuned RoBERTa model (Table A1, Figures A3 andA5). 6 It is an interesting feature of our model that we can use these probabilities as assistance for deciding when to follow the classifier. If, for example, one feeds the model with a particular trademark and the model returns that the trademark has a 90% chance of being inherently distinctive, one also learns that the model's "confidence" in this prediction is very high (about 93%). If, however, the model returns for another trademark a 50% chance of being inherently distinctive, one also knows that the model is much less confident in this prediction (about 86%).
We would like to understand which words in a trademark application are highly responsible for affecting our model's prediction toward distinctiveness on an instance as well as data level. To this extent, we explore separate instances and different Nice classes. To get an intuition, we provide examples highlighting words according to how much they contribute to the model's prediction. We use a metric from the explainable AI or model explanation literature, called SHAP score (also known as Shapley value), which summarizes the contribution of an input feature to a model prediction (Lundberg & Lee, 2017;Molnar, 2024).
For example, Figure 4 shows the text input for the mark "ZIPSCENE." 7 Words that are highlighted in light gray have a higher SHAP score, indicating 6
We run robustness checks in Appendix A.3 on whether the intuition holds that the more difficult it becomes to locate a trademark on the Abercrombie spectrum, the lower the model confidence is. As a rough proxy for how difficult it is to locate a trademark on the Abercrombie spectrum, we use the length of the trademark application procedure. Appendix A.3 shows that there is a negative relationship between length of procedure and model confidence.
We use all the features mentioned in Table 2. Highlighting gives SHAP scores aggregated at word level. We remove the special tokens "[SEP]" and "[CLS]" so that they are not attributed to model probability through SHAP.
that they contribute to a finding distinctiveness. The contribution of "ZIP" outweighs the negative contribution of "SCENE," thus resulting in the model's decision for distinctiveness (see also Figure A13). In Appendix 10.1, we expand this analysis beyond the single example of ZIPSCENE. We analyze SHAP scores for samples having different levels of predicted model probabilities of distinctiveness. We observe the influence of mark shifts from negative to positive as the probability of distinctiveness increases.
To further investigate the importance of mark names to effect model predictions, we investigate the influential words on a Nice class level using their SHAP scores. For this analysis, we focus on Nice classes 12 (vehicles), 16 (paper/cardboard), 32 (beverages), 43 (food/drinks), and 25 (clothing), as they provide welldefined categories of goods or services. Table 5 presents the top five words, their contributions, and associated marks for each of the Nice categories. From the table, it becomes clear that mark names are essential for determining the model's predictions, and are the most important feature. For the selected Nice classes, all but one of the top words contributing toward distinctiveness are directly associated with mark names.
Although the contribution of mark names toward deciding distinctiveness is essential from the model's perspective, it is also important to note that the model can also focus on other parts of the input text. To test the performance of our model in this regard, we explore whether the model probabilities change when the same mark has different descriptions (statement text). Ideally, if no relationship between mark name and description existed, an increase in model probabilities would be indicative of arbitrary, distinctive mark. If some relationship between mark name and description existed, model probabilities could indicate either a suggestive (distinctive) or a descriptive (non-distinctive) mark. As an example, let us consider the mark "WATERMELON" and pair it with two different descriptions, one describing computer hardware and the other describing retail and wholesale fruit distributorship. Figure 5 presents our results. As expected, our model probabilities decrease from 0.67 in the context of computer services to 0.49 in the context of fruit distributorship. From the SHAP scores, we can see that words such as "stores" and "wholesale" have a high impact on the mark not being distinctive.
F I G U R E 4 Model explanation results for ZIPSCENE. Word contribution toward distinctiveness for ZIPSCENE. P(distinctive) = 0.91. Light gray implies a positive contribution toward distinctive; and dark gray implies a negative contribution toward non-distinctive. SHAP scores for BERT models are on a sub-word level. We aggregate the SHAP scores for a word to produce the figure above. "ZIP" and "SCENE" have high contributions, although in opposite directions.
T A B L E 5 Top five words from NICE classes 12,16,32,43,and , 16, 32, 43, and 25. The marks in the parentheses correspond to the mark in which these words were present. In some cases, the word might be present in more than one mark. Thus, for each Nice class in the table, we determine the average SHAP score of a word in the dataset and select the top five words. SHAP contribution displays the word's contribution toward distinctiveness. Class 12 represents Vehicles; apparatus for locomotion by land, air or water. Class 16 represents paper and cardboard; printed matter; bookbinding material; photographs; stationery and office requisites, except furniture; adhesives for stationery or household purposes; drawing materials and materials for artists; paintbrushes; instructional and teaching materials; plastic sheets, films and bags for wrapping and packaging; printers' type, printing blocks. Class 25 represents clothing, footwear, and headwear. Class 32 represents beers; non-alcoholic beverages; mineral and aerated waters; fruit beverages and fruit juices; syrups and other preparations for making non-alcoholic beverages. Class 43 represents services for providing food and drink; temporary accommodation.
To extend our understanding of how SHAP can be used to understand our model's performance, we test how the model performs when giving made up words, as this can be an indicator of fanciful marks. We use the marks "EMERLWONAT," which is an anagram of "WATERMELON," and a madeup word "XADERMAC" paired with the fruit distributorship. In both cases, the model is able to predict the distinctiveness with high probabilities. Figures 6 andA13 show the SHAP scores for these marks. In both cases, the mark names have a high impact on predicting distinctiveness. 8 Although the mark name is the most important feature, the model performance would be suboptimal if we just used the mark name as an input. Training the model only with marks produces an AUC of 0.69, seven points below the model performance. Similarly, training the model without marks produces an AUC of 0.71, which is again lower than the model performance in Table 3. As a consequence, we think that both mark name and other features are important inputs when analyzing which features of a trademark application contribute to our model's prediction.
As Section 4 showed, our machine-learning model does not perform particularly well when applying a uniform decision threshold to all trademark applications. But in subsets of the data where the model is confident in its predictions, it performs astonishingly well-with over 90% accuracy. In this section, we explore the implications our model may have for trademark practice and theory. First, we outline the contours of a decision-support system based on our model that may pave a road toward a robot trademark clerk (not judge). Second, we discuss the implications the varying confidence in locating trademarks on the Abercrombie spectrum has on the Abercrombie spectrum itself.
The results from Table 4 show that our model performs very well for trademarks with a high or low predicted probability. Hence, it could potentially serve as a backbone for a trademark decision-support system. This may not only be of interest to USPTO examining attorneys dealing with trademark applications 8 Since "EMERLWONAT" is an anagram of "WATERMELON," it would be interesting to compare the attention distribution between "EMERLWONAT" and "WATERMELON." Figure A15 shows the attention score distribution for "EMERLWONAT." Compared to before, words attend to the mark name more strongly. Similarly, "fruit" attends to the mark name more strongly, but not as strongly as "computers." The average attention score of the mark name increased from 0.1082 ("WATERMELON") to 0.1207 ("EMERLWONAT"). and courts dealing with trademark infringement lawsuits. It could also be interesting for trademark attorneys and brand developers interested in assessing chances of registering a particular trademark. The bar plot displays high negative attribution toward "featuring," "wholesale," and "stores," resulting in lower probability of distinctiveness than "WATERMELON" in the context of computer hardware.
Figure 7 shows a possible framework for such a decision-support system, illustrating the example of ZIPSCENE, a trademark that is part of our dataset (and which is registered as inherently distinctive). In this framework, users need to input the name of the trademark, its respective Nice classes, possible wordplays (for detecting pseudo marks), and a description of the trademark. The backend calculations involve text pre-processing, English translation of the trademark, and fetching the Nice class description. The model churns out the calibrated probability, which the user can refer to Table 4.
Of course, we do not claim that our model is sufficient to replace a human decision-maker. As Section 6 demonstrates, we are aware of the many limitations a model such as ours would have to overcome before being able to replace a trademark specialist. But in our view, determining whether a trademark is distinctive is an excellent candidate for an automated decision-support system. Locating a trademark on the Abercrombie spectrum is a repetitive task that human experts have to perform nearly 740,000 times per year at the USPTO. Average trademark pendency has increased over the last years at the USPTO, and the agency has announced plans to hire an additional 140 trademark examining attorneys in 2023 and 2024 to speed up registration procedures (Fromer & McKenna, 2024, p. 30). As trademark applications are a high-volume business, trademark offices around the world have started to explore machine-learning techniques to speed up the trademark registration process (see Gangjee, 2021, p. 184).
Once the decision-support system we envision has been set up, it would be easy to use and would give clear recommendations about which trademark application the human expert should focus his or her attention on. The system could also become an intuitive assistant to trademark examining attorneys, as the system would arguably develop similar background knowledge as trademark experts who have seen many trademark applications over a period of several years. We envision that the system would not only save trademark examining attorneys time; it could also increase the overall quality of their decisions. They would receive help from a machine to distinguish clear-cut cases, which are easy to decide and only require a quick look from the examining attorney, from cases that require more attention from the examining attorney. To assess the effectiveness of such a system, one could ideally test the system in a controlled policy experiment with trademark offices (for a policy experiment of the USPTO in patent law, see Pairolero et al., 2022).
We want to add that such a trademark decision-support system would not only inform the trademark specialist about the likelihood that a trademark is inherently distinctive. As we have shown in Section 4.3, AI explanation tools such as SHAP can be used to analyze the features of a trademark application that drive our model's prediction. If a trademark decision-support system would become equipped with an explanation algorithm such as SHAP, this could point out to trademark examining attorneys and other decision-makers which factors lend the most weight to distinctiveness in the model, and would allow them to introspect on whether that coincides with their own experience and intuition. Furthermore, model explanations might help trademark scholars further develop our understanding of the drivers of distinctiveness.
F I G U R E 7 A framework for a decision-support system for the trademark "ZIPSCENE."
Our study not only paves the road toward decision-support systems in trademark law. Perhaps more importantly, it also enables us to shed light on the limitations of the Abercrombie spectrum and the normative challenges resulting from these limitations.
A simplistic view on how the Abercrombie spectrum gets applied is that the human decision-maker learns about the definitions of the five tiers of the spectrum, gets experience by learning earlier decisions along the spectrum, and then applies the spectrum with the same confidence to all the different trademarks. However, such a view is naïve as the decision-maker will find it easier to determine the inherent distinctiveness of a trademark in clear-cut cases than in murky cases. In reality, the human decision-maker's confidence will be higher for both clearly distinctive and non-distinctive trademarks, whereas it will be lower for all the trademarks between these two extremes (see Figure 8).
It is interesting that the performance of our machine-learning model roughly matches this conception of a human decision-maker's varying confidence in determining trademark distinctiveness. We argue that both human decisionmakers and our model may perform reasonably well in clear-cut cases (where the probability of inherent trademark distinctiveness is either very high or very low), but that they perform much worse for middle-ground trademarks between these extremes.
While human decision-makers and our model may therefore share a similar varying confidence in determining trademark distinctiveness, they exhibit very different levels of transparency. As far as human decision-makers are concerned, their varying confidence in determining trademark distinctiveness is not easily observable. The USPTO examining attorney will not reveal how confident she was when deciding that a trademark is only descriptive rather than suggestive. And even though judges may sometimes indicate that they are less confident in their assessment of a trademark's distinctiveness, this information is neither provided in a systematic nor in a quantifiable or otherwise reliable manner.
By contrast, our model provides information about its own confidence jointly with its prediction. As described in Section 5.1, if one asks the model to predict whether a particular trademark is inherently distinctive, the model can assess its confidence in its own prediction. In our view, it is not only our model that is less confident in determining the distinctiveness of a murky, middle-ground trademark compared with clear-cut cases. Arguably, human decisionmakers will have a similar variance in confidence. Humans may perform well in cases where the model prediction is clear-cut (i.e., where the model has high confidence in whether a trademark is clearly distinctive or clearly not distinctive), while they may perform much worse in the murky areas between such clear cases. After all, our model was trained on actual decisions by human decisionmakers (USPTO examining attorneys). In this case, there may be an entire area of trademarks where both human decision-makers and machine-learning models perform poorly when deciding whether the trademark is inherently distinctive. These are the not-clear-cut cases, and the notion that determining trademark distinctiveness is a binary choice conceals the considerable variance decisionmakers face when having to decide whether a middle-ground trademark is distinctive.
From our perspective, this raises the question whether the Abercrombie spectrum is the right framework for deciding whether murky, middle-ground signs should be protected as trademarks after all. 9 If decision-makers do not have high confidence in applying the Abercrombie spectrum test to such marks, perhaps one should actually refrain from the Abercrombie spectrum test. Rather, one could rely on other normative grounds when deciding whether such signs should be protected as trademarks. When it is not clear, for example, whether a sign is truly inherently distinctive, perhaps trademark law should refrain from providing protection to such signs as overbroad protection may chill speech or unnecessarily clutter the linguistic space of potential trademarks (see Beebe & Fromer, 2018;Ouellette, 2014, p. 360). This could also push trademark owners toward more clearly arbitrary or fanciful marks (Tushnet, 2017, p. 922). One could also rely on the conceptual distance between the primary meaning of a 9
We are not the first to question whether the Abercrombie spectrum should be used to decide cases that are at the boundary between "descriptive or generic" and "arbitrary or fanciful" territory. For a discussion on whether suggestive marks should be treated as descriptive marks (i.e., requiring acquired secondary meaning) under the Abercrombie spectrum, see Linford (2015). Our focus in this article is on the varying confidence of decision-makers in locating a trademark on the Abercrombie spectrum, depending on whether it is a clear-cut or a murky case. This problem may occur in all categories of the Abercrombie spectrum. mark and the goods or services for which it is being used (Fromer, 2022). Or one could activate a refined version of a secondary meaning test in such cases. Our article is not the place to develop an alternative test of trademark distinctiveness. Yet, our empirical results cast some doubt on Abercrombie's central role in trademark jurisprudence.
We have presented a machine-learning model that is trained on a semantic understanding of existing trademark applications to predict whether a trademark is inherently distinctive. Our machine-learning model looks at the trademark, its description and other features. It then tries to understand the meaning of the trademark. Arguably, our method mimics how an experienced USPTO examining attorney who is familiar with the USPTO decision practice determines trademark distinctiveness.
In our view, using a natural-language model such as BERT that has learned the semantic relations among words and sentences is an important step toward determining trademark distinctiveness with quantitative approaches. Ten years ago, Ouellette (2014) proposed to determine trademark strength through conducting Google searches. She argued and demonstrated through case studies that the stronger a trademark is, the more top search results for the trademark will appear on Google. Courts have resisted such approaches on the ground that search results do not provide adequate information about the context in which a trademark is used (Gilson & Gilson, 2024, § 2.05[11]; In Re Bayer Aktiengesellschaft, 2007, p. 967; U.S. Patent and Trademark Office, 2024b, §710.01(b)). Natural language processing tools take a step toward incorporating the semantic context of a trademark application.
Of course, our approach is not without limitations. First, as described in Section 3.1, while our approach allows us to predict whether a trademark is inherently distinctive, it does not allow us to locate the trademark on the five-tier Abercrombie spectrum in a fine-grained way. As trademark distinctiveness is a binary decision in trademark registration, our approach still delivers the relevant prediction in this administrative process. Second, our classifier is only trained with real-world trademark applications that have reached the USPTO in the past. We have not trained the classifier on trademark candidates that a company may have considered during its brand development process, but decided to never put forward to the USPTO. If such trademark candidates differed systematically from the observable trademark applications, we cannot exclude that our model would perform differently on such trademark candidates. However, we have tested our model with court decisions on trademark distinctiveness for outof-sample predictions. Appendix A.2 reports results from this validation. Third, while using BERT as a natural-language processing approach enables our study to understand the semantic features of trademark applications, choosing this machine-learning approach over others has implications for the interpretability of our results. With the emergence of large autoregressive language models such as GPT, one might want to use these models instead of BERT. The potential advantage of GPT over BERT is that it has access to a larger implicit general knowledge base and can reason over different factors. BERT, on the other hand, can be fine-tuned for our specific task and provides calibrated prediction probabilities. Our model is built specifically to follow USPTO decisions using information that is specified in the statute. GPT, meanwhile, would make a binary guess based on its general knowledge, rather than on what the USPTO has previously decided. In line with that, our model performs better at this task than recent baselines using GPT prompts (e.g., Guha et al., 2022). The varying confidence in predicting Abercrombie (see Section 4.2) led us to challenge the notion that the Abercrombie spectrum is a good approach to decide about trademark registrability for murky, middle-ground trademarks (see Section 5.2). Using GPT prompts for our task as in Goodhue and Wei (2023) or Guha et al. (2022) would not easily produce the relevant model precision metrics that are necessary to engage in such normative debate. In addition, using a model such as BERT enables us to use AI explanation tools to explore why our model makes particular predictions. Hence, while GPT-type models open up exciting research avenues for trademark scholarship, the BERT approach is overall better-suited to our approach and aims.
Fourth, it should be noted that the model's confidence scores are concentrated in the middle of the distribution. This means that the model is not confident for most of the trademarks, in which case the system would not be very useful for the decision-maker. Given that, for example, only 2.5% of the data points get a confidence score over 90%, that puts limits on the system's performance. Note that the low confidence could be due to subjectivity, noise, and other irreducible errors in the USPTO process. In that case, no model would be able to confidently predict large swaths of the data.
Fifth, fine-tuning models and deriving explanations on the entire dataset can be computationally expensive. In our study, we used Nvidia V-100 GPUs to fine-tune a comparatively small language model; each epoch taking 2 h. For larger models such as RoBERTa, this time increased to 3.5 h. Further, calculating SHAP values on an instance level is quick, but to have a global view of the features in the dataset is computationally challenging. As an example, for Nice class 12, which had only 3890 examples in our test dataset, it took us 2 h. For Nice classes having multiple thousands of examples across the years, this approach can be computationally infeasible.
Finally, our framework does not capture the trademark registration process in its full complexity. While our system arguably mimics how an experienced USPTO examining attorney would determine trademark distinctiveness, it takes distinctiveness of trademarks as determined by the USPTO in the training set as a given (see Lim, 2022aLim, , pp. 1358Lim, -1359)). Determining ground truth is a recurring problem for trademark distinctiveness, and our article presents two approaches (USPTO decisions in the main analysis and court decisions in Appendix A.2) to tackle this problem. For a project that aims at presenting a decision-support system and exploring the foundations of the Abercrombie test, these approaches seem the right starting points. Also, our model only covers word marks. An interesting extension of this work would be to expand the analysis to special form trademarks (including design elements or colors, as in the case of logos) and image trademarks by either translating images into text descriptors or by applying multimodal machine-learning techniques that work directly on both text and images. Furthermore, current trademark doctrine has many subtle rules for special cases, such as double entendre, composites, misspelling, foreign equivalents, or trademarks with similar pronunciation, which our current framework does not capture (see Goodhue & Xing, 2023). Relatedly, if trademark examination was increasingly automated, and if brand development and trademark applications became increasingly automated as well, one could envision a future in which machines apply for trademarks which are then checked by machines. The recent trend to apply to "nonsense marks" (see Fromer & McKenna, 2024) due to the particular design of the Amazon Brand Registry may point toward such a future. We think that it is for trademark law and procedure to provide a framework that minimizes negative effects of dubious trademark registration strategies (see Fromer & McKenna, 2024, pp. 67-68). A machine-learning model that makes predictions based on earlier decisions by trademark examining attorneys cannot solve larger policy problems that the current trademark system faces. Yet, such a model could be used to identify potential candidates for nonsense trademark applications in a more refined way than by counting the number of consonants or vowels that appear in a word in a row (see Fromer & McKenna, 2024, p. 41). Thereby, machine-learning models could become part of the solution to nonsense marks. Overall, we view our framework not as a system that is ready to be used in the real world, but as a prototype that enables us to explore the extent to which the Abercrombie spectrum can be automated and what kind of normative tradeoffs such automatization would entail.
We started our journey with a vision to automate the Abercrombie spectrum in trademark law. Trademark scholars are often skeptical of such visions, as it is unclear whether artificial intelligence technologies will ever be able to reflect the holistic and human-centric approaches on which many complex trademark doctrines are built (Moerland & Freitas, 2021, p. 284). If trademark doctrines such as inherent distinctiveness require a subjective and nuanced evaluation (see Moerland & Freitas, 2021, p. 288;Katyal & Kesari, 2020, p. 526;Gangjee, 2021, p. 190;Lim, 2022aLim, , pp. 1362Lim, -1364)), it is not clear whether we will ever be able to turn trademark procedure over to the machines.
We view our study as an endeavor to explore these challenges not in the abstract, but through an engineering approach of trial and error. We have presented a prototype for a robot trademark clerk and suggest trying out such a clerk in the field under controlled conditions. In this process, we have observed that our framework may provide helpful empirical input to discussions trademark scholars have only conducted on a theoretical level so far. In their discussion of proximity and distance costs, for example, Hemel and Ouellette (2021) provide a conceptual framework to think about the implications of trademark distinctiveness. But they lack an empirical exploration validating their framework. Our study provides a roadmap on how to quantify distance costs. In our baseline Xgboost model (see Appendix A.6), for example, we use the distance between text embeddings of the trademark and its description as a predictive feature. Furthermore, our efforts on how to infuse the Abercrombie spectrumwhich is a highly nuanced, complicated trademark doctrine that may be applied in an inconsistent and subjective manner-with current machine-learning methods may indicate how challenging it is to define proper benchmarks for large language models (see Guha et al., 2023).
More importantly, our framework has raised doubts whether the Abercrombie spectrum should be the guiding principle for trademark registrability when trademark distinctiveness is uncertain, where the model-predicted probabilities are close to 50%. Our model performs poorly on these middle-ground trademarks, and we expect that this mimics the low confidence that human trademark experts have when deciding such cases. The Abercrombie spectrum has always focused on how consumers perceive a particular sign and then provided normative prescriptions on how such signs should be treated by trademark decision-makers. But when trademark decision-makers have low confidence on how consumers perceive a range of signs, it is perhaps a mistake to let the Abercrombie spectrum derive normative prescriptions from a reality of sign usage that is murky and blurred.
The Abercrombie spectrum serves ultimately only as a heuristic to aid trademark examiners, lawyers, and judges in quickly determining whether, as a matter of trademark policy, a mark merits protection without any showing of acquired distinctiveness. It substitutes for an involved cost-benefit analysis of whether granting protection would benefit competition and consumers. When it was first formulated, the Abercrombie test was recognized as a means of reducing this analysis to what were often more tractable and easily-answered questions: Is the mark a new, coined term? Does the mark have any semantic relation to its products or services? But when those questions are not easily answered, the Abercrombie test may do more harm than good. By focusing on the formal, semantic classification of a mark, it may distract from the core policy question of whether it is optimal to grant exclusive rights in a mark without any additional showing of acquired distinctiveness. In such cases, it may make sense to set the heuristic aside and return to and confront directly the underlying policy question otherwise latent in the Abercrombie test. By demonstrating the substantial limitations of the Abercrombie heuristic, machine-learning methods may teach us to separate those situations in which machinelearning models can effectively do the work of the Abercrombie test and of humans applying it from those situations in which the test should simply not be applied by humans or machines. Automating Abercrombie would therefore not turn trademark procedure over to the machines, but potentially free trademark procedure too beyond the Abercrombie test when it produces inconclusive results.