A COMPARISON OF ALGORITHMS FOR THE EXTRACTION OF KEYWORDS IN A PATENT DATABASE
Thiago V. Reginaldo1; Daniel L. B. Lucindo1; Magali R. G. Meireles1; Zenilton K. Patrocínio Júnior1; Paulo E. M. de Almeida2
1 Pontifical Catholic University of Minas Gerais; 2 Federal Center for Technological Education
doi:10.20906/CPS/CILAMCE2017-0715
Resumo
Patents are one of the oldest forms of intellectual capital protection and are an important source of information for measuring the technological advancement of a specific domain of knowledge. However, patents, available on large digital databases, are complex legal documents, which generally contain more details and descriptions than scientific articles. The information contained in the patents is distributed in a large number of fields that can be accessed in the portals made available by the governmental authorities or through search tools. Classification systems used by these offices group the documents into sections, classes, subclasses, groups and subgroups. An automatic system for identifying the set of keywords or key terms related to the main contents of patent groups would enable advances in the issues associated with the automatic generation of titles in patent subgroups. To implement this strategy, it is essential to identify the keywords that represent the patent since the keywords are not pre-defined by the patent authors. This work presents results of an empirical experiment, which compares the application of three algorithms for the extraction of keywords in patents, namely: X2, TF.IDF and Wiki-TFIDF. The patent database used for this experiment consists of 100 patents from the United States Patent and Trademark Office (USPTO) classified by the Cooperative Patent Classification (CPC) system. The patents are from G06K 7/1443 subgroup of the subclass GO6K, called "Recognition of data, presentation of data, record carriers; handling record carriers". To validate the methods, the ground truth was defined as the USPTO description for the selected subgroup. The texts of the patents were preprocessed, withdrawing special characters, stop words, and stemming the words. The Wiki-TFIDF method also uses the title database of the Wikipedia main namespace (22-Jun-2017), which lists all the articles titles on the website and provides external information that refines the choice of keywords. In this method, th
Palavras-chave: Keywords extraction; Information Systems; Knowledge Organization; Patents