Keywords

Eg:keywords/machine-learningCount:1,881,699Links:works

A keyword is a short phrase saying what a work is about: gut-microbiota, alphafold, urban-heat-island. Keywords are OpenAlex’s most specific aboutness signal, much finer than topics. There are about 1.9 million keywords, and 419 million works carry at least one, typically five or six. A keyword’s OpenAlex ID is a readable slug, so a keyword looks like https://openalex.org/keywords/machine-learning; fetch one at api.openalex.org/keywords/machine-learning.

About

A model reads each work’s title, abstract and venue and writes the keywords a careful indexer would: what the work is about, not what kind of document it is. When a record says nothing about its subject (a bare dataset entry, a table of contents), the model returns no keywords rather than guess. That is why about 11% of works have none.

The model is a small open model (Qwen3-4B) trained to copy a frontier model (Claude Opus 5.5) acting as an indexer on about a million works. Its keywords are then folded into one vocabulary: spelling variants, plurals and synonyms are merged under one heading (neanderthals also covers “Neandertals”), and a keyword enters the vocabulary only once it describes at least 100 works.

Each keyword on a work carries a score: the model’s confidence in that keyword, from 0 to 1. Keywords are listed best first. On the keyword object itself, works_count and cited_by_count roll those assignments up across the corpus.

On 400 random works, an independent AI judge (OpenAI’s GPT-6 Astra) rated 91% of these keywords accurate, against 44% for the previous, topic-derived ones. They match the keywords authors chose for their own papers about twice as often. Benchmarks, code, prompts, the vocabulary and the model weights are all open: openalex-keywords (v3), with the weights in the v3.0 release.

The vocabulary changes

The keyword list is not fixed. New fields emerge, and keywords are added for them. Keywords get merged when they turn out to mean the same thing, and split when one label covers two meanings. If a keyword is wrong, or two keywords should be one, tell us. Store keyword IDs with the date you fetched them, and expect some to change.

Each keyword has a one-sentence description and, where one exists, a link to its Wikidata item in ids. Each keyword also lists its related topics (see topics): the topics its works fall in, each with the share of the keyword’s works in that topic. Keywords sit beside the topic tree rather than under it: one keyword can relate to several topics, with different weights. Next, we’ll look at joining keywords to other outside vocabularies like MeSH, and vector search over keywords.

Good uses

  • Find what text search misses. A keyword finds works whatever words their authors used, including works with no abstract. Search does this for you: the title, abstract and keywords search (openalex.org’s default) matches a work when your words are in its title or abstract or a phrase you typed names one of its keywords (full recipe: Finding papers with keywords):

    https://api.openalex.org/works?search.title_abstract_keywords="antimicrobial resistance"
  • Map a field. Filter to any set of works, then group_by=keywords.id to see what it is made of; add group_by=publication_year on a keyword filter to see its trend.

  • Find experts. Filter on a keyword, then group_by=authorships.author.id.

  • Tag your own text. The text aboutness endpoint returns keywords for any title and abstract, from the same model and vocabulary.

Keywords are specific by design; for broad summaries (“how much of this institution’s work is chemistry?”), use topics, subfields and fields.

Attributes

This is the canonical dictionary of every attribute on a keyword object. Attributes shared with other entities are documented once on Common attributes; keyword-specific notes are below.

id

String. The OpenAlex ID for this keyword. Unlike most entities, a keyword’s ID is a readable slug rather than a letter-and-number code, e.g. https://openalex.org/keywords/machine-learning. See Common attributes.

display_name

String. The keyword’s human-readable label, e.g. machine learning. See Common attributes.

description

String. One sentence saying what the keyword means, written by an AI model from the keyword and works that carry it. Machine-made: it can be wrong, and it improves as the vocabulary changes.

ids

Object. External identifiers for this keyword, as URIs. Keyword-specific keys: openalex and, when a confident match exists, wikidata (about a fifth of keywords, which cover most keyword uses). Matches are machine-made. See Common attributes.

primary_topic

Object. The keyword’s main related topic, when one topic holds at least 20% of the keyword’s works: id, display_name, score, and the topic’s subfield, field and domain, the same shape as a work’s primary_topic. score is the share of the keyword’s works that fall in this topic. null for broad keywords whose works spread across many topics (about 29% of keywords, e.g. “quantum fluctuations” or most place names); those still list their topics.

topics

List. The keyword’s related topics: every topic holding at least 7% of the keyword’s works, best first, each with the same fields as primary_topic. A keyword links to about two topics on average. Links come from the topics of the works that carry the keyword and are recomputed daily. On a judged sample, 92% of primary topics and 80% of the other links were acceptable.

works_count

Integer. The number of works tagged with this keyword, across the whole corpus. A keyword filter searches only the core corpus by default, so add corpus=all to match this count: /works?filter=keywords.id:remote-work&corpus=all. See Common attributes.

cited_by_count

Integer. The total citations received by all works tagged with this keyword. See Common attributes.

works_api_url

String. A ready-made Works API URL returning every work tagged with this keyword, e.g. https://api.openalex.org/works?filter=keywords.id:keywords/machine-learning. A convenience link: it’s the same query you’d build with the keywords.id filter.

created_date

String. The date this keyword was added to OpenAlex (YYYY-MM-DD). See Common attributes.

updated_date

String. The ISO 8601 UTC timestamp of the last change to this keyword object. See Common attributes.

In the API

The Keywords endpoint is at api.openalex.org/keywords. Fetch a single keyword by its slug ID (/keywords/machine-learning) or a list, and filter, search, sort, group, and page over the fields above.

To find the works carrying a keyword, filter on the Works endpoint:

https://api.openalex.org/works?filter=keywords.id:machine-learning

To list every keyword related to a topic, filter keywords by topic: topics.id matches any related topic, primary_topic.id only the primary one.

https://api.openalex.org/keywords?filter=topics.id:T10283&sort=works_count:desc

For the full list of filterable, sortable, and groupable fields see the Keywords API reference; for all endpoints see the endpoints index.

Last updated

View as Markdown