Skip to content
Wikimedia Enterprise
Documentation menu

Data Model and Data Dictionary

View as Markdown (opens in a new tab)

Find a detailed reference of every field and every data schema in the API Reference Models. This page goes into more detail for three specific data schemas:

The Wikimedia project article model

The main set of Wikimedia projects that can be accessed via Wikimedia Enterprise are the 'textual Wikimedia projects'. These are MediaWiki instances that hold some textual knowledge: Wikipedia, Wikibooks, Wikivoyage, Wikiversity, Wiktionary, Wikiquote and Wikisource. All of these projects follow the same data schema, regardless of project or language. This is the Wikimedia project article model. This same data schema is delivered as JSON by the three main Wikimedia Enterprise API Collections:

APIReturnsWhat one unit is
SnapshotA compressed file for a whole project, downloaded on a scheduleOne article per line of NDJSON
On-demandA JSON array of exact matches to an article name queryOne article per JSON object in the top-level array
RealtimeA stream you hold open, or an hourly batch fileOne article per SSE or NDJSON event

Each of these collections provides access to the same data, but through different delivery models. Use the Snapshot API to download hundreds of thousands of articles as full project bundles or chunks of a Wikimedia project. Retrieve a single article object, or a small set of article objects, using the On-demand API. The Realtime API sends the latest revision of an article object every time a revision happens. Output from all three API collections can be parsed and reused with minimal to no changes on the user side, because of this identical article schema.

Warning:

Realtime API endpoints are the only endpoints that carry the event.date_published, event.partition and event.offset fields.

The below table explains possible use cases for the properties in the article model. Almost all fields are optional and omitempty=true, meaning they will not show up in the API response if their value is empty.

What it doesFields
Identifies the pagename, identifier, url, namespace, in_language, is_part_of
Holds article contentabstract, article_body, image
Links to other contentmain_entity, additional_entities, categories, templates, redirects
Provides Reuse informationlicense
Places it in timeevent, version, previous_version, date_created, date_modified, date_previously_modified
Shows community input and signals credibilityprotection, visibility, watchers_count, version

The article_body object holds the contents of the actual article: titles, paragraphs, links, etc. The article_body is provided both as HTML and wikitext. Either or both of the HTML and wikitext versions of the article_body can be absent, if the event.type is delete or visibility-change, or if the article body is empty and thus omitted.

Article revisions

Wikimedia project articles are constantly updated. The event, version and previous_version objects provide all of the metadata about when and what changes have been made to a Wikimedia article.

event describes the change as it passed through the Wikimedia APIs. identifier is a UUID for the change itself, and type is update, delete, or visibility-change. Realtime responses add three more fields - date_published, partition, and offset - which are used to reconnect to the Realtime API stream.

version describes the latest revision to the article. version.identifier is a revision ID, it changes with every edit. version.identifier is only unique within its own Wikimedia project, and increases monotonically within that project. Keeping the article with the highest version.identifier when you have duplicate articles ensures you only keep the latest revision of the article. version holds important information about the edit and the editor, like editor.name and comment, as well as credibility signals such as revertrisk, referencerisk and referenceneed.

previous_version only contains the prior revision's identifier and its number_of_characters.

{
"event": {
"identifier": "e69c5020-5b60-4a03-98c2-9c572fe0a0f6",
"type": "update",
"date_created": "2026-04-10T16:05:39.751737Z",
"date_published": "2026-04-10T16:31:57.033Z",
"partition": 4,
"offset": 3593806
},
"version": {
"identifier": 1041549311,
"comment": "Reformat 1 archive link.",
"scores": {
"revertrisk": {
"probability": {
"false": 0.8429209291934967,
"true": 0.1570790708065033
}
},
"referencerisk": {
"reference_risk_score": 0
},
"referenceneed": {
"reference_need_score": 0.16544117647058823
}
},
"editor": {
"identifier": 27823944,
"name": "GreenC bot",
"edit_count": 3600514,
"groups": [
"bot",
"templateeditor",
"*",
"user",
"autoconfirmed"
],
"is_bot": true,
"date_started": "2016-03-13T16:05:44Z"
},
"number_of_characters": 240501,
"size": {
"value": 240598,
"unit_text": "B"
},
"maintenance_tags": {}
},
"previous_version": {
"identifier": 1041549002,
"number_of_characters": 305917
}
}

Credibility signals are included in every payload

Every article revision payload has fields that act as credibility signals, so you can decide, in your own pipeline, how to judge the reliability and trust of an edit.

  • Signals from the article: protection, visibility, watchers_count.
  • Signals from the revision: version.scores (RevertRisk, ReferenceRisk, ReferenceNeed), version.maintenance_tags, version.size, version.tags, version.comment, version.is_minor_edit, version.is_flagged_stable, version.is_breaking_news, version.noindex, version.number_of_characters.
  • Signals from the editor: version.editor, including group membership, account age, and edit count.

Understand what each signal means and where to set thresholds: Tuning Credibility Signals. Learn more about the value of Credibility Signals: Credibility Signals overview.

Structured Contents article model

structured_content is the same article with the article contents parsed instead of delivered raw. Wikitext and HTML are hard to consume at scale, so the Structured Contents endpoints break an article into machine-readable JSON for easier ingestion.

This schema keeps other metadata about the article unchanged, such as identifier, version and event, so it can be used in conjunction with the same articles retrieved from non-Structured Contents endpoints.

Fields
Addsdescription, infoboxes, sections, references, tables
Dropsarticle_body, categories, templates, redirects, protection, visibility, watchers_count, previous_version, date_previously_modified

If you need fields that are dropped by the Structured Contents article model, you can call the Article Lookup endpoint with the same article name and specify the fields you need in the fields request body parameter.

curl 'https://api.enterprise.wikimedia.com/v2/articles/{name}' \
--request POST \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer YOUR_SECRET_TOKEN' \
--data '{
"fields": [
"categories",
"templates",
"redirects",
"protection",
"visibility",
"watchers_count",
"previous_version",
"date_previously_modified"
]
}'
Note:

Structured Contents endpoints are in beta: experimental, not covered by an SLA, and not recommended for production. Fields and shapes can change as the parsers improve.

Structured Contents infoboxes, sections, and lists

Structured Contents does not model an infobox and a section as two different things. Both are modeled using the same part object, and a part can contain more subparts without a fixed depth.

A part carries a type (infobox, section, field, list, list_item, paragraph,table, or image), an optional name, and then whichever payload its type calls for: value for a single string, values for a list, images, links, citations, or table_references. Subparts of a part are found in the has_parts object.

To correctly capture all parts and subparts, handle the ingestion of this data recursively. Every article has different depth properties that can't be predicted, and new revisions of an article can change the depth of its parts. Don't use code that ingests parts through has_parts[0].has_parts[1], but recursively ingest has_parts.

If a part has a table or reference as a subpart, it is referenced within that part with a reference identifier. A section points at a table in the table_references field, and a citation inside a paragraph points at a reference by identifier. The link can be one-sided in either direction: some citations might have no matching reference, and some tables might not be referenced from within a section.

Wikidata article model

wikidata_article is what the Wikidata endpoints return. It reuses the same set of metadata fields as the Wikimedia project article model: name, identifier, url, namespace, is_part_of, license, event and date fields. The article body is replaced by the entity object, which contains the Wikidata item or property itself. In Wikidata snapshots and hourly batches every line is one Wikidata article.

The following fields are contained in the entity object:

  • labels, descriptions, and aliases name and describe the Wikidata entity per Wikimedia language code
  • sitelinks link the entity to every matching Wikimedia project page
  • statements define the Wikidata item or object itself through subject-predicate-object statements. Every statement is keyed by property ID (P31, P569), and holds statements with a rank, a value, and optional qualifiers and references.

Please keep in mind the following differences between the Wikimedia textual article model and the Wikidata article model:

  • The version object is smaller than on a Wikimedia project article. A Wikidata revision carries identifier, comment, editor, tags, is_minor_edit and a scores object with RevertRisk only, so ReferenceRisk and ReferenceNeed do not apply to Wikidata responses. See wikidata_version. previous_version is the wikidata_previous_version object, the previous revision's identifier and namespace.
  • The statements object is a map, not a fixed array. Iterate over its PID keys to collect all statement data. Index it directly when you know the PID you want.

For what Wikidata is and how it relates to the other projects, see the Data Primer.

Note:

Wikidata endpoints are in beta: experimental, not covered by an SLA, and not recommended for production.

Where every field is defined

The following pages give detailed reference information about every field and data model used in the Wikimedia Enterprise APIs.

  • API reference - A searchable Scalar site that references every model and every field, with the endpoints that return them. Permalinks to specific data models are structured as follows: /docs/api/#models/<name>.
  • OpenAPI definition - reference file for all Wikimedia Enterprise APIs written following the OpenAPI specification. Can be used with Postman, Insomnia, Swagger, and other API discovery and code generation tools.
  • Markdown for agents - append .md to any /docs/ URL for a plain-text copy of that guide, and /llms.txt indexes the site for language models.

If you are looking for the free Wikimedia APIs

Wikimedia Enterprise is a commercial product of the Wikimedia Foundation, built for organizations requesting Wikimedia content at high volume. It is not the only way to get that content, and might not fit your specific use case.

The Wikimedia Foundation also publishes free APIs and data dumps: the Wikimedia REST API and the API catalog at api.wikimedia.org, the MediaWiki Action API that comes with every MediaWiki instance, and the public dumps. They are free, they cover every project and every namespace, and they are often the best tool for research, one-off analysis, and low-volume reuse. These endpoints are subject to rate limits.

Wikimedia Enterprise APIs should be used if your integration is meant for commercial use, if you require high-volume data retrieval and throughput, or if you often need to ensure data freshness. Wikimedia Enterprise also provides personalized user support, commercial SLA options, and derived features such as the parsed Structured Contents output, and credibility signals computed per revision. Wikimedia Enterprise APIs can only be accessed after you've signed up for a free account. Authentication is done with a JWT bearer token rather than an API key. A free account covers monthly snapshots and a request quota, so you can read the real payloads before deciding.

To learn where your use case fits into the Wikimedia API ecosystem, the Data Primer explains how the projects and their APIs relate, and Getting Started gets you making your first calls in a few minutes.

Note:

If you are a Wikimedia community member, you can get exclusive access to Wikimedia Enterprise APIs. Request community access on Meta.

See also