Documentation menu
Data Model and Data Dictionary
Find a detailed reference of every field and every data schema in the API Reference Models. This page goes into more detail for three specific data schemas:
- the Wikimedia project article model
- the Structured Contents article model
- the Wikidata article model
The Wikimedia project article model
The main set of Wikimedia projects that can be accessed via Wikimedia Enterprise are the 'textual Wikimedia projects'. These are MediaWiki instances that hold some textual knowledge: Wikipedia, Wikibooks, Wikivoyage, Wikiversity, Wiktionary, Wikiquote and Wikisource. All of these projects follow the same data schema, regardless of project or language. This is the Wikimedia project article model. This same data schema is delivered as JSON by the three main Wikimedia Enterprise API Collections:
| API | Returns | What one unit is |
|---|---|---|
| Snapshot | A compressed file for a whole project, downloaded on a schedule | One article per line of NDJSON |
| On-demand | A JSON array of exact matches to an article name query | One article per JSON object in the top-level array |
| Realtime | A stream you hold open, or an hourly batch file | One article per SSE or NDJSON event |
Each of these collections provides access to the same data, but through different delivery models. Use the Snapshot API to download hundreds of thousands of articles as full project bundles or chunks of a Wikimedia project. Retrieve a single article object, or a small set of article objects, using the On-demand API. The Realtime API sends the latest revision of an article object every time a revision happens. Output from all three API collections can be parsed and reused with minimal to no changes on the user side, because of this identical article schema.
Realtime API endpoints are the only endpoints that carry the event.date_published, event.partition and event.offset fields.
The below table explains possible use cases for the properties in the article model. Almost all fields are optional and omitempty=true, meaning they will not show up in the API response if their value is empty.
| What it does | Fields |
|---|---|
| Identifies the page | name, identifier, url, namespace, in_language, is_part_of |
| Holds article content | abstract, article_body, image |
| Links to other content | main_entity, additional_entities, categories, templates, redirects |
| Provides Reuse information | license |
| Places it in time | event, version, previous_version, date_created, date_modified, date_previously_modified |
| Shows community input and signals credibility | protection, visibility, watchers_count, version |
The article_body object holds the contents of the actual article: titles, paragraphs, links, etc. The article_body is provided both as HTML and wikitext. Either or both of the HTML and wikitext versions of the article_body can be absent, if the event.type is delete or visibility-change, or if the article body is empty and thus omitted.
Article revisions
Wikimedia project articles are constantly updated. The event, version and previous_version objects provide all of the metadata about when and what changes have been made to a Wikimedia article.
event describes the change as it passed through the Wikimedia APIs. identifier is a UUID for the change itself, and type is update, delete, or visibility-change. Realtime responses add three more fields - date_published, partition, and offset - which are used to reconnect to the Realtime API stream.
version describes the latest revision to the article. version.identifier is a revision ID, it changes with every edit. version.identifier is only unique within its own Wikimedia project, and increases monotonically within that project. Keeping the article with the highest version.identifier when you have duplicate articles ensures you only keep the latest revision of the article. version holds important information about the edit and the editor, like editor.name and comment, as well as credibility signals such as revertrisk, referencerisk and referenceneed.
previous_version only contains the prior revision's identifier and its number_of_characters.
{ "event": { "identifier": "e69c5020-5b60-4a03-98c2-9c572fe0a0f6", "type": "update", "date_created": "2026-04-10T16:05:39.751737Z", "date_published": "2026-04-10T16:31:57.033Z", "partition": 4, "offset": 3593806 }, "version": { "identifier": 1041549311, "comment": "Reformat 1 archive link.", "scores": { "revertrisk": { "probability": { "false": 0.8429209291934967, "true": 0.1570790708065033 } }, "referencerisk": { "reference_risk_score": 0 }, "referenceneed": { "reference_need_score": 0.16544117647058823 } }, "editor": { "identifier": 27823944, "name": "GreenC bot", "edit_count": 3600514, "groups": [ "bot", "templateeditor", "*", "user", "autoconfirmed" ], "is_bot": true, "date_started": "2016-03-13T16:05:44Z" }, "number_of_characters": 240501, "size": { "value": 240598, "unit_text": "B" }, "maintenance_tags": {} }, "previous_version": { "identifier": 1041549002, "number_of_characters": 305917 }}Credibility signals are included in every payload
Every article revision payload has fields that act as credibility signals, so you can decide, in your own pipeline, how to judge the reliability and trust of an edit.
- Signals from the article:
protection,visibility,watchers_count. - Signals from the revision:
version.scores(RevertRisk, ReferenceRisk, ReferenceNeed),version.maintenance_tags,version.size,version.tags,version.comment,version.is_minor_edit,version.is_flagged_stable,version.is_breaking_news,version.noindex,version.number_of_characters. - Signals from the editor:
version.editor, including group membership, account age, and edit count.
Understand what each signal means and where to set thresholds: Tuning Credibility Signals. Learn more about the value of Credibility Signals: Credibility Signals overview.
Structured Contents article model
structured_content is the same article with the article contents parsed instead of delivered raw. Wikitext and HTML are hard to consume at scale, so the Structured Contents endpoints break an article into machine-readable JSON for easier ingestion.
This schema keeps other metadata about the article unchanged, such as identifier, version and event, so it can be used in conjunction with the same articles retrieved from non-Structured Contents endpoints.
| Fields | |
|---|---|
| Adds | description, infoboxes, sections, references, tables |
| Drops | article_body, categories, templates, redirects, protection, visibility, watchers_count, previous_version, date_previously_modified |
If you need fields that are dropped by the Structured Contents article model, you can call the Article Lookup endpoint with the same article name and specify the fields you need in the fields request body parameter.
curl 'https://api.enterprise.wikimedia.com/v2/articles/{name}' \ --request POST \ --header 'Content-Type: application/json' \ --header 'Authorization: Bearer YOUR_SECRET_TOKEN' \ --data '{ "fields": [ "categories", "templates", "redirects", "protection", "visibility", "watchers_count", "previous_version", "date_previously_modified" ]}'Structured Contents endpoints are in beta: experimental, not covered by an SLA, and not recommended for production. Fields and shapes can change as the parsers improve.
Structured Contents infoboxes, sections, and lists
Structured Contents does not model an infobox and a section as two different things. Both are modeled using the same part object, and a part can contain more subparts without a fixed depth.
A part carries a type (infobox, section, field, list, list_item, paragraph,table, or image), an optional name, and then whichever payload its type calls for: value for a single string, values for a list, images, links, citations, or table_references. Subparts of a part are found in the has_parts object.
To correctly capture all parts and subparts, handle the ingestion of this data recursively. Every article has different depth properties that can't be predicted, and new revisions of an article can change the depth of its parts. Don't use code that ingests parts through has_parts[0].has_parts[1], but recursively ingest has_parts.
If a part has a table or reference as a subpart, it is referenced within that part with a reference identifier. A section points at a table in the table_references field, and a citation inside a paragraph points at a reference by identifier. The link can be one-sided in either direction: some citations might have no matching reference, and some tables might not be referenced from within a section.
Wikidata article model
wikidata_article is what the Wikidata endpoints return. It reuses the same set of metadata fields as the Wikimedia project article model: name, identifier, url, namespace, is_part_of, license, event and date fields. The article body is replaced by the entity object, which contains the Wikidata item or property itself. In Wikidata snapshots and hourly batches every line is one Wikidata article.
The following fields are contained in the entity object:
labels,descriptions, andaliasesname and describe the Wikidata entity per Wikimedia language codesitelinkslink the entity to every matching Wikimedia project pagestatementsdefine the Wikidata item or object itself through subject-predicate-object statements. Every statement is keyed by property ID (P31,P569), and holds statements with arank, avalue, and optionalqualifiersandreferences.
Please keep in mind the following differences between the Wikimedia textual article model and the Wikidata article model:
- The
versionobject is smaller than on a Wikimedia project article. A Wikidata revision carriesidentifier,comment,editor,tags,is_minor_editand ascoresobject with RevertRisk only, so ReferenceRisk and ReferenceNeed do not apply to Wikidata responses. Seewikidata_version.previous_versionis thewikidata_previous_versionobject, the previous revision'sidentifierandnamespace. - The
statementsobject is a map, not a fixed array. Iterate over its PID keys to collect all statement data. Index it directly when you know the PID you want.
For what Wikidata is and how it relates to the other projects, see the Data Primer.
Wikidata endpoints are in beta: experimental, not covered by an SLA, and not recommended for production.
Where every field is defined
The following pages give detailed reference information about every field and data model used in the Wikimedia Enterprise APIs.
- API reference - A searchable Scalar site that references every model and every field, with the endpoints that return them. Permalinks to specific data models are structured as follows:
/docs/api/#models/<name>. - OpenAPI definition - reference file for all Wikimedia Enterprise APIs written following the OpenAPI specification. Can be used with Postman, Insomnia, Swagger, and other API discovery and code generation tools.
- Markdown for agents - append
.mdto any/docs/URL for a plain-text copy of that guide, and /llms.txt indexes the site for language models.
If you are looking for the free Wikimedia APIs
Wikimedia Enterprise is a commercial product of the Wikimedia Foundation, built for organizations requesting Wikimedia content at high volume. It is not the only way to get that content, and might not fit your specific use case.
The Wikimedia Foundation also publishes free APIs and data dumps: the Wikimedia REST API and the API catalog at api.wikimedia.org, the MediaWiki Action API that comes with every MediaWiki instance, and the public dumps. They are free, they cover every project and every namespace, and they are often the best tool for research, one-off analysis, and low-volume reuse. These endpoints are subject to rate limits.
Wikimedia Enterprise APIs should be used if your integration is meant for commercial use, if you require high-volume data retrieval and throughput, or if you often need to ensure data freshness. Wikimedia Enterprise also provides personalized user support, commercial SLA options, and derived features such as the parsed Structured Contents output, and credibility signals computed per revision. Wikimedia Enterprise APIs can only be accessed after you've signed up for a free account. Authentication is done with a JWT bearer token rather than an API key. A free account covers monthly snapshots and a request quota, so you can read the real payloads before deciding.
To learn where your use case fits into the Wikimedia API ecosystem, the Data Primer explains how the projects and their APIs relate, and Getting Started gets you making your first calls in a few minutes.
If you are a Wikimedia community member, you can get exclusive access to Wikimedia Enterprise APIs. Request community access on Meta.
See also
- Getting Started - sign up, get a token, make the first call.
- Identifiers - what each
{identifier}expects, and the difference between page IDs and revision IDs. - Fields, Filters, and Limit - request only the parts of this model you need.
- Tuning Credibility Signals - putting the trust fields to work.
- API reference - every endpoint, every model, every field.