<!-- Wikimedia Enterprise docs. Canonical page: https://enterprise.wikimedia.com/docs/data-dictionary/ -->

# Data Model and Data Dictionary

Find a detailed reference of every field and every data schema in the [API Reference Models](https://enterprise.wikimedia.com/docs/api/#models). This page goes into more detail for three specific data schemas:

- the [Wikimedia project article model](https://enterprise.wikimedia.com/docs/api/#models/article)
- the [Structured Contents article model](https://enterprise.wikimedia.com/docs/api/#models/structured-content)
- the [Wikidata article model](https://enterprise.wikimedia.com/docs/api/#models/wikidata-article)

## The Wikimedia project article model

The main set of Wikimedia projects that can be accessed via Wikimedia Enterprise are the 'textual Wikimedia projects'. These are MediaWiki instances that hold some textual knowledge: Wikipedia, Wikibooks, Wikivoyage, Wikiversity, Wiktionary, Wikiquote and Wikisource.
All of these projects follow the same data schema, regardless of project or language. This is the [Wikimedia project article model](https://enterprise.wikimedia.com/docs/api/#models/article). This same data schema is delivered as JSON by the three main Wikimedia Enterprise API Collections:

| API                                                           | Returns                                                         | What one unit is                                   |
| ------------------------------------------------------------- | --------------------------------------------------------------- | -------------------------------------------------- |
| [Snapshot](https://enterprise.wikimedia.com/docs/snapshot/)   | A compressed file for a whole project, downloaded on a schedule | One article per line of NDJSON                     |
| [On-demand](https://enterprise.wikimedia.com/docs/on-demand/) | A JSON array of exact matches to an article name query          | One article per JSON object in the top-level array |
| [Realtime](https://enterprise.wikimedia.com/docs/realtime/)   | A stream you hold open, or an hourly batch file                 | One article per SSE or NDJSON event                |

Each of these collections provides access to the same data, but through different delivery models. Use the Snapshot API to download hundreds of thousands of articles as full project bundles or chunks of a Wikimedia project. Retrieve a single article object, or a small set of article objects, using the On-demand API. The Realtime API sends the latest revision of an article object every time a revision happens. Output from all three API collections can be parsed and reused with minimal to no changes on the user side, because of this identical article schema.

> **Warning:** Realtime API endpoints are the only endpoints that carry the `event.date_published`, `event.partition` and `event.offset` fields.

The below table explains possible use cases for the properties in the article model. Almost all fields are optional and `omitempty=true`, meaning they will not show up in the API response if their value is empty.

| What it does                                      | Fields                                                                                                                                                                                                                                                                                                                                                                                         |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Identifies the page**                           | `name`, `identifier`, `url`, [`namespace`](https://enterprise.wikimedia.com/docs/api/#models/article-namespace), [`in_language`](https://enterprise.wikimedia.com/docs/api/#models/language), [`is_part_of`](https://enterprise.wikimedia.com/docs/api/#models/project)                                                                                                                        |
| **Holds article content**                         | `abstract`, [`article_body`](https://enterprise.wikimedia.com/docs/api/#models/article-body), [`image`](https://enterprise.wikimedia.com/docs/api/#models/image)                                                                                                                                                                                                                               |
| **Links to other content**                        | [`main_entity`](https://enterprise.wikimedia.com/docs/api/#models/entity), [`additional_entities`](https://enterprise.wikimedia.com/docs/api/#models/entity), [`categories`](https://enterprise.wikimedia.com/docs/api/#models/category), [`templates`](https://enterprise.wikimedia.com/docs/api/#models/template), [`redirects`](https://enterprise.wikimedia.com/docs/api/#models/redirect) |
| **Provides Reuse information**                    | [`license`](https://enterprise.wikimedia.com/docs/api/#models/license)                                                                                                                                                                                                                                                                                                                         |
| **Places it in time**                             | [`event`](https://enterprise.wikimedia.com/docs/api/#models/event), [`version`](https://enterprise.wikimedia.com/docs/api/#models/version), [`previous_version`](https://enterprise.wikimedia.com/docs/api/#models/previous-version), `date_created`, `date_modified`, `date_previously_modified`                                                                                              |
| **Shows community input and signals credibility** | [`protection`](https://enterprise.wikimedia.com/docs/api/#models/protection), [`visibility`](https://enterprise.wikimedia.com/docs/api/#models/visibility), `watchers_count`, `version`                                                                                                                                                                                                        |

The **`article_body`** object holds the contents of the actual article: titles, paragraphs, links, etc. The `article_body` is provided both as HTML and wikitext. Either or both of the HTML and wikitext versions of the `article_body` can be absent, if the `event.type` is `delete` or `visibility-change`, or if the article body is empty and thus omitted.

### Article revisions

Wikimedia project articles are constantly updated. The `event`, `version` and `previous_version` objects provide all of the metadata about when and what changes have been made to a Wikimedia article.

[**`event`**](https://enterprise.wikimedia.com/docs/api/#models/event) describes the change as it passed through the Wikimedia APIs. `identifier` is a UUID for the change itself, and `type` is `update`, `delete`, or `visibility-change`. Realtime responses add three more fields - `date_published`, `partition`, and `offset` - which are used to [reconnect to the Realtime API stream](https://enterprise.wikimedia.com/docs/realtime/#reconnecting-to-the-realtime-api).

[**`version`**](https://enterprise.wikimedia.com/docs/api/#models/version) describes the latest revision to the article. `version.identifier` is a revision ID, it changes with every edit. `version.identifier` is only unique within its own Wikimedia project, and increases monotonically within that project. Keeping the article with the highest `version.identifier` when you have duplicate articles ensures you only keep the latest revision of the article. `version` holds important information about the edit and the editor, like `editor.name` and `comment`, as well as credibility signals such as `revertrisk`, `referencerisk` and `referenceneed`.

[**`previous_version`**](https://enterprise.wikimedia.com/docs/api/#models/previous-version) only contains the prior revision's `identifier` and its `number_of_characters`.

```json
{
  "event": {
    "identifier": "e69c5020-5b60-4a03-98c2-9c572fe0a0f6",
    "type": "update",
    "date_created": "2026-04-10T16:05:39.751737Z",
    "date_published": "2026-04-10T16:31:57.033Z",
    "partition": 4,
    "offset": 3593806
  },
  "version": {
    "identifier": 1041549311,
    "comment": "Reformat 1 archive link.",
    "scores": {
        "revertrisk": {
            "probability": {
                "false": 0.8429209291934967,
                "true": 0.1570790708065033
            }
        },
        "referencerisk": {
            "reference_risk_score": 0
        },
        "referenceneed": {
            "reference_need_score": 0.16544117647058823
        }
    },
    "editor": {
        "identifier": 27823944,
        "name": "GreenC bot",
        "edit_count": 3600514,
        "groups": [
            "bot",
            "templateeditor",
            "*",
            "user",
            "autoconfirmed"
        ],
        "is_bot": true,
        "date_started": "2016-03-13T16:05:44Z"
    },
    "number_of_characters": 240501,
    "size": {
        "value": 240598,
        "unit_text": "B"
    },
    "maintenance_tags": {}
  },
  "previous_version": {
    "identifier": 1041549002,
    "number_of_characters": 305917
  }
}
```

### Credibility signals are included in every payload

Every article revision payload has fields that act as credibility signals, so you can decide, in your own pipeline, how to judge the reliability and trust of an edit.

- Signals from the article: [`protection`](https://enterprise.wikimedia.com/docs/api/#models/protection), [`visibility`](https://enterprise.wikimedia.com/docs/api/#models/visibility), `watchers_count`.
- Signals from the revision: [`version.scores`](https://enterprise.wikimedia.com/docs/api/#models/scores) (RevertRisk, ReferenceRisk, ReferenceNeed), [`version.maintenance_tags`](https://enterprise.wikimedia.com/docs/api/#models/maintenance-tags), [`version.size`](https://enterprise.wikimedia.com/docs/api/#models/size), `version.tags`, `version.comment`, `version.is_minor_edit`, `version.is_flagged_stable`, `version.is_breaking_news`, `version.noindex`, `version.number_of_characters`.
- Signals from the editor: [`version.editor`](https://enterprise.wikimedia.com/docs/api/#models/editor), including group membership, account age, and edit count.

Understand what each signal means and where to set thresholds: [Tuning Credibility Signals](https://enterprise.wikimedia.com/docs/credibility-signals/). Learn more about the value of Credibility Signals: [Credibility Signals overview](https://enterprise.wikimedia.com/api/credibility-signals/).

## Structured Contents article model

[`structured_content`](https://enterprise.wikimedia.com/docs/api/#models/structured-content) is the same article with the article contents parsed instead of delivered raw. Wikitext and HTML are hard to consume at scale, so the [Structured Contents](https://enterprise.wikimedia.com/api/structured-contents/) endpoints break an article into machine-readable JSON for easier ingestion.

This schema keeps other metadata about the article unchanged, such as `identifier`, `version` and `event`, so it can be used in conjunction with the same articles retrieved from non-Structured Contents endpoints.

|           | Fields                                                                                                                                                                                                                                                                                                       |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Adds**  | `description`, [`infoboxes`](https://enterprise.wikimedia.com/docs/api/#models/part), [`sections`](https://enterprise.wikimedia.com/docs/api/#models/part), [`references`](https://enterprise.wikimedia.com/docs/api/#models/reference), [`tables`](https://enterprise.wikimedia.com/docs/api/#models/table) |
| **Drops** | `article_body`, `categories`, `templates`, `redirects`, `protection`, `visibility`, `watchers_count`, `previous_version`, `date_previously_modified`                                                                                                                                                         |

If you need fields that are dropped by the Structured Contents article model, you can call the [Article Lookup](https://enterprise.wikimedia.com/docs/api/#tag/articles/GET/v2/articles/{name}) endpoint with the same article name and specify the fields you need in the `fields` request body parameter.

```bash
curl 'https://api.enterprise.wikimedia.com/v2/articles/{name}' \
  --request POST \
  --header 'Content-Type: application/json' \
  --header 'Authorization: Bearer YOUR_SECRET_TOKEN' \
  --data '{
  "fields": [
    "categories",
    "templates",
    "redirects",
    "protection",
    "visibility",
    "watchers_count",
    "previous_version",
    "date_previously_modified"
  ]
}'
```

> **Note:** Structured Contents endpoints are in beta: experimental, not covered by an SLA, and not recommended for production. Fields and shapes can change as the parsers improve.

### Structured Contents infoboxes, sections, and lists

Structured Contents does not model an infobox and a section as two different things. Both are modeled using the same [`part`](https://enterprise.wikimedia.com/docs/api/#models/part) object, and a part can contain more subparts without a fixed depth.

A part carries a `type` (`infobox`, `section`, `field`, `list`, `list_item`, `paragraph`,`table`, or `image`), an optional `name`, and then whichever payload its type calls for: `value` for a single string, `values` for a list, [`images`](https://enterprise.wikimedia.com/docs/api/#models/image), [`links`](https://enterprise.wikimedia.com/docs/api/#models/link), [`citations`](https://enterprise.wikimedia.com/docs/api/#models/citation), or [`table_references`](https://enterprise.wikimedia.com/docs/api/#models/table-reference). Subparts of a part are found in the `has_parts` object.

**To correctly capture all parts and subparts, handle the ingestion of this data recursively**. Every article has different depth properties that can't be predicted, and new revisions of an article can change the depth of its parts. Don't use code that ingests parts through `has_parts[0].has_parts[1]`, but recursively ingest `has_parts`.

If a part has a table or reference as a subpart, it is referenced within that part with a reference identifier. A section points at a table in the `table_references` field, and a [`citation`](https://enterprise.wikimedia.com/docs/api/#models/citation) inside a paragraph points at a [`reference`](https://enterprise.wikimedia.com/docs/api/#models/reference) by identifier. The link can be one-sided in either direction: some citations might have no matching reference, and some tables might not be referenced from within a section.

## Wikidata article model

[`wikidata_article`](https://enterprise.wikimedia.com/docs/api/#models/wikidata-article) is what the [Wikidata endpoints](https://enterprise.wikimedia.com/docs/api/#tag/wikidata) return. It reuses the same set of metadata fields as the [Wikimedia project article model](https://enterprise.wikimedia.com/docs/api/#models/article): `name`, `identifier`, `url`, `namespace`, `is_part_of`, `license`, `event` and date fields. The article body is replaced by the [`entity`](https://enterprise.wikimedia.com/docs/api/#models/wikidata-entity) object, which contains the Wikidata item or property itself. In [Wikidata snapshots](https://enterprise.wikimedia.com/docs/snapshot/#wikidata-snapshots) and [hourly batches](https://enterprise.wikimedia.com/docs/realtime/#wikidata-batches) every line is one Wikidata article.

The following fields are contained in the `entity` object:

- `labels`, `descriptions`, and `aliases` name and describe the Wikidata entity per Wikimedia language code
- `sitelinks` link the entity to every matching Wikimedia project page
- [`statements`](https://enterprise.wikimedia.com/docs/api/#models/wikidata-statement-map) define the Wikidata item or object itself through subject-predicate-object statements. Every statement is keyed by property ID (`P31`, `P569`), and holds [statements](https://enterprise.wikimedia.com/docs/api/#models/wikidata-statement) with a `rank`, a `value`, and optional `qualifiers` and `references`.

Please keep in mind the following differences between the Wikimedia textual article model and the Wikidata article model:

- **The `version` object is smaller than on a Wikimedia project article.** A Wikidata revision carries `identifier`, `comment`, `editor`, `tags`, `is_minor_edit` and a `scores` object with RevertRisk only, so ReferenceRisk and ReferenceNeed do not apply to Wikidata responses. See [`wikidata_version`](https://enterprise.wikimedia.com/docs/api/#models/wikidata-version). `previous_version` is the [`wikidata_previous_version`](https://enterprise.wikimedia.com/docs/api/#models/wikidata-previous-version) object, the previous revision's `identifier` and `namespace`.
- **The `statements` object is a map, not a fixed array.** Iterate over its PID keys to collect all statement data. Index it directly when you know the PID you want.

For what Wikidata is and how it relates to the other projects, see the [Data Primer](https://enterprise.wikimedia.com/docs/data-primer/).

> **Note:** Wikidata endpoints are in beta: experimental, not covered by an SLA, and not recommended for production.

## Where every field is defined

The following pages give detailed reference information about every field and data model used in the Wikimedia Enterprise APIs.

- **[API reference](https://enterprise.wikimedia.com/docs/api/)** - A searchable Scalar site that references every model and every field, with the endpoints that return them. Permalinks to specific data models are structured as follows: `/docs/api/#models/<name>`.
- **[OpenAPI definition](https://enterprise.wikimedia.com/docs/api/wme-api.yaml)** - reference file for all Wikimedia Enterprise APIs written following the OpenAPI specification. Can be used with Postman, Insomnia, Swagger, and other API discovery and code generation tools.
- **Markdown for agents** - append `.md` to any `/docs/` URL for a plain-text copy of that guide, and [/llms.txt](https://enterprise.wikimedia.com/llms.txt) indexes the site for language models.

## If you are looking for the free Wikimedia APIs

Wikimedia Enterprise is a commercial product of the Wikimedia Foundation, built for organizations requesting Wikimedia content at high volume. It is not the only way to get that content, and might not fit your specific use case.

The Wikimedia Foundation also publishes free APIs and data dumps: the [Wikimedia REST API and the API catalog](https://api.wikimedia.org/wiki/API_catalog) at api.wikimedia.org, the MediaWiki [Action API](https://www.mediawiki.org/wiki/API:Main_page) that comes with every MediaWiki instance, and the [public dumps](https://dumps.wikimedia.org/). They are free, they cover every project and every namespace, and they are often the best tool for research, one-off analysis, and low-volume reuse. These endpoints are subject to [rate limits](https://www.mediawiki.org/wiki/Wikimedia_APIs/Rate_limits).

Wikimedia Enterprise APIs should be used if your integration is meant for commercial use, if you require high-volume data retrieval and throughput, or if you often need to ensure data freshness. Wikimedia Enterprise also provides personalized user support, commercial SLA options, and derived features such as the parsed Structured Contents output, and credibility signals computed per revision. Wikimedia Enterprise APIs can only be accessed after you've signed up for a free account. Authentication is done with a [JWT bearer token](https://enterprise.wikimedia.com/docs/authentication/) rather than an API key. A [free account](https://dashboard.enterprise.wikimedia.com/signup/?utm_source=website\&utm_medium=cta\&utm_campaign=signup\&utm_content=docs-data-model) covers monthly snapshots and a request quota, so you can read the real payloads before deciding.

To learn where your use case fits into the Wikimedia API ecosystem, the [Data Primer](https://enterprise.wikimedia.com/docs/data-primer/) explains how the projects and their APIs relate, and [Getting Started](https://enterprise.wikimedia.com/docs/) gets you making your first calls in a few minutes.

> **Note:** If you are a Wikimedia community member, you can get exclusive access to Wikimedia Enterprise APIs. [Request community access on Meta](https://meta.wikimedia.org/wiki/Wikimedia_Enterprise#Access).

## See also

- [Getting Started](https://enterprise.wikimedia.com/docs/) - sign up, get a token, make the first call.
- [Identifiers](https://enterprise.wikimedia.com/docs/identifiers/) - what each `{identifier}` expects, and the difference between page IDs and revision IDs.
- [Fields, Filters, and Limit](https://enterprise.wikimedia.com/docs/filtering/) - request only the parts of this model you need.
- [Tuning Credibility Signals](https://enterprise.wikimedia.com/docs/credibility-signals/) - putting the trust fields to work.
- [API reference](https://enterprise.wikimedia.com/docs/api/) - every endpoint, every model, every field.
