You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将维基百科/维基媒体Dump XML文件转文本?是否有工具包?

Hey there! Let's tackle your questions about converting Wikipedia/Wikimedia XML dumps to text using Python—this is a super common task, and there are some solid tools to make it way easier.

Available Python Toolkits

First, let's cover the go-to libraries that handle the heavy lifting for you:

  • mwparserfromhell: A dedicated library for parsing Wikipedia's wikitext syntax. It can strip out all the messy markup (templates, links, tables, etc.) to give you clean plain text, and it's highly customizable if you need fine-grained control.
  • wikiextractor: A lightweight, fast tool built specifically for extracting and cleaning text from Wikimedia XML dumps. It's perfect for batch processing large dumps without writing tons of custom code.
Method 1: Using mwparserfromhell with XML Dump Parsing

This approach gives you full control over how you process each page. We'll use mwxml (another handy library) to read the XML dump efficiently, even if it's compressed.

First, install the dependencies:

pip install mwparserfromhell mwxml

Here's a sample script to extract page titles and plain text:

import mwxml
import mwparserfromhell

def process_wikimedia_dump(dump_file_path):
    # Iterate through each page in the XML dump
    for page in mwxml.Dump.from_file(open(dump_file_path, "rb")):
        # Only process main namespace pages (skip talk pages, user pages, etc.)
        if page.namespace == 0:
            # Grab the latest revision of the page
            for revision in page:
                wikitext_content = revision.text
                if wikitext_content:
                    # Parse the raw wikitext
                    parsed_content = mwparserfromhell.parse(wikitext_content)
                    # Strip all wiki markup to get plain text
                    plain_text = parsed_content.strip_code()
                    # Print or save the results (truncated here for brevity)
                    print(f"Page Title: {page.title}")
                    print(f"Plain Text Preview: {plain_text[:500]}...\n")
                    break  # Stop after the latest revision

if __name__ == "__main__":
    # Replace with your dump file path (works with .bz2 compressed files directly)
    process_wikimedia_dump("enwiki-latest-pages-articles.xml.bz2")

A quick note: mwxml handles compressed dumps (like .bz2) natively, so you don't need to unzip them first. The strip_code() method removes almost all wiki-specific markup, but you can tweak it if you want to keep certain elements (like links) by passing additional parameters.

Method 2: Fast Batch Processing with wikiextractor

If you're dealing with a huge dump and just want clean text output without custom logic, wikiextractor is your best bet. It automatically splits the output into manageable files and handles compression.

Install it first:

pip install wikiextractor

You can run it directly from the command line (this is the most common way):

wikiextractor your-wikimedia-dump.xml.bz2 --output extracted-text --bytes 1M --compress
  • --output: Directory where the cleaned text files will be saved
  • --bytes: Maximum size per output file (adjust based on your needs)
  • --compress: Compress output files with gzip to save space

If you want to run it within a Python script, use subprocess:

import subprocess

subprocess.run([
    "wikiextractor",
    "your-wikimedia-dump.xml.bz2",
    "--output", "extracted-text",
    "--bytes", "1M",
    "--compress"
])
Quick Clarification: Wikipedia vs. Wikimedia Dumps

Just to clear things up: Wikipedia is a part of the Wikimedia Foundation, so Wikipedia XML dumps are a subset of Wikimedia dumps. Both tools above work seamlessly with other Wikimedia project dumps (like Wiktionary, Wikibooks, etc.)—just swap out the input XML file path, and you're good to go.

Pro tip: For large-scale processing, stick with wikiextractor—it's optimized for speed and low memory usage. If you need to extract specific sections, modify the text before saving, or retain certain markup elements, mwparserfromhell gives you the flexibility to do that.

内容的提问来源于stack exchange,提问作者Amit Sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 06:24:45