You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从英文维基百科XML Dump提取标题与内容用于Tantivy_py测试?

处理维基百科XML Dump提取标题和内容的方案

Python 实现(快速开发,低内存占用)

针对92GB的超大XML文件,必须用流式解析避免加载整个文件到内存。结合xml.etree.ElementTree的iterparse和专门清理维基标记的mwparserfromhell库,可以高效提取条目页面(命名空间为0的内容,过滤掉讨论页、用户页等非条目内容)。

步骤1:安装依赖

pip install mwparserfromhell

步骤2:流式解析脚本

import xml.etree.ElementTree as ET
import mwparserfromhell

# 维基百科XML的命名空间
MEDIAWIKI_NS = {"mw": "http://www.mediawiki.org/xml/export-0.10/"}

def extract_wiki_content(xml_path, output_txt):
    with open(output_txt, 'w', encoding='utf-8') as out_file:
        # 仅监听page元素的结束事件,流式处理
        for event, elem in ET.iterparse(
            xml_path,
            events=('end',),
            tag='{' + MEDIAWIKI_NS['mw'] + '}page'
        ):
            # 过滤非条目页面(仅保留namespace=0的内容)
            ns_node = elem.find('{' + MEDIAWIKI_NS['mw'] + '}ns')
            if ns_node is None or ns_node.text != '0':
                elem.clear()
                # 清除父节点引用,释放内存
                while elem.getprevious() is not None:
                    del elem.getparent()[0]
                continue

            # 提取标题
            title_node = elem.find('{' + MEDIAWIKI_NS['mw'] + '}title')
            title = title_node.text.strip() if title_node else ''

            # 提取页面内容(最新版本的text)
            revision_node = elem.find('{' + MEDIAWIKI_NS['mw'] + '}revision')
            text_node = revision_node.find('{' + MEDIAWIKI_NS['mw'] + '}text') if revision_node else None
            raw_content = text_node.text.strip() if text_node and text_node.text else ''

            # 清理维基标记(链接、模板、格式符等)
            clean_content = mwparserfromhell.parse(raw_content).strip_code() if raw_content else ''

            # 写入文件,用分隔符区分不同页面
            out_file.write(f"Title: {title}\n")
            out_file.write(f"Content: {clean_content}\n")
            out_file.write("---\n")

            # 释放当前元素的内存
            elem.clear()
            while elem.getprevious() is not None:
                del elem.getparent()[0]

if __name__ == "__main__":
    # 替换为你的XML Dump路径和输出文件路径
    extract_wiki_content("enwiki-latest-pages-articles.xml", "wiki_extracted.txt")

优化:处理压缩Dump

如果下载的是.bz2或.xz压缩版Dump,可以直接流式解压处理,无需先解压整个92GB文件:

import bz2  # 对应.bz2压缩包
# import lzma  # 对应.xz压缩包

def extract_compressed_wiki(compressed_path, output_txt):
    with bz2.BZ2File(compressed_path, 'r') as xml_file, open(output_txt, 'w', encoding='utf-8') as out_file:
        context = ET.iterparse(
            xml_file,
            events=('end',),
            tag='{' + MEDIAWIKI_NS['mw'] + '}page'
        )
        # 后续逻辑和上面的extract_wiki_content一致,省略重复代码

C++ 实现(极致性能,低内存开销)

如果需要更高的处理速度,用C++结合libxml2的SAX解析(纯流式,不加载整个文件到内存)是最优选择。

核心思路

  • 用libxml2的SAX回调函数,逐元素解析XML
  • 仅保留命名空间为0的条目页面
  • 清理维基标记(可结合正则或字符串替换实现)

示例代码框架

#include <iostream>
#include <fstream>
#include <string>
#include <libxml/parser.h>
#include <libxml/xmlsax.h>

// 全局状态变量,记录当前解析的元素和内容
std::string current_element;
std::string current_title;
std::string current_content;
bool is_entry_page = false;
std::ofstream* output_file;

// 清理维基标记的工具函数
std::string clean_wiki_markup(const std::string& text) {
    std::string cleaned = text;
    // 移除[[链接]]格式
    size_t pos;
    while ((pos = cleaned.find("[[", 0)) != std::string::npos) {
        size_t end = cleaned.find("]]", pos);
        if (end != std::string::npos) cleaned.erase(pos, end - pos + 2);
        else break;
    }
    // 移除{{模板}}格式
    while ((pos = cleaned.find("{{", 0)) != std::string::npos) {
        size_t end = cleaned.find("}}", pos);
        if (end != std::string::npos) cleaned.erase(pos, end - pos + 2);
        else break;
    }
    return cleaned;
}

// SAX回调:开始解析元素
void start_element(void* ctx, const xmlChar* name, const xmlChar** attrs) {
    current_element = reinterpret_cast<const char*>(name);
    if (current_element == "page") {
        // 重置页面状态
        current_title.clear();
        current_content.clear();
        is_entry_page = false;
    }
}

// SAX回调:结束解析元素
void end_element(void* ctx, const xmlChar* name) {
    std::string elem_name = reinterpret_cast<const char*>(name);
    if (elem_name == "ns") {
        // 判断是否为条目页面
        is_entry_page = (current_content == "0");
    } else if (elem_name == "page" && is_entry_page) {
        // 写入清理后的内容
        std::string cleaned = clean_wiki_markup(current_content);
        *output_file << "Title: " << current_title << "\n";
        *output_file << "Content: " << cleaned << "\n";
        *output_file << "---\n";
    }
    current_element.clear();
    current_content.clear();
}

// SAX回调:读取元素文本内容
void characters(void* ctx, const xmlChar* ch, int len) {
    std::string text(reinterpret_cast<const char*>(ch), len);
    if (current_element == "title") {
        current_title += text;
    } else if (current_element == "ns") {
        current_content = text;
    } else if (current_element == "text") {
        current_content += text;
    }
}

int main() {
    output_file = new std::ofstream("wiki_extracted.txt", std::ios::out | std::ios::binary);
    if (!output_file->is_open()) {
        std::cerr << "Failed to open output file" << std::endl;
        return 1;
    }

    // 初始化SAX处理器
    xmlSAXHandler sax_handler = {0};
    sax_handler.startElement = start_element;
    sax_handler.endElement = end_element;
    sax_handler.characters = characters;

    // 流式解析XML文件
    if (xmlSAXUserParseFile(&sax_handler, nullptr, "enwiki-latest-pages-articles.xml") != 0) {
        std::cerr << "XML parsing failed" << std::endl;
        return 1;
    }

    output_file->close();
    delete output_file;
    xmlCleanupParser();
    return 0;
}

编译运行

需要链接libxml2库:

g++ -o wiki_extractor wiki_extractor.cpp -lxml2
./wiki_extractor

替代方案:使用现成工具

如果不想自己写代码,用WikiExtractor(Python编写的维基百科提取工具)可以快速得到处理好的纯文本内容:

  • 直接从命令行运行,自动流式处理大文件
  • 自动清理维基标记,过滤非条目页面
  • 支持分块输出,避免生成单个超大文本文件

运行示例:

python wikiextractor.py --output extracted --bytes 100M enwiki-latest-pages-articles.xml

运行后会在extracted目录下生成多个小文件,每个文件包含若干页面的标题和纯文本内容,格式适合直接用于全文搜索测试。

内容的提问来源于stack exchange,提问作者chetxn04

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 04:51:01