如何从英文维基百科XML Dump提取标题与内容用于Tantivy_py测试?
处理维基百科XML Dump提取标题和内容的方案
Python 实现(快速开发,低内存占用)
针对92GB的超大XML文件,必须用流式解析避免加载整个文件到内存。结合xml.etree.ElementTree的iterparse和专门清理维基标记的mwparserfromhell库,可以高效提取条目页面(命名空间为0的内容,过滤掉讨论页、用户页等非条目内容)。
步骤1:安装依赖
pip install mwparserfromhell
步骤2:流式解析脚本
import xml.etree.ElementTree as ET import mwparserfromhell # 维基百科XML的命名空间 MEDIAWIKI_NS = {"mw": "http://www.mediawiki.org/xml/export-0.10/"} def extract_wiki_content(xml_path, output_txt): with open(output_txt, 'w', encoding='utf-8') as out_file: # 仅监听page元素的结束事件,流式处理 for event, elem in ET.iterparse( xml_path, events=('end',), tag='{' + MEDIAWIKI_NS['mw'] + '}page' ): # 过滤非条目页面(仅保留namespace=0的内容) ns_node = elem.find('{' + MEDIAWIKI_NS['mw'] + '}ns') if ns_node is None or ns_node.text != '0': elem.clear() # 清除父节点引用,释放内存 while elem.getprevious() is not None: del elem.getparent()[0] continue # 提取标题 title_node = elem.find('{' + MEDIAWIKI_NS['mw'] + '}title') title = title_node.text.strip() if title_node else '' # 提取页面内容(最新版本的text) revision_node = elem.find('{' + MEDIAWIKI_NS['mw'] + '}revision') text_node = revision_node.find('{' + MEDIAWIKI_NS['mw'] + '}text') if revision_node else None raw_content = text_node.text.strip() if text_node and text_node.text else '' # 清理维基标记(链接、模板、格式符等) clean_content = mwparserfromhell.parse(raw_content).strip_code() if raw_content else '' # 写入文件,用分隔符区分不同页面 out_file.write(f"Title: {title}\n") out_file.write(f"Content: {clean_content}\n") out_file.write("---\n") # 释放当前元素的内存 elem.clear() while elem.getprevious() is not None: del elem.getparent()[0] if __name__ == "__main__": # 替换为你的XML Dump路径和输出文件路径 extract_wiki_content("enwiki-latest-pages-articles.xml", "wiki_extracted.txt")
优化:处理压缩Dump
如果下载的是.bz2或.xz压缩版Dump,可以直接流式解压处理,无需先解压整个92GB文件:
import bz2 # 对应.bz2压缩包 # import lzma # 对应.xz压缩包 def extract_compressed_wiki(compressed_path, output_txt): with bz2.BZ2File(compressed_path, 'r') as xml_file, open(output_txt, 'w', encoding='utf-8') as out_file: context = ET.iterparse( xml_file, events=('end',), tag='{' + MEDIAWIKI_NS['mw'] + '}page' ) # 后续逻辑和上面的extract_wiki_content一致,省略重复代码
C++ 实现(极致性能,低内存开销)
如果需要更高的处理速度,用C++结合libxml2的SAX解析(纯流式,不加载整个文件到内存)是最优选择。
核心思路
- 用libxml2的SAX回调函数,逐元素解析XML
- 仅保留命名空间为0的条目页面
- 清理维基标记(可结合正则或字符串替换实现)
示例代码框架
#include <iostream> #include <fstream> #include <string> #include <libxml/parser.h> #include <libxml/xmlsax.h> // 全局状态变量,记录当前解析的元素和内容 std::string current_element; std::string current_title; std::string current_content; bool is_entry_page = false; std::ofstream* output_file; // 清理维基标记的工具函数 std::string clean_wiki_markup(const std::string& text) { std::string cleaned = text; // 移除[[链接]]格式 size_t pos; while ((pos = cleaned.find("[[", 0)) != std::string::npos) { size_t end = cleaned.find("]]", pos); if (end != std::string::npos) cleaned.erase(pos, end - pos + 2); else break; } // 移除{{模板}}格式 while ((pos = cleaned.find("{{", 0)) != std::string::npos) { size_t end = cleaned.find("}}", pos); if (end != std::string::npos) cleaned.erase(pos, end - pos + 2); else break; } return cleaned; } // SAX回调:开始解析元素 void start_element(void* ctx, const xmlChar* name, const xmlChar** attrs) { current_element = reinterpret_cast<const char*>(name); if (current_element == "page") { // 重置页面状态 current_title.clear(); current_content.clear(); is_entry_page = false; } } // SAX回调:结束解析元素 void end_element(void* ctx, const xmlChar* name) { std::string elem_name = reinterpret_cast<const char*>(name); if (elem_name == "ns") { // 判断是否为条目页面 is_entry_page = (current_content == "0"); } else if (elem_name == "page" && is_entry_page) { // 写入清理后的内容 std::string cleaned = clean_wiki_markup(current_content); *output_file << "Title: " << current_title << "\n"; *output_file << "Content: " << cleaned << "\n"; *output_file << "---\n"; } current_element.clear(); current_content.clear(); } // SAX回调:读取元素文本内容 void characters(void* ctx, const xmlChar* ch, int len) { std::string text(reinterpret_cast<const char*>(ch), len); if (current_element == "title") { current_title += text; } else if (current_element == "ns") { current_content = text; } else if (current_element == "text") { current_content += text; } } int main() { output_file = new std::ofstream("wiki_extracted.txt", std::ios::out | std::ios::binary); if (!output_file->is_open()) { std::cerr << "Failed to open output file" << std::endl; return 1; } // 初始化SAX处理器 xmlSAXHandler sax_handler = {0}; sax_handler.startElement = start_element; sax_handler.endElement = end_element; sax_handler.characters = characters; // 流式解析XML文件 if (xmlSAXUserParseFile(&sax_handler, nullptr, "enwiki-latest-pages-articles.xml") != 0) { std::cerr << "XML parsing failed" << std::endl; return 1; } output_file->close(); delete output_file; xmlCleanupParser(); return 0; }
编译运行
需要链接libxml2库:
g++ -o wiki_extractor wiki_extractor.cpp -lxml2 ./wiki_extractor
替代方案:使用现成工具
如果不想自己写代码,用WikiExtractor(Python编写的维基百科提取工具)可以快速得到处理好的纯文本内容:
- 直接从命令行运行,自动流式处理大文件
- 自动清理维基标记,过滤非条目页面
- 支持分块输出,避免生成单个超大文本文件
运行示例:
python wikiextractor.py --output extracted --bytes 100M enwiki-latest-pages-articles.xml
运行后会在extracted目录下生成多个小文件,每个文件包含若干页面的标题和纯文本内容,格式适合直接用于全文搜索测试。
内容的提问来源于stack exchange,提问作者chetxn04
相关产品推荐
相关产品推荐

