You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按章节提取Wikipedia文本及对应引用构建关联数据集?

解决方案:按章节提取维基百科文本及对应引用

不用手动解析HTML,推荐使用专门处理维基百科wikitext的mwparserfromhell工具,结合维基百科API获取完整内容,就能实现按章节关联引用的需求。以下是具体实现步骤和代码:

步骤说明

  1. 通过维基百科API获取目标文章的完整wikitext(包含章节结构、引用标记和参考文献)
  2. 用mwparserfromhell解析wikitext,提取指定层级(H2/H3)的章节
  3. 提取每个章节内的引用标记,建立引用编号/名称与实际链接的映射
  4. 整理成你需要的数据集格式

代码实现

首先安装依赖:

pip install mwparserfromhell requests

核心代码:

import mwparserfromhell
import requests

def get_wiki_wikitext(title):
    """调用维基百科API获取文章完整wikitext"""
    url = "https://en.wikipedia.org/w/api.php"
    params = {
        "action": "query",
        "format": "json",
        "titles": title,
        "prop": "revisions",
        "rvprop": "content",
        "rvslots": "main"
    }
    resp = requests.get(url, params=params)
    page_data = next(iter(resp.json()["query"]["pages"].values()))
    return page_data["revisions"][0]["slots"]["main"]["content"]

def extract_section_dataset(wikitext):
    """解析wikitext,生成章节-文本-引用的数据集"""
    parsed_content = mwparserfromhell.parse(wikitext)
    dataset = []
    
    # 提取所有H2、H3层级的章节
    for section in parsed_content.get_sections(levels=[2, 3]):
        # 获取章节标题
        section_title = section.filter_headings()[0].title.strip()
        # 获取章节文本(保留原始内容,去除wikitext标签但保留引用上下文)
        section_text = section.strip_code(keep_template_params=False).strip()
        
        # 提取当前章节内的引用标签
        section_refs = section.filter_tags(matches=lambda tag: tag.tag == "ref")
        ref_ids = []
        for ref_tag in section_refs:
            # 引用可能有自定义name,或者用自动编号
            ref_name = ref_tag.get("name", None)
            ref_ids.append(ref_name if ref_name else str(len(ref_ids)+1))
        
        # 建立全局引用映射:引用ID -> 外部链接
        ref_map = {}
        # 先尝试找参考文献章节(==References==)
        ref_section = parsed_content.get_sections(matches="References", levels=[2])
        if ref_section:
            ref_parsed = mwparserfromhell.parse(str(ref_section[0]))
            all_ref_tags = ref_parsed.filter_tags(matches=lambda tag: tag.tag == "ref")
            for idx, ref_tag in enumerate(all_ref_tags, 1):
                # 提取引用中的外部链接
                ext_links = ref_tag.filter_external_links()
                if ext_links:
                    ref_id = ref_tag.get("name", str(idx))
                    ref_map[ref_id] = str(ext_links[0].url)
        else:
            # 处理用{{reflist}}模板的情况,遍历所有全局ref标签
            all_global_refs = parsed_content.filter_tags(matches=lambda tag: tag.tag == "ref")
            for idx, ref_tag in enumerate(all_global_refs, 1):
                ext_links = ref_tag.filter_external_links()
                if ext_links:
                    ref_id = ref_tag.get("name", str(idx))
                    ref_map[ref_id] = str(ext_links[0].url)
        
        # 匹配当前章节的引用链接
        section_references = [ref_map[rid] for rid in ref_ids if rid in ref_map]
        
        # 添加到数据集
        dataset.append({
            "title": section_title,
            "text": section_text,
            "references": section_references
        })
    return dataset

# 使用示例
if __name__ == "__main__":
    # 替换为你需要的维基百科文章标题
    article_title = "Python (programming language)"
    wikitext = get_wiki_wikitext(article_title)
    section_dataset = extract_section_dataset(wikitext)
    # 打印第一个章节的结果
    print(section_dataset[0])

注意事项

  • 针对非英文维基百科,修改API的url参数即可(比如中文维基用https://zh.wikipedia.org/w/api.php)
  • 部分引用使用{{cite web}}等模板,需要额外解析模板参数提取URL,可以通过ref_tag.filter_templates()获取模板后提取url参数
  • 若引用是维基内部链接,可通过https://en.wikipedia.org/wiki/ + 链接标题拼接成完整URL

内容的提问来源于stack exchange,提问作者Nitzan Barzilay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 09:03:26