You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取维基百科页面所有纯文本章节?Python实现方案问询

解决方法:利用MediaWiki API获取结构化章节内容

你可以通过MediaWiki官方API结合专门的维基文本解析库,优雅地提取维基百科页面的所有子章节内容,同时自动过滤引用标记、图片说明等冗余信息。以下是具体实现:

步骤说明

  1. 使用MediaWiki API的action=parse接口,获取目标页面的章节结构和原始维基文本(wikitext)。
  2. 借助mwparserfromhell库解析维基文本,该库专门用于处理MediaWiki格式的内容,可以轻松移除引用、图片、表格等冗余元素。
  3. 遍历章节结构,提取每个子章节的标题和清理后的内容。

代码实现

首先安装依赖:

pip install requests mwparserfromhell

然后编写Python代码:

import requests
import mwparserfromhell

# MediaWiki API基础配置
API_URL = "https://en.wikipedia.org/w/api.php"
PAGE_TITLE = "Artificial Intelligence"

def get_wikipedia_sections(title):
    params = {
        "action": "parse",
        "page": title,
        "prop": "sections|wikitext",
        "format": "json",
        "redirects": True
    }
    response = requests.get(API_URL, params=params)
    data = response.json()
    
    # 解析章节列表
    sections = data["parse"]["sections"]
    wikitext = data["parse"]["wikitext"]["*"]
    
    # 解析维基文本
    parsed_wikitext = mwparserfromhell.parse(wikitext)
    
    # 移除冗余元素:引用、图片、表格、注释
    parsed_wikitext.remove(parsed_wikitext.filter_references())
    parsed_wikitext.remove(parsed_wikitext.filter_images())
    parsed_wikitext.remove(parsed_wikitext.filter_tables())
    parsed_wikitext.remove(parsed_wikitext.filter_comments())
    
    # 提取每个章节的内容
    section_contents = {}
    for section in sections:
        section_title = section["line"]
        # 根据章节索引提取对应内容
        section_text = parsed_wikitext.get_sections(matches=section_title, include_lead=False)[0].strip()
        section_contents[section_title] = str(section_text)
    
    return section_contents

# 调用函数并输出结果
ai_sections = get_wikipedia_sections(PAGE_TITLE)
for title, content in ai_sections.items():
    print(f"=== {title} ===")
    print(content[:500] + "..." if len(content) > 500 else content)
    print("\n")

代码解释

  • action=parse:告诉API解析指定页面,返回结构化的章节信息和原始维基文本。
  • prop=sections|wikitext:同时请求章节列表和页面的维基文本内容。
  • mwparserfromhell.parse():将维基文本转换为可操作的对象,方便过滤冗余元素。
  • filter_references()/filter_images()等方法:精准移除引用标记、图片说明等不需要的内容,无需手动处理正则表达式。

这种方法既避免了仅能提取引言的局限,也解决了HTML解析带来的冗余问题,完全基于API的结构化数据处理,是更优雅的维基百科内容提取方案。

内容的提问来源于stack exchange,提问作者blindeyes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 12:10:20