如何按章节提取Wikipedia文本及对应引用构建关联数据集?
解决方案:按章节提取维基百科文本及对应引用
不用手动解析HTML,推荐使用专门处理维基百科wikitext的mwparserfromhell工具,结合维基百科API获取完整内容,就能实现按章节关联引用的需求。以下是具体实现步骤和代码:
步骤说明
- 通过维基百科API获取目标文章的完整wikitext(包含章节结构、引用标记和参考文献)
- 用
mwparserfromhell解析wikitext,提取指定层级(H2/H3)的章节 - 提取每个章节内的引用标记,建立引用编号/名称与实际链接的映射
- 整理成你需要的数据集格式
代码实现
首先安装依赖:
pip install mwparserfromhell requests
核心代码:
import mwparserfromhell import requests def get_wiki_wikitext(title): """调用维基百科API获取文章完整wikitext""" url = "https://en.wikipedia.org/w/api.php" params = { "action": "query", "format": "json", "titles": title, "prop": "revisions", "rvprop": "content", "rvslots": "main" } resp = requests.get(url, params=params) page_data = next(iter(resp.json()["query"]["pages"].values())) return page_data["revisions"][0]["slots"]["main"]["content"] def extract_section_dataset(wikitext): """解析wikitext,生成章节-文本-引用的数据集""" parsed_content = mwparserfromhell.parse(wikitext) dataset = [] # 提取所有H2、H3层级的章节 for section in parsed_content.get_sections(levels=[2, 3]): # 获取章节标题 section_title = section.filter_headings()[0].title.strip() # 获取章节文本(保留原始内容,去除wikitext标签但保留引用上下文) section_text = section.strip_code(keep_template_params=False).strip() # 提取当前章节内的引用标签 section_refs = section.filter_tags(matches=lambda tag: tag.tag == "ref") ref_ids = [] for ref_tag in section_refs: # 引用可能有自定义name,或者用自动编号 ref_name = ref_tag.get("name", None) ref_ids.append(ref_name if ref_name else str(len(ref_ids)+1)) # 建立全局引用映射:引用ID -> 外部链接 ref_map = {} # 先尝试找参考文献章节(==References==) ref_section = parsed_content.get_sections(matches="References", levels=[2]) if ref_section: ref_parsed = mwparserfromhell.parse(str(ref_section[0])) all_ref_tags = ref_parsed.filter_tags(matches=lambda tag: tag.tag == "ref") for idx, ref_tag in enumerate(all_ref_tags, 1): # 提取引用中的外部链接 ext_links = ref_tag.filter_external_links() if ext_links: ref_id = ref_tag.get("name", str(idx)) ref_map[ref_id] = str(ext_links[0].url) else: # 处理用{{reflist}}模板的情况,遍历所有全局ref标签 all_global_refs = parsed_content.filter_tags(matches=lambda tag: tag.tag == "ref") for idx, ref_tag in enumerate(all_global_refs, 1): ext_links = ref_tag.filter_external_links() if ext_links: ref_id = ref_tag.get("name", str(idx)) ref_map[ref_id] = str(ext_links[0].url) # 匹配当前章节的引用链接 section_references = [ref_map[rid] for rid in ref_ids if rid in ref_map] # 添加到数据集 dataset.append({ "title": section_title, "text": section_text, "references": section_references }) return dataset # 使用示例 if __name__ == "__main__": # 替换为你需要的维基百科文章标题 article_title = "Python (programming language)" wikitext = get_wiki_wikitext(article_title) section_dataset = extract_section_dataset(wikitext) # 打印第一个章节的结果 print(section_dataset[0])
注意事项
- 针对非英文维基百科,修改API的
url参数即可(比如中文维基用https://zh.wikipedia.org/w/api.php) - 部分引用使用
{{cite web}}等模板,需要额外解析模板参数提取URL,可以通过ref_tag.filter_templates()获取模板后提取url参数 - 若引用是维基内部链接,可通过
https://en.wikipedia.org/wiki/+ 链接标题拼接成完整URL
内容的提问来源于stack exchange,提问作者Nitzan Barzilay
相关产品推荐
相关产品推荐

