You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则表达式匹配维基百科HTML内容异常求助

解决维基百科HTML标题间内容提取问题

问题原因

  • 正则默认处于单行模式,.不会匹配换行符,导致你的正则只能匹配单行内容,出现截断或匹配失败的情况。
  • 正则写法存在语法错误:[\n|.]+里的|是字符集内的普通字符,并非逻辑或的含义,完全没必要添加。

正则修复方案

如果坚持用正则,需启用re.DOTALL模式(让.匹配包括换行在内的所有字符),同时使用非贪婪匹配避免过度匹配:

import wikipedia
import re

cancers = wikipedia.page("List_of_cancer_types")
# 匹配两个h2标签之间的内容,非贪婪模式+DOTALL
regex = r'<h2 id="Bone_and_muscle_sarcoma".*?>(.*?)<h2 id="Brain_and_nervous_system"'
pattern = re.compile(regex, re.DOTALL)
result = re.findall(pattern, cancers.html())
if result:
    print(result[0])

更可靠的HTML解析方案:用BeautifulSoup

正则处理HTML极易出现边界问题,推荐使用专业的HTML解析库:

  1. 先安装依赖库:
pip install beautifulsoup4
  1. 编写解析代码:
import wikipedia
from bs4 import BeautifulSoup

cancers = wikipedia.page("List_of_cancer_types")
soup = BeautifulSoup(cancers.html(), 'html.parser')

# 定位目标标题标签
start_h2 = soup.find('h2', id='Bone_and_muscle_sarcoma')
end_h2 = soup.find('h2', id='Brain_and_nervous_system')

# 提取两个标题间的所有内容
content = []
current_element = start_h2.next_sibling
while current_element != end_h2:
    if current_element.name:
        content.append(str(current_element))
    current_element = current_element.next_sibling

# 输出完整内容
print(''.join(content))

# 若仅需提取列表项文本:
target_ul = start_h2.find_next('ul')
for li in target_ul.find_all('li'):
    print(li.get_text(strip=True))

该方案能精准处理HTML结构,不会因标签格式细微变化导致解析失败。

内容的提问来源于stack exchange,提问作者MattJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 01:00:15