Python正则表达式匹配维基百科HTML内容异常求助
解决维基百科HTML标题间内容提取问题
问题原因
- 正则默认处于单行模式,
.不会匹配换行符,导致你的正则只能匹配单行内容,出现截断或匹配失败的情况。 - 正则写法存在语法错误:
[\n|.]+里的|是字符集内的普通字符,并非逻辑或的含义,完全没必要添加。
正则修复方案
如果坚持用正则,需启用re.DOTALL模式(让.匹配包括换行在内的所有字符),同时使用非贪婪匹配避免过度匹配:
import wikipedia import re cancers = wikipedia.page("List_of_cancer_types") # 匹配两个h2标签之间的内容,非贪婪模式+DOTALL regex = r'<h2 id="Bone_and_muscle_sarcoma".*?>(.*?)<h2 id="Brain_and_nervous_system"' pattern = re.compile(regex, re.DOTALL) result = re.findall(pattern, cancers.html()) if result: print(result[0])
更可靠的HTML解析方案:用BeautifulSoup
正则处理HTML极易出现边界问题,推荐使用专业的HTML解析库:
- 先安装依赖库:
pip install beautifulsoup4
- 编写解析代码:
import wikipedia from bs4 import BeautifulSoup cancers = wikipedia.page("List_of_cancer_types") soup = BeautifulSoup(cancers.html(), 'html.parser') # 定位目标标题标签 start_h2 = soup.find('h2', id='Bone_and_muscle_sarcoma') end_h2 = soup.find('h2', id='Brain_and_nervous_system') # 提取两个标题间的所有内容 content = [] current_element = start_h2.next_sibling while current_element != end_h2: if current_element.name: content.append(str(current_element)) current_element = current_element.next_sibling # 输出完整内容 print(''.join(content)) # 若仅需提取列表项文本: target_ul = start_h2.find_next('ul') for li in target_ul.find_all('li'): print(li.get_text(strip=True))
该方案能精准处理HTML结构,不会因标签格式细微变化导致解析失败。
内容的提问来源于stack exchange,提问作者MattJ
相关产品推荐
相关产品推荐

