使用Requests和BeautifulSoup爬取网页时过滤li内加粗文本及提取标题下内容
解决步骤
1. 过滤li标签内的加粗标题
只需在提取strong标签时检查其是否存在li祖先节点,即可排除列表内的加粗内容,修改后的提取代码如下:
import requests from bs4 import BeautifulSoup url = 'https://www.emirates.com/pk/english/help/covid-19/dubai-travel-requirements/tourists/' r = requests.get(url) soup = BeautifulSoup(r.content, 'html.parser') headers = [] for strong_tag in soup.find_all('strong'): # 父级链无li标签的strong才判定为标题 if not strong_tag.find_parents('li'): header_text = strong_tag.get_text(strip=True) # 同步做去重处理,避免重复标题入库 if header_text and header_text not in headers: headers.append(header_text)
2. 提取标题对应正文内容
原有代码的核心问题是未重置内容变量、提前单节点插入列表,修改后的代码如下:
main_data = [] for header in headers: # 适配文本首尾空格/换行问题,精准定位标题节点 target = soup.find(['h3', 'p'], string=lambda t: t and header in t.strip()) if not target: continue content_parts = [] # 遍历后续兄弟节点,直到遇到下一个带strong的标题节点 for sib in target.find_next_siblings(): if sib.find('strong'): break sib_text = sib.get_text(strip=True) if sib_text: content_parts.append(sib_text) # 拼接当前标题下所有正文,统一存入结果列表 full_content = '\n'.join(content_parts) main_data.append([full_content])
执行后main_data的每个子列表对应一个标题下的全部正文内容,符合预期输出格式。
内容的提问来源于stack exchange,提问作者Lopez
相关产品推荐
相关产品推荐

