You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Requests和BeautifulSoup爬取网页时过滤li内加粗文本及提取标题下内容

解决步骤

1. 过滤li标签内的加粗标题

只需在提取strong标签时检查其是否存在li祖先节点,即可排除列表内的加粗内容,修改后的提取代码如下:

import requests
from bs4 import BeautifulSoup

url = 'https://www.emirates.com/pk/english/help/covid-19/dubai-travel-requirements/tourists/'
r = requests.get(url)
soup = BeautifulSoup(r.content, 'html.parser')
headers = []

for strong_tag in soup.find_all('strong'):
    # 父级链无li标签的strong才判定为标题
    if not strong_tag.find_parents('li'):
        header_text = strong_tag.get_text(strip=True)
        # 同步做去重处理,避免重复标题入库
        if header_text and header_text not in headers:
            headers.append(header_text)

2. 提取标题对应正文内容

原有代码的核心问题是未重置内容变量、提前单节点插入列表,修改后的代码如下:

main_data = []
for header in headers:
    # 适配文本首尾空格/换行问题,精准定位标题节点
    target = soup.find(['h3', 'p'], string=lambda t: t and header in t.strip())
    if not target:
        continue
    content_parts = []
    # 遍历后续兄弟节点,直到遇到下一个带strong的标题节点
    for sib in target.find_next_siblings():
        if sib.find('strong'):
            break
        sib_text = sib.get_text(strip=True)
        if sib_text:
            content_parts.append(sib_text)
    # 拼接当前标题下所有正文,统一存入结果列表
    full_content = '\n'.join(content_parts)
    main_data.append([full_content])

执行后main_data的每个子列表对应一个标题下的全部正文内容,符合预期输出格式。

内容的提问来源于stack exchange,提问作者Lopez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 15:27:02