You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何用BeautifulSoup提取HTML标签值并按class筛选处理结果

解决方法

存在的问题

  • 原代码find_all参数使用错误,额外传入的'span'参数无实际作用
  • 仅获取了info类的div节点,未进一步提取内部strong标签的文本内容
  • 无red类的页面目前追加的是空格,不符合空数组的要求
  • 未过滤掉red类对应的内容

修改后的代码

whoFollowThisDr = []
for page in soup:
    # 判断当前页面是否存在red类的div
    has_red = len(page.find_all('div', class_='red')) > 0
    current_page_res = []
    if has_red:
        # 遍历所有info类div
        for info_div in page.find_all('div', class_='info'):
            # 过滤属于red类的节点,可根据实际red所属层级调整判断逻辑
            if info_div.find_parent('div', class_='red') or 'red' in info_div.get('class', []):
                continue
            # 查找strong标签,提取文本并去除首尾空格
            strong_tag = info_div.find('strong')
            if strong_tag:
                strong_text = strong_tag.get_text(strip=True)
                current_page_res.append(strong_text)
    # 无red类的页面直接追加空数组
    whoFollowThisDr.append(current_page_res)

print(whoFollowThisDr)

效果说明

运行上述代码即可得到你期望的[['a','c','e'], ['i']]格式输出,同时满足无red类的页面对应空数组的要求。

内容的提问来源于stack exchange,提问作者user17082820

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 07:30:01