You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过单个循环同时提取网页中p标签与ul标签的内容?

单个循环抓取同一层级p与ul标签内容的解决方案

你可以通过调整find_all的参数,同时选中目标层级下的p和ul标签,再在循环内区分标签类型处理内容,实现单个循环完成抓取需求。

修改后的代码:

import requests
from bs4 import BeautifulSoup as bs

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36",
    "Accept-Encoding": "gzip, deflate, br",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9",
    "DNT": "1",
    "Connection": "close",
    "Upgrade-Insecure-Requests": "1"
}

source = requests.get('https://insights.blackcoffer.com/how-small-business-can-survive-the-coronavirus-crisis/', headers=headers)
page = source.content
soup = bs(page, 'html.parser')

information = ''
# 同时选取同一层级下的p和ul标签
for section in soup.find('div', class_='td-post-content').find_all(['p', 'ul']):
    if section.name == 'p':
        # 处理p标签内容
        content = section.text.strip()
    else:
        # 处理ul标签,遍历所有li提取内容
        li_texts = [li.text.strip() for li in section.find_all('li')]
        content = '\n- ' + '\n- '.join(li_texts)
    
    if content:  # 跳过空内容
        if information:
            information += '\n\n' + content
        else:
            information = content

print(information)

关键说明:

  • 将find_all('p')改为find_all(['p', 'ul']),让BeautifulSoup一次性抓取目标容器下所有同级的p和ul标签
  • 循环内通过section.name判断当前标签类型:
    • 若为p标签,直接提取文本
    • 若为ul标签,遍历内部的li标签,将每个li的文本整理为带列表符号的格式
  • 加入空内容判断,避免输出多余的换行

内容的提问来源于stack exchange,提问作者Mr. Basu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 06:01:01