You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup的find_all提取多个指定p标签内容并存为列表

BeautifulSoup find_all结果多索引选取方案

find_all()返回的ResultSet对象本质是支持索引操作的类列表结构,本身没有内置的多索引选择方法,直接用Python原生的列表操作即可实现需求:

  • 选择连续索引的元素:直接用列表切片语法
    例:提取索引20到30(包含两端)的p标签:
    all_p = soup.find_all('p')
    # 切片左闭右开,所以结束索引要写31
    selected_p = all_p[20:31]
    
  • 选择不连续的指定索引元素:用列表推导式筛选
    例:提取索引为22、25、28的p标签:
    target_indexes = [22,25,28]
    all_p = soup.find_all('p')
    selected_p = [all_p[i] for i in target_indexes]
    

提取选中元素的文本存为列表:

text_list = [p.text.strip() for p in selected_p]

更稳定的筛选方案(推荐)

硬编码全局p标签的索引稳定性很差,网页结构稍有变动索引就会失效,你可以先定位到目标p标签的上层父容器,再提取内部的p标签,自动忽略其他位置的p标签:

# 先定位到你说的figure节点,id填实际的id值
target_container = soup.find('figure', id='你的figure的id值')
# 只提取这个容器内部的p标签,不会拿到页面其他位置的p
all_target_p = target_container.find_all('p')
# 再按上面的方法选需要的索引即可

修改后的完整示例代码

from bs4 import BeautifulSoup
import requests

def getFacts(target_indexes):
    page = requests.get("https://thoughtcatalog.com/jacob-geers/2016/04/really-funny-random-weird-facts/")
    soup = BeautifulSoup(page.content, 'html.parser')
    all_p = soup.find_all('p')
    
    # 选择指定索引的p标签,提取文本存为列表
    fact_list = [all_p[i].text.rstrip('\n') for i in target_indexes]
    
    print(f'共选中{len(fact_list)}条内容')
    # 写入文件,每条占一行
    with open('out.txt', 'w', encoding='utf-8') as f:
        f.write('\n'.join(fact_list))
    
    # 读取验证
    with open('out.txt', 'r', encoding='utf-8') as file:
        content = file.read()
    print(content)
    return fact_list

# 调用示例:提取索引22、25、28、30的内容
getFacts([22,25,28,30])

内容的提问来源于stack exchange,提问作者noidss

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 17:45:03