You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取全部h2、h3、p标签并存储为DataFrame

问题解答

是的,你需要使用find_all方法提取所有目标条目,你现有代码只能拿到单条数据的原因是:直接从所有条目的总容器columns is-multiline中提取标签,BeautifulSoup默认只会返回匹配到的第一个元素。

修正逻辑

  • 先定位所有条目的总容器,再调用find_all获取总容器下所有单条条目的子容器列表
  • 遍历每个子容器,分别提取对应字段的内容,存入临时列表
  • 最后将临时列表转换为DataFrame即可

修正后参考代码

from urllib.request import urlopen
from bs4 import BeautifulSoup
import re
import pandas as pd

# 替换为实际爬取的URL
url = "你的目标网站地址"
html = urlopen(url)
soup = BeautifulSoup(html, 'html.parser')

# 定位所有条目的总容器
total_container = soup.find(class_ = re.compile('columns is-multiline'))
# 替换为单条条目对应的class名称,可通过浏览器F12审查元素获取
item_containers = total_container.find_all(class_ = "单条条目class名")

data_list = []
for item in item_containers:
    # 增加异常捕获,跳过字段缺失的异常条目
    try:
        position = item.h2.text.strip()
        company = item.h3.text.strip()
        city_state = item.find_all('p')[-2].text.strip()
        data_list.append({
            "职位": position,
            "公司": company,
            "所在城市": city_state
        })
    except (AttributeError, IndexError):
        continue

# 转换为DataFrame
df = pd.DataFrame(data_list)
# 可按需导出为csv文件
# df.to_csv("爬取结果.csv", index=False, encoding="utf-8-sig")
print(df)

注意事项

  • 需先通过浏览器开发者工具确认单条条目的class属性,替换代码中对应占位内容
  • 如果不同条目的p标签数量不一致,不要使用[-2]这种索引方式,可通过p标签的class属性或者文本特征定位,提升代码稳定性

内容的提问来源于stack exchange,提问作者user16835025

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 01:18:02