You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在网页爬取中选择单个标签填充DataFrame?Python爬表疑问

如何从多个匹配的表格标签中选择单个来填充DataFrame?

嘿,我来帮你搞定这个问题!你现在用find_all抓到了一堆class为table m-b-0的表格,想要挑其中一个转成Pandas DataFrame对吧?这里有几个实用的方法:

方法1:直接取第一个匹配的表格

如果目标表格是页面里第一个符合class的表格,你可以直接用find()代替find_all()——它会返回第一个匹配到的元素,省得再处理列表索引:

from bs4 import BeautifulSoup
import requests
import pandas as pd

res = requests.get("http://portal.stf.jus.br/processos/listarPartes.asp?termo=paulo%20salim%20maluf")
soup = BeautifulSoup(res.text, "lxml")

# 直接获取第一个匹配的表格
target_table = soup.find('table', {'class': 'table m-b-0'})

# 用Pandas解析表格成DataFrame
df = pd.read_html(str(target_table))[0]
print(df.head())

要是你已经用了find_all()存到parts里,那直接取parts[0]也能拿到第一个表格。

方法2:取特定位置的表格

如果你知道目标表格是第N个(注意Python列表是从0开始计数的),直接用索引就行:

# 比如取第二个表格
target_table = parts[1]
df = pd.read_html(str(target_table))[0]

方法3:通过表格内的独特内容精准定位

如果页面里的表格结构类似,没法靠位置区分,你可以找表格里的独特标识(比如某个表头文本、特定内容)来定位:

# 假设目标表格里有个表头叫"Processo",先找到这个表头元素,再向上找父表格
header = soup.find('th', string='Processo')
target_table = header.find_parent('table', {'class': 'table m-b-0'})

df = pd.read_html(str(target_table))[0]

这里要注意,pd.read_html()接受HTML字符串作为输入,所以得把BeautifulSoup的表格对象转成字符串str(target_table),它返回的是一个DataFrame列表,取索引0就是我们要的表格数据啦。

内容的提问来源于stack exchange,提问作者Reinaldo Chaves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:12:42