如何在网页爬取中选择单个标签填充DataFrame?Python爬表疑问
如何从多个匹配的表格标签中选择单个来填充DataFrame?
嘿,我来帮你搞定这个问题!你现在用find_all抓到了一堆class为table m-b-0的表格,想要挑其中一个转成Pandas DataFrame对吧?这里有几个实用的方法:
方法1:直接取第一个匹配的表格
如果目标表格是页面里第一个符合class的表格,你可以直接用find()代替find_all()——它会返回第一个匹配到的元素,省得再处理列表索引:
from bs4 import BeautifulSoup import requests import pandas as pd res = requests.get("http://portal.stf.jus.br/processos/listarPartes.asp?termo=paulo%20salim%20maluf") soup = BeautifulSoup(res.text, "lxml") # 直接获取第一个匹配的表格 target_table = soup.find('table', {'class': 'table m-b-0'}) # 用Pandas解析表格成DataFrame df = pd.read_html(str(target_table))[0] print(df.head())
要是你已经用了find_all()存到parts里,那直接取parts[0]也能拿到第一个表格。
方法2:取特定位置的表格
如果你知道目标表格是第N个(注意Python列表是从0开始计数的),直接用索引就行:
# 比如取第二个表格 target_table = parts[1] df = pd.read_html(str(target_table))[0]
方法3:通过表格内的独特内容精准定位
如果页面里的表格结构类似,没法靠位置区分,你可以找表格里的独特标识(比如某个表头文本、特定内容)来定位:
# 假设目标表格里有个表头叫"Processo",先找到这个表头元素,再向上找父表格 header = soup.find('th', string='Processo') target_table = header.find_parent('table', {'class': 'table m-b-0'}) df = pd.read_html(str(target_table))[0]
这里要注意,pd.read_html()接受HTML字符串作为输入,所以得把BeautifulSoup的表格对象转成字符串str(target_table),它返回的是一个DataFrame列表,取索引0就是我们要的表格数据啦。
内容的提问来源于stack exchange,提问作者Reinaldo Chaves
相关产品推荐
相关产品推荐

