BeautifulSoup爬取表格报ResultSet无find_all属性错误排查
报错根因
find_all()返回的是ResultSet对象,本质是匹配到的标签元素列表,列表本身没有find()/find_all()这类单标签元素的方法,你直接对soup.find_all('table')返回的列表调用find_all('tbody')必然触发这个报错。
你代码另外还有3个问题:
- 页面存在多个用于布局的无关table,直接查找所有table会匹配到非目标内容,增加不必要的遍历和报错概率
- BeautifulSoup的属性参数写法错误:只有
class这类和Python保留字冲突的属性才需要加下划线(即class_),普通属性比如align直接写align="left"即可,不需要下划线 - 表格包含表头行,这类行没有
align=left的td节点,直接取.text会触发空指针报错,需要加判空逻辑
修正代码
import requests from bs4 import BeautifulSoup url = 'https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfpcd/classification.cfm?start_search=1&submission_type_id=&devicename=&productcode=&deviceclass=&thirdparty=&panel=®ulationnumber=&implant_flag=&life_sustain_support_flag=&summary_malfunction_reporting=&sortcolumn=deviceclassdesc&pagenum=10' # 加请求头模拟浏览器,避免被站点反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" } r = requests.get(url, headers=headers) r.encoding = r.apparent_encoding soup = BeautifulSoup(r.text, 'html.parser') # 直接定位宽度95%的产品列表目标表格,不需要遍历所有table product_table = soup.find('table', width="95%") # 取tbody下所有行 rows = product_table.find('tbody').find_all('tr') for row in rows: # 跳过无目标td的表头/空行 product_td = row.find('td', align="left") if not product_td: continue # 提取产品编码 product_code = product_td.text.strip() # 提取详情页相对链接 detail_relative_url = product_td.find('a')['href'] if product_td.find('a') else '' # 提取当前行其余列内容 row_other_content = [col.text.strip() for col in row.find_all('td')[1:]] print(f"产品编码:{product_code}") print(f"详情页相对路径:{detail_relative_url}") print(f"行其余字段:{row_other_content}\n")
后续爬取注意事项
- 拿到的详情页链接是相对路径,需要拼接前缀
https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfpcd/组成完整URL后再发请求 - 批量爬取时建议加1-3秒的随机延时,不要高频请求,避免被站点封禁IP
- 如果后续爬取量大,可以提前配置代理池降低被封概率
内容的提问来源于stack exchange,提问作者user15267842
相关产品推荐
相关产品推荐

