使用BeautifulSoup抓取网页表格报错,求正确实现方法及标签排查
解决BeautifulSoup抓取ABS网站表格数据的报错问题
问题背景
我正在开发Python脚本,用BeautifulSoup抓取ABS网站的住宅房价指数表格数据,目标页面是:https://www.abs.gov.au/statistics/economy/price-indexes-and-inflation/residential-property-price-indexes-eight-capital-cities/latest-release,目标表格见截图(https://i.sstatic.net/hryHf.png)。
尝试的代码:
import requests from bs4 import BeautifulSoup res=requests.get("https://www.abs.gov.au/statistics/economy/price-indexes-and-inflation/residential-property-price-indexes-eight-capital-cities/latest-release") soup=BeautifulSoup(res.text,"html.parser") table=soup.find_all('table', class_='chart-data-table has-chart responsive-enabled double-headers') print(table.get_text())
运行后报错:
Traceback (most recent call last): File "/Users/ryanngan/PycharmProjects/Webscraping/seek.py", line 6, in <module> print(table.get_text()) File "/Users/ryanngan/PycharmProjects/Webscraping/venv/lib/python3.9/site-packages/bs4/element.py", line 2289, in __getattr__ raise AttributeError( AttributeError: ResultSet object has no attribute 'get_text'. You're probably treating a list of elements like a single element. Did you call find_all() when you meant to call find()?
报错原因
soup.find_all()返回的是ResultSet(结果列表),并非单个元素,因此无法直接调用.get_text()方法。另外你指定的表格类名是正确的,但需要确认页面中目标表格的数量。
修正方案
1. 针对单个目标表格
用soup.find()替代soup.find_all(),直接获取单个表格元素:
import requests from bs4 import BeautifulSoup res = requests.get("https://www.abs.gov.au/statistics/economy/price-indexes-and-inflation/residential-property-price-indexes-eight-capital-cities/latest-release") soup = BeautifulSoup(res.text, "html.parser") # 获取单个目标表格 table = soup.find('table', class_='chart-data-table has-chart responsive-enabled double-headers') if table: # 提取并格式化表格文本 print(table.get_text(strip=True, separator='\n')) else: print("未找到目标表格")
2. 针对多个同类型表格
如果页面存在多个符合条件的表格,遍历结果列表逐个处理:
import requests from bs4 import BeautifulSoup res = requests.get("https://www.abs.gov.au/statistics/economy/price-indexes-and-inflation/residential-property-price-indexes-eight-capital-cities/latest-release") soup = BeautifulSoup(res.text, "html.parser") # 获取所有符合条件的表格 tables = soup.find_all('table', class_='chart-data-table has-chart responsive-enabled double-headers') for idx, table in enumerate(tables): print(f"第{idx+1}个表格数据:") # 遍历表格行,提取单元格文本 for row in table.find_all('tr'): cells = [cell.get_text(strip=True) for cell in row.find_all(['th', 'td'])] print('\t'.join(cells))
内容的提问来源于stack exchange,提问作者ryantl
相关产品推荐
相关产品推荐

