使用Beautiful Soup爬取ERCOT网站表格时出现AttributeError问题求助
问题根因
出现该错误是因为table = soup.find("table", attrs={"class": "tableStyle"})这行代码没有找到匹配的表格,返回了None值,后续调用.tbody属性自然会触发AttributeError。
定位步骤
- 先验证请求有效性:打印
requests.get(URL).status_code确认返回状态码为200,同时打印返回的页面内容前500字符,确认拿到的是目标页面而非反爬拦截页。绝大多数情况下该错误是因为未加请求头,爬虫请求被网站拦截导致返回内容不包含目标表格。 - 验证表格class匹配性:在返回的页面文本中搜索
tableStyle关键词,确认class名称没有拼写错误、大小写匹配。 - 验证DOM结构:如果表格存在,检查原始HTML是否真的包含
<tbody>标签,部分浏览器会自动补全tbody标签,但原始页面源码中可能没有该层级,直接调用table.tbody也会返回None。
修复方案
方案1:手动解析BeautifulSoup对象
import requests from bs4 import BeautifulSoup import pandas as pd # 新增请求头模拟浏览器访问,绕过基础反爬策略 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } URL = "http://www.ercot.com/content/cdr/html/real_time_spp" response = requests.get(URL, headers=headers) # 请求异常时直接抛出错误,避免后续解析无效内容 response.raise_for_status() soup = BeautifulSoup(response.text, "lxml") table = soup.find("table", attrs={"class": "tableStyle"}) if not table: raise ValueError("未找到匹配的表格,请检查页面内容或class配置") # 直接从table节点下查找所有tr,跳过tbody层级,兼容无tbody的场景 table_rows = table.find_all("tr") # 提取表头 header = [th.get_text(strip=True) for th in table_rows[0].find_all("th")] # 提取行内容 data = [] for tr in table_rows[1:]: row = [td.get_text(strip=True) for td in tr.find_all("td")] if row: data.append(row) # 转DataFrame df = pd.DataFrame(data, columns=header)
方案2:直接用pandas读取(更简便)
pandas内置的read_html方法可以自动识别页面内的表格,无需手动解析DOM:
import pandas as pd headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } # storage_options参数用于传递请求头配置 df_list = pd.read_html("http://www.ercot.com/content/cdr/html/real_time_spp", storage_options=headers) # 取第一个匹配的表格即可 df = df_list[0]
内容的提问来源于stack exchange,提问作者AJK
相关产品推荐
相关产品推荐

