为何BeautifulSoup无法识别目标网站中的表格?
爬取香港赛马会骑师排名表格失败问题解决
问题描述
尝试使用BeautifulSoup和Pandas爬取香港赛马会骑师排名页面的表格并保存为Excel,但运行代码时输出Table not found,但网页中明确存在目标表格。原代码如下:
from bs4 import BeautifulSoup import pandas as pd # Send a GET request to the URL url = "https://racing.hkjc.com/racing/information/English/Jockey/JockeyRanking.aspx" response = requests.get(url) # Parse the HTML content using BeautifulSoup soup = BeautifulSoup(response.content, "html.parser") # Find the table element and extract the data table = soup.find("table", {"class": "table_bd "}) if table is None: print("Table not found.") else: df = pd.read_html(str(table))[0] # Save the data to an Excel spreadsheet df.to_excel("hkjc.xlsx", index=False)
问题原因
- 依赖库缺失:代码中调用
requests.get()但未导入requests库,若未手动补充导入会直接报错;即使临时补全,后续仍可能因其他问题找不到表格 - CSS类名匹配错误:原代码中查找表格时使用的
class值为"table_bd "(末尾带空格),但网页中目标表格的实际class属性值是"table_bd"(无末尾空格),BeautifulSoup的精确字典匹配会因这个空格导致匹配失败 - 反爬拦截导致内容不完整:网站可能会检查请求头,缺少
User-Agent的请求会被判定为非浏览器访问,返回的HTML内容不完整,从而无法找到目标表格
修复方案
方案一:修正原代码逻辑
修复后的完整代码如下:
from bs4 import BeautifulSoup import pandas as pd import requests # 补充导入缺失的requests库 url = "https://racing.hkjc.com/racing/information/English/Jockey/JockeyRanking.aspx" # 添加请求头模拟浏览器访问,避免反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 发送带请求头的GET请求 response = requests.get(url, headers=headers) response.encoding = "utf-8" # 确保HTML内容编码正确 # 解析HTML soup = BeautifulSoup(response.content, "html.parser") # 修正class匹配,使用class_参数并去掉末尾空格 table = soup.find("table", class_="table_bd") if table is None: print("Table not found.") else: df = pd.read_html(str(table))[0] df.to_excel("hkjc_jockey_ranking.xlsx", index=False) print("数据已成功保存至hkjc_jockey_ranking.xlsx")
方案二:直接使用Pandas爬取(更简洁)
Pandas的read_html方法可以直接处理网页表格,无需BeautifulSoup,代码更简洁:
import pandas as pd import requests url = "https://racing.hkjc.com/racing/information/English/Jockey/JockeyRanking.aspx" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) # 直接读取页面中的所有表格,目标表格为第一个 dfs = pd.read_html(response.text) dfs[0].to_excel("hkjc_jockey_ranking.xlsx", index=False) print("数据保存成功")
内容的提问来源于stack exchange,提问作者NNBananas
相关产品推荐
相关产品推荐

