You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何BeautifulSoup无法识别目标网站中的表格?

爬取香港赛马会骑师排名表格失败问题解决

问题描述

尝试使用BeautifulSoup和Pandas爬取香港赛马会骑师排名页面的表格并保存为Excel,但运行代码时输出Table not found,但网页中明确存在目标表格。原代码如下:

from bs4 import BeautifulSoup
import pandas as pd

# Send a GET request to the URL
url = "https://racing.hkjc.com/racing/information/English/Jockey/JockeyRanking.aspx"
response = requests.get(url)

# Parse the HTML content using BeautifulSoup
soup = BeautifulSoup(response.content, "html.parser")

# Find the table element and extract the data
table = soup.find("table", {"class": "table_bd "})
if table is None:
    print("Table not found.")
else:
    df = pd.read_html(str(table))[0]
    # Save the data to an Excel spreadsheet
    df.to_excel("hkjc.xlsx", index=False)

问题原因

  • 依赖库缺失:代码中调用requests.get()但未导入requests库,若未手动补充导入会直接报错;即使临时补全,后续仍可能因其他问题找不到表格
  • CSS类名匹配错误:原代码中查找表格时使用的class值为"table_bd "(末尾带空格),但网页中目标表格的实际class属性值是"table_bd"(无末尾空格),BeautifulSoup的精确字典匹配会因这个空格导致匹配失败
  • 反爬拦截导致内容不完整:网站可能会检查请求头,缺少User-Agent的请求会被判定为非浏览器访问,返回的HTML内容不完整,从而无法找到目标表格

修复方案

方案一:修正原代码逻辑

修复后的完整代码如下:

from bs4 import BeautifulSoup
import pandas as pd
import requests  # 补充导入缺失的requests库

url = "https://racing.hkjc.com/racing/information/English/Jockey/JockeyRanking.aspx"
# 添加请求头模拟浏览器访问,避免反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 发送带请求头的GET请求
response = requests.get(url, headers=headers)
response.encoding = "utf-8"  # 确保HTML内容编码正确

# 解析HTML
soup = BeautifulSoup(response.content, "html.parser")
# 修正class匹配,使用class_参数并去掉末尾空格
table = soup.find("table", class_="table_bd")

if table is None:
    print("Table not found.")
else:
    df = pd.read_html(str(table))[0]
    df.to_excel("hkjc_jockey_ranking.xlsx", index=False)
    print("数据已成功保存至hkjc_jockey_ranking.xlsx")

方案二:直接使用Pandas爬取(更简洁)

Pandas的read_html方法可以直接处理网页表格,无需BeautifulSoup,代码更简洁:

import pandas as pd
import requests

url = "https://racing.hkjc.com/racing/information/English/Jockey/JockeyRanking.aspx"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

response = requests.get(url, headers=headers)
# 直接读取页面中的所有表格,目标表格为第一个
dfs = pd.read_html(response.text)
dfs[0].to_excel("hkjc_jockey_ranking.xlsx", index=False)
print("数据保存成功")

内容的提问来源于stack exchange,提问作者NNBananas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 19:07:40