You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取求助:如何提取表格主体数据?

解决动态表格爬取问题

问题根源

这个页面的表格数据是JavaScript动态渲染的,直接用requests请求只能拿到静态HTML框架,实际表格数据是页面加载完成后通过JS从后端接口获取并插入的,所以你用BeautifulSoup只能提取到表头,拿不到真实数据。

另外你的代码还有几处语法错误:

  • 重复导入了pandas、requests、BeautifulSoup,可清理冗余代码
  • 循环中使用了未定义的变量team,应该改为PAX
  • id_='passengerTrafficStat'是错误定位,表格内的<td>标签并没有这个id,且静态页面中本身就没有数据内容

解决方案1:直接请求数据接口(推荐)

通过浏览器开发者工具抓包,能找到页面加载表格数据的API接口,直接请求该接口可获取结构化JSON数据,无需解析HTML。

示例代码:

import requests
import pandas as pd

# 对应2023年8月的统计数据接口
api_url = "https://www.immd.gov.hk/eng/resources/stat/passenger-statistics-202308.json"
response = requests.get(api_url)
data = response.json()

# 提取表格行数据与表头
rows = data['passengerTrafficStat']['rows']
headers = [header['label'] for header in data['passengerTrafficStat']['headers']]

# 转为DataFrame并处理
df = pd.DataFrame(rows, columns=headers)
print(df)
# 保存为Excel文件
df.to_excel('passenger_statistics_202308.xlsx', index=False)

解决方案2:用Selenium模拟浏览器渲染(适合新手)

如果不想查找接口,可使用Selenium启动真实浏览器,等待页面完全加载后再提取表格数据。

先安装依赖:

pip install selenium

示例代码:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

# 初始化Chrome浏览器(需提前下载对应版本的ChromeDriver,配置好路径)
service = Service('chromedriver.exe')  # Windows系统;Linux/Mac请修改为对应路径
driver = webdriver.Chrome(service=service)

url = 'https://www.immd.gov.hk/eng/facts/passenger-statistics.html?d=20230830'
driver.get(url)

# 等待表格加载完成
wait = WebDriverWait(driver, 10)
table = wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'table-passengerTrafficStat')))

# 提取表头与表格数据
headers = [th.text.strip() for th in table.find_elements(By.TAG_NAME, 'th')]
data_rows = []
for row in table.find_elements(By.TAG_NAME, 'tr'):
    cols = row.find_elements(By.TAG_NAME, 'td')
    if cols:
        data_rows.append([col.text.strip() for col in cols])

# 转为DataFrame并输出
df = pd.DataFrame(data_rows, columns=headers)
print(df)

# 关闭浏览器
driver.quit()

# 保存为Excel文件
df.to_excel('passenger_statistics_202308.xlsx', index=False)

补充说明

  • 接口URL规律:若要爬取其他月份数据,只需将接口URL中的202308改为对应年月(如202309对应2023年9月)即可
  • Selenium方案需注意浏览器驱动与浏览器版本匹配,避免兼容性问题

内容的提问来源于stack exchange,提问作者Nicole

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 22:57:42