Python网页爬取求助:如何提取表格主体数据?
解决动态表格爬取问题
问题根源
这个页面的表格数据是JavaScript动态渲染的,直接用requests请求只能拿到静态HTML框架,实际表格数据是页面加载完成后通过JS从后端接口获取并插入的,所以你用BeautifulSoup只能提取到表头,拿不到真实数据。
另外你的代码还有几处语法错误:
- 重复导入了
pandas、requests、BeautifulSoup,可清理冗余代码 - 循环中使用了未定义的变量
team,应该改为PAX id_='passengerTrafficStat'是错误定位,表格内的<td>标签并没有这个id,且静态页面中本身就没有数据内容
解决方案1:直接请求数据接口(推荐)
通过浏览器开发者工具抓包,能找到页面加载表格数据的API接口,直接请求该接口可获取结构化JSON数据,无需解析HTML。
示例代码:
import requests import pandas as pd # 对应2023年8月的统计数据接口 api_url = "https://www.immd.gov.hk/eng/resources/stat/passenger-statistics-202308.json" response = requests.get(api_url) data = response.json() # 提取表格行数据与表头 rows = data['passengerTrafficStat']['rows'] headers = [header['label'] for header in data['passengerTrafficStat']['headers']] # 转为DataFrame并处理 df = pd.DataFrame(rows, columns=headers) print(df) # 保存为Excel文件 df.to_excel('passenger_statistics_202308.xlsx', index=False)
解决方案2:用Selenium模拟浏览器渲染(适合新手)
如果不想查找接口,可使用Selenium启动真实浏览器,等待页面完全加载后再提取表格数据。
先安装依赖:
pip install selenium
示例代码:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd # 初始化Chrome浏览器(需提前下载对应版本的ChromeDriver,配置好路径) service = Service('chromedriver.exe') # Windows系统;Linux/Mac请修改为对应路径 driver = webdriver.Chrome(service=service) url = 'https://www.immd.gov.hk/eng/facts/passenger-statistics.html?d=20230830' driver.get(url) # 等待表格加载完成 wait = WebDriverWait(driver, 10) table = wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'table-passengerTrafficStat'))) # 提取表头与表格数据 headers = [th.text.strip() for th in table.find_elements(By.TAG_NAME, 'th')] data_rows = [] for row in table.find_elements(By.TAG_NAME, 'tr'): cols = row.find_elements(By.TAG_NAME, 'td') if cols: data_rows.append([col.text.strip() for col in cols]) # 转为DataFrame并输出 df = pd.DataFrame(data_rows, columns=headers) print(df) # 关闭浏览器 driver.quit() # 保存为Excel文件 df.to_excel('passenger_statistics_202308.xlsx', index=False)
补充说明
- 接口URL规律:若要爬取其他月份数据,只需将接口URL中的
202308改为对应年月(如202309对应2023年9月)即可 - Selenium方案需注意浏览器驱动与浏览器版本匹配,避免兼容性问题
内容的提问来源于stack exchange,提问作者Nicole
相关产品推荐
相关产品推荐

