You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python抓取MUFAP网站利率:pandas.read_html无表报错求助

解决pandas.read_html无法抓取目标网站利率数据的问题

问题原因

pandas的read_html()方法仅能识别页面中标准HTML 标签包裹的表格数据。目标网站的利率数据大概率属于以下两种情况之一:

  1. 用
    等非表格标签配合CSS模拟表格结构;
  2. 数据通过JavaScript动态渲染,静态HTML中不存在完整的表格内容。
    这两种情况都会导致read_html()无法识别表格,抛出ValueError: No tables found错误。

解决方案

根据页面加载方式,提供两种可行的抓取方案:

方案1:静态页面抓取(requests + BeautifulSoup)

如果数据是静态渲染的,只是用非标准结构展示,可先获取页面源码,再用BeautifulSoup定位数据元素:

# 先安装依赖库
pip install requests beautifulsoup4 pandas
import requests
from bs4 import BeautifulSoup
import pandas as pd

# 请求页面(添加User-Agent模拟浏览器访问)
url = "https://www.mufap.com.pk/industry.php?tab=03"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
response = requests.get(url, headers=headers)
response.encoding = "utf-8"

# 解析页面源码
soup = BeautifulSoup(response.text, "html.parser")

# 定位目标表格(根据页面实际结构调整,示例为查找带有class的table标签)
target_table = soup.find("table", class_="table")

# 提取表头和数据行
if target_table:
    # 获取表头
    column_headers = [th.text.strip() for th in target_table.find_all("th")]
    # 获取数据行
    data_rows = []
    for row in target_table.find_all("tr")[1:]:  # 跳过表头行
        row_data = [td.text.strip() for td in row.find_all("td")]
        if row_data:  # 过滤空行
            data_rows.append(row_data)
    # 转为DataFrame
    df = pd.DataFrame(data_rows, columns=column_headers)
    print(df)
else:
    print("未找到目标表格,请检查页面结构")

方案2:动态页面抓取(Selenium)

如果数据是通过JavaScript动态加载的,静态HTML中没有完整表格,需要用Selenium模拟浏览器加载页面:

# 安装依赖库
pip install selenium pandas
# 需同时安装对应浏览器的驱动(比如ChromeDriver,版本需与浏览器匹配)
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import pandas as pd
import time

# 配置Chrome无头模式(不显示浏览器窗口)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--disable-gpu")

# 启动浏览器并访问页面
driver = webdriver.Chrome(options=chrome_options)
driver.get("https://www.mufap.com.pk/industry.php?tab=03")
time.sleep(3)  # 等待页面动态加载完成

# 用read_html提取页面中的表格
try:
    tables = pd.read_html(driver.page_source)
    # 假设第一个表格是目标数据,可根据实际情况调整索引
    df = tables[0]
    print(df)
except ValueError:
    print("页面加载后仍未找到表格,请检查页面结构")

# 关闭浏览器
driver.quit()

注意事项

  • 若方案1无法找到表格,说明数据是动态加载的,优先使用方案2;
  • 需定期检查目标网站的页面结构,若网站更新HTML结构,需调整代码中的元素定位规则;
  • 遵守网站的robots.txt协议,避免频繁请求给服务器造成压力。

内容的提问来源于stack exchange,提问作者Sumair

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 20:08:23