You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas.read_html抓取网页表格仅获取表头无数据该如何解决

pandas.read_html仅获取表头无表格内容的解决方案

问题原因

  • pandas发起请求时未携带浏览器标识,被目标网站的反爬策略拦截,仅返回了表格的表头结构,未返回实际课程数据
  • 部分场景下是因为表格内容由JavaScript动态渲染,静态HTML源码中不存在实际行数据,read_html无法直接解析到

解决步骤

方案1:添加请求头模拟浏览器请求(优先测试,对应当前场景90%可解决)

  1. 先安装依赖库requests
pip install requests
  1. 修改代码逻辑:先用requests带请求头获取页面源码,再传给pandas解析
import pandas as pd
import requests

term_codes = {"fall":"10", "spring":"20", "summer":"30"}

# year must be last number in school year: 2021-2022 so we pick 2022
year = "2022"
department = "CSCI"
term_code = year + term_codes["fall"]
url = "https://courselist.wm.edu/courselist/courseinfo/searchresults?term_code=" + term_code + "&term_subj=" + department + "&attr=0&attr2=0&levl=0&status=0&ptrm=0&search=Search"

# 模拟浏览器请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

def findCourseTable():
    # 先发起请求获取页面内容
    resp = requests.get(url, headers=headers)
    resp.raise_for_status() # 检查请求是否成功
    dfs = pd.read_html(resp.text)
    # 注意:该网站返回的表格列表中,实际课程数据在索引为1的表,索引0是表头占位
    df = dfs[1]
    print(df)
    df.to_csv(r'courses.csv', index=False)

if __name__ == "__main__":
    findCourseTable()

方案2:处理动态渲染的表格

如果方案1返回的源码中还是没有实际表格数据,说明内容是JS动态加载的,使用无头浏览器渲染页面后再提取:

  1. 安装selenium相关依赖
pip install selenium webdriver-manager
  1. 对应代码示例
import pandas as pd
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
import time

term_codes = {"fall":"10", "spring":"20", "summer":"30"}
year = "2022"
department = "CSCI"
term_code = year + term_codes["fall"]
url = "https://courselist.wm.edu/courselist/courseinfo/searchresults?term_code=" + term_code + "&term_subj=" + department + "&attr=0&attr2=0&levl=0&status=0&ptrm=0&search=Search"

def findCourseTable():
    # 启动无头Chrome
    options = webdriver.ChromeOptions()
    options.add_argument("--headless=new")
    options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")
    driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
    driver.get(url)
    time.sleep(2) # 等待页面渲染完成
    # 提取页面源码解析
    dfs = pd.read_html(driver.page_source)
    df = dfs[1]
    print(df)
    df.to_csv(r'courses.csv', index=False)
    driver.quit()

if __name__ == "__main__":
    findCourseTable()

注意事项

  • 拿到返回的表格列表后,可以先打印len(dfs)确认有几个表格,再逐个查看内容,匹配你需要的数据集
  • 请求头的User-Agent可以更换为你自己本地浏览器的标识,避免被拦截

内容的提问来源于stack exchange,提问作者jbcallv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 10:09:03