使用pandas.read_html抓取网页表格仅获取表头无数据该如何解决
pandas.read_html仅获取表头无表格内容的解决方案
问题原因
- pandas发起请求时未携带浏览器标识,被目标网站的反爬策略拦截,仅返回了表格的表头结构,未返回实际课程数据
- 部分场景下是因为表格内容由JavaScript动态渲染,静态HTML源码中不存在实际行数据,read_html无法直接解析到
解决步骤
方案1:添加请求头模拟浏览器请求(优先测试,对应当前场景90%可解决)
- 先安装依赖库
requests
pip install requests
- 修改代码逻辑:先用requests带请求头获取页面源码,再传给pandas解析
import pandas as pd import requests term_codes = {"fall":"10", "spring":"20", "summer":"30"} # year must be last number in school year: 2021-2022 so we pick 2022 year = "2022" department = "CSCI" term_code = year + term_codes["fall"] url = "https://courselist.wm.edu/courselist/courseinfo/searchresults?term_code=" + term_code + "&term_subj=" + department + "&attr=0&attr2=0&levl=0&status=0&ptrm=0&search=Search" # 模拟浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } def findCourseTable(): # 先发起请求获取页面内容 resp = requests.get(url, headers=headers) resp.raise_for_status() # 检查请求是否成功 dfs = pd.read_html(resp.text) # 注意:该网站返回的表格列表中,实际课程数据在索引为1的表,索引0是表头占位 df = dfs[1] print(df) df.to_csv(r'courses.csv', index=False) if __name__ == "__main__": findCourseTable()
方案2:处理动态渲染的表格
如果方案1返回的源码中还是没有实际表格数据,说明内容是JS动态加载的,使用无头浏览器渲染页面后再提取:
- 安装selenium相关依赖
pip install selenium webdriver-manager
- 对应代码示例
import pandas as pd from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager import time term_codes = {"fall":"10", "spring":"20", "summer":"30"} year = "2022" department = "CSCI" term_code = year + term_codes["fall"] url = "https://courselist.wm.edu/courselist/courseinfo/searchresults?term_code=" + term_code + "&term_subj=" + department + "&attr=0&attr2=0&levl=0&status=0&ptrm=0&search=Search" def findCourseTable(): # 启动无头Chrome options = webdriver.ChromeOptions() options.add_argument("--headless=new") options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36") driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) driver.get(url) time.sleep(2) # 等待页面渲染完成 # 提取页面源码解析 dfs = pd.read_html(driver.page_source) df = dfs[1] print(df) df.to_csv(r'courses.csv', index=False) driver.quit() if __name__ == "__main__": findCourseTable()
注意事项
- 拿到返回的表格列表后,可以先打印
len(dfs)确认有几个表格,再逐个查看内容,匹配你需要的数据集 - 请求头的User-Agent可以更换为你自己本地浏览器的标识,避免被拦截
内容的提问来源于stack exchange,提问作者jbcallv
相关产品推荐
相关产品推荐

