使用BeautifulSoup通过ID获取表格数据返回None的解决办法
解决fbref表格爬取返回None的问题
问题原因
fbref这类体育数据网站存在反爬机制,直接用requests.get()获取的页面内容要么不包含动态渲染的stats_standard表格,要么请求被识别为非浏览器请求,导致返回的HTML里没有目标表格。
解决方案一:添加请求头模拟浏览器
给请求添加浏览器标识的请求头,让服务器判定为正常浏览器访问:
import requests from bs4 import BeautifulSoup url = 'https://fbref.com/en/comps/9/2023-2024/stats/2023-2024-Premier-League-Stats' # 模拟Chrome浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } page = requests.get(url, headers=headers) soup = BeautifulSoup(page.text, 'html.parser') table = soup.find('table', id='stats_standard') print(table) # 此时应能获取到表格对象
解决方案二:用Selenium获取动态渲染页面
如果添加请求头后仍无法获取表格,说明表格是通过JavaScript动态加载的,需要模拟浏览器完整渲染页面:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup url = 'https://fbref.com/en/comps/9/2023-2024/stats/2023-2024-Premier-League-Stats' # 配置Chrome无头模式(可选,无需打开浏览器窗口) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--disable-gpu') driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 获取完全渲染后的页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') table = soup.find('table', id='stats_standard') print(table) driver.quit()
提取表格数据
拿到表格对象后,可解析表头和内容:
if table: # 提取表头 headers = [th.text.strip() for th in table.find('thead').find_all('th')] # 提取表格行数据 rows = [] for tr in table.find('tbody').find_all('tr'): row = [td.text.strip() for td in tr.find_all('td')] if row: # 跳过空行 rows.append(row) # 输出结果 print(headers) print(rows)
内容的提问来源于stack exchange,提问作者i love whisky
相关产品推荐
相关产品推荐

