Python 3中BeautifulSoup定位页面表格失败的问题排查
无法通过requests+BeautifulSoup获取目标页面的价格历史表格
原因分析
- 动态渲染导致静态HTML无目标内容:目标价格历史表格并非页面初始HTML的一部分,而是浏览器加载页面后,通过JavaScript调用后端接口获取数据再动态生成的。
requests仅能获取页面初始静态源码,自然找不到该表格。 - 选择器语法错误:原选择器
price-history-chart > div > div:nth-child(1) > div > div > table缺少ID选择器前缀#,正确写法应为#price-history-chart > div > div:nth-child(1) > div > div > table。但即使修正选择器,因静态源码中无目标表格,仍无法定位。
解决方案
方案1:直接调用后端API(高效推荐)
通过浏览器开发者工具抓包,找到页面加载价格历史数据的API接口,直接用requests请求接口获取数据,无需依赖浏览器渲染。
示例代码:
import requests import pandas as pd # 替换为实际抓包得到的价格历史API地址 API_URL = "https://www.dekudeals.com/items/buy-the-game-i-have-a-gun-sheesh-man-digital-deluxe-mega-chad-edition/price-history?format=digital" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } # 请求API获取价格数据 response = requests.get(API_URL, headers=headers) price_history = response.json() # 转换为DataFrame格式 df = pd.DataFrame(price_history) print(df)
方案2:优化Selenium代码(浏览器渲染方案)
若必须使用浏览器渲染,可优化代码,直接等待目标表格加载完成后再获取,避免无效等待。
示例代码:
import pandas as pd from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from selenium import webdriver from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.support import expected_conditions as EC service = Service(executable_path=".../Driver/chromedriver") driver = webdriver.Chrome(service=service) target_url = "https://www.dekudeals.com/items/buy-the-game-i-have-a-gun-sheesh-man-digital-deluxe-mega-chad-edition?format=digital" driver.get(target_url) # 等待目标表格加载完成 wait = WebDriverWait(driver, 25) target_table = wait.until( EC.presence_of_element_located( (By.CSS_SELECTOR, "#price-history-chart > div > div:nth-child(1) > div > div > table") ) ) # 直接读取表格数据 price_df = pd.read_html(target_table.get_attribute('outerHTML'))[0] print(price_df) driver.quit()
内容的提问来源于stack exchange,提问作者Łukasz Zedler
相关产品推荐
相关产品推荐

