You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Selenium+BeautifulSoup4获取JS动态加载的页面数据?

如何提取动态加载的网页库存数据?

我正在练习用Python抓取网页数据,目标网站是sandbox.oxylabs.io/products。目前已经用基础方法提取了大部分数据,但遇到一个动态加载的库存元素——这个元素会在初始页面加载完成后才出现,现有代码无法获取到它。

我尝试获取的库存状态元素,加载后显示为“Out of Stock”,但它并没有出现在print(b)的输出里。我试过用CSS选择器和.find()方法,也知道库存充足的元素类是in-stock,缺货的是out-of-stock,但因为元素没加载,用if语句设置变量时总是返回“Out of Stock”。

现有代码如下:

import pandas as pd
from bs4 import BeautifulSoup
from selenium import webdriver

options = webdriver.FirefoxOptions()
driver = webdriver.Firefox(options=options)
driver.get('https://sandbox.oxylabs.io/products')

results = []
other_results = []
status = []
content = driver.page_source
soup = BeautifulSoup(content, 'html.parser')

for element in soup.find_all(attrs={'class': 'product-card'}):
    name = element.find('h4')
    if name not in results:
        results.append(name.text)
        
for b in soup.find_all(attrs={'class': 'product-card'}):
# Note the use of 'attrs' to again select an element with the specified class.
    name2 = b.find(attrs={'class': 'price-wrapper'})
    # stock = b.find(attrs={'class': 'price-wrapper'}).find_next_sibling('p')
    # status.append(stock.text)
    other_results.append(name2.text)
    print(b)

df = pd.DataFrame({'Names': results, 'Prices': other_results, 'Stock': status})
df.to_csv('products.csv', index=False, encoding='utf-8')

解决方案:使用Selenium显式等待等待动态元素加载

问题核心是你调用driver.page_source的时机过早——页面还没完成动态内容的加载,库存元素尚未渲染到HTML中。需要用Selenium的显式等待机制,等待目标元素加载完成后再解析页面。

修改后的代码:

import pandas as pd
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.FirefoxOptions()
driver = webdriver.Firefox(options=options)
driver.get('https://sandbox.oxylabs.io/products')

# 显式等待:等待所有库存元素加载完成,最长等待10秒
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, '.product-card .in-stock, .product-card .out-of-stock')))

# 动态内容加载完成后,再获取页面源码
content = driver.page_source
soup = BeautifulSoup(content, 'html.parser')

results = []
other_results = []
status = []

# 单次遍历商品卡片,同时提取所有数据
for card in soup.find_all(attrs={'class': 'product-card'}):
    # 提取商品名称
    name = card.find('h4').text.strip()
    results.append(name)
    
    # 提取价格
    price = card.find(attrs={'class': 'price-wrapper'}).text.strip()
    other_results.append(price)
    
    # 提取库存状态
    stock_element = card.find(attrs={'class': ['in-stock', 'out-of-stock']})
    status.append(stock_element.text.strip() if stock_element else 'Unknown')

# 生成数据文件并关闭浏览器
df = pd.DataFrame({'Names': results, 'Prices': other_results, 'Stock': status})
df.to_csv('products.csv', index=False, encoding='utf-8')
driver.quit()

关键修改说明:

  1. 显式等待机制:通过WebDriverWait等待所有库存元素(in-stock/out-of-stock类)出现,确保动态内容完全加载。
  2. 合并遍历逻辑:无需两次遍历商品卡片,一次遍历即可完成名称、价格、库存的提取,提升效率。
  3. 精准定位库存元素:直接通过目标类名列表定位库存元素,避免依赖兄弟节点的不稳定定位方式。
  4. 异常兜底处理:如果库存元素未找到,默认填充“Unknown”,避免程序崩溃。

内容的提问来源于stack exchange,提问作者mattie malling

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 23:15:13