使用BeautifulSoup findAll获取网页b标签返回空列表,本地文件正常
解决warframe.market页面无法抓取动态渲染的price类b标签问题
问题描述
我尝试抓取warframe.market某物品页面中所有带有price类的<b>标签,但使用BeautifulSoup的findAll方法返回了空列表。不过将页面保存为本地HTML文件后,执行相同代码却能正常获取标签。相关代码如下:
from bs4 import BeautifulSoup from urllib.request import Request, urlopen url = 'https://warframe.market/items/nami_skyla_prime_blueprint' req = Request(url, headers={'User-Agent': 'Mozilla/5.0'}) webpage = urlopen(req).read() soup = BeautifulSoup(webpage, 'html.parser') tags = soup.findAll('b') print(tags)
问题原因
核心问题是目标页面的内容由JavaScript动态渲染生成。你用urllib获取的只是服务器返回的原始HTML骨架,那些带有price类的<b>标签是页面加载完成后,通过前端JS动态插入到DOM中的,并没有包含在初始的HTTP响应里。而本地HTML文件是浏览器加载完成后的完整页面,自然能解析到目标标签。
解决方案
方案1:用Selenium模拟浏览器渲染
Selenium可以模拟完整的浏览器加载流程,等待JS执行完毕后再获取完整页面内容,从而拿到动态生成的标签。
先安装依赖:
pip install selenium
示例代码:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import time url = 'https://warframe.market/items/nami_skyla_prime_blueprint' # 配置无头模式(可选,无需打开可视化浏览器窗口) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--user-agent=Mozilla/5.0') driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 等待页面加载完成(可根据网络情况调整时长,或用显式等待替代固定sleep) time.sleep(3) page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 精准定位带price类的b标签 price_tags = soup.find_all('b', class_='price') for tag in price_tags: print(tag.get_text(strip=True)) driver.quit()
方案2:调用官方API(高效稳定推荐)
warframe.market提供了公开的API接口,直接调用API获取数据比爬取页面更可靠,还能避免动态渲染的限制。
示例代码:
import requests # 物品名称需转为下划线分隔的小写格式 item_name = 'nami_skyla_prime_blueprint' api_url = f'https://api.warframe.market/v1/items/{item_name}/orders' headers = { 'User-Agent': 'Mozilla/5.0', 'Accept': 'application/json' } response = requests.get(api_url, headers=headers) data = response.json() # 提取所有出售订单的价格信息 for order in data['payload']['orders']: if order['order_type'] == 'sell': print(f"售价: {order['platinum']} 白金")
内容的提问来源于stack exchange,提问作者Youtipie
相关产品推荐
相关产品推荐

