使用lxml爬取TCGPlayer时Xpath返回空数组问题求助
解决TCGPlayer卡牌价格爬取返回空数组的问题
问题原因
- 绝对XPath不可靠:你使用的是从根节点开始的绝对路径,网站结构只要有微小调整(比如新增/删除一个容器节点),这个路径就会直接失效,无法定位到目标元素。
- 动态内容渲染:TCGPlayer的部分页面内容是通过JavaScript动态加载的,
requests获取的只是原始HTML源码,不包含JS渲染后的价格数据。 - 反爬检测:网站可能识别出你的请求来自脚本而非浏览器,返回的内容不完整或被拦截。
解决方案
1. 稳定选择器+请求头模拟浏览器
先给请求添加浏览器请求头规避反爬,同时改用基于元素属性(如data-testid、class)的相对选择器,替代脆弱的绝对XPath。
修改后的代码:
from lxml import html import requests def clean_text(element): all_text = element.text_content() cleaned = ' '.join(all_text.split()) return cleaned # 添加浏览器请求头,模拟真实访问 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } url = "https://www.tcgplayer.com/product/231462/pokemon-first-partner-pack-pikachu?Language=English" page = requests.get(url, headers=headers) tree = html.fromstring(page.content) # 用相对XPath定位价格元素,依赖页面稳定的属性标识 price_elements = tree.xpath('//section[@data-testid="product-pricing"]//span[contains(@class, "price-point__data")]') if price_elements: price = clean_text(price_elements[0]) print(f"价格:{price}") else: print("未找到价格元素,可能内容是动态加载的")
2. 用requests-html渲染动态内容
如果上述方法仍无法获取数据,说明价格是JS动态生成的,可以使用requests-html库(内置轻量Chromium内核,能自动渲染JS,无需手动打开多个浏览器窗口)。
先安装库:
pip install requests-html
示例代码:
from requests_html import HTMLSession def clean_text(text): return ' '.join(text.split()) session = HTMLSession() headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } url = "https://www.tcgplayer.com/product/231462/pokemon-first-partner-pack-pikachu?Language=English" r = session.get(url, headers=headers) # 触发JS渲染,加载动态内容 r.html.render() # 定位价格元素 price_element = r.html.find('section[data-testid="product-pricing"] span.price-point__data', first=True) if price_element: price = clean_text(price_element.text) print(f"价格:{price}") else: print("未找到价格元素")
3. 优先使用官方API
TCGPlayer提供官方API(需注册获取API密钥),直接调用API查询价格是最稳定、合规的方式,能彻底避免页面爬取的反爬和结构变更问题。你可以在其开发者平台注册后,通过产品ID直接获取价格数据。
注意事项
- 爬取时添加请求间隔(如
time.sleep(1)),避免频繁请求触发反爬机制。 - 遵守网站
robots.txt规则,确保爬取行为合规。
内容的提问来源于stack exchange,提问作者Mitchell Prior
相关产品推荐
相关产品推荐

