You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python lxml requests循环xpath查询仅最后一条结果存入列表bug修复

问题根源

代码存在4个直接导致运行异常的问题:

  • 读取txt文件内链接时未去除行尾换行符、首尾空白字符,传入requests的URL包含非法字符,多数请求实际无法命中正确的卡牌详情页
  • lxml库的xpath()方法不存在返回None的情况:匹配成功时返回元素/文本列表,匹配失败时返回空列表[],原有逻辑中if code.xpath(xpathPrice) == None的判断永远不会触发,价格匹配逻辑完全失效
  • 未配置合法请求头,Cardmarket站点有基础反爬规则,默认的requests请求标识会被拦截,返回反爬验证页而非目标卡牌内容,自然无法匹配到预设xpath
  • 原有逻辑重复调用xpath执行相同查询,既浪费性能,也容易在页面结构波动时出现取值不一致的问题
修复后可运行代码
from lxml import html
import requests
import time

price_list = []
name_list = []
# 配置浏览器请求头,绕过基础反爬校验
request_headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
}
# 注意:绝对路径xpath适配性极差,页面小幅改版就会失效,建议后续替换为基于属性、相邻文本的相对定位写法
xpath_price_primary = '/html/body/main/div[4]/section[2]/div/div[2]/div[1]/div/div[1]/div/div[2]/dl/dd[6]/text()'
xpath_price_secondary = '/html/body/main/div[4]/section[2]/div/div[2]/div[1]/div/div[1]/div/div[2]/dl/dd[5]/text()'
xpath_card_name = '/html/body/main/div[3]/div[1]/h1/text()'

with open('sample.txt', 'r', encoding='utf-8') as file:
    for line in file:
        target_url = line.strip()
        # 跳过空行
        if not target_url:
            continue
        # 增加超时配置,避免请求卡死
        resp = requests.get(target_url, headers=request_headers, timeout=15)
        # 校验请求状态码,非200状态直接跳过并打印提示
        if resp.status_code != 200:
            print(f"链接{target_url}请求失败,状态码:{resp.status_code}")
            name_list.append([])
            price_list.append([])
            continue
        page_dom = html.fromstring(resp.content)
        # 单次执行xpath查询,避免重复调用
        card_name = page_dom.xpath(xpath_card_name)
        card_price = page_dom.xpath(xpath_price_primary)
        # 匹配空列表时走备用xpath
        if len(card_price) == 0:
            card_price = page_dom.xpath(xpath_price_secondary)
        name_list.append(card_name)
        price_list.append(card_price)
        # 增加1秒请求间隔,避免请求过频被IP封禁
        time.sleep(1)

output_str = '----- Name ------------ Preis -----\n\n'
print(name_list)
print(price_list)
for idx in range(len(name_list)):
    # 增加空值兜底,避免匹配失败时解包报错
    current_name = name_list[idx][0].strip() if len(name_list[idx]) > 0 else "未匹配到卡牌名称"
    current_price = price_list[idx][0].strip() if len(price_list[idx]) > 0 else "未匹配到卡牌价格"
    output_str += f"{current_name} --> {current_price}\n"

print(output_str)
优化建议
  • 尽快替换写死的绝对路径xpath:可以通过定位包含“Price”类文本的dt标签,再取相邻dd节点的方式写相对xpath,页面结构小幅调整时不会失效
  • 若爬取量级较大,可以增加代理池、随机请求间隔配置,降低被反爬拦截的概率
  • 可以将匹配失败的链接单独写入错误日志文件,方便后续补爬

内容的提问来源于stack exchange,提问作者Sayonight

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 07:48:32