Python lxml requests循环xpath查询仅最后一条结果存入列表bug修复
问题根源
代码存在4个直接导致运行异常的问题:
- 读取txt文件内链接时未去除行尾换行符、首尾空白字符,传入requests的URL包含非法字符,多数请求实际无法命中正确的卡牌详情页
- lxml库的
xpath()方法不存在返回None的情况:匹配成功时返回元素/文本列表,匹配失败时返回空列表[],原有逻辑中if code.xpath(xpathPrice) == None的判断永远不会触发,价格匹配逻辑完全失效 - 未配置合法请求头,Cardmarket站点有基础反爬规则,默认的requests请求标识会被拦截,返回反爬验证页而非目标卡牌内容,自然无法匹配到预设xpath
- 原有逻辑重复调用xpath执行相同查询,既浪费性能,也容易在页面结构波动时出现取值不一致的问题
修复后可运行代码
from lxml import html import requests import time price_list = [] name_list = [] # 配置浏览器请求头,绕过基础反爬校验 request_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" } # 注意:绝对路径xpath适配性极差,页面小幅改版就会失效,建议后续替换为基于属性、相邻文本的相对定位写法 xpath_price_primary = '/html/body/main/div[4]/section[2]/div/div[2]/div[1]/div/div[1]/div/div[2]/dl/dd[6]/text()' xpath_price_secondary = '/html/body/main/div[4]/section[2]/div/div[2]/div[1]/div/div[1]/div/div[2]/dl/dd[5]/text()' xpath_card_name = '/html/body/main/div[3]/div[1]/h1/text()' with open('sample.txt', 'r', encoding='utf-8') as file: for line in file: target_url = line.strip() # 跳过空行 if not target_url: continue # 增加超时配置,避免请求卡死 resp = requests.get(target_url, headers=request_headers, timeout=15) # 校验请求状态码,非200状态直接跳过并打印提示 if resp.status_code != 200: print(f"链接{target_url}请求失败,状态码:{resp.status_code}") name_list.append([]) price_list.append([]) continue page_dom = html.fromstring(resp.content) # 单次执行xpath查询,避免重复调用 card_name = page_dom.xpath(xpath_card_name) card_price = page_dom.xpath(xpath_price_primary) # 匹配空列表时走备用xpath if len(card_price) == 0: card_price = page_dom.xpath(xpath_price_secondary) name_list.append(card_name) price_list.append(card_price) # 增加1秒请求间隔,避免请求过频被IP封禁 time.sleep(1) output_str = '----- Name ------------ Preis -----\n\n' print(name_list) print(price_list) for idx in range(len(name_list)): # 增加空值兜底,避免匹配失败时解包报错 current_name = name_list[idx][0].strip() if len(name_list[idx]) > 0 else "未匹配到卡牌名称" current_price = price_list[idx][0].strip() if len(price_list[idx]) > 0 else "未匹配到卡牌价格" output_str += f"{current_name} --> {current_price}\n" print(output_str)
优化建议
- 尽快替换写死的绝对路径xpath:可以通过定位包含“Price”类文本的dt标签,再取相邻dd节点的方式写相对xpath,页面结构小幅调整时不会失效
- 若爬取量级较大,可以增加代理池、随机请求间隔配置,降低被反爬拦截的概率
- 可以将匹配失败的链接单独写入错误日志文件,方便后续补爬
内容的提问来源于stack exchange,提问作者Sayonight
相关产品推荐
相关产品推荐

