You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用Web Scrapy抓取Li类元素及Xpath提取VIN失败

Scrapy抓取Li元素及VIN提取问题解决办法

一、VIN提取失败的修复方案

针对你用response.xpath(".//div[@class='vinDisplay']/text()").get()提取失败的情况,结合页面结构给出几种可行写法:

  • 处理文本空白:如果VIN文本前后有空格、换行,用normalize-space()过滤后提取:
response.xpath("normalize-space(.//div[@class='vinDisplay']/text())").get()
  • 检查子节点嵌套:如果VIN实际在div.vinDisplay的子标签(比如<span>)里,把路径改成匹配所有子节点文本:
response.xpath(".//div[@class='vinDisplay']//text()").get()
  • 排查动态加载:如果页面是JS渲染的,静态HTML里找不到这个div,就得用动态渲染工具,比如scrapy-playwright,示例代码:
async def parse(self, response):
    page = await self.playwright.new_page()
    await page.goto(response.url)
    # 等待元素加载完成
    await page.locator("div.vinDisplay").wait_for()
    vin = await page.locator("div.vinDisplay").text_content()
    await page.close()
    # 后续处理vin数据

二、Li元素抓取的通用写法

  • 抓取页面所有Li元素的文本内容:
all_li_texts = response.xpath("//li/text()").getall()
  • 抓取指定class的Li元素并提取内部内容:
target_list = response.xpath("//li[@class='your-target-class']")
for li in target_list:
    item = {}
    item['title'] = li.xpath(".//h3/text()").get()
    item['detail'] = li.xpath(".//p/text()").get()
    yield item

内容的提问来源于stack exchange,提问作者Yonatan Gebreyesus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 20:15:33