You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取含大量重复类名页面的技术求助

解决Scrapy爬取宝可梦维基时类名重复的信息提取问题

嘿,刚接触Scrapy就碰到这种DOM结构重复的问题太正常了!我之前爬宝可梦维基的时候也踩过同样的坑,给你分享几个精准提取目标信息的方法:

核心思路:利用上下文关联定位

页面里重复的pi-smart-group其实是一个个信息分组,每个分组都有对应的标题(比如“属性”“种类”)。我们要做的就是通过标题锁定目标分组,再提取分组内的内容,而不是直接抓取所有类名匹配的元素。


方法1:用CSS选择器的:has()精准定位分组

Scrapy基于的Parsel库支持:has()伪类,可以轻松找到包含特定标题的分组。比如提取“属性”“种类”“特性”的代码示例:

# 提取宝可梦属性
type_section = response.css('section.pi-smart-group:has(h3:contains("属性"))')
pokemon_types = type_section.css('div.pi-smart-data-value a::text').getall()

# 提取种类
species_section = response.css('section.pi-smart-group:has(h3:contains("种类"))')
pokemon_species = species_section.css('div.pi-smart-data-value ::text').get().strip()

# 提取特性
ability_section = response.css('section.pi-smart-group:has(h3:contains("特性"))')
pokemon_abilities = ability_section.css('div.pi-smart-data-value ::text').getall()
# 清理特性文本里的空格和换行
pokemon_abilities = [text.strip() for text in pokemon_abilities if text.strip()]

方法2:批量提取所有目标字段(更高效)

如果要一次性提取多个字段,可以先定义字段和对应标题的映射,循环处理:

# 定义需要提取的字段与页面标题的对应关系
target_fields = {
    '属性': 'types',
    '种类': 'species',
    '特性': 'abilities',
    '身高': 'height',
    '体重': 'weight'
}

pokemon_info = {}
for title, key in target_fields.items():
    # 定位到对应标题的分组
    section = response.css(f'section.pi-smart-group:has(h3:contains("{title}"))')
    if not section:
        pokemon_info[key] = None
        continue
    # 根据字段类型选择提取方式(多值/单值)
    if title in ['属性', '特性']:
        content = section.css('div.pi-smart-data-value ::text').getall()
        pokemon_info[key] = [text.strip() for text in content if text.strip()]
    else:
        pokemon_info[key] = section.css('div.pi-smart-data-value ::text').get().strip()

备选方案:用XPath轴定位(兼容更多场景)

如果担心:has()的兼容性(其实Parsel完全支持),可以用XPath的兄弟轴来关联标题和内容:

# 用XPath提取身高
height = response.xpath('//h3[contains(text(), "身高")]/../following-sibling::section[@class="pi-smart-group-body"]//div[@class="pi-smart-data-value"]/text()').get().strip()

这个XPath的逻辑是:先找到包含“身高”的h3,回到它的父节点(标题栏),再找到后面紧邻的内容栏(pi-smart-group-body),最后提取文本。


小技巧提示

  1. 先在浏览器开发者工具里测试选择器:按Ctrl+F输入CSS/XPath,确认是否精准定位到目标元素;
  2. 处理文本时一定要用strip(),去掉页面里多余的空格、换行符;
  3. 部分宝可梦有隐藏特性,提取特性时要注意区分常规特性和隐藏特性的结构。

内容的提问来源于stack exchange,提问作者aria

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 23:42:30