You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy提取标签值与选择器?非结构化网站爬取遇阻

解决Scrapy中非规范页面的标签遍历与条件判断问题

一、先解决核心的Selector属性判断错误问题

你之前的判断失效,是因为对Scrapy Selector的返回值理解有误:

  • products[5].xpath('@class')返回的是SelectorList对象(包含匹配到的Selector的列表),不是直接的字符串值,所以直接和'list'比较永远不成立。
  • 要获取实际的属性值,必须用.get()(取第一个匹配项)或.getall()(取所有匹配项)方法提取内容。

正确的属性判断写法

# 遍历products中的每个标签
for item in products:
    # 获取标签名(如div、ul、table)
    tag_name = item.xpath('name()').get()
    # 获取class属性值(不存在则返回None)
    class_attr = item.xpath('@class').get()
    # 获取id属性值
    id_attr = item.xpath('@id').get()
    
    # 按标签类型+属性做条件判断
    if tag_name == 'div' and class_attr == 'a':
        # 执行div.class=a的操作,比如提取内容
        content = ' '.join([txt.strip() for txt in item.xpath('.//text()').getall() if txt.strip()])
        print("处理div.a内容:", content)
    elif tag_name == 'table' and id_attr == 'b':
        # 执行table#b的操作,比如处理表格行
        rows = item.xpath('.//tr[position()>1]')  # 跳过表头
        for row in rows:
            cols = row.xpath('.//td/text()').getall()
            print("表格行数据:", cols)
    elif tag_name == 'ul' and class_attr == 'list':
        print("匹配到ul.list标签")

单个元素的验证写法

在Scrapy Shell中测试单个元素的正确判断:

>>> products = response.xpath('//*[@class="r"]//*')
>>> # 获取第5个元素的class属性值
>>> class_val = products[5].xpath('@class').get()
>>> if class_val == 'list':
...     print("it works!")
...
it works!

二、处理非规范页面的标签关联问题(无容器的顺序式结构)

针对你提到的「两组标签无嵌套/容器,仅按顺序排列(比如标题和对应表格)」的场景,可以用以下两种方法关联元素:

方法1:通过索引匹配对应元素

如果两组标签数量完全对应(每个标题对应一个表格,顺序一一对应):

# 获取所有标题标签
titles = response.xpath('//h3[@class="title"]')
# 获取所有对应的数据表格
tables = response.xpath('//table[@class="data-table"]')

# 遍历索引,关联标题和表格
for idx in range(len(titles)):
    title_text = titles[idx].xpath('.//text()').get().strip()
    table = tables[idx]
    # 处理当前标题对应的表格数据
    print(f"处理标题: {title_text}")
    rows = table.xpath('.//tr')
    # ...后续表格处理逻辑

方法2:用XPath的following-sibling轴定位相邻元素

如果标签是「标题紧跟对应表格」的结构,用following-sibling可以精准获取当前元素之后的第一个目标元素:

# 遍历所有标题
for title in response.xpath('//h3[@class="title"]'):
    title_text = title.xpath('.//text()').get().strip()
    # 获取当前标题之后的第一个data-table表格
    target_table = title.xpath('following-sibling::table[@class="data-table"][1]')
    # 处理表格数据
    if target_table:
        rows = target_table.xpath('.//tr')
        print(f"标题{title_text}对应的表格有{len(rows)}行")

内容的提问来源于stack exchange,提问作者user2115136

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 09:22:07