You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取网页时多分段内容合并问题求助

Scrapy爬取多section内容问题修复

问题分析

你遇到的核心问题有三个:

  1. 循环内每次赋值section_text会直接覆盖之前的内容,最终仅保留最后一个section的数据
  2. 代码中' '.join('a::text').getall()和' '.join('ul::text').getall()是错误写法,不能直接用字符串调用getall(),必须基于当前section节点提取对应元素文本
  3. 每个section内的p、a、ul文本没有合并,而是被逐个覆盖,导致单一section的内容也不完整

修正后的代码

def parse_instructions(self, response):
    title = response.xpath('//*[@id="d-article"]/div[1]/div[1]/h1/text()').get()
    description = response.xpath('//*[@id="ency_summary"]/p/text()').getall()
    joined_description = ' '.join(description)
    sections = response.css('section div.section:not([class*=" "])')
    
    # 初始化列表存储所有section的内容
    all_section_texts = []
    
    for section in sections:
        # 提取当前section下的p、a、ul文本内容
        p_text = section.css('p::text').getall()
        a_text = section.css('a::text').getall()
        ul_text = section.css('ul::text').getall()
        
        # 合并当前section的所有文本并去重空格
        section_content = ' '.join(p_text + a_text + ul_text).strip()
        if section_content:  # 过滤空内容的section
            all_section_texts.append(section_content)
    
    yield {
        "title": title,
        "description": joined_description,
        # 返回所有section的内容列表,若需要合并成单个字符串可替换为' '.join(all_section_texts)
        "section_text": all_section_texts,
    }

关键修正点

  • 用all_section_texts列表收集每个section的内容,彻底避免循环内的变量覆盖问题
  • 正确基于当前section节点提取p、a、ul的文本,合并后加入列表
  • 新增空内容过滤逻辑,避免无效数据
  • 可选:如果需要将所有section内容合并为一个字符串,直接把"section_text"的值改为' '.join(all_section_texts)即可

内容的提问来源于stack exchange,提问作者Sairam S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 01:55:12