You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫提取文章内容遇问题:目标Div含冗余内容,求代码修复

问题描述

使用Python的Scrapy框架爬取指定文章内容时,目标div包含大量冗余元素,当前代码运行后Article_content返回空值,需要精准提取纯文章内容。

目标页面:https://www.fodors.com/world/asia/india/experiences/news/things-you-need-to-know-before-you-visit-india

用户原有代码:

# Content = {}
# header,paragraphs  = "",[]
# for element in response.xpath('//*[@class="entry-content content-single container "]/*'):
#     tag = element.re(r"<(\w+)\s")  # get the tag name
#     # if its a paragraph add it to the paragraph list
#     if tag[0] == "p":              
#         paragraphs += element.xpath(".//text()").getall()
#     # if it's a heading place the heading and paragraphs in the
#     # dictionary and start a new heading with the current text.
#     elif tag[0] == "h3":
#         Content[header] = ''.join(paragraphs).strip()
#         header = ' '.join(element.xpath(".//text()").getall()).strip()
#         paragraphs = []


Article_Content = response.xpath('//*[@class="entry-content content-single container "]/text()')
Content = '\n'.join(Article_Content.getall()).strip()

yield{
    'Category':Category,
    'Headlines':Headlines,
    'Author': Author,
    'Source': Source,
    'Publication Date': Published_Date,
    'Feature_Image': Feature_Image,
    'Article Content': Content
}
修复方案

原代码返回空值的核心原因://*[@class="entry-content content-single container "]/text()仅提取该div下的直接文本节点,但文章的实际内容全部嵌套在<p>、<h3>等子标签中,因此无法获取到有效内容。

方案1:提取合并后的纯文本内容

直接获取目标div下所有子节点的文本,过滤空白内容后合并:

# 提取目标div下所有非空白文本节点
all_text_nodes = response.xpath('//*[@class="entry-content content-single container "]//text()').getall()
# 过滤空内容,清理每个文本段的前后空白
clean_content = '\n'.join([text.strip() for text in all_text_nodes if text.strip()])

yield {
    'Category': Category,
    'Headlines': Headlines,
    'Author': Author,
    'Source': Source,
    'Publication Date': Published_Date,
    'Feature_Image': Feature_Image,
    'Article Content': clean_content
}

方案2:结构化提取(保留标题与段落的对应关系)

如果需要保留文章的层级结构(如<h3>标题对应下方的<p>段落),可以改进循环逻辑,避免用正则解析标签名,直接通过XPath判断标签类型:

content_struct = {}
current_heading = ""
current_paras = []

# 遍历目标div下的所有直接子元素
for elem in response.xpath('//*[@class="entry-content content-single container "]/*'):
    # 处理h3标题
    if elem.xpath('self::h3'):
        # 保存上一组标题和段落
        if current_heading and current_paras:
            content_struct[current_heading] = '\n'.join(current_paras).strip()
        # 更新当前标题
        current_heading = elem.xpath('.//text()').get().strip()
        current_paras = []
    # 处理p段落
    elif elem.xpath('self::p'):
        para_text = elem.xpath('.//text()').get().strip()
        if para_text:
            current_paras.append(para_text)

# 处理最后一组未保存的内容
if current_heading and current_paras:
    content_struct[current_heading] = '\n'.join(current_paras).strip()

# 转换为带格式的纯文本(或直接返回字典结构)
formatted_content = ""
for heading, paras in content_struct.items():
    formatted_content += f"### {heading}\n{paras}\n\n"

yield {
    'Category': Category,
    'Headlines': Headlines,
    'Author': Author,
    'Source': Source,
    'Publication Date': Published_Date,
    'Feature_Image': Feature_Image,
    'Article Content': formatted_content.strip()
}

内容的提问来源于stack exchange,提问作者Info Rewind

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 01:40:51