Scrapy爬虫提取文章内容遇问题:目标Div含冗余内容,求代码修复
问题描述
使用Python的Scrapy框架爬取指定文章内容时,目标div包含大量冗余元素,当前代码运行后Article_content返回空值,需要精准提取纯文章内容。
目标页面:https://www.fodors.com/world/asia/india/experiences/news/things-you-need-to-know-before-you-visit-india
用户原有代码:
# Content = {} # header,paragraphs = "",[] # for element in response.xpath('//*[@class="entry-content content-single container "]/*'): # tag = element.re(r"<(\w+)\s") # get the tag name # # if its a paragraph add it to the paragraph list # if tag[0] == "p": # paragraphs += element.xpath(".//text()").getall() # # if it's a heading place the heading and paragraphs in the # # dictionary and start a new heading with the current text. # elif tag[0] == "h3": # Content[header] = ''.join(paragraphs).strip() # header = ' '.join(element.xpath(".//text()").getall()).strip() # paragraphs = [] Article_Content = response.xpath('//*[@class="entry-content content-single container "]/text()') Content = '\n'.join(Article_Content.getall()).strip() yield{ 'Category':Category, 'Headlines':Headlines, 'Author': Author, 'Source': Source, 'Publication Date': Published_Date, 'Feature_Image': Feature_Image, 'Article Content': Content }
修复方案
原代码返回空值的核心原因://*[@class="entry-content content-single container "]/text()仅提取该div下的直接文本节点,但文章的实际内容全部嵌套在<p>、<h3>等子标签中,因此无法获取到有效内容。
方案1:提取合并后的纯文本内容
直接获取目标div下所有子节点的文本,过滤空白内容后合并:
# 提取目标div下所有非空白文本节点 all_text_nodes = response.xpath('//*[@class="entry-content content-single container "]//text()').getall() # 过滤空内容,清理每个文本段的前后空白 clean_content = '\n'.join([text.strip() for text in all_text_nodes if text.strip()]) yield { 'Category': Category, 'Headlines': Headlines, 'Author': Author, 'Source': Source, 'Publication Date': Published_Date, 'Feature_Image': Feature_Image, 'Article Content': clean_content }
方案2:结构化提取(保留标题与段落的对应关系)
如果需要保留文章的层级结构(如<h3>标题对应下方的<p>段落),可以改进循环逻辑,避免用正则解析标签名,直接通过XPath判断标签类型:
content_struct = {} current_heading = "" current_paras = [] # 遍历目标div下的所有直接子元素 for elem in response.xpath('//*[@class="entry-content content-single container "]/*'): # 处理h3标题 if elem.xpath('self::h3'): # 保存上一组标题和段落 if current_heading and current_paras: content_struct[current_heading] = '\n'.join(current_paras).strip() # 更新当前标题 current_heading = elem.xpath('.//text()').get().strip() current_paras = [] # 处理p段落 elif elem.xpath('self::p'): para_text = elem.xpath('.//text()').get().strip() if para_text: current_paras.append(para_text) # 处理最后一组未保存的内容 if current_heading and current_paras: content_struct[current_heading] = '\n'.join(current_paras).strip() # 转换为带格式的纯文本(或直接返回字典结构) formatted_content = "" for heading, paras in content_struct.items(): formatted_content += f"### {heading}\n{paras}\n\n" yield { 'Category': Category, 'Headlines': Headlines, 'Author': Author, 'Source': Source, 'Publication Date': Published_Date, 'Feature_Image': Feature_Image, 'Article Content': formatted_content.strip() }
内容的提问来源于stack exchange,提问作者Info Rewind
相关产品推荐
相关产品推荐

