如何用Scrapy(Python)筛选结果并提取纯净作者名?
解决作者姓名提取与爬取结果筛选问题
我来帮你搞定这两个核心问题——提取纯净的作者姓名,以及按特定字符串筛选爬取结果。咱们一步步拆解:
一、提取纯净的作者姓名
针对你提到的两种格式("by Elbert" 和 "Elbert (author)"),我整理了两种可靠的处理方式,你可以根据需求选择:
方法1:字符串简单处理(适合固定格式场景)
先去除字符串前后空白,再针对性去掉多余前缀后缀,逻辑直白易懂:
def clean_author(author_str): if not author_str: return "" # 先清理首尾空白 cleaned = author_str.strip() # 去掉开头的"by "前缀 if cleaned.startswith("by "): cleaned = cleaned[3:] # 去掉末尾的" (author)"后缀 if cleaned.endswith(" (author)"): cleaned = cleaned[:-9] return cleaned.strip()
方法2:正则表达式(更灵活,适配格式变种)
如果以后出现类似格式的变体(比如"by Elbert Smith"或"Elbert (writer)"),正则表达式能更健壮地提取姓名:
import re def clean_author(author_str): if not author_str: return "" # 匹配两种格式,精准提取姓名部分 match = re.search(r'(?:by )?([\w\s]+?)(?: \(author\))?$', author_str.strip()) return match.group(1).strip() if match else ""
在你的parse函数里,只需要把原始提取的作者字段传入这个清理函数:
author_raw = quote.xpath('.//div[@class="author"]/text()').extract_first() author = clean_author(author_raw)
二、基于特定字符串筛选爬取结果
假设你想筛选标题包含指定关键词、作者是特定人物,或者标签包含目标词的结果,只需要在循环中加入条件判断即可。举几个常见场景的例子:
场景1:筛选标题包含"Python"的内容
TARGET_KEYWORD = "Python" for quote in response.xpath('//div[@class="book"]'): title = quote.xpath('./div[@class="title"]/text()').extract_first() or "" # 不包含关键词就跳过当前条目 if TARGET_KEYWORD not in title: continue # 后续提取逻辑不变...
场景2:多条件筛选(作者是Elbert且标签包含"programming")
TARGET_AUTHOR = "Elbert" TARGET_TAG = "programming" for quote in response.xpath('//div[@class="book"]'): # 先提取并清理所有字段 author_raw = quote.xpath('.//div[@class="author"]/text()').extract_first() or "" author = clean_author(author_raw) tags_list = quote.xpath('.//div[@class="keywords"]/span[@class="tag"]/text()').extract() # 同时满足两个条件才保留 if author != TARGET_AUTHOR or TARGET_TAG not in tags_list: continue # 后续写入CSV和yield结果的逻辑...
修改后的完整代码
把上面的逻辑整合到你的代码中,最终版本如下(还修复了表头与列顺序不匹配的小问题):
# -- coding: utf-8 -- import csv import re def clean_author(author_str): if not author_str: return "" match = re.search(r'(?:by )?([\w\s]+?)(?: \(author\))?$', author_str.strip()) return match.group(1).strip() if match else "" def parse(self, response): # 可自定义筛选条件,留空则不启用对应筛选 TARGET_TITLE_KEYWORD = "" TARGET_AUTHOR = "" TARGET_TAG = "" with open('quotes-data.csv', 'w') as output_file: csv_writer = csv.writer(output_file, delimiter='\t', quotechar="'") # 修正表头顺序,与后续写入的row对应 csv_writer.writerow(['index', 'author', 'title', 'description', 'tags']) i = 1 for quote in response.xpath('//div[@class="book"]'): title = quote.xpath('./div[@class="title"]/text()').extract_first() or "" author_raw = quote.xpath('.//div[@class="author"]/text()').extract_first() or "" author = clean_author(author_raw) description = quote.xpath('.//div[@class="description"]/text()').extract_first() or "" tags_list = quote.xpath('.//div[@class="keywords"]/span[@class="tag"]/text()').extract() tags = ' '.join(tags_list) # 组合筛选条件 should_keep = True if TARGET_TITLE_KEYWORD and TARGET_TITLE_KEYWORD not in title: should_keep = False if TARGET_AUTHOR and author != TARGET_AUTHOR: should_keep = False if TARGET_TAG and TARGET_TAG not in tags_list: should_keep = False if not should_keep: continue # 格式化字段并写入CSV author = f'"{author}"' description = f'"{description}"' tags = f'"{tags}"' row = [i, author, title, description, tags] csv_writer.writerow(row) i += 1 # 返回格式化后的字典结果 yield { 'title': title.strip(), 'author': author.strip('"'), 'tags': tags.strip('"'), 'description': description.strip('"') }
内容的提问来源于stack exchange,提问作者Anss Sheikh
相关产品推荐
相关产品推荐

