You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy(Python)筛选结果并提取纯净作者名?

解决作者姓名提取与爬取结果筛选问题

我来帮你搞定这两个核心问题——提取纯净的作者姓名,以及按特定字符串筛选爬取结果。咱们一步步拆解:

一、提取纯净的作者姓名

针对你提到的两种格式("by Elbert" 和 "Elbert (author)"),我整理了两种可靠的处理方式,你可以根据需求选择:

方法1:字符串简单处理(适合固定格式场景)

先去除字符串前后空白,再针对性去掉多余前缀后缀,逻辑直白易懂:

def clean_author(author_str):
    if not author_str:
        return ""
    # 先清理首尾空白
    cleaned = author_str.strip()
    # 去掉开头的"by "前缀
    if cleaned.startswith("by "):
        cleaned = cleaned[3:]
    # 去掉末尾的" (author)"后缀
    if cleaned.endswith(" (author)"):
        cleaned = cleaned[:-9]
    return cleaned.strip()

方法2:正则表达式(更灵活,适配格式变种)

如果以后出现类似格式的变体(比如"by Elbert Smith"或"Elbert (writer)"),正则表达式能更健壮地提取姓名:

import re

def clean_author(author_str):
    if not author_str:
        return ""
    # 匹配两种格式,精准提取姓名部分
    match = re.search(r'(?:by )?([\w\s]+?)(?: \(author\))?$', author_str.strip())
    return match.group(1).strip() if match else ""

在你的parse函数里,只需要把原始提取的作者字段传入这个清理函数:

author_raw = quote.xpath('.//div[@class="author"]/text()').extract_first()
author = clean_author(author_raw)

二、基于特定字符串筛选爬取结果

假设你想筛选标题包含指定关键词、作者是特定人物,或者标签包含目标词的结果,只需要在循环中加入条件判断即可。举几个常见场景的例子:

场景1:筛选标题包含"Python"的内容

TARGET_KEYWORD = "Python"

for quote in response.xpath('//div[@class="book"]'):
    title = quote.xpath('./div[@class="title"]/text()').extract_first() or ""
    # 不包含关键词就跳过当前条目
    if TARGET_KEYWORD not in title:
        continue
    # 后续提取逻辑不变...

场景2:多条件筛选(作者是Elbert且标签包含"programming")

TARGET_AUTHOR = "Elbert"
TARGET_TAG = "programming"

for quote in response.xpath('//div[@class="book"]'):
    # 先提取并清理所有字段
    author_raw = quote.xpath('.//div[@class="author"]/text()').extract_first() or ""
    author = clean_author(author_raw)
    tags_list = quote.xpath('.//div[@class="keywords"]/span[@class="tag"]/text()').extract()
    
    # 同时满足两个条件才保留
    if author != TARGET_AUTHOR or TARGET_TAG not in tags_list:
        continue
    # 后续写入CSV和yield结果的逻辑...

修改后的完整代码

把上面的逻辑整合到你的代码中,最终版本如下(还修复了表头与列顺序不匹配的小问题):

# -- coding: utf-8 --
import csv
import re

def clean_author(author_str):
    if not author_str:
        return ""
    match = re.search(r'(?:by )?([\w\s]+?)(?: \(author\))?$', author_str.strip())
    return match.group(1).strip() if match else ""

def parse(self, response):
    # 可自定义筛选条件,留空则不启用对应筛选
    TARGET_TITLE_KEYWORD = ""
    TARGET_AUTHOR = ""
    TARGET_TAG = ""

    with open('quotes-data.csv', 'w') as output_file:
        csv_writer = csv.writer(output_file, delimiter='\t', quotechar="'")
        # 修正表头顺序,与后续写入的row对应
        csv_writer.writerow(['index', 'author', 'title', 'description', 'tags'])
        i = 1
        for quote in response.xpath('//div[@class="book"]'):
            title = quote.xpath('./div[@class="title"]/text()').extract_first() or ""
            author_raw = quote.xpath('.//div[@class="author"]/text()').extract_first() or ""
            author = clean_author(author_raw)
            description = quote.xpath('.//div[@class="description"]/text()').extract_first() or ""
            tags_list = quote.xpath('.//div[@class="keywords"]/span[@class="tag"]/text()').extract()
            tags = ' '.join(tags_list)
            
            # 组合筛选条件
            should_keep = True
            if TARGET_TITLE_KEYWORD and TARGET_TITLE_KEYWORD not in title:
                should_keep = False
            if TARGET_AUTHOR and author != TARGET_AUTHOR:
                should_keep = False
            if TARGET_TAG and TARGET_TAG not in tags_list:
                should_keep = False
            
            if not should_keep:
                continue
            
            # 格式化字段并写入CSV
            author = f'"{author}"'
            description = f'"{description}"'
            tags = f'"{tags}"'
            row = [i, author, title, description, tags]
            csv_writer.writerow(row)
            i += 1
            # 返回格式化后的字典结果
            yield {
                'title': title.strip(),
                'author': author.strip('"'),
                'tags': tags.strip('"'),
                'description': description.strip('"')
            }

内容的提问来源于stack exchange,提问作者Anss Sheikh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:05:53