You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Newspaper3k爬虫导出csv仅保存最后一条数据的问题求解

问题根因

你现有代码的写入逻辑是在所有URL遍历结束后才执行一次,遍历过程中article变量会被新爬取的内容不断覆盖,最终写入时仅保留了最后一次爬取成功的新闻数据,同时你用追加模式打开输出文件还每次重复写入表头,也会导致后续追加内容出现多余表头行的问题。

代码调整要点
  • 提前初始化输出csv文件,仅写入一次表头
  • 把写入行的逻辑移动到爬取循环内部,每成功爬取一条就立即写入
  • 爬取失败的URL直接跳过,不执行写入操作
  • 新增编码指定避免特殊字符乱码,优化异常提示的稳定性
调整后可正常运行的完整代码
from newspaper import Config
from newspaper import Article
from newspaper import ArticleException
import csv

USER_AGENT = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:78.0) Gecko/20100101 Firefox/78.0'

config = Config()
config.browser_user_agent = USER_AGENT
config.request_timeout = 10

# 提前初始化输出文件,仅写入一次表头
with open('test2.csv', 'w', newline='', encoding='utf-8') as csvfile:
    headers = ['article title', 'article text']
    writer = csv.DictWriter(csvfile, lineterminator='\n', fieldnames=headers)
    writer.writeheader()

    # 遍历爬取逻辑
    with open('test1.csv', 'r', encoding='utf-8') as file:
        csv_file = file.readlines()
        for url in csv_file:
            url = url.strip()
            if not url:
                continue
            try:
                article = Article(url, config=config)
                article.download()
                article.parse()
                print(article.title)
                # 处理正文换行,避免写入csv后格式错乱
                cleaned_text = article.text.replace('\n', ' ')
                print(cleaned_text)
                # 爬取成功立即写入
                writer.writerow({'article title': article.title,
                                'article text': cleaned_text})
            except ArticleException:
                print('***FAILED TO DOWNLOAD***', url)

内容的提问来源于stack exchange,提问作者Robbie Voort

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 15:06:06