You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫存储CSV仅写入单行数据问题求助

解决Python爬取cryptonews.net数据到CSV时重复写入/仅存单行的问题

问题根源分析

你的代码存在几个核心逻辑错误,直接导致了该bug:

  • 变量名冲突:循环页码变量page被requests.get(url)的响应结果覆盖,导致后续页码循环异常。
  • 数据收集逻辑错误:每次爬取单条数据时直接覆盖data变量,未将所有数据存入列表,最终写入CSV时只会使用最后一次的data值。
  • 写入时机错误:爬取完所有页面后才打开文件,且用无意义的range(100)循环重复写入同一个data。
  • 页面元素定位错误:soup.find_all("main")仅能获取页面的main容器,而非每条新闻条目,导致实际只爬取到单条(或未正确获取所有新闻)。

修复后的完整代码

import requests
from bs4 import BeautifulSoup
import csv

# 初始化存储所有新闻数据的列表
all_news_data = []
header = ['Title', 'Tag', 'UTC', 'Web_Address', 'Image_Src']

# 循环爬取第0到第9页(共10页)
for page_num in range(0, 10):
    url = f"https://cryptonews.net/?page={page_num}"
    response = requests.get(url)
    soup = BeautifulSoup(response.content, "html.parser")
    
    # 定位页面内所有新闻条目(根据网站实际DOM结构,新闻条目class为news-item)
    news_items = soup.find_all("div", class_="news-item")
    
    for item in news_items:
        # 提取字段并添加异常处理,避免单个字段缺失导致程序崩溃
        title = item.find('a', class_="title").text.strip() if item.find('a', class_="title") else "N/A"
        tag = item.find('span', class_="etc-mark").text.strip() if item.find('span', class_="etc-mark") else "N/A"
        datetime = item.find('span', class_="datetime").text.strip() if item.find('span', class_="datetime") else "N/A"
        # 修正网页地址提取逻辑:取a标签的href属性而非文本
        web_address = item.find('a', class_="title")['href'] if item.find('a', class_="title") else "N/A"
        # 修正图片地址提取逻辑:取img标签的src属性
        img_src = item.find('img')['src'] if item.find('img') else "N/A"
        
        # 将单条新闻数据加入总列表
        all_news_data.append([title, tag, datetime, web_address, img_src])

# 批量写入CSV文件
with open('crypto.csv', 'w', newline='', encoding='utf-8') as crypto_file:
    writer = csv.writer(crypto_file)
    writer.writerow(header)
    # 一次性写入所有收集到的新闻数据
    writer.writerows(all_news_data)

关键优化点说明

  • 变量名规范:将循环页码变量改为page_num,避免与请求响应对象冲突。
  • 数据存储逻辑:用all_news_data列表存储所有爬取结果,确保每条数据都被保留,不会被覆盖。
  • 元素定位修正:使用div.news-item定位单条新闻,而非main标签,确保能获取页面内所有新闻条目。
  • 字段提取修正:网页地址提取a.title的href属性,图片地址提取img标签的src属性,匹配网站实际结构。
  • 写入逻辑优化:爬取完成后一次性打开文件批量写入,提升效率同时避免重复写入问题。
  • 异常防护:添加if...else判断,避免因单个字段缺失导致程序中断。

内容的提问来源于stack exchange,提问作者Abdul Rehman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 16:21:12