You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup从CSV批量爬取URL数据仅获最后一条,如何修正?

解决爬虫仅保留最后一条URL数据的问题

Hey, I see the issue here! The problem is that you're opening the output.csv file inside your loop with the 'w' mode—this overwrites the entire file every time you process a new URL, which is why only the last entry stays. Let's fix this properly:

核心问题分析

当你使用open('output.csv', 'w')时:

  • 'w'是写入模式,会清空文件原有内容并创建新文件(如果不存在)
  • 每次循环都执行这个操作,前面爬取的数据都会被覆盖

两种解决方案

方案1:将文件打开操作移到循环外(推荐)

这种方式更高效,只打开/关闭文件一次,避免重复IO开销:

import requests
from bs4 import BeautifulSoup
import csv

# 先打开输出文件,写入表头
with open('output.csv', 'w', newline='', encoding='utf-8') as output_file:
    csv_writer = csv.writer(output_file)
    csv_writer.writerow(['Ngoname', 'CEO', 'City', 'Address', 'Phone', 'Mobile', 'E-mail'])

    # 读取URL列表并循环爬取
    with open('urls.csv' , 'r', encoding='utf-8') as csv_file:
        csv_reader = csv.reader(csv_file)
        for line in csv_reader:
            url = line[0]
            try:
                r = requests.get(url).text
                soup = BeautifulSoup(r,'lxml')
                
                # 提取数据(增加异常处理,避免单个URL爬取失败中断整个程序)
                ngoname = soup.find('h1').text if soup.find('h1') else 'N/A'
                
                ceo_tag = soup.find('h2', class_='')
                ceo_name = ceo_tag.split(':')[1].strip() if ceo_tag else 'N/A'
                
                spans = soup.find_all('span')
                city = spans[5].text if len(spans)>=6 else 'N/A'
                address = spans[6].text if len(spans)>=7 else 'N/A'
                phone = spans[7].text if len(spans)>=8 else 'N/A'
                mobile = spans[8].text if len(spans)>=9 else 'N/A'
                email = spans[9].text if len(spans)>=10 else 'N/A'
                
                # 打印并写入数据
                print(f"Processed: {ngoname}")
                csv_writer.writerow([ngoname, ceo_name, city, address, phone, mobile, email])
                
            except Exception as e:
                print(f"Failed to process {url}: {str(e)}")
                # 可以写入错误信息到文件,方便排查
                csv_writer.writerow(['Error', str(e), url, 'N/A', 'N/A', 'N/A', 'N/A'])

方案2:使用追加模式('a')

如果必须在循环内打开文件,可以用'a'追加模式,但要注意只写一次表头:

import requests
from bs4 import BeautifulSoup
import csv
import os

# 检查文件是否存在,不存在则写入表头
if not os.path.exists('output.csv'):
    with open('output.csv', 'w', newline='', encoding='utf-8') as output_file:
        csv_writer = csv.writer(output_file)
        csv_writer.writerow(['Ngoname', 'CEO', 'City', 'Address', 'Phone', 'Mobile', 'E-mail'])

# 读取URL并循环爬取
with open('urls.csv' , 'r', encoding='utf-8') as csv_file:
    csv_reader = csv.reader(csv_file)
    for line in csv_reader:
        url = line[0]
        try:
            r = requests.get(url).text
            soup = BeautifulSoup(r,'lxml')
            
            # 提取数据(同方案1的提取逻辑)
            ngoname = soup.find('h1').text if soup.find('h1') else 'N/A'
            ceo_tag = soup.find('h2', class_='')
            ceo_name = ceo_tag.split(':')[1].strip() if ceo_tag else 'N/A'
            spans = soup.find_all('span')
            city = spans[5].text if len(spans)>=6 else 'N/A'
            address = spans[6].text if len(spans)>=7 else 'N/A'
            phone = spans[7].text if len(spans)>=8 else 'N/A'
            mobile = spans[8].text if len(spans)>=9 else 'N/A'
            email = spans[9].text if len(spans)>=10 else 'N/A'
            
            # 追加写入数据
            with open('output.csv', 'a', newline='', encoding='utf-8') as output_file:
                csv_writer = csv.writer(output_file)
                csv_writer.writerow([ngoname, ceo_name, city, address, phone, mobile, email])
                
            print(f"Processed: {ngoname}")
            
        except Exception as e:
            print(f"Failed to process {url}: {str(e)}")
            with open('output.csv', 'a', newline='', encoding='utf-8') as output_file:
                csv_writer = csv.writer(output_file)
                csv_writer.writerow(['Error', str(e), url, 'N/A', 'N/A', 'N/A', 'N/A'])

额外优化点

  1. 增加异常处理:避免单个URL爬取失败(比如网络错误、元素找不到)导致整个程序终止
  2. 指定编码:添加encoding='utf-8'避免中文乱码
  3. 使用newline='':避免CSV文件出现多余空行
  4. 复用spans变量:不用每次都调用soup.find_all('span'),提升效率

内容的提问来源于stack exchange,提问作者harshan p

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:57:12