You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup多URL爬虫如何向CSV追加写入新行?

问题原因

  • 核心逻辑缩进错误:你把页面解析、CSV写入的代码全部放在了for url in urls循环的外部,循环内仅完成了请求和soup对象初始化,每一轮循环都会覆盖上一轮的soup对象,最终循环结束后soup只保留了最后一个URL的页面内容,自然只能采集到最后一个地址的结果。
  • 文件打开模式如果用w(写入模式),每次打开都会清空原有内容,若要追加内容需用a(追加模式),结合你的使用场景,更建议把所有采集逻辑放到循环内,首次写入表头后续追加数据即可。

修正后可运行代码

from bs4 import BeautifulSoup
import requests
from csv import writer

urls = ['https://example.com/1', 'https://example.com/2']
# 标记是否已经写入表头,避免重复写入
header_writed = False

for url in urls:
    my_url = requests.get(url)
    html = my_url.content
    soup = BeautifulSoup(html,'html.parser')

    lists = soup.find_all('div', class_="profile-info-holder")
    links = soup.find_all('a', class_="intercept")

    # 循环内处理每个页面的数据写入,用a模式追加内容
    with open('multi.csv', 'a', encoding='utf8', newline='') as f:
        thewriter = writer(f)
        if not header_writed:
            header = ['Name', 'Location', 'Link', 'Link2', 'Link3']
            thewriter.writerow(header)
            header_writed = True

        for list_item in lists:
            name = list_item.find('div', class_="profile-name").text
            location = list_item.find('div', class_="profile-location").text
            # 补充判断避免页面链接不足3个时报索引错误
            social1 = links[0].get('href') if len(links)>=1 else ''
            social2 = links[1].get('href') if len(links)>=2 else ''
            social3 = links[2].get('href') if len(links)>=3 else ''

            info = [name, location, social1, social2, social3]
            thewriter.writerow(info)

额外优化提示

  • 不要用list作为循环变量名,list是Python内置关键字,覆盖后会引发未知错误,上述代码已替换为list_item。
  • 补充了links长度不足的判断,避免部分页面社交链接少于3个时抛出索引越界异常。
  • 新增的表头写入标记,确保CSV文件只会在第一次写入时添加表头,不会重复写入表头行。

内容的提问来源于stack exchange,提问作者alexeidos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 16:15:03