Python BeautifulSoup多URL爬虫如何向CSV追加写入新行?
问题原因
- 核心逻辑缩进错误:你把页面解析、CSV写入的代码全部放在了
for url in urls循环的外部,循环内仅完成了请求和soup对象初始化,每一轮循环都会覆盖上一轮的soup对象,最终循环结束后soup只保留了最后一个URL的页面内容,自然只能采集到最后一个地址的结果。 - 文件打开模式如果用
w(写入模式),每次打开都会清空原有内容,若要追加内容需用a(追加模式),结合你的使用场景,更建议把所有采集逻辑放到循环内,首次写入表头后续追加数据即可。
修正后可运行代码
from bs4 import BeautifulSoup import requests from csv import writer urls = ['https://example.com/1', 'https://example.com/2'] # 标记是否已经写入表头,避免重复写入 header_writed = False for url in urls: my_url = requests.get(url) html = my_url.content soup = BeautifulSoup(html,'html.parser') lists = soup.find_all('div', class_="profile-info-holder") links = soup.find_all('a', class_="intercept") # 循环内处理每个页面的数据写入,用a模式追加内容 with open('multi.csv', 'a', encoding='utf8', newline='') as f: thewriter = writer(f) if not header_writed: header = ['Name', 'Location', 'Link', 'Link2', 'Link3'] thewriter.writerow(header) header_writed = True for list_item in lists: name = list_item.find('div', class_="profile-name").text location = list_item.find('div', class_="profile-location").text # 补充判断避免页面链接不足3个时报索引错误 social1 = links[0].get('href') if len(links)>=1 else '' social2 = links[1].get('href') if len(links)>=2 else '' social3 = links[2].get('href') if len(links)>=3 else '' info = [name, location, social1, social2, social3] thewriter.writerow(info)
额外优化提示
- 不要用
list作为循环变量名,list是Python内置关键字,覆盖后会引发未知错误,上述代码已替换为list_item。 - 补充了links长度不足的判断,避免部分页面社交链接少于3个时抛出索引越界异常。
- 新增的表头写入标记,确保CSV文件只会在第一次写入时添加表头,不会重复写入表头行。
内容的提问来源于stack exchange,提问作者alexeidos
相关产品推荐
相关产品推荐

