如何将多个URL抓取内容保存为CSV?Python爬虫问题求助
问题分析与解决
你的代码每次循环都会重新创建并覆盖results.csv,最终只保留最后一个URL的抓取结果——原因是在for循环内部使用wb模式打开文件,这个模式会清空原有内容再写入新数据,循环1500次就会覆盖1499次。
解决思路
把结果文件的打开操作移到循环外面,使用ab(追加二进制)模式,这样每次请求到的数据都会追加到文件末尾,而不是覆盖原有内容。
根据接口返回的CSV格式,分两种处理方式:
情况1:每个接口返回单条餐厅数据(无重复表头)
如果接口返回的是单条数据行(不带表头),可以先手动写入一次表头(如果需要),然后循环追加每条数据:
import csv import requests # 打开结果文件,用追加模式 with open("results.csv", "ab") as result_file: with open(r'C:\Users\xxx\Desktop\Brand.csv') as f: reader = csv.reader(f) # 可选:如果需要表头,先写入一次(替换成接口返回的实际表头字段) # header = b"businessId,callToActions,city,country,...\n" # result_file.write(header) for row in reader: url = "https://locator.uberall.com/api/storefinders/04AxJXeSBFk4qtBbhQ9JCK1987mtnF/locations/" + row[0] + "?v=20230110&language=de&fieldMask=businessId&fieldMask=callToActions&fieldMask=city&fieldMask=country&fieldMask=descriptionLong&fieldMask=descriptionShort&fieldMask=distance&fieldMask=email&fieldMask=fax&fieldMask=googlePlaceId&fieldMask=id&fieldMask=identifier&fieldMask=keywords&fieldMask=lat&fieldMask=lng&fieldMask=name&fieldMask=openingHours&fieldMask=openingHoursNotes&fieldMask=phone&fieldMask=photos&fieldMask=province&fieldMask=socialPost&fieldMask=specialOpeningHours&fieldMask=streetAndNumber&fieldMask=timezone&fieldMask=zip&fieldMask=brands&fieldMask=customItems&fieldMask=events&fieldMask=languages&fieldMask=menus&fieldMask=paymentOptions&fieldMask=people&fieldMask=products&fieldMask=services&fieldMask=socialProfiles&identifier=true" print(url) r = requests.get(url) # 追加当前请求的数据 result_file.write(r.content) # 手动加换行,避免多条数据粘在一起 result_file.write(b"\n")
情况2:每个接口返回包含表头的完整CSV
如果接口返回的是带表头的完整CSV,需要跳过后续请求的表头,只保留第一次的表头:
import csv import requests first_write = True with open("results.csv", "ab") as result_file: with open(r'C:\Users\xxx\Desktop\Brand.csv') as f: reader = csv.reader(f) for row in reader: url = "https://locator.uberall.com/api/storefinders/04AxJXeSBFk4qtBbhQ9JCK1987mtnF/locations/" + row[0] + "?v=20230110&language=de&fieldMask=businessId&fieldMask=callToActions&fieldMask=city&fieldMask=country&fieldMask=descriptionLong&fieldMask=descriptionShort&fieldMask=distance&fieldMask=email&fieldMask=fax&fieldMask=googlePlaceId&fieldMask=id&fieldMask=identifier&fieldMask=keywords&fieldMask=lat&fieldMask=lng&fieldMask=name&fieldMask=openingHours&fieldMask=openingHoursNotes&fieldMask=phone&fieldMask=photos&fieldMask=province&fieldMask=socialPost&fieldMask=specialOpeningHours&fieldMask=streetAndNumber&fieldMask=timezone&fieldMask=zip&fieldMask=brands&fieldMask=customItems&fieldMask=events&fieldMask=languages&fieldMask=menus&fieldMask=paymentOptions&fieldMask=people&fieldMask=products&fieldMask=services&fieldMask=socialProfiles&identifier=true" print(url) r = requests.get(url) content = r.content if not first_write: # 跳过后续内容的表头,只保留数据行 lines = content.split(b"\n") if len(lines) > 1: content = b"\n".join(lines[1:]) + b"\n" result_file.write(content) first_write = False
额外提示
- 加异常处理:用
try-except包裹请求逻辑,避免单个请求失败导致整个程序中断 - 限速:1500次请求容易触发反爬,建议每次请求后加
time.sleep(1)(需要导入time模块)
内容的提问来源于stack exchange,提问作者theDemnex
相关产品推荐
相关产品推荐

