Python使用requests库批量处理URL列表数据的实现方法
Python批量处理URL列表实现方案
你的代码问题出在两点:一是URL列表定义格式错误,二是没有用循环结构遍历列表执行处理逻辑,按以下方式修改即可:
- 修正URL列表写法:你注释的
myString存在两个语法错误:列表末尾误用大括号}闭合,且多个URL被拼接在同一个字符串内。正确写法是每个URL作为独立的字符串元素,用方括号[]包裹列表,元素间用逗号分隔。 - 用
for循环遍历URL列表,把你已经调试通的单URL处理逻辑放到循环体内,每次循环取一个URL发起请求即可。 - 补充两个必要的容错处理:第一,仅当响应状态码为200时才执行内容解析,3xx/4xx/其他状态码直接走你预留的分支逻辑,避免非成功响应的页面结构不匹配,导致字符串分割报错;第二,增加请求超时设置和异常捕获,避免单个URL请求失败直接中断整个批量任务。
修改后可直接运行的完整代码:
import requests # 正确定义URL列表,每个URL为独立字符串元素 myString = [ 'https://{redacted by me}/vessels/amadi_9682552_10003796/', 'https://{redacted by me}/vessels/akebono-maru_9554729_2866687/', 'https://{redacted by me}/vessels/amani_9661869_9276632/', 'https://{redacted by me}/vessels/aman-sendai_9134323_2017277/', 'https://{redacted by me}/vessels/al-aamriya_9338266_25273/' ] # 遍历列表批量处理 for url in myString: print(f"当前处理: {url}") try: # 设置10秒超时,避免请求无限等待 r = requests.get(url, timeout=10) if r.status_code == 200: print(f"请求成功,状态码: {r.status_code}") data = r.content.decode("utf-8") # 提取目标字段 next_port_locode = data.split('locode: "')[1].split('"')[0].strip() next_port_iso2 = data.split('iso2: "')[1].split('"')[0].strip() next_port_name = data.split('iso2: "')[1].split('name: "')[1].split('"')[0].strip() next_port_eta = data.split('eta: moment("')[1].split('"')[0].strip() next_port_latitude = float(data.split('latitude: ')[1].split(',')[0].strip()) next_port_longitude = float(data.split('longitude: ')[1].split('\n')[0].strip()) datajson = { "next_port_locode": next_port_locode, "next_port_iso2": next_port_iso2, "next_port_name": next_port_name, "next_port_eta": next_port_eta, "next_port_latitude": next_port_latitude, "next_port_longitude": next_port_longitude, } print(f"字段提取结果: {datajson}") # 提交数据到接口 requests.post("https://{redacted by me}/api/Moments", json=datajson, timeout=10) elif r.status_code == 300: print(f"状态码300,URL: {url},待补充重定向处理逻辑") elif r.status_code == 404: print(f"状态码404,URL: {url},待补充失效链接处理逻辑") else: print(f"异常状态码{r.status_code},URL: {url}") except Exception as e: print(f"处理{url}失败,错误信息: {str(e)}") # 出错跳过当前URL,继续处理下一个 continue
后续新增待处理URL时,直接往myString列表里追加对应的字符串元素即可,不需要修改循环处理逻辑。
内容的提问来源于stack exchange,提问作者AlwaysResearching
相关产品推荐
相关产品推荐

