You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用requests库批量处理URL列表数据的实现方法

Python批量处理URL列表实现方案

你的代码问题出在两点:一是URL列表定义格式错误,二是没有用循环结构遍历列表执行处理逻辑,按以下方式修改即可:

  • 修正URL列表写法:你注释的myString存在两个语法错误:列表末尾误用大括号}闭合,且多个URL被拼接在同一个字符串内。正确写法是每个URL作为独立的字符串元素,用方括号[]包裹列表,元素间用逗号分隔。
  • 用for循环遍历URL列表,把你已经调试通的单URL处理逻辑放到循环体内,每次循环取一个URL发起请求即可。
  • 补充两个必要的容错处理:第一,仅当响应状态码为200时才执行内容解析,3xx/4xx/其他状态码直接走你预留的分支逻辑,避免非成功响应的页面结构不匹配,导致字符串分割报错;第二,增加请求超时设置和异常捕获,避免单个URL请求失败直接中断整个批量任务。

修改后可直接运行的完整代码:

import requests

# 正确定义URL列表,每个URL为独立字符串元素
myString = [
    'https://{redacted by me}/vessels/amadi_9682552_10003796/',
    'https://{redacted by me}/vessels/akebono-maru_9554729_2866687/',
    'https://{redacted by me}/vessels/amani_9661869_9276632/',
    'https://{redacted by me}/vessels/aman-sendai_9134323_2017277/',
    'https://{redacted by me}/vessels/al-aamriya_9338266_25273/'
]

# 遍历列表批量处理
for url in myString:
    print(f"当前处理: {url}")
    try:
        # 设置10秒超时,避免请求无限等待
        r = requests.get(url, timeout=10)
        if r.status_code == 200:
            print(f"请求成功,状态码: {r.status_code}")
            data = r.content.decode("utf-8")
            
            # 提取目标字段
            next_port_locode = data.split('locode: "')[1].split('"')[0].strip()
            next_port_iso2 = data.split('iso2: "')[1].split('"')[0].strip()
            next_port_name = data.split('iso2: "')[1].split('name: "')[1].split('"')[0].strip()
            next_port_eta = data.split('eta: moment("')[1].split('"')[0].strip()
            next_port_latitude = float(data.split('latitude: ')[1].split(',')[0].strip())
            next_port_longitude = float(data.split('longitude: ')[1].split('\n')[0].strip())

            datajson = {
                "next_port_locode": next_port_locode,
                "next_port_iso2": next_port_iso2,
                "next_port_name": next_port_name,
                "next_port_eta": next_port_eta,
                "next_port_latitude": next_port_latitude,
                "next_port_longitude": next_port_longitude,
            }
            print(f"字段提取结果: {datajson}")
            # 提交数据到接口
            requests.post("https://{redacted by me}/api/Moments", json=datajson, timeout=10)
        elif r.status_code == 300:
            print(f"状态码300,URL: {url},待补充重定向处理逻辑")
        elif r.status_code == 404:
            print(f"状态码404,URL: {url},待补充失效链接处理逻辑")
        else:
            print(f"异常状态码{r.status_code},URL: {url}")
    except Exception as e:
        print(f"处理{url}失败,错误信息: {str(e)}")
        # 出错跳过当前URL,继续处理下一个
        continue

后续新增待处理URL时,直接往myString列表里追加对应的字符串元素即可,不需要修改循环处理逻辑。

内容的提问来源于stack exchange,提问作者AlwaysResearching

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 05:31:00