You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python从谷歌云存储URL中提取bucket与blob字段值

实现方案

方案1:字符串分割实现(新手友好,无需正则)

针对固定结构的GCS存储URL,直接用字符串分割就能快速拿到结果,理解成本更低:

url = "https://storage.cloud.google.com/test_bucket/first_test_user/0.jpg"
# 按/分割URL得到数组
url_parts = url.split("/")
# 数组索引从0开始计数,第4个元素就是第3个/和第4个/之间的bucket内容
bucket = url_parts[3]
# 拼接第4个/之后的路径,再去掉末尾的.jpg后缀
path_after_4th_slash = "/".join(url_parts[4:])
blob = path_after_4th_slash.removesuffix(".jpg")

# 验证输出
print(bucket)  # test_bucket
print(blob)  # first_test_user/0

如果使用Python3.9以下版本,不支持removesuffix方法,可以替换为:

blob = path_after_4th_slash[:-4]

方案2:正则表达式实现

如果需要适配更多URL变体,用正则的捕获组可以直接提取目标内容:

import re

url = "https://storage.cloud.google.com/test_bucket/first_test_user/0.jpg"
# 正则规则:第一个捕获组拿bucket,第二个捕获组拿.jpg之前的blob内容
pattern = r"https://storage\.cloud\.google\.com/([^/]+)/(.+)\.jpg"
match_result = re.fullmatch(pattern, url)
if match_result:
    bucket = match_result.group(1)
    blob = match_result.group(2)

# 验证输出
print(bucket)  # test_bucket
print(blob)  # first_test_user/0

正则规则说明:

  • [^/]+:匹配所有不包含/的字符,刚好匹配不允许带斜杠的bucket名称
  • (.+)\.jpg:匹配任意字符直到遇到.jpg后缀为止,拿到完整的blob路径

内容的提问来源于stack exchange,提问作者Bernardo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 00:39:00