如何使用Python从谷歌云存储URL中提取bucket与blob字段值
实现方案
方案1:字符串分割实现(新手友好,无需正则)
针对固定结构的GCS存储URL,直接用字符串分割就能快速拿到结果,理解成本更低:
url = "https://storage.cloud.google.com/test_bucket/first_test_user/0.jpg" # 按/分割URL得到数组 url_parts = url.split("/") # 数组索引从0开始计数,第4个元素就是第3个/和第4个/之间的bucket内容 bucket = url_parts[3] # 拼接第4个/之后的路径,再去掉末尾的.jpg后缀 path_after_4th_slash = "/".join(url_parts[4:]) blob = path_after_4th_slash.removesuffix(".jpg") # 验证输出 print(bucket) # test_bucket print(blob) # first_test_user/0
如果使用Python3.9以下版本,不支持removesuffix方法,可以替换为:
blob = path_after_4th_slash[:-4]
方案2:正则表达式实现
如果需要适配更多URL变体,用正则的捕获组可以直接提取目标内容:
import re url = "https://storage.cloud.google.com/test_bucket/first_test_user/0.jpg" # 正则规则:第一个捕获组拿bucket,第二个捕获组拿.jpg之前的blob内容 pattern = r"https://storage\.cloud\.google\.com/([^/]+)/(.+)\.jpg" match_result = re.fullmatch(pattern, url) if match_result: bucket = match_result.group(1) blob = match_result.group(2) # 验证输出 print(bucket) # test_bucket print(blob) # first_test_user/0
正则规则说明:
[^/]+:匹配所有不包含/的字符,刚好匹配不允许带斜杠的bucket名称(.+)\.jpg:匹配任意字符直到遇到.jpg后缀为止,拿到完整的blob路径
内容的提问来源于stack exchange,提问作者Bernardo
相关产品推荐
相关产品推荐

