如何用Python正则表达式从URL提取Artifactory相关信息?
用Python正则高效提取Artifactory URL关键信息
你完全可以通过Python正则一次性提取出Artifactory实例名称、仓库名称和制品名称,无需多次拆分组合,利用正则捕获组就能直接搞定。
正则匹配思路
观察你提供的URL结构,共性是:
- 以
https://开头,后跟Artifactory实例域名(可能带端口,但示例中不需要端口) - 固定路径
/artifactory/作为分隔符 - 之后依次是仓库名称和制品的完整路径
基于这个结构,我们可以写出精准匹配的正则表达式。
代码实现
import re # 编译正则模式,三个捕获组分别对应实例、仓库、制品 url_pattern = re.compile(r"https://([^:/]+)(?::\d+)?/artifactory/([^/]+)/(.*)") # 测试示例URL test_urls = [ "https://artifactory.intuit.veg.com:443/artifactory/annual-budget-local/manifests-approved/1.0.0/annual-chart/po09ij/annual-f3c.tgz", "https://artifactory.skopeo.marvel.org/artifactory/bulletins_virtual/manifests-approved/po09ij/annual-f3c.tgz" ] for url in test_urls: match_result = url_pattern.match(url) if match_result: # 提取三个捕获组的内容 artifactory_instance = match_result.group(1) repo_name = match_result.group(2) artifact_name = match_result.group(3) print(f"Artifactory实例: {artifactory_instance}") print(f"仓库名称: {repo_name}") print(f"制品名称: {artifact_name}") print("---")
捕获组说明
([^:/]+):匹配Artifactory实例名称,[^:/]表示匹配除了:和/之外的所有字符,刚好避开端口号部分(?::\d+)?:可选的端口匹配,用?:标记为非捕获组,不影响我们提取目标内容([^/]+):匹配仓库名称,直到遇到下一个/为止(.*):匹配剩余的所有路径,也就是完整的制品名称
运行结果
Artifactory实例: artifactory.intuit.veg.com 仓库名称: annual-budget-local 制品名称: manifests-approved/1.0.0/annual-chart/po09ij/annual-f3c.tgz --- Artifactory实例: artifactory.skopeo.marvel.org 仓库名称: bulletins_virtual 制品名称: manifests-approved/po09ij/annual-f3c.tgz
这个方案一步到位,比多次split更简洁高效,也能准确适配你给出的URL格式。
内容的提问来源于stack exchange,提问作者Rafa S
相关产品推荐
相关产品推荐

