Bash脚本URL字符串处理:移除IP后未知数量无关路径段
解决URL路径精简的正则方案
嘿,我来帮你搞定这个折腾了你两小时的正则问题!核心思路其实很直接:先精准抓住URL的协议+IP/域名部分,然后跳过中间所有未知数量的dontcare路径段,直接把keyword及后面的内容接在域名后面就行。
正则表达式 & 替换逻辑
基础正则(固定keyword)
用这个正则匹配目标URL,然后替换成指定格式:
- 正则模式:
^(https?://[^/]+)/.*?/(keyword/.*)$ - 替换模板:
$1/$2(不同语言里可能用\1/\2,比如Python)
代码示例(Python)
import re # 示例URL original_url = "https://192.168.3.10/dontcareA/dontcareB/dontcareX/keyword/user/profile?id=123" # 正则模式 pattern = r"^(https?://[^/]+)/.*?/(keyword/.*)$" # 执行替换 processed_url = re.sub(pattern, r"\1/\2", original_url) print(processed_url) # 输出结果:https://192.168.3.10/keyword/user/profile?id=123
正则部分详解
让我拆解一下每个部分的作用,方便你理解和调整:
^(https?://[^/]+):捕获组1,匹配URL开头的协议(http/https)加上IP/域名部分。[^/]+表示匹配到第一个/之前的所有字符,完美适配任意格式的IP或域名。/.*?/:非贪婪匹配中间的任意路径段。.*?会尽可能少地匹配字符,直到遇到后面的/(keyword/.*),这样不管中间有多少个dontcare段,都会被跳过。(keyword/.*)$:捕获组2,从keyword开始,匹配到URL结尾的所有内容(包括后续路径、参数等)。
动态keyword的适配
如果你的keyword是可变的,可以用正则转义处理特殊字符,避免匹配出错:
import re target_keyword = "my-custom-keyword" original_url = "https://x.xx.xxx.xxx/dontcare1/dontcare2/my-custom-keyword/rest/of/string" # 用re.escape转义keyword里的特殊字符 pattern = rf"^(https?://[^/]+)/.*?/({re.escape(target_keyword)}/.*)$" processed_url = re.sub(pattern, r"\1/\2", original_url) print(processed_url) # 输出:https://x.xx.xxx.xxx/my-custom-keyword/rest/of/string
内容的提问来源于stack exchange,提问作者Daniele Foti
相关产品推荐
相关产品推荐

