You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式提取字符串中的非恒定字符段?

解决方法

这里提供两种简单的方式来获取目标字符串:

方法一:正则替换

直接构造正则匹配需要移除的片段,将其替换为空,一步得到结果。

正则表达式可以写成:

^archives\/latest\/pipelines|\/pages

替换规则是将匹配到的内容替换为空字符串。

以Python为例:

import re
original_str = "archives/latest/pipelines/my-page/pages/content/page1/"
result = re.sub(r'^archives\/latest\/pipelines|\/pages', '', original_str)
# 输出结果: /my-page/content/page1/

或者用捕获组精准提取需保留的部分:

^archives\/latest\/pipelines(\/[^\/]+)\/pages(.*)

替换为\1\2(不同语言捕获组引用符号可能有差异,比如Python可用\1或\g<1>):

import re
original_str = "archives/latest/pipelines/my-page/pages/content/page1/"
match = re.match(r'^archives\/latest\/pipelines(\/[^\/]+)\/pages(.*)', original_str)
if match:
    result = match.group(1) + match.group(2)
# 输出结果: /my-page/content/page1/

方法二:字符串分割拼接

无需正则,直接通过字符串操作分步处理:

  1. 先移除开头的固定前缀archives/latest/pipelines/
  2. 再将剩余字符串中的/pages/替换为/

以Python为例:

original_str = "archives/latest/pipelines/my-page/pages/content/page1/"
# 第一步:移除开头前缀
temp_str = original_str.replace("archives/latest/pipelines/", "")
# 第二步:替换pages/部分
result = temp_str.replace("/pages/", "/")
# 输出结果: /my-page/content/page1/

内容的提问来源于stack exchange,提问作者Shijith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 13:45:09