You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写正则匹配除HTML标签(含</>编码)外的文本适配Ruby gsub

实现方案

因为Ruby正则不支持*SKIP/*F专属语法,我们采用优先匹配标签直接保留、剩余内容做自定义处理的逻辑实现需求,可同时兼容普通<...>标签和&amp;lt;...&amp;gt;格式的编码标签。

替换标签外空格的代码示例

# 你的输入文本
input = 'Hi &amp;lt;p class="hello"&amp;gt;
there, how are you
&amp;lt;/p&amp;gt; <span class="test">other content</span>'

result = input.gsub(/(?:<[^>]*>|&amp;lt;(?:(?!&amp;gt;).)*&amp;gt;)|[^<&]+/) do |match|
  # 匹配到标签直接原样返回,不修改内部空格
  if match.start_with?('<') || match.start_with?('&amp;lt;')
    match
  else
    # 非标签内容,按需替换空格,示例为删除所有空格,可自定义替换规则
    match.gsub(/\s+/, '')
  end
end

# 输出结果:Hi&amp;lt;p class="hello"&amp;gt;there,howareyou&amp;lt;/p&amp;gt;<span class="test">othercontent</span>
puts result

仅提取标签外文本的代码示例

如果不需要替换,只需要提取所有标签外的文本片段,可以用scan方法实现:

extracted = input.scan(/(?:<[^>]*>|&amp;lt;(?:(?!&amp;gt;).)*&amp;gt;)|([^<&]+)/).flatten.compact
# 提取结果为 ["Hi ", "\nthere, how are you\n", "other content"]

注意事项

如果你的编码标签是单层转义格式(文本里直接显示为&lt;而非&amp;lt;),只需要把正则中的&amp;lt;替换为&lt;、&amp;gt;替换为&gt;即可。该方案适合结构规范的类HTML文本,极端场景(如属性值内嵌标签边界符号)可根据实际情况调整正则的匹配规则。

内容的提问来源于stack exchange,提问作者raquelhortab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 12:15:04