You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中如何去除字符串指定前缀WordMatch(content=和末尾)字符

字符串清理实现方案

核心逻辑

  • 固定前缀WordMatch(content=长度为18位,直接从第18位开始截取内容
  • 去除字符串末尾最后一个)字符即可得到目标HTML内容
  • 也可使用正则表达式一次性匹配提取中间内容,适配前后存在多余空格等异常场景

各常用后端语言实现示例

Python 实现

def clean_html(raw_str):
    # 方法1:字符串截取,性能更高
    prefix = "WordMatch(content="
    if raw_str.startswith(prefix) and raw_str.endswith(")"):
        return raw_str[len(prefix):-1]
    # 方法2:正则提取,兼容性更强
    import re
    match_res = re.match(r"^WordMatch\(content=(.*)\)$", raw_str, re.DOTALL)
    return match_res.group(1) if match_res else raw_str

Node.js 实现

function cleanHtml(rawStr) {
    // 方法1:字符串截取
    const prefix = "WordMatch(content="
    if (rawStr.startsWith(prefix) && rawStr.endsWith(")")) {
        return rawStr.slice(prefix.length, -1)
    }
    // 方法2:正则提取
    const matchRes = rawStr.match(/^WordMatch\(content=([\s\S]*)\)$/)
    return matchRes ? matchRes[1] : rawStr
}

PHP 实现

function clean_html($raw_str) {
    $prefix = "WordMatch(content=";
    // 方法1:字符串截取
    if (str_starts_with($raw_str, $prefix) && str_ends_with($raw_str, ")")) {
        return substr($raw_str, strlen($prefix), -1);
    }
    // 方法2:正则提取
    if (preg_match('/^WordMatch\(content=(.*)\)$/s', $raw_str, $matches)) {
        return $matches[1];
    }
    return $raw_str;
}

额外注意事项

如果清理后的内容包含</>这类HTML实体字符,需要额外做实体解码后再写入文件,Python示例如下:

import html
raw_input_str = "你的原始待处理字符串"
cleaned_str = clean_html(raw_input_str)
final_html = html.unescape(cleaned_str)
# 写入本地文件
with open("output.html", "w", encoding="utf-8") as f:
    f.write(final_html)

内容的提问来源于stack exchange,提问作者Evan Gertis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 22:24:03