Java中如何去除字符串指定前缀WordMatch(content=和末尾)字符
字符串清理实现方案
核心逻辑
- 固定前缀
WordMatch(content=长度为18位,直接从第18位开始截取内容 - 去除字符串末尾最后一个
)字符即可得到目标HTML内容 - 也可使用正则表达式一次性匹配提取中间内容,适配前后存在多余空格等异常场景
各常用后端语言实现示例
Python 实现
def clean_html(raw_str): # 方法1:字符串截取,性能更高 prefix = "WordMatch(content=" if raw_str.startswith(prefix) and raw_str.endswith(")"): return raw_str[len(prefix):-1] # 方法2:正则提取,兼容性更强 import re match_res = re.match(r"^WordMatch\(content=(.*)\)$", raw_str, re.DOTALL) return match_res.group(1) if match_res else raw_str
Node.js 实现
function cleanHtml(rawStr) { // 方法1:字符串截取 const prefix = "WordMatch(content=" if (rawStr.startsWith(prefix) && rawStr.endsWith(")")) { return rawStr.slice(prefix.length, -1) } // 方法2:正则提取 const matchRes = rawStr.match(/^WordMatch\(content=([\s\S]*)\)$/) return matchRes ? matchRes[1] : rawStr }
PHP 实现
function clean_html($raw_str) { $prefix = "WordMatch(content="; // 方法1:字符串截取 if (str_starts_with($raw_str, $prefix) && str_ends_with($raw_str, ")")) { return substr($raw_str, strlen($prefix), -1); } // 方法2:正则提取 if (preg_match('/^WordMatch\(content=(.*)\)$/s', $raw_str, $matches)) { return $matches[1]; } return $raw_str; }
额外注意事项
如果清理后的内容包含</>这类HTML实体字符,需要额外做实体解码后再写入文件,Python示例如下:
import html raw_input_str = "你的原始待处理字符串" cleaned_str = clean_html(raw_input_str) final_html = html.unescape(cleaned_str) # 写入本地文件 with open("output.html", "w", encoding="utf-8") as f: f.write(final_html)
内容的提问来源于stack exchange,提问作者Evan Gertis
相关产品推荐
相关产品推荐

