You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java HTML剥离方法修复:如何处理单独的<或>字符?

修正HTML内容清理方法以处理单独的<或>字符

你的原始代码如下:

//Method for Strip HTML
public static String stripHtml(String inStr) {
  boolean inTag = false;
  char c;
  StringBuffer outStr = new StringBuffer();
  int len = inStr.length();
  for (int i = 0; i < len; i++) {
    c = inStr.charAt(i);
    if (c == '&lt;') {
      inTag = true;
    }
    if (!inTag) {
      outStr.append(c);
    }
    if (c == '&gt;') {
      inTag = false;
    }
  }
  //Print to show that the this method is removing the necessary characters
  System.out.println(outStr);
  return outStr.toString();
}

核心问题分析

  1. 判断条件错误:代码中用c == '&lt;'和c == '&gt;'判断字符,这是HTML转义字符的写法,但Java字符串中<和>是单个字符,直接写'<'和'>'即可。原条件导致实际遇到</>时无法触发标签判断,所以单独的<会被保留在输出中。
  2. 逻辑缺陷:当遇到未闭合的<(无对应>)时,inTag会一直保持true,导致<之后的所有字符都无法添加到结果,不符合你保留后续字符的需求。

修改后的代码

调整遍历逻辑为索引跳转模式,同时修正判断条件,确保单独的</>也能被正确移除:

//Method for Strip HTML
public static String stripHtml(String inStr) {
  StringBuffer outStr = new StringBuffer();
  int len = inStr.length();
  int i = 0;
  while (i < len) {
    char c = inStr.charAt(i);
    if (c == '<') {
      // 查找当前<之后的第一个>位置
      int endTagIndex = inStr.indexOf('>', i);
      if (endTagIndex != -1) {
        // 跳过整个<...>区间
        i = endTagIndex + 1;
      } else {
        // 无闭合>,仅跳过当前<
        i++;
      }
    } else if (c == '>') {
      // 单独出现的>直接跳过
      i++;
    } else {
      // 正常字符添加到结果
      outStr.append(c);
      i++;
    }
  }
  System.out.println(outStr);
  return outStr.toString();
}

验证效果

  • 输入app<html>le → 输出apple(符合期望)
  • 输入app<le → 输出apple(符合期望)
  • 输入app>le → 输出apple(额外支持单独>的处理)

内容的提问来源于stack exchange,提问作者Pedro Aguilar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 14:55:20