Java正则表达式需求:匹配无中文整行,排除含中文行
解决方案
你当前的正则错误地排除了部分标点符号,且未启用多行模式来准确匹配整行,导致无法筛选出完全不含中文的行。以下是调整后的实现方案:
方案1:通过替换移除含中文的行
使用正则匹配所有包含中文字符的整行并替换为空,再清理多余空行:
String example = "He is 45 today\n你今天45岁了\n这个句子没有数字\nOk I see"; // 替换所有含中文的整行为空 String tempResult = example.replaceAll("(?m)^.*\\p{IsHan}.*$", ""); // 移除空行并整理格式 String finalResult = tempResult.replaceAll("(?m)^\\s*$", "").trim(); System.out.println(finalResult);
方案2:直接提取不含中文的行
通过Pattern和Matcher直接查找并提取符合要求的行:
import java.util.regex.Matcher; import java.util.regex.Pattern; public class Main { public static void main(String[] args) { String example = "He is 45 today\n你今天45岁了\n这个句子没有数字\nOk I see"; Pattern pattern = Pattern.compile("(?m)^[^\\p{IsHan}]+$"); Matcher matcher = pattern.matcher(example); StringBuilder resultBuilder = new StringBuilder(); while (matcher.find()) { resultBuilder.append(matcher.group()).append("\n"); } // 移除末尾多余换行 String finalResult = resultBuilder.length() > 0 ? resultBuilder.substring(0, resultBuilder.length() - 1) : ""; System.out.println(finalResult); } }
正则关键说明
(?m):启用多行模式,让^匹配每行开头、$匹配每行结尾^.*\p{IsHan}.*$:匹配包含至少一个中文字符的整行(用于移除)^[^\p{IsHan}]+$:匹配完全不含中文字符的非空整行(用于提取)
两种方案都能实现你的需求:保留He is 45 today和Ok I see,排除所有含中文的行。
内容的提问来源于stack exchange,提问作者Hasen
相关产品推荐
相关产品推荐

