如何在Scala/Java中匹配以#开头的单词?正则表达式求解
正确提取以#开头单词的正则表达式方案
你之前用split的思路不对,split是通过匹配分隔符拆分字符串,而你需要的是直接提取符合规则的目标内容,应该用正则的匹配查找功能,而非分割。
正确正则表达式
使用#\w+即可,规则解释:
#:匹配开头的井号\w+:匹配后续一个或多个字母、数字、下划线(\w等价于[a-zA-Z0-9_])
Scala 实现代码
替换原代码的正则定义和提取逻辑,用findAllIn方法收集所有匹配项:
val regex = "#\\w+".r val text = "#shouldMatch1 #shouldMatch2 notMatch nope#shouldMatch3 nooope()#shouldMatch4" regex.findAllIn(text).toList shouldBe List("#shouldMatch1", "#shouldMatch2", "#shouldMatch3", "#shouldMatch4")
Java 实现代码
Java 中使用Pattern和Matcher遍历匹配结果:
import java.util.ArrayList; import java.util.List; import java.util.regex.Matcher; import java.util.regex.Pattern; public class Main { public static void main(String[] args) { Pattern pattern = Pattern.compile("#\\w+"); String text = "#shouldMatch1 #shouldMatch2 notMatch nope#shouldMatch3 nooope()#shouldMatch4"; Matcher matcher = pattern.matcher(text); List<String> matches = new ArrayList<>(); while (matcher.find()) { matches.add(matcher.group()); } // 输出结果应为 [#shouldMatch1, #shouldMatch2, #shouldMatch3, #shouldMatch4] System.out.println(matches); } }
为什么之前的方案失效?
你用的[^#\\w+]是一个否定字符集,匹配**不是#、字母、数字、下划线、+**的字符,用它做split分隔符时,会把这些字符作为拆分点,但非目标内容(比如notMatch)里没有符合该字符集的内容,所以会被完整保留下来,导致结果不符合预期。
内容的提问来源于stack exchange,提问作者Hejwo
相关产品推荐
相关产品推荐

