You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scala/Java中匹配以#开头的单词?正则表达式求解

正确提取以#开头单词的正则表达式方案

你之前用split的思路不对,split是通过匹配分隔符拆分字符串,而你需要的是直接提取符合规则的目标内容,应该用正则的匹配查找功能,而非分割。

正确正则表达式

使用#\w+即可,规则解释:

  • #:匹配开头的井号
  • \w+:匹配后续一个或多个字母、数字、下划线(\w等价于[a-zA-Z0-9_])

Scala 实现代码

替换原代码的正则定义和提取逻辑,用findAllIn方法收集所有匹配项:

val regex = "#\\w+".r

val text = "#shouldMatch1 #shouldMatch2 notMatch nope#shouldMatch3 nooope()#shouldMatch4"

regex.findAllIn(text).toList shouldBe List("#shouldMatch1", "#shouldMatch2", "#shouldMatch3", "#shouldMatch4")

Java 实现代码

Java 中使用Pattern和Matcher遍历匹配结果:

import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class Main {
    public static void main(String[] args) {
        Pattern pattern = Pattern.compile("#\\w+");
        String text = "#shouldMatch1 #shouldMatch2 notMatch nope#shouldMatch3 nooope()#shouldMatch4";
        Matcher matcher = pattern.matcher(text);
        
        List<String> matches = new ArrayList<>();
        while (matcher.find()) {
            matches.add(matcher.group());
        }
        
        // 输出结果应为 [#shouldMatch1, #shouldMatch2, #shouldMatch3, #shouldMatch4]
        System.out.println(matches);
    }
}

为什么之前的方案失效?

你用的[^#\\w+]是一个否定字符集,匹配**不是#、字母、数字、下划线、+**的字符,用它做split分隔符时,会把这些字符作为拆分点,但非目标内容(比如notMatch)里没有符合该字符集的内容,所以会被完整保留下来,导致结果不符合预期。

内容的提问来源于stack exchange,提问作者Hejwo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 05:18:22