You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Java提取拼接词的单词序列?求开源库及实现方案

拼接词拆分相关问题解答

一、任务专业名称

这类拆分拼接词(比如驼峰式变量名、PascalCase命名)的任务,专业上叫做驼峰式拆分(Camel Case Splitting),更通用的术语是复合词分割(Compound Word Segmentation),用这两个关键词检索相关资料会更精准。

二、可用的Java开源库

  • Apache Commons Lang:Lang包中的StringUtils类提供了splitByCharacterTypeCamelCase方法,可直接拆分驼峰式命名。示例代码:
    import org.apache.commons.lang3.StringUtils;
    
    public class CamelSplitExample {
        public static void main(String[] args) {
            String input = "NullPointException";
            String[] parts = StringUtils.splitByCharacterTypeCamelCase(input);
            // 输出结果:["Null", "Point", "Exception"]
            for (String part : parts) {
                System.out.println(part);
            }
        }
    }
    
  • Guava:借助CaseFormat类转换命名风格间接实现拆分,比如将PascalCase转为空格分隔的字符串后拆分:
    import com.google.common.base.CaseFormat;
    
    public class GuavaSplitExample {
        public static void main(String[] args) {
            String input = "NullPointException";
            String spaced = CaseFormat.UPPER_CAMEL.to(CaseFormat.LOWER_SPACE, input);
            String[] parts = spaced.split(" ");
            // 输出结果:["null", "point", "exception"],可按需调整大小写
            for (String part : parts) {
                System.out.println(part);
            }
        }
    }
    

三、Lingpipe能否完成该任务

Lingpipe主打自然语言处理中的句子拆分、命名实体识别等场景,没有专门针对驼峰式拼接词的拆分工具。如果硬要用Lingpipe实现,需要自行训练模型或结合词典做匹配,成本远高于通用工具,不推荐。

四、Java实现类似Python wordninja的功能

wordninja核心是基于词典的动态规划算法,找出拼接词中概率最高的拆分路径。Java中可按以下思路实现:

  1. 准备英文单词词典(可从开源词典文件读取,或内置常用单词集合);
  2. 用动态规划从后往前计算每个位置的最优拆分,记录最大概率的拆分路径;
  3. 根据路径拆分出单词。

简单实现框架:

import java.util.*;

public class WordNinjaLike {
    private static final Set<String> DICTIONARY = new HashSet<>(Arrays.asList("null", "point", "exception")); // 可补充更多单词

    public static List<String> split(String input) {
        input = input.toLowerCase();
        int n = input.length();
        List<String>[] dp = new List[n + 1];
        dp[0] = new ArrayList<>();

        for (int i = 1; i <= n; i++) {
            for (int j = 0; j < i; j++) {
                String substring = input.substring(j, i);
                if (DICTIONARY.contains(substring) && dp[j] != null) {
                    List<String> temp = new ArrayList<>(dp[j]);
                    temp.add(substring);
                    if (dp[i] == null || temp.size() > dp[i].size()) {
                        dp[i] = temp;
                    }
                }
            }
        }
        return dp[n] != null ? dp[n] : Collections.singletonList(input);
    }

    public static void main(String[] args) {
        List<String> parts = split("NullPointException");
        System.out.println(parts);
    }
}

也可直接寻找开源项目中的同类型自定义工具,核心逻辑均围绕词典与动态规划展开。

内容的提问来源于stack exchange,提问作者Link

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 14:21:02