You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用StringTokenizer检测行中连续分隔符并填充空值

处理带连续制表符的.dat文件字段拆分问题

我逐行读取.dat文件,想用制表符\t作为分隔符拆分字段,但部分非必填字段为空,连续制表符对应的空字符串需要被识别并保留。当前用StringTokenizer的代码无法处理这种场景,导致字段数量不一致,丢失空值。

当前代码

StringTokenizer stringTokenizer = new StringTokenizer(line, "\t");
ArrayList<String> al = new ArrayList<>();
while (stringTokenizer.hasMoreTokens()) {
    al.add(stringTokenizer.nextToken());
}
System.out.println(al.size() + " >> " + al);

输入行示例

R	900081458	22222-22-2			1	-1	1	0	0	1
R	245047685	7250-46-6			0	-1	0	0	0	0
R	245048731	13755-29-8	237-340-6	0	-1	0	0	0	0
R	245047201	1080-12-2	214-096-9	0	-1	0	0	0	0
R	1	118725-24-9	612-118-00-5	405-080-4	0	0	0	0	0	0

当前输出(丢失空值,字段数不符)

9 >> [R, 900081458, 22222-22-2, 1, -1, 1, 0, 0, 1]
9 >> [R, 245047685, 7250-46-6, 0, -1, 0, 0, 0, 0]
10 >> [R, 245048731, 13755-29-8, 237-340-6, 0, -1, 0, 0, 0, 0]
10 >> [R, 245047201, 1080-12-2, 214-096-9, 0, -1, 0, 0, 0, 0]
11 >> [R, 1, 118725-24-9, 612-118-00-5, 405-080-4, 0, 0, 0, 0, 0, 0]

期望输出(保留空值,字段数一致,示例用"BLANK"标识空字符串)

11 >> [R, 900081458, 22222-22-2, "BLANK", "BLANK", 1, -1, 1, 0, 0, 1]
11 >> [R, 245047685, 7250-46-6, "BLANK", "BLANK", 0, -1, 0, 0, 0, 0]
11 >> [R, 245048731, 13755-29-8, 237-340-6, "BLANK", 0, -1, 0, 0, 0, 0]
11 >> [R, 245047201, 1080-12-2, 214-096-9, "BLANK", 0, -1, 0, 0, 0, 0]
11 >> [R, 1, 118725-24-9, 612-118-00-5, 405-080-4, 0, 0, 0, 0, 0, 0]

问题原因

StringTokenizer会忽略连续的分隔符,不会生成对应的空字符串,因此无法满足保留空字段的需求。

解决方案

方法1:使用String.split方法

String.split支持正则表达式,设置第二个参数为负数可保留末尾空字符串,同时识别连续分隔符之间的空值。

String[] fields = line.split("\t", -1);
ArrayList<String> al = new ArrayList<>(Arrays.asList(fields));

// 将空字符串替换为"BLANK"
for (int i = 0; i < al.size(); i++) {
    if (al.get(i).isEmpty()) {
        al.set(i, "\"BLANK\"");
    }
}

// 补全字段数至11个
while (al.size() < 11) {
    al.add("\"BLANK\"");
}

System.out.println(al.size() + " >> " + al);

方法2:使用Apache Commons CSV库(专业TSV处理)

若允许引入第三方库,Apache Commons CSV可完美处理TSV格式,自动保留空字段。

先添加Maven依赖:

<dependency>
    <groupId>org.apache.commons</groupId>
    <artifactId>commons-csv</artifactId>
    <version>1.10.0</version>
</dependency>

代码示例:

CSVFormat format = CSVFormat.TDF.withHeader().withSkipHeaderRecord();
try (CSVParser parser = CSVParser.parse(new File("your-file.dat"), StandardCharsets.UTF_8, format)) {
    for (CSVRecord record : parser) {
        ArrayList<String> al = new ArrayList<>();
        // 遍历补全所有11个字段
        for (int i = 0; i < 11; i++) {
            String value = record.get(i);
            al.add(value.isEmpty() ? "\"BLANK\"" : value);
        }
        System.out.println(al.size() + " >> " + al);
    }
} catch (IOException e) {
    e.printStackTrace();
}

方法3:手动遍历字符串拆分(无依赖)

不想用正则或第三方库时,可手动逐字符处理,识别连续制表符并添加空字符串:

ArrayList<String> al = new ArrayList<>();
StringBuilder sb = new StringBuilder();

for (char c : line.toCharArray()) {
    if (c == '\t') {
        al.add(sb.toString());
        sb.setLength(0);
    } else {
        sb.append(c);
    }
}
// 添加最后一个字段
al.add(sb.toString());

// 替换空字符串并补全字段数
for (int i = 0; i < al.size(); i++) {
    if (al.get(i).isEmpty()) {
        al.set(i, "\"BLANK\"");
    }
}
while (al.size() < 11) {
    al.add("\"BLANK\"");
}

System.out.println(al.size() + " >> " + al);

内容的提问来源于stack exchange,提问作者Vasilis Iak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 02:55:17