You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java动态获取PRN文件列宽实现通用分块解析方案咨询

PRN固定宽度文件动态列宽解析实现

现有PRN解析逻辑通过硬编码方式传入列宽参数,无法适配不同PRN文件的列数、列宽差异,需要实现列宽自动识别,输出整数数组形式的列宽集合,完成通用解析。

原有实现代码

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.FileWriter;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;

public class PrnParser {

    public void parsePrn(String inputReader, String outputWriter) {

        try (BufferedReader br = new BufferedReader(new FileReader("..//data.prn"))); {
            String line;
            while ((line = br.readLine()) != null) {

                if (line.trim().equals("")) {
                    continue;
                }
                // 手动传入固定列宽做分块解析
                String[] parts = splitStringToChunks(line, 16, 22, 9, 14, 13, 8);
                for (String str : parts) {
                    System.out.println(str);
                }
                System.out.println();
            }

        } catch (IOException ex) {
            System.out.println(ex.getMessage());
        }
    }

    public String[] splitStringToChunks(String inputString, Integer... chunkSizes) {
        List<String> list = new ArrayList<>();
        int chunkStart = 0, chunkEnd = 0;
        for (int length : chunkSizes) {
            chunkStart = chunkEnd;
            chunkEnd = chunkStart + length;
            String dataChunk = inputString.substring(chunkStart, chunkEnd);
            list.add(dataChunk.trim());
        }
        return list.toArray(new String[0]);
    }

}

原有代码存在的问题

  • 列宽参数硬编码,仅能适配单种格式的PRN文件,无法通用
  • 注意:存在隐藏运行bug:try-with-resources声明资源的右括号后多写了分号,会导致BufferedReader在进入逻辑块前就被关闭,读取数据时直接抛出IO异常
  • 未对行长度不足的场景做容错,调用substring截取时容易触发索引越界异常

动态列宽识别实现方案

PRN是典型的固定宽度文本格式,列与列之间通过连续空格分隔,只需要扫描第一个非空表头行的连续空格分隔位置,即可反推出每一列的宽度,核心逻辑如下:

  • 跳过文件开头的空行,读取第一个非空行作为表头识别基准
  • 遍历表头字符,识别连续空格段的起止位置,两个相邻非空内容段之间的长度即为对应列的宽度
  • 单独处理最后一列,裁剪行尾多余空格,避免列宽计算偏大

列宽自动识别方法代码

/**
 * 从表头行自动识别每列宽度
 * @param headerLine PRN文件的表头行
 * @return 列宽整数数组
 */
public Integer[] detectColumnWidths(String headerLine) {
    List<Integer> widths = new ArrayList<>();
    int currentColStart = 0;
    boolean inSpaceSegment = false;
    int spaceStartIdx = -1;

    for (int i = 0; i < headerLine.length(); i++) {
        char c = headerLine.charAt(i);
        if (Character.isWhitespace(c)) {
            if (!inSpaceSegment) {
                inSpaceSegment = true;
                spaceStartIdx = i;
            }
        } else {
            if (inSpaceSegment) {
                int colWidth = spaceStartIdx - currentColStart;
                if (colWidth > 0) {
                    widths.add(colWidth);
                }
                currentColStart = i;
                inSpaceSegment = false;
            }
        }
    }

    // 处理最后一列,裁剪行尾多余空格
    if (currentColStart < headerLine.length()) {
        int lastColEnd = headerLine.length();
        while (lastColEnd > currentColStart && Character.isWhitespace(headerLine.charAt(lastColEnd - 1))) {
            lastColEnd--;
        }
        widths.add(lastColEnd - currentColStart);
    }

    return widths.toArray(new Integer[0]);
}

改造后的完整解析逻辑

需要新增java.util.Arrays导入,改造后的解析方法如下:

public void parsePrn(String prnFilePath) {
    try (BufferedReader br = new BufferedReader(new FileReader(prnFilePath))) {
        String line;
        Integer[] columnWidths = null;

        // 跳过开头空行,读取表头识别列宽
        while ((line = br.readLine()) != null) {
            if (line.trim().isEmpty()) {
                continue;
            }
            columnWidths = detectColumnWidths(line);
            break;
        }

        if (columnWidths == null || columnWidths.length == 0) {
            throw new IOException("未识别到有效PRN文件表头");
        }

        // 计算总列宽,用于行长度容错
        int totalWidth = Arrays.stream(columnWidths).mapToInt(Integer::intValue).sum();

        // 逐行解析数据
        while ((line = br.readLine()) != null) {
            if (line.trim().isEmpty()) {
                continue;
            }
            // 长度不足的行末尾补空格,避免substring越界
            if (line.length() < totalWidth) {
                line = String.format("%-" + totalWidth + "s", line);
            }
            String[] parts = splitStringToChunks(line, columnWidths);
            for (String colVal : parts) {
                System.out.println(colVal);
            }
            System.out.println();
        }

    } catch (IOException ex) {
        System.out.println(ex.getMessage());
    }
}

可选优化点

  • 多行列宽校准:如果遇到表头和数据行存在轻微偏移的PRN文件,可以读取前3~5行非空数据,联合计算列分隔位置的交集,提升识别准确率
  • 自定义分隔符支持:部分PRN文件可能用制表符、特殊字符做列分隔,可以扩展识别逻辑支持自定义分隔规则
  • 大文件优化:GB级大文件场景下不需要加载全量内容,仅读取开头前几行即可完成列宽识别,IO开销极低

内容的提问来源于stack exchange,提问作者Vijay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 21:15:37