You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java解析.prn文件转换为HTML时排版异常问题求解

PRN固定宽度文件转HTML对齐异常解决方案

问题核心

你之前两种方案失效的根本原因是完全没用到PRN「固定列宽」的核心属性:

  • 按2个及以上空白拆分的方案:只要字段内容本身带连续空格、或者短字段间填充空格不足2个,就会拆错列,和源文件对应不上
  • 按单空格拆分直接塞HTML的方案:浏览器默认会折叠所有连续空白,再加上默认用非等宽字体,排版必然错乱,和控制台等宽环境下的显示效果完全是两回事

可落地实现方案

根据你的输出需求二选一即可:

方案1:零解析成本 原生格式还原

不需要做任何列拆分,直接利用HTML的<pre>标签保留原始文本的空白和排版,搭配等宽字体就能1:1还原PRN的原始显示效果,实现成本最低:

package main;

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.FileWriter;
import java.io.IOException;

public class PrnToHtml {
    public static void main(String[] args) {
        try (BufferedReader reader = new BufferedReader(new FileReader(".\\Workbook2.prn"));
             FileWriter writer = new FileWriter("output_prn.html")) {
            // 写入HTML头部,配置等宽字体,用pre标签保留原始格式
            writer.write("""
                <html>
                <head>
                    <style>pre { font-family: Consolas, Monaco, monospace; font-size: 14px; }</style>
                </head>
                <body>
                    <h3>PRN 转换结果</h3>
                    <pre>
                """);
            String currentLine;
            while ((currentLine = reader.readLine()) != null) {
                // 转义HTML特殊字符,避免<>&等字符破坏页面结构
                String escaped = currentLine.replace("&", "&amp;")
                        .replace("<", "&lt;")
                        .replace(">", "&gt;");
                writer.write(escaped + "\n");
            }
            writer.write("""
                    </pre>
                </body>
                </html>
                """);
        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}
  • 优点:实现极简,100%还原源文件排版,不会出现列错位
  • 缺点:输出是纯文本块,不是结构化表格,后续如果要做列排序、筛选类交互会比较麻烦

方案2:固定列宽识别 结构化表格输出

如果你需要输出标准HTML表格做后续处理,不要靠空格拆分,先扫描所有行确定每一列的起始和结束位置,再按固定位置切割每一行的字段,从根源避免拆分错误:

package main;

import java.io.BufferedReader;
import java.io.FileReader;
import java.io.FileWriter;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;

public class PrnToHtmlTable {
    public static void main(String[] args) {
        List<String> allLines = new ArrayList<>();
        int maxLineLen = 0;
        // 第一步:读取所有行,记录最长行长度
        try (BufferedReader reader = new BufferedReader(new FileReader(".\\Workbook2.prn"))) {
            String line;
            while ((line = reader.readLine()) != null) {
                allLines.add(line);
                maxLineLen = Math.max(maxLineLen, line.length());
            }
        } catch (IOException e) {
            e.printStackTrace();
            return;
        }

        // 第二步:识别列分割位置(整列全为空格的位置就是列间隙)
        boolean[] isAllSpace = new boolean[maxLineLen];
        for (int i = 0; i < maxLineLen; i++) {
            isAllSpace[i] = true;
            for (String line : allLines) {
                if (i < line.length() && line.charAt(i) != ' ') {
                    isAllSpace[i] = false;
                    break;
                }
            }
        }

        // 提取每一列的起止下标
        List<int[]> colRanges = new ArrayList<>();
        int colStart = -1;
        for (int i = 0; i < maxLineLen; i++) {
            if (!isAllSpace[i] && colStart == -1) {
                colStart = i;
            } else if (isAllSpace[i] && colStart != -1) {
                colRanges.add(new int[]{colStart, i});
                colStart = -1;
            }
        }
        if (colStart != -1) {
            colRanges.add(new int[]{colStart, maxLineLen});
        }

        // 第三步:按列范围切割每一行,输出HTML表格
        try (FileWriter writer = new FileWriter("output_prn_table.html")) {
            writer.write("""
                <html>
                <head>
                    <style>
                        table { border-collapse: collapse; font-family: sans-serif; }
                        td { border: 1px solid #ccc; padding: 4px 8px; white-space: nowrap; }
                    </style>
                </head>
                <body>
                    <h3>PRN 转结构化表格</h3>
                    <table>
                """);
            for (String line : allLines) {
                writer.write("<tr>");
                for (int[] range : colRanges) {
                    int start = range[0], end = range[1];
                    String field;
                    if (start >= line.length()) {
                        field = "";
                    } else {
                        field = line.substring(start, Math.min(end, line.length())).trim();
                    }
                    // 转义特殊字符
                    field = field.replace("&", "&amp;")
                            .replace("<", "&lt;")
                            .replace(">", "&gt;");
                    writer.write("<td>" + field + "</td>");
                }
                writer.write("</tr>\n");
            }
            writer.write("""
                    </table>
                </body>
                </html>
                """);
        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}
  • 优点:输出标准结构化HTML表格,列对齐完全准确,支持后续表格交互操作
  • 注意:如果PRN存在跨列的合并表头,在列识别完成后手动调整列范围即可

避坑提醒

  • 所有嵌入HTML的原始文本都要做HTML特殊字符转义,避免源文件里的<、>、&字符破坏页面结构
  • 如果不用<pre>标签,一定要给对应容器加white-space: pre样式,同时搭配等宽字体,否则连续空格还是会被浏览器折叠
  • 永远不要靠空格数量拆分PRN文件,固定宽度文件的拆分依据是「字符位置」不是「分隔符」

内容的提问来源于stack exchange,提问作者Vijay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 09:03:34