You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从字节数组填充与提取数据:多PDF字节数组拆分方法咨询

拆分包含多个PDF的byte[]数组为单独PDF字节数组

嘿,这个问题其实可以利用PDF文件的格式特性轻松解决——每个标准PDF都有明确的起始和结束标识,咱们就靠这个把合并的字节数组拆分开来。

核心思路:利用PDF的固定标识

标准PDF文件有两个关键特征:

  • 起始标记:所有PDF都以%PDF-开头(对应的ASCII字节是 37, 80, 68, 70, 45)
  • 结束标记:所有PDF都以%%EOF结尾(对应的ASCII字节是 37, 37, 69, 79, 70),注意部分PDF可能在%%EOF后面跟着换行符(ASCII 10)或回车符(ASCII 13),处理时要考虑这种情况。

具体实现步骤(以Java为例)

下面是一段可直接参考的代码逻辑,它会遍历目标字节数组,定位每个PDF的起止位置并截取:

import java.util.ArrayList;
import java.util.List;

public class PdfSplitter {
    public static List<byte[]> splitMultiplePdfs(byte[] allBytes) {
        List<byte[]> pdfList = new ArrayList<>();
        if (allBytes == null || allBytes.length == 0) {
            return pdfList;
        }

        byte[] startMarker = "%PDF-".getBytes();
        byte[] endMarker = "%%EOF".getBytes();
        int currentIndex = 0;
        int totalLength = allBytes.length;

        while (currentIndex <= totalLength - startMarker.length) {
            // 寻找下一个PDF的起始位置
            int startIndex = findMarkerIndex(allBytes, startMarker, currentIndex);
            if (startIndex == -1) {
                break; // 没有更多PDF了
            }

            // 从起始位置开始寻找结束标记
            int endIndex = findMarkerIndex(allBytes, endMarker, startIndex);
            if (endIndex == -1) {
                break; // 找不到结束标记,终止处理
            }
            // 调整结束索引,包含%%EOF以及可能的后续换行/回车
            endIndex += endMarker.length;
            // 检查后面是否有换行/回车,一并包含(可选,根据实际情况调整)
            while (endIndex < totalLength && (allBytes[endIndex] == 10 || allBytes[endIndex] == 13)) {
                endIndex++;
            }

            // 截取当前PDF的字节数组
            byte[] singlePdf = new byte[endIndex - startIndex];
            System.arraycopy(allBytes, startIndex, singlePdf, 0, singlePdf.length);
            pdfList.add(singlePdf);

            // 更新当前索引,继续寻找下一个PDF
            currentIndex = endIndex;
        }

        return pdfList;
    }

    // 辅助方法:从指定起始位置开始寻找标记字节数组的索引
    private static int findMarkerIndex(byte[] source, byte[] marker, int startPos) {
        int sourceLen = source.length;
        int markerLen = marker.length;

        for (int i = startPos; i <= sourceLen - markerLen; i++) {
            boolean match = true;
            for (int j = 0; j < markerLen; j++) {
                if (source[i + j] != marker[j]) {
                    match = false;
                    break;
                }
            }
            if (match) {
                return i;
            }
        }
        return -1;
    }

    // 测试用例
    public static void main(String[] args) {
        // 假设allBytes是包含多个PDF的字节数组
        byte[] allBytes = ...;
        List<byte[]> pdfs = splitMultiplePdfs(allBytes);
        System.out.println("拆分出 " + pdfs.size() + " 个PDF文件");
    }
}

注意事项

  • 如果你的场景中PDF之间夹杂了其他无关字节,可能需要额外的校验逻辑(比如验证截取后的字节数组是否是有效的PDF)
  • 部分非标准PDF可能存在标识不规范的情况,这时候可以结合PDF的其他结构特征(比如交叉引用表)来辅助验证,但绝大多数场景下,上述方法足够解决问题

内容的提问来源于stack exchange,提问作者user2848242

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:27:57