写入MJPEG二进制文件后读取异常:实际数据与预期长度不符
我通过Servlet的POST方法接收Motion JPEG(MJPEG)流,提取流中的JPEG图像并写入MJPEG文件,文件结构如下:
--END Content-Type: image/jpeg Content-Length: <Length of following JPEG image in bytes> <JPEG image data> --END
每个图像以--END开头,上一个图像的--END后开始下一个图像。
文件写入的Java实现代码如下:
private static final String BOUNDARY = "END"; private static final Charset CHARSET = StandardCharsets.UTF_8; final byte[] PRE_BOUNDARY = ("--" + BOUNDARY).getBytes(CHARSET); final byte[] CONTENT_TYPE = "Content-Type: image/jpeg".getBytes(CHARSET); final byte[] CONTENT_LENGTH = "Content-Length: ".getBytes(CHARSET); final byte[] LINE_FEED = "\n".getBytes(CHARSET); try (OutputStream out = new BufferedOutputStream(new FileOutputStream(targetFile))) { for every image received { onImage(image, out); } } public void onImage(byte[] image, OutputStream out) { try { out.write(PRE_BOUNDARY); out.write(LINE_FEED); out.write(CONTENT_TYPE); out.write(LINE_FEED); out.write(CONTENT_LENGTH); out.write(String.valueOf(image.length).getBytes(CHARSET)); out.write(LINE_FEED); out.write(LINE_FEED); out.write(image); out.write(LINE_FEED); out.write(LINE_FEED); } catch (IOException e) { e.printStackTrace(); } }
现在我需要读取该MJPEG文件并处理其中的图像,为此编写了如下读取代码:
package de.supportgis.stream; import java.io.BufferedReader; import java.io.FileInputStream; import java.io.FileNotFoundException; import java.io.IOException; import java.io.InputStream; import java.io.InputStreamReader; import java.io.Reader; import java.nio.charset.Charset; import java.nio.charset.StandardCharsets; public class MediaConverter { private static final String BOUNDARY = "END"; private static final Charset CHARSET = StandardCharsets.UTF_8; private static final int READ_PRE_BOUNDARY = 1; private static final int READ_CONTENT_TYPE = 2; private static final int READ_CONTENT_LENGTH = 3; private static final int READ_CONTENT = 4; public static void createMovieFromMJPEG(String file) throws FileNotFoundException, IOException { char LINE_FEED = '\n'; char[] PRE_BOUNDARY = new String("--" + BOUNDARY + LINE_FEED).toCharArray(); try (InputStream in = new FileInputStream(file); Reader reader = new InputStreamReader(in, CHARSET); Reader buffer = new BufferedReader(reader)) { int r; StringBuffer content_buf = new StringBuffer(); int mode = READ_PRE_BOUNDARY; long content_length = 0; int[] cmdBuf = new int[PRE_BOUNDARY.length]; int boundaryPointer = 0; int counter = 0; while ((r = reader.read()) != -1) { System.out.print((char)r); counter++; if (mode == READ_PRE_BOUNDARY) { if (r == PRE_BOUNDARY[boundaryPointer]) { boundaryPointer++; if (boundaryPointer >= PRE_BOUNDARY.length - 1) { // Read a PRE_BOUNDARY mode = READ_CONTENT_TYPE; boundaryPointer = 0; } } } else if (mode == READ_CONTENT_TYPE) { if (r != LINE_FEED) { content_buf.append((char)r); } else { if (content_buf.length() == 0) { // leading line break, ignore... } else { mode = READ_CONTENT_LENGTH; content_buf.setLength(0); } } } else if (mode == READ_CONTENT_LENGTH) { if (r != LINE_FEED) { content_buf.append((char)r); } else { if (content_buf.length() == 0) { // leading line break, ignore... } else { String number = content_buf.substring(content_buf.lastIndexOf(":") + 1).trim(); content_length = Long.valueOf(number); content_buf.setLength(0); mode = READ_CONTENT; } } } else if (mode == READ_CONTENT) { char[] fileBuf = new char[(int)content_length]; reader.read(fileBuf); System.out.println(fileBuf); mode = READ_PRE_BOUNDARY; } } } } public static void main(String[] args) { try { createMovieFromMJPEG("video.mjpeg"); } catch (IOException e) { e.printStackTrace(); } } }
目前该读取器尚未能生成可用的JPEG图像,我遇到的问题是:读取了Content-Length字段指定的值,期望读取对应长度的字节以获取完整图像数据,但实际读取到的fileBuf中包含了完整图像、后续元数据以及下一张图像的部分字节,读取数据量远超预期。请问我在二进制数据的读写和编码处理中犯了什么错误?
你的代码核心错误在于用字符流(Reader)处理二进制图像数据,以及读取逻辑中的几个关键疏漏,具体如下:
1. 字符流与二进制流的误用
JPEG是纯二进制格式数据,而InputStreamReader这类字符流会把字节按照指定编码(这里是UTF-8)转换为字符,这个过程会直接破坏二进制数据:
- UTF-8对部分字节序列有特定解析规则,遇到无法解析的字节会替换为
�(替换字符),导致图像字节丢失或篡改。 - 字符流的
read()方法返回的是字符的Unicode码点,而非原始字节值,把这些字符转回字节时无法还原原始图像数据。 - 用
char[]存储图像数据,本质是把二进制字节当成字符处理,完全不符合JPEG的二进制存储要求。
修正方案:全程使用字节流(InputStream)处理文件,仅在解析文本头(Boundary、Content-Type、Content-Length)时临时将字节转换为字符串。
2. 读取Content-Length后的换行未处理
根据你的写入逻辑,Content-Length行之后有两个换行符,才会开始图像数据。你的读取代码在解析完Content-Length后直接进入READ_CONTENT模式,没有跳过这两个换行的字节,导致读取的图像数据开头包含了换行字节,后续读取位置偏移,把本该属于元数据和下一张图的边界也读进来。
3. reader.read(fileBuf)不能保证读取到指定长度
字符流的read(char[])方法不一定会填满整个数组,它返回实际读取的字符数。你直接假设它会读取content_length个字符,不仅会导致数据读取不完整,更关键的是二进制字节和字符是1:N或N:1的对应关系,用字符数匹配字节数本身就是错误逻辑。
修正后的读取代码示例
package de.supportgis.stream; import java.io.BufferedInputStream; import java.io.FileInputStream; import java.io.IOException; import java.io.InputStream; import java.nio.charset.StandardCharsets; public class MediaConverter { private static final String BOUNDARY = "END"; private static final byte[] PRE_BOUNDARY = ("--" + BOUNDARY + "\n").getBytes(StandardCharsets.UTF_8); private static final byte[] CONTENT_TYPE_PREFIX = "Content-Type: image/jpeg".getBytes(StandardCharsets.UTF_8); private static final byte[] CONTENT_LENGTH_PREFIX = "Content-Length: ".getBytes(StandardCharsets.UTF_8); private static final byte LINE_FEED = '\n'; private static final int READ_PRE_BOUNDARY = 1; private static final int READ_CONTENT_TYPE = 2; private static final int READ_CONTENT_LENGTH = 3; private static final int READ_CONTENT = 4; public static void createMovieFromMJPEG(String file) throws IOException { try (InputStream in = new BufferedInputStream(new FileInputStream(file))) { int mode = READ_PRE_BOUNDARY; long contentLength = 0; int boundaryPointer = 0; byte[] lineBuffer = new byte[1024]; int lineIndex = 0; while (true) { int b = in.read(); if (b == -1) break; switch (mode) { case READ_PRE_BOUNDARY: if (b == PRE_BOUNDARY[boundaryPointer]) { boundaryPointer++; if (boundaryPointer == PRE_BOUNDARY.length) { boundaryPointer = 0; mode = READ_CONTENT_TYPE; } } else { boundaryPointer = 0; } break; case READ_CONTENT_TYPE: if (b == LINE_FEED) { String line = new String(lineBuffer, 0, lineIndex, StandardCharsets.UTF_8).trim(); if (line.startsWith(new String(CONTENT_TYPE_PREFIX, StandardCharsets.UTF_8))) { mode = READ_CONTENT_LENGTH; } lineIndex = 0; } else if (lineIndex < lineBuffer.length - 1) { lineBuffer[lineIndex++] = (byte) b; } break; case READ_CONTENT_LENGTH: if (b == LINE_FEED) { String line = new String(lineBuffer, 0, lineIndex, StandardCharsets.UTF_8).trim(); if (line.startsWith(new String(CONTENT_LENGTH_PREFIX, StandardCharsets.UTF_8))) { String lengthStr = line.substring(CONTENT_LENGTH_PREFIX.length).trim(); contentLength = Long.parseLong(lengthStr); // 跳过Content-Length后的空行(两个换行) in.read(); mode = READ_CONTENT; } lineIndex = 0; } else if (lineIndex < lineBuffer.length - 1) { lineBuffer[lineIndex++] = (byte) b; } break; case READ_CONTENT: // 循环读取保证获取完整的图像字节 byte[] imageData = new byte[(int) contentLength]; int bytesRead = 0; while (bytesRead < contentLength) { int read = in.read(imageData, bytesRead, (int) (contentLength - bytesRead)); if (read == -1) break; bytesRead += read; } // 此处可添加图像处理逻辑,比如保存为单独JPEG文件 // saveImage(imageData, "output_" + System.currentTimeMillis() + ".jpg"); // 跳过图像后的两个换行 in.read(); in.read(); mode = READ_PRE_BOUNDARY; break; } } } } // 示例:保存图像数据到文件 // private static void saveImage(byte[] imageData, String filePath) throws IOException { // try (OutputStream out = new FileOutputStream(filePath)) { // out.write(imageData); // } // } public static void main(String[] args) { try { createMovieFromMJPEG("video.mjpeg"); } catch (IOException e) { e.printStackTrace(); } } }
关键修正点说明
- 全程使用
InputStream字节流处理文件,仅在解析文本头时临时转换为字符串,保证二进制数据不被破坏。 - 严格处理每个阶段的换行符,包括Content-Length后、图像数据后的空行,确保读取位置准确。
- 读取图像数据时用循环保证读取到指定长度的字节,避免单次读取不足导致的数据不完整。
- 边界匹配改为字节级对比,避免字符转换带来的错误。
内容的提问来源于stack exchange,提问作者user3792852

