如何将numpy int16矩阵转换为Android/Java可读的低占用格式?
Hey there! Let's work through this problem—you've got a 7000×2048 int16 NumPy matrix that's only ~40MB in its native binary form, but CSV/JSON blow it up way too much for your Android app. Here are the most practical, space-efficient formats that work smoothly with Android/Java:
1. Raw Binary (Smallest Possible Size)
This is the closest you can get to the original NumPy disk size—no extra overhead, just the raw int16 bytes. The only catch is you need to track the matrix dimensions (rows/cols) separately (either in a tiny side file or embedded directly in the binary).
Python Export
Save the raw bytes directly from NumPy, with embedded dimensions for convenience:
import numpy as np import struct matrix = np.random.randint(-32768, 32767, size=(7000, 2048), dtype=np.int16) # Write 2 int32 values (rows, cols) first, then the raw int16 bytes with open("matrix_with_dims.raw", "wb") as f: f.write(struct.pack("<ii", matrix.shape[0], matrix.shape[1])) matrix.tofile(f)
Android/Java Reading
Use ByteBuffer for fast, direct file access:
import java.io.RandomAccessFile; import java.nio.ByteBuffer; import java.nio.ByteOrder; public short[][] readRawMatrix(String filePath) throws Exception { RandomAccessFile file = new RandomAccessFile(filePath, "r"); ByteBuffer buffer = ByteBuffer.allocate((int) file.length()); buffer.order(ByteOrder.LITTLE_ENDIAN); // Match NumPy's default endianness file.getChannel().read(buffer); buffer.flip(); // Read dimensions first int rows = buffer.getInt(); int cols = buffer.getInt(); // Read the int16 matrix short[][] matrix = new short[rows][cols]; for (int i = 0; i < rows; i++) { for (int j = 0; j < cols; j++) { matrix[i][j] = buffer.getShort(); } } file.close(); return matrix; }
Pros: Tiny file size (~40MB), ultra-fast read/write.
Cons: No built-in structure—you have to handle dimensions manually.
2. Protocol Buffers (Protobuf)
Google's cross-platform binary format is perfect for structured, type-safe data with minimal overhead. It's slightly bigger than raw binary but adds schema validation and easy maintainability.
Step 1: Define a Protobuf Schema
Create a matrix.proto file:
syntax = "proto3"; message Int16Matrix { int32 rows = 1; int32 cols = 2; repeated sint16 data = 3; // sint16 is optimized for signed integers }
Python Export
Compile the proto, then serialize your matrix:
import numpy as np from matrix_pb2 import Int16Matrix matrix = np.random.randint(-32768, 32767, size=(7000, 2048), dtype=np.int16) proto_matrix = Int16Matrix() proto_matrix.rows = matrix.shape[0] proto_matrix.cols = matrix.shape[1] proto_matrix.data.extend(matrix.flatten().tolist()) # Save to file with open("matrix.protobin", "wb") as f: f.write(proto_matrix.SerializeToString())
Android/Java Reading
Add the Protobuf dependency to your build.gradle, compile the proto, then read:
import com.yourpackage.Int16Matrix; import java.io.FileInputStream; public short[][] readProtobufMatrix(String filePath) throws Exception { Int16Matrix protoMatrix = Int16Matrix.parseFrom(new FileInputStream(filePath)); int rows = protoMatrix.getRows(); int cols = protoMatrix.getCols(); short[][] matrix = new short[rows][cols]; int idx = 0; for (int i = 0; i < rows; i++) { for (int j = 0; j < cols; j++) { matrix[i][j] = (short) protoMatrix.getData(idx++); } } return matrix; }
Pros: Structured, type-safe, cross-platform, good compression (~45-50MB).
Cons: Requires defining a schema and compiling proto files.
3. MsgPack
A binary alternative to JSON that's much smaller and faster. It supports NumPy types directly with the right libraries, so no need to flatten to lists manually.
Python Export
Use msgpack-numpy for seamless NumPy serialization:
import numpy as np import msgpack_numpy as mpack matrix = np.random.randint(-32768, 32767, size=(7000, 2048), dtype=np.int16) # Serialize (includes shape and dtype info automatically) packed = mpack.packb(matrix) with open("matrix.msgpack", "wb") as f: f.write(packed)
Android/Java Reading
Use the msgpack-java library to deserialize:
import org.msgpack.core.MessagePack; import org.msgpack.core.MessageUnpacker; import java.io.FileInputStream; public short[][] readMsgPackMatrix(String filePath) throws Exception { MessageUnpacker unpacker = MessagePack.newDefaultUnpacker(new FileInputStream(filePath)); // MsgPack-numpy stores arrays as a map with "shape", "dtype", "data" unpacker.unpackMapHeader(); int rows = 0, cols = 0; short[] flatData = null; while (unpacker.hasNext()) { String key = unpacker.unpackString(); switch (key) { case "shape": unpacker.unpackArrayHeader(); rows = unpacker.unpackInt(); cols = unpacker.unpackInt(); break; case "data": int dataLen = unpacker.unpackArrayHeader(); flatData = new short[dataLen]; for (int i = 0; i < dataLen; i++) { flatData[i] = unpacker.unpackShort(); } break; case "dtype": // Skip dtype if you don't need validation unpacker.unpackString(); break; } } // Reshape into 2D array short[][] matrix = new short[rows][cols]; int idx = 0; for (int i = 0; i < rows; i++) { System.arraycopy(flatData, idx, matrix[i], 0, cols); idx += cols; } return matrix; }
Pros: No schema needed, smaller than JSON (~50-60MB), supports NumPy types natively.
Cons: Slightly less structured than Protobuf, requires handling NumPy's map format.
4. Apache Parquet
A columnar storage format designed for efficient compression and analytics. It's great if you might need to access only parts of the matrix later (e.g., specific columns).
Python Export
Use pyarrow to write the matrix to Parquet:
import numpy as np import pyarrow as pa import pyarrow.parquet as pq matrix = np.random.randint(-32768, 32767, size=(7000, 2048), dtype=np.int16) # Convert to PyArrow table (one column per matrix column) table = pa.table({"col_" + str(i): matrix[:, i] for i in range(matrix.shape[1])}) pq.write_table(table, "matrix.parquet")
Android/Java Reading
Use the arrow-android library to read the Parquet file:
import org.apache.arrow.memory.RootAllocator; import org.apache.arrow.vector.IntVector; import org.apache.arrow.vector.VectorSchemaRoot; import org.apache.arrow.vector.ipc.ArrowReader; import org.apache.arrow.vector.ipc.file.ArrowFileReader; import java.io.FileInputStream; public short[][] readParquetMatrix(String filePath) throws Exception { RootAllocator allocator = new RootAllocator(Long.MAX_VALUE); ArrowReader reader = new ArrowFileReader(new FileInputStream(filePath), allocator); VectorSchemaRoot root = reader.getVectorSchemaRoot(); reader.loadNextBatch(); int rows = root.getRowCount(); int cols = root.getFieldVectors().size(); short[][] matrix = new short[rows][cols]; for (int j = 0; j < cols; j++) { IntVector colVector = (IntVector) root.getVector(j); for (int i = 0; i < rows; i++) { matrix[i][j] = (short) colVector.get(i); } } reader.close(); allocator.close(); return matrix; }
Pros: Excellent compression (~40-45MB), column-wise access support, ideal for large datasets.
Cons: Slightly more complex setup on Android, overkill if you always need the full matrix.
Quick Recommendation
- Smallest size + fastest speed: Raw binary (with embedded dimensions)
- Structured & type-safe: Protocol Buffers
- No schema needed: MsgPack
- Partial column access: Apache Parquet
内容的提问来源于stack exchange,提问作者A. Sam

