Java中0、1字符串转byte[]最优方法及文件读取代码优化问询
Hey there! Let's break down your two Java questions and figure out the most efficient solutions for you.
The core idea here is to process the binary string in chunks of 8 characters (since one byte holds 8 bits). We also need to handle cases where the string length isn't a multiple of 8—usually by padding with leading zeros (adjust this based on your specific needs if you require trailing zeros instead).
Here's a clean, efficient implementation:
public static byte[] binaryStringToByteArray(String binaryStr) { // Calculate the padded length to ensure it's a multiple of 8 int strLength = binaryStr.length(); int paddedLength = (strLength + 7) / 8 * 8; // Pad with leading zeros if needed String paddedBinary = String.format("%" + paddedLength + "s", binaryStr).replace(' ', '0'); byte[] byteArray = new byte[paddedLength / 8]; for (int i = 0; i < paddedBinary.length(); i += 8) { // Extract 8-character chunk String byteChunk = paddedBinary.substring(i, i + 8); // Convert chunk to byte: use Integer.parseInt for unsigned range (0-255), then cast to byte byteArray[i / 8] = (byte) Integer.parseInt(byteChunk, 2); } return byteArray; }
Key optimizations:
- Batch processing: We process 8 characters at a time, minimizing loop iterations.
- Efficient padding:
String.formatprovides a concise way to pad the string to the required length. - Handling unsigned bytes: By using
Integer.parseInt(which handles 0-255) before casting tobyte, we avoid issues with Java's signed byte range (-128 to 127) if you need unsigned values.
Your existing code uses a BufferedReader but loads the entire file into a string first, which can be inefficient (and memory-heavy) for large files. Let's optimize this with two approaches depending on your file size:
Option 1: Stream processing (ideal for large files)
This approach processes the file line-by-line and character-by-character, without loading the entire file into memory:
import java.nio.file.Files; import java.nio.file.Paths; import java.nio.charset.StandardCharsets; import java.io.BufferedReader; import java.util.ArrayList; import java.util.List; public void transform(String inputFile) { List<Byte> byteList = new ArrayList<>(); StringBuilder currentByteChunk = new StringBuilder(8); // Pre-size to 8 for efficiency try (BufferedReader br = Files.newBufferedReader(Paths.get(inputFile), StandardCharsets.UTF_8)) { String line; while ((line = br.readLine()) != null) { String cleanedLine = line.trim(); // Remove any whitespace in the line for (char c : cleanedLine.toCharArray()) { if (c != '0' && c != '1') { throw new IllegalArgumentException("Invalid character found: " + c); } currentByteChunk.append(c); // Convert to byte once we have 8 characters if (currentByteChunk.length() == 8) { byteList.add((byte) Integer.parseInt(currentByteChunk.toString(), 2)); currentByteChunk.setLength(0); // Reset the builder } } } // Handle any remaining characters (pad with leading zeros if needed) if (currentByteChunk.length() > 0) { while (currentByteChunk.length() < 8) { currentByteChunk.insert(0, '0'); // Pad leading zeros; use append() for trailing } byteList.add((byte) Integer.parseInt(currentByteChunk.toString(), 2)); } // Convert List<Byte> to byte[] byte[] result = new byte[byteList.size()]; for (int i = 0; i < result.length; i++) { result[i] = byteList.get(i); } // Do something with your result here (e.g., write to a file) } catch (Exception e) { // Replace with proper logging or custom exception handling in production code e.printStackTrace(); } }
Option 2: One-time read (ideal for small files)
If your file is small enough to fit in memory, this is a more concise approach that reuses the method from Question 1:
import java.nio.file.Files; import java.nio.file.Paths; import java.nio.charset.StandardCharsets; public void transform(String inputFile) { try { // Read all content, remove whitespace, then convert String binaryContent = new String(Files.readAllBytes(Paths.get(inputFile)), StandardCharsets.UTF_8) .replaceAll("\\s+", ""); // Remove any spaces/newlines byte[] result = binaryStringToByteArray(binaryContent); // Process the result here } catch (Exception e) { e.printStackTrace(); } }
Key improvements over your original code:
- Auto-resource management: Uses try-with-resources to automatically close the
BufferedReader, preventing resource leaks. - Memory efficiency: Stream processing avoids loading the entire file into memory, which is critical for large files.
- Input validation: Adds checks for invalid characters (non-0/1) to catch bad input early.
- Flexibility: Two options to fit different file size scenarios.
内容的提问来源于stack exchange,提问作者Sirp

