求助:如何用R语言的readBin读取指定二进制文件?
I get it, translating that NumPy binary read logic over to R can feel a bit tricky at first—let's break this down step by step so you can replicate exactly what your Python code does.
First, let's map your NumPy dtype to R's equivalent types for readBin:
np.uint32→ 32-bit unsigned integer: useinteger()withsize = 4andsigned = FALSEnp.uint16→ 16-bit unsigned integer: useinteger()withsize = 2andsigned = FALSEnp.int16→ 16-bit signed integer: useinteger()withsize = 2andsigned = TRUE(this is the default, but it's good to be explicit)
Each record in your binary file takes up 12 bytes total (4 + 2 + 2 + 2 + 2), so we'll use that to calculate how many records we need to read.
Here's the complete R code to replicate your Python workflow:
# Set your filename filename <- "filename.DAT" # Open a connection to the binary file con <- file(filename, "rb") # Calculate total number of records: file size / bytes per record file_size <- file.info(filename)$size bytes_per_record <- 4 + 2 + 2 + 2 + 2 # 12 bytes total num_records <- file_size %/% bytes_per_record # Read each variable in order, matching the dtype specification var1 <- readBin(con, what = integer(), n = num_records, size = 4, signed = FALSE) var2 <- readBin(con, what = integer(), n = num_records, size = 2, signed = FALSE) var3 <- readBin(con, what = integer(), n = num_records, size = 2, signed = TRUE) var4 <- readBin(con, what = integer(), n = num_records, size = 2, signed = TRUE) var5 <- readBin(con, what = integer(), n = num_records, size = 2, signed = TRUE) # Close the file connection close(con) # Combine into a data frame (or tibble if you prefer) data <- data.frame( var1 = var1, var2 = var2, var3 = var3, var4 = var4, var5 = var5 ) # Optional: If you want to verify, check the first few rows head(data)
Key Notes:
- The
rbmode infile()opens the file for binary reading—critical for this task. - We calculate
num_recordsto make sure we read exactly the right number of entries for each variable. - The
signedparameter is crucial for unsigned types (var1andvar2)—if you omit it, R will treat them as signed integers, which could give wrong values for numbers larger than 2^31-1 (for uint32) or 2^15-1 (for uint16). - If you're working with very large files, you might want to read in chunks instead of all at once, but for most cases, this approach works smoothly.
Quick Verification Tip
If you have a small test file, you can export a sample from Python and compare with R's output to ensure consistency. For example, in Python:
import numpy as np dt = np.dtype([('var1', np.uint32), ('var2', np.uint16), ('var3', np.int16), ('var4', np.int16), ('var5', np.int16)]) test_data = np.array([(123456, 65535, -32768, 1000, -500)], dtype=dt) test_data.tofile("test.DAT")
Then read it in R with the code above—you should get the exact same values.
内容的提问来源于stack exchange,提问作者Jehol

