C++解析二进制文件面试任务求助:提取日期、时长等信息
Hey there! It sounds like you’ve made a solid start extracting phone numbers and the "TO" tag—great work. Let’s walk through how to unlock the rest of the data in those binary files:
1. First, Lock Down the File/Record Structure
Your hunch about identifying header markers like 01 02 03 04 is spot-on. Here’s how to nail this:
- Grab your sample call record and open multiple entries (if you have them) in Hex Fiend. Look for repeating byte sequences that act as record separators or entry headers. These patterns will let you split the raw binary into discrete call records, so you don’t have to parse the entire file as one blob.
- Once you spot a consistent marker, note its length and position. For example, if every call record starts with
0x01020304, you can iterate through the byte array and split whenever you hit this 4-byte sequence.
2. Decode Non-ASCII Fields (Dates, Durations)
Those unreadable bytes aren’t garbage—they’re almost certainly binary-encoded numerical data. Here’s how to crack them:
- Date/Time: Most call logs use either:
- A Unix timestamp: A 4-byte (or 8-byte) integer counting seconds/milliseconds since 1970. Take a known call date from your sample, convert it to a Unix timestamp, then search Hex Fiend for those bytes (remember to check endianness—little-endian is common on most systems, so the least significant byte comes first).
- Packed BCD: Binary-Coded Decimal, where each nibble (4 bits) represents a digit. For example,
0x240515might mean 2024-05-15. Use your sample date to cross-reference these packed values.
- Duration: This is usually a small integer (2-byte or 4-byte) representing seconds or milliseconds. Match a known call duration (e.g., a 2-minute call = 120 seconds) to the corresponding hex bytes in your sample record.
3. Refine Your Data Extraction Logic
Instead of replacing all non-printable ASCII with spaces, try targeted extraction once you have the record structure:
- Once you’ve mapped a record’s layout (e.g.,
[4-byte header][11-byte phone number][1-byte "TO" tag][4-byte timestamp][2-byte duration]), extract specific byte ranges directly instead of filtering the whole file. This preserves the raw bytes you need to decode dates/durations. - For example, in C++, read the file into a
std::vector<uint8_t>, then for each record, pull the timestamp bytes starting at offset X and convert them to an integer:// Example: Convert 4-byte little-endian timestamp to time_t uint8_t* timestamp_ptr = &raw_data[record_start + 16]; // Adjust offset as needed uint32_t timestamp = (timestamp_ptr[3] << 24) | (timestamp_ptr[2] << 16) | (timestamp_ptr[1] << 8) | timestamp_ptr[0]; time_t call_time = static_cast<time_t>(timestamp);
4. Cross-Validate with Your Sample
This is the most critical step. Take your known sample call record (with confirmed phone number, date, duration) and map every piece of known info to its exact byte positions in Hex Fiend. Create a "record template" that you can reuse for every entry in the file. This template will eliminate guesswork and let you parse consistent fields reliably.
Quick Recap of Your Action Plan
- Identify record separators/headers via Hex Fiend comparison.
- Map known fields (phone number, "TO") to their byte offsets in a single record.
- Use sample data to reverse-engineer how dates and durations are encoded (timestamp vs BCD, endianness, byte length).
- Build a C++ parser that reads each record per your template, converting raw bytes to usable data types.
内容的提问来源于stack exchange,提问作者Elementah Official

