如何用sed和awk从十六进制文件提取指定长度的英文字符串?
Absolutely! You can pull this off with a combination of sed and awk (or even just awk for a cleaner, more straightforward solution). The core idea is first converting the hex data to readable ASCII, then filtering out only the words that match your criteria—5 to 25 characters long, containing only A-Z or a-z. Let’s walk through a couple of practical approaches:
Approach 1: Using awk (Recommended for Simplicity)
If your hex file is a raw stream of hex characters (no extra formatting like addresses or spaces), start by converting it to ASCII with xxd, then use awk to snip out valid words:
xxd -r -p your_hex_file | awk '{ while (match($0, /[A-Za-z]{5,25}/, matches)) { print matches[0] $0 = substr($0, RSTART + RLENGTH) } }'
Quick breakdown:
xxd -r -p: Converts raw hex data to plain ASCII (the-pflag skips line numbering and extra formatting clutter).- The
awkloop: Scans each line of decoded ASCII, finds every sequence of 5-25 letters, prints it, then removes that sequence from the line to hunt for the next match.
If your hex file is in hexdump format (with address columns like 00000000 48 65 6C...), first strip out non-hex parts with sed before decoding:
sed 's/^[0-9A-Fa-f]* //; s/ .*//; s/ //g' your_hexdump_file | xxd -r -p | awk '{ while (match($0, /[A-Za-z]{5,25}/, matches)) { print matches[0] $0 = substr($0, RSTART + RLENGTH) } }'
Breakdown of the sed pre-processing:
s/^[0-9A-Fa-f]* //: Removes the leading address column from each line.s/ .*//: Trims the ASCII preview section at the end of hexdump lines.s/ //g: Removes spaces between hex pairs to create a clean raw hex stream.
Approach 2: Using sed (More Verbose, But Doable)
While sed is less intuitive for multiple matches per line, you can still make it work with a loop:
xxd -r -p your_hex_file | sed -n ' :loop s/\([A-Za-z]\{5,25\}\)/\1\n/g t loop s/[^A-Za-z\n]//g /^[A-Za-z]\{5,25\}$/p '
Quick breakdown:
:loop+t loop: Repeatedly replaces valid word sequences with themselves plus a newline, forcing every match onto its own line.s/[^A-Za-z\n]//g: Deletes all non-letter, non-newline characters to clean up noise./^[A-Za-z]\{5,25\}$/p: Prints only lines that are exactly 5-25 letters long, filtering out any leftover short sequences.
Both methods ensure no hex content is printed—only the valid words you’re targeting.
内容的提问来源于stack exchange,提问作者user7249767

