使用Unix命令(无第三方工具)从XML文件检索指定内容
Got it, let's break this down into two straightforward solutions using only standard Unix tools—no third-party libraries or apps required. I’ll walk you through each part with concrete examples.
Standard Unix utilities like grep, sed, and awk can handle basic XML field extraction, even without proper XML parsers. Here are common approaches:
Example: Extract all <NAME> field values
Using grep (Perl-compatible regex for precise matching)
grep -oP '(?<=<NAME>).*?(?=</NAME>)' XML_file.xml
-ooutputs only the matched part-Penables Perl regex (supports lookbehind/lookahead to avoid capturing the tags themselves)(?<=<NAME>)= positive lookbehind (start after opening tag),(?=</NAME>)= positive lookahead (end before closing tag)
Using sed (for more control over multi-line content)
If your XML fields span multiple lines, grep might fail—use sed to capture the full field block first:
# Extract entire <NAME> blocks, then strip tags sed -n '/<NAME>/,/<\/NAME>/p' XML_file.xml | sed 's/<NAME>//;s/<\/NAME>//'
Using awk (great for structured text parsing)
awk -F'<NAME>|</NAME>' '{if ($2 != "") print $2}' XML_file.xml
-Fsets field separators to the opening/closing tags- Prints the second field (the content between the tags)
Note: These methods work best for well-formed, flat XML. If your XML has nested tags, you’ll need more advanced logic (covered in part 2).
search_file.txt & Extract <ADR> to output.txt For this task, we’ll use awk (the most reliable tool for multi-block matching and conditional extraction). Let’s assume your XML structure looks like this for each <H> record:
<H> <NAME>John</NAME> <CODE>MANG</CODE> <ID>102</ID> <ADR>123 Main St, Cityville</ADR> </H>
Full awk Script Solution
Create a shell script or run this command directly:
awk -F',' ' BEGIN { # Load all match criteria from search_file.txt into an array while ((getline line < "search_file.txt") > 0) { split(line, match_vals, ",") # Create a unique key for each name/code/ID combo match_key = match_vals[1] "|" match_vals[2] "|" match_vals[3] target_records[match_key] = 1 } close("search_file.txt") in_h_block = 0 current_match_key = "" current_adr = "" } # Trigger when entering an <H> record block /<H>/ { in_h_block = 1 current_match_key = "" current_adr = "" next } # Trigger when exiting an <H> block—check if we need to save the ADR /<\/H>/ { if (in_h_block && current_match_key in target_records) { print current_adr >> "output.txt" } in_h_block = 0 next } # Process content inside an <H> block in_h_block { # Extract NAME and build the match key if (/<NAME>/) { sub(/.*<NAME>/, "", $0) sub(/<\/NAME>.*/, "", $0) current_match_key = $0 "|" } # Extract CODE and append to the match key else if (/<CODE>/) { sub(/.*<CODE>/, "", $0) sub(/<\/CODE>.*/, "", $0) current_match_key = current_match_key $0 "|" } # Extract ID and finalize the match key else if (/<ID>/) { sub(/.*<ID>/, "", $0) sub(/<\/ID>.*/, "", $0) current_match_key = current_match_key $0 } # Extract ADR content (handles multi-line ADRs too) else if (/<ADR>/) { current_adr = "" # Read until we hit the closing </ADR> tag while (getline > 0) { if (/<\/ADR>/) { sub(/<\/ADR>.*/, "", $0) current_adr = current_adr $0 break } current_adr = current_adr $0 } } } ' XML_file.xml
How This Works:
- Load Match Criteria: First, it reads every line from
search_file.txt(e.g.,John,MANG,102) and stores each combo as a unique key in an array. - Traverse XML Blocks: It flags when it enters/exits an
<H>record, resetting variables for each new block. - Build Match Key: Inside each
<H>block, it extracts the<NAME>,<CODE>, and<ID>values to build a key that matches the ones fromsearch_file.txt. - Extract & Save ADR: If the block’s key matches a target, it extracts the full
<ADR>content (even if it spans multiple lines) and appends it tooutput.txt.
Alternative: Grep + Sed Combo (Simpler but Less Reliable)
If your XML fields are always in the same order and never span lines, you can use a simpler pipeline:
# Extract all <H> blocks to a temp file first sed -n '/<H>/,/<\/H>/p' XML_file.xml > temp_h_blocks.txt # Loop through each line in search_file.txt and match while IFS=, read -r name code id; do # Find the matching block, then extract the ADR grep -A 100 -B 100 "<NAME>$name</NAME>" temp_h_blocks.txt | \ grep -A 100 -B 100 "<CODE>$code</CODE>" | \ grep -A 100 -B 100 "<ID>$id</ID>" | \ grep -oP '(?<=<ADR>).*?(?=</ADR>)' >> output.txt done < search_file.txt # Clean up temp file rm temp_h_blocks.txt
Note: This fails if fields are out of order or span lines—stick with the awk method for robustness.
内容的提问来源于stack exchange,提问作者Munu

