使用Bash提取XML标签对内容并分别存储至多个文件
Bash Script to Extract Blocks from XML into Separate Files
Got it, here's a practical Bash script using awk that splits each <note> section from your XML file into its own numbered file (1.xml, 2.xml, etc.). It’s built to handle thousands of notes without running into resource issues, since it properly closes each file after processing.
The Script (extract_notes.sh)
#!/bin/bash # Check if input file is provided if [ $# -ne 1 ]; then echo "Usage: $0 <input-xml-file>" exit 1 fi INPUT_FILE="$1" awk ' /<note>/ { in_note = 1 count++ filename = count ".xml" # Print the opening note tag to the file print > filename next } in_note { # Print current line to the active note file print > filename # Close the file when we hit the closing note tag if (/<\/note>/) { in_note = 0 close(filename) } } ' "$INPUT_FILE" echo "Extraction complete! Check the current directory for files 1.xml, 2.xml, etc."
How It Works
Let’s break down the key parts:
- Input Check: First, the script verifies you’ve passed an XML file as input. If not, it shows a quick usage guide.
- Awk Logic:
- When it finds a line with
<note>, it activates a capture flag (in_note), increments a counter, and creates the output file (e.g.,1.xml). - While the capture flag is active, every line gets written to the current numbered file.
- When it hits
</note>, it turns off the capture flag and closes the file—critical for avoiding open file limits with thousands of notes.
- When it finds a line with
How to Use
- Save the script as
extract_notes.sh. - Make it executable:
chmod +x extract_notes.sh. - Run it with your XML file:
./extract_notes.sh your_input.xml.
Each full <note> block (including opening and closing tags) will be saved to its own file, starting from 1.xml upwards.
内容的提问来源于stack exchange,提问作者Vinay Kumar Bommana
相关产品推荐
相关产品推荐

