如何用Python修改并枚举XML标签?为<sentence>标签添加自增ID属性
Got it, this is a straightforward task once you break it down. Here's exactly how you can implement auto-incrementing IDs for all your <sentence> tags using BeautifulSoup:
Step-by-Step Solution
First, make sure you have the right tools installed. You'll need beautifulsoup4 and lxml (since XML parsing requires a proper XML parser, not the default HTML one). If you haven't installed them yet, run:
pip install beautifulsoup4 lxml
Then, use this code snippet to process your XML file:
from bs4 import BeautifulSoup # 1. Read your XML corpus file with open("your_corpus.xml", "r", encoding="utf-8") as input_file: xml_data = input_file.read() # 2. Parse the XML with the lxml-xml parser (critical for proper XML handling) soup = BeautifulSoup(xml_data, "lxml-xml") # 3. Initialize a counter for the IDs current_id = 1 # 4. Loop through every <sentence> tag and assign the auto-incrementing ID for sentence_tag in soup.find_all("sentence"): # Assign the ID (convert to string since XML attributes are text) sentence_tag["id"] = str(current_id) # Increment the counter for the next sentence current_id += 1 # 5. Save the modified XML back to a file (use a new filename to avoid overwriting original) with open("your_corpus_with_ids.xml", "w", encoding="utf-8") as output_file: output_file.write(soup.prettify())
Key Notes to Keep in Mind
- Use the correct parser: Always specify
"lxml-xml"when parsing XML with BeautifulSoup. Using the default HTML parser can mess up XML-specific syntax (like self-closing tags or namespace attributes). - Avoid overwriting your original file: I recommend saving to a new filename first (like
your_corpus_with_ids.xml) so you can verify the changes before replacing the original. - Handle existing IDs (optional): If some
<sentence>tags already have IDs and you don't want to overwrite them, add a check in the loop:for sentence_tag in soup.find_all("sentence"): if "id" not in sentence_tag.attrs: sentence_tag["id"] = str(current_id) current_id += 1 - Prettify for readability: The
prettify()method formats the output XML with indentation, making it easier to read. If you don't care about formatting, you can just writestr(soup)instead.
That's it! This will give you sequentially numbered IDs starting from 1 for all your <sentence> tags.
内容的提问来源于stack exchange,提问作者Paul Kremershof
相关产品推荐
相关产品推荐

