You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python修改并枚举XML标签?为<sentence>标签添加自增ID属性

How to Add Auto-Incrementing IDs to Tags in XML with BeautifulSoup

Got it, this is a straightforward task once you break it down. Here's exactly how you can implement auto-incrementing IDs for all your <sentence> tags using BeautifulSoup:

Step-by-Step Solution

First, make sure you have the right tools installed. You'll need beautifulsoup4 and lxml (since XML parsing requires a proper XML parser, not the default HTML one). If you haven't installed them yet, run:

pip install beautifulsoup4 lxml

Then, use this code snippet to process your XML file:

from bs4 import BeautifulSoup

# 1. Read your XML corpus file
with open("your_corpus.xml", "r", encoding="utf-8") as input_file:
    xml_data = input_file.read()

# 2. Parse the XML with the lxml-xml parser (critical for proper XML handling)
soup = BeautifulSoup(xml_data, "lxml-xml")

# 3. Initialize a counter for the IDs
current_id = 1

# 4. Loop through every <sentence> tag and assign the auto-incrementing ID
for sentence_tag in soup.find_all("sentence"):
    # Assign the ID (convert to string since XML attributes are text)
    sentence_tag["id"] = str(current_id)
    # Increment the counter for the next sentence
    current_id += 1

# 5. Save the modified XML back to a file (use a new filename to avoid overwriting original)
with open("your_corpus_with_ids.xml", "w", encoding="utf-8") as output_file:
    output_file.write(soup.prettify())

Key Notes to Keep in Mind

  • Use the correct parser: Always specify "lxml-xml" when parsing XML with BeautifulSoup. Using the default HTML parser can mess up XML-specific syntax (like self-closing tags or namespace attributes).
  • Avoid overwriting your original file: I recommend saving to a new filename first (like your_corpus_with_ids.xml) so you can verify the changes before replacing the original.
  • Handle existing IDs (optional): If some <sentence> tags already have IDs and you don't want to overwrite them, add a check in the loop:
    for sentence_tag in soup.find_all("sentence"):
        if "id" not in sentence_tag.attrs:
            sentence_tag["id"] = str(current_id)
            current_id += 1
    
  • Prettify for readability: The prettify() method formats the output XML with indentation, making it easier to read. If you don't care about formatting, you can just write str(soup) instead.

That's it! This will give you sequentially numbered IDs starting from 1 for all your <sentence> tags.

内容的提问来源于stack exchange,提问作者Paul Kremershof

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:01:16