You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XML标注数据集转换为BRAT格式的技术实现咨询

Got it, let's walk through how to convert your XML-annotated text to BRAT's standoff format properly. I've dealt with this exact task for medical NLP datasets before, so here's a practical, tested solution:

Step 1: Remember BRAT's Core Requirements

First, BRAT needs two paired files for every document:

  • A plain text file (.txt) with all XML tags stripped out (this is the base text BRAT displays)
  • An annotation file (.ann) with line-by-line standoff entries, following the format: T<ID> <TYPE> <START_OFFSET> <END_OFFSET> <ANNOTATED_TEXT>

The tricky part is calculating accurate character offsets—since removing XML tags shifts the position of your annotated text, you can't just use positions from the original XML.

Step 2: Python Script to Handle Conversion

This script uses regex to locate annotations and tracks character positions as it builds the tag-free text, ensuring offsets are 100% accurate:

import re

def xml_to_brat(xml_input, txt_out_path, ann_out_path):
    # Regex to match annotation tags and capture CUI + annotated text
    annotation_pattern = re.compile(r'<annotation cui="(.*?)">(.*?)</annotation>')
    
    plain_text_parts = []
    brat_annotations = []
    current_text_length = 0
    annotation_id = 1

    # Iterate through every matched annotation in the XML
    for match in annotation_pattern.finditer(xml_input):
        cui, annotated_phrase = match.groups()
        
        # Add the text before the current annotation to our plain text
        plain_text_parts.append(xml_input[current_text_length:match.start()])
        # Calculate where the annotated phrase starts in the final plain text
        start_offset = len(''.join(plain_text_parts))
        # Add the annotated phrase to the plain text
        plain_text_parts.append(annotated_phrase)
        # Calculate the end offset
        end_offset = start_offset + len(annotated_phrase)
        
        # Build the BRAT annotation line (use CUI as type if preferred)
        brat_annotations.append(f"T{annotation_id} annotation {start_offset} {end_offset} {annotated_phrase}")
        annotation_id += 1
        
        # Update our position in the original XML text
        current_text_length = match.end()

    # Add any remaining text after the last annotation
    plain_text_parts.append(xml_input[current_text_length:])
    final_plain_text = ''.join(plain_text_parts)

    # Write output files
    with open(txt_out_path, 'w', encoding='utf-8') as txt_file:
        txt_file.write(final_plain_text)
    
    with open(ann_out_path, 'w', encoding='utf-8') as ann_file:
        ann_file.write('\n'.join(brat_annotations))

# Test with your example input
sample_xml = 'Treatment of <annotation cui="C0267055">Erosive Esophagitis</annotation> in patients'
xml_to_brat(sample_xml, 'patient_note.txt', 'patient_note.ann')

Step 3: Customization & Verification

  • Use CUI as annotation type: If you want to label annotations with their CUI (e.g., C0267055 instead of annotation), modify the annotation line to:
    brat_annotations.append(f"T{annotation_id} {cui} {start_offset} {end_offset} {annotated_phrase}")
    
  • Check output: For your sample input, the script will generate:
    • patient_note.txt: Treatment of Erosive Esophagitis in patients
    • patient_note.ann: T1 annotation 14 33 Erosive Esophagitis
      Which perfectly matches the BRAT format you provided.

Edge Cases to Keep In Mind

  • Duplicate phrases: The script processes annotations in order, so even if the same text appears multiple times, each gets a unique ID and correct offsets.
  • Multi-line text: Line breaks are counted as characters, so offsets will work even for wrapped or multi-paragraph XML input.

内容的提问来源于stack exchange,提问作者max

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:35:02