You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在存在XML违规的类XML文本文件中提取第n个<后的单词?

Got it, let's work through this problem together. Since your pseudo-XML file has violations that break standard XML parsers, we can't rely on those tools—so we need to use raw string processing to pull out the word after the nth < tag start, even with inconsistent line formatting. Here are two solid approaches:

方案1:逐字符遍历(稳健,适配不规则格式)

This method avoids regex pitfalls and works directly with the raw text, making it great for messy, non-compliant files where tags might have weird spacing or invalid syntax.

步骤:

  • Read the entire file as a single string (since line breaks don't follow rules, line-by-line processing won't work reliably)
  • Track how many < characters we've found
  • When we hit the nth <, skip any leading whitespace after it, then collect characters until we hit a non-tag-name character (like a space, >, or invalid symbol)
  • Exit early once we've found our target to save resources

代码示例(Python):

def get_nth_tag_name(raw_text, n):
    tag_count = 0
    for idx, char in enumerate(raw_text):
        if char == '<':
            tag_count += 1
            if tag_count == n:
                # Skip any whitespace right after the < (e.g., <  user>)
                current_pos = idx + 1
                while current_pos < len(raw_text) and raw_text[current_pos].isspace():
                    current_pos += 1
                # Collect valid XML tag name characters (letters, numbers, _, -, :)
                tag_name = []
                while current_pos < len(raw_text):
                    c = raw_text[current_pos]
                    if c.isalnum() or c in ['_', '-', ':']:
                        tag_name.append(c)
                        current_pos += 1
                    else:
                        break
                return ''.join(tag_name)
    # Return None if there aren't enough < characters
    return None

# Usage
with open('broken_xml_file.txt', 'r', encoding='utf-8') as f:
    file_content = f.read()

target_tag = get_nth_tag_name(file_content, 3)  # Get the 3rd tag's name
print(target_tag)

方案2:正则表达式(简洁,适合大部分场景)

If your file doesn't have super edge-case syntax (like nested comments or malformed CDATA blocks), regex is a quicker way to extract all tag names at once, then pick the nth one.

代码示例(Python):

import re

def get_nth_tag_name_regex(raw_text, n):
    # Regex to match tag names, skipping comments, CDATA, and processing instructions
    # Adjust the pattern if you need to include closing tags (remove the (?!/) part)
    tag_pattern = r'<\s*([a-zA-Z0-9_\-:]+)(?!/)'
    all_tag_names = re.findall(tag_pattern, raw_text)
    
    if 0 < n <= len(all_tag_names):
        return all_tag_names[n-1]  # Lists are 0-indexed, so subtract 1
    else:
        return None

# Usage
with open('broken_xml_file.txt', 'r', encoding='utf-8') as f:
    file_content = f.read()

target_tag = get_nth_tag_name_regex(file_content, 3)
print(target_tag)

关键注意事项

  • Including closing tags: If you want to count closing tags (like </user>) and extract their names, remove the (?!/) part from the regex pattern, or adjust the string traversal logic to not skip tags starting with </.
  • Large files: If your file is too big to load into memory all at once, modify the string traversal method to read the file in chunks, tracking the tag count and partial tag name across chunks.
  • Invalid tag characters: If your pseudo-XML uses non-standard characters in tag names, update the character collection logic (in the traversal method) or regex pattern to include those characters.

内容的提问来源于stack exchange,提问作者Bhanuchander Udhayakumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:34:56