XSLT需求:匹配节点并截取特定字符后的节点内容(保留标签)
Alright, let's tackle this problem head-on. When you need to extract content from an XML node starting at a specific character position while keeping all nested elements and tags intact, raw string slicing won't work—it'll break the XML structure. Instead, we need to work with the XML's tree structure using a proper parser.
Here's a practical solution using Python's lxml library, which excels at handling XML/HTML parsing and manipulation:
Step 1: Install lxml
First, make sure you have the library installed:
pip install lxml
Step 2: Full Implementation Code
This function will locate your target node, calculate text positions across nested elements, and slice the content while preserving all tags:
from lxml import etree def slice_node_from_position(target_node, start_char_pos): """ Slice content from an XML node starting at a specific character position, preserving all nested elements and tags. Args: target_node: The lxml Element node to slice start_char_pos: The 0-indexed character position to start slicing from Returns: String representation of the sliced node """ total_text_length = 0 # Create a new element to hold our sliced content (matches original node's tag/attributes) sliced_element = etree.Element(target_node.tag, attrib=target_node.attrib) # Handle the node's own leading text (before any child elements) if target_node.text: text_len = len(target_node.text) if total_text_length + text_len <= start_char_pos: total_text_length += text_len else: remaining_pos = start_char_pos - total_text_length sliced_element.text = target_node.text[remaining_pos:] total_text_length += text_len # Iterate through all child nodes (elements and text) for child in target_node: if isinstance(child, etree._ElementUnicodeResult): # This is a text node text_len = len(child) if total_text_length + text_len <= start_char_pos: total_text_length += text_len continue remaining_pos = start_char_pos - total_text_length # Append the sliced text to the last element or the main node if sliced_element.getchildren(): sliced_element.getchildren()[-1].tail = child[remaining_pos:] else: sliced_element.text = child[remaining_pos:] total_text_length += text_len else: # This is an element node child_text_len = len(child.text) if child.text else 0 if total_text_length + child_text_len <= start_char_pos: # Skip this element's text, add its tail to total length total_text_length += child_text_len if child.tail: total_text_length += len(child.tail) continue elif total_text_length >= start_char_pos: # We're past the start position, add the entire element sliced_element.append(child) else: # Need to slice the element's text, then keep its children remaining_pos = start_char_pos - total_text_length sliced_child = etree.Element(child.tag, attrib=child.attrib) sliced_child.text = child.text[remaining_pos:] if child.text else None # Add all nested children of the original element for grandchild in child: sliced_child.append(grandchild) # Keep the element's tail text sliced_child.tail = child.tail sliced_element.append(sliced_child) total_text_length += child_text_len # Convert the sliced element back to a formatted XML string return etree.tostring(sliced_element, encoding='unicode', pretty_print=True) # Example Usage # Parse your XML content xml_data = '''<?xml version="1.0" encoding="utf-8" ?> <body> <div class="todo"> <p class="commission">1. shop at the <b>supermarket</b> when the weather <i>is nice</i></p> <p class="letter"><i>a)</i> buy chocolate <b>biscuits</b></p> <p class="number"><b>2.</b> after finishing the shopping, call mom</p> </div> </body>''' tree = etree.fromstring(xml_data) # Locate your target node (using XPath to find the p tag with class="commission") target_node = tree.xpath('//p[@class="commission"]')[0] # Slice starting at character position 4 (skips "1. " which is 3 characters) sliced_result = slice_node_from_position(target_node, start_char_pos=4) print(sliced_result)
Step 3: What This Does
- Preserves Structure: Unlike string slicing, this method respects XML's nested elements—so
<b>supermarket</b>and<i>is nice</i>stay intact. - Handles Text Across Nodes: It calculates the total text length across all text nodes and element content to find exactly where your start position falls.
- Matches Original Node: The sliced output keeps the original node's tag and attributes (like
class="commission").
Example Output
Running the code above will produce:
<p class="commission">shop at the <b>supermarket</b> when the weather <i>is nice</i></p>
You can adjust the start_char_pos value to slice from any position—even if it falls in the middle of a nested element's text. For example, setting start_char_pos=15 would slice starting at "the" in "the supermarket".
内容的提问来源于stack exchange,提问作者H725

