如何用XPath选取位于w:ins/w:del之间的指定w:r元素?
Fixing Your XPath for Targeting Specific w:r Elements in Nokogiri
Let's break down what's off with your original XPath and build the correct one step by step, tailored for your Nokogiri workflow.
What's Wrong with Your Current Expression
Your existing XPath //w:r[. = ' ' and preceding-sibling::w:ins and following-sibling::w:del] has two critical gaps:
- It uses
preceding-sibling::w:insandfollowing-sibling::w:del, which match any precedingw:insor followingw:del—not the immediately adjacent sibling elements you need. - It doesn't account for the "either/or" rule (previous/next can be
w:insorw:del), and checking. = ' 'might not reliably catch whitespace-onlyw:tnodes inside thew:r.
The Correct XPath Expression
Here's the adjusted query that meets all your requirements:
//w:r[ w:t[normalize-space(.) = '' and string-length(.) > 0] and preceding-sibling::*[1][self::w:ins or self::w:del] and following-sibling::*[1][self::w:ins or self::w:del] ]
Let's Break Down Each Part
w:t[normalize-space(.) = '' and string-length(.) > 0]: Ensures thew:rcontains aw:telement with nothing but whitespace (handles single spaces, multiple spaces, tabs, etc.). Thestring-lengthcheck ensures we don't match emptyw:tnodes with no content at all.preceding-sibling::*[1][self::w:ins or self::w:del]: Targets the immediately preceding sibling element (using*[1]to grab the closest one) and verifies it's eitherw:insorw:del.following-sibling::*[1][self::w:ins or self::w:del]: Does the same for the immediately following sibling element.
Using This in Nokogiri
Don't forget to register the WordprocessingML namespace when querying—Nokogiri doesn't auto-detect it by default:
require 'nokogiri' # Load your XML content (file or string) xml_content = File.read('your_word_doc.xml') doc = Nokogiri::XML(xml_content) # Register the 'w' namespace for WordprocessingML namespace = { 'w' => 'http://schemas.openxmlformats.org/wordprocessingml/2006/main' } # Run the XPath query target_elements = doc.xpath("//w:r[w:t[normalize-space(.) = '' and string-length(.) > 0] and preceding-sibling::*[1][self::w:ins or self::w:del] and following-sibling::*[1][self::w:ins or self::w:del]]", namespace)
Testing Against Your Examples
- Example 1: This will select the first and third whitespace-only
w:relements (the ones you want), and skip the second one (since its previous sibling is a regularw:r, notw:ins/w:del). - Example 2: It won't select the whitespace
w:rbecause its next sibling is a regularw:r, notw:insorw:del.
内容的提问来源于stack exchange,提问作者chell
相关产品推荐
相关产品推荐

