You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XPath选取位于w:ins/w:del之间的指定w:r元素?

Fixing Your XPath for Targeting Specific w:r Elements in Nokogiri

Let's break down what's off with your original XPath and build the correct one step by step, tailored for your Nokogiri workflow.

What's Wrong with Your Current Expression

Your existing XPath //w:r[. = ' ' and preceding-sibling::w:ins and following-sibling::w:del] has two critical gaps:

  • It uses preceding-sibling::w:ins and following-sibling::w:del, which match any preceding w:ins or following w:del—not the immediately adjacent sibling elements you need.
  • It doesn't account for the "either/or" rule (previous/next can be w:ins or w:del), and checking . = ' ' might not reliably catch whitespace-only w:t nodes inside the w:r.

The Correct XPath Expression

Here's the adjusted query that meets all your requirements:

//w:r[
  w:t[normalize-space(.) = '' and string-length(.) > 0]
  and preceding-sibling::*[1][self::w:ins or self::w:del]
  and following-sibling::*[1][self::w:ins or self::w:del]
]

Let's Break Down Each Part

  • w:t[normalize-space(.) = '' and string-length(.) > 0]: Ensures the w:r contains a w:t element with nothing but whitespace (handles single spaces, multiple spaces, tabs, etc.). The string-length check ensures we don't match empty w:t nodes with no content at all.
  • preceding-sibling::*[1][self::w:ins or self::w:del]: Targets the immediately preceding sibling element (using *[1] to grab the closest one) and verifies it's either w:ins or w:del.
  • following-sibling::*[1][self::w:ins or self::w:del]: Does the same for the immediately following sibling element.

Using This in Nokogiri

Don't forget to register the WordprocessingML namespace when querying—Nokogiri doesn't auto-detect it by default:

require 'nokogiri'

# Load your XML content (file or string)
xml_content = File.read('your_word_doc.xml')
doc = Nokogiri::XML(xml_content)

# Register the 'w' namespace for WordprocessingML
namespace = { 'w' => 'http://schemas.openxmlformats.org/wordprocessingml/2006/main' }

# Run the XPath query
target_elements = doc.xpath("//w:r[w:t[normalize-space(.) = '' and string-length(.) > 0] and preceding-sibling::*[1][self::w:ins or self::w:del] and following-sibling::*[1][self::w:ins or self::w:del]]", namespace)

Testing Against Your Examples

  • Example 1: This will select the first and third whitespace-only w:r elements (the ones you want), and skip the second one (since its previous sibling is a regular w:r, not w:ins/w:del).
  • Example 2: It won't select the whitespace w:r because its next sibling is a regular w:r, not w:ins or w:del.

内容的提问来源于stack exchange,提问作者chell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:15:40