You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Nokogiri XPath中使用正则表达式?如何用XPath获取标签后的文本?

Hey there! Let’s work through your two Nokogiri XPath questions in detail—since you noted you tried some approaches without success, I’ll focus on practical, actionable solutions that cover common pitfalls.

1. Using Regular Expressions in Nokogiri XPath

Nokogiri relies on libxml2, which defaults to XPath 1.0 (which doesn’t have native regex support). But don’t worry—there are two reliable ways to add regex matching to your queries:

Option 1: Combine XPath with Ruby’s Native Regex

This is the most straightforward approach, especially if you’ve struggled with XPath-only regex attempts. First fetch candidate nodes with XPath, then filter them using Ruby’s select method and regex:

require 'nokogiri'

# Sample HTML
doc = Nokogiri::HTML(<<~HTML)
  <ul>
    <li>Product A - 2024</li>
    <li>Product B - 2023</li>
    <li>Service C</li>
  </ul>
HTML

# Get all <li> tags whose text contains a 4-digit year
matching_nodes = doc.xpath('//li').select { |node| node.text.match?(/\d{4}/) }
matching_nodes.each { |node| puts node.text.strip }
# Output: Product A - 2024, Product B - 2023

Option 2: Use XPath 2.0’s matches() Function (Advanced)

If you need regex directly in your XPath expression, you can enable XPath 2.0 support (via the nokogiri-xpath2 gem) or register a custom Ruby function for XPath. Here’s the custom function approach:

# Register a namespace for Ruby functions
Nokogiri::XML::XPath.register_namespace('ruby', 'http://www.ruby-lang.org/xmlns/ruby/1.0')

# Use the custom regex match function in XPath
matching_nodes = doc.xpath('//li[ruby:match(text(), "\d{4}")]')
matching_nodes.each { |node| puts node.text.strip }

Stick with Option 1 unless you specifically need regex in the XPath string—it’s less prone to environment-specific issues.

2. Fetching Arbitrary Text After a Specific Tag

This is a common pain point, usually because text after a tag might be a sibling text node, nested in other elements, or spread across multiple nodes. Let’s cover the most common scenarios you might have tried:

Scenario 1: Get All Text After a Tag (Including Nested Elements)

If you want every piece of text that comes after a specific tag (even inside other elements):

<!-- Sample HTML -->
<div class="content">
  <h2>Introduction</h2>
  This is text right after the heading.
  <p>Here’s a paragraph with more text.</p>
  And even more text outside any tag!
</div>
doc = Nokogiri::HTML(above_html)
target_heading = doc.at_xpath('//h2[text()="Introduction"]')

# Fetch all following text nodes and join them
text_after = target_heading.xpath('following::text()').map(&:text).join.strip
puts text_after
# Output: This is text right after the heading. Here’s a paragraph with more text. And even more text outside any tag!

Scenario 2: Get Only the Immediate Sibling Text Node

If you just want the plain text directly after the tag (not inside subsequent elements):

<!-- Sample HTML -->
<div>
  <label>Name:</label>
  John Doe
  <input type="text">
</div>
target_label = doc.at_xpath('//label[text()="Name:"]')

# Target the first sibling text node after the label
immediate_text = target_label.xpath('following-sibling::text()[1]').text.strip
puts immediate_text
# Output: John Doe

Scenario 3: Get Text Between Two Tags

If you need to capture text from your target tag up to another specific tag:

<!-- Sample HTML -->
<div>
  <h3>Start Here</h3>
  Capture this text
  <span>and this span content</span>
  stop here
  <h3>End Here</h3>
</div>
start_node = doc.at_xpath('//h3[text()="Start Here"]')
end_node = doc.at_xpath('//h3[text()="End Here"]')

text_between = []
current_node = start_node.next_sibling

# Traverse nodes until we hit the end tag
while current_node && current_node != end_node
  if current_node.text?
    text_between << current_node.text.strip
  elsif current_node.element?
    text_between << current_node.text.strip
  end
  current_node = current_node.next_sibling
end

puts text_between.join(' ')
# Output: Capture this text and this span content stop here

A common mistake here is forgetting that text outside tags is its own text node—using following-sibling::text() instead of just following::text() can limit your results if you need nested content.


内容的提问来源于stack exchange,提问作者anonymous1997

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:38:18