如何在Nokogiri XPath中使用正则表达式?如何用XPath获取标签后的文本?
Hey there! Let’s work through your two Nokogiri XPath questions in detail—since you noted you tried some approaches without success, I’ll focus on practical, actionable solutions that cover common pitfalls.
Nokogiri relies on libxml2, which defaults to XPath 1.0 (which doesn’t have native regex support). But don’t worry—there are two reliable ways to add regex matching to your queries:
Option 1: Combine XPath with Ruby’s Native Regex
This is the most straightforward approach, especially if you’ve struggled with XPath-only regex attempts. First fetch candidate nodes with XPath, then filter them using Ruby’s select method and regex:
require 'nokogiri' # Sample HTML doc = Nokogiri::HTML(<<~HTML) <ul> <li>Product A - 2024</li> <li>Product B - 2023</li> <li>Service C</li> </ul> HTML # Get all <li> tags whose text contains a 4-digit year matching_nodes = doc.xpath('//li').select { |node| node.text.match?(/\d{4}/) } matching_nodes.each { |node| puts node.text.strip } # Output: Product A - 2024, Product B - 2023
Option 2: Use XPath 2.0’s matches() Function (Advanced)
If you need regex directly in your XPath expression, you can enable XPath 2.0 support (via the nokogiri-xpath2 gem) or register a custom Ruby function for XPath. Here’s the custom function approach:
# Register a namespace for Ruby functions Nokogiri::XML::XPath.register_namespace('ruby', 'http://www.ruby-lang.org/xmlns/ruby/1.0') # Use the custom regex match function in XPath matching_nodes = doc.xpath('//li[ruby:match(text(), "\d{4}")]') matching_nodes.each { |node| puts node.text.strip }
Stick with Option 1 unless you specifically need regex in the XPath string—it’s less prone to environment-specific issues.
This is a common pain point, usually because text after a tag might be a sibling text node, nested in other elements, or spread across multiple nodes. Let’s cover the most common scenarios you might have tried:
Scenario 1: Get All Text After a Tag (Including Nested Elements)
If you want every piece of text that comes after a specific tag (even inside other elements):
<!-- Sample HTML --> <div class="content"> <h2>Introduction</h2> This is text right after the heading. <p>Here’s a paragraph with more text.</p> And even more text outside any tag! </div>
doc = Nokogiri::HTML(above_html) target_heading = doc.at_xpath('//h2[text()="Introduction"]') # Fetch all following text nodes and join them text_after = target_heading.xpath('following::text()').map(&:text).join.strip puts text_after # Output: This is text right after the heading. Here’s a paragraph with more text. And even more text outside any tag!
Scenario 2: Get Only the Immediate Sibling Text Node
If you just want the plain text directly after the tag (not inside subsequent elements):
<!-- Sample HTML --> <div> <label>Name:</label> John Doe <input type="text"> </div>
target_label = doc.at_xpath('//label[text()="Name:"]') # Target the first sibling text node after the label immediate_text = target_label.xpath('following-sibling::text()[1]').text.strip puts immediate_text # Output: John Doe
Scenario 3: Get Text Between Two Tags
If you need to capture text from your target tag up to another specific tag:
<!-- Sample HTML --> <div> <h3>Start Here</h3> Capture this text <span>and this span content</span> stop here <h3>End Here</h3> </div>
start_node = doc.at_xpath('//h3[text()="Start Here"]') end_node = doc.at_xpath('//h3[text()="End Here"]') text_between = [] current_node = start_node.next_sibling # Traverse nodes until we hit the end tag while current_node && current_node != end_node if current_node.text? text_between << current_node.text.strip elsif current_node.element? text_between << current_node.text.strip end current_node = current_node.next_sibling end puts text_between.join(' ') # Output: Capture this text and this span content stop here
A common mistake here is forgetting that text outside tags is its own text node—using following-sibling::text() instead of just following::text() can limit your results if you need nested content.
内容的提问来源于stack exchange,提问作者anonymous1997

