能否用Selenium和Nokogiri根据邻近标签定位元素并提取内容?
Your pseudocode is spot-on—this kind of text-based element targeting is totally feasible with both tools. Let’s break it down into working code, covering both static HTML parsing (Nokogiri alone) and dynamic page scenarios (combining Selenium + Nokogiri).
Nokogiri-Only Solution (For Static HTML)
If you already have the parsed HTML document (like your sample snippet), here’s how to turn your pseudocode into functional Ruby:
require 'nokogiri' # Replace this with your actual parsed document source html = <<~HTML <div> <h2>Details</h2> <ul> <li><b>Height:</b>6 ft</li> <li><b>Weight:</b>6 kg</li> <li><b>Age:</b>6</li> </ul> </div> HTML doc = Nokogiri::HTML(html) # Locate the div containing the "Details" heading div_of_interest = doc.at('div:has(h2:text("Details"))') # Find the list item that includes "Weight:" weight_li = div_of_interest.at('li:contains("Weight:")') # Extract just the value after "Weight:" selected_text = weight_li.text.sub(/Weight:\s*/, '').strip puts selected_text # Outputs: "6 kg"
How this works:
div:has(h2:text("Details")): Uses Nokogiri’s extended CSS selectors to zero in on the div with an<h2>child containing exactly "Details".li:contains("Weight:"): Targets the specific list item holding the weight data.- The regex
sub(/Weight:\s*/, '')cleans up the text by removing "Weight:" and any trailing whitespace, leaving only the value.
Combining with Selenium (For Dynamic Pages)
If your content loads dynamically (e.g., via JavaScript), use Selenium to fetch the page first, then pass the source to Nokogiri:
require 'selenium-webdriver' require 'nokogiri' # Initialize Selenium (use :firefox, :edge, etc., as needed) driver = Selenium::WebDriver.for :chrome begin # Navigate to your target page driver.get('https://your-target-url.com') # Parse the loaded page source with Nokogiri doc = Nokogiri::HTML(driver.page_source) # Reuse the same logic as the static solution div_of_interest = doc.at('div:has(h2:text("Details"))') weight_li = div_of_interest.at('li:contains("Weight:")') selected_text = weight_li.text.sub(/Weight:\s*/, '').strip puts selected_text # Outputs: "6 kg" ensure driver.quit # Always close the driver to avoid resource leaks end
Alternative: Use Selenium Directly (No Nokogiri Needed)
You can also locate the element and extract the value using Selenium’s built-in methods, with XPath:
require 'selenium-webdriver' driver = Selenium::WebDriver.for :chrome begin driver.get('https://your-target-url.com') # Use XPath to find the exact li element weight_element = driver.find_element( xpath: '//div[h2[text()="Details"]]/ul/li[contains(text(), "Weight:")]' ) # Clean up the text to get just the weight value selected_text = weight_element.text.gsub(/Weight:\s*/, '').strip puts selected_text # Outputs: "6 kg" ensure driver.quit end
Pro Tips:
- For even cleaner parsing (no regex), target the
<b>tag directly and grab its next sibling text:weight_li.at('b').next_sibling.text.strip—this works if the HTML structure is consistent. - If text might have case variations or extra spaces, adjust the regex to be more flexible (e.g.,
sub(/Weight:\s*/i, '')for case-insensitive matching).
Content of the question originates from Stack Exchange, question author Joe Morano

