如何用Jsoup/Selenium获取::before标签后数字及标签内信息?已尝试方案无效
Hey there! Let's work through this problem together. The key issue here is that pseudo-elements like ::before aren't part of the actual DOM structure—they're rendered by CSS, which is why tools like Jsoup (that only parse static HTML source) can't access their content directly. Let's break down solutions for both your requirements:
::before伪元素后的数字数据 Since ::before content is generated by the browser's CSS engine, you'll need a tool that can interact with the rendered page (like Selenium) to grab it. Here's a concrete Selenium example using Java (adjust for your language if needed):
First, locate the element that the ::before pseudo-element is attached to, then use JavaScript to fetch its computed style:
// 1. 初始化Selenium驱动(比如ChromeDriver) WebDriver driver = new ChromeDriver(); driver.get("your-target-url"); // 2. 定位到承载::before的父元素(替换成你实际的选择器) WebElement targetElement = driver.findElement(By.cssSelector(".your-element-class")); // 3. 通过JavascriptExecutor获取::before的content属性 JavascriptExecutor js = (JavascriptExecutor) driver; String beforeContent = (String) js.executeScript( "return window.getComputedStyle(arguments[0], ':before').getPropertyValue('content');", targetElement ); // 4. 清理内容(通常content会带引号,需要去掉)并转成数字 String cleanNumber = beforeContent.replaceAll("\"", "").trim(); int targetNumber = Integer.parseInt(cleanNumber); System.out.println("从::before获取到的数字:" + targetNumber); // 记得关闭驱动 driver.quit();
为什么Jsoup不行?
Jsoup parses the raw HTML source sent by the server, not the rendered page. Since ::before content isn't present in the raw HTML, Jsoup can't access it at all—this is why your Xpath approach with Jsoup didn't work.
This should be straightforward with either Jsoup or Selenium, assuming your selector/Xpath was correct. Let's cover both tools:
用Jsoup的示例
If the page is static (no JS rendering needed), Jsoup works perfectly. For example, if your HTML looks like this:
<div class="content-box"> <span class="data-value">12345</span> <p>Additional info here</p> </div>
You can grab the tag content with:
Document doc = Jsoup.connect("your-target-url").get(); // 获取span标签内的数字 String tagContent = doc.select(".content-box .data-value").text(); int tagNumber = Integer.parseInt(tagContent); // 获取p标签内的文本 String additionalInfo = doc.select(".content-box p").text(); System.out.println("标签内数字:" + tagNumber); System.out.println("附加信息:" + additionalInfo);
If your earlier Xpath with Jsoup failed, double-check your expression—maybe you used a syntax that Jsoup doesn't fully support (Jsoup's Xpath support is limited compared to full DOM parsers). Using CSS selectors with Jsoup is usually more reliable.
用Selenium的示例
If the page is dynamically rendered (content loads after JS runs), use Selenium to wait for the element to load first:
WebDriver driver = new ChromeDriver(); driver.get("your-target-url"); // 等待元素加载完成(显式等待) WebElement tagElement = new WebDriverWait(driver, Duration.ofSeconds(10)) .until(ExpectedConditions.visibilityOfElementLocated(By.cssSelector(".data-value"))); String tagContent = tagElement.getText(); int tagNumber = Integer.parseInt(tagContent); System.out.println("标签内数字:" + tagNumber); driver.quit();
内容的提问来源于stack exchange,提问作者user7251661

