You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在HtmlUnit提取HtmlElement文本时保留换行并拆分信息

Fixing Missing Line Breaks in HtmlUnit for Extracting Owner Information

Hey there! Let's work through this HtmlUnit issue you're hitting—where getTextContent() is squashing line breaks and merging your owner info into one messy string. The old WebView workaround is outdated (HtmlUnit dumped that dependency ages ago, hence the compile error), but we've got two better, modern solutions to get those three separate strings cleanly.

The most reliable approach is to locate each piece of information in its own specific HTML element, instead of grabbing text from a parent container that mixes all three. This way, you don't have to deal with line breaks at all—each value comes pre-isolated.

Here's how to implement it:

try (WebClient webClient = new WebClient()) {
    // Disable unnecessary features to speed up loading (adjust if JS/CSS is needed for content)
    webClient.getOptions().setCssEnabled(false);
    webClient.getOptions().setJavaScriptEnabled(false);

    // Load the target page
    HtmlPage page = webClient.getPage("https://taxtest.navajocountyaz.gov/Pages/WebForm1.aspx?p=1&apn=103-03-122");

    // Extract Owner Name (adjust XPath/CSS selector to match the actual page structure)
    HtmlElement ownerNameEl = page.getFirstByXPath("//*[contains(text(), 'Owner Name')]/following-sibling::node()");
    String ownerName = ownerNameEl.getTextContent().trim();

    // Extract Street Address
    HtmlElement streetAddrEl = page.getFirstByXPath("//*[contains(text(), 'Street Address')]/following-sibling::node()");
    String streetAddress = streetAddrEl.getTextContent().trim();

    // Extract City, State & ZIP
    HtmlElement cityStateZipEl = page.getFirstByXPath("//*[contains(text(), 'City, State, ZIP')]/following-sibling::node()");
    String cityStateZip = cityStateZipEl.getTextContent().trim();

    // Verify the results
    System.out.println("Owner Name: " + ownerName);
    System.out.println("Street Address: " + streetAddress);
    System.out.println("City/State/ZIP: " + cityStateZip);
} catch (IOException e) {
    e.printStackTrace();
}

Notes for this approach:

  • Inspect the target page's HTML (using browser dev tools) to refine the XPath/CSS selectors. For example, if the owner name is inside a <span class="owner-name">, use page.querySelector(".owner-name") instead of XPath.
  • This method is far more maintainable—if the page's styling changes but the element structure stays intact, your code won't break.

Solution 2: Process Raw Text with Line Breaks

If you absolutely have to extract text from a parent container that holds all three values, you can access the raw, uncompressed text (including line breaks) by working with the element's XML content or child text nodes.

Here's an example using regex to split the raw text:

try (WebClient webClient = new WebClient()) {
    HtmlPage page = webClient.getPage("https://taxtest.navajocountyaz.gov/Pages/WebForm1.aspx?p=1&apn=103-03-122");

    // Get the parent element containing all owner info (adjust selector to match)
    HtmlElement ownerContainer = page.getFirstByXPath("//div[@class='owner-details']");

    // Get the raw XML content, which preserves line breaks
    String rawContent = ownerContainer.asXml();

    // Use regex to extract the three values (tweak pattern if needed)
    Pattern ownerPattern = Pattern.compile("(Johnson Tommy A & Nell H Cprs)\\s+?(133 Maricopa Dr)\\s+?(Winslow AZ 86047-2013)");
    Matcher matcher = ownerPattern.matcher(rawContent);

    if (matcher.find()) {
        String ownerName = matcher.group(1).trim();
        String streetAddress = matcher.group(2).trim();
        String cityStateZip = matcher.group(3).trim();
        
        // Use your extracted values here
    }
} catch (IOException e) {
    e.printStackTrace();
}

Caveat:

  • Regex is fragile—if the page's text formatting changes (extra spaces, newlines), your pattern might fail. Use this only if you can't target individual elements.

Why the Old WebView Solution Fails

HtmlUnit removed its WebView dependency years ago (starting with version 2.x, which includes both 2.47.1 and 2.69.0 you're using). That old workaround was for ancient HtmlUnit versions (1.x), so it's completely obsolete now—you can safely ignore it.


内容的提问来源于stack exchange,提问作者God's Gift To Java

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 18:15:37