如何在HtmlUnit提取HtmlElement文本时保留换行并拆分信息
Hey there! Let's work through this HtmlUnit issue you're hitting—where getTextContent() is squashing line breaks and merging your owner info into one messy string. The old WebView workaround is outdated (HtmlUnit dumped that dependency ages ago, hence the compile error), but we've got two better, modern solutions to get those three separate strings cleanly.
Solution 1: Target Individual Elements Directly (Recommended)
The most reliable approach is to locate each piece of information in its own specific HTML element, instead of grabbing text from a parent container that mixes all three. This way, you don't have to deal with line breaks at all—each value comes pre-isolated.
Here's how to implement it:
try (WebClient webClient = new WebClient()) { // Disable unnecessary features to speed up loading (adjust if JS/CSS is needed for content) webClient.getOptions().setCssEnabled(false); webClient.getOptions().setJavaScriptEnabled(false); // Load the target page HtmlPage page = webClient.getPage("https://taxtest.navajocountyaz.gov/Pages/WebForm1.aspx?p=1&apn=103-03-122"); // Extract Owner Name (adjust XPath/CSS selector to match the actual page structure) HtmlElement ownerNameEl = page.getFirstByXPath("//*[contains(text(), 'Owner Name')]/following-sibling::node()"); String ownerName = ownerNameEl.getTextContent().trim(); // Extract Street Address HtmlElement streetAddrEl = page.getFirstByXPath("//*[contains(text(), 'Street Address')]/following-sibling::node()"); String streetAddress = streetAddrEl.getTextContent().trim(); // Extract City, State & ZIP HtmlElement cityStateZipEl = page.getFirstByXPath("//*[contains(text(), 'City, State, ZIP')]/following-sibling::node()"); String cityStateZip = cityStateZipEl.getTextContent().trim(); // Verify the results System.out.println("Owner Name: " + ownerName); System.out.println("Street Address: " + streetAddress); System.out.println("City/State/ZIP: " + cityStateZip); } catch (IOException e) { e.printStackTrace(); }
Notes for this approach:
- Inspect the target page's HTML (using browser dev tools) to refine the XPath/CSS selectors. For example, if the owner name is inside a
<span class="owner-name">, usepage.querySelector(".owner-name")instead of XPath. - This method is far more maintainable—if the page's styling changes but the element structure stays intact, your code won't break.
Solution 2: Process Raw Text with Line Breaks
If you absolutely have to extract text from a parent container that holds all three values, you can access the raw, uncompressed text (including line breaks) by working with the element's XML content or child text nodes.
Here's an example using regex to split the raw text:
try (WebClient webClient = new WebClient()) { HtmlPage page = webClient.getPage("https://taxtest.navajocountyaz.gov/Pages/WebForm1.aspx?p=1&apn=103-03-122"); // Get the parent element containing all owner info (adjust selector to match) HtmlElement ownerContainer = page.getFirstByXPath("//div[@class='owner-details']"); // Get the raw XML content, which preserves line breaks String rawContent = ownerContainer.asXml(); // Use regex to extract the three values (tweak pattern if needed) Pattern ownerPattern = Pattern.compile("(Johnson Tommy A & Nell H Cprs)\\s+?(133 Maricopa Dr)\\s+?(Winslow AZ 86047-2013)"); Matcher matcher = ownerPattern.matcher(rawContent); if (matcher.find()) { String ownerName = matcher.group(1).trim(); String streetAddress = matcher.group(2).trim(); String cityStateZip = matcher.group(3).trim(); // Use your extracted values here } } catch (IOException e) { e.printStackTrace(); }
Caveat:
- Regex is fragile—if the page's text formatting changes (extra spaces, newlines), your pattern might fail. Use this only if you can't target individual elements.
Why the Old WebView Solution Fails
HtmlUnit removed its WebView dependency years ago (starting with version 2.x, which includes both 2.47.1 and 2.69.0 you're using). That old workaround was for ancient HtmlUnit versions (1.x), so it's completely obsolete now—you can safely ignore it.
内容的提问来源于stack exchange,提问作者God's Gift To Java

