如何提取特定class前的<p>标签数据?Jsoup冗余数据问题求助
Got it, let's tackle that redundant data problem you're facing! Right now, your code grabs every <p> tag on the page, which is why you're getting extra stuff you don't need. Here's how to extract only the <p> tags that come before a specific class element, and clean up any unwanted content along the way.
Step 1: Target the Right Parent Container
First, narrow down your search to the <div class="des"> container since that's where your target <p> tags live. This avoids grabbing random <p> tags from other parts of the page:
Element desContainer = document.selectFirst("div.des"); if (desContainer == null) { // Handle case where the container doesn't exist System.out.println("Des container not found!"); return; }
Step 2: Collect
Tags Before Your Specific Class Element
Let's say the element you want to stop at has a class like target-stop-class (replace this with your actual target class). We'll iterate through the container's children, collecting <p> tags until we hit that specific element:
List<Element> desiredPs = new ArrayList<>(); for (Element child : desContainer.children()) { // Stop as soon as we encounter the specific class element if (child.hasClass("target-stop-class")) { break; } // Only add <p> tags to our list if ("p".equals(child.tagName())) { desiredPs.add(child); } }
Alternatively, if you prefer using Jsoup's selector shorthand, you can directly grab all <p> tags that come before your target element:
Element stopElement = desContainer.selectFirst(".target-stop-class"); Elements desiredPs = stopElement != null ? stopElement.previousElementSiblings().select("p") : new Elements();
Step 3: Clean Up Redundant Content in
Tags
If your <p> tags have extra elements like <span class="hint"> that you don't want, remove those before getting the text:
for (Element p : desiredPs) { // Remove unwanted child elements (adjust the selector to match your redundant content) p.select(".hint").remove(); // Trim whitespace to get clean, tidy text System.out.println(p.text().trim()); }
Example Output
Using your sample HTML, if we stop at a hypothetical element with class next-term-section, this would output:
- (with) all, do
...without the extra <span class="hint"><em>smth</em></span> content cluttering the result.
内容的提问来源于stack exchange,提问作者Sveta Tulova

