You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取特定class前的<p>标签数据?Jsoup冗余数据问题求助

Fixing Redundant Data When Parsing

Tags with Jsoup

Got it, let's tackle that redundant data problem you're facing! Right now, your code grabs every <p> tag on the page, which is why you're getting extra stuff you don't need. Here's how to extract only the <p> tags that come before a specific class element, and clean up any unwanted content along the way.

Step 1: Target the Right Parent Container

First, narrow down your search to the <div class="des"> container since that's where your target <p> tags live. This avoids grabbing random <p> tags from other parts of the page:

Element desContainer = document.selectFirst("div.des");
if (desContainer == null) {
    // Handle case where the container doesn't exist
    System.out.println("Des container not found!");
    return;
}

Step 2: Collect

Tags Before Your Specific Class Element

Let's say the element you want to stop at has a class like target-stop-class (replace this with your actual target class). We'll iterate through the container's children, collecting <p> tags until we hit that specific element:

List<Element> desiredPs = new ArrayList<>();
for (Element child : desContainer.children()) {
    // Stop as soon as we encounter the specific class element
    if (child.hasClass("target-stop-class")) {
        break;
    }
    // Only add <p> tags to our list
    if ("p".equals(child.tagName())) {
        desiredPs.add(child);
    }
}

Alternatively, if you prefer using Jsoup's selector shorthand, you can directly grab all <p> tags that come before your target element:

Element stopElement = desContainer.selectFirst(".target-stop-class");
Elements desiredPs = stopElement != null ? stopElement.previousElementSiblings().select("p") : new Elements();

Step 3: Clean Up Redundant Content in

Tags

If your <p> tags have extra elements like <span class="hint"> that you don't want, remove those before getting the text:

for (Element p : desiredPs) {
    // Remove unwanted child elements (adjust the selector to match your redundant content)
    p.select(".hint").remove();
    // Trim whitespace to get clean, tidy text
    System.out.println(p.text().trim());
}

Example Output

Using your sample HTML, if we stop at a hypothetical element with class next-term-section, this would output:

  1. (with) all, do

...without the extra <span class="hint"><em>smth</em></span> content cluttering the result.

内容的提问来源于stack exchange,提问作者Sveta Tulova

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:07:22