You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Jsoup爬取Purina官网返回空白结果问题求助

Fixing Empty Document Issues When Scraping Purina.ru with Jsoup

Hey, let's break down why you're getting empty tags and blank text when trying to scrape purina.ru. I see a couple of key issues in your code and some common anti-scraping hurdles you might be hitting—here's how to fix them:

1. You Forgot to Actually Fetch the Document!

Looking at your code, you set up the Connection object but never executed the request to get the HTML. That's why you're getting empty output—you haven't retrieved any content yet. Add the get() call to fetch the document:

import org.jsoup.nodes.Document;
import org.jsoup.Jsoup;
import org.jsoup.Connection;
import java.io.IOException;

public class getText {
    public static void main(String args[]) throws IOException {
        String url = "https://www.purina.ru/";
        Connection connection = Jsoup.connect(url)
                // Use a full, modern User-Agent to mimic a real browser
                .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36");
        
        // Critical line: Execute the request and get the document
        Document doc = connection.get();
        
        // Now test your outputs
        System.out.println("Page Title: " + doc.title());
        // Print first 500 characters to avoid spamming your console
        System.out.println("Page Text Preview: " + doc.text().substring(0, 500));
    }
}

2. Add More Browser-Like Request Headers

Purina's site might be blocking requests that don't look like they're coming from a real browser. Beyond User-Agent, add these common headers to make your request more authentic:

Connection connection = Jsoup.connect(url)
        .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
        .header("Accept", "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8")
        .header("Accept-Language", "ru-RU,ru;q=0.8,en-US;q=0.5,en;q=0.3")
        .header("Referer", "https://www.google.com/")
        .header("DNT", "1")
        .header("Upgrade-Insecure-Requests", "1");

3. Handle Cookies if Needed

Some sites set cookies on the first request and expect them to be present in subsequent requests. If adding headers still doesn't work, try fetching cookies first then reusing them:

// First request to grab cookies
Connection.Response initialResponse = Jsoup.connect(url)
        .userAgent("your-full-user-agent")
        .execute();

// Second request with the retrieved cookies
Document doc = Jsoup.connect(url)
        .cookies(initialResponse.cookies())
        .userAgent("your-full-user-agent")
        .get();

4. Check for Dynamic Content

If all the above fails, the site might be rendering content with JavaScript. Jsoup can't execute JS—it only parses static HTML. For dynamic sites, use a tool that mimics a browser like HtmlUnit or Selenium. Here's a quick HtmlUnit example:

import com.gargoylesoftware.htmlunit.WebClient;
import com.gargoylesoftware.htmlunit.html.HtmlPage;
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;

public class getText {
    public static void main(String[] args) throws Exception {
        // Initialize WebClient to enable JS
        try (WebClient webClient = new WebClient()) {
            webClient.getOptions().setJavaScriptEnabled(true);
            webClient.getOptions().setCssEnabled(false);
            webClient.getOptions().setThrowExceptionOnScriptError(false); // Ignore JS errors
            
            HtmlPage page = webClient.getPage("https://www.purina.ru/");
            String renderedHtml = page.asXml();
            
            // Parse the rendered HTML with Jsoup
            Document doc = Jsoup.parse(renderedHtml);
            System.out.println(doc.title());
            System.out.println(doc.text().substring(0, 500));
        }
    }
}

Start with fixing the missing get() call first—that's almost certainly the immediate issue. If that still returns empty content, work through adding headers and handling cookies. If none of that works, dynamic content is probably the culprit.

内容的提问来源于stack exchange,提问作者Thami

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:45:27