使用Jsoup爬取Purina官网返回空白结果问题求助
Hey, let's break down why you're getting empty tags and blank text when trying to scrape purina.ru. I see a couple of key issues in your code and some common anti-scraping hurdles you might be hitting—here's how to fix them:
1. You Forgot to Actually Fetch the Document!
Looking at your code, you set up the Connection object but never executed the request to get the HTML. That's why you're getting empty output—you haven't retrieved any content yet. Add the get() call to fetch the document:
import org.jsoup.nodes.Document; import org.jsoup.Jsoup; import org.jsoup.Connection; import java.io.IOException; public class getText { public static void main(String args[]) throws IOException { String url = "https://www.purina.ru/"; Connection connection = Jsoup.connect(url) // Use a full, modern User-Agent to mimic a real browser .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"); // Critical line: Execute the request and get the document Document doc = connection.get(); // Now test your outputs System.out.println("Page Title: " + doc.title()); // Print first 500 characters to avoid spamming your console System.out.println("Page Text Preview: " + doc.text().substring(0, 500)); } }
2. Add More Browser-Like Request Headers
Purina's site might be blocking requests that don't look like they're coming from a real browser. Beyond User-Agent, add these common headers to make your request more authentic:
Connection connection = Jsoup.connect(url) .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") .header("Accept", "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8") .header("Accept-Language", "ru-RU,ru;q=0.8,en-US;q=0.5,en;q=0.3") .header("Referer", "https://www.google.com/") .header("DNT", "1") .header("Upgrade-Insecure-Requests", "1");
3. Handle Cookies if Needed
Some sites set cookies on the first request and expect them to be present in subsequent requests. If adding headers still doesn't work, try fetching cookies first then reusing them:
// First request to grab cookies Connection.Response initialResponse = Jsoup.connect(url) .userAgent("your-full-user-agent") .execute(); // Second request with the retrieved cookies Document doc = Jsoup.connect(url) .cookies(initialResponse.cookies()) .userAgent("your-full-user-agent") .get();
4. Check for Dynamic Content
If all the above fails, the site might be rendering content with JavaScript. Jsoup can't execute JS—it only parses static HTML. For dynamic sites, use a tool that mimics a browser like HtmlUnit or Selenium. Here's a quick HtmlUnit example:
import com.gargoylesoftware.htmlunit.WebClient; import com.gargoylesoftware.htmlunit.html.HtmlPage; import org.jsoup.Jsoup; import org.jsoup.nodes.Document; public class getText { public static void main(String[] args) throws Exception { // Initialize WebClient to enable JS try (WebClient webClient = new WebClient()) { webClient.getOptions().setJavaScriptEnabled(true); webClient.getOptions().setCssEnabled(false); webClient.getOptions().setThrowExceptionOnScriptError(false); // Ignore JS errors HtmlPage page = webClient.getPage("https://www.purina.ru/"); String renderedHtml = page.asXml(); // Parse the rendered HTML with Jsoup Document doc = Jsoup.parse(renderedHtml); System.out.println(doc.title()); System.out.println(doc.text().substring(0, 500)); } } }
Start with fixing the missing get() call first—that's almost certainly the immediate issue. If that still returns empty content, work through adding headers and handling cookies. If none of that works, dynamic content is probably the culprit.
内容的提问来源于stack exchange,提问作者Thami

