求助:使用Jsoup无法获取目标网站指定脚本的解决办法
Hey there! Let's figure out why Jsoup isn't grabbing that script with math formulas you need, and how to fix it.
First, let's break down the common reasons this happens:
- Dynamic content loading: The script with math formulas is probably loaded after the initial page renders (via JavaScript, like async calls or frameworks such as React/Vue). Jsoup only fetches the static HTML sent from the server—it can't execute JavaScript to load content that appears later.
- External script sources: Your target script might be hosted as an external file (linked via the
srcattribute) instead of being inline. Jsoup doesn't automatically fetch external scripts unless you explicitly tell it to. - Anti-scraping blocks: Some sites reject simple HTTP requests like the ones Jsoup sends. They might require specific headers (like a user-agent), cookies, or session tokens to serve the full content.
Let's go through practical solutions:
1. Use a Headless Browser to Render Dynamic Content
Since Jsoup can't run JavaScript, you'll need a tool that mimics a real browser to capture fully rendered content. Here are two popular options:
Selenium with Headless Chrome
This is a robust choice for most dynamic sites:
// Set up headless Chrome ChromeOptions options = new ChromeOptions(); options.addArguments("--headless=new"); WebDriver driver = new ChromeDriver(options); // Load the target page driver.get(url); // Wait for the script to load (adjust timeout based on the site) new WebDriverWait(driver, Duration.ofSeconds(10)).until( ExpectedConditions.presenceOfElementLocated(By.tagName("script")) ); // Get the fully rendered HTML (including dynamically loaded scripts) String fullRenderedHtml = driver.getPageSource(); // Parse with Jsoup if you want to filter the script Document doc = Jsoup.parse(fullRenderedHtml); Elements targetScripts = doc.select("script:contains(MathJax)"); // Replace with your script's unique identifier driver.quit();
HtmlUnit
A lightweight headless browser that integrates smoothly with Jsoup and can execute basic JavaScript.
2. Explicitly Fetch External Scripts
If your target script is linked via a src attribute:
- Use Jsoup to parse the initial HTML and extract the script's absolute URL.
- Fetch the script content directly with Jsoup (or another HTTP client like OkHttp):
Document initialDoc = Jsoup.connect(url).get(); // Select the script tag matching your target (adjust the selector to fit your case) Element scriptTag = initialDoc.select("script[src*=math-script]").first(); if (scriptTag != null) { String scriptUrl = scriptTag.attr("abs:src"); // Get the absolute URL of the script // Fetch the script content (ignore content type since scripts are often application/javascript) String scriptContent = Jsoup.connect(scriptUrl) .ignoreContentType(true) .execute() .body(); // Now you have the full script content! }
3. Add Proper Request Headers to Bypass Blocks
Many sites block requests that don't have a valid user-agent or other required headers. Copy these from your browser's DevTools (Network tab) and add them to your Jsoup request:
Document doc = Jsoup.connect(url) .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") .cookie("sessionId", "your-session-cookie-if-required") .referrer("https://www.example.com") // Some sites check the referrer header .get();
4. Check for Inline Scripts Loaded Dynamically
If the script is inline but only appears after JS runs, a headless browser (like Selenium) will ensure you capture all rendered content—including these hidden inline scripts.
- Always check the site's
robots.txtand terms of service before scraping to make sure you're allowed to access the content. - If the math formulas use libraries like MathJax or KaTeX, the target script might be part of their CDN. Fetching the specific CDN URL directly could also work.
内容的提问来源于stack exchange,提问作者zDarK_NighT_

