求助:如何用Jsoup提取网页JavaScript中的键值对数据
edge_sidecar_to_children Links from JavaScript-Stored Page Data Got it, let's work through this problem together. Jsoup is fantastic for HTML parsing, but since your target data lives in JavaScript code (not plain HTML or JSON), we need to pair it with other tools or adjust our approach. Here are three practical methods to grab those edge_sidecar_to_children links:
Method 1: Jsoup + JSON Parsing (Manual JS Extraction)
First, use Jsoup to fetch the page's HTML and isolate the <script> tag holding your target data. Then, strip out the surrounding JavaScript syntax to turn the data into valid JSON.
- Step 1: Fetch the page and grab the relevant script content
Document doc = Jsoup.connect("your-target-url").get(); Elements scripts = doc.select("script"); String targetScript = ""; for (Element script : scripts) { if (script.data().contains("edge_sidecar_to_children")) { targetScript = script.data(); break; } } - Step 2: Use regex to extract the JSON-like object from the JS
Look for patterns likeedge_sidecar_to_children = { ... };and pull out the curly-braced content. Tweak the regex if the syntax varies:Pattern pattern = Pattern.compile("edge_sidecar_to_children\\s*=\\s*({.*?});", Pattern.DOTALL); Matcher matcher = pattern.matcher(targetScript); String jsonStr = ""; if (matcher.find()) { jsonStr = matcher.group(1) .replace("'", "\"") // Convert single quotes to double quotes for valid JSON .replaceAll(",\\s*}", "}"); // Fix trailing commas if present } - Step 3: Parse the JSON and extract links
Use a library like Gson or Jackson to parse the JSON string, then traverse theedge_sidecar_to_childrenstructure to get your links:Gson gson = new Gson(); Map<String, Object> dataMap = gson.fromJson(jsonStr, new TypeToken<Map<String, Object>>(){}.getType()); // Traverse the map to access edge_sidecar_to_children entries and pull out links
Method 2: Use a Headless Browser (Selenium/Playwright)
If the JavaScript data is dynamically generated or the structure is too messy to parse manually, a headless browser will run the page's JS and let you access the fully built edge_sidecar_to_children object directly.
For example, with Playwright:
try (Playwright playwright = Playwright.create()) { Browser browser = playwright.chromium().launch(new BrowserType.LaunchOptions().setHeadless(true)); Page page = browser.newPage(); page.navigate("your-target-url"); // Execute JS to grab the object directly Object sidecarData = page.evaluate("return edge_sidecar_to_children;"); // Convert the object to a JSON string or map and process links String jsonStr = new Gson().toJson(sidecarData); // ... parse and extract as needed }
This method skips regex headaches and works even for data loaded asynchronously.
Method 3: Check for Undocumented API Endpoints
You mentioned API limits, but sometimes the page loads this data via a background XHR/fetch request. Use your browser's DevTools (Network tab) to hunt for requests that return JSON containing edge_sidecar_to_children. If you find that endpoint, you might be able to call it directly (just watch for authentication or rate limits).
A quick tip: When converting JS to JSON, always validate the resulting string with a tool like JSONLint to catch syntax issues (like unescaped quotes or trailing commas) that break parsers.
内容的提问来源于stack exchange,提问作者Saif

