如何用Jsoup解析含多表格的网页提取指定数据?
Hey there! Let's walk through how to solve this problem—extracting model names and detail page URLs from the AMD A4 Series list page, then parsing each detail page using Jsoup. Since you're new to Jsoup but know HTML/CSS, we'll break this into clear, actionable steps.
Step 1: Parse the Main List Page to Get Model Names & Detail URLs
First, let's target the main page (https://www.cpu-world.com/info/AMD/AMD_A4-Series.html). If you inspect the page's HTML, you'll see CPU models are listed in a structured table where each row has a link to the detail page. Here's how to extract those key details:
// Run this in a background thread (critical for Android!) try { Document mainDoc = Jsoup.connect("https://www.cpu-world.com/info/AMD/AMD_A4-Series.html") .timeout(5000) .get(); // Target the table containing CPU models (verify selector via browser dev tools if needed) Elements cpuRows = mainDoc.select("table.info > tbody > tr:not(:first-child)"); List<CpuModel> cpuModels = new ArrayList<>(); for (Element row : cpuRows) { // Grab the link element that holds the model name and detail URL Element link = row.selectFirst("td:first-child > a"); if (link != null) { String modelName = link.text().trim(); // Convert relative URLs to full absolute URLs with absUrl() String detailUrl = link.absUrl("href"); cpuModels.add(new CpuModel(modelName, detailUrl)); } } // Now pass the list of models to parse their detail pages parseDetailPages(cpuModels); } catch (IOException e) { e.printStackTrace(); // Add user-facing error handling here (e.g., show a toast in Android) } // Helper class to store model name and URL static class CpuModel { String name; String detailUrl; CpuModel(String name, String detailUrl) { this.name = name; this.detailUrl = detailUrl; } }
Key Notes for Step 1:
- Selector Breakdown:
table.info > tbody > tr:not(:first-child)targets all data rows in the main table, skipping the header row. Adjust this selector if the page's HTML structure changes—always double-check with your browser's dev tools. - absUrl("href"): This fixes relative links (like
/CPUs/K10/AMD-A4-Series%20A4-3300.html) to full absolute URLs, which Jsoup needs to load detail pages correctly.
Step 2: Parse Each Detail Page (Refining Your Existing Code)
You already have a solid start on detail page parsing—let's polish it and integrate it with the model list we extracted. Here's an optimized version:
private void parseDetailPages(List<CpuModel> cpuModels) { for (CpuModel model : cpuModels) { try { Document detailDoc = Jsoup.connect(model.detailUrl) .timeout(3000) .get(); // Target the spec table (your original selector is correct) Element specTable = detailDoc.selectFirst("table.spec_table"); if (specTable == null) { continue; // Skip if the spec table isn't found for this model } Elements rows = specTable.select("tr"); // Skip the header row using sublist() for cleaner code for (Element row : rows.subList(1, rows.size())) { Elements cols = row.select("td"); if (cols.size() >= 2) { // Avoid null errors by checking column count String rowName = cols.get(0).text().trim(); String rowValue = cols.get(1).text().trim(); // Do something with the data—log it, store in a database, etc. System.out.println(model.name + ": " + rowName + " = " + rowValue); } } } catch (IOException e) { e.printStackTrace(); // Handle individual page errors without stopping the entire process } } }
Improvements to Your Original Code:
- Simplified Iteration: Using
rows.subList(1, rows.size())skips the header row without needing an iterator, making the code more readable. - Null Safety: Added checks for the spec table and column count to prevent unexpected crashes if a page's structure differs slightly.
- Model Association: Ties each spec entry to its parent CPU model, so you can easily track which specs belong to which model.
Critical Android Considerations
Since you're working in Android:
- Network Permission: Add this line to your
AndroidManifest.xmlto allow internet access:<uses-permission android:name="android.permission.INTERNET"/> - Background Threads: Never run network operations (like Jsoup's
connect().get()) on the main thread. Use Kotlin Coroutines,AsyncTask(deprecated but functional for simple cases), or aThreadPoolExecutorto handle background work. - User Feedback: Replace raw stack trace prints with user-friendly error messages (e.g., a toast saying "Failed to load CPU data") to improve the app experience.
内容的提问来源于stack exchange,提问作者Ziggy

