ObservableValue新旧值校验:网页URL采集终止逻辑实现求助
Fixing Duplicate URL Termination for Your Web Scraper
Let's break down why your current approach isn't working and walk through a solid solution.
What's Wrong With Your Current Code?
- JavaFX Property Behavior: The
SimpleStringPropertyyou're using doesn't fire change events when you set the same value twice. So even if you calledsetPostURL()with an identical URL, yourChangeListenerwould never get notified—this is why you're never seeing "BREAK" in the output. - Misaligned Logic: Your goal is to stop scraping when you hit duplicate pages (where the entire list of URLs matches a previous page), but your listener is only checking if consecutive individual URLs are the same. Even on a duplicate page, you're looping through the URLs one by one, so adjacent values will rarely be identical.
The Fix: Track Page URL Lists Internally
The most reliable way to detect duplicate pages is to store the full list of URLs from the previous page and compare it directly with the current page's URLs. Here's how to modify your Pagination class:
import javafx.beans.property.SimpleStringProperty; import javafx.beans.property.StringProperty; import java.util.List; import java.util.Objects; public class Pagination { private final StringProperty postURL = new SimpleStringProperty(); private List<String> previousPageUrls; // Track URLs from the last scraped page public String getPostURL() { return postURL.get(); } public void setPostURL(String value) { postURL.set(value); } public StringProperty postURLProperty() { return postURL; } public void gather(int maxPages) { for (int i = 0; i < maxPages; i++) { // Fix the first page URL (your original code uses trang-0.htm, but the actual first page is xa-hoi.htm) String pageUrl = i == 0 ? "http://dantri.com.vn/xa-hoi.htm" : "http://dantri.com.vn/xa-hoi/trang-" + i + ".htm"; List<String> currentPageUrls = getAllURLToPage(pageUrl); // Check if current page is identical to the last one if (previousPageUrls != null && Objects.equals(currentPageUrls, previousPageUrls)) { System.out.println("Detected duplicate page—stopping scrape."); break; } // Publish each URL from the current page for (String url : currentPageUrls) { setPostURL(url); } // Update the previous page URL list for next iteration previousPageUrls = currentPageUrls; } } // Your existing method to extract URLs from a page private List<String> getAllURLToPage(String pageUrl) { // Replace this with your actual URL extraction logic return List.of(); } }
Updated Main Method
Now your main method can focus on handling scraped URLs, while the termination logic lives where it belongs—inside the scraper class:
public static void main(String[] args) { Pagination pagination = new Pagination(); // Keep the listener to process each scraped URL (e.g., save to DB, log) pagination.postURLProperty().addListener((observable, oldValue, newValue) -> { System.out.println("Collected URL: " + newValue); }); // Start scraping with a reasonable max page limit pagination.gather(10000); }
Optional: External Termination Trigger
If you need to trigger termination from outside the Pagination class, you can add a boolean property to signal when to stop:
Add this to the Pagination class:
import javafx.beans.property.SimpleBooleanProperty; import javafx.beans.property.BooleanProperty; private final BooleanProperty scrapeStopped = new SimpleBooleanProperty(false); public BooleanProperty scrapeStoppedProperty() { return scrapeStopped; }
Then update the gather loop condition:
for (int i = 0; i < maxPages && !scrapeStopped.get(); i++) { // ... existing code ... if (previousPageUrls != null && Objects.equals(currentPageUrls, previousPageUrls)) { scrapeStopped.set(true); break; } // ... existing code ... }
And listen for it in main:
pagination.scrapeStoppedProperty().addListener((observable, oldValue, newValue) -> { if (newValue) { System.out.println("BREAK"); } });
内容的提问来源于stack exchange,提问作者Touya Akira
相关产品推荐
相关产品推荐

