如何用Stormcrawler精准抓取网站特定子目录?
Great question—this is such a common pain point when targeting specific site sections, and Stormcrawler has a clean way to handle it without resorting to full-site crawling or missing indirect links to your target subdirectory.
Optimal Configuration Approach
The trick is balancing two needs: letting the crawler access non-/world/ pages that link back to your target subdirectory, while avoiding unnecessary crawling of unrelated site sections. Here's how to set it up:
1. Allow Site-Wide URL Access (But Only When Linked)
First, update your crawler-conf.yaml to let the crawler reach any page on www.world.com. This ensures it can access pages like /bye/new when they’re linked from /world/new:
urlfilters.config: - "+www.world.com/.*" # Permit all URLs from the target domain - "-.*" # Block external sites entirely
2. Restrict Link Extraction to Only /world/ URLs
Next, configure the link extractor to only pull links that point to your /world/ subdirectory. This means even when the crawler visits a non-/world/ page (like /bye/new), it won’t queue up other /bye/ pages—only the /world/ links from that page get added to the crawl queue:
linkextractor.config: - "+.*/world/.*" # Keep only links targeting the /world/ subdirectory - "-.*" # Discard all other links
3. Seed the Crawler with Your Target Subdirectory
Start with https://www.world.com/world/ as your seed URL. This kickstarts the crawl on your target section, and the crawler will only visit non-/world/ pages when they’re explicitly linked from /world/ pages—then extract any relevant /world/ links from those pages.
Why This Works
- No missed indirect
/world/pages: When/world/newlinks to/bye/new, the crawler fetches/bye/new, extracts the/world/nextlink inside it, and adds/world/nextto the queue. - No full-site crawl bloat: The crawler never randomly traverses unrelated subdirectories—it only visits non-/world/ pages that are directly linked from your target section.
- No custom code required: All setup uses Stormcrawler’s built-in configuration tools, so you don’t need to write custom filters.
Optional: Prevent Re-Fetching Duplicate Non-/world/ Pages
If multiple /world/ pages link to the same non-/world/ page (like /bye/new), you can ensure the crawler doesn’t re-fetch it unnecessarily. Stormcrawler’s deduplication is enabled by default, but you can verify/tweak the cache settings:
status.updater.cache.spec: "maximumSize=1000000,expireAfterAccess=24h"
This tracks fetched URLs and skips re-fetching them within the 24-hour window (adjust the time as needed).
内容的提问来源于stack exchange,提问作者Ranadip

