You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Stormcrawler精准抓取网站特定子目录?

Great question—this is such a common pain point when targeting specific site sections, and Stormcrawler has a clean way to handle it without resorting to full-site crawling or missing indirect links to your target subdirectory.

Optimal Configuration Approach

The trick is balancing two needs: letting the crawler access non-/world/ pages that link back to your target subdirectory, while avoiding unnecessary crawling of unrelated site sections. Here's how to set it up:

1. Allow Site-Wide URL Access (But Only When Linked)

First, update your crawler-conf.yaml to let the crawler reach any page on www.world.com. This ensures it can access pages like /bye/new when they’re linked from /world/new:

urlfilters.config:
  - "+www.world.com/.*"  # Permit all URLs from the target domain
  - "-.*"                 # Block external sites entirely

2. Restrict Link Extraction to Only /world/ URLs

Next, configure the link extractor to only pull links that point to your /world/ subdirectory. This means even when the crawler visits a non-/world/ page (like /bye/new), it won’t queue up other /bye/ pages—only the /world/ links from that page get added to the crawl queue:

linkextractor.config:
  - "+.*/world/.*"  # Keep only links targeting the /world/ subdirectory
  - "-.*"            # Discard all other links

3. Seed the Crawler with Your Target Subdirectory

Start with https://www.world.com/world/ as your seed URL. This kickstarts the crawl on your target section, and the crawler will only visit non-/world/ pages when they’re explicitly linked from /world/ pages—then extract any relevant /world/ links from those pages.

Why This Works

  • No missed indirect /world/ pages: When /world/new links to /bye/new, the crawler fetches /bye/new, extracts the /world/next link inside it, and adds /world/next to the queue.
  • No full-site crawl bloat: The crawler never randomly traverses unrelated subdirectories—it only visits non-/world/ pages that are directly linked from your target section.
  • No custom code required: All setup uses Stormcrawler’s built-in configuration tools, so you don’t need to write custom filters.

Optional: Prevent Re-Fetching Duplicate Non-/world/ Pages

If multiple /world/ pages link to the same non-/world/ page (like /bye/new), you can ensure the crawler doesn’t re-fetch it unnecessarily. Stormcrawler’s deduplication is enabled by default, but you can verify/tweak the cache settings:

status.updater.cache.spec: "maximumSize=1000000,expireAfterAccess=24h"

This tracks fetched URLs and skips re-fetching them within the 24-hour window (adjust the time as needed).

内容的提问来源于stack exchange,提问作者Ranadip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 07:12:40