You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy递归折叠二级链接及多语言维基百科链接网络爬取问询

Hey there! Let's break down how to solve your two Scrapy challenges step by step—they’re both perfect for building a solid Wikipedia/Wikidata scraping workflow.

需求一:在Scrapy中实现二级链接的递归折叠

This just means controlling your crawler to only go two levels deep (starting page → direct linked pages, no further recursion). Here are two straightforward ways to pull this off:

  • Option 1: Set a global depth limit
    The easiest way is to add this line to your settings.py file, or directly in your crawler class:

    DEPTH_LIMIT = 2
    

    Scrapy will automatically handle the rest—stopping recursion once it hits the second level of pages.

  • Option 2: Manual depth control in callbacks
    If you need more flexibility (like only limiting depth for specific link types), check the current page's depth in your callback and decide whether to spawn new requests:

    def parse(self, response):
        # Process data from the current (first-level) page here
        current_depth = response.meta.get('depth', 0)
        
        # Only crawl second-level links (stop after that)
        if current_depth < 1:
            for link in response.css('a::attr(href)').getall():
                # Filter out any unwanted links here if needed
                yield response.follow(
                    link,
                    callback=self.parse_secondary_page,
                    meta={'depth': current_depth + 1}
                )
    
    def parse_secondary_page(self, response):
        # Process data from the second-level page here
        # No new requests spawned here = recursion "folds" at level 2
        pass
    
需求二:爬取跨所有语言的维基百科链接网络(结合Wikidata)

This workflow ties Wikipedia pages to their shared Wikidata IDs (which uniquely identify topics across languages). Here's how to implement your requested flow:

Wikidata links live in the sidebar of every Wikipedia page—we can target that specifically, then gather other page links to crawl:

def parse_source_wiki(self, response):
    # Grab the Wikidata link from the source page
    wikidata_link = response.css('li#t-wikidata a::attr(href)').get()
    if not wikidata_link:
        return
    
    # Extract the unique Wikidata ID (e.g., Q12345)
    wikidata_id = wikidata_link.split('/')[-1]
    base_item = {
        'source_wiki_url': response.url,
        'wikidata_id': wikidata_id
    }

    # Crawl all other valid Wikipedia links on the page
    for wiki_link in response.css('a[href^="/wiki/"]::attr(href)').getall():
        # Skip non-content pages (talk, user, help, etc.)
        excluded_prefixes = ['Special:', 'Talk:', 'User:', 'Help:', 'Portal:']
        if any(prefix in wiki_link for prefix in excluded_prefixes):
            continue
        
        full_link = response.urljoin(wiki_link)
        # Pass the Wikidata ID to the next callback via meta
        yield response.follow(
            full_link,
            callback=self.parse_target_wiki,
            meta={'base_item': base_item.copy()}
        )

Step 2: Process target Wikipedia pages in the new callback

In this callback, we can verify that the target page maps to the same Wikidata ID (confirming it's a cross-language version of the same topic) and save the data:

def parse_target_wiki(self, response):
    base_item = response.meta['base_item']
    
    # Get the Wikidata ID from the target page
    target_wikidata_link = response.css('li#t-wikidata a::attr(href)').get()
    target_wikidata_id = target_wikidata_link.split('/')[-1] if target_wikidata_link else None

    # Same Wikidata ID = cross-language match
    if target_wikidata_id == base_item['wikidata_id']:
        base_item['target_wiki_url'] = response.url
        # Extract the language code (e.g., 'en' from en.wikipedia.org)
        base_item['target_language'] = response.url.split('.')[0].split('//')[-1]
        yield base_item
    # Optional: Handle pages that don't match the Wikidata ID here

Quick Pro Tips

  • Don't forget to set a DOWNLOAD_DELAY = 2 in settings.py and use a realistic USER_AGENT—Wikipedia has anti-scraping measures, so play nice!
  • If you want to scale this to all languages for a topic, you could use the Wikidata API to fetch all linked Wikipedia pages directly, but the above recursive approach works great for following link networks naturally.

内容的提问来源于stack exchange,提问作者tat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:28:03