Scrapy递归折叠二级链接及多语言维基百科链接网络爬取问询
Hey there! Let's break down how to solve your two Scrapy challenges step by step—they’re both perfect for building a solid Wikipedia/Wikidata scraping workflow.
This just means controlling your crawler to only go two levels deep (starting page → direct linked pages, no further recursion). Here are two straightforward ways to pull this off:
Option 1: Set a global depth limit
The easiest way is to add this line to yoursettings.pyfile, or directly in your crawler class:DEPTH_LIMIT = 2Scrapy will automatically handle the rest—stopping recursion once it hits the second level of pages.
Option 2: Manual depth control in callbacks
If you need more flexibility (like only limiting depth for specific link types), check the current page's depth in your callback and decide whether to spawn new requests:def parse(self, response): # Process data from the current (first-level) page here current_depth = response.meta.get('depth', 0) # Only crawl second-level links (stop after that) if current_depth < 1: for link in response.css('a::attr(href)').getall(): # Filter out any unwanted links here if needed yield response.follow( link, callback=self.parse_secondary_page, meta={'depth': current_depth + 1} ) def parse_secondary_page(self, response): # Process data from the second-level page here # No new requests spawned here = recursion "folds" at level 2 pass
This workflow ties Wikipedia pages to their shared Wikidata IDs (which uniquely identify topics across languages). Here's how to implement your requested flow:
Step 1: Extract the Wikidata link from the source Wikipedia page
Wikidata links live in the sidebar of every Wikipedia page—we can target that specifically, then gather other page links to crawl:
def parse_source_wiki(self, response): # Grab the Wikidata link from the source page wikidata_link = response.css('li#t-wikidata a::attr(href)').get() if not wikidata_link: return # Extract the unique Wikidata ID (e.g., Q12345) wikidata_id = wikidata_link.split('/')[-1] base_item = { 'source_wiki_url': response.url, 'wikidata_id': wikidata_id } # Crawl all other valid Wikipedia links on the page for wiki_link in response.css('a[href^="/wiki/"]::attr(href)').getall(): # Skip non-content pages (talk, user, help, etc.) excluded_prefixes = ['Special:', 'Talk:', 'User:', 'Help:', 'Portal:'] if any(prefix in wiki_link for prefix in excluded_prefixes): continue full_link = response.urljoin(wiki_link) # Pass the Wikidata ID to the next callback via meta yield response.follow( full_link, callback=self.parse_target_wiki, meta={'base_item': base_item.copy()} )
Step 2: Process target Wikipedia pages in the new callback
In this callback, we can verify that the target page maps to the same Wikidata ID (confirming it's a cross-language version of the same topic) and save the data:
def parse_target_wiki(self, response): base_item = response.meta['base_item'] # Get the Wikidata ID from the target page target_wikidata_link = response.css('li#t-wikidata a::attr(href)').get() target_wikidata_id = target_wikidata_link.split('/')[-1] if target_wikidata_link else None # Same Wikidata ID = cross-language match if target_wikidata_id == base_item['wikidata_id']: base_item['target_wiki_url'] = response.url # Extract the language code (e.g., 'en' from en.wikipedia.org) base_item['target_language'] = response.url.split('.')[0].split('//')[-1] yield base_item # Optional: Handle pages that don't match the Wikidata ID here
Quick Pro Tips
- Don't forget to set a
DOWNLOAD_DELAY = 2insettings.pyand use a realisticUSER_AGENT—Wikipedia has anti-scraping measures, so play nice! - If you want to scale this to all languages for a topic, you could use the Wikidata API to fetch all linked Wikipedia pages directly, but the above recursive approach works great for following link networks naturally.
内容的提问来源于stack exchange,提问作者tat

