Scrapy异步循环Meta Item混淆导致爬取异常求助
Fixing Scrapy Async Loop Data Confusion & Duplicate Entries
Alright, let's tackle this problem head-on. The core issue here is that you're reusing the same LocationItem object across multiple asynchronous requests, which causes data to get overwritten as loops run. Here's what's happening and how to fix it:
Why Your Code Is Breaking
- Shared Object References: In your
parseStatesmethod, you modify the samelocationitem for every city in the list. Since Python uses reference semantics, all thescrapy.Requestobjects you yield point to the exact same Item instance. By the time the requests are processed, thecityfield has already been updated to the last value in the loop, leading to duplicate city names and missing data. - Asynchronous Scheduling: Scrapy queues requests as you yield them, so it's normal for several state requests to be queued before the first
parseStatescallback runs. The scheduling order isn't the root problem—but the shared Item makes this behavior cause visible data errors.
Corrected Code
First, add the copy module to handle deep copies of your Item:
import copy import scrapy import xml.etree.ElementTree as ET from your_project.items import LocationItem # Replace with your actual Item import def parse(self, response): statesList = ["AC","AL","AM","BA","CE","DF","ES","GO","MA","MG","MS","MT","PA","PB","PE","PR","RJ","RN","RO","RS","SC","SE","SP"] for state in statesList: linkState = f'https://institucional.xpi.com.br/api/Escritorios/FilialListarCidadesV2?vSiglaEstado={state}' # Create a new Item for each state location = LocationItem() location['state'] = state yield scrapy.Request(url=linkState, callback=self.parseStates, meta={'item': location}) def parseStates(self, response): original_location = response.meta['item'] root = ET.fromstring(response.body) cityList = [city.text for city in root.iter('{http://schemas.datacontract.org/2004/07/XP.Portal.Entities}Nome')] for city in cityList: # Create a deep copy of the original state Item for each city location = copy.deepcopy(original_location) location['city'] = city state = location['state'] # Fix URL encoding: use & instead of & (Scrapy handles proper URL encoding) linkCity = f'https://institucional.xpi.com.br/api/Escritorios/FilialListarPorEstadoCidadeV2?vSiglaEstado={state}&vNomeCidade={city.replace(" ","%20")}' yield scrapy.Request(url=linkCity, callback=self.parseCities, meta={'item': location}) def parseCities(self, response): location = response.meta['item'] state = location['state'] city = location['city'] root = ET.fromstring(response.body) mailList = [elem.text for elem in root.iter('{http://schemas.datacontract.org/2004/07/XP.Portal.Entities}EmailPadronizadoSocioResponsavel')] companyList = [elem.text for elem in root.iter('{http://schemas.datacontract.org/2004/07/XP.Portal.Entities}RazaoSocial')] contactList = [elem.text for elem in root.iter('{http://schemas.datacontract.org/2004/07/XP.Portal.Entities}SocioResponsavel')] telList = [elem.text for elem in root.iter('{http://schemas.datacontract.org/2004/07/XP.Portal.Entities}Telefone')] for i in range(len(mailList)): # Ensure we're using the unique state/city from the copied Item write(state, city, companyList[i], contactList[i], mailList[i], telList[i])
Key Fixes Explained
- Deep Copy Items: Using
copy.deepcopy(original_location)creates a completely independent instance of yourLocationItemfor each city. This means changes to one Item won't affect others in the queue. - Clean URL Construction: Replaced
&with&in the city URL—&is an HTML entity, not needed in raw URLs. Using f-strings makes the URL easier to read and maintain. - Isolated State/City Data: Each request now carries its own unique Item with the correct state and city values, eliminating duplicate entries and missing data.
内容的提问来源于stack exchange,提问作者gnark
相关产品推荐
相关产品推荐

