Scrapy新手遍历房源广告报错:list对象无xpath属性
解决Scrapy中AttributeError: 'list' object has no attribute 'xpath'的问题
问题根源
你的代码存在两个关键错误:
ads = response.xpath('//li[@class="property initial"]').getall(),末尾多了一个逗号,导致ads变成包含字符串列表的元组,而非直接的列表- 使用
getall()会返回纯字符串列表,而非Scrapy的Selector对象,字符串无法调用xpath()方法
修正后的代码
class ScrapAiaSpider(scrapy.Spider): name = 'scrap_aia' allowed_domains = ['aia-immobilier.fr'] start_urls = ['http://aia-immobilier.fr/fr/ventes'] def parse(self, response): # 去掉末尾逗号,保留SelectorList对象(不调用getall()) ads = response.xpath('//li[@class="property initial"]') for ad in ads: yield { 'type': ad.xpath('./div[@class="titles"]/h2/text()').get(), # 修正子节点xpath:原节点是li,内部价格等信息在子div而非子li中 'price': ad.xpath('./div[@class="price"]/div/text()').get(), 'room': ad.xpath('./div[@class="room"]/div/text()').get(), 'area': ad.xpath('./div[@class="area"]/div/text()').get() }
补充说明
- 直接使用
response.xpath()返回的SelectorList可以遍历每个Selector元素,每个元素支持继续调用xpath()进行子节点查询 - 同时修正了价格、房间、面积的xpath路径:原代码中错误地将子节点写成
li,实际页面结构中这些信息是在div标签下,修正后才能正确提取数据
内容的提问来源于stack exchange,提问作者kévin Goncalves
相关产品推荐
相关产品推荐

