Scrapy CrawlSpider报错‘str对象无iter属性’的原因与解决方法
Scrapy CrawlSpider 报错:AttributeError: 'str' object has no attribute 'iter'
报错信息
AttributeError: 'str' object has no attribute 'iter' 2024-03-15 14:01:19 [scrapy.core.engine] INFO: Closing spider (finished)
爬虫代码
class AuctionSpider(CrawlSpider): name = "auction" allowed_domains = ["auct.co.th"] start_urls = ["https://www.auct.co.th/products"] rules = (Rule(LinkExtractor(restrict_xpaths="//div[@class='p-2 card']/text()"), callback="parse_item", follow=True),) def parse_item(self, response): yield { 'auction_date': response.xpath("//b[@id ='product_auction_date']/text()").get(), 'price_start': response.xpath("//b[@id ='product_price_start']/text()").get(), 'order': response.xpath("//b[@id ='product_order']/b/text()").get(), 'product_title': response.xpath("//div[@class ='col-md-12']/b/text()").get(), 'product_regis_id': response.xpath("//div[@class ='col-sm-12 col-md-12 col-xl-12']/b/text()").get(), 'total_drive': response.xpath("//b[@id='product_total_drive']/text()").get(), 'product_gear': response.xpath("//b[@id='product_gear']/text()").get(), 'product_color': response.xpath("//b[@id='product_color']/text()").get(), 'cc': response.xpath("//b[@id='product_engin_cc']/text()").get(), 'regis_year': response.xpath("//b[@id='product_regis_year']/text()").get(), 'build_year': response.xpath("//b[@id='product_build_year']/text()").get(), 'gas_type': response.xpath("//b[@id='product_gas_type']/text()").get(), 'vin_no': response.xpath("//b[@id='product_body_number']/text()").get(), 'engine_no': response.xpath("//b[@id='product_engin_number']/text()").get(), 'endtax': response.xpath("//b[@id='product_endtax']/text()").get(), 'stock': response.xpath("//b[@id='product_oderstock']/text()").get(), 'price': response.xpath("//b[@id='product_price_other']/text()").get(), 'gadget': response.xpath("//b[@id='product_gadget']/text()").get(), 'remark': response.xpath("//b[@id='product_remark']/text()").get(), }
错误原因
LinkExtractor的restrict_xpaths参数要求传入节点元素的XPath表达式,用于限定提取链接的范围。但你写的//div[@class='p-2 card']/text()会提取该div下的文本内容(字符串类型),Scrapy内部会尝试对这个字符串调用iter()方法遍历节点,而字符串没有该属性,因此触发AttributeError。
修复方案
修正
restrict_xpaths的XPath:
将XPath修改为指向包含目标链接的节点,而非文本。目标链接通常在该div下的<a>标签内,可调整为:rules = (Rule(LinkExtractor(restrict_xpaths="//div[@class='p-2 card']"), callback="parse_item", follow=True),)或者更精确地指向
<a>标签:rules = (Rule(LinkExtractor(restrict_xpaths="//div[@class='p-2 card']/a"), callback="parse_item", follow=True),)优化
parse_item中的XPath(可选):order字段的XPath//b[@id='product_order']/b/text()多了一层<b>嵌套,实际应为//b[@id='product_order']/text(),否则会无法提取到内容。
内容的提问来源于stack exchange,提问作者Spatdy
相关产品推荐
相关产品推荐

