如何使用Scrapy爬取详情页而非列表页?附示例代码
如何用Scrapy爬取目标网站详情页内容
现有代码的问题
allowed_domains配置错误,目标域名是ssr1.scrape.center,原配置的域名会导致合法请求被Scrapy过滤- 详情页解析的XPath存在语法错误(如
address字段的XPath开头缺失//) - 提取详情链接时使用
extract()[0]存在索引越界风险,建议用更安全的get()或extract_first() - 下一页链接拼接逻辑有漏洞,当没有下一页时会生成无效URL
修正后的完整代码
import scrapy from scrapytutorial.items import ScrapytutorialItem class FirstprojectSpider(scrapy.Spider): name = 'firstproject' # 修正允许的域名 allowed_domains = ['ssr1.scrape.center'] # 替换为目标网站首页,或直接删除(因重写了start_requests) start_urls = ['https://ssr1.scrape.center'] def start_requests(self): # 遍历1-10页列表页 for page in range(1, 11): url = f'https://ssr1.scrape.center/page/{page}' yield scrapy.Request( url=url, headers={ 'user_agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/105.0.0.0 Safari/537.36', 'host': 'ssr1.scrape.center', 'Referer': 'https://ssr1.scrape.center/' }, callback=self.parse ) def parse(self, response): """处理列表页,提取详情页链接""" # 优化XPath,通过class定位电影卡片更稳定 movie_list = response.xpath('//div[contains(@class, "el-card")]') for card in movie_list: # 安全提取详情链接,避免索引越界 detail_path = card.xpath('.//a[@class="name"]/@href').get() if detail_path: # 使用response.follow自动拼接完整URL,代码更简洁 yield response.follow( url=detail_path, callback=self.getdetail ) # 修正下一页链接提取逻辑 next_page_path = response.xpath('//ul[@class="el-pager"]/li[@class="btn-next"]/a/@href').get() if next_page_path and next_page_path != '#': yield response.follow( url=next_page_path, callback=self.parse ) def getdetail(self, response): """解析详情页内容""" items = ScrapytutorialItem() # 优化XPath,通过class定位元素,同时处理空值情况 items['name'] = response.xpath('//h2[@class="m-b-sm"]/text()').get().strip() if response.xpath('//h2[@class="m-b-sm"]/text()').get() else None items['address'] = response.xpath('//div[@class="m-v-sm info"]/span[1]/text()').get().strip() if response.xpath('//div[@class="m-v-sm info"]/span[1]/text()').get() else None items['times'] = response.xpath('//div[@class="m-v-sm info"]/span[3]/text()').get().strip() if response.xpath('//div[@class="m-v-sm info"]/span[3]/text()').get() else None items['timestamp'] = response.xpath('//div[@class="m-v-sm info"]/span[last()]/text()').get().strip() if response.xpath('//div[@class="m-v-sm info"]/span[last()]/text()').get() else None items['plot'] = response.xpath('//div[@class="m-v-sm synopsis"]/p/text()').get().strip() if response.xpath('//div[@class="m-v-sm synopsis"]/p/text()').get() else None yield items
关键优化点说明
- 域名配置:
allowed_domains必须匹配目标网站域名,否则Scrapy会拦截跨域请求 - 链接提取:用
get()替代extract()[0],元素不存在时返回None,避免抛出索引错误 - XPath稳定性:优先通过class属性定位元素,比依赖固定id更能适配页面结构的微调
- URL拼接:使用
response.follow自动处理基础URL拼接,无需手动拼接完整链接,代码更简洁
内容的提问来源于stack exchange,提问作者Aloha
相关产品推荐
相关产品推荐

