Scrapy爬虫多页抓取求助:列表页转详情页提取数据失败
解决TrekEarth会员爬虫跳转详情页的问题
我帮你排查了代码里的几个关键问题,这就是你的爬虫无法正常运行的原因,以及修正后的实现方案:
你的代码存在的核心问题
- 循环提前终止:在
parse方法里,你用了return request,这会导致循环只处理第一行会员数据就直接返回,后面的会员都不会被处理,应该用yield来生成请求 - 详情页链接提取错误:
row.xpath('./td[2]//a/@href/text()')是错误的写法——@href属性本身就存储了链接值,不需要再加/text(),这样写会返回None,导致请求无效 - 相对路径未转绝对路径:页面里的会员详情链接是相对路径(比如
/members/monareng/),直接用的话Scrapy无法正确发起请求,需要用response.urljoin()转换为绝对URL - Item类需提前定义:确保你已经在
items.py里正确定义了FinalItem,包含name、other_url、groups这些字段
修正后的完整爬虫代码
import scrapy from your_project_name.items import FinalItem # 替换成你的实际项目名称 class ScrapeMembers(scrapy.Spider): name = 'trekearth_members' start_urls = [ 'https://www.trekearth.com/members/page1.htm?sort_by=md' ] def parse(self, response): # 遍历会员表格的每一行(跳过表头行) for row in response.xpath('//table[@class="member-table"]//tr[position() > 1]'): item = FinalItem() # 提取会员名称 item['name'] = row.xpath('./td[2]//a/text()').extract_first() # 提取详情页相对链接并转为绝对URL relative_url = row.xpath('./td[2]//a/@href').extract_first() if relative_url: detail_url = response.urljoin(relative_url) # 携带Item发起详情页请求 yield scrapy.Request(detail_url, callback=self.parse_page2, meta={'item': item}) # 恢复分页爬取(如果需要爬取多页会员) next_page = response.xpath('//div[@class="page-nav-btm"]/ul/li[last()]/a/@href').extract_first() if next_page: next_page_url = response.urljoin(next_page) yield scrapy.Request(next_page_url, callback=self.parse) def parse_page2(self, response): item = response.meta['item'] # 存储详情页URL item['other_url'] = response.url # 提取会员所属的所有组(如果只需要第一个组就用extract_first()) item['groups'] = response.xpath('//div[@class="groups-btm"]/ul/li/text()').extract() # 返回封装好的Item yield item
关键修改说明
- 替换return为yield:确保循环能处理所有会员行,每个详情页请求都被生成
- 修正链接提取逻辑:去掉
@href后的/text(),并通过response.urljoin()转换为绝对路径,保证请求有效 - 恢复分页逻辑:如果你需要爬取多页会员,保留这段分页代码,爬虫会自动翻页处理后续页面
- 调整groups提取方式:用
extract()代替extract_first()可以获取会员所属的所有群组,如果你只需要第一个群组,再改回extract_first()即可 - 类名优化:把
ScrapeMovies改成ScrapeMembers,更贴合爬虫的功能,当然你也可以保留原来的名字
额外注意事项
- 确保
items.py里的FinalItem定义正确,示例如下:
import scrapy class FinalItem(scrapy.Item): name = scrapy.Field() other_url = scrapy.Field() groups = scrapy.Field()
- 运行爬虫时使用命令:
scrapy crawl trekearth_members -o members.json(可以替换成你想要的输出格式,比如csv)
内容的提问来源于stack exchange,提问作者Mrowkacala
相关产品推荐
相关产品推荐

