在Google Colab运行Clutch爬虫无更多结果的技术问询
Clutch网站爬虫XPath结果数量不足的优化方案
问题背景
在Google Colab中运行小型爬虫,目标Clutch网站(https://clutch.co/pt/it-services)采用CloudFlare防护,已通过curl_cffi绕过并爬取到部分数据,但当前XPath语句返回的结果数量不足,需优化路径以获取更多数据。当前实现代码如下:
%pip install -q curl_cffi %pip install -q fake-useragent %pip install -q lxml from curl_cffi import requests from fake_useragent import UserAgent ua = UserAgent() headers = {'User-Agent': ua.safari} resp = requests.get('https://clutch.co/pt/it-services', headers=headers, impersonate="safari15_3") resp.status_code # 验证请求内容 from IPython.display import HTML HTML(resp.text) from lxml.html import fromstring tree = fromstring(resp.text) data = [] for company in tree.xpath('//ul/li[starts-with(@id, "provider")]'): data.append({ "name": company.xpath('./@data-title')[0].strip(), "location": company.xpath('.//span[@class = "locality"]')[0].text, "wage": company.xpath('.//div[@data-content = "<i>Avg. hourly rate</i>"]/span/text()')[0].strip(), "min_project_size": company.xpath('.//div[@data-content = "<i>Min. project size</i>"]/span/text()')[0].strip(), "employees": company.xpath('.//div[@data-content = "<i>Employees</i>"]/span/text()')[0].strip(), "description": company.xpath('.//blockquote//p')[0].text, "website_link": (company.xpath('.//a[contains(@class, "website-link__item")]/@href') or ['Not Available'])[0], }) import pandas as pd from pandas import json_normalize df = json_normalize(data, max_level=0) df
优化方案
1. 放宽顶层节点选择范围
原XPath //ul/li[starts-with(@id, "provider")] 限定了必须在ul下的li,但页面中可能存在其他容器包裹公司项。改为直接匹配所有符合id前缀的li,避免容器限制:
//li[starts-with(@id, "provider")]
2. 处理分页/动态加载数据
Clutch默认只加载首屏数据,后续数据需要分页请求。可以循环请求分页URL(如https://clutch.co/pt/it-services?page=2、page=3等),合并多页数据:
# 示例:爬取前3页数据 data = [] for page in range(1, 4): url = f'https://clutch.co/pt/it-services?page={page}' resp = requests.get(url, headers=headers, impersonate="safari15_3") if resp.status_code != 200: continue tree = fromstring(resp.text) for company in tree.xpath('//li[starts-with(@id, "provider")]'): # 数据提取逻辑... data.append(company_info)
3. 增强XPath容错性
原代码中直接用[0]提取元素,若某公司的某个字段缺失会导致报错中断,同时可能过滤掉部分结构有差异的公司项。修改所有字段的提取逻辑,增加默认值:
for company in tree.xpath('//li[starts-with(@id, "provider")]'): company_info = { "name": company.xpath('./@data-title')[0].strip() if company.xpath('./@data-title') else 'Not Available', "location": company.xpath('.//span[@class="locality"]/text()')[0].strip() if company.xpath('.//span[@class="locality"]/text()') else 'Not Available', "wage": company.xpath('.//div[@data-content="<i>Avg. hourly rate</i>"]/span/text()')[0].strip() if company.xpath('.//div[@data-content="<i>Avg. hourly rate</i>"]/span/text()') else 'Not Available', "min_project_size": company.xpath('.//div[@data-content="<i>Min. project size</i>"]/span/text()')[0].strip() if company.xpath('.//div[@data-content="<i>Min. project size</i>"]/span/text()') else 'Not Available', "employees": company.xpath('.//div[@data-content="<i>Employees</i>"]/span/text()')[0].strip() if company.xpath('.//div[@data-content="<i>Employees</i>"]/span/text()') else 'Not Available', "description": company.xpath('.//blockquote//p/text()')[0].strip() if company.xpath('.//blockquote//p/text()') else 'Not Available', "website_link": company.xpath('.//a[contains(@class, "website-link__item")]/@href')[0] if company.xpath('.//a[contains(@class, "website-link__item")]/@href') else 'Not Available', } data.append(company_info)
4. 验证页面实际结构
通过HTML(resp.text)查看返回的页面源码,确认是否有部分公司项的id前缀不是provider,或者字段的HTML结构有差异(比如data-content的格式不同),针对性调整XPath。
内容的提问来源于stack exchange,提问作者zero
相关产品推荐
相关产品推荐

