You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Google Colab运行Clutch爬虫无更多结果的技术问询

Clutch网站爬虫XPath结果数量不足的优化方案

问题背景

在Google Colab中运行小型爬虫,目标Clutch网站(https://clutch.co/pt/it-services)采用CloudFlare防护,已通过curl_cffi绕过并爬取到部分数据,但当前XPath语句返回的结果数量不足,需优化路径以获取更多数据。当前实现代码如下:

%pip install -q curl_cffi
%pip install -q fake-useragent
%pip install -q lxml

from curl_cffi import requests
from fake_useragent import UserAgent

ua = UserAgent()    
headers = {'User-Agent': ua.safari}
resp = requests.get('https://clutch.co/pt/it-services', headers=headers, impersonate="safari15_3")
resp.status_code

# 验证请求内容
from IPython.display import HTML
HTML(resp.text)

from lxml.html import fromstring
tree = fromstring(resp.text)

data = []

for company in tree.xpath('//ul/li[starts-with(@id, "provider")]'):
    data.append({
        "name": company.xpath('./@data-title')[0].strip(),
        "location": company.xpath('.//span[@class = "locality"]')[0].text,
        "wage": company.xpath('.//div[@data-content = "<i>Avg. hourly rate</i>"]/span/text()')[0].strip(),
        "min_project_size": company.xpath('.//div[@data-content = "<i>Min. project size</i>"]/span/text()')[0].strip(),
        "employees": company.xpath('.//div[@data-content = "<i>Employees</i>"]/span/text()')[0].strip(),
        "description": company.xpath('.//blockquote//p')[0].text,
        "website_link": (company.xpath('.//a[contains(@class, "website-link__item")]/@href') or ['Not Available'])[0],
    })

import pandas as pd
from pandas import json_normalize
df = json_normalize(data, max_level=0)
df

优化方案

1. 放宽顶层节点选择范围

原XPath //ul/li[starts-with(@id, "provider")] 限定了必须在ul下的li,但页面中可能存在其他容器包裹公司项。改为直接匹配所有符合id前缀的li,避免容器限制:

//li[starts-with(@id, "provider")]

2. 处理分页/动态加载数据

Clutch默认只加载首屏数据,后续数据需要分页请求。可以循环请求分页URL(如https://clutch.co/pt/it-services?page=2、page=3等),合并多页数据:

# 示例:爬取前3页数据
data = []
for page in range(1, 4):
    url = f'https://clutch.co/pt/it-services?page={page}'
    resp = requests.get(url, headers=headers, impersonate="safari15_3")
    if resp.status_code != 200:
        continue
    tree = fromstring(resp.text)
    for company in tree.xpath('//li[starts-with(@id, "provider")]'):
        # 数据提取逻辑...
        data.append(company_info)

3. 增强XPath容错性

原代码中直接用[0]提取元素,若某公司的某个字段缺失会导致报错中断,同时可能过滤掉部分结构有差异的公司项。修改所有字段的提取逻辑,增加默认值:

for company in tree.xpath('//li[starts-with(@id, "provider")]'):
    company_info = {
        "name": company.xpath('./@data-title')[0].strip() if company.xpath('./@data-title') else 'Not Available',
        "location": company.xpath('.//span[@class="locality"]/text()')[0].strip() if company.xpath('.//span[@class="locality"]/text()') else 'Not Available',
        "wage": company.xpath('.//div[@data-content="<i>Avg. hourly rate</i>"]/span/text()')[0].strip() if company.xpath('.//div[@data-content="<i>Avg. hourly rate</i>"]/span/text()') else 'Not Available',
        "min_project_size": company.xpath('.//div[@data-content="<i>Min. project size</i>"]/span/text()')[0].strip() if company.xpath('.//div[@data-content="<i>Min. project size</i>"]/span/text()') else 'Not Available',
        "employees": company.xpath('.//div[@data-content="<i>Employees</i>"]/span/text()')[0].strip() if company.xpath('.//div[@data-content="<i>Employees</i>"]/span/text()') else 'Not Available',
        "description": company.xpath('.//blockquote//p/text()')[0].strip() if company.xpath('.//blockquote//p/text()') else 'Not Available',
        "website_link": company.xpath('.//a[contains(@class, "website-link__item")]/@href')[0] if company.xpath('.//a[contains(@class, "website-link__item")]/@href') else 'Not Available',
    }
    data.append(company_info)

4. 验证页面实际结构

通过HTML(resp.text)查看返回的页面源码,确认是否有部分公司项的id前缀不是provider,或者字段的HTML结构有差异(比如data-content的格式不同),针对性调整XPath。

内容的提问来源于stack exchange,提问作者zero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 22:55:04