You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Capterra产品页指定字段与功能提取方案

Capterra产品页面指定字段爬取方案

问题根因

直接调用父容器的.text属性时,解析工具会把容器内所有层级的文本节点无分隔拼接,不会按<li>这类子标签自动拆分,因此会出现字段混杂、功能文本粘连的问题。另外原代码遗漏了BeautifulSoup的导入语句,直接运行会抛出报错。

完整实现代码

from selenium import webdriver
from bs4 import BeautifulSoup
import time

# 初始化驱动并访问目标页面
driver = webdriver.Firefox()
driver.get("https://www.capterra.com/p/81310/AMCS/")
# 等待页面动态内容渲染完成
time.sleep(3)

# 用页面源码初始化BeautifulSoup解析器后即可关闭浏览器
soup = BeautifulSoup(driver.page_source, 'html.parser')
driver.quit()

# 提取企业所属国家
company_info_block = soup.find("ul", class_="nb-type-md nb-list-undecorated undefined")
company_country = ""
for item in company_info_block.find_all("li"):
    text = item.text.strip()
    if text.startswith("Located in"):
        company_country = text.replace("Located in ", "")
        break

# 提取官方网站URL
official_site = ""
for item in company_info_block.find_all("li"):
    text = item.text.strip()
    if text.startswith("http"):
        official_site = text
        break

# 提取结构化产品功能清单
feature_block = soup.find("div", class_="nb-col-count-1 sm:nb-col-count-2 md:nb-col-count-3 nb-col-gap-xl nb-my-0 nb-mx-auto")
feature_list = []
for item in feature_block.find_all("li"):
    feature_name = item.text.strip()
    if feature_name:
        feature_list.append(feature_name)

# 结果验证输出
print(f"企业所属国家:{company_country}")
print(f"官方网站地址:{official_site}")
print(f"产品功能数量:{len(feature_list)}")
print(f"产品功能清单:{feature_list}")

关键逻辑说明

  • 所有字段均通过遍历父容器下的独立<li>子节点提取,从根源避免父容器直接取文本导致的内容拼接问题
  • 相比全页面模糊匹配XPath的方式,在指定信息块内遍历匹配的稳定性更高,不会误抓取页面其他位置的相似文本
  • 产品功能最终输出为标准列表结构,每个元素对应一个独立功能项,无需额外做字符串分割处理
  • 代码中添加了固定等待时长适配Capterra的前端动态渲染逻辑,网络环境较差时可适当调大等待秒数,避免元素定位失败

内容的提问来源于stack exchange,提问作者BQuist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 23:15:50