You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python爬取多标签下从Capacity开头的指定段落数据

问题修正说明

你的代码存在两个核心错误,直接修改即可得到目标内容:

  • 选择器参数错误:你要定位的规格面板div的id属性值为tab-specification,你误将该div的class属性值填入了id筛选条件,导致匹配逻辑失效
  • 选择范围错误:直接选择整个div容器会返回容器内的表格等所有内容,需要进一步定位到容器内包含Capacity内容的p标签
修正后的可运行代码
import requests
from bs4 import BeautifulSoup

url = "https://khusheimstore.com/product/makita-cordless-4-mode-impact-driver-for-18vli-ion-dtp140z-dtp140z-220/"
headers = {"Accept-Language": "en-US, en;q=0.5"}
results = requests.get(url, headers=headers)

soup = BeautifulSoup(results.text, "html.parser")

# 先定位规格面板div
spec_panel = soup.find('div', attrs={"id":"tab-specification"})
# 定位面板下包含Capacity内容的p标签
target_p = spec_panel.find('p', string=lambda text: text and 'Capacity' in text.strip())

# 提取目标内容,如需按行分割可加.split('\n')处理
capacity_content = target_p.get_text(strip=False)

# 存入列表
Techspec = []
Techspec.append(capacity_content)

# 测试打印内容
print(capacity_content)

运行以上代码即可得到从Capacity开始到p标签末尾的全部内容,如果你需要将内容处理为结构化的键值对,可按换行符拆分后逐行解析即可。

内容的提问来源于stack exchange,提问作者Zainab Alkamal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 07:15:04