You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取:调用BeautifulSoup的find函数出现TypeError报错

解决Flipkart商品名称爬取的TypeError问题

问题场景

要爬取Flipkart商品页(https://www.flipkart.com/apple-iphone-14-midnight-128-gb/p/itm9e6293c322a84)中<span class="B_NuCI">标签内的商品名称,原代码运行时触发以下报错:

TypeError: slice indices must be integers or None or have an __index__ method

错误原因

原代码中这一行是核心问题:

soup=BeautifulSoup(html_content,"html.parser").prettify()

prettify()方法会将BeautifulSoup对象转换为格式化的字符串,而后续调用soup.find()时,实际是在对字符串调用方法——字符串的find()用于查找子串,和BeautifulSoup对象的find()方法逻辑完全不同,参数不匹配导致了类型错误。

修正后的代码

import requests
from bs4 import BeautifulSoup

url = "https://www.flipkart.com/apple-iphone-14-midnight-128-gb/p/itm9e6293c322a84"
# 添加请求头模拟浏览器访问,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

r = requests.get(url, headers=headers)
html_content = r.content
# 去掉prettify(),保留BeautifulSoup对象
soup = BeautifulSoup(html_content, "html.parser")
# 查找目标标签
name_tag = soup.find("span", class_="B_NuCI")
# 提取标签内的文本内容
if name_tag:
    print(name_tag.get_text(strip=True))
else:
    print("未找到目标商品名称标签")

关键改动说明

  • 移除.prettify():确保soup是BeautifulSoup对象,才能正常使用find()等解析方法
  • 添加headers:模拟浏览器请求,避免Flipkart的反爬机制拦截请求,导致无法获取正确的HTML内容
  • 增加判空逻辑:防止标签不存在时触发新的报错,提升代码健壮性
  • 使用get_text(strip=True):提取文本并自动去除首尾空白字符,得到更干净的结果

内容的提问来源于stack exchange,提问作者ANISH KUMAR RA2111026010011

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 13:40:40