Python网页爬取:调用BeautifulSoup的find函数出现TypeError报错
解决Flipkart商品名称爬取的TypeError问题
问题场景
要爬取Flipkart商品页(https://www.flipkart.com/apple-iphone-14-midnight-128-gb/p/itm9e6293c322a84)中<span class="B_NuCI">标签内的商品名称,原代码运行时触发以下报错:
TypeError: slice indices must be integers or None or have an __index__ method
错误原因
原代码中这一行是核心问题:
soup=BeautifulSoup(html_content,"html.parser").prettify()
prettify()方法会将BeautifulSoup对象转换为格式化的字符串,而后续调用soup.find()时,实际是在对字符串调用方法——字符串的find()用于查找子串,和BeautifulSoup对象的find()方法逻辑完全不同,参数不匹配导致了类型错误。
修正后的代码
import requests from bs4 import BeautifulSoup url = "https://www.flipkart.com/apple-iphone-14-midnight-128-gb/p/itm9e6293c322a84" # 添加请求头模拟浏览器访问,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } r = requests.get(url, headers=headers) html_content = r.content # 去掉prettify(),保留BeautifulSoup对象 soup = BeautifulSoup(html_content, "html.parser") # 查找目标标签 name_tag = soup.find("span", class_="B_NuCI") # 提取标签内的文本内容 if name_tag: print(name_tag.get_text(strip=True)) else: print("未找到目标商品名称标签")
关键改动说明
- 移除
.prettify():确保soup是BeautifulSoup对象,才能正常使用find()等解析方法 - 添加
headers:模拟浏览器请求,避免Flipkart的反爬机制拦截请求,导致无法获取正确的HTML内容 - 增加判空逻辑:防止标签不存在时触发新的报错,提升代码健壮性
- 使用
get_text(strip=True):提取文本并自动去除首尾空白字符,得到更干净的结果
内容的提问来源于stack exchange,提问作者ANISH KUMAR RA2111026010011
相关产品推荐
相关产品推荐

