BeautifulSoup获取商品色板异常:仅返回单个结果或空值
解决BeautifulSoup抓取Sanmar衬衫色板数据的问题
问题背景
我用BeautifulSoup抓取Sanmar网站(目标URL:https://www.sanmar.com/p/2383_RoyClsNvy?text=s508)上衬衫的价格和色板数据,价格抓取正常,但色板要么只返回第一个,要么返回空值。使用了Zenrows,但不影响问题。以下是两段有问题的异步代码:
第一段代码(仅返回第一个色板)
async def call_url(url): try: response = await client.get_async(url) if (response.ok): soup = BeautifulSoup(response.text, "html.parser") prices = soup.find_all(class_='price')[0].get_text() price1 = prices[15] price2 = prices[16] price3 = prices[17] price4 = prices[18] price5 = prices[19] price = price1 + price2 + price3 + price4 + price5 newprice = "=" + price + "+ 8" return { "style": soup.find_all(class_= 'product-style-number')[0].get_text(), "price": price, "new price": newprice, "colors": soup.find_all(class_='swatches')[0].get_text() } except Exception as e: pass
第二段代码(返回空值)
async def call_url(url): try: response = await client.get_async(url) if (response.ok): soup = BeautifulSoup(response.text, "html.parser") prices = soup.find_all(class_='price')[0].get_text() price1 = prices[15] price2 = prices[16] price3 = prices[17] price4 = prices[18] price5 = prices[19] price = price1 + price2 + price3 + price4 + price5 newprice = "=" + price + "+ 8" colors = soup.find_all('div', attrs={'class':'swatch-name'})[0].get_text() for color in colors: swatch = [color] return { "style": soup.find_all(class_= 'product-style-number')[0].get_text(), "price": price, "new price": newprice, "colors": swatch } except Exception as e: pass
问题原因
- 第一段代码:只取了第一个
.swatches元素的文本,且直接用get_text()会把子元素文本混在一起,无法区分单个色板。 - 第二段代码:只取了第一个
.swatch-name元素的文本,之后的循环是遍历该文本的每个字符,最后swatch只保留最后一个字符,导致返回异常。
修复后的代码
async def call_url(url): try: response = await client.get_async(url) if response.ok: soup = BeautifulSoup(response.text, "html.parser") # 保留原有价格抓取逻辑 prices = soup.find_all(class_='price')[0].get_text() price1 = prices[15] price2 = prices[16] price3 = prices[17] price4 = prices[18] price5 = prices[19] price = price1 + price2 + price3 + price4 + price5 newprice = "=" + price + "+ 8" # 修复色板抓取逻辑:获取所有色板名称 color_elements = soup.find_all('div', class_='swatch-name') # 提取每个色板的文本,去除空白,过滤空值 colors = [elem.get_text(strip=True) for elem in color_elements if elem.get_text(strip=True)] return { "style": soup.find(class_='product-style-number').get_text(strip=True), "price": price, "new price": newprice, "colors": colors } except Exception as e: # 打印异常便于排查 print(f"抓取出错:{e}") return None
关键修复点
- 改用
soup.find_all('div', class_='swatch-name')获取全部色板元素,而非仅第一个。 - 用列表推导式遍历元素,提取文本并清理空白,同时过滤无效的空文本。
- 把
soup.find_all()[0]替换为soup.find(),因为product-style-number是唯一元素,更高效。 - 添加异常打印,方便后续排查问题。
内容的提问来源于stack exchange,提问作者Kyle
相关产品推荐
相关产品推荐

