使用Beautiful Soup从给定HTML结构中正确提取商品价格的问题咨询
调整方案
你原来的代码错误出在find方法的属性筛选参数格式错误,没有指定筛选的属性为class,修正方法如下:
正确代码
from bs4 import BeautifulSoup # 替换为实际爬取到的HTML内容 html = """ <div class="a-section a-spacing-small a-spacing-top-small"> <span class="a-declarative" data-action="show-all-offers-display" data-show-all-offers-display="{}"> <a class="a-link-normal" href="/gp/offer-listing/B08HLZXHZY/ref=dp_olp_NEW_mbc?ie=UTF8&condition=NEW"> <span>Neu (3) ab </span><span class="a-size-base a-color-price">1.930,99 €</span> </a> </span> <span class="a-size-base a-color-base">& <b>Kostenlose Lieferung</b></span> </div> """ # 初始化BeautifulSoup对象,指定解析器 soup = BeautifulSoup(html, 'html.parser') # 写法1:使用class_参数筛选类名(推荐,避免和Python关键字class冲突) price_text = soup.find("span", class_="a-size-base a-color-price").text.strip() # 写法2:使用属性字典筛选 # price_text = soup.find("span", {"class": "a-size-base a-color-price"}).text.strip() print(price_text)
补充说明
- 代码里加的
strip()方法可以自动去除价格前后的空格、 对应的空白字符等冗余内容 - 初始化BeautifulSoup时需要指定解析器,常用的有内置的
html.parser、第三方的lxml等,不指定会抛出警告 - 如果页面存在多个相同类名的价格标签,可以把
find换成find_all批量提取所有价格
内容的提问来源于stack exchange,提问作者Boolty
相关产品推荐
相关产品推荐

