Python爬虫:如何去除列表元素内的重复价格字符串?
解决亚马逊商品价格爬取重复问题
你的代码抓取到重复价格,是因为当前定位的父元素包含了两个相同的价格子元素,text方法会把所有子元素的文本拼接在一起,导致重复输出。以下是两种可行的修改方案:
方案1:精准定位单个价格元素
亚马逊页面中,span.a-offscreen是专门存储商品实际价格的隐藏元素,仅包含一次价格文本,直接定位这个元素就能避免重复:
import requests from bs4 import BeautifulSoup headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:94.0) Gecko/20100101 Firefox/94.0', 'Accept-Language': 'en-US, en;q=0.5' } amazn = requests.get("https://a.co/d/cTzyJwv", headers=headers) amazn_src = amazn.content soup = BeautifulSoup(amazn_src, "lxml") # 定位单个价格元素 gpu_s3r = soup.find_all("span", {"class": "a-offscreen"}) gpu_s3r_ap = [] for item in gpu_s3r: gpu_s3r_ap.append(item.text.strip()) print(gpu_s3r_ap) # 输出:['$389.00']
方案2:处理重复字符串(备选)
如果不想修改元素定位逻辑,可以在提取文本后对重复内容进行处理,取字符串的前半部分即可:
import requests from bs4 import BeautifulSoup headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:94.0) Gecko/20100101 Firefox/94.0', 'Accept-Language': 'en-US, en;q=0.5' } amazn = requests.get("https://a.co/d/cTzyJwv", headers=headers) amazn_src = amazn.content soup = BeautifulSoup(amazn_src, "lxml") gpu_s3r = soup.find_all("span",{"class":"a-price aok-align-center reinventPricePriceToPayMargin priceToPay"}) gpu_s3r_ap =[] for item in gpu_s3r: price_text = item.text.strip() # 取字符串前半部分去除重复 unique_price = price_text[:len(price_text)//2] gpu_s3r_ap.append(unique_price) print(gpu_s3r_ap) # 输出:['$389.00']
内容的提问来源于stack exchange,提问作者jwad shaqir
相关产品推荐
相关产品推荐

