使用Python requests爬取Spotify头像时重复输出URL如何去重
问题解决方法
出现重复输出的原因是目标页面中存在2个class为bg lazy-image的div元素,二者的data-src属性值完全相同,遍历所有匹配元素时就会重复打印。
有两种常用修改方案:
方案1:直接取第一个匹配元素(更高效,适合当前场景)
你只需要获取用户头像这一个URL,不需要遍历所有匹配结果,直接取第一个返回的元素即可,同时注意不要用list作为变量名,它是Python内置关键字,容易引发冲突,修改后的代码如下:
import requests from bs4 import BeautifulSoup user_urls = ["https://open.spotify.com/user/0n7zzdkxmt0ldpo1kqugwca67", "https://open.spotify.com/user/1l23d3k5yq2v9ey191zp8uqxr", ] for url in user_urls: response = requests.get(url) html_content = response.content soup = BeautifulSoup(html_content, "html.parser") # 用find代替find_all,直接取第一个匹配元素 avatar_div = soup.find("div",{"class":"bg lazy-image"}) # 增加非空判断避免页面结构变动时报错 if avatar_div: print(avatar_div.get("data-src"))
方案2:用集合去重(适合可能存在多个不同URL需要去重的通用场景)
如果需要保留遍历所有匹配元素的逻辑,也可以通过集合存储已经打印过的URL,避免重复输出:
import requests from bs4 import BeautifulSoup user_urls = ["https://open.spotify.com/user/0n7zzdkxmt0ldpo1kqugwca67", "https://open.spotify.com/user/1l23d3k5yq2v9ey191zp8uqxr", ] # 初始化集合存储已经打印过的URL printed_urls = set() for url in user_urls: response = requests.get(url) html_content = response.content soup = BeautifulSoup(html_content, "html.parser") for avatar_div in soup.find_all("div",{"class":"bg lazy-image"}): img_url = avatar_div.get("data-src") if img_url not in printed_urls: print(img_url) printed_urls.add(img_url)
内容的提问来源于stack exchange,提问作者Ahmet Bilgin
相关产品推荐
相关产品推荐

