使用BeautifulSoup抓取页面图片失败,仅加载图标可下载
问题:使用BeautifulSoup抓取tvnz剧集图片仅下载到加载图标,无法获取实际剧集图片
我尝试使用BeautifulSoup抓取tvnz网站科幻奇幻分类页面中的剧集图片,但运行代码后仅能下载旋转加载图标。查看页面请求标签时能看到其他图片的请求,且这些图片也包含在页面HTML的img标签中,却无法下载,请问原因是什么?
import re import requests from bs4 import BeautifulSoup site = 'https://www.tvnz.co.nz/categories/sci-fi-and-fantasy' response = requests.get(site) soup = BeautifulSoup(response.text, 'html.parser') image_tags = soup.find_all('img') urls = [img['src'] for img in image_tags] for url in urls: filename = re.search(r'/([\w_-]+[.](jpg|gif|png))$', url) if not filename: print("Regular expression didn't match with the url: {}".format(url)) continue with open(filename.group(1), 'wb') as f: if 'http' not in url: url = '{}{}'.format(site, url) response = requests.get(url) f.write(response.content) print("Download complete, downloaded images can be found in current directory!")
原因分析
- 懒加载机制导致
src属性为占位图:该网站采用图片懒加载策略,页面初始返回的img标签src指向的是旋转加载图标,实际剧集图片的真实URL通常存储在data-src、data-lazy这类自定义属性中。只有当图片进入视口时,前端JS才会把真实URL替换到src里,你的代码只提取src自然只能拿到占位图。 - 缺少请求头被网站拦截:直接用
requests.get()发起请求时,没有携带浏览器标识(User-Agent)等必要请求头,网站可能识别出这是非浏览器请求,返回的HTML中没有填充真实图片URL,或者请求真实图片时被拒绝返回。 - 相对路径拼接错误:处理相对路径时,你用完整的页面路径
site和图片URL拼接,会导致生成错误的图片地址。比如真实图片路径是https://www.tvnz.co.nz/path/to/image.jpg,你的代码会拼成https://www.tvnz.co.nz/categories/sci-fi-and-fantasy/path/to/image.jpg,请求这个错误地址自然拿不到正确图片。
内容的提问来源于stack exchange,提问作者Jonathan Devereux
相关产品推荐
相关产品推荐

