You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup抓取页面图片失败,仅加载图标可下载

问题:使用BeautifulSoup抓取tvnz剧集图片仅下载到加载图标,无法获取实际剧集图片

我尝试使用BeautifulSoup抓取tvnz网站科幻奇幻分类页面中的剧集图片,但运行代码后仅能下载旋转加载图标。查看页面请求标签时能看到其他图片的请求,且这些图片也包含在页面HTML的img标签中,却无法下载,请问原因是什么?

import re
import requests
from bs4 import BeautifulSoup
site = 'https://www.tvnz.co.nz/categories/sci-fi-and-fantasy'
response = requests.get(site)
soup = BeautifulSoup(response.text, 'html.parser')
image_tags = soup.find_all('img')
urls = [img['src'] for img in image_tags]
for url in urls:
    filename = re.search(r'/([\w_-]+[.](jpg|gif|png))$', url)
    if not filename:
         print("Regular expression didn't match with the url: {}".format(url))
         continue
    with open(filename.group(1), 'wb') as f:
        if 'http' not in url:
            url = '{}{}'.format(site, url)
        response = requests.get(url)
        f.write(response.content)
print("Download complete, downloaded images can be found in current directory!")

原因分析

  • 懒加载机制导致src属性为占位图:该网站采用图片懒加载策略,页面初始返回的img标签src指向的是旋转加载图标,实际剧集图片的真实URL通常存储在data-src、data-lazy这类自定义属性中。只有当图片进入视口时,前端JS才会把真实URL替换到src里,你的代码只提取src自然只能拿到占位图。
  • 缺少请求头被网站拦截:直接用requests.get()发起请求时,没有携带浏览器标识(User-Agent)等必要请求头,网站可能识别出这是非浏览器请求,返回的HTML中没有填充真实图片URL,或者请求真实图片时被拒绝返回。
  • 相对路径拼接错误:处理相对路径时,你用完整的页面路径site和图片URL拼接,会导致生成错误的图片地址。比如真实图片路径是https://www.tvnz.co.nz/path/to/image.jpg,你的代码会拼成https://www.tvnz.co.nz/categories/sci-fi-and-fantasy/path/to/image.jpg,请求这个错误地址自然拿不到正确图片。

内容的提问来源于stack exchange,提问作者Jonathan Devereux

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 07:25:19