Python Beautiful Soup 提取网页图片src属性 运行后输出文本文件为空
代码问题排查及修复方案
常见问题原因
- 反爬拦截无有效返回:直接调用
requests.get()未携带请求头UA,大多数站点会拦截无标识的爬虫请求,返回403状态码或空白页面,导致BeautifulSoup无法解析到有效标签。可先打印请求状态码验证:print(requests.get(stripped_line).status_code),正常应返回200。 - CSS选择器匹配失效:
#main img要求页面存在id为main的父节点,若目标站点结构不符、或图片为前端JS动态渲染生成,静态HTML解析无法拿到对应元素。可打印请求返回的源码验证是否存在目标img标签:print(requests.get(stripped_line).text)。 - 文件名非法无法创建:
stripped_line为完整网页链接,包含/、:、?等系统文件名禁止的字符,会导致文件创建失败,内容无写入路径。需先处理链接生成合法文件名。 - 路径变量异常:若代码未提前定义
workingpath变量、或路径权限不足,会导致文件生成到未知目录或创建失败,可打印completeName验证生成路径是否符合预期。 - 图片链接属性非
src:多数站点懒加载图片的真实链接存放在data-src、data-original等自定义属性中,src多为占位空白图,可打印inp.attrs查看所有属性确认链接存放字段。
修复后参考代码
import requests from bs4 import BeautifulSoup import os # 提前定义工作路径 workingpath = "./output" os.makedirs(workingpath, exist_ok=True) # 配置请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" } with open("links.txt", "r") as a_file: for line in a_file: stripped_line = line.strip() if not stripped_line: continue # 发送带请求头的请求 resp = requests.get(stripped_line, headers=headers) resp.encoding = resp.apparent_encoding # 解决乱码问题 soup = BeautifulSoup(resp.text, "html.parser") # 生成合法文件名 safe_filename = stripped_line.replace("https://", "").replace("http://", "").replace("/", "_").replace(":", "").replace("?", "_") + ".txt" completeName = os.path.join(workingpath, safe_filename) with open(completeName, "w", encoding="utf-8") as f_out: # 可根据实际页面结构调整选择器 for inp in soup.select('#main img'): # 优先取懒加载属性,取不到再取src img_url = inp.attrs.get("data-src", inp.attrs.get("src", "")) if img_url: print(img_url, file=f_out) print(img_url)
内容的提问来源于stack exchange,提问作者Katherine Elizabeth Kath
相关产品推荐
相关产品推荐

