使用Python的BeautifulSoup无法获取完整图片源的技术求助
解决图片源地址截断的问题
看起来你遇到的问题是爬取到的图片地址被截断了,这通常是因为网站为了懒加载或性能优化,把完整的图片地址放在了其他属性里,而src属性仅存储了小尺寸占位图或截断的临时地址。我来帮你调整代码解决这个问题:
先查看img标签的完整内容
先不要直接提取src,而是打印整个img对象,看看它的所有属性,找到存储真实地址的字段:import requests from bs4 import BeautifulSoup url = "http://www.thehindu.com/entertainment/movies/why-does-kollywood-lack-financial-transparency/article23432094.ece" res = requests.get(url) soup = BeautifulSoup(res.content, 'html.parser') body = soup.find("body") imageparentobject = body.find("div", class_="lead-img-cont") image = imageparentobject.find("img", "lead-img") # 打印整个img标签,查看所有属性 print(image)运行后你会看到类似这样的输出(示例):
<img class="lead-img" src="http://www.the..." data-src="http://www.thehindu.com/xxx/完整的图片地址.jpg" alt="文章封面图">提取正确的属性值
从输出里找到存储完整地址的属性(常见的有data-src、data-original、data-url等),然后修改代码提取这个属性:# 假设完整地址在data-src属性中 print(image['data-src'])额外处理:如果遇到相对路径
万一图片地址是相对路径,可以用urljoin拼接成完整URL:from urllib.parse import urljoin # Python3版本 # from urlparse import urljoin # Python2版本 full_image_url = urljoin(url, image['data-src']) print(full_image_url)
另外提醒下:你的代码里print image['src']是Python2的写法,如果使用Python3需要改成print(image['src'])哦。
内容的提问来源于stack exchange,提问作者Sachin Shetti
相关产品推荐
相关产品推荐

