You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用SimplifiedDoc获取img标签的title与alt属性?

如何从simplified_scrapy的img元素中获取title和alt属性?

原代码

from simplified_scrapy import SimplifiedDoc, req
url = 'https://google-ranking-check.org'
html = req.get(url)
doc = SimplifiedDoc(html)
imgs = doc.listImg(url = url)
print([img.url for img in imgs])

imgs = doc.selects('img')
for img in imgs:
  print (img)
  print("####")

执行输出

['https://google-ranking-check.org/assets/img/googleranking.png']
{'tag': 'img', 'class': 'img-fluid rounded-circle', 'src': 'assets/img/googleRanking.PNG', 'alt': 'google ranking check', 'title': 'google ranking check', 'html': ''}
####

解决方案

从输出能看到,doc.selects('img')返回的每个img是字典结构,直接通过键名或dict.get()方法就能获取属性值:

方法1:针对doc.selects('img')返回的字典对象

from simplified_scrapy import SimplifiedDoc, req
url = 'https://google-ranking-check.org'
html = req.get(url)
doc = SimplifiedDoc(html)

imgs = doc.selects('img')
for img in imgs:
    # 直接通过键名访问,属性不存在会抛出KeyError
    print(f"alt: {img['alt']}, title: {img['title']}")
    
    # 推荐用get方法,属性不存在时返回默认值(避免报错)
    alt_text = img.get('alt', '无alt属性')
    title_text = img.get('title', '无title属性')
    print(f"alt: {alt_text}, title: {title_text}")

方法2:针对doc.listImg()返回的封装对象

如果用doc.listImg()获取图片,返回的是封装后的对象,可以直接通过点语法访问属性:

from simplified_scrapy import SimplifiedDoc, req
url = 'https://google-ranking-check.org'
html = req.get(url)
doc = SimplifiedDoc(html)

imgs = doc.listImg(url = url)
for img in imgs:
    print(f"alt: {img.alt}, title: {img.title}")

内容的提问来源于stack exchange,提问作者Philipp Schmid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 11:32:14