如何使用SimplifiedDoc获取img标签的title与alt属性?
如何从simplified_scrapy的img元素中获取title和alt属性?
原代码
from simplified_scrapy import SimplifiedDoc, req url = 'https://google-ranking-check.org' html = req.get(url) doc = SimplifiedDoc(html) imgs = doc.listImg(url = url) print([img.url for img in imgs]) imgs = doc.selects('img') for img in imgs: print (img) print("####")
执行输出
['https://google-ranking-check.org/assets/img/googleranking.png'] {'tag': 'img', 'class': 'img-fluid rounded-circle', 'src': 'assets/img/googleRanking.PNG', 'alt': 'google ranking check', 'title': 'google ranking check', 'html': ''} ####
解决方案
从输出能看到,doc.selects('img')返回的每个img是字典结构,直接通过键名或dict.get()方法就能获取属性值:
方法1:针对doc.selects('img')返回的字典对象
from simplified_scrapy import SimplifiedDoc, req url = 'https://google-ranking-check.org' html = req.get(url) doc = SimplifiedDoc(html) imgs = doc.selects('img') for img in imgs: # 直接通过键名访问,属性不存在会抛出KeyError print(f"alt: {img['alt']}, title: {img['title']}") # 推荐用get方法,属性不存在时返回默认值(避免报错) alt_text = img.get('alt', '无alt属性') title_text = img.get('title', '无title属性') print(f"alt: {alt_text}, title: {title_text}")
方法2:针对doc.listImg()返回的封装对象
如果用doc.listImg()获取图片,返回的是封装后的对象,可以直接通过点语法访问属性:
from simplified_scrapy import SimplifiedDoc, req url = 'https://google-ranking-check.org' html = req.get(url) doc = SimplifiedDoc(html) imgs = doc.listImg(url = url) for img in imgs: print(f"alt: {img.alt}, title: {img.title}")
内容的提问来源于stack exchange,提问作者Philipp Schmid
相关产品推荐
相关产品推荐

