Python中如何从bs4.element.Tag中提取全部独立文本与href属性?
解决方案
问题原因
BeautifulSoup的get_text()方法默认会将当前节点所有后代的文本内容拼接为单个字符串返回,不会自动拆分不同子节点的独立文本get('属性名')方法仅能读取当前节点的对应属性,无法自动遍历子节点收集属性值
你需要的独立文本和href属性全部存储在目标div内部的<a>标签中,只要先提取所有子<a>标签再遍历取值即可。
实现代码
from urllib.request import urlopen from bs4 import BeautifulSoup html = urlopen("https://apptopia.com/store-insights/top-charts/google-play/comics/united-states").read() # 解析HTML建议用html.parser而非xml解析器,避免出现解析异常 soup = BeautifulSoup(html, 'html.parser') app_info_lst = soup.find_all("div", {"class": "media-object app-link-block"}) # 处理第一个元素示例 target = app_info_lst[0] # 提取内部所有a标签 a_tags = target.find_all('a') # 1. 获取独立文本列表 text_list = [a.get_text(strip=True) for a in a_tags] print(text_list) # 输出:['WEBTOON', 'WEBTOON ENTERTAINMENT'] # 2. 获取所有href属性列表 href_list = [a.get('href') for a in a_tags] print(href_list) # 输出:['https://apptopia.com/google-play/app/com.naver.linewebtoon/intelligence', '/publishers/google_play/2457079']
扩展说明
如果需要同时收集外层div的href属性,只需要在href_list前追加target.get('href')即可。
内容的提问来源于stack exchange,提问作者Boomshakalaka
相关产品推荐
相关产品推荐

