You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何从bs4.element.Tag中提取全部独立文本与href属性?

解决方案

问题原因

  • BeautifulSoup的get_text()方法默认会将当前节点所有后代的文本内容拼接为单个字符串返回,不会自动拆分不同子节点的独立文本
  • get('属性名')方法仅能读取当前节点的对应属性,无法自动遍历子节点收集属性值
    你需要的独立文本和href属性全部存储在目标div内部的<a>标签中,只要先提取所有子<a>标签再遍历取值即可。

实现代码

from urllib.request import urlopen
from bs4 import BeautifulSoup

html = urlopen("https://apptopia.com/store-insights/top-charts/google-play/comics/united-states").read()
# 解析HTML建议用html.parser而非xml解析器,避免出现解析异常
soup = BeautifulSoup(html, 'html.parser')

app_info_lst = soup.find_all("div", {"class": "media-object app-link-block"})

# 处理第一个元素示例
target = app_info_lst[0]
# 提取内部所有a标签
a_tags = target.find_all('a')

# 1. 获取独立文本列表
text_list = [a.get_text(strip=True) for a in a_tags]
print(text_list)
# 输出:['WEBTOON', 'WEBTOON ENTERTAINMENT']

# 2. 获取所有href属性列表
href_list = [a.get('href') for a in a_tags]
print(href_list)
# 输出:['https://apptopia.com/google-play/app/com.naver.linewebtoon/intelligence', '/publishers/google_play/2457079']

扩展说明

如果需要同时收集外层div的href属性,只需要在href_list前追加target.get('href')即可。

内容的提问来源于stack exchange,提问作者Boomshakalaka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 02:36:00