You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup提取标签内的目标文本内容?

解决BeautifulSoup提取标签文本的问题

嘿,我来帮你搞定这个问题!你遇到的核心问题是:soup.findAll()返回的是一个标签结果集(ResultSet),本质是个列表,而不是单个标签元素,所以直接在它后面加.text或者.string肯定没用~

问题原因拆解

你当前得到的[<span class="tierRank"> Master </span>]是一个包含单个span标签的列表,你需要先从列表中取出这个标签,再提取它的文本内容。

两种可行的解决方案

方案1:用find()代替findAll()(推荐,因为你只需要一个结果)

find()方法会直接返回匹配到的第一个标签元素,而不是列表,这样就能直接调用.text来提取文本了,再用strip()去掉前后的空格:

theurl = "http://na.op.gg/summoner/userName=Darshan"
thepage = urlopen(theurl)
soup = BeautifulSoup(thepage,"html.parser")
rank_tag = soup.find('span', {"class":"tierRank"})
if rank_tag:  # 避免找不到标签时报错
    print(rank_tag.text.strip())  # 输出:Master

方案2:继续用findAll(),但要取列表中的第一个元素

如果你坚持用findAll(),记得先判断列表不为空,再通过索引取出第一个标签,然后提取文本:

theurl = "http://na.op.gg/summoner/userName=Darshan"
thepage = urlopen(theurl)
soup = BeautifulSoup(thepage,"html.parser")
rank_list = soup.findAll('span', {"class":"tierRank"})
if rank_list:  # 确保列表有内容,防止索引越界
    print(rank_list[0].text.strip())  # 输出:Master

补充说明

  • .text和.get_text()是等价的,都能提取标签内的所有文本内容;.string则适用于标签内只有文本没有子标签的情况,这里用.text更稳妥。
  • 加上strip()是因为原标签里的文本前后有空格,去掉后更干净。

内容的提问来源于stack exchange,提问作者Naveen Manoharan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:53:04