You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup抓取并拼接<br>标签内的文本?

解决方法

你的代码报错的原因是:find_all("br")返回的是标签列表,列表没有.text属性,而且目标区域的文本并非存放在<br>标签内,而是在该标签前后的文本节点中。以下是两种可行的抓取方案:

方案一:使用get_text()提取并格式化文本

直接通过目标元素的get_text()方法提取所有文本,指定分隔符替换<br>,再清理多余空白:

from bs4 import BeautifulSoup
import requests

url = "https://www.booklooker.de/Bücher/Bastian-Sick+Der-Dativ-ist-dem-Genitiv-sein-Tod-Ein-Wegweiser-durch-den-Irrgarten-der-deutschen/id/A02ArCkS01ZZy"
page = requests.get(url)
souped = BeautifulSoup(page.content, "html.parser")

# 定位目标元素
desc_element = souped.find(class_="propertyItem_13")
# 提取文本,用换行替换<br>,同时去除首尾空白
raw_desc = desc_element.get_text(separator="\n", strip=True)
# 清理多余空行
cleaned_desc = "\n".join(line.strip() for line in raw_desc.split("\n") if line.strip())

print(cleaned_desc)

方案二:遍历文本节点提取

利用stripped_strings遍历所有非空白文本节点,再拼接成完整内容:

from bs4 import BeautifulSoup
import requests

url = "https://www.booklooker.de/Bücher/Bastian-Sick+Der-Dativ-ist-dem-Genitiv-sein-Tod-Ein-Wegweiser-durch-den-Irrgarten-der-deutschen/id/A02ArCkS01ZZy"
page = requests.get(url)
souped = BeautifulSoup(page.content, "html.parser")

desc_element = souped.find(class_="propertyItem_13")
# 提取所有清理过首尾空白的文本片段
text_fragments = [text for text in desc_element.stripped_strings]
# 用换行拼接成完整描述
cleaned_desc = "\n".join(text_fragments)

print(cleaned_desc)

这两种方法都能准确抓取截图所示区域的所有文本,并且输出格式整洁。

内容的提问来源于stack exchange,提问作者Benedikt Faude

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 02:40:41