如何用BeautifulSoup提取<br>标签间文本及解决爬取异常
问题1:提取
中被
分隔的目标文本
问题分析
你当前代码的问题在于:
soup.select()返回的是匹配元素的列表,不能直接调用.replace()或.text方法- 未处理文本中的双引号和
<br>分隔的结构
解决方法
可以通过以下步骤实现:
- 用
select_one()定位到目标<p>元素(避免列表操作) - 遍历
<p>的子节点,找到包含目标前缀的文本片段 - 清理文本中的双引号和前缀内容
代码示例
from bs4 import BeautifulSoup html = ''' <div id="foo"> <p> " Data 1 : Lorem" <br> <br> " Data 2 : Ipsum" <br> </p> </div> ''' soup = BeautifulSoup(html, 'html.parser') # 定位目标p元素 p_element = soup.select_one('div#foo > p:-soup-contains("Data 1 : ")') result = "" # 遍历p元素的所有子节点 for content in p_element.contents: # 筛选出包含目标前缀的文本节点 if isinstance(content, str) and "Data 1 : " in content: # 清理双引号、前缀和多余空格 result = content.strip().strip('"').replace("Data 1 : ", "").strip() print(result)
另一种简化方式(按
拆分文本)
利用get_text(separator='\n')将<br>转换为换行符,再拆分处理:
text = p_element.get_text(separator='\n') # 过滤空行并清理每行内容 lines = [line.strip().strip('"') for line in text.split('\n') if line.strip()] for line in lines: if line.startswith("Data 1 : "): result = line.replace("Data 1 : ", "").strip() print(result)
问题2:爬取页面时触发异常的解决
问题分析
你当前代码的核心问题是:
soup.select()返回的是元素列表,列表没有.text属性,直接调用会触发AttributeError- Selenium的定位方法(如
find_element)返回单个元素,而BeautifulSoup的select返回列表,两者行为不同
解决方法
改用select_one()获取单个匹配元素,再判断元素是否存在后提取文本:
修正后的代码
try: # 用select_one获取单个匹配元素 p_element = soup.select_one('div#ui-accordion-1-panel-1 > div.tab-content-wrapper > p:-soup-contains("Collection")') if p_element: # 提取完整文本(可根据需求调整strip参数) collection = p_element.get_text().strip() else: collection = "" print("No Collection") except Exception as e: collection = "" print(f"Error occurred: {e}")
关键说明
select_one()返回第一个匹配的元素(或None),避免了列表索引操作- 增加元素存在性判断,防止因找不到元素导致的报错
- 保留异常捕获以便排查未知问题
内容的提问来源于stack exchange,提问作者ismaouste
相关产品推荐
相关产品推荐

