You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup处理HTML后,如何去除冗余空格提取规范地址?

解决BeautifulSoup提取地址时冗余空格的问题

原代码提取出的文本包含大量换行和多余空格,无法得到规整的地址格式,以下是几种可行的解决方法:

方法1:使用正则替换多余空白字符

提取文本后,用正则表达式把所有连续的空白字符(包括换行、空格、制表符等)替换成单个空格:

from bs4 import BeautifulSoup
import re

raw_text = """<div style="margin: 0 0 15px 0;">
                <b>Location:</b><br>


                        K23<br>

                    4225 Oaknoll Circle<br>Duluth, GA 30096

            </div>"""

soup = BeautifulSoup(raw_text, "html.parser")
location_text = soup.find("b", text="Location:")
parent_location = location_text.parent
parent_location.find("b").extract()  # 移除Location标签
location = parent_location.text.strip()
# 替换所有连续空白为单个空格
clean_location = re.sub(r'\s+', ' ', location)
print("location:", clean_location)

方法2:利用BeautifulSoup的stripped_strings属性

BeautifulSoup提供了stripped_strings属性,会自动忽略空白节点,返回去除前后空白的有效文本片段,直接用空格拼接即可:

from bs4 import BeautifulSoup

raw_text = """<div style="margin: 0 0 15px 0;">
                <b>Location:</b><br>


                        K23<br>

                    4225 Oaknoll Circle<br>Duluth, GA 30096

            </div>"""

soup = BeautifulSoup(raw_text, "html.parser")
location_text = soup.find("b", text="Location:")
parent_location = location_text.parent
parent_location.find("b").extract()
# 用stripped_strings获取所有有效文本并拼接
clean_location = ' '.join(parent_location.stripped_strings)
print("location:", clean_location)

方法3:遍历文本节点过滤空白

手动遍历父节点下的所有文本子节点,过滤掉空文本后拼接:

from bs4 import BeautifulSoup

raw_text = """<div style="margin: 0 0 15px 0;">
                <b>Location:</b><br>


                        K23<br>

                    4225 Oaknoll Circle<br>Duluth, GA 30096

            </div>"""

soup = BeautifulSoup(raw_text, "html.parser")
location_text = soup.find("b", text="Location:")
parent_location = location_text.parent
parent_location.find("b").extract()
# 收集所有非空且去除前后空白的文本
text_parts = [text.strip() for text in parent_location.find_all(text=True) if text.strip()]
clean_location = ' '.join(text_parts)
print("location:", clean_location)

以上三种方法都能得到目标格式:K23 4225 Oaknoll Circle Duluth, GA 30096

内容的提问来源于stack exchange,提问作者Ben David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 16:46:01