使用BeautifulSoup处理HTML后,如何去除冗余空格提取规范地址?
解决BeautifulSoup提取地址时冗余空格的问题
原代码提取出的文本包含大量换行和多余空格,无法得到规整的地址格式,以下是几种可行的解决方法:
方法1:使用正则替换多余空白字符
提取文本后,用正则表达式把所有连续的空白字符(包括换行、空格、制表符等)替换成单个空格:
from bs4 import BeautifulSoup import re raw_text = """<div style="margin: 0 0 15px 0;"> <b>Location:</b><br> K23<br> 4225 Oaknoll Circle<br>Duluth, GA 30096 </div>""" soup = BeautifulSoup(raw_text, "html.parser") location_text = soup.find("b", text="Location:") parent_location = location_text.parent parent_location.find("b").extract() # 移除Location标签 location = parent_location.text.strip() # 替换所有连续空白为单个空格 clean_location = re.sub(r'\s+', ' ', location) print("location:", clean_location)
方法2:利用BeautifulSoup的stripped_strings属性
BeautifulSoup提供了stripped_strings属性,会自动忽略空白节点,返回去除前后空白的有效文本片段,直接用空格拼接即可:
from bs4 import BeautifulSoup raw_text = """<div style="margin: 0 0 15px 0;"> <b>Location:</b><br> K23<br> 4225 Oaknoll Circle<br>Duluth, GA 30096 </div>""" soup = BeautifulSoup(raw_text, "html.parser") location_text = soup.find("b", text="Location:") parent_location = location_text.parent parent_location.find("b").extract() # 用stripped_strings获取所有有效文本并拼接 clean_location = ' '.join(parent_location.stripped_strings) print("location:", clean_location)
方法3:遍历文本节点过滤空白
手动遍历父节点下的所有文本子节点,过滤掉空文本后拼接:
from bs4 import BeautifulSoup raw_text = """<div style="margin: 0 0 15px 0;"> <b>Location:</b><br> K23<br> 4225 Oaknoll Circle<br>Duluth, GA 30096 </div>""" soup = BeautifulSoup(raw_text, "html.parser") location_text = soup.find("b", text="Location:") parent_location = location_text.parent parent_location.find("b").extract() # 收集所有非空且去除前后空白的文本 text_parts = [text.strip() for text in parent_location.find_all(text=True) if text.strip()] clean_location = ' '.join(text_parts) print("location:", clean_location)
以上三种方法都能得到目标格式:K23 4225 Oaknoll Circle Duluth, GA 30096
内容的提问来源于stack exchange,提问作者Ben David
相关产品推荐
相关产品推荐

