You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决BeautifulSoup .text属性丢失换行符问题:多结构HTML产品描述文本提取方案

如何用BeautifulSoup提取结构不统一的产品描述文本(保留换行)

我懂你现在的麻烦——要从结构乱七八糟的HTML里扒产品描述,一会儿文本在div里,一会儿在p里,甚至还有span,用普通的.text提取要么把所有文字挤成一团,要么把该有的换行搞丢了,太闹心了!

下面给你两个实用方案,从最简单的快速解决,到更精准的自定义处理:

最简单的快速解决方案:用get_text()自定义分隔符

BeautifulSoup的get_text()方法比.text灵活太多,你只需要给它指定分隔符,就能轻松保留基本的换行结构。直接替换你代码里的source.text为:

description = source.get_text(separator='\n', strip=True)
  • separator='\n':让不同元素里的文本用换行符分隔开,避免挤在一起
  • strip=True:自动去掉每个文本块前后的多余空格和换行,让结果更干净

这个方法几乎不需要改动现有代码,对于大多数产品描述场景来说足够好用,唯一的小缺点是可能会给内联元素(比如span)也加上换行,但整体影响不大。

更精准的方案:只给块级元素保留换行

如果你想要更贴近原网页的段落结构,比如只在p、div这类块级元素结束时添加换行,避免内联元素带来的多余换行,可以写一个简单的递归函数来收集文本:

def extract_text_with_newlines(soup):
    text_parts = []
    for element in soup.descendants:
        # 处理纯文本节点
        if isinstance(element, str):
            stripped_text = element.strip()
            if stripped_text:  # 跳过空文本
                text_parts.append(stripped_text)
        # 遇到块级元素或换行标签,添加换行
        elif element.name in ['p', 'div', 'br', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6']:
            text_parts.append('\n')
    # 合并文本,去掉连续的多余换行
    cleaned_text = '\n'.join([part for part in text_parts if part.strip() or part == '\n'])
    return cleaned_text

然后在你的代码里调用这个函数就行:

description = extract_text_with_newlines(source)

这个方法会精准识别段落类元素,保留原本的阅读结构,提取出来的文本更接近原网页的排版逻辑。

整合到你的代码里

修改后的完整代码大概是这样:

session = AsyncHTMLSession()
path = 'F:\\Users\\Zé\\Products\\'
all_products = os.listdir(path)

# 定义文本提取函数
def extract_text_with_newlines(soup):
    text_parts = []
    for element in soup.descendants:
        if isinstance(element, str):
            stripped_text = element.strip()
            if stripped_text:
                text_parts.append(stripped_text)
        elif element.name in ['p', 'div', 'br', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6']:
            text_parts.append('\n')
    cleaned_text = '\n'.join([part for part in text_parts if part.strip() or part == '\n'])
    return cleaned_text

for product_id in all_products:
    with open(path + product_id + "\\" + product_id + ".txt") as current_product_info:
        description_url = json.loads(current_product_info.read())['descriptionModule']['descriptionUrl']
        headers = {'referer': "https://aliexpress.com/item/" + product_id + ".html"}
        source = await LnN.get_source_code(description_url, False, session=session, additional_headers=headers)
        # 提取整理后的描述文本
        description = extract_text_with_newlines(source)
        # 这里可以添加保存或处理description的逻辑,比如写入文件

内容的提问来源于stack exchange,提问作者José Guedes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 13:22:39