解决BeautifulSoup .text属性丢失换行符问题:多结构HTML产品描述文本提取方案
如何用BeautifulSoup提取结构不统一的产品描述文本(保留换行)
我懂你现在的麻烦——要从结构乱七八糟的HTML里扒产品描述,一会儿文本在div里,一会儿在p里,甚至还有span,用普通的.text提取要么把所有文字挤成一团,要么把该有的换行搞丢了,太闹心了!
下面给你两个实用方案,从最简单的快速解决,到更精准的自定义处理:
最简单的快速解决方案:用get_text()自定义分隔符
BeautifulSoup的get_text()方法比.text灵活太多,你只需要给它指定分隔符,就能轻松保留基本的换行结构。直接替换你代码里的source.text为:
description = source.get_text(separator='\n', strip=True)
separator='\n':让不同元素里的文本用换行符分隔开,避免挤在一起strip=True:自动去掉每个文本块前后的多余空格和换行,让结果更干净
这个方法几乎不需要改动现有代码,对于大多数产品描述场景来说足够好用,唯一的小缺点是可能会给内联元素(比如span)也加上换行,但整体影响不大。
更精准的方案:只给块级元素保留换行
如果你想要更贴近原网页的段落结构,比如只在p、div这类块级元素结束时添加换行,避免内联元素带来的多余换行,可以写一个简单的递归函数来收集文本:
def extract_text_with_newlines(soup): text_parts = [] for element in soup.descendants: # 处理纯文本节点 if isinstance(element, str): stripped_text = element.strip() if stripped_text: # 跳过空文本 text_parts.append(stripped_text) # 遇到块级元素或换行标签,添加换行 elif element.name in ['p', 'div', 'br', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6']: text_parts.append('\n') # 合并文本,去掉连续的多余换行 cleaned_text = '\n'.join([part for part in text_parts if part.strip() or part == '\n']) return cleaned_text
然后在你的代码里调用这个函数就行:
description = extract_text_with_newlines(source)
这个方法会精准识别段落类元素,保留原本的阅读结构,提取出来的文本更接近原网页的排版逻辑。
整合到你的代码里
修改后的完整代码大概是这样:
session = AsyncHTMLSession() path = 'F:\\Users\\Zé\\Products\\' all_products = os.listdir(path) # 定义文本提取函数 def extract_text_with_newlines(soup): text_parts = [] for element in soup.descendants: if isinstance(element, str): stripped_text = element.strip() if stripped_text: text_parts.append(stripped_text) elif element.name in ['p', 'div', 'br', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6']: text_parts.append('\n') cleaned_text = '\n'.join([part for part in text_parts if part.strip() or part == '\n']) return cleaned_text for product_id in all_products: with open(path + product_id + "\\" + product_id + ".txt") as current_product_info: description_url = json.loads(current_product_info.read())['descriptionModule']['descriptionUrl'] headers = {'referer': "https://aliexpress.com/item/" + product_id + ".html"} source = await LnN.get_source_code(description_url, False, session=session, additional_headers=headers) # 提取整理后的描述文本 description = extract_text_with_newlines(source) # 这里可以添加保存或处理description的逻辑,比如写入文件
内容的提问来源于stack exchange,提问作者José Guedes
相关产品推荐
相关产品推荐

