如何提取嵌套网页文本并保留换行符?技术实现求助
问题描述
我需要从无明显规律、无可用类的深度嵌套网站中提取文本,得编写通用逻辑适配多场景。示例输入如下:
<div><span>Hello<br>World</span>, how are you doing?</div> <span><span>This<br><br><br>is difficult</span>, at<br>least for me.</span>
期望提取结果:
- 第一个元素:
Hello<br>World, how are you doing? - 第二个元素:
This<br><br><br>is difficult, at<br>least for me.
要求保留文本及换行符并按原顶级元素分组。
我用BeautifulSoup写了代码,但输出存在重复元素问题,求解决方案:
def is_visible_text(element): if isinstance(element, NavigableString): # Remove non-visible characters using regex text = re.sub(r'[-]', '', element) return text.strip() != '' return False def extract_deepest_text_elements(element): if isinstance(element, NavigableString) and is_visible_text(element): return [element] if element.name in ['br']: return [element] # List to hold extracted text <br> elements extracted_elements = [] # Process child elements first for child in element.contents: extracted_elements.extend(extract_deepest_text_elements(child)) return extracted_elements def refine_content(input_file, output_file): with open(input_file, 'r', encoding='utf-8') as file: content = file.read() soup = BeautifulSoup(content, 'html.parser') new_body_content = soup.new_tag('div') # Start with the highest-order elements (div, span, p) elements = soup.find_all(['div', 'span', 'p']) for elem in elements: while elem: deepest_elements = extract_deepest_text_elements(elem) if deepest_elements: for element in deepest_elements: new_body_content.append(element) new_body_content.append(soup.new_tag('br')) # Ensure BRs after text # Move up to the parent element elem = elem.parent if elem.parent and elem.parent.name != 'body' else None new_soup = BeautifulSoup('<html><body></body></html>', 'html.parser') new_soup.body.append(new_body_content) with open(output_file, 'w', encoding='utf-8') as file: file.write(new_soup.prettify())
解决方案
你的代码出现重复元素,核心原因是遍历了所有div/span/p元素后,又递归向上处理父元素,导致同一文本被多次提取。解决思路是只处理顶级目标元素(直接在body下或文档根级的div/span/p),避免重复处理子元素和父元素。
修改后的代码如下:
import re from bs4 import BeautifulSoup, NavigableString def is_visible_text(element): if isinstance(element, NavigableString): # 移除不可见字符 text = re.sub(r'[\u200B-\u200D\uFEFF]', '', element) return text.strip() != '' return False def extract_element_content(element): """提取单个元素内的所有可见文本和br标签,保留原顺序""" content_parts = [] for child in element.descendants: if isinstance(child, NavigableString) and is_visible_text(child): # 避免连续文本节点重复添加空格 if content_parts and isinstance(content_parts[-1], NavigableString): content_parts.append(child.strip()) else: content_parts.append(child) elif child.name == 'br': content_parts.append(child) # 将内容拼接成完整HTML片段 temp_soup = BeautifulSoup("<div></div>", "html.parser") for part in content_parts: temp_soup.div.append(part) return temp_soup.div.decode_contents() def refine_content(input_file, output_file): with open(input_file, 'r', encoding='utf-8') as file: content = file.read() soup = BeautifulSoup(content, 'html.parser') new_body = soup.new_tag('body') # 优先处理body下的顶级目标元素,避免递归遍历子元素 top_level_elements = [] if soup.body: top_level_elements = soup.body.find_all(['div', 'span', 'p'], recursive=False) # 若没有body标签,直接取文档根级的目标元素 if not top_level_elements: top_level_elements = [elem for elem in soup.contents if hasattr(elem, 'name') and elem.name in ['div', 'span', 'p']] for elem in top_level_elements: extracted_content = extract_element_content(elem) # 按原元素类型保存提取结果 result_elem = soup.new_tag(elem.name) result_elem.append(BeautifulSoup(extracted_content, 'html.parser')) new_body.append(result_elem) # 生成最终HTML new_soup = BeautifulSoup("<html></html>", "html.parser") new_soup.html.append(new_body) with open(output_file, 'w', encoding='utf-8') as file: file.write(new_soup.prettify())
关键修改说明:
- 避免重复提取:通过
recursive=False只获取顶级元素,不再遍历所有子元素和父元素 - 内容提取优化:遍历元素的所有后代节点,精准收集可见文本和br标签,保留原有顺序
- 结果分组匹配:每个顶级元素对应一个独立的提取结果,完全符合期望的分组格式
测试示例输入后,输出结果如下:
<html> <body> <div> Hello<br/>World, how are you doing? </div> <span> This<br/><br/><br/>is difficult, at<br/>least for me. </span> </body> </html>
内容的提问来源于stack exchange,提问作者Clms
相关产品推荐
相关产品推荐

