如何修改BeautifulSoup代码移除class为gmail_quote的div标签?
问题
现有一段Python代码,可扫描HTML_bak文件夹中的.htm文件,提取指定style属性的div标签并保存至HTML文件夹:
import os from bs4 import BeautifulSoup # Get a list of all .htm files in the HTML_bak folder html_files = [file for file in os.listdir('HTML_bak') if file.endswith('.htm')] # Loop through each HTML file for file_name in html_files: input_file_path = os.path.join('HTML_bak', file_name) output_file_path = os.path.join('HTML', file_name) # Read the input file with errors='ignore' with open(input_file_path, 'r', encoding='utf-8', errors='ignore') as input_file: input_content = input_file.read() # Parse the input content using BeautifulSoup with html5lib parser soup = BeautifulSoup(input_content, 'html5lib') main_content = soup.find('div', style='position:initial;float:left;text-align:left;overflow-wrap:break-word !important;width:98%;margin-left:5px;background-color:#FFFFFF;color:black;') # Overwrite the output file with modified content with open(output_file_path, 'w', encoding='utf-8') as output_file: output_file.write(str(main_content))
但目标div标签内部存在class为gmail_quote的div标签(示例:<div class="gmail_quote">2010/2/11 some text here .... </div>)需要移除,同时要保留指定div内<body bgColor=#ffffff>后的内容。示例HTML内容如下:
<html><body style="background-color:#FFFFFF;"><div></div></body></html><article style="width:100%;float:left; position:left;background-color:#FFFFFF; margin: 0mm 0mm 0mm 0mm; "><style> @media print { pre { overflow-x:break-word; white-space:pre; white-space:hp-pre-wrap; white-space:-moz-pre-wrap; white-space:-o-pre-wrap; white-space:-pre-wrap; white-space:pre-wrap; word-wrap:break-word;} }pre { overflow-x:break-word; white-space:pre; white-space:hp-pre-wrap; white-space:-moz-pre-wrap; white-space:-o-pre-wrap; white-space:-pre-wrap; white-space:pre-wrap; word-wrap:break-word;} @page {size: auto; margin: 12mm 4mm 12mm 6mm; } </style> <div style="position:initial;float:left;background-color:transparent;text-align:left;width:100%;margin-left:5px;"> <html><head><meta http-equiv="Content-Type" content="text/html;charset=UTF-8;"><style> .hdrfldname{color:black;font-size:20px; line-height:120%;} .hdrfldtext{overflow-wrap:break-word;color:black;font-size:20px;line-height:120%;} </style></head> <body bgColor=#ffffff> <div style="position:initial;float:left;text-align:left;font-weight:normal;width:100%;background-color:#eee9e9;"> <span class='hdrfldname'>SUBJECT: </span><span class='hdrfldtext'>lorem ipsum</span><br> <span class='hdrfldname'>FROM: </span><span class='hdrfldtext'>lorem ipsum</span><br> <span class='hdrfldname'>TO: </span><span class='hdrfldtext'>lorem ipsum</span><br> <span class='hdrfldname'>DATE: </span><span class='hdrfldtext'>2010/02/12 09:10</span><br> </div></body></html> </div> <div style="position:initial;float:left;text-align:left;overflow-wrap:break-word !important;width:98%;margin-left:5px;background-color:#FFFFFF;color:black;"><br> <html><head><meta http-equiv="Content-Type" content="text/html;charset=UTF-8;"><style> pre { overflow-x:break-word; white-space:pre; white-space:hp-pre-wrap; white-space:-moz-pre-wrap; white-space:-o-pre-wrap; white-space:-pre-wrap; white-space:pre-wrap; word-wrap:break-word;} </style></head><body bgColor=#ffffff> <div> lorem ipsum </div> <div class="gmail_quote">2010/2/11 lorem ipsum<span dir="ltr"><<a style="max-width:100%;" href="lorem ipsum">lorem ipsum</a>></span><br> </body></html> </div> </article> <div> <br></div>
请问如何修改上述代码?
解决方案
修改后的代码如下,核心是定位目标div内的指定<body>标签,提取其内容并移除其中的gmail_quote类元素:
import os from bs4 import BeautifulSoup # 获取HTML_bak文件夹下所有.htm文件 html_files = [file for file in os.listdir('HTML_bak') if file.endswith('.htm')] # 遍历每个HTML文件 for file_name in html_files: input_file_path = os.path.join('HTML_bak', file_name) output_file_path = os.path.join('HTML', file_name) # 读取输入文件,忽略编码错误 with open(input_file_path, 'r', encoding='utf-8', errors='ignore') as input_file: input_content = input_file.read() # 用html5lib解析器解析内容 soup = BeautifulSoup(input_content, 'html5lib') # 定位目标div标签 main_content = soup.find('div', style='position:initial;float:left;text-align:left;overflow-wrap:break-word !important;width:98%;margin-left:5px;background-color:#FFFFFF;color:black;') if main_content: # 找到目标div内的指定body标签 target_body = main_content.find('body', attrs={'bgColor': '#ffffff'}) if target_body: # 移除所有class为gmail_quote的div标签 for quote_div in target_body.find_all('div', class_='gmail_quote'): quote_div.decompose() # 将处理后的body内容作为输出 output_content = str(target_body) else: # 未找到body标签时,保留原div内容 output_content = str(main_content) else: # 未找到目标div时,输出空内容 output_content = "" # 写入处理后的内容到输出文件 with open(output_file_path, 'w', encoding='utf-8') as output_file: output_file.write(output_content)
关键修改说明
- 精准定位内容范围:通过
main_content.find('body', attrs={'bgColor': '#ffffff'})锁定目标div内的指定body元素,确保只保留该body包含的内容。 - 移除指定元素:使用
find_all批量获取所有class为gmail_quote的div,调用decompose()方法彻底删除这些元素及其内部内容。 - 容错处理:增加对目标div和body标签不存在的判断,避免运行时出现异常报错。
内容的提问来源于stack exchange,提问作者misaligar
相关产品推荐
相关产品推荐

