如何用Python为HTML字符串逐行添加\n换行符(不使用正则)
解决方法
问题原因
原代码针对块级开放标签的替换逻辑仅能匹配无属性的<标签名>格式,带属性的<标签名 属性="值">无法被简单字符串替换匹配,因此ol开放标签后没有生成换行。
实现思路
无需正则,通过手动遍历转义后的HTML字符串识别开放标签,提取标签名后匹配块级标签列表,符合条件的在标签结束位置后插入换行。
完整修改代码
TAGS = ['p', 'h1', 'h2', 'h3', 'h4', 'li', 'img','ol'] SINGLE_LINE_TAGS = ['ul', 'ol'] INLINE_TAGS = ['strong', 'i', 'u', 'em'] html = '''<ol class="X5LH0c"><li class="TrT0Xe" id="hello">Create A Bakery Business Plan. ... </li><li class="TrT0Xe" id="hello">Choose A Location For Your Bakery Business. ... </li><li class="TrT0Xe">Get All Licenses Required To Open A Bakery Business In India. ... </li><li class="TrT0Xe">Get Manpower Required To Open A Bakery. ... </li><li class="TrT0Xe">Buy Equipment Needed To Start A Bakery Business.</li></ol>''' # 处理所有闭合标签的换行 for tag in TAGS: html = html.replace('</{}>'.format(tag), '</{}>\n'.format(tag)) # 处理块级标签闭合标签的重复换行(可选,避免重复添加换行) for tag in SINGLE_LINE_TAGS: html = html.replace('</{}>\n\n'.format(tag), '</{}>\n'.format(tag)) # 处理自闭合标签换行 html = html.replace(' />', ' />\n') # 处理带属性的块级开放标签换行 def add_newline_after_block_open_tags(html, block_tags): result = [] i = 0 html_len = len(html) while i < html_len: # 匹配转义后的标签开头 < if html[i:i+4] == '<': tag_start_idx = i i += 4 # 提取标签名 tag_name_buf = [] while i < html_len and html[i] not in (' ', '&'): tag_name_buf.append(html[i]) i += 1 tag_name = ''.join(tag_name_buf) # 跳过闭合标签、注释标签 if tag_name.startswith('/') or tag_name.startswith('!'): result.append(html[tag_start_idx:i]) continue # 查找当前标签的结束位置 > while i < html_len and html[i:i+4] != '>': i += 1 if i + 4 > html_len: result.append(html[tag_start_idx:]) break i += 4 # 写入完整标签 result.append(html[tag_start_idx:i]) # 块级标签追加换行 if tag_name in block_tags: result.append('\n') else: result.append(html[i]) i += 1 return ''.join(result) html = add_newline_after_block_open_tags(html, SINGLE_LINE_TAGS) print(html)
输出效果
运行后可得到预期格式,带属性的ol开放标签后会自动插入换行,其他原有逻辑保持不变。
内容的提问来源于stack exchange,提问作者nguyenphanhoaiduc
相关产品推荐
相关产品推荐

