如何用Python移除Markdown中无嵌套子列表的无序列表
问题描述
需要移除Markdown中没有无序列表子项的顶级无序列表条目及后续空行。
现有Markdown内容:
# Known * Languages with Latin alphabet: * English * French * Portuguese * Languages with Greek alphabet: * Greek * Languages with Armenian alphabet: * Languages with Ethiopic script: * Languages with Tamil script: * Languages with Hiragana characters: * Japanese * Languages with Klingon script: # Wanted **These wanted languages:** * Languages with Sumero-Akkadian script: * Languages with Hangul characters: * Korean * Languages with Linear A script:
期望处理后结果:
# Known * Languages with Latin alphabet: * English * French * Portuguese * Languages with Greek alphabet: * Greek * Languages with Hiragana characters: * Japanese # Wanted **These wanted languages:** * Languages with Hangul characters: * Korean
原尝试的Python代码未达到预期效果:
import re file_name = "Language support" def remove_empty_lists(): # It will read the file with open("{}.md".format(file_name), "r") as f: # It will search all the lines of the whole text lines = f.readlines() for line in lines: # It will find all the lines that contain "* Language with ..." if re.search(r"\* Languages with", line): # then if these contained lines have the symbol ":" if re.search(r":$", line): # It will find the empy new line if re.search(r"^$\n", line): # Then finally it will remove the list and the new line lines.remove(line) # It will rewrite the same file with open("{}.md".format(file_name), "w") as f: for line in lines: f.write(line)
原代码存在的问题
- 逻辑错误:判断空行的条件
re.search(r"^$\n", line)是检查当前行是否为空行,但当前行是顶级列表项(比如* Languages with Armenian alphabet:),不可能是空行,所以这个判断永远不成立,根本不会执行删除操作。 - 遍历列表时直接修改原列表:
lines.remove(line)会导致遍历过程中跳过部分元素,因为列表长度变化了。 - 没有处理列表项后续的空行:即使删除了空列表项,也没移除它后面的空行。
正确实现方案
思路
- 逐行遍历,跟踪当前是否处于需要检查的顶级无序列表项状态。
- 当遇到顶级列表项(以
* Languages with开头且以:结尾),记录该位置,然后检查下一行是否是缩进的子列表项(以*开头)。 - 如果下一行不是子列表项,则标记当前列表项和后续的空行需要删除。
- 最后将不需要删除的内容写入文件。
代码实现
import re file_name = "Language support.md" def remove_empty_lists(): with open(file_name, "r") as f: lines = [line.rstrip('\n') for line in f] # 统一处理换行符,避免混乱 keep_lines = [] i = 0 while i < len(lines): line = lines[i] # 匹配顶级列表项:以* Languages with开头,以:结尾 if re.match(r"\* Languages with.*:$", line): # 检查下一行是否是子列表项(缩进2空格+*) has_child = False # 跳过当前行后的空行,看是否有子项 j = i + 1 while j < len(lines): next_line = lines[j].strip() if next_line == "": j += 1 continue # 子列表项以 * 开头(注意前面的两个空格) if re.match(r" \*.*", lines[j]): has_child = True break if has_child: keep_lines.append(line) i += 1 else: # 跳过当前空列表项,以及后续的所有空行 i += 1 while i < len(lines) and lines[i].strip() == "": i += 1 continue else: keep_lines.append(line) i += 1 # 将处理后的内容写入文件,每行加换行符 with open(file_name, "w") as f: f.write('\n'.join(keep_lines) + '\n') if __name__ == "__main__": remove_empty_lists()
代码解释
- 先把所有行的换行符统一处理,用
rstrip('\n')去掉换行,最后写入时再统一添加,避免不同系统换行符问题。 - 使用
while循环遍历,方便控制索引,处理需要跳过的行。 - 遇到目标顶级列表项时,向后检查是否存在缩进的子列表项:跳过中间的空行,看第一个非空行是否是子项格式。
- 如果没有子项,就跳过当前列表项以及后续的所有空行;如果有子项,就保留当前行,继续遍历。
- 最后把需要保留的行用
'\n'.join拼接,写入文件。
内容的提问来源于stack exchange,提问作者Oo'-
相关产品推荐
相关产品推荐

