如何在无导入的Python代码中检测HTML标签缺失的闭合尖括号>
问题分析与修复
你的程序无法检测HTML中缺失的>,核心问题是遇到<后未处理找不到对应>的场景:
- 处理开始标签时,若
content.find(">", i+1)返回-1(无匹配的>),程序直接跳过该标签,未标记无效; - 处理结束标签时,同样未检查是否找到
>,即便缺失>,程序仍会尝试提取标签名,不会触发错误。
修改后的代码
filePath = input("Enter the file path of your HTML file: ") with open(filePath, "r") as file: content = file.read() tagList = [] mismatch = False for i in range(len(content)): if content[i] == "<": # 处理开始标签 if content[i : i + 2] != "</": endOfTag = content.find(">", i + 1) # 检测是否缺失> if endOfTag == -1: print("Invalid HTML, missing '>' in opening tag") mismatch = True break tag = content[i : endOfTag + 1] tagType = tag[1 : -1].split()[0] tagList.append(tagType) # 处理闭合标签 else: endOfClosingTag = content.find(">", i + 1) # 检测是否缺失> if endOfClosingTag == -1: print("Invalid HTML, missing '>' in closing tag") mismatch = True break closingTagType = content[i + 2 : endOfClosingTag] if tagList and tagList[-1] == closingTagType: tagList.pop() else: print("Invalid HTML, mismatched opening and closing tags") mismatch = True break if not mismatch: if len(tagList) == 0: print("Valid HTML") else: print("Invalid HTML, mismatched opening and closing tags")
关键修改点
- 开始标签逻辑:新增
endOfTag == -1判断,找不到>时直接输出错误并终止程序; - 闭合标签逻辑:先获取闭合标签的
>位置,检查是否缺失,若缺失则标记错误; - 最终输出逻辑:调整为仅在未触发
mismatch时,根据tagList判断标签匹配状态。
内容的提问来源于stack exchange,提问作者MooMerr
相关产品推荐
相关产品推荐

