Python移除HTML标签遇FileNotFoundError,路径拼接错误如何解决?
解决HTML文件处理时的FileNotFoundError问题
我编写了一段用于移除HTML文件中style脚本和HTML标签,并将处理后内容保存至新文件夹的Python代码,但运行时出现FileNotFoundError,错误信息显示路径拼接错误(示例:'/content/CS13309_Archivos_HTML/Files307.html')。尝试改用BeautifulSoup仍出现相同错误,请问如何修改代码以正确读取文件?
原代码
import os, re #Paths src_path = "/content/CS13309_Archivos_HTML/Files" dest_path = "/content/CS13309_Archivos_HTML/Nuevos_Archivos" files = os.listdir(src_path) for file in files: # open the file using its name file_name = file.split(".") open_file = open(src_path + file_name[0] + ".html", "r",encoding='latin1') # parse the file for style scripts and html tags style_scripts = re.compile('<style.*?>.*?</style>') html_tags = re.compile('<.*?>') stripped_data = re.sub(style_scripts, '', open_file.read()) stripped_data = re.sub(html_tags, '', stripped_data) # create a new file without style scripts and html tags new_file = open(dest_path + file_name[0] + 'NuevosArchivos.txt', 'w+') # write the content from the old file to the new file new_file.write(stripped_data) # close the opened file open_file.close() new_file.close()
报错信息
FileNotFoundError: [Errno 2] No such file or directory: '/content/CS13309_Archivos_HTML/Files307.html'
错误原因
从报错路径能直接看出问题:路径拼接时缺少目录分隔符。src_path是/content/CS13309_Archivos_HTML/Files,直接拼接文件名会变成Files307.html,正确路径应该是Files/307.html。另外手动用split(".")拆分文件名也有隐患,比如文件名包含多个.时会出错。
修改后的代码
使用Python标准路径处理工具os.path.join()避免拼接错误,同时用with语句自动管理文件生命周期:
import os import re # 定义源文件夹和目标文件夹路径 src_path = "/content/CS13309_Archivos_HTML/Files" dest_path = "/content/CS13309_Archivos_HTML/Nuevos_Archivos" # 确保目标文件夹存在,不存在则自动创建 os.makedirs(dest_path, exist_ok=True) # 遍历源文件夹下的所有文件 for file in os.listdir(src_path): # 只处理HTML文件,跳过其他类型文件 if not file.endswith(".html"): continue # 拼接源文件的完整路径 src_file_full_path = os.path.join(src_path, file) # 读取并处理文件内容 with open(src_file_full_path, 'r', encoding='latin1') as src_file: raw_content = src_file.read() # 移除所有<style>块(包含内部换行) cleaned_content = re.sub(r'<style.*?>.*?</style>', '', raw_content, flags=re.DOTALL) # 移除所有HTML标签 cleaned_content = re.sub(r'<.*?>', '', cleaned_content) # 生成目标文件名:去掉原文件的.html后缀,添加自定义后缀 file_base_name = os.path.splitext(file)[0] dest_file_name = f"{file_base_name}NuevosArchivos.txt" # 拼接目标文件的完整路径 dest_file_full_path = os.path.join(dest_path, dest_file_name) # 将处理后的内容写入目标文件 with open(dest_file_full_path, 'w', encoding='latin1') as dest_file: dest_file.write(cleaned_content)
关键优化点
- 路径拼接:用
os.path.join()替代手动字符串拼接,自动适配不同操作系统的路径分隔符,彻底解决路径错误问题 - 文件夹检查:新增
os.makedirs(dest_path, exist_ok=True),确保目标文件夹存在,避免写入时出现文件夹不存在的错误 - 文件名处理:用
os.path.splitext()安全拆分文件名和后缀,避免手动split(".")在多后缀文件名上的失效 - 文件管理:使用
with语句自动管理文件上下文,无需手动调用close(),避免资源泄漏 - 正则优化:添加
re.DOTALL标志,确保
相关产品推荐
相关产品推荐

