You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python移除HTML标签遇FileNotFoundError,路径拼接错误如何解决?

解决HTML文件处理时的FileNotFoundError问题

我编写了一段用于移除HTML文件中style脚本和HTML标签,并将处理后内容保存至新文件夹的Python代码,但运行时出现FileNotFoundError,错误信息显示路径拼接错误(示例:'/content/CS13309_Archivos_HTML/Files307.html')。尝试改用BeautifulSoup仍出现相同错误,请问如何修改代码以正确读取文件?

原代码

import os, re 

#Paths
src_path = "/content/CS13309_Archivos_HTML/Files" 
dest_path = "/content/CS13309_Archivos_HTML/Nuevos_Archivos" 

files = os.listdir(src_path)

for file in files: 
    # open the file using its name 
    file_name = file.split(".") 
    open_file = open(src_path + file_name[0] + ".html", "r",encoding='latin1') 

    # parse the file for style scripts and html tags
    style_scripts = re.compile('<style.*?>.*?</style>')
    html_tags = re.compile('<.*?>')
    stripped_data = re.sub(style_scripts, '', open_file.read())
    stripped_data = re.sub(html_tags, '', stripped_data)

    # create a new file without style scripts and html tags
    new_file = open(dest_path + file_name[0] + 'NuevosArchivos.txt', 'w+')

    # write the content from the old file to the new file
    new_file.write(stripped_data)

    # close the opened file 
    open_file.close() 
    new_file.close() 

报错信息

FileNotFoundError: [Errno 2] No such file or directory: '/content/CS13309_Archivos_HTML/Files307.html'

错误原因

从报错路径能直接看出问题:路径拼接时缺少目录分隔符。src_path是/content/CS13309_Archivos_HTML/Files,直接拼接文件名会变成Files307.html,正确路径应该是Files/307.html。另外手动用split(".")拆分文件名也有隐患,比如文件名包含多个.时会出错。

修改后的代码

使用Python标准路径处理工具os.path.join()避免拼接错误,同时用with语句自动管理文件生命周期:

import os
import re

# 定义源文件夹和目标文件夹路径
src_path = "/content/CS13309_Archivos_HTML/Files" 
dest_path = "/content/CS13309_Archivos_HTML/Nuevos_Archivos" 

# 确保目标文件夹存在,不存在则自动创建
os.makedirs(dest_path, exist_ok=True)

# 遍历源文件夹下的所有文件
for file in os.listdir(src_path):
    # 只处理HTML文件,跳过其他类型文件
    if not file.endswith(".html"):
        continue
    
    # 拼接源文件的完整路径
    src_file_full_path = os.path.join(src_path, file)
    
    # 读取并处理文件内容
    with open(src_file_full_path, 'r', encoding='latin1') as src_file:
        raw_content = src_file.read()
        # 移除所有<style>块(包含内部换行)
        cleaned_content = re.sub(r'<style.*?>.*?</style>', '', raw_content, flags=re.DOTALL)
        # 移除所有HTML标签
        cleaned_content = re.sub(r'<.*?>', '', cleaned_content)
    
    # 生成目标文件名:去掉原文件的.html后缀,添加自定义后缀
    file_base_name = os.path.splitext(file)[0]
    dest_file_name = f"{file_base_name}NuevosArchivos.txt"
    # 拼接目标文件的完整路径
    dest_file_full_path = os.path.join(dest_path, dest_file_name)
    
    # 将处理后的内容写入目标文件
    with open(dest_file_full_path, 'w', encoding='latin1') as dest_file:
        dest_file.write(cleaned_content)

关键优化点

  • 路径拼接:用os.path.join()替代手动字符串拼接,自动适配不同操作系统的路径分隔符,彻底解决路径错误问题
  • 文件夹检查:新增os.makedirs(dest_path, exist_ok=True),确保目标文件夹存在,避免写入时出现文件夹不存在的错误
  • 文件名处理:用os.path.splitext()安全拆分文件名和后缀,避免手动split(".")在多后缀文件名上的失效
  • 文件管理:使用with语句自动管理文件上下文,无需手动调用close(),避免资源泄漏
  • 正则优化:添加re.DOTALL标志,确保
相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 15:25:23