You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Open()读取XML(.csproj)文件时出现无效字符致格式错误

解决读取.csproj文件时的BOM字符问题

我用Python读取.csproj格式的XML文件,执行file_content=f.read().replace('\n','')时,标签前出现了这类无效字符,导致XML格式不合法报错。相关代码如下:

def parse_csproj(repository_root, repository_name, source):

    # If the file is not XML, then a ParseError
    # is raised. Return None to indicate that the
    # file is not a csproj file.
    try:
        p = repository_root / repository_name / source
        with p.open() as f:
            file_content=f.read().replace('\n','') #I put a breakpoint here and I see invalid characters before <xml tag
            root=ET.fromstring(str(file_content))
            #root = ET.parse(f).getroot()
    except Exception as ex:
        return None

    # Extract the list of NuGet packages that are
    # used by the csproj file.
    refs = []
    for element in root.findall('.//{*}PackageReference'): #You need to add {*} otherwise it won't work. Or you 
        #will need to access based on namespace. Because notice in .csproj, the root element contains a namespace.
        ref = PackageReference(
            ecosystem = 'NuGet',
            repository_root = repository_root,
            repository_name = repository_name,
            source = source,
            package_name = element.attrib['Include'],
            version = element.attrib['Version'])
        refs.append(ref)

    return refs 

注:是UTF-8 BOM(字节顺序标记),Windows环境下生成的.csproj文件常带有这个标记,直接读取会残留该字符导致XML解析失败。

解决方案

  • 方法一:指定编码读取文件
    修改文件打开方式,使用utf-8-sig编码,它会自动识别并去除UTF-8文件开头的BOM标记:

    with p.open('r', encoding='utf-8-sig') as f:
        file_content = f.read().replace('\n','')
    
  • 方法二:直接使用ET.parse读取(更推荐)
    无需手动读取文件内容再解析,ET.parse可以直接处理带BOM的文件,同时简化代码:

    def parse_csproj(repository_root, repository_name, source):
        try:
            p = repository_root / repository_name / source
            root = ET.parse(p).getroot()  # 直接传入文件路径解析
        except Exception as ex:
            return None
    
        refs = []
        for element in root.findall('.//{*}PackageReference'):
            ref = PackageReference(
                ecosystem = 'NuGet',
                repository_root = repository_root,
                repository_name = repository_name,
                source = source,
                package_name = element.attrib['Include'],
                version = element.attrib['Version'])
            refs.append(ref)
        return refs 
    

内容的提问来源于stack exchange,提问作者nikhil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 12:35:08