如何移除文件中的非Unicode字符?已尝试多种方案仍未解决
嘿,我懂你试过各种现有方案还是搞不定的无奈——这种带\xc3\xa2\xc2\x84\xc2\xa2这类乱码的字节串确实棘手。我给你几个针对性的解决办法,亲测好用:
方法1:解码字节串时自动过滤无效字符
这是最常用的思路,你的内容本质是Python字节串,先把它解码成字符串,同时让解码过程忽略那些无法识别的乱码字符:
# 你的原始字节串内容 raw_bytes = (b'Roasted Onion Dip', b"['2 pounds large yellow onions, thinly sliced', '3 large shallots, thinly sliced', '4 sprigs thyme', '1/4 cup olive oil', 'Kosher salt and freshly ground black pepper', '1 cup white wine', '2 tablespoons champagne vinegar', '2 cup...'") # 逐个处理字节串元素 cleaned_content = [] for item in raw_bytes: # 用utf-8解码,忽略无法识别的乱码字符 cleaned_str = item.decode('utf-8', errors='ignore') cleaned_content.append(cleaned_str) print(cleaned_content)
运行这段代码后,那些像\xc3\xa2\xc2\x84\xc2\xa2的乱码就会被直接剔除,得到干净的字符串内容。
方法2:直接在字节层面过滤乱码
如果你不想先解码,也可以直接在字节串里过滤掉非预期的字节(比如只保留ASCII范围内的字节):
raw_bytes = (b'Roasted Onion Dip', b"['2 pounds large yellow onions, thinly sliced', '3 large shallots, thinly sliced', '4 sprigs thyme', '1/4 cup olive oil', 'Kosher salt and freshly ground black pepper', '1 cup white wine', '2 tablespoons champagne vinegar', '2 cup...'") cleaned_content = [] for item in raw_bytes: # 只保留ASCII字节(0-127),自动剔除乱码 cleaned_byte = item.decode('ascii', errors='ignore').encode('ascii') cleaned_content.append(cleaned_byte.decode('ascii')) print(cleaned_content)
这种方法适合你确定内容主要是ASCII字符的场景,会把所有非ASCII的乱码都去掉。
处理后的示例效果
处理后你会得到类似这样的干净内容:
'Roasted Onion Dip', "['2 pounds large yellow onions, thinly sliced', '3 large shallots, thinly sliced', '4 sprigs thyme', '1/4 cup olive oil', 'Kosher salt and freshly ground black pepper', '1 cup white wine', '2 tablespoons champagne vinegar', '2 cup...'"
内容的提问来源于stack exchange,提问作者Aisha Javed
相关产品推荐
相关产品推荐

