You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取含0xc0异常字节文件报错及Pandas读取TSV解码失败求助

Fixing UnicodeDecodeError and Mismatched Fields When Reading TSV Files with 0xc0 Byte

Let’s break down your problem and fix it step by step:

First, the root causes:

  • The byte 0xc0 is not a valid UTF-8 start byte (UTF-8 requires multi-byte sequences starting with 0xc0 to be paired with a continuation byte, which is missing here), hence the UnicodeDecodeError.
  • ASCII only supports bytes 0-127, so 0xc0 (decimal 192) is way outside its range—this is why ASCII decoding fails too.
  • The "expected 11 fields, saw 12" error is a side effect of the encoding failure: when pandas can’t decode bytes correctly, it may misinterpret characters as tab separators, leading to incorrect field counts.

1. Use a "safe" single-byte encoding to avoid decoding errors

The most reliable workaround for files with invalid UTF-8 bytes is to use latin-1 (also called iso-8859-1). This encoding maps every single byte to a unique Unicode character without throwing errors, letting pandas read the file correctly first.

Update your read command to:

import pandas as pd
df = pd.read_table(fn, na_filter=False, encoding='latin-1', on_bad_lines='skip')

Note: error_bad_lines=False is deprecated in newer pandas versions—use on_bad_lines='skip' instead, or on_bad_lines='warn' if you want to log which lines are being skipped.

2. Resolve the mismatched fields issue

In most cases, fixing the encoding will automatically fix the field count problem, since pandas can now correctly identify tab separators. If you still see mismatches:

  • Use on_bad_lines='warn' to get details about problematic lines, so you can manually inspect what’s causing extra fields.
  • Use a custom handler function to adjust bad lines on the fly. For example, if you want to keep only the first 11 fields:
    def fix_bad_line(line):
        # Split the line into fields and truncate to 11
        return line.split('\t')[:11]
    
    df = pd.read_table(fn, na_filter=False, encoding='latin-1', on_bad_lines=fix_bad_line)
    

3. Clean up the 0xc0 byte after reading

Once loaded, the 0xc0 byte will appear as the character À (per latin-1 mapping). You can replace it with whatever makes sense for your data:

# Remove the character entirely
df = df.replace('À', '', regex=True)

# Or replace it with a known correct character (e.g., 'A')
df = df.replace('À', 'A', regex=True)

If you know 0xc0 is an encoding mistake (e.g., it should be a different byte), fix the file at the byte level first:

# Read raw bytes, replace 0xc0 with the correct byte value
with open(fn, 'rb') as f:
    raw_content = f.read().replace(b'\xc0', b'YOUR_CORRECT_BYTE')

# Save fixed content to a new file
with open('fixed_' + fn, 'wb') as f:
    f.write(raw_content)

# Now read the fixed file with UTF-8
df = pd.read_table('fixed_' + fn, na_filter=False)

内容的提问来源于stack exchange,提问作者Nikhil VJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:00:53