You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas读取含冒号的TXT文件报错,如何仅按首次分隔符拆分?

问题描述

有一个TXT格式的数据集,内容示例如下:

beer_name: Legbiter
beer_id: 19827
brewery_name: Strangford Lough Brewing Company Ltd
brewery_id: 10093
style: English Pale Ale
abv: 4.8
date: 1357729200
user_name: AgentMunky
user_id: agentmunky.409755
appearance: 4.0
aroma: 3.75
palate: 3.5
taste: 3.5
overall: 3.75
rating: 3.64
text: Poured from a 12 ounce bottle into a pilsner glass.A: A finger of creamy head with clear-dark amber body.S: Rich brown sugar. Malty...T: Slight sugars, dry malt, vague hops. Big malty-brown with sugar.M: Dry and slightly astringent before a boring endtaste.O: Solid beer. Drinkable and interesting. Still vaguely bland.
review: True

使用以下代码尝试转换为DataFrame时出错:

rb_file_data = pd.read_csv(os.path.join(MATCHED_BEER_DIR, 'ratings_with_text_rb.txt'), sep=":", header=None, names=["Key", "Value"])

错误原因是text字段内容包含冒号,导致解析失败,报错信息:

ParserError: Error tokenizing data. C error: Expected 2 fields in line 34, saw 7

希望保留text字段,询问是否可以仅按每行首次出现的冒号拆分,或其他解决方法。

解决方案

方法1:用正则分隔符匹配首个冒号

直接修改read_csv的参数,用正则表达式指定只匹配每行第一个冒号,同时兼容冒号后的空格:

import pandas as pd
import os

rb_file_data = pd.read_csv(
    os.path.join(MATCHED_BEER_DIR, 'ratings_with_text_rb.txt'),
    sep=r':\s*',  # 匹配首个冒号及后续可选空格
    header=None,
    names=["Key", "Value"],
    engine='python',  # 必须用python引擎,C引擎不支持正则分隔符
    usecols=[0, 1]  # 强制只取前两列,避免多余拆分
)

方法2:手动逐行拆分(更灵活)

如果正则方法有兼容性问题,直接手动读取每行并控制拆分逻辑:

import pandas as pd
import os

data = []
with open(os.path.join(MATCHED_BEER_DIR, 'ratings_with_text_rb.txt'), 'r', encoding='utf-8') as f:
    for line in f:
        line = line.strip()
        if not line:
            continue  # 跳过空行
        # 定位首个冒号的位置
        colon_pos = line.find(':')
        if colon_pos == -1:
            continue  # 跳过无冒号的无效行
        key = line[:colon_pos].strip()
        value = line[colon_pos+1:].strip()
        data.append({'Key': key, 'Value': value})

rb_file_data = pd.DataFrame(data)

方法3:整理为结构化记录(如果数据按评分组)

要是你的数据是多条啤酒评价按空行分隔的(示例只展示了一条),可以把每条评价整理成DataFrame的一行,字段作为列:

import pandas as pd
import os
from itertools import groupby

records = []
current_record = {}

with open(os.path.join(MATCHED_BEER_DIR, 'ratings_with_text_rb.txt'), 'r', encoding='utf-8') as f:
    # 按空行分组识别单条评价
    for is_empty, lines in groupby(f, lambda x: x.strip() == ''):
        if not is_empty:
            for line in lines:
                line = line.strip()
                colon_pos = line.find(':')
                if colon_pos == -1:
                    continue
                key = line[:colon_pos].strip()
                value = line[colon_pos+1:].strip()
                current_record[key] = value
            records.append(current_record)
            current_record = {}

rb_file_data = pd.DataFrame(records)

内容的提问来源于stack exchange,提问作者Eboyer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 13:43:18