Pandas读取含冒号的TXT文件报错,如何仅按首次分隔符拆分?
问题描述
有一个TXT格式的数据集,内容示例如下:
beer_name: Legbiter beer_id: 19827 brewery_name: Strangford Lough Brewing Company Ltd brewery_id: 10093 style: English Pale Ale abv: 4.8 date: 1357729200 user_name: AgentMunky user_id: agentmunky.409755 appearance: 4.0 aroma: 3.75 palate: 3.5 taste: 3.5 overall: 3.75 rating: 3.64 text: Poured from a 12 ounce bottle into a pilsner glass.A: A finger of creamy head with clear-dark amber body.S: Rich brown sugar. Malty...T: Slight sugars, dry malt, vague hops. Big malty-brown with sugar.M: Dry and slightly astringent before a boring endtaste.O: Solid beer. Drinkable and interesting. Still vaguely bland. review: True
使用以下代码尝试转换为DataFrame时出错:
rb_file_data = pd.read_csv(os.path.join(MATCHED_BEER_DIR, 'ratings_with_text_rb.txt'), sep=":", header=None, names=["Key", "Value"])
错误原因是text字段内容包含冒号,导致解析失败,报错信息:
ParserError: Error tokenizing data. C error: Expected 2 fields in line 34, saw 7
希望保留text字段,询问是否可以仅按每行首次出现的冒号拆分,或其他解决方法。
解决方案
方法1:用正则分隔符匹配首个冒号
直接修改read_csv的参数,用正则表达式指定只匹配每行第一个冒号,同时兼容冒号后的空格:
import pandas as pd import os rb_file_data = pd.read_csv( os.path.join(MATCHED_BEER_DIR, 'ratings_with_text_rb.txt'), sep=r':\s*', # 匹配首个冒号及后续可选空格 header=None, names=["Key", "Value"], engine='python', # 必须用python引擎,C引擎不支持正则分隔符 usecols=[0, 1] # 强制只取前两列,避免多余拆分 )
方法2:手动逐行拆分(更灵活)
如果正则方法有兼容性问题,直接手动读取每行并控制拆分逻辑:
import pandas as pd import os data = [] with open(os.path.join(MATCHED_BEER_DIR, 'ratings_with_text_rb.txt'), 'r', encoding='utf-8') as f: for line in f: line = line.strip() if not line: continue # 跳过空行 # 定位首个冒号的位置 colon_pos = line.find(':') if colon_pos == -1: continue # 跳过无冒号的无效行 key = line[:colon_pos].strip() value = line[colon_pos+1:].strip() data.append({'Key': key, 'Value': value}) rb_file_data = pd.DataFrame(data)
方法3:整理为结构化记录(如果数据按评分组)
要是你的数据是多条啤酒评价按空行分隔的(示例只展示了一条),可以把每条评价整理成DataFrame的一行,字段作为列:
import pandas as pd import os from itertools import groupby records = [] current_record = {} with open(os.path.join(MATCHED_BEER_DIR, 'ratings_with_text_rb.txt'), 'r', encoding='utf-8') as f: # 按空行分组识别单条评价 for is_empty, lines in groupby(f, lambda x: x.strip() == ''): if not is_empty: for line in lines: line = line.strip() colon_pos = line.find(':') if colon_pos == -1: continue key = line[:colon_pos].strip() value = line[colon_pos+1:].strip() current_record[key] = value records.append(current_record) current_record = {} rb_file_data = pd.DataFrame(records)
内容的提问来源于stack exchange,提问作者Eboyer
相关产品推荐
相关产品推荐

