使用pandas read_csv的delim_whitespace参数时如何设置最大分隔次数
delim_whitespace=True等价于sep='\s+',会将所有连续空白识别为分隔符,本身不支持设置最大分隔次数,你可以用以下方法实现需求:
读取后拆分(最简洁兼容)
先把每行内容完整读为单列,再调用字符串拆分方法指定最大拆分次数即可:
import pandas as pd from io import StringIO # 把这里的data替换为你的results变量即可 data = """ 0 a b this is my first file 1 c d this is my second file 2 e f this is my third file 3 g h this is my fourth file 4 i j this is my fifth file """ # 先读取所有内容为单列 df = pd.read_csv(StringIO(data), header=None, names=['raw']) # 按空格最多拆分3次,expand=True将拆分结果展开为独立列 df = df['raw'].str.split(n=3, expand=True) # 按需重命名列名 df.columns = ['序号', '第二列', '第三列', '文本内容']
运行后得到的df就符合你要的结构,第四列会完整保留后面的所有文本内容。
内容的提问来源于stack exchange,提问作者PicxyB
相关产品推荐
相关产品推荐

