如何将Python中的指定OHLC字符串数据转换为DataFrame?
将OHLC字符串快速转换为Pandas DataFrame的方法
方法一:利用pandas.read_csv(推荐,简洁高效)
先修正你的变量定义(原代码存在语法问题):
ohlc = "1664788799|38444.9|38569.2|38327.85|38412.35|0|0|,1664789099|38408.35|38587.3|38394.05|38586.15|0|0|,1664789399|38589.55|38641.6|38420.05|38422.35|0|0|"
通过io.StringIO将字符串模拟为文件对象,配合read_csv直接解析:
import pandas as pd from io import StringIO # 清理数据:移除条目末尾多余的|,将逗号分隔替换为换行 cleaned_data = ohlc.replace('|,', '\n').rstrip('|') # 读取为DataFrame并指定列名、数据类型 df = pd.read_csv( StringIO(cleaned_data), sep='|', names=['timestamp', 'open', 'high', 'low', 'close', 'volume', 'OI'], dtype={ 'timestamp': int, 'open': float, 'high': float, 'low': float, 'close': float, 'volume': int, 'OI': int } ) # 可选:将时间戳转换为可读性更强的datetime格式 df['timestamp'] = pd.to_datetime(df['timestamp'], unit='s')
方法二:手动拆分字符串(适合小量数据)
如果不想引入StringIO,可以直接拆分字符串后构造DataFrame:
import pandas as pd # 按逗号拆分每个K线条目,再按|拆分字段(去掉末尾多余的空值) rows = [row.split('|')[:-1] for row in ohlc.split(',')] # 构造DataFrame并转换数据类型 df = pd.DataFrame(rows, columns=['timestamp', 'open', 'high', 'low', 'close', 'volume', 'OI']).astype({ 'timestamp': int, 'open': float, 'high': float, 'low': float, 'close': float, 'volume': int, 'OI': int }) # 可选转换时间戳格式 df['timestamp'] = pd.to_datetime(df['timestamp'], unit='s')
两种方法最终都会得到结构正确的DataFrame,推荐第一种方法,尤其是处理大量数据时,read_csv的性能和容错性更好。
内容的提问来源于stack exchange,提问作者user3698161
相关产品推荐
相关产品推荐

