C++多线程解析大格式文本文件的性能优化可行性咨询
大文件解析性能优化咨询
我需要解析大小≥100MB的特定格式大文本文件,提取预设token供后续API处理。文件格式示例如下:
=======DB VERSION 2.0======= $BLOCK$ block1_name $TREE$ tree1_name /** Info dump related to a data structure **/ $END_TREE$ $SUB_BLOCK$ sub1_name /** Info dump related to a data structure **/ $SUB_BLOCK$ $BLOCK_END$ $BLOCK$ block3_name $TREE$ tree1_name /** Info dump related to a data structure **/ $END_TREE$ $TREE$ tree2_name /** Info dump related to a data structure **/ $END_TREE$ $SUB_BLOCK$ sub1_name /** Info dump related to a data structure **/ $SUB_BLOCK$ $BLOCK_END$ $BLOCK$ block2_name $TREE$ tree1_name /** Info dump related to a data structure **/ $END_TREE$ $TREE$ tree2_name /** Info dump related to a data structure **/ $END_TREE$ $SUB_BLOCK$ sub1_name /** Info dump related to a data structure **/ $SUB_BLOCK$ $BLOCK_END$
文件内包含多个信息块,且API要求单个块必须在单线程内完整处理,因此无法拆分文件。当前使用getline函数解析速度极慢,尝试过mio内存映射文件但无法将块保留为完整chunk,想咨询是否可通过多线程分别处理每个块来提升性能。
现有类似代码如下:
void parseFile(){ ifstream file("file.db"); string line; while(getline(file, line)) { Utility::cleanString(line); // Removes trailing spaces if (line == BLOCK_START){ openContext(line); } else if (line == BLOCK_END){ closeContext(line); } else if (line == TREE_START){ processTree(line); } else if (line == SUB_BLOCK){ processSubBlock(line); } } }
内容的提问来源于stack exchange,提问作者Tejas Sharma
相关产品推荐
相关产品推荐

