C++逐词读取文本文件时符号与单词粘连拆分异常问题求解
问题描述
实现逐行读取文本文件、逐词解析内容的逻辑,预期效果为找到指定目标词后,跳过当前行剩余所有单词直接读取下一行;实际运行时出现单词与相邻符号组合输出的异常,拆分结果不符合预期。
现有实现代码
void SymbolScanning::symScanning() { std::string s; std::ifstream myfile; myfile.open("SymbolRead.txt"); if(myfile.is_open()) { while(std::getline(myfile, s)) { std::istringstream iss(s); std::string word; while(iss >> word) { std::cout << word << std::endl; // if desired word found, skip remainnig word and jump // over next line. } } } else cout<<"File is not open"; }
实际运行表现
输入行内容:
line (60 * SCALE, 70 * SCALE, 106 * SCALE, 70 * SCALE);
实际拆分输出:
line , (60 , * , SCALE, , 70 , * , SCALE, , 106 , * , SCALE, , 70 , * , SCALE);
预期拆分规则
标点符号需要和相邻的单词、数字拆分,作为独立单元输出:
(60拆分为(、60SCALE,拆分为SCALE、,SCALE);拆分为SCALE、)、;
待读取文件样例
/* version: v1p1 */ library("aacbnsfp_90spwd_9t45g_rxy") { SCALE = 1.0 / 10.0; symbol ("PQXN67_0P5_9IJT16R") { circle (75 * SCALE, -40 * SCALE, 5 * SCALE); line (20 * SCALE, -20 * SCALE, 80 * SCALE, -20 * SCALE); line (20 * SCALE, 0 * SCALE, 80 * SCALE, 0 * SCALE); line (240 * SCALE, -20 * SCALE, 200 * SCALE, -20 * SCALE); ...... ...... } /* end of symbol "PQXN67_0P5_9IJT16R" */
问题原因
istringstream的>>运算符默认仅将*空白字符(空格、制表符、换行)*作为分隔符,括号、逗号、分号等标点不属于空白字符,会和相邻的字母、数字拼接为同一个字符串被读取,因此出现符号和单词/数字粘连的问题。同时现有代码未实现「找到目标词后跳过当前行剩余内容」的逻辑。
修复方案
逐字符遍历每行内容,将字母、数字、下划线、小数点等普通字符拼接为词元,遇到标点时先输出已拼接的词元,再将标点作为独立词元输出;检测到目标词时直接终止当前行的遍历,自动进入下一行读取流程。
修复后完整代码如下:
#include <fstream> #include <string> #include <iostream> #include <cctype> // 自定义需要匹配的目标词 const std::string TARGET_WORD = "line"; void SymbolScanning::symScanning() { std::ifstream myfile("SymbolRead.txt"); if (!myfile.is_open()) { std::cout << "File is not open"; return; } std::string line; while (std::getline(myfile, line)) { std::string current_token; bool skip_rest_line = false; for (char c : line) { if (skip_rest_line) break; // 遇到空白,输出已拼接的词元 if (std::isspace(c)) { if (!current_token.empty()) { std::cout << current_token << std::endl; if (current_token == TARGET_WORD) { skip_rest_line = true; } current_token.clear(); } continue; } // 遇到需要拆分的标点,先输出已存词元,再单独输出标点 if (c == '(' || c == ')' || c == ',' || c == ';' || c == '*' || c == '/' || c == '=' || c == '{' || c == '}' || c == '"' || c == '+' || c == '-') { if (!current_token.empty()) { std::cout << current_token << std::endl; if (current_token == TARGET_WORD) { skip_rest_line = true; current_token.clear(); continue; } current_token.clear(); } std::cout << std::string(1, c) << std::endl; continue; } // 普通字符加入当前词元 current_token += c; } // 处理行末尾剩余的词元 if (!skip_rest_line && !current_token.empty()) { std::cout << current_token << std::endl; } } myfile.close(); }
以上代码运行后,会按预期将粘连的符号和单词/数字拆分,匹配到目标词后会立刻跳过当前行剩余内容,直接读取下一行。
内容的提问来源于stack exchange,提问作者tushar
相关产品推荐
相关产品推荐

