关于使用C++提取文本句子及按关键词拆分TXT文件的技术求助
C++文本处理问题解答
问题1:如何用C++从文本中提取句子
首先,提取句子的核心是识别句子的分隔符——一般来说句子是以.、!、?这类标点结尾的(如果是复杂场景,可能还要考虑缩写里的句号,比如Mr.,但简单需求可以先忽略这类特殊情况)。
给你一个基础实现思路:
- 读取整个文本内容(或逐字符读取)
- 把字符累积到临时字符串,直到遇到句子分隔符
- 处理掉句子前后的多余空格,将其作为完整句子保存
- 跳过连续的分隔符或空白字符,继续处理下一个句子
示例代码:
#include <iostream> #include <fstream> #include <string> #include <vector> #include <cctype> using namespace std; vector<string> extractSentences(const string& text) { vector<string> sentences; string currentSentence; const char sentenceDelimiters[] = ".!?"; for (char c : text) { // 将当前字符加入临时句子 currentSentence += c; // 检查是否是句子分隔符 bool isDelimiter = false; for (char delimiter : sentenceDelimiters) { if (c == delimiter) { isDelimiter = true; break; } } if (isDelimiter) { // 去除句子前后的空白字符 size_t start = currentSentence.find_first_not_of(" \t\n\r"); size_t end = currentSentence.find_last_not_of(" \t\n\r"); if (start != string::npos && end != string::npos) { sentences.push_back(currentSentence.substr(start, end - start + 1)); } currentSentence.clear(); } } // 处理最后一个没有结尾分隔符的句子(如果存在) if (!currentSentence.empty()) { size_t start = currentSentence.find_first_not_of(" \t\n\r"); size_t end = currentSentence.find_last_not_of(" \t\n\r"); if (start != string::npos && end != string::npos) { sentences.push_back(currentSentence.substr(start, end - start + 1)); } } return sentences; } int main() { ifstream inputFile("input.txt"); if (!inputFile.is_open()) { cerr << "无法打开输入文件!" << endl; return 1; } // 读取整个文件内容 string text((istreambuf_iterator<char>(inputFile)), istreambuf_iterator<char>()); inputFile.close(); vector<string> sentences = extractSentences(text); // 输出提取到的句子 cout << "提取到的句子:" << endl; for (const string& sentence : sentences) { cout << "- " << sentence << endl; } return 0; }
这个代码会把文本中的句子提取到vector里,你可以根据需求进一步处理这些句子。
问题2:按关键词拆分TXT文件(修正你的代码)
先说说你现有代码的几个问题:
- 你只判断了
buffer == "Error",但需求是句子以error、warning、information开头,还没考虑大小写不匹配的情况(比如文件里是小写error,但你写的是大写Error) - 没有打开输出文件,所以只能在控制台输出,没法写入到对应文件
- 逻辑错误:你代码里找到"Error"后会跳行读取下一行输出,但实际需求是每一行本身就以关键词开头,不需要跳行
我帮你改写了代码,实现按关键词拆分的功能:
#include <iostream> #include <fstream> #include <string> #include <algorithm> // 用于忽略大小写的判断 using namespace std; // 辅助函数:检查字符串是否以指定前缀开头(忽略大小写) bool startsWithIgnoreCase(const string& line, const string& prefix) { if (line.length() < prefix.length()) { return false; } string lineLower = line.substr(0, prefix.length()); string prefixLower = prefix; transform(lineLower.begin(), lineLower.end(), lineLower.begin(), ::tolower); transform(prefixLower.begin(), prefixLower.end(), prefixLower.begin(), ::tolower); return lineLower == prefixLower; } int main() { ifstream inputFile("try.txt"); // 打开三个输出文件,分别对应三种类型 ofstream errorFile("error.txt"); ofstream warningFile("warning.txt"); ofstream infoFile("information.txt"); // 检查所有文件是否成功打开 if (!inputFile.is_open() || !errorFile.is_open() || !warningFile.is_open() || !infoFile.is_open()) { cerr << "无法打开文件,请检查路径是否正确!" << endl; return 1; } string line; // 逐行读取输入文件 while (getline(inputFile, line)) { // 检查当前行以哪个关键词开头 if (startsWithIgnoreCase(line, "error")) { errorFile << line << endl; cout << "写入error.txt:" << line << endl; } else if (startsWithIgnoreCase(line, "warning")) { warningFile << line << endl; cout << "写入warning.txt:" << line << endl; } else if (startsWithIgnoreCase(line, "information")) { infoFile << line << endl; cout << "写入information.txt:" << line << endl; } // 如果都不匹配,可以选择忽略或者写入其他文件 } // 手动关闭文件(虽然程序结束会自动关闭,但这是好习惯) inputFile.close(); errorFile.close(); warningFile.close(); infoFile.close(); cout << "文件拆分完成!" << endl; cin.get(); return 0; }
代码说明:
- 新增了
startsWithIgnoreCase函数,用来忽略大小写判断行开头的关键词,避免因为大小写不一致导致匹配失败 - 同时打开了三个输出文件,每读取一行就判断属于哪一类,写入对应的文件
- 增加了文件打开失败的错误提示,方便排查问题
- 保留了控制台输出,方便你确认哪些内容被写入了文件
如果你不需要忽略大小写,直接用line.find(prefix) == 0判断即可(比如line.find("error") == 0)。
内容的提问来源于stack exchange,提问作者muza
相关产品推荐
相关产品推荐

