C++解析文件名生成CSV:移除分组列冗余词汇求助
问题解决:C++解析CSV文件名并清理组名冗余词汇
问题背景
需用C++解析file_list.csv中的文件名,生成包含StudentList、GroupName、QuestionsAnswered列的CSV文件。现有代码无法去除GroupName列中的"Group""question"等冗余词汇,需修正。
现有代码问题
原正则表达式R"(([a-zA-Z]+)_?(\d+)_?(\d+)?.*)"无法适配所有文件名格式:
- 无法识别
Group a_1_xxx这类带空格的组名,会错误将"Group"提取为组名 - 处理
d_question1_xxx时,会把"d"或"question"误判为组名 - 未处理大小写差异(如
Group/group、QUESTION/question)
修改后的完整代码
#include <iostream> #include <fstream> #include <sstream> #include <vector> #include <unordered_map> #include <regex> #include <algorithm> #include <cctype> // 分割字符串并去除首尾空格 std::vector<std::string> splitString(const std::string& s, char delimiter) { std::vector<std::string> tokens; std::stringstream ss(s); std::string token; while (std::getline(ss, token, delimiter)) { // 移除首尾空格 token.erase(token.begin(), std::find_if(token.begin(), token.end(), [](unsigned char ch) { return !std::isspace(ch); })); token.erase(std::find_if(token.rbegin(), token.rend(), [](unsigned char ch) { return !std::isspace(ch); }).base(), token.end()); if (!token.empty()) { tokens.push_back(token); } } return tokens; } // 清理组名中的冗余词汇 std::string cleanGroupName(std::string group) { // 转换为小写用于匹配 std::string lowerGroup = group; std::transform(lowerGroup.begin(), lowerGroup.end(), lowerGroup.begin(), ::tolower); // 移除冗余关键词 std::vector<std::string> redundantWords = {"group", "question"}; for (const auto& word : redundantWords) { size_t pos = lowerGroup.find(word); while (pos != std::string::npos) { // 删除对应位置的原字符串内容 group.erase(pos, word.length()); lowerGroup.erase(pos, word.length()); // 清理删除后可能出现的多余下划线或空格 group.erase(std::remove_if(group.begin(), group.end(), [](unsigned char ch) { return ch == '_' || std::isspace(ch); }), group.end()); lowerGroup = group; std::transform(lowerGroup.begin(), lowerGroup.end(), lowerGroup.begin(), ::tolower); pos = lowerGroup.find(word); } } // 最后清理首尾的下划线和空格 group.erase(group.begin(), std::find_if(group.begin(), group.end(), [](unsigned char ch) { return ch != '_' && !std::isspace(ch); })); group.erase(std::find_if(group.rbegin(), group.rend(), [](unsigned char ch) { return ch != '_' && !std::isspace(ch); }).base(), group.end()); // 首字母大写统一格式(可选) if (!group.empty()) { group[0] = std::toupper(group[0]); } return group; } int main() { const std::string inputFile = "file_list.csv"; const std::string outputFile = "output.csv"; std::ifstream inFile(inputFile); if (!inFile.is_open()) { std::cerr << "无法打开输入文件: " << inputFile << std::endl; return 1; } std::unordered_map<std::string, std::pair<std::string, std::vector<int>>> studentInfo; // 适配所有测试文件名格式的正则表达式 std::regex pattern(R"((?:(Group|group)\s*)?([a-zA-Z]+)(?:_question|_QUESTION\s*)?(\d+)_?(\d+).*)"); std::string line; while (std::getline(inFile, line)) { std::smatch match; if (std::regex_match(line, match, pattern)) { std::string rawGroup = match[2]; std::string studentID = match[4]; std::string questionNumberStr = match[3]; int questionNumber = -1; if (!questionNumberStr.empty()) { questionNumber = std::stoi(questionNumberStr); } if (!studentID.empty()) { // 清理组名 std::string cleanedGroup = cleanGroupName(rawGroup); // 若清理后为空,保留原始标识(如d) if (cleanedGroup.empty()) { cleanedGroup = rawGroup; std::transform(cleanedGroup.begin(), cleanedGroup.end(), cleanedGroup.begin(), ::toupper); } if (studentInfo.find(studentID) == studentInfo.end()) { studentInfo[studentID] = std::make_pair(cleanedGroup, std::vector<int>()); } // 避免重复添加同一题号 if (questionNumber != -1) { auto& questions = studentInfo[studentID].second; if (std::find(questions.begin(), questions.end(), questionNumber) == questions.end()) { questions.push_back(questionNumber); } } } } } inFile.close(); std::ofstream outFile(outputFile); if (!outFile.is_open()) { std::cerr << "无法打开输出文件: " << outputFile << std::endl; return 1; } outFile << "StudentList,GroupName,QuestionsAnswered" << std::endl; for (const auto& entry : studentInfo) { std::vector<int> questions = entry.second.second; std::sort(questions.begin(), questions.end()); outFile << entry.first << "," << entry.second.first << ","; for (size_t i = 0; i < questions.size(); ++i) { outFile << questions[i]; if (i < questions.size() - 1) { outFile << ","; } } outFile << std::endl; } outFile.close(); std::cout << "输出文件生成成功: " << outputFile << std::endl; return 0; }
关键修改说明
- 正则表达式优化:新增正则可匹配带"Group"前缀、"question"关键词的文件名,准确提取组名、题号和学号,覆盖所有测试格式。
- 组名清理逻辑:
cleanGroupName函数统一处理大小写,移除冗余关键词,同时清理多余的下划线和空格,确保组名简洁规范。 - 题号去重:添加题号存在性检查,避免同一学生同一题号重复记录。
- 空格处理:分割字符串时自动去除首尾空格,适配文件名中存在的冗余空格(如
d_3_2211011096 .txt)。
测试输出示例
针对给出的file_list.csv内容,输出文件部分内容如下:
StudentList,GroupName,QuestionsAnswered 2211011228,D,3 2211012211,D,3 2211011054,D,3 2211011096,D,3 2111011034,D,1,2,3 2211011032,A,1,2,3 2211011024,C,1
内容的提问来源于stack exchange,提问作者Athos
相关产品推荐
相关产品推荐

