You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++解析文件名生成CSV:移除分组列冗余词汇求助

问题解决:C++解析CSV文件名并清理组名冗余词汇

问题背景

需用C++解析file_list.csv中的文件名,生成包含StudentList、GroupName、QuestionsAnswered列的CSV文件。现有代码无法去除GroupName列中的"Group""question"等冗余词汇,需修正。

现有代码问题

原正则表达式R"(([a-zA-Z]+)_?(\d+)_?(\d+)?.*)"无法适配所有文件名格式:

  • 无法识别Group a_1_xxx这类带空格的组名,会错误将"Group"提取为组名
  • 处理d_question1_xxx时,会把"d"或"question"误判为组名
  • 未处理大小写差异(如Group/group、QUESTION/question)

修改后的完整代码

#include <iostream>
#include <fstream>
#include <sstream>
#include <vector>
#include <unordered_map>
#include <regex>
#include <algorithm>
#include <cctype>

// 分割字符串并去除首尾空格
std::vector<std::string> splitString(const std::string& s, char delimiter) {
    std::vector<std::string> tokens;
    std::stringstream ss(s);
    std::string token;
    while (std::getline(ss, token, delimiter)) {
        // 移除首尾空格
        token.erase(token.begin(), std::find_if(token.begin(), token.end(), [](unsigned char ch) {
            return !std::isspace(ch);
        }));
        token.erase(std::find_if(token.rbegin(), token.rend(), [](unsigned char ch) {
            return !std::isspace(ch);
        }).base(), token.end());
        if (!token.empty()) {
            tokens.push_back(token);
        }
    }
    return tokens;
}

// 清理组名中的冗余词汇
std::string cleanGroupName(std::string group) {
    // 转换为小写用于匹配
    std::string lowerGroup = group;
    std::transform(lowerGroup.begin(), lowerGroup.end(), lowerGroup.begin(), ::tolower);

    // 移除冗余关键词
    std::vector<std::string> redundantWords = {"group", "question"};
    for (const auto& word : redundantWords) {
        size_t pos = lowerGroup.find(word);
        while (pos != std::string::npos) {
            // 删除对应位置的原字符串内容
            group.erase(pos, word.length());
            lowerGroup.erase(pos, word.length());
            // 清理删除后可能出现的多余下划线或空格
            group.erase(std::remove_if(group.begin(), group.end(), [](unsigned char ch) {
                return ch == '_' || std::isspace(ch);
            }), group.end());
            lowerGroup = group;
            std::transform(lowerGroup.begin(), lowerGroup.end(), lowerGroup.begin(), ::tolower);
            pos = lowerGroup.find(word);
        }
    }

    // 最后清理首尾的下划线和空格
    group.erase(group.begin(), std::find_if(group.begin(), group.end(), [](unsigned char ch) {
        return ch != '_' && !std::isspace(ch);
    }));
    group.erase(std::find_if(group.rbegin(), group.rend(), [](unsigned char ch) {
        return ch != '_' && !std::isspace(ch);
    }).base(), group.end());

    // 首字母大写统一格式(可选)
    if (!group.empty()) {
        group[0] = std::toupper(group[0]);
    }
    return group;
}

int main() {
    const std::string inputFile = "file_list.csv";
    const std::string outputFile = "output.csv";

    std::ifstream inFile(inputFile);
    if (!inFile.is_open()) {
        std::cerr << "无法打开输入文件: " << inputFile << std::endl;
        return 1;
    }

    std::unordered_map<std::string, std::pair<std::string, std::vector<int>>> studentInfo;

    // 适配所有测试文件名格式的正则表达式
    std::regex pattern(R"((?:(Group|group)\s*)?([a-zA-Z]+)(?:_question|_QUESTION\s*)?(\d+)_?(\d+).*)");

    std::string line;
    while (std::getline(inFile, line)) {
        std::smatch match;
        if (std::regex_match(line, match, pattern)) {
            std::string rawGroup = match[2];
            std::string studentID = match[4];
            std::string questionNumberStr = match[3];

            int questionNumber = -1;
            if (!questionNumberStr.empty()) {
                questionNumber = std::stoi(questionNumberStr);
            }

            if (!studentID.empty()) {
                // 清理组名
                std::string cleanedGroup = cleanGroupName(rawGroup);
                // 若清理后为空,保留原始标识(如d)
                if (cleanedGroup.empty()) {
                    cleanedGroup = rawGroup;
                    std::transform(cleanedGroup.begin(), cleanedGroup.end(), cleanedGroup.begin(), ::toupper);
                }

                if (studentInfo.find(studentID) == studentInfo.end()) {
                    studentInfo[studentID] = std::make_pair(cleanedGroup, std::vector<int>());
                }
                // 避免重复添加同一题号
                if (questionNumber != -1) {
                    auto& questions = studentInfo[studentID].second;
                    if (std::find(questions.begin(), questions.end(), questionNumber) == questions.end()) {
                        questions.push_back(questionNumber);
                    }
                }
            }
        }
    }

    inFile.close();

    std::ofstream outFile(outputFile);
    if (!outFile.is_open()) {
        std::cerr << "无法打开输出文件: " << outputFile << std::endl;
        return 1;
    }

    outFile << "StudentList,GroupName,QuestionsAnswered" << std::endl;

    for (const auto& entry : studentInfo) {
        std::vector<int> questions = entry.second.second;
        std::sort(questions.begin(), questions.end());

        outFile << entry.first << "," << entry.second.first << ",";
        for (size_t i = 0; i < questions.size(); ++i) {
            outFile << questions[i];
            if (i < questions.size() - 1) {
                outFile << ",";
            }
        }
        outFile << std::endl;
    }

    outFile.close();
    std::cout << "输出文件生成成功: " << outputFile << std::endl;

    return 0;
}

关键修改说明

  1. 正则表达式优化:新增正则可匹配带"Group"前缀、"question"关键词的文件名,准确提取组名、题号和学号,覆盖所有测试格式。
  2. 组名清理逻辑:cleanGroupName函数统一处理大小写,移除冗余关键词,同时清理多余的下划线和空格,确保组名简洁规范。
  3. 题号去重:添加题号存在性检查,避免同一学生同一题号重复记录。
  4. 空格处理:分割字符串时自动去除首尾空格,适配文件名中存在的冗余空格(如d_3_2211011096 .txt)。

测试输出示例

针对给出的file_list.csv内容,输出文件部分内容如下:

StudentList,GroupName,QuestionsAnswered
2211011228,D,3
2211012211,D,3
2211011054,D,3
2211011096,D,3
2111011034,D,1,2,3
2211011032,A,1,2,3
2211011024,C,1

内容的提问来源于stack exchange,提问作者Athos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 09:53:16