You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Lexicographical排序程序求助:重复词替换功能实现问题

问题分析与修正方案

先说说你现有代码里的几个核心问题:

  • 重复替换逻辑里j固定为1,这意味着你每次都拿第i行和第1行比,而不是和前一行或者相邻行比,完全不符合需求逻辑
  • 用字符逐个比对的方式找公共部分,没有考虑单词边界,会把单词中间的相同字符也当成重复前缀处理,比如"apple"和"app"会被错误截断
  • 没有处理字符串越界的情况,当其中一行比另一行短的时候,访问l[i][end]会直接触发程序崩溃
  • 替换逻辑只是简单删除前缀插"-",没有按单词粒度去替换,和你需求里的“前一行出现过的单词替换为‘-’”不匹配

修正后的实现思路

要实现“字典序排序后,将前一行出现过的单词替换为‘-’”,正确的步骤应该是:

  1. 逐行读取文件内容并排序(这部分你已经做好了)
  2. 把每行拆分成独立的单词列表
  3. 从第二行开始,和前一行的单词列表逐个对比:
    • 遇到和前一行相同的单词,就替换为"-"
    • 遇到第一个不同的单词后,后面的所有单词都保留原样(字典序排序后,前缀相同的单词只会集中在相邻行)
  4. 把处理后的单词列表重新拼接成字符串输出

完整修正代码

#include <iostream>
#include <list>
#include <fstream>
#include <vector>
#include <algorithm>
#include <sstream>
#include <string>

using namespace std;

// 把一行字符串拆分成单词列表
vector<string> splitLine(const string& line) {
    vector<string> words;
    stringstream ss(line);
    string word;
    while (ss >> word) {
        words.push_back(word);
    }
    return words;
}

// 把单词列表拼接成一行字符串
string joinWords(const vector<string>& words) {
    string result;
    for (size_t i = 0; i < words.size(); ++i) {
        if (i != 0) {
            result += " ";
        }
        result += words[i];
    }
    return result;
}

int main() {
    list<string> lista;
    string line;
    ifstream myfile("./file.txt");

    // 修复原代码的eof判断问题:原写法会多读一行空行
    while (getline(myfile, line)) {
        lista.push_back(line);
    }

    // 排序并转成vector
    vector<string> sortedLines;
    lista.sort();
    copy(begin(lista), end(lista), back_inserter(sortedLines));

    cout << "=== 排序后原始内容 ===" << endl;
    for (const auto& l : sortedLines) {
        cout << l << endl;
    }

    // 处理空文件情况
    if (sortedLines.empty()) {
        return 0;
    }

    vector<string> processedLines;
    processedLines.push_back(sortedLines[0]); // 第一行直接保留
    vector<string> prevWords = splitLine(sortedLines[0]);

    for (size_t i = 1; i < sortedLines.size(); ++i) {
        vector<string> currWords = splitLine(sortedLines[i]);
        vector<string> newWords;

        size_t minLen = min(prevWords.size(), currWords.size());
        size_t sameCount = 0;

        // 统计前缀相同的单词数量
        while (sameCount < minLen && currWords[sameCount] == prevWords[sameCount]) {
            newWords.push_back("-");
            sameCount++;
        }

        // 追加剩余的不同单词
        for (size_t j = sameCount; j < currWords.size(); ++j) {
            newWords.push_back(currWords[j]);
        }

        processedLines.push_back(joinWords(newWords));
        prevWords = currWords; // 更新前一行的单词列表
    }

    cout << "\n=== 处理后内容 ===" << endl;
    for (const auto& l : processedLines) {
        cout << l << endl;
    }

    return 0;
}

关键改动说明

  1. 修复了文件读取逻辑:原代码while(!myfile.eof())会导致多读一行空行,改成while(getline(myfile, line))才是正确的写法
  2. 新增了单词拆分与拼接的辅助函数,确保按单词粒度处理内容
  3. 按完整单词对比替换,避免了字符级对比带来的错误
  4. 处理了空文件、单行文件等边界情况,避免程序崩溃

举个例子,如果你的file.txt内容是:

apple banana
banana apple
apple cherry
orange

排序后输出:

=== 排序后原始内容 ===
apple banana
apple cherry
banana apple
orange

处理后输出:

=== 处理后内容 ===
apple banana
- cherry
banana apple
orange

内容的提问来源于stack exchange,提问作者operator_skrzyni_biegow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 19:15:51