You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何动态实现从文件中提取并输出最长的10个单词?

嘿,这事儿不难搞定!我帮你梳理下思路,既能移除那些死板的假设,又能轻松输出最长的10个单词~

第一步:先干掉那些死板的假设

你之前的两个限制(最长200字符、单词间最多60空格)完全没必要,咱们用语言自带的文本处理能力就能动态解决:

  • 处理任意数量的空格:不管单词之间是1个空格还是100个,直接用「按任意空白字符分割」的方法就行——比如Python里不带参数的split(),或者C++里的cin >> word,都会自动跳过所有空白(包括空格、制表符、换行),精准提取每个单词。
  • 支持任意长度的单词:不用预设最大长度,直接读取整个单词就行,编程语言本身会动态分配内存存储,不管单词是1个字符还是1000个字符都能处理。

第二步:实现「取最长10个单词」的逻辑

核心思路是收集所有单词,然后按长度排序取前10;如果是超大文件,用堆结构优化内存会更高效。下面给你两种常用语言的实现示例:

Python 实现(简单直观,适合中小文件)

import heapq

def print_top_10_longest_words(file_path):
    all_words = []
    # 逐行读取文件
    with open(file_path, 'r', encoding='utf-8') as f:
        for line in f:
            # split()无参数时,自动分割任意空白字符,忽略空字符串
            words = line.strip().split()
            all_words.extend(words)
    
    # 用heapq的nlargest直接取长度最大的10个单词,比手动排序更高效
    top_10_words = heapq.nlargest(10, all_words, key=lambda word: len(word))
    
    # 逐行输出每个单词
    for word in top_10_words:
        print(word)

# 调用示例,替换成你的文件路径
print_top_10_longest_words("your_input_file.txt")

C++ 实现(内存高效,适合超大文件)

如果你的文件特别大,不想把所有单词都存进内存,可以用最小堆来维护「当前最长的10个单词」,这样内存占用始终保持在10个单词的大小:

#include <iostream>
#include <fstream>
#include <vector>
#include <string>
#include <queue>
#include <algorithm>

using namespace std;

// 自定义堆的比较规则:让长度最短的单词留在堆顶,方便替换
struct CompareWord {
    bool operator()(const string& a, const string& b) {
        if (a.size() == b.size()) {
            return a < b; // 长度相同时,字典序大的优先(可选规则)
        }
        return a.size() > b.size();
    }
};

void print_top_10_longest_words(const string& file_path) {
    ifstream input_file(file_path);
    if (!input_file.is_open()) {
        cerr << "Failed to open file!" << endl;
        return;
    }

    priority_queue<string, vector<string>, CompareWord> min_heap;
    string current_word;

    // 自动跳过所有空白字符,读取每个单词
    while (input_file >> current_word) {
        if (min_heap.size() < 10) {
            min_heap.push(current_word);
        } else {
            // 如果当前单词比堆顶的单词更长,就替换堆顶
            if (current_word.size() > min_heap.top().size()) {
                min_heap.pop();
                min_heap.push(current_word);
            }
            // 要是你想包含长度并列的单词,可以在这里调整逻辑,比如长度相等也加入堆(不过要注意堆的大小控制)
        }
    }

    // 把堆里的单词转成vector,逆序后输出(因为堆顶是第10长的单词)
    vector<string> top_10;
    while (!min_heap.empty()) {
        top_10.push_back(min_heap.top());
        min_heap.pop();
    }
    reverse(top_10.begin(), top_10.end());

    for (const string& word : top_10) {
        cout << word << endl;
    }

    input_file.close();
}

int main() {
    print_top_10_longest_words("your_input_file.txt");
    return 0;
}

额外说明

  • 关于并列第10长的单词:上面的代码默认会取最先遇到的或者按字典序筛选的,如果你想把所有并列的都输出,可以在排序/堆处理后,额外检查后续单词的长度是否和第10个相等,然后一起输出——这个逻辑可以根据你的需求灵活调整。
  • 编码问题:如果你的文件是非UTF-8编码(比如GBK),记得在打开文件时指定对应的编码格式,避免乱码。

内容的提问来源于stack exchange,提问作者Exiler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:34:30