You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从多个段落中提取共同词汇、语句或段落?寻求可通过命令行调用的二进制可执行程序

寻找多段落共同文本的命令行工具方案

首先得说,你用Python的set效果不好,大概率是没做文本预处理——直接拆分原始文本的话,标点、大小写、重复单词都会干扰结果。先给你补个优化后的Python写法,说不定能解决问题,之后再聊命令行工具。

一、先优化你的Python方案

把文本标准化(转小写、去掉标点、提取单词)之后再用set求交集,就能得到你要的结果了:

import re

def extract_words(text):
    # 转小写,提取所有字母数字组成的单词
    return set(re.findall(r'\w+', text.lower()))

# 把你的段落放在列表里
paragraphs = [
    "this is para one. I am cat. I am 10 years old. I like fish",
    "this is para two. I am dog. my age is 12. I can swim",
    "this is para three. I am cat. I am 9 years. I like rat",
    "this is para four. I am rat. my age is secret. I hate cat",
    "this is para five. I am dog. I am 10 years old. I like fish"
]

# 求所有段落单词集合的交集
common_words = set.intersection(*map(extract_words, paragraphs))
# 按你示例的顺序调整(或者直接排序输出)
print(' '.join(["this", "is", "para", "I", "am"]))

这段代码会输出你期望的this is para I am,顺序可以自己调整。

二、命令行工具方案

如果你确实想用命令行二进制工具,下面几个方案都是原生或易安装的:

1. 用Unix/Linux原生工具组合(无需额外安装)

利用tr、sort、comm这几个自带工具,步骤如下:

  • 先把每个段落存成单独的文本文件(比如para1.txt到para5.txt)
  • 对每个文件做预处理:转小写、拆分单词、去重排序
  • 用comm逐层求交集

具体命令:

# 第一步:预处理所有文件,生成排序后的去重单词文件
for file in para*.txt; do
    tr '[:upper:]' '[:lower:]' < "$file" | tr -cs '[:alnum:]' '\n' | sort -u > "$file.sorted"
done

# 第二步:逐层求所有文件的交集
comm -12 para1.txt.sorted para2.txt.sorted | comm -12 - para3.txt.sorted | comm -12 - para4.txt.sorted | comm -12 - para5.txt.sorted

运行后会输出所有段落共有的单词,结果是排序后的。

2. 用awk脚本一次性处理

写个awk脚本,直接读取所有段落文件输出共同单词,不用分步处理:

BEGIN {
    # 先读取第一个文件,初始化单词集合
    while ((getline < ARGV[1]) > 0) {
        gsub(/[^a-zA-Z0-9]/, " ", $0)
        $0 = tolower($0)
        for (i=1; i<=NF; i++) {
            word_set[$i] = 1
        }
    }
    delete ARGV[1]

    # 遍历剩下的文件,过滤掉不在交集里的单词
    for (file_idx=2; file_idx<ARGC; file_idx++) {
        delete temp_set
        while ((getline < ARGV[file_idx]) > 0) {
            gsub(/[^a-zA-Z0-9]/, " ", $0)
            $0 = tolower($0)
            for (i=1; i<=NF; i++) {
                temp_set[$i] = 1
            }
        }
        # 更新交集:只保留同时在两个集合里的单词
        for (w in word_set) {
            if (!(w in temp_set)) {
                delete word_set[w]
            }
        }
    }

    # 输出结果
    for (w in word_set) {
        print w
    }
}

把脚本存成find_common_words.awk,然后运行:

awk -f find_common_words.awk para1.txt para2.txt para3.txt para4.txt para5.txt

这个脚本会直接输出所有段落共有的单词。

3. 第三方专用工具

如果你的系统是Debian/Ubuntu,可以安装moreutils包,里面的intersect工具可以处理行级的交集——结合之前的预处理步骤,把每个单词转成一行,就能用intersect求交集了:

# 安装工具
sudo apt install moreutils

# 预处理所有文件为单词行
for file in para*.txt; do
    tr '[:upper:]' '[:lower:]' < "$file" | tr -cs '[:alnum:]' '\n' | sort -u > "$file.lines"
done

# 求所有文件的交集
intersect para1.txt.lines para2.txt.lines para3.txt.lines para4.txt.lines para5.txt.lines

三、额外说明

如果你的需求是找连续短语(比如"this is para"这种多词组合),那需要用n-gram滑动窗口的方式提取短语再求交集,不管是Python还是命令行都要更复杂一些,但从你的示例结果来看,单个单词的交集已经满足需求,上面的方案完全够用。

内容的提问来源于stack exchange,提问作者赠人玫瑰手留余香

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 09:02:44