You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于permutations()的文件单词打乱程序及函数实现咨询

完整解决方案:用itertools.permutations实现文件单词打乱

我来帮你补全这个程序,先明确核心需求:处理文件中的每个单词,长度≤3的保持原样,长度>3的通过itertools.permutations生成打乱后的版本(支持两种模式:生成所有可能排列,或每个单词随机输出一个打乱版本),最终输出符合预期的结果。


版本1:生成所有可能的唯一排列

这个版本会把每个长单词的所有不重复排列都列出来,按原单词的顺序整理到结果中:

import itertools as it

def get_permuted_lines(word_list):
    '''该函数接收文件中按出现顺序排列的单词列表,返回按相同顺序排列的所有打乱后单词的列表'''
    final_list = []
    for word in word_list:
        # 长度≤3的单词直接保留原词
        if len(word) <= 3:
            final_list.append(word)
            continue
        
        # 生成该单词的所有字母排列,转成字符串后去重(避免重复字母导致的冗余排列)
        all_permutations = it.permutations(word)
        unique_perms = set(''.join(perm) for perm in all_permutations)
        
        # 将去重后的排列批量加入结果列表
        final_list.extend(unique_perms)
    
    return final_list

def process_file(input_path, output_path):
    # 读取文件内容,按空格分割单词(如果需要处理带标点的文本,可以用正则优化)
    with open(input_path, 'r', encoding='utf-8') as f:
        content = f.read()
    words = content.split()
    
    # 获取处理后的单词列表并写入输出文件
    processed_words = get_permuted_lines(words)
    with open(output_path, 'w', encoding='utf-8') as f:
        f.write('\n'.join(processed_words))

# 示例调用
if __name__ == "__main__":
    process_file("input.txt", "output_all_permutations.txt")

版本2:每个单词生成一个随机打乱版本

如果你的需求是生成类似“单词打乱但可读性保留”的效果(比如常见的首尾字母不变、中间打乱,这里用全排列实现随机打乱),可以用这个版本,每个单词只输出一个随机打乱后的结果:

import itertools as it
import random as rdm
import re

def get_permuted_lines(word_list):
    '''该函数接收文件中按出现顺序排列的单词列表,返回按相同顺序排列的每个单词的随机打乱版本列表'''
    final_list = []
    for word in word_list:
        word_len = len(word)
        if word_len <= 3:
            final_list.append(word)
            continue
        
        # 生成所有唯一排列,并排除原单词(可选,若不需要排除可删除此行)
        all_perms = set(''.join(perm) for perm in it.permutations(word))
        all_perms.discard(word)
        
        # 随机选一个打乱版本,若没有可用排列(比如全相同字母的单词)则保留原词
        if all_perms:
            final_list.append(rdm.choice(list(all_perms)))
        else:
            final_list.append(word)
    
    return final_list

def process_file(input_path, output_path):
    # 用正则分割单词,处理带标点的文本(比如"hello,"会被拆成"hello",避免标点干扰)
    with open(input_path, 'r', encoding='utf-8') as f:
        words = re.findall(r'\b\w+\b', f.read())
    
    # 获取处理后的单词列表,按原格式用空格分隔写入文件
    processed_words = get_permuted_lines(words)
    with open(output_path, 'w', encoding='utf-8') as f:
        f.write(' '.join(processed_words))

# 示例调用
if __name__ == "__main__":
    process_file("input.txt", "output_random_perm.txt")

关键逻辑说明

  1. 短单词处理:严格遵循需求,长度≤3的单词直接保留原词,不做任何打乱操作。
  2. 排列去重:itertools.permutations会生成所有字母排列,但如果单词有重复字母(比如"apple"),会产生大量重复结果,用set去重可以有效避免冗余。
  3. 文本兼容性:第二个版本用正则re.findall(r'\b\w+\b')分割单词,能更好地处理带标点的真实文本,避免标点和单词混在一起。

内容的提问来源于stack exchange,提问作者Raj Kumar Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:19:00