如何从单词列表生成有意义句子?现有实现无有效结果求解决
问题描述
需求:从一组(理想状态下至少100个)单词中生成符合语法、具备语义的句子。
已尝试方案及问题:
- 语法检查API仅能修正近似正确的句子,无法处理随机单词组合;
- 使用NLTK生成的句子完全无意义,示例:
door to me go
me go our game
go our game eight
game eight plays go
eight plays go me
my dog is blue you
dog is blue you her
is blue you her run. - 直接组合学方法因可能性过多不可行;
- 尝试结合
itertools排列与多进程,通过Tatoeba数据集验证有效性,但运行1小时未得到任何有效句子(用户无需所有可能句子,支持提前终止)。
附当前Python实现代码:
import hashlib import itertools import multiprocessing import time import random from nltk import pos_tag from nltk.corpus import wordnet as wn import nltk # NLTK setup nltk.download('averaged_perceptron_tagger') nltk.download('wordnet') # Force data load pos_tag(['test']) wn.synsets('dog') def hash_sentence(sentence): """Hashes a sentence using SHA-256.""" return hashlib.sha256(sentence.lower().encode()).hexdigest() def generate_permutations(shared_list, words, max_length=7): """Generates filtered permutations of words and adds them to a shared list.""" for r in range(2, max_length + 1): for perm in itertools.permutations(words, r): with shared_list.get_lock(): # Synchronize access shared_list.append(' '.join(perm)) with shared_list.get_lock(): shared_list.append(None) # Signal that generation is complete def validate_sentences(shared_list, hashed_sentences, output_file): """Validates sentences from a shared list against hashed sentences and writes valid ones to a file.""" with open(output_file, 'a', encoding='utf-8') as file: while True: time.sleep(0.1) # Small delay to allow generator to populate list with shared_list.get_lock(): if not shared_list: continue sentence = shared_list.pop(0) if sentence is None: break if hash_sentence(sentence) in hashed_sentences: print(f"Valid sentence: {sentence}") file.write(sentence + '\n') def shuffle_list(shared_list): """Shuffles the sentences in the shared list every 3 minutes.""" while True: time.sleep(180) # Sleep for 3 minutes with shared_list.get_lock(): # Shuffle all items except the termination signal (None) if it exists sentences = [s for s in shared_list if s is not None] random.shuffle(sentences) shared_list[:] = sentences if None in shared_list: shared_list.append(None) def load_hashed_sentences(filepath): """Loads and hashes sentences from the Tatoeba dataset.""" hashed_sentences = set() with open(filepath, 'r', encoding='utf-8') as file: next(file) # Skip the header if present for line in file: parts = line.strip().split('\t') if parts[1] == 'eng': sentence_hash = hash_sentence(parts[2]) hashed_sentences.add(sentence_hash) return hashed_sentences def main(): manager = multiprocessing.Manager() shared_list = manager.list() # Shared list managed by the manager tsv_path = r"C:\Users\BaseC\Code\French\Sentence_Gen\eng_sentences.tsv" hashed_sentences = load_hashed_sentences(tsv_path) words = "my dog is blue you her run she orange cow door to me go our game eight plays go me A broad river runs through the city".split(' ') output_file = 'valid_sentences.txt' generator_process = multiprocessing.Process(target=generate_permutations, args=(shared_list, words, 7)) validator_process = multiprocessing.Process(target=validate_sentences, args=(shared_list, hashed_sentences, output_file)) shuffler_process = multiprocessing.Process(target=shuffle_list, args=(shared_list,)) generator_process.start() validator_process.start() shuffler_process.start() generator_process.join() validator_process.join() shuffler_process.join() if __name__ == "__main__": main()
问题分析
现有方案的核心缺陷:
- 全排列组合空间过大:即使100个单词生成7词句子,排列数约为8.6e12,随机生成的句子几乎不可能匹配Tatoeba中的现成句子,概率趋近于0;
- 未利用语法规则约束:代码导入了NLTK词性标注工具但未使用,生成的句子完全忽略词性搭配(比如动词放末尾、代词误用),必然无意义;
- 验证方式过于死板:要求生成的句子与Tatoeba数据集完全一致,而实际需要的是「符合语法语义」的句子,而非现成句子的复刻。
解决方案
针对编程新手,推荐两种可行路径,复杂度从低到高:
路径1:词性分类+规则生成+语法验证
先对单词做词性分类,基于基础语法规则生成候选句子,再用语法检查工具验证有效性,大幅缩小生成范围。
实现步骤:
- 用NLTK对单词做词性标注,分类为代词、动词、名词、形容词、介词等;
- 定义简单的语法句式模板,比如:
- 代词 + 动词 + 形容词 + 名词
- 形容词 + 名词 + 动词 + 介词 + 名词
- 从分类后的单词池中随机选取符合词性的单词填充模板;
- 用
language-tool-python检查语法,保留通过验证的句子。
示例代码:
import nltk import random import language_tool_python # 初始化工具 nltk.download('averaged_perceptron_tagger') tool = language_tool_python.LanguageTool('en-US') def pos_tag_words(words): """给单词做词性标注并分类""" tagged = nltk.pos_tag(words) # 简化词性映射 word_classes = { 'pronoun': [], # 代词:I, you, she, my, her... 'verb': [], # 动词:run, plays, go, runs... 'noun': [], # 名词:dog, game, river, cow... 'adj': [], # 形容词:blue, orange, broad... 'prep': [] # 介词:to, through... } for word, tag in tagged: if tag.startswith('PR'): word_classes['pronoun'].append(word) elif tag.startswith('VB'): word_classes['verb'].append(word) elif tag.startswith('NN'): word_classes['noun'].append(word) elif tag.startswith('JJ'): word_classes['adj'].append(word) elif tag.startswith('IN'): word_classes['prep'].append(word) # 去重 for key in word_classes: word_classes[key] = list(set(word_classes[key])) return word_classes def generate_candidate_sentences(word_classes, num_candidates=100): """基于句式模板生成候选句子""" templates = [ "{pronoun} {verb} the {adj} {noun}", "{adj} {noun} {verb} {prep} {noun}", "{pronoun} {verb} {pronoun}'s {noun}", "The {adj} {noun} {verb} {prep} {pronoun}" ] candidates = [] for _ in range(num_candidates): template = random.choice(templates) try: sentence = template.format( pronoun=random.choice(word_classes['pronoun']), verb=random.choice(word_classes['verb']), noun=random.choice(word_classes['noun']), adj=random.choice(word_classes['adj']), prep=random.choice(word_classes['prep']) ) # 首字母大写,末尾加句号 sentence = sentence.capitalize() + '.' candidates.append(sentence) except IndexError: # 某类单词为空时跳过 continue return candidates def filter_valid_sentences(candidates): """用语法工具过滤有效句子""" valid = [] for sent in candidates: matches = tool.check(sent) if len(matches) == 0: valid.append(sent) return valid def main(): words = "my dog is blue you her run she orange cow door to me go our game eight plays go me A broad river runs through the city".split(' ') word_classes = pos_tag_words(words) candidates = generate_candidate_sentences(word_classes, 200) valid_sentences = filter_valid_sentences(candidates) print("生成的有效句子:") for sent in valid_sentences[:10]: # 输出前10个 print(sent) # 保存到文件 with open('valid_sentences.txt', 'w', encoding='utf-8') as f: for sent in valid_sentences: f.write(sent + '\n') if __name__ == "__main__": main()
路径2:用预训练语言模型生成(推荐)
利用预训练语言模型(比如DistilGPT2)直接基于给定单词生成通顺的句子,无需手动处理语法规则,效果更好。
实现步骤:
- 安装
transformers和torch库; - 使用
pipeline创建文本生成任务; - 构造提示词,让模型基于给定单词生成句子;
- 过滤生成结果,保留符合要求的句子。
示例代码:
from transformers import pipeline, AutoTokenizer, AutoModelForCausalLM import torch def generate_sentences_with_llm(words, num_sentences=10): """用预训练语言模型生成句子""" # 使用轻量版DistilGPT2,适合新手 model_name = "distilgpt2" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained(model_name) # 添加pad token(DistilGPT2默认没有) tokenizer.pad_token = tokenizer.eos_token # 构造提示词:让模型用给定单词生成句子 prompt = f"Use these words to make meaningful sentences: {', '.join(words)}\n1. " generator = pipeline( "text-generation", model=model, tokenizer=tokenizer, torch_dtype=torch.float32, device=-1 # 使用CPU,有GPU的话改成0 ) outputs = generator( prompt, max_new_tokens=100, num_return_sequences=num_sentences, temperature=0.7, # 控制生成随机性,0.7适中 top_p=0.9, pad_token_id=tokenizer.eos_token_id ) # 提取生成的句子,整理格式 sentences = [] for output in outputs: text = output['generated_text'].split('\n')[1:] # 提取第一条之后的句子 for sent in text: if sent.strip() and sent[-1] in '.!?': # 保留完整句子 sentences.append(sent.strip()) # 去重 sentences = list(set(sentences)) return sentences[:num_sentences] def main(): words = "my dog is blue you her run she orange cow door to me go our game eight plays go me A broad river runs through the city".split(' ') valid_sentences = generate_sentences_with_llm(words, 10) print("生成的有效句子:") for i, sent in enumerate(valid_sentences, 1): print(f"{i}. {sent}") with open('valid_sentences.txt', 'w', encoding='utf-8') as f: for sent in valid_sentences: f.write(sent + '\n') if __name__ == "__main__": main()
关键优化点说明
- 缩小生成范围:通过词性分类或预训练模型的语法知识,避免无意义的随机排列;
- 灵活验证:用语法检查工具或模型自身的语义能力替代死板的数据集匹配;
- 提前终止:两种方案都可以设置生成数量(比如生成10个有效句子就停止),符合需求。
内容的提问来源于stack exchange,提问作者user24883689
相关产品推荐
相关产品推荐

