UTF-16字符组合生成脚本故障排查与功能修复请求
修复Discord违规词生成脚本的三个核心问题
1. 无法跳过亚洲字符的修复
原脚本未精准识别中日韩等亚洲字符的Unicode范围,导致过滤规则失效。通过检查字符编码是否属于CJK相关区块,可实现精准过滤:
def is_asian_char(char): code = ord(char) # 覆盖中日韩主要字符区块 return ( (0x4E00 <= code <= 0x9FFF) or # CJK统一表意文字 (0x3400 <= code <= 0x4DBF) or # CJK扩展A (0x20000 <= code <= 0x2A6DF) or # CJK扩展B (0xF900 <= code <= 0xFAFF) or # CJK兼容表意文字 (0x2F800 <= code <= 0x2FA1F) # CJK兼容扩展 )
生成特殊字符时,跳过返回True的字符即可。
2. 文件未写入内容的修复
原脚本缺少断点续存的读取逻辑和正确的文件写入模式。先读取目标文件已有的条目数,从该数值开始计数,每次生成有效组合后立即追加写入:
def get_existing_count(file_path): try: with open(file_path, 'r', encoding='utf-16') as f: return sum(1 for _ in f) except FileNotFoundError: return 0 # 写入逻辑示例 with open(output_file, 'a', encoding='utf-16') as f: f.write(f"{combination}\n")
使用'a'模式确保追加写入,utf-16编码匹配脚本需求。
3. 无限循环的修复
递归函数未设置终止条件,导致无法停止。维护一个计数器,当生成的组合数达到1000时立即终止所有递归:
def generate_all_utf16_combinations(base_word, special_chars, current_count, max_count, output_file): # 先检查是否已达上限,终止递归 if current_count >= max_count: return current_count # 生成基础单词全排列 for perm in permutations(base_word): perm_str = ''.join(perm) # 写入基础排列 with open(output_file, 'a', encoding='utf-16') as f: f.write(f"{perm_str}\n") current_count += 1 if current_count >= max_count: return current_count # 插入特殊字符及递归逻辑 for i in range(len(perm_str) + 1): for char in special_chars: if is_asian_char(char): continue new_str = perm_str[:i] + char + perm_str[i:] with open(output_file, 'a', encoding='utf-16') as f: f.write(f"{new_str}\n") current_count += 1 if current_count >= max_count: return current_count # 递归插入更多特殊字符 current_count = generate_all_utf16_combinations(new_str, special_chars, current_count, max_count, output_file) if current_count >= max_count: return current_count return current_count
每次递归调用前检查计数器,达到上限后立即返回,终止后续递归。
完整修复后的脚本
from itertools import permutations def is_asian_char(char): code = ord(char) return ( (0x4E00 <= code <= 0x9FFF) or (0x3400 <= code <= 0x4DBF) or (0x20000 <= code <= 0x2A6DF) or (0xF900 <= code <= 0xFAFF) or (0x2F800 <= code <= 0x2FA1F) ) def get_existing_count(file_path): try: with open(file_path, 'r', encoding='utf-16') as f: return sum(1 for _ in f) except FileNotFoundError: return 0 def generate_all_utf16_combinations(base_word, special_chars, current_count, max_count, output_file): if current_count >= max_count: return current_count # 生成基础单词的全排列 for perm in permutations(base_word): perm_str = ''.join(perm) # 写入基础排列 with open(output_file, 'a', encoding='utf-16') as f: f.write(f"{perm_str}\n") current_count += 1 if current_count >= max_count: return current_count # 插入单个特殊字符 for i in range(len(perm_str) + 1): for char in special_chars: if is_asian_char(char): continue new_str = perm_str[:i] + char + perm_str[i:] with open(output_file, 'a', encoding='utf-16') as f: f.write(f"{new_str}\n") current_count += 1 if current_count >= max_count: return current_count # 递归插入更多特殊字符 current_count = generate_all_utf16_combinations(new_str, special_chars, current_count, max_count, output_file) if current_count >= max_count: return current_count return current_count # 主逻辑 if __name__ == "__main__": target_word = "banana" # 示例特殊字符集合(可按需扩展) special_chars = "!@#$%^&*()_+-=[]{}|;:'\",.<>?/~`" output_file = "bad_words.txt" max_results = 1000 existing_count = get_existing_count(output_file) if existing_count >= max_results: print(f"已达到最大生成数量{max_results},无需继续生成") else: generate_all_utf16_combinations(target_word, special_chars, existing_count, max_results, output_file) print(f"生成完成,共写入{max_results - existing_count}条新内容,总条目数:{get_existing_count(output_file)}")
内容的提问来源于stack exchange,提问作者jose vazquez
相关产品推荐
相关产品推荐

