You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

UTF-16字符组合生成脚本故障排查与功能修复请求

修复Discord违规词生成脚本的三个核心问题

1. 无法跳过亚洲字符的修复

原脚本未精准识别中日韩等亚洲字符的Unicode范围,导致过滤规则失效。通过检查字符编码是否属于CJK相关区块,可实现精准过滤:

def is_asian_char(char):
    code = ord(char)
    # 覆盖中日韩主要字符区块
    return (
        (0x4E00 <= code <= 0x9FFF) or  # CJK统一表意文字
        (0x3400 <= code <= 0x4DBF) or  # CJK扩展A
        (0x20000 <= code <= 0x2A6DF) or  # CJK扩展B
        (0xF900 <= code <= 0xFAFF) or  # CJK兼容表意文字
        (0x2F800 <= code <= 0x2FA1F)  # CJK兼容扩展
    )

生成特殊字符时,跳过返回True的字符即可。

2. 文件未写入内容的修复

原脚本缺少断点续存的读取逻辑和正确的文件写入模式。先读取目标文件已有的条目数,从该数值开始计数,每次生成有效组合后立即追加写入:

def get_existing_count(file_path):
    try:
        with open(file_path, 'r', encoding='utf-16') as f:
            return sum(1 for _ in f)
    except FileNotFoundError:
        return 0

# 写入逻辑示例
with open(output_file, 'a', encoding='utf-16') as f:
    f.write(f"{combination}\n")

使用'a'模式确保追加写入,utf-16编码匹配脚本需求。

3. 无限循环的修复

递归函数未设置终止条件,导致无法停止。维护一个计数器,当生成的组合数达到1000时立即终止所有递归:

def generate_all_utf16_combinations(base_word, special_chars, current_count, max_count, output_file):
    # 先检查是否已达上限,终止递归
    if current_count >= max_count:
        return current_count
    
    # 生成基础单词全排列
    for perm in permutations(base_word):
        perm_str = ''.join(perm)
        # 写入基础排列
        with open(output_file, 'a', encoding='utf-16') as f:
            f.write(f"{perm_str}\n")
        current_count += 1
        if current_count >= max_count:
            return current_count
        
        # 插入特殊字符及递归逻辑
        for i in range(len(perm_str) + 1):
            for char in special_chars:
                if is_asian_char(char):
                    continue
                new_str = perm_str[:i] + char + perm_str[i:]
                with open(output_file, 'a', encoding='utf-16') as f:
                    f.write(f"{new_str}\n")
                current_count += 1
                if current_count >= max_count:
                    return current_count
                # 递归插入更多特殊字符
                current_count = generate_all_utf16_combinations(new_str, special_chars, current_count, max_count, output_file)
                if current_count >= max_count:
                    return current_count
    return current_count

每次递归调用前检查计数器,达到上限后立即返回,终止后续递归。

完整修复后的脚本

from itertools import permutations

def is_asian_char(char):
    code = ord(char)
    return (
        (0x4E00 <= code <= 0x9FFF) or
        (0x3400 <= code <= 0x4DBF) or
        (0x20000 <= code <= 0x2A6DF) or
        (0xF900 <= code <= 0xFAFF) or
        (0x2F800 <= code <= 0x2FA1F)
    )

def get_existing_count(file_path):
    try:
        with open(file_path, 'r', encoding='utf-16') as f:
            return sum(1 for _ in f)
    except FileNotFoundError:
        return 0

def generate_all_utf16_combinations(base_word, special_chars, current_count, max_count, output_file):
    if current_count >= max_count:
        return current_count
    
    # 生成基础单词的全排列
    for perm in permutations(base_word):
        perm_str = ''.join(perm)
        # 写入基础排列
        with open(output_file, 'a', encoding='utf-16') as f:
            f.write(f"{perm_str}\n")
        current_count += 1
        if current_count >= max_count:
            return current_count
        
        # 插入单个特殊字符
        for i in range(len(perm_str) + 1):
            for char in special_chars:
                if is_asian_char(char):
                    continue
                new_str = perm_str[:i] + char + perm_str[i:]
                with open(output_file, 'a', encoding='utf-16') as f:
                    f.write(f"{new_str}\n")
                current_count += 1
                if current_count >= max_count:
                    return current_count
                # 递归插入更多特殊字符
                current_count = generate_all_utf16_combinations(new_str, special_chars, current_count, max_count, output_file)
                if current_count >= max_count:
                    return current_count
    return current_count

# 主逻辑
if __name__ == "__main__":
    target_word = "banana"
    # 示例特殊字符集合(可按需扩展)
    special_chars = "!@#$%^&*()_+-=[]{}|;:'\",.<>?/~`"
    output_file = "bad_words.txt"
    max_results = 1000
    
    existing_count = get_existing_count(output_file)
    if existing_count >= max_results:
        print(f"已达到最大生成数量{max_results},无需继续生成")
    else:
        generate_all_utf16_combinations(target_word, special_chars, existing_count, max_results, output_file)
        print(f"生成完成,共写入{max_results - existing_count}条新内容,总条目数:{get_existing_count(output_file)}")

内容的提问来源于stack exchange,提问作者jose vazquez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 00:13:36