域名生成器出现重复变量域名问题求助
问题原因
你用了random.choices()方法,这是有放回抽样,会导致同一个词被多次选中,所以出现像https://web.login-group.claim.claim.com里claim重复的情况。
修复方案
把random.choices()换成random.sample()(无放回抽样),同时处理源数据重复、抽样长度超限等问题,确保每个域名里的变量仅出现一次。修改后的代码如下:
import random with open("dictionaries/Combined.txt") as i: Test = [line.rstrip() for line in i] # 先对源数据去重,避免本身存在重复词 Test = list(set(Test)) delimiters = ['', '-', '.'] web = 'web' HTTP = ['http://', 'https://'] suffix = ['odoo', 'info', 'com'] output = [] # 限制子域名数量不超过源数据长度,防止抽样报错 max_subdomain = min(4, len(Test)) valid_subdomain_counts = [count for count in [2,3,4] if count <= max_subdomain] for i in range(100): for subdomain_count in valid_subdomain_counts: http = random.choice(HTTP) # 无放回抽样,保证选中的子域名词不重复 selected = random.sample(Test, k=subdomain_count) # 确保web不会和选中的词重复(如果源数据里有web) if web in selected: selected = random.sample(Test, k=subdomain_count) data = [web] + selected random.shuffle(data) delims = (random.choices(delimiters, k=len(data)-1) + ['.' + random.choice(suffix)]) address = ''.join([a+b for a, b in zip(data, delims)]) webs = http + address output.append(webs) for o in output: print(o)
关键修改点
- 替换抽样方法:用
random.sample()替代random.choices(),实现无放回抽样,避免重复选词 - 源数据去重:将
Test转成集合再转回列表,清除源文件里的重复词 - 限制抽样长度:确保子域名数量不超过源数据的有效长度,防止报错
- 规避web重复:增加判断,保证
web不会和选中的子域名词重复
内容的提问来源于stack exchange,提问作者Tsukuru
相关产品推荐
相关产品推荐

