You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Pandas为csv每行生成随机字母数字串时全列值重复

问题修复方案

错误原因

循环内file['short'] = result_str的写法是对整个short列全量赋值,每一次循环都会覆盖所有行的取值,最终所有行都会保留最后一次循环生成的随机值。

修复代码

方案1:修改原有循环逻辑

仅给当前遍历到的行赋值即可:

import random
import string
import pandas as pd

file = pd.read_csv('url short.csv')
letters = string.ascii_lowercase + string.digits
# 提前初始化列
file['short'] = ''

for i in range(len(file)):
    result_str = ''.join(random.choice(letters) for _ in range(4))
    print(result_str)
    # 仅给索引为i的行赋值
    file.loc[i, 'short'] = result_str
    
print(file)

方案2:更高效的无循环写法(自动保证唯一性)

如果要严格保证随机值不重复,推荐用集合去重的逻辑实现:

import random
import string
import pandas as pd

file = pd.read_csv('url short.csv')
letters = string.ascii_lowercase + string.digits
code_length = 4
required_count = len(file)

# 生成足量不重复的4位随机串
unique_codes = set()
while len(unique_codes) < required_count:
    cur_code = ''.join(random.choice(letters) for _ in range(code_length))
    unique_codes.add(cur_code)

# 批量赋值到列
file['short'] = list(unique_codes)
print(file)

注意事项

4位小写字母+数字的组合总共有36⁴=1679616种可能,如果你的CSV文件行数超过该值,将无法生成符合要求的唯一值,需要加长随机串的长度。

内容的提问来源于stack exchange,提问作者Rohit Chabukswar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 02:36:03