将DataFrame中标签添加至Set对象时结果异常,如何获取不重复标签
解决DataFrame列标签存入集合时出现单个字符的问题
问题场景
想要把DataFrame中practice列的完整标签去重后存入Set对象,执行代码后却得到了由单个字符组成的集合,不符合预期。
错误代码及结果
types = set() for t in frame4['practice']: types.update(t) types
返回结果:
{'1', '3', 'A', 'B', 'C', 'D', 'E', 'F', 'G', 'I', 'L', 'M', 'N', 'O', 'P', 'S', 'T', 'W', 'Z', '_', 'a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'k', 'l', 'm', 'n', 'o', 'p', 'r', 's', 't', 'u', 'v', 'w', 'y'}
DataFrame的practice列示例(已移除NaN,存在重复标签)
2 Identifier_Cookie_or_similar_Tech_1stParty 3 Identifier_IP_Address_1stParty 4 Identifier_Cookie_or_similar_Tech_1stParty 8 Identifier_Cookie_or_similar_Tech_3rdParty 10 Demographic_3rdParty ... 21612 Demographic_1stParty 21613 Demographic_3rdParty 21614 Identifier_Cookie_or_similar_Tech_1stParty 21615 Identifier_Cookie_or_similar_Tech_3rdParty 21616 Identifier_Cookie_or_similar_Tech_1stParty Name: practice, Length: 10201, dtype: object
错误原因
set.update()方法会将传入的可迭代对象拆解为单个元素添加到集合中。这里每个t是字符串类型,而字符串属于字符的可迭代对象,因此update(t)会把标签字符串拆成单个字符逐个加入集合,最终得到的就是所有字符的集合,而非完整标签的集合。
正确解法
方法1:使用set.add()替代update()
add()方法用于向集合添加单个元素,不会拆解传入的对象。修改代码如下:
types = set() for t in frame4['practice']: types.add(t)
方法2:直接将列转换为集合(更简洁)
利用Python和Pandas的特性,直接将practice列转为集合,自动去重:
# 方式一:直接转集合 types = set(frame4['practice']) # 方式二:先通过Pandas的unique()获取唯一值再转集合 types = set(frame4['practice'].unique())
以上两种方法都能得到预期的无重复完整标签集合,例如:
{'Identifier_Cookie_or_similar_Tech_1stParty', 'Identifier_IP_Address_1stParty', 'Identifier_Cookie_or_similar_Tech_3rdParty', 'Demographic_3rdParty', 'Demographic_1stParty'}
内容的提问来源于stack exchange,提问作者Angelos Zinonos
相关产品推荐
相关产品推荐

