You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python在CSV文件中查找属性?附示例文件内容

使用Python在CSV文件中查找属性的方案

我来帮你梳理下如何用Python在attributes.csv和db.csv这两个文件里实现属性查找的功能,这里提供两种常用方案:用Python内置的csv模块(无需额外安装依赖),或者用更高效便捷的pandas库,你可以根据数据规模和需求来选择。

先明确文件结构

首先得确认两个CSV的格式:

  • attributes.csv:第一行是属性表头(Ethnics, HairColor, EyeColor, HairStyle),后续每行是一组合法的属性组合
  • db.csv:第一行是用户信息表头(Name, Id, Ethnics, HairColor, HairStyle, EyeColor),后续每行是一个用户的详细数据

两个文件都是空格分隔的CSV,所以读取时要指定分隔符为空格。


方法一:使用Python内置csv模块

适合小数据量场景,无需额外安装库。

1. 读取并解析CSV文件

先把两个文件的内容读取到内存,整理成方便操作的格式:

import csv

# 读取attributes.csv,收集每个属性的合法值集合
attributes = {}
with open('attributes.csv', 'r', encoding='utf-8') as f:
    # 指定空格分隔,跳过初始空格避免解析错误
    reader = csv.DictReader(f, delimiter=' ', skipinitialspace=True)
    # 初始化每个属性的空集合
    for attr in reader.fieldnames:
        attributes[attr] = set()
    # 遍历每行,收集属性值
    for row in reader:
        for attr in reader.fieldnames:
            attributes[attr].add(row[attr])

# 读取db.csv,存储用户记录
user_records = []
with open('db.csv', 'r', encoding='utf-8') as f:
    reader = csv.DictReader(f, delimiter=' ', skipinitialspace=True)
    for row in reader:
        # 将Id转为整数(可选,方便后续数值操作)
        row['Id'] = int(row['Id'])
        user_records.append(row)

2. 实现常见查找需求

需求1:查找符合单个属性值的用户

比如找出所有Ethnics为Asian的用户:

def find_users_by_attr(records, attr_name, target_value):
    """根据单个属性值筛选用户"""
    return [user for user in records if user.get(attr_name) == target_value]

# 示例:查找亚洲用户
asian_users = find_users_by_attr(user_records, "Ethnics", "Asian")
print("亚洲用户列表:")
for user in asian_users:
    print(f"姓名:{user['Name']},ID:{user['Id']}")

需求2:验证用户属性是否合法

检查某个用户的属性是否在attributes.csv的合法值范围内:

def is_attr_valid(attr_name, value, valid_attrs):
    """验证属性值是否合法"""
    if attr_name not in valid_attrs:
        return False
    return value in valid_attrs[attr_name]

# 示例:检查Ana的HairColor是否合法
ana = next(user for user in user_records if user['Name'] == 'Ana')
print(f"\nAna的HairColor是否合法?{is_attr_valid('HairColor', ana['HairColor'], attributes)}")

需求3:查找符合多属性条件的用户

比如找出Ethnics为European且EyeColor为Green的用户:

def find_users_by_multi_attr(records, conditions):
    """根据多个属性条件筛选用户,conditions是属性名:目标值的字典"""
    matched = []
    for user in records:
        # 检查所有条件是否都满足
        if all(user.get(attr) == value for attr, value in conditions.items()):
            matched.append(user)
    return matched

# 示例:查找欧洲且眼睛为绿色的用户
target_users = find_users_by_multi_attr(user_records, {"Ethnics": "European", "EyeColor": "Green"})
print("\n欧洲且眼睛为绿色的用户:")
for user in target_users:
    print(f"姓名:{user['Name']},ID:{user['Id']}")

方法二:使用pandas库(推荐用于大数据量)

pandas是Python的数据分析利器,处理CSV文件更简洁高效,适合数据量较大的场景。首先需要安装:

pip install pandas

1. 读取CSV文件

import pandas as pd

# 读取attributes.csv,sep='\s+'匹配任意数量的空格
attrs_df = pd.read_csv('attributes.csv', sep='\s+')
# 收集每个属性的合法值集合
valid_attributes = {col: set(attrs_df[col].unique()) for col in attrs_df.columns}

# 读取db.csv
db_df = pd.read_csv('db.csv', sep='\s+')
# 将Id列转为整数类型
db_df['Id'] = db_df['Id'].astype(int)

2. 实现查找需求

需求1:单个属性筛选

# 查找Ethnics为Asian的用户
asian_users = db_df[db_df['Ethnics'] == 'Asian']
print("亚洲用户:")
print(asian_users[['Name', 'Id']])

需求2:属性合法性验证

# 给用户数据新增一列,标记HairColor是否合法
db_df['HairColor_Valid'] = db_df['HairColor'].isin(valid_attributes['HairColor'])
print("\n用户HairColor合法性检查:")
print(db_df[['Name', 'HairColor', 'HairColor_Valid']])

需求3:多属性条件筛选

# 查找欧洲且眼睛为绿色的用户,注意用&连接条件,每个条件加括号
target_users = db_df[(db_df['Ethnics'] == 'European') & (db_df['EyeColor'] == 'Green')]
print("\n欧洲且眼睛为绿色的用户:")
print(target_users[['Name', 'Id', 'EyeColor']])

内容的提问来源于stack exchange,提问作者jack

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:55:05