使用Pandas读取CSV时,如何在迭代中按列名访问数据?
Pandas读取CSV后按列名迭代访问的解决方法
问题场景
读取CSV文件后,希望在循环中通过列名(如product_url)访问每行数据,尝试了以下代码但无法实现预期效果:
import pandas as pd products = pd.read_csv(f'./storage/temp/{file_name}', delimiter=',', on_bad_lines='skip').to_dict() for index, product in products: print(product.product_url)
CSV示例数据:
product_url,image_url,categories,attributes,description,product_name,reviews,stock,sku,product_gift,price,price_with_discount,categories_urls http://localhost:8000/?product=aw-bellies,http://localhost:8000/wp-content/uploads/2022/10/red-as-454-aw-11-original-imaeebfwsdf6jdf6-324x324.jpeg,Footwear,"""[{""""Ideal For"""":""""Women""""}","{""""Occasion"""":""""Casual""""}","{""""Color"""":""""Red""""}]""""","""Key Features of AW Bellies Sandals Wedges Heel Casuals",AW Bellies Price: Rs. 499 Material: Synthetic Lifestyle: Casual Heel Type: Wedge Warranty Type: Manufacturer Product Warranty against manufacturing defects: 30 days Care instructions: Allow your pair of shoes to air and de-odorize at regular basis; use shoe bags to prevent any stains or mildew; dust any dry dirt from the surface using a clean cloth; do not use polish or shiner,"Specifications of AW Bellies General Ideal For Women Occasion Casual Shoe Details Color Red Outer Material Patent Leather Heel Height 1 inch Number of Contents in Sales Package Pack of 1 In the Box One Pair Of Shoes""""","""AW Bellies""""",,1,2,,499,499,http://localhost:8000/?product_cat=footwear
期望的访问方式:
for product in products: print(product.product_url)
错误原因
to_dict()默认参数为'dict',会将DataFrame转换为以列名为键、对应列数据为值的字典,并非按行存储的结构,因此循环时无法拿到每行数据。- 即使转换为行字典,普通字典也不能用
.语法访问键,需使用[]语法。
解决方案
提供三种可行实现方式,按需选择:
方法1:使用itertuples()(支持.语法访问列名)
itertuples()返回每行数据的命名元组,可直接通过列名作为属性访问:
import pandas as pd df = pd.read_csv(f'./storage/temp/{file_name}', delimiter=',', on_bad_lines='skip') for product in df.itertuples(index=False): print(product.product_url)
index=False表示不返回行索引,仅返回行数据。
方法2:转换为行字典列表(用[]访问键)
用to_dict('records')将DataFrame转换为每行对应一个字典的列表,循环时通过键名访问:
import pandas as pd products = pd.read_csv(f'./storage/temp/{file_name}', delimiter=',', on_bad_lines='skip').to_dict('records') for product in products: print(product['product_url'])
方法3:使用iterrows()(返回索引和行Series)
iterrows()返回每行的索引和Series对象,Series可用.或[]访问列名:
import pandas as pd df = pd.read_csv(f'./storage/temp/{file_name}', delimiter=',', on_bad_lines='skip') for index, product in df.iterrows(): print(product.product_url) # 或 product['product_url']
内容的提问来源于stack exchange,提问作者DaniD
相关产品推荐
相关产品推荐

