如何从pandas read_csv读取的结果中提取指定行或字符串值
操作方法
1. 读取CSV的注意事项
你给出的CSV示例是用空格分隔两列,默认pd.read_csv用逗号作为分隔符会识别错误,建议读取时指定分隔符参数:
import pandas as pd # \s+ 代表匹配任意数量的空格作为分隔符 news_headlines = pd.read_csv('/content/sample_data/crypto_headlines.csv', sep='\s+')
2. 提取指定内容
你需要的发布日期为20130511的第二行内容,按需求选择对应代码即可:
提取整行数据
# 注意pandas索引从0开始计数,第二行对应iloc[1] # 如果publishdate字段是字符串类型,把20130511改为'20130511' target_row = news_headlines[news_headlines['publishdate'] == 20130511].iloc[1]
仅提取标题字符串
target_headline = news_headlines[news_headlines['publishdate'] == 20130511]['headlinetext'].iloc[1]
3. 批量字符串处理方法
不需要单独提取每个字符串,pandas提供了.str系列接口可以直接对整列字符串做批量处理,正好适配你要做的大小写转换、特殊字符去除需求:
- 全部转为小写:
news_headlines['clean_headline'] = news_headlines['headlinetext'].str.lower()
- 去除特殊字符(仅保留英文字母、数字和空格):
news_headlines['clean_headline'] = news_headlines['clean_headline'].str.replace(r'[^a-zA-Z0-9\s]', '', regex=True)
- 去除首尾多余空格:
news_headlines['clean_headline'] = news_headlines['clean_headline'].str.strip()
内容的提问来源于stack exchange,提问作者izzypt
相关产品推荐
相关产品推荐

