如何在Pandas DataFrame中基于指定短语提取文本中的温度数值
Pandas 文本列提取温度数值实现方案
核心用Pandas内置的str.extract方法配合正则表达式实现,无需自定义循环,是Pandas场景下最简洁的写法:
1. 提取所有匹配的温度数值(不分最高/最低)
如果不需要区分high near对应高温、low near对应低温,只需要提取所有匹配的温度值,直接用如下代码:
# 假设存储天气文本的列名为 weather,提取结果存入 temperature 列 df['temperature'] = df['weather'].str.extract(r'(?:high|low) near (\d+)', expand=False).astype('Int64')
2. 分开提取最高温和最低温
如果需要分别存储高温、低温的数值,用如下代码:
# 提取最高温 df['high_temp'] = df['weather'].str.extract(r'high near (\d+)', expand=False).astype('Int64') # 提取最低温 df['low_temp'] = df['weather'].str.extract(r'low near (\d+)', expand=False).astype('Int64')
代码说明
- 正则表达式中
(?:high|low)是非捕获组,只做匹配不提取,后面的(\d+)是捕获组,会提取high near或者low near后面连续的数字 expand=False表示返回Series而非DataFrame,直接赋值给新列更方便- 用
Int64类型(注意首字母大写)是为了兼容没有匹配到数值的空值场景,避免数值被强制转成float类型
测试示例
你可以用如下测试代码验证效果:
import pandas as pd # 构造测试数据 data = { 'weather': [ 'Sunny, with a high near 82. Light and variable wind becoming northwest 5 to 7 mph in the afternoon.', 'A 50 percent chance of showers. Partly sunny, with a high near 61.', 'Clear, with a low near 48. Northeast wind around 3 mph.', 'Rainy, high near 59, low near 42.' ] } df = pd.DataFrame(data) # 提取温度 df['high_temp'] = df['weather'].str.extract(r'high near (\d+)', expand=False).astype('Int64') df['low_temp'] = df['weather'].str.extract(r'low near (\d+)', expand=False).astype('Int64') print(df)
运行后输出结果如下:
weather high_temp low_temp 0 Sunny, with a high near 82. Light and variable... 82 <NA> 1 A 50 percent chance of showers. Partly sunny,... 61 <NA> 2 Clear, with a low near 48. Northeast wind arou... <NA> 48 3 Rainy, high near 59, low near 42. 59 42
内容的提问来源于stack exchange,提问作者Guayule
相关产品推荐
相关产品推荐

