如何向pandas DataFrame指定列子集增量添加列表值?循环赋值报长度不匹配错误
问题描述
背景
我可以通过以下代码将天气API(meteostat)返回的一行值赋值给dataframe的指定列子集:
Morel_df = pd.read_csv('C:/git/NOAA Weather/Morel_TreeSpecies_TEST_reformat_short.csv') Test_Morel_df = Morel_df x = morelParser('C:/git/NOAA Weather/Morel_TreeSpecies_TEST_reformat_short.csv') location_vector = x.__next__() point = (location_vector[0], location_vector[1], location_vector[2]) date = datetime.strptime(Test_Morel_df.loc[i, 'Date'], "%m/%d/%Y %H:%M") temp_weather_df = callMeteostat(date, point) temp_weather_list = temp_weather_df.values.tolist() flattened_weather_list = [element for innerlist in temp_weather_list for element in innerlist] i = 0 print(len(Test_Morel_df.loc[i, 'tavg_Before_30':'tsun_After_29'])) print(len(flattened_weather_list)) Test_Morel_df.loc[i, 'tavg_Before_30':'tsun_After_29'] = flattened_weather_list Test_Morel_df.to_csv('C:/git/NOAA Weather/TEST_MOREL_WEATHER_SHORT.csv', index = False)
为了遍历主dataframe的每一行,编写了如下for循环复用逻辑:
Morel_df = pd.read_csv('C:/git/NOAA Weather/Morel_TreeSpecies_TEST_reformat_short.csv') Test_Morel_df = Morel_df x = morelParser('C:/git/NOAA Weather/Morel_TreeSpecies_TEST_reformat_short.csv') for i in range(len(Test_Morel_df)): location_vector = x.__next__() point = (location_vector[0], location_vector[1], location_vector[2]) date = datetime.strptime(Test_Morel_df.loc[i, 'Date'], "%m/%d/%Y %H:%M") temp_weather_df = callMeteostat(date, point) temp_weather_list = temp_weather_df.values.tolist() flattened_weather_list = [element for innerlist in temp_weather_list for element in innerlist] print(len(Test_Morel_df.loc[i, 'tavg_Before_30':'tsun_After_29'])) print(len(flattened_weather_list)) Test_Morel_df.loc[i, 'tavg_Before_30':'tsun_After_29'] = flattened_weather_list Test_Morel_df.to_csv('C:/git/NOAA Weather/TEST_MOREL_WEATHER_SHORT.csv', index = False)
运行后触发报错:
ValueError('Must have equal len keys and value ')
经打印排查确认错误原因:for循环中得到的flattened_weather_list长度比单步运行时少10个值。
相关函数代码
def callMeteostat(input_date, input_location): observation_dates = [] observation_dates = list(dateRange(input_date)) observation_location = convertLocation(input_location[0], input_location[1], input_location[2]) weather_data_call_before = Daily(observation_location, observation_dates[0], observation_dates[1]) weather_data_before = weather_data_call_before.fetch() return weather_data_before def morelParser(morel_data_file_path): Morel_df = pd.read_csv(morel_data_file_path) for i in range(len(Morel_df)): input_location_list = (Morel_df.loc[i, 'DDLat'], Morel_df.loc[i, 'DDLon'], Morel_df.loc[i, 'Elevation_meters'], i) yield input_location_list return False def dateRange(input_date): before_date = input_date - dt.timedelta(days = 30) after_date = input_date + dt.timedelta(days = 30) return before_date, after_date def convertLocation(input_lat, input_lon, input_elev): input_location = Point(input_lat, input_lon, input_elev) return input_location
完整报错回溯
ValueError Traceback (most recent call last) c:\git\NOAA Weather\weather.py in <module> 266 print(len(Test_Morel_df.loc[0, 'tavg_Before_30':'tsun_After_29'])) 267 print(len(flat_list)) --> 268 Test_Morel_df.loc[1, 'tavg_Before_30':'tsun_After_29'] = flat_list 269 270 Test_Morel_df.to_csv('C:/git/NOAA Weather/TEST_MOREL_WEATHER_SHORT.csv', index = False) ~\miniconda3\envs\HistoricalWeatherData\lib\site packages\pandas\core\indexing.py in __setitem__(self, key, value) 721 722 iloc = self if self.name == "iloc" else self.obj.iloc --> 723 iloc._setitem_with_indexer(indexer, value, self.name) 724 725 def _validate_key(self, key, axis: int): ~\miniconda3\envs\HistoricalWeatherData\lib\site-packages\pandas\core\indexing.py in _setitem_with_indexer(self, indexer, value, name) 1728 if take_split_path: 1729 # We have to operate column-wise -> 1730 self._setitem_with_indexer_split_path(indexer, value, name) 1731 else: 1732 self._setitem_single_block(indexer, value, name) ~\miniconda3\envs\HistoricalWeatherData\lib\site-packages\pandas\core\indexing.py in _setitem_with_indexer_split_path(self, indexer, value, name) 1806 1807 else: -> 1808 raise ValueError( 1809 "Must have equal len keys and value " 1810 "when setting with an iterable" ValueError: Must have equal len keys and value when setting with an iterable
解决方案
错误原因
- meteostat的
Daily接口不会返回无观测数据的日期行,单步测试用的第一条数据对应地点的61天时间范围刚好没有数据缺失,长度匹配;循环到其他行时,部分日期无观测数据,返回的dataframe行数不足,扁平化后的列表长度自然小于目标列的长度,触发赋值错误。 - 现有代码存在冗余问题:用
morelParser生成器二次读取CSV文件,如果两次读取的文件内容不一致会导致坐标和行错位,另外直接用Test_Morel_df = Morel_df是浅拷贝,修改测试表会同步修改原始表。
修复代码
- 首先修改
callMeteostat函数,补全缺失的日期行,确保返回数据长度固定:
def callMeteostat(input_date, input_location): start_date, end_date = dateRange(input_date) # 生成完整的61天日期序列 full_date_range = pd.date_range(start=start_date, end=end_date, freq='D') observation_location = convertLocation(input_location[0], input_location[1], input_location[2]) weather_data_call = Daily(observation_location, start_date, end_date) weather_data = weather_data_call.fetch() # 重新索引到完整日期范围,缺失值默认填充NaN,可加fill_value参数自定义填充值 weather_data = weather_data.reindex(full_date_range) return weather_data
- 优化主循环逻辑,移除冗余的生成器调用,修改浅拷贝问题:
import pandas as pd import datetime as dt from meteostat import Daily, Point Morel_df = pd.read_csv('C:/git/NOAA Weather/Morel_TreeSpecies_TEST_reformat_short.csv') Test_Morel_df = Morel_df.copy() # 改为深拷贝,避免修改原始表 for i in range(len(Test_Morel_df)): # 直接从当前行取坐标,无需二次读文件 lat = Test_Morel_df.loc[i, 'DDLat'] lon = Test_Morel_df.loc[i, 'DDLon'] elev = Test_Morel_df.loc[i, 'Elevation_meters'] point = (lat, lon, elev) date = dt.datetime.strptime(Test_Morel_df.loc[i, 'Date'], "%m/%d/%Y %H:%M") temp_weather_df = callMeteostat(date, point) # 用内置flatten方法简化代码 flattened_weather_list = temp_weather_df.values.flatten().tolist() # 修复后两者长度完全一致 print(f"行{i}:目标列长度{len(Test_Morel_df.loc[i, 'tavg_Before_30':'tsun_After_29'])},返回数据长度{len(flattened_weather_list)}") Test_Morel_df.loc[i, 'tavg_Before_30':'tsun_After_29'] = flattened_weather_list Test_Morel_df.to_csv('C:/git/NOAA Weather/TEST_MOREL_WEATHER_SHORT.csv', index = False)
内容的提问来源于stack exchange,提问作者Jon。
相关产品推荐
相关产品推荐

