Python中Max()函数获取XML数据返回错误值的排查求助
问题排查:2019年4月食物分发数据最大值计算错误
问题描述
需求为找出2019年4月期间,分发食物磅数最多、服务个人数量最多、服务家庭数量最多的地区,但运行代码后,max函数返回的服务个人数和家庭数均为999,与XML中的实际更高数值不符。原代码如下:
import requests from lxml import objectify URL = 'https://data.virginia.gov/api/views/xvir-sctz/rows.xml?accessType=DOWNLOAD' response = requests.get(URL).content import requests from lxml import objectify #parse xml file root = objectify.fromstring(response) #print data print(response) #create emtpy lists to store wanted data localities = [] pounds_of_food_distributed = [] individuals_served = [] households_served = [] #assign data values to varibles localities = root.xpath('//response/row/row/locality/text()') pounds_of_food_distributed = root.xpath('//response/row/row/pounds_of_food_distributed/text()') individuals_served = root.xpath('//response/row/row/individuals_served/text()') households_served = root.xpath('//response/row/row/households_served/text()') #for loop to iterate through data set for row in root.iterchildren(): if row.get("month") == ("April") and row.get("Year") == "2019": localities.append(row.get("Locality")) pounds_of_food_distributed.append(row.get("Pounds")) individuals_served.append(row.get("Individuals")) households_served.append(row.get("Households")) #Indexing to retrieve max value of each requested variable max_pounds_index = pounds_of_food_distributed.index(max(pounds_of_food_distributed)) max_individuals_index = individuals_served.index(max(individuals_served)) max_households_index = households_served.index(max(households_served)) #print values print("Locality with greatest distribution of food in pounds: ", localities[max_pounds_index]) print("Greatest number of individuals served: ", individuals_served[max_individuals_index]) print("Greatest number of households served: ", households_served[max_households_index])
运行结果:
Locality with greatest distribution of food in pounds: Colonial Heights City Greatest number of individuals served: 999.0 Greatest number of households served: 999.0
错误原因分析
- 重复导入模块:代码冗余导入
requests和lxml.objectify,虽不影响结果但不符合规范。 - 全量数据与筛选数据混合:先通过XPath加载所有时间段的全量数据到列表,后续又尝试追加筛选数据,导致列表包含无关数据,干扰最大值计算。
- 节点遍历错误:
root.iterchildren()遍历的是根节点的直接子节点(外层<row>),但实际数据存储在**内层<row>**节点中,导致筛选条件失效,追加的筛选数据为空。 - 数据类型不匹配:XPath获取的
text()结果是字符串类型,max()函数按字符串字典序比较(如"999"的ASCII码大于"1000"的首字符"1"),而非数值大小,导致错误的最大值。 - 属性名不匹配:循环中使用的
"Locality"、"Pounds"等属性名,与XML实际的locality、pounds_of_food_distributed属性名不一致,无法获取正确数据。
修正后的代码
import requests from lxml import objectify URL = 'https://data.virginia.gov/api/views/xvir-sctz/rows.xml?accessType=DOWNLOAD' response = requests.get(URL).content # 解析XML root = objectify.fromstring(response) # 初始化空列表,仅存储2019年4月的数据 localities = [] pounds = [] individuals = [] households = [] # 遍历所有内层数据行 for row in root.xpath('//response/row/row'): # 匹配2019年4月的数据(严格对应XML属性名) if row.get('month') == 'April' and row.get('year') == '2019': localities.append(row.get('locality')) # 转换数值类型并处理异常值 pounds_val = row.get('pounds_of_food_distributed') pounds.append(int(pounds_val) if pounds_val.isdigit() else 0) indiv_val = row.get('individuals_served') individuals.append(int(indiv_val) if indiv_val.isdigit() else 0) house_val = row.get('households_served') households.append(int(house_val) if house_val.isdigit() else 0) # 计算最大值索引 max_pounds_idx = pounds.index(max(pounds)) max_indiv_idx = individuals.index(max(individuals)) max_house_idx = households.index(max(households)) # 输出结果 print(f"2019年4月分发食物磅数最多的地区: {localities[max_pounds_idx]},磅数: {pounds[max_pounds_idx]}") print(f"2019年4月服务个人数量最多的地区: {localities[max_indiv_idx]},人数: {individuals[max_indiv_idx]}") print(f"2019年4月服务家庭数量最多的地区: {localities[max_house_idx]},家庭数: {households[max_house_idx]}")
关键修正点说明
- 移除重复导入,精简代码结构。
- 直接遍历内层
<row>节点,确保正确获取数据属性。 - 严格匹配XML中的属性名(如
month、year),避免数据获取失败。 - 将数值型数据转换为整数,确保
max()函数按数值大小比较。 - 添加简单异常处理,避免非数值数据导致程序报错。
内容的提问来源于stack exchange,提问作者Frost4820
相关产品推荐
相关产品推荐

