如何从格式化文本提取指定辐照度值?正则与BeautifulSoup遇阻求助
提取辐射数据中的指定数值及代码问题
原始数据
******************************************************************************* * * * int. normalized values of : * * --------------------------- * * % of irradiance at ground level * * % of direct irr. % of diffuse irr. % of enviro. irr * * 0.488 0.418 0.093 * * reflectance at satellite level * * atm. intrin. ref. background ref. pixel reflectance * * 0.127 0.146 0.170 * * * * int. absolute values of * * ----------------------- * * irr. at ground level (w/m2/mic) * * direct solar irr. atm. diffuse irr. environment irr * * 592.299 507.010 113.283 * * rad at satel. level (w/m2/sr/mic) * * atm. intrin. rad. background rad. pixel radiance * * 58.837 67.355 78.685 * * * * * * sol. spect (in w/m2/mic) * * 2054.457 * * * *******************************************************************************
需求
提取direct solar irr.、atm. diffuse irr.和environment irr.对应的数值。
第一次尝试:正则表达式方法
尝试用正则匹配,但返回None:
import re def extract_values(text): pattern = r"direct solar irr\.\s*atm. diffuse irr\.\s*environment irr\s*([\d\.]+)\s*([\d\.]+)\s*([\d\.]+)" match = re.search(pattern, text) if match: return { "direct solar irr.": match.group(1), "atm. diffuse irr.": match.group(2), "environment irr.": match.group(3) }
问题原因:正则没考虑文本换行和冗余空格,且标签与数值不在同一行,同时原数据中environment和irr.之间有两个空格,导致匹配失败。
第二次尝试:BeautifulSoup方法
运行时抛出ValueError: could not convert string to float: '*'错误:
def extract_values(text): soup = BeautifulSoup(text, 'html.parser') # Get all text elements lines = [line.strip() for line in soup.get_text().splitlines() if line.strip() != ""] # Identify the line after the "direct solar irr." label for i, line in enumerate(lines): if "direct solar irr." in line: # Look for the next line with a number for subsequent_line in lines[i+1:]: if re.search(r'\d', subsequent_line): # Check if the line has a digit values = subsequent_line.split() return { "direct solar irr.": float(values[0]), "atm. diffuse irr.": float(values[1]), "environment irr.": float(values[2]) }
问题原因:处理后的行列表包含带*的分隔线行,这类行虽含数字,但分割后会得到*字符,转float时触发错误。
解决方案
修正正则表达式方法
调整正则,允许标签和数值跨多行,匹配任意空白(含换行):
import re def extract_values(text): # 匹配标签行与下一行的数值,支持跨换行和任意空白 pattern = r"direct solar irr\.\s+atm\. diffuse irr\.\s+environment\s+irr\.\s*\n\s*([\d\.]+)\s+([\d\.]+)\s+([\d\.]+)" match = re.search(pattern, text, re.DOTALL) if match: return { "direct solar irr.": float(match.group(1)), "atm. diffuse irr.": float(match.group(2)), "environment irr.": float(match.group(3)) }
修正BeautifulSoup方法
过滤带*的行,只处理纯数值行:
from bs4 import BeautifulSoup import re def extract_values(text): soup = BeautifulSoup(text, 'html.parser') lines = [line.strip() for line in soup.get_text().splitlines() if line.strip() != ""] for i, line in enumerate(lines): if "direct solar irr." in line: for subsequent_line in lines[i+1:]: # 仅匹配由三个数值组成的行 if re.fullmatch(r'\s*[\d\.]+\s+[\d\.]+\s+[\d\.]+\s*', subsequent_line): values = subsequent_line.split() return { "direct solar irr.": float(values[0]), "atm. diffuse irr.": float(values[1]), "environment irr.": float(values[2]) }
简洁文本处理方法
无需BeautifulSoup,直接遍历文本行提取:
import re def extract_values(text): lines = [line.strip() for line in text.splitlines() if line.strip()] target_idx = None for idx, line in enumerate(lines): if "direct solar irr." in line: target_idx = idx + 1 break if target_idx is not None and target_idx < len(lines): # 提取该行所有数字 nums = re.findall(r'[\d\.]+', lines[target_idx]) if len(nums) >=3: return { "direct solar irr.": float(nums[0]), "atm. diffuse irr.": float(nums[1]), "environment irr.": float(nums[2]) }
内容的提问来源于stack exchange,提问作者mark
相关产品推荐
相关产品推荐

