You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从格式化文本提取指定辐照度值?正则与BeautifulSoup遇阻求助

提取辐射数据中的指定数值及代码问题

原始数据

*******************************************************************************
*                                                                             *
*                         int. normalized  values  of  :                      *
*                         ---------------------------                         *
*                      % of irradiance at ground level                        *
*     % of direct  irr.    % of diffuse irr.    % of enviro. irr              *
*               0.488               0.418               0.093                 *
*                       reflectance at satellite level                        *
*     atm. intrin. ref.   background  ref.  pixel  reflectance                *
*               0.127               0.146               0.170                 *
*                                                                             *
*                         int. absolute values of                             *
*                         -----------------------                             *
*                      irr. at ground level (w/m2/mic)                        *
*     direct solar irr.    atm. diffuse irr.    environment  irr              *
*             592.299             507.010             113.283                 *
*                      rad at satel. level (w/m2/sr/mic)                      *
*     atm. intrin. rad.    background  rad.    pixel  radiance                *
*              58.837              67.355              78.685                 *
*                                                                             *
*                                                                             *
*                      sol. spect (in w/m2/mic)                               *
*                                2054.457                                     *
*                                                                             *
*******************************************************************************

需求

提取direct solar irr.、atm. diffuse irr.和environment irr.对应的数值。

第一次尝试:正则表达式方法

尝试用正则匹配,但返回None:

import re

def extract_values(text):
    pattern = r"direct solar irr\.\s*atm. diffuse irr\.\s*environment irr\s*([\d\.]+)\s*([\d\.]+)\s*([\d\.]+)"
    match = re.search(pattern, text)
    if match:
       return {
            "direct solar irr.": match.group(1),
            "atm. diffuse irr.": match.group(2),
            "environment irr.": match.group(3)
        }

问题原因:正则没考虑文本换行和冗余空格,且标签与数值不在同一行,同时原数据中environment和irr.之间有两个空格,导致匹配失败。

第二次尝试:BeautifulSoup方法

运行时抛出ValueError: could not convert string to float: '*'错误:

def extract_values(text):
    soup = BeautifulSoup(text, 'html.parser')

    # Get all text elements
    lines = [line.strip() for line in soup.get_text().splitlines() if line.strip() != ""]

    # Identify the line after the "direct solar irr." label
    for i, line in enumerate(lines):
        if "direct solar irr." in line:
            # Look for the next line with a number
            for subsequent_line in lines[i+1:]:
                if re.search(r'\d', subsequent_line):  # Check if the line has a digit
                    values = subsequent_line.split()
                    return {
                        "direct solar irr.": float(values[0]),
                        "atm. diffuse irr.": float(values[1]),
                        "environment irr.": float(values[2])
                    }

问题原因:处理后的行列表包含带*的分隔线行,这类行虽含数字,但分割后会得到*字符,转float时触发错误。

解决方案

修正正则表达式方法

调整正则,允许标签和数值跨多行,匹配任意空白(含换行):

import re

def extract_values(text):
    # 匹配标签行与下一行的数值,支持跨换行和任意空白
    pattern = r"direct solar irr\.\s+atm\. diffuse irr\.\s+environment\s+irr\.\s*\n\s*([\d\.]+)\s+([\d\.]+)\s+([\d\.]+)"
    match = re.search(pattern, text, re.DOTALL)
    if match:
        return {
            "direct solar irr.": float(match.group(1)),
            "atm. diffuse irr.": float(match.group(2)),
            "environment irr.": float(match.group(3))
        }

修正BeautifulSoup方法

过滤带*的行,只处理纯数值行:

from bs4 import BeautifulSoup
import re

def extract_values(text):
    soup = BeautifulSoup(text, 'html.parser')
    lines = [line.strip() for line in soup.get_text().splitlines() if line.strip() != ""]

    for i, line in enumerate(lines):
        if "direct solar irr." in line:
            for subsequent_line in lines[i+1:]:
                # 仅匹配由三个数值组成的行
                if re.fullmatch(r'\s*[\d\.]+\s+[\d\.]+\s+[\d\.]+\s*', subsequent_line):
                    values = subsequent_line.split()
                    return {
                        "direct solar irr.": float(values[0]),
                        "atm. diffuse irr.": float(values[1]),
                        "environment irr.": float(values[2])
                    }

简洁文本处理方法

无需BeautifulSoup,直接遍历文本行提取:

import re

def extract_values(text):
    lines = [line.strip() for line in text.splitlines() if line.strip()]
    target_idx = None
    for idx, line in enumerate(lines):
        if "direct solar irr." in line:
            target_idx = idx + 1
            break
    if target_idx is not None and target_idx < len(lines):
        # 提取该行所有数字
        nums = re.findall(r'[\d\.]+', lines[target_idx])
        if len(nums) >=3:
            return {
                "direct solar irr.": float(nums[0]),
                "atm. diffuse irr.": float(nums[1]),
                "environment irr.": float(nums[2])
            }

内容的提问来源于stack exchange,提问作者mark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 17:56:03