You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Bash环境下从.mhtml文件提取指定数字1,621.41?

解决方案

方法1:直接用Xidel解析MHTML文件

Xidel原生支持解析MHTML格式,无需额外转换,只要针对目标数字的HTML节点特征编写选择器即可:

# 若已知数字所在元素的class/id,直接定位提取
xidel your_file.mhtml -e '//span[@class="target-amount"]/text()'

# 若不确定元素结构,用正则匹配提取带逗号和小数点的数字格式
xidel your_file.mhtml -e '//text()' | grep -oP '\d{1,3}(,\d{3})*\.\d{2}'

方法2:先转MHTML为HTML再处理

用pandoc将MHTML转换为标准HTML,再用熟悉的工具提取:

# 转换MHTML到HTML
pandoc -s your_file.mhtml -o converted.html

# 用Xidel结合正则提取
xidel converted.html -e '//*[contains(text(),"金额")]/text()' | grep -oP '\d{1,3}(,\d{3})*\.\d{2}'

# 或用sed直接匹配提取
sed -n 's/.*\(\d\{1,3\},\d\{3\}\.\d\{2\}\).*/\1/p' converted.html

方法3:Python脚本精准解析

针对复杂HTML结构,用Python结合解析库处理:

  1. 安装依赖:
pip install beautifulsoup4 mhtml
  1. 编写提取脚本(extract_amount.py):
from mhtml import parse_mhtml
from bs4 import BeautifulSoup
import re

# 读取本地MHTML文件
with open('your_file.mhtml', 'rb') as f:
    mhtml_data = f.read()

# 解析MHTML为HTML内容
parsed_pages = parse_mhtml(mhtml_data)
html_content = parsed_pages[0].content.decode('utf-8')

# 解析HTML并定位目标数字
soup = BeautifulSoup(html_content, 'html.parser')
# 示例:查找包含"总计"文本的元素,提取其相邻节点的数字
total_label = soup.find(text=re.compile('总计'))
if total_label:
    amount_text = total_label.find_next('div').text
    match = re.search(r'\d{1,3}(,\d{3})*\.\d{2}', amount_text)
    if match:
        print(match.group())
  1. 运行脚本:
python extract_amount.py

内容的提问来源于stack exchange,提问作者foopeen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 20:35:24