如何在Bash环境下从.mhtml文件提取指定数字1,621.41?
解决方案
方法1:直接用Xidel解析MHTML文件
Xidel原生支持解析MHTML格式,无需额外转换,只要针对目标数字的HTML节点特征编写选择器即可:
# 若已知数字所在元素的class/id,直接定位提取 xidel your_file.mhtml -e '//span[@class="target-amount"]/text()' # 若不确定元素结构,用正则匹配提取带逗号和小数点的数字格式 xidel your_file.mhtml -e '//text()' | grep -oP '\d{1,3}(,\d{3})*\.\d{2}'
方法2:先转MHTML为HTML再处理
用pandoc将MHTML转换为标准HTML,再用熟悉的工具提取:
# 转换MHTML到HTML pandoc -s your_file.mhtml -o converted.html # 用Xidel结合正则提取 xidel converted.html -e '//*[contains(text(),"金额")]/text()' | grep -oP '\d{1,3}(,\d{3})*\.\d{2}' # 或用sed直接匹配提取 sed -n 's/.*\(\d\{1,3\},\d\{3\}\.\d\{2\}\).*/\1/p' converted.html
方法3:Python脚本精准解析
针对复杂HTML结构,用Python结合解析库处理:
- 安装依赖:
pip install beautifulsoup4 mhtml
- 编写提取脚本(extract_amount.py):
from mhtml import parse_mhtml from bs4 import BeautifulSoup import re # 读取本地MHTML文件 with open('your_file.mhtml', 'rb') as f: mhtml_data = f.read() # 解析MHTML为HTML内容 parsed_pages = parse_mhtml(mhtml_data) html_content = parsed_pages[0].content.decode('utf-8') # 解析HTML并定位目标数字 soup = BeautifulSoup(html_content, 'html.parser') # 示例:查找包含"总计"文本的元素,提取其相邻节点的数字 total_label = soup.find(text=re.compile('总计')) if total_label: amount_text = total_label.find_next('div').text match = re.search(r'\d{1,3}(,\d{3})*\.\d{2}', amount_text) if match: print(match.group())
- 运行脚本:
python extract_amount.py
内容的提问来源于stack exchange,提问作者foopeen
相关产品推荐
相关产品推荐

