如何提取JSON中days_unresolved字段的纯数值并更新该字段?
解决方案:提取
days_unresolved字段的纯数值 没问题,这事儿好办!我们可以通过两种实用方法来处理这个字段,把带HTML标签的字符串转换成纯数字:
方法1:正则表达式(简单快速)
如果你的HTML格式比较固定(比如始终是<font>标签包裹数字或<3这类内容),用正则匹配数字是最快的方式。我们只需要在遍历数据的循环里加入提取数字的逻辑即可:
修改后的完整代码
import urllib import json import re # 导入正则模块 url = 'http://www.webiron.com/abuse_feed//?format=json' response = urllib.urlopen(url) data_json = json.loads(response.read()) for i in data_json: i['LogEvent'] = 'Trial' i['EvtLen'] = 213 # 处理days_unresolved字段 raw_str = i['days_unresolved'] # 匹配字符串中的数字(包括<3里的3) num_match = re.search(r'\d+', raw_str) if num_match: # 转成整数类型 i['days_unresolved'] = int(num_match.group()) else: # 处理没有数字的异常情况,可根据需求调整默认值 i['days_unresolved'] = 0 print(json.dumps(data_json, indent=6))
方法2:HTML解析库BeautifulSoup(更稳健)
如果担心未来HTML标签结构变化(比如换成<span>或者其他嵌套标签),用专业的HTML解析库会更可靠。首先需要安装依赖:
pip install beautifulsoup4
修改后的完整代码
import urllib import json from bs4 import BeautifulSoup # 导入BeautifulSoup url = 'http://www.webiron.com/abuse_feed//?format=json' response = urllib.urlopen(url) data_json = json.loads(response.read()) for i in data_json: i['LogEvent'] = 'Trial' i['EvtLen'] = 213 # 处理days_unresolved字段 raw_html = i['days_unresolved'] # 解析HTML soup = BeautifulSoup(raw_html, 'html.parser') # 提取纯文本并去除空格 text_content = soup.get_text(strip=True) # 处理<3这种特殊格式,去掉开头的< if text_content.startswith('<'): text_content = text_content[1:] # 转成整数,异常情况设默认值 try: i['days_unresolved'] = int(text_content) except ValueError: i['days_unresolved'] = 0 print(json.dumps(data_json, indent=6))
修改后的输出示例片段
[ { "incidents_reported": 3, "attacker_ip": "178.137.88.8", "event_time": "2018-05-15 19:30:09.832568-07", "event_emails": [ "hostmaster@kyivstar.net", "abuse@kyivstar.net", "noc@kyivstar.net" ], "entry_type": "report", "EvtLen": 213, "emails_deliverable": "Yes", "LogEvent": "Trial", "event_msg": "Fake Referrer Log SPAM Bot", "days_unresolved": 3 }, { "incidents_reported": 52, "attacker_ip": "221.229.166.171", "event_time": "2018-05-15 19:29:45.039281-07", "event_emails": [ "anti-spam@ns.chinanet.cn.net" ], "entry_type": "report", "EvtLen": 213, "emails_deliverable": "No", "LogEvent": "Trial", "event_msg": "Abusive network connectivity", "days_unresolved": 3 } ]
小提示
- 正则方法适合固定格式的简单场景,代码量少;
- BeautifulSoup方法更通用,能应对各种HTML结构变化,推荐在生产环境使用。
内容的提问来源于stack exchange,提问作者Sun-IT
相关产品推荐
相关产品推荐

