You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析HTML格式Telegram消息并在发送前优化内容展示

解决方案

你可以在Jenkins流水线调用脚本解析处理Jira输出的HTML内容后再推送给Telegram,具体操作如下:

工具选择

使用Python的BeautifulSoup库做HTML解析,比正则匹配更稳定,不会出现格式误判问题。

处理代码示例

  1. 首先安装依赖:
    pip install beautifulsoup4
  2. 内容处理逻辑参考:
from bs4 import BeautifulSoup

def format_telegram_msg(jira_raw_html):
    # 解析原始HTML
    soup = BeautifulSoup(jira_raw_html, "html.parser")
    
    # 需求1:删除所有URL链接,保留链接内的文本/图片说明
    for a_tag in soup.find_all("a"):
        # 去掉a标签本身,保留内部内容
        a_tag.unwrap()
    
    # 需求2可选:不需要保留任何HTML标签的话,直接提取纯文本即可
    # pure_text = soup.get_text(separator="\n", strip=True)
    # return pure_text
    
    # 需求2可选:保留Telegram支持的少量格式标签
    allowed_tags = ["b", "i", "u", "s", "code", "pre"]
    for tag in soup.find_all(True):
        if tag.name not in allowed_tags:
            tag.unwrap()
        # 清空所有标签的多余属性,避免冗余内容
        tag.attrs = {}
    
    return str(soup)

流水线集成

将上述处理逻辑封装为可执行脚本,在Jenkins流水线中调用脚本获取处理后的内容,再调用Telegram Bot的发送接口即可,处理后的内容不会携带多余的HTML标签和URL链接,Telegram端展示整齐。

内容的提问来源于stack exchange,提问作者Salvador Arreola Rojas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 21:15:04