You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用单行Python命令修复curl获取的网页标题HTML实体编码

单行Python命令处理HTML实体编码

完整网页解码保存

直接解码整个网页的HTML实体并保存到文件:

curl "%URLLink%" --insecure -L | python -c "import sys, html; print(html.unescape(sys.stdin.read()))" > "%HTMLFile%"

仅提取并解码标题(无额外依赖)

用正则匹配标题标签后解码,无需安装第三方库:

curl "%URLLink%" --insecure -L | python -c "import sys, html, re; title_match = re.search(r'<title>(.*?)</title>', sys.stdin.read(), re.IGNORECASE); print(html.unescape(title_match.group(1)) if title_match else '未找到标题')"

仅提取并解码标题(精准HTML解析)

若网页结构复杂,推荐用BeautifulSoup解析(需先执行pip install beautifulsoup4完成安装):

curl "%URLLink%" --insecure -L | python -c "import sys, html, bs4; soup = bs4.BeautifulSoup(sys.stdin.read(), 'html.parser'); print(html.unescape(soup.title.string) if soup.title else '未找到标题')"

内容的提问来源于stack exchange,提问作者redking

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 20:22:34