You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python或正则提取单引号与<=之间的加粗文本

解决方案

问题原因

你写的正则(?=')(.*)(?= <=)使用了贪婪匹配模式,会从文本中第一个单引号开始,一直匹配到最后一个<=的前一位,无法精准匹配每个符合要求的分段内容,也没有针对性提取加粗标签内的文本。

方案1:直接使用正则提取

可以用非贪婪匹配结合捕获组,精准定位<strong>标签内、且后续跟着<=的内容:

import re

# 你的原始文本
raw_text = """[Text(447.1153846153846, 471.625, '<strong>the</strong> <= 0.5
entropy = 0.97
samples = 100.0%
value = [0.399, 0.601]
class = True News'), Text(238.46153846153845, 336.875, '<strong>donald</strong> <= 0.5
entropy = 0.921
samples = 83.7%
value = [0.336, 0.664]
class = True News'), Text(119.23076923076923, 202.125, '<strong>hillary</strong> <= 0.5
entropy = 0.981
samples = 55.6%
value = [0.42, 0.58]
class = True News'), Text(59.61538461538461, 67.375, '\n  (...)  \n'), Text(178.84615384615384, 67.375, '\n  (...)  \n'), Text(357.6923076923077, 202.125, '<strong>hillary</strong> <= 0.5
entropy = 0.663
samples = 28.2%
value = [0.172, 0.828]
class = True News'), Text(298.0769230769231, 67.375, '\n  (...)  \n'), Text(417.30769230769226, 67.375, '\n  (...)  \n'), Text(655.7692307692307, 336.875, '<strong>trumps</strong> <= 0.5
entropy = 0.859
samples = 16.3%
value = [0.718, 0.282]
class = Fake News'), Text(596.1538461538462, 202.125, '<strong>hillary</strong> <= 0.5
entropy = 0.821
samples = 15.7%
value = [0.744, 0.256]
class = Fake News'), Text(536.5384615384615, 67.375, '\n  (...)  \n'), Text(655.7692307692307, 67.375, '\n  (...)  \n'), Text(715.3846153846154, 202.125, 'entropy = 0.0
samples = 0.6%
value = [0.0, 1.0]
class = True News')]"""

# 正则匹配
result = re.findall(r"'<strong>(.*?)<\/strong> <= ", raw_text)
print(result)

运行输出结果为:
['the', 'donald', 'hillary', 'hillary', 'trumps', 'hillary']

方案2:先提取Text参数再匹配(更稳定)

如果文本结构后续有变动,可以先提取每个Text对象的第三个字符串参数,再从参数里提取目标内容,避免误匹配:

import re
from ast import literal_eval

# 先把文本里的Text占位替换成元组,用literal_eval安全解析
parseable_text = raw_text.replace("Text(", "(")
text_list = literal_eval(parseable_text)

result = []
for item in text_list:
    content = item[2]
    match = re.search(r"<strong>(.*?)<\/strong> <= ", content)
    if match:
        result.append(match.group(1))
print(result)

输出结果和方案1完全一致,容错性更高。

内容的提问来源于stack exchange,提问作者DHH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 16:15:03