You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将带有序号的句子组拆分并转换为指定JSON格式

句子拆分任务解决方案

核心逻辑

拆分误判的根源是混淆了行首的序号标记和句子内部的数字+点内容,因此只需限定仅匹配行首位置的序号规则,即可完全规避句子内部内容的干扰。

具体实现方案

场景1:仍保留原始HTML结构

直接使用HTML解析器提取有序列表内容,完全不需要正则拆分,准确率100%。以Python为例,使用BeautifulSoup库的实现代码如下:

from bs4 import BeautifulSoup
import json

# 输入的HTML片段
html_content = '''
<blockquote>
<ol>
<li>Go to the dining room. Click on the cabinet to take the whisky bottle.</li>
<li>Go to the kitchen. Click on the fridge. Jason gets a lemonade.</li>
</ol>
</blockquote>
'''
soup = BeautifulSoup(html_content, 'html.parser')
li_tags = soup.find('ol').find_all('li')
result = []
for index, li in enumerate(li_tags, start=1):
    result.append({"id": str(index), "text": li.get_text(strip=True)})

# 输出标准JSON
print(json.dumps(result, ensure_ascii=False, indent=2))

场景2:仅留存纯文本内容,无原始HTML

使用带行首定位的正则做分割,正则表达式写为^\d{1,4}\.\s,开启多行匹配模式后,仅会匹配每一行开头1-4位数字+点+空格的序号格式,不会命中句子内部的同类字符组合。拆分后按顺序取序号作为id,剩余内容去除首尾空白后作为text字段即可。

标准输出结果

处理后得到的合法JSON格式如下:

[
  {"id": "1", "text": "Go to the dining room. Click on the cabinet to take the whisky bottle."},
  {"id": "2", "text": "Go to the kitchen. Click on the fridge. Jason gets a lemonade."}
]

注:原需求给出的JSON片段缺少外层数组括号、键未加双引号,不符合JSON规范,以上为修正后的标准格式。

内容的提问来源于stack exchange,提问作者Plugz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 02:24:04