You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多文本差异对比生成超集:替代split逐词对比的实现方案咨询

文本合并的替代实现方案

需求说明

需要将以下三个句子:

  1. Supply and install the material for renovation services. (Size 100 cm X 200 cm)
  2. Supply and install the material for renovation services. (500 sq.m)
  3. Renovation services - Supply and install the material.

合并为指定输出:

Supply and install the material for renovation services. (Size 100 cm X 200 cm) (500 sq.m) Renovation services - Supply and install the material.

目前已通过split逐词对比实现,现寻求正则表达式或现成库的替代方案。

正则表达式实现思路

核心逻辑是先识别重复的核心文本,再提取所有附加信息,最后拼接成目标内容:

  1. 用正则匹配出重复的核心句段:Supply and install the material for renovation services.
  2. 提取前两句中核心文本后的括号内容,统一收集
  3. 将核心文本、所有附加信息、第三句按顺序拼接

Python示例代码

import re

sentences = [
    "Supply and install the material for renovation services. (Size 100 cm X 200 cm)",
    "Supply and install the material for renovation services. (500 sq.m)",
    "Renovation services - Supply and install the material."
]

# 匹配核心文本与括号内容的正则规则
pattern = r"(Supply and install the material for renovation services\.)\s*\((.*?)\)"

core_text = None
attach_info = []

# 处理前两句,提取核心和附加信息
for s in sentences[:2]:
    match_result = re.match(pattern, s)
    if match_result:
        core_text = match_result.group(1)
        attach_info.append(f"({match_result.group(2)})")

# 拼接生成最终结果
final_output = f"{core_text} {' '.join(attach_info)} {sentences[2]}"
print(final_output)

现成库辅助实现

如果需要处理更复杂的文本去重合并场景,可以借助以下工具:

  • fuzzywuzzy:基于模糊匹配识别近似重复的文本片段,适合核心文本存在细微差异的情况
  • spaCy:通过自然语言处理技术提取文本主体与附属属性,实现更智能的语义级合并

fuzzywuzzy示例代码(Python)

from fuzzywuzzy import process
import re

sentences = [
    "Supply and install the material for renovation services. (Size 100 cm X 200 cm)",
    "Supply and install the material for renovation services. (500 sq.m)",
    "Renovation services - Supply and install the material."
]

# 先去掉括号内容,提取前两句的核心候选
core_candidates = [re.sub(r"\s*\(.*?\)", "", s).strip() for s in sentences[:2]]
# 识别重复的核心文本
core_text = process.extractOne(core_candidates[0], core_candidates)[0]

# 收集所有括号内的附加信息
attach_info = re.findall(r"\(.*?\)", ' '.join(sentences[:2]))

# 拼接结果
final_output = f"{core_text} {' '.join(attach_info)} {sentences[2]}"
print(final_output)

内容的提问来源于stack exchange,提问作者Amit Pathak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 17:00:18