You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在JavaScript中提取字符串中的中文标点符号?

提取字符串中的中文标点符号的简便实现

核心思路

既然中文标点数量有限(约20种),直接枚举所有目标标点构建匹配规则,是比Unicode范围匹配更简洁直观的方案,完全避开复杂的字符集范围配置。

优化后的实现代码

const str = "你好,让我们试试这个分词效果,你说怎么样?Let's try Intl.Segmenter, should we ?"
// 枚举常用中文标点集合,可根据需求增减
const chinesePunctuations = ',。?!;:“”‘’《》【】()——……·、'
// 构建匹配中文标点的正则表达式
const chinesePunctRegex = new RegExp(`[${chinesePunctuations}]`, 'u')

const segmenterZH = new Intl.Segmenter('zh', { granularity: 'grapheme' })
const segments = segmenterZH.segment(str)

for (const segment of segments) {
  if (chinesePunctRegex.test(segment.segment)) {
    console.log(`${segment.index}:${segment.segment}`)
  }
}

关于补充测试的说明

  • /\p{sc=Han}/u 专门匹配中文字符(表意文字),本身就不包含标点符号,符合预期;
  • /\p{scx=Han}/u 仅能匹配《》这类明确归属中文脚本的标点,大部分中文标点(如全角逗号、问号)属于全角符号范畴,并不归属于Han脚本扩展,因此会遗漏,这个正则确实不适合用来匹配所有中文标点。

内容的提问来源于stack exchange,提问作者Qiulang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 15:33:26