如何使用Python按空格拆分字符串中的中英文词汇片段
中英混合字符串拆分方案
失效根因
使用isalpha+isspace判断无效的核心原因:Python的str.isalpha()方法会对所有Unicode定义的字母类字符返回True,中文字符也属于该类别,因此会被误判为符合规则,导致整个字符串被划入第一部分,最终得到["Coffee Hello 咖啡 咖啡", ""]的错误结果。
正则实现代码(Python示例)
import re def split_str(content): # 匹配首个中文字符的位置 first_cn = re.search(r"[\u4e00-\u9fff]", content) if not first_cn: return content, "" split_idx = first_cn.start() # 拆分后去除首尾多余空格,不需要可直接去掉strip() part_en = content[:split_idx].strip() part_cn = content[split_idx:].strip() return part_en, part_cn # 测试效果 test_content = "Coffee Hello 咖啡 咖啡" print(split_str(test_content)) # 输出:('Coffee Hello', '咖啡 咖啡')
拓展说明
- 如果需要匹配生僻汉字,可将正则规则扩展为
r"[\u4e00-\u9fff\u3400-\u4dbf]" - 如果拆分逻辑反过来,要在首个英文字符位置拆分,只需将正则替换为
r"[a-zA-Z]"即可
内容的提问来源于stack exchange,提问作者Harshdeep Kanhai
相关产品推荐
相关产品推荐

