You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则捕获文本首个单词并实现标题与对应单词的输出?

解决方案:捕获标题首个单词并对应打印

一、捕获首个单词的正确实现

你之前的代码用line[:10]截取固定长度字符,这是错误的。要精准捕获每行的首个单词,有两种可靠方法:

方法1:正则表达式匹配

使用re.match从行首匹配连续的单词字符(支持大小写字母、数字、下划线),如果标题首个单词仅包含字母,可以把\w+换成[a-zA-Z]+。

import re

first_word = []
for line in messy_info:
    # 先去除行首行尾的空白字符,避免空行或开头空格干扰
    cleaned_line = line.strip()
    # 匹配行首的单词
    match = re.match(r"^\w+", cleaned_line)
    if match:
        # group()获取匹配到的完整单词
        first_word.append(match.group())

print(first_word)

方法2:字符串分割法

无需正则,直接用split(maxsplit=1)按第一个空格分割字符串,取第一部分即可:

first_word = []
for line in messy_info:
    cleaned_line = line.strip()
    if cleaned_line:  # 跳过空行
        # maxsplit=1确保只分割一次,避免后续空格影响
        word = cleaned_line.split(maxsplit=1)[0]
        first_word.append(word)

print(first_word)

二、打印标题及对应首个单词的正确代码

你之前的代码存在变量未定义(word)、正则分割逻辑错误的问题,修正后的实现如下:

正则版本

import re

for line in messy_info:
    cleaned_line = line.strip()
    match = re.match(r"^\w+", cleaned_line)
    if match:
        first_word = match.group()
        print(cleaned_line)
        print(f"First Word: {first_word}")
        print()
    else:
        print("--- not a match ---")
        print()

字符串分割版本

for line in messy_info:
    cleaned_line = line.strip()
    if cleaned_line:
        first_word = cleaned_line.split(maxsplit=1)[0]
        print(cleaned_line)
        print(f"First Word: {first_word}")
        print()
    else:
        print("--- not a match ---")
        print()

关键修改点

  1. 用strip()处理每行的首尾空白,避免空行或开头空格导致的匹配错误
  2. 去掉无意义的line[:100]截取,直接使用清理后的完整行内容
  3. 修正了未定义变量word的问题,直接从当前行提取首个单词
  4. 用maxsplit=1限制分割次数,确保只取第一个空格前的内容

内容的提问来源于stack exchange,提问作者Alyce1988

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 08:00:30