You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将带编号的提取字符串转换为指定格式的Pandas DataFrame?

问题描述

已提取字符串如下:

extractedString = "1) No structured exercise.

2) Above ideal body Mass index.

3) Cancer gene testing.

4) Suboptimal vitamin D.

5) Slight anaemia."

需要将其转换为指定格式的Pandas DataFrame,最终输出如下:

0
0       No structured exercise
1  Above ideal body Mass index
2          Cancer gene testing
3         Suboptimal vitamin D
4               Slight anaemia

最优实现方式

方法1:正则匹配提取(适配性最强)

用正则精准提取每个条目的核心内容,不受序号格式、空格变化的影响,代码如下:

import pandas as pd
import re

extractedString = "1) No structured exercise.\n\n2) Above ideal body Mass index.\n\n3) Cancer gene testing.\n\n4) Suboptimal vitamin D.\n\n5) Slight anaemia."

# 匹配规则:跳过数字+括号,提取到末尾句号前的文本
pattern = r'\d+\s*\)\s*(.*?)\.'
items = re.findall(pattern, extractedString)

# 转为DataFrame
df = pd.DataFrame(items, columns=[0])
print(df)

说明:正则表达式\d+\s*\)\s*(.*?)\.会自动处理:

  • 序号后的空白字符
  • 条目末尾的句号
  • 不同长度的序号数字

方法2:分割字符串清理(格式固定时更高效)

如果字符串格式完全统一,可通过分割+文本清理快速处理,代码如下:

import pandas as pd

extractedString = "1) No structured exercise.\n\n2) Above ideal body Mass index.\n\n3) Cancer gene testing.\n\n4) Suboptimal vitamin D.\n\n5) Slight anaemia."

# 按换行分割,过滤空行后逐个清理内容
items = [line.split(')', 1)[1].strip().rstrip('.') 
         for line in extractedString.split('\n') 
         if line.strip()]

df = pd.DataFrame(items, columns=[0])
print(df)

说明:

  1. split('\n')拆分所有行,if line.strip()过滤空行
  2. split(')', 1)[1]按第一个括号拆分,取后半段内容
  3. strip()去掉前后空白,rstrip('.')移除末尾句号

两种方法都能得到目标结果,正则法更适合格式有变动的场景,分割法在格式固定时执行效率更高。

内容的提问来源于stack exchange,提问作者WhoamI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 05:25:06