You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将含固定关键词的字符串解析为Pandas DataFrame

需求说明

我有一组字符串需要循环处理并转换成Pandas DataFrame,每个字符串都包含固定关键词,要将这些关键词设为列名,关键词后至下一个关键词前的文本作为对应列的值,所有字符串的列关键词一致。

待解析字符串示例

所有字符串均包含Wilhering Abbey、The buildings、The Rococo关键词,示例内容如下:

Wilhering Abbey is a Cistercian monastery in Wilhering in Upper Austria about 8 km from the city of Linz The buildings re-constructed in the 18th century are known for their spectacular Rococo The Rococo style in the German-speaking world This photograph depicts the interior of the church

理想输出的DataFrame格式

| Wilhering Abbey       | The buildings                               | The Rococo style
|:----------------------|:--------------------------------------------|:-------------------------------------------
| is a Cistercian       | re-constructed                               | style in the German-speaking 
| monastery in Wilhering| in the 18th century are known                | world This photograph depicts   
| in Upper Austria      | for their spectacular Rococo                 | the interior of the church   
| about 8 km from       |                                              |
| the city of Linz      |                                              |

已尝试的代码

result = "Wilhering Abbey is a Cistercian monastery in Wilhering in Upper Austria about 8 km from the city of Linz The buildings re-constructed in the 18th century are known for their spectacular Rococo The Rococo style in the German-speaking world This photograph depicts the interior of the church"

keywords = ['Wilhering Abbey', 'The buildings', 'The Rococo style']
pattern = '|'.join(['(' + i + ')' for i in keywords])
lst = re.split(pattern, result)

解决方案

实现思路

当前的re.split会拆分出关键词本身,导致内容混乱。我们需要先精准提取每个关键词对应的完整文本,再将长文本拆分为多行,最后转换为DataFrame。

完整代码

import re
import pandas as pd

def split_text_into_lines(text, max_words_per_line=5):
    """将文本按指定单词数拆分为多行,保证每行长度相近"""
    words = text.split()
    return [' '.join(words[i:i+max_words_per_line]) for i in range(0, len(words), max_words_per_line)]

# 待处理字符串
target_str = "Wilhering Abbey is a Cistercian monastery in Wilhering in Upper Austria about 8 km from the city of Linz The buildings re-constructed in the 18th century are known for their spectacular Rococo The Rococo style in the German-speaking world This photograph depicts the interior of the church"
keywords = ['Wilhering Abbey', 'The buildings', 'The Rococo style']

# 构建正则匹配模式:提取每个关键词及其后续内容(直到下一个关键词或字符串结尾)
pattern = ""
for idx, kw in enumerate(keywords):
    if idx < len(keywords) - 1:
        pattern += f"({kw})(.*?)(?={keywords[idx+1]})"
    else:
        pattern += f"({kw})(.*)"

# 执行匹配,获取关键词与对应内容的配对
matches = re.findall(pattern, target_str)

# 整理数据:拆分每个列的内容为多行
column_data = {}
for kw, content in matches:
    clean_content = content.strip()
    column_data[kw] = split_text_into_lines(clean_content)

# 转换为DataFrame,空值填充为空字符串
df = pd.DataFrame.from_dict(column_data, orient="index").T.fillna("")
df.columns = keywords

# 打印结果(输出为markdown格式)
print(df.to_markdown(index=False))

代码说明

  1. 正则匹配:动态构建模式,精准捕获每个关键词对应的后续文本,避免拆分错误。
  2. 文本拆分:自定义函数按指定单词数拆分长文本,让每行内容长度均匀。
  3. DataFrame转换:通过from_dict将字典转为DataFrame,orient="index"确保列名对应关键词,fillna填充空行的空白内容。

内容的提问来源于stack exchange,提问作者Munrock

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 15:10:24