如何用Python将含固定关键词的字符串解析为Pandas DataFrame
需求说明
我有一组字符串需要循环处理并转换成Pandas DataFrame,每个字符串都包含固定关键词,要将这些关键词设为列名,关键词后至下一个关键词前的文本作为对应列的值,所有字符串的列关键词一致。
待解析字符串示例
所有字符串均包含Wilhering Abbey、The buildings、The Rococo关键词,示例内容如下:
Wilhering Abbey is a Cistercian monastery in Wilhering in Upper Austria about 8 km from the city of Linz The buildings re-constructed in the 18th century are known for their spectacular Rococo The Rococo style in the German-speaking world This photograph depicts the interior of the church
理想输出的DataFrame格式
| Wilhering Abbey | The buildings | The Rococo style |:----------------------|:--------------------------------------------|:------------------------------------------- | is a Cistercian | re-constructed | style in the German-speaking | monastery in Wilhering| in the 18th century are known | world This photograph depicts | in Upper Austria | for their spectacular Rococo | the interior of the church | about 8 km from | | | the city of Linz | |
已尝试的代码
result = "Wilhering Abbey is a Cistercian monastery in Wilhering in Upper Austria about 8 km from the city of Linz The buildings re-constructed in the 18th century are known for their spectacular Rococo The Rococo style in the German-speaking world This photograph depicts the interior of the church" keywords = ['Wilhering Abbey', 'The buildings', 'The Rococo style'] pattern = '|'.join(['(' + i + ')' for i in keywords]) lst = re.split(pattern, result)
解决方案
实现思路
当前的re.split会拆分出关键词本身,导致内容混乱。我们需要先精准提取每个关键词对应的完整文本,再将长文本拆分为多行,最后转换为DataFrame。
完整代码
import re import pandas as pd def split_text_into_lines(text, max_words_per_line=5): """将文本按指定单词数拆分为多行,保证每行长度相近""" words = text.split() return [' '.join(words[i:i+max_words_per_line]) for i in range(0, len(words), max_words_per_line)] # 待处理字符串 target_str = "Wilhering Abbey is a Cistercian monastery in Wilhering in Upper Austria about 8 km from the city of Linz The buildings re-constructed in the 18th century are known for their spectacular Rococo The Rococo style in the German-speaking world This photograph depicts the interior of the church" keywords = ['Wilhering Abbey', 'The buildings', 'The Rococo style'] # 构建正则匹配模式:提取每个关键词及其后续内容(直到下一个关键词或字符串结尾) pattern = "" for idx, kw in enumerate(keywords): if idx < len(keywords) - 1: pattern += f"({kw})(.*?)(?={keywords[idx+1]})" else: pattern += f"({kw})(.*)" # 执行匹配,获取关键词与对应内容的配对 matches = re.findall(pattern, target_str) # 整理数据:拆分每个列的内容为多行 column_data = {} for kw, content in matches: clean_content = content.strip() column_data[kw] = split_text_into_lines(clean_content) # 转换为DataFrame,空值填充为空字符串 df = pd.DataFrame.from_dict(column_data, orient="index").T.fillna("") df.columns = keywords # 打印结果(输出为markdown格式) print(df.to_markdown(index=False))
代码说明
- 正则匹配:动态构建模式,精准捕获每个关键词对应的后续文本,避免拆分错误。
- 文本拆分:自定义函数按指定单词数拆分长文本,让每行内容长度均匀。
- DataFrame转换:通过
from_dict将字典转为DataFrame,orient="index"确保列名对应关键词,fillna填充空行的空白内容。
内容的提问来源于stack exchange,提问作者Munrock
相关产品推荐
相关产品推荐

