如何基于正则表达式拆分含前置空格标识的Pandas DataFrame?
问题描述
我有一个包含问题和结果的CSV文件,编写了一段代码将其转换为DataFrame用于分析,但部分内容未能成功拆分。原因是原代码用startswith('<Q')判断问题行开头,无法处理那些开头带有空格的“ <Q”标识。
原代码:
def start_and_finish_points(df): df_indices_start = [] df_indices_end = [] rows = df.iloc[:, 0].to_list() for i, row in enumerate(rows): if str(row).startswith('<Q'): df_indices_start.append(i) if str(row).endswith('++'): df_indices_end.append(i) return df_indices_start, df_indices_end start, finish = start_and_finish_points(df)
问题行示例(开头带空格):
698 <Q8> To what extent are you concerned about of the following.................Climate change 700 <Q11e> How often d...
目标DataFrame列数据:
698 <Q8> To what extent are you concerned about of the following.................Climate change 699 700 All respondents 704 Unweighted row 705 Effective sample size 706 Total 707 1: Not at all concerned 710 2 713 3 716 4 719 5: Very concerned 722 Not applicable 725 Total: Concerned 728 Total: Not Concerned 731 Net % concerned 733 95% lower case or +, 99% UPPER CASE or ++ 735 <Q11e> How often do you access local greenspaces (e.g. parks, community gardens)? 736 737 All respondents 741 Unweighted row 742 Effective sample size 743 Total 744 Hardly ever or never 747 Some of the time 750 Often 753 (Prefer not to say) Name: nan, dtype: object
如何用正则表达式通用化开头匹配逻辑,处理字符串开头的空格?
解决方案
用正则表达式的re.match()替代startswith(),可以轻松匹配开头带任意数量空格的<Q标识。修改后的代码如下:
import re def start_and_finish_points(df): df_indices_start = [] df_indices_end = [] rows = df.iloc[:, 0].to_list() # 匹配开头任意空白字符(包括0个)后接<Q的正则模式 question_pattern = re.compile(r'^\s*<Q') for i, row in enumerate(rows): row_str = str(row) # 匹配问题行开头 if question_pattern.match(row_str): df_indices_start.append(i) # 保留原有的末尾++匹配逻辑 if row_str.endswith('++'): df_indices_end.append(i) return df_indices_start, df_indices_end start, finish = start_and_finish_points(df)
细节优化
如果需要更精准匹配完整的问题标识(比如<Q8>、<Q11e>),避免误匹配其他包含<Q的行,可以把正则表达式改成:
question_pattern = re.compile(r'^\s*<Q\w+>')
这个模式会严格匹配开头任意空格后接<Q+字母数字+>的完整标识。
另外,如果++行末尾可能带有空格,也可以用正则处理:
if re.search(r'\+\+\s*$', row_str): df_indices_end.append(i)
内容的提问来源于stack exchange,提问作者elksie5000
相关产品推荐
相关产品推荐

