You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于正则表达式拆分含前置空格标识的Pandas DataFrame?

问题描述

我有一个包含问题和结果的CSV文件,编写了一段代码将其转换为DataFrame用于分析,但部分内容未能成功拆分。原因是原代码用startswith('<Q')判断问题行开头,无法处理那些开头带有空格的“ <Q”标识。

原代码:

def start_and_finish_points(df):
    df_indices_start = []
    df_indices_end = []
    rows = df.iloc[:, 0].to_list()
    for i, row in enumerate(rows):
        if str(row).startswith('&lt;Q'):
            df_indices_start.append(i)
        if str(row).endswith('++'):
            df_indices_end.append(i)    
    return df_indices_start, df_indices_end
start, finish = start_and_finish_points(df) 

问题行示例(开头带空格):

698 &lt;Q8&gt; To what extent are you concerned about of the following.................Climate change
    700  &lt;Q11e&gt; How often d...

目标DataFrame列数据:

698    &lt;Q8&gt; To what extent are you concerned about of the following.................Climate change
699                                                                                               
700                                                                                All respondents
704                                                                                 Unweighted row
705                                                                          Effective sample size
706                                                                                          Total
707                                                                        1: Not at all concerned
710                                                                                              2
713                                                                                              3
716                                                                                              4
719                                                                              5: Very concerned
722                                                                                 Not applicable
725                                                                               Total: Concerned
728                                                                           Total: Not Concerned
731                                                                                Net % concerned
733                                                      95% lower case or +, 99% UPPER CASE or ++
735             &lt;Q11e&gt; How often do you access local greenspaces  (e.g. parks, community gardens)?
736                                                                                               
737                                                                                All respondents
741                                                                                 Unweighted row
742                                                                          Effective sample size
743                                                                                          Total
744                                                                           Hardly ever or never
747                                                                               Some of the time
750                                                                                          Often
753                                                                            (Prefer not to say)
Name: nan, dtype: object

如何用正则表达式通用化开头匹配逻辑,处理字符串开头的空格?


解决方案

用正则表达式的re.match()替代startswith(),可以轻松匹配开头带任意数量空格的<Q标识。修改后的代码如下:

import re

def start_and_finish_points(df):
    df_indices_start = []
    df_indices_end = []
    rows = df.iloc[:, 0].to_list()
    # 匹配开头任意空白字符(包括0个)后接&lt;Q的正则模式
    question_pattern = re.compile(r'^\s*&lt;Q')
    
    for i, row in enumerate(rows):
        row_str = str(row)
        # 匹配问题行开头
        if question_pattern.match(row_str):
            df_indices_start.append(i)
        # 保留原有的末尾++匹配逻辑
        if row_str.endswith('++'):
            df_indices_end.append(i)    
    return df_indices_start, df_indices_end

start, finish = start_and_finish_points(df) 

细节优化

如果需要更精准匹配完整的问题标识(比如<Q8>、<Q11e>),避免误匹配其他包含<Q的行,可以把正则表达式改成:

question_pattern = re.compile(r'^\s*&lt;Q\w+&gt;')

这个模式会严格匹配开头任意空格后接<Q+字母数字+>的完整标识。

另外,如果++行末尾可能带有空格,也可以用正则处理:

if re.search(r'\+\+\s*$', row_str):
    df_indices_end.append(i)

内容的提问来源于stack exchange,提问作者elksie5000

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 19:56:13