You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算DataFrame每行最长字符链及指定重复模式的最长长度?

Got it, let's tackle these two Pandas questions one by one. I'll walk you through practical, easy-to-implement solutions with code examples.

1. 计算DataFrame每行的最长连续字符链长度

假设你的DataFrame有一列文本数据(比如叫text),我们可以写一个自定义函数来处理每个字符串,然后用apply方法把它应用到每行。

实现思路

遍历字符串的每个字符,跟踪当前连续字符的长度,同时记录遇到的最大长度:

  • 如果当前字符和前一个相同,就增加当前长度;
  • 否则重置当前长度为1;
  • 最后返回记录的最大长度(同时处理空字符串或非字符串的边界情况)。

代码示例

import pandas as pd

def get_longest_consecutive_char_length(s):
    # 处理空字符串或非字符串类型的情况
    if not isinstance(s, str) or len(s) == 0:
        return 0
    max_length = 1
    current_length = 1
    for i in range(1, len(s)):
        if s[i] == s[i-1]:
            current_length += 1
            if current_length > max_length:
                max_length = current_length
        else:
            current_length = 1
    return max_length

# 测试用示例DataFrame
df = pd.DataFrame({
    'text': ['aaabbbcc', 'ababab', 'aaaaa', '', 'xxxyyyzzzz']
})

# 应用函数生成新列
df['longest_consecutive_char'] = df['text'].apply(get_longest_consecutive_char_length)
print(df)

运行结果

text  longest_consecutive_char
0    aaabbbcc                         3
1     ababab                         1
2       aaaaa                         5
3                                     0
4  xxxyyyzzzz                         4
2. 计算每行中连续"a"的最长出现长度

这个需求更具体,只关注字符a的连续序列,这里提供两种常用解法:

方法一:正则表达式(简洁高效)

用正则匹配所有连续的a序列,然后计算这些序列的长度,取最大值即可;如果没有匹配到a,返回0。

代码示例

import re
import pandas as pd

def get_longest_consecutive_a_length(s):
    if not isinstance(s, str):
        return 0
    # 匹配所有连续的'a'序列
    a_sequences = re.findall(r'a+', s)
    if not a_sequences:
        return 0
    # 返回最长序列的长度
    return max(len(seq) for seq in a_sequences)

# 测试用示例DataFrame
df = pd.DataFrame({
    'text': ['aaabbbccaaa', 'ababab', 'aaaaa', 'xxxyyy', 'aabbaaa']
})

df['longest_consecutive_a'] = df['text'].apply(get_longest_consecutive_a_length)
print(df)

运行结果

text  longest_consecutive_a
0  aaabbbccaaa                       3
1    ababab                           1
2      aaaaa                           5
3     xxxyyy                           0
4    aabbaaa                           3

方法二:遍历统计(直观易懂)

和第一个问题的思路类似,但只在字符是a时累加长度,遇到其他字符就重置当前长度,同时记录最大的连续a长度。

代码示例

import pandas as pd

def get_longest_consecutive_a_length(s):
    if not isinstance(s, str) or len(s) == 0:
        return 0
    max_a_length = 0
    current_a_length = 0
    for char in s:
        if char == 'a':
            current_a_length += 1
            if current_a_length > max_a_length:
                max_a_length = current_a_length
        else:
            current_a_length = 0
    return max_a_length

# 应用到示例DataFrame,结果和方法一一致
df['longest_consecutive_a'] = df['text'].apply(get_longest_consecutive_a_length)

内容的提问来源于stack exchange,提问作者Victor A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:53:03