You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中按指定字符的第N次出现分割列并提取内容

Fixing Your Pandas ID Column Extraction Issue

Hey there! Let's work through this problem to get the exact result you're looking for. First, let's break down why your previous attempts didn't work, then walk through two reliable solutions.

Why Your Original Methods Failed

1. Regular Expression Misstep

Your regex pattern r'([^.]*,[^,]*)' doesn't match the structure of your IDs at all—you're using commas instead of periods, which is why it didn't capture the right content. Additionally, the "string methods not callable" error likely happened because your ID column wasn't stored as a string type. Pandas can't call .str methods on non-string data (like numeric types or mixed types).

2. List Comprehension Error

Your list comprehension tried to treat the entire Series of split results as a single item, then joined everything with spaces. That's why you didn't get the per-row results you wanted. You need to handle each row's split values individually, not the whole Series at once.

Solution 1: Regular Expression Extraction

First, make sure your ID column is a string, then use a regex that specifically captures the first three period-separated segments.

# Ensure the ID column is string type (critical for str methods)
df['ID'] = df['ID'].astype(str)

# Regex pattern to capture first 3 segments (before the 3rd period)
pattern = r'^([^.]+\.[^.]+\.[^.]+)'
df['test'] = df['ID'].str.extract(pattern, expand=False)
  • ^ anchors the match to the start of the string
  • [^.]+ matches one or more characters that aren't a period
  • \. matches a literal period
  • Repeating this three times captures exactly the first three segments joined by periods.

Solution 2: Split & Join (Simpler for This Case)

This method is more intuitive for splitting by periods and keeping the first three parts:

# Again, ensure ID is string type
df['ID'] = df['ID'].astype(str)

# Split each ID by periods, take first 3 elements, then rejoin with periods
df['test'] = df['ID'].str.split('.').str[:3].str.join('.')
  • .str.split('.') splits each ID into a list of segments
  • .str[:3] grabs the first three elements of each list
  • .str.join('.') puts those three segments back together with periods.

Both methods will give you the expected results:

IDtest
AB.156483.15645431.1561313513AB.156483.15645431
CD.15615a.4651d15351.1512.1.21CD.15615a.4651d15351

内容的提问来源于stack exchange,提问作者Thomas Byrnes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 23:09:04