You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效移除DataFrame中facet_cls列的'-'及后续字符?

Fastest Ways to Clean the facet_cls Column in a Large Pandas DataFrame

Since you're dealing with a massive dataset, we need to prioritize vectorized string operations (avoid slow apply() loops at all costs—they’ll crawl on big data). Here are the most efficient methods to strip everything from the first '-' onwards in the facet_cls column:

This method is optimized exactly for your use case: it splits the string at the first '-' and returns only the part before it. It’s consistently one of the fastest options for this task:

df['facet_cls'] = df['facet_cls'].str.partition('-')[0]

It stops processing immediately after the first '-', so there’s no unnecessary work done on the rest of the string.

Method 2: Using str.split() with n=1

If you prefer split syntax, limiting the number of splits to 1 ensures we only grab the segment before the first '-':

df['facet_cls'] = df['facet_cls'].str.split('-', n=1).str[0]

This is nearly as fast as partition, but slightly less efficient in most real-world benchmarks.

Method 3: Using str.extract() with Regex

A regex-based approach that captures all characters up to the first '-':

df['facet_cls'] = df['facet_cls'].str.extract(r'^([^-]+)', expand=False)

This works reliably, but regex operations can have marginal overhead compared to the split/partition methods for extremely large datasets.

Quick Notes:

  • All these methods leave values without '-' completely unchanged, which aligns perfectly with your input/output example.
  • For datasets with millions of rows, str.partition() is usually the clear winner in speed tests.
  • Never use df['facet_cls'].apply(lambda x: x.split('-')[0])—this is a hidden loop under the hood and will be drastically slower for big data.

内容的提问来源于stack exchange,提问作者EJ Kang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:51:52