You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python字符串中间文本提取:从Pandas Product列提取商品名称

Solution to Extract Product Names

Got it, let's get this sorted for you! Your goal is to strip the "5 Guys" prefix and then remove everything starting from the first digit or hyphen—here's a working implementation that does exactly that:

Step-by-Step Code

import pandas as pd

# Your original data
fg = pd.DataFrame({
    'Product': [' 5 Guys Greasy Burger 3/5LB (24) [51656]',
                '5 Guys Super Strawberry Shake - (3/4) OZ (9) [5645654]',
                '5 Guys Giant Loaded Double Cheese Burger 1/2LB Buns - 8Z Cups (22) [564654]']
})

# 1. Remove the "5 Guys" prefix (handles leading/trailing spaces around it)
fg['Product'] = fg['Product'].str.replace(r'^\s*5 Guys\s*', '', regex=True)

# 2. Strip everything from the first digit OR hyphen to the end, then clean up extra spaces
fg['Product'] = fg['Product'].str.replace(r'(\d|-).*$', '', regex=True).str.strip()

# Check the result
print(fg['Product'])

Output

0                      Greasy Burger
1            Super Strawberry Shake
2    Giant Loaded Double Cheese Burger
Name: Product, dtype: object

Why Your Previous Code Didn't Work

  • str.strip('5 Guys') doesn't remove the fixed prefix—it removes any combination of the characters 5, , G, u, y, s from the start/end of the string, which can accidentally alter your product names.
  • Your str.replace(r'\[d+\]') had two issues: missing the replacement value (you need to specify what to replace it with, like ''), and the regex was incorrect (\[d+\] matches literal [d+] instead of digits—you'd need \d+ for numbers, but even that wouldn't handle the hyphen case).

Breakdown of the Regex

  1. Prefix Removal: ^\s*5 Guys\s*

    • ^ = matches the start of the string
    • \s* = matches any number of spaces (handles leading spaces before "5 Guys")
    • 5 Guys = the exact prefix we want to remove
    • \s* = matches any spaces after "5 Guys"
  2. Truncate After First Digit/Hyphen: (\d|-).*$

    • (\d|-) = matches the first occurrence of either a digit (\d) or a hyphen (-)
    • .*$ = matches everything from that position to the end of the string
    • .str.strip() = cleans up any leftover trailing spaces after truncation

内容的提问来源于stack exchange,提问作者chasedcribbet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 21:08:11