You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改代码参数提取DataFrame中【】内不同长度的数字?

Fixing the Number Extraction from Your DataFrame Column

The problem with your current code is that (\d{4}) specifically matches exactly 4 digits—so it misses shorter numbers like 123 or 1 entirely. To extract any length of digits wrapped in 【】, we need to adjust the regular expression to match one or more digits instead of a fixed count.

Step-by-Step Solution

  1. Update the regex pattern: Replace \d{4} with \d+, which matches 1 or more consecutive digits. We also anchor it to the 【 and 】 delimiters to ensure we only extract numbers inside those brackets.
  2. Modified code:
    df['num'] = df['News'].str.extract('【(\d+)】')
    

Example Verification

If your df["News"] has entries like:

  • 【123】文本文本
  • 【1234】文本文本文本
  • 【1】文本文本文本

Running the modified code will populate df['num'] with:

  • 123
  • 1234
  • 1

Bonus Note

If some cells have multiple 【】 blocks and you need all extracted numbers, use str.extractall instead and reshape the result (but from your description, it looks like each cell has one number block, so str.extract works perfectly).

内容的提问来源于stack exchange,提问作者Arthur Law

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:39:08