如何修改代码参数提取DataFrame中【】内不同长度的数字?
Fixing the Number Extraction from Your DataFrame Column
The problem with your current code is that (\d{4}) specifically matches exactly 4 digits—so it misses shorter numbers like 123 or 1 entirely. To extract any length of digits wrapped in 【】, we need to adjust the regular expression to match one or more digits instead of a fixed count.
Step-by-Step Solution
- Update the regex pattern: Replace
\d{4}with\d+, which matches 1 or more consecutive digits. We also anchor it to the【and】delimiters to ensure we only extract numbers inside those brackets. - Modified code:
df['num'] = df['News'].str.extract('【(\d+)】')
Example Verification
If your df["News"] has entries like:
【123】文本文本【1234】文本文本文本【1】文本文本文本
Running the modified code will populate df['num'] with:
12312341
Bonus Note
If some cells have multiple 【】 blocks and you need all extracted numbers, use str.extractall instead and reshape the result (but from your description, it looks like each cell has one number block, so str.extract works perfectly).
内容的提问来源于stack exchange,提问作者Arthur Law
相关产品推荐
相关产品推荐

