You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何构建适配符号互换场景的文本1-1亿数值检测器

Let's fix this numerical detector properly! Your current approach only checks digit length, which misses numbers with thousands separators (like 5.000.000) and doesn't account for actual value ranges accurately (e.g., 600700800 is 9 digits but way over 100 million). Here's a step-by-step solution tailored to your locale's number formatting:

Step 1: Define a helper function to extract and validate numbers

We need to:

  • Match numbers in your locale's format (. as thousands separator, , as decimal point)
  • Convert them to standard numerical values
  • Check if they fall within the 1,000,000 to 100,000,000 range
import re
import pandas as pd

def extract_valid_million_value(text):
    # Regex pattern to match numbers in your locale's format:
    # - Matches pure digits OR digits with thousands separators (e.g., 5.000.000)
    # - Optional decimal part (e.g., 2.000.000,00)
    pattern = r'\b(?:\d{1,3}(?:\.\d{3})+|\d+)(?:,\d+)?\b'
    matches = re.findall(pattern, text)
    
    for match in matches:
        # Convert locale-specific format to standard numeric string
        normalized_num = match.replace('.', '').replace(',', '.')
        try:
            num = float(normalized_num)
            # Check if value is between 1M and 100M (inclusive)
            if 1000000 <= num <= 100000000:
                # Return integer if it's a whole number (matches your example output)
                return int(num) if num.is_integer() else num
        except ValueError:
            # Skip invalid conversions (just in case)
            continue
    # No valid number found, return 0
    return 0

Step 2: Apply the function to your DataFrame

Let's test this with your sample data:

# Sample data
data = {
    'id': [1, 2, 3, 4, 5],
    'Body': [
        "You 're get 4500000",
        "Congrats, you receive 500000",
        "Congrats, you receive 5.000.000",
        "Congrats, you get 2.000.000,00!",
        "Your verification code is 600700800"
    ]
}
df = pd.DataFrame(data)

# Add the 'millons' column
df['millons'] = df['Body'].apply(extract_valid_million_value)

Step 3: Verify the output

Running this will give you exactly the expected result:

idBodymillons
1You 're get 45000004500000
2Congrats, you receive 5000000
3Congrats, you receive 5.000.0005000000
4Congrats, you get 2.000.000,00!2000000
5Your verification code is 6007008000

Why this works better than your original approach:

  • Handles locale-specific formatting: It correctly parses numbers with . as thousands separators and , as decimals.
  • Value-based validation: Instead of just checking digit length, it verifies the actual numeric value falls in your target range (so 600700800 is correctly marked as 0 because it's over 100 million).
  • Robustness: Includes error handling for invalid conversions, so malformed number strings won't break the process.

内容的提问来源于stack exchange,提问作者Nabih Bawazir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:23:28