英国电话号码DataFrame验证:代码语法错误及分类标记失败
Fixing UK Phone Number Validation in Pandas DataFrame
Let’s start by breaking down the issues in your code, then build a working solution to classify UK phone numbers as mobile, home (landline), or invalid.
What’s Wrong With Your Current Code?
- Mixed Pandas/PySpark Syntax:
withColumnis a PySpark DataFrame method, not a Pandas one. Pandas usesapply/mapfor row/column-level operations. - Invalid Assignment:
df['Phonenumber']=df(df.withColumn(...))is incorrect syntax—you’re trying to call the DataFrame like a function, which doesn’t work. - Missing Regex Logic: You haven’t defined any regular expressions to distinguish mobile vs landline numbers, which is the core of the validation.
- Incorrect Conditional Syntax: You can’t use a raw
if/elseblock directly when assigning a new column in Pandas; you need to wrap logic in a function or lambda.
Working Solution
First, let’s define standard UK phone number patterns and a validation function, then apply it to your DataFrame.
Step 1: Define Validation Logic with Regex
UK phone numbers follow consistent patterns:
- Mobile numbers: Start with
07,+447, or00447(total of 11 digits after the country code/prefix) - Landline (home) numbers: Start with
01,02,+441,+442, or00441/00442(9-10 digits after the prefix) - We’ll first clean numbers to remove non-digit characters (spaces, parentheses, hyphens) to handle common formatting variations.
import pandas as pd import re def validate_uk_phone(phone): # Clean the number: remove all non-digit/non-+ characters cleaned_phone = re.sub(r'[^\d+]', '', str(phone)) # Regex patterns for UK mobile and landline numbers mobile_pattern = r'^(\+447|07|00447)\d{9}$' landline_pattern = r'^(\+44[12]|0[12]|0044[12])\d{9,10}$' # Check against patterns if re.match(mobile_pattern, cleaned_phone): return 'Mobile Number' elif re.match(landline_pattern, cleaned_phone): return 'Home Number' else: return 'Invalid Number'
Step 2: Apply to Your DataFrame
Use Pandas apply to run the validation function on every value in your Phonenumber column and create a new results column:
# Example DataFrame (replace with your actual data) df = pd.DataFrame({ 'Phonenumber': [ '07123 456 789', '+44 7987 654 321', '01234 567 890', '+44 20 1234 5678', '123456789', '0800 123 456' ] }) # Add the validation column df['Phone_Number_Validity'] = df['Phonenumber'].apply(validate_uk_phone) # Show the results display(df)
Expected Output
| Phonenumber | Phone_Number_Validity |
|---|---|
| 07123 456 789 | Mobile Number |
| +44 7987 654 321 | Mobile Number |
| 01234 567 890 | Home Number |
| +44 20 1234 5678 | Home Number |
| 123456789 | Invalid Number |
| 0800 123 456 | Invalid Number |
Notes
- If you need to support more formatting variations (e.g., parentheses like
(07123) 456789), the cleaning step already handles this by stripping non-digit/non-+ characters. - To adjust the regex for stricter/laxer matching (e.g., allow extension numbers), modify the patterns accordingly.
内容的提问来源于stack exchange,提问作者shama khan
相关产品推荐
相关产品推荐

