手机号标准化处理与多维度重复值检测的技术实现问询
手机号标准化处理与多维度重复值检测的技术实现问询
Hey everyone, hope you're all doing well! I’ve got a data processing task I’m stuck on and would love some guidance from this community.
Here’s the breakdown of my problem:
I’m working with a dataset that has three columns:
- id: Contains IDs, but there are duplicate entries for some IDs
- phone_number: Phone numbers tied to each ID, captured in messy, inconsistent formats that need cleaning to a single standard. A single ID might end up with multiple unique phone numbers after cleaning.
- phone_random: Random phone numbers linked to IDs from a separate source
I need to create three output columns based on this data:
- output_1: Flags whether an ID has more than one unique cleaned phone number
- output_2: Flags whether an ID’s cleaned phone number matches any cleaned phone number from a different ID
- output_3: Flags whether an ID’s cleaned phone number matches any cleaned number from the
phone_randomcolumn (even if it’s from the same ID)
Here’s a sample of my data in R:
id <- c(1, 2, 3, 1, 4, 5,2,3,7) phone_number <- c("4121234567","3137894561", "1234567788","(412)123-45%67", "919-789-1$122","(123)1112233", "(412)1234567","1234567788", "123-11%12233") phone_random<- c("na","4121234567", "","na", "123-1112233", "na","","919-789-1$122","") df <- data.frame(id, phone_number,phone_random) # Preview first 6 rows df %>% head()
The preview output looks like this:
id phone_number phone_random 1 1 4121234567 na 2 2 3137894561 4121234567 3 3 1234567788 4 1 (412)123-45%67 na 5 4 919-789-1$122 123-1112233 6 5 (123)1112233 na
Feel free to ask if you need any extra details about the task—I can share more context if needed. Thank you so much for your help in advance!
备注:内容来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

