Data Studio中REGEXP_MATCH函数求助:数据清洗正则表达式编写
Got it, let's work through this data cleaning challenge in Google Data Studio together. Based on your sample data and desired output, here's a step-by-step solution using nested REGEXP_REPLACE functions (tailored to Data Studio's RE2 regex syntax):
Final Formula
Replace YourFieldName with the actual name of your data field in Data Studio:
REGEXP_REPLACE( REGEXP_REPLACE( REGEXP_REPLACE( REGEXP_REPLACE(YourFieldName, '^\\d+\\s*-\\s*', ''), '-', '' ), '[éè]', 'e' ), 's$', '' )
Breakdown of Each Step
Let's walk through what each part does, using your sample data as reference:
Remove leading number & hyphen prefix
- Regex:
^\\d+\\s*-\\s* - What it matches: One or more digits at the start of the string, followed by optional spaces, a hyphen, and more optional spaces.
- Example:
3 - Pre-apple→ becomesPre-apple
- Regex:
Remove all hyphens in the string
- Regex:
- - What it does: Eliminates any hyphens to combine parts like "Pre-apple" into "Preapple"
- Example:
Pré-apples→ becomesPréapples
- Regex:
Normalize accented characters
- Regex:
[éè] - What it does: Replaces accented e characters (é, è) with a standard lowercase e to unify variations like
PrèappleandPré-apples - Example:
Prèapple→ becomesPreapple
- Regex:
Remove trailing plural "s"
- Regex:
s$ - What it does: Targets only the "s" at the end of the string to convert plural forms like "Apples" to "Apple"
- Example:
Apples→ becomesApple
- Regex:
Testing Against Your Sample Data
Let's verify this works for all your examples:
1- Apple→Apple2- Apples→Apple3 - Pre-apple→Preapple4- Pré-apples→Preapple5-Prèapple→Preapple
All outputs match exactly what you're looking for! If you have additional accented characters or edge cases, you can expand the character sets in the regex (e.g., add à to [éè] as [éèà] and replace with a if needed).
内容的提问来源于stack exchange,提问作者h_data

