在R中使用多正则表达式模式进行数据过滤的技术咨询
Hey there! Let's tackle this data filtering regex problem step by step. You need to match four specific string patterns, so I'll break down each one, then combine them into a single regex that works for all cases.
Breakdown of Each Pattern's Regex
Let's start with the regex for each individual pattern, so you understand exactly what each part does:
- 10-digit consecutive numbers:
^\d{10}$- This matches exactly 10 digits, no more, no less. The
^and$make sure we're matching the entire string—so we won't catch something like1250126681abc(which has extra characters after the digits).\d{10}is shorthand for "match 10 digits in a row".
- This matches exactly 10 digits, no more, no less. The
- 13-digit consecutive numbers:
^\d{13}$- Same logic as the 10-digit pattern, just scaled to 13 digits. Perfect for your example
9781626724266.
- Same logic as the 10-digit pattern, just scaled to 13 digits. Perfect for your example
- Starts with "id" followed by 9 digits:
^id\d{9}$^idensures the string starts with the literal characters "id", then\d{9}matches exactly 9 digits right after. Again,$makes sure there's nothing after those 9 digits—soid975448501xwon't be a match.
- 10-character mix of uppercase letters and digits:
^[A-Z0-9]{10}$[A-Z0-9]defines the allowed characters (uppercase A-Z and 0-9),{10}enforces exactly 10 of them. This will catch your exampleB004TLHNOCbut reject lowercase versions likeb004tlhnoc(if you need to allow lowercase, we can adjust this with a flag, but your example uses uppercase so we kept it strict).
Combined Regex
To match any of these four patterns in your dataset, we can combine them using the | (OR) operator. To make it cleaner, we can wrap all the patterns in a non-capturing group so we only need to use the start/end anchors once:
^(?:\d{10}|\d{13}|id\d{9}|[A-Z0-9]{10})$
If you prefer to write it without the group (repeating the anchors), this works too:
^\d{10}$|^\d{13}$|^id\d{9}$|^[A-Z0-9]{10}$
Example Usage (Python with Pandas)
If you're working with a pandas DataFrame (a common use case for this kind of filtering), here's how you'd apply this regex to get your expected results:
import pandas as pd # Sample dataset (mix of valid and invalid entries) sample_data = pd.DataFrame({ 'identifier': [ '1250126681', '9781626724266', 'id975448501', 'B004TLHNOC', '12345', 'id1234567', 'b004tlhnoc', '12345678901' ] }) # Filter to keep only valid entries filtered_data = sample_data[sample_data['identifier'].str.match(r'^(?:\d{10}|\d{13}|id\d{9}|[A-Z0-9]{10})$')] print(filtered_data)
Running this will return only the first four valid entries in the sample list—exactly what you're looking for.
Quick Notes
- If you need to match these patterns as substrings (e.g., within a longer text like "Order ID: 1250126681"), just remove the
^and$anchors from the regex. But based on your examples, it seems like you want to match full strings, so the anchors are important here. - If you're using a different tool (like Excel, SQL, or R), the regex syntax might be slightly different, but the core pattern logic stays the same.
内容的提问来源于stack exchange,提问作者user3614783

