如何用正则表达式结合PHP从人员信息文件提取姓名与邮箱?
Great question! Let's walk through how to extract names and email addresses from your file using PHP and regular expressions. First, let's break down the pattern in your text: each entry starts with a name, followed by various fields (Location, Expertise, etc.), and ends with an Email: [address] line.
Step 1: Craft the Regular Expression
We'll use a regex with named capture groups to easily pull out the name and email. Here's the pattern we'll use:
/(?P<name>.*?)\s+Location:.*?Email:\s*(?P<email>[^\s]+)/s
Let's break this down:
(?P<name>.*?): Non-greedily captures everything from the start of an entry up to the firstLocation:— this gives us the name. The non-greedy*?ensures we don't accidentally capture extra content from subsequent fields.\s+Location:: Matches the start of the Location field, which marks the end of the name..*?: Skips over all the intermediate fields (Expertise, Website, Tel) without capturing them.Email:\s*(?P<email>[^\s]+): Captures the email address immediately afterEmail:.[^\s]+matches all non-whitespace characters (perfect since email addresses don't contain spaces)./s: The "dotall" modifier, which makes.match newline characters — important if your entries span multiple lines.
Step 2: PHP Implementation
Now let's put this into code to read the file and extract the data:
// Replace 'your-contacts.txt' with your actual file path $filePath = 'your-contacts.txt'; // Read the entire file content (use this for smaller files; see note below for large files) $fileContent = file_get_contents($filePath); // Our regex pattern with named groups $regexPattern = '/(?P<name>.*?)\s+Location:.*?Email:\s*(?P<email>[^\s]+)/s'; // Execute the match and organize results by entry preg_match_all($regexPattern, $fileContent, $matches, PREG_SET_ORDER); // Clean up and format the results $contacts = []; foreach ($matches as $entry) { // Remove extra spaces from the name (e.g., multiple spaces between first/last names) $cleanName = trim(preg_replace('/\s+/', ' ', $entry['name'])); $contacts[] = [ 'name' => $cleanName, 'email' => $entry['email'] ]; } // Example: Print the extracted contacts print_r($contacts);
Notes for Edge Cases
- Large Files: If your file is extremely large (hundreds of thousands of entries),
file_get_contentsmight use too much memory. Instead, read the file line by line and accumulate content until you hit anEmail:line, then process that entry and reset the accumulator. - Complex Names: If names include special characters (hyphens, apostrophes), the regex still works since
.*?captures any character except newlines (and the/smodifier handles newlines if needed). - Strict Email Validation: If you need to enforce strict email format rules, replace
[^\s]+with a standard email regex like[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}— but for your given example,[^\s]+is sufficient and faster.
Testing this with your sample content will return:
Array ( [0] => Array ( [name] => Coulthard Sally Coulthard [email] => sally@veterinaryphysio.co.uk ) [1] => Array ( [name] => Kate Haynes [email] => katehaynesphysio@yahoo.co.uk ) )
内容的提问来源于stack exchange,提问作者James_Inger
相关产品推荐
相关产品推荐

