如何用PHP从纯文本提取有序问答?现有方案需优化
Alright, let's break down your problem step by step. You're trying to extract structured medical Q&A (questions numbered 12-15, each with 4-5 options) from plain text using a custom PHP function, but you're getting inconsistent results—especially when questions contain numbers or letters like A/B. After 4 hours of troubleshooting, you're wondering what's wrong with your current approach, how to fix it, and if regex would deliver more accurate data. Let's dive in:
1. What's Wrong With Your Current Custom Function?
Your function's core issues stem from ignoring the inherent structure of your Q&A content and relying on fragile, one-size-fits-all markers:
- Over-reliance on fixed start/end triggers: Using
?or:as hard stop points breaks instantly if a question contains those characters mid-sentence (e.g., a question like "Is Vitamin A deficiency linked to night blindness?") or if an option includes punctuation. The function will truncate content incorrectly instead of recognizing the actual end of the question. - No awareness of Q&A patterns: It can't distinguish between a question's numbered identifier (like
12.) and numbers within the question text, or between option labels (likeA.) and letters in medical terms (like "Vitamin A"). This leads to misalignment and broken extractions. - Poor handling of whitespace/line breaks: Plain text Q&A often has line breaks between questions or options, but your function treats the entire string as a single block. This throws off
strposcalculations when content spans multiple lines. - Inflexible logic: The function only checks for two possible question endings, but real-world medical questions might have punctuation in the middle or even no trailing punctuation (though less common here), leading to missed or incorrect extractions.
2. How to Optimize Your Approach (Even Without Regex)
If you want to stick with a custom PHP function, refactor it to work with your content's structured patterns instead of fixed markers:
- Split text into individual Q&A blocks first: Use question numbers (e.g.,
12.,13.) as delimiters to split the full text into separate chunks for each question. You can usestrtokor combineexplodewith checks for digit-plus-dot patterns to isolate each block. - Separate questions from options within each block: For each Q&A chunk, find the first occurrence of an option label (
A.,B., etc.). Everything before that point is the question; everything after is the options set. - Normalize whitespace first: Clean up extra spaces, line breaks, and tabs using
preg_replace('/\s+/', ' ', $string)to make string operations more reliable across inconsistent formatting.
3. Can Regex Give More Accurate Results?
Absolutely—regex is built for exactly this kind of structured text extraction, and it will be far more reliable than your current function. It can target the unique patterns of your Q&A content (numbered questions, lettered options) without being thrown off by internal numbers or letters.
Example PHP Regex Solution
Here's a practical implementation that handles your medical Q&A format:
// Replace this with your actual plain text Q&A content $plainText = "12. What is the primary function of Vitamin D? A. Bone mineralization B. Red blood cell production C. Immune system regulation D. Nerve signal transmission 13. Which of the following is a contraindication for metformin: A. Type 2 diabetes B. Renal impairment C. Obesity D. Hypertension E. Hyperlipidemia 14. A 55-year-old patient presents with chest pain radiating to the left arm—what is the initial diagnostic test? A. ECG B. Chest X-ray C. Blood glucose D. Troponin levels 15. What is the mechanism of action of beta-blockers? A. Block adrenaline receptors B. Inhibit ACE enzyme C. Increase heart rate D. Dilate blood vessels"; // Pattern to match full Q&A blocks (handles multi-line content) $qaPattern = '/(\d+)\.\s*(.*?)\s*(A\..*?)(?=\n\d+\.|$)/s'; preg_match_all($qaPattern, $plainText, $matches, PREG_SET_ORDER); $extractedQA = []; foreach ($matches as $match) { $questionNumber = $match[1]; $questionContent = trim($match[2]); // Split options into key-value pairs (A: text, B: text, etc.) $optionPattern = '/([A-E])\.\s*(.*?)(?=\s*[A-E]\.|$)/s'; preg_match_all($optionPattern, $match[3], $optionMatches, PREG_SET_ORDER); $options = []; foreach ($optionMatches as $opt) { $options[$opt[1]] = trim($opt[2]); } $extractedQA[] = [ 'number' => $questionNumber, 'question' => $questionContent, 'options' => $options ]; } // Output structured results print_r($extractedQA);
How This Regex Works
(\d+)\.\s*: Matches the question's numbered identifier (e.g.,12.) and any following whitespace.(.*?): Non-greedily captures the question content until it hits the first option label (avoids over-capturing into options).(A\..*?): Captures all options starting fromA.until the next question number or the end of the text.- The
smodifier lets.match line breaks, so it handles questions/options that span multiple lines. - The option regex splits each option by its lettered label, creating clean key-value pairs.
This approach ignores numbers/letters within question text, handles inconsistent formatting, and reliably separates questions from options—even when questions contain punctuation or medical terms with letters like "A".
内容的提问来源于stack exchange,提问作者Editor

