求助:如何用PHP将文本文件解析为4列(第三列可选存在)
Hey there! Let's work through this parsing problem together. The core issue here is handling rows that either have 3 visible columns (with the 4th being implied as empty for the third column) or 4 columns—plus that tricky edge case where the third and fourth columns are merged with a slash.
First, let's break down the consistent patterns in your text:
- Column 1: The identifier (like
246/RD/2010or996 /RD/2015) – always the first element, followed by a space. - Column 2: A date in
dd.mm.yyyyformat – present in every row as the second element. - Column 3: Optional code (like
211/Por1049/P) – sometimes missing entirely. - Column 4: The final date in
dd.mm.yyyyformat – always present, either as a standalone element or merged with Column 3 via a slash.
Solution: Use a Flexible Regular Expression
Replacing spaces (like you tried earlier) won't work because spacing is inconsistent, and some rows merge columns with slashes. Instead, a regex that matches all possible row variations is the way to go. Here's a PHP implementation:
// Replace this with your actual file reading logic $sampleText = <<<TEXT 246/RD/2010 05.01.2010 211/P 12.11.2010 247/RD/2010 05.01.2010 195/P 09.11.2010 248/RD/2010 05.01.2010 13.10.2010 251/RD/2010 05.01.2010 274/P 08.12.2010 996 /RD/2015 19.01.2015 1049/P/04.12.2015 148934/RD/2010 13.10.2010 28.01.2011 TEXT; // Split text into individual rows $rows = explode("\n", trim($sampleText)); // Regex pattern to cover all row variations $parsePattern = '/^(.+?)\s+(\d{2}\.\d{2}\.\d{4})\s+(?:([^\s\/]+\/[^\s\/]+)\/?(\d{2}\.\d{2}\.\d{4})|(\d{2}\.\d{2}\.\d{4}))$/'; $parsedResults = []; foreach ($rows as $row) { if (preg_match($parsePattern, $row, $matches)) { $col1 = trim($matches[1]); $col2 = $matches[2]; // Assign column 3 if it exists, else leave empty $col3 = !empty($matches[3]) ? $matches[3] : ''; // Column 4 comes from either the merged case or the no-col3 case $col4 = !empty($matches[4]) ? $matches[4] : $matches[5]; $parsedResults[] = [ 'column_1' => $col1, 'column_2' => $col2, 'column_3' => $col3, 'column_4' => $col4 ]; } } // Example: Output parsed data echo "<pre>"; print_r($parsedResults); echo "</pre>";
How the Regex Works
Let's unpack the pattern to make it clear:
^(.+?)\s+(\d{2}\.\d{2}\.\d{4})\s+: Captures Column 1 (non-greedy match until the first date) and Column 2 (thedd.mm.yyyydate).(?: ... ): A non-capturing group that handles two scenarios for the remaining content:([^\s\/]+\/[^\s\/]+)\/?(\d{2}\.\d{2}\.\d{4}): Matches both merged columns (like1049/P/04.12.2015) and separate columns (like211/P 12.11.2010), capturing Column 3 and Column 4.|(\d{2}\.\d{2}\.\d{4}): The fallback case where there's no Column 3 – directly captures Column 4.
Reading from a File
If you're working with a physical file instead of a sample string, use this file reading logic instead:
$fileHandle = fopen('your-file-path.txt', 'r'); while (($row = fgets($fileHandle)) !== false) { $row = trim($row); if (empty($row)) continue; // Skip empty lines // Apply the preg_match logic from above here } fclose($fileHandle);
This solution will reliably parse all the row variations in your sample data into the 4 columns you need, even when Column 3 is missing or merged with Column 4.
内容的提问来源于stack exchange,提问作者AlexF

