You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何用PHP将文本文件解析为4列(第三列可选存在)

PHP解析可变列数的文本文件为固定4列

Hey there! Let's work through this parsing problem together. The core issue here is handling rows that either have 3 visible columns (with the 4th being implied as empty for the third column) or 4 columns—plus that tricky edge case where the third and fourth columns are merged with a slash.

First, let's break down the consistent patterns in your text:

  • Column 1: The identifier (like 246/RD/2010 or 996 /RD/2015) – always the first element, followed by a space.
  • Column 2: A date in dd.mm.yyyy format – present in every row as the second element.
  • Column 3: Optional code (like 211/P or 1049/P) – sometimes missing entirely.
  • Column 4: The final date in dd.mm.yyyy format – always present, either as a standalone element or merged with Column 3 via a slash.

Solution: Use a Flexible Regular Expression

Replacing spaces (like you tried earlier) won't work because spacing is inconsistent, and some rows merge columns with slashes. Instead, a regex that matches all possible row variations is the way to go. Here's a PHP implementation:

// Replace this with your actual file reading logic
$sampleText = <<<TEXT
246/RD/2010 05.01.2010 211/P 12.11.2010
247/RD/2010 05.01.2010 195/P 09.11.2010
248/RD/2010 05.01.2010 13.10.2010
251/RD/2010 05.01.2010 274/P 08.12.2010
996 /RD/2015 19.01.2015 1049/P/04.12.2015
148934/RD/2010 13.10.2010 28.01.2011
TEXT;

// Split text into individual rows
$rows = explode("\n", trim($sampleText));

// Regex pattern to cover all row variations
$parsePattern = '/^(.+?)\s+(\d{2}\.\d{2}\.\d{4})\s+(?:([^\s\/]+\/[^\s\/]+)\/?(\d{2}\.\d{2}\.\d{4})|(\d{2}\.\d{2}\.\d{4}))$/';

$parsedResults = [];
foreach ($rows as $row) {
    if (preg_match($parsePattern, $row, $matches)) {
        $col1 = trim($matches[1]);
        $col2 = $matches[2];
        // Assign column 3 if it exists, else leave empty
        $col3 = !empty($matches[3]) ? $matches[3] : '';
        // Column 4 comes from either the merged case or the no-col3 case
        $col4 = !empty($matches[4]) ? $matches[4] : $matches[5];
        
        $parsedResults[] = [
            'column_1' => $col1,
            'column_2' => $col2,
            'column_3' => $col3,
            'column_4' => $col4
        ];
    }
}

// Example: Output parsed data
echo "<pre>";
print_r($parsedResults);
echo "</pre>";

How the Regex Works

Let's unpack the pattern to make it clear:

  • ^(.+?)\s+(\d{2}\.\d{2}\.\d{4})\s+: Captures Column 1 (non-greedy match until the first date) and Column 2 (the dd.mm.yyyy date).
  • (?: ... ): A non-capturing group that handles two scenarios for the remaining content:
    1. ([^\s\/]+\/[^\s\/]+)\/?(\d{2}\.\d{2}\.\d{4}): Matches both merged columns (like 1049/P/04.12.2015) and separate columns (like 211/P 12.11.2010), capturing Column 3 and Column 4.
    2. |(\d{2}\.\d{2}\.\d{4}): The fallback case where there's no Column 3 – directly captures Column 4.

Reading from a File

If you're working with a physical file instead of a sample string, use this file reading logic instead:

$fileHandle = fopen('your-file-path.txt', 'r');
while (($row = fgets($fileHandle)) !== false) {
    $row = trim($row);
    if (empty($row)) continue; // Skip empty lines
    // Apply the preg_match logic from above here
}
fclose($fileHandle);

This solution will reliably parse all the row variations in your sample data into the 4 columns you need, even when Column 3 is missing or merged with Column 4.

内容的提问来源于stack exchange,提问作者AlexF

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 20:37:33