iOS中NSRegularExpression解析印地语SRT字幕缺失匹配问题求助
Root Cause of the Problem
Your regex fails to parse Hindi subtitles because the fourth capturing group ([\s\p{P}]) only matches whitespace and punctuation characters. Hindi uses Devanagari script characters (letters, combining marks, numbers) that aren't included in this set—so the regex can't fully capture the subtitle text. This leads to incorrect boundary detection between subtitle entries, resulting in missing matches, especially towards the end of the file.
Corrected Regular Expression
Update your regex to include all Unicode characters relevant to subtitle text and handle cross-platform line endings:
(\d+)\R([\d:,.]+)\s+-->\s+([\d:,.]+)\R([\s\S]*?)(?=\R{2,}|$)
Key Improvements:
\Rinstead of\n: Matches any line break sequence (CR, LF, CRLF) for cross-platform compatibility.[\s\S]*?for subtitle text: Non-greedily matches all characters (whitespace and non-whitespace) until two consecutive line breaks or the end of the string, covering all languages including Hindi.
Swift Implementation:
let regexPattern = #"(\d+)\R([\d:,.]+)\s+-->\s+([\d:,.]+)\R([\s\S]*?)(?=\R{2,}|$)"# do { let regex = try NSRegularExpression(pattern: regexPattern, options: []) let range = NSRange(subtitleString.startIndex..., in: subtitleString) let matches = regex.matches(in: subtitleString, options: [], range: range) // Process matches as needed } catch { print("Regex initialization error: \(error.localizedDescription)") }
Ensuring Multi-Language Parsing Accuracy
To make your parser robust across all languages:
- Use Unicode property classes: If you need granular control instead of
[\s\S], leverage ICU regex properties like\p{L}(all letters),\p{M}(combining marks),\p{N}(numbers) to cover non-Latin scripts. - Test with diverse scripts: Validate your parser with subtitles in Devanagari, Chinese, Arabic, Cyrillic, and other non-Latin writing systems.
- Handle encoding correctly: Always read SRT files using UTF-8 encoding (or detect the encoding dynamically if needed) to avoid garbled text.
- Avoid over-reliance on regex: Regex can struggle with edge cases like multi-line subtitles or malformed SRT entries. A dedicated parser is often more reliable.
Alternative Solutions (No Regex)
If regex continues to cause issues, use a more straightforward parsing approach:
Custom SRT Parser
Split the subtitle string into individual entries first, then parse each entry:
func parseSRT(_ subtitleString: String) -> [SubtitleEntry] { let entries = subtitleString.components(separatedBy: "\n\n") var parsedEntries: [SubtitleEntry] = [] for entry in entries where !entry.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty { let lines = entry.components(separatedBy: .newlines).filter { !$0.isEmpty } guard lines.count >= 2 else { continue } let index = Int(lines[0]) ?? 0 let timestampLine = lines[1] let subtitleText = lines[2...].joined(separator: "\n") // Parse timestamp into start/end times let timestampParts = timestampLine.components(separatedBy: " --> ") if timestampParts.count == 2 { let startTime = parseTimestamp(timestampParts[0]) let endTime = parseTimestamp(timestampParts[1]) parsedEntries.append(SubtitleEntry(index: index, startTime: startTime, endTime: endTime, text: subtitleText)) } } return parsedEntries } // Helper to convert timestamp string to TimeInterval func parseTimestamp(_ timestamp: String) -> TimeInterval { let components = timestamp.components(separatedBy: [":", ","]) guard components.count == 4 else { return 0 } let hours = Double(components[0]) ?? 0 let minutes = Double(components[1]) ?? 0 let seconds = Double(components[2]) ?? 0 let milliseconds = Double(components[3]) ?? 0 return hours * 3600 + minutes * 60 + seconds + milliseconds / 1000 } // Struct to hold subtitle data struct SubtitleEntry { let index: Int let startTime: TimeInterval let endTime: TimeInterval let text: String }
Use a Dedicated Library
For production code, consider using a well-maintained Swift library like:
- SwiftSRT: Lightweight library specifically built for parsing SRT files.
- SubtitleKit: Supports multiple subtitle formats and handles Unicode text seamlessly.
These libraries handle edge cases and multi-language support out of the box, reducing the need for custom regex or parsing logic.
内容的提问来源于stack exchange,提问作者drishya

