C#中如何将字符串转换为SEPA格式?第三方工具集成问询
Great question! When integrating with SEPA-compliant tools, you don’t have to manually code every character replacement from the EPC conversion table. There’s a more efficient way to normalize and clean strings to meet SEPA’s strict character requirements. Let’s break down the approach:
Key Background: SEPA Allowed Characters
SEPA only accepts a limited ASCII charset:
- Uppercase/lowercase letters (A-Z, a-z)
- Digits (0-9)
- Specific symbols:
',-,?,:,(,),/,\,,, and spaces
Any other characters (like HTML entities, Unicode smart quotes, accented letters) need to be converted or removed.
Step-by-Step C# Implementation
Here’s a reusable helper method that handles HTML entity decoding, Unicode normalization, and character filtering in one go:
using System; using System.Globalization; using System.Text; using System.Text.RegularExpressions; using System.Net; // For WebUtility in .NET Core/.NET 5+ public static class SepaStringConverter { // Regex to catch any characters NOT in the SEPA allowed set private static readonly Regex _invalidSepaChars = new Regex(@"[^A-Za-z0-9'\-?:()/\\, ]", RegexOptions.Compiled); public static string ToSepaCompliantString(string input) { if (string.IsNullOrWhiteSpace(input)) return input?.Trim() ?? string.Empty; // Step 1: Decode HTML entities (e.g., ’ → ') string decodedInput = WebUtility.HtmlDecode(input); // Step 2: Normalize Unicode characters to ASCII equivalents // FormKD breaks down compatibility characters (like smart apostrophes, accented letters) string normalized = decodedInput.Normalize(NormalizationForm.FormKD); StringBuilder cleanedBuilder = new StringBuilder(); foreach (char c in normalized) { UnicodeCategory charCategory = CharUnicodeInfo.GetUnicodeCategory(c); // Keep allowed characters: letters, digits, approved symbols if (charCategory is UnicodeCategory.UppercaseLetter or UnicodeCategory.LowercaseLetter or UnicodeCategory.DecimalDigitNumber || "'-?:()/\\, ".Contains(c)) { cleanedBuilder.Append(c); } // Replace non-breaking spaces with regular spaces else if (c == '\u00A0') { cleanedBuilder.Append(' '); } // All other characters are discarded (matches SEPA's requirement to remove invalid chars) } // Step 3: Final cleanup - remove any remaining invalid chars and normalize spaces string finalString = _invalidSepaChars.Replace(cleanedBuilder.ToString(), string.Empty); finalString = Regex.Replace(finalString, @"\s+", " ").Trim(); return finalString; } }
How It Works
HTML Entity Decoding:
UsesWebUtility.HtmlDecodeto convert entities like’to their actual Unicode characters (in your example,O’CONNORbecomesO’CONNOR).Unicode Normalization:
NormalizationForm.FormKDdecomposes complex Unicode characters into their basic ASCII equivalents. For example:- Smart apostrophes
’→ regular apostrophes' - Accented letters like
é→e(since SEPA doesn’t allow accented characters)
- Smart apostrophes
Character Filtering:
Explicitly keeps only SEPA-approved characters, discards everything else, and cleans up extra spaces for consistency.
Usage with EPPlus
When writing to your XLSX file, just pass your raw string through the helper method first:
using OfficeOpenXml; // Initialize your EPPlus package and worksheet using (var package = new ExcelPackage(new FileInfo("sepa-data.xlsx"))) { var worksheet = package.Workbook.Worksheets.Add("SEPA Transactions"); // Convert customer name to SEPA format before writing string rawCustomerName = "O’CONNOR"; string sepaCompliantName = SepaStringConverter.ToSepaCompliantString(rawCustomerName); worksheet.Cells["A1"].Value = sepaCompliantName; // Writes "O'CONNOR" package.Save(); }
Customization Tips
If you need to handle edge cases from the EPC conversion table (e.g., specific symbol mappings), you can add a manual replacement step before normalization:
// Add this before normalization in the helper method string mappedInput = decodedInput .Replace("`", "'") // Map backticks to apostrophes .Replace("\"", string.Empty); // Remove double quotes (not allowed in SEPA)
内容的提问来源于stack exchange,提问作者korulis

