如何在C#中无需数据库将DAT文件转PDF并避免重复记录
Got it, let's tackle this problem step by step. You want to cut out the database middleman, convert DAT files straight to PDF in C#, and eliminate those duplicate records that were messing up your Crystal Reports output. Here's how I'd approach it:
Before even thinking about PDFs, let's get rid of duplicate records right when loading the DAT data into memory. First, define a model that matches your DAT file's structure (adjust fields to fit your actual data):
// Match this to your DAT file's columns public class DatRecord { public int RecordId { get; set; } public string CustomerName { get; set; } public decimal TransactionAmount { get; set; } // Add other fields as needed }
Then write a method to load all DAT files, parse their contents, and deduplicate using a unique identifier (or combination of fields that define a unique record):
using System.IO; using System.Linq; using System.Collections.Generic; public List<DatRecord> LoadAndDeduplicateDats(string[] datFilePaths) { var allRawRecords = new List<DatRecord>(); foreach (var filePath in datFilePaths) { var lines = File.ReadAllLines(filePath); // Skip header line if your DAT files have one; remove Skip(1) if not foreach (var line in lines.Skip(1)) { // Replace ',' with your actual delimiter (e.g., '\t' for tab-separated) var fields = line.Split(','); var record = new DatRecord { RecordId = int.Parse(fields[0]), CustomerName = fields[1].Trim(), TransactionAmount = decimal.Parse(fields[2]) // Map remaining fields here }; allRawRecords.Add(record); } } // Deduplicate: Group by your unique key (use a composite key like new { r.RecordId, r.CustomerName } if needed) var deduplicatedRecords = allRawRecords .GroupBy(r => r.RecordId) .Select(group => group.First()) // Keep only the first instance of each unique record .ToList(); return deduplicatedRecords; }
You don't need a database to feed data into Crystal Reports—or you can even ditch Crystal Reports entirely for a lighter solution. Here are two options:
Option A: Keep Using Crystal Reports (With In-Memory Data)
Crystal Reports supports binding to in-memory DataTables or lists, so you can skip inserting data into a database entirely. Convert your deduplicated list to a DataTable and bind it to your report:
using System.Data; using System.Reflection; using CrystalDecisions.CrystalReports.Engine; using CrystalDecisions.Shared; // Helper to convert list to DataTable for Crystal Reports public DataTable ConvertListToDataTable<T>(List<T> data) { var props = typeof(T).GetProperties(); var dt = new DataTable(); // Add columns matching your model properties foreach (var prop in props) { dt.Columns.Add(prop.Name, Nullable.GetUnderlyingType(prop.PropertyType) ?? prop.PropertyType); } // Populate rows foreach (var item in data) { var row = dt.NewRow(); foreach (var prop in props) { row[prop.Name] = prop.GetValue(item) ?? DBNull.Value; } dt.Rows.Add(row); } return dt; } // Export to PDF using Crystal Reports public void ExportCrystalReportToPdf(List<DatRecord> deduplicatedRecords, string outputPdfPath) { var dataTable = ConvertListToDataTable(deduplicatedRecords); var report = new YourCustomCrystalReport(); // Replace with your actual report class report.SetDataSource(dataTable); // Configure export settings var exportOpts = report.ExportOptions; exportOpts.ExportFormatType = ExportFormatType.PortableDocFormat; exportOpts.DestinationOptions = new DiskFileDestinationOptions { DiskFileName = outputPdfPath }; exportOpts.ExportDestinationType = ExportDestinationType.DiskFile; report.Export(); }
Option B: Use a Lightweight PDF Library (No Crystal Reports Dependency)
If you want to avoid Crystal Reports' runtime dependencies, use a library like iTextSharp (or iText 7 for .NET Core) to build PDFs directly from your in-memory data:
First, add the iTextSharp NuGet package, then use this code:
using iTextSharp.text; using iTextSharp.text.pdf; using System.IO; public void GeneratePdfFromRecords(List<DatRecord> records, string outputPath) { using (var fs = new FileStream(outputPath, FileMode.Create)) { var document = new Document(PageSize.A4, 20, 20, 20, 20); PdfWriter.GetInstance(document, fs); document.Open(); // Add a table to hold your data var pdfTable = new PdfPTable(3); // Match column count to your data fields pdfTable.WidthPercentage = 100; // Add header cells pdfTable.AddCell(new Phrase("Record ID")); pdfTable.AddCell(new Phrase("Customer Name")); pdfTable.AddCell(new Phrase("Amount")); // Add data rows foreach (var record in records) { pdfTable.AddCell(new Phrase(record.RecordId.ToString())); pdfTable.AddCell(new Phrase(record.CustomerName)); pdfTable.AddCell(new Phrase(record.TransactionAmount.ToString("C"))); } document.Add(pdfTable); document.Close(); } }
- Define a solid unique key: Make sure the field/combination you use for grouping (e.g.,
RecordId) truly identifies unique records—using the wrong key will either leave duplicates or delete valid data. - Validate DAT data: Add checks during parsing (e.g., ensure field counts match expected values, validate numeric data) to avoid corrupted records breaking your workflow.
- Test with sample data: Run your deduplication logic against a small set of DAT files with known duplicates to confirm it works as expected before scaling.
内容的提问来源于stack exchange,提问作者Vinoothna Y

