基于正则匹配计数筛选Excel导入数据目标列的技术问询
Got it, let's fix this up for you! The issue with your current code is that a Set only tracks unique column indices, not how many times each column matched or which rows contributed to those matches. Here's how you can modify it to get both per-column match counts and the corresponding rows:
First, instead of a Set, we'll use an object (or Map) to keep track of each column's match count and the rows that matched. This way, we can accumulate counts per column and store the relevant rows at the same time.
Modified Code
columnsWithDescription() { const refDesRegex = [/resistor/i, /capacitor/i, /res/i, /cap/i]; // Use an object to track each column's data: { count: number, rows: Array } const columnStats = {}; for (const row of this.data) { for (let cellIndex = 0; cellIndex < row.length; cellIndex++) { const cellValue = row[cellIndex]; // Check if this cell matches any of the regex patterns const isMatch = refDesRegex.some(regex => regex.test(cellValue)); if (isMatch) { // Initialize the column entry if it doesn't exist if (!columnStats[cellIndex]) { columnStats[cellIndex] = { count: 0, rows: [] }; } // Increment the count and add the row to the list columnStats[cellIndex].count++; columnStats[cellIndex].rows.push([...row]); // Clone the row to avoid reference issues } } } // Optional: Convert the object to an array of entries for easier iteration const columnStatsArray = Object.entries(columnStats).map(([colIndex, stats]) => ({ columnNumber: parseInt(colIndex), matchCount: stats.count, matchingRows: stats.rows })); return columnStatsArray; // Or return columnStats if you prefer the object format }
Key Changes Explained
- Replaced
SetwithcolumnStatsobject: This object uses column indices as keys, and each value holds two pieces of data:count: The number of times the column matched any of your regex patternsrows: An array of all rows where this column had a matching value
- Early match check with
some(): Instead of looping through all regex patterns for every cell, we useArray.some()to stop checking as soon as one regex matches (more efficient) - Row cloning: We use
[...row]to create a copy of the row, so any later changes tothis.datawon't affect the stored rows - Optional conversion to array: The
columnStatsArrayconverts the object into an array of objects with clear labels (columnNumber,matchCount,matchingRows) for easier use later
How to Use the Result
Once you call this function, you'll get an array like this:
[ { columnNumber: 2, matchCount: 15, matchingRows: [ [row1data], [row2data], ... ] }, { columnNumber: 5, matchCount: 8, matchingRows: [ [rowAdata], [rowBdata], ... ] } ]
You can iterate over this array to analyze each column's performance, or filter columns based on match count (e.g., only keep columns with more than 5 matches to reduce false positives).
内容的提问来源于stack exchange,提问作者Banani720

