You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用sed替换多行换行符并添加分隔符,适配文本文件表格导入?

How to Convert Your Artist/Venue .txt File to Spreadsheet-Friendly Format

Got it, let's break down how to fix this file so it imports cleanly into a spreadsheet. The core challenge is handling variable artist counts and messy line breaks—here are two straightforward approaches depending on your comfort level with tools:

Option 1: Use a Text Editor with Regex (No Coding Needed)

This works great if you're using Notepad++, VS Code, or any editor that supports regular expression find-and-replace.

Step 1: Clean Up Extra Line Breaks

First, get rid of all those random empty lines between entries:

  • Open your .txt file in the editor
  • Enable Regular Expression mode in the find-and-replace tool
  • Find: \n\s*\n (for Windows systems, use \r\n\s*\r\n instead)
  • Replace with: \n
  • Click "Replace All"—repeat until no more empty lines are left.

Step 2: Restructure Entries into 4 Columns

Now we'll reformat each line to split into your desired four fields (Artists, Venue, Location, Date) using a tab character as the separator (tabs work better than commas here since artists already use commas):

  • Keep regex mode enabled
  • Find pattern: ^(.*?),\s*([^,]+),\s*([^,]+),\s*([^,]+),\s*([^\r\n]+)$
  • Replace with: $1\t$2\t$3, $4\t$5
  • Hit "Replace All"

Let's break down what this does:

  • $1: Captures all the artists (everything up to the 4th-to-last comma)
  • $2: Grabs the venue name
  • $3, $4: Merges City and State into a single "Location" field
  • $5: Pulls the date
  • \t: Inserts a tab between each field, which spreadsheets will recognize as column separators

Step 3: Import to Spreadsheet

Open Excel, Google Sheets, or your preferred tool, then import the cleaned file. When prompted, select Tab as the delimiter—your data will pop into four neat columns automatically.

Option 2: Python Script (For Batch/Automated Processing)

If you have multiple files or want to automate this, a quick Python script will handle it reliably:

import re

# Replace with your input file path
input_path = "your_input_file.txt"
# Replace with your desired output file path
output_path = "cleaned_for_spreadsheet.txt"

# Read and clean the file content
with open(input_path, "r", encoding="utf-8") as f:
    content = f.read()

# Remove extra empty lines
clean_content = re.sub(r"\n\s*\n", "\n", content).strip()

processed_lines = []
for line in clean_content.split("\n"):
    # Split line into parts, stripping whitespace from each
    parts = [segment.strip() for segment in line.split(",")]
    
    # Extract fields: handle variable artist counts
    artists = ", ".join(parts[:-4])  # All parts except last 4 are artists
    venue = parts[-4]
    location = f"{parts[-3]}, {parts[-2]}"  # Merge City + State
    date = parts[-1]
    
    # Combine into tab-separated line
    processed_lines.append("\t".join([artists, venue, location, date]))

# Write the cleaned data to a new file
with open(output_path, "w", encoding="utf-8") as f:
    f.write("\n".join(processed_lines))

Quick Note:

If any artist names, venues, or locations contain commas (e.g., "Doe, Jane"), the split method will break. For those edge cases, you'd need to use a CSV parser that handles quoted fields—but this script works perfectly for the format you described.

内容的提问来源于stack exchange,提问作者r123454321

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:59:07