You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HIVE字符串路径清理:移除特定字符间非纯字母内容

Clean Paths to Retain Only Letter-Only Segments

Got it, let's tackle this path-cleaning task head-on. The core ask is simple: we need to strip out any path segments wrapped in / that contain non-alphabetic characters (like numbers, hyphens, underscores), and keep only the segments made up entirely of letters.

Approach Breakdown

Here's a step-by-step way to implement this logic:

  1. Split the input path into individual segments using / as the separator.
  2. Filter out any segments that aren't 100% alphabetic (a-z, A-Z only).
  3. Reconstruct the cleaned path, making sure to preserve the leading / and match the trailing slash behavior from your examples.

Example Implementation (Python)

def clean_target_path(input_path):
    # Split path into segments, skip empty strings from leading/trailing slashes
    segments = [seg for seg in input_path.split('/') if seg]
    # Keep only segments with pure letters
    valid_segments = [seg for seg in segments if seg.isalpha()]
    # Rebuild the path with leading slash, add trailing slash if needed (matches your examples)
    cleaned_path = '/' + '/'.join(valid_segments)
    # Add trailing slash only if there are valid segments and the original had a non-letter segment at the end
    if valid_segments and not input_path.endswith(valid_segments[-1]):
        cleaned_path += '/'
    return cleaned_path

# Test with your sample paths
sample_paths = [
    "/keywordOneA/keywordTwoA/393r-mr49-j5n65_9e8e77g77b8",
    "/keywordOneA/keywordTwoA/111-4444-jjjj_1b1b1b1b1b1b1b",
    "/keywordOneA/keywordTwoB/393r-mr49-j5n65_9e8e77g77b8/keywordThreeA"
]

for idx, path in enumerate(sample_paths, 1):
    print(f"字符串{idx}处理后: {clean_target_path(path)}")

Output Result

字符串1处理后: /keywordOneA/keywordTwoA/
字符串2处理后: /keywordOneA/keywordTwoA/
字符串3处理后: /keywordOneA/keywordTwoB/keywordThreeA

How This Works

  • split('/') breaks the path into parts (e.g., ["", "keywordOneA", "keywordTwoA", "393r-mr49-j5n65_9e8e77g77b8"] for the first sample).
  • We filter out empty strings to avoid extra slashes in the final output.
  • seg.isalpha() checks if a segment has only letters—this automatically rejects any segment with numbers, symbols, or mixed characters.
  • The trailing slash logic ensures we match your expected output: if the original path ended with a non-letter segment, we add a trailing slash to the cleaned path; if it ended with a valid letter segment, we skip the trailing slash.

If you're using a different language, the core logic translates directly. For example, in JavaScript, you'd use split('/'), filter(seg => /^[a-zA-Z]+$/.test(seg)), then join('/') with similar leading/trailing slash handling.

内容的提问来源于stack exchange,提问作者Developer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:01:37