Skip to main content
Question

Does Google SecOps YARA-L 2.0 support Levenshtein distance or fuzzy string matching?

  • October 8, 2026
  • 4 replies
  • 52 views

NASEEF
Forum|alt.badge.img+9

Hi Google SecOps Community,

i would like to confirm whether YARA-L supports Levenshtein distance (edit distance) or any equivalent fuzzy string-matching functionality.

Use Case:

We have a detection that identifies potential data exfiltration through Microsoft 365 Exchange emails by comparing the username portion of the sender's corporate email address with the username portion of the recipient's personal email address.

The objective is to identify sender and recipient usernames that are identical or differ by fewer than three character edits.

For example:

  • [removed by moderator] → [removed by moderator] (distance = 1)

  • [removed by moderator] → [removed by moderator] (distance = 2)

Questions:

  1. Does Google SecOps YARA-L 2.0 provide a built-in function for calculating Levenshtein distance between two strings?

  2. If Levenshtein distance is not supported natively, is there another supported YARA-L function or approach for fuzzy string matching?

  3. Can this comparison be implemented directly within a YARA-L detection rule using extracted sender and recipient usernames?

  4. If native support is unavailable, what is Google's recommended approach for implementing this detection? Would upstream enrichment be required?

  5. Are there any existing Google SecOps detection rules or examples that implement similar username-similarity logic?

Any guidance, documentation, or working YARA-L examples would be greatly appreciated.

Thanks!

4 replies

Eoved
Forum|alt.badge.img+10
  • Bronze 4
  • October 8, 2026

Hi,
I think YARA-L does not currently have a built-in function for Levenshtein distance, edit distance, or fuzzy string matching.
However, you can try the following YARA-L 2.0 rule for the exact-match scenario:

rule email_exfiltration_to_personal_account {
meta:
author = "Eoved For Community"
description = "Detects outbound corporate email sent to an external personal/freemail address where sender and recipient usernames match."
mitre_attack_tactic = "TA0010"
mitre_attack_technique = "T1048.003"
severity = "MEDIUM"
priority = "MEDIUM"

events:
$e.metadata.event_type = "EMAIL_TRANSACTION"

// 1. Direct domain checks on UDM event fields
$e.principal.user.email_addresses = /@yourcompany\.com$/ //or NOT your corporate domain
$e.target.user.email_addresses = /@(gmail\.com|yahoo\.com|outlook\.com|hotmail\.com|proton\.me|icloud\.com)$/

// 2. Extract and lowercase usernames directly from event fields
$matching_user = strings.to_lower(re.capture($e.principal.user.email_addresses, "^([^@]+)@"))
$matching_user = strings.to_lower(re.capture($e.target.user.email_addresses, "^([^@]+)@"))

// 3. Exclude empty captures (prevents "" = "" matches)
$matching_user != ""

outcome:
$risk_score = 65
$sender = array_distinct($e.principal.user.email_addresses)
$recipient = array_distinct($e.target.user.email_addresses)
$matched_username = array_distinct($matching_user)

condition:
$e
}

 


NASEEF
Forum|alt.badge.img+9
  • Author
  • Silver 2
  • October 8, 2026

thank you Eoved


codyro
Forum|alt.badge.img
  • Bronze 1
  • October 8, 2026

Hey! I also went down this rabbit hole somewhat recently and came to the same conclusion as ​@Eoved that there is no levenshtein function or anything similar. 

For a use-case I was working on recently, I ended up building the character comparison manually which resulted in a substitution check (so technically Hamming distance, not true Levenshtein). 

Here’s where I landed if it helps jump start your journey hah. Let me know if you find any other cool ways to approach this issue.

The rule currently only fires on 1 character substitution set in the condition ($total_mismatches = 1) but you could obviously increase the tolerance to a larger range if you wanted to see a higher number of potential incorrect characters. 

rule potential_bec_typosquat {

meta:
author = "Citreno"
description = "Detects potential BEC lookalikes using positional character analysis (10 chars)."
data_source = "ProofPoint On Demand"
tactic = "TA0001"
technique = "T1566"

events:
$outbound.metadata.log_type = "PROOFPOINT_ON_DEMAND"
$outbound.metadata.event_type = "EMAIL_TRANSACTION"
$outbound.network.direction = "OUTBOUND"
$internal_user = strings.to_lower($outbound.principal.user.email_addresses)
$out_d = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@(.+)$`))

$inbound.metadata.log_type = "PROOFPOINT_ON_DEMAND"
$inbound.metadata.event_type = "EMAIL_TRANSACTION"
$inbound.network.direction = "INBOUND"
$internal_user = strings.to_lower($inbound.target.user.email_addresses)
$in_d = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@(.+)$`))

$inbound.network.email.subject = /^\s*(\[[^\]]*\]\s*)?(re|fw|fwd)\s*:/ nocase

$outbound.metadata.event_timestamp.seconds <= $inbound.metadata.event_timestamp.seconds

$out_d != ""
$in_d != ""
$out_d != $in_d

$o1 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@(.)`))
$o2 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@.{1}(.)`))
$o3 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@.{2}(.)`))
$o4 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@.{3}(.)`))
$o5 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@.{4}(.)`))
$o6 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@.{5}(.)`))
$o7 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@.{6}(.)`))
$o8 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@.{7}(.)`))
$o9 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@.{8}(.)`))
$o10 = strings.to_lower(re.capture($outbound.target.user.email_addresses, `@.{9}(.)`))

$i1 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@(.)`))
$i2 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@.{1}(.)`))
$i3 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@.{2}(.)`))
$i4 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@.{3}(.)`))
$i5 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@.{4}(.)`))
$i6 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@.{5}(.)`))
$i7 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@.{6}(.)`))
$i8 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@.{7}(.)`))
$i9 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@.{8}(.)`))
$i10 = strings.to_lower(re.capture($inbound.principal.user.email_addresses, `@.{9}(.)`))

match:
$internal_user, $out_d, $in_d over 2d

outcome:
$total_mismatches = max(
if($o1 != $i1 and $o1 != "", 1, 0) +
if($o2 != $i2 and $o2 != "", 1, 0) +
if($o3 != $i3 and $o3 != "", 1, 0) +
if($o4 != $i4 and $o4 != "", 1, 0) +
if($o5 != $i5 and $o5 != "", 1, 0) +
if($o6 != $i6 and $o6 != "", 1, 0) +
if($o7 != $i7 and $o7 != "", 1, 0) +
if($o8 != $i8 and $o8 != "", 1, 0) +
if($o9 != $i9 and $o9 != "", 1, 0) +
if($o10 != $i10 and $o10 != "", 1, 0)
)

$first_outbound_time = min($outbound.metadata.event_timestamp.seconds)
$first_inbound_time = min($inbound.metadata.event_timestamp.seconds)
$last_outbound_time = max($outbound.metadata.event_timestamp.seconds)
$last_inbound_time = max($inbound.metadata.event_timestamp.seconds)

condition:
$outbound and $inbound and $total_mismatches = 1
}

NASEEF
Forum|alt.badge.img+9
  • Author
  • Silver 2
  • October 8, 2026

Good work codyro, much appreciated. but I’m checking email prefixes, which can often be longer than 11 characters, so we may need to create a larger drop-down. But this is good.